Construction body for preparing circular RNA, method and application thereof
By optimizing the self-splicing design of type II introns, splitting intron fragments and introducing complementary pairing sequences, the problems of scar sequence residue and low efficiency in circular RNA preparation in existing technologies were solved, and efficient, GTP-free circular RNA preparation and expression were achieved.
Patent Information
- Application Number
- CN202510830820.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-19
- Publication Date
- 2025-09-09
AI Technical Summary
In the existing technology for preparing circular RNA, type I introns require longer original exon sequences, resulting in residual scar sequences, and the splicing efficiency of type II introns is low, which cannot meet the needs of efficient preparation of circular RNA.
By using an optimized type II intron design, the 5' and 3' intron fragments were split into two fragments, which were self-spliced in vitro to form circular RNA. By designing complementary pairing sequences and modifying the EBS sequence, the splicing efficiency was improved and the participation of GTP was avoided.
The preparation of circular RNA without scar sequences was achieved, and the splicing efficiency was increased from 10% to about 50%, and even reached 98%, which is suitable for expression and application in eukaryotic cells.
Smart Images

Figure BDA0005459434230000571 
Figure BDA0005459434230000581 
Figure BDA0005459434230000591
Abstract
Description
Reference to electronically submitted sequence listing
[0001] This application incorporates by reference the sequence listing filed with this application as a text file entitled "TPI04718 - Sequence listing" created on June 19, 2025 and having a size of 175 KB bytes. 1. Technical Field
[0002] The present invention relates to the field of molecular biology, and more particularly to constructs, methods, and applications for preparing circular RNAs that can be used to express target proteins in eukaryotic cells or to perform corresponding functions in the form of non-coding RNAs. 2. Background Technology
[0003] Circular RNA (circRNA) is a type of circular RNA molecule formed by head-to-tail connection. In recent years, literature has reported that circular RNA can regulate gene transcription, neutralize miRNA activity and the binding of RNA-binding proteins, and can also serve as templates for translation to generate proteins (Yang, Y. et al., “Extensive translation of circular RNAs driven by N(6)-methyladenosine,” Cell Research, 27(5): 626-641 (2017); Abe, N. et al., “Rolling Circle Translation of Circular RNA in Living Human Cells”, Scientific Reports, 5: 16435 (2015); Gao, X. et al., “Circular RNA-encoded oncogenic E-cadherin variant promotes glioblastoma tumorigenicity through activation of EGFR-STAT3 signaling,” Nature Cell Biology, 23(3): 278-291 (2021); Pamudurti, NR. et al., “Translation of CircRNAs,” Molecular Cell, 66(1): 9-21 (2017)). Compared with linear RNA, circular RNA is not easily recognized by the RNA degradation system due to its head-to-tail covalent closed loop structure, so it has greater stability and has the potential and prospects to become a new generation of RNA drug platform.
[0004] Currently, there are three main methods for preparing circular RNA in vitro. One method involves connecting the 5' and 3' ends of linear RNA end-to-end through an RNA ligation reaction catalyzed by a nucleic acid ligase. The RNA ligase is an exogenous protein, such as T4 RNA ligase. Another method involves chemical ligation, which connects the 5' and 3' ends of RNA using cyanide bromide and morpholino derivatives. Another, more advanced method involves obtaining end-to-end connected circular RNA through a ribozyme-catalyzed RNA splicing reaction. This method expresses circular RNA by designing an expression framework containing a ribozyme sequence with self-splicing function.
[0005] Currently, ribozymes capable of RNA self-splicing are generally divided into two major categories, known as type I and type II introns. Literature reports indicate that both types of introns can self-splice under appropriate reaction conditions, joining two RNA fragments together. While the splicing products of these two types of ribozymes are similar, the structures and splicing mechanisms of the ribozymes themselves differ significantly.
[0006] The group I intron has a 9-helical structure, and catalytic splicing requires the hydroxyl group (pG-OH) in the external guanosine phosphate to trigger the reaction, and is highly dependent on the exon sequences located at both ends of the group I intron.
[0007] Group II introns rely on their own hydroxyl groups within the nucleic acid sequence to trigger splicing. This splicing mechanism is closer to the splicing reaction mediated by the spliceosome, which better mimics the splicing of higher organisms.
[0008] The above structural differences determine that group I intron self-splicing requires a longer original exon sequence, also known as the scar sequence.
[0009] Previous studies have shown that circular RNA can be prepared in vitro using both types of intronic ribozymes, but the efficiency is low (Puttaraju, M. and Been, MD., "Group I permuted intron-exon (PIE) sequences self-splice to produce circular exons," Nucleic Acids Research, 20(20): 5357-64 (1992); Mikheeva, S. et al., "Use of an engineered ribozyme to produce a circular human exon," Nucleic Acids Research, 25(24): 5085-94 (1997)).
[0010] An article by Wesselhoeft et al. reports a method for improving RNA cyclization efficiency by optimizing a construct containing a type I intron (Wesselhoeft, RA. et al., "Engineering circular RNA for potent and stable translation in eukaryotic cells," Nature Communications, 9(1): 2629(2018)). A related patent application (WO 2019 / 236673 A1) discloses a construct containing a type I intron for forming a circular coding RNA. Wesselhoeft et al. rearranged the type I intron and the exons at both ends thereof, and constructed a target protein (POI) with a ribosome entry site (IRES) into this framework, and then obtained a circular coding RNA that can translate the target protein by self-splicing in the presence of GTP. By selecting different type I introns and performing design modifications, the RNA cyclization efficiency is improved. Specifically, the technique first deletes some of the Td gene of the T4 phage, retaining sequences that allow for proper folding and thus maintain ribozyme activity, including introns and a portion of exons. The gene is then split into two, with the 3' intron and exon fragment 2 (E2) constructed into the 5' end of the IRES-POI, and the exon fragment 1 (E1) and the 5' intron constructed into the 3' end of the IRES-POI. In the presence of GTP and magnesium ions, self-splicing occurs to produce circular RNA. However, Wesselhoeft et al. discovered that the 5' and 3' splice sites could not be efficiently spliced due to the insertion of the target gene. To address this issue, Wesselhoeft et al. inserted complementary "homology arms" near the splice sites, thereby improving splicing efficiency. Based on existing literature (Mikheeva, S. et al., (1997), supra), another type I intron, Anabaena, was selected and found to have higher splicing efficiency than the Td intron. Similar design modifications were then made to the intron, further improving splicing efficiency. The article finally verified that the expression framework can effectively translate the target protein.
[0011] However, the design of Wesselhoeft et al. has the following disadvantages: 1. When using group I introns, they must contain a long original exon sequence. Therefore, the expression product will contain a section of original exogenous sequence (scar sequence). When preparing the target sequence into circular RNA, it is usually desirable to remove this sequence that does not belong to the target sequence to facilitate subsequent applications. 2. Group I introns require GTP to provide energy during self-splicing. On the other hand, the splicing efficiency of group II introns in previous literature is low (about 10%) (Mikheeva, S. et al., (1997), supra). Therefore, there is still a need in the art for improved constructs and methods for preparing circular RNA. 3. Summary of the Invention
[0012] The inventors of the present application have created a methodology for preparing circular RNA through self-splicing of group II introns through screening and design optimization, overcoming the above problems.
[0013] Therefore, the present invention provides a polynucleotide construct having self-splicing activity in vitro, which comprises the following operably linked elements from 5' to 3': (a) 3' intron fragment; (b) exon fragment 2 (E2); (c) target sequence; (d) exon fragment 1 (E1); (e) 5' intron fragment, in: The 5' intron fragment and the 3' intron fragment are each a fragment of a group II intron, wherein the 5' intron fragment is located on the 5' side of the 3' intron fragment in the group II intron, The E1 is a fragment of the 5' exon of the group II intron, and its length is ≥ 0 nucleotides. The E2 is a fragment of the 3' exon of the group II intron, and its length is ≥ 0 nucleotides. The target sequence is empty, or is a protein coding sequence, a non-coding sequence, or a combination of the two; And wherein the group II intron is BR23 or CL.
[0014] The present invention also provides a polynucleotide construct having self-splicing activity in vitro, which comprises the following operably linked elements from 5' to 3': (a) 3' intron fragment; (b) exon fragment 2 (E2); (c) linker sequence; (d) target sequence; (e) linker sequence; (f) exon fragment 1 (E1); (g) 5' intron fragment, in: The 5' intron fragment and the 3' intron fragment are each a fragment of a group II intron, wherein the 5' intron fragment is located on the 5' side of the 3' intron fragment in the group II intron, The E1 is a fragment of the 5' exon of the group II intron, and its length is ≥ 0 nucleotides. The E2 is a fragment of the 3' exon of the group II intron, and its length is ≥ 0 nucleotides. The target sequence is empty, or is a protein coding sequence, a non-coding sequence, or a combination of the two; And wherein the group II intron is BR23 or CL.
[0015] The present invention also provides a polynucleotide construct having self-splicing activity in vitro, which comprises the following operably linked elements from 5' to 3': (a) 5' homology arm; (b) 3' intron fragment; (c) exon fragment 2 (E2); (d) target sequence; (e) exon fragment 1 (E1); (f) 5' intron fragment; (g) 3' homology arm, in: The 5' intron fragment and the 3' intron fragment are each a fragment of a group II intron, wherein the 5' intron fragment is located on the 5' side of the 3' intron fragment in the group II intron, The E1 is a fragment of the 5' exon of the group II intron, and its length is ≥ 0 nucleotides. The E2 is a fragment of the 3' exon of the group II intron, and its length is ≥ 0 nucleotides. The target sequence is empty, or is a protein coding sequence, a non-coding sequence, or a combination of the two; And wherein the group II intron is BR23 or CL.
[0016] The present invention also provides a polynucleotide construct having self-splicing activity in vitro, which comprises the following operably linked elements from 5' to 3': (a) 5' homology arm; (b) 3' intron fragment; (c) exon fragment 2 (E2); (d) linker sequence; (e) target sequence; (f) linker sequence; (g) exon fragment 1 (E1); (h) 5' intron fragment; (i) 3' homology arm, in: The 5' intron fragment and the 3' intron fragment are each a fragment of a group II intron divided into two fragments, the 5' intron fragment is located on the 5' side of the 3' intron fragment in the group II intron, The E1 is a fragment of the 5' exon of the group II intron, and its length is ≥ 0 nucleotides. The E2 is a fragment of the 3' exon of the group II intron, and its length is ≥ 0 nucleotides. The target sequence is empty, or is a protein coding sequence, a non-coding sequence, or a combination of the two; And wherein the group II intron is BR23 or CL.
[0017] In some embodiments, the polynucleotide construct is an RNA polynucleotide construct.
[0018] In some embodiments, the polynucleotide construct is capable of forming a circular RNA of a target sequence in vitro.
[0019] In some embodiments, the polynucleotide construct is capable of forming a circular RNA of a target sequence in vivo.
[0020] The present invention provides circular RNA produced by the polynucleotide constructs of the present invention. In some embodiments, the circular RNA is at least 500 nucleotides in length, at least 1,000 nucleotides in length, or at least 1,500 nucleotides in length.
[0021] The present invention provides a method for producing circular RNA using the polynucleotide construct of the present invention.
[0022] The present invention provides a method for producing circular RNA, comprising: preparing a vector comprising the following operably linked elements from 5' to 3': (a) 3' intron fragment; (b) exon fragment 2 (E2); (c) target sequence; (d) exon fragment 1 (E1); (e) 5' intron fragment, in: The 5' intron fragment and the 3' intron fragment are each a fragment of a group II intron, wherein the 5' intron fragment is located on the 5' side of the 3' intron fragment in the group II intron, The E1 is a fragment of the 5' exon of the group II intron, and its length is ≥ 0 nucleotides. The E2 is a fragment of the 3' exon of the group II intron, and its length is ≥ 0 nucleotides. The target sequence is empty, or is a protein coding sequence, a non-coding sequence, or a combination of the two; And wherein the group II intron is BR23 or CL.
[0023] The present invention also provides a method for producing circular RNA, comprising: preparing a vector comprising the following operably linked elements from 5' to 3': (a) 3' intron fragment; (b) exon fragment 2 (E2); (c) linker sequence; (d) target sequence; (e) linker sequence; (f) exon fragment 1 (E1); (g) 5' intron fragment, in: The 5' intron fragment and the 3' intron fragment are each a fragment of a group II intron, wherein the 5' intron fragment is located on the 5' side of the 3' intron fragment in the group II intron, The E1 is a fragment of the 5' exon of the group II intron, and its length is ≥ 0 nucleotides. The E2 is a fragment of the 3' exon of the group II intron, and its length is ≥ 0 nucleotides. The target sequence is empty, or is a protein coding sequence, a non-coding sequence, or a combination of the two; And wherein the group II intron is BR23 or CL.
[0024] The present invention also provides a method for producing circular RNA, comprising: preparing a vector comprising the following operably linked elements from 5' to 3': (a) 5' homology arm; (b) 3' intron fragment; (c) exon fragment 2 (E2); (d) target sequence; (e) exon fragment 1 (E1); (f) 5' intron fragment; (g) 3' homology arm, in: The 5' intron fragment and the 3' intron fragment are each a fragment of a group II intron, wherein the 5' intron fragment is located on the 5' side of the 3' intron fragment in the group II intron, The E1 is a fragment of the 5' exon of the group II intron, and its length is ≥ 0 nucleotides. The E2 is a fragment of the 3' exon of the group II intron, and its length is ≥ 0 nucleotides. The target sequence is empty, or is a protein coding sequence, a non-coding sequence, or a combination of the two; And wherein the group II intron is BR23 or CL.
[0025] The present invention also provides a method for producing circular RNA, comprising: preparing a vector comprising the following operably linked elements from 5' to 3': (a) 5' homology arm; (b) 3' intron fragment; (c) exon fragment 2 (E2); (d) linker sequence; (e) target sequence; (f) linker sequence; (g) exon fragment 1 (E1); (h) 5' intron fragment; (i) 3' homology arm, in: The 5' intron fragment and the 3' intron fragment are each a fragment of a group II intron divided into two fragments, the 5' intron fragment is located on the 5' side of the 3' intron fragment in the group II intron, The E1 is a fragment of the 5' exon of the group II intron, and its length is ≥ 0 nucleotides. The E2 is a fragment of the 3' exon of the group II intron, and its length is ≥ 0 nucleotides. The target sequence is empty, or is a protein coding sequence, a non-coding sequence, or a combination of the two; And wherein the group II intron is BR23 or CL.
[0026] The present invention provides a method for expressing a protein in a cell, comprising transfecting the circular RNA of the present invention into the cell.
[0027] The present invention provides a method for expressing a protein in a cell, comprising (a) transfecting the circular RNA of the present invention into the cell, or (b) allowing the polynucleotide construct of the present invention to undergo a self-splicing cyclization reaction to form a circular RNA, and transfecting the circular RNA into the cell; wherein, preferably, the cell is a eukaryotic cell.
[0028] The constructs, methods and uses of the present invention have at least the following advantages: 1. Produce circular RNA without scar sequences, which is more conducive to orderly application; 2. The self-splicing reaction of the polynucleotide to form circular RNA does not require GTP (e.g., in some embodiments, only Mg ions, Na ions are required); and / or 3. The splicing efficiency of group II introns has been greatly improved, from 10% to about 50%, and even up to 98% splicing efficiency.
[0029] In a specific embodiment, the length of E1 and / or E2 is 0-20 nucleotides, preferably 0-10 nucleotides, such as 0, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10 nucleotides.
[0030] In a specific embodiment, the 5' intron fragment and the 3' intron fragment divide the group II intron into two fragments from the unpaired region. In a specific embodiment, the unpaired region is selected from the linear region between two adjacent domains of the group II intron or the loop region of the stem-loop structure of domain 4.
[0031] In a specific embodiment, the group II intron comprises one or more nucleotide modifications relative to its wild-type form, wherein the modification is selected from one or more of deletion, substitution, and addition.
[0032] In a specific embodiment, the 5' intron fragment and the 3' intron fragment each contain one or more pairs of complementary paired sequences. In a preferred embodiment, the length of the complementary paired sequences is greater than 20 nucleotides.
[0033] In a specific embodiment, the 5' intron fragment and / or the 3' intron fragment contains one or more affinity tag sequences, and the affinity tag sequences are selected from one or more of the following groups: a probe binding sequence, an MS2 binding site, a PP7 binding site, and a streptavidin binding site.
[0034] In a specific embodiment, wherein the E1 and E2 are 0, and the modification comprises modifying one or more EBS sequences of the type II intron so that the EBS sequences are complementary to one or more regions of corresponding length in the target sequence at at least 60% of the nucleotide positions. The EBS sequence is selected from one or more of EBS1, EBS2, and EBS3, preferably any two, more preferably EBS1 and EBS3. In a preferred embodiment, the modification is to modify the two EBS sequences of the type II intron, preferably EBS1 and EBS3, so that the EBS sequences are complementary to two regions of corresponding length in the target sequence at at least 60% of the nucleotide positions. In a preferred embodiment, the modification is to modify the two EBS sequences of the type II intron, preferably EBS1' and EBS3', so that the EBS sequences are complementary to two regions of corresponding length in the target sequence at at least 60% of the nucleotide positions. In a preferred embodiment, the modification is to modify two EBS sequences of the group II intron, preferably EBS1" and EBS3", so that the EBS sequences are complementary to two regions of corresponding length in the target sequence at at least 60% of the nucleotide positions. In another preferred embodiment, the modification is to modify the δ or δ" sequence of the group II intron, wherein the δ or δ" sequence is complementary to a region of corresponding length in the target sequence at at least 60% of the nucleotide positions; preferably, the region is located at one end of the target sequence.
[0035] In a preferred embodiment, the two regions of corresponding lengths in the target sequence are located at both ends of the target sequence.
[0036] In a specific embodiment, the modification is deletion of part or all of domain 4, such as deletion of the IEP sequence in domain 4, preferably deletion of the entire domain 4.
[0037] In a specific embodiment, the type II intron is a type II intron derived from a microorganism. Preferably, the type II intron has in vitro self-splicing activity. In a specific embodiment, the type II intron is a type II intron from Clostridium, such as Clostridium tetani, or Bacillus, such as Bacillus thuringiensis. In a specific embodiment, the type II intron is a type II intron contained in the nucleotide sequence of SEQ ID NO: 1 or 2.
[0038] In a specific embodiment, the protein non-coding sequence is selected from one or more of the following groups: a spacer sequence such as any one of SEQ ID NO: 4-6, an A and / or T-rich sequence, a polyA sequence, a polyA-C sequence, a polyC sequence, a polyU sequence, an IRES, a ribosome binding site, an adaptor sequence, an RNA scaffold, a riboswitch, a ribozyme other than a self-splicing ribozyme, a small RNA, a translation regulatory sequence, and a protein binding site.
[0039] In specific embodiments, the polynucleotide construct is capable of forming a circular RNA of a target sequence in vitro.
[0040] In specific embodiments, the polynucleotide construct is capable of forming a circular RNA of a target sequence in vivo.
[0041] In a second aspect, the present invention provides a circular RNA produced by the construct of the first aspect. Preferably, the circular RNA does not contain any other sequence that does not belong to the target sequence, such as E2 and E1 sequences.
[0042] In specific embodiments, such as those in which the target sequence is a protein-coding sequence, the circular RNA is at least 500 nucleotides long, preferably at least 1,000 nucleotides long, and preferably at least 1,500 nucleotides long. In those in which the target sequence is a non-coding RNA, the target sequence may be shorter.
[0043] In a third aspect, the present invention provides a method for expressing a protein in a cell, comprising transfecting the circular RNA of the second aspect into the cell.
[0044] In a fourth aspect, the present invention provides a method for expressing a protein in a cell, comprising subjecting the construct of the first aspect to a self-splicing cyclization reaction to form a circular RNA, and transfecting the circular RNA into the cell.
[0045] In specific embodiments of the third and fourth aspects, the cell is a eukaryotic cell.
[0046] The constructs, methods and uses of the present invention have at least the following advantages: 1. In the preferred technical solution, circular RNA without scar sequence can be produced, which is more conducive to orderly application; 2. During the self-splicing reaction of circular RNA, GTP is not required, and only Mg and Na ions are required; 3. The splicing efficiency of group II introns has been greatly improved, from 10% to about 50%, and even up to 98% splicing efficiency. 4. Illustrative Implementations Group 1 1. A polynucleotide construct having self-splicing activity in vitro, comprising the following operably linked elements from 5' to 3': (1) 3' intron fragment; (2) exon segment 2 (E2); (3) target sequence; (4) exon segment 1 (E1); (5) 5' intron fragment, wherein the 5' intron fragment and the 3' intron fragment are obtained by splitting a group II intron into two fragments, the 5' intron fragment is located at the 5' side of the 3' intron fragment in the group II intron, The E1 is a fragment of the 5' exon of the group II intron, and its length is ≥ 0 nucleotides. The E2 is a fragment of the 3' exon of the group II intron, and its length is ≥ 0 nucleotides. The target sequence is empty, or is a protein coding sequence and / or a non-coding sequence; And wherein the group II intron is BR23 or CL. 2. The polynucleotide construct of clause 1, wherein the length of E1 and / or E2 is 0-20 nucleotides, preferably 0-10 nucleotides, such as 0, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10 nucleotides. 3. The polynucleotide construct of clause 1, wherein the 5' intron fragment and the 3' intron fragment are obtained by splitting a group II intron into two fragments from an unpaired region, and the unpaired region is preferably selected from a linear region between two adjacent domains of a group II intron or a loop region of a stem-loop structure of domain 4. 4. The polynucleotide construct of clause 1, wherein the group II intron comprises one or more nucleotide modifications relative to its wild-type form, the modifications being selected from one or more of deletion, substitution, and addition. 5. The polynucleotide construct of clause 4, wherein said E1 and E2 are 0, and said modification comprises modifying one or more EBS sequences of said group II intron so that said EBS sequences are complementary to one or more regions of corresponding length in said target sequence at at least 60% of the nucleotide positions. 6. The polynucleotide construct of clause 5, wherein the modification is to modify two EBS sequences of the group II intron, such as EBS1 and EBS3, so that the EBS sequences are complementary to two regions of corresponding length in the target sequence at at least 60% of the nucleotide positions; preferably, the two regions are located at both ends of the target sequence. 7. The polynucleotide construct of clause 4, wherein the modification comprises deleting part or all of domain 4, such as deleting the IEP sequence in domain 4, preferably deleting the entire domain 4. 8. The polynucleotide construct of clause 1, wherein the non-coding sequence is selected from the group consisting of: any one of the spacer sequences SEQ ID NO: 4-6, a polyA sequence, a polyA-C sequence, a polyC sequence, a polyU sequence, an IRES, a ribosome binding site, an aptamer sequence, an RNA scaffold, a riboswitch, a ribozyme other than a self-splicing ribozyme, a small RNA binding site, a translation regulatory sequence, and a protein binding site. 9. A circular RNA produced by the polynucleotide construct of any one of clauses 1 to 8. 10. The circular RNA according to Item 9, which does not contain any other sequence that does not belong to the target sequence, such as all or part of the E2 or E1 sequence. 11. A method for expressing a protein in a cell, comprising (a) transfecting the circular RNA of clause 9 or 10 into the cell, or (b) subjecting the construct of any one of clauses 1 to 8 to self-splicing circularization reaction to form a circular RNA, and transfecting the circular RNA into the cell; Preferably, the cell is a eukaryotic cell. Group 2
[0047] Implementation Plan 1
[0048] A polynucleotide construct having self-splicing activity, comprising the following operably linked elements from 5' to 3': (a) 3' intron fragment; (b) exon fragment 2 (E2); (c) target sequence; (d) exon fragment 1 (E1); (e) 5' intron fragment, in: The 5' intron fragment and the 3' intron fragment are each a fragment of a group II intron, wherein the 5' intron fragment is located on the 5' side of the 3' intron fragment in the group II intron, The E1 is a fragment of the 5' exon of the group II intron, and its length is ≥ 0 nucleotides. The E2 is a fragment of the 3' exon of the group II intron, and its length is ≥ 0 nucleotides. The target sequence is empty, or is a protein coding sequence, a non-coding sequence, or a combination of the two; And wherein the group II intron is BR23 or CL.
[0049] Implementation Plan 2
[0050] A polynucleotide construct having self-splicing activity, comprising the following operably linked elements from 5' to 3': (a) 3' intron fragment; (b) exon fragment 2 (E2); (c) linker sequence; (d) target sequence; (e) linker sequence; (f) exon fragment 1 (E1); (g) 5' intron fragment, in: The 5' intron fragment and the 3' intron fragment are each a fragment of a group II intron, wherein the 5' intron fragment is located on the 5' side of the 3' intron fragment in the group II intron, The E1 is a fragment of the 5' exon of the group II intron, and its length is ≥ 0 nucleotides. The E2 is a fragment of the 3' exon of the group II intron, and its length is ≥ 0 nucleotides. The target sequence is empty, or is a protein coding sequence, a non-coding sequence, or a combination of the two; And wherein the group II intron is BR23 or CL.
[0051] Implementation Plan 3
[0052] A polynucleotide construct having self-splicing activity, comprising the following operably linked elements from 5' to 3': (a) 5' homology arm; (b) 3' intron fragment; (c) exon fragment 2 (E2); (d) target sequence; (e) exon fragment 1 (E1); (f) 5' intron fragment; (g) 3' homology arm, in: The 5' intron fragment and the 3' intron fragment are each a fragment of a group II intron, wherein the 5' intron fragment is located on the 5' side of the 3' intron fragment in the group II intron, The E1 is a fragment of the 5' exon of the group II intron, and its length is ≥ 0 nucleotides. The E2 is a fragment of the 3' exon of the group II intron, and its length is ≥ 0 nucleotides. The target sequence is empty, or is a protein coding sequence, a non-coding sequence, or a combination of the two; And wherein the group II intron is BR23 or CL.
[0053] Implementation Plan 4
[0054] A polynucleotide construct having self-splicing activity, comprising the following operably linked elements from 5' to 3': (a) 5' homology arm; (b) 3' intron fragment; (c) exon fragment 2 (E2); (d) linker sequence; (e) target sequence; (f) linker sequence; (g) exon fragment 1 (E1); (h) 5' intron fragment; (i) 3' homology arm, in: The 5' intron fragment and the 3' intron fragment are each a fragment of a group II intron divided into two fragments, the 5' intron fragment is located on the 5' side of the 3' intron fragment in the group II intron, The E1 is a fragment of the 5' exon of the group II intron, and its length is ≥ 0 nucleotides. The E2 is a fragment of the 3' exon of the group II intron, and its length is ≥ 0 nucleotides. The target sequence is empty, or is a protein coding sequence, a non-coding sequence, or a combination of the two; And wherein the group II intron is BR23 or CL.
[0055] Implementation Plan 5
[0056] The polynucleotide construct of any one of embodiments 1-4, wherein the polynucleotide construct has self-splicing activity in vitro.
[0057] Implementation Plan 6
[0058] The polynucleotide construct of any one of embodiments 1-5, wherein the length of E1 and / or E2 is 0-20 nucleotides, preferably 0-10 nucleotides, such as 0, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10 nucleotides.
[0059] Implementation Plan 7
[0060] The polynucleotide construct of any one of embodiments 1-6, wherein the 5' intron fragment and the 3' intron fragment are obtained by splitting a group II intron into two fragments from an unpaired region, for example, an unpaired region of a linear region between two adjacent domains of a group II intron.
[0061] Implementation Plan 8
[0062] The polynucleotide construct of any one of embodiments 1 to 6, wherein the 5' intron fragment and the 3' intron fragment are obtained by splitting the group II intron from the loop region of the domain 1 stem-loop structure.
[0063] Implementation Plan 9
[0064] The polynucleotide construct of any one of embodiments 1 to 6, wherein the 5' intron fragment and the 3' intron fragment are obtained by splitting the group II intron from the loop region of the domain 2 stem-loop structure.
[0065] Implementation Plan 10
[0066] The polynucleotide construct of any one of embodiments 1 to 6, wherein the 5' intron fragment and the 3' intron fragment are obtained by splitting the group II intron from the loop region of the domain 3 stem-loop structure.
[0067] Implementation Plan 11
[0068] The polynucleotide construct of any one of embodiments 1 to 6, wherein the 5' intron fragment and the 3' intron fragment are obtained by splitting the group II intron from the loop region of the domain 4 stem-loop structure.
[0069] Implementation Plan 12
[0070] The polynucleotide construct of any one of embodiments 1 to 6, wherein the 5' intron fragment and the 3' intron fragment are obtained by splitting the group II intron from the loop region of the domain 5 stem-loop structure.
[0071] Implementation Plan 13
[0072] The polynucleotide construct of any one of embodiments 1 to 6, wherein the 5' intron fragment and the 3' intron fragment are obtained by splitting the group II intron from the loop region of the domain 6 stem-loop structure.
[0073] Implementation Plan 14
[0074] The polynucleotide construct of any one of embodiments 1-6, wherein the 5' intron fragment and the 3' intron fragment are obtained by splitting the group II intron from the linear region between domain 1 and domain 2.
[0075] Implementation Plan 15
[0076] The polynucleotide construct of any one of embodiments 1-6, wherein the 5' intron fragment and the 3' intron fragment are obtained by splitting the group II intron from the linear region between domain 2 and domain 3.
[0077] Implementation Plan 16
[0078] The polynucleotide construct of any one of embodiments 1-6, wherein the 5' intron fragment and the 3' intron fragment are obtained by splitting the group II intron from the linear region between domain 3 and domain 4.
[0079] Implementation Plan 17
[0080] The polynucleotide construct of any one of embodiments 1-6, wherein the 5' intron fragment and the 3' intron fragment are obtained by splitting the group II intron from the linear region between domain 4 and domain 5.
[0081] Implementation Plan 18
[0082] The polynucleotide construct of any one of embodiments 1-6, wherein the 5' intron fragment and the 3' intron fragment are obtained by splitting the group II intron from the linear region between domain 5 and domain 6.
[0083] Implementation Plan 19
[0084] The polynucleotide construct of any one of embodiments 1-18, wherein the group II intron comprises one or more nucleotide modifications relative to its wild-type form, the modifications being selected from one or more of deletion, substitution, and addition.
[0085] Implementation Plan 20
[0086] The polynucleotide construct of embodiment 19, wherein the modification comprises modifying one or more EBS sequences of the group II intron, wherein the EBS sequences are complementary to one or more regions of corresponding length in the target sequence at at least 60% of the nucleotide positions.
[0087] Implementation Plan 21
[0088] The polynucleotide construct of embodiment 19, wherein the modification is to modify two EBS sequences of the group II intron, such as EBS1 and EBS3, wherein the EBS sequences are complementary to two regions of corresponding length in the target sequence at at least 60% of the nucleotide positions; preferably, the two regions are located at both ends of the target sequence.
[0089] The polynucleotide construct of embodiment 19, wherein the modification is modification of the EBS1 and / or δ sequence of the group II intron or modification of the EBS1' and / or δ" sequence, wherein the EBS1 and / or δ sequence is complementary to a region of corresponding length in the target sequence at least 60% of the nucleotides, optionally the modification is modification of the EBS1 and / or δ sequence and its upstream sequence, wherein the EBS1 and / or δ sequence and its upstream sequence are complementary to a region of corresponding length in the target sequence at least 60% of the nucleotides. In some embodiments, the region of corresponding length in the target sequence is IBS3, IBS3', IBS3 with a downstream sequence, IBS3' with a downstream sequence. In some embodiments, the δ sequence and its upstream sequence comprise a nucleic acid sequence selected from the following group: (a) wherein the modification is modification of the δ or δ" sequence of the group II intron, wherein the δ or δ" sequence is complementary to a region of corresponding length in the target sequence at at least 60% of the nucleotide positions; preferably, the region is located at one end of the target sequence. In some embodiments, the δ sequence and its upstream sequence comprise SEQ ID NO: 127, (b) SEQ ID NO: 128, (c) SEQ ID NO: 129, (d) SEQ ID NO: 130. In some embodiments, the IBS3 and its downstream comprise a nucleic acid sequence selected from the group consisting of: (a) SEQ ID NO: 131, (b) SEQ ID NO: 132, (c) SEQ ID NO: 133, (d) SEQ ID NO: 134.
[0090] Implementation Plan 22
[0091] The polynucleotide construct of embodiment 19, wherein the modification comprises deleting part or all of domain 4, such as deleting an intron-encoded protein (IEP) sequence in domain 4, preferably deleting the entire domain 4.
[0092] Implementation Plan 23
[0093] The polynucleotide construct of embodiment 19, wherein the modification comprises deleting an open reading frame (ORF).
[0094] Implementation Plan 24
[0095] The polynucleotide construct of any one of embodiments 1-23, wherein the polynucleotide construct is capable of forming a nearly scarless circular RNA of the target sequence.
[0096] Implementation Plan 25
[0097] The polynucleotide construct of embodiment 24, wherein the nearly scarless circular RNA has a scar region of equal to or less than 1 nucleotide, 2 nucleotides, 3 nucleotides, 4 nucleotides, 5 nucleotides, 6 nucleotides, 7 nucleotides, 8 nucleotides, 9 nucleotides, 10 nucleotides, 11 nucleotides, 12 nucleotides, 13 nucleotides, 14 nucleotides, 15 nucleotides, 16 nucleotides, 17 nucleotides, 18 nucleotides, 19 nucleotides, or 20 nucleotides in length.
[0098] Implementation Plan 26
[0099] The polynucleotide construct of any one of embodiments 1-23, wherein the polynucleotide construct is capable of forming a scarless circular RNA of the target sequence.
[0100] Implementation Plan 27
[0101] The polynucleotide construct of any one of embodiments 1-26, wherein E1 and E2 are each 0 nucleotides in length.
[0102] Implementation Plan 28
[0103] The polynucleotide construct of any one of embodiments 1-26, wherein the length of E1 is 0 nucleotides.
[0104] Implementation Plan 29
[0105] The polynucleotide construct of any one of embodiments 1-26, wherein the length of E2 is 0 nucleotides.
[0106] Implementation Plan 30
[0107] The polynucleotide construct of any one of embodiments 1-29, wherein the group II intron is a group II intron derived from a microorganism such as Clostridium tetani or a Bacillus species such as Bacillus thuringiensis.
[0108] Implementation Plan 31
[0109] The polynucleotide construct of any one of embodiments 1-30, wherein the non-coding sequence is selected from the group consisting of a spacer sequence of SEQ ID NOs: 4-6, a polyA sequence, a polyA-C sequence, a polyC sequence, a polyU sequence, an IRES, a ribosome binding site, an aptamer sequence, an RNA scaffold, a riboswitch, a ribozyme other than a self-splicing ribozyme, an antisense oligonucleotide (ASO), a scaffold, a small RNA binding site, a translation regulatory sequence, and a protein binding site.
[0110] Implementation Plan 32
[0111] The polynucleotide construct of any one of embodiments 1-31, wherein the group II intron comprises a nucleic acid sequence selected from the group consisting of: (a) SEQ ID NO: 143; or (b) SEQ ID NO:144.
[0112] Implementation Plan 32-1
[0113] The polynucleotide construct of embodiment 32, wherein the group II intron consists essentially of the following nucleic acid sequence: SEQ ID NO: 143 or SEQ ID NO: 144.
[0114] Implementation Plan 32-2
[0115] The polynucleotide construct of embodiment 32, wherein the group II intron consists of a nucleic acid sequence selected from the group consisting of SEQ ID NO: 143 or SEQ ID NO: 144.
[0116] Implementation Plan 33
[0117] The polynucleotide construct of any one of embodiments 1-32, wherein the polynucleotide construct is an RNA polynucleotide construct.
[0118] Implementation Plan 34
[0119] The polynucleotide construct of embodiment 33, wherein the 3' intron fragment comprises a nucleic acid sequence selected from the group consisting of: (a) a nucleic acid sequence that is at least 95% identical to SEQ ID NO: 145; (b) a nucleic acid sequence that is at least 98% identical to SEQ ID NO: 145; (c) a nucleic acid sequence that is at least 99% identical to SEQ ID NO: 145; (d) SEQ ID NO: 145; (e) a nucleic acid sequence that is at least 95% identical to SEQ ID NO: 146; (f) a nucleic acid sequence that is at least 98% identical to SEQ ID NO: 146; (g) a nucleic acid sequence that is at least 99% identical to SEQ ID NO: 146; (h) SEQ ID NO:146.
[0120] Implementation Plan 34-1
[0121] The polynucleotide construct of embodiment 34, wherein the 3' intron fragment consists essentially of a nucleic acid sequence that is at least 95% identical to SEQ ID NO: 145 or SEQ ID NO: 146, a nucleic acid sequence that is at least 98% identical to SEQ ID NO: 145 or SEQ ID NO: 146, a nucleic acid sequence that is at least 99% identical to SEQ ID NO: 145 or SEQ ID NO: 146, or SEQ ID NO: 145 or SEQ ID NO: 146.
[0122] Implementation Plan 34-2
[0123] The polynucleotide construct of embodiment 34, wherein the 3' intron fragment consists of a nucleic acid sequence that is at least 95% identical to SEQ ID NO: 145 or SEQ ID NO: 146, a nucleic acid sequence that is at least 98% identical to SEQ ID NO: 145 or SEQ ID NO: 146, a nucleic acid sequence that is at least 99% identical to SEQ ID NO: 145 or SEQ ID NO: 146, or SEQ ID NO: 145 or SEQ ID NO: 146.
[0124] Implementation Plan 35
[0125] The polynucleotide construct of embodiment 33 or 34, wherein said E2 comprises a nucleic acid sequence selected from the group consisting of: (a) SEQ ID NO: 53; (b) SEQ ID NO: 54; (c) SEQ ID NO: 55; (d) SEQ ID NO: 56; (e) SEQ ID NO: 57; (f) SEQ ID NO: 58; (g) SEQ ID NO: 59; (h) SEQ ID NO: 60; (i) SEQ ID NO: 61; (j) SEQ ID NO: 62; (k) SEQ ID NO:63.
[0126] Implementation Plan 35-1
[0127] The polynucleotide construct of embodiment 35, wherein said E2 consists essentially of a nucleic acid sequence selected from the group consisting of SEQ ID NO: 53 to SEQ ID NO: 63.
[0128] Implementation Plan 35-2
[0129] The polynucleotide construct of embodiment 35, wherein the E2 consists of a nucleic acid sequence selected from the group consisting of SEQ ID NO: 53 to SEQ ID NO: 63.
[0130] Implementation Plan 36
[0131] The polynucleotide construct of any one of embodiments 33-35, wherein the E1 comprises a nucleic acid sequence selected from the group consisting of: (a) SEQ ID NO: 64; (b) SEQ ID NO: 65; (c) SEQ ID NO: 66; (d) SEQ ID NO: 67; (e) SEQ ID NO: 68; (f) SEQ ID NO: 69; (g) SEQ ID NO: 70; (h) SEQ ID NO: 71; (i) SEQ ID NO: 72; (j) SEQ ID NO: 73; (k) SEQ ID NO:74.
[0132] Implementation Plan 36-1
[0133] The polynucleotide construct of embodiment 36, wherein the E1 consists essentially of a nucleic acid sequence selected from the group consisting of SEQ ID NO: 64 to SEQ ID NO: 74.
[0134] Implementation Plan 36-2
[0135] The polynucleotide construct of embodiment 36, wherein the E1 consists of a nucleic acid sequence selected from the group consisting of SEQ ID NO: 64 to SEQ ID NO: 74.
[0136] Implementation Plan 37
[0137] The polynucleotide construct of any one of embodiments 33-36, wherein the 5' intron fragment comprises a nucleic acid sequence selected from the group consisting of: (a) a nucleic acid sequence that is at least 95% identical to SEQ ID NO: 147; (b) a nucleic acid sequence that is at least 98% identical to SEQ ID NO: 147; (c) a nucleic acid sequence that is at least 99% identical to SEQ ID NO: 147; (d) SEQ ID NO: 147; (e) a nucleic acid sequence that is at least 95% identical to SEQ ID NO: 148; (f) a nucleic acid sequence that is at least 98% identical to SEQ ID NO: 148; (g) a nucleic acid sequence that is at least 99% identical to SEQ ID NO: 148; and (h) SEQ ID NO:148.
[0138] Implementation Plan 37-1
[0139] The polynucleotide construct of embodiment 37, wherein the 5' intron fragment consists essentially of a nucleic acid sequence that is at least 95% identical to SEQ ID NO: 147 or SEQ ID NO: 148, a nucleic acid sequence that is at least 98% identical to SEQ ID NO: 147 or SEQ ID NO: 148, a nucleic acid sequence that is at least 99% identical to SEQ ID NO: 147 or SEQ ID NO: 148, or SEQ ID NO: 147 or SEQ ID NO: 148.
[0140] Implementation Plan 37-2
[0141] The polynucleotide construct of embodiment 37, wherein the 5' intron fragment consists of a nucleic acid sequence that is at least 95% identical to SEQ ID NO: 147 or SEQ ID NO: 148, a nucleic acid sequence that is at least 98% identical to SEQ ID NO: 147 or SEQ ID NO: 148, a nucleic acid sequence that is at least 99% identical to SEQ ID NO: 147 or SEQ ID NO: 148, or SEQ ID NO: 147 or SEQ ID NO: 148.
[0142] Implementation Plan 38
[0143] The polynucleotide construct of any one of embodiments 3-37, wherein the 5' homology arm comprises the nucleic acid sequence of SEQ ID NO: 105.
[0144] Implementation Plan 38-1
[0145] The polynucleotide construct of embodiment 38, wherein the 5' homology arm consists essentially of the nucleic acid sequence of SEQ ID NO: 105.
[0146] Implementation Plan 38-2
[0147] The polynucleotide construct of embodiment 38, wherein the 5' homology arm consists of the nucleic acid sequence of SEQ ID NO: 105.
[0148] Implementation Plan 39
[0149] The polynucleotide construct of any one of embodiments 3-38, wherein the 3' homology arm comprises the nucleic acid sequence of SEQ ID NO: 106.
[0150] Implementation Plan 39-1
[0151] The polynucleotide construct of embodiment 39, wherein the 3' homology arm consists essentially of the nucleic acid sequence of SEQ ID NO: 106.
[0152] Implementation Plan 39-2
[0153] The polynucleotide construct of embodiment 39, wherein the 3' homology arm consists of the nucleic acid sequence of SEQ ID NO: 106.
[0154] Implementation Plan 40
[0155] The polynucleotide construct of any one of embodiments 3-39, wherein the 5' homology arm or the 3' homology arm is 15-60 nucleotides in length.
[0156] Implementation Plan 41
[0157] The polynucleotide construct of any one of embodiments 3-40, wherein the 5' homology arm or 3' homology arm sequence has at most 10% base mismatches.
[0158] Implementation Plan 42
[0159] The polynucleotide construct of any one of embodiments 1-41, wherein the target sequence comprises a 5' arm sequence selected from the group consisting of: (a) SEQ ID NO: 89; (b) SEQ ID NO: 90; (c) SEQ ID NO: 91; (d) SEQ ID NO: 92; (e) SEQ ID NO: 93; (f) SEQ ID NO: 94; (g) SEQ ID NO: 95; (h) SEQ ID NO:96.
[0160] Implementation Plan 43
[0161] The polynucleotide construct of any one of embodiments 1-42, wherein the target sequence comprises a 3' arm sequence selected from the group consisting of: (a) SEQ ID NO: 97; (b) SEQ ID NO: 98; (c) SEQ ID NO: 99; (d) SEQ ID NO: 100; (e) SEQ ID NO: 101; (f) SEQ ID NO: 102; (g) SEQ ID NO: 103; (h) SEQ ID NO: 104.
[0162] Implementation Plan 44
[0163] The polynucleotide construct of any one of embodiments 1-43, wherein the target sequence comprises Formula I: TI-(L) n -Z1(I) in: TI is a modified translation initiation element containing an internal ribosome entry site (IRES)-like polynucleotide sequence or a natural IRES sequence, and Z1 is an expression sequence encoding a therapeutic product; L is the linker sequence; A1 and B1 are a pair of sequences capable of cyclizing the RNA polynucleotide; n is an integer selected from 0-2.
[0164] Implementation Plan 45
[0165] The polynucleotide construct of embodiment 44, wherein Z1 comprises a nucleic acid sequence selected from the group consisting of: (a) a nucleic acid sequence that is at least 95% identical to SEQ ID NO: 107; (b) a nucleic acid sequence that is at least 98% identical to SEQ ID NO: 107; (c) a nucleic acid sequence that is at least 99% identical to SEQ ID NO: 107; (d) SEQ ID NO: 107; (e) a nucleic acid sequence that is at least 95% identical to SEQ ID NO: 108; (f) a nucleic acid sequence that is at least 98% identical to SEQ ID NO: 108; (g) a nucleic acid sequence that is at least 99% identical to SEQ ID NO: 108; (h) SEQ ID NO: 108; (i) a nucleic acid sequence that is at least 95% identical to SEQ ID NO: 109; (j) a nucleic acid sequence that is at least 98% identical to SEQ ID NO: 109; (k) a nucleic acid sequence that is at least 99% identical to SEQ ID NO: 109; (1) SEQ ID NO: 109; (m) a nucleic acid sequence that is at least 95% identical to SEQ ID NO: 110; (n) a nucleic acid sequence that is at least 98% identical to SEQ ID NO: 110; (o) a nucleic acid sequence that is at least 99% identical to SEQ ID NO: 110; (p) SEQ ID NO: 110; (q) a nucleic acid sequence that is at least 95% identical to SEQ ID NO: 111; (r) a nucleic acid sequence that is at least 98% identical to SEQ ID NO: 111; (s) a nucleic acid sequence that is at least 99% identical to SEQ ID NO: 111; (t) SEQ ID NO: 111; (u) a nucleic acid sequence that is at least 95% identical to SEQ ID NO: 112; (v) a nucleic acid sequence that is at least 98% identical to SEQ ID NO: 112; (w) a nucleic acid sequence that is at least 99% identical to SEQ ID NO: 112; (x) SEQ ID NO: 112; (y) a nucleic acid sequence that is at least 95% identical to SEQ ID NO: 149; (z) a nucleic acid sequence that is at least 98% identical to SEQ ID NO: 149; (w) a nucleic acid sequence that is at least 99% identical to SEQ ID NO: 149; (x) SEQ ID NO:149.
[0166] Implementation Plan 45-1
[0167] The polynucleotide construct of embodiment 45, wherein Z1 consists essentially of a nucleic acid sequence selected from the group consisting of a nucleic acid sequence having at least 95% identity to any one of SEQ ID NO: 107 to SEQ ID NO: 112 and SEQ ID NO: 149, a nucleic acid sequence having at least 98% identity to any one of SEQ ID NO: 107 to SEQ ID NO: 112 and SEQ ID NO: 149, a nucleic acid sequence having at least 99% identity to any one of SEQ ID NO: 107 to SEQ ID NO: 112 and SEQ ID NO: 149, and any one of SEQ ID NO: 107 to SEQ ID NO: 112 and SEQ ID NO: 149.
[0168] Implementation Plan 45-2
[0169] The polynucleotide construct of embodiment 45, wherein Z1 consists of a nucleic acid sequence selected from the group consisting of a nucleic acid sequence having at least 95% identity to any one of SEQ ID NO: 107 to SEQ ID NO: 112 and SEQ ID NO: 149, a nucleic acid sequence having at least 98% identity to any one of SEQ ID NO: 107 to SEQ ID NO: 112 and SEQ ID NO: 149, a nucleic acid sequence having at least 99% identity to any one of SEQ ID NO: 107 to SEQ ID NO: 112 and SEQ ID NO: 149, and any one of SEQ ID NO: 107 to SEQ ID NO: 112 and SEQ ID NO: 149.
[0170] Implementation Plan 46
[0171] The polynucleotide construct of embodiment 44, wherein Z1 comprises a nucleic acid sequence encoding an amino acid sequence selected from the group consisting of: (a) SEQ ID NO: 113; (b) SEQ ID NO: 114; (c) SEQ ID NO: 115; (d) SEQ ID NO: 116; (e) SEQ ID NO: 117; (f) SEQ ID NO: 118; and (g) SEQ ID NO:150.
[0172] Implementation Plan 46-1
[0173] The polynucleotide construct of embodiment 46, wherein said Z1 consists essentially of a nucleic acid sequence encoding an amino acid sequence selected from the group consisting of SEQ ID NO: 113 to SEQ ID NO: 118 and SEQ ID NO: 150.
[0174] Implementation Plan 46-2
[0175] The polynucleotide construct of embodiment 46, wherein Z1 consists of a nucleic acid sequence encoding an amino acid sequence selected from the group consisting of SEQ ID NO: 113 to SEQ ID NO: 118 and SEQ ID NO: 150.
[0176] Implementation Plan 47
[0177] The polynucleotide construct of any one of embodiments 1-46, comprising modified RNA nucleotides and / or modified nucleosides.
[0178] Implementation Plan 48
[0179] The polynucleotide construct of any one of embodiments 1-47, comprising 10%-100% modified RNA nucleotides and / or modified nucleosides.
[0180] Implementation Plan 49
[0181] The polynucleotide construct of any one of embodiments 47-48, wherein at least one of the modified RNA nucleotides and / or modified nucleosides is m5C (5-methylcytidine).
[0182] Implementation Plan 50
[0183] The polynucleotide construct of any one of embodiments 47-48, wherein at least one of said modified RNA nucleotides and / or modified nucleosides is m5U (5-methyluridine).
[0184] Implementation Plan 51
[0185] The polynucleotide construct of any one of embodiments 47-48, wherein at least one of the modified RNA nucleotides and / or modified nucleosides is m6A (N6-methyladenosine).
[0186] Implementation Plan 52
[0187] The polynucleotide construct of any one of embodiments 47-48, wherein at least one of the modified RNA nucleotides and / or modified nucleosides is Y (pseudouridine).
[0188] Implementation Plan 53
[0189] The polynucleotide construct of any one of embodiments 47-48, wherein at least one of the modified RNA nucleotides and / or modified nucleosides is mlA (1-methyladenosine).
[0190] Implementation Plan 54
[0191] The polynucleotide construct of any one of embodiments 47-53, wherein at least one of said modified RNA nucleotides and / or modified nucleosides is introduced during in vitro transcription (IVT).
[0192] Implementation Plan 55
[0193] The polynucleotide construct of any one of embodiments 47-48, wherein the modified nucleoside is selected from the group consisting of m5C (5-methylcytidine), m5U (5-methyluridine), m6A (N6-methyladenosine), s2U (2-thiouridine), Y (pseudouridine), Um (2'-O-methyluridine), m1A (1-methyladenosine), m2A (2-methyladenosine), Am (2'-O-methyladenosine), ms2 m6A (2-methylthio-N6-methyladenosine), i6A (N6-isopentenyladenosine), ms2i6A (2-methylthio-N6-isopentenyladenosine), io6A (N6-(cis-hydroxyisopentenyl)adenosine), ms2io6A (2-methylthio-N6-(cis-hydroxyisopentenyl)adenosine), g6A (N6-glycylcarbamoyladenosine), t6A (N6-threonylcarbamoyladenosine), ms2t6A (2-methylthio-N6-threonylcarbamoyladenosine), m6t6A (N6-methyl-N6-threonylcarbamoyladenosine), hn6A (N6-hydroxy Norvalylcarbamoyladenosine), ms2hn6A (2-methylthio-N6-hydroxynorvalylcarbamoyladenosine), Ar(p)(2'-O-ribosyladenosine (phosphate)), I(inosine), m1I(1-methylinosine), m1hn(1,2'-O-dimethylinosine), m3C(3-methylcytidine), Cm(2'-O-methylcytidine), s2C(2-thiocytidine), ac4C(N4-acetylcytidine), (5-formylcytidine), m5Cm(5,2'-O-dimethylcytidine), ac4Cm(N4-acetyl-2'-O-methylcytidine), k2C(lysine), m! G (1-methylguanosine), m2G (N2-methylguanosine), m7G (7-methylguanosine), Gm (2'-0-methylguanosine), m2 2G (N2,N2-dimethylguanosine), m2Gm (N2,2'-O-dimethylguanosine), m2 aGm (N2, N2, 2'-O-trimethylguanosine), Gr(p) (2'-O-ribosylguanosine (phosphate)), yW (hybutosine), oayW (peroxyhybutosine), OHyW (hydroxyhybutosine), OHyW* (undermodified hydroxyhybutosine), imG (hybutosine), mimG (methylhybutosine), Q (braid), oQ (epoxybraid), galQ (galactosyl-braid), manQ (mannosyl-braid), preQo (7-cyano-7-deazaguanosine), preQi (7-aminomethyl-7-deazaguanosine), G+ (archauridine), D (dihydrouridine), m5Um (5,2'-0-dimethyluridine), s4U (4-thiouridine), m5s2U (5-methyl-2-thiouridine), s2Um (2-thio-2'-0-methyluridine), acp3U (3-(3-amino-3-carboxypropyl)uridine), ho5U (5-hydroxyuridine), mo5U (5-methoxyuridine), cmo5U (uridine 5-oxyacetic acid), mcmo5U (uridine 5-oxyacetic acid methyl ester), chm5U (5-(carboxyhydroxymethyl)uridine), mchm5U (5-(carboxyhydroxymethyl)uridine methyl ester), mcm5U (5-methoxycarbonylmethyluridine), mcm5Um (5-methoxycarbonylmethyl-2'-0-methyluridine), m cm5s2U (5-methoxycarbonylmethyl-2-thiouridine), nm5S2U (5-aminomethyl-2-thiouridine), mnm5U (5-methylaminomethyluridine), mnm5s2U (5-methylaminomethyl-2-thiouridine), mnm5se2U (5-methylaminomethyl-2-selenoyluridine), ncm5U (5-carbamoylmethyluridine), ncm5Um (5-carbamoylmethyl-2'-O-methyluridine), cmnm5U (5-carboxymethylaminomethyluridine), cmnm5Um (5-carboxymethylaminomethyl-2'-O-methyluridine), cmnm5s2U (5-carboxymethylaminomethyl-2-thiouridine), m6 2A (N6, N6-dimethyladenosine), Im (2'-0-methylinosine), m4C (N4-methylcytidine), m4Cm (N4, 2'-0-dimethylcytidine), hm5C (5-hydroxymethylcytidine), m3U (3-methyluridine), cm5U (5-carboxymethyluridine), m6Am (N6, 2'-O-dimethyladenosine), m6 2Am (N6, N6, 0-2'-trimethyladenosine), m2,7G (N2, 7-dimethylguanosine), m2,2,7G (N2, N2, 7-trimethylguanosine), m3Um (3,2'-0-dimethyluridine), m5D (5-methyldihydrouridine), f5Cm (5-formyl-2'-0-methylcytidine), m'Gm (l, 2'-0-dimethylguanosine), m'Am (l,2'-0-dimethyladenosine), rm5U (5-tauromethyluridine), τm5s2U (5-tauromethyl-2-thiouridine), imG-14 (4-demethylwyoside), imG2 (isowyoside), ac6A (N6-acetyladenosine), pyridin-4-one ribonucleoside, 5-aza-uridine, 2-thio-5-aza-uridine, 2-thiouridine, 4-thiouridine, 2-thiouridine, 5-hydroxyuridine, 3-methyluridine, 5-carboxymethyl-uridine, 1-carboxymethyl-pseudouridine, 5-propynyl-uridine, 1-propynyl-pseudouridine, 5-tauromethyluridine, 1-tauromethyl-pseudouridine, 5-tauromethyl-2-thiouridine uridine, 1-methyl-pseudouridine, 4-thio-1-methyl-pseudouridine, 2-thio-1-methyl-pseudouridine, 1-methyl-1-deaza-pseudouridine, 2-thio-1-methyl-1-deaza-pseudouridine, dihydrouridine, dihydropseudouridine, 2-thio-dihydrouridine, 2-thio-dihydropseudouridine, 2-methoxyuridine, 2-methoxy-4-thio-uridine, 4-methoxy-pseudouridine, 4-methoxy-2-thio-pseudouridine, 5-azacytidine, pseudoisocytidine, 3-methyl-cytidine, N4-acetylcytidine, 5-formylcytidine, N4-methylcytidine, 5- Hydroxymethylcytidine, 1-methyl-pseudoisocytidine, pyrrolo-cytidine, pyrrolo-pseudoisocytidine, 2-thio-cytidine, 2-thio-5-methyl-cytidine, 4-thio-pseudoisocytidine, 4-thio-1-methyl-pseudoisocytidine, 4-thio-1-methyl-1-deaza-pseudoisocytidine, 1-methyl-1-deaza-pseudoisocytidine, zebularine, 5-aza-zebularine, 5-methyl-zebularine, 5-aza-2-thio-zebularine, 2-thio-zebularine, 2-methoxy-cytidine, 2-methoxy-5-methyl-cytidine, 4-methoxy-pseudoisocytidine, 4-methoxy-1-methyl-pseudoisocytidine , 2-aminopurine, 2,6-diaminopurine, 7-deaza-adenine, 7-deaza-8-aza-adenine, 7-deaza-2-aminopurine, 7-deaza-8-aza-2-aminopurine, 7-deaza-2,6-diaminopurine, 7-deaza-8-aza-2,6-diaminopurine, 1-methyladenosine, N6-methyladenosine, N6-isopentenyladenosine, N6-(cis-hydroxyisopentenyl)adenosine, 2-methylthio-N6-(cis-hydroxyisopentenyl)adenosine, N6-glycylcarbamoyladenosine, N6-threonylcarbamoyladenosine, 2-methylthio-N6-threonylcarbamoyladenosine, N6,N6-dimethyladenosine, 7-methyladenine, 2-methylthio-adenine, 2-methoxy-adenine, inosine, 1-methyl-inosine, wyosine, wyobutine, 7-deaza-guanosine, 7-deaza-8-aza-guanosine, 6-thio-guanosine, 6-thio-7-deaza-guanosine, 6-thio-7-deaza-8-aza-guanosine, 7-methyl-guanosine, 6-thio-7-methyl guanosine, 7-methylinosine, 6-methoxyguanosine, 1-methylguanosine, N2-methylguanosine, N2,N2-dimethylguanosine, 8-oxoguanosine, 7-methyl-8-oxoguanosine, 1-methyl-6-thioguanosine, N2-methyl-6-thioguanosine, N2,N2-dimethyl-6-thioguanosine, 5-methylcytosine, pseudouridine, 1-methylpseudouridine.
[0194] Implementation Plan 56
[0195] The circular RNA produced by the polynucleotide construct of any one of embodiments 1-55, for example, the circular RNA is at least 500 nucleotides in length, at least 1,000 nucleotides in length, or at least 1,500 nucleotides in length.
[0196] Implementation Plan 57
[0197] The circular RNA of embodiment 56 does not contain any other sequence that does not belong to the target sequence, such as all or part of the E2 and E1 sequences.
[0198] Implementation Plan 58
[0199] A method for producing circular RNA, comprising: preparing a vector comprising the following operably linked elements from 5' to 3': (a) 3' intron fragment; (b) exon fragment 2 (E2); (c) target sequence; (d) exon fragment 1 (E1); (e) 5' intron fragment, in: The 5' intron fragment and the 3' intron fragment are each a fragment of a group II intron, wherein the 5' intron fragment is located on the 5' side of the 3' intron fragment in the group II intron, The E1 is a fragment of the 5' exon of the group II intron, and its length is ≥ 0 nucleotides. The E2 is a fragment of the 3' exon of the group II intron, and its length is ≥ 0 nucleotides. The target sequence is empty, or is a protein coding sequence, a non-coding sequence, or a combination of the two; And wherein the group II intron is BR23 or CL.
[0200] Implementation Plan 59
[0201] A method for producing circular RNA, comprising: preparing a vector comprising the following operably linked elements from 5' to 3': (a) 3' intron fragment; (b) exon fragment 2 (E2); (c) linker sequence; (d) target sequence; (e) linker sequence; (f) exon fragment 1 (E1); (g) 5' intron fragment, in: The 5' intron fragment and the 3' intron fragment are each a fragment of a group II intron, wherein the 5' intron fragment is located on the 5' side of the 3' intron fragment in the group II intron, The E1 is a fragment of the 5' exon of the group II intron, and its length is ≥ 0 nucleotides. The E2 is a fragment of the 3' exon of the group II intron, and its length is ≥ 0 nucleotides. The target sequence is empty, or is a protein coding sequence, a non-coding sequence, or a combination of the two; And wherein the group II intron is BR23 or CL.
[0202] Implementation Plan 60
[0203] A method for producing circular RNA, comprising: preparing a vector comprising the following operably linked elements from 5' to 3': (a) 5' homology arm; (b) 3' intron fragment; (c) exon fragment 2 (E2); (d) target sequence; (e) exon fragment 1 (E1); (f) 5' intron fragment; (g) 3' homology arm, in: The 5' intron fragment and the 3' intron fragment are each a fragment of a group II intron, wherein the 5' intron fragment is located on the 5' side of the 3' intron fragment in the group II intron, The E1 is a fragment of the 5' exon of the group II intron, and its length is ≥ 0 nucleotides. The E2 is a fragment of the 3' exon of the group II intron, and its length is ≥ 0 nucleotides. The target sequence is empty, or is a protein coding sequence, a non-coding sequence, or a combination of the two; And wherein the group II intron is BR23 or CL.
[0204] Implementation Plan 61
[0205] A method for producing circular RNA, comprising: preparing a vector comprising the following operably linked elements from 5' to 3': (a) 5' homology arm; (b) 3' intron fragment; (c) exon fragment 2 (E2); (d) linker sequence; (e) target sequence; (f) linker sequence; (g) exon fragment 1 (E1); (h) 5' intron fragment; (i) 3' homology arm, in: The 5' intron fragment and the 3' intron fragment are each a fragment of a group II intron, wherein the 5' intron fragment is located on the 5' side of the 3' intron fragment in the group II intron, The E1 is a fragment of the 5' exon of the group II intron, and its length is ≥ 0 nucleotides. The E2 is a fragment of the 3' exon of the group II intron, and its length is ≥ 0 nucleotides. The target sequence is empty, or is a protein coding sequence, a non-coding sequence, or a combination of the two; And wherein the group II intron is BR23 or CL.
[0206] Implementation Plan 62
[0207] A method for expressing a protein in a cell, comprising (a) transfecting the circular RNA of any one of embodiments 58-61 into the cell, or (b) subjecting the polynucleotide construct of any one of embodiments 1-57 to self-splicing cyclization reaction to form a circular RNA, and transfecting the circular RNA into the cell; wherein, preferably, the cell is a eukaryotic cell.
[0208] Implementation Plan 63
[0209] A method for expressing a protein in a cell, comprising (a) transfecting the circular RNA of any one of embodiments 58-61 into the cell, or (b) allowing the construct of any one of embodiments 1-57 to undergo self-splicing circularization reaction to form a circular RNA, and transfecting the circular RNA into the cell; wherein the cell is preferably a hepatocyte, an epithelial cell, a hematopoietic cell, an epithelial cell, an endothelial cell, a lung cell, a bone cell, a stem cell, a mesenchymal cell, a neural cell (e.g., a meningeal cell, a astrocyte, a motor neuron, a dorsal root ganglion cell, anterior horn motor neuron), a photoreceptor cell (e.g., a rod cell, a cone cell), a retinal pigment epithelial cell, a secretory cell, a cardiac cell, an adipocyte, a vascular smooth muscle cell, a cardiomyocyte, a skeletal muscle cell, a β cell, a pituitary cell, a synovial lining cell, an ovarian cell, a testicular cell, a fibroblast, a B cell, a T cell, a dendritic cell, a macrophage, a reticulocyte, a leukocyte, a granulocyte, a tumor cell, a NK cell, a hepatic stellate cell (starlet cell), a glial cell, a glial cell, a stellate ... cell), HEK293, HEK293T, HeLa, MCF7, PC3, A549, NCI-H727, HCT-116, MCF10A, HPReC, FHC, immortalized cell lines, primary cells, yeast cells, Saccharomyces cerevisiae, Pichia pastoris, bacterial cells, Escherichia coli, insect cells, Spodoptera frugiperda sf9, Mimic Sf9, sf21, Drosophila S2.
[0210] Implementation Plan 64
[0211] The polynucleotide construct, circular RNA or method according to any of the preceding embodiments, wherein the 5' intron fragment and the 3' intron fragment each comprise one or more pairs of complementary sequences. In a preferred embodiment, the length of the complementary sequences is greater than 20 nucleotides.
[0212] Implementation Plan 65
[0213] The polynucleotide construct, circular RNA or method according to any of the preceding embodiments, wherein the 5' intron fragment and / or the 3' intron fragment comprises one or more affinity tag sequences, wherein the affinity tag sequence is selected from the group consisting of a probe binding sequence, an MS2 binding site, a PP7 binding site, and a streptavidin binding site.
[0214] Implementation Plan 66
[0215] The polynucleotide construct, circular RNA or method according to any one of the preceding embodiments, wherein the EBS sequence is selected from one or more of EBS1, EBS2, EBS3, preferably two, more preferably EBS1 and EBS3.
[0216] Implementation Plan 67
[0217] The polynucleotide construct, circular RNA, or method of any of the preceding embodiments, wherein one or more EBS sequences of the group II intron, preferably EBS1 and EBS3, are modified, wherein the EBS sequences are complementary to two regions of corresponding length in the target sequence at at least 60% of the nucleotide positions. In a preferred embodiment, the two regions of corresponding length in the target sequence are located at either end of the target sequence.
[0218] Implementation Plan 68
[0219] The polynucleotide construct, circular RNA or method of any one of the preceding embodiments, wherein the polynucleotide construct is capable of forming a circular RNA of a target sequence in vitro.
[0220] Implementation Plan 69
[0221] The polynucleotide construct, circular RNA or method of any one of the preceding embodiments, wherein the polynucleotide construct is capable of forming a circular RNA of a target sequence in vivo. 5. Description of the Figures
[0222] Embodiments of the present invention are described with reference to the accompanying drawings.
[0223] Figure 1 It is a flow chart introducing the method of the present invention, showing the process starting from natural self-splicing ribozymes, through design and modification, and finally reacting to obtain circular RNA.
[0224] Figure 2 Figures AB illustrate the screening process for group II introns in Example 1. (A) A DNA construct containing a Gluc coding sequence fragment and E1-group II intron (self-splicing ribozyme)-E2 is prepared. Using this DNA construct as a template, linear RNA is produced by in vitro transcription and purified. If the linear RNA produces two fragments of different sizes (the excised intron and the remaining portion of the construct) through in vitro self-splicing, in vitro self-splicing activity is demonstrated, and the group II intron and its flanking E1 and E2 sequences can be used as cRNAzyme precursors for designing cRNAzyme constructs. (B) In vitro self-splicing reaction conditions for screening cRNAzyme precursors.
[0225] Figure 3The figures are gel electrophoresis images of two group II introns confirmed to have self-splicing activity according to the method of Example 1. The names of the group II introns are indicated by three-letter codes on the respective electrophoretograms.
[0226] Figure 4 Figures AC show the design of cRNAzyme constructs using Cte as an example and the comparative experimental results between different schemes. (A) cRNAzyme construct design; (B) Circularization efficiency determined by gel electrophoresis after constructs were generated by segmenting the IICte intron at different positions; (C) Graphs showing the results of experiments verifying the successful formation of circular RNA using different methods.
[0227] Figure 5 AC show the results obtained under different conditions during the optimization process. (A) The circularization efficiency of the cRNAzyme construct was improved by optimizing the reaction conditions and modifying the sequence; (B) Gel electrophoresis results of the circularization products under different reaction conditions using Cte as an exemplary self-splicing ribozyme, with the lower bar graph showing the quantitative circularization efficiency PC% (Circularization efficiency (PC%) = circular / (circular + linear) × 100%); (C) Gel electrophoresis of the circularization products generated by three constructs using Cte as an exemplary ribozyme and Renilla luciferase (Rluc) as the insert fragment, with different spacer sequences added; the lower bar graph shows the quantitative circularization efficiency PC%.
[0228] Figure 6 AB relates to the improved constructs prepared in Example 4 that can eliminate scar sequences. (A) Schematic diagram of the construct structure. (B) Gel electrophoresis and sequencing results of the cyclized products of the three target sequences at different magnesium ion concentrations.
[0229] Figure 7 Shown are gel electrophoresis results of circular RNAs generated when target sequences of varying lengths were inserted.
[0230] Figure 8 Figures AB show the intracellular expression results of circular RNAs with different target sequences generated using the constructs and methods of the present invention. (A) GFP expression was detected by Western blotting after transfection of cells using a "scarless" construct and circular RNA with GFP as the target sequence. (B) Gluc expression was detected by microplate reader after transfection of cells using a "scarless" construct and circular RNA with Gluc as the target sequence.
[0231] Figure 9 This is a schematic diagram of the structure of group II introns.
[0232] Figure 10(a.) Branching pathway, (b.) hydrolysis pathway of group II introns are shown.
[0233] Figure 11 The splicing mechanism of group I and group II introns is shown.
[0234] Figure 12A Schematic diagram of a nearly scarless system designed based on the interactions between IBS1 and EBS1, IBS2 and EBS2, and IBS3 and EBS3. The autocatalytically self-splicing group II intron is split into two fragments at the D4 domain, and a custom exon containing E1, E2, and the target sequence is inserted between the split introns. Arrows indicate the interactions between IBS1 and EBS1, IBS2 and EBS2, and IBS3 and EBS3.
[0235] Figure 12B Schematic diagram of the design of a near-scarless system based on the interaction between δ and IBS3. The autocatalytically self-splicing group II intron is split into two fragments from the D4 domain, and a custom exon containing E1, E2, and the target sequence is inserted between the split introns. Arrows indicate the interactions between IBS1 and EBS1, IBS2 and EBS2, and IBS3 and δ.
[0236] Figure 12C Schematic diagram of a scarless system designed based on the interaction between IBS1' and EBS1. The autocatalytic, self-splicing group II intron is split into two fragments at the D4 domain, and the target sequence is inserted between the split introns. Arrows indicate the interactions between IBS1' and EBS1, and IBS3' and EBS3. IBS1' is a region on the target sequence that functions similarly to IBS1. IBS3' is a region on the target sequence that functions similarly to IBS3.
[0237] Figure 12D Schematic diagram of a scarless system designed based on the interaction between δ and IBS3'. The autocatalytic self-splicing group II intron is split into two fragments from the D4 domain, and the target sequence is inserted between the split introns. Arrows indicate the interactions between IBS1' and EBS1, and IBS3' and δ.
[0238] Figure 13 This is the result of in vitro synthesis of circRNAs and analysis on agarose gels. IBS1 was mutated to prevent self-splicing. After intravenous transfection, circularized RNA was confirmed by poly A tailing and RNase R treatment.
[0239] Figure 14 yes Figure 4 Updated Figure C, showing experimental results verifying successful circular RNA formation by different methods and diagrams of linear and circular RNA constructs.
[0240] Figure 15 Shows the span Figure 14 Sanger sequencing output of RT-PCR of the splicing junction of the circRNA samples depicted in lanes 1 and 3.
[0241] Figure 16 yes Figure 6 Updated Figure B, which shows a map of the construct, gel electrophoresis images of the cyclization products of the three target sequences at different magnesium ion concentrations, and sequencing results.
[0242] Figure 17 Shown are the circularization efficiencies using different spacers in front of the CVB3 IRES.
[0243] Figure 18 Luminescent signals from luciferase protein expression from circular RNA derived from different cRNAzyme variants are shown.
[0244] Figure 19 The figure shows the transfection and translation of circRNAs at different doses. CircRNAs containing two different spacers were gel-purified and transfected at different doses into three cell lines cultured in 24-well plates. Luciferase activity was measured 24 hours after transfection.
[0245] Figure 20 The circularization efficiency of cRNAzyme variant CV4 containing different genes is shown. Different RNAs were transcribed in vitro and circularized in the in vitro transcription reaction. Gene 1: Gluc, Gene 2: EGFP, Gene 3: RBD, Gene 4: Rluc, Gene 5: Fluc, Gene 6: saCAS9.
[0246] Figure 21 Figure 2 shows the results of a time course experiment of circRNA translation. CircRNA encoding the Rluc gene was transfected into 293T cells (500 ng of circRNA was used in each transfection), and luciferase activity was measured 6, 12, and 24 hours after transfection.
[0247] Figure 22A , B shows the production of linear mRNA and circRNA after transfection into cells.
[0248] Figure 23 Comparison of protein production from linear mRNAs and circRNAs generated using the PIE protocol or the novel CirCode system.
[0249] Figure 24HPLC purification of CVB3-Gluc circRNA from a spin column purified sample after IVT is shown. The upper panel is an HPLC chromatogram indicating peaks for precursor, circular, and intronic RNA, respectively. The lower panel shows an agarose gel of the input fraction and the collected fractions.
[0250] Figure 25 Shown are the amounts of cell death caused by transfection of unpurified circRNA and purified circRNA compared to mock transfection.
[0251] Figure 26 Shown is a comparison of unpurified and purified circRNAs that stimulate innate immune responses by inducing RIG-I and IFN-B1.
[0252] Figure 27 Amplification of circRNA generation is demonstrated.
[0253] Figure 28 Shown is the analysis of four batches of CVB3-Gluc circRNA purified from HPLC using capillary electrophoresis using an Agilent 2100 Bioanalyzer.
[0254] Figure 29 Schematic diagram of the CircRNA-LNP complex and the particle size of CircRNAGluc-LNP.
[0255] Figure 30 Shown are Gaussia luciferase activities measured in mouse serum 24 hours after injection of circRNAGluc-LNPs in different formulations.
[0256] Figure 31 Representative IVIS images of BALB / c mice administered 20 ug of CircRNAGluc-LNPs using two formulations via the intramuscular (im) route are shown. Relative luminescence graphs are shown, and the scale of luminescence is indicated.
[0257] Figure 32 HPLC-purified circRNA-RBD and circRNA-RBD dimers were shown and analyzed by capillary electrophoresis (lower left) and agarose gel electrophoresis using an Agilent 2100 bioanalyzer.
[0258] Figure 33 The particle size and encapsulation efficiency of CircRNARBD-LNP complex are shown.
[0259] Figure 34Schematic diagram of the circRNARBD-LNP vaccination process and serum collection schedule in BALB / c mice is shown.
[0260] Figure 35 Results of RBD binding to B cells are shown. A flow cytometry antibody panel was designed to identify naive B cells (CD19+IgD+CD27-), total memory B cells (CD19+CD27+) (including unconverted IgD+ populations and converted IgD- populations), plasma cells (CD19+IgD-CD38+CD27+), and transitional B cells (CD19+IgD dark CD38+). To determine whether LNP-circRNA-RBP vaccination induces activation and expansion of antigen-specific B cells, we measured the frequency of RBD-binding B cells using Alexa 647-labeled RBD (RBD-Alexa 647). We found that a large proportion of RBD-specific lymphocytes (i.e., CD19+CD27+ B lymphocytes) were detected among memory B cells, including converted RBD-specific B cell populations (CD19+CD27+IgD-RBD+) and unconverted RBD-specific memory B cell populations (CD19+CD27+IgD+RBD+).
[0261] Figure 36 Figure 3 Antibody responses elicited by circRNA RBD-LNP vaccination are shown. Serum was collected 2 weeks after boost and assessed for RBD-specific IgG1, IgG2a, and IgG2c by ELISA.
[0262] Figure 37 Inhibition of RBD binding to hACE2 overexpressing cell lines is shown.
[0263] Figure 38 The ratios between IgG2a / IgG1 and IgG2c / IgG1 are shown.
[0264] Figure 39 Neutralizing antibody responses elicited by circRNA RBD-LNP vaccination are shown. Pseudovirus neutralization titers were obtained for sera collected 2 weeks after boosting.
[0265] Figure 40 A schematic diagram of the group II introns is shown with IBS1, IBS2, IBS3, EBS1, EBS2, EBS3 (shown in bold).
[0266] Figure 41 Schematic diagram of the structure of a group II intron with IBS1, IBS2, IBS3, EBS1, EBS2, EBS3, and δ (shown in bold).
[0267] Figures 42A-42CShown is the analysis of in vitro transcribed circRNA species from group II introns BR-23 and CL. Figure 42A Agarose gel electrophoresis results of BR23 and CL circRNA samples are shown, showing distinct bands of precursor, intronic product, and circRNA. Figure 42B A table is provided summarizing the relative percentages of circular RNAs, nick sites, and precursors in BR23 and CL samples based on capillary electrophoresis peak areas. Figure 42C Capillary electrophoresis results are shown, corresponding to BR23 and CL samples, with the peaks of circRNA, precursor, and nick sites marked.
[0268] Figure 43 Western Blot results are shown, demonstrating protein expression of circular RNAs generated by self-splicing of BR-23 and CL-based cRNAzymes. 6. Specific Implementation Methods 6.1. Definitions
[0269] As used herein, "substantially free" with respect to a specified component is used herein to mean that none of the specified component has been purposefully formulated into the composition and / or is present only as a contaminant or in trace amounts. Thus, the total amount of the specified component resulting from any accidental contamination of the composition is well below 0.1%, preferably below 0.05%, and more preferably below 0.01%. Most preferred are compositions in which the specified component is present in an undetectable amount using standard analytical methods.
[0270] As used herein in the specification, "a" or "an" may mean one or more. As used herein in one or more claims, when used in conjunction with the word "comprising," the word "a" or "an" may mean one or more than one.
[0271] As used herein, the term "or" in the claims is used to mean "and / or" unless explicitly stated to refer only to alternatives or the alternatives are mutually exclusive, although this disclosure supports a definition referring only to alternatives and "and / or." As used herein, "another" or "another" can mean at least a second or more.
[0272] As used herein, the term "about" is used to indicate that a value includes the inherent variation of error for the device or method used to determine the value, or the variation that exists between study subjects. In some embodiments, "about" means that the variation is ±5%, ±4%, ±3%, ±2%, ±1%, ±0.5%, ±0.2%, or ±0.1% of the value to which "about" is referred. In some embodiments, "about" means that the variation is ±1%, ±0.5%, ±0.2%, or ±0.1% of the value to which "about" is referred.
[0273] The term "cRNAzyme" is used herein to refer to a linear ribonucleic acid (RNA) that is capable of generating circular RNA via an autocatalytic back-splicing reaction.
[0274] The term "cRNAzyme construct" is a linear RNA construct having cRNAzyme activity.
[0275] The term "EBS" is used herein to refer to an exon binding sequence that interacts (eg, forms a complementary pair) with an intron binding sequence (IBS) in an exon region to trigger splicing upon its own hydroxyl group within the EBS nucleic acid sequence.
[0276] The term "EBS1" is used herein to refer to exon binding sequence 1. Figure 9 and 41 .
[0277] The term "EBS2" is used herein to refer to exon binding sequence 2. Figure 9 and 41 .
[0278] The term "EBS3" is used herein to refer to exon binding sequence 3. Figure 9 and 41 .
[0279] The term "EBS1'" is used herein to refer to a modified EBS1 sequence that interacts with IBS1'. The interaction between EBS1' and IBS1' is similar to the interaction between EBS1 and IBS1. Figure 12C and 12D .
[0280] The term "EBS3'" is used herein to refer to a modified EBS3 sequence that interacts with IBS3'. The interaction between EBS3' and IBS3' is similar to the interaction between EBS3 and IBS3. Figure 12C .
[0281] The term "domain 1" or "D1" is used herein to refer to the stem-loop structure of domain 1 of a group II intron. The term "domain 2" or "D2" is used herein to refer to the stem-loop structure of domain 2 of a group II intron. The term "domain 3" or "D3" is used herein to refer to the stem-loop structure of domain 3 of a group II intron. The term "domain 4" or "D4" is used herein to refer to the stem-loop structure of domain 4 of a group II intron. The term "domain 5" or "D5" is used herein to refer to the stem-loop structure of domain 5 of a group II intron. The term "domain 6" or "D6" is used herein to refer to the stem-loop structure of domain 6 of a group II intron. A stem-loop structure is a type of RNA secondary structure that can be determined by any suitable polynucleotide folding algorithm. Some programs are based on calculations of minimum Gibbs free energy. An example of one such algorithm is mFold, and is described by Zuker and Stiegler (Nucleic Acids Res. 9 (1981), 133-148). Another exemplary folding algorithm is the online web server RNAfold developed by the Institute of Theoretical Chemistry at the University of Vienna using a centroid structure prediction algorithm (e.g., AR Gruber et al., 2008, Cell 106).(1):23-24; and PA Carr and GM Church, 2009, Nature Biotechnology 27(12):1151-62). Additional algorithms can be found in U.S. Provisional Patent Application No. 61 / 836,080 (Attorney Docket No. 44790.11.2022; Broad reference number BI-2013 / 004A), which is incorporated herein by reference. Group II introns primarily consist of six stem-loop structures, known as domains 1-6 (D1-D6). These six domains are arranged sequentially and contain multiple exon-binding sequences (EBSs), such as EBS1, EBS2, and EBS3. These EBSs interact with intron-binding sequences (IBSs) within the exon region, forming complementary pairs and triggering splicing via the hydroxyl groups within the EBS nucleic acid sequence.
[0282] As used herein, the term "group II intron" is used herein to refer to RNA molecules encoded by group II introns that have similar secondary and tertiary structures. Group II intron RNA molecules typically have six domains. See Figure 9 and Figure 41 Domain 4 (also referred to as domain IV) of the group II intron RNA comprises a nucleotide sequence encoding a "group II intron-encoded protein."
[0283] The term "IBS" is used herein to refer to an intron binding sequence, which interacts with an exon binding sequence (EBS) to position splice sites.
[0284] The term "IBS1" is used herein to refer to intron binding sequence 1, which interacts with exon binding sequence 1 (EBS1) to locate splice sites.
[0285] The term "IBS1'" is used herein to refer to a region of a target sequence that functions similarly to IBS1.
[0286] The term "IBS2" is used herein to refer to intron binding sequence 2, which interacts with exon binding sequence 2 (EBS2) to locate splice sites.
[0287] The term "IBS3" is used herein to refer to intron binding sequence 3, which interacts with exon binding sequence 3 (EBS3) to locate splice sites.
[0288] The term "IBS3'" is used herein to refer to a region of a target sequence that functions similarly to IBS3.
[0289] The term "delta" is used herein to refer to a region on domain 1 of a group II intron that is a single nucleotide immediately upstream of EBS1. δ pairs with IBS3 and the interaction between δ and IBS3 is referred to as delta-IBS3 pairing. Figure 41 、 12B .
[0290] The term "δ" (delta") is used herein to refer to a region on domain 1 of a group II intron that is a single nucleotide immediately upstream of EBS1'. δ" pairs with IBS3' and the interaction between δ" and IBS3' is referred to as δ"-IBS3' pairing. See Figure 41 、 12D .
[0291] The term "IVT" is used herein to refer to in vitro transcription, which is a general method for producing RNA in vitro by synthesizing RNA from a DNA template using RNA polymerase, ribonucleotides, and appropriate buffer conditions.
[0292] As used herein, the term "portion," when used with respect to a polypeptide or peptide, refers to a fragment of the polypeptide or peptide. In some embodiments, a "portion" of a polypeptide or peptide retains at least one function and / or activity of the full-length polypeptide or peptide from which it is derived. For example, in some embodiments, if the full-length polypeptide binds a given ligand, then a portion of the full-length polypeptide also binds the same ligand.
[0293] The terms "protein" and "polypeptide" are used interchangeably herein.
[0294] When used with respect to a protein, gene, nucleic acid or polynucleotide in a cell or organism, the term "exogenous" refers to a protein, gene, nucleic acid or polynucleotide that has been introduced into a cell or organism by artificial or natural means; or when used with respect to a cell, the term refers to a cell that has been isolated by artificial or natural means and subsequently introduced into a cell population or organism. An exogenous nucleic acid can be from a different organism or cell, or it can be one or more additional copies of a nucleic acid that naturally occurs within an organism or cell. An exogenous cell can be from a different organism, or it can be from the same organism. As a non-limiting example, an exogenous nucleic acid is a nucleic acid that is in a chromosomal location that is different from its location in a natural cell, or is otherwise flanked by a nucleic acid sequence that is different from a nucleic acid sequence found in nature. The term "exogenous" can be used interchangeably with the term "heterologous".
[0295] "Expression construct" or "expression cassette" is used to refer to a nucleic acid molecule capable of directing transcription. An expression construct includes at least one or more transcriptional control elements (such as a promoter, enhancer, or functionally equivalent structures thereof) that direct gene expression in one or more desired cell types, tissues, or organs. Additional elements, such as transcription termination signals, may also be included.
[0296] A "vector" or "construct" (sometimes called a gene delivery system or gene transfer "vehicle") refers to a macromolecule or complex of molecules that contains a polynucleotide or protein expressed by the polynucleotide to be delivered to a host cell in vitro or in vivo.
[0297] A "plasmid," a common type of vector, is an extrachromosomal DNA molecule that is separate from the chromosomal DNA and capable of replicating independently of the chromosomal DNA. In certain instances, it is circular and double-stranded.
[0298] The terms "nucleic acid sequence," "polynucleotide," and "oligonucleotide" are used interchangeably herein and refer to polymers or oligomers of pyrimidine and / or purine bases, such as cytosine, thymine and uracil, adenine and guanine, respectively (see Albert L. Lehninger, Principles of Biochemistry, 793-800 (Worth Pub. 1982)), unless otherwise specified or the context indicates otherwise. The terms encompass any deoxyribonucleotide, ribonucleotide, or peptide nucleic acid component and any chemical variants thereof, such as methylated, hydroxymethylated, or glycosylated forms of these bases. The composition of the polymer or oligomer may be heterogeneous or homogeneous, may be isolated from naturally occurring sources, or may be produced artificially or synthetically. Furthermore, the nucleic acid may be DNA or RNA, or a mixture thereof, and may exist permanently or transiently in single-stranded or double-stranded form, including homoduplexes, heteroduplexes, and hybrid states. Nucleic acids or nucleic acid sequences can comprise other types of nucleic acid structures, such as, for example, DNA / RNA helices, peptide nucleic acids (PNAs), morpholino nucleic acids (see, for example, Braasch and Corey, Biochemistry, 4 / (14): 4503-4510 (2002) and U.S. Patent No. 5,034,506), locked nucleic acids (LNAs; see Wahlestedt et al., Proc. Natl. Acad. Sci. USA, 97: 5633-5638 (2000)), cyclohexenyl nucleic acids (see Wang, Am. Chem. Soc., 122: 8595-8602 (2000)), and / or ribozymes. The terms "nucleic acid," "nucleic acid sequence," "polynucleotide," and "oligonucleotide" can also encompass chains comprising non-natural nucleotides, modified nucleotides, and / or non-nucleotide building blocks (e.g., "nucleotide analogs") that can exhibit the same function as natural nucleotides. The term "DNA sequence" is used herein to refer to a nucleic acid comprising a series of DNA bases.
[0299] The terms "polypeptide" and "protein" are used interchangeably herein and refer to a polymeric form of amino acids comprising at least two or more consecutive amino acids chemically or biochemically modified or derivatized amino acids. As used herein, the term "peptide" refers to a class of short polypeptides. The term peptide can refer to a polymer of amino acids (natural or non-naturally occurring) having a length of up to about 100 amino acids. For example, the length of a peptide can be from about 1 to about 10, from about 10 to about 25, from about 25 to about 50, from about 50 to about 75, from about 75 to about 100 amino acid residues. In some embodiments, the length of the peptide can be about 100, about 200, about 300, about 400, about 500, about 600, about 700, about 800, about 900, about 1000, about 1250, about 1500, about 1750, about 2000, about 2250, about 2500, about 2750, about 3000, about 3250, about 3500, about 3750, about 4000, about 4250, about 4500, about 4750, about 5000 amino acid residues.
[0300] The nomenclature of nucleotides, nucleic acids, nucleosides, and amino acids used herein is in accordance with the International Union of Pure and Applied Chemistry (IUPAC) standards (see, eg, bioinformatics.org / smsylupac.html).
[0301] The term "identity" when referring to nucleic acid sequences or protein sequences is used to indicate the similarity between two sequences. Sequence similarity or identity can be determined using standard techniques known in the art, including, but not limited to, the local sequence identity algorithm of Smith and Waterman, Adv. Appl. Math. 2, 482 (1981), the sequence identity alignment algorithm of Needleman and Wunsch, J Mol. Biol. 48, 443 (1970), the search for similarity method of Pearson and Lipman, Proc. Natl. Acad. Sci. USA 85, 2444 (1988), computerized implementations of these algorithms (GAP, BESTFIT, FASTA, and TFASTA in the Wisconsin Genetics Software Package of Genetics Computer Group, 575 Science Drive, Madison, Wisconsin), the Best Fit program or test described by Devereux et al., Nucl. Acid Res. 12, 387-395 (1984). Another algorithm is the BLAST algorithm described in Altschul et al., J Mol. Biol. 215, 403-410, (1990) and Karlin et al., Proc. Natl. Acad. Sci. USA 90, 5873-5787 (1993). A particularly useful BLAST program is the WU-BLAST-2 program, which is available from Altschul et al., Methods in Enzymology, 266, 460-480 (1996); blast.wustl / edu / blast / README.html. WU-BLAST-2 uses several search parameters, which are optionally set to default values. The parameters are dynamic values and are established by the program itself based on the composition of the specific sequence and the composition of the specific database for which the target sequence is being searched; however, the values can be adjusted to increase sensitivity. Further, another useful algorithm is gapped BLAST as reported in Altschul et al., (1997) Nucleic Acids Res. 25, 3389-3402. Unless otherwise indicated, percent identity is determined herein using the algorithm available on the internet at: blast.ncbi.nlm.nih.gov / Blast.cgi.
[0302] The terms "internal ribosome entry site," "internal ribosome entry site sequence," "IRES," and "IRES sequence region" are used interchangeably herein and refer to cis-elements of viral or human cellular RNAs (e.g., messenger RNA (mRNA) and / or circRNA) that bypass the canonical eukaryotic cap-dependent translation initiation step. The canonical cap-dependent mechanism used by the vast majority of eukaryotic mRNAs requires an mRNA at the 5' end of the mRNA. 7 The IRES is composed of a G-cap, the initiator Met-tRNAmet, more than a dozen initiation factor proteins, directional scanning, and GTP hydrolysis to position a translation-competent ribosome at the start codon. The IRES is typically composed of a long and highly structured 5-UTR that mediates the binding of the translation initiation complex and catalyzes the formation of functional ribosomes.
[0303] The term "IRES-like sequence" or "internal ribosome entry site-like sequence" refers to a synthetic nucleotide sequence that exhibits the function of a natural IRES. In some embodiments, an IRES-like sequence can recruit ribosomal components to mediate cap-independent translation.
[0304] When referring to a nucleic acid sequence, the terms "coding sequence", "coding sequence region", "coding region" and "CDS" can be used interchangeably herein to refer to the portion of a DNA or RNA sequence that is or can be translated into a protein. The terms "reading frame", "open reading frame" and "ORF" can be used interchangeably herein to refer to a nucleotide sequence that begins with a start codon (e.g., ATG) and ends with a stop codon (e.g., TAA, TAG or TGA) in some embodiments. An open reading frame can contain introns and exons, and therefore, all CDSs are ORFs, but not all ORFs are CDSs.
[0305] The terms "complementary" and "complementarity" refer to the relationship between two nucleic acid sequences or nucleic acid monomers that have the ability to form one or more hydrogen bonds with each other through traditional Watson-Crick base pairing or other non-traditional types of pairing. The degree of complementarity between two nucleic acid sequences can be indicated by the percentage of nucleotides in a nucleic acid sequence that can form hydrogen bonds (e.g., Watson-Crick base pairing) with a second nucleic acid sequence (e.g., about 50%, about 60%, about 70%, about 80%, about 90%, and 100% complementary). Two nucleic acid sequences are "fully complementary" if all consecutive nucleotides of a nucleic acid sequence will form hydrogen bonds with the same number of consecutive nucleotides in a second nucleic acid sequence. Two nucleic acid sequences are "substantially complementary" if the degree of complementarity between the two nucleic acid sequences is at least 60% (e.g., at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 97%, at least 98%, at least 99%, or 100%) over a region of at least 8 nucleotides (e.g., at least 9, at least 10, at least 11, at least 12, at least 13, at least 14, at least 15, at least 16, at least 17, at least 18, at least 19, at least 20, at least 21, at least 22, at least 23, at least 24, at least 25, at least 30, at least 35, at least 40, at least 45, at least 50, or more nucleotides) or if the two nucleic acid sequences hybridize under at least moderate, or in some embodiments, high, stringency conditions. Exemplary moderate stringency conditions include overnight incubation at 37°C in a solution comprising 20% formamide, 5% SSC (150 mM NaCl, 15 mM trisodium citrate), 50 mM sodium phosphate (pH 7.6), 5x Denhardt's solution, 10% dextran sulfate, and 20 mg / ml denatured sheared salmon sperm DNA, followed by washing the filter in 1*SSC at about 37°C-50°C, or substantially similar conditions, e.g., as described in Sambrook, J., Molecular Cloning: A Laboratory Manual, Cold Spring Harbor Laboratory Press; 4th ed. (June 15, 2012).High stringency conditions are those that use, for example, (1) low ionic strength and high temperature for washing, such as 0.015 M sodium chloride / 0.0015 M sodium citrate / 0.1% sodium dodecyl sulfate (SDS) at 50°C; (2) a denaturing agent, such as formamide, e.g., 50% (v / v) formamide with 0.1% bovine serum albumin (BSA) / 0.1% Ficoll / 0.1% polyvinylpyrrolidone (PVP) / 50 mM sodium phosphate buffer, pH 6.5, containing 750 mM sodium chloride and 75 mM sodium citrate, at 42°C during hybridization; or (3) 50% formamide, 5xSSC (0.75 M NaCl, 0.075 M sodium citrate), 50 mM sodium phosphate (pH 6.5), 0.1% bovine serum albumin (BSA) / 0.1% Ficoll / 0.1% polyvinylpyrrolidone (PVP) at 42°C during hybridization. 6.8), 0.1% sodium pyrophosphate, 5x Denhardt's solution, sonicated salmon sperm DNA (50 pg / ml), 0.1% SDS, and 10% dextran sulfate, and washes (i) in 0.2*SSC at 42° C., (ii) in 50% formamide at 55° C., and (iii) in 0.1*SSC (optionally in combination with EDTA) at 55° C. Additional details and explanations of hybridization reaction stringency are provided, for example, in Sambrook, supra, and Ausubel et al., eds., Short Protocols in Molecular Biology, 5th ed., John Wiley & Sons, Inc., Hoboken, NJ (2002).
[0306] The term "hybridization" or "hybridized" when referring to nucleic acid sequences is the association formed between and / or among sequences that have complementarity.
[0307] The term "control elements" collectively refers to promoter regions, polyadenylation signals, transcription termination sequences, upstream regulatory domains, replication origins, internal ribosome entry sites (IRES), enhancers, splice junctions, etc., which together provide for replication, transcription, post-transcriptional processing and translation of coding sequences in recipient cells. Not all of these control elements need to be present as long as the selected coding sequence can be replicated, transcribed and translated in appropriate host cells.
[0308] The term "promoter" is used herein to refer to a nucleotide region comprising a DNA regulatory sequence, wherein the regulatory sequence is derived from a gene that can bind to RNA polymerase and allow the initiation of transcription of a downstream (3' direction) coding sequence. It can comprise genetic elements to which regulatory proteins and molecules can bind, such as RNA polymerase and other transcription factors, to initiate specific transcription of a nucleic acid sequence. The phrases "operably positioned," "operably connected," "under control," and "under transcriptional control" mean that a promoter is in correct functional position and / or orientation relative to a nucleic acid sequence to control transcription initiation and / or expression of the sequence.
[0309] "Enhancer" means a nucleic acid sequence that, when located proximal to a promoter, confers increased transcriptional activity relative to the transcriptional activity produced by the promoter in the absence of the enhancer domain.
[0310] "Operably linked" with respect to nucleic acid molecules means that two or more nucleic acid molecules (eg, a nucleic acid molecule to be transcribed, a promoter, and a functional effector element) are linked in a manner that permits transcription of the nucleic acid molecules.
[0311] The term "homology" refers to the percent identity between the nucleic acid residues of two polynucleotides or the amino acid residues of two polypeptides. The correspondence between one sequence and another can be determined by techniques known in the art. For example, homology can be determined by aligning the sequence information and using readily available computer programs to directly compare the sequence information between the two polypeptides. As determined using the above methods, two polynucleotides (e.g., DNA) or two polypeptide sequences are "substantially homologous" to each other when at least about 80%, preferably at least about 90%, and most preferably at least about 95% of the nucleotides or amino acids are matched, respectively, over a defined length of the molecule.
[0312] The term "scar" refers to a region of a certain length in the circular product that does not include the target sequence. Scarless cirRNA contains a scar sequence of 0 nucleotides. Near-scarless cirRNA contains a scar sequence equal to or less than 20 nucleotides in length.
[0313] "Treatment" or "treatment of a disease or condition" refers to the implementation of a regimen or treatment plan that may include administering one or more drugs or agents to a patient in an effort to alleviate the signs or symptoms of the disease or the recurrence of the disease. Desirable therapeutic effects include a reduction in the rate of disease progression, improvement or alleviation of the disease state, as well as regression, increased survival, improved quality of life, or improved prognosis. Relief or prevention can occur before the onset of signs or symptoms of a disease or condition as well as after their onset. Furthermore, "treating" or "treatment" does not require complete relief of signs or symptoms and does not require a cure.
[0314] As used throughout this application, the terms "therapeutic benefit" or "therapeutically effective" refer to anything that, with respect to the medical treatment of a condition, promotes or enhances the health of a subject. This includes, but is not limited to, a reduction in the frequency, severity, or rate of progression of signs or symptoms of a disease. For example, treatment of cancer can involve, for example, a reduction in tumor size, a reduction in tumor aggressiveness, a reduction in cancer growth rate, or a reduction in metastasis or recurrence rate. Treatment of cancer can also refer to prolonging the survival of a subject suffering from cancer.
[0315] The phrase "pharmaceutically or pharmacologically acceptable" means that the molecular entities and compositions do not produce adverse, allergic or other untoward reactions, as the case may be, when administered to animals (e.g., humans). For animal (e.g., human) administration, it will be understood that the formulations should meet sterility, pyrogenicity, general safety and purity standards as required by, for example, the FDA's Office of Biological Standards.
[0316] As used herein, "pharmaceutically acceptable carrier" includes any and all aqueous biocompatible solvents (e.g., saline solutions, phosphate buffered saline, parenteral vehicles such as sodium chloride, Ringer's dextrose, etc.), antioxidants, preservatives (e.g., antibacterial or antifungal agents, antioxidants, chelating agents, and inert gases), isotonic agents, such similar materials, and combinations thereof, as known to those of ordinary skill in the art. The pH and exact concentrations of the various components of the pharmaceutical composition are adjusted according to well-known parameters.
[0317] As used herein and unless otherwise indicated, the term "about" means within ±10% of a given value or range. In certain embodiments, the term "about" encompasses the exact number recited. 6.2. Ribozymes and Group II Introns
[0318] Ribozymes are themselves RNA nucleic acid molecules. Because such nucleic acid sequences have enzymatic activity, they are called ribozymes. For example, some intron sequences in mitochondria or bacteria can directly catalyze splicing without relying on the spliceosome. These are called "ribozymes with self-splicing activity", "self-splicing ribozymes" or "self-splicing introns". Self-splicing introns that can complete splicing without any protein include type I and type II. As mentioned above, the two types of introns have obvious differences in structure and self-splicing reaction mechanism. See Figure 10 The present invention particularly relates to type II self-splicing introns, or "type II introns" for short.
[0319] The method of using self-splicing ribozymes to prepare circular RNA has the following advantages: 1) Reduce the use of biological and chemical reagents. The use of ribozymes can effectively reduce the contamination of exogenous biological products (such as ligases) and other chemical reagents during the preparation process. When ribozymes are used to catalyze self-splicing, only a few reagents, such as Tris-HCl buffer, Mg ions, sodium ions and GTP, need to be used in the reaction system. In the case of the present invention, due to the use of type II introns, GTP can also be omitted. In contrast, when using ligase to perform a ligation reaction to prepare circular RNA, in addition to the ligase itself, on the one hand, the preservation of the enzyme requires the use of corresponding chemical reagents, such as Tris-HCl buffer, KCl, DTT, EDTA, glycerol, etc., and on the other hand, the reaction system also requires the participation of chemical reagents, such as Mg ions, DTT, ATP, etc. The reduction in the types of reagents can save costs and simplify operations. 2) Ease of use. As mentioned above, the reaction requires only a few reagents; the circularization reaction can be completed in a single step on a PCR instrument by simply adding a buffer containing GTP (for type I self-splicing introns only) and ions to the RNA. In contrast, ligation reactions using RNA ligase require at least additional ligase. 3) Simple design. For circular RNAs with larger molecular weights, such as those containing coding sequences, direct ligation is inefficient and often requires the introduction of exogenous DNA splints. This requires precise pairing of RNA and DNA, increasing the complexity of design and operation.
[0320] Figure 9 A schematic diagram of the secondary structure of group II introns is shown in Figure 2. Figure 9 As shown, group II introns primarily consist of six stem-loop structures, designated domains 1-6 (D1-D6). These six domains are arranged sequentially and contain multiple exon-binding sequences (EBSs), such as EBS1, EBS2, and EBS3. These EBSs interact with intron-binding sequences (IBSs) within the exon region, forming complementary pairs. This triggers splicing by hydroxyl groups within the EBS nucleic acid sequence. This splicing mechanism is closer to splicing mediated by the spliceosome and more similar to splicing in higher organisms.
[0321] In a preferred embodiment, the type II intron is derived from the kingdom Microbial (domain Bacteria). In a specific embodiment, the type II intron is derived from Clostridium, such as Clostridium tetani, or Bacillus, such as Bacillus thuringiensis. As will be appreciated by those skilled in the art, the key to the present invention lies in the design of the constructs and methods, which are applicable to a variety of type II introns. The practice of the present invention is not limited to a specific type of type II intron, as long as the type II intron has self-splicing cyclization activity in vitro, which activity can be confirmed by those skilled in the art through conventional means.
[0322] In some embodiments of the present invention, the group II intron can be a wild-type group II intron or a modified group II intron. The modified group II intron comprises one or more nucleotide substitutions, deletions, and / or additions. Preferably, the modification does not affect the self-splicing activity of the group II intron, particularly the self-splicing activity in vitro. 6.3. Constructs of the Invention
[0323] In the context of the present invention, a naturally occurring self-splicing ribozyme may be referred to as a self-splicing ribozyme or a cRNAzyme precursor, and a rearranged and modified self-splicing ribozyme may be referred to as a cRNAzyme. Furthermore, a cRNAzyme linked to a target sequence, such as a protein-coding sequence or a non-protein-coding sequence, may be referred to as a cRNAzyme construct, i.e., a polynucleotide construct of the present invention.
[0324] Specifically, a sequence (E1-intron-E2) consisting of a natural type II intron and its two flanking exon fragments (E1, E2) is divided into two to form two fragments, namely a first fragment with a structure of E1-5' intron fragment, and a second fragment with a structure of 3' intron fragment-E2. The 5' intron fragment was originally located at the 5' end of the 3' intron fragment and was adjacent to each other. When constructing cRNAzyme, the first and second fragments are interchanged and reconnected. The rearranged sequence structure is "3' intron fragment-E2-E1-5' intron fragment". A sequence having this structure and having self-splicing activity is called cRNAzyme. The self-splicing activity is preferably an activity that undergoes self-splicing and causes the target protein sequence inserted therein to form a circular RNA. The self-splicing activity is preferably an activity that undergoes self-splicing in vitro.
[0325] When a cRNAzyme is used to catalyze the formation of a target protein into a circular RNA, the target protein sequence, including the target protein coding sequence and / or non-coding sequence, is incorporated into the cRNAzyme between E2 and E1, thereby forming a cRNAzyme construct. The cRNAzyme construct can be transcribed into RNA, which then undergoes self-splicing through the cRNAzyme structural elements contained therein, resulting in the formation of a circular RNA containing the target protein sequence.
[0326] In general, the principle of designing cRNAzyme constructs based on group II introns (cRNAzyme precursors) is to retain the maximum cyclization efficiency while keeping the overall length as short as possible. After the self-splicing cyclization reaction occurs, the intron part will be cut out, such as Figure 1As shown. The circular RNA product obtained no longer contains the intron portion. Therefore, the circular RNA product has fewer total nucleotides than the linear cRNAzyme construct structure in which no splicing reaction occurs. Based on this, the circular RNA product and the cRNAzyme construct can be distinguished by agarose gel electrophoresis. In the context of the present invention, the cyclization efficiency (PC) is defined as the percentage of circular RNA to the sum of linear RNA and circular RNA. The specific quantitative method adopts the semi-quantitative method commonly used in the art, and is determined based on the intensity of the bands in the gel electrophoresis diagram.
[0327] Based on the above principles, the length of E1 and / or E2 is preferably no more than 20 nucleotides, for example no more than 10 nucleotides, such as 1, 2, 3, 4, 5, 6, 7, 8, 9 or 10 nucleotides. In a special embodiment, E1 and E2 can be 0.
[0328] Also based on the above principles, the intron sequence in the cRNAzyme construct, such as the 5' intron fragment and / or the 3' intron fragment, and / or the exon sequence, such as E1 and / or E2, can contain one or more nucleotide modifications relative to its naturally occurring wild-type sequence, such as the addition, deletion, or substitution of one or more nucleotides.
[0329] The stem-loop structure is a type of RNA secondary structure that can be determined by any suitable polynucleotide folding algorithm. Some programs are based on the calculation of the minimum Gibbs free energy. An example of such an algorithm is mFold, and is described by Zuker and Stiegler (Nucleic Acids Res.9 (1981), 133-148). Another exemplary folding algorithm is the online web server RNAfold developed by the Institute of Theoretical Chemistry of the University of Vienna using a centroid structure prediction algorithm (e.g., AR Gruber et al., 2008, Cell 106). (1): 23-24; and PA Carr and GM Church, 2009, Nature Biotechnology 27 (12): 1151-62). Additional algorithms can be found in U.S. Provisional Patent Application No. 61 / 836,080 (Agent Docket No. 44790.11.2022; Pan Reference No. BI-2013 / 004A), which is incorporated herein by reference.
[0330] In one embodiment, in order to shorten the sequence, a portion of the sequence or nucleotides can be deleted without affecting the activity. For example, the intron-encoded protein (IEP) sequence in the group II intron domain 4 can be deleted. The IEP sequence or similar structure in domain 4 is present in all group II introns and encodes a protein with reverse transcriptase activity, which can catalyze the intron to act as a reverse transcription factor and move in its genome through an RNA intermediate. This function is required for natural group II introns to perform retrotransposition in the genome, but this function is not required for in vitro transcription, so part or all of this sequence in domain 4 can be deleted in the construct of the present invention.
[0331] E1 and E2 usually need to include an IBS sequence so that they interact with the EBS sequence contained in the intron to achieve self-splicing. In one embodiment of the present invention, the E1 and E2 sequences can be 0. The advantage of this is that the final circularized RNA no longer contains any other sequences except the target sequence. In this case, in order to ensure that the EBS in the intron can still have an "IBS" sequence to pair with it, the EBS sequence of the intron needs to be modified so that it complements and pairs with a sequence in the target sequence, thereby interacting. In other words, a sequence in the target sequence is regarded as "IBS" and interacts with the modified EBS sequence in the intron to ensure the completion of self-splicing.
[0332] Therefore, in one embodiment of the present invention, the group II intron is a modified group II intron, specifically a modified group II intron in the EBS region. The modification can be a substitution of one or more nucleotides, specifically a substitution of one or more nucleotides in the EBS region, so that the modified EBS region is complementary to a region of corresponding length in the target sequence. The expression "complementary pairing" means that the two sequences can be complementary to each other after being transcribed into RNA, and the pairing covers the pairing mode of G and U in RNA. The length of the modified EBS may be 3-20 nucleotides, preferably 5-15 nucleotides, more preferably 6-10 nucleotides, for example 6, 7, 8, 9 or 10 nucleotides.
[0333] The target sequence region that is complementary to the modified EBS can be present at any position in the target sequence, as long as it can be paired with the EBS and form a secondary structure that can promote self-splicing. Generally speaking, the sequences at both ends of the target sequence can be used as the basis for the design of the modified EBS. This is because the sequences at both ends are located in the positions of the original E1 and E2 in the construct, and the positions of E1 and E2 are also the positions of the IBS sequence that originally interacted with the EBS. Therefore, in specific embodiments, the modified EBS regions are modified EBS1 and EBS3 regions. In specific embodiments, the region in the target sequence that is complementary to the modified EBS region is located at the 3' and / or 5' end of the target sequence.
[0334] Since the purpose of this complementary pairing is to ensure interaction between the EBS and the target sequence segment acting as an IBS, a certain degree of mismatch can be tolerated as long as the interaction exists. In some embodiments, the modified EBS region is complementary to a region of a corresponding length in the target sequence at at least 60%, such as at least 70%, at least 80%, at least 90%, at least 95%, or 100% of the nucleotide positions, or has at least 60% identity, such as at least 70%, at least 80%, at least 90%, at least 95%, or 100% identity, with the complementary pairing sequence of a region of a corresponding length in the target sequence.
[0335] In another embodiment, the 5' intron fragment and the 3' intron fragment may contain one or more complementary paired sequences. Such paired sequences shorten the spatial distance between the 5' intron fragment and the 3' intron fragment, thereby promoting the cyclization reaction. In a preferred embodiment, the complementary paired sequences are at least about 20 nucleotides in length.
[0336] The target sequence in the construct can include any sequence that is desired to be prepared into a circular RNA. The target sequence can be a protein coding sequence, a protein non-coding sequence, or a combination of the two. In other words, the target sequence can contain a variety of elements. The protein coding sequence can encode any protein, for example, selected from functional proteins, antigenic proteins, signal peptides, tag proteins, etc.
[0337] For example, the protein non-coding sequence included in the target sequence can be a spacer sequence, such as an AT-rich sequence, which can adjust the flexibility of the sequence. Such a spacer sequence can be located at any position in the target sequence, such as at one end of the target sequence, adjacent to E1 and / or E2.
[0338] For example, the protein non-coding sequence included in the target sequence can be a translation regulatory sequence, such as an internal ribosome entry site (IRES). The IRES that can be used in the present invention can be from any source. In some embodiments, the circular RNA can include the IRES disclosed in WO2023231959 or WO2020186991, the contents of which are incorporated herein by reference in their entirety.
[0339] The cRNAzymes and cRNAzyme constructs of the present invention are prepared intact in the form of DNA, and then transcribe and self-splicing to form the desired circular RNA. The BR23 intron is derived from an uncultured cyanobacterium isolated from porcine intestine (GenBank CAMEFH010000052.1, contig ERZ10407562.16455a). BR23 belongs to subgroup IIB and folds into a canonical six-domain structure (DI-DVI). The branch point adenosine in DVI is retained, which catalyzes the two-step transesterification pathway characteristic of self-splicing group II introns.
[0340] In some embodiments, provided herein is a cRNAzyme derived from the BR23 intron (SEQ ID NO: 139). In order to minimize the size of the cassette while maintaining catalytic ability, the intron is divided into a 5' intron fragment (SEQ ID NO: 147, amino acids 1-633 of SEQ ID NO: 139) and a 3' intron fragment (SEQ ID ID: 145, amino acids 749-845 of SEQ ID ID: 139). The cRNAzyme containing these two intron fragments retains all the tertiary contacts required for post-transcriptional domain-domain docking, thereby allowing the two fragments to reassemble and perform cis self-splicing. Alternative cleavage sites can be placed in the loops of DI, DII, DIII, DIV, DV, DVI or in the short domain-interconnecting linker; as long as the rupture of the excision does not destroy the indispensable secondary structure stem, there is no upper or lower position restriction for the precise connection.
[0341] Because the boundaries between intron domains are structurally tolerated, the junction where the wild-type BR23 intron is segmented can be moved upstream or downstream by approximately ±30 nucleotides without significantly disrupting tertiary contacts. Thus, in some embodiments, the 5' intron segment variant can extend to nucleotide 663 of SEQ ID NO: 139 (i.e., include up to an additional 30 nt), or be truncated to nucleotide 603 of SEQ ID NO: 139 (removing up to 30 nt from the 3' end). Similarly, the 3' intron segment variant can begin as early as nucleotide 719 of SEQ ID NO: 139 (extending by 30 nt) or as late as nucleotide 779 of SEQ ID NO: 139 (shortening by 30 nt). Furthermore, within each segment, the cRNAzyme can tolerate up to approximately 10% interspersed nucleotide substitutions, deletions, or insertions (i.e., ≥90% identity to SEQ ID NO: 145 or SEQ ID NO: 147), provided that: the catalytic AGC triplet in DV and the branch point adenosine in DVI are conserved; the base-paired stems flanking DV and DVI retain canonical Watson-Crick pairing (allowing compensatory mutations); and EBS1 and EBS3 within the 5' segment remain complementary to the corresponding engineered IBS sites in the target. Substitutions in non-conserved peripheral loops, bulges, or linker regions are generally well tolerated.
[0342] In some embodiments, the BR23 intron-based cRNAzyme provided herein comprises a 5' to 3' 3' intron fragment (identity ≥90% with SEQ ID NO: 145), an optional downstream exon fragment (E2), a target sequence, an optional upstream exon fragment (E1), and a 5' intron fragment (identity ≥90% with SEQ ID NO: 147). During transcription, the two intron fragments reconstitute an active ribozyme that self-cleaves and connects E1 to E2, or in the absence of E1 and E2, to the 5' and 3' ends of the target sequence. When the length of E1 and / or E2 is 0-20nt (preferably 0-10nt; most preferably 0nt), the resulting circular RNA has at most one nearly scarless connection, typically 1-3nt, or completely scarless if both E1 and E2 are absent. In some embodiments, the cRNAzyme provided herein comprises a short linker sequence flanking the target sequence. These linkers (1-50 nt, typically 4-12 nt) provide structural flexibility without extending the scar. The cassette can also be flanked by a pair of complementary 5' and 3' homology arms. If desired, linkers can be inserted again on either side of the target.
[0343] The CL intron is located on the circular chromosome of Subdoligranulum variabile (NZ_CP102293.1), an anaerobic, butyrate-producing member of the human gut microbiota that thrives at 37°C. CL belongs to subgroup IIB, folds into a canonical six-domain structure (DI-DVI), and has a branch point adenosine in DVI that drives a two-step transesterification pathway. In some embodiments, provided herein are cRNAzymes derived from the CL intron (SEQ ID NO: 140). In order to minimize the size of the cassette while maintaining catalytic capacity, the intron is divided into a 5' intron fragment (SEQ ID NO: 148, amino acids 1-661 of SEQ ID NO: 140) and a 3' intron fragment (SEQ ID ID: 146, amino acids 807-931 of SEQ ID ID: 140). The cRNAzyme containing these two intronic fragments retains all tertiary contacts required for post-transcriptional domain-domain docking, allowing the two fragments to reassemble and undergo cis-autosplicing. Alternative cleavage sites can be placed in the loops of DI, DII, DIII, DIV, DV, DVI, or in short interdomain linkers; there are no upper or lower position restrictions for the precise connection, as long as the excised break does not disrupt the indispensable secondary structure stem.
[0344] Because the boundaries between intron domains are structurally tolerated, the junction where the wild-type CL intron is segmented can be moved upstream or downstream by approximately ±30 nucleotides without significantly disrupting tertiary contacts. Thus, in some embodiments, the 5' intron segment variant can extend to nucleotide 691 of SEQ ID NO: 140 (i.e., include up to an additional 30 nt), or be truncated to nucleotide 631 of SEQ ID NO: 140 (removing up to 30 nt from the 3' end). Similarly, the 3' intron segment variant can begin as early as nucleotide 777 of SEQ ID NO: 140 (extending by 30 nt) or as late as nucleotide 837 of SEQ ID NO: 140 (shortening by 30 nt). Furthermore, within each segment, the cRNAzyme can tolerate up to approximately 10% interspersed nucleotide substitutions, deletions, or insertions (i.e., ≥90% identity to SEQ ID NO: 146 or SEQ ID NO: 148), provided that: the catalytic AGC triplet in DV and the branch point adenosine in DVI are conserved; the base-paired stems flanking DV and DVI retain canonical Watson-Crick pairing (allowing compensatory mutations); and EBS1 and EBS3 within the 5' segment remain complementary to the corresponding engineered IBS sites in the target. Substitutions in non-conserved peripheral loops, bulges, or linker regions are generally well tolerated.
[0345] In some embodiments, the CL intron-based cRNAzyme provided herein comprises a 5' to 3' 3' intron fragment (identity ≥90% with SEQ ID NO: 146), an optional downstream exon fragment (E2), a target sequence, an optional upstream exon fragment (E1), and a 5' intron fragment (identity ≥90% with SEQ ID NO: 148). During transcription, the two intron fragments reconstitute an active ribozyme that self-cleaves and connects E1 to E2, or in the absence of E1 and E2, to the 5' and 3' ends of the target sequence. When the length of E1 and / or E2 is 0-20nt (preferably 0-10nt; most preferably 0nt), the resulting circular RNA has at most one nearly scarless connection, typically 1-3nt, or completely scarless if both E1 and E2 are absent. In some embodiments, the cRNAzyme provided herein comprises a short linker sequence flanking the target sequence. These linkers (1-50 nt, typically 4-12 nt) provide structural flexibility without extending the scar. The cassette can also be flanked by a pair of complementary 5' and 3' homology arms. If desired, linkers can be inserted again on either side of the target. 6.4. Self-splicing reaction system
[0346] Self-splicing of group II introns requires high salinity and, unlike group I introns, does not require the introduction of GTP.
[0347] In a specific embodiment of the present invention, the self-splicing buffer used in the self-splicing reaction contains 10mM-100mM, such as 10mM, 20mM, 30mM, 40mM, 50mM, 60mM, 70mM, 80mM, 90mM, 100mM divalent magnesium ions, such as MgCl2. The self-splicing buffer may contain 10mM-100mM, such as 10mM, 20mM, 30mM, 40mM, 50mM, 60mM, 70mM, 80mM, 90mM, 100mM NaCl.
[0348] In a preferred embodiment, the self-splicing reaction of the present invention is carried out in vitro for about 5 min to about 1 h, such as about 5 min, about 10 min, about 15 min, about 20 min, about 25 min, about 30 min, about 35 min, about 40 min, about 45 min, about 50 min, about 55 min, or about 1 h. In preferred embodiments, the constructs of the invention are capable of achieving a circularization rate of at least 30%, such as at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, at least 95%. Target sequence
[0349] In some embodiments, the target sequence is empty. In some embodiments, the target sequence is a protein coding sequence. In some embodiments, the target sequence is a non-coding sequence.
[0350] In some embodiments, the target sequence encodes a therapeutic product.
[0351] In specific embodiments, the therapeutic product is a polypeptide, protein, enzyme, or antibody. In specific embodiments, the therapeutic product comprises one or more polypeptides, proteins, enzymes, antibodies, or combinations thereof.
[0352] In specific embodiments, the protein or enzyme is associated with a disease having pathological manifestations traceable to genetic alterations and / or protein dysregulation.
[0353] In specific embodiments, the polypeptide or protein is similar to a weakened or dead form of a pathogenic agent, which can be a microorganism, such as a bacterium, virus, fungus, parasite, or one or more toxins and / or one or more proteins (e.g., surface proteins (i.e., antigens)) of such a microorganism. In specific embodiments, the therapeutic product is an antigen or agent that can stimulate the body's immune system to recognize the agent as a foreign invader, generate antibodies against the agent, destroy the agent, and generate memory for the agent. In specific embodiments, the therapeutic product is an antigen or agent that can induce vaccine-induced memory and / or enable the immune system to act quickly to protect the body from future encounters with any of these agents.
[0354] In some embodiments, the therapeutic product is derived from an infectious agent. In some embodiments, the infectious agent is selected from a member of the group consisting of a viral strain and a bacterial strain.
[0355] In any of the embodiments provided herein, the infectious agent is a strain selected from the group consisting of adenovirus; herpes simplex virus type 1; herpes simplex virus type 2; encephalitis virus, papillomavirus, varicella-zoster virus; Epstein-Barr virus; human cytomegalovirus; human herpes virus type 8; human papillomavirus; BK virus; JC virus; smallpox virus; poliovirus; hepatitis B virus; human bocavirus; parvovirus B19; human astrovirus; norwalk virus; coxsackievirus; hepatitis A virus; poliovirus; rhinovirus; severe acute respiratory syndrome virus; hepatitis C virus; yellow fever virus; dengue virus; West Nile virus; rubella virus; hepatitis E virus; human immunodeficiency virus (HIV) ); influenza virus; Guanarito virus; Junin virus; Lassa virus; Machupo virus; Sabia virus; Crimean-Congo hemorrhagic fever virus; Ebola virus; Marburg virus; measles virus; mumps virus; parainfluenza virus; respiratory syncytial virus (RSV); human metapneumovirus; Hendra virus; Nipah virus; rabies virus; hepatitis D virus; rotavirus; orbivirus; Colorado tick fever virus; Banna virus; human enterovirus; hantavirus; West Nile virus; coronavirus, severe acute respiratory syndrome (SARS)-related coronavirus (SARS-CoV), SARS-CoV-2 virus (associated with COVID-19); Middle East respiratory syndrome coronavirus; Japanese encephalitis virus; vesicular exanthema virus; Eastern equine encephalitis virus; and influenza virus. In some embodiments, the infectious agent is a bacterial strain selected from the group consisting of tuberculosis (Mycobacterium tuberculosis), clindamycin-resistant Clostridium difficile, fluoroquinolone-resistant Clostridium difficile, methicillin-resistant Staphylococcus aureus (MRSA), multidrug-resistant Enterococcus faecalis, multidrug-resistant Enterococcus faecium, multidrug-resistant Pseudomonas aeruginosa, multidrug-resistant Acinetobacter baumannii, and vancomycin-resistant Staphylococcus aureus (VRSA).
[0356] In some embodiments, the infectious agent is associated with birds, pigs, horses, dogs, humans, or non-human primates.
[0357] In some embodiments, the antibodies include but are not limited to monoclonal antibodies, polyclonal antibodies, recombinantly produced antibodies, human antibodies, humanized antibodies, chimeric antibodies, synthetic antibodies, tetrameric antibodies comprising two heavy chains and two light chain molecules, antibody light chain monomers, antibody heavy chain monomers, antibody light chain dimers, antibody heavy chains, antibody heavy chain dimers, antibody light chain-heavy chain pairs, intrabody, heteroconjugate antibodies, monovalent antibodies, antigen-binding fragments of full-length antibodies, and the above-mentioned fusion proteins. Such antigen-binding fragments include but are not limited to single domain antibodies (heavy chain antibodies (VHH) or variable domains of nanobodies), Fab, F(ab')2, and scFv (single chain variable fragments).
[0358] In specific embodiments, the nucleic acids (e.g., polynucleotides) and nucleic acid sequences disclosed herein can be codon-optimized, for example, via any codon optimization technique known to those skilled in the art (see, e.g., Quax et al., 2015, Mol Cell 59:149-161 for review).
[0359] In some embodiments, the target sequence encodes an aptamer sequence. In some embodiments, the target sequence encodes a single-stranded DNA or RNA (ssDNA or ssRNA) molecule that selectively binds to a specific target (including proteins, peptides, carbohydrates, small molecules, toxins, and even living cells).
[0360] In some embodiments, the target sequence encodes a ribozyme, which is a ribonucleic acid (RNA) enzyme that can catalyze a chemical reaction.
[0361] In some embodiments, the target sequence encodes an antisense oligonucleotide (ASO) that binds specifically to the target RNA sequence and regulates protein expression through several different mechanisms.
[0362] In some embodiments, the target sequence encodes Decoy, which is a short sequence that is identical to or has homology to a miRNA binding site or protein binding site in an endogenous target.
[0363] In some embodiments, the target sequence encodes an RNA scaffold, which is an RNA sequence designed to co-localize an enzyme in an engineered biological pathway in vivo through interaction between the protein docking domain of the scaffold and its affinity protein-enzyme fusion. 6.6. Vectors, linear RNA, precursor RNA, and circular RNA
[0364] In some embodiments, the RNA polynucleotides provided herein are single-stranded RNA. In some embodiments, the polynucleotides are linear RNA. In some embodiments, precursor RNAs are provided herein. In some embodiments, the RNA polynucleotides provided are encoded by a vector. In some embodiments, the precursor RNA is a linear RNA produced by in vitro transcription of a vector provided herein.
[0365] In some embodiments, the RNA polynucleotide is a circular RNA or can be used to make a circular RNA polynucleotide. In some embodiments, circular RNA is provided herein. In some embodiments, the circular RNA is a circular RNA produced by a vector provided herein. In some embodiments, the circular RNA is a circular RNA produced by cyclization of a precursor RNA provided herein. circular RNA
[0366] Circular RNA (also known as "circRNA" or "cRNA") is a single-stranded RNA that is linked head-to-tail. CircRNAs have long been considered a ubiquitous class of non-coding RNAs in eukaryotic cells. CircRNAs, typically generated by reverse splicing, have been found to be highly stable.
[0367] In some embodiments, splint connection can be used to generate circular RNA. Splint connection involves the use of an oligonucleotide splint, which hybridizes with both ends of a linear RNA to bring the ends of the linear RNA together for connection. The 5-phosphate and 3-OH at the end of the hybridization-directed RNA of the splint (which can be a deoxyribonucleotide or a ribonucleotide) are connected. As described above, subsequent connection can be performed using chemical or enzymatic techniques. Enzymatic connection can be performed, for example, with T4 DNA ligase (requiring a DNA splint), T4 RNA ligase 1 (requiring an RNA splint), or T4 RNA ligase 2 (DNA or RNA splint). If the structure of the hybridized splint-RNA complex interferes with enzyme activity, chemical connection (such as with BrCN or EDC) is more effective than enzymatic connection in some cases (see, for example, Dolinnaya et al. Nucleic Acids Res, 2 / (23): 5403-5407 (1993); Petkovic et al., Nucleic Acids Res, 43(4): 2454-2465 (2015)).
[0368] In some embodiments, the RNA polynucleotide (e.g., circular RNA) can have any length or size. In some embodiments, the RNA polynucleotide has a length between 300 and 10,000, 400 and 9,000, 500 and 8,000, 600 and 7,000, 700 and 6,000, 800 and 5,000, 900 and 5,000, 1,000 and 5,000, 1,100 and 5,000, 1,200 and 5,000, 1,300 and 5,000, 1,400 and 5,000, and / or 1,500 and 5,000 nucleotides.
[0369] In some embodiments, the length of the RNA polynucleotide (e.g., circular RNA) is at least 300nt, 400nt, 500nt, 600nt, 700nt, 800nt, 900nt, 1000nt, 1100nt, 1200nt, 1300nt, 1400nt, 1500nt, 2000nt, 2500nt, 3000nt, 3500nt, 4000nt, 4500nt, or 5000nt. In some embodiments, the length of the RNA polynucleotide is no more than 3000nt, 3500nt, 4000nt, 4500nt, 5000nt, 6000nt, 7000nt, 8000nt, 9000nt, or 10000nt.
[0370] In some embodiments, the length of the RNA polynucleotide (e.g., circular RNA) is about 300nt, 400nt, 500nt, 600nt, 700nt, 800nt, 900nt, 1000nt, 1100nt, 1200nt, 1300nt, 1400nt, 1500nt, 2000nt, 2500nt, 3000nt, 3500nt, 4000nt, 4500nt, 5000nt, 6000nt, 7000nt, 8000nt, 9000nt or 10000nt.
[0371] In some embodiments, the RNA polynucleotide (e.g., circular RNA) is at least 500, 600, 700, 800, 900, 1000, 1500, 2000, 2500, 3000, 3500, 4000, 4500, 5000, 5500, 6000, 6500, 7000, 7500, 8000, 8500, 9000, 9500, or 10000 nt in length. The RNA polynucleotide (e.g., circular RNA) can be unmodified, partially modified, or fully modified.
[0372] In some embodiments, the circular RNA provided herein has a higher functional stability than an mRNA comprising the same expressed sequence. In some embodiments, the circular RNA provided herein has a higher functional stability than an mRNA comprising the same expressed sequence, 5moU modification, optimized UTR, cap and / or polyA tail.
[0373] In some embodiments, provided herein are circular RNA polynucleotides having a functional half-life of at least 5 hours, 10 hours, 15 hours, 20 hours, 30 hours, 40 hours, 50 hours, 60 hours, 70 hours, or 80 hours. In some embodiments, provided herein are circular RNA polynucleotides having a functional half-life of 5-80, 10-70, 15-60, and / or 20-50 hours. In some embodiments, provided herein are circular RNA polynucleotides having a functional half-life greater than that of an equivalent linear RNA polynucleotide encoding the same protein (e.g., at least 1.5 times, at least 2 times). In some embodiments, functional half-life can be assessed by detecting functional protein synthesis.
[0374] In some embodiments, the circular RNA polynucleotides provided herein have a half-life of at least 5 hours, 10 hours, 15 hours, 20 hours, 30 hours, 40 hours, 50 hours, 60 hours, 70 hours, or 80 hours. In some embodiments, the circular RNA polynucleotides provided herein have a half-life of 5-80, 10-70, 15-60, and / or 20-50 hours. In some embodiments, the half-life of the circular RNA polynucleotides provided herein is greater than the half-life of the equivalent linear RNA polynucleotide encoding the same protein (e.g., at least 1.5 times, at least 2 times).
[0375] In some embodiments, circular RNA provided herein can have an expression amplitude higher than equivalent linear mRNA, for example, 24 hours after RNA is applied to cells, there is a higher expression amplitude. In some embodiments, circular RNA provided herein has an expression amplitude higher than mRNA comprising the same expression sequence, 5moU modified, optimized UTR, cap and / or polyA tail. In some embodiments, circular RNA provided herein can have a stability higher than equivalent linear mRNA. In some embodiments, this can be shown by measuring the presence and density of receptors in vitro or in vivo at the time point of 1 week measurement after electroporation. In some embodiments, this can be shown by measuring the presence of RNA via qPCR or ISH.
[0376] In some embodiments, the circular RNA polynucleotides provided herein comprise modified RNA nucleotides and / or modified nucleosides. In some embodiments, the modified nucleosides are m5 C (5-methylcytidine). In another embodiment, the modified nucleoside is m 5 U (5-methyluridine). In another embodiment, the modified nucleoside is m 6 A(N 6 -methyladenosine). In another embodiment, the modified nucleoside is s 2 U (2-thiouridine). In another embodiment, the modified nucleoside is Y (pseudouridine). In another embodiment, the modified nucleoside is Um (2'-O-methyluridine). In other embodiments, the modified nucleoside is m ! A(1-methyladenosine); m 2 A(2-methyladenosine); Am(2'-0-methyladenosine); ms 2 m 6 A(2-methylthio-N 6 -methyladenosine); i 6 A(N 6 -isopentenyl adenosine); ms2i6A (2-methylthio-N 6 Isopentyl adenosine); 6 A(N 6 -(cis-hydroxyisopentenyl)adenosine); ms 2 io 6 A(2-methylthio-N 6 -(cis-hydroxyisopentenyl)adenosine); g 6 A(N 6 -glycylaminoformyladenosine); t 6 A(N 6 -threonylcarbamoyladenosine); ms 2 t 6 A(2-methylthio-N 6 -threonylcarbamoyladenosine); m 6 t 6 A(N 6 -methyl-N 6 -threonylcarbamoyladenosine); hn 6 A(N 6 -hydroxynorvalylcarbamoyladenosine); ms 2 hn 6 A(2-methylthio-N 6 -hydroxynorvalylcarbamoyladenosine); Ar(p)(2'-O-ribosyladenosine(phosphate)); I(inosine); m 1 I(1-methylinosine);m 1 hn(1,2'-O-dimethylinosine);m 3C(3-methylcytidine); Cm(2'-0-methylcytidine); s 2 C(2-thiocytidine); ac 4 C(N 4 -acetylcytidine); (5-formylcytidine); m 5 Cm(5,2'-O-dimethylcytidine);ac 4 Cm(N 4 -acetyl-2'-O-methylcytidine); k 2 C (lysine); m ! G(1-methylguanosine); m 2 G(N 2 -methylguanosine); m 7 G(7-methylguanosine); Gm(2′-0-methylguanosine); m 2 2G(N 2 ,N 2 -dimethylguanosine); m 2 Gm(N 2 ,2'-O-dimethylguanosine); m 2 aGm(N 2 ,N 2 , 2'-O-trimethylguanosine); Gr(p)(2'-O-ribosylguanosine (phosphate)); yW(whitingoside); oayW(peroxywhitingoside); OHyW(hydroxywhitingoside); OHyW*(undermodified hydroxywhitingoside); imG(whitingoside); mimG(methylwhitingoside); Q(braidoside); oQ(epoxybraidoside); galQ(galactosyl-braidoside); manQ(mannosyl-braidoside); preQo(7-cyano-7-deazaguanosine); preQi(7-aminomethyl-7-deazaguanosine); G + (uridine); D (dihydrouridine); m 5 Um(5,2'-0-dimethyluridine);s 4 U (4-thiouridine); m 5 s 2 U (5-methyl-2-thiouridine); 2 Um (2-thio-2'-0-methyluridine); acp 3 U (3-(3-amino-3-carboxypropyl) uridine); ho 5 U (5-hydroxyuridine); mo 5 U (5-methoxyuridine); cmo 5 U (uridine 5-oxyacetic acid); mcmo 5 U (uridine 5-oxyacetate methyl ester); chm 5 U(5-(carboxyhydroxymethyl)uridine)); mchm 5U (5-(carboxyhydroxymethyl)uridine methyl ester); mcm 5 U (5-methoxycarbonylmethyluridine); mcm 5 Um (5-methoxycarbonylmethyl-2'-0-methyluridine); mcm 5 s 2 U (5-methoxycarbonylmethyl-2-thiouridine); nm 5 S 2 U (5-aminomethyl-2-thiouridine); mnm 5 U (5-methylaminomethyluridine); mnm 5 s 2 U (5-methylaminomethyl-2-thiouridine); mnm 5 se 2 U (5-methylaminomethyl-2-selenoyluridine); ncm 5 U (5-carbamoylmethyluridine); ncm 5 Um (5-carbamoylmethyl-2'-O-methyluridine); cmnm 5 U (5-carboxymethylaminomethyluridine); cmnm 5 Um (5-carboxymethylaminomethyl-2'-0-methyluridine); cmnm 5 s 2 U (5-carboxymethylaminomethyl-2-thiouridine); m 6 2A(N 6 ,N 6 -dimethyladenosine); Im(2'-0-methylinosine); m 4 C(N 4 -methylcytidine); m 4 Cm(N 4 ,2'-0-dimethylcytidine);hm 5 C(5-hydroxymethylcytidine); m 3 U (3-methyluridine); cm 5 U (5-carboxymethyluridine); m 6 Am(N 6 ,2'-O-dimethyladenosine); m 6 2Am(N 6 ,N 6 ,0-2'-trimethyladenosine); m 2,7 G(N 2 ,7-dimethylguanosine); m 2,2,7 G(N 2 ,N 2 ,7-trimethylguanosine); m 3 Um (3,2'-0-dimethyluridine); m 5 D(5-methyldihydrouridine); f5 Cm (5-formyl-2'-0-methylcytidine); m'Gm (1,2'-0-dimethylguanosine); m'Am (1,2'-0-dimethyladenosine); rm 5 U (5-tauromethyluridine); τm5s2U (5-tauromethyl-2-thiouridine); imG-14 (4-demethylwyoside); imG2 (isowyoside); or ac 6 A(N 6 -acetyladenosine).
[0377] In some embodiments, the modified nucleoside may include a compound selected from the group consisting of pyridin-4-one ribonucleoside, 5-aza-uridine, 2-thio-5-aza-uridine, 2-thio-uridine, 4-thio-pseudouridine, 2-thio-pseudouridine, 5-hydroxyuridine, 3-methyluridine, 5-carboxymethyl-uridine, 1-carboxymethyl-pseudouridine, 5-propynyl-uridine, 1-propynyl-pseudouridine, 5-taurinemethyluridine, 1-taurinemethyl-pseudouridine, 5-taurinemethyl-2-thio-uridine, 1-taurinemethyl-4-thio-uridine, 5-methyl-uridine, 1-methyl-pseudouridine. Uridine, 4-thio-1-methyl-pseudouridine, 2-thio-1-methyl-pseudouridine, 1-methyl-1-deaza-pseudouridine, 2-thio-1-methyl-1-deaza-pseudouridine, dihydrouridine, dihydropseudouridine, 2-thio-dihydrouridine, 2-thio-dihydropseudouridine, 2-methoxyuridine, 2-methoxy-4-thio-uridine, 4-methoxy-pseudouridine, 4-methoxy-2-thio-pseudouridine, 5-azacytidine, pseudoisocytidine, 3-methyl-cytidine, N4-acetylcytidine, 5-formylcytidine, N4-methylcytidine, 5-hydroxymethylcytidine, 1-methyl-pseudoisocytidine, pyrrole cytidine, pyrrolo-pseudoisocytidine, 2-thiocytidine, 2-thio-5-methyl-cytidine, 4-thio-pseudoisocytidine, 4-thio-1-methyl-pseudoisocytidine, 4-thio-1-methyl-1-deaza-pseudoisocytidine, 1-methyl-1-deaza-pseudoisocytidine, zebulline, 5-aza-zebulline, 5-methyl-zebulline, 5-aza-2-thio-zebulline, 2-thio-zebulline, 2-methoxy-cytidine, 2-methoxy-5-methyl-cytidine, 4-methoxy-pseudoisocytidine, 4-methoxy-1-methyl-pseudoisocytidine, 2-aminopurine, 2,6-diaminopurine adenosine, 7-deaza-adenine, 7-deaza-8-aza-adenine, 7-deaza-2-aminopurine, 7-deaza-8-aza-2-aminopurine, 7-deaza-2,6-diaminopurine, 7-deaza-8-aza-2,6-diaminopurine, 1-methyladenosine, N6-methyladenosine, N6-isopentenyladenosine, N6-(cis-hydroxyisopentenyl)adenosine, 2-methylthio-N6-(cis-hydroxyisopentenyl)adenosine, N6-glycylcarbamoyladenosine, N6-threonylcarbamoyladenosine, 2-methylthio-N6-threonylcarbamoyladenosine, N6,N6-dimethyladenosine, 7-methyladenine, 2-methylthio-adenine, 2-methoxy-adenine, inosine, 1-methyl-inosine, wyosine, wyobutosine, 7-deaza-guanosine, 7-deaza-8-aza-guanosine, 6-thio-guanosine, 6-thio-7-deaza-guanosine, 6-thio-7-deaza-8-aza-guanosine, 7-methyl-guanosine, 6-thio-7-methyl-guanosine, 7-methylinosine, 6-methoxy-guanosine, 1-methylguanosine, N2-methylguanosine, N2,N2-dimethylguanosine, 8-oxo-guanosine, 7-methyl-8-oxo-guanosine, 1-methyl-6-thio-guanosine, N2-methyl-6-thio-guanosine and N2,N2-dimethyl-6-thio-guanosine. In another embodiment, the modifications are independently selected from the group consisting of 5-methylcytosine, pseudouridine, and 1-methylpseudouridine.
[0378] In some embodiments, the polynucleotide can be codon optimized. A codon-optimized sequence can be a sequence in which codons in a polynucleotide encoding a therapeutic product have been replaced to increase the expression, stability, and / or activity of the therapeutic product. Factors affecting codon optimization include, but are not limited to, one or more of the following: (i) changes in codon bias or synthetically constructed bias tables between two or more organisms or genes; (ii) changes in the degree of codon bias within an organism, gene, or gene set; (iii) systematic changes in codons, including background; (iv) changes in the codons for which tRNAs are decoded; (v) changes in codons based on GC% overall or in one position of a triplet; (vi) changes in the degree of similarity to a reference sequence (e.g., a naturally occurring sequence); (vii) changes in the codon frequency cutoff; (viii) structural properties of mRNA transcribed from a DNA sequence; (ix) prior knowledge of the function of the DNA sequence on which the design of the codon replacement set is based; and / or (x) systematic changes in the codon set for each amino acid. In some embodiments, codon-optimized polynucleotides can minimize ribozyme collisions and / or limit structural interference between expression sequences and the IRES.
[0379] In some embodiments, a polynucleotide construct having self-splicing activity comprises the following operably linked elements from 5' to 3': (a) 3' intron fragment; (b) exon fragment 2 (E2); (c) target sequence; (d) exon fragment 1 (E1); (d) 5' intron fragment. In one embodiment, the 5' intron fragment and the 3' intron fragment are each fragments of a type II intron. In one embodiment, the 5' intron fragment is located on the 5' side of the 3' intron fragment in the type II intron. In one embodiment, the E1 is the 5' adjacent exon fragment of the type II intron, and its length is ≥0 nucleotides. In one embodiment, the E2 is the 3' adjacent exon fragment of the type II intron, and its length is ≥0 nucleotides, and the target sequence is empty, or is a protein coding sequence, a non-coding sequence, or a combination thereof.
[0380] In some embodiments, a polynucleotide construct having self-splicing activity comprises, from 5′ to 3′, the following operably linked elements: (a) a 3′ intron segment; (b) exon fragment 2 (E2); (c) linker sequence; (d) target sequence; (e) linker sequence; (f) exon fragment 1 (E1); (g) 5' intron fragment. In one embodiment, the 5' intron fragment and the 3' intron fragment are each a fragment of a group II intron. In one embodiment, the 5' intron fragment is located on the 5' side of the 3' intron fragment in the group II intron. In one embodiment, the E1 is a 5' adjacent exon fragment of the group II intron, and its length is ≥0 nucleotides. In one embodiment, the E2 is a 3' adjacent exon fragment of the group II intron, and its length is ≥0 nucleotides, and the target sequence is empty, or is a protein coding sequence, a non-coding sequence, or a combination thereof.
[0381] In some embodiments, a polynucleotide construct having self-splicing activity comprises the following operably linked elements from 5' to 3': (a) a 5' homology arm; (b) a 3' intron fragment; (c) exon fragment 2 (E2); (d) a target sequence; (e) exon fragment 1 (E1); (f) a 5' intron fragment; (g) a 3' homology arm. In one embodiment, the 5' intron fragment is located on the 5' side of the 3' intron fragment in the group II intron. In one embodiment, the E1 is the 5' adjacent exon fragment of the group II intron, and its length is ≥0 nucleotides. In one embodiment, the E2 is the 3' adjacent exon fragment of the group II intron, and its length is ≥0 nucleotides, and the target sequence is empty, or is a protein coding sequence, a non-coding sequence, or a combination thereof.
[0382] In some embodiments, a polynucleotide construct having self-splicing activity comprises the following operably linked elements from 5' to 3': (a) a 5' homology arm; (b) 3' intron fragment; (c) exon fragment 2 (E2); (d) linker sequence; (e) target sequence; (f) linker sequence; (g) exon fragment 1 (E1); (h) 5' intron fragment; (i) 3' homology arm. In one embodiment, the 5' intron fragment is located on the 5' side of the 3' intron fragment in the group II intron. In one embodiment, the E1 is the 5' adjacent exon fragment of the group II intron, and its length is ≥0 nucleotides. In one embodiment, the E2 is the 3' adjacent exon fragment of the group II intron, and its length is ≥0 nucleotides, and the target sequence is empty, or is a protein coding sequence, a non-coding sequence or a combination thereof.
[0383] In some embodiments, the polynucleotide construct has self-splicing activity in vitro.
[0384] In some embodiments, the length of E1 and / or E2 is 0-20 nucleotides. In a preferred embodiment, the length of E1 and / or E2 is 0-10 nucleotides. In one embodiment, the length of E1 and / or E2 is 0, 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 nucleotides.
[0385] In some embodiments, the 5' intron fragment and the 3' intron fragment are obtained by splitting a group II intron into two fragments from an unpaired region, for example, an unpaired region of a linear region between two adjacent domains of a group II intron.
[0386] In some embodiments, the 5' intron fragment and the 3' intron fragment are obtained by segmenting a group II intron from the loop region of the domain 1 stem-loop structure.
[0387] In some embodiments, the 5' intron fragment and the 3' intron fragment are obtained by splitting the group II intron from the loop region of the domain 2 stem-loop structure.
[0388] In some embodiments, the 5' intron fragment and the 3' intron fragment are obtained by splitting the group II intron from the loop region of the domain 3 stem-loop structure.
[0389] In some embodiments, the 5' intron fragment and the 3' intron fragment are obtained by splitting a group II intron from the loop region of the domain 4 stem-loop structure.
[0390] In some embodiments, the 5' intron fragment and the 3' intron fragment are obtained by splitting the group II intron from the loop region of the domain 5 stem-loop structure.
[0391] In some embodiments, the 5' intron fragment and the 3' intron fragment are obtained by splitting the group II intron from the loop region of the domain 6 stem-loop structure.
[0392] In some embodiments, the 5' intron fragment and the 3' intron fragment are obtained by splitting the group II intron from the linear region between domain 1 and domain 2.
[0393] In some embodiments, the 5' intron fragment and the 3' intron fragment are obtained by splitting the group II intron from the linear region between domain 2 and domain 3.
[0394] In some embodiments, the 5' intron fragment and the 3' intron fragment are obtained by splitting the group II intron from the linear region between domain 3 and domain 4.
[0395] In some embodiments, the 5' intron fragment and the 3' intron fragment are obtained by splitting the group II intron from the linear region between domain 4 and domain 5.
[0396] In some embodiments, the 5' intron fragment and the 3' intron fragment are obtained by splitting the group II intron from the linear region between domain 5 and domain 6.
[0397] In some embodiments, the group II intron comprises one or more nucleotide modifications relative to its wild-type form, wherein the modification is selected from one or more of deletion, substitution, and addition.
[0398] In some embodiments, the modification comprises modifying one or more EBS sequences of the group II intron, wherein the EBS sequence is complementary to one or more regions of corresponding length in the target sequence at at least 60%, 61%, 62%, 63%, 64%, 65%, 66%, 67%, 68%, 69%, 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 100% of the nucleotide positions, respectively.
[0399] In some embodiments, the modification is to modify two EBS sequences of the group II intron, such as EBS1 and EBS3, wherein the EBS sequences are complementary to two regions of corresponding length in the target sequence at at least 60%, 61%, 62%, 63%, 64%, 65%, 66%, 67%, 68%, 69%, 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 100% of the nucleotide positions; preferably, the two regions are located at both ends of the target sequence.
[0400] In some embodiments, the modification is to modify two EBS sequences of the group II intron, such as EBS1' and EBS3', wherein the EBS sequences are complementary to two regions of corresponding length in the target sequence at at least 60%, 61%, 62%, 63%, 64%, 65%, 66%, 67%, 68%, 69%, 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 100% of the nucleotide positions; preferably, the two regions are located at both ends of the target sequence.
[0401] In some embodiments, the modification is to modify the EBS1 and / or δ sequence of the group II intron or to modify the EBS1' and / or δ" sequence, wherein the EBS1 and / or δ sequence is at least 60%, 61%, 62%, 63%, 64%, 65%, 66%, 67%, 68%, 69%, 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 100%, 101%, 102%, 103%, 104%, 105%, 106%, 107%, 108%, 109%, 110%, 111%, 112%, 113%, 114%, 115%, 116%, 117%, 118%, 119%, 120%, 121%, 122%, 123%, 124%, 125%, 126%, 127%, 128%, 129%, 130%, 131%, 132%, 133%, 134%, 135 93%, 94%, 95%, 96%, 97%, 98%, 99%, 100% of the nucleotides are complementary, optionally the modification is to modify the EBS1 and / or δ sequence and its upstream sequence, wherein the EBS1 and / or δ sequence and its upstream sequence are complementary to the region of corresponding length in the target sequence at least 60% of the nucleotides. In some embodiments, the region of corresponding length in the target sequence is IBS3, IBS3', IBS3 with downstream sequence, IBS3' with downstream sequence. In some embodiments, the δ sequence and its upstream sequence comprise a nucleic acid sequence selected from the group consisting of: (a) SEQ ID NO: 127, (b) SEQ ID NO: 128, (c) SEQ ID NO: 129, (d) SEQ ID NO: 130. In some embodiments, the IBS3 and its downstream comprise a nucleic acid sequence selected from the group consisting of: (a) SEQ ID NO: 131, (b) SEQ ID NO: 132, (c) SEQ ID NO: 133, (d) SEQ ID NO: 134. Figure 6 and 16 .
[0402] In some embodiments, the modification comprises deleting part or all of domain 4, such as deleting an intron-encoded protein (IEP) sequence in domain 4, preferably deleting all of domain 4.
[0403] In some embodiments, the modification comprises deleting an open reading frame (ORF).
[0404] In some embodiments, the polynucleotide construct is capable of forming a nearly scarless circular RNA of the target sequence.
[0405] In some embodiments, the nearly scarless circular RNA has a scar region that is equal to or less than 1 nucleotide, 2 nucleotides, 3 nucleotides, 4 nucleotides, 5 nucleotides, 6 nucleotides, 7 nucleotides, 8 nucleotides, 9 nucleotides, 10 nucleotides, 11 nucleotides, 12 nucleotides, 13 nucleotides, 14 nucleotides, 15 nucleotides, 16 nucleotides, 17 nucleotides, 18 nucleotides, 19 nucleotides, or 20 nucleotides in length.
[0406] In some embodiments, the polynucleotide construct is capable of forming a scarless circular RNA of the target sequence.
[0407] In some embodiments, E1 and E2 are each 0 nucleotides in length. In some embodiments, E1 is 0 nucleotides in length. In some embodiments, E2 is 0 nucleotides in length.
[0408] In some embodiments, the group II intron is a group II intron derived from a microorganism such as Clostridium tetani or a Bacillus species such as Bacillus thuringiensis.
[0409] In some embodiments, the non-coding sequence is selected from the group consisting of a spacer sequence SEQ ID NO: 4-6, a polyA sequence, a polyA-C sequence, a polyC sequence, a polyU sequence, an IRES, a ribosome binding site, an adaptor sequence, an RNA scaffold, a riboswitch, a ribozyme other than a self-splicing ribozyme, an antisense oligonucleotide (ASO), a scaffold, a small RNA binding site, a translation regulatory sequence, and a protein binding site.
[0410] In some embodiments, the group II intron has the following nucleic acid sequence: SEQ ID NO: 143 or SEQ ID NO: 144.
[0411] In some embodiments, the group II intron comprises a nucleic acid sequence that is at least 95% identical to SEQ ID NO: 143; or a nucleic acid sequence that is at least 95% identical to SEQ ID NO: 144.
[0412] In some embodiments, the group II intron comprises a nucleic acid sequence that is at least 98% identical to SEQ ID NO: 143; or a nucleic acid sequence that is at least 98% identical to SEQ ID NO: 144.
[0413] In some embodiments, the group II intron comprises a nucleic acid sequence that is at least 99% identical to SEQ ID NO: 143; or a nucleic acid sequence that is at least 99% identical to SEQ ID NO: 144.
[0414] In some embodiments, the group II intron consists essentially of the following nucleic acid sequence: SEQ ID NO: 143 or SEQ ID NO: 144.
[0415] In some embodiments, the group II intron consists of the following nucleic acid sequence: SEQ ID NO: 143-SEQ ID NO: 144.
[0416] In some embodiments, the group II intron consists of the following nucleic acid sequence: SEQ ID NO: 143-SEQ ID NO: 144.
[0417] In some embodiments, the polynucleotide construct is an RNA polynucleotide construct.
[0418] In some embodiments, the 3' intron fragment comprises a nucleic acid sequence selected from the group consisting of: (a) a nucleic acid sequence that is at least 95% identical to SEQ ID NO: 145; (b) a nucleic acid sequence that is at least 98% identical to SEQ ID NO: 145; (c) a nucleic acid sequence that is at least 99% identical to SEQ ID NO: 145; (d) SEQ ID NO: 145; (e) a nucleic acid sequence that is at least 95% identical to SEQ ID NO: 146; (f) a nucleic acid sequence that is at least 98% identical to SEQ ID NO: 146; (g) a nucleic acid sequence that is at least 99% identical to SEQ ID NO: 146; and (h) SEQ ID NO: 146.
[0419] In some embodiments, the 3' intron fragment consists essentially of a nucleic acid sequence selected from the group consisting of a nucleic acid sequence that is at least 95% identical to SEQ ID NO: 145 or SEQ ID NO: 146, a nucleic acid sequence that is at least 98% identical to SEQ ID NO: 145 or SEQ ID NO: 146, a nucleic acid sequence that is at least 99% identical to SEQ ID NO: 145 or SEQ ID NO: 146, and SEQ ID NO: 145 or SEQ ID NO: 146.
[0420] In some embodiments, the 3' intron fragment consists of a nucleic acid sequence selected from the group consisting of a nucleic acid sequence that is at least 95% identical to SEQ ID NO: 145 or SEQ ID NO: 146, a nucleic acid sequence that is at least 98% identical to SEQ ID NO: 145 or SEQ ID NO: 146, a nucleic acid sequence that is at least 99% identical to SEQ ID NO: 145 or SEQ ID NO: 146, and SEQ ID NO: 145 or SEQ ID NO: 146.
[0421] In some embodiments, the E2 comprises a nucleic acid sequence selected from the group consisting of: (a) SEQ ID NO: 53; (b) SEQ ID NO: 54; (c) SEQ ID NO: 55; (d) SEQ ID NO: 56; (e) SEQ ID NO: 57; (f) SEQ ID NO: 58; (g) SEQ ID NO: 59; (h) SEQ ID NO: 60; (i) SEQ ID NO: 61; (j) SEQ ID NO: 62; (k) SEQ ID NO: 63.
[0422] In some embodiments, the E2 consists essentially of a nucleic acid sequence selected from the group consisting of SEQ ID NO:53-SEQ ID NO:63.
[0423] In some embodiments, the E2 consists of a nucleic acid sequence selected from the group consisting of SEQ ID NO:53-SEQ ID NO:63.
[0424] In some embodiments, the E1 comprises a nucleic acid sequence selected from the group consisting of SEQ ID NO:64; SEQ ID NO:65; SEQ ID NO:66; SEQ ID NO:67; SEQ ID NO:68; SEQ ID NO:69; SEQ ID NO:70; SEQ ID NO:71; SEQ ID NO:72; SEQ ID NO:73; SEQ ID NO:74.
[0425] In some embodiments, the E1 consists essentially of a nucleic acid sequence selected from the group consisting of SEQ ID NO: 64-SEQ ID NO: 74.
[0426] In some embodiments, the E1 consists of a nucleic acid sequence selected from the group consisting of SEQ ID NO: 64 to SEQ ID NO: 74. In some embodiments, the 5' intron fragment comprises a nucleic acid sequence selected from the group consisting of: (a) a nucleic acid sequence having at least 95% identity to SEQ ID NO: 147; (b) a nucleic acid sequence having at least 98% identity to SEQ ID NO: 147;
[0427] (c) a nucleic acid sequence that is at least 99% identical to SEQ ID NO: 147; (d) SEQ ID NO: 147; (e) a nucleic acid sequence that is at least 95% identical to SEQ ID NO: 148; (f) a nucleic acid sequence that is at least 98% identical to SEQ ID NO: 148; (g) a nucleic acid sequence that is at least 99% identical to SEQ ID NO: 148; and (h) SEQ ID NO: 148.
[0428] In some embodiments, the 5' intron fragment consists essentially of a nucleic acid sequence selected from the group consisting of a nucleic acid sequence that is at least 95% identical to SEQ ID NO: 147 or SEQ ID NO: 148, a nucleic acid sequence that is at least 98% identical to SEQ ID NO: 147 or SEQ ID NO: 148, a nucleic acid sequence that is at least 99% identical to SEQ ID NO: 147 or SEQ ID NO: 148, and SEQ ID NO: 147 or SEQ ID NO: 148.
[0429] In some embodiments, the 5' intron fragment consists of a nucleic acid sequence selected from the group consisting of a nucleic acid sequence that is at least 95% identical to SEQ ID NO: 147 or SEQ ID NO: 148, a nucleic acid sequence that is at least 98% identical to SEQ ID NO: 147 or SEQ ID NO: 148, a nucleic acid sequence that is at least 99% identical to SEQ ID NO: 147 or SEQ ID NO: 148, and SEQ ID NO: 147 or SEQ ID NO: 148.
[0430] In some embodiments, the 5' homology arm comprises a nucleic acid sequence of SEQ ID NO: 105. In some embodiments, the 5' homology arm comprises a nucleic acid sequence that is at least 95% identical to SEQ ID NO: 105. In some embodiments, the 5' homology arm comprises a nucleic acid sequence that is at least 98% identical to SEQ ID NO: 105. In some embodiments, the 5' homology arm comprises a nucleic acid sequence that is at least 99% identical to SEQ ID NO: 105.
[0431] In some embodiments, the 5' homology arm consists essentially of the nucleic acid sequence of SEQ ID NO:105.
[0432] In some embodiments, the 5' homology arm consists of the nucleic acid sequence of SEQ ID NO:105.
[0433] In some embodiments, the 3' homology arm comprises a nucleic acid sequence of SEQ ID NO: 106. In some embodiments, the 3' homology arm comprises a nucleic acid sequence that is at least 95% identical to SEQ ID NO: 106. In some embodiments, the 3' homology arm comprises a nucleic acid sequence that is at least 98% identical to SEQ ID NO: 106. In some embodiments, the 3' homology arm comprises a nucleic acid sequence that is at least 99% identical to SEQ ID NO: 106.
[0434] In some embodiments, the 3' homology arm consists essentially of the nucleic acid sequence of SEQ ID NO: 106.
[0435] In some embodiments, the 3' homology arm consists of the nucleic acid sequence of SEQ ID NO:106.
[0436] In some embodiments, the length of the 5' homology arm or the 3' homology arm is 15-60 nucleotides. In some embodiments, the length of the 5' homology arm or the 3' homology arm is 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60 nucleotides.
[0437] In some embodiments, the 5' homology arm or 3' homology arm sequence has at most 1%, 2%, 3%, 4%, 5%, 6%, 7%, 8%, 9%, 10% base mismatches.
[0438] In some embodiments, the target sequence comprises a 5' arm sequence selected from the group consisting of: (a) SEQ ID NO:89; (b) SEQ ID NO:90; (c) SEQ ID NO:91; (d) SEQ ID NO:92; (e) SEQ ID NO:93; (f) SEQ ID NO:94; (g) SEQ ID NO:95; (h) SEQ ID NO:96.
[0439] In some embodiments, the target sequence comprises a 3' arm sequence selected from the group consisting of: (a) SEQ ID NO:97; (b) SEQ ID NO:98; (c) SEQ ID NO:99; (d) SEQ ID NO:100; (e) SEQ ID NO:101; (f) SEQ ID NO:102; (g) SEQ ID NO:103; (h) SEQ ID NO:104.
[0440] In some embodiments, the target sequence comprises Formula I: TI-(L)n-Z1(I), wherein: TI is a modified translation initiation element comprising an internal ribosome entry site (IRES)-like polynucleotide sequence or a natural IRES sequence, Z1 is an expression sequence encoding a therapeutic product; L is a linker sequence; A1 and B1 are a pair of sequences capable of cyclizing the RNA polynucleotide; and n is an integer selected from 0-2.
[0441] In some embodiments, Z1 comprises a nucleic acid sequence selected from the group consisting of: a nucleic acid sequence having at least 95% identity to SEQ ID NO: 107; a nucleic acid sequence having at least 98% identity to SEQ ID NO: 107; a nucleic acid sequence having at least 99% identity to SEQ ID NO: 107; SEQ ID NO: 107; a nucleic acid sequence having at least 95% identity to SEQ ID NO: 108; a nucleic acid sequence having at least 98% identity to SEQ ID NO: 108; a nucleic acid sequence having at least 99% identity to SEQ ID NO: 108; SEQ ID NO: 108; a nucleic acid sequence having at least 95% identity to SEQ ID NO: 109; a nucleic acid sequence having at least 98% identity to SEQ ID NO: 109; a nucleic acid sequence having at least 99% identity to SEQ ID NO: 109; SEQ ID NO: 109; a nucleic acid sequence having at least 95% identity to SEQ ID NO: 110; a nucleic acid sequence having at least 95% identity to SEQ ID NO: 111; a nucleic acid sequence having at least 98% identity to SEQ ID NO: 112; a nucleic acid sequence having at least 99% identity to SEQ ID NO: 113. NO:110 has at least 98% identity with a nucleic acid sequence; a nucleic acid sequence having at least 99% identity with SEQ ID NO:110; SEQ ID NO:110; a nucleic acid sequence having at least 95% identity with SEQ ID NO:111; a nucleic acid sequence having at least 98% identity with SEQ ID NO:111; a nucleic acid sequence having at least 99% identity with SEQ ID NO:111; SEQ ID NO:111; a nucleic acid sequence having at least 95% identity with SEQ ID NO:112; a nucleic acid sequence having at least 98% identity with SEQ ID NO:112; a nucleic acid sequence having at least 99% identity with SEQ ID NO:112; SEQ ID NO:112; a nucleic acid sequence having 95% identity with SEQ ID NO:149; a nucleic acid sequence having 98% identity with SEQ ID NO:149; a nucleic acid sequence having 99% identity with SEQ ID NO:149; SEQ ID NO:149.
[0442] In some embodiments, Z1 consists essentially of a nucleic acid sequence selected from the group consisting of a nucleic acid sequence that is at least 95% identical to any one of SEQ ID NO:107-SEQ ID NO:112 and SEQ ID NO:149, a nucleic acid sequence that is at least 98% identical to any one of SEQ ID NO:107-SEQ ID NO:112 and SEQ ID NO:149, a nucleic acid sequence that is at least 99% identical to any one of SEQ ID NO:107-SEQ ID NO:112 and SEQ ID NO:149, and any one of SEQ ID NO:107-SEQ ID NO:112 and SEQ ID NO:149.
[0443] In some embodiments, Z1 consists of a nucleic acid sequence selected from the following groups: a nucleic acid sequence having at least 95% identity with any one of SEQ ID NO:107-SEQ ID NO:112 and SEQ ID NO:149, a nucleic acid sequence having at least 98% identity with any one of SEQ ID NO:107-SEQ ID NO:112 and SEQ ID NO:149, a nucleic acid sequence having at least 99% identity with any one of SEQ ID NO:107-SEQ ID NO:112 and SEQ ID NO:149, and any one of SEQ ID NO:107-SEQ ID NO:112 and SEQ ID NO:149.
[0444] In some embodiments, Z1 comprises a nucleic acid sequence encoding an amino acid sequence selected from the group consisting of: (a) SEQ ID NO: 113; (b) SEQ ID NO: 114; (c) SEQ ID NO: 115; (d) SEQ ID NO: 116; (e) SEQ ID NO: 117; (f) SEQ ID NO: 118; and SEQ ID NO: 150.
[0445] In some embodiments, Z1 consists essentially of a nucleic acid sequence encoding an amino acid sequence selected from the group consisting of SEQ ID NO: 113-SEQ ID NO: 118 and SEQ ID NO: 150.
[0446] In some embodiments, Z1 consists of a nucleic acid sequence encoding an amino acid sequence selected from the group consisting of SEQ ID NO: 113-SEQ ID NO: 118 and SEQ ID NO: 150.
[0447] In some embodiments, the polynucleotide construct comprises modified RNA nucleotides and / or modified nucleosides.
[0448] In some embodiments, the polynucleotide construct comprises 10%-100% modified RNA nucleotides and / or modified nucleosides. In some embodiments, the polynucleotide construct comprises 10%, 11%, 12%, 13%, 14%, 15%, 16%, 17%, 18%, 19%, 20%, 21%, 22%, 23%, 24%, 25%, 26%, 27%, 28%, 29%, 30%, 31%, 32%, 33%, 34%, 35%, 36%, 37%, 38%, 39%, 40%, 31%, 32%, 33%, 34%, 35%, 36%, 37%, 38%, 39%, 40%, 41%, 42%, 43%, 44%, 45%, 46%, 47%, 48%, 49%, 50%, 51%, 52%, 53%, 54%, 55%, 56%, 57%, 58%, 59%, 60%, 61%, 62%, 63%, 64%, 65%, 66%, 67%, 68%, 69%, 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 100%, 101%, 102%, %, 51%, 52%, 53%, 54%, 55%, 56%, 57%, 58%, 59%, 60%, 61%, 62%, 63%, 64%, 65%, 66%, 67%, 68%, 69%, 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 100% modified RNA nucleotides and / or modified nucleosides.
[0449] In some embodiments, the modified RNA nucleotides and / or modified nucleosides are m5C (5-methylcytidine). In some embodiments, the polynucleotide construct of any one of embodiments 47-48, wherein at least one of the modified RNA nucleotides and / or modified nucleosides is m5U (5-methyluridine).
[0450] In some embodiments, the modified RNA nucleotide and / or modified nucleoside is m6A (N6-methyladenosine).
[0451] In some embodiments, the modified RNA nucleotide and / or modified nucleoside is Y (pseudouridine).
[0452] In some embodiments, the modified RNA nucleotide and / or modified nucleoside is mlA (1-methyladenosine).
[0453] In some embodiments, the modified RNA nucleotides and / or modified nucleosides are introduced during in vitro transcription (IVT).
[0454] In some embodiments, the modified nucleoside is selected from the group consisting of m5C (5-methylcytidine), m5U (5-methyluridine), m6A (N6-methyladenosine), s2U (2-thiouridine), Y (pseudouridine), Um (2'-O-methyluridine), m1A (1-methyladenosine), m2A (2-methyladenosine), Am (2'-O-methyladenosine), ms2 m6A (2-methylthio-N6-methyladenosine), i6A (N6-isopentenyladenosine), ms2i6A (2-methylthio-N6-isopentenyladenosine), io6A (N6-(cis-hydroxyisopentenyl)adenosine), ms2io6A (2-methylthio-N6-(cis-hydroxyisopentenyl)adenosine), g6A (N6-glycylcarbamoyladenosine), t6A (N6-threonylcarbamoyladenosine), ms2t6A (2-methylthio-N6-threonylcarbamoyladenosine), m6t6A (N6-methyl-N6-threonylcarbamoyladenosine), hn6A (N6-hydroxy Norvalylcarbamoyladenosine), ms2hn6A (2-methylthio-N6-hydroxynorvalylcarbamoyladenosine), Ar(p)(2'-O-ribosyladenosine (phosphate)), I(inosine), m1I(1-methylinosine), m1hn(1,2'-O-dimethylinosine), m3C(3-methylcytidine), Cm(2'-O-methylcytidine), s2C(2-thiocytidine), ac4C(N4-acetylcytidine), (5-formylcytidine), m5Cm(5,2'-O-dimethylcytidine), ac4Cm(N4-acetyl-2'-O-methylcytidine), k2C(lysine), m! G (1-methylguanosine), m2G (N2-methylguanosine), m7G (7-methylguanosine), Gm (2'-0-methylguanosine), m2 2G (N2,N2-dimethylguanosine), m2Gm (N2,2'-O-dimethylguanosine), m2 aGm (N2, N2, 2'-O-trimethylguanosine), Gr(p) (2'-O-ribosylguanosine (phosphate)), yW (hybutosine), oayW (peroxyhybutosine), OHyW (hydroxyhybutosine), OHyW* (undermodified hydroxyhybutosine), imG (hybutosine), mimG (methylhybutosine), Q (braid), oQ (epoxybraid), galQ (galactosyl-braid), manQ (mannosyl-braid), preQo (7-cyano-7-deazaguanosine), preQi (7-aminomethyl-7-deazaguanosine), G+ (archauridine), D (dihydrouridine), m5Um (5,2'-0-dimethyluridine), s4U (4-thiouridine), m5s2U (5-methyl-2-thiouridine), s2Um (2-thio-2'-0-methyluridine), acp3U (3-(3-amino-3-carboxypropyl)uridine), ho5U (5-hydroxyuridine), mo5U (5-methoxyuridine), cmo5U (uridine 5-oxyacetic acid), mcmo5U (uridine 5-oxyacetic acid methyl ester), chm5U (5-(carboxyhydroxymethyl)uridine), mchm5U (5-(carboxyhydroxymethyl)uridine methyl ester), mcm5U (5-methoxycarbonylmethyluridine), mcm5Um (5-methoxycarbonylmethyl-2'-0-methyluridine), m cm5s2U (5-methoxycarbonylmethyl-2-thiouridine), nm5S2U (5-aminomethyl-2-thiouridine), mnm5U (5-methylaminomethyluridine), mnm5s2U (5-methylaminomethyl-2-thiouridine), mnm5se2U (5-methylaminomethyl-2-selenoyluridine), ncm5U (5-carbamoylmethyluridine), ncm5Um (5-carbamoylmethyl-2'-O-methyluridine), cmnm5U (5-carboxymethylaminomethyluridine), cmnm5Um (5-carboxymethylaminomethyl-2'-O-methyluridine), cmnm5s2U (5-carboxymethylaminomethyl-2-thiouridine), m6 2A (N6, N6-dimethyladenosine), Im (2'-0-methylinosine), m4C (N4-methylcytidine), m4Cm (N4, 2'-0-dimethylcytidine), hm5C (5-hydroxymethylcytidine), m3U (3-methyluridine), cm5U (5-carboxymethyluridine), m6Am (N6, 2'-O-dimethyladenosine), m6 2Am (N6, N6, 0-2'-trimethyladenosine), m2,7G (N2, 7-dimethylguanosine), m2,2,7G (N2, N2, 7-trimethylguanosine), m3Um (3,2'-0-dimethyluridine), m5D (5-methyldihydrouridine), f5Cm (5-formyl-2'-0-methylcytidine), m'Gm (l, 2'-0-dimethylguanosine), m'Am (l,2'-0-dimethyladenosine), rm5U (5-tauromethyluridine), τm5s2U (5-tauromethyl-2-thiouridine), imG-14 (4-demethylwyoside), imG2 (isowyoside), ac6A (N6-acetyladenosine), pyridin-4-one ribonucleoside, 5-aza-uridine, 2-thio-5-aza-uridine, 2-thiouridine, 4-thio-pseudouridine, 2-thiouridine, 5-hydroxyuridine, 3-methyluridine, 5-carboxymethyl-uridine, 1-carboxymethyl-pseudouridine, 5-propynyl-uridine, 1-propynyl-pseudouridine, 5-tauromethyluridine, 1-tauromethyl-pseudouridine, 5-tauromethyl- 2-thio-uridine, l-taurine methyl-4-thio-uridine, 5-methyl-uridine, 1-methyl-pseudouridine, 4-thio-1-methyl-pseudouridine, 2-thio-1-methyl-pseudouridine, 1-methyl-1-deaza-pseudouridine, 2-thio-1-methyl-1-deaza-pseudouridine, dihydrouridine, dihydropseudouridine, 2-thio-dihydrouridine, 2-thio-dihydropseudouridine, 2-methoxyuridine, 2-methoxy-4-thio-uridine, 4-methoxy-pseudouridine, 4-methoxy-2-thio-pseudouridine, 5-azacytidine, pseudoisocytidine, 3-methyl-cytidine, N4-acetylcytidine, 5-formylcytidine, N4-methyl cytidine, 5-hydroxymethylcytidine, 1-methyl-pseudoisocytidine, pyrrolo-cytidine, pyrrolo-pseudoisocytidine, 2-thio-cytidine, 2-thio-5-methyl-cytidine, 4-thio-pseudoisocytidine, 4-thio-1-methyl-pseudoisocytidine, 4-thio-1-methyl-1-deaza-pseudoisocytidine, 1-methyl-1-deaza-pseudoisocytidine, zebulline, 5-aza-zebulline, 5-methyl-zebulline, 5-aza-2-thio-zebulline, 2-thio-zebulline, 2-methoxy-cytidine, 2-methoxy-5-methyl-cytidine, 4-methoxy-pseudoisocytidine, 4-methoxy-1-methyl-pseudoisocytidine, 2- Aminopurine, 2,6-diaminopurine, 7-deaza-adenine, 7-deaza-8-aza-adenine, 7-deaza-2-aminopurine, 7-deaza-8-aza-2-aminopurine, 7-deaza-2,6-diaminopurine, 7-deaza-8-aza-2,6-diaminopurine, 1-methyladenosine, N6-methyladenosine, N6-isopentenyladenosine, N6-(cis-hydroxyisopentenyl)adenosine, 2-methylthio-N6-(cis-hydroxyisopentenyl)adenosine, N6-glycylcarbamoyladenosine, N6-threonylcarbamoyladenosine, 2-methylthio-N6-threonylcarbamoyladenosine, N6,N6-dimethyladenosine, 7-methyladenine, 2-methylthio-adenine, 2-methoxy-adenine, inosine, 1-methyl-inosine, wyosine, wyobutine, 7-deaza-guanosine, 7-deaza-8-aza-guanosine, 6-thio-guanosine, 6-thio-7-deaza-guanosine, 6-thio-7-deaza-8-aza-guanosine, 7-methyl-guanosine, 6-thio-7-methyl guanosine, 7-methylinosine, 6-methoxyguanosine, 1-methylguanosine, N2-methylguanosine, N2,N2-dimethylguanosine, 8-oxoguanosine, 7-methyl-8-oxoguanosine, 1-methyl-6-thioguanosine, N2-methyl-6-thioguanosine, N2,N2-dimethyl-6-thioguanosine, 5-methylcytosine, pseudouridine, 1-methylpseudouridine.
[0455] In some embodiments, the circular RNA is at least 500 nucleotides in length, at least 1,000 nucleotides in length, or at least 1,500 nucleotides in length.
[0456] In some embodiments, the circular RNA does not contain any other sequence that does not belong to the target sequence, such as all or part of the E2 or E1 sequence.
[0457] In some embodiments, provided herein is a method for making circular RNA, the method comprising: preparing a vector comprising the following operably linked elements from 5' to 3': (a) 3' intron fragment; (b) exon fragment 2 (E2); (c) target sequence; (d) exon fragment 1 (E1); (d) 5' intron fragment. In one embodiment, the 5' intron fragment and the 3' intron fragment are each fragments of a type II intron. In one embodiment, the 5' intron fragment is located on the 5' side of the 3' intron fragment in the type II intron. In one embodiment, the E1 is a 5' adjacent exon fragment of the type II intron, and its length is ≥0 nucleotides. In one embodiment, the E2 is a 3' adjacent exon fragment of the type II intron, and its length is ≥0 nucleotides, and the target sequence is empty, or is a protein coding sequence, a non-coding sequence, or a combination thereof.
[0458] In some embodiments, provided herein is a method for making circular RNA, the method comprising: preparing a vector comprising the following operably linked elements from 5' to 3': (a) a 3' intron fragment; (b) exon fragment 2 (E2); (c) a linker sequence; (d) a target sequence; (e) a linker sequence; (f) exon fragment 1 (E1); (g) a 5' intron fragment. In one embodiment, the 5' intron fragment is located on the 5' side of the 3' intron fragment in the type II intron. In one embodiment, the E1 is the 5' adjacent exon fragment of the type II intron, and its length is ≥0 nucleotides. In one embodiment, the E2 is the 3' adjacent exon fragment of the type II intron, and its length is ≥0 nucleotides, and the target sequence is empty, or is a protein coding sequence, a non-coding sequence, or a combination thereof.
[0459] In some embodiments, provided herein is a method for making circular RNA, the method comprising: preparing a vector comprising the following operably linked elements from 5' to 3': (a) a 5' homology arm; (b) a 3' intron fragment; (c) exon fragment 2 (E2); (d) a target sequence; (e) exon fragment 1 (E1); (f) a 5' intron fragment; (g) a 3' homology arm. In one embodiment, the 5' intron fragment is located on the 5' side of the 3' intron fragment in the type II intron. In one embodiment, the E1 is a 5' adjacent exon fragment of the type II intron, and its length is ≥0 nucleotides. In one embodiment, the E2 is a 3' adjacent exon fragment of the type II intron, and its length is ≥0 nucleotides, and the target sequence is empty, or is a protein coding sequence, a non-coding sequence, or a combination thereof.
[0460] In some embodiments, provided herein is a method for making circular RNA, the method comprising: preparing a vector comprising the following operably linked elements from 5' to 3': (a) a 5' homology arm; (b) a 3' intron fragment; (c) exon fragment 2 (E2); (d) a linker sequence; (e) a target sequence; (f) a linker sequence; (g) exon fragment 1 (E1); (h) a 5' intron fragment; (i) a 3' homology arm. In one embodiment, the 5' intron fragment is located on the 5' side of the 3' intron fragment in the type II intron. In one embodiment, the E1 is a 5' adjacent exon fragment of the type II intron, and its length is ≥0 nucleotides. In one embodiment, the E2 is a 3' adjacent exon fragment of the type II intron, and its length is ≥0 nucleotides, and the target sequence is empty, or is a protein coding sequence, a non-coding sequence, or a combination thereof.
[0461] In some embodiments, provided herein is a method for expressing a protein in a cell, comprising (a) transfecting the circular RNA of any one of embodiments 58-61 into the cell, or (b) subjecting the polynucleotide construct of any one of embodiments 1-57 to a self-splicing cyclization reaction to form a circular RNA, and transfecting the circular RNA into the cell; wherein, preferably, the cell is a eukaryotic cell.
[0462] In some embodiments, provided herein is a method for expressing a protein in a cell, comprising (a) transfecting the circular RNA of any one of embodiments 58-61 into the cell, or (b) causing the construct of any one of embodiments 1-57 to undergo self-splicing circularization reaction to form a circular RNA, and transfecting the circular RNA into the cell; wherein, preferably, the cell is a hepatocyte, an epithelial cell, a hematopoietic cell, an epithelial cell, an endothelial cell, a lung cell, a bone cell, a stem cell, a mesenchymal cell, a neural cell (e.g., a meningeal cell, a astrocyte, a motor neuron, a dorsal root ganglion cell, anterior horn motor neuron), a photoreceptor cell (e.g., a rod cell, a cone cell), a retinal pigment epithelial cell, or a glial cell. cells, secretory cells, cardiac cells, adipocytes, vascular smooth muscle cells, cardiomyocytes, skeletal muscle cells, β cells, pituitary cells, synovial lining cells, ovarian cells, testicular cells, fibroblasts, B cells, T cells, dendritic cells, macrophages, reticulocytes, leukocytes, granulocytes, tumor cells, NK cells, hepatic stellate cells, HEK293, HEK293T, HeLa, MCF7, PC3, A549, NCI-H727, HCT-116, MCF10A, HPReC, FHC, immortalized cell lines, primary cells, yeast cells, Saccharomyces cerevisiae, Pichia pastoris, bacterial cells, Escherichia coli, insect cells, Spodoptera frugiperda sf9, Mimic Sf9, sf21, Drosophila S2.
[0463] In some embodiments, the polynucleotide construct, circular RNA, or method of any of the preceding embodiments, wherein the 5' intron fragment and the 3' intron fragment each comprise one or more pairs of complementary sequences. In a preferred embodiment, the length of the complementary sequences is greater than 20 nucleotides.
[0464] In some embodiments, the polynucleotide construct, circular RNA or method of any of the preceding embodiments, wherein the 5' intron fragment and / or the 3' intron fragment comprises one or more affinity tag sequences, wherein the affinity tag sequence is selected from the group consisting of a probe binding sequence, an MS2 binding site, a PP7 binding site, and a streptavidin binding site.
[0465] In some embodiments, the polynucleotide construct, circular RNA or method of any one of the preceding embodiments, wherein the EBS sequence is selected from one or more of EBS1, EBS2, EBS3, preferably two, more preferably EBS1 and EBS3.
[0466] In some embodiments, the polynucleotide construct, circular RNA or method of any of the preceding embodiments, wherein one or more EBS sequences of the group II intron, preferably EBS1 and EBS3, are modified, wherein the EBS sequences are complementary to two regions of corresponding length in the target sequence at at least 60% of the nucleotide positions, respectively.
[0467] In a preferred embodiment, the two regions of corresponding lengths in the target sequence are located at both ends of the target sequence.
[0468] In a preferred embodiment, the polynucleotide construct, circular RNA or method according to any one of the preceding embodiments, wherein the polynucleotide construct is capable of forming a circular RNA of a target sequence in vitro.
[0469] In a preferred embodiment, the polynucleotide construct, circular RNA or method of any one of the preceding embodiments, wherein the polynucleotide construct is capable of forming a circular RNA of a target sequence in vivo. Purification
[0470] In some embodiments, the polynucleotide comprises a purification tag. In preferred embodiments, the purification tag is a 15-40 nt polynucleotide annealed to an oligonucleotide conjugated to a purification matrix. Purification matrices include, but are not limited to, magnetic resins or beads, silicone resins, Sephadex resins, affinity resins, nanoparticles, and nanomaterial-coated surfaces.
[0471] In some embodiments, the purification tag is an intron tag. In some embodiments, the purification tag is a 5' intron tag. In some embodiments, the purification tag is a 3' intron tag.
[0472] The circular RNA produced by the construct or method of the present invention can be purified. For example, the purification method is selected from one or more of the following: enzyme treatment; chromatography, including but not limited to affinity column chromatography, reversed-phase silica gel column liquid chromatography, and gel exclusion liquid chromatography; electrophoresis, including but not limited to gel electrophoresis such as agarose gel electrophoresis, and capillary electrophoresis.
[0473] Before transfecting the circular RNA product into the cell, it is preferred to remove uncircularized linear RNA, dsRNA and other unwanted components as much as possible through purification. The phosphate groups at both ends of linear RNA and some dsRNA will activate the RIG-1 signaling pathway, causing a strong immune response in the cell, leading to the degradation of exogenous RNA, and affecting the function of circular RNA in the cell. Methods for removing linear RNA include enzymatic treatment, such as treatment with RNase R; chromatography, such as high performance liquid chromatography (HPLC). Methods for removing terminal phosphate groups include treatment with alkaline phosphatase, such as alkaline phosphatase (calf intestinal alkaline phosphatase; CIP) derived from bovine small intestine. Administration and delivery The circular RNA generated by the construct of the present invention or the method can be delivered into a cell or an animal using any of a variety of delivery systems. For example, the delivery system is selected from one or more of the following groups: liposomes, polyethyleneimine (PEI), metal organic framework materials (MOF), liposome nanomaterials (LNP), polycations, blood glycoproteins, erythrocyte transport vehicles, gold nanoparticles (AuNP) vehicles, magnetic nanomaterial vehicles, carbon nanotubes, graphene molecule vehicles, quantum dot materials vehicles, up-conversion nanocrystals, layered hydroxide materials vehicles, silicon crystal nanomaterials, calcium phosphate. In some embodiments, circular RNA can be transfected into cells using, for example, liposome transfection or electroporation. 6.5.2. Target cells
[0474] In some embodiments, the target cell lacks a protein or enzyme of interest. For example, in the case where it is desired to deliver nucleic acid to hepatocytes, hepatocytes represent the target cells. In some embodiments, the compositions of the present disclosure transfect target cells (i.e., non-target cells are not transfected) on the basis of identification. The compositions of the present disclosure can also be prepared to preferentially target a variety of target cells and / or express in a variety of target cells, including but not limited to hepatocytes, epithelial cells, hematopoietic cells, epithelial cells, endothelial cells, pneumocytes, bone cells, stem cells, mesenchymal cells, neural cells (e.g., meningeal cells, astrocytes, motor neurons, dorsal root ganglion cells, anterior horn motor neurons), photoreceptor cells (e.g., rods, cones), retinal pigment epithelial cells, secretory cells, cardiac cells, adipocytes, blood vessels, umbilical cord cells, pelvic floor ... Tumor cells, smooth muscle cells, cardiomyocytes, skeletal muscle cells, beta cells, pituitary cells, synovial lining cells, ovarian cells, testicular cells, fibroblasts, B cells, T cells, dendritic cells, macrophages, reticulocytes, leukocytes, granulocytes, tumor cells, NK cells, hepatic stellate cells, HEK293, HEK293T, HeLa, MCF7, PC3, A549, NCI-H727, HCT-116, MCF10A, HPReC, FHC, other immortalized cell lines, primary cell lines.
[0475] In some embodiments, the compositions of the present disclosure can also be optimized for various yeast cells, including but not limited to Saccharomyces cerevisiae and Pichia pastoris.
[0476] In some embodiments, the compositions of the present disclosure may also be optimized for a variety of bacterial cells, including but not limited to E. coli.
[0477] In some embodiments, the compositions of the present disclosure can also be optimized for a variety of insect cells, including but not limited to Spodoptera frugiperda sf9, Mimic Sf9, sf21, and Drosophila S2.
[0478] The compositions of the present disclosure can be prepared to be preferentially distributed to target cells and / or optimized for target cells, such as in the heart, lungs, kidneys, liver and spleen. In some embodiments, the compositions of the present disclosure are distributed to cells in the liver to promote the delivery and subsequent expression of the circRNA contained therein by cells in the liver (e.g., hepatocytes). Targeted cells can act as biological "reservoirs" or "depots" that can produce and systemically excrete functional proteins or enzymes. Therefore, in one embodiment of the present disclosure, the transfer vector can target hepatocytes and / or be preferentially distributed to cells in the liver when delivered. In embodiments, after transfection of the target hepatocytes, the circRNA loaded in the vector is translated, and a functional protein product is produced, excreted and distributed systemically. In other embodiments, cells other than hepatocytes (e.g., cells of the lungs, spleen, heart, eyes or central nervous system) can be used as storage locations for protein production.
[0479] In one embodiment, the composition of the present disclosure promotes the endogenous production of one or more functional proteins and / or enzymes in the subject. In an embodiment of the present disclosure, the transfer vehicle comprises a circRNA encoding a defective protein or enzyme. After such a composition is distributed in the target tissue and subsequently transfected into such a target cell, the exogenous circRNA loaded into the transfer vehicle (e.g., lipid nanoparticles) can be translated in vivo to produce a functional protein or enzyme (e.g., a protein or enzyme that the subject lacks) encoded by the circRNA administered exogenously. Therefore, the composition of the present disclosure utilizes the ability of the subject to translate exogenous or recombinantly prepared circRNA to produce endogenously translated proteins or enzymes, thereby producing (and excreting, where applicable) functional proteins or enzymes. The protein or enzyme expressed or translated can also be characterized by comprising a natural post-translational modification in vivo, which may not generally be present in the recombinantly prepared protein or enzyme, thereby further reducing the immunogenicity of the translated protein or enzyme.
[0480] Administration of circRNA encoding a defective protein or enzyme avoids the need to deliver the nucleic acid to a specific organelle within the target cell. Instead, after transfection of the target cell and delivery of the nucleic acid to the cytoplasm of the target cell, the circRNA content of the transfer vector can be translated and the functional protein or enzyme can be expressed.
[0481] In some embodiments, circular RNA includes one or more miRNA binding sites. In some embodiments, circular RNA includes one or more miRNA binding sites identified by a miRNA present in one or more non-target cells or non-target cell types (e.g., Kupffer cells) and not present in one or more target cells or target cell types (e.g., hepatocytes). In some embodiments, circular RNA includes one or more miRNA binding sites identified by a miRNA present in one or more non-target cells or non-target cell types (e.g., Kupffer cells) with an increased concentration compared to one or more target cells or target cell types (e.g., hepatocytes). MiRNA is believed to work by pairing with a complementary sequence within an RNA molecule, thereby causing gene silencing. 6.5.3. Pharmaceutical Composition / Administration
[0482] In some embodiments, provided herein are compositions (e.g., pharmaceutical compositions) comprising therapeutic agents provided herein. In some embodiments, the therapeutic agent is a circular RNA polynucleotide provided herein. In some embodiments, the therapeutic agent is a vector provided herein. In some embodiments, the therapeutic agent is a cell comprising a circular RNA or vector provided herein. In some embodiments, the composition further comprises a pharmaceutically acceptable carrier. In some embodiments, provided herein are compositions comprising a therapeutic agent provided herein with a combination of other pharmaceutically active agents or drugs. In preferred embodiments, the pharmaceutical composition comprises a cell or a colony thereof provided herein.
[0483] With respect to pharmaceutical compositions, a pharmaceutically acceptable carrier can be any of the conventionally used carriers and is limited only by chemicophysical considerations (such as solubility and lack of reactivity with one or more active agents) and the route of administration. The pharmaceutically acceptable carriers described herein (e.g., vehicles, adjuvants, excipients, and diluents) are well known to those skilled in the art and are readily available to the public. Preferably, a pharmaceutically acceptable carrier is a carrier that is chemically inert to the one or more therapeutic agents and a carrier that has no harmful side effects or toxicity under the conditions of use.
[0484] The choice of carrier will depend in part on the particular therapeutic agent, as well as the particular method used to administer the therapeutic agent.Thus, there is a wide variety of suitable formulations for the pharmaceutical compositions provided herein.
[0485] In some embodiments, the pharmaceutical composition comprises a preservative. In some embodiments, suitable preservatives may include, for example, methylparaben, propylparaben, sodium benzoate, and benzalkonium chloride. Optionally, a mixture of two or more preservatives may be used. The preservative or mixture thereof is typically present in an amount of about 0.0001% to about 2% by weight of the total composition.
[0486] In some embodiments, the pharmaceutical composition comprises a buffer. In some embodiments, suitable buffers can include, for example, citric acid, sodium citrate, phosphoric acid, potassium phosphate, and a variety of other acids and salts. Optionally, a mixture of two or more buffers can be used. The buffer or mixture thereof is typically present in an amount of about 0.001% to about 4% by weight of the total composition.
[0487] In some embodiments, the concentration of the therapeutic agent in the pharmaceutical composition can vary, for example, less than about 1%, or at least about 1%, 2%, 3%, 4%, 5%, 6%, 7%, 8%, 9%, 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, or about 50% or more by weight, and can be selected primarily based on the fluid volume and viscosity according to the particular mode of administration selected.
[0488] The following formulations for oral, aerosol, parenteral (e.g., subcutaneous, intravenous, intraarterial, intramuscular, intradermal, intraperitoneal, and intrathecal) and topical administration are exemplary only and are in no way limiting. More than one route may be used to administer the therapeutic agents provided herein, and in some cases, a particular route may provide a more immediate and more effective response than another route.
[0489] Formulations suitable for oral administration may include or consist of: (a) liquid solutions, such as an effective amount of a therapeutic agent dissolved in a diluent such as water, saline, or orange juice; (b) capsules, sachets, tablets, lozenges, and troches, each containing a predetermined amount of the active ingredient in the form of a solid or particulate form; (c) powders; (d) suspensions in appropriate liquids; and (e) suitable emulsions. Liquid formulations may include diluents such as water and alcohols (e.g., ethanol, benzyl alcohol, and polyvinyl alcohol), with or without the addition of a pharmaceutically acceptable surfactant. Capsule forms may be of the conventional hard or soft shell gelatin type containing, for example, a surfactant, a lubricant, and an inert filler such as lactose, sucrose, calcium phosphate, and corn starch. Tablet forms can include one or more of the following: lactose, sucrose, mannitol, corn starch, potato starch, alginic acid, microcrystalline cellulose, gum arabic, gelatin, guar gum, colloidal silicon dioxide, croscarmellose sodium, talc, magnesium stearate, calcium stearate, zinc stearate, stearic acid and other excipients, colorants, diluents, buffers, disintegrants, wetting agents, preservatives, flavorings and other pharmacologically compatible excipients. Lozenge forms can contain the therapeutic agent with a flavoring (typically sucrose, gum arabic or tragacanth). Candy lozenges can contain the therapeutic agent with an inert matrix (such as gelatin and glycerin, or sucrose and gum arabic, emulsions, gels, etc.), and also contain such excipients as known in the art.
[0490] Formulations suitable for parenteral administration include aqueous and non-aqueous isotonic sterile injection solutions which may contain antioxidants, buffers, bacteriostats and solutes which render the formulation isotonic with the blood of the intended recipient; and aqueous and non-aqueous sterile suspensions which may include suspending agents, solubilizers, thickening agents, stabilizers and preservatives. In some embodiments, the therapeutic agents provided herein can be administered in a physiologically acceptable diluent in a pharmaceutical carrier, such as a sterile liquid or liquid mixture including water, saline, aqueous dextrose and related sugar solutions, alcohols such as ethanol or hexadecyl alcohol, glycols such as propylene glycol or polyethylene glycol, dimethyl sulfoxide, glycerol, ketals such as 2,2-dimethyl-1,3-dioxolane-4-methanol, ethers, poly(ethylene glycol) 400, oils, fatty acids, fatty acid esters or glycerides or acetylated fatty acid glycerides, with or without the addition of pharmaceutically acceptable surfactants (such as soaps or detergents), suspending agents (such as pectin, carbomer, methylcellulose, hydroxypropyl methylcellulose or carboxymethylcellulose) or emulsifiers and other pharmaceutical adjuvants.
[0491] In some embodiments, the oil that can be used in parenteral preparations comprises petroleum, animal oil, vegetable oil or synthetic oil.The specific example of oil comprises peanut oil, soybean oil, sesame oil, cottonseed oil, corn oil, olive oil, petrolatum and mineral oil.The suitable fatty acid that is used for using in parenteral preparations comprises oleic acid, stearic acid and isostearic acid.Ethyl oleate and isopropyl myristate are examples of suitable fatty acid esters.
[0492] Suitable soaps for use in some embodiments of parenteral formulations include fatty alkali metal, ammonium and triethanolamine salts, and suitable detergents include (a) cationic detergents such as, for example, dimethyldialkylammonium halides and alkylpyridinium halides; (b) anionic detergents such as, for example, alkyl, aryl and olefin sulfonates, alkyl, olefin, ether and monoglyceride sulfates and sulfosuccinates; (c) nonionic detergents such as, for example, fatty amine oxides, fatty acid alkanolamides and polyoxyethylene polypropylene copolymers; (d) amphoteric detergents such as, for example, alkyl-b-aminopropionates and 2-alkyl-imidazoline quaternary ammonium salts; and (e) mixtures thereof.
[0493] In some embodiments, the parenteral formulation will contain, for example, from about 0.5% to about 25% therapeutic agent by weight in the solution. Preservatives and buffers can be used. In order to minimize or eliminate the irritation at the injection site, such compositions can contain one or more nonionic surfactants with, for example, a hydrophilic-lipophilic balance (HLB) value from about 12 to about 17. The amount of surfactant in such a formulation will generally range from about 5% to about 15% by weight. Suitable surfactants include polyethylene glycol, sorbitan fatty acid esters (such as sorbitan monooleate) and high molecular weight adducts of ethylene oxide and a hydrophobic matrix, which is formed by the condensation of propylene oxide and propylene glycol. The parenteral formulation can be provided in a unit dose or multi-dose sealed container (such as an ampoule or a vial) and can be stored under freeze-dried (lyophilized) conditions, requiring only the addition of a sterile liquid excipient (for example, water) for injection just before use. Extemporaneous injection solutions and suspensions may be prepared from sterile powders, granules and tablets of the kind previously described.
[0494] In some embodiments, provided herein are injectable formulations. The requirements for effective pharmaceutical carriers for injectable compositions are well known to those of ordinary skill in the art (see, e.g., Pharmaceutics and Pharmacy Practice, JB Lippincott Company, Philadelphia, Pennsylvania, Banker and Chalmers, eds., pp. 238-250 (1982), and ASHP Handbook on Injectable Drugs, Toissel, 4th edition, pp. 622-630 (1986)).
[0495] In some embodiments, provided herein are topical formulations. Topical formulations (including topical formulations that can be used for transdermal drug delivery) are suitable for application to the skin in the context of certain embodiments provided herein. In some embodiments, the therapeutic agent, alone or in combination with other suitable components, can be made into aerosol formulations to be applied via inhalation. These aerosol formulations can be placed in pressurized acceptable propellants, such as dichlorodifluoromethane, propane, nitrogen, etc. They can also be formulated as medicines for non-pressurized preparations, such as in a sprayer or nebulizer. Such spray formulations can also be used for spraying mucous membranes.
[0496] In some embodiments, the therapeutic agents provided herein can be formulated as inclusion complexes, such as cyclodextrin inclusion complexes, or as liposomes. Liposomes can be used to target therapeutic agents to specific tissues. Liposomes can also be used to increase the half-life of therapeutic agents. Many methods can be used to prepare liposomes, such as those described in, for example, Szoka et al., Ann. Rev. Biophys. Bioeng., 9,467 (1980) and U.S. Patents 4,235,871, 4,501,728, 4,837,028, and 5,019,369.
[0497] In some embodiments, the therapeutic agent provided herein is formulated in a timed release, delayed release or sustained release delivery system so that the delivery of the composition occurs before the sensitization of the part to be treated and there is enough time to cause sensitization. Such a system can avoid repeated administration of the therapeutic agent, thereby increasing the convenience to the subject and the physician, and can be particularly suitable for certain composition embodiments provided herein. In one embodiment, the compositions of the present disclosure are formulated so that they are suitable for the extended release of the circRNA contained therein. Such extended release compositions can be conveniently administered to the subject with an extended dosing interval. For example, in one embodiment, the compositions of the present disclosure are administered to the subject twice a day, once a day, or once every other day. In an embodiment, the compositions of the present disclosure are administered to the subject twice a week, once a week, once every ten days, once every two weeks, once every three weeks, once every four weeks, once a month, once every six weeks, once every eight weeks, once every three months, once every four months, once every six months, once every eight months, once every nine months, or once a year.
[0498] In some embodiments, the target cell produces a protein encoded by the polynucleotides described herein for a sustained period of time. For example, after administration, the protein can produce more than one hour, more than four hours, more than six hours, more than 12 hours, more than 24 hours, more than 48 hours, or more than 72 hours. In some embodiments, the therapeutic product is expressed at a peak level about six hours after administration. In some embodiments, the expression of the therapeutic product is at least sustained at a therapeutic level. In some embodiments, after administration, the therapeutic product is expressed at least at a therapeutic level for more than one hour, more than four hours, more than six hours, more than 12 hours, more than 24 hours, more than 48 hours, or more than 72 hours. In some embodiments, the therapeutic product can be detected at a therapeutic level in patient serum or tissue (e.g., liver or lung). In some embodiments, a certain level of detectable therapeutic product comes from the continuous expression of the circRNA composition over a period of more than one hour, more than four hours, more than six hours, more than 12 hours, more than 24 hours, more than 48 hours, or more than 72 hours after administration.
[0499] In some embodiments, the protein encoded by the polynucleotide as herein described is produced at a level higher than normal physiological levels. Compared with the control, protein levels can increase. In some embodiments, the control is the baseline physiological level of the therapeutic product in normal individuals or in normal individual colonies. In other embodiments, the control is the baseline physiological level of the therapeutic product in the individuality with related protein or polypeptide defect or in the individual colony with related protein or polypeptide defect. In some embodiments, the control can be the normal level of related protein or polypeptide in the individuality of the compositions. In other embodiments, the control is the expression level of the therapeutic product described in one or more comparable time points after other treatment interventions (for example, after directly injecting the corresponding therapeutic product).
[0500] In some embodiments, a certain level of protein encoded by the polynucleotides described herein can be detected 3 days, 4 days, 5 days or 1 week or longer after administration. Increased levels of secreted protein can be observed in serum and / or tissues (e.g., liver or lung).
[0501] In some embodiments, the methods result in a prolonged circulating half-life of a protein encoded by a polynucleotide described herein. For example, the protein can be detected hours or days longer than the half-life observed via subcutaneous injection of the protein or mRNA encoding the protein. In some embodiments, the half-life of the protein is 1 day, 2 days, 3 days, 4 days, 5 days, or 1 week or longer.
[0502] Many types of release delivery systems are available and known to those of ordinary skill in the art. They include polymer-based systems such as poly (lactide-glycolide), copolyoxalates, polycaprolactones, polyesteramides, polyorthoesters, polyhydroxybutyric acid, and polyanhydrides. The aforementioned drug-containing polymer microcapsules are described in, for example, U.S. Patent No. 5,075,109. Delivery systems also include non-polymer systems, which are lipids, including sterols such as cholesterol, cholesterol esters, and fatty acids, or neutral fats such as monoglycerides, diglycerides, and triglycerides; hydrogel release systems; silicone rubber systems; peptide-based systems; wax coatings; compressed tablets using conventional binders and excipients; partially fused implants; etc. Specific examples include, but are not limited to: (a) erosion systems in which the active composition is contained within a matrix, such as those described in U.S. Patents 4,452,775, 4,667,014, 4,748,034, and 5,239,660; and (b) diffusion systems in which the active ingredient permeates from a polymer at a controlled rate, such as those described in U.S. Patents 3,832,253 and 3,854,480. In addition, pump-based hardware delivery systems can be used, some of which are suitable for implantation.
[0503] In some embodiments, the therapeutic agent can be conjugated directly or indirectly to the targeting moiety via a linking moiety. Methods for conjugating therapeutic agents to targeting moieties are known in the art. See, for example, Wadwa et al., J. Drug Targeting 3:111 (1995) and U.S. Patent No. 5,087,616.
[0504] In some embodiments, the therapeutic agent provided herein is formulated into a reservoir form so that the mode in which the therapeutic agent is released into the body to which it is applied is controlled with respect to time and position in the body (see, for example, U.S. Patent No. 4,450,150). The reservoir form of the therapeutic agent can be, for example, an implantable composition comprising a therapeutic agent and a porous or non-porous material, such as a polymer, wherein the therapeutic agent is encapsulated by the material or diffused into the entire material and / or diffused by degradation of the non-porous material. The reservoir is then implanted in the desired position in the body and the therapeutic agent is released from the implant at a predetermined rate. 6.5.4. Purpose
[0505] Depending on the type of target sequence, the circular RNA generated by the constructs or methods of the present invention can be used to achieve a variety of applications. For example, when the target sequence comprises or consists of a protein coding sequence, the resulting circular RNA can be used to express the protein. The circular RNA of the present invention can also be used to modulate miRNA activity, neutralize RNA binding proteins, express aptamers, and other functions. 7. Determination
[0506] Various IRES-like sequence variants, endogenous IRES sequence variants, or combinations thereof can be tested for their ability to attract eukaryotic ribosomal translation initiation complexes and / or promote translation initiation. The following assays are described for IRES-like sequences, but can be similarly performed for endogenous IRES sequences, combinations of IRES-like sequences and endogenous IRES sequences, or sequences comprising one or more IRES-like sequences or endogenous IRES sequences. 7.1. Determining the secondary structure and cleavage sites of group II introns
[0507] The stem-loop structure is a type of RNA secondary structure that can be determined by any suitable polynucleotide folding algorithm. Some programs are based on the calculation of the minimum Gibbs free energy. An example of such an algorithm is mFold, and is described by Zuker and Stiegler (Nucleic Acids Res.9 (1981), 133-148). Another exemplary folding algorithm is the online web server RNAfold developed by the Institute of Theoretical Chemistry of the University of Vienna using a centroid structure prediction algorithm (e.g., AR Gruber et al., 2008, Cell 106). (1): 23-24; and PA Carr and GM Church, 2009, Nature Biotechnology 27 (12): 1151-62). Additional algorithms can be found in U.S. Provisional Patent Application No. 61 / 836,080 (Agent Docket No. 44790.11.2022; Pan Reference No. BI-2013 / 004A), which is incorporated herein by reference. Type II introns mainly consist of six stem-loop structures, called domains 1-6 (D1-D6). The six domains are arranged in sequence and contain multiple exon binding sequences (EBS), such as EBS1, EBS2, and EBS3. These EBS sequences interact with the intron binding sequences (IBS) in the exon region, such as complementary pairing, and trigger splicing by relying on the hydroxyl groups within the EBS nucleic acid sequence. The exemplary structure of the secondary structure of type II introns is shown in Figure 9 In one embodiment, an online prediction tool or prediction software is used to identify group II introns. An example of such an online prediction tool is the online web server "http: / / webapps2.ucalgary.ca / " created by the Zimmerly Laboratory at the University of Calgary.
[0508] In one embodiment, the autocatalytic self-splicing type II intron is split into two fragments from the D1 domain, and the target sequence is inserted between the split intron fragments. In one embodiment, the autocatalytic self-splicing type II intron is split into two fragments from the D2 domain, and the target sequence is inserted between the split intron fragments. In one embodiment, the autocatalytic self-splicing type II intron is split into two fragments from the D3 domain, and the target sequence is inserted between the split intron fragments.
[0509] In a preferred embodiment, the autocatalytic self-splicing group II intron is split into two fragments from the D4 domain, and the target sequence is inserted between the split intron fragments, as shown in FIG12 .
[0510] In one embodiment, the autocatalytic self-splicing group II intron is split into two fragments from the D5 domain, and the target sequence is inserted between the split intron fragments. In one embodiment, the autocatalytic self-splicing group II intron is split into two fragments from the D6 domain, and the target sequence is inserted between the split intron fragments. a. In vitro circRNA production
[0511] Precursor RNA is generated by in vitro transcription and then circularized by the cRNAzyme system.
[0512] The vectors provided herein can be made using standard techniques of molecular biology. For example, the various elements of the vectors provided herein can be obtained using recombinant methods, such as by screening cDNA and genomic libraries from cells or by obtaining the polynucleotides from known vectors comprising the polynucleotides.
[0513] Various elements of the vectors provided herein can also be produced synthetically (rather than cloned) based on known sequences. The complete sequence can be assembled from overlapping oligonucleotides prepared by standard methods and assembled into the complete sequence. See, for example, Edge, Nature (1981) 292:756; Nambai et al., Science (1984) 223:1299; and Jay et al., J. Biol. Chem. (1984) 259:631 1.
[0514] Thus, a specific nucleotide sequence can be obtained from a vector containing the desired sequence, or synthesized in whole or in part using various oligonucleotide synthesis techniques known in the art (such as site-directed mutagenesis and polymerase chain reaction (PCR) techniques, where appropriate). One method for obtaining the nucleotide sequence encoding the desired vector element is by annealing a complementary set of overlapping synthetic oligonucleotides generated in a conventional automated polynucleotide synthesizer, followed by ligation with an appropriate DNA ligase and amplification of the ligated nucleotide sequence via PCR. See, for example, Jayaraman et al., Proc. Natl. Acad. Sci. USA (1991) 88: 4084-4088. In addition, oligonucleotide-directed synthesis (Jones et al., Nature (1986) 54:75-82), oligonucleotide-directed mutagenesis of pre-existing nucleotide regions (Riechmann et al., Nature (1988) 332:323-327 and Verhoeyen et al., Science (1988) 239:1534-1536), and enzymatic filling of gapped oligonucleotides using T4 DNA polymerase (Queen et al., Proc. Natl. Acad. Sci. USA (1989) 86:10029-10033) can be used.
[0515] Can generate precursor RNA provided herein by hatching carrier provided herein under the condition that allows the precursor RNA transcribed by carrier encoding.For example, in some embodiments, by the carrier that comprises its 5 ' duplex forming region and / or the RNA polymerase promoter upstream of expression sequence provided herein and compatible RNA polymerase are hatched together under the condition that allows in vitro transcription to synthesize precursor RNA.In some embodiments, described carrier is hatched in cell by phage RNA polymerase, or is hatched in the nucleus of cell by host RNA polymerase P.
[0516] In some embodiments, provided herein are methods for generating precursor RNA by in vitro transcription using a vector provided herein as a template (eg, a vector provided herein having an RNA polymerase promoter located upstream of a 5' homology region).
[0517] In some embodiments, the resulting precursor RNA can be used to generate circular RNA (e.g., a circular RNA polynucleotide provided herein).
[0518] Therefore, in some embodiments, provided herein is a method for making circular RNA. In some embodiments, the method includes synthesizing a precursor RNA by transcribing (e.g., uncontrolled transcription) using a carrier provided herein as a template, and incubating the resulting precursor RNA under conditions suitable for cyclization to form a circular RNA.
[0519] In some embodiments, the composition comprising the circular RNA has been purified. The circular RNA can be purified by any known method commonly used in the art (such as column chromatography, gel filtration chromatography, and size exclusion chromatography). In some embodiments, purification comprises one or more of the following steps: phosphatase treatment, HPLC size exclusion purification, and RNase R digestion. In some embodiments, purification comprises the following steps in order: RNase R digestion, phosphatase treatment, and HPLC size exclusion purification. In some embodiments, purification comprises reverse phase HPLC. In some embodiments, the purified composition contains less double-stranded RNA, DNA splints, triphosphorylated RNA, phosphatase proteins, protein ligases, capping enzymes, and / or nicked RNA compared to unpurified RNA. 7.2. Assessment of Expression Levels and / or Activity of Therapeutic Products
[0520] The level of therapeutic product (such as polypeptide, protein, antibody or enzyme) can be determined by any method known in the art or as described herein. For example, the level of therapeutic product (such as polypeptide, protein, antibody or enzyme) in tissue sample can be determined by using, for example, Northern blot, PCR analysis, real-time PCR analysis or any other technical assessment (such as, quantification) of protein in sample. In one embodiment, the level of therapeutic product (such as polypeptide, protein, antibody or enzyme) in tissue sample can be determined by assessing (such as, quantification) the mRNA of protein in sample. The level of therapeutic product (such as polypeptide, protein, antibody or enzyme) in tissue sample can also be determined by using, for example, immunohistochemical analysis, Western blot, ELISA, immunoprecipitation, flow cytometry analysis or any other technical assessment (such as, quantification) of therapeutic product in sample or protein expression level. In a special embodiment, the method for the amount of therapeutic product present in the tissue sample of the patient (such as, in human serum) can be quantified, and / or the method for detecting the correction of protein level after treatment with circRNA or the preparation comprising circRNA is determined. 7.3. Screening assay
[0521] Various IRES-like sequence variants, endogenous IRES sequence variants, or combinations thereof can be tested for their ability to attract eukaryotic ribosomal translation initiation complexes and / or promote translation initiation. The following assays are described for IRES-like sequences, but can be similarly performed for endogenous IRES sequences, combinations of IRES-like sequences and endogenous IRES sequences, or sequences comprising one or more IRES-like sequences or endogenous IRES sequences. 7.3.1. Expression levels
[0522] For example, the impact of IRES-like sequences disclosed herein, endogenous IRES sequences or variants thereof or combinations thereof on the expression level of therapeutic products can be assessed. In some embodiments, the impact can be assessed by in vitro translation of RNA polynucleotides or circular RNAs comprising IRES-like sequences disclosed herein, endogenous IRES sequences or variants thereof or combinations thereof in a cell-free system. In some embodiments, the impact can be assessed by transfecting / transforming cells with RNA polynucleotides or circular RNAs comprising IRES-like sequences disclosed herein, endogenous IRES sequences or variants thereof or combinations thereof and expressing the RNA polynucleotides or circular RNAs. The expression level of therapeutic products can be assessed according to any method known in the art and / or described herein (e.g., immunohistochemical analysis, Western blot, ELISA, immunoprecipitation, and flow cytometry analysis).
[0523] In some embodiments, the ability of IRES-like sequences, endogenous IRES sequences, or variants thereof, or combinations thereof, identified based on their effect on the expression level of a therapeutic product (e.g., expression of a marker such as a fluorescent protein or luciferase) to promote, facilitate, or modulate (e.g., increase or decrease) the expression of a therapeutic product in vitro can be further assessed. In some embodiments, the ability of IRES-like sequences, endogenous IRES sequences, or variants thereof, or combinations thereof, to promote, facilitate, or modulate (e.g., increase or decrease) the expression of a therapeutic product in vivo can be further assessed in an animal model. In some embodiments, the ability of IRES-like sequences, endogenous IRES sequences, or variants thereof, or combinations thereof to promote, facilitate, or modulate (e.g., increase or decrease) the expression of a therapeutic product in a specific tissue or organ can be further assessed in an animal model.
[0524] Non-limiting illustrative examples of various assays or methods are provided below.
[0525] In vitro transfection and translation
[0526] The desired amount of circular RNA can be transfected into cells (e.g., prokaryotic or eukaryotic cells) using a transfection reagent such as Lipofectamine 3000 (Invitrogen). In vitro translation of the desired amount of circular RNA can be performed in a cell-free lysate. After incubation, the cell lysate can be collected to analyze protein expression.
[0527] Western blotting
[0528] Cells transfected with an RNA polynucleotide or circular RNA comprising an IRES-like sequence, an endogenous IRES sequence, or a variant thereof, or a combination thereof, according to the methods or techniques described herein can be lysed. For example, by lysing with 4%-20% ExpressPlus TM PAGE gel Total cell lysates are resolved by electrophoresis. The proteins are transferred to a membrane (e.g., a PVDF membrane) for detection using antibodies against the protein. A secondary antibody conjugated to the primary antibody can be used to stain the membrane for detection. Positive controls such as known Gtx, Rsv, CrPV, PSIV, or TSV IRES can be used.
[0529] Luciferase assay
[0530] Cells transfected with an RNA polynucleotide or circular RNA comprising an IRES-like sequence, an endogenous IRES sequence, or a variant thereof, or a combination thereof, of the present disclosure can be lysed according to the methods or techniques described herein. Luciferase reporter assays (e.g., from Promega TMA dual-luciferase reporter assay system (e.g., a luciferase reporter assay system) is used to generate a luminescent signal. Luminescence is measured, for example, by a Bio-Tek synergy H1.
[0531] Flow cytometry
[0532] Cells transfected with RNA polynucleotides or circular RNAs comprising an IRES-like sequence, an endogenous IRES sequence, or a variant thereof, or a combination thereof according to the methods or techniques described herein are collected for flow cytometry, for example, using a BD FACSAria II. To select single cells, SSC-A and FSC-A can be used to select 293T cells. Two rounds of selection for single cells can use SSC-W and FSC-H, and FSC-W and FSC-H. FITC-A and FSC-A can be used to select GFP-positive cells, and expression levels can be determined by fluorescence levels.
[0533] ELISA
[0534] ELISA can be performed according to Bull World Health Organ. 54(2):129-39(1976) (PMID:798633).
[0535] Optionally, the therapeutic product can be derivatized with other compounds and have a derivatization group that promotes the separation of the compound. The non-limiting example of a derivatization group includes biotin, fluorescein, digoxin (digoxygenin), green fluorescent protein, isotope, polyhistidine, magnetic beads, glutathione S transferase (GST), a photoactivatable cross-linking agent or any combination thereof. Optionally, the expression level of the therapeutic product can be assessed by measuring the level of the derivatization group. 7.3.2. Biological activity
[0536] The biological activity or function of the IRES-like sequences, endogenous IRES sequences, or variants thereof, or combinations thereof disclosed herein can be assessed according to any method known in the art and / or described herein (e.g., in vivo imaging and PET). For example, the biological activity can include inhibition of tumor growth, and the biological activity can be assessed in animal models such as cell line-derived xenografts (CDX) and patient-derived xenografts (PDX) models.
[0537] Non-limiting illustrative examples of various assays or methods are provided below.
[0538] In vivo imaging
[0539] To detect in vivo expression, female BALB / c mice aged 6-8 weeks can be used to administer RNA polynucleotides or circular RNAs (e.g., circular RNAs encoding luciferase) comprising an IRES-like sequence, an endogenous IRES sequence, or a variant thereof, or a combination thereof of the present disclosure. Administration can be via intramuscular (im), subcutaneous (sc), or intranasal (in) routes. At some time after administration, the animal is injected intraperitoneally (ip) with a luciferase substrate. Fluorescent signals are collected, for example, using an IVIS spectrometer (PerkinElmer). For in vitro imaging, tissues (including brain, heart, liver, spleen, lung, kidney, and muscle) from the animal are immediately collected, and the fluorescent signals of each tissue are measured, for example, using an IVIS imager. Fluorescent signals in regions of interest (ROIs) are quantified, for example, by using Living Image 3.0.
[0540] PET
[0541] Tumor xenografts are formed in mice using tumor cell lines. PET scans are performed before and after administration of an RNA polynucleotide or circular RNA comprising an IRES-like sequence, an endogenous IRES sequence, or a variant thereof, or a combination thereof, of the present disclosure to the animal. For example, a PET scan can be performed 1 hour after administration of 3.7 MBq to 7.4 MBq. A second PET scan can be performed at an appropriate time point after further administration.
[0542] For example, the assay can be used to determine the competitive binding capacity of an expressed therapeutic product in a biological system.
[0543] For example, the assays can be used to determine the enzymatic capacity of an expressed therapeutic product in a biological system.
[0544] For example, the assays can be used to determine the inhibitory activity of an expressed therapeutic product in a biological system.
[0545] Assays performed in cell-free systems (e.g., those that can be derived with purified or semi-purified proteins) are generally preferred as "primary" screens because they can be generated to allow rapid development and relatively easy detection of changes in the molecular target mediated by the test compound. Furthermore, in in vitro systems, the effects of the cytotoxicity or bioavailability of the test compound can generally be ignored, with the assay focusing primarily on the effect of the therapeutic product on the molecular target.
[0546] sequence
[0547] Table 1.: Sequence Listing Part 1 * Regarding the sequence of SEQ ID NO: 2 (group II intron Cte_original), nucleotides from exons are represented by capital letters. Nucleotides from exons that interact with the EBS region are represented by Underlined uppercase letters The six domains of group II introns are represented by Underlined lowercase letters Indicates that the EBS area is composed of Italic bold underlined lowercase letters instruct.
[0548] Table 2.: Group II intron sequences
[0549] Table 3. 3' intron fragments
[0550] Table 4: E2
[0551] Table 5.: E1
[0552] Table 6. 5' intron fragments
[0553] Table 7. 5' arms of target sequences
[0554] Table 8. 3' arms of target sequences
[0555] Table 9. Homology arm sequences
[0556] Table 10. Target sequences
[0557] Table 11.: Amino acid sequences
[0558] Table 12.: EBS1
[0559] Table 13. IBS1
[0560] Table 14.: δ and its upstream sequences
[0561] Table 15. IBS1 and its downstream sequences 8. Examples
[0562] In order to more fully understand and apply the present invention, the present invention will be described in detail below with reference to the embodiments and accompanying drawings. The embodiments are intended to illustrate the present invention by way of example only and are not intended to limit the scope of the present invention. The scope of the present invention is specifically defined by the appended claims. 8.1. Example 1. Screening of Group II Introns
[0563] This example relates to a method for confirming the in vitro self-splicing ability of natural group II introns.
[0564] First, according to the native sequence direct synthesis DNA sequence of type II intron (Genewiz, Suzhou), the synthetic DNA sequence, except type II intron sequence itself, also includes the exon E1, E2 sequence naturally present in flank, particularly all or part of the intron binding region in the exon next-door neighbour of type II intron. The DNA sequence is cloned into the modified expression vector psiCHECK-2 (Promega, C8021) containing Rluc (Renilla luciferase) coding sequence, T7 promoter and T7 terminator by molecular biology methods. Specifically, psiCHECK-2 endonuclease XhoI (New England Biolabs (NEB)) is subjected to single enzyme digestion, and then the synthetic DNA sequence is cloned into the vector through enzyme digestion using DNA seamless assembly method (AB clonal Technology, Wuhan), positioned at 3 ' downstream of Rluc, to obtain corresponding construct. The skeleton of this expression vector is the psiCHECK-2 comprising T7 promoter and T7 terminator.
[0565] Using universal primers targeting the T7 promoter and T7 terminator, PCR amplification was performed from the above vector to obtain template DNA for transcription. The PCR reaction conditions were: 95°C for 30 seconds, 60°C for 20 seconds, and 72°C for 60 seconds, for 23-25 cycles. The template DNA obtained by PCR was extracted with a 1:1 ratio of phenol to chloroform and then purified by precipitation with 2.5 volumes of anhydrous ethanol.
[0566] Purified template DNA was transcribed in vitro using T7 RNA polymerase (NEB or Promega) according to the manufacturer's instructions. The transcript was digested with DNase I at 37°C for 30 minutes to degrade the PCR template. The transcript was then purified by column purification to obtain highly pure RNA.
[0567] The column-purified transcript RNA was added to the auto-splicing buffer (10, 20, 50, or 100 mM MgCl2, 50 mM NaCl, 40 mM Tris-HCl, pH = 7.5) for auto-splicing reaction. The reaction conditions were 95°C for 1 min, 75°C to 45°C (-0.5°C, 15 sec / cycle, 60 cycles in total), 45°C hold and buffer addition, 45°C for 5 min, 53°C for 15-30 min (see Figure 2 B) After the in vitro self-splicing reaction, 200 ng of the product was analyzed by electrophoresis using a 1.5% agarose gel to detect the self-splicing efficiency of the group II intron.
[0568] If self-splicing occurs successfully, two RNAs of different sizes will be produced. The unspliced RNA is larger and is located at the top of the gel electrophoresis diagram; the spliced RNA is smaller and is located at the bottom of the gel electrophoresis diagram. For example, Figure 3 The electrophoresis results of two type II introns identified by the above method as being capable of self-splicing are shown, namely, type II intron Bth from Bacillus thuringiensis and type II intron Cte from Clostridium tetani ( Figure 3 ). Figure 3 The arrows in the figure respectively show the unspliced and spliced RNAs separated by electrophoresis. The group II intron and its flanking exon sequences confirmed by the method of this example (such as SEQ ID NO: 1 or SEQ ID NO: 2, which respectively contain the group II introns of Bth and Cte and their flanking 6 nucleotides of E1 and E2) are used as precursors for preparing the self-splicing ribozyme construct cRNAzyme, or cRNAzyme precursors.
[0569] This plasmid contains the native group II intron, EGFP, and T7 promoter / terminator sequences and was synthesized by Sangon Biotech. Template DNA was obtained by digesting the plasmid with PmeI or XbaI (Thermo Fisher), followed by purification with 2.5 volumes of anhydrous ethanol precipitation. The purified template was transcribed in vitro using T7 RNA polymerase (NEB or Promega) according to the manufacturer's recommended conditions. The transcript was treated with DNase I for 30 minutes at 37°C to degrade the template and then purified by column chromatography to produce highly pure RNA. For analysis, 200 ng of RNA product was electrophoresed on a 1.5% agarose gel to assess the efficiency of group II intron self-splicing. If self-splicing occurs successfully, two RNAs of different sizes will be produced. The unspliced RNA is larger and appears at the top of the gel electrophoresis pattern; the spliced RNA is smaller and appears at the bottom of the gel electrophoresis pattern. The group II intron and its flanking exon sequences identified by this method can be used as precursors for preparing self-splicing ribozyme constructs, cRNAzymes. 8.2. Example 2. Preparation of expression constructs containing group II introns
[0570] Based on the type II intron cRNAzyme precursors with self-splicing properties obtained through screening, self-splicing ribozyme cRNAzyme constructs were further prepared. As mentioned above, the general principles for designing cRNAzyme constructs are to minimize the total length of the intron sequence and the E1 and E2 sequences and to maximize the cyclization rate.
[0571] This example describes in detail the process of designing and preparing cRNAzyme constructs using the Cte cRNAzyme precursor screened in Example 1.
[0572] Based on the Cte cRNAzyme precursor sequence (SEQ ID NO:2, 1,028 nucleotides long, including the Cte intron itself and exon sequences of 6 nucleotides each at either end, i.e., E1 and E2, both of which are 6 nucleotides long), the 310 nucleotide intron-encoded protein (IEP) sequence in domain 4 (nucleotide positions 625-934 of SEQ ID NO:2) was deleted. Sequences that support proper folding and autosplicing activity, including a small number of exons (6 nucleotides each at either end, i.e., IBS1 and IBS3), were retained to create the E1-CteΔIEP-E2 sequence. The E1-CteΔIEP-E2 sequence was then split at a position within the intron. The two fragments were swapped, with the first fragment consisting of E1 and the 5' intron fragment inserted into the 3' end of the Rluc insert, and the second fragment consisting of the 3' intron fragment and E2 inserted into the 5' end of the Rluc insert. On this basis, AATACCTTACTTAATAGTAACAA TAGAAAATC (SEQ ID NO: 14) was inserted into the 5' end of the newly formed fragment, and AAGCTAGATCATATTACTATTAAGTAAGGTATT (SEQ ID NO: 15) was inserted into the 3' end to obtain the cRNAzyme_Cte construct. The two inserted sequences of SEQ ID NO: 14 and SEQ ID NO: 15 act as "homologous arms" to bring the 5' and 3' splice sites closer to each other, thereby improving splicing efficiency. When splitting the intron, three different split positions were tried, located in the loop region in domain 1 (between positions 369 and 370), domain 3 (between positions 560 and 561), and domain 4 (between positions 825 and 826), respectively. Thus, three cRNAzymes were formed, namely cRNAzyme_Cte V1, cRNAzyme_Cte V2, and cRNAzyme_Cte V3. In the presence of cations, a self-splicing cyclization reaction was performed in vitro to test the cyclization activity of the obtained cRNAzyme. Figure 4 b As can be seen from FIG. 2 , the splicing effects are different after segmentation at different positions, among which the third segmentation method (V3) has the best splicing effect, and was used for subsequent experiments. The sequence of the construct is shown in SEQ ID NO: 16.
[0573] According to the fragment size, Figure 4The band shown as a circle in the gel electrophoresis diagram of b represents circular RNA. The sequence was verified by RT-PCR and Sanger sequencing, and it was determined that the band contained a sequence that spanned the E1E2 junction, proving that E1 and E2 had been connected together. However, it is still not possible to confirm that it is a circular RNA, because the band may also be a transcript from reverse splicing, or other sequences. In order to confirm the successful formation of circular RNA, based on the characteristic of covalent closure of circular RNA head and tail, the following three methods were used to verify its structure: Figure 4 c).
[0574] Method I
[0575] The splice site in Cte was mutated to lose its ability to cyclize, as a control for a linear RNA of the same length but unable to cyclize. Specifically, the splice site in SEQ ID NO: 2 (nucleotides 1-26) was mutated, and the C at position 3 was changed to A, the T at position 5 was changed to G, the G at position 17 was mutated to T, the C at position 18 was mutated to T, the A at position 21 was mutated to C, and the T at position 26 was mutated to G. Those skilled in the art will understand that this mutation is intended to destroy the ability to cyclize, and other mutations of different numbers, positions, and types can also be performed to achieve similar purposes. The mutated Cte is referred to as Cte-mut (SEQ ID NO: 3). Polyadenylation is then used to add a polyA tail to this mutated linear RNA. Since circular RNA is closed at both ends and has no 3' end, it cannot be tailed, while linear RNA can be added with hundreds of adenylate nucleotides through this reaction, and the change in RNA size before and after tailing can be distinguished by agarose gel electrophoresis.
[0576] The specific steps include: (1) Purified circular RNA or control linear RNA was tailed with poly(A) tailing enzyme (NEB) at 37°C for 30 min. (2) Column purification of RNA.
[0577] like Figure 4 As shown in the upper panel of c, lanes 3 and 4 represent products with polyA tails added, while lanes 1 and 2 represent products without polyA tails. In the absence of circularization, the bands of the Cte- and Cte-mut-based cRNAzyme precursor RNAs shift upward after polyA addition, indicating molecular enlargement and successful polyA addition. In contrast, the bands of the presumed uncircularized RNAs are essentially located and of the same size with and without polyA addition.
[0578] Method II
[0579] Digest with RNase R. RNase R is a 3'-5' exoribonuclease that degrades linear RNA molecules, but not circular RNA. Agarose gel electrophoresis can be used to determine if the RNA has been digested.
[0580] The specific steps include: (1) Purified circular RNA or control linear RNA was tailed with poly(A) tailing enzyme (NEB) at 37°C for 30 min. (2) RNase R (Lucigen) was added to digest the linear RNA at 37°C for 30 min. (3) Column purification of RNA.
[0581] like Figure 4 As shown in the upper panel of C, lanes 5 and 6 show the results after RNase R treatment, while lanes 1-4 show the results without RNase R treatment. It can be seen that the large linear RNA band disappears after RNase R treatment (lanes 5 and 6), while the band in lane 5, presumed to be circular RNA, still exists.
[0582] Method III
[0583] Digestion is performed using RNase H. RNase H is an endoribonuclease that specifically hydrolyzes RNA in DNA-RNA hybrid chains. Because linear RNA and circular RNA have different structures, they can be cut into fragments of varying lengths by RNaseH after binding to the same DNA probe. Agarose gel electrophoresis can be used to distinguish the lengths of the RNA fragments produced by the cuts, thereby inferring the original structure of the RNA. Specifically, for circular RNA, two DNA probes are used to bind to the RNA and then cut, resulting in two bands. In contrast, if there is no circularization and the RNA remains linear, the same method should yield three bands. (1) The DNA probe is bound to the RNA by annealing at 95°C for 2 min, followed by a gradual cooling to 25°C. The probe sequences are shown in SEQ ID NOs: 7 and 8. (2) DNA / RNA duplexes were digested with RNase H (NEB) at 37°C for 30 min. (3) Column purification of RNA. Figure 4 As shown in the lower panel C, the product of this application obtained two bands.
[0584] The results obtained by the above three methods all confirmed the successful formation of circular RNA and verified the cyclization activity of the constructed cRNAzyme_Cte. 8.3. Example 3. Improving self-splicing efficiency by optimizing the reaction system and modifying the construct
[0585] In order to increase the final circular RNA yield, it is first necessary to improve the circularization efficiency ( Figure 5 a). The inventors optimized the circularization efficiency of the expression construct from two aspects.
[0586] Optimization of reaction conditions
[0587] The inventors tried a variety of different ion concentrations in the reaction system (50mM and 100mM NaCl; 2mM, 5mM, 10mM, 20mM Mg2+) and a variety of different reaction time combinations (5 minutes, 15 minutes, 30 minutes) to determine the optimal reaction system.
[0588] It was found that in a reaction system of 20 mM Mg2+ and 50 mM NaCl, the cyclization efficiency could be increased from the current 30% to more than 60% under the conditions of 15 and 30 minutes ( Figure 5 B).
[0589] Sequence optimization
[0590] The sequence was further modified based on RNA secondary structure. Specifically, after inserting different target sequences, some sequences may not be effectively spliced due to structural reasons. In such cases, splicing efficiency can be improved by adding spacer sequences to the target sequence to increase structural flexibility. For example, the spacer sequence can be an AT-rich sequence.
[0591] Based on Example 2, three different spacer sequences were inserted into the front end of Rluc in cRNAzyme_Cte by molecular cloning, namely spacer sequence 1 of SEQ ID No: 4, spacer sequence 2 of SEQ ID NO: 5, and spacer sequence 3 of SEQ ID NO: 6, to obtain three further optimized constructs. These three precursor RNAs with spacer sequences were circularized in vitro in the optimal self-splicing reaction system determined in Example 2 (10 mM Mg, 50 mM NaCl, 30 minutes reaction time). Figure 5 The results of C show that the cyclization efficiency after adding three spacer sequences is approximately 60%, 80%, and 98%, respectively. Figure 5 c). 8.4. Example 4. Preparation of “scarless” circular RNA without scar sequence
[0592] The circular RNA obtained using the aforementioned method still contains a small amount of non-target sequences, namely sequences from exons E1 and E2. To remove these sequences, the construct can be further modified.
[0593] When preparing the cRNAzyme construct, the two ends of the target sequence are respectively E2 and E1 of shorter length, which are derived from the intron binding sequence (IBS) of the exon region flanking the type II intron, generally between 0-20 nucleotides in length. When forming circular RNA, it is not desirable to include sequences other than the non-target sequence, such as the exon sequences E1 and E2. If E1 and E2 are directly removed, the self-splicing cyclization process will be affected due to the lack of IBS sequences that interact with the EBS or δ sequences in the intron. The inventors of the present invention have innovatively thought that a part of the target sequence can be directly regarded as an "IBS" sequence, and by modifying the EBS or δ in the intron, it can interact with the region regarded as "IBS" in the target sequence. Such a method, while ensuring that the cRNAzyme construct has self-splicing function, gets rid of the dependence on the exon sequences E1 and E2 and removes them from the construct and the final cyclization product.
[0594] Still taking Cte ribozyme as an example, different target sequences (GFP, Gluc and 2A peptide) are used to illustrate the design concept of the cRNAzyme construct of this embodiment.
[0595] Based on the 6 nucleotide sequences at both ends of each target sequence, EBS1 in the group II intron and the upstream sequence of EBS1 (including δ) are replaced with sequences that are at least partially complementary to these two 6 nucleotide sequences. Specifically, the upstream sequence of EBS1 (including δ) is made to complementarily pair with the 6 nucleotides at the 5' end of the target sequence when it is in a linear state (such as the state before the cRNAzyme construct is self-splicing and cyclized), and EBS1 is made to complementarily pair with the 6 nucleotides at the 3' end of the target sequence when it is in a linear state. The specific modified sequences are shown in the updated Figure 6 The right panel of Figure B shows the modified EBS1 and upstream sequences of EBS1 (including δ) designed for three different target sequences. It is worth noting that the modified EBS1 and upstream sequences of EBS1 (including δ) do not have to be perfectly complementary to the target sequence fragments serving as "IBS1" and "IBS3". A certain proportion of mismatches or less robust pairings such as A and G, or G and U can be tolerated. Generally speaking, the upstream sequences of EBS1 and EBS1 (including δ) used to replace the corresponding sequences in group II introns are complementary to the corresponding regions of the target sequences at at least 60% of the nucleotide positions, or are at least 60% identical to the complementary pairing sequences of the corresponding regions of the target sequences.
[0596] From the electrophoresis results, we can see that the transformation method effectively generated circular RNA ( Figure 6B, left). Sanger sequencing results confirmed that this modification can completely eliminate the exogenous scar sequence ( Figure 6 B, right). Figure 6 As shown in Figure B, after cyclization, the cRNAzyme construct with the modified EBS and upstream sequences of EBS1 (including δ) has the target sequence at both ends (two shaded 6-nucleotide sequences, serving as "IBS1" and "IBS3") directly connected end to end, without being separated by E1 and E2 in the middle. This is more conducive to the subsequent application of the generated circular RNA. The construct modified in this way is called a "scarless" construct, and the circular RNA after cyclization is called "scarless" RNA. 8.5. Example 5. Circularization results of target sequences of different lengths
[0597] Based on the methods described in Examples 2 and 3, the inventors also tested target fragments of varying lengths. These target fragments included a 555-nucleotide Gluc construct, whose nucleotide sequence is shown in SEQ ID NO:9; a 936-nucleotide Rluc1 construct, whose nucleotide sequence is shown in SEQ ID NO:10; and a 1,160-nucleotide Rluc2 construct, whose nucleotide sequence is shown in SEQ ID NO:11. No spacer sequence was added to the Gluc and Rluc1 constructs; the Rluc2 construct consisted of a Cat1 IRES sequence and Rluc.
[0598] It was found that all the tested fragments could be efficiently self-spliced to obtain circular RNA products ( Figure 7 ). 8.6. Example 6. Expression of target genes using the constructs of the present invention
[0599] On this basis, the expression of circular RNA products with different target sequences after transfection of cells was further tested.
[0600] To minimize immunodegradation caused by linear RNA, the RNA product was treated with the following three steps before transfection. (1) RNase R treatment to digest linear RNA. The reaction conditions are 37°C for 30 min. (2) HPLC purification to remove small linear RNA. HPLC conditions: Gel exclusion chromatography column: Waters BEH450A, column temperature: 40°C, flow rate: 1 min / ml, elution conditions: 0-30 min, 100% buffer A (10 mM Tris, 0.5 mM EDTA, DEPC water). CIP treatment was performed to remove phosphate groups from the linear RNA terminals. The reaction conditions were the addition of quick CIP (NEB) and incubation at 37°C for 30 min.
[0601] The target RNA was transfected using lipo RNAmax (Invitrogen). The transfection conditions were as per the instructions of the supplier, and the transfection time was 24 hours.
[0602] Using the construction methods described in Examples 2 and 3 (using spacer sequence 2), and the modifications described in Example 4 that enable traceless cyclization, different target sequences were tested. To facilitate detection of protein expression, constructs containing fluorescent protein coding sequences, including IRES-GFP (SEQ ID NO: 12) and IRES-Gluc (SEQ ID NO: 13), were constructed. The addition of IRES can initiate non-canonical translation that is independent of the cap structure, enabling the coding sequence in the circular RNA to be translated into protein.
[0603] Different methods were used to detect protein expression for different target sequences.
[0604] In the case where the translation product is GFP, the cells were lysed with RIPA lysis buffer (Beyotime) by observing fluorescence under a microscope, and then protein expression was detected by Western blotting. It was determined that the expression of GFP was obtained by the method of the present invention ( Figure 8 A).
[0605] In the case where the translation product is luciferase, cells were first lysed with Passive Lysis Buffer (Promega), and then protein expression was detected by microplate reader using a luciferase detection kit (Promega). Figure 8 B). 8.7. Example 7. In vitro design and production of circular RNA
[0606] To efficiently generate circular RNAs, we exploited the autocatalytic splicing reaction of group II introns, mobile genetic elements found primarily in bacterial and organelle genomes (Lambowitz, AM and Zimmerly, S. (2011). Group II introns: mobile ribozymes that invade DNA. Cold Spring Harb Perspect Biol 3, a003616). All group II introns have six domains (D1 to D6), of which domain 1 (D1) is the largest domain and contains several short exon binding sites (EBS) that determine splicing specificity ( Figure 12A). Domains 2 and 3 play a key role in assembling active intron structures and stimulating splicing reactions, while D4 is a stem-loop structure in which the long loop contains the ORF of the mature enzyme. The highly conserved D5 is the key to the self-splicing active site (Rybak-Wolf et al. (2015). Circular RNAs in the Mammalian Brain Are Highly Abundant, Conserved, and Dynamically Expressed. Mol Cell 58, 870-885.), while D6 contains a protruding adenosine as a branching site ( Figure 12A ). In vitro self-splicing of group II introns only requires the correct folding of the intron RNA structure and Mg2+ (Peebles, CL et al. (1986). A self-splicing RNA excises an intron lariat. Cell 44, 213-223), whereas in vivo splicing requires the assistance of mature enzymes (Pyle, AM (2016). Group II Intron Self-Splicing. Annu Rev Biophys 45, 183-205).
[0607] Based on this domain configuration, we split the type II self-splicing intron of the surface layer protein from Clostridium tetani (McNeil, BA, Simon, DM and Zimmerly, S. (2014). Alternative splicing of a group II intron in a surface layer protein gene in Clostridium tetani. Nucleic Acids Res 42, 1959-1969) at D4 to generate a split intron system containing custom exons flanked by upstream D5-D6 and downstream D1-D2-D3 of the intron. A portion of the D4 stem was isolated and placed at each end of the resulting RNA, thereby forming a complementary structure to assist in the folding of the active intron ( Figure 12A , green line). Two short 6nt sequences, IBS3 (intron binding site 3) and IBS1, are included at each junction between the intron and the custom exon to provide long-range interactions with the intron's EBS3 and EBS1. After uncontrolled transcription in vitro, the resulting RNA precursor can be self-splicing by the group II intron to produce a circular RNA of the custom exon and a branched intron RNA. By including an IRES sequence upstream of the target gene (GOI), the resulting circular RNA can function as an mRNA to guide protein synthesis through cap-independent translation ( Figure 12AWe named this system circular coding RNA (CirCode), which can be used as a universal platform to generate any given circRNA for protein translation.
[0608] To validate this design, we included the IRES and ORF of the Renilla luciferase gene in a custom exon and translated the D5-D5-exon-D1. We found that the resulting RNA precursor could indeed self-splice in vitro to produce an additional band corresponding to the circRNA, while mutation of the splice junction (at IBS1) failed to produce a circRNA ( Figure 1 B). We further used two different methods to confirm that the additional band below the precursor RNA was indeed a circRNA, as it could not be extended by poly A polymerase tailing reaction and was resistant to digestion by RNase R treatment ( Figure 13 ). In addition, we purified the circRNA from gel and used RNase H digestion of the linker ( Figure 14 ) and direct sequencing ( Figure 15 ) confirmed the circular identity. Finally, we optimized the reaction conditions by using different concentrations of MgCl2 and NaCl in the circularization buffer after the IVT step (see Methods) and found that 10-20 mM Mg2+ with 50-100 mM NaCl was the optimal reaction condition for RNA circularization ( Figure 5 B) We also observed some degree of RNA circularization (about 30%) even when the circularization buffer did not contain any Mg2+, likely because RNA circularization occurred cotranscriptionally in the IVT buffer containing 24 mM MgCl2.
[0609] Platform optimization for circRNA production and translation
[0610] Previous self-splicing systems that used group I introns to generate circRNAs also introduced unrelated sequences from T4 phage or Anabaena into the final product. This “scar” sequence is typically about 80-180 nt long, which limits the design flexibility of the target circRNA and may introduce some unwanted effects during drug development. Our initial design used two short sequences (IBS1 and IBS3) for intron-exon recognition, which left a shorter 12nt scar. To reduce the potential interference of the scar sequence, we further modified the design by changing the exon binding site in the D1 domain so that EBS1 and EBS3 form base pairs with the 3' and 5' ends of the circular exon, respectively ( Figure 16, left). This new design enables the generation of "traceless circRNAs" without extraneous sequences. The only sequence requirement is an "AG" at the 5' end of the circular exon, which makes the design of circular RNAs fully programmable. We validated this design by generating two different circRNAs encoding EGFP and Rluc-P2A ( Figure 16 , right), and observed the effective circularization of traceless circRNA under different ionic conditions. The sequencing of the final circRNA product further confirmed the traceless self-splicing of circRNA ( Figure 16 In summary, we designed the CirCode system, which can efficiently generate scarless circRNAs of any sequence, providing a universal platform for using circRNAs as diverse therapeutic tools.
[0611] Previous studies using the PIE approach have shown that adding a short spacer before the IRES may contribute to the correct folding of the IRES and / or the active structure of the intron (Wesselhoeft, RA et al., 2018, supra), so we introduced several forms of spacer sequences at each end of the circular exon to optimize their circularization efficiency ( Figure 17 We found that different spacers indeed affected RNA circularization efficiency, which ranged from 47% (for SP5) to 83% (for SP4). Furthermore, we transfected two circRNAs with different spacers into three cell lines at different doses and observed a dose-dependent increase in protein production as judged by luciferase activity assay ( Figure 18 Interestingly, the effect of the spacer sequence on translation efficiency appeared to be minimal and dependent on circRNA dose and cell line. We further examined the time course of protein production after circRNA transfection and found that active protein could be detected within six hours after transfection, with peak expression at 24 hours ( Figure 19 ). 8.8. Example 8. circRNA Mediates Extended Protein Production
[0612] The main advantage of circular mRNA is its excellent stability due to the lack of free ends, so compared with its linear counterpart, circRNA should have a good shelf life for protein expression. To test this directly, we synthesized linear and circular mRNA encoding Gaussia luciferase (Gluc) and stored them in parallel in pure water at room temperature for different days before transfection into 293T cells. We found that the activity of circRNA in directing protein translation remained essentially unchanged over two weeks, while linear mRNA lost about half of its activity by the third day of storage ( Figure 3A), indicating that circRNAs can be stably stored at room temperature. Furthermore, we found that protein production from linear mRNAs declined rapidly after transfection into cells, dropping by >80% by day 2, whereas translation from circRNAs persisted for 5–7 days ( Figure 3 B). The prolonged protein translation is also consistent with our previous results using backsplicing circRNA reporters (Wang, Y. and Wang, Z. (2015). Efficient backsplicing produces translatable circular mRNAs. RNA 21, 172-179). It is also worth noting that in this experiment, protein production from circRNAs may also be inhibited at the end because we did not change the culture medium during the entire week of the experiment ( Figure 16 ).
[0613] We further compared protein production from linear mRNAs and circRNAs generated using either the PIE protocol or the novel CirCode system. Capped and unmodified linear mRNAs with identical Gluc coding sequences were generated using IVT and transfected into 293T cells in parallel with two types of circRNAs containing the CVB3 IRES and Gluc ORF. We found that protein expression from both circRNAs was more robust compared to that from the linear mRNAs ( Figure 23 , supra), supporting previous reports using PIE circRNAs (Wesselhoeft, RA et al., 2018, supra). When we measured accumulated luciferase activity over a 6-day time span, an even greater increase in protein production was observed in the circRNAs ( Figure 23 , below), which may be due to the excellent stability of circRNA. In addition, circRNAs generated from two different protocols showed similar abilities in directing protein translation ( Figure 24 ). 8.9. Example 9. Purified circRNA can guide robust translation of target proteins
[0614] It was found that mRNA purity is a key factor in protein production and induction of innate immunity, because the removal of dsRNA by HPLC can eliminate immune activation and improve the translation of linear nucleoside-modified mRNA (Kariko, K. et al., (2011). Generating the optimal mRNA for therapy: HPLC purification eliminates immune activation and improves translation of nucleoside-modified, protein-encoding mRNA. Nucleic Acids Res 39, e142). However, there is some controversy about the immunogenicity of circRNA. Although early reports showed that in vitro synthesized circRNAs are more likely to induce cellular immune responses than linear RNAs (Chen, YG et al., (2017). Sensing Self and Foreign Circular RNAs by Intron Identity. Mol Cell 67, 228-238e225), it was later reported that purifying circRNAs from IVT and cyclization reaction byproducts (including dsRNA, linear RNA fragments, and triphosphate RNA) can eliminate the cytotoxicity and immunogenicity of circRNAs (Wesselhoeft, RA et al., (2019). RNA Circularization Diminishes Immunogenicity and Can Extend Translation Duration In Vivo. Mol Cell 74, 508-520e504; Breuer, J. et al., (2022). What goes around comes around: artificial circular RNAs bypass cellular antiviral responses. Mol Ther Nucleic Acids). Recent studies have also shown that sequence identity and structure are the main determinants of circRNA cellular immunity, as circRNAs produced by different methods show different immunogenicity (Liu, CX et al., (2022). RNA circles with minimized immunogenicity as potent PKR inhibitors. Mol Cell 82, 420-434e426.).To examine whether circRNAs generated using the CirCode platform could induce innate immune responses and cytotoxicity, we purified the circRNAs by gel purification or HPLC (. Figure 3 D), and measured whether circRNA could induce cellular immune responses after circRNA transfection. We found that transfection of unpurified circRNA led to substantial cell death compared with mock transfection, while purified circRNA showed no detectable cytotoxicity ( Figure 25 Furthermore, the purified circRNAs showed significantly reduced immunogenicity compared to unpurified circRNAs that stimulated innate immune responses by inducing RIG-I and IFN-B1 (interferon-β1). Figure 26 ), supporting previous observations that the immunogenicity of circRNAs is primarily caused by byproducts of IVT reactions (Wesselhoeft, RA et al., 2019, supra). 8.10. Example 10. LNP-encapsulated circRNA directs robust protein production in mice
[0615] An important question for the therapeutic application of circRNAs is whether the production of circRNAs can be reliably amplified and how reproducible it is between different production batches. Because the CirCode platform uses self-splicing introns for RNA circularization without the involvement of RNA ligase or other co-activators, the amplification procedure is relatively simple. To test the scalability of the system, we scaled up the IVT and circularization reaction system 50-fold (from 20 μl to 1 ml) and found that the high circularization efficiency (approximately 70%) remained essentially unchanged, while the total amount of RNA product reached 7.5 mg in a single reaction ( Figure 27 Furthermore, we found that IVT / circularization and HPLC purification of circRNAs were highly reproducible between different batches ( Figure 28 ), laying the foundation for the in vivo application of circRNA.
[0616] We further generated lipid nanoparticle (LNP)-encapsulated circRNAs for their in vivo delivery ( Figure 29). circRNA in aqueous solution is filled with ionizable cationic lipids, which form nanoparticles with other lipid components (such as DMG-PEG2000 and cholesterol) (see Methods), achieving an encapsulation efficiency of approximately 95% with an effective diameter of approximately 80 nm. We tested three different ionizable cationic lipids in our formulation to encapsulate circRNA encoding Gluc, and the resulting LNPs were injected into BALC / c mice by intramuscular (IM) or intraperitoneal (IP) injection (n=3 per experimental group). Three formulations using different ionizable cationic lipids (MC3, SM-102, and ALC-0315) were tested in this experiment. Serum luciferase luminescence assay ( Figure 30 ) or bioluminescence imaging of animals ( Figure 31 Luciferase expression was measured using a luciferase assay. We found robust expression of luciferase from two different LNP-circRNA formulations, demonstrating that circRNAs generated by our procedure can reliably induce protein expression in vivo. 8.11. Example 11. Generation of Potential CircRNA Vaccines for SARS-CoV-2
[0617] We further tested the application of circRNA in mRNA therapy by engineering circRNA encoding the receptor binding domain (RBD) of the S protein from SARS-CoV-2, which can potentially be used to produce mRNA vaccines. Based on previous reports, two different antigens were designed and constructed into the circRNA ( Figure 32, on) (Dai, L. and Gao, GF (2021). Viral targets for vaccines against COVID-19. Nat Rev Immunol 21, 73-82). The first used a single RBD fused to a foldon from T4 minor fibrin (fibritin), which enables trimerization of RBD (Meier, S. et al., (2004). Foldon, the natural trimerization domain of T4 fibritin, dissociates into a monomeric A-state form containing a stable beta-hairpin: atomic details of trimer dissociation and local beta-hairpin stability from residual dipolar couplings. J Mol Biol 344, 1051-1069). Another method uses an RBD dimer that has been shown to induce strong antibody production in experimental studies (Dai, L. et al. (2020). A Universal Design of Betacoronavirus Vaccines against COVID-19, MERS, and SARS. Cell 182, 722-733e711; Dai, L. et al., 2021, supra). We designed the coding sequences of the two proteins using a modified IRES and constructed a circRNA vector. The production and purification of the resulting circRNA were verified ( Figure 32 , lower left), and further validated by transfection into 293 cells and detected by western blot antigen production for the translation of circRNA-encoded proteins ( Figure 32 , lower right).
[0618] We used two different formulations to generate LNP-circRNA particles and achieved high encapsulation efficiency (>90%) with typical nanoparticle sizes of 90-100 nm ( Figure 33 Representative results using SM-102 formulation are shown). LNP-circRNA was inoculated into BALB / c mice with two separate IM injections, and blood samples were collected 2 weeks after each inoculation for further analysis ( Figure 34We tested several formulations and circRNA designs, however, for better comparison with published results, results using the SM-102 formulation of LNP-circRNA-RBP are shown. Using flow cytometry, we found that the majority of memory B cells (i.e., CD19+CD27+ B lymphocytes) had robust RBD antibody expression two weeks after the second vaccination ( Figure 35 、 36 ), indicating a durable immune response induced by antigen. Specifically, >70% of CD19+CD27+ B cells expressed RBD-specific antibodies (i.e., RBD+), and approximately 50% of CD19+CD27+ B cells also expressed IgD+RBD+ ( Figure 36 ), indicating a strong antibody response. As a control, injection of LNP alone did not induce activation of RBD antibodies, indicating that a specific response would occur after vaccination.
[0619] To test the activity of RBD antibodies in mouse serum, we next performed an antibody blocking assay to examine whether mouse serum could block the binding of fluorescently labeled RBD to 293T cells stably expressing ACE2 ( Figure 37 ). We found that the serum could effectively block the binding of all RBD variants to the cell surface at a moderate dilution (1:20). This protection is very impressive given that we used a fairly high concentration of RBD (0.5μg / ml). Even at a very high dilution ratio (1:200), the serum still completely blocked the binding of the wild-type RBD and significantly reduced the RBD binding of the delta variant. However, the serum did not block the binding of the omikon RBD, which is consistent with the finding that current mRNA vaccines based on the wild-type S protein have weak protection against the omikon variant.
[0620] We further measured the titers of neutralizing antibodies against RBD after each vaccination and found robust production of IgG against RBD ( Figures 38-39 ). Antibody titers approached 105 two weeks after the first injection and reached nearly 108 after the second injection, which is at least 10 times higher than the results reported in mice using linear mRNA encapsulated by LNPs with similar formulations or using another circRNA design. In addition, the IgG2 / IgG1 ratio was approximately 0.75, which is similar to that previously reported using linear mRNA vaccines. Finally, we measured antibody titers using a protection assay with SARS-CoV-2 pseudovirus and found strong protection with titers approaching 105 ( Figure 40Overall, our preliminary data in mouse models demonstrate the robust performance of the CirCode platform for designing and generating circRNA vaccines against SARS-CoV-2, which can be potentially extended to vaccine development for other viral pathogens.
[0621] Plasmid construction
[0622] Fragments of the group II intron and IRES sequence from Clostridium tetani (CTE) were chemically synthesized from GENEWIZ, and the different protein-coding fragments were amplified by PCR. These fragments were cloned into an NheI- and XbaI-digested backbone containing the T7 RNA polymerase promoter and terminator by Gibson assembly. Fragments of the group II intron and IRES sequence from BR23 and CL were similarly constructed, and the different protein-coding fragments were amplified by PCR.
[0623] RNA synthesis and circularization
[0624] In the presence of unmodified NTPs, T7 RiboMAX TM Large-scale RNA production system MEGAscript TM RNA was in vitro transcribed from an XbaI-digested, linearized plasmid DNA template using a T7 (Thermo, AMB1335-5). After DNase I treatment, the RNA product was purified using an RNA Clean and Concentrator Kit (ZYMO research, R1019) column to remove excess NTPs and other salts in the IVT buffer, as well as potential small RNA fragments generated during IVT. In some experiments, the purified RNA was further circularized in fresh circularization buffer. The RNA was first heated to 75°C for 5 minutes and rapidly cooled to 45°C. A buffer containing the specified magnesium and sodium concentrations (50 mM Tris-HCl (pH 7.5), 50 / 100 mM NaCl, and 0–40 mM MgCl2) was then added to the final concentrations, followed by heating at 53°C for the specified time for circularization. The best optimized reaction conditions (including magnesium and sodium concentrations and incubation time at 53°C) were selected for further experiments.
[0625] Example 12. circRNA identification
[0626] For poly A tailing and RNase R treatment, total RNA from IVT was purified using an RNA cleanup column and then treated with Escherichia coli (E. coli) Poly A polymerase (NEB, M0276S) according to the manufacturer's instructions. This step adds a poly A tail to the free end of the unspliced RNA precursor. After poly A tailing, the purified RNA was digested with RNase R exoribonuclease (Lucigen, RNR07520) according to the manufacturer's instructions, and the enriched circRNA was purified by column purification.
[0627] For the RNase H nicking assay, we incubated a 24nt ssDNA probe at a 1:20 ratio with the RNase R-enriched circRNA at 65°C for 5 minutes. RNase H buffer was immediately added to the DNA-RNA mixture. The mixture was then slowly cooled to room temperature. After annealing, RNase H (Thermo Scientific, EN0201) was added to the mixture at 37°C for 20 minutes. The sequence of the ssDNA probe was 5'-TGGTGCTCGTAGGAGTAGTGAAAG-3'.
[0628] RNA gel electrophoresis and purification
[0629] RNA was run on a low-melting-point agarose gel (Sigma-Aldrich, A4018) at 120 V using ice-cold DEPC-treated MOPS buffer. After electrophoresis, the circRNA lane was excised and purified using the Zymoclean Gel RNA Recovery Kit (ZYMO Research, R1011). Prior to transfection, the column-purified circRNA was treated with phosphatase (NEB, M0525S) to remove potential 5′ phosphates that could be immunogenic.
[0630] Measurement of circRNA translation products
[0631] The cells were seeded into 24-well plates one day before transfection. Purified circRNA was transfected into cells using Lipofectamine Messenger Max (Invitrogen, LMRNA001) according to the manufacturer's manual. After transfection, the cells were cultured at 37°C for 24 h. Cell lysates and supernatants were collected for use. Luminescence assays were performed using a reporter assay system (Promega, E1910).
[0632] RT-PCR of circular RNA
[0633] The RNA purified from the IVT column was electrophoresed on an agarose gel. The band corresponding to the circRNA was excised and extracted using the Zymoclean Gel RNA Extraction Kit (Zymogen, R1011). The purified RNA was reverse transcribed into cDNA using a PrimeScript RT kit (TAKARA, RR037B) with random primers, and PCR was then performed using primers that can amplify the transcript across the splice junction. The PCR product was subjected to Sanger sequencing to confirm the reverse splicing of the circular RNA.
[0634] HPLC purification of circular RNA
[0635] To obtain high-quality circRNAs, DNase I-treated RNA purified from IVT spin columns was resolved by high-performance liquid chromatography (HPLC). FPLC purification was performed to remove condensates, aggregates, and small RNA fragments. FPLC conditions: size exclusion chromatography (Sepax, 215950-30030), flow rate: 10 ml / min, elution conditions: 0 to 30 minutes, elution buffer: 75 mM PB buffer containing 10 mM Tris, 0.5 mM EDTA, and water for injection (pH 7.4). Second, affinity chromatography purification was performed to remove precursor and spliced intronic RNAs. Affinity chromatography columns: HiTap NHS-activated HP (Cytiva, 17071601) coupled to a specific ligand; flow rate: 1.5 ml / min, elution conditions: 0 to 15 minutes, binding buffer: 15 mM LiCl, 10 mM Tris, 0.5 mM EDTA, and water for injection. Finally, an optional RNase R treatment was performed to remove nicked RNA. The reaction was carried out at 37°C for 30 minutes. After column purification using the RNA Clean and Concentrate Kit (ZYMO Research, R1019), the purified circular RNA was ready for further experiments.
[0636] Analysis of circular RNA by capillary electrophoresis
[0637] Circular RNA purified from HPLC was further analyzed by capillary electrophoresis in RNA mode using an Agilent 2100 Bioanalyzer. Samples were diluted to appropriate concentrations and analyzed according to the manufacturer's instructions.
[0638] LNP production method
[0639] As previously described (Corbett, KS et al., (2020). SARS-CoV-2mRNA vaccine design enabled by prototype pathogen preparedness. Nature 586, 567-571; Polack, FP et al., (2020). Safety and Efficacy of the BNT162b2mRNA Covid-19 Vaccine. N Engl J Med 383, 2603-2615), circular RNA is encapsulated in lipid nanoparticles via NanoAssemblr Ignite system. In brief, the circRNA aqueous solution of pH 4.0 is rapidly mixed with a lipid mixture dissolved in ethanol, and the lipid mixture contains different ionizable cationic lipids, distearoylphosphatidylcholine (DSPC), DMG-PEG2000 and cholesterol. The ratio of the lipid mixture was: for Formulation 1, MC3:DSPC:cholesterol:PEG-2000 = 50:10:38.5:1.5; for Formulation 2, SM-102:DSPC:cholesterol:PEG-2000 = 50:10:38.5:1.5; for Formulation 3, ALC-0315:DSPC:cholesterol:ALC-0159 = 46.3:9.4:42.7:1.6. The resulting LNP mixture was then dialyzed against PBS and stored at a concentration of 0.5 μg / μl at -80°C for further use.
[0640] Administration of LNP-circRNA to mice
[0641] 8-week-old female BALB / C mice were purchased from Shanghai Model Organisms Center. 20 μg of circRNA-LNP in PBS was administered intramuscularly to mice using a 3 / 10 insulin syringe (BD biosciences). Serum was collected 24 hours after LNP administration, and 50 μl of serum was used for in vitro luciferase activity assay. Bioluminescence imaging was performed using an IVIS spectrometer (Roper Scientific). 24 hours after Gluc-LNP injection, 2 mg / kg of coelenterazine (Coelenterazine) (MedChemExpress, MCE) was administered intraperitoneally to mice. The mice were then anesthetized with 2.5% isoflurane (RWD Life Science Co.) in the chamber after receiving the substrate and placed on the imaging platform while maintaining 2% isoflurane via a nose cone. Mice were imaged 5 minutes after substrate injection with an exposure time of 30 seconds to ensure that signals were effectively and fully obtained.
[0642] For immunogenicity studies, 8-week-old female BALB / c mice (Shanghai Model Organisms Center, lnc) were used. For primary and booster injections, 10 μg of circRNA-RBD-LNP was diluted in 50 μl 1XPBS and administered intramuscularly to the same hind legs of the mice. The mice in the control group received PBS and empty LNP. Two weeks after the primary and booster injections, blood samples were collected. The spleens of mice were also harvested at the endpoint (two weeks after the boost) for immunostaining and flow cytometry.
[0643] CircRNA molecules carrying different IRESs were transfected into C2C12 (mouse myoblast cell line), A673 (human rhabdomyosarcoma cell line), and L6 (rat myoblast cell line) cells using Lipofectamine MessengerMAX (Invitrogen) and seeded at a density of 10,000 cells per well in 96-well plates. Each well received 100 ng of circRNA. After incubation at 37°C for 24 hours, the cells were assayed for HGF expression using ELISA (Elabscience, E-EL-H0084c) according to the manufacturer's instructions, and cell viability was assessed using an ATP assay. For the ATP assay, the culture medium was aspirated and 100 μL of a 1:1 mixture of complete culture medium and detection reagent was added to each well. The plate was mixed thoroughly, incubated at room temperature for 10 minutes, and then luminescence was measured using a BioTeksynergy H1 microplate reader. 8.13. Example 13 Preparation and Verification of Expression Constructs Comprising Group II Introns BR23 and CL
[0644] In this example, fragments of the group II intron and IRES sequences in BR23 and CL were constructed, and different protein coding fragments were amplified by PCR. The 3' intron fragment sequence used in BR23 is shown in SEQ ID NO: 145, the 3' intron fragment sequence used in CL is shown in SEQ ID NO: 146, the E2 sequence used in BR23 and CL is shown in SEQ ID NO: 58, and the 5' arm sequence (linker sequence) of the target sequence used in BR23 and CL is shown in SEQ ID NO: 92. The target sequence used in BR23 and CL is shown in SEQ ID NO: 138, which consists of a non-coding sequence and an RBD (SEQ ID NO: 141). Those skilled in the art will understand that the non-coding sequence and RBD are merely exemplary, and those skilled in the art will fully understand that they can be replaced with any target sequence as needed. The 3' arm sequence (linker sequence) of the target sequence used in BR23 and CL is shown in SEQ ID NO: 100, and the 5' intron fragment sequence used in BR23 is shown in SEQ ID NO: NO:147, and the sequence of the 5' intron fragment used in CL is shown in SEQ ID NO:148. The sequence of the group II intron and IRES sequence in BR23, from 3' to 5', is as follows: 3' intron fragment, E2 fragment, 5' arm sequence, target sequence, 3' arm sequence, E1 fragment, and 5' intron fragment. The sequence of the group II intron and IRES sequence in CL, from 3' to 5', is as follows: 3' intron fragment, E2 fragment, 5' arm sequence, target sequence, 3' arm sequence, E1 fragment, and 5' intron fragment. The full-length RBD protein was detected in the BR23 and CL circRNA samples by Western blotting as described in this application. The protein band was visualized and confirmed to be approximately 150 kDa, consistent with the expected molecular weight. β-actin was used as a loading control to ensure equal protein loading between samples. The results are shown in Figure 2. Figure 43 shown.
Claims
1. A polynucleotide construct having self-splicing activity, comprising the following operably linked elements from 5' to 3': (a) 3' intron fragment; (b) exon fragment 2 (E2); (c) target sequence; (d) exon fragment 1 (E1); (e) 5' intron fragment, in: The 5' intron fragment and the 3' intron fragment are each a fragment of a group II intron, wherein the 5' intron fragment is located on the 5' side of the 3' intron fragment in the group II intron, The E1 is a fragment of the 5' exon of the group II intron, and its length is ≥ 0 nucleotides. The E2 is a fragment of the 3' exon of the group II intron, and its length is ≥ 0 nucleotides. The target sequence is empty, or is a protein coding sequence, a non-coding sequence, or a combination of the two; And wherein the group II intron is CL.
2. A polynucleotide construct having self-splicing activity, comprising the following operably linked elements from 5' to 3': (a) 3' intron fragment; (b) exon fragment 2 (E2); (c) linker sequence; (d) target sequence; (e) linker sequence; (f) exon fragment 1 (E1); (g) 5' intron fragment, in: The 5' intron fragment and the 3' intron fragment are each a fragment of a group II intron, wherein the 5' intron fragment is located on the 5' side of the 3' intron fragment in the group II intron, The E1 is a fragment of the 5' exon of the group II intron, and its length is ≥ 0 nucleotides. The E2 is a fragment of the 3' exon of the group II intron, and its length is ≥ 0 nucleotides. The target sequence is empty, or is a protein coding sequence, a non-coding sequence, or a combination of the two; And wherein the group II intron is CL.
3. A polynucleotide construct having self-splicing activity, comprising the following operably linked elements from 5' to 3': (a) 5' homology arm; (b) 3' intron fragment; (c) exon fragment 2 (E2); (d) target sequence; (e) exon fragment 1 (E1); (f) 5' intron fragment; (g) 3' homology arm, in: The 5' intron fragment and the 3' intron fragment are each a fragment of a group II intron, wherein the 5' intron fragment is located on the 5' side of the 3' intron fragment in the group II intron, The E1 is a fragment of the 5' exon of the group II intron, and its length is ≥ 0 nucleotides. The E2 is a fragment of the 3' exon of the group II intron, and its length is ≥ 0 nucleotides. The target sequence is empty, or is a protein coding sequence, a non-coding sequence, or a combination of the two; And wherein the group II intron is CL.
4. A polynucleotide construct having self-splicing activity, comprising the following operably linked elements from 5' to 3': (a) 5' homology arm; (b) 3' intron fragment; (c) exon fragment 2 (E2); (d) linker sequence; (e) target sequence; (f) linker sequence; (g) exon fragment 1 (E1); (h) 5' intron fragment; (i) 3' homology arm, in: The 5' intron fragment and the 3' intron fragment are each a fragment of a group II intron divided into two fragments, the 5' intron fragment is located on the 5' side of the 3' intron fragment in the group II intron, The E1 is a fragment of the 5' exon of the group II intron, and its length is ≥ 0 nucleotides. The E2 is a fragment of the 3' exon of the group II intron, and its length is ≥ 0 nucleotides. The target sequence is empty, or is a protein coding sequence, a non-coding sequence, or a combination of the two; And wherein the group II intron is CL.
5. The polynucleotide construct of any one of claims 1 to 4, wherein the polynucleotide construct has self-splicing activity in vitro.
6. The polynucleotide construct of any one of claims 1-5, wherein the E1 and / or E2 has a length of 0-20 nucleotides, preferably a length of 0-10 nucleotides, such as a length of 0, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10 nucleotides. 7 . The polynucleotide construct according to claim 1 , wherein the 5′ intron fragment and the 3′ intron fragment are obtained by splitting a group II intron from an unpaired region into two fragments.
8. The polynucleotide construct of any one of claims 1 to 6, wherein the 5' intron fragment and the 3' intron fragment are obtained by splitting a group II intron from the loop region of the domain 1 stem-loop structure.
9. The polynucleotide construct of any one of claims 1 to 6, wherein the 5' intron fragment and the 3' intron fragment are obtained by splitting a group II intron from the loop region of the domain 2 stem-loop structure.
10. The polynucleotide construct of any one of claims 1 to 6, wherein the 5' intron fragment and the 3' intron fragment are obtained by splitting a group II intron from the loop region of the domain 3 stem-loop structure.
Citation Information
Patent Citations
Apparatus for obtaining combustible gas.
US1111995A
Method of making an inflatable balloon catheter
US3832253A
Drug-delivery system
US3854480A
Method of encapsulating biologically active materials in lipid vesicles
US4235871A
Biodegradable, implantable drug delivery depots, and method for preparing and using the same
US4450150A