Synthetic circular RNA compositions and methods of using the same
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- SHANGHAI CIRCODE BIOMED CO LTD
- Filing Date
- 2023-05-29
- Publication Date
- 2026-06-04
AI Technical Summary
Existing gene therapy and vaccination methods using DNA face risks of integration into the genome and generation of anti-DNA antibodies, and natural IRES elements are lengthy, complex, and prone to immune rejection, lacking systematic control over protein expression.
Development of synthetic IRES elements and methods for predicting and controlling protein expression using engineered IRES-like sequences in circular RNAs, optimized for cap-independent translation and immune compatibility.
Enhances protein expression control and stability, minimizing immune rejection and genomic integration risks, suitable for therapeutic applications.
Smart Images

Figure 2023231959000001 
Figure 2023231959000002
Abstract
Description
Technical Field
[0001] 1. Related Applications This application claims the priority and benefit of the specification of International Application No. PCT / CN2022 / 095949, filed on May 30, 2022, the content of which is incorporated herein by reference in its entirety.
[0002] 2. Incorporation by Reference of Sequence Listing The content of the electronic sequence listing (Family 03_sql_0529.xml, size: 12,831,515 bytes, and creation date: May 29, 2023) is incorporated herein by reference in its entirety.
[0003] 3. Technical Field The present disclosure relates to compositions of matters, methods, processes, kits, and devices for the selection, design, preparation, manufacture, formulation, and / or use of polynucleotides having an internal ribosome entry site (IRES) sequence, an IRES-like sequence, or a combination thereof. The present disclosure also relates to compositions of matters, methods, processes, kits, and devices for the selection, design, preparation, manufacture, formulation, and / or use of circular polynucleotides (e.g., circular RNAs) comprising the IRES sequence, the IRES-like sequence, or a combination thereof. Further, the present disclosure relates to methods for improving the expression, functional stability, immunogenicity, manufacture, and / or half-life of a therapeutic product encoded by a circular RNA.
Background Art
[0004] 4. Background Art Gene therapy and gene vaccination provide highly specific and individualized treatment options for various diseases, such as genetic diseases, autoimmune diseases, cancer, and inflammatory diseases. In the field of gene therapy and gene vaccination, not only DNA but also RNA may be used as nucleic acid molecules for administration. DNA therapy is stable and easy to manipulate, but there are risks of undesirable results, such as integration into the genome and generation of anti-DNA antibodies. Furthermore, the expression of the encoded protein may be limited by dependence on the presence of specific transcription factors that regulate DNA transcription. In the absence of such factors, DNA transcription is hindered and the level of the translated protein decreases.
[0005] By using RNA instead of DNA gene therapy, the risks of undesirable results such as integration into the genome and generation of anti-DNA antibodies can be minimized or avoided. In particular, circular RNA is useful for the design and production of stable forms of RNA. Circular RNA is useful for in vivo applications, especially in the areas of RNA-based gene expression control, protein production, and therapies including protein replacement therapy and vaccination.
[0006] The translation of circular RNA is promoted by cap-independent translation. Therefore, the appropriate design and selection of cap-independent translation initiation elements are important for the control of protein expression. Internal ribosome entry site (IRES) elements are useful for cap-independent gene expression in eukaryotic cells. However, naturally occurring IRES elements may have 1) a long length of nucleic acid residues, 2) contain a complex secondary structure, and / or 3) be prone to host cell immune rejection reactions, none of which are favorable for in vivo applications such as protein replacement therapy.
[0007] Numerous natural IRES sequences have been identified that have been shown to promote cap-independent translation, but the identification and development of novel IRES elements have been secondary and serendipitous. Natural IRES elements do not allow for fine-tuning control of protein expression, and little systematic methodology has been shown for predicting functional novel IRES elements. No systematic approach has been shown for the identification and development of synthetic IRES elements that enable efficient cap-independent protein expression.
[0008] Accordingly, there is a need in the art for approaches and methodologies for developing synthetic IRES elements, as well as methodologies for predicting and controlling the efficiency of protein expression from synthetic IRES elements. There is also a need in the art to provide polynucleotides (e.g., circular RNAs) having synthetic IRES elements suitable for use as pharmaceuticals or vaccines, such as for application to gene therapy and / or gene vaccination. The present disclosure addresses these unmet needs.
SUMMARY OF THE INVENTION
[0009] 5. Summary of the Invention As described herein, a composition of matter (e.g., a synthetic internal ribosome entry site (IRES) sequence, an IRES-like sequence, or a combination thereof, and a circular RNA) and methods for the selection, design, preparation, manufacture, formulation, and / or use of a polynucleotide having an IRES sequence, an IRES-like sequence, or a combination thereof, and a circular RNA are described.
[0010] As used herein, Formula I: 5’-(A1) 0-1 -(L) n -TI-(L) n -Z1-(L) n -(B1) 0-1 -3’(I) (wherein, TI is an engineered translation initiation element that includes an IRES-like polynucleotide sequence, and the IRES-like polynucleotide sequence includes a nucleic acid sequence that is about 90%, 95%, 97%, 98%, 99%, or 100% identical or at least about 90%, 95%, 97%, 98%, 99%, or 100% identical to SEQ ID NOs: 1025 to 14161 or SEQ ID NOs: 14412 to 15341. Z1 is an expression sequence encoding a therapeutic product. Each L is independently a linker sequence. A1 and B1 are each independently a sequence capable of circularizing an RNA polynucleotide or A1 and B1 each independently include a nucleotide derivative capable of joining the 5'-end and the 3'-end via a 3'-to-5' phosphodiester bond to circularize an RNA polynucleotide. n is an integer selected from 0 to 2) and provides an RNA polynucleotide comprising the construct of.
[0011] As used herein, Formula II: 5'-(A1) 0-1 -(L) n -Z1 B -(L) n -TI-(L) n -Z1 A- (L) n -(B1) 0-1 -3'(II) (wherein, TI is an engineered translation initiation element that includes an IRES-like polynucleotide sequence, and the IRES-like polynucleotide sequence includes a nucleic acid sequence that is about 90%, 95%, 97%, 98%, 99%, or 100% identical or at least about 90%, 95%, 97%, 98%, 99%, or 100% identical to SEQ ID NOs: 1025 to 14161 or SEQ ID NOs: 14412 to 15341. Z1 A is the first part of an expression sequence encoding a therapeutic product. Z1B is the second part of the expression sequence encoding the therapeutic agent, each L is independently a linker sequence, A1 and B1 are each independently a sequence capable of circularizing an RNA polynucleotide, or A1 and B1 each independently contain a nucleotide derivative capable of joining the 5'-end and the 3'-end via a 3'-to-5' phosphodiester bond to circularize the RNA polynucleotide, n is an integer selected from 0 to 2) An RNA polynucleotide comprising the construct of
[0012] Also provided herein is Formula III: 5'-(3'intron fragment)-(E2)-(L) n -TI-(L) n -Z1-(L) n -(E1)-(5'intron fragment)-3'(III) (wherein, TI is a modified translation initiation element comprising an IRES-like polynucleotide sequence, and the IRES-like polynucleotide sequence comprises a nucleic acid sequence that is about 90%, 95%, 97%, 98%, 99%, or 100% identical or at least about 90%, 95%, 97%, 98%, 99%, or 100% identical to SEQ ID NOs: 1025 to 14161 or SEQ ID NOs: 14412 to 15341, Z1 is an expression sequence encoding a therapeutic agent, each L is independently a linker sequence, the 5'intron fragment and the 3'intron fragment are each a fragment of a Group II intron, and the 5'intron fragment is located 5' to the 3'intron fragment within the Group II intron, E1 is a 5' flanking exon fragment of a Group II intron having a length of ≧0 nucleotides, E2 is a 3' flanking exon fragment of a Group II intron having a length of ≧0 nucleotides, and n is an integer selected from 0 to 2) To provide an RNA polynucleotide comprising the construct.
[0013] Also, herein, Formula IV: 5’-(3’ intron fragment)-(E2)-(L) n -Z1 B -(L) n -TI-(L) n -Z1 A- -(L) n -(E1)-(5’ intron fragment)-3’ (IV) (wherein, TI is a modified translation initiation element comprising an IRES-like polynucleotide sequence, and the IRES-like polynucleotide sequence comprises a nucleic acid sequence that is about 90%, 95%, 97%, 98%, 99%, or 100% identical or at least about 90%, 95%, 97%, 98%, 99%, or 100% identical to SEQ ID NOs: 1025 to 14161 or SEQ ID NOs: 14412 to 15341, Z1 A is the first part of the expression sequence encoding a therapeutic agent, Z1 B is the second part of the expression sequence encoding a therapeutic agent, each L is independently a linker sequence, the 5’ intron fragment and the 3’ intron fragment are each a fragment of a Group II intron, and the 5’ intron fragment is located 5’ to the 3’ intron fragment within the Group II intron, the E1 is a 5’ adjacent exon fragment of a Group II intron having a length of ≧0 nucleotides, the E2 is a 3’ adjacent exon fragment of a Group II intron having a length of ≧0 nucleotides, and n is an integer selected from 0 to 2) To provide an RNA polynucleotide comprising the construct.
[0014] Also, herein, Formula V: 5’-TI-(L) n -Z1-3’ (V) (wherein, TI is a modified translation initiation element containing an IRES-like polynucleotide sequence, and the IRES-like polynucleotide sequence contains a nucleic acid sequence that is about 90%, 95%, 97%, 98%, 99%, or 100% identical or at least about 90%, 95%, 97%, 98%, 99%, or 100% identical to SEQ ID NOs: 1025 to 14161 or SEQ ID NOs: 14412 to 15341. Z1 is an expression sequence encoding a therapeutic agent. Each L is independently a linker sequence, and n is an integer selected from 0 to 2) provided is an RNA polynucleotide comprising a construct of
[0015] Also provided herein is an RNA polynucleotide comprising a modified translation initiation element (TI) containing an IRES-like polynucleotide sequence, wherein the IRES-like polynucleotide sequence contains a nucleic acid sequence that is about 90%, 95%, 97%, 98%, 99%, or 100% identical or at least about 90%, 95%, 97%, 98%, 99%, or 100% identical to SEQ ID NOs: 1025 to 14161 or SEQ ID NOs: 14412 to 15341.
[0016] Also provided herein is an RNA polynucleotide comprising a construct of Formula I, Formula II, Formula III, Formula IV, or Formula V: 5’-(A1) 0-1 -(L) n -TI-(L) n -Z1-(L) n -(B1) 0-1 -3’(I), 5’-(A1) 0-1 -(L) n -Z1 B -(L) n -TI-(L) n -Z1 A- (L) n -(B1) 0-1 -3’(II), 5’-(3’ intron fragment)-(E2)-(L) n -TI-(L) n -Z1-(L) n-(E1)-(5’ intron fragment)-3’(III), 5’-(3’ intron fragment)-(E2)-(L) n -Z1 B -(L) n -TI-(L) n -Z1 A- -(L) n -(E1)-(5’ intron fragment)-3’(IV), or 5’-TI-(L) n -Z1-3’(V) (wherein, TI is a modified translation initiation element containing an IRES-like polynucleotide sequence, and the IRES-like polynucleotide sequence contains a nucleic acid sequence that is about 90%, 95%, 97%, 98%, 99%, or 100% identical or at least about 90%, 95%, 97%, 98%, 99%, or 100% identical to SEQ ID NO: 1025-14161 or SEQ ID NO: 14412-15341, Z1 is an expression sequence encoding a therapeutic agent, Z1 A is the first part of the expression sequence encoding a therapeutic agent, Z1 B is the second part of the expression sequence encoding a therapeutic agent, each L is independently a linker sequence, A1 and B1 are each independently a sequence capable of circularizing an RNA polynucleotide, or A1 and B1 each independently contain a nucleotide derivative capable of linking the 5’ end and the 3’ end via a 3’ to 5’ phosphodiester bond to circularize an RNA polynucleotide, the 5’ intron fragment and the 3’ intron fragment are each a fragment of a group II intron, and the 5’ intron fragment is located on the 5’ side of the 3’ intron fragment within the group II intron, the E1 is a 5’ adjacent exon fragment of a group II intron having a length of ≧0 nucleotides, the E2 is a 3’ adjacent exon fragment of a group II intron having a length of ≧0 nucleotides, and n is an integer selected from 0 to 2) Provided is an RNA polynucleotide comprising the construct of .
[0017] Disclosed herein is a method for generating an internal ribosome entry site (IRES)-like polynucleotide sequence, comprising: (a) generating a polynucleotide query sequence consisting of X nucleic acid residues, where X is an integer of 3 or more; (b) generating X - Y + 1 overlapping polynucleotide fragment sequences within the polynucleotide query sequence, where each polynucleotide fragment sequence consists of Y nucleic acid residues, the first position of each polynucleotide fragment sequence is n, the last position of the same polynucleotide fragment sequence is Y + n - 1, and n represents each positive integer between 1 and X - Y + 1; (c) determining an enrichment score for each polynucleotide fragment sequence of (b); (d) determining a numerical score for the polynucleotide query sequence by summing the enrichment scores of each polynucleotide fragment sequence of (c); and (e) identifying the polynucleotide query sequence as a modified IRES-like polynucleotide sequence according to a reference value. Also provided is a method comprising the above steps.
[0018] Disclosed herein is a method for determining an enrichment score for a polynucleotide fragment sequence, comprising: i) generating an expression plasmid library, wherein each expression plasmid in the library contains a different polynucleotide fragment sequence and a reporter gene; ii) contacting a population of cells with the expression plasmid library; iii) quantifying the expression level of the reporter gene corresponding to each expression plasmid in the expression plasmid library. iv) Based on the protein expression levels of iii), dividing the entire cell population into a first population and a second population, and v) Determining an enrichment score for the polynucleotide fragment sequence using the following system of equations: [Number] [Number] (where f1 is the frequency of the polynucleotide fragment sequence in the first population, f2 is the frequency of the same polynucleotide fragment sequence in the second population, N1 is the size of the first population, and N2 is the size of the second population) provided by determining.
[0019] Provided herein is an RNA polynucleotide comprising a modified translation initiation element (TI), wherein the TI comprises an internal ribosome entry site (IRES)-like polynucleotide sequence generated by the methods of the present disclosure.
[0020] Provided herein are polypeptides expressed by the RNA polynucleotides of the present disclosure; DNA vectors encoding the RNA polynucleotides of the present disclosure or suitable for synthesizing the RNA polynucleotides of the invention; cells comprising the RNA polynucleotides, polypeptides, or DNA vectors of the present disclosure; and compositions comprising the RNA polynucleotides, polypeptides, DNA vectors, or cells of the present disclosure and a pharmaceutically acceptable carrier; and methods for their manufacture.
[0021] Provided herein is a method of regulating protein expression in a subject in need thereof, comprising administering a therapeutically effective amount of the RNA polynucleotides, polypeptides, DNA vectors, or cells of the present disclosure to the subject.
[0022] A method for treating or preventing a disease or disorder in a subject in need thereof, the method comprising administering to the subject a therapeutically effective amount of an RNA polynucleotide, polypeptide, DNA vector, or cell of the present disclosure.
[0023] The present disclosure provides the use of an RNA polynucleotide, polypeptide, DNA vector, or cell of the present disclosure in the manufacture of a medicament for regulating the expression of a protein or for treating or preventing a disease or disorder in a subject in need thereof.
[0024] The present disclosure provides an RNA polynucleotide, polypeptide, DNA vector, or cell of the present disclosure for regulating the expression of a protein or for treating or preventing a disease or disorder in a subject in need thereof.
[0025] In some embodiments, the nucleic acids (e.g., polynucleotides) and nucleic acid sequences disclosed herein may be codon-optimized (e.g., for expression of a therapeutic agent) via any codon-optimization technique known in the art (see, e.g., Quax el al., 2015, Mol Cell 59: 149-161).
[0026] 6. Exemplary Embodiments Embodiment 1
[0027] Formula I: 5’-(A1) 0-1 -(L) n -TI-(L) n -Z1-(L) n -(B1) 0-1 -3’(I) (wherein, TI is a modified translation initiation element containing an IRES-like polynucleotide sequence, and the IRES-like polynucleotide sequence contains a nucleic acid sequence that is about 90%, 95%, 97%, 98%, 99%, or 100% identical or at least about 90%, 95%, 97%, 98%, 99%, or 100% identical to SEQ ID NOs: 1025 to 14161 or SEQ ID NOs: 14412 to 15341. Z1 is an expression sequence encoding a therapeutic agent. Each L is independently a linker sequence. A1 and B1 are each independently a sequence capable of circularizing an RNA polynucleotide, or A1 and B1 each independently contain a nucleotide derivative capable of joining the 5'-end and the 3'-end via a 3' to 5' phosphodiester bond to circularize the RNA polynucleotide. n is an integer selected from 0 to 2) An RNA polynucleotide containing the construct of is provided.
[0028] Embodiment 1-1
[0029] Formula I-1: 5'-(A1) 0-1 -(L) n -TI-(L) n -Z1-(L) n -(B1) 0-1 -3'(I-1) (Wherein, TI is a modified translation initiation element containing an IRES-like polynucleotide sequence, and the IRES-like polynucleotide sequence contains a nucleic acid sequence that is about 90%, 95%, 97%, 98%, 99%, or 100% identical or at least about 90%, 95%, 97%, 98%, 99%, or 100% identical to SEQ ID NOs: 1025 to 14161 or SEQ ID NOs: 14412 to 15341. Z1 is an expression sequence encoding a therapeutic agent. Each L is independently a linker sequence. A1 and B1 are each independently a sequence capable of circularizing an RNA polynucleotide, and n is an integer selected from 0 to 2) An RNA polynucleotide comprising the construct of is provided.
[0030] Embodiment 1-2
[0031] Formula I-2: 5’-(A1) 0-1 -(L) n -TI-(L) n -Z1-(L) n -(B1) 0-1 -3’(I-2) (wherein TI is a modified translation initiation element comprising an IRES-like polynucleotide sequence, and the IRES-like polynucleotide sequence comprises a nucleic acid sequence that is about 90%, 95%, 97%, 98%, 99%, or 100% identical or at least about 90%, 95%, 97%, 98%, 99%, or 100% identical to SEQ ID NO: 1025 to 14161 or SEQ ID NO: 14412 to 15341, Z1 is an expression sequence encoding a therapeutic agent, Each L is independently a linker sequence, A1 and B1 each independently comprise a nucleotide derivative capable of joining the 5'-end and the 3'-end via a 3' to 5' phosphodiester bond to circularize the RNA polynucleotide, n is an integer selected from 0 to 2) An RNA polynucleotide comprising the construct of is provided.
[0032] Embodiment 2
[0033] Formula II: 5’-(A1) 0-1 -(L) n -Z1 B -(L) n -TI-(L) n -Z1 A- (L) n -(B1) 0-1 -3’(II) (wherein TI is a modified translation initiation element containing an IRES-like polynucleotide sequence, and the IRES-like polynucleotide sequence contains a nucleic acid sequence that is about 90%, 95%, 97%, 98%, 99%, or 100% identical or at least about 90%, 95%, 97%, 98%, 99%, or 100% identical to SEQ ID NOs: 1025 to 14161 or SEQ ID NOs: 14412 to 15341. Z1 A is the first part of the expression sequence encoding the therapeutic agent. Z1 B is the second part of the expression sequence encoding the therapeutic agent. Each L is independently a linker sequence. A1 and B1 are each independently a sequence capable of circularizing an RNA polynucleotide, or A1 and B1 each independently contain a nucleotide derivative capable of joining the 5'-end and the 3'-end via a 3'-to-5' phosphodiester bond to circularize the RNA polynucleotide. n is an integer selected from 0 to 2) An RNA polynucleotide containing the construct of is provided.
[0034] Embodiment 2-1
[0035] Formula II-1: 5'-(A1) 0-1 -(L) n -Z1 B -(L) n -TI-(L) n -Z1 A- (L) n -(B1) 0-1 -3'(II-1) (wherein, TI is a modified translation initiation element containing an IRES-like polynucleotide sequence, and the IRES-like polynucleotide sequence contains a nucleic acid sequence that is about 90%, 95%, 97%, 98%, 99%, or 100% identical or at least about 90%, 95%, 97%, 98%, 99%, or 100% identical to SEQ ID NOs: 1025 to 14161 or SEQ ID NOs: 14412 to 15341. Z1 Ais the first part of the expression sequence encoding the therapeutic agent, Z1 B is the second part of the expression sequence encoding the therapeutic agent, each L is independently a linker sequence, A1 and B1 are each independently a sequence capable of circularizing the RNA polynucleotide, and n is an integer selected from 0 to 2). An RNA polynucleotide comprising the construct of
[0036] Embodiment 2-2
[0037] Formula II-2: 5’-(A1) 0-1 -(L) n -Z1 B -(L) n -TI-(L) n -Z1 A- (L) n -(B1) 0-1 -3’(II-2) (wherein, TI is a modified translation initiation element comprising an IRES-like polynucleotide sequence, and the IRES-like polynucleotide sequence comprises a nucleic acid sequence that is about 90%, 95%, 97%, 98%, 99%, or 100% identical or at least about 90%, 95%, 97%, 98%, 99%, or 100% identical to SEQ ID NOs: 1025 to 14161 or SEQ ID NOs: 14412 to 15341, Z1 A is the first part of the expression sequence encoding the therapeutic agent, Z1 B is the second part of the expression sequence encoding the therapeutic agent, each L is independently a linker sequence, A1 and B1 each independently comprise a nucleotide derivative capable of joining the 5’ end and the 3’ end via a 3’ to 5’ phosphodiester bond to circularize the RNA polynucleotide, n is an integer selected from 0 to 2). An RNA polynucleotide comprising the construct of
[0038] Embodiment 3
[0039] Formula III: 5’-(3’ intron fragment)-(E2)-(L) n -TI-(L) n -Z1-(L) n -(E1)-(5’ intron fragment)-3’(III) (In the formula, TI is a modified translation initiation element containing an IRES-like polynucleotide sequence, and the IRES-like polynucleotide sequence contains a nucleic acid sequence that is about 90%, 95%, 97%, 98%, 99%, or 100% identical or at least about 90%, 95%, 97%, 98%, 99%, or 100% identical to SEQ ID NOs: 1025 to 14161 or SEQ ID NOs: 14412 to 15341, Z1 is an expression sequence encoding a therapeutic agent, each L is independently a linker sequence, the 5’ intron fragment and the 3’ intron fragment are each a fragment of a Group II intron, the 5’ intron fragment is located on the 5’ side of the 3’ intron fragment within the Group II intron, the E1 is a 5’ adjacent exon fragment of a Group II intron having a length of ≧0 nucleotides, the E2 is a 3’ adjacent exon fragment of a Group II intron having a length of ≧0 nucleotides, and n is an integer selected from 0 to 2) An RNA polynucleotide comprising the construct of is provided.
[0040] Embodiment 4
[0041] Formula IV: 5’-(3’ intron fragment)-(E2)-(L) n -Z1 B -(L) n -TI-(L) n -Z1 A- -(L) n -(E1)-(5’ intron fragment)-3’(IV) (wherein TI is a modified translation initiation element containing an IRES-like polynucleotide sequence, and the IRES-like polynucleotide sequence contains a nucleic acid sequence that is about 90%, 95%, 97%, 98%, 99%, or 100% identical or at least about 90%, 95%, 97%, 98%, 99%, or 100% identical to SEQ ID NOs: 1025 to 14161 or SEQ ID NOs: 14412 to 15341, Z1 A is the first part of the expression sequence encoding the therapeutic agent, Z1 B is the second part of the expression sequence encoding the therapeutic agent, each L is independently a linker sequence, the 5' intron fragment and the 3' intron fragment are each a fragment of a group II intron, and the 5' intron fragment is located on the 5' side of the 3' intron fragment within the group II intron, the E1 is a 5' adjacent exon fragment of a group II intron having a length of ≧0 nucleotides, the E2 is a 3' adjacent exon fragment of a group II intron having a length of ≧0 nucleotides, and n is an integer selected from 0 to 2) An RNA polynucleotide comprising the construct of
[0042] Embodiment 5
[0043] The RNA polynucleotide of Embodiment 3 or 4, further comprising a 5' homology arm at the 5' end of the 3' intron fragment.
[0044] Embodiment 6
[0045] The RNA polynucleotide of Embodiment 3 or 4, further comprising a 3' homology arm at the 3' end of the 5' intron fragment.
[0046] Embodiment 7
[0047] The RNA polynucleotide of embodiment 3 or 4, further comprising a 5' homology arm at the 5' end of the 3' intron fragment and further comprising a 3' homology arm at the 3' end of the 5' intron fragment.
[0048] Embodiment 8
[0049] The E1 and the E2 are each independently 0 to 20 nucleotides in length, preferably 0 to 10 nucleotides in length, for example 0, 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 nucleotides in length, the RNA polynucleotide of any one of embodiments 3 to 7.
[0050] Embodiment 9
[0051] The 5' intron fragment and the 3' intron fragment are obtained by splitting the group II intron into two fragments at the unpaired region, and the unpaired region is preferably selected from the linear region between two adjacent domains of the group II intron and the loop region of the stem-loop structure of domain 4 of the group II intron, the RNA polynucleotide of any one of embodiments 3 to 8.
[0052] Embodiment 10
[0053] Formula IV: 5'-TI-(L) n -Z1-3'(IV) (Wherein, TI is a modified translation initiation element comprising an IRES-like polynucleotide sequence, and the IRES-like polynucleotide sequence comprises a nucleic acid sequence that is about 90%, 95%, 97%, 98%, 99%, or 100% identical or at least about 90%, 95%, 97%, 98%, 99%, or 100% identical to SEQ ID NO: 1025 to 14161 or SEQ ID NO: 14412 to 15341, Z1 is an expression sequence encoding a therapeutic agent, Each L is independently a linker sequence, and n is an integer selected from 0 to 2) An RNA polynucleotide comprising the construct.
[0054] Embodiment 11
[0055] An RNA polynucleotide comprising a modified translation initiation element (TI) comprising an IRES-like polynucleotide sequence, wherein the IRES-like polynucleotide sequence is about 90%, 95%, 97%, 98%, 99%, or 100% identical or at least about 90%, 95%, 97%, 98%, 99%, or 100% identical to SEQ ID NO: 1025-14161 or SEQ ID NO: 14412-15341 and comprises a nucleic acid sequence.
[0056] Embodiment 12
[0057] The RNA polynucleotide according to any one of Embodiments 1 to 11, wherein the RNA polynucleotide is circularized by a ligation reaction.
[0058] Embodiment 13
[0059] The RNA polynucleotide according to Embodiment 12, wherein the RNA polynucleotide is circularized in the presence of T4 ligase.
[0060] Embodiment 14
[0061] The RNA polynucleotide according to Embodiment 12, wherein the RNA polynucleotide is circularized in the absence of T4 ligase.
[0062] Embodiment 15
[0063] The RNA polynucleotide according to any one of Embodiments 1 to 11, wherein the RNA polynucleotide is circularized by a splicing reaction.
[0064] Embodiment 16
[0065] The RNA polynucleotide according to Embodiment 15, wherein the RNA polynucleotide is circularized in the presence of a spliceosome.
[0066] Embodiment 17
[0067] The RNA polynucleotide of Embodiment 15, wherein the RNA polynucleotide is circularized in the absence of a spliceosome.
[0068] Embodiment 18
[0069] The RNA polynucleotide according to any one of Embodiments 1 to 11, wherein the RNA polynucleotide is circularized by a self-splicing reaction.
[0070] Embodiment 19
[0071] The RNA polynucleotide according to any one of Embodiments 1 to 18, wherein the length of the IRES-like polynucleotide sequence is 6 to 12 residues.
[0072] Embodiment 20
[0073] The RNA polynucleotide according to any one of Embodiments 1 to 19, wherein the TI further comprises a second IRES-like polynucleotide sequence.
[0074] Embodiment 20-1
[0075] The RNA polynucleotide of Embodiment 20, wherein the first IRES-like polynucleotide sequence and the second IRES-like polynucleotide sequence are the same.
[0076] Embodiment 20-2
[0077] The RNA polynucleotide of Embodiment 20, wherein the first IRES-like polynucleotide sequence and the second IRES-like polynucleotide sequence are different.
[0078] Embodiment 21
[0079] The second IRES-like polynucleotide sequence is an RNA polynucleotide of Embodiment 20, 20-1, or 20-2, comprising a nucleic acid sequence that is about 90%, 95%, 97%, 98%, 99% or 100% identical or at least about 90%, 95%, 97%, 98%, 99% or 100% identical to SEQ ID NO: 1025-14161 or SEQ ID NO: 14412-15341.
[0080] Embodiment 22
[0081] The TI of any one of Embodiments 1-21 further comprises a natural IRES sequence. The TI of any one of Embodiments 1-21 further comprises an IRES sequence isolated from or derived from a natural IRES sequence.
[0082] Embodiment 23
[0083] The IRES sequence of any one of Embodiments 1-22 is an RNA sequence capable of binding to a eukaryotic ribosome.
[0084] In some embodiments, the natural IRES sequence is an endogenous IRES sequence isolated from or derived from Homo sapiens. In some embodiments, the natural IRES sequence is an endogenous IRES sequence isolated from or derived from a human tissue or a human sample.
[0085] In some embodiments, the natural IRES sequence is isolated from or derived from IRES sequences of viruses, mammals, and Drosophila.
[0086] In some embodiments, the natural IRES sequence is isolated from or derived from picornavirus complementary DNA (cDNA), encephalomyocarditis virus (EMCV) cDNA, poliovirus cDNA, ABPV IGRpred, AEV, ALPV IGRpred, BQCV IGRpred, BVDV1 1-385, BVDV1 29-391, CrPV 5NCR, CrPV IGR, crTMV IREScp, crTMV_IRESmp75, crTMV_IRESmp228, crTMV IREScp, crTMV IREScp, CSFV, CVB3, DCV IGR, EMCV-R, EoPV_5NTR, ERAV_245-96l, ERBV_l62-920, EV7l_l-748, FeLV-Notch2, FMDV type C, GBV-A, GBV-B, GBV-C, gypsy_env, gypsyD5, gypsyD2, HAV HM175, HCV type la, HiPVJGRpred, HIV-1, HoCVlJGRpred, HRV-2, IAPVJGRpred, idefix, KBV IGRpred, LINE-l_ORFl_-lOl_to_-l, LINE-l_ORFl_-302_to_-202, LINE-l_ORF2_-l38_to_-86, LINE- 1_ORF 1_-44_to_- 1, PSIV IGR, PV typel Mahoney, PV_type3_Leon, REV-A, RhPV 5NCR, RhPV IGR, SINV l IGRpred, SV40 661-830, TMEV, TMV_UI_IRESmp228, TRV 5NTR, TrV IGR, or TSV IGR sequence.
[0087] In some embodiments, the natural IRES sequence of Drosophila is the Antennapedia gene of Drosophila melanogaster.
[0088] In some embodiments, the native IRES sequence is isolated from or derived from the AML1 / RUNX1, Antp-D, Antp-DE, Antp-CDE, Apaf-1, Apaf-1, AQP4, AT1R var1, AT1R_var2, AT1R_var3, AT1R_var4, BAG1_p36delta236nt, BAG1_p36, BCL2, BiP_-222_-3, C-IAP1 285-1399, c-IAP1 13 13-1462, c-jun, c-myc, Cat-1_224, CCND1, DAP5, eIF4G, eIF4GI-ext, eIF4GII, eIF4GII-long, ELG1, ELH, FGF1A, FMR1, Gtx-133-141, Gtx-1-166, Gtx-1-120, Gtx-1-196, Hairless, HAP4, HIF1a, hSNM1, Hsp101, hsp70, hsp70, Hsp90, IGF2_leader2, Kv1.4_1.2, L-myc, LamB1 -335 -1, LEF1, MNT 75-267, MNT 36-160, MTG8a, MYB, MYT2 997-1152, n-MYC, NDST1, NDST2, NDST3, NDST4L, NDST4S, NRF_-653_-17, NtHSF1, ODC1, p27kip1, p53_128-269, PDGF2 / c-sis, Pim-1, PITSLRE_p58, Rbm3, Reaper, Scamper, TFIID, TIF4631, Ubx_1-966, Ubx_373-961, UNR, Ure2, UtrA, VEGF-A-133-1, XIAP 5-464, XIAP 305-466, or YAP1 sequence.
[0089] In some embodiments, the native IRES sequence is isolated from or derived from the synthetic (GAAA)16, (PPT19)4, KMI1, KMI1, KMI2, KMI2, KMIX, XI, or X2 IRES sequence.
[0090] Embodiment 24
[0091] The L is an RNA polynucleotide of any one of Embodiments 1 to 23, including a nucleic acid sequence encoding a 5’UTR, 3’UTR, polyA sequence, polyA-C sequence, polyC sequence, polyU sequence, polyG sequence, ribosome binding site, aptamer, riboswitch, ribozyme, small RNA binding site, translation regulatory element (e.g., Kozak sequence), protein binding site (e.g., PTBP1 or HUR), unnatural nucleotide, or non-nucleotide chemical linker sequence.
[0092] Embodiment 25
[0093] The L is an RNA polynucleotide of any one of Embodiments 1 to 24, with a length of about 3 to about 100 nucleotide residues.
[0094] Embodiment 26
[0095] The L is an RNA polynucleotide of any one of Embodiments 1 to 25, including the nucleic acid sequence of RCC, where R is guanine or adenine.
[0096] Embodiment 27
[0097] The RNA polynucleotide is a single-stranded RNA polynucleotide of any one of Embodiments 1 to 26.
[0098] Embodiment 28
[0099] The RNA polynucleotide is a circular RNA polynucleotide of any one of Embodiments 1 to 27.
[0100] Embodiment 29
[0101] The RNA polynucleotide is a linear RNA polynucleotide of any one of Embodiments 1 to 27.
[0102] Embodiment 30
[0103] Any one of the RNA polynucleotides of Embodiments 1 to 29 that can be circularized in the absence of an enzyme.
[0104] Embodiment 31
[0105] Any one of the RNA polynucleotides of Embodiments 1 to 30, wherein the RNA polynucleotide is an isolated RNA polynucleotide or a synthetic RNA polynucleotide.
[0106] Embodiment 32
[0107] A polypeptide expressed by any one of the RNA polynucleotides of Embodiments 1 to 31.
[0108] Embodiment 33
[0109] A DNA vector encoding any one of the RNA polynucleotides of Embodiments 1 to 31.
[0110] Embodiment 34
[0111] A DNA vector suitable for synthesizing any one of the RNA polynucleotides of Embodiments 1 to 31.
[0112] Embodiment 35
[0113] A cell comprising any one of the RNA polynucleotides of Embodiments 1 to 31, the polypeptide of Embodiment 32, or the DNA vector of Embodiment 33 or 34.
[0114] Embodiment 36
[0115] A composition comprising any one of the RNA polynucleotides of Embodiments 1 to 31, the polypeptide of Embodiment 32, the DNA vector of Embodiment 33 or 34, or the cell of Embodiment 35 and a pharmaceutically acceptable carrier.
[0116] Embodiment 37
[0117] A method for producing a cell population, the method comprising contacting the cells of the population with any one of the RNA polynucleotides of Embodiments 1 to 31, the polypeptide of Embodiment 32, or the DNA vector of Embodiment 33 or 34.
[0118] Embodiment 38
[0119] A method for generating an internal ribosome entry site (IRES)-like polynucleotide sequence, comprising: (a) generating a polynucleotide query sequence consisting of X nucleic acid residues, where X is an integer of 3 or more; (b) generating X - Y + 1 overlapping polynucleotide fragment sequences within the polynucleotide query sequence, where each polynucleotide fragment sequence consists of Y nucleic acid residues, the first position of each polynucleotide fragment sequence is n, the last position of the same polynucleotide fragment sequence is Y + n - 1, and n represents each positive integer between 1 and X - Y + 1; (c) determining an enrichment score for each polynucleotide fragment sequence of (b); (d) determining a numerical score for the polynucleotide query sequence by summing the enrichment scores of each polynucleotide fragment sequence of (c); and (e) identifying the polynucleotide query sequence as a modified IRES-like polynucleotide sequence according to a reference value. A method comprising the steps above.
[0120] Embodiment 38-1
[0121] The method of Embodiment 38, wherein the reference value is calculated according to the following formula: reference value = (23.75×X) - 84.85 - Z, where Z represents any number between 1 and 10,000, and X represents the length of the IRES-like polynucleotide sequence.
[0122] Embodiment 38-2
[0123] The method of Embodiment 38-1, wherein Z represents any number between 50 and 5000.
[0124] Embodiment 38-3
[0125] The method of Embodiment 38-1, wherein Z represents any number between 175 and 2500.
[0126] Embodiment 38-4
[0127] The method of Embodiment 38-1, wherein Z represents any number between 150 and 200.
[0128] Embodiment 39
[0129] The method of Embodiment 38, wherein the reference value is characteristic of the absence of the expression of the therapeutic agent.
[0130] Embodiment 40
[0131] The method of Embodiment 38, wherein the reference value is the average score of all numerical scores of two or more IRES-like polynucleotide sequences, two or more natural IRES sequences, or a combination thereof.
[0132] Embodiment 41
[0133] The method of Embodiment 38, wherein the reference value is the average score of all numerical scores of two or more IRES-like polynucleotide sequences, two or more natural IRES sequences, or a combination thereof.
[0134] Embodiment 42
[0135] The method of embodiment 38, wherein the reference value is about 95%, about 90%, about 85%, about 80%, about 75%, about 70%, about 65%, about 60%, about 55%, about 50%, about 45%, about 40%, about 35%, about 30%, about 25%, about 20%, about 15%, about 10%, about 9%, about 8%, about 7%, about 6%, about 5%, about 4%, about 3%, about 2%, about 1%, about 0.9%, about 0.8%, about 0.7%, about 0.6%, about 0.5%, about 0.4%, about 0.3%, about 0.2%, or about 0.1% of the highest numerical score of two or more IRES-like polynucleotide sequences, two or more natural IRES sequences, or a combination thereof.
[0136] Embodiment 43
[0137] The method of embodiment 38, wherein the reference value is about 50%, about 45%, about 40%, about 35%, about 30%, about 25%, about 20%, about 15%, or about 10% of the highest numerical score of two or more IRES-like polynucleotide sequences, two or more natural IRES sequences, or a combination thereof.
[0138] Embodiment 44
[0139] The method of embodiment 38, wherein the reference value is at least about 95%, about 90%, about 85%, about 80%, about 75%, about 70%, about 65%, about 60%, about 55%, about 50%, about 45%, about 40%, about 35%, about 30%, about 25%, about 20%, about 15%, about 10%, about 9%, about 8%, about 7%, about 6%, about 5%, about 4%, about 3%, about 2%, about 1%, about 0.9%, about 0.8%, about 0.7%, about 0.6%, about 0.5%, about 0.4%, about 0.3%, about 0.2%, or about 0.1% higher than the lowest numerical score of two or more IRES-like polynucleotide sequences, two or more natural IRES sequences, or a combination thereof.
[0140] Embodiment 45
[0141] The method of embodiment 38, wherein the reference value is at least about 95%, about 90%, about 85%, about 80%, about 75%, about 70%, about 65%, about 60%, about 55%, or about 50% higher than the minimum numerical score of two or more IRES-like polynucleotide sequences, two or more natural IRES sequences, or a combination thereof.
[0142] Embodiment 46
[0143] The method of embodiment 38, wherein the reference value is 0 or more.
[0144] Embodiment 47
[0145] The method of any one of embodiments 38 to 46, wherein X is an integer selected from 3 to 300.
[0146] Embodiment 48
[0147] The method of any one of embodiments 38 to 47, wherein X is an integer selected from 3 to 100.
[0148] Embodiment 49
[0149] The method of any one of embodiments 38 to 48, wherein X is an integer selected from 5 to 100.
[0150] Embodiment 50
[0151] The method of any one of embodiments 38 to 49, wherein X is an integer selected from 6 to 100.
[0152] Embodiment 51
[0153] The method of any one of embodiments 38 to 50, wherein X is an integer of 5 or more.
[0154] Embodiment 52
[0155] The method of any one of embodiments 38 to 51, wherein X is an integer of 6 or more.
[0156] Embodiment 53
[0157] The polynucleotide query sequence is generated within a DNA vector suitable for synthesizing the polynucleotide query sequence, by any one of the methods of Embodiments 38 to 52.
[0158] Embodiment 54
[0159] Any one of the methods of Embodiments 38 to 53, wherein the polynucleotide query sequence is chemically synthesized.
[0160] Embodiment 55
[0161] Any one of the methods of Embodiments 38 to 54, wherein the length of the overlapping polynucleotide fragment sequence within the polynucleotide query sequence is 5, 6, 7, 8, 9, or 10 nucleic acid residues.
[0162] Embodiment 56
[0163] The enrichment score for the polynucleotide fragment sequence is i) generating an expression plasmid library, wherein each expression plasmid of the library contains a different polynucleotide fragment sequence and a reporter gene, ii) contacting a population of cells with the expression plasmid library, iii) quantifying the expression level of the reporter gene corresponding to each expression plasmid of the expression plasmid library, iv) dividing the entire population of cells into a first population and a second population based on the protein expression level of iii), and v) determining the enrichment score of the polynucleotide fragment sequence using the following system of equations:
Equation
Equation
[0164] Embodiment 57
[0165] The protein expression level of the first population is within the upper 0 to 10 percent of the protein expression level of the entire population, and the protein expression level of the second population is within the lower 10 to 90 percent of the protein expression level of the entire population. The method of Embodiment 56.
[0166] Embodiment 58
[0167] The protein expression level of the first population is within the upper 50.1 percent of the protein expression level of the entire population, and the protein expression level of the second population is within the lower 49.9 percent of the protein expression level of the entire population. The method of Embodiment 56 or 57.
[0168] Embodiment 59
[0169] The protein expression level of the first population is within the upper 10 percent of the protein expression level of the entire population, and the protein expression level of the second population is within the lower 90 percent of the protein expression level of the entire population. The method of Embodiment 56 or 57.
[0170] Embodiment 60
[0171] The method according to any one of Embodiments 38 to 59, wherein the length of the polynucleotide fragment sequence is 5 nucleic acid residues.
[0172] Embodiment 60-1
[0173] The method of Embodiment 60, wherein the polynucleotide fragment sequence is selected from SEQ ID NOs: 1 to 1024 having the enrichment score shown in Table 1.
[0174] Embodiment 61
[0175] The method of any one of Embodiments 38 to 59, wherein the length of the polynucleotide fragment sequence is 6 nucleic acid residues.
[0176] Embodiment 62
[0177] The method of any one of Embodiments 38 to 59, wherein the length of the polynucleotide fragment sequence is 7 nucleic acid residues.
[0178] Embodiment 63
[0179] The method of any one of Embodiments 38 to 59, wherein the length of the polynucleotide fragment sequence is 8 nucleic acid residues.
[0180] Embodiment 64
[0181] The method of any one of Embodiments 38 to 59, wherein the length of the polynucleotide fragment sequence is 9 nucleic acid residues.
[0182] Embodiment 65
[0183] The method of any one of Embodiments 38 to 59, wherein the length of the polynucleotide fragment sequence is 10 nucleic acid residues.
[0184] Embodiment 66
[0185] The method of any one of Embodiments 38 to 59, wherein the length of the polynucleotide fragment sequence is 11 nucleic acid residues.
[0186] Embodiment 67
[0187] The method of any one of Embodiments 38 to 59, wherein the length of the polynucleotide fragment sequence is 12 nucleic acid residues.
[0188] Embodiment 68
[0189] The method according to any one of Embodiments 38 to 59, wherein the length of the polynucleotide fragment sequence is 13 to 100 nucleic acid residues.
[0190] Embodiment 69
[0191] An RNA polynucleotide comprising a modified translation initiation element (TI), wherein the TI comprises an internal ribosome entry site (IRES)-like polynucleotide sequence generated by the method according to any one of Embodiments 38 to 68.
[0192] Embodiment 70
[0193] A polypeptide expressed by the RNA polynucleotide of Embodiment 69.
[0194] Embodiment 71
[0195] A DNA vector encoding the RNA polynucleotide of Embodiment 69.
[0196] Embodiment 72
[0197] A DNA vector suitable for synthesizing the RNA polynucleotide of Embodiment 69.
[0198] Embodiment 73
[0199] A cell comprising the RNA polynucleotide of Embodiment 69, the polypeptide of Embodiment 70, or the DNA vector of Embodiment 71 or 72.
[0200] Embodiment 74
[0201] A composition comprising the RNA polynucleotide of Embodiment 69, the polypeptide of Embodiment 70, the DNA vector of Embodiment 71 or 72, or the cell of Embodiment 73 and a pharmaceutically acceptable carrier.
[0202] Embodiment 75
[0203] A method of producing a population of cells, the method comprising contacting the cells of the population with the RNA polynucleotide of Embodiment 69, the polypeptide of Embodiment 70, or the DNA vector of Embodiment 71 or 72.
[0204] Embodiment 76
[0205] A method of regulating protein expression in a subject in need thereof, the method comprising administering to the subject a therapeutically effective amount of any one of the polynucleotides of Embodiments 1 to 31 and 69, the polypeptide of Embodiment 32 or 70, the DNA vector of Embodiments 33, 34, 71, or 72, or the cell of Embodiment 35 or 73.
[0206] Embodiment 77
[0207] A method of treating or preventing a disease or disorder in a subject in need thereof, the method comprising administering to the subject a therapeutically effective amount of any one of the polynucleotides of Embodiments 1 to 31 and 69, the polypeptide of Embodiment 32 or 69, the DNA vector of Embodiments 33, 34, 71, or 72, or the cell of Embodiment 35 or 73.
[0208] Embodiment 78
[0209] Use of any one of the polynucleotides of Embodiments 1 to 31 and 69, the polypeptide of Embodiment 32 or 70, the DNA vector of Embodiments 33, 34, 71, or 72, or the cell of Embodiment 35 or 73, in the manufacture of a medicament for regulating protein expression or treating or preventing a disease or disorder in a subject in need thereof.
[0210] Embodiment 79
[0211] For regulating the expression of a protein or treating or preventing a disease or disorder in a subject in need thereof, any one of the polynucleotides of Embodiments 1 to 31 and 69, the polypeptide of Embodiment 32 or 70, the DNA vector of Embodiments 33, 34, 71, or 72, or the cell of Claim 35 or 73.
Brief Description of the Drawings
[0212] 7. Brief Description of the Drawings
Figure 1
[0213]
Figure 2
[0214]
Figure 3
[0215]
Figure 4
[0216]
Figure 5
[0217]
Figure 6
[0218]
Figure 7
[0219]
Figure 8
[0220]
Figure 9
[0221]
Figure 10A
[0222]
Figure 10B
[0223]
Figure 11
[0224]
Figure 12A
[0225]
Figure 12B
[0226]
Figure 12C
Figure 12D
Figure 12E
[0227]
Figure 13A
[0228]
Figure 13B
[0229]
Figure 14A
[0230]
Figure 14B
[0231]
Figure 14C
[0232]
Figure 14D
[0233]
Figure 14E
[0234]
Figure 15
[0235]
Figure 16
BEST MODE FOR CARRYING OUT THE INVENTION
[0236] 8. BEST MODE FOR CARRYING OUT THE INVENTION The present disclosure overcomes problems associated with current technologies by providing a novel methodology for identifying and generating synthetic internal ribosome entry site (IRES)-like sequences. The present disclosure is based, at least in part, on the development of an algorithmic method and a high-throughput reporter assay that can systematically screen and quantify the IRES activity of RNA sequences that can promote circular RNA translation. The inventors systematically screened a pool of random polynucleotide sequences that promote protein expression and used a systematic ranking method to distinguish nucleic acid sequence motifs that enhance protein expression (i.e., IRES-like sequences) from nucleic acid sequence motifs that reduce protein expression (i.e., non-IRES sequences), or to distinguish nucleic acid sequence motifs that enhance protein expression more potently (i.e., IRES-like sequences) from other nucleic acid sequence motifs. This helps to identify novel sequences with IRES activity that may function more efficiently than natural IRES sequences. This also helps to identify universal and species-independent IRES elements.
[0237] The discovery of these nucleic acid sequence motifs enables the generation of synthetic IRES-like sequences for use in enhancing protein expression in eukaryotic cells. An increase in protein expression is desirable for improving gene expression and / or protein expression and production in therapeutic applications including, but not limited to, protein replacement therapy and vaccination. Furthermore, the use of a combination of IRES-like sequences and non-IRES sequences enables fine-tuning control and regulation of protein expression. The use of IRES-like sequences in combination with non-IRES sequences is also advantageous for precisely controlling peptide expression ratios in the production and manufacture of multimeric proteins and / or polycistronic cassettes for therapeutic applications such as antibody production.
[0238] Accordingly, the present disclosure also provides a polynucleotide comprising an IRES-like sequence and an expression sequence encoding a therapeutic agent of interest, which may be suitable for use as a pharmaceutical or a vaccine, such as for application in gene therapy and / or gene vaccination. In some embodiments, the present disclosure provides a polynucleotide comprising an IRES-like sequence and an expression sequence encoding a therapeutic agent of interest for use in improving the production and manufacture of a protein. In some embodiments, the polynucleotide is an RNA polynucleotide. In some embodiments, the RNA polynucleotide is a circular RNA polynucleotide. Overall, the present disclosure provides an improved polynucleotide comprising an IRES-like sequence, which overcomes the disadvantages of the prior art by providing a cost-effective, systematic, and direct approach for regulating protein expression.
[0239] This specification provides non-natural RNAs that have group II intron self-splicing activity and can form circular RNAs by self-splicing. Circular RNAs (circRNAs) are single-stranded RNAs that are head-to-tail linked. As is known in the art, circRNAs can be produced in vitro from precursor RNAs, which refer to linear RNA molecules from which circRNAs are directly generated, via chemical means or enzymatic activity, regardless of the method of circularization. For example, the 5' and 3' ends of a linear nucleic acid can be chemically linked by the catalysis of bromocyanide and morpholinyl derivatives, or can be ligated head-to-tail by the activity of a nucleic acid ligase. CircRNAs can also be produced by splicing. When a linear precursor undergoes splicing, a portion of the molecule is excised, resulting in a circRNA with a total number of nucleotides fewer than that of the precursor RNA.
[0240] circRNAs can also be produced by RNA splicing catalyzed by ribozymes. As used herein and understood in the art, the term "ribozyme" refers to an RNA molecule having enzymatic activity. Some ribozymes can catalyze self-splicing independently of the spliceosome, and these are referred to as "ribozymes with self-splicing activity", "self-splicing ribozymes", or "self-splicing introns". Naturally occurring self-splicing ribozymes can be divided into group I and group II introns. The splicing products of the two categories of ribozymes are similar, but the structures of the ribozymes themselves and the splicing mechanisms are very different. Group I introns have a nine-helix structure and require the external hydroxyl group of guanosine monophosphate (pG-OH) to induce the reaction during catalytic splicing, and are highly dependent on the sequences of the exons located at both ends of the group I intron. Group II introns rely on their own hydroxyl groups within the nucleotide sequence to trigger splicing. (See Figure 8). This splicing mechanism is similar to the splicing reaction mediated by the spliceosome and better simulates splicing in higher organisms. The terms "group I intron self-splicing activity" or "group I intron activity" refer to the self-splicing activity derived from group I introns, and the terms "group II intron self-splicing activity" or "group II intron activity" refer to the self-splicing activity derived from group II introns. Methods for preparing circRNAs based on the self-splicing activity of group II introns have advantages such as at least reducing the use of biological and chemical reagents (such as ligases and related reagents), ease of operation, and simplicity of design.
[0241] RNA designed to have ribozyme self-splicing activity forms circRNA by self-splicing and is also referred to herein as "cRNAzyme". In some embodiments, cRNAzyme may have in vitro self-splicing activity. Disclosed herein is a novel non-natural RNA having Group II intron self-splicing activity and forming circRNA, i.e., "Group II cRNAzyme", by self-splicing. Further provided herein are vectors containing polynucleotides encoding these Group II cRNAzymes, methods for preparing the Group II cRNAzymes disclosed herein by transcribing these vectors, and the use of these Group II cRNAzymes for creating circRNA.
[0242] Before further describing the present disclosure, it should be understood that the present disclosure is not limited to the specific embodiments described herein, and it should also be understood that the technical terms used herein are for the purpose of describing specific embodiments and are not intended to be limiting.
[0243] 8.1. Definitions As used herein, "essentially free of" with respect to a particular component means that the particular component is not intentionally formulated in the composition and / or is present only as a contaminant or in trace amounts. Thus, the total amount of the particular component resulting from unintentional contamination of the composition is less than 0.1%, preferably less than 0.05%, more preferably less than 0.01%. Most preferably, the composition is one in which the particular component is not detected at all by standard analytical methods.
[0244] As used herein, "a" or "an" may mean one or more. When used in combination with the word "comprising" in the claims herein, the words "a" or "an" may mean one or more.
[0245] As used herein, the term "or" in a claim is used to mean "and / or" unless explicitly indicated to refer to only alternative cases or when the alternatives are mutually exclusive, although the disclosure supports definitions that refer only to alternatives and "and / or". As used herein, "another" or "additional" may mean at least a second or more.
[0246] The term "about" as used herein is used to indicate that a value includes inherent variations due to, for example, errors of a device, the method used to determine the value, or variations that exist between the subjects of study. In some embodiments, "about" means that the variation is ±10%, ±9%, ±8%, ±7%, ±6%, ±5%, ±4%, ±3%, ±2%, ±1%, ±0.5%, ±0.2%, or ±0.1% of the value that "about" refers to. In some embodiments, "about" means that the variation is ±10%, ±9%, ±8%, ±7%, ±6%, ±5%, ±4%, ±3%, or ±2% of the value that "about" refers to. In some embodiments, "about" means that the variation is ±10% of the value that "about" refers to. In some embodiments, "about" means that the variation is ±5%, ±4%, ±3%, ±2%, or ±1% of the value that "about" refers to. In some embodiments, "about" means that the variation is ±1%, ±0.5%, ±0.2%, or ±0.1% of the value that "about" refers to.
[0247] Unless explicitly indicated to the contrary, this specification uses the terms "average" and "mean" interchangeably.
[0248] As used herein, the term "portion" when used with respect to a polypeptide or peptide refers to a fragment of the polypeptide or peptide. In some embodiments, a "portion" of a polypeptide or peptide retains at least one function and / or activity of the full-length polypeptide or peptide from which it is derived. For example, in some embodiments, if a full-length polypeptide binds to a particular ligand, a portion of that full-length polypeptide also binds to the same ligand.
[0249] Unless otherwise indicated, the terms "protein" and "polypeptide" are used interchangeably herein.
[0250] The term "exogenous", when used in reference to a protein, gene, nucleic acid, or polynucleotide within a cell or organism, refers to a protein, gene, nucleic acid, or polynucleotide that has been introduced into the cell or organism by artificial or natural means. Alternatively, when used in reference to a cell, the term refers to a cell that has been isolated and then introduced into a cell population or organism by artificial or natural means. Exogenous nucleic acids may be derived from a different organism or cell, or may be one or more additional copies of a nucleic acid that naturally exists within the organism or cell. Exogenous cells may be from a different organism or from the same organism. By way of non-limiting example, an exogenous nucleic acid is a nucleic acid at a chromosomal location different from its location in a native cell, or a nucleic acid adjacent to a nucleic acid sequence different from the nucleic acid sequence found in nature. The term "exogenous" is used in the same sense as the term "heterologous".
[0251] "Expression construct" or "expression cassette" means a nucleic acid molecule capable of inducing transcription. An expression construct includes at least one or more transcriptional control elements (such as a promoter, enhancer, or a functionally equivalent structure thereof) that induce gene expression in one or more desired cell types, tissues, or organs. Additional elements such as a transcription termination signal may also be included.
[0252] "Vector" or "construct" (sometimes referred to as a gene delivery system or gene transfer "vehicle") refers to a macromolecule or molecular complex that contains a polynucleotide or a protein expressed by the polynucleotide and is delivered to a host cell either in vitro or in vivo.
[0253] A "plasmid", which is a common type of vector, is an extrachromosomal DNA molecule separate from chromosomal DNA and can replicate independently of chromosomal DNA. In some cases, it is circular and double-stranded.
[0254] As used herein, the term "cRNAzyme" is used to refer to a linear ribonucleic acid (RNA) capable of generating circular RNA by a self-catalytic back-splicing reaction.
[0255] The term "cRNAzyme construct" is a linear RNA construct having cRNAzyme activity.
[0256] As used herein, the term "EBS" is used to refer to an exon binding sequence that interacts with an intron binding sequence (IBS) within an exon region (e.g., forms a complementary pair) and induces splicing by its own hydroxyl group within the EBS.
[0257] As used herein, the term "IBS" is used to refer to an intron binding sequence that interacts with an exon binding sequence (EBS) (e.g., forms a complementary pair) to identify a splicing site.
[0258] As used herein, the term "EBS1" is used to refer to exon binding sequence 1. In some embodiments, EBS1 comprises a nucleic acid sequence selected from the group consisting of (a) UAGGGC, (b) UUAUGG, (c) UCAACG, and (d) UGUGGC.
[0259] As used herein, the term "EBS2" is used to refer to exon binding sequence 2.
[0260] As used herein, the term "EBS3" is used to refer to exon binding sequence 3.
[0261] As used herein, the term "EBS1'" is used to refer to a modified EBS1 sequence that interacts with IBS1'. The interaction between EBS1' and IBS1' is similar to the interaction between EBS1 and IBS1. In some embodiments, EBS1' comprises a nucleic acid sequence selected from the group consisting of (a) UAGGGC, (b) UUAUGG, (c) UCAACG, and (d) UGUGGC.
[0262] As used herein, the term "EBS3'" is used to refer to a modified EBS3 sequence that interacts with IBS3'. The interaction between EBS3' and IBS3' is similar to the interaction between EBS3 and IBS3.
[0263] As used herein, the term "IBS1" is used to refer to intron binding sequence 1 that interacts with exon binding sequence 1 (EBS1) to specify a splicing site. In some embodiments, IBS1 comprises a nucleic acid sequence selected from the group consisting of (a) GCCCUG, (b) CCAUGG, (c) CGUUGA, and (d) GCCAUA.
[0264] As used herein, the term "IBS1'" is used to refer to a region on a target sequence having a function similar to that of IBS1.
[0265] As used herein, the term "IBS2" is used to refer to intron binding sequence 2 that interacts with exon binding sequence 2 (EBS2) to specify a splicing site.
[0266] As used herein, the term "IBS3" is used to refer to intron binding sequence 3 that interacts with exon binding sequence 3 (EBS3) to specify a splicing site. In some embodiments, IBS3 and its downstream sequence comprise a nucleic acid sequence selected from the group consisting of (a) AGCAAA, (b) AGCAGU, (c) AGAGAA, and (d) AGCAAA.
[0267] As used herein, the term "IBS3'" is used to refer to a region on a target sequence having a function similar to that of IBS3.
[0268] As used herein, the term "δ" (delta) is used to refer to the region of domain 1 of group II intron, which is a single nucleotide immediately upstream of EBS1. δ pairs with IBS3, and the interaction between δ and IBS3 is called δ-IBS3 pair formation. In some embodiments, the δ sequence and its upstream region include a nucleic acid sequence selected from the group consisting of (a) UGUGCU, (b) AAUGCU, (c) UGCUCU, and (d) UGUGCU.
[0269] As used herein, the term "δ''" (delta double prime) is used to refer to the region of domain 1 of group II intron, which is a single nucleotide immediately upstream of EBS1'. δ'' pairs with IBS3', and the interaction between δ'' and IBS3' is called δ''-IBS3' pair formation.
[0270] As used herein, the terms "group II intron" and "group II intron" are used interchangeably herein and refer to RNA molecules encoded by group II introns and sharing similar secondary and tertiary structures. Group II intron RNA molecules typically have six domains. Group II introns mainly include six stem-loop structures called domains 1 to 6 (D1 to D6), and the six domains are arranged in sequence and include multiple exon binding sequences (EBS) such as EBS1, EBS2, and EBS3.
[0271] Group II introns may contain modifications of one or more nucleotides compared to the wild type, and the modifications are selected from one or more of deletions, substitutions, and additions. In some embodiments, the modifications include modifications of one or more EBS sequences of the Group II intron, and each of the EBS sequences is complementary to one or more regions of corresponding length within the target sequence at at least 60% of the nucleotide positions. In some embodiments, the modifications are modifications of two EBS sequences of the Group II intron, such as EBS1 and EBS3, and each of the EBS sequences is complementary to two regions of corresponding length within the target sequence at at least 60% of the nucleotide positions, and preferably, the two regions are located at both ends of the target sequence, respectively. In some embodiments, the modifications are modifications of two EBS sequences of the Group II intron, such as EBS1’ and EBS3’, and each of the EBS sequences is complementary to two regions of corresponding length within the target sequence at at least 60% of the nucleotide positions, and preferably, the two regions are located at both ends of the target sequence, respectively. In some embodiments, the modifications are modifications of EBS1 and / or the δ sequence of the Group II intron, or modifications of EBS1’ and / or the δ” sequence, and the EBS1 and / or the δ sequence is complementary to a region of corresponding length of the target sequence at at least 60% of the nucleotides, and optionally, the modifications are modifications of the EBS1 and / or the δ sequence and its upstream sequence, and the EBS1 and / or the δ sequence and its upstream are complementary to a region of corresponding length within the target sequence at at least 60% of the nucleotides.In some embodiments, the modification is a modification of EBS1 and / or the δ sequence of group II introns, or a modification of EBS1' and / or the δ" sequence, wherein the EBS1 and / or δ sequence is complementary to a region of the corresponding length of the target sequence at at least 60%, 61%, 62%, 63%, 64%, 65%, 66%, 67%, 68%, 69%, 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% of the nucleotides. Optionally, the modification is a modification of EBS1 and / or the δ sequence and its upstream sequence, and the EBS1 and / or the δ sequence and its upstream sequence are complementary to a region of the corresponding length within the target sequence at at least 60% of the nucleotides. In some embodiments, the region of the corresponding length within the target sequence is IBS3, IBS3', IBS3 having a downstream sequence, or IBS3' having a downstream sequence. In some embodiments, the modification includes a deletion of part or all of domain 4, such as a deletion of the intron-encoded protein (IEP) sequence of domain 4, preferably a deletion of the entire domain 4. In some embodiments, the modification includes a deletion of an open reading frame (ORF).
[0272] As used herein, the terms "Domain 1" or "D1" are used to refer to the stem-loop structure of Domain 1 of Group II introns. As used herein, the terms "Domain 2" or "D2" are used to refer to the stem-loop structure of Domain 2 of Group II introns. As used herein, the terms "Domain 3" or "D3" are used to refer to the stem-loop structure of Domain 3 of Group II introns. As used herein, the terms "Domain 4" or "D4" are used to refer to the stem-loop structure of Domain 4 of Group II introns. As used herein, the terms "Domain 5" or "D5" are used to refer to the stem-loop structure of Domain 5 of Group II introns. As used herein, the terms "Domain 6" or "D6" are used to refer to the stem-loop structure of Domain 6 of Group II introns. The stem-loop structure is a type of RNA secondary structure and can be determined by an appropriate polynucleotide folding algorithm. Some programs are based on the calculation of the minimum Gibbs free energy. A representative algorithm is mFold (Zuker and Stiegler, Nucleic Acids Res. 9, 133-148 (1981)). Other exemplary folding algorithms are known in the art, as described, for example, in AR Gruber et al., Cell 106, 23-24 (2008), PA Carr and GM Church, Nature Biotechnology 27, 1151-62 (2009), and International Application Publication No. WO 2014 / 093709, the contents of which are incorporated herein by reference.
[0273] As used herein, the terms "5' intron fragment" and "3' intron fragment" are used to refer to intron fragments obtained by splitting a Group II intron at the loop region of the stem-loop structure of Domains 1, 2, 3, 4, 5, or 6.
[0274] The terms "nucleic acid sequence", "polynucleotide", and "oligonucleotide" are used interchangeably herein unless otherwise specified or the context indicates otherwise, and each refers to a polymer or oligomer of pyrimidine and / or purine bases such as cytosine, thymine, uracil, adenine, and guanine (see Albert L. Lehninger, Principles of Biochemistry, at 793-800 (Worth Pub. 1982)). These terms include deoxyribonucleotides, ribonucleotides, or peptide nucleic acid components, and chemical variants such as methylated, hydroxymethylated, or glycosylated forms of these bases. The polymer or oligomer may have a heterogeneous or homogeneous composition and may be isolated from a naturally occurring source, or may be produced artificially or synthetically. Further, the nucleic acid is DNA, RNA, or a mixture thereof and may exist permanently or transiently in single-stranded or double-stranded form, including homoduplex, heteroduplex, and hybrid states. A nucleic acid or nucleic acid sequence may include, for example, other types of nucleic acid structures such as DNA / RNA helices, peptide nucleic acids (PNA), morpholino nucleic acids (see, e.g., Braasch and Corey, Biochemistry, 4 / (14):4503-4510 (2002) and U.S. Patent No. 5,034,506), locked nucleic acids (LNA, see Wahlestedt et al., Proc. Natl. Acad. Sci. U.S.A., 97:5633-5638 (2000)), cyclohexenyl nucleic acids (see Wang, Am. Chem. Soc., 122:8595-8602 (2000)), and / or ribozymes. The terms "nucleic acid", "nucleic acid sequence", "polynucleotide", and "oligonucleotide" may also include chains containing non-natural nucleotides, modified nucleotides, and / or non-nucleotide components (e.g., "nucleotide analogs") that can perform the same functions as natural nucleotides. As used herein, the term "DNA sequence" is used to refer to a nucleic acid containing a series of DNA bases.
[0275] As is understood in the art, nucleic acid strands are inherently directional, the carbon atoms in the sugar ring are numbered from 1' to 5', the "5' end" has a free hydroxyl (or phosphate) on the 5' carbon, and the "3' prime end" has a free hydroxyl (or phosphate) on the 3' carbon. As used herein and as is understood in the art, a nucleic acid having specific sequence elements "from 5' to 3'" means that these sequence elements are arranged linearly from the 5' end to the 3' end of the nucleic acid.
[0276] As used herein, with respect to an oligonucleotide, "complementary" means that when the nucleobase sequence of the oligonucleotide and the nucleobase sequence of another nucleic acid are aligned in opposite directions, at least 70% of the nucleobases of the oligonucleotide or one or more regions thereof can hydrogen bond to the nucleobases of the other nucleic acid or one or more regions thereof. Complementary nucleobases mean nucleobases that can form hydrogen bonds with each other. Complementary nucleobase pairs include, but are not limited to, adenine (A) and thymine (T), adenine (A) and uracil (U), cytosine (C) and guanine (G), and 5-methylcytosine ( m C) and guanine (G). Complementary oligonucleotides and / or nucleic acids need not have nucleobase complementarity at each nucleoside. Rather, some mismatches are tolerated. As used herein, with respect to an oligonucleotide, "fully complementary" or "100% complementary" means that the oligonucleotide is complementary to another oligonucleotide or nucleic acid at each nucleoside of the oligonucleotide.
[0277] As used herein, the term "nucleobase" means an unmodified nucleobase or a modified nucleobase. As used herein, an "unmodified nucleobase" is adenine (A), thymine (T), cytosine (C), uracil (U), and guanine (G). As used herein, a "modified nucleobase" is a group of atoms other than unmodified A, T, C, U, or G that can pair with at least one unmodified nucleobase. "5-Methylcytosine" is a modified nucleobase. A universal base is a modified nucleobase that can form a pair with any of the five unmodified nucleobases. As used herein, "nucleobase sequence" means the order of consecutive nucleobases in a nucleic acid or oligonucleotide, independent of sugar or internucleoside bond modifications.
[0278] As used herein, the term "nucleoside" means a compound containing a nucleobase and a sugar moiety. The nucleobase and the sugar moiety are each independently either unmodified or modified. As used herein, a "modified nucleoside" means a nucleoside containing a modified nucleobase and / or a modified sugar moiety. Modified nucleosides include abasic nucleosides that lack a nucleobase. A "linked nucleoside" is a nucleoside linked in a continuous sequence (i.e., there are no additional nucleosides between the linked nucleosides).
[0279] The terms "polypeptide" and "protein" are used interchangeably herein, unless otherwise specified or the context suggests otherwise, and refer to a polymeric form of amino acids comprising at least two or more contiguous amino acids, including chemically or biochemically modified or derivatized amino acids. As used herein, the term "peptide" refers to the class of short polypeptides. The term peptide may refer to a polymer of amino acids (natural or non-natural) having a length of up to about 100 amino acids. For example, a peptide may be about 1 to about 10, about 10 to about 25, about 25 to about 50, about 50 to about 75, about 75 to about 100 amino acid residues in length. In some embodiments, a peptide may be about 100, about 200, about 300, about 400, about 500, about 600, about 700, about 800, about 900, about 1000, about 1250, about 1500, about 1750, about 2000, about 2250, about 2500, about 2750, about 3000, about 3250, about 3500, about 3750, about 4000, about 4250, about 4500, about 4750, about 5000 amino acid residues in length.
[0280] The nomenclature for nucleotides, nucleic acids, nucleosides, and amino acids used herein conforms to the standards of the International Union of Pure and Applied Chemistry (IUPAC) (see, e.g., bioinformatics.org / smsylupac.html).
[0281] When referring to a nucleic acid sequence or a protein sequence, the term "identity" is used to indicate the similarity between two sequences. The similarity or identity of sequences can be determined using standard techniques known in the art. Examples include, but are not limited to, the local sequence identity algorithm of Smith & Waterman, Adv. Appl. Math. 2, 482 (1981), the sequence identity alignment algorithm of Needleman & Wunsch, J Mol. Biol. 48, 443 (1970), the similarity search method of Pearson & Lipman, Proc. Natl. Acad. Sci. USA 85, 2444 (1988), the computer implementations of these algorithms (GAP, BESTFIT, FASTA, and TFASTA in the Wisconsin Genetics Software Package, Genetics Computer Group, 575 Science Drive, Madison, WI), the best fit sequence program described in Devereux et al., Nucl. Acid Res. 12, 387-395 (1984), and the inspection. Another algorithm is the BLAST algorithm described in Altschul et al., J Mol. Biol. 215, 403-410, (1990) and Karlin et al., Proc. Natl. Acad. Sci. USA 90, 5873-5787 (1993). A particularly useful BLAST program is the WU-BLAST-2 program obtained from Altschul et al., Methods in Enzymology, 266, 460-480 (1996); blast.wustl / edu / blast / README.html. WU-BLAST-2 uses several search parameters, which are set to default values optionally. The parameters are dynamic values and are set by the program itself according to the composition of a particular sequence and the composition of a particular database to search for the target sequence, but the values may be adjusted to increase sensitivity.Furthermore, the gapped BLAST reported by Altschul et al., (1997) Nucleic Acids Res. 25, 3389-3402 is also a useful algorithm. Unless otherwise specified, the percent identity is determined using the algorithm available at the Internet address: blast.ncbi.nlm.nih.gov / Blast.cgi.
[0282] The terms "internal ribosome entry site", "internal ribosome entry site sequence", "IRES", and "IRES sequence region" are used interchangeably herein and refer to cis-elements of viral or human cellular RNAs (e.g., messenger RNA (mRNA) and / or circRNA) that bypass the process of canonical eukaryotic cap-dependent translation initiation. In the canonical cap-dependent mechanism used by most eukaryotic mRNAs, an m 7 G cap, initiator Met-tRNAmet, more than 12 initiation factor proteins, a directional scan, and GTP hydrolysis are required to position a translationally competent ribosome at the start codon. An IRES typically consists of a long and highly structured 5-UTR that mediates the binding of the translation initiation complex and catalyzes the formation of a functional ribosome.
[0283] The term "IRES-like sequence" or "internal ribosome entry site-like sequence" refers to a synthetic nucleotide sequence that exhibits the function of a native IRES. In some embodiments, the IRES-like sequence can recruit ribosomal components to mediate cap-independent translation.
[0284] When referring to a nucleic acid sequence, the terms "coding sequence", "coding sequence region", "coding region", and "CDS" are used interchangeably herein and may refer to a portion of a DNA or RNA sequence that is translated or potentially translatable into a protein. As used herein, the terms "reading frame", "open reading frame", and "ORF" may be used interchangeably to refer to a nucleotide sequence that begins with a start codon (e.g., ATG) and, in some embodiments, ends with a stop codon (e.g., TAA, TAG, or TGA). An open reading frame may include introns and exons, and thus all CDSs are ORFs, but not all ORFs are CDSs.
[0285] The term "X-mer" refers to "X" nucleic acid residues / nucleotides (where "X" is an integer) in a polynucleotide sequence (e.g., an IRES-like sequence). For example, "10-mer" refers to a polynucleotide sequence of the disclosure having 10 nucleic acid residues / nucleotides.
[0286] The terms "complementary" and "complementarity" refer to the relationship between two nucleic acid sequences or nucleic acid monomers that have the ability to form hydrogen bonds with each other by traditional Watson-Crick base pairing or other non-traditional types of pairing. The degree of complementarity between two nucleic acid sequences is indicated by the proportion of nucleotides in the nucleic acid sequence that can form hydrogen bonds (e.g., Watson-Crick base pairing) with the second nucleic acid sequence (e.g., about 50%, about 60%, about 70%, about 80%, about 90%, and 100% complementary). Two nucleic acid sequences are "fully complementary" if all consecutive nucleotides of one nucleic acid sequence hydrogen bond with the same number of consecutive nucleotides within the second nucleic acid sequence. The degree of complementarity between two nucleic acid sequences is at least 60% (e.g., at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 97%, at least 98%, at least 99%, or 100%) (e.g., at least 9, at least 10, at least 11, at least 12, at least 13, at least 14, at least 15, at least 16, at least 17, at least 18, at least 19, at least 20, at least 21, at least 22, at least 23, at least 24, at least 25, at least 30, at least 35, at least 40, at least 45, at least 50, or more nucleotides) over a region of at least 8 nucleotides, or if the two nucleic acid sequences hybridize under at least moderate, or in some embodiments, high stringency conditions, the two nucleic acid sequences are "substantially complementary".Examples of medium stringency conditions include incubating overnight at 37°C in a solution containing 20% formamide, 5% SSC (150 mM sodium chloride (NaCl), 15 mM trisodium citrate), 50 mM sodium phosphate (pH 7.6), 5× Denhardt's solution, 10% dextran sulfate, and 20 mg / mL denatured sonicated salmon sperm DNA, followed by washing the filter in 1× SSC at approximately 37 - 50°C or under substantially similar conditions, such as those described in Sambrook, J., Molecular Cloning: A Laboratory Manual, Cold Spring Harbor Laboratory Press; 4th edition (June 15, 2012) for medium stringency. High stringency conditions include, for example, (1) using 0.015 M sodium chloride / 0.0015 M sodium citrate / 0.1% sodium dodecyl sulfate (SDS) at 50°C for washing, where the ionic strength is low and the washing temperature is high; (2) using a denaturing agent such as 50% (v / v) formamide combined with 0.1% bovine serum albumin (BSA) / 0.1% Ficoll / 0.1% polyvinylpyrrolidone (PVP) / 50 mM sodium phosphate buffer at pH 6.5 containing 750 mM sodium chloride and 75 mM sodium citrate at 42°C during hybridization; or (3) using 50% formamide, 5× SSC (0.75 M sodium chloride (NaCl), 0.075 M sodium citrate), 50 mM sodium phosphate (pH 6.8), 0.1% sodium pyrophosphate, 5× Denhardt's solution, sonicated salmon sperm DNA (50 pg / mL), 0.1% SDS, and 10% dextran sulfate at 42°C, and washing with (i) 0.2× SSC at 42°C, (ii) 50% formamide at 55°C, and (iii) 0.1× SSC (optionally in combination with EDTA) at 55°C.Additional details and explanations regarding the stringency of hybridization reactions are provided, for example, in Sambrook, supra, and Ausubel et al., eds., Short Protocols in Molecular Biology, 5th ed., John Wiley & Sons, Inc., Hoboken, N.J. (2002).
[0287] The term "animal" refers to humans and non-human animals, including non-human primates (e.g., bonobos, chimpanzees, gorillas, monkeys), as well as other animals such as birds, pigs, horses, dogs, cats, rabbits, mice, rats, cows, sheep, goats, and deer.
[0288] The terms "hybridization" or "hybridized," when referring to nucleic acid sequences, refer to associations formed between and / or among complementary sequences.
[0289] The term "control element" collectively refers to promoter regions, polyadenylation signals, transcription termination sequences, upstream regulatory domains, origins of replication, internal ribosome entry sites (IRES), enhancers, splice junctions, and the like, which generally indicate replication, transcription, post-transcriptional processing, and translation of the coding sequence in the recipient cell. Not all of these control elements need to be present as long as the selected coding sequence is capable of being replicated, transcribed, and translated in the appropriate host cell.
[0290] As used herein, the term "promoter" is used to refer to a nucleotide region that contains a DNA regulatory sequence, which regulatory sequence is derived from a gene that binds RNA polymerase and enables the initiation of transcription of a downstream (3' direction) coding sequence. It may contain genetic elements to which regulatory proteins and molecules such as RNA polymerase and other transcription factors bind to initiate the specific transcription of a nucleic acid sequence. The phrases "operatively positioned", "operatively linked", "under control", and "under transcriptional control" mean that the promoter is in the correct functional position and / or orientation with respect to a nucleic acid sequence and controls the initiation of transcription and / or expression of that sequence.
[0291] "Enhancer" means a nucleic acid sequence that, when placed near a promoter, increases the transcriptional activity compared to the transcriptional activity that occurs from the promoter in the absence of the enhancer domain.
[0292] With respect to a nucleic acid molecule, "operatively linked" means that two or more nucleic acid molecules (e.g., a nucleic acid molecule to be transcribed, a promoter, and a functional effector element) are linked in such a way as to enable the transcription of the nucleic acid molecule.
[0293] The term "homology" refers to the percentage of identity between the nucleic acid residues of two polynucleotides or the amino acid residues of two polypeptides. The correspondence between one sequence and another can be determined by techniques known in the art. For example, sequence information can be aligned and homology can be determined by directly comparing the sequence information between two polypeptides using readily available computer programs. Two polynucleotides (e.g., DNA) or two polypeptide sequences are "substantially homologous" to each other if at least about 80%, preferably at least about 90%, and most preferably at least about 95% of the nucleotides or amino acids, respectively, over the defined length of the molecule match as determined using the above methods.
[0294] "Treatment" or "treating a disease or condition" refers to implementing a protocol or treatment plan that involves administering one or more agents or active ingredients to a patient to alleviate the signs or symptoms of a disease or the recurrence of a disease. Desirable effects of treatment include a reduced rate of disease progression, improvement or alleviation of the medical condition, remission, extended survival rate, improved quality of life, or improved prognosis. Alleviation or prevention can be carried out either before or after the signs or symptoms of a disease or condition appear. "Treatment" or "therapy" does not require complete alleviation of signs or symptoms, nor does it require a cure.
[0295] As used throughout this disclosure, the terms "therapeutic benefit" or "therapeutically effective" refer to anything that promotes or enhances the health of a subject with respect to the medical treatment of a condition. This includes, but is not limited to, a reduction in the frequency, severity, or rate of progression of the signs or symptoms of a disease. For example, treatment of cancer can include, for example, a reduction in the size of a tumor, a reduction in the invasiveness of a tumor, a decrease in the cancer growth rate, or a decrease in the metastasis or recurrence rate. Treatment of cancer may also refer to extending the survival period of a subject afflicted with cancer.
[0296] The phrase "pharmaceutically or pharmacologically acceptable" refers to molecular entities and compositions that do not cause adverse reactions, allergic reactions, or other undesirable reactions when administered, as necessary, to animals such as humans. It is understood that in the case of administration to animals (e.g., humans), the formulation must meet the standards of sterility, pyrogenicity, general safety, and purity required by, for example, the FDA's Office of Biological Standards.
[0297] As used herein, the term "pharmaceutically acceptable carrier" includes any aqueous biocompatible solvent known to those skilled in the art (e.g., parenteral vehicles such as physiological saline, phosphate buffered saline, sodium chloride, Ringer's dextrose), antioxidants, preservatives (e.g., antibacterial or antifungal agents, antioxidants, chelating agents, and inert gases), substances such as isotonic agents, and combinations thereof. The pH and exact concentrations of the various components in the pharmaceutical composition are adjusted according to well-known parameters.
[0298] As used herein, the term "bulged adenosine" (also known as bulge A) refers to a residue within the intron sequence as the initiating nucleophile. Conventional group II bulged adenosines are located in domain 6 (D6 or DVI) of the group II intron. In some group II introns, the bulged adenosine is 7 or 8 nucleotides away from the 3' splicing site. The bulged adenosine is usually conserved and plays a central role in the splicing process. During this process, the 2'-hydroxyl of the bulged adenosine attacks the 5' splice site, followed by a nucleophilic attack on the 3' splice site by the 3'-OH of the upstream exon. As a result, a branched intron lariat linked by a 2'-phosphodiester bond at the bulged adenosine site is formed (Van der Veen et al., The EMBO Journal, 6(12):3827-3831 (1987); Jacquier et al., Journal of molecular biology, 219(3): 415-428 (1991); Daniels et al., Journal of molecular biology, 256(1):31-49 (1996)).
[0299] As used herein, the term "atypical bulge adenosine" (also known as atypical bulge A) refers to the region within D6 of group IIB intron C.te.I1 (Cte) found in the human pathogen Clostridium tetani. There is no obvious bulge adenosine in D6 of Cte. Instead, there is a loop region (see FIGS. 2A-2E), which functions as a nucleophile in the first step (branching pathway) of splicing (McNeil et al., RNA, 20(6):855-866 (2014)).
[0300] As used herein, the term "catalytic triad" refers to a highly conserved region (AGC) that forms a triple helix known as a "catalytic triplex" by forming "base triples" with other nucleotides. This catalytic triplex forms a binding pocket for two active site magnesium ions (Chan et al. Nature Comm. 9.1 (2018): 1-10).
[0301] As used herein, the term "scar" refers to a non-target sequence region within a circRNA splicing product. As used herein, the term "scarless splicing" refers to the self-splicing of a cRNAzyme that generates a circRNA that does not contain additional sequence elements outside of the target sequence. Thus, "scarless" circRNA means that it does not contain a scar and consists only of the target sequence. As used herein, the term "splicing that is almost scarless" refers to the self-splicing of a cRNAzyme that generates a circRNA that contains 20 nucleotides or less outside of the target sequence. An "almost scarless" circRNA is a circRNA that results from the "almost scarless" self-splicing of a cRNAzyme and contains only 20 nucleotides or less outside of the target sequence.
[0302] The term "in vitro transcription" or "IVT" refers to a versatile method of synthesizing RNA from a DNA template using RNA polymerase, ribonucleotides, and appropriate buffer conditions to generate RNA in vitro.
[0303] As used herein, the term "resulting target sequence" refers to the target sequence formed within a circRNA upon self-splicing of the RNA (or cRNAzyme) provided herein.
[0304] 8.2. Cap-independent translation initiation elements In some embodiments, an RNA polynucleotide is provided that includes a modified translation initiation element (TI) comprising an internal ribosome entry site (IRES)-like polynucleotide sequence. In some embodiments, the IRES-like polynucleotide sequence is capable of mediating cap-independent translation initiation.
[0305] Translation initiation of mRNA in eukaryotic cells is a complex process involving the coordinated interaction of multiple factors (Pain (1996) Eur. J. Biochem. 236, 747-771). In the case of most mRNAs, the first step is to recruit the ribosomal 40S subunit to the mRNA at or near the capped 5' end. The binding of 40S to mRNA is greatly facilitated by the cap-binding protein complex eIF4F. The factor eIF4F is composed of three subunits: the RNA helicase eIF4A, the cap-binding protein eIF4E, and the multi-adapter protein eIF4G, which functions as a scaffold for the proteins within the complex and has binding sites for eIF4E, eIF4a, eIF3, and the poly(A)-binding protein.
[0306] Circular RNAs are a type of single-stranded covalently closed-loop RNAs that do not contain the 5' cap which is well-known to be required for cap-dependent translation. Thus, in circular RNA translation, alternative mechanisms for initiating cap-independent translation are utilized, such as the use of internal ribosome entry site (IRES) sequences recognized by ribosomes. See, for example, Wesselhoeft, R. A. et al., Nat. Commun. 9, 2629 (2018) and Chinese Patent Application No. 2021 / 10594352.4.
[0307] 8.2.1. Natural internal ribosome entry site (IRES) sequences In some embodiments, the RNA polynucleotides, precursor RNAs, and circular RNAs provided herein include a natural internal ribosome entry site (IRES) sequence. In some embodiments, the IRES sequence is an RNA sequence capable of binding to a ribosome (e.g., a eukaryotic ribosome). The IRES sequence enables the translation of one or more open reading frames (e.g., open reading frames forming an expression sequence) from the circular RNA. The IRES attracts the ribosome (e.g., a eukaryotic ribosome) translation initiation complex and promotes translation initiation.
[0308] A number of natural IRES sequences are available, including sequences derived from or isolated from various viruses such as the leader sequences of picornaviruses such as encephalomyocarditis virus (EMCV) UTR, poliovirus leader sequence, hepatitis A virus leader sequence, hepatitis C virus IRES, human rhinovirus type 2 IRES, foot-and-mouth disease virus IRES element, dicistrovirus IRES, and the like.
[0309] In some embodiments, the native IRES sequence is the IRES sequence of Taura syndrome virus, the IRES sequence of the cricket paralysis virus, Theiler's encephalomyelitis virus, simian virus 40, the IRES sequence of the red imported fire ant (Solenopsis invicta) virus 1, the IRES sequence of the bird cherry-oat aphid (Rhopalosiphum padi) virus, reticuloendotheliosis virus, human poliovirus type 1, the IRES sequence of the brown marmorated stink bug midgut (Plautia stali intestine) virus, Kashmir bee virus, human rhinovirus type 2, the IRES sequence of the glassy-winged sharpshooter (Homalodisca coagulata) virus 1, human immunodeficiency virus type 1, the IRES sequence of the minute pirate bug P virus, hepatitis C virus, hepatitis A virus, hepatitis A virus HA16, GB hepatitis virus, foot-and-mouth disease virus, human enterovirus 71, equine rhinitis virus, the IRES sequence of Ectrapis obliqua picornapicoma-like virus, encephalomyocarditis virus, Drosophila C virus, human coxsackievirus B3, Brassicaceae tobamovirus, cricket paralysis virus, bovine viral diarrhea virus 1, black queen cell virus, aphid lethal paralysis virus, avian encephalomyelitis virus, acute bee paralysis virus, hibiscus chlorotic ringspot virus, classical swine fever virus, tobacco etch virus, turnip crinkle virus, EMCV-A, EMCV-B, EMCV-Bf, EMCV-Cf, EMCVpEC9, picobirnavirus, HCVQC64, human cosavirus E / D, human cosavirus F, human cosavirus JMY, rhinovirus NAT001, HRV14, HRV89, HRV C-02, HRV-A21, salivirus ASHI, salivirus FHB, salivirus NG-J1, human parechovirus 1, kobuvirus B, Yc-3, roseolovirus M-7, shambe virus A, pasi virus A, pasi virus A2, echovirus E14, human parechovirus 5, aichi virus, hopivirus, CVA10, enterovirus C, enterovirus D, enterovirus J, human pegivirus 2, GBV-C GT110, GBV-C K1737, GBV-C Iowa, pegivirus A1220, pasi virus A3, sapelovirus, roseolovirus B, Bakunsa virus, tremovirus A, porcine pasi virus 1, PLV-CHN, pasi virus A, sisi virus, hepaci virus K, hepaci virus A, BVDV1, border disease virus, BVDV2, CSFV-PK15C, SF573 dicistravirus, Hubei Picoma-like virus, CRPV, salivirus ABN5, salivirus ABN2, salivirus A02394, salivirus AGUT, salivirus ACH, salivirus ASZ1, salivirus FHB, coxsackievirus (e.g., CVA3, CVA12, CVB1, CVB3, CVB5), echovirus 7, enterovirus A71, and / or EV24, or is isolated from or derived from them.
[0310] In some embodiments, the native IRES sequence is isolated from or derived from a eukaryotic IRES element selected from human FGF2, human SFTPA1, human AML1 / RUNX1, Drosophila Antennapedia, human AQP4, human AT1R, human BAG-1, human BCL2, human BiP, human c-IAP1, human c-myc, human eIF4G, mouse NDST4L, human LEF1, mouse HIF1 alpha, human n.myc, mouse Gtx, human p27kipl, human PDGF2 / c-sis, human p53, human Pim-1, mouse Rbm3, Drosophila Reaper, canine Scamper, Drosophila Ubx, human UNR, mouse UtrA, human VEGF-A, human XIAP, Drosophila Hairless, yeast (S. cerevisiae) TFIID, and yeast YAP1.
[0311] In some embodiments, the native IRES sequence is an endogenous IRES sequence derived from or isolated from Homo sapiens. In some embodiments, the native IRES sequence is an endogenous IRES sequence derived from or isolated from human tissue or a human sample.
[0312] In some embodiments, the IRES sequence is isolated from or derived from cellular IRES elements selected from AML1 / RUNX1, Antp-D, Antp-DE, Antp-CDE, ATlR var1, ATlR_var2, ATlR_var3, ATlR_var4, BAGl_p36delta236nt, BAGl_p36, BiP_-222_-3, C-IAP1 285-1399, c IAP1 13 13-1462, c-jun, Cat-l_224, CCND1, eIF4GI-ext, eIF4GII, eIF4GII-long, FGF1A, FMR1, Gtx-l33-l4l, Gtx-l-l66, Gtx-l-l20, Gtx-l-l96, HAP4, HIFla, hSNMl, HsplOl, hsp70, hsp70, Hsp90, IGF2_leader2, L-myc, MNT 75-267, MNT 36-160, MTG8a, MYB, MYT2 997-1 152, NRF_-653_-l7, NtHSF1, ODC1, p27kipl, p53_l28-269, PDGF2 / c-sis, PITSLRE_p58, Rbm3, Reaper, SCaMPER, TFIID, TIF4631, Ubx_l-966, Ubx_373-96l, UNR, Ure2, XIAP 5-464, XIAP 305-466, YAP1, (GAAA)16, (PPT19)4, and XI.
[0313] In some embodiments, the IRES sequence is isolated from or derived from viral IRES elements selected from ABPV IGRpred, AEV, ALPV IGRpred, BQCV IGRpred, BVDV1 1-385, BVDV1 29-391, CrPV 5NCR, CrPV IGR, crTMV_IRESmp228, CSFV, DCV IGR, EoPV_5NTR, ERBV_l62-920, EV7l_l-748, FMDV type C, GBV-A, GBV-C, HAV HM175, HiPVJGRpred, HIV-1, HoCVlJGRpred, IAPVJGRpred, idefix, KBV IGRpred, PSIV IGR, PV type1 Mahoney, PV_type3_Leon, REV-A, RhPV 5NCR, RhPV IGR, SINV l IGRpred, SV40 661-830, TMEV, TMV_UI_IRESmp228, TRV 5NTR, TrV IGR, TSV, and IGR.
[0314] In some embodiments, the IRES sequences of the present disclosure include sequences isolated from or derived from natural IRES sequences. Exemplary natural IRES sequences include, but are not limited to, those listed in Tables 98-100. In some embodiments, the IRES sequence comprises a nucleic acid sequence that is at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 100% identical to the nucleic acid sequence of the natural IRES sequence of Tables 98-100 or a functional portion thereof.
[0315] 8.2.2. IRES-like sequences In some embodiments, the RNA polynucleotides, precursor RNAs, and circular RNAs provided herein include an IRES-like sequence. As used herein, the term "IRES-like sequence" refers to a synthetic or artificial IRES sequence having the function of a native IRES sequence. In some embodiments, the IRES-like sequence is an RNA sequence capable of recruiting ribosomes (e.g., eukaryotic ribosomes). In some embodiments, the IRES-like sequence is an RNA sequence capable of mediating cap-independent translation initiation. An IRES-like sequence can be identified by any of the methods disclosed herein.
[0316] In some embodiments, the length of the IRES-like sequence is 3 nucleic acid residues or more. In some embodiments, the length of the IRES-like sequence is 3 to 300 nucleic acid residues. In some embodiments, the length of the IRES-like sequence is 3 to 200 nucleic acid residues. In some embodiments, the length of the IRES-like sequence is 4 to 200 nucleic acid residues. In some embodiments, the length of the IRES-like sequence is 5 to 200 nucleic acid residues. In some embodiments, the length of the IRES-like sequence is 6 to 200 nucleic acid residues. In some embodiments, the length of the IRES-like sequence is 7 to 200 nucleic acid residues. In some embodiments, the length of the IRES-like sequence is 3 to 100 nucleic acid residues. In some embodiments, the length of the IRES-like sequence is 4 to 100 nucleic acid residues. In some embodiments, the length of the IRES-like sequence is 5 to 100 nucleic acid residues. In some embodiments, the length of the IRES-like sequence is 6 to 100 nucleic acid residues. In some embodiments, the length of the IRES-like sequence is 7 to 100 nucleic acid residues. In some embodiments, the length of the IRES-like sequence is 3 to 50 nucleic acid residues. In some embodiments, the length of the IRES-like sequence is 4 to 50 nucleic acid residues. In some embodiments, the length of the IRES-like sequence is 5 to 50 nucleic acid residues. In some embodiments, the length of the IRES-like sequence is 6 to 50 nucleic acid residues. In some embodiments, the length of the IRES-like sequence is 7 to 50 nucleic acid residues. In some embodiments, the length of the IRES-like sequence is 3 to 40 nucleic acid residues. In some embodiments, the length of the IRES-like sequence is 4 to 40 nucleic acid residues. In some embodiments, the length of the IRES-like sequence is 5 to 40 nucleic acid residues. In some embodiments, the length of the IRES-like sequence is 6 to 40 nucleic acid residues. In some embodiments, the length of the IRES-like sequence is 7 to 40 nucleic acid residues. In some embodiments, the length of the IRES-like sequence is 3 to 30 nucleic acid residues. In some embodiments, the length of the IRES-like sequence is 4 to 30 nucleic acid residues. In some embodiments, the length of the IRES-like sequence is 5 to 30 nucleic acid residues. In some embodiments, the length of the IRES-like sequence is 6 to 30 nucleic acid residues.In some embodiments, the length of the IRES-like sequence is 7 to 30 nucleic acid residues. In some embodiments, the length of the IRES-like sequence is 3 to 20 nucleic acid residues. In some embodiments, the length of the IRES-like sequence is 4 to 20 nucleic acid residues. In some embodiments, the length of the IRES-like sequence is 5 to 20 nucleic acid residues. In some embodiments, the length of the IRES-like sequence is 6 to 20 nucleic acid residues. In some embodiments, the length of the IRES-like sequence is 7 to 20 nucleic acid residues. In some embodiments, the length of the IRES-like sequence is 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, or 100 nucleic acid residues. In some embodiments, the length of the IRES-like sequence is 5 nucleic acid residues. In some embodiments, the length of the IRES-like sequence is 6 nucleic acid residues. In some embodiments, the length of the IRES-like sequence is 7 nucleic acid residues. In some embodiments, the length of the IRES-like sequence is 8 nucleic acid residues. In some embodiments, the length of the IRES-like sequence is 9 nucleic acid residues. In some embodiments, the length of the IRES-like sequence is 10 nucleic acid residues. In some embodiments, the length of the IRES-like sequence is 11 nucleic acid residues. In some embodiments, the length of the IRES-like sequence is 12 nucleic acid residues.
[0317] Exemplary IRES-like sequences of the present disclosure include, but are not limited to, those listed in Tables 2-96, Table 110, and Table 112. In some embodiments, the IRES-like sequence comprises a nucleic acid sequence that is at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 100% identical to the IRES-like sequence or a functional portion thereof in Tables 2-96, Table 110, and Table 112.
[0318] In some embodiments, the IRES-like sequence or a portion thereof described herein can combine with the native IRES sequence or a portion thereof described herein to form additional IRES-like sequences. In some embodiments, a complete native IRES sequence is combined with a complete IRES-like sequence. In some embodiments, a portion of the native IRES sequence is combined with a complete IRES-like sequence. In some embodiments, a complete native IRES sequence is combined with a portion of the IRES-like sequence. In some embodiments, a portion of the native IRES sequence is combined with a portion of the IRES-like sequence. Exemplary sequences combining a portion of the native IRES sequence and a portion of the IRES-like sequence include, but are not limited to, those listed in Table 97. 8.2.2.1. Method for Generating IRES-like Sequences The present disclosure provides a method for generating an IRES-like polynucleotide sequence, comprising: (a) generating a polynucleotide query sequence (i.e., an X-mer) consisting of X nucleic acid residues, wherein X is an integer of 3 or more; (b) generating X-Y+1 overlapping polynucleotide fragment sequences within the polynucleotide query sequence, wherein each polynucleotide fragment sequence consists of Y nucleic acid residues, the first position of each polynucleotide fragment sequence is n, the last position of the same polynucleotide fragment sequence is Y+n-1, and n represents each positive integer between 1 and X-Y+1. Step of determining the enrichment score of each polynucleotide fragment sequence of (c)(b); Step of determining the numerical score of the polynucleotide query sequence by summing the enrichment scores of each polynucleotide fragment sequence of (d)(c), and Step of identifying the polynucleotide query sequence as a modified IRES-like polynucleotide sequence according to a reference value A method is provided that includes the above steps.
[0319] In some embodiments, the reference value is determined as described, for example, in the generation of the polynucleotide query sequence in 8.2.2.1.a and the generation of polynucleotide sub-sequences that overlap within the query sequence in 8.2.2.1.b.
[0320] In some embodiments, the reference value is characteristic of the absence of therapeutic agent expression.
[0321] In some embodiments, the reference value is the average score of all the numerical scores of two or more IRES-like polynucleotide sequences, two or more natural IRES sequences, or a combination thereof.
[0322] In some embodiments, the reference value is about 95%, about 90%, about 85%, about 80%, about 75%, about 70%, about 65%, about 60%, about 55%, about 50%, about 45%, about 40%, about 35%, about 30%, about 25%, about 20%, about 15%, about 10%, about 9%, about 8%, about 7%, about 6%, about 5%, about 4%, about 3%, about 2%, about 1%, about 0.9%, about 0.8%, about 0.7%, about 0.6%, about 0.5%, about 0.4%, about 0.3%, about 0.2%, or about 0.1% of the highest numerical score of two or more IRES-like polynucleotide sequences, two or more natural IRES sequences, or a combination thereof.
[0323] In some embodiments, the reference value is about 50%, about 45%, about 40%, about 35%, about 30%, about 25%, about 20%, about 15%, or about 10% of the highest numerical score of two or more IRES-like polynucleotide sequences, two or more natural IRES sequences, or combinations thereof.
[0324] In some embodiments, the reference value is at least about 95%, about 90%, about 85%, about 80%, about 75%, about 70%, about 65%, about 60%, about 55%, about 50%, about 45%, about 40%, about 35%, about 30%, about 25%, about 20%, about 15%, about 10%, about 9%, about 8%, about 7%, about 6%, about 5%, about 4%, about 3%, about 2%, about 1%, about 0.9%, about 0.8%, about 0.7%, about 0.6%, about 0.5%, about 0.4%, about 0.3%, about 0.2%, or about 0.1% higher than the lowest numerical score of two or more IRES-like polynucleotide sequences, two or more natural IRES sequences, or combinations thereof.
[0325] In some embodiments, the reference value is at least about 95%, about 90%, about 85%, about 80%, about 75%, about 70%, about 65%, about 60%, about 55%, or about 50% higher than the lowest numerical score of two or more IRES-like polynucleotide sequences, two or more natural IRES sequences, or combinations thereof.
[0326] In some embodiments, the reference value is 0 or more.
[0327] In some embodiments, the numerical score of the polynucleotide query sequence is calculated using a system of equations, and the numerical score of the polynucleotide query sequence is 0 or more.
[0328] Regarding the number of nucleic acid residues in the polynucleotide query sequence (i.e., X), it is described in detail herein, for example, in 8.2.2.1.a Generation of Polynucleotide Query Sequences and 8.2.2.1.b Generation of Polynucleotide Subsequences Duplicating within the Query Sequence.
[0329] In some embodiments, X is an integer selected from 3 to 300. In some embodiments, X is an integer selected from 3 to 100. In some embodiments, X is an integer selected from 5 to 100. In some embodiments, X is an integer selected from 6 to 100. In some embodiments, X is an integer of 5 or more. In some embodiments, X is an integer of 6 or more.
[0330] In some embodiments, the polynucleotide query sequence is generated within a DNA vector suitable for synthesizing the polynucleotide query sequence. In some embodiments, the polynucleotide query sequence is chemically synthesized.
[0331] An enrichment score for a polynucleotide fragment sequence is determined by: i) generating an expression plasmid library, wherein each expression plasmid of the library contains a different polynucleotide fragment sequence and a reporter gene; ii) contacting a population of cells with the expression plasmid library; iii) quantifying the expression level of the reporter gene corresponding to each expression plasmid of the expression plasmid library; iv) dividing the entire population of cells into a first population and a second population based on the protein expression level of iii); and v) determining the enrichment score of the polynucleotide fragment sequence.
[0332] The enrichment score of the polynucleotide fragment sequence can be determined according to the methods described herein, for example, in accordance with the determination of the enrichment score in 8.2.2.1.c.
[0333] The numerical score of the polynucleotide query sequence can be determined according to the methods described herein, for example, in accordance with the determination of the numerical score of the polynucleotide sequence and the identification of the IRES-like sequence in 8.2.2.1.d.
[0334] 8.2.2.1.a. Generation of Polynucleotide Query Sequence A random polynucleotide query sequence library is known in the art or can be generated by any suitable method described herein. In some embodiments, a random polynucleotide query sequence (e.g., an RNA polynucleotide) is generated using any DNA vector suitable for synthesizing the polynucleotide query sequence. In some embodiments, a random polynucleotide query sequence (e.g., an RNA polynucleotide) is generated by chemical synthesis, error-prone PCR, or transcriptome reverse transcription PCR. For example, to generate a random decamer sequence library, the Klenow fragment (NEB) can be used to generate a foldback primer [Chemical Formula] and the resulting DNA can be digested with BsmBI and ligated to pcircGFP-BsmBI digested with BsmBI. The ligation product can be transformed into ElectroMax DH-5α (Invitrogen) to generate E. coli clones (a total of 2 million clones) that achieve a coverage of approximately 2-fold of all possible DNA decamers.
[0335] In some embodiments, the length of the random polynucleotide query sequence is 3 residues or more. In some embodiments, the length of the random polynucleotide query sequence is 3 to 300 nucleic acid residues. In some embodiments, the length of the random polynucleotide query sequence is 3 to 200 nucleic acid residues. In some embodiments, the length of the random polynucleotide query sequence is 3 to 100 nucleic acid residues. In some embodiments, the length of the random polynucleotide query sequence is 5 to 100 nucleic acid residues. In some embodiments, the length of the random polynucleotide query sequence is 6 to 100 nucleic acid residues. In some embodiments, the length of the random polynucleotide query sequence is 7 to 100 nucleic acid residues. In some embodiments, the length of the random polynucleotide query sequence is 7 to 90 nucleic acid residues. In some embodiments, the length of the random polynucleotide query sequence is 7 to 80 nucleic acid residues. In some embodiments, the length of the random polynucleotide query sequence is 7 to 70 nucleic acid residues. In some embodiments, the length of the random polynucleotide query sequence is 7 to 60 nucleic acid residues. In some embodiments, the length of the random polynucleotide query sequence is 7 to 50 nucleic acid residues. In some embodiments, the length of the random polynucleotide query sequence is 7 to 40 nucleic acid residues. In some embodiments, the length of the random polynucleotide query sequence is 7 to 30 nucleic acid residues. In some embodiments, the length of the random polynucleotide query sequence is 7 to 20 nucleic acid residues. In some embodiments, the length of the random polynucleotide query sequence is 7 to 15 nucleic acid residues. In some embodiments, the length of the random polynucleotide query sequence is 6 to 90 nucleic acid residues. In some embodiments, the length of the random polynucleotide query sequence is 6 to 80 nucleic acid residues. In some embodiments, the length of the random polynucleotide query sequence is 6 to 70 nucleic acid residues. In some embodiments, the length of the random polynucleotide query sequence is 6 to 60 nucleic acid residues. In some embodiments, the length of the random polynucleotide query sequence is 6 to 50 nucleic acid residues. In some embodiments, the length of the random polynucleotide query sequence is 6 to 40 nucleic acid residues.In some embodiments, the length of the random polynucleotide query sequence is 6 to 30 nucleic acid residues. In some embodiments, the length of the random polynucleotide query sequence is 6 to 20 nucleic acid residues. In some embodiments, the length of the random polynucleotide query sequence is 6 to 15 nucleic acid residues.
[0336] In some embodiments, the length of the random polynucleotide query sequence is at least 5 nucleic acid residues. In some embodiments, the length of the random polynucleotide query sequence is at least 6 nucleic acid residues. In some embodiments, the length of the random polynucleotide query sequence is at least 7 nucleic acid residues. In some embodiments, the length of the random polynucleotide query sequence is at least 8 nucleic acid residues. In some embodiments, the length of the random polynucleotide query sequence is at least 9 nucleic acid residues. In some embodiments, the length of the random polynucleotide query sequence is at least 10 nucleic acid residues. In some embodiments, the length of the random polynucleotide query sequence is at least 11 nucleic acid residues. In some embodiments, the length of the random polynucleotide query sequence is at least 12 nucleic acid residues. In some embodiments, the length of the random polynucleotide query sequence is at least 13 nucleic acid residues. In some embodiments, the length of the random polynucleotide query sequence is at least 14 nucleic acid residues. In some embodiments, the length of the random polynucleotide query sequence is at least 15 nucleic acid residues. In some embodiments, the length of the random polynucleotide query sequence is at least 16 nucleic acid residues. In some embodiments, the length of the random polynucleotide query sequence is at least 17 nucleic acid residues. In some embodiments, the length of the random polynucleotide query sequence is at least 18 nucleic acid residues. In some embodiments, the length of the random polynucleotide query sequence is at least 19 nucleic acid residues. In some embodiments, the length of the random polynucleotide query sequence is at least 20 nucleic acid residues.
[0337] In some embodiments, the length of the random polynucleotide query sequence is 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, or 100 nucleic acid residues. In some embodiments, the length of the random polynucleotide query sequence is 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, or 100 nucleic acid residues. In some embodiments, the length of the random polynucleotide query sequence is 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, or 100 nucleic acid residues.In some embodiments, the length of the random polynucleotide query sequence is 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, or 100 nucleic acid residues. In some embodiments, the length of the random polynucleotide query sequence is 5 nucleic acid residues. In some embodiments, the length of the random polynucleotide query sequence is 6 nucleic acid residues. In some embodiments, the length of the random polynucleotide query sequence is 7 nucleic acid residues. In some embodiments, the length of the random polynucleotide query sequence is 8 nucleic acid residues. In some embodiments, the length of the random polynucleotide query sequence is 9 nucleic acid residues. In some embodiments, the length of the random polynucleotide query sequence is 10 nucleic acid residues. In some embodiments, the length of the random polynucleotide query sequence is 11 nucleic acid residues. In some embodiments, the length of the random polynucleotide query sequence is 12 nucleic acid residues. In some embodiments, the length of the random polynucleotide query sequence is 13 nucleic acid residues. In some embodiments, the length of the random polynucleotide query sequence is 14 nucleic acid residues. In some embodiments, the length of the random polynucleotide query sequence is 15 nucleic acid residues. In some embodiments, the length of the random polynucleotide query sequence is 16 nucleic acid residues. In some embodiments, the length of the random polynucleotide query sequence is 17 nucleic acid residues. In some embodiments, the length of the random polynucleotide query sequence is 18 nucleic acid residues. In some embodiments, the length of the random polynucleotide query sequence is 19 nucleic acid residues. In some embodiments, the length of the random polynucleotide query sequence is 20 nucleic acid residues.
[0338] In some embodiments, a random polynucleotide query sequence library is constructed using primers having a random region. In some embodiments, the primer having a random region is a hairpin primer. In some embodiments, the random polynucleotide query sequence library is constructed using the method described in Fan et. al. 2020 (doi: https: / / doi.org / 10.1101 / 473207). In some embodiments, the random polynucleotide query sequence library is constructed using a hairpin primer that includes a random region. In some embodiments, the hairpin primer is extended by a DNA polymerase or a DNA polymerase fragment. In some embodiments, the DNA polymerase fragment is a Klenow fragment. In some embodiments, the DNA polymerase is Taq polymerase.
[0339] In some embodiments, the length of the random region of a primer or foldback primer is a 5-mer, 6-mer, 7-mer, 8-mer, 9-mer, 10-mer, 11-mer, 12-mer, 13-mer, 14-mer, 15-mer, 16-mer, 17-mer, 18-mer, 19-mer, 20-mer, 21-mer, 22-mer, 23-mer, 24-mer, 25-mer, 26-mer, 27-mer, 28-mer, 29-mer, 30-mer, 31-mer, 32-mer, 33-mer, 34-mer, 35-mer, 36-mer, 37-mer, 38-mer, 39-mer, 40-mer, 41-mer, 42-mer, 43-mer, 44-mer, or 45-mer. mer, 45mer, 46mer, 47mer, 48mer, 49mer, 50mer, 51mer, 52mer, 54mer, 55mer, 56mer, 57mer, 58mer, 59mer, 60mer, 61mer, 62mer, 63mer, 64mer, 65mer, 66mer, 67mer, 68mer, 69mer, 70mer, 81mer, 82mer, 83mer, 84mer, 85mer, 86mer, 87mer, 88mer, 89mer, 90mer, 91mer, 92mer, 93mer, 94mer, 95mer, 96mer, 97mer, 98mer, 99mer, or 100mer. In some embodiments, the length of the random region of the turn-around primer is a 10-mer. In some embodiments, the length of the random region of the turn-around primer is an 11-mer. In some embodiments, the length of the random region of the turn-around primer is a 12-mer. In some embodiments, the length of the random region of the turn-around primer is a 13-mer. In some embodiments, the length of the random region of the turn-around primer is a 14-mer. In some embodiments, the length of the random region of the turn-around primer is a 15-mer.
[0340] In some embodiments, the turn-around primer comprises or consists of the sequence ATTCCGTCTCAAGTAANNNNNNNNNNATCATGGAGACGCACTGTTTTTTTCAGTGCGTCTCCATGA (SEQ ID NO: 15342), where "N" represents any nucleic acid residue.
[0341] In some embodiments, the obtained PCR product is ligated upstream of an expression cassette (e.g., one encoding a reporter protein). In some embodiments, the PCR product obtained from the hairpin primer is digested with a restriction enzyme and ligated into an expression vector containing an expression cassette (e.g., one encoding a reporter protein).
[0342] Any suitable expression vector known in the art or described herein can be used. One of ordinary skill in the art is considered to have sufficient ability to construct an expression vector by standard recombinant techniques (e.g., see Sambrook et al., 2001 (supra) and Ausubel et al., 1996 (supra), both of which are incorporated herein by reference in their entirety) for the expression of the random polynucleotide query sequences of the present disclosure. Vectors include plasmids, cosmids, viruses (bacteriophages, animal viruses, and plant viruses), and retroviral vectors (e.g., derived from Moloney murine leukemia virus vector (MoMLV), MSCV, SFFV, MPSV, SNV, etc.), lentiviral vectors (e.g., derived from HIV-1, HIV-2, SIV, BIV, FIV, etc.), adenovirus (Ad) vectors including replication-competent, replication-deficient, and gutless types, adeno-associated virus (AAV) vectors, simian virus 40 (SV-40) vectors, bovine papillomavirus vectors, Epstein-Barr virus vectors, yeast vectors, bovine papillomavirus (BPV) vectors, herpesvirus vectors, vaccinia virus vectors, Harvey murine sarcoma virus vectors, mouse mammary tumor virus vectors, Rous sarcoma virus vectors, parvovirus vectors, poliovirus vectors, vesicular stomatitis virus vectors, Maraba virus vectors, and artificial chromosomes (e.g., YAC) such as group B adenovirus Enadenotucirev vectors, but are not limited thereto.
[0343] In some embodiments, the expression vector is the pcircGFP-BsmBI vector (Fan et. al. 2020 (doi: https: / / doi.org / 10.1101 / 473207)). In some embodiments, the restriction digest enzyme is BsmBI.
[0344] Any suitable reporter protein known in the art or described herein (e.g., a reporter protein encoded by an expression cassette) can be used. In some embodiments, cells containing the constructs of the present disclosure (e.g., an RNA polynucleotide having an IRES-like sequence and an expression cassette encoding a reporter protein) can be identified in vitro or in vivo by including a marker in the expression vector. Such markers impart an identifiable change to the cell and enable easy identification of cells containing the expression vector. Generally, a selectable marker is a marker that confers a property that enables selection. A positive selectable marker is a marker that enables selection by the presence of the marker, and a negative selectable marker is a marker that prevents selection by the presence of the marker. Examples of positive selectable markers include drug resistance markers (e.g., genes conferring resistance to neomycin, puromycin, hygromycin, DHFR, GPT, zeocin, and histidinol). Other types of markers including screenable markers such as GFP are also contemplated.
[0345] One of ordinary skill in the art is likely to know methods of using selectable markers in combination with FACS analysis. The marker used is considered unimportant as long as it can be expressed simultaneously with the nucleic acid encoding the gene product. Further examples of selectable and screenable markers are well known to those of ordinary skill in the art. In some embodiments, the reporter protein is a fluorescent protein. In some embodiments, the reporter protein is green fluorescent protein (GFP).
[0346] 8.2.2.1.b. Generation of overlapping polynucleotide sub-sequences within the query sequence In some embodiments, a population of cells is contacted with an expression plasmid library. In some embodiments, a random polynucleotide query sequence expression plasmid library is transfected into a cell line to generate a cell population. In some embodiments, the transfected cell line is a human cell line. In some embodiments, the transfected cell line is a human HEK293T cell line.
[0347] In some embodiments, the expression level of a reporter gene corresponding to each expression plasmid of the expression plasmid library is quantified. In some embodiments, a random polynucleotide query sequence expression plasmid library is transfected into a cell line to generate a cell population, and the level of reporter protein expression from a plurality of cells within the population is determined.
[0348] In some embodiments, the cells into which the plasmid library has been introduced as described above are collected and classified by reporter signal intensity (e.g., fluorescence intensity). In some embodiments, the reporter signal intensity is determined by fluorescence-activated cell sorting (FACS).
[0349] In some embodiments, the reporter signal intensities of a plurality of cells within the population are ranked. In some embodiments, the plurality of cells are separated into two or more populations. In some embodiments, the plurality of cells are separated into at least a first population and a second population. In some embodiments, the plurality of cells are divided into two populations (e.g., high and low expression of the reporter protein). In some embodiments, the plurality of cells are divided into three populations (e.g., high, medium, and low expression of the reporter protein). In some embodiments, the plurality of cells are divided into four populations (e.g., high, medium, low, and no expression of the reporter protein).
[0350] In some embodiments, the high-expression population is determined by reporter protein signal intensity. In some embodiments, the high-expression population consists of cells having reporter protein signal intensities in the top 0 to 0.1%, 0 to 0.2%, 0 to 0.3%, 0 to 0.4%, 0 to 0.5%, 0 to 0.6%, 0 to 0.7%, 0 to 0.8%, 0 to 0.9%, 0 to 1%, 0 to 1.1%, 0 to 1.2%, 0 to 1.3%, 0 to 1.4%, 0 to 1.5%, 0 to 1.6%, 0 to 1.7%, 0 to 1.8%, 0 to 1.9%, 0 to 2%, 0 to 2.1%, 0 to 2.2%, 0 to 2.3%, 0 to 2.4%, 0 to 2.5%, 0 to 2.6%, 0 to 2.7%, 0 to 2.8%, 0 to 2.9%, 0 to 3%, 0 to 3.1%, 0 to 3.2%, 0 to 3.3%, 0 to 3.4%, 0 to 3.5%, 0 to 3.6%, 0 to 3.7%, 0 to 3.8%, 0 to 3.9%, 0 to 4%, 0 to 4.1%, 0 to 4.2%, 0 to 4.3%, 0 to 4.4%, 0 to 4.5%, 0 to 4.6%, 0 to 4.7%, 0 to 4.8%, 0 to 4.9%, 0 to 5%, 0 to 5.1%, 0 to 5.2%, 0 to 5.3%, 0 to 5.4%, 0 to 5.5%, 0 to 5.6%, 0 to 5.7%, 0 to 5.8%, 0 to 5.9%, 0 to 6%, 0 to 6.1%, 0 to 6.2%, 0 to 6.3%, 0 to 6.4%, 0 to 6.5%, 0 to 6.6%, 0 to 6.7%, 0 to 6.8%, 0 to 6.9%, 0 to 7%, 0 to 7.1%, 0 to 7.2%, 0 to 7.3%, 0 to 7.4%, 0 to 7.5%, 0 to 7.6%, 0 to 7.7%, 0 to 7.8%, 0 to 7.9%, 0 to 8%, 0 to 8.1%, 0 to 8.2%, 0 to 8.3%, 0 to 8.4%, 0 to 8.5%, 0 to 8.6%, 0 to 8.7%, 0 to 8.8%, 0 to 8.9%, 0 to 9%, 0 to 9.1%, 0 to 9.2%, 0 to 9.3%, 0 to 9.4%, 0 to 9.5%, 0 to 9.6%, 0 to 9.7%, 0 to 9.8%, 0 to 9.9%, or 0 to 10% of the total population of cells ranked by reporter protein signal intensity. In some embodiments, the high-expression population consists of cells having reporter protein signal intensities in the top approximately 0.1%, approximately 0.2%, approximately 0.3%, approximately 0.4%, approximately 0.5%, approximately 0.6%, approximately 0.7%, approximately 0.8%, approximately 0.9%, approximately 1%, approximately 2%, approximately 3%, approximately 4%, approximately 5%, approximately 6%, approximately 7%, approximately 8%, approximately 9%, or approximately 10% of the total population of cells ranked by reporter protein signal intensity.In some embodiments, the high-expression population consists of cells having reporter protein signal intensities of 0 to 10% of the total population of cells ranked by reporter signal intensity.
[0351] In some embodiments, the low-expression population is determined by reporter protein signal intensity. In some embodiments, the low-expression population is the bottom 0-0.1%, 0-0.2%, 0-0.3%, 0-0.4%, 0-0.5%, 0-0.6%, 0-0.7%, 0-0.8%, 0-0.9%, 0-1%, 0-1.1%, 0-1.2%, 0-1.3%, 0-1.4%, 0-1.5%, 0-1.6%, 0-1.7%, 0-1.8%, 0-1.9%, 0-2%, 0-2.1%, 0-2.2%, 0-2.3%, 0-2.4%, 0-2.5%, 0-2.6%, 0-2.7%, 0-2.8%, 0-2.9%, 0-3%, 0-3.1%, 0-3.2%, 0-3.3%, 0-3.4%, 0-3.5%, 0-3.6%, 0-3.7%, 0-3.8%, 0-3.9%, 0-4%, 0-4.1%, 0-4.2%, 0-4.3%, 0-4.4%, 0-4.5%, 0-4.6%, 0-4.7%, 0-4.8%, 0-4.9%, 0-5%, 0-5.1%, 0-5.2%, 0-5.3%, 0-5.4%, 0-5.5%, 0-5.6%, 0-5.7%, 0-5.8%, 0-5.9%, 0-6%, 0-6.1%, 0-6.2%, 0-6.3%, 0-6.4%, 0-6.5%, 0-6.6%, 0-6.7%, 0-6.8%, 0-6.9%, 0-7%, 0-7.1%, 0-7.2%, 0-7.3%, 0-7.4%, 0-7.5%, 0-7.6%, 0-7.7%, 0-7.8%, 0-7.9%, 0-8%, 0-8.1%, 0-8.2%, 0-8.3%, 0-8.4%, 0-8.5%, 0-8.6%, 0-8.7%, 0-8.8%, 0-8.9%, 0-9%, 0-9.1%, 0-9.2%, 0-9.3%, 0-9.4%, 0-9.5%, 0-9.6%, 0-9.7%, 0-9.8%, 0-9.Comprising cells having a reporter protein signal intensity of 9%, 0-10%, 0-11%, 0-12%, 0-13%, 0-14%, 0-15%, 0-16%, 0-17%, 0-18%, 0-19%, 0-20%, 0-21%, 0-22%, 0-23%, 0-24%, 0-25%, 0-30%, 0-35%, 0-40%, 0-45%, 0-50%, 0-55%, 0-60%, 0-65%, 0-70%, 0-75%, 0-80%, 0-85%, 0-90%, 5-10%, 5-11%, 5-12%, 5-13%, 5-14%, 5-15%, 5-16%, 5-17%, 5-18%, 5-19%, 5-20%, 5-21%, 5-22%, 5-23%, 5-24%, 5-25%, 5-30%, 5-35%, 5-40%, 5-45%, 5-50%, 5-55%, 5-60%, 5-65%, 5-70%, 5-75%, 5-80%, 5-85%, or 10-90%. In some embodiments, the low-expression population consists of cells having a reporter protein signal intensity in the lower about 0.1%, about 0.2%, about 0.3%, about 0.4%, about 0.5%, about 0.6%, about 0.7%, about 0.8%, about 0.9%, about 1%, about 2%, about 3%, about 4%, about 5%, about 6%, about 7%, about 8%, about 9%, about 10%, about 15%, about 20%, about 25%, about 30%, about 35%, about 40%, about 45%, about 50%, about 55%, about 60%, about 65%, about 70%, about 75%, about 80%, about 85%, or about 90% of the total population of cells ranked by reporter signal intensity.
[0352] In some embodiments, the medium-expression population is determined by reporter protein signal intensity. In some embodiments, the low-expression population consists of cells having reporter signal intensity between about 49% - 51%, 48% - 52%, 47% - 53%, 46% - 54%, 45% - 55%, 44% - 56%, 43% - 57%, 42% - 58%, 41% - 59%, 40% - 60%, 39% - 61%, 38% - 62%, 37% - 63%, 36% - 64%, 35% - 65%, 34% - 66%, 33% - 67%, 32% - 68%, 31% - 69%, 30% - 70%, 29% - 71%, 28% - 72%, 27% - 73%, 26% - 74%, or 25% - 75% of the total population of cells ranked by reporter signal intensity. In some embodiments, the selected expression population consists of cells having reporter protein signal intensity between about 49% - 51%, 48% - 52%, 47% - 53%, 46% - 54%, 45% - 55%, 44% - 56%, 43% - 57%, 42% - 58%, 41% - 59%, 40% - 60%, 39% - 61%, 38% - 62%, 37% - 63%, 36% - 64%, 35% - 65%, 34% - 66%, 33% - 67%, 32% - 68%, 31% - 69%, 30% - 70%, 29% - 71%, 28% - 72%, 27% - 73%, 26% - 74%, or 25% - 75% of the total population of cells ranked by reporter signal intensity.
[0353] In some embodiments, the non-expressing population is determined by reporter protein signal intensity. In some embodiments, the non-expressing population consists of cells having reporter protein signal intensity in the bottom 0-5%, 0-10%, 0-15%, 0-20%, 0-25%, 0-30%, 0-35%, 0-40%, 0-45%, or 0-50% of the total population of cells ranked by reporter signal intensity. In some embodiments, the low-expressing population consists of cells having reporter protein signal intensity in the bottom approximately 0.1%, approximately 0.2%, approximately 0.3%, approximately 0.4%, approximately 0.5%, approximately 0.6%, approximately 0.7%, approximately 0.8%, approximately 0.9%, approximately 1%, approximately 2%, approximately 3%, approximately 4%, approximately 5%, approximately 6%, approximately 7%, approximately 8%, approximately 9%, approximately 10%, approximately 11%, approximately 12%, approximately 13%, approximately 14%, approximately 15%, approximately 16%, approximately 17%, approximately 18%, approximately 19%, approximately 20%, approximately 21%, approximately 22%, approximately 23%, approximately 24%, approximately 25%, approximately 30%, approximately 35%, approximately 40%, approximately 45%, or approximately 50% of the total population of cells ranked by reporter signal intensity.
[0354] In some embodiments, the high-expression population consists of cells having reporter protein signal intensities in the top 0-10% of the total population of cells ranked by reporter signal intensity, and the low-expression population consists of cells having reporter protein signal intensities in the bottom 0-90% of the total population of cells ranked by reporter signal intensity. In some embodiments, the high-expression population consists of cells having reporter protein signal intensities in the top 0-5% of the total population of cells ranked by reporter signal intensity, and the low-expression population consists of cells having reporter protein signal intensities in the bottom 0-80% of the total population of cells ranked by reporter signal intensity. In some embodiments, the high-expression population consists of cells having reporter protein signal intensities in the top 0-1% of the total population of cells ranked by reporter signal intensity, and the low-expression population consists of cells having reporter protein signal intensities in the bottom 0-75% of the total population of cells ranked by reporter signal intensity. In some embodiments, the high-expression population consists of cells having reporter protein signal intensities in the top 0-0.5% of the total population of cells ranked by reporter signal intensity, and the low-expression population consists of cells having reporter protein signal intensities in the bottom 0-75% of the total population of cells ranked by reporter signal intensity. In some embodiments, the high-expression population consists of cells having reporter protein signal intensities in the top approximately 0.5% of the total population of cells ranked by reporter signal intensity, and the low-expression population consists of cells having reporter protein signal intensities in the bottom approximately 75% of the total population of cells ranked by reporter signal intensity.
[0355] In some embodiments, the reference value for identifying a polynucleotide query sequence as a modified IRES-like polynucleotide sequence is determined according to the reporter signal intensity, the high-expression population, and / or the low-expression population.
[0356] In some embodiments, the reference value for identifying a polynucleotide query sequence as a modified IRES-like polynucleotide sequence is determined using the following formula: Reference value = (23.75 × length of the polynucleotide query sequence) - 84.85 - Z (where Z represents any number between 1 and 10,000, and X represents the length of the IRES-like polynucleotide sequence). In some embodiments, Z represents any number between 50 and 5,000. In some embodiments, Z represents any number between 175 and 2,500. In some embodiments, Z represents any number between 150 and 200.
[0357] In some embodiments, Z represents a number of at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 20, at least 30, at least 40, at least 50, at least 60, at least 70, at least 80, at least 90, at least 100, at least 200, at least 300, at least 400, at least 500, at least 600, at least 700, at least 800, at least 900, at least 1,000, at least 1,500, at least 2,000, at least 2,500, at least 3,000, at least 3,500, at least 4,000, at least 4,500, at least 5,000, at least 5,500, at least 6,000, at least 6,500, at least 7,000, at least 7,500, at least 8,000, at least 8,500, at least 9,000, at least 9,500, or at least 10,000.
[0358] In some embodiments, Z represents a number of about 1, about 2, about 3, about 4, about 5, about 6, about 7, about 8, about 9, about 10, about 20, about 30, about 40, about 50, about 60, about 70, about 80, about 90, about 100, about 200, about 300, about 400, about 500, about 600, about 700, about 800, about 900, about 1000, about 1500, about 2000, about 2500, about 3000, about 3500, about 4000, about 4500, about 5000, about 5500, about 6000, about 6500, about 7000, about 7500, about 8000, about 8500, about 9000, about 9500, or about 10000.
[0359] In some embodiments, a random polynucleotide query sequence is recovered from a sorted cell population and analyzed for enriched sequences. In some embodiments, the length of the random sequences recovered from the sorted cell population is 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, or 100 nucleic acid residues.
[0360] In some embodiments, an X-mer random polynucleotide query sequence (i.e., a random polynucleotide query sequence of X nucleic acid residues) is recovered from high-expression and low-expression populations. In some embodiments, the X-mer random polynucleotide sequences are from high-expression and non-expressing populations.
[0361] In some embodiments, a random polynucleotide query sequence of tetramers to 100-mers, tetramers to 90-mers, tetramers to 80-mers, tetramers to 70-mers, tetramers to 60-mers, tetramers to 50-mers, tetramers to 40-mers, tetramers to 30-mers, tetramers to 20-mers, tetramers to 19-mers, tetramers to 18-mers, tetramers to 17-mers, tetramers to 16-mers, tetramers to 15-mers, tetramers to 14-mers, tetramers to 13-mers, tetramers to 12-mers, tetramers to 11-mers, tetramers to 10-mers, tetramers to 9-mers, tetramers to 8-mers, tetramers to 7-mers, pentamers to 100-mers, pentamers to 90-mers, pentamers to 80-mers, pentamers to 70-mers, pentamers to 60-mers, pentamers to 50-mers, pentamers to 40-mers, pentamers to 30-mers, pentamers to 20-mers, pentamers to 19-mers, pentamers to 18-mers, pentamers to 17-mers, pentamers to 16-mers, pentamers to 15-mers, pentamers to 14-mers, pentamers to 13-mers, pentamers to 12-mers, pentamers to 11-mers, pentamers to 10-mers, pentamers to 9-mers, pentamers to 8-mers, pentamers to 7-mers, hexamers to 100-mers, hexamers to 90-mers, hexamers to 80-mers, hexamers to 70-mers, hexamers to 60-mers, hexamers to 50-mers, hexamers to 40-mers, hexamers to 30-mers, hexamers to 20-mers, hexamers to 19-mers, hexamers to 18-mers, hexamers to 17-mers, hexamers to 16-mers, hexamers to 15-mers, hexamers to 14-mers, hexamers to 13-mers, hexamers to 12-mers, hexamers to 11-mers, hexamers to 10-mers, hexamers to 9-mers, hexamers to 8-mers, hexamers to 7-mers, heptamers to 100-mers, heptamers to 90-mers, heptamers to 80-mers, heptamers to 70-mers, heptamers to 60-mers, heptamers to 50-mers, heptamers to 40-mers, heptamers to 30-mers, heptamers to 20-mers, heptamers to 19-mers, heptamers to 18-mers, heptamers to 17-mers, heptamers to 16-mers, heptamers to 15-mers, heptamers to 14-mers, heptamers to 13-mers, heptamers to 12-mers, heptamers to 11-mers, heptamers to 10-mers, heptamers to 9-mers, or heptamers to 8-mers (i.e., a random polynucleotide query sequence of X nucleic acid residues) is recovered from a high-expression population and a low-expression population.In some embodiments, the random polynucleotide sequences of tetramers to 100-mers, tetramers to 90-mers, tetramers to 80-mers, tetramers to 70-mers, tetramers to 60-mers, tetramers to 50-mers, tetramers to 40-mers, tetramers to 30-mers, tetramers to 20-mers, tetramers to 19-mers, tetramers to 18-mers, tetramers to 17-mers, tetramers to 16-mers, tetramers to 15-mers, tetramers to 14-mers, tetramers to 13-mers, tetramers to 12-mers, tetramers to 11-mers, tetramers to 10-mers, tetramers to 9-mers, tetramers to 8-mers, tetramers to 7-mers, pentamers to 100-mers, pentamers to 90-mers, pentamers to 80-mers, pentamers to 70-mers, pentamers to 60-mers, pentamers to 50-mers, pentamers to 40-mers, pentamers to 30-mers, pentamers to 20-mers, pentamers to 19-mers, pentamers to 18-mers, pentamers to 17-mers, pentamers to 16-mers, pentamers to 15-mers, pentamers to 14-mers, pentamers to 13-mers, pentamers to 12-mers, pentamers to 11-mers, pentamers to 10-mers, pentamers to 9-mers, pentamers to 8-mers, pentamers to 7-mers, hexamers to 100-mers, hexamers to 90-mers, hexamers to 80-mers, hexamers to 70-mers, hexamers to 60-mers, hexamers to 50-mers, hexamers to 40-mers, hexamers to 30-mers, hexamers to 20-mers, hexamers to 19-mers, hexamers to 18-mers, hexamers to 17-mers, hexamers to 16-mers, hexamers to 15-mers, hexamers to 14-mers, hexamers to 13-mers, hexamers to 12-mers, hexamers to 11-mers, hexamers to 10-mers, hexamers to 9-mers, hexamers to 8-mers, hexamers to 7-mers, heptamers to 100-mers, heptamers to 90-mers, heptamers to 80-mers, heptamers to 70-mers, heptamers to 60-mers, heptamers to 50-mers, heptamers to 40-mers, heptamers to 30-mers, heptamers to 20-mers, heptamers to 19-mers, heptamers to 18-mers, heptamers to 17-mers, heptamers to 16-mers, heptamers to 15-mers, heptamers to 14-mers, heptamers to 13-mers, heptamers to 12-mers, heptamers to 11-mers, heptamers to 10-mers, heptamers to 9-mers, or heptamers to 8-mers are from high-expression populations and populations with no expression.
[0362] In some embodiments, an X-mer random polynucleotide sequence (XN) is recovered from a population and extended by adding random nucleotides (r) of a vector (e.g., a DNA vector) sequence to the 5′ end and / or the 3′ end to generate a polynucleotide query sequence (e.g., 5′-r-XN-r-3′). In some embodiments, the added random nucleotides of the vector sequence are 1, 2, 3, 4, or 5 nucleic acid residues in length. In some embodiments, the random nucleotides added to each end are 1 nucleic acid residue in length (e.g., 5′-1r-XN-1r-3′). In some embodiments, the added random nucleotides are 2 nucleic acid residues in length (e.g., 5′-2r-XN-2r-3′).
[0363] As an example, in some embodiments, a 10-mer random polynucleotide sequence is recovered from a high-expression population and a low-expression population. In some embodiments, the 10-mer random polynucleotide sequence is from a high-expression population and a non-expressing population.
[0364] In some embodiments, a 10-mer random polynucleotide sequence (10N) is recovered from a population and extended by adding random nucleotides (r) of a vector (e.g., a DNA vector) sequence to the 5′ end and / or the 3′ end to generate a polynucleotide query sequence (e.g., 5′-r-10N-r-3′). In some embodiments, the added random nucleotides of the vector sequence are 1, 2, 3, 4, or 5 nucleic acid residues in length. In some embodiments, the random nucleotides added to each end are 1 nucleic acid residue in length (e.g., 5′-1r-10N-1r-3′). In some embodiments, the added random nucleotides are 2 nucleic acid residues in length (e.g., 5′-2r-10N-2r-3′).
[0365] In some embodiments, a series of overlapping polynucleotide fragment sequences (e.g., overlapping sub-sequences) are generated from a polynucleotide query sequence. In some embodiments, in the X - Y + 1 overlapping polynucleotide fragment sequences within the polynucleotide query sequence, X is the length of the polynucleotide query sequence, each polynucleotide fragment sequence consists of Y nucleic acid residues in length, the first position of each polynucleotide fragment sequence is n, the last position of the same polynucleotide fragment sequence is Y + n - 1, and n represents each positive integer between 1 and X - Y + 1.
[0366] In some embodiments, the length of the overlapping polynucleotide fragment sequence is 2 or more nucleic acid residues. In some embodiments, the length of the overlapping polynucleotide fragment sequence is 2 to 299 nucleic acid residues. In some embodiments, the length of the overlapping polynucleotide fragment sequence is 2 to 200 nucleic acid residues. In some embodiments, the length of the overlapping polynucleotide fragment sequence is 2 to 100 nucleic acid residues. In some embodiments, the length of the overlapping polynucleotide fragment sequence is 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, or 100 nucleic acid residues. In some embodiments, the length of the overlapping polynucleotide fragment sequence is 5 nucleic acid residues. In some embodiments, the length of the overlapping polynucleotide fragment sequence is 6 nucleic acid residues. In some embodiments, the length of the overlapping polynucleotide fragment sequence is 7 nucleic acid residues. In some embodiments, the length of the overlapping polynucleotide fragment sequence is 8 nucleic acid residues. In some embodiments, the length of the overlapping polynucleotide fragment sequence is 9 nucleic acid residues. In some embodiments, the length of the overlapping polynucleotide fragment sequence is 10 nucleic acid residues. In some embodiments, the length of the overlapping polynucleotide fragment sequence is 11 nucleic acid residues. In some embodiments, the length of the overlapping polynucleotide fragment sequence is 12 nucleic acid residues.
[0367] 8.2.2.1.c. Determination of Enrichment Score In some embodiments, an enrichment score for a polynucleotide fragment sequence is determined by: i) generating an expression plasmid library as described above, wherein each expression plasmid in the library comprises a different polynucleotide fragment sequence and a reporter gene; ii) contacting a population of cells with the expression plasmid library; iii) quantifying the expression level of the reporter gene corresponding to each expression plasmid in the expression plasmid library; iv) dividing the entire population of cells into a first population and a second population based on the protein expression levels of iii); and v) determining the enrichment score for the polynucleotide fragment sequence.
[0368] The enrichment score can be calculated using any suitable statistical analysis method known in the art or described herein. Exemplary methods for calculating the enrichment score include, but are not limited to, the Z-test, odds ratio, t-test, or Fisher's exact test. An exemplary method for calculating the enrichment score is described in Wang, Z. et al. (Cell 119, 831 (2004)), the contents of which are hereby incorporated by reference in their entirety.
[0369] In some embodiments, the enrichment score for each overlapping polynucleotide fragment sequence in the two populations is calculated using the Z-test, thereby generating a Z-score. In some embodiments, the enrichment score (e.g., Z-score) for each overlapping polynucleotide fragment sequence is calculated using the Z-test. In some embodiments, the Z-test formula used to calculate the Z-score is represented by the following system of equations.
Equation
Equation
[0370] In some embodiments, the overlapping polynucleotide fragment sequences are hexamers (length of 6 nucleic acid residues or hexamers). In some embodiments, f1 is the frequency of each hexamer sequence in population 1, f2 is the frequency of the same hexamer sequence in population 2, N1 is the size of population 1 (i.e., the total number of hexamers in population 1), and N2 is the size of population 2 (i.e., the total number of hexamers in population 2).
[0371] In some embodiments, the overlapping polynucleotide fragment sequences are pentamers (length of 5 nucleic acid residues or pentamers). In some embodiments, f1 is the frequency of each pentamer sequence in population 1, f2 is the frequency of the same pentamer sequence in population 2, N1 is the size of population 1 (i.e., the total number of pentamers in population 1), and N2 is the size of population 2 (i.e., the total number of pentamers in population 2).
[0372] In some embodiments, the overlapping polynucleotide fragment sequences are pentamers (length of 5 nucleic acid residues or pentamers). In some embodiments, f1 is the frequency of each pentamer sequence in population 1, f2 is the frequency of the same pentamer sequence in population 2, N1 is the size of population 1, and N2 is the size of population 2.
[0373] In some embodiments, population 1 is the high-expression population described herein. In some embodiments, population 1 is the medium-expression population described herein. In some embodiments, population 1 is the low-expression population described herein. In some embodiments, population 1 is the non-expression population described herein.
[0374] In some embodiments, population 2 is the high-expression population described herein. In some embodiments, population 2 is the medium-expression population described herein. In some embodiments, population 2 is the low-expression population described herein. In some embodiments, population 2 is the non-expression population described herein.
[0375] In some embodiments, Group 1 is the high-expression group described herein, and Group 2 is the medium-expression group described herein. In some embodiments, Group 1 is the high-expression group described herein, and Group 2 is the low-expression group described herein. In some embodiments, Group 1 is the high-expression group described herein, and Group 2 is the non-expression group described herein.
[0376] In some embodiments, Group 1 is the medium-expression group described herein, and Group 2 is the low-expression group described herein. In some embodiments, Group 1 is the medium-expression group described herein, and Group 2 is the non-expression group described herein. In some embodiments, Group 1 is the medium-expression group described herein, and Group 2 is the high-expression group described herein.
[0377] In some embodiments, Group 1 is the low-expression group described herein, and Group 2 is the non-expression group described herein. In some embodiments, Group 1 is the low-expression group described herein, and Group 2 is the medium-expression group described herein. In some embodiments, Group 1 is the low-expression group described herein, and Group 2 is the high-expression group described herein.
[0378] In some embodiments, Group 1 is the non-expression group described herein, and Group 2 is the low-expression group described herein. In some embodiments, Group 1 is the non-expression group described herein, and Group 2 is the medium-expression group described herein. In some embodiments, Group 1 is the non-expression group described herein, and Group 2 is the high-expression group described herein.
[0379] Examples of overlapping polynucleotide fragment sequences (e.g., pentamers) and enrichment scores (e.g., Z-scores) are shown in Table 1.
[0380] As will be understood by those skilled in the art, the ranking of the overlapping polynucleotide fragments described herein provides a way to design or generate a functional synthetic IRES-like sequence. Conversely, without wishing to be bound by theory, it is believed that including a polynucleotide fragment motif ranked highest within the synthetic IRES-like sequence at a high frequency increases the likelihood of IRES-like function. Polynucleotide fragment motifs ranked highest within the IRES-like sequence include, but are not limited to, SEQ ID NOs: 1, 2, 3, 4, 5, 6, 8, 9, 10, 11, 12, 13, 14, 15, 17, 22, 23, 25, 58, and 62.
[0381] Conversely, without wishing to be bound by theory, it is believed that including a polynucleotide fragment motif ranked lowest within the synthetic IRES-like sequence at a high frequency decreases the likelihood of IRES-like function.
[0382] 8.2.2.1.d. Determination of Numerical Score of Polynucleotide Sequence and Identification of IRES-like Sequence In some embodiments, the numerical score of the polynucleotide query sequence is determined by summing the enrichment scores (e.g., Z-scores) of each overlapping polynucleotide fragment sequence. In some embodiments, the numerical score of the polynucleotide query sequence is determined by averaging the enrichment scores (e.g., Z-scores) of each overlapping polynucleotide fragment sequence.
[0383] In some embodiments, the numerical score is normally distributed according to the following system of equations z = (X 数値スコア - μ) / σ = 2.06, and thus X = z × σ + μ. In some embodiments, a polynucleotide query sequence for which σ is greater than or equal to z is identified as an IRES-like sequence.
[0384] In some embodiments, a polynucleotide query sequence for which σ is less than or equal to X is identified as a non-IRES sequence.
[0385] In some embodiments, a polynucleotide query sequence having a numerical score of 0 or greater is identified as an IRES-like sequence. In some embodiments, a polynucleotide query sequence having a numerical score of 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 1100, 1200, 1300, 1400, 1500, 1600, 1700, 1800, 1900, 2000, 2100, 2200, 2300, or 2500 or greater is identified as an IRES-like sequence.
[0386] In some embodiments, a polynucleotide query sequence having a numerical score less than 0 is identified as a non-IRES sequence. In some embodiments, a polynucleotide query sequence having a numerical score of less than 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 1100, 1200, 1300, 1400, 1500, 1600, 1700, 1800, 1900, 2000, 2100, 2200, 2300, or 2500 is identified as an IRES-like sequence.
[0387] In some embodiments, a polynucleotide query sequence having a numerical score ranked in the top 5% of all polynucleotide query sequences of equal length is considered an IRES-like sequence. In some embodiments, a polynucleotide query sequence having a numerical score ranked in the top 10% of all polynucleotide query sequences of equal length is considered an IRES-like sequence. In some embodiments, a polynucleotide query sequence having a numerical score ranked in the top 15% of all polynucleotide query sequences of equal length is considered an IRES-like sequence. In some embodiments, a polynucleotide query sequence having a numerical score ranked in the top 20% of all polynucleotide query sequences of equal length is considered an IRES-like sequence. In some embodiments, a polynucleotide query sequence having a numerical score ranked in the top 25% of all polynucleotide query sequences of equal length is considered an IRES-like sequence.
[0388] In some embodiments, a polynucleotide query sequence having a numerical score ranked in the bottom 5% of all polynucleotide query sequences of equal length is considered a non-IRES-like sequence. In some embodiments, a polynucleotide query sequence having a numerical score ranked in the bottom 10% of all polynucleotide query sequences of equal length is considered a non-IRES-like sequence. In some embodiments, a polynucleotide query sequence having a numerical score ranked in the bottom 15% of all polynucleotide query sequences of equal length is considered a non-IRES-like sequence. In some embodiments, a polynucleotide query sequence having a numerical score ranked in the bottom 20% of all polynucleotide query sequences of equal length is considered a non-IRES-like sequence. In some embodiments, a polynucleotide query sequence having a numerical score ranked in the bottom 25% of all polynucleotide query sequences of equal length is considered a non-IRES-like sequence.
[0389] In some embodiments, a polynucleotide query sequence having a numerical score ranked in the top 10% of all polynucleotide query sequences of equal length is considered an IRES-like sequence. In some embodiments, a polynucleotide query sequence having a numerical score ranked in the top 5% of all polynucleotide query sequences of equal length is considered an IRES-like sequence. In some embodiments, a polynucleotide query sequence having a numerical score ranked in the top 2% of all polynucleotide query sequences of equal length is considered an IRES-like sequence. Exemplary IRES-like sequences and numerical scores are shown in Tables 2 to 96, Table 110, and Table 112. Thus, exemplary IRES-like sequences of the present disclosure include, but are not limited to, those listed in Tables 2 to 96, Table 110, and Table 112. In some embodiments, the IRES-like sequence comprises a nucleic acid sequence that is at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 100% identical to the nucleic acid sequence of the IRES-like sequence of Tables 2 to 96, Table 110, and Table 112 or a functional portion thereof.
[0390] 8.2.2.2. Method for Optimizing IRES-Like Sequences A genetic algorithm is a search algorithm that generates new subjects using random perturbations of a list of random subsets (see Schmitt, Lothar M (2001), Theory of Genetic Algorithms, Theoretical Computer Science (259), pp. 1-61). Genetic algorithms are one area of a research field called evolutionary computation, which mimics the biological processes of reproduction and natural selection to derive the "most suitable" solutions (Goldberg, D.E. (1989), Genetic Algorithms in Search, Optimization, and Machine Learning. Reading: Addison-Wesley). A genetic algorithm is an iterative optimization procedure that repeatedly applies operators (such as selection, hybridization, mutation, etc.) to a group of solutions until some convergence criterion is met. Each iteration step in which a new population is obtained is called a generation.
[0391] In some embodiments, a genetic algorithm is used to generate a top sequence pool (Figure 3, (130) and Figure 4, (230)).
[0392] In some embodiments, the genetic algorithm procedure includes the following steps as shown in Figure 3: a) randomly generate an initial population of a defined length as the current parent population (the size of the initial population is N(100)); b) use an evaluation function (also known as a fitness function, where a more fit array has a higher score) (e.g., Z-score) to evaluate the objective function of each individual array in the initial population (110); c) select top sequences to form a top sequence pool (size = D) (130); d) use a selection operator to select some of the sequences in the top sequence pool (the more a sequence fits, the higher the probability of being selected) (120); and e) generate an offspring population by a hybridization operator (121) and / or a mutation operator (122), and repeat steps b) - e).
[0393] In some embodiments, the genetic algorithm procedure includes the following steps as shown in FIG. 4: a) randomly generating an initial population of a defined length as the current parent population (the size of the initial population is N (200)); b) using an evaluation function (also known as a fitness function, where a more fit array has a higher score) (e.g., Z - score) to evaluate the objective function of each individual array within the initial population (210); c) selecting top sequences to form a top sequence pool (size = D) (230); d) using a selection operator to select some of the sequences within the top sequence pool (the more a sequence fits, the higher the likelihood of being selected) (220); and e) generating an offspring population by means of a crossover operator (221) and / or a mutation operator (222), and repeating steps b) - e).
[0394] In some embodiments, the initial population has a defined length of 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, or 100 nucleic acid residues. In some embodiments, the initial population has a defined length of between 5 and 12 nucleic acid residues. In some embodiments, the initial population has a defined length of between 12 and 20 nucleic acid residues. In some embodiments, the initial population has a defined length of between 20 and 100 nucleic acid residues. In some embodiments, the initial population has a defined length of between 100 and 200 nucleic acid residues. In some embodiments, the initial population has a defined length of between 200 and 300 nucleic acid residues. In some embodiments, the initial population has a defined length of between 300 and 400 nucleic acid residues. In some embodiments, the initial population has a defined length of between 400 and 500 nucleic acid residues. In some embodiments, the initial population has a defined length of between 500 and 600 nucleic acid residues. In some embodiments, the initial population has a defined length of between 600 and 700 nucleic acid residues. In some embodiments, the initial population has a defined length of between 500 and 600 nucleic acid residues. In some embodiments, the initial population has a defined length of between 600 and 700 nucleic acid residues. In some embodiments, the initial population has a defined length of between 700 and 800 nucleic acid residues. In some embodiments, the initial population has a defined length of between 800 and 900 nucleic acid residues. In some embodiments, the initial population has a defined length of between 900 and 1000 nucleic acid residues.
[0395] In some embodiments, the genetic algorithm is iterated until the top sequence pool stabilizes and no longer changes over many generations. In some embodiments, the genetic algorithm is iterated until the top sequence pool stabilizes and no longer changes over 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, or 100 generations.
[0396] In some embodiments, the genetic algorithm is iterated until the top sequence pool stabilizes and no longer changes over 20 generations.
[0397] In some embodiments, the processes of selection, crossover, and mutation are continued until the number of offspring is the same as the initial population, such that the offspring generation consists entirely of new offspring and the parent generation is completely replaced.
[0398] In some embodiments, the processes of selection, hybridization, and mutation allow for the high-scoring sequences of the parent generation to survive into the offspring generation.
[0399] In some embodiments, a genetic algorithm is used to select the top-ranked sequences.
[0400] In some embodiments, a genetic algorithm is used to select the lowest-ranked sequences.
[0401] In some embodiments, hybridization and mutation are applied to the sequence pool of the current generation to create the next generation. In some embodiments, random addition or deletion is applied to the sequence pool of the current generation to create the next generation. In the hybridization step, the functions of a subset of pairs are combined to form a new subset. In a preferred embodiment, the hybridization is the same operator as crossover.
[0402] In a preferred embodiment, a race search that uses a t-test to determine the probability that a subset is at least a user-specified threshold better than the current best subset is suitable.
[0403] In some embodiments, hybrids and mutations are applied to the sequence pool of the current generation to create the next generation. In some embodiments, random additions or deletions are applied to the sequence pool of the current generation to create the next generation. In the hybrid process, the functions of pairs of subsets are combined to form new subsets. In a preferred embodiment, the hybrid is the same operator as crossover.
[0404] In a preferred embodiment, a race search that uses a t-test to determine the probability that a subset is at least a user-specified threshold better than the current best subset is suitable.
[0405] 8.3. Therapeutic agent In some embodiments, the polynucleotides of the present disclosure encode a therapeutic agent.
[0406] In some embodiments, the therapeutic agent is a polypeptide, protein, enzyme, antibody, or a combination thereof. In some embodiments, the therapeutic agent comprises one or more polypeptides, proteins, enzymes, antibodies, or a combination thereof.
[0407] In some embodiments, the therapeutic agent is a protein or enzyme. In some embodiments, the protein or enzyme is associated with a genetic disease (e.g., a disease in which a change in a gene (e.g., a mutation) and / or abnormal regulation of a protein plays a role in the onset, development, and / or manifestation of the disease).
[0408] In some embodiments, the therapeutic agent is a polypeptide or a protein. In some embodiments, the polypeptide or protein resembles an attenuated or non-viable form of an agent that causes a disease (e.g., an infectious agent such as a pathogen) that can be selected from one or more components of microorganisms such as bacteria, viruses, fungi, parasites, or toxins, proteins (e.g., surface proteins), and / or cell walls. In some embodiments, the therapeutic agent is an antigen or an agent that can stimulate the body's immune system to recognize the antigen or agent, generate antibodies against the antigen or agent, destroy the antigen or agent, and / or develop immunological memory of the antigen or agent. In some embodiments, the therapeutic agent is an antigen or an agent that induces and / or enhances vaccine-induced memory and / or enables the immune system to respond quickly and effectively to the antigen or agent in subsequent encounters.
[0409] In some embodiments, the therapeutic agent is derived from an infectious agent or a part, component, and / or product thereof (e.g., cell wall, genomic sequence, membrane, capsid, protein, lipid, glycan, toxin). In some embodiments, the infectious agent is selected from viruses, bacteria, fungi, protozoa, and helminths. In some embodiments, the infectious agents selected are viruses and bacteria.
[0410] In some embodiments, the infectious agent is a virus selected from the group consisting of adenovirus, herpes simplex type 1, herpes simplex type 2, encephalitis virus, papillomavirus, varicella-zoster virus, Epstein-Barr virus, human cytomegalovirus, human herpesvirus 8, BK virus, JC virus, smallpox, poliovirus, human bocavirus, parvovirus B19, human astrovirus, norovirus, coxsackievirus, hepatitis A virus, hepatitis B virus, hepatitis C virus, hepatitis D virus, hepatitis E virus, rhinovirus, severe acute respiratory syndrome (SARS) virus, yellow fever virus, dengue virus, West Nile virus, rubella virus, human immunodeficiency virus (HIV), influenza virus, Guanarito virus, Junin virus, Lassa fever virus, Machupo virus, Sabia virus, Crimean-Congo hemorrhagic fever virus, Ebola virus, Marburg virus, measles virus, mumps virus, parainfluenza virus, respiratory syncytial virus (RSV), human metapneumovirus, Hendra virus, Nipah virus, rabies virus, rotavirus, orbivirus, coltivirus, bunyavirus, human enterovirus, hantavirus, West Nile virus, coronavirus, SARS-related coronavirus (SARS-CoV), SARS-CoV-2 virus (COVID-19-related), Middle East respiratory syndrome coronavirus, Japanese encephalitis virus, vesicular exanthema virus, and eastern equine encephalitis.
[0411] In some embodiments, the infectious agent is a bacterium selected from Mycobacterium tuberculosis, Clostridium difficile resistant to clindamycin, Clostridium difficile resistant to fluoroquinolone, methicillin-resistant Staphylococcus aureus (MRSA), multidrug-resistant Enterococcus faecalis, multidrug-resistant Enterococcus faecium, multidrug-resistant Pseudomonas aeruginosa, multidrug-resistant Acinetobacter baumannii, and vancomycin-resistant Staphylococcus aureus (VRSA).
[0412] In some embodiments, the infectious agent is associated with humans, non-human primates, or other animals such as birds, pigs, horses, dogs, cats, rabbits, mice, rats, cows, sheep, goats, and deer.
[0413] In some embodiments, the therapeutic agent is an antibody. In some embodiments, the antibodies include monoclonal antibodies, polyclonal antibodies, recombinant antibodies, human antibodies, humanized antibodies, chimeric antibodies, synthetic antibodies, tetrameric antibodies comprising two heavy chain molecules and two light chain molecules, antibody light chain monomers, antibody heavy chain monomers, antibody light chain dimers, antibody heavy chains, antibody heavy chain dimers, antibody light chain-heavy chain pairs, intrabodies, heteroconjugate antibodies, monovalent antibodies, antigen-binding fragments of full-length antibodies, and fusion proteins as described above, but are not limited thereto. The antigen-binding fragments include, but are not limited to, single domain antibodies (variable domains of heavy chain antibodies (VHH) or nanobodies), Fabs, Fab’s, F(ab’)2s, Fds, Fvs, scFvs (single-chain variable fragments).
[0414] 8.3.1. Method for producing therapeutic agent The present disclosure further provides a method for producing a protein intracellularly, the method comprising contacting a cell with an RNA polynucleotide (e.g., a circular RNA molecule) described herein, a DNA vector encoding or suitable for synthesizing the RNA polynucleotide, whereby a nucleic acid sequence encoding the expressed protein is translated and the protein is produced intracellularly.
[0415] In some embodiments, the method for producing a protein intracellularly comprises contacting the cell with an RNA polynucleotide (e.g., circular RNA) sequence described herein or a vector comprising an RNA polynucleotide (e.g., circular RNA) sequence under conditions where the protein-coding nucleic acid sequence is translated and the protein is produced intracellularly. Also provided are proteins produced by the disclosed methods.
[0416] In some embodiments, protein production is tissue-specific. For example, the protein can be selectively produced in one or more tissues of muscle, liver, kidney, brain, lung, skin, pancreas, blood, or heart.
[0417] In some embodiments, the protein is recursively expressed intracellularly.
[0418] In some embodiments, the protein is produced intracellularly at least about 10%, at least about 20%, or at least about 30% longer than when the protein-coding nucleic acid sequence is provided to the cell using a viral vector encoding linear RNA or as linear RNA.
[0419] In some embodiments, the protein is produced intracellularly at a level that is at least about 10%, at least about 20%, or at least about 30% higher than when the protein-coding nucleic acid sequence is provided to the cell using a viral vector or as linear RNA.
[0420] In some embodiments, the RNA polynucleotides (e.g., circular RNAs) described herein (including its components such as IRES-like sequences) promote cap-independent translation activity from circRNAs. Standard translation by cap-independent mechanisms may be reduced in some human diseases. Thus, the use of circRNAs for expressing proteins may be particularly useful for treating such diseases. In some embodiments, the use of the circRNAs described herein promotes cap-independent translation activity from circRNAs under conditions where cap-dependent translation is reduced or turned off intracellularly.
[0421] As described above, when the IRES is in-frame with the protein-coding nucleic acid sequence and there is no stop codon in the protein-coding sequence, translation of the protein-coding nucleic acid sequence can occur in an infinite loop (i.e., recursively). Thus, in some embodiments, a method of producing a protein intracellularly produces a linked protein.
[0422] Any prokaryotic or eukaryotic host cell described herein may be contacted with a recombinant circRNA molecule or a vector containing a circRNA molecule. The host cell may be a mammalian cell such as a human cell. In some embodiments, the cell is in vivo. In some embodiments, the cell is in vitro. In some embodiments, the cell is ex vivo. In some embodiments, the cell is of a mammal such as a human.
[0423] In some embodiments, regardless of the selected cell type, 5'-cap-dependent translation intracellularly is impaired (e.g., reduced, diminished, inhibited, or completely absent). In some embodiments, substantially no 5'-cap-dependent translation occurs intracellularly.
[0424] A recombinant circular RNA molecule, a DNA molecule encoding the same, or a vector containing the same can be introduced into cells by any method, such as transfection, transformation, or transduction. The terms "transfection", "transformation", and "transduction" are used interchangeably herein and refer to introducing one or more exogenous polynucleotides into a host cell using a physical or chemical method. Many transfection techniques are known in the art, such as calcium phosphate DNA coprecipitation (e.g., Murray E. J. (ed.), Methods in Molecular Biology, Vol. 7, Gene Transfer and Expression Protocols, Humana Press (1991)); DEAE-dextran; electroporation; transfection via cationic liposomes; particle bombardment with tungsten particles (Johnston, Nature, 346: 776-777 (1990)); strontium phosphate DNA coprecipitation (Brash et al., Mol. Cell. Biol., 7: 2031-2034 (1987); and gene delivery using magnetic nanoparticles (see Dobson, J., Gene Ther, 13 (4): 283-7 (2006)).
[0425] 8.3.2. Post-translational modification of therapeutic agents In some embodiments, the therapeutic agent (e.g., a protein) is post-translationally modified. In some embodiments, the post-translational modification is specific to the cell type, tissue type, or organ in which the therapeutic agent is produced or delivered.
[0426] In some embodiments, the post-translational modification is acetylation, SUMOylation, glycosylation, tyrosine sulfation, phosphorylation, ADP-ribosylation, prenylation, myristoylation, palmitoylation, ubiquitination, centrinylation, and / or ubiquitin-like protein modification.
[0427] In some embodiments, the therapeutic agent is modified after translation in vitro, ex vivo, or in vivo.
[0428] In some embodiments, the therapeutic agent is modified after translation when expressed in cells (e.g., primary cells, cell lines derived from primary cells, immortalized cells). In some embodiments, the cells are immortalized cells (e.g., human immortalized cells).
[0429] In some embodiments, the therapeutic agent is modified after translation in a living body (e.g., bacteria, animals).
[0430] In some embodiments, the therapeutic agent is modified after translation in an animal (e.g., the animals described herein). In some embodiments, the therapeutic agent is modified after translation in a human.
[0431] In some embodiments, the post-translational modification can be detected by techniques known in the art, including gel electrophoresis, Western blotting, Eastern blotting, immunoprecipitation, mass spectrometry, chromatography, and flow cytometry. Analysis of the modified protein is usually performed by electrophoresis and autoradiography, and specificity is improved by immunoprecipitating the protein of interest before electrophoresis.
[0432] In some embodiments, the detection of the post-translational modification involves in vivo labeling of a cellular substrate pool with a radioactive substrate or a substrate precursor molecule, such that a radioactive labeled moiety including, but not limited to, phosphate, fatty acid acyl (e.g., myristoyl or palmitoyl), centrin, methyl, acetyl, hydroxyl, iodine, flavin, ubiquitin, or ADP-ribose is incorporated into the therapeutic agent. In some embodiments, the detection of the post-translational modification may include enzymatically incorporating a labeled moiety into the therapeutic agent in vitro to estimate the state of modification in vivo. In some embodiments, the labeled moiety includes, but is not limited to, a radioisotope, a luminescent enzyme (e.g., horseradish peroxidase, HRP), or a fluorescent label.
[0433] In some embodiments, the post-translational modification can be detected by analyzing a change in the electrophoretic mobility of the modified therapeutic agent compared to the unmodified therapeutic agent.
[0434] In some embodiments, the post-translational modification can be detected by thin layer chromatography of radioactive labeled fatty acids extracted from the therapeutic agent.
[0435] In some embodiments, the post-translational modification can be detected by partitioning the therapeutic agent into a detergent-rich layer or a detergent layer by phase separation, and by the effect of enzymatic treatment of the therapeutic agent on the partitioning between an aqueous environment and a detergent-rich environment.
[0436] 8.4. RNA Polynucleotide In some embodiments, an RNA polynucleotide is provided that includes a modified translation initiation element (TI) that includes an IRES-like polynucleotide sequence.
[0437] In some embodiments, herein, Formula I: 5'-(A1) 0-1 -(L) n -TI-(L) n -Z1-(L) n -(B1) 0-1An RNA polynucleotide comprising a construct of -3'(I) is provided.
[0438] In some embodiments, in Formula I, TI is a modified translation initiation element comprising an IRES-like polynucleotide sequence, Z1 is an expression sequence encoding a therapeutic agent, each L is independently a linker sequence, and n is a positive integer (e.g., an integer selected from 0-2).
[0439] In some embodiments, in Formula I, A1 and B1 are each independently a sequence capable of circularizing the RNA polynucleotide. In some embodiments, in Formula I, A1 and B1 are a pair of homologous sequences capable of spontaneous cleavage and circularization of the RNA polynucleotide.
[0440] In some embodiments, in Formula I, A1 and B1 each independently comprise a nucleotide derivative capable of joining the 5' end and the 3' end via a 3' to 5' phosphodiester bond to circularize the RNA polynucleotide.
[0441] In some embodiments, in Formula I, A1 and B1 each independently comprise a nucleotide triphosphate derivative capable of joining the 5' end and the 3' end via a 3' to 5' phosphodiester bond to circularize the RNA polynucleotide.
[0442] In some embodiments, herein, Formula II: 5'-(A1) 0-1 -(L) n -Z1 B -(L) n -TI-(L) n -Z1 A -(L) n -(B1) 0-1 An RNA polynucleotide comprising a construct of -3'(II) is provided.
[0443] In some embodiments, in Formula II, TI is a modified translation initiation element comprising an IRES-like polynucleotide sequence, Z1 Ais the first part of the expression sequence encoding the therapeutic agent, Z1 B is the second part of the expression sequence encoding the therapeutic agent, each L is independently a linker sequence, and n is a positive integer (e.g., an integer selected from 0 to 2).
[0444] In some embodiments, in Formula II, A1 and B1 are each independently a sequence capable of circularizing the RNA polynucleotide. In some embodiments, in Formula I, A1 and B1 are a pair of homologous sequences capable of spontaneous cleavage and circularization of the RNA polynucleotide.
[0445] In some embodiments, in Formula II, A1 and B1 each independently include a nucleotide derivative capable of joining the 5' end and the 3' end via a 3' to 5' phosphodiester bond to circularize the RNA polynucleotide.
[0446] In some embodiments, in Formula II, A1 and B1 each independently include a nucleotide triphosphate derivative capable of joining the 5' end and the 3' end via a 3' to 5' phosphodiester bond to circularize the RNA polynucleotide.
[0447] In some embodiments, herein, Formula III: 5'-(3' intron fragment)-(E2)-(L) n -TI-(L) n -Z1-(L) n -(E1)-(5' intron fragment)-3' (III) is provided, an RNA polynucleotide comprising the construct.
[0448] In some embodiments, in Formula III, TI is a modified translation initiation element comprising an IRES-like polynucleotide sequence, Z1 is an expression sequence encoding a therapeutic agent, each L is independently a linker sequence, and n is a positive integer (e.g., an integer selected from 0 to 2).
[0449] In some embodiments, in Formula III, the 5' intron fragment and the 3' intron fragment are each a fragment of a Group II intron, and the 5' intron fragment is located on the 5' side of the 3' intron fragment within the Group II intron.
[0450] In some embodiments, in Formula III, E1 is a 5' flanking exon fragment of a Group II intron having a length of ≧0 nucleotides.
[0451] In some embodiments, in Formula III, E2 is a 3' flanking exon fragment of a Group II intron having a length of ≧0 nucleotides.
[0452] In some embodiments, in Formula III, E1 is a 5' flanking exon fragment of a Group II intron having a length of ≧0 nucleotides, and E2 is a 3' flanking exon fragment of a Group II intron having a length of ≧0 nucleotides.
[0453] In some embodiments, herein, a RNA polynucleotide comprising a construct of Formula IV: 5'-(3' intron fragment)-(E2)-(L) n -Z1 B -(L) n -TI-(L) n -Z1 A- -(L) n -(E1)-(5' intron fragment)-3' (IV) is provided.
[0454] In some embodiments, in Formula IV, TI is a modified translation initiation element comprising an IRES-like polynucleotide sequence, and Z1 A is the first part of an expression sequence encoding a therapeutic agent, and Z1 B is the second part of an expression sequence encoding a therapeutic agent, each L is independently a linker sequence, and n is a positive integer (e.g., an integer selected from 0 to 2).
[0455] In some embodiments, in Formula IV, the 5'intron fragment and the 3'intron fragment are each a fragment of a Group II intron, and the 5'intron fragment is located on the 5'side of the 3'intron fragment within the Group II intron.
[0456] In some embodiments, in Formula IV, E1 is a 5'adjacent exon fragment of a Group II intron having a length of ≧0 nucleotides.
[0457] In some embodiments, in Formula IV, E2 is a 3'adjacent exon fragment of a Group II intron having a length of ≧0 nucleotides.
[0458] In some embodiments, in Formula IV, E1 is a 5'adjacent exon fragment of a Group II intron having a length of ≧0 nucleotides, and E2 is a 3'adjacent exon fragment of a Group II intron having a length of ≧0 nucleotides.
[0459] In some embodiments, Formula V: 5'-TI-(L) n -Z1-3'(V) provides an RNA polynucleotide comprising a construct.
[0460] In some embodiments, in Formula V, TI is a modified translation initiation element comprising an IRES-like polynucleotide sequence, Z1 is an expression sequence encoding a therapeutic agent, each L is independently a linker sequence, and n is a positive integer (eg, an integer selected from 0 to 2).
[0461] In some embodiments, in Formula III or IV, the RNA polynucleotide further comprises a 5'homology arm at the 5'terminus of the 3'intron fragment.
[0462] In some embodiments, in Formula III or IV, the RNA polynucleotide further comprises a 3'homology arm at the 3'terminus of the 5'intron fragment.
[0463] In some embodiments, in Formula III or IV, the RNA polynucleotide further comprises a 5' homology arm at the 5' end of the 3' intron fragment and a 3' homology arm at the 3' end of the 5' intron fragment.
[0464] In some embodiments, in Formula III or IV, E1 and E2 are each independently 0 to 20 nucleotides in length.
[0465] In some embodiments, in Formula III or IV, the 5' intron fragment and the 3' intron fragment are obtained by splitting a Group II intron into two fragments at an unpaired region, and the unpaired region is preferably selected from a linear region between two adjacent domains of the Group II intron and a loop region of the stem-loop structure of Domain 4 of the Group II intron.
[0466] In some embodiments, examples of the 5' homology arm of the present disclosure include, but are not limited to, those listed in Table 109.
[0467] In some embodiments, the 5' homology arm of the present disclosure comprises a nucleic acid sequence that is at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 100% identical to the 5' homology arm sequence of Table 109 or a functional portion thereof. In some embodiments, the 5' homology arm of the present disclosure comprises a combination of a plurality (e.g., two, three, or four) of nucleic acid sequences that are at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 100% identical to the 5' homology arm sequence of Table 109 or a functional portion thereof. In some embodiments, the 5' homology arm of the present disclosure comprises a nucleic acid sequence selected from Table 109 or a functional portion thereof. In some embodiments, the 5' homology arm of the present disclosure comprises a combination of a plurality (e.g., two, three, or four) of nucleic acid sequences selected from Table 109 or a functional portion thereof.
[0468] In some embodiments, the 5' homology arm of the present disclosure is a nucleic acid sequence that is at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 100% identical to the 5' homology arm sequence of Table 109 or a functional portion thereof. In some embodiments, the 5' homology arm of the present disclosure is a combination of a plurality (e.g., two, three, or four) of nucleic acid sequences that are at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 100% identical to the 5' homology arm sequence of Table 109 or a functional portion thereof. In some embodiments, the 5' homology arm of the present disclosure is a nucleic acid sequence selected from Table 109 or a functional portion thereof. In some embodiments, the 5' homology arm of the present disclosure is a combination of a plurality (e.g., two, three, or four) of nucleic acid sequences or functional portions thereof selected from Table 109.
[0469] In some embodiments, examples of the 3' homology arm of the present disclosure include, but are not limited to, those listed in Table 109.
[0470] In some embodiments, the 3' homology arm of the present disclosure comprises a nucleic acid sequence that is at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 100% identical to the 3' homology arm sequence of Table 109 or a functional portion thereof. In some embodiments, the 3' homology arm of the present disclosure comprises a combination of a plurality (e.g., two, three, or four) of nucleic acid sequences that are at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 100% identical to the 3' homology arm sequence of Table 109 or a functional portion thereof. In some embodiments, the 3' homology arm of the present disclosure comprises a nucleic acid sequence selected from Table 109 or a functional portion thereof. In some embodiments, the 3' homology arm of the present disclosure comprises a combination of a plurality (e.g., two, three, or four) of nucleic acid sequences selected from Table 109 or a functional portion thereof.
[0471] In some embodiments, the 3' homology arm of the present disclosure is a nucleic acid sequence that is at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 100% identical to the 3' homology arm sequence of Table 109 or a functional portion thereof. In some embodiments, the 3' homology arm of the present disclosure is a combination of a plurality (e.g., two, three, or four) of nucleic acid sequences that are at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 100% identical to the 3' homology arm sequence of Table 109 or a functional portion thereof. In some embodiments, the 3' homology arm of the present disclosure is a nucleic acid sequence selected from Table 109 or a functional portion thereof. In some embodiments, the 3' homology arm of the present disclosure is a combination of a plurality (e.g., two, three, or four) of nucleic acid sequences selected from Table 109 or a functional portion thereof.
[0472] In some embodiments, examples of the E1 sequence of the present disclosure include, but are not limited to, those listed in Table 103. In some embodiments, the E1 sequence comprises a nucleic acid sequence that is at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 100% identical to the nucleic acid sequence of the E1 sequence of Table 103 or a functional portion thereof. In some embodiments, the E1 sequence is a nucleic acid sequence that is at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 100% identical to the nucleic acid sequence of the E1 sequence of Table 103 or a functional portion thereof.
[0473] In some embodiments, the E2 sequences of the present disclosure include, but are not limited to, those listed in Table 104. In some embodiments, the E2 sequence comprises a nucleic acid sequence that is at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 100% identical to the nucleic acid sequence of the E2 sequence in Table 104 or a functional portion thereof. In some embodiments, the E2 sequence is a nucleic acid sequence that is at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 100% identical to the nucleic acid sequence of the E2 sequence in Table 104 or a functional portion thereof.
[0474] In some embodiments, the Group II intron sequences of the present disclosure include, but are not limited to, those listed in Table 105. In some embodiments, the Group II intron sequence comprises a nucleic acid sequence that is at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 100% identical to the nucleic acid sequence of the Group II intron sequence in Table 105 or a functional portion thereof. In some embodiments, the Group II intron sequence is a nucleic acid sequence that is at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 100% identical to the nucleic acid sequence of the Group II intron sequence in Table 105 or a functional portion thereof.
[0475] In some embodiments, the 3' intron fragment sequences of the present disclosure include, but are not limited to, those listed in Table 106. In some embodiments, the 3' intron fragment sequence comprises a nucleic acid sequence that is at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 100% identical to the nucleic acid sequence of the 3' intron fragment sequence of Table 106 or a functional portion thereof. In some embodiments, the 3' intron fragment sequence is a nucleic acid sequence that is at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 100% identical to the nucleic acid sequence of the 3' intron fragment sequence of Table 106 or a functional portion thereof.
[0476] In some embodiments, the 5' intron fragment sequences of the present disclosure include, but are not limited to, those listed in Table 107. In some embodiments, the 5' intron fragment sequence comprises a nucleic acid sequence that is at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 100% identical to the nucleic acid sequence of the 5' intron fragment sequence of Table 107 or a functional portion thereof. In some embodiments, the 5' intron fragment sequence is a nucleic acid sequence that is at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 100% identical to the nucleic acid sequence of the 5' intron fragment sequence of Table 107 or a functional portion thereof.
[0477] In some embodiments, in any one of the formulas described herein, n is an integer selected from 0 to 5. In some embodiments, n is an integer selected from 0 to 2. In some embodiments, n is 2. In some embodiments, n is 1. In some embodiments, n is 0.
[0478] In some embodiments, in any one of the formulas described herein, TI further comprises a second IRES-like polynucleotide sequence. In some embodiments, the first IRES-like polynucleotide sequence and the second IRES-like polynucleotide sequence are the same. In some embodiments, the first IRES-like polynucleotide sequence and the second IRES-like polynucleotide sequence are different.
[0479] In some embodiments, in any one of the formulas described herein, the IRES-like polynucleotide sequence is determined by any of the methods described herein. In some embodiments, the IRES-like polynucleotide sequence is about 90%, 95%, 97%, 98%, 99%, or 100% identical or at least about 90%, 95%, 97%, 98%, 99%, or 100% identical to a nucleic acid sequence of SEQ ID NO: 1025-14161 or SEQ ID NO: 14412-15341. In embodiments, the IRES-like polynucleotide is about 90%, 95%, 97%, 98%, 99%, or 100% identical or at least about 90%, 95%, 97%, 98%, 99%, or 100% identical to a nucleic acid sequence of SEQ ID NO: 1, 2, 3, 4, 5, 6, 8, 9, 10, 11, 12, 13, 14, 15, 17, 22, 23, 25, 58, or 62.
[0480] In some embodiments, in any of the formulas described herein, the length of the IRES-like sequence is 3 residues or more. In some embodiments, the length of the IRES-like sequence is from 3 to 300 nucleic acid residues. In some embodiments, the length of the IRES-like sequence is from 3 to 200 nucleic acid residues. In some embodiments, the length of the IRES-like sequence is 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, or 100 nucleic acid residues. In some embodiments, the length of the IRES-like sequence is 5 nucleic acid residues. In some embodiments, the length of the IRES-like sequence is 6 nucleic acid residues. In some embodiments, the length of the IRES-like sequence is 7 nucleic acid residues. In some embodiments, the length of the IRES-like sequence is 8 nucleic acid residues. In some embodiments, the length of the IRES-like sequence is 9 nucleic acid residues. In some embodiments, the length of the IRES-like sequence is 10 nucleic acid residues. In some embodiments, the length of the IRES-like sequence is 11 nucleic acid residues. In some embodiments, the length of the IRES-like sequence is 12 nucleic acid residues.
[0481] In some embodiments, in any of the formulas described herein, TI (i.e., the modified translation initiation element containing the IRES-like polynucleotide sequence) further comprises an IRES sequence isolated or derived from a native IRES sequence. In some embodiments, the native IRES sequence comprises a nucleic acid sequence that is at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identical to any of the native IRES sequences in Tables 98-100 or a functional portion thereof.
[0482] 8.4.1. Expression sequence In some embodiments, the RNA polynucleotides provided herein include an expression sequence encoding a therapeutic agent (Z1). In some embodiments, the expression sequence encodes a reporter protein. In some embodiments, the expression sequence encodes a therapeutic protein. Exemplary expression sequences include, but are not limited to, those listed in Tables 101 and 102.
[0483] In some embodiments, the RNA polynucleotide includes one expression sequence. In some embodiments, the polynucleotide includes multiple expression sequences, e.g., 2, 3, 4, or 5 expression sequences.
[0484] In some embodiments, the RNA polynucleotide encodes a protein composed of subunits encoded by multiple genes. For example, the protein is a heterodimer, and each strand or subunit of the protein is encoded by a separate gene. It is also possible for multiple RNA polynucleotides to be delivered by a transport vehicle, with each RNA polynucleotide encoding a separate subunit of the protein. Alternatively, a single RNA polynucleotide may be designed to encode multiple subunits. In some embodiments, separate RNA polynucleotide molecules encoding individual subunits may be administered by separate transport vehicles.
[0485] 8.4.2. Linker In some embodiments, the RNA polynucleotides provided herein include a linker sequence (L). In some embodiments, the length of the linker sequence is from 3 to 300 nucleic acid residues. In some embodiments, the linker sequence is a nucleic acid sequence having a length of about 3 to 10, 10 to 20, 20 to 30, 30 to 40, 40 to 50, 50 to 60, 60 to 70, 70 to 80, 80 to 90, 90 to 100, 100 to 125, 125 to 150, 150 to 175, 175 to 200, 200 to 225, 225 to 250, 250 to 275, or 275 to 300. In some embodiments, the length of the linker sequence is about 3N nucleic acid residues, where N is an integer selected from 1 to 100. In some embodiments, the length of the linker sequence is about 3N nucleic acid residues, where N is an integer selected from 1 to 50. In some embodiments, the length of the linker sequence is about 3N nucleic acid residues, where N is an integer selected from 1 to 20. In some embodiments, the length of the linker sequence is about 3N nucleic acid residues, where N is an integer selected from 1 to 10. In some embodiments, the length of the linker sequence is about 3N nucleic acid residues, where N is 1, 2, 3, 4, or 5. In some embodiments, the length of the linker sequence is about 3N nucleic acid residues, where N is 1, 2, or 3. In some embodiments, the length of the linker sequence is about 3 nucleic acid residues.
[0486] In some embodiments, the linker sequence includes a nucleic acid sequence of RCC, where R is guanine or adenine. In some embodiments, the linker includes a nucleic acid sequence of RCCRCC, where R is guanine or adenine. In some embodiments, the linker includes a nucleic acid sequence of RCCRCCRCC, where R is guanine or adenine.
[0487] In some embodiments, the linker comprises a nucleic acid sequence encoding a 5’UTR, 3’UTR, polyA sequence, polyA-C sequence, polyC sequence, polyU sequence, polyG sequence, ribosome binding site, aptamer, riboswitch, ribozyme, small RNA binding site, translation regulatory element (e.g., Kozak sequence), protein binding site (e.g., PTBP1 or HUR), unnatural nucleotide, or non-nucleotide chemical linker sequence.
[0488] In some embodiments, the RNA polynucleotide provided herein comprises a 3’UTR. In some embodiments, the 3’UTR is derived from human β-globin, human α-globin, Xenopus laevis β-globin, Xenopus laevis α-globin, human prolactin, human GAP-43, human eEFlal, human Tau, human TNFa, dengue virus, hantavirus small mRNA, bunyavirus small mRNA, turnip yellow mosaic virus, hepatitis C virus, rubella virus, tobacco mosaic virus, human IL-8, human actin, human GAPDH, human tubulin, hibiscus chlorotic ringspot virus, woodchuck hepatitis virus post-transcriptional regulatory element, sindbis virus, turnip crinkle virus, tobacco etch virus, or Venezuelan equine encephalitis virus.
[0489] In some embodiments, the RNA polynucleotide provided herein comprises a 5’UTR. In some embodiments, the 5’UTR is derived from human β-globin, Xenopus laevis β-globin, human α-globin, Xenopus laevis α-globin, rubella virus, tobacco mosaic virus, mouse Gtx, dengue virus, heat shock protein 70 kDa protein 1A, tobacco alcohol dehydrogenase, tobacco etch virus, turnip crinkle virus, or adenovirus tripartite leader.
[0490] In some embodiments, the RNA polynucleotide provided herein comprises a polyA region. In some embodiments, the length of the polyA region is at least 30 nucleotides or at least 60 nucleotides.
[0491] In some embodiments, the RNA polynucleotides described herein are circularized by a ligation reaction. In some embodiments, the RNA polynucleotides described herein are circularized in the presence of T4 ligase. In some embodiments, the RNA polynucleotides described herein are circularized in the absence of T4 ligase.
[0492] In some embodiments, the RNA polynucleotides described herein are circularized by a splicing reaction. In some embodiments, the RNA polynucleotides described herein are circularized in the presence of spliceosomes. In some embodiments, the RNA polynucleotides described herein are circularized in the absence of spliceosomes.
[0493] In some embodiments, the RNA polynucleotides described herein are circularized by a self-splicing reaction.
[0494] 8.4.3. Vectors, Linear RNA, Precursor RNA, and Circular RNA In some embodiments, the RNA polynucleotides provided herein are single-stranded RNA. In some embodiments, the polynucleotide is linear RNA. In some embodiments, precursor RNA is provided herein. In some embodiments, the RNA polynucleotide is encoded by a vector. In some embodiments, the precursor RNA is linear RNA produced by in vitro transcription of the vector provided herein.
[0495] In some embodiments, the RNA polynucleotide is a circular RNA or is useful for creating a circular RNA polynucleotide. In some embodiments, circular RNAs are provided herein. In some embodiments, the circular RNA is a circular RNA produced by a vector provided herein. In some embodiments, the circular RNA is a circular RNA produced by circularization of a precursor RNA provided herein.
[0496] Circular RNA Circular RNA (also referred to as "circRNA" or "cRNA") is a single-stranded RNA that is covalently linked head-to-tail. circRNAs are recognized as a class of non-coding RNAs that are widely present in eukaryotic cells. It has been found that circRNAs, which are usually generated by backsplicing, are very stable.
[0497] Circular RNAs are known in the art or can be generated by any suitable method disclosed herein. Exemplary methods for generating circular RNAs are described in WO 2021 / 263124, Wesselhoeft, R. A. et al., Nat. Commun. 9, 2629 (2018), and Chinese Patent Application No. 2021 / 10594352.4, the contents of each of which are incorporated herein by reference in their entirety.
[0498] Circular RNAs can be generated by splicing methods other than those in mammals. For example, linear RNAs containing various types of introns, such as self-splicing group I introns, self-splicing group II introns, spliceosomal introns, and tRNA introns, can be circularized. In particular, group I and group II introns have the advantage that they can be easily used for the generation of circular RNAs in vitro and in vivo because they are capable of undergoing self-splicing by self-catalytic ribozyme activity.
[0499] Alternatively, circular RNAs can be produced in vitro from linear RNAs by chemically or enzymatically ligating the 5' and 3' ends of the RNA. Chemical ligation can be performed, for example, by activating nucleotide phosphomonoester groups using cyanogen bromide (BrCN) or ethyl-3-(3-dimethylaminopropyl)carbodiimide (EDC) to enable the formation of phosphodiester bonds (Sokolova, FEBS Lett, 232: 153-155 (1988); Dolinnaya et al., Nucleic Acids Res., 19: 3067-3072 (1991); Fedorova, Nucleosides Nucleotides Nucleic Acids, 15: 1137-1147 (1996)). Alternatively, RNA can also be circularized using an enzymatic ligation reaction. Examples of ligases that can be used include T4 DNA ligase (T4 Dnl), T4 RNA ligase 1 (T4 Rnl 1), and T4 RNA ligase 2 (T4 Rnl2).
[0500] In some embodiments, circular RNAs can be generated using splint ligation. In splint ligation, oligonucleotide splints that hybridize to both ends of a linear RNA are used to ligate the two ends of the linear RNA. Hybridization of the splint, which can be either a deoxyribooligonucleotide or a ribooligonucleotide, orients the 5-phosphate and 3-OH of the RNA ends for ligation. Subsequent ligation can be performed using either chemical or enzymatic methods as described above. Enzymatic ligation can be performed, for example, using T4 DNA ligase (requires a DNA splint), T4 RNA ligase 1 (requires an RNA splint), or T4 RNA ligase 2 (DNA or RNA splint). Chemical ligation, such as with BrCN or EDC, may be more efficient than enzymatic ligation if the structure of the hybridized splint RNA complex interferes with enzyme activity (see, for example, Dolinnaya et al. Nucleic Acids Res, 2 / (23): 5403-5407 (1993); Petkovic et al., Nucleic Acids Res, 43(4): 2454-2465 (2015)).
[0501] Circular RNAs are generally more stable than linear RNAs, mainly because they lack the free ends required for degradation by exonucleases. To further improve stability, additional modifications may be added to the recombinant circRNAs described herein. Further, other types of modifications may improve circularization efficiency, purification of circRNAs, and / or protein expression from circRNAs. For example, recombinant circRNAs may be modified to include "homology arms" (i.e., nucleotides 9-19 in length located at the 5' and 3' ends of the precursor RNA for the purpose of bringing the 5' and 3' splice sites closer together), spacer sequences (i.e., linker sequences), and / or phosphorothioate (PS) caps (Wesselhoeft et al., Nat. Commun., 9: 2629 (2018)). Recombinant circRNAs may be modified to include 2'-O-methylfluoro- or -O-methoxyethyl conjugates, phosphorothioate backbones, or 2',4'-cyclic 2'-O-ethyl modifications to enhance stability (Holdt et al., Front Physiol., 9: 1262 (2018); Kratzfeldt et al., Nature, 438(7068): 685-9 (2005); and Crooke et al., CellMetab., 27(4): 714-739 (2018)). Recombinant circRNA molecules may also include one or more modifications that reduce the innate immunogenicity of the circRNA molecule in the host, such as at least one N6-methyladenosine (m 6 A).
[0502] In some embodiments, the recombinant circular RNA molecule is encoded by a nucleic acid comprising at least two introns and at least one exon. In some embodiments, the DNA sequence encoding the circular RNA molecule comprises a sequence encoding at least two introns and at least one exon. As used herein, the term "exon" refers to a nucleic acid sequence present within a gene that is represented in the mature form of the RNA molecule after introns have been removed during transcription. Exons may be translated into proteins (e.g., in the case of messenger RNA (mRNA)). As used herein, the term "intron" refers to a nucleic acid sequence present within a particular gene that is removed by RNA splicing during the maturation of the final RNA product. Introns are generally present between exons. During transcription, introns are removed from the precursor messenger RNA (pre-mRNA), and exons are joined by RNA splicing. In some embodiments, the recombinant circular RNA molecule comprises a nucleic acid sequence comprising one or more exons and one or more introns.
[0503] Thus, circular RNAs can be generated by splicing of either endogenous or exogenous introns as described in WO 2017 / 222911 pamphlet, the content of which is hereby incorporated by reference in its entirety. As used herein, the term "endogenous intron" means an intron sequence that is specific to the host cell in which the circRNA is generated. For example, when a circRNA is expressed in human cells, the human intron is an endogenous intron. "Exogenous intron" means an intron that is heterologous to the host cell in which the circRNA is generated. For example, when a circRNA is expressed in human cells, a bacterial intron is an exogenous intron. Numerous intron sequences from a variety of organisms and viruses are known and include sequences derived from genes encoding proteins, ribosomal RNA (rRNA), or transfer RNA (tRNA). Representative intron sequences are available in various databases, including the Group I intron sequence and structure database (ma.whu.edu.cn / gissd / ), the database of bacterial Group II introns (webapps2.ucalgary.ca / ~groupii / index.html), the database of mobile Group II introns (fp.ucalgary.ca / group2introns), the yeast intron database (emblS16 heidelberg.de / Extemallnfo / seraphin / yidb.html), the Ares Lab yeast intron database (compbio. soe.ucsc.edu / yeast_introns.html), the U12 intron database (genome.crg.es / cgibin / ul2db / ul2db.cgi), and the exon intron database (bpg.utoledo.edu / ~afedorov / lab / eid.html).
[0504] In some embodiments, the RNA polynucleotide (e.g., circular RNA) may be of any length or size. In some embodiments, the length of the RNA polynucleotide is 300 to 10,000, 400 to 9,000, 500 to 8,000, 600 to 7,000, 700 to 6,000, 800 to 5,000, 900 to 5,000, 1,000 to 5,000, 1,100 to 5,000, 1,200 to 5,000, 1,300 to 5,000, 1,400 to 5,000, and / or 1,500 to 5,000 nucleotides.
[0505] In some embodiments, the length of the RNA polynucleotide (e.g., circular RNA) is at least 300 nt, 400 nt, 500 nt, 600 nt, 700 nt, 800 nt, 900 nt, 1,000 nt, 1,100 nt, 1,200 nt, 1,300 nt, 1,400 nt, 1,500 nt, 2,000 nt, 2,500 nt, 3,000 nt, 3,500 nt, 4,000 nt, 4,500 nt, or 5,000 nt. In some embodiments, the length of the RNA polynucleotide is 3,000 nt, 3,500 nt, 4,000 nt, 4,500 nt, 5,000 nt, 6,000 nt, 7,000 nt, 8,000 nt, 9,000 nt, or 10,000 nt or less.
[0506] In some embodiments, the length of the RNA polynucleotide (e.g., circular RNA) is about 300 nt, 400 nt, 500 nt, 600 nt, 700 nt, 800 nt, 900 nt, 1,000 nt, 1,100 nt, 1,200 nt, 1,300 nt, 1,400 nt, 1,500 nt, 2,000 nt, 2,500 nt, 3,000 nt, 3,500 nt, 4,000 nt, 4,500 nt, 5,000 nt, 6,000 nt, 7,000 nt, 8,000 nt, 9,000 nt, or 10,000 nt.
[0507] In some embodiments, the length of the RNA polynucleotide (e.g., circular RNA) is at least 500, 600, 700, 800, 900, 1000, 1500, 2000, 2500, 3000, 3500, 4000, 4500, 5000, 5500, 6000, 6500, 7000, 7500, 8000, 8500, 9000, 9500, or 10000 nucleotides. The RNA polynucleotide (e.g., circular RNA) may be unmodified, partially modified, or fully modified.
[0508] In some embodiments, the circular RNAs provided herein have higher functional stability than mRNAs containing the same expression sequences. In some embodiments, the circular RNAs provided herein have higher functional stability than mRNAs containing the same expression sequences, 5moU modifications, optimized UTRs, caps, and / or polyA tails.
[0509] In some embodiments, the circular RNA polynucleotides provided herein have a functional half-life of at least 5 hours, 10 hours, 15 hours, 20 hours, 30 hours, 40 hours, 50 hours, 60 hours, 70 hours, or 80 hours. In some embodiments, the circular RNA polynucleotides provided herein have a functional half-life of 5 - 80, 10 - 70, 15 - 60, and / or 20 - 50 hours. In some embodiments, the circular RNA polynucleotides provided herein have a longer (e.g., at least 1.5-fold longer, at least 2-fold longer) functional half-life than equivalent linear RNA polynucleotides encoding the same protein. In some embodiments, the functional half-life can be evaluated through detection of functional protein synthesis.
[0510] In some embodiments, the circular RNA polynucleotides provided herein have a half-life of at least 5 hours, 10 hours, 15 hours, 20 hours, 30 hours, 40 hours, 50 hours, 60 hours, 70 hours, or 80 hours. In some embodiments, the half-life of the circular RNA polynucleotides provided herein is 5 - 80, 10 - 70, 15 - 60, and / or 20 - 50 hours. In some embodiments, the circular RNA polynucleotides provided herein have a longer half-life (e.g., at least 1.5-fold, at least 2-fold longer) than the half-life of equivalent linear RNA polynucleotides encoding the same protein.
[0511] In some embodiments, the circular RNAs provided herein may have a higher expression level than equivalent linear mRNAs, e.g., a higher expression level 24 hours after administration of the RNA to cells. In some embodiments, the circular RNAs provided herein have a higher expression level than mRNAs containing the same expression sequence, 5moU modification, optimized UTR, cap, and / or polyA tail. In some embodiments, the circular RNAs provided herein may have higher stability than equivalent linear mRNAs. In some embodiments, this may be demonstrated by measuring the presence and density of receptors over a one-week period, either in vitro or in vivo. In some embodiments, this may be demonstrated by measuring the presence of the RNA via qPCR or ISH.
[0512] In some embodiments, the circular RNA polynucleotides provided herein contain modified RNA nucleotides and / or modified nucleosides. In some embodiments, said modified nucleoside is m 5 C (5-methylcytidine). In some embodiments, said modified nucleoside is m 5 U (5-methyluridine). In some embodiments, said modified nucleoside is m 6 A (N 6 -methyladenosine). In some embodiments, said modified nucleoside is s 2It is U (2-thiouridine). In some embodiments, the modified nucleoside is Y (pseudouridine). In some embodiments, the modified nucleoside is Um (2'-O-methyluridine). In some embodiments, the modified nucleoside is m ! A (1-methyladenosine), m 2 A (2-methyladenosine), Am (2'-O-methyladenosine), ms 2 m 6 A (2-methylthio-N 6 -methyladenosine), i 6 A (N 6 -isopentenyladenosine), ms 2 i 6 A (2-methylthio-N 6 isopentenyladenosine), io 6 A (N 6 -(cis-hydroxyisopentenyl)adenosine), ms 2 io 6 A (2-methylthio-N 6 -(cis-hydroxyisopentenyl)adenosine), g 6 A (N6-glycylcarbamoyladenosine), t 6 A (N 6 -threonylcarbamoyladenosine), ms 2 t 6 A (2-methylthio-N 6 -threonylcarbamoyladenosine), m 6 t 6 A (N 6 -methyl-N 6 -threonylcarbamoyladenosine), hn 6 A (N 6 -hydroxynorvalylcarbamoyladenosine), ms 2 hn 6 A (2-methylthio-N 6 -hydroxynorvalylcarbamoyladenosine), Ar(p) (2'-O-ribosyladenosine (phosphate)), I (inosine), m 1 I (1-methylinosine), m 1 hn (1,2'-O-dimethylinosine), m 3C(3-methylcytidine), Cm(2'-O-methylcytidine), s 2 C(2-thiocytidine), ac 4 C(N 4 -acetylcytidine), (5-formylcytidine), m 5 Cm(5,2'-O-dimethylcytidine), ac 4 Cm(N 4 -acetyl-2'-O-methylcytidine), k 2 C(lysidine), m ! G(1-methylguanosine), m 2 G(N2-methylguanosine), m 7 G(7-methylguanosine), Gm(2'-O-methylguanosine), m 2 G(N 2 ,N 2 -dimethylguanosine), m 2 Gm(N 2 ,2'-O-dimethylguanosine), m 2 aGm(N 2 ,N 2 ,2'-O-trimethylguanosine), Gr(p)(2'-O-ribosylguanosine (phosphate)), yW(wybutosine), oayW(peroxywybutosine), OHyW(hydroxywybutosine), OHyW*(undermodified hydroxywybutosine), imG(wiosine), mimG(methylwiosine), Q(queuosine), oQ(epoxyqueuosine), galQ(galactosylqueuosine), manQ(mannosylqueuosine), preQo(7-cyano-7-deazaguanosine), preQi(7-aminomethyl-7-deazaguanosine), G + (archaeosine), D(dihydrouridine), m 5 Um(5,2'-O-dimethyluridine), s 4 U(4-thiouridine), m 5 s 2 U(5-methyl-2-thiouridine), s 2 Um(2-thio-2'-O-methyluridine), acp 3 U(3-(3-amino-3-carboxypropyl)uridine), ho 5 U(5-hydroxyuridine), mo5 U (5-methoxyuridine), cmo 5 U (uridine 5-oxyacetic acid), mcmo 5 U (methyl uridine 5-oxyacetate), chm 5 U (5-(carboxyhydroxymethyl)uridine), mchm 5 U (methyl 5-(carboxyhydroxymethyl)uridine), mcm 5 U (5-methoxycarbonylmethyluridine), mcm 5 Um (5-methoxycarbonylmethyl-2'-O-methyluridine), mcm 5 s 2 U (5-methoxycarbonylmethyl-2-thiouridine), nm 5 S 2 U (5-aminomethyl-2-thiouridine), mnm 5 U (5-methylaminomethyluridine), mnm 5 s 2 U (5-methylaminomethyl-2-thiouridine), mnm 5 se 2 U (5-methylaminomethyl-2-selenouridine), ncm 5 U (5-carbamoylmethyluridine), ncm 5 Um (5-carbamoylmethyl-2'-O-methyluridine), cmnm 5 U (5-carboxymethylaminomethyluridine), cmnm 5 Um (5-carboxymethylaminomethyl-2'-O-methyluridine), cmnm 5 s 2 U (5-carboxymethylaminomethyl-2-thiouridine), m 6 2A (N 6 ,N 6 -dimethyladenosine), Im (2'-O-methylinosine), m 4 C (N 4 -methylcytidine), m 4 Cm (N 4 ,2'-O-dimethylcytidine), hm 5 C (5-hydroxymethylcytidine), m 3 U (3-methyluridine), cm 5U (5-carboxymethyluridine), m 6 Am(N 6 , 2'-O-dimethyladenosine), m 6 2Am(N 6 , N 6 , 0-2'-trimethyladenosine), m 2,7 G(N 2 , 7-dimethylguanosine), m 2,2,7 G(N 2 , N 2 , 7-trimethylguanosine), m 3 Um(3, 2'-O-dimethyluridine), m 5 D(5-methyldihydrouridine), f 5 Cm(5-formyl-2'-O-methylcytidine), m'Gm(1, 2'-O-dimethylguanosine), m'Am(1, 2'-O-dimethyladenosine), rm 5 U(5-taurinomethyluridine), τm5s2U(5-taurinomethyl-2-thiouridine), imG-14(4-demethylwyosine), imG2(isowyosine), or ac 6 A(N 6 -acetyladenosine).
[0513] In some embodiments, the modified nucleoside is pyridin-4-one ribonucleoside, 5-azauridine, 2-thio-5-azauridine, 2-thiouridine, 4-thio-pseudouridine, 2-thio-pseudouridine, 5-hydroxyuridine, 3-methyluridine, 5-carboxymethyluridine, 1-carboxymethyl-pseudouridine, 5-propynyluridine, 1-propynyl-pseudouridine, 5-taurinomethyluridine, 1-taurinomethyl-pseudouridine, 5-taurinomethyl-2-thio-uridine, l-taurinomethyl-4-thio-uridine, 5-methyl-uridine, 1-methyl-pseudouridine, 4-thio-1-methyl-pseudouridine, 2-thio-1-methyl-pseudouridine, 1-methyl-1-deaza-pseudouridine, 2-thio-1-methyl-1-deaza-pseudouridine, dihydrouridine, dihydro-pseudouridine, 2-thio-dihydrouridine, 2-thiodihydropseudouridine, 2-methoxyuridine, 2-methoxy-4-thiouridine, 4-methoxypseudouridine, 4-methoxy-2-thiopseudouridine, 5-azacytidine, pseudoisocytidine, 3-methylcytidine, N4-acetylcytidine, 5-formylcytidine, N4-methylcytidine, 5-hydroxymethylcytidine, 1-methyl-pseudo-isocytidine, pyrrolocytidine, pyrrolo-pseudo-isocytidine, 2-thiocytidine, 2-thio-5-methylcytidine, 4-thio-pseudo-isocytidine, 4-thio-1-methyl-pseudo-isocytidine, 4-thio-1-methyl-1-deaza-pseudo-isocytidine, 1-methyl-1-deaza-pseudo-isocytidine, zebularine, 5-aza-zebularine, 5-methyl-zebularine, 5-aza-2-thio-zebularine, 2-thio-zebularine, 2-methoxy-cytidine, 2-methoxy-5-methyl-cytidine, 4-methoxy-pseudo-isocytidine, 4-methoxy-1-methyl-pseudo-isocytidine, 2-aminopurine, 2,6-diaminopurine, 7-deaza-adenine, 7-deaza-8-aza-adenine, 7-deaza-2-aminopurine, 7-deaza-8-aza-2-aminopurine, 7-deaza-2,It may include 6-diaminopurine, 7-deaza-8-aza-2,6-diaminopurine, 1-methyladenosine, N6-methyladenosine, N6-isopentenyladenosine, N6-(cis-hydroxyisopentenyl)adenosine, 2-methylthio-N6-(cis-hydroxyisopentenyl)adenosine, N6-glycinylcarbamoyladenosine, N6-threonylcarbamoyladenosine, 2-methylthio-N6-threonylcarbamoyladenosine, N6,N6-dimethyladenosine, 7-methyladenine, 2-methylthioadenine, 2-methoxyadenine, inosine, 1-methylinosine, wyosine, wybutosine, 7-deazaguanosine, 7-deaza-8-azaguanosine, 6-thioguanosine, 6-thio-7-deazaguanosine, 6-thio-7-deaza-8-azaguanosine, 7-methylguanosine, 6-thio-7-methylguanosine, 7-methylinosine, 6-methoxyguanosine, 1-methylguanosine, N2-methylguanosine, N2,N2-dimethylguanosine, 8-oxoguanosine, 7-methyl-8-oxoguanosine, 1-methyl-6-thioguanosine, N2-methyl-6-thioguanosine, and N2,N2-dimethyl-6-thioguanosine. In some embodiments, the modification is independently selected from the group consisting of 5-methylcytosine, pseudouridine, and 1-methyl-pseudouridine.,
[0514] In some embodiments, the polynucleotide may be codon-optimized. A codon-optimized sequence may be a sequence in which codons within the polynucleotide encoding the therapeutic agent are replaced in order to enhance the expression, stability, and / or activity of the therapeutic agent. Factors affecting codon optimization include, but are not limited to, one or more of the following: (i) changes in codon bias between two or more organisms or genes or synthetically constructed bias tables; (ii) variation in the degree of codon bias within an organism, gene, or gene set; (iii) systematic mutation of codons including context; (iv) mutation of codons according to decoding tRNA; (v) mutation of codons by GC% (overall or at one position of the triplet); (vi) variation in similarity to a reference sequence (e.g., a naturally occurring sequence); (vii) variation in codon frequency cut-off; (viii) structural properties of mRNA transcribed from a DNA sequence; (ix) prior knowledge regarding the function of the DNA sequence underlying the design of a codon substitution set; and / or (x) systematic mutation of codon sets for each amino acid. In some embodiments, a polynucleotide with optimized codons can minimize ribozyme collisions and / or limit structural interference between the expression sequence and the IRES.
[0515] 8.4.4. RNA Polynucleotide Production The vectors provided herein can be made using standard techniques of molecular biology. For example, the various elements of the vectors provided herein can be obtained using recombinant methods such as by screening cDNA and genomic libraries from cells or by using polynucleotides from vectors known to contain polynucleotides.
[0516] The various elements of the vectors provided herein can be generated synthetically rather than cloned based on known sequences. The complete sequences can be assembled from overlapping oligonucleotides prepared by standard methods. See, for example, Edge, Nature (1981) 292:756; Nambair et al, Science (1984) 223:1299; and Jay et al, J. Biol. Chem. (1984) 259:6311.
[0517] Accordingly, a particular nucleotide sequence can be obtained from a vector containing the desired sequence or, where appropriate, can be synthesized completely or partially using various oligonucleotide synthesis techniques known in the art such as site-directed mutagenesis or polymerase chain reaction (PCR) techniques. One method for obtaining a nucleotide sequence encoding a desired vector element is to anneal a complementary set of overlapping synthetic oligonucleotides generated by a conventional automated polynucleotide synthesizer, subsequently ligate with an appropriate DNA ligase, and amplify the ligated nucleotide sequence by PCR. See, for example, Jayaraman et al, Proc. Natl. Acad. Sci. USA (1991) 88:4084-4088. Additionally, oligonucleotide-directed synthesis (Jones et al, Nature (1986) 54:75-82), oligonucleotide-directed mutagenesis of existing nucleotide regions (Riechmann et al, Nature (1988) 332:323-327 and Verhoeyen et al., Science (1988) 239:1534-1536), and enzymatic filling of gap oligonucleotides using T4 DNA polymerase (Queen et al, Proc. Natl. Acad. Sci. USA (1989) 86:10029-10033) can also be used.
[0518] The precursor RNA provided herein can be generated by incubating the vector provided herein under conditions that permit transcription of the precursor RNA encoded by the vector. For example, in some embodiments, the vector provided herein (including an RNA polymerase promoter and / or expression sequence upstream of the 5' duplex-forming region) is incubated with a compatible RNA polymerase enzyme under conditions that permit in vitro transcription, whereby the precursor RNA is synthesized. In some embodiments, the vector is incubated intracellularly by a bacteriophage RNA polymerase or in the nucleus of a cell by a host RNA polymerase P.
[0519] In some embodiments, provided herein is a method for generating precursor RNA by performing in vitro transcription using as a template the vector provided herein (e.g., the vector provided herein having an RNA polymerase promoter located upstream of the 5' homologous region).
[0520] In some embodiments, the resulting precursor RNA can be used to generate circular RNA (e.g., the circular RNA polynucleotide provided herein).
[0521] Accordingly, in some embodiments, provided herein is a method for making circular RNA. In some embodiments, the method includes synthesizing precursor RNA by transcription (e.g., runoff transcription) using as a template the vector provided herein, and incubating the resulting precursor RNA under conditions suitable for circularization to form circular RNA.
[0522] In some embodiments, a composition comprising circular RNA is purified. The circular RNA can be purified by any known method widely used in the art, such as column chromatography, gel filtration chromatography, size exclusion chromatography, etc. In some embodiments, the purification comprises one or more steps of phosphatase treatment, HPLC size exclusion purification, and RNase R digestion. In some embodiments, the purification comprises the steps of RNase R digestion, phosphatase treatment, and HPLC size exclusion purification in sequence. In some embodiments, the purification comprises reverse phase HPLC. In some embodiments, the purified composition comprises less double-stranded RNA, DNA sprites, triphosphorylated RNA, phosphatase protein, protein ligase, capping enzyme, and / or nicked RNA than the unpurified RNA.
[0523] 8.5. Formulation and Delivery In some embodiments, the polynucleotides or circular RNAs disclosed herein can be formulated using liposomes, lipoplexes, lipid nanoparticles, polymer-based delivery systems, and viral vectors. In some embodiments, the polynucleotides or circular RNAs can be formulated into lipid nanoparticles as described in WO 2012 / 170930 pamphlet, the content of which is incorporated herein by reference in its entirety. In some embodiments, the lipid can be a cleavable lipid such as those described in WO 2012 / 170889 pamphlet, the content of which is incorporated herein by reference in its entirety. In some embodiments, the pharmaceutical composition of the polynucleotide or circular RNA may include at least one of the PEGylated lipids described in WO 2012 / 099755 pamphlet incorporated herein by reference. In some embodiments, the lipid nanoparticle formulation can be formulated by the methods described in WO 2011 / 127255 pamphlet or WO 2008 / 103276 pamphlet, the content of each of these documents is incorporated herein by reference in its entirety. The lipid nanoparticles may be coated or conjugated with a copolymer such as, but not limited to, a block copolymer such as the branched polyether-polyamide block copolymer described in WO 2013 / 012476 incorporated herein by reference in its entirety. Liposomes, lipoplexes, or lipid nanoparticles can be used to improve the efficiency of protein production by the polynucleotides or circular RNAs. These formulations are for increasing cell transfection by the polynucleotides and circular RNAs, extending the in vivo or in vitro half-life of the polynucleotides and circular RNAs, and / or enabling sustained release.
[0524] In some embodiments, the polynucleotides or circular RNAs disclosed herein encode a protein composed of subunits encoded by multiple genes. For example, the protein is a heterodimer, and each strand or subunit of the protein is encoded by a separate gene. Multiple polynucleotide or circular RNA molecules can be delivered in a delivery vehicle, and each polynucleotide or circular RNA can encode a separate subunit of the protein. Alternatively, a single polynucleotide or circular RNA molecule may be modified to encode multiple subunits. In some embodiments, separate polynucleotide or circular RNA molecules encoding individual subunits may be administered in separate delivery vehicles.
[0525] The present invention contemplates the differential targeting of target cells and tissues by both passive and active targeting means. The phenomenon of passive targeting takes advantage of the natural distribution pattern of the delivery vehicle in vivo without relying on the use of additional excipients or means to enhance the recognition of the delivery vehicle by the target cells. For example, delivery vehicles that are phagocytosed by reticuloendothelial cells are likely to accumulate in the liver or spleen and thus can provide a means of passively directing the delivery of the composition to such target cells.
[0526] Alternatively, the present disclosure contemplates active targeting, which involves the use of targeting moieties that can be attached (either covalently or non-covalently) to the transport vehicle to facilitate the localization of such transport vehicles in specific target cells or target tissues. For example, targeting can be mediated by including one or more endogenous targeting moieties within or on the transport vehicle to facilitate distribution to target cells or tissues. Recognition of the target moiety by the target tissue actively promotes the tissue distribution and intracellular uptake of the transport vehicle and / or its contents in target cells and tissues (e.g., inclusion of apolipoprotein E target ligand within or on the transport vehicle promotes recognition and binding of the transport vehicle to the endogenous low density lipoprotein receptor expressed by hepatocytes). As provided herein, the composition can include moieties that can enhance the affinity of the composition for target cells. The target moiety may be bound to the outer bilayer of the lipid particle during or after formulation. These methods are well known in the art. Additionally, some lipid particle formulations employ fusogenic polymers such as PEAA, hemagglutinin, other lipopeptides (see U.S. Patent No. 6,417,326, the entire content of which is incorporated herein by reference), and other features useful for in vivo and / or intracellular delivery. In some other embodiments, the compositions of the present disclosure exhibit improved transfection efficiency and / or improved selectivity for target cells or tissues of interest. Accordingly, compositions are envisioned that include one or more moieties (e.g., peptides, aptamers, oligonucleotides, vitamins, or other molecules) that can enhance the affinity of the composition and its nucleic acid contents for target cells or tissues. Suitable moieties may be bound or linked to the surface of the transport vehicle, if desired. In some embodiments, the target moiety can spread on the surface of the transport vehicle or be encapsulated within the transport vehicle. Suitable moieties are selected based on physical, chemical, or biological properties (e.g., selective affinity and / or recognition of target cell surface markers or features). Cell-specific target sites and their corresponding target ligands are diverse.An appropriate target moiety is selected such that the unique properties of the target cell are utilized, thereby enabling the composition to distinguish between target cells and non-target cells. For example, the compositions of the present disclosure include surface markers (e.g., apolipoprotein B or apolipoprotein E), thereby selectively enhancing the recognition or affinity for hepatocytes (e.g., by recognition and binding via receptors for such surface markers). As an example, using galactose as a targeting moiety, the compositions of the present invention are expected to be directed to hepatocytes, or using a mannose-containing sugar residue as a targeting ligand, the compositions of the present invention are expected to be directed to hepatic endothelial cells (e.g., mannose-containing sugar residues that can preferentially bind to the asialoglycoprotein receptor present on hepatocytes) (see Hillery A M, et al. “Drug Delivery and Targeting: For Pharmacists and Pharmaceutical Scientists” (2002) Taylor & Francis, Inc.). Thus, the presentation of such target moieties conjugated to moieties present in a delivery vehicle (e.g., lipid nanoparticles) facilitates the recognition and uptake of the compositions of the present disclosure in target cells and tissues. Examples of suitable target moieties include one or more peptides, proteins, aptamers, vitamins, and oligonucleotides.
[0527] In some embodiments, the polynucleotides or circular RNAs disclosed herein are formulated according to the process described in US Patent Application Publication No. 2018 / 0153822. In some embodiments, the present disclosure provides a process for encapsulating a polynucleotide or circular RNA into lipid nanoparticles, comprising forming lipids into pre-formed lipid nanoparticles (i.e., those formed in the absence of RNA) and then combining the pre-formed lipid nanoparticles with RNA. In some embodiments, the formulation process results in an RNA formulation that potentially has better tolerance both in vitro and in vivo, higher efficacy (peptide or protein expression), and higher effectiveness (improvement of biologically relevant endpoints), compared to the same RNA formulation prepared without the step of pre-forming lipid nanoparticles (e.g., directly combining lipids with RNA).
[0528] In some embodiments, the delivery vehicle is formulated and / or targeted as described in Shobaki N et al., Int J Nanomedicine 2018;13:8395-8410. In some embodiments, the delivery vehicle is composed of three lipids. In some embodiments, the delivery vehicle is composed of four lipids. In some embodiments, the delivery vehicle is composed of five lipids. In some embodiments, the delivery vehicle is composed of six lipids.
[0529] In the case of certain cationic lipid nanoparticle formulations of RNA, it is necessary to heat the RNA in a buffer (e.g., citrate buffer) in order to achieve a high degree of encapsulation of the RNA. In these processes or methods, heating after formulation (after nanoparticle formation) does not improve the encapsulation efficiency of the RNA into the lipid nanoparticles, so heating needs to be performed prior to the formulation process (i.e., heating of the individual components). In contrast, in some embodiments, the heating order of the RNA is seen not to affect the encapsulation rate of the RNA. In some embodiments, it is not necessary to heat (i.e., maintain at ambient temperature) one or more of a solution containing pre-formed lipid nanoparticles, a solution containing RNA, and a mixed solution containing RNA encapsulated in lipid nanoparticles, either before or after the formulation process.
[0530] RNA is provided in a solution that is mixed with a lipid solution, such that the RNA is encapsulated within lipid nanoparticles. Suitable RNA solutions can be any aqueous solution containing RNA encapsulated at various concentrations. For example, suitable RNA solutions can contain RNA at concentrations of about 0.01 mg / mL, 0.05 mg / mL, 0.06 mg / mL, 0.07 mg / mL, 0.08 mg / mL, 0.09 mg / mL, 0.1 mg / mL, 0.15 mg / mL, 0.2 mg / mL, 0.3 mg / mL, 0.4 mg / mL, 0.5 mg / mL, 0.6 mg / mL, 0.7 mg / mL, 0.8 mg / mL, 0.9 mg / mL, or 1.0 mg / mL or greater. In some embodiments, suitable RNA solutions can contain RNA at concentrations in the range of about 0.01 - 1.0 mg / mL, 0.01 - 0.9 mg / mL, 0.01 - 0.8 mg / mL, 0.01 - 0.7 mg / mL, 0.01 - 0.6 mg / mL, 0.01 - 0.5 mg / mL, 0.01 - 0.4 mg / mL, 0.01 - 0.3 mg / mL, 0.01 - 0.2 mg / mL, 0.01 - 0.1 mg / mL, 0.05 - 1.0 mg / mL, 0.05 - 0.9 mg / mL, 0.05 - 0.8 mg / mL, 0.05 - 0.7 mg / mL, 0.05 - 0.6 mg / mL, 0.05 - 0.5 mg / mL, 0.05 - 0.4 mg / mL, 0.05 - 0.3 mg / mL, 0.05 - 0.2 mg / mL, 0.05 - 0.1 mg / mL, 0.1 - 1.0 mg / mL, 0.2 - 0.9 mg / mL, 0.3 - 0.8 mg / mL, 0.4 - 0.7 mg / mL, or 0.5 - 0.6 mg / mL.
[0531] Typically, suitable RNA solutions can also contain a buffer and / or a salt. Generally, buffers can include HEPES, ammonium sulfate, Tris, sodium bicarbonate, sodium citrate, sodium acetate, potassium phosphate, or sodium phosphate. In some embodiments, suitable concentrations of the buffer can be in the range of about 0.1 mM - 100 mM, 0.5 mM - 90 mM, 1.0 mM - 80 mM, 2 mM - 70 mM, 3 mM - 60 mM, 4 mM - 50 mM, 5 mM - 40 mM, 6 mM - 30 mM, 7 mM - 20 mM, 8 mM - 15 mM, or 9 - 12 mM.
[0532] Exemplary salts include sodium chloride, magnesium chloride, and potassium chloride. In some embodiments, the appropriate concentration of salt in the RNA solution can be in the range of about 1 mM to 500 mM, 5 mM to 400 mM, 10 mM to 350 mM, 15 mM to 300 mM, 20 mM to 250 mM, 30 mM to 200 mM, 40 mM to 190 mM, 50 mM to 180 mM, 50 mM to 170 mM, 50 mM to 160 mM, 50 mM to 150 mM, or 50 mM to 100 mM.
[0533] In some embodiments, the appropriate RNA solution can have a pH in the range of about 3.5 to 6.5, 3.5 to 6.0, 3.5 to 5.5, 3.5 to 5.0, 3.5 to 4.5, 4.0 to 5.5, 4.0 to 5.0, 4.0 to 4.9, 4.0 to 4.8, 4.0 to 4.7, 4.0 to 4.6, or 4.0 to 4.5.
[0534] To prepare an RNA solution suitable for the present disclosure, various methods can be used. In some embodiments, the RNA may be directly dissolved in the buffer described herein. In some embodiments, an RNA solution may be generated by mixing an RNA stock solution with a buffer solution before mixing with a lipid solution for encapsulation. In some embodiments, an RNA solution may be generated by mixing an RNA stock solution with a buffer solution immediately prior to mixing with a lipid solution for encapsulation.
[0535] According to the present invention, the lipid solution comprises a mixture of lipids suitable for forming a transport vehicle for encapsulating RNA. In some embodiments, the appropriate lipid solution is ethanol-based. For example, the appropriate lipid solution includes a mixture of the desired lipids dissolved in pure ethanol (i.e., 100% ethanol). In some embodiments, the appropriate lipid solution is isopropyl alcohol-based. In some embodiments, the appropriate lipid solution is dimethyl sulfoxide-based. In some embodiments, the appropriate lipid solution is a mixture of suitable solvents including, but not limited to, ethanol, isopropyl alcohol, and dimethyl sulfoxide.
[0536] A suitable lipid solution may contain a mixture of desired lipids at various concentrations. In some embodiments, a suitable lipid solution contains a mixture of desired lipids at a total concentration in the range of about 0.1 to 100 mg / mL, 0.5 to 90 mg / mL, 1.0 to 80 mg / mL, 1.0 to 70 mg / mL, 1.0 to 60 mg / mL, 1.0 to 50 mg / mL, 1.0 to 40 mg / mL, 1.0 to 30 mg / mL, 1.0 to 20 mg / mL, 1.0 to 15 mg / mL, 1.0 to 10 mg / mL, 1.0 to 9 mg / mL, 1.0 to 8 mg / mL, 1.0 to 7 mg / mL, 1.0 to 6 mg / mL, or 1.0 to 5 mg / mL.
[0537] Any lipids can be mixed in any ratio suitable for encapsulating RNA. In some embodiments, a suitable lipid solution contains a mixture of desired lipids including a cationic lipid, a helper lipid (e.g., a non-cationic lipid and / or a cholesterol lipid), and / or a PEGylated lipid. In some embodiments, a suitable lipid solution contains a mixture of desired lipids including one or more cationic lipids, one or more helper lipids (e.g., a non-cationic lipid and / or a cholesterol lipid), and one or more PEGylated lipids.
[0538] In some embodiments, the polynucleotides or circular RNAs disclosed herein are formulated using viral vectors. The viral vectors can be derived from various viruses such as adenovirus, adeno-associated virus, lentivirus (e.g., HIV, FIV, and EIAV), and herpes virus. Examples of commercially available viral vectors include pSilencer adeno (Ambion, Austin, TX) and pLenti6 / BLOCK-iT™-DEST (Invitrogen, Carlsbad, CA). The selection of viral vectors, the method of expressing polynucleotides or circular RNAs from the vectors, and the method of delivering the viral vectors are within the ordinary skill of those in the art. In some embodiments, the viral vector is a recombinant AAV (rAAV) vector known in the art (PMID: 30245471, PMID: 33614232).
[0539] The present invention also provides a delivery system comprising the polynucleotides or circular RNAs disclosed herein. In some embodiments, the delivery system is any one of a liposome, a nanoparticle, a polymer-based delivery system, or a ligand-conjugate delivery system. In some embodiments, the ligand-conjugate delivery system comprises one or more of an antibody, a peptide, a sugar moiety, or a combination thereof.
[0540] In some embodiments, the delivery system of the present disclosure comprises nanoparticles comprising the polynucleotides or circular RNAs disclosed herein.
[0541] In some embodiments, the nanoparticles include polymeric nanoparticles, lipid-polymeric nanoparticles, metallic nanoparticles, carbon nanotube-based nanoparticles, nanocrystals, or polymer micelles. In some embodiments, the polymeric nanoparticles include multi-block copolymers, diblock copolymers, polymer micelles, or hyperbranched polymers. In some embodiments, the polymeric nanoparticles include multi-block copolymers or diblock copolymers. In some embodiments, the polymeric nanoparticles are pH-responsive. In some embodiments, the polymeric nanoparticles further include a buffering component.
[0542] In some embodiments, the delivery system includes liposomes. A liposome is a spherical vesicle having at least one lipid bilayer and, in some embodiments, an aqueous core. In some embodiments, the lipid bilayer of the liposome can include phospholipids. An exemplary non-limiting example of a phospholipid is phosphatidylcholine, but the lipid bilayer may include additional lipids such as phosphatidylethanolamine. Liposomes can be of a multi-layered structure composed of multiple lamellar phase lipid bilayers or of a single-layered structure composed of a single lipid bilayer. Liposomes can be produced within a specific size range that is a viable target for phagocytosis. The size of the liposomes is in the range of 20 nm to 100 nm, 100 nm to 400 nm, 1 μM or more, or 200 nm to 3 μM. Examples of lipidoids and lipid-based formulations are described in U.S. Patent Application Publication No. 2009 / 0023673. In some embodiments, one or more lipids are one or more cationic lipids.
[0543] In some embodiments, the liposomes or nanoparticles of the present disclosure include micelles. A micelle is an aggregate of surfactant molecules. Exemplary micelles include aggregates of amphiphilic polymers, polymers, or copolymers in an aqueous solution, where the hydrophilic heads are in contact with the surrounding solvent and the hydrophobic tail regions are sequestered at the center of the micelle.
[0544] In some embodiments, the nanoparticles include nanocrystals. Exemplary nanocrystals are crystalline particles having at least one dimension less than 1000 nanometers, preferably less than 100 nanometers.
[0545] In some embodiments, the nanoparticles include polymeric nanoparticles. In some embodiments, the polymer includes a multiblock copolymer, diblock copolymer, polymer micelle, or hyperbranched polymer. In some embodiments, the particles include one or more cationic polymers. In some embodiments, the cationic polymer is chitosan, protamine, polylysine, polyhistidine, polyarginine, or poly(ethylene)imine. In some embodiments, the one or more polymers include a buffering component, a degradable component, a hydrophilic component, a cleavable linking component, or a combination thereof.
[0546] In some embodiments, the nanoparticles or a portion thereof are degradable. In some embodiments, the lipids and / or polymers of the nanoparticles are degradable.
[0547] In some embodiments, any of these delivery systems of the present disclosure can include a buffering component. In some embodiments, any of the present disclosure can include a buffering component and a degradable component. In some embodiments, any of the present disclosure can include a buffering component and a hydrophilic component. In some embodiments, any of the present disclosure can include a buffering component and a cleavable linking component. In some embodiments, any of the present disclosure can include a buffering component, a degradable component, and a hydrophilic component. In some embodiments, any of the present disclosure can include a buffering component, a degradable component, and a cleavable linking component. In some embodiments, any of the present disclosure can include a buffering component, a hydrophilic component, and a cleavable linking component. In some embodiments, any of the present disclosure can include a buffering component, a degradable component, a hydrophilic component, and a cleavable linking component. In some embodiments, the particles are composed of one or more polymers including any combination of the foregoing components.
[0548] In some embodiments, the delivery system includes a ligand-conjugate delivery system. In some embodiments, the ligand-conjugate delivery system includes one or more of an antibody, a peptide, a sugar moiety, a lipid, or a combination thereof.
[0549] In some embodiments, the polynucleotide or circular RNA disclosed herein is conjugated, complexed, or encapsulated in one or more lipids or polymers of the delivery system. In some embodiments, the polynucleotide or circular RNA can be encapsulated within the hollow core of a nanoparticle. Alternatively or additionally, the polynucleotide or circular RNA can be incorporated into the lipid or polymer-based shell of the delivery system, for example, via intercalation. Alternatively or additionally, the polynucleotide or circular RNA can be attached to the surface of the delivery system. In some embodiments, the polynucleotide or circular RNA is conjugated to one or more lipids or polymers of the delivery system, for example, by a covalent bond.
[0550] In some embodiments, the ligand-conjugate delivery system further includes a targeting agent. In some embodiments, the targeting agent includes a peptide ligand, a nucleotide ligand, a polysaccharide ligand, a fatty acid ligand, a lipid ligand, a small molecule ligand, an antibody, an antibody fragment, an antibody mimetic, or an antibody mimetic fragment.
[0551] In some embodiments, the delivery system of the present disclosure includes a polymeric delivery system. In some embodiments, the polymeric delivery system includes a blend polymer. In some embodiments, the blend polymer is a copolymer comprising a degradable component and a hydrophilic component. In some embodiments, the degradable component of the blend polymer is polyester, poly(orthoester), poly(ethyleneimine), poly(caprolactone), polyanhydride, poly(acrylic acid), polyglycolide, or poly(urethane). In some embodiments, the degradable component of the blend polymer is poly(lactic acid) (PLA) or poly(lactic acid-glycolic acid copolymer) (PLGA). In some embodiments, the hydrophilic component of the blend polymer is polyalkylene glycol or polyalkylene oxide. In some embodiments, the polyalkylene glycol is polyethylene glycol (PEG). In some embodiments, the polyalkylene oxide is polyethylene oxide (PEO).
[0552] In some embodiments, the delivery system of the present disclosure is polymeric nanoparticles. The polymeric nanoparticles comprise one or more polymers. In some embodiments, the one or more polymers comprise polyester, poly(orthoester), poly(ethyleneimine), poly(caprolactone), polyanhydride, poly(acrylic acid), polyglycolide, or poly(urethane). In some embodiments, the one or more polymers comprise poly(lactic acid) (PLA) or poly(lactic acid-glycolic acid copolymer) (PLGA). In some embodiments, the one or more polymers comprise poly(lactic acid-glycolic acid) (PLGA). In some embodiments, the one or more polymers comprise poly(lactic acid) (PLA). In some embodiments, the one or more polymers comprise polyalkylene glycol or polyalkylene oxide. In some embodiments, the polyalkylene glycol is polyethylene glycol (PEG), or the polyalkylene oxide is polyethylene oxide (PEO).
[0553] In some embodiments, the polymeric nanoparticles comprise poly(lactic-co-glycolic acid) (PLGA) polymer. In some embodiments, the PLGA nanoparticles further comprise a targeting agent described herein.
[0554] In some embodiments, the delivery system of the present disclosure is nanoparticles having an average characteristic dimension of less than about 500 nm, 400 nm, 300 nm, 250 nm, 200 nm, 180 nm, 150 nm, 120 nm, 100 nm, 90 nm, 80 nm, 70 nm, 60 nm, 50 nm, 40 nm, 30 nm, or 20 nm. In some embodiments, the average characteristic dimension of the nanoparticles is 10 nm, 20 nm, 30 nm, 40 nm, 50 nm, 60 nm, 70 nm, 80 nm, 90 nm, 100 nm, 120 nm, 150 nm, 180 nm, 200 nm, 250 nm, or 300 nm. In some embodiments, the average characteristic dimension of the nanoparticles is 10-500 nm, 10-400 nm, 10-300 nm, 10-250 nm, 10-200 nm, 10-150 nm, 10-100 nm, 10-75 nm, 10-50 nm, 50-500 nm, 50-400 nm, 50-300 nm, 50-200 nm, 50-150 nm, 50-100 nm, 50-75 nm, 100-500 nm, 100-400 nm, 100-300 nm, 100-250 nm, 100-200 nm, 100-150 nm, 150-500 nm, 150-400 nm, 150-300 nm, 150-250 nm, 150-200 nm, 200-500 nm, 200-400 nm, or 200-300 nm.
[0555] In some embodiments, the target cells lack a protein or enzyme of interest. For example, when it is desired to deliver a nucleic acid to hepatocytes, the hepatocytes represent the target cells. In some embodiments, the compositions of the present disclosure differentially transfect target cells (i.e., do not transfect non-target cells). The compositions of the invention can be prepared to preferentially target a variety of target cells including, but not limited to, hepatocytes, epithelial cells, hematopoietic cells, epithelial cells, endothelial cells, lung cells, bone cells, stem cells, mesenchymal cells, nerve cells (e.g., meningeal cells, astrocytes, motor neurons, dorsal root ganglion cells, and anterior horn motor neurons), photoreceptor cells (e.g., rod cells and cone cells), retinal pigment epithelial cells, secretory cells, heart cells, adipocytes, vascular smooth muscle cells, cardiomyocytes, skeletal muscle cells, beta cells, pituitary cells, synovial cells, ovarian cells, testicular cells, fibroblasts, B cells, T cells, NK cells, dendritic cells, macrophages, reticulocytes, white blood cells, granulocytes, tumor cells (including tumor cell lines such as Hela, MCF7, PC3, A549, NCI-H727, and HCT-116), immortalized cells (such as MCF10A, HEK293, HEK293T), and primary cell lines (such as hepatic stellate cells, HPrEC, FHC).
[0556] The compositions of the present disclosure can be prepared to preferentially distribute to target cells within specific organs such as the heart, lungs, kidneys, liver, and spleen, as described herein. In some embodiments, the compositions of the present disclosure distribute within the cells of the liver and promote the delivery and subsequent expression of the circular RNAs contained therein by the cells of the liver (e.g., hepatocytes). The target cells can function as a biological "reservoir" or "depot" (storage site) that produces a functional protein or enzyme and can excrete it systemically. Thus, in some embodiments, the transport vehicle can target hepatocytes upon delivery and / or preferentially distribute to the cells of the liver. In some embodiments, after transfection of the target hepatocytes, the circular RNA loaded onto the vehicle is translated, a functional protein product is produced, excreted, and distributed systemically. In some embodiments, cells other than hepatocytes (e.g., cells of the lung, spleen, heart, eye, or central nervous system) can function as a storage site for protein production.
[0557] In some embodiments, the compositions of the present disclosure promote the endogenous production of one or more functional proteins and / or enzymes by a subject. In some embodiments, the transport vehicle comprises a circular RNA encoding a deficient protein or enzyme. When such a composition is distributed to a target tissue and subsequently transfected into such target cells, the exogenous circular RNA loaded onto the transport vehicle (e.g., lipid nanoparticles) is translated in vivo, and a functional protein or enzyme encoded by the exogenously administered circular RNA (e.g., the protein or enzyme that the subject is deficient in) can be produced. Thus, the compositions of the present disclosure utilize the ability of a subject to translate exogenously or recombinantly prepared circular RNAs to produce endogenously translated proteins or enzymes, thereby producing (and, where applicable, excreting) functional proteins or enzymes. The expressed or translated protein or enzyme is also characterized by the inclusion in vivo of native post-translational modifications that are often not present in recombinantly prepared proteins or enzymes, thereby further reducing the immunogenicity of the translated protein or enzyme.
[0558] Administering a circular RNA encoding a defective protein or enzyme obviates the need to deliver the nucleic acid to a specific organelle within the target cell. Rather, if the target cell is transfected and the nucleic acid is delivered to the cytoplasm of the target cell, the circular RNA content of the transport vehicle can be translated and a functional protein or enzyme can be expressed.
[0559] In some embodiments, the circular RNA comprises one or more miRNA binding sites. In some embodiments, the circular RNA comprises one or more miRNA binding sites that are recognized by one or more miRNAs that are present in one or more non-target cells or non-target cell types (e.g., Kupffer cells) and not present in one or more target cells or target cell types (e.g., hepatocytes). In some embodiments, the circular RNA comprises one or more miRNA binding sites that are recognized by one or more miRNAs that are present at a high concentration in one or more non-target cells or non-target cell types (e.g., Kupffer cells) as compared to one or more target cells or target cell types (e.g., hepatocytes). miRNAs are thought to function in pairs with complementary sequences within the RNA molecule to effect gene silencing.
[0560] 8.6. Pharmaceutical Compositions In some embodiments, compositions (e.g., pharmaceutical compositions) comprising a therapeutic agent provided herein are provided herein. In some embodiments, the therapeutic agent is a circular RNA polynucleotide provided herein. In some embodiments, the therapeutic agent is a vector provided herein. In some embodiments, the therapeutic agent is a cell comprising a circular RNA or vector provided herein. In some embodiments, the composition further comprises a pharmaceutically acceptable carrier. In some embodiments, the compositions provided herein comprise a therapeutic agent provided herein in combination with other pharmaceutically active agents or drugs. In a preferred embodiment, the pharmaceutical composition comprises a cell or population thereof provided herein.
[0561] Regarding pharmaceutical compositions, the pharmaceutically acceptable carriers can be any of the conventionally used carriers, and are limited only by physicochemical considerations such as solubility and lack of reactivity with the active agent, as well as the route of administration. The pharmaceutically acceptable carriers described herein, such as vehicles, adjuvants, excipients, and diluents, are well known to those skilled in the art and are generally readily available. The pharmaceutically acceptable carrier is preferably chemically inert to the therapeutic agent and has no harmful side effects or toxicity under the conditions of use.
[0562] The selection of the carrier is determined in part not only by the particular therapeutic agent but also by the particular method used to administer the therapeutic agent. Accordingly, there are various suitable formulations for the pharmaceutical compositions provided herein.
[0563] In some embodiments, the pharmaceutical composition contains a preservative. In some embodiments, suitable preservatives include, for example, methylparaben, propylparaben, sodium benzoate, and benzalkonium chloride. Optionally, a mixture of two or more preservatives may be used. The preservative or its mixture is usually present in an amount of about 0.0001% to about 2% of the total weight of the composition.
[0564] In some embodiments, the pharmaceutical composition contains a buffering agent. In some embodiments, suitable buffering agents include, for example, citric acid, sodium citrate, phosphoric acid, potassium phosphate, and various other acids and salts. Optionally, a mixture of two or more buffering agents may be used. The buffering agent or its mixture is usually present in an amount of about 0.001% to about 4% of the total weight of the composition.
[0565] In some embodiments, the concentration of the therapeutic agent in the pharmaceutical composition can be varied, for example, to less than about 1% by weight, or at least about 1%, 2%, 3%, 4%, 5%, 6%, 7%, 8%, 9%, 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, or about 50% or more, and can be selected mainly by the fluid volume and viscosity depending on the particular mode of administration selected.
[0566] The following formulations for oral, aerosol, parenteral (e.g., subcutaneous, intravenous, intraarterial, intramuscular, intradermal, intraperitoneal, and intrathecal), and topical administration are merely illustrative and are in no way limiting. Multiple routes can be used to administer the therapeutic agents described herein, and in some cases, a particular route may result in a more rapid and effective response than other routes.
[0567] Formulations suitable for oral administration can include or consist of (a) liquid solutions such as an effective amount of a therapeutic agent dissolved in a diluent such as water, saline, or orange juice, (b) capsules, sachets, tablets, lozenges, and troches containing a predetermined amount of the active ingredient as a solid or granule, (c) powders, (d) suspensions in a suitable liquid, and (e) suitable emulsions. Liquid formulations can include diluents such as water and alcohols (e.g., ethanol, benzyl alcohol, polyethylene alcohol), with or without the addition of pharmaceutically acceptable surfactants. Capsule forms can be of the ordinary hard or soft shell gelatin type, including, for example, surfactants, lubricants, and inert fillers such as lactose, sucrose, calcium phosphate, and corn starch. Tablet forms can include one or more of lactose, sucrose, mannitol, corn starch, potato starch, alginic acid, microcrystalline cellulose, acacia, gelatin, guar gum, colloidal silicon dioxide, croscarmellose sodium, talc, magnesium stearate, calcium stearate, zinc stearate, stearic acid, other excipients, coloring agents, diluents, buffering agents, disintegrating agents, wetting agents, preservatives, flavoring agents, and other pharmacologically compatible excipients. Lozenge forms can include a flavoring (usually sucrose, acacia, or tragacanth) with the therapeutic agent. Troches can include gelatin and glycerin, or an inert base such as sucrose and acacia, emulsions, gels, and the like, with the therapeutic agent, along with excipients known in the art.
[0568] Formulations suitable for parenteral administration may include aqueous and non-aqueous isotonic sterile injection solutions that may contain antioxidants, buffers, bacteriostats, and solutes that render the solution isotonic with the blood of the intended recipient, as well as aqueous and non-aqueous sterile suspensions that may contain suspending agents, solubilizing agents, thickening agents, stabilizers, and preservatives. In some embodiments, the therapeutic agents provided herein are administered in a sterile liquid or mixture of liquids, including water, physiological saline, aqueous dextrose, and related sugar solutions, alcohols such as ethanol or hexadecyl alcohol, glycols such as propylene glycol or polyethylene glycol, dimethyl sulfoxide, glycerol, ketals such as 2,2-dimethyl-1,3-dioxolan-4-methanol, ethers, poly(ethylene glycol) 400, oils, fatty acids, fatty acid esters or glycerides, or acetylated fatty acid glycerides (with or without the addition of pharmaceutically acceptable surfactants such as soaps or detergents, suspending agents such as pectin, carbomer, methylcellulose, hydroxypropylmethylcellulose, carboxymethylcellulose, or emulsifying agents and other pharmaceutical adjuvants) in a physiologically acceptable diluent in a pharmaceutical carrier.
[0569] In some embodiments, oils that can be used in parenteral formulations include petroleum, animal oils, vegetable oils, or synthetic oils. Specific examples of oils include peanut oil, soybean oil, sesame oil, cottonseed oil, corn oil, olive oil, petrolatum, and mineral oil. Fatty acids suitable for parenteral formulations include oleic acid, stearic acid, and isostearic acid. Examples of suitable fatty acid esters include ethyl oleate and isopropyl myristate.
[0570] Suitable soaps for use in some embodiments of the parenteral formulation include aliphatic alkali metals, ammonium, and triethanolamine salts, and suitable detergents include (a) cationic detergents such as, for example, dimethyldialkylammonium halides and alkylpyridinium halides, (b) anionic detergents such as, for example, alkyl, aryl, olefin sulfonates, alkyl, olefin, ether, monoglyceride sulfates, and sulfosuccinates, (c) nonionic detergents such as, for example, aliphatic amine oxides, fatty acid alkanolamides, and polyoxyethylene polypropylene copolymers, (d) amphoteric detergents such as, for example, alkyl-b-aminopropionates and 2-alkyl-imidazoline quaternary ammonium salts, and (e) mixtures thereof.
[0571] In some embodiments, the parenteral formulation contains the therapeutic agent, for example, in a weight of about 0.5% to about 25% in solution. Preservatives and buffers may be used. To minimize or eliminate irritation at the injection site, such a composition may contain, for example, one or more nonionic surfactants having a hydrophilic-lipophilic balance (HLB) of about 12 to about 17. The amount of surfactant in such a formulation typically ranges, for example, from about 5% to about 15% by weight. Suitable surfactants include polyethylene glycol, sorbitan fatty acid esters such as sorbitan monooleate, and high molecular weight adducts of hydrophobic bases formed by the condensation of propylene oxide and propylene glycol with ethylene oxide. The parenteral formulation can be provided in a sealed container of a unit dose or multiple doses such as an ampoule or vial, and can be stored in a lyophilized (freeze-dried) state by simply adding a sterile liquid excipient such as water for injection immediately before use. Immediate injection solutions and suspensions can be prepared from sterile powders, granules, and tablets of the aforementioned types.
[0572] In some embodiments, injectable formulations are provided herein. The requirements for an effective pharmaceutical carrier for injectable compositions are well known to those of ordinary skill in the art (see, e.g., Pharmaceutics and Pharmacy Practice, J.B. Lippincott Company, Philadelphia, PA, Banker and Chalmers, eds., pages 238-250 (1982) and ASHP Handbook on Injectable Drugs, Toissel, 4th ed, pages 622-630 (1986)).
[0573] In some embodiments, topical formulations are provided herein. Topical formulations, including those useful for transdermal drug delivery, are suitable for application to the skin in the context of the particular embodiments provided herein. In some embodiments, the therapeutic agent can be formulated into an aerosol formulation for administration by inhalation, either alone or in combination with other suitable components. These aerosol formulations can be placed in a pressurized acceptable propellant such as dichlorodifluoromethane, propane, nitrogen, and the like. They can also be formulated as pharmaceuticals for non-pressurized formulations such as nebulizers or atomizers. Such spray formulations can also be used for spraying on mucous membranes.
[0574] In some embodiments, the therapeutic agents provided herein can be formulated as inclusion complexes such as cyclodextrin inclusion complexes, or as liposomes. Liposomes can be useful for targeting the therapeutic agent to specific tissues. Liposomes can also be used to extend the half-life of the therapeutic agent. There are numerous methods for preparing liposomes, for example, as described in Szoka et al, Ann. Rev. Biophys. Bioeng., 9, 467 (1980) and U.S. Patent Nos. 4,235,871, 4,501,728, 4,837,028, and 5,019,369.
[0575] In some embodiments, the therapeutic agents provided herein are formulated in a timed-release, delayed-release, or sustained-release delivery system such that delivery of the composition occurs over a sufficient period of time prior to eliciting sensitization of the site to be treated. Such systems can avoid repeated administration of the therapeutic agent, thereby improving convenience for the subject and the physician, and may be particularly suitable for the specific composition embodiments provided herein. In some embodiments, the compositions of the present disclosure are formulated to be suitable for sustained release (slow release) of the circRNA contained therein. Such sustained-release compositions are convenient to administer to a subject at longer dosing intervals. For example, in some embodiments, the compositions of the present disclosure are administered to a subject twice a day, daily, or every other day. In some embodiments, the compositions of the present disclosure are administered to a subject twice a week, once a week, every 10 days, every two weeks, every three weeks, every four weeks, once a month, every six weeks, every eight weeks, every three months, every four months, every six months, every eight months, every nine months, or annually.
[0576] In some embodiments, the protein encoded by the polynucleotide described herein is produced by the target cells over a period of time. For example, the protein may be produced for 1 hour or more, 4 hours or more, 6 hours or more, 12 hours or more, 24 hours or more, 48 hours or more, or 72 hours or more after administration. In some embodiments, the therapeutic agent is expressed at peak levels at about 6 hours after administration. In some embodiments, the expression of the therapeutic agent is sustained at least at therapeutic levels. In some embodiments, the therapeutic agent is expressed at least at therapeutic levels for 1 hour or more, 4 hours or more, 6 hours or more, 12 hours or more, 24 hours or more, 48 hours or more, or 72 hours or more after administration. In some embodiments, the therapeutic agent is detectable at therapeutic levels in the serum or tissue (e.g., liver or lung) of the patient. In some embodiments, the detectable level of the therapeutic agent is due to continuous expression from the circRNA composition over a period of 1 hour or more, 4 hours or more, 6 hours or more, 12 hours or more, 24 hours or more, 48 hours or more, or 72 hours or more after administration.
[0577] In some embodiments, the protein encoded by the polynucleotide described herein is produced at levels that exceed normal physiological levels. The level of the protein may be increased compared to a control group. In some embodiments, the control is the baseline physiological level of the therapeutic agent in a normal individual or a population of normal individuals. In some embodiments, the control is the baseline physiological level of the therapeutic agent in an individual having a deficiency of the relevant protein or polypeptide, or in a population of individuals having a deficiency of the relevant protein or polypeptide. In some embodiments, the control may be the normal level of the relevant protein or polypeptide in the individual to whom the composition is administered. In some embodiments, the control is the expression level of the therapeutic agent at the time of direct injection of the corresponding therapeutic agent at one or more comparable time points, for example, in other therapeutic interventions.
[0578] In some embodiments, the level of the protein encoded by the polynucleotide described herein is detectable 3, 4, 5 days, or more than 1 week after administration. An increase in the secreted protein level may be observed in serum and / or tissue (e.g., liver or lung).
[0579] In some embodiments, the method results in a prolonged circulating half-life of the protein encoded by the polynucleotide described herein. For example, the protein may be detected for several hours or days longer than the half-life observed by subcutaneous injection of the protein or mRNA encoding the protein. In some embodiments, the half-life of the protein is 1, 2, 3, 4, 5 days, or more than 1 week.
[0580] Many types of release delivery systems are available and are known to those skilled in the art. These include polymeric systems such as poly(lactide-glycolide), copolioxalate, polycaprolactone, polyesteramide, polyorthoester, polyhydroxybutyrate, and polyanhydride. Microcapsules of the aforementioned polymers containing a drug are described, for example, in U.S. Patent No. 5,075,109. Delivery systems include non-polymeric systems that are lipids containing sterols such as cholesterol, cholesterol esters, and fatty acids or neutral fats such as mono, di, triglycerides, hydrogel release systems, silastic systems, peptide-based systems, wax coatings, compressed tablets using conventional binders and excipients, partially fused implants, and the like. Specific examples include, but are not limited to: (a) an erosion system in which an active composition is contained within a matrix as described in U.S. Patent Nos. 4,452,775, 4,667,014, 4,748,034, and 5,239,660; and (b) a diffusion system in which an active ingredient penetrates at a rate controlled by a polymer as described in U.S. Patent Nos. 3,832,253 and 3,854,480. Additionally, a pump-based hardware delivery system can be used. Some of which are adapted for implantation.
[0581] In some embodiments, the therapeutic agent can be conjugated directly or indirectly via a linking moiety to a target moiety. Methods for attaching the therapeutic agent to the target moiety are known in the art. See, for example, Wadwa et al, J, Drug Targeting 3:111 (1995) and U.S. Patent No. 5,087,616.
[0582] In some embodiments, the therapeutic agents provided herein are formulated in depot form (see, e.g., U.S. Patent No. 4,450,150) such that the manner in which the therapeutic agent is released into the administered body is controlled with respect to time and location within the body. The depot form of the therapeutic agent is, for example, an implantable composition comprising the therapeutic agent and a porous or non-porous material such as a polymer, and the therapeutic agent is encapsulated by the material, diffuses throughout the material, and / or involves the degradation of the non-porous material. The depot is then implanted at the desired location in the body, and the therapeutic agent is released from the implant at a predetermined rate.
[0583] 8.7. Methods of Use In some aspects, provided herein are methods for treating and / or preventing, for example, a cancerous condition, the method comprising introducing the pharmaceutical composition provided herein into a subject in need thereof (e.g., a subject suffering from cancer). In some embodiments, the pharmaceutical composition comprises a circular RNA polynucleotide provided herein. In some embodiments, the pharmaceutical composition comprises a vector provided herein. In some embodiments, the pharmaceutical composition comprises a cell (e.g., a human cell) comprising a polynucleotide provided herein (e.g., a circular RNA or vector provided herein).
[0584] Accordingly, in some embodiments, provided herein are methods for treating and / or preventing a disease in a subject (e.g., a mammalian subject such as a human subject). Without being bound by a particular theory or mechanism, the circular RNAs provided herein can be used to express a therapeutic protein for protein replacement therapy.
[0585] In some embodiments, the therapeutic agents provided herein are administered simultaneously with one or more additional therapeutic agents (e.g., as part of the same pharmaceutical composition or as a separate pharmaceutical composition). In some embodiments, the therapeutic agents provided herein can be administered first, followed by administration of one or more additional therapeutic agents, or vice versa. Alternatively, the therapeutic agents provided herein and one or more additional therapeutic agents can be administered simultaneously.
[0586] In some embodiments, the therapeutic agent is a cell or cell population comprising a circular RNA or vector provided herein that expresses a protein encoded by the circular RNA or vector. In some embodiments, the administered cells are of the same species as the subject being treated. In some embodiments, the administered cells are autologous cells to the subject being treated.
[0587] In some embodiments, the subject is a mammal. In some embodiments, the mammals referred to herein can be any mammal including, but not limited to, rodent mammals such as mice and hamsters, or lagomorph mammals such as rabbits. The mammal may belong to the order Carnivora including Felidae (cats) and Canidae (dogs). The mammal may belong to the order Artiodactyla including Bovidae (cows) and Suidae (pigs), or the order Perissodactyla including Equidae (horses). The mammal may belong to the order Primates, the order Ceboids or Simoids (monkeys), or the order Anthropoidea (humans and apes). Preferably, the mammal is a human.
[0588] A therapeutically effective amount of an RNA polynucleotide (e.g., circular RNA) can be administered by various routes, including parenteral administration such as intravenous, intraperitoneal, intramuscular, intrasternal, or intra-articular injection or infusion.
[0589] 8.7.1. Gene Therapy Administering a therapeutically effective amount of an RNA polynucleotide or circular RNA disclosed herein to a subject (e.g., a human) to treat, for example, a disease or disorder associated with protein misfolding and / or protein denaturation diseases, and a disease or disorder associated with a genetic mutation, compositions and methods for treating a disease or disorder are described, wherein the therapeutic agent is involved in the onset, development, and / or expression of the disease or disorder. In some embodiments, the RNA polynucleotide or circular RNA described herein comprises an IRES-like sequence, an endogenous IRES sequence or a variant thereof, or a combination thereof. In some embodiments, the IRES-like sequence, endogenous IRES sequence or a variant thereof, or a combination thereof is optimized for the expression of a therapeutic agent as disclosed herein. In some embodiments, the RNA polynucleotide or circular RNA provided herein can be used, for example, in gene therapy. In some embodiments, the RNA polynucleotide or circular RNA described herein can be used in combination with a CRISPR-Cas system, zinc finger nuclease (ZFN), transcription activator-like effector nuclease (TALEN), meganuclease, and Puf for RNA editing.
[0590] 8.7.2. Vaccine As used herein, a therapeutically effective amount of an RNA polynucleotide or circular RNA disclosed herein is administered to a subject (e.g., a human) to stimulate an immune response against an antigen or agent (e.g., an infectious agent such as a pathogen), generate antibodies against the antigen or agent, destroy the antigen or agent, and / or develop immune memory of the antigen or agent. Compositions and methods are described for this purpose. In some embodiments, the RNA polynucleotides described herein include an IRES-like sequence, an endogenous IRES sequence or variant thereof, or a combination thereof. In some embodiments, the RNA polynucleotides described herein include an expression sequence encoding a therapeutic agent. In some embodiments, the IRES-like sequence, endogenous IRES sequence or variant thereof, or a combination thereof, is optimized for expression of the therapeutic agent in a subject as disclosed herein. In some embodiments, the RNA polynucleotides described herein can be used directly in the manufacture of a vaccine (e.g., an RNA vaccine).
[0591] 8.7.3. Antibody-based Therapy In some embodiments, the therapeutic agent is a polypeptide or protein. In some embodiments, the polypeptide or protein is a disease-causing agent (e.g., a pathogen) that is similar to an attenuated or non-viable form of a microorganism such as a bacterium, virus, fungus, parasite, or one or more components of such a microorganism such as a toxin, protein (e.g., a surface protein), and / or cell wall. In some embodiments, the therapeutic agent is an antigen or agent that can stimulate the body's immune system to recognize the antigen or agent, generate antibodies against the antigen or agent, destroy the antigen or agent, and / or develop immune memory of the antigen or agent. In some embodiments, the therapeutic agent is an antigen or agent that induces and / or enhances vaccine-induced memory and / or enables the immune system to respond quickly and effectively to the antigen or agent upon subsequent encounter.
[0592] In some embodiments, the therapeutic agent is derived from an infectious agent or a part, component, and / or product thereof (e.g., cell wall, genomic sequence, membrane, capsid, protein, lipid, glycan, toxin). In some embodiments, the infectious agent is selected from viruses, bacteria, fungi, protozoa, and helminths. In some embodiments, the infectious agents selected are viruses and bacteria.
[0593] In some embodiments, the infectious agent is a virus selected from the group consisting of adenovirus, herpes simplex type 1, herpes simplex type 2, encephalitis virus, papillomavirus, varicella-zoster virus, Epstein-Barr virus, human cytomegalovirus, human herpesvirus 8, BK virus, JC virus, smallpox, poliovirus, human bocavirus, parvovirus B19, human astrovirus, Norwalk virus, coxsackievirus, hepatitis A virus, hepatitis B virus, hepatitis C virus, hepatitis D virus, hepatitis E virus, rhinovirus, severe acute respiratory syndrome (SARS) virus, yellow fever virus, dengue virus, West Nile virus, rubella virus, human immunodeficiency virus (HIV), influenza virus, Guanarito virus, Furin virus, Lassa fever virus, Machupo virus, Sabia virus, Crimean-Congo hemorrhagic fever virus, Ebola virus, Marburg virus, measles virus, mumps virus, parainfluenza virus, respiratory syncytial virus (RSV), human metapneumovirus, Hendra virus, Nipah virus, rabies virus, rotavirus, Orbivirus, Coltivirus, Banna virus, human enterovirus, hantavirus, West Nile virus, coronavirus, SARS-related coronavirus (SARS-CoV), SARS-CoV-2 virus (COVID-19 related), Middle East respiratory syndrome coronavirus, Japanese encephalitis virus, vesicular exanthema virus, and eastern equine encephalitis.
[0594] In some embodiments, the infectious agent is a bacterium selected from Mycobacterium tuberculosis, Clostridium difficile resistant to clindamycin, Clostridium difficile resistant to fluoroquinolone, methicillin-resistant Staphylococcus aureus (MRSA), multidrug-resistant Enterococcus faecalis, multidrug-resistant Enterococcus faecium, multidrug-resistant Pseudomonas aeruginosa, multidrug-resistant Acinetobacter baumannii, and vancomycin-resistant Staphylococcus aureus (VRSA).
[0595] In some embodiments, the infectious agent is associated with humans, non-human primates, or other animals such as birds, pigs, horses, dogs, cats, rabbits, mice, rats, cows, sheep, goats, and deer.
[0596] As will be appreciated by those skilled in the art, the terms "antibody therapy" or "antibody-based therapy" are used interchangeably herein to refer to a method of treating a disease or disorder in a subject, the method comprising the administration of one or more antibodies.
[0597] In some embodiments, the RNA polynucleotides described herein comprise an expression sequence encoding a therapeutic agent. In some embodiments, the therapeutic agent is a polypeptide or a protein. In some embodiments, the polypeptide or protein is a therapeutic antibody. In some embodiments, the RNA polynucleotides and compositions comprising the RNA polynucleotides described herein can be used in the manufacturing process of an antibody for use in antibody therapy.
[0598] Exemplary therapeutic antibodies include antibodies used in the treatment of cancer. Non-limiting examples of therapeutic antibodies include 131I-tositumomab (follicular lymphoma, B-cell lymphoma, leukemia), 3F8 (neuroblastoma), 8H9, Abagovomab (ovarian cancer), Adecatumumab (prostate cancer and breast cancer), Afutuzumab (lymphoma), Alacizumab pegol, Alemtuzumab (B-cell chronic lymphocytic leukemia, T-cell lymphoma), Amatuximab, AME-133v (follicular lymphoma, cancer), AMG 102 (advanced renal cell carcinoma), Anatumomab mafenatox (non-small cell lung cancer), Apolizumab (solid tumors, leukemia, non-Hodgkin lymphoma, lymphoma), Bavituximab (cancer, viral infections), Bectumomab (non-Hodgkin lymphoma), Belimumab (non-Hodgkin lymphoma), Bevacizumab (colon cancer, breast cancer, brain and central nervous system tumors, lung cancer, hepatocellular carcinoma, kidney cancer, breast cancer, pancreatic cancer, bladder cancer, sarcoma, melanoma, esophageal cancer; gastric cancer, metastatic renal cell carcinoma; kidney cancer, glioblastoma, liver cancer, proliferative diabetic retinopathy, macular degeneration), Bivatuzumab mertansine (squamous cell carcinoma), Blinatumomab, Brentuximab vedotin (blood cancer), Cantuzumab (colon cancer, gastric cancer, pancreatic cancer, non-small cell lung cancer), Cantuzumab mertansine (colorectal cancer), Cantuzumab ravtansine (cancer), Capromab pendetide (prostate cancer), Carlumab, Catumaxomab (ovarian cancer, fallopian tube tumors, peritoneal tumors), Cetuximab (metastatic colorectal cancer and head and neck cancer), Citatuzumab bogatoxBogatox (ovarian cancer and other solid tumors), Cixutumumab (solid tumors), Clivatuzumab tetraxetan (pancreatic cancer), CNTO 328 (B-cell non-Hodgkin lymphoma, multiple myeloma, Castleman disease, ovarian cancer), CNTO 95 (malignant melanoma), Conatumumab, Dacetuzumab (blood cancer), Dalotuzumab, Denosumab (multiple myeloma, giant cell tumor of bone, breast cancer, prostate cancer, osteoporosis), Detumomab (lymphoma), Drozitumab, Ecromeximab (malignant melanoma), Edrecolomab (colon cancer), Elotuzumab (multiple myeloma), Elsilimomab, Enavatuzumab, Ensituximab, Epratuzumab (autoimmune diseases, systemic lupus erythematosus, non-Hodgkin lymphoma, leukemia), Ertumaxomab (breast cancer), Ertumaxomab (breast cancer), Etaracizumab (malignant melanoma, prostate cancer, ovarian cancer), Farletuzumab (ovarian cancer), FBTA05 (chronic lymphocytic leukemia), Ficlatuzumab (cancer), Figitumumab (adrenocortical carcinoma, non-small cell lung cancer), Flanvotumab (malignant melanoma), Galiximab (B-cell lymphoma), Galiximab (non-Hodgkin lymphoma), Ganitumumab, GC1008 (advanced renal cell carcinoma; malignant melanoma, pulmonary fibrosis), Gemtuzumab (leukemia), Gemtuzumab ozogamicin (acute myeloid leukemia), Girentuximab (clear cell renal cell carcinoma), Glembatumumab vedotinvedotin) (melanoma, breast cancer), GS6624 (idiopathic pulmonary fibrosis and solid tumors), HuC242-DM4 (colon cancer, gastric cancer, pancreatic cancer), HuHMFG1 (breast cancer), HuN901-DM1 (myeloma), Ibritumomab (relapsed or refractory low-grade, follicular, or transformed B-cell non-Hodgkin lymphoma (NHL)), Icrucumab, ID09C3 (non-Hodgkin lymphoma), Indatuximab ravtansine, Inotuzumab ozogamicin, Intetumumab (solid tumors (prostate cancer, melanoma)), Ipilimumab (sarcoma, melanoma, lung cancer, ovarian cancer, leukemia, lymphoma, brain and central nervous system tumors, testicular cancer, prostate cancer, pancreatic cancer, breast cancer), Iratumumab (Hodgkin lymphoma), Labetuzumab (colon cancer), Lexatumumab, Lintuzumab, Lorvotuzumab mertansine, Lucatumumab (multiple myeloma, non-Hodgkin lymphoma, Hodgkin lymphoma), Lumiliximab (chronic lymphocytic leukemia), Mapatumumab (colon cancer, myeloma), Matuzumab (lung cancer, cervical cancer, esophageal cancer), MDX-060 (Hodgkin lymphoma, lymphoma), MEDI 522 (solid tumors, leukemia, lymphoma, small intestine cancer, melanoma), Mitumomab (small cell lung cancer), Mogamulizumab, MORab-003 (ovarian cancer, fallopian tube cancer, peritoneal cancer), MORab-009 (pancreatic cancer, mesothelioma, ovarian cancer, non-small cell lung cancer, fallopian tube cancer, peritoneal cavity cancer), Moxetumomab pasudotox, MT103 (non-Hodgkin lymphoma), Nakolomab tafenatox (colon cancer), Naptumomab estafenatox (colon cancer),estafenatox) (non-small cell lung cancer, renal cell carcinoma), Narnatumab, Necitumumab (non-small cell lung cancer), Nimotuzumab (squamous cell carcinoma, head and neck cancer, nasopharyngeal carcinoma, glioma), Nimotuzumab (squamous cell carcinoma, glioma, solid tumor, lung cancer), Olaratumab, Onartuzumab (cancer), Oportuzumab monatox, Oregovomab (ovarian cancer), Oregovomab (ovarian cancer, fallopian tube cancer, peritoneal cavity cancer), PAM4 (pancreatic cancer), Panitumumab (colon cancer, lung cancer, breast cancer; bladder cancer; ovarian cancer), Patritumab, Pemtumomab, Pertuzumab (breast cancer, ovarian cancer, lung cancer, prostate cancer), Pritumumab (brain tumor), Racotumomab, Radretumab, Ramucirumab (solid tumor), Rilotumumab (solid tumor), Rituximab (urticaria, rheumatoid arthritis, ulcerative colitis, chronic encephalitis, non-Hodgkin lymphoma, lymphoma, chronic lymphocytic leukemia), Robatumumab, Samalizumab, SGN-30 (Hodgkin lymphoma, lymphoma), SGN-40 (non-Hodgkin lymphoma, multiple myeloma, leukemia, chronic lymphocytic leukemia), Sibrotuzumab, Siltuximab, Tabalumab (B-cell cancer), Tacatuzumab tetraxetan, Taplitumomab paptoxPaptox), Tenatumomab, Teprotumumab (hematological malignancies), TGN1412 (chronic lymphocytic leukemia, rheumatoid arthritis), Ticilimumab (= Tremelimumab), Tigatuzumab, TNX-650 (Hodgkin lymphoma), Tositumomab (follicular lymphoma, B-cell lymphoma, leukemia, myeloma), Trastuzumab (breast cancer, endometrial cancer, solid tumors), TRBS07 (melanoma), Tremelimumab, TRU-016 (chronic lymphocytic leukemia), TRU-016 (non-Hodgkin lymphoma), Tucotuzumab celmoleukin, Ublituximab, Urelumab, Veltuzumab (non-Hodgkin lymphoma), Veltuzumab (IMMU-106) (non-Hodgkin lymphoma), Volociximab (renal cell carcinoma, pancreatic cancer, melanoma), Votumumab (colorectal tumors), WX-G250 (renal cell carcinoma), Zalutumumab (head and neck cancer, squamous cell carcinoma), and Zanolimumab (T-cell lymphoma).
[0599] Exemplary therapeutic antibodies include antibodies used for the treatment of immune disorders. Non-limiting examples of therapeutic antibodies include efalizumab (psoriasis), epratuzumab (autoimmune diseases, systemic lupus erythematosus, non-Hodgkin lymphoma, leukemia), etrolizumab (inflammatory bowel disease), fontolizumab (Crohn's disease), ixekizumab (autoimmune diseases), mepolizumab (eosinophilic granulomatosis with polyangiitis, asthma, eosinophilic gastroenteritis, Churg-Strauss syndrome, eosinophilic esophagitis), milatuzumab (multiple myeloma and other hematological malignancies), pooled immunoglobulin (primary immunodeficiency), priliximab (Crohn's disease, multiple sclerosis), rituximab (urticaria, rheumatoid arthritis, ulcerative colitis, chronic encephalitis, non-Hodgkin lymphoma, lymphoma, chronic lymphocytic leukemia), rontalizumab (systemic lupus erythematosus), ruplizumab (rheumatic diseases), sarilumab (rheumatoid arthritis, ankylosing spondylitis), vedolizumab (Crohn's disease, ulcerative colitis), visilizumab (Crohn's disease, ulcerative colitis), reslizumab (inflammation of the airway, skin, and gastrointestinal tract), adalimumab (rheumatoid arthritis, Crohn's disease, ankylosing spondylitis, psoriatic arthritis), aselizumab (critically ill patients), atinumab (treatment of the nervous system), atlizumab (rheumatoid arthritis, systemic juvenile idiopathic arthritis), bertilimumab (severe allergic diseases), besilesomab (inflammatory lesions and metastases), BMS-945429, ALD518 (cancer and rheumatoid arthritis), briakinumab (psoriasis, rheumatoid arthritis, inflammatory bowel disease, multiple sclerosis), brodalumab (inflammatory diseases), canakinumab (rheumatoid arthritis), canakinumab (cryopyrin-associated periodic syndromes...
Claims
1. RNA polynucleotide comprising the following constructs of formula I, formula II, formula III, formula IV, or formula V: 5'-(A1) 0-1 -(L) n -TI-(L) n -Z1- (L) n - (B1) 0-1 -3'(I), 5'-(A1) 0-1 -(L) n -Z1 B -(L) n -TI-(L) n -Z1 A- (L) n - (B1) 0-1 -3'(II), 5'-(3' intron fragment)-(E2)-(L) n -TI-(L) n -Z1-(L) n - (E1) - (5' intron fragment) - 3' (III), 5'-(3' intron fragment)-(E2)-(L) n -Z1 B - (L) n -TI-(L) n -Z1 A- - (L) n -(E1)-(5' intron fragment)-3'(IV), or 5'-TI-(L) n -Z1-3'(V) And, TI is a modified translation initiation element containing an IRES-like polynucleotide sequence, wherein the IRES-like polynucleotide sequence includes a nucleic acid sequence that is approximately 90%, 95%, 97%, 98%, 99%, or 100% identical to SEQ ID NOs. 1025-14161 or SEQ ID NOs. 14412-15341, or at least approximately 90%, 95%, 97%, 98%, 99%, or 100% identical. Z1 is an expression sequence that codes for a therapeutic drug. Z1 A This is the first part of the expression sequence that codes for the therapeutic agent, Z1 B This is the second part of the expression sequence that codes for the therapeutic agent, Each L is an independent linker sequence. A1 and B1 are sequences that can independently cyclically form the RNA polynucleotide, or A1 and B1 each independently contain a nucleotide derivative capable of linking the 5' end and the 3' end via a phosphodiester bond from 3' to 5' in order to cyclically form the RNA polynucleotide. The 5' intron fragment and the 3' intron fragment are each fragments of a group II intron, and the 5' intron fragment is located on the 5' side of the 3' intron fragment within the group II intron. E1 is a 5' adjacent exon fragment of the group II intron, having a length of ≥0 nucleotides. E2 is a 3' adjacent exon fragment of the group II intron having a length of ≥0 nucleotides, and n is an integer selected from 0 to 2. RNA polynucleotides.
2. i) A 5' homology arm is provided at the 5' end of the 3' intron fragment, ii) A 3' homology arm at the 3' end of the 5' intron fragment, or iii) A 5' homology arm at the 5' end of the 3' intron fragment and a 3' homology arm at the 3' end of the 5' intron fragment The RNA polynucleotide according to claim 1, further comprising:
3. The RNA polynucleotide according to claim 1, wherein E1 and E2 are each independently 0 to 20 nucleotides in length.
4. The RNA polynucleotide according to claim 1, wherein the 5' intron fragment and the 3' intron fragment are obtained by splitting a group II intron into two fragments at an unpaired region, the unpaired region is preferably selected from a linear region between two adjacent domains of the group II intron and a loop region of a stem-loop structure of domain 4 of the group II intron.
5. RNA polynucleotide comprising a modified translation initiation element (TI) containing an IRES-like polynucleotide sequence, wherein the IRES-like polynucleotide sequence comprises a nucleic acid sequence that is approximately 90%, 95%, 97%, 98%, 99%, or 100% identical to SEQ ID NOs. 1025-14161 or SEQ ID NOs. 14412-15341, or at least approximately 90%, 95%, 97%, 98%, 99%, or 100% identical.
6. The RNA polynucleotide according to claim 1, wherein the length of the IRES-like polynucleotide sequence is 6 to 12 residues.
7. The RNA polynucleotide according to claim 1, wherein the TI further comprises a second IRES-like polynucleotide sequence.
8. The RNA polynucleotide according to claim 7, wherein the second IRES-like polynucleotide sequence includes nucleic acid sequences that are approximately 90%, 95%, 97%, 98%, 99%, or 100% identical to SEQ ID NOs. 1025-14161 or SEQ ID NOs. 14412-15341, or at least approximately 90%, 95%, 97%, 98%, 99%, or 100% identical.
9. i) Each L independently comprises a 5'UTR, a 3'UTR, a polyA sequence, a polyA-C sequence, a polyC sequence, a polyU sequence, a polyG sequence, a ribosome binding site, an aptamer, a riboswitch, a ribozyme, a small RNA (small RNA) binding site, a translation regulatory element (e.g., a Kozak sequence), a protein binding site (e.g., PTBP1 or HUR), a non-natural nucleotide, or a non-nucleotide chemical linker. ii) Each L independently has a length of approximately 3 to approximately 100 nucleotide residues, and / or iii) Each L independently contains the nucleic acid sequence of RCC, and R is guanine or adenine. The RNA polynucleotide according to claim 1.
10. The RNA polynucleotide according to claim 1, wherein the RNA polynucleotide is a single-stranded RNA polynucleotide, a cyclic RNA polynucleotide, and / or a linear RNA polynucleotide.
11. The RNA polynucleotide according to claim 1, which can be cyclized in the absence of an enzyme.
12. A polypeptide expressed by the RNA polynucleotide described in claim 1.
13. A DNA vector encoding an RNA polynucleotide as described in claim 1.
14. A cell comprising the RNA polynucleotide described in claim 1.
15. A composition comprising the RNA polynucleotide described in claim 1 and a pharmaceutically acceptable carrier.
16. A method for producing a population of cells, comprising contacting the cells of the population with the RNA polynucleotide described in claim 1.
17. A method for generating an internal ribosome entry site (IRES)-like polynucleotide sequence, (a) A step of generating a polynucleotide query sequence having a length of X nucleic acid residues, wherein X is an integer of 3 or more. (b) A step of generating X-Y+1 duplicate polynucleotide fragment sequences within the polynucleotide query sequence, Each polynucleotide fragment sequence consists of Y nucleic acid residues in length. The starting position of each polynucleotide fragment sequence is n, the ending position of the same polynucleotide fragment sequence is Y+n-1, and n represents each positive integer between 1 and X - Y + 1. (c) A step of determining the enrichment score of each polynucleotide fragment sequence in (b), (d)(c) A step of determining the numerical score of the polynucleotide query sequence by summing the enrichment scores of each polynucleotide fragment sequence, and (e) A step of identifying the polynucleotide query sequence as a modified IRES-like polynucleotide sequence according to a reference value. Methods that include...
18. RNA polynucleotide comprising a modified translation initiation element (TI), wherein the TI comprises an internal ribosome entry site (IRES)-like polynucleotide sequence generated by the method of claim 17.
19. A method for regulating the expression of a protein in a subject that requires regulation of protein expression, comprising administering a therapeutically effective amount of the polynucleotide described in claim 1 to the subject.
20. A method for treating or preventing a disease or disorder in a subject requiring treatment or prevention of the said disease or disorder, comprising administering a therapeutically effective amount of the polynucleotide described in claim 1 to the subject.