Synthetic circular RNA compositions and methods of use thereof
By generating and identifying engineered translation initiation elements with IRES-like polynucleotide sequences, combined with the design of circular RNA, the difficulties in the development and prediction of IRES components in the prior art are solved, and efficient fine-tuning control and stability improvement of protein expression are achieved.
Patent Information
- Application Number
- CN202380044375.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2022-05-30
- Filing Date
- 2023-05-29
- Publication Date
- 2025-06-03
AI Technical Summary
It is difficult to effectively develop and predict novel synthetic IRES elements in the prior art, and natural IRES elements have shortcomings in protein expression regulation and immune rejection.
By generating and identifying engineered translation initiation elements with IRES-like polynucleotide sequences, combined with the design of circular RNA, fine-tuning control and prediction of protein expression is achieved.
It provides stable and functional RNA polynucleotides that can be applied in gene therapy and gene vaccination, improves the efficiency and stability of protein expression, and reduces the risk of immune rejection.
Smart Images

Figure CN120092091A_ABST
Abstract
Description
[0001] 1. Related Applications
[0002] This application claims the priority and benefit of PCT application PCT / CN2022 / 095949, filed on May 30, 2022, the content of which is incorporated herein by reference in its entirety.
[0003] 2. Incorporation of Sequence Listing by Reference
[0004] The content of the electronic sequence listing (Family 03_sql_0529.xml; size: 12831515 bytes; creation date: May 29, 2023) is incorporated herein by reference in its entirety. 3. Technical Field
[0005] The present disclosure relates to compositions of matter, methods, processes, kits, and devices for selecting, designing, preparing, manufacturing, formulating, and / or using polynucleotides having an internal ribosome entry site (IRES) sequence, an IRES-like sequence, or a combination thereof. The present disclosure also relates to compositions of matter, methods, processes, kits, and devices for selecting, designing, preparing, manufacturing, formulating, and / or using circular polynucleotides (e.g., circular RNAs) comprising an IRES sequence, an IRES-like sequence, or a combination thereof. In addition, the present disclosure relates to methods for improving the expression, functional stability, immunogenicity, manufacture, and / or half-life of therapeutic products encoded by said circular RNAs. 4. Background Art
[0006] Gene therapy and gene vaccination provide highly specific and individualized treatment options for various diseases such as genetic genetic diseases, autoimmune diseases, cancer, and inflammatory diseases. In the context of gene therapy and gene vaccination, DNA and RNA can be used as nucleic acid molecules for administration. Although DNA therapy is stable and easy to manipulate, there is a risk of adverse consequences such as genomic integration and the generation of anti-DNA antibodies. Additionally, due to the dependence on the presence of specific transcription factors that regulate DNA transcription, the expression of the encoded protein may be limited. In the absence of such factors, DNA transcription is hindered, leading to low levels of the translated protein.
[0007] By using RNA instead of DNA gene therapy, the risk of adverse consequences such as genomic integration and the generation of anti-DNA antibodies can be minimized or avoided. In particular, circular RNAs can be used to design and generate stable forms of RNA. Circular RNAs can also be used for in vivo applications, especially in the fields of RNA-based gene expression control, protein manufacture, and therapeutics, including protein replacement therapy and vaccination.
[0008] Translation of circular RNAs is promoted by cap-independent translation. Thus, designing and selecting appropriate cap-independent translation initiation elements is important for controlling protein expression. Internal ribosome entry site (IRES) elements can be used for cap-independent gene expression in eukaryotic cells. However, naturally-occurring IRES elements may 1) have a large length of nucleic acid residues, 2) contain complex secondary structures, and / or 3) be prone to host cell immune rejection, each of which is not preferred for in vivo applications such as protein replacement therapy.
[0009] Although many natural IRES sequences that show promotion of cap-independent translation have been identified, the identification and development of novel IRES elements have been serendipitous and accidental. Natural IRES elements do not allow for fine-tuning control of protein expression, and there are few systematic methods for predicting functional novel IRES elements. No systematic method has emerged for identifying and developing synthetically-generated IRES elements that can provide efficient cap-independent protein expression.
[0010] Accordingly, there is a need in the art for methods and approaches for developing synthetic IRES elements, as well as methods for predicting and controlling the efficiency of protein expression from synthetic IRES elements. There is also a need in the art to provide polynucleotides (e.g., circular RNAs) having synthetic IRES elements that are suitable for use as drugs or vaccines for applications such as in gene therapy and / or gene vaccination. The present disclosure addresses these unmet needs. 5. Summary of the Invention
[0012] Compositions and methods of matter (e.g., synthetic internal ribosome entry site (IRES) sequences, IRES-like sequences, or combinations thereof, and circular RNAs) are described for selecting, designing, preparing, manufacturing, formulating, and / or using polynucleotides having an IRES sequence, IRES-like sequence, or combination thereof, and circular RNAs as described herein.
[0013] Provided herein are RNA polynucleotides comprising a Form I construct:
[0014] 5’-(A1) 0-1 -(L) n -TI-(L) n -Z1-(L) n -(B1) 0-1 -3’(I)
[0015] Wherein:
[0016] TI is an engineered translation initiation element comprising an IRES-like polynucleotide sequence, wherein said IRES-like polynucleotide sequence comprises a nucleic acid sequence that is about or at least about 90%, 95%, 97%, 98%, 99% or 100% identical to SEQ ID NO: 1025 - 14161 or SEQ ID NO: 14412 - 15341;
[0017] Z1 is an expression sequence encoding a therapeutic product;
[0018] Each L is independently a linker sequence;
[0019] A1 and B1 are each independently a sequence capable of cyclizing said RNA polynucleotide, or
[0020] A1 and B1 each independently comprise a nucleotide derivative capable of joining the 5'-end and the 3'-end by a 3' to 5' phosphodiester linkage to cyclize said RNA polynucleotide; and
[0021] n is an integer selected from 0 to 2.
[0022] The present disclosure also provides an RNA polynucleotide comprising a construct of Formula II:
[0023] 5’-(A1) 0-1 -(L) n -Z1 B -(L) n -TI-(L) n -Z1 A- (L) n -(B1) 0-1 -3’(II)
[0024] wherein:
[0025] TI is an engineered translation initiation element comprising an IRES-like polynucleotide sequence, wherein said IRES-like polynucleotide sequence comprises a nucleic acid sequence that is about or at least about 90%, 95%, 97%, 98%, 99% or 100% identical to SEQ ID NO: 1025 - 14161 or SEQ ID NO: 14412 - 15341;
[0026] Z1 A is the first part of an expression sequence encoding a therapeutic product;
[0027] Z1 B is the second part of an expression sequence encoding a therapeutic product;
[0028] Each L is independently a linker sequence;
[0029] A1 and B1 are each independently a sequence capable of cyclizing the RNA polynucleotide, or
[0030] A1 and B1 each independently comprise nucleotide derivatives capable of linking the 5'-end and the 3'-end through a 3'-to-5' phosphodiester linkage to cyclize the RNA polynucleotide; and
[0031] n is an integer selected from 0 to 2.
[0032] The present disclosure also provides an RNA polynucleotide comprising a construct of Formula III:
[0033] 5’-(3’ intron fragment)-(E2)-(L) n -TI-(L) n -Z1-(L) n -(E1)-(5’ intron fragment)-3’(III) wherein:
[0034] TI is an engineered translation initiation element comprising an IRES-like polynucleotide sequence, wherein the IRES-like polynucleotide sequence comprises a nucleic acid sequence that is about or at least about 90%, 95%, 97%, 98%, 99% or 100% identical to SEQ ID NO: 1025-14161 or SEQ ID NO: 14412-15341;
[0035] Z1 is an expression sequence encoding a therapeutic product;
[0036] Each L is independently a linker sequence;
[0037] The 5' intron fragment and the 3' intron fragment are each a fragment of a type II intron, wherein the 5' intron fragment is on the 5' side of the 3' intron fragment in the type II intron;
[0038] E1 is the 5’-adjacent exon fragment of the type II intron, and its length ≥ 0 nucleotides;
[0039] E2 is the 3’-adjacent exon fragment of the type II intron, and its length ≥ 0 nucleotides; and
[0040] n is an integer selected from 0 to 2.
[0041] The present disclosure also provides an RNA polynucleotide comprising a construct of Formula IV:
[0042] 5’-(3’ intron fragment)-(E2)-(L) n -Z1 B -(L) n -TI-(L) n -Z1 A- -(L)n -(E1)-(5'-intron fragment)-3'(IV)
[0043] Wherein:
[0044] TI is an engineered translation initiation element comprising an IRES-like polynucleotide sequence, wherein the IRES-like polynucleotide sequence comprises a nucleic acid sequence that is about or at least about 90%, 95%, 97%, 98%, 99% or 100% identical to SEQ ID NO: 1025-14161 or SEQ ID NO: 14412-15341;
[0045] Z1 A is the first part of the expression sequence encoding a therapeutic product;
[0046] Z1 B is the second part of the expression sequence encoding a therapeutic product;
[0047] Each L is independently a linker sequence;
[0048] The 5'-intron fragment and the 3'-intron fragment are each a fragment of a type II intron, wherein the 5'-intron fragment is located 5' to the 3'-intron fragment in the type II intron;
[0049] The E1 is the 5'-adjacent exon fragment of the type II intron, and its length ≥ 0 nucleotides;
[0050] The E2 is the 3'-adjacent exon fragment of the type II intron, and its length ≥ 0 nucleotides; and
[0051] n is an integer selected from 0 to 2.
[0052] The present invention also provides an RNA polynucleotide comprising the construct of formula V,
[0053] 5'-TI-(L) n -Z1-3'(V)
[0054] Wherein:
[0055] TI is an engineered translation initiation element comprising an IRES-like polynucleotide sequence, wherein the IRES-like polynucleotide sequence comprises a nucleic acid sequence that is about or at least about 90%, 95%, 97%, 98%, 99% or 100% identical to SEQ ID NO: 1025-14161 or SEQ ID NO: 14412-15341;
[0056] Z1 is the expression sequence encoding a therapeutic product;
[0057] Each L is independently a linker sequence; and
[0058] n is an integer selected from 0 to 2.
[0059] The present disclosure also provides the following RNA polynucleotides, which comprise an engineered translation initiation element (TI) containing an IRES-like polynucleotide sequence, wherein the IRES-like polynucleotide sequence comprises a nucleic acid sequence that is about or at least about 90%, 95%, 97%, 98%, 99% or 100% identical to SEQ ID NO: 1025-14161 or SEQ ID NO: 14412-15341.
[0060] The present disclosure also provides an RNA polynucleotide comprising a construct of Formula I, Formula II, Formula III, Formula IV or Formula V:
[0061] 5’-(A1) 0-1 -(L) n -TI-(L) n -Z1-(L) n -(B1) 0-1 -3’(I),
[0062] 5’-(A1) 0-1 -(L) n -Z1 B -(L) n -TI-(L) n -Z1 A- (L) n -(B1) 0-1 -3’(II),
[0063] 5’-(3’ intron fragment)-(E2)-(L) n -TI-(L) n -Z1-(L) n -(E1)-(5’ intron fragment)-3’(III),
[0064] 5’-(3’ intron fragment)-(E2)-(L) n -Z1 B -(L) n -TI-(L) n -Z1 A- -(L) n -(E1)-(5’ intron fragment)-3’(IV), or 5’-TI-(L) n -Z1-3’(V),
[0065] Wherein:
[0066] TI is an engineered translation initiation element comprising an IRES-like polynucleotide sequence, wherein the IRES-like polynucleotide sequence comprises a nucleic acid sequence that is about or at least about 90%, 95%, 97%, 98%, 99% or 100% identical to SEQ ID NO: 1025-14161 or SEQ ID NO: 14412-15341;
[0067] Z1 is an expression sequence encoding a therapeutic product;
[0068] Z1 A is the first part of the expression sequence encoding a therapeutic product;
[0069] Z1 B is the second part of the expression sequence encoding a therapeutic product;
[0070] Each L is independently a linker sequence;
[0071] A1 and B1 are each independently a sequence capable of cyclizing the RNA polynucleotide, or
[0072] A1 and B1 each independently comprise nucleotide derivatives capable of ligating the 5'-end and 3'-end by a 3' to 5' phosphodiester linkage to cyclize the RNA polynucleotide;
[0073] The 5'-intron fragment and the 3'-intron fragment are each a fragment of a type II intron, wherein the 5'-intron fragment is located on the 5'-side of the 3'-intron fragment in the type II intron;
[0074] E1 is the 5'-adjacent exon fragment of the type II intron, and its length ≥ 0 nucleotides;
[0075] E2 is the 3'-adjacent exon fragment of the type II intron, and its length ≥ 0 nucleotides; and
[0076] n is an integer selected from 0 to 2.
[0077] Provided herein is a method for generating an internal ribosome entry site (IRES)-like polynucleotide sequence, the method comprising the following steps:
[0078] (a) Generating a polynucleotide query sequence consisting of X nucleic acid residues in length, wherein X is an integer greater than or equal to 3;
[0079] (b) Generating X - Y + 1 overlapping polynucleotide fragment sequences within the polynucleotide query sequence,
[0080] wherein each polynucleotide fragment sequence consists of Y nucleic acid residues in length,
[0081] where the first position of each polynucleotide fragment sequence is n and the last position of the same polynucleotide fragment sequence is Y + n - 1, and
[0082] where n represents each positive integer between 1 and X - Y + 1;
[0083] (c) determining the enrichment score of each polynucleotide fragment sequence of (b);
[0084] (d) determining the numerical score of the polynucleotide query sequence by summing the enrichment scores of each polynucleotide fragment sequence of (c); and
[0085] (e) identifying the polynucleotide query sequence as an engineered IRES-like polynucleotide sequence based on a reference value.
[0086] Provided herein is a method for determining the enrichment score of a polynucleotide fragment sequence, the method comprising
[0087] i) generating a library of expression plasmids, wherein each expression plasmid in the library contains a different polynucleotide fragment sequence and a reporter gene;
[0088] ii) contacting a cell population with the library of expression plasmids;
[0089] iii) quantifying the expression level of the reporter gene corresponding to each expression plasmid of the library of expression plasmids;
[0090] iv) dividing the total cell population into a first population and a second population based on the protein expression level of iii);
[0091] v) determining the enrichment score of the polynucleotide fragment sequence using the following system of equations:
[0092]
[0093] where f 1 is the frequency of the polynucleotide fragment sequence in the first population,
[0094] where f 2 is the frequency of the same polynucleotide fragment sequence in the second population,
[0095] where N 1 is the size of the first population, and
[0096] where N 2 is the size of the second population.
[0097] The present disclosure provides an RNA polynucleotide comprising an engineered translation initiation element (TI), wherein the TI comprises an internal ribosome entry site (IRES)-like polynucleotide sequence generated by the methods of the present disclosure.
[0098] The present disclosure provides a polypeptide expressed by the RNA polynucleotide of the present disclosure, a DNA vector encoding or suitable for synthesizing the RNA polynucleotide of the present disclosure, a cell comprising the RNA polynucleotide, polypeptide, or DNA vector of the present disclosure, and a composition comprising the RNA polynucleotide, polypeptide, DNA vector, or cell of the present disclosure and a pharmaceutically acceptable carrier, and methods for their preparation.
[0099] The present disclosure provides a method of modulating protein expression in a subject in need thereof, the method comprising administering to the subject a therapeutically effective amount of the RNA polynucleotide, polypeptide, DNA vector, or cell of the present disclosure.
[0100] The present disclosure provides a method of treating or preventing a disease or disorder in a subject in need thereof, the method comprising administering to the subject a therapeutically effective amount of the RNA polynucleotide, polypeptide, DNA vector, or cell of the present disclosure.
[0101] The present disclosure provides the use of the RNA polynucleotide, polypeptide, DNA vector, or cell of the present disclosure in the manufacture of a medicament for modulating protein expression or treating or preventing a disease or disorder in a subject in need thereof.
[0102] The present disclosure provides the RNA polynucleotide, polypeptide, DNA vector, or cell of the present disclosure for modulating the expression of a protein or treating or preventing a disease or disorder in a subject in need thereof.
[0103] In some embodiments, the nucleic acids (e.g., polynucleotides) and nucleic acid sequences disclosed herein can be codon-optimized (e.g., for the expression of a therapeutic product), for example, by any codon-optimization technique known in the art (see, e.g., Quax et al., 2015, Mol Cell 59:149-161).
[0104] 6. Illustrative Embodiments
[0105] Embodiment 1
[0106] An RNA polynucleotide comprising a construct of Formula I:
[0107] 5’-(A1) 0-1 -(L) n -TI-(L) n -Z1-(L) n -(B1) 0-1-3’(I)
[0108] Wherein:
[0109] TI is an engineered translation initiation element comprising an IRES-like polynucleotide sequence, wherein said IRES-like polynucleotide sequence comprises a nucleic acid sequence that is about or at least about 90%, 95%, 97%, 98%, 99% or 100% identical to SEQ ID NO: 1025-14161 or SEQ ID NO: 14412-15341;
[0110] Z1 is an expression sequence encoding a therapeutic product;
[0111] Each L is independently a linker sequence;
[0112] A1 and B1 are each independently a sequence capable of cyclizing said RNA polynucleotide, or
[0113] A1 and B1 each independently comprise a nucleotide derivative capable of ligating the 5' and 3' ends via a 3' to 5' phosphodiester linkage to cyclize said RNA polynucleotide; and
[0114] n is an integer selected from 0 to 2.
[0115] Embodiment 1-1
[0116] An RNA polynucleotide comprising a construct of Formula I-1:
[0117] 5’-(A1) 0-1 -(L) n -TI-(L) n -Z1-(L) n -(B1) 0-1 -3’(I-1)
[0118] Wherein:
[0119] TI is an engineered translation initiation element comprising an IRES-like polynucleotide sequence, wherein said IRES-like polynucleotide sequence comprises a nucleic acid sequence that is about or at least about 90%, 95%, 97%, 98%, 99% or 100% identical to SEQ ID NO: 1025-14161 or SEQ ID NO: 14412-15341;
[0120] Z1 is an expression sequence encoding a therapeutic product;
[0121] Each L is independently a linker sequence;
[0122] A1 and B1 are each independently a sequence capable of cyclizing said RNA polynucleotide; and
[0123] n is an integer selected from 0 to 2.
[0124] Embodiments 1-2
[0125] An RNA polynucleotide comprising a construct of Formula I-2:
[0126] 5’-(A1) 0-1 -(L) n -TI-(L) n -Z1-(L) n -(B1) 0-1 -3’(I-2)
[0127] Wherein:
[0128] TI is an engineered translation initiation element comprising an IRES-like polynucleotide sequence, wherein the IRES-like polynucleotide sequence comprises a nucleic acid sequence that is about or at least about 90%, 95%, 97%, 98%, 99% or 100% identical to SEQ ID NO: 1025-14161 or SEQ ID NO: 14412-15341;
[0129] Z1 is an expression sequence encoding a therapeutic product;
[0130] Each L is independently a linker sequence;
[0131] A1 and B1 each independently comprise nucleotide derivatives capable of ligating the 5' and 3' ends by a 3' to 5' phosphodiester linkage to circularize the RNA polynucleotide; and
[0132] n is an integer selected from 0 to 2.
[0133] Embodiment 2
[0134] An RNA polynucleotide comprising a construct of Formula II:
[0135] 5’-(A1) 0-1 -(L) n -Z1 B -(L) n -TI-(L) n -Z1 A- (L) n -(B1) 0-1 -3’(II)
[0136] Wherein:
[0137] TI is an engineered translation initiation element comprising an IRES-like polynucleotide sequence, wherein said IRES-like polynucleotide sequence comprises a nucleic acid sequence that is about or at least about 90%, 95%, 97%, 98%, 99% or 100% identical to SEQ ID NO: 1025-14161 or SEQ ID NO: 14412-15341;
[0138] Z1 A is the first part of an expression sequence encoding a therapeutic product;
[0139] Z1 B is the second part of an expression sequence encoding a therapeutic product;
[0140] Each L is independently a linker sequence;
[0141] A1 and B1 are each independently a sequence capable of circularizing said RNA polynucleotide, or
[0142] A1 and B1 each independently comprise nucleotide derivatives capable of joining the 5' and 3' ends by 3' to 5' phosphodiester linkages to circularize said RNA polynucleotide; and
[0143] n is an integer selected from 0 to 2.
[0144] Embodiment 2-1
[0145] An RNA polynucleotide comprising a construct of Formula II-1:
[0146] 5’-(A1) 0-1 -(L) n -Z1 B -(L) n -TI-(L) n -Z1 A- (L) n -(B1) 0-1 -3’(II-1)
[0147] Wherein:
[0148] TI is an engineered translation initiation element comprising an IRES-like polynucleotide sequence, wherein said IRES-like polynucleotide sequence comprises a nucleic acid sequence that is about or at least about 90%, 95%, 97%, 98%, 99% or 100% identical to SEQ ID NO: 1025-14161 or SEQ ID NO: 14412-15341;
[0149] Z1 A is the first part of an expression sequence encoding a therapeutic product;
[0150] Z1B is the second part of the expression sequence encoding a therapeutic product;
[0151] Each L is independently a linker sequence;
[0152] A1 and B1 are each independently a sequence capable of cyclizing the RNA polynucleotide; and
[0153] n is an integer selected from 0 to 2.
[0154] Embodiment 2-2
[0155] An RNA polynucleotide comprising a construct of formula II-2:
[0156] 5’-(A1) 0-1 -(L) n -Z1 B -(L) n -TI-(L) n -Z1 A- (L) n -(B1) 0-1 -3’(II-2)
[0157] Wherein:
[0158] TI is an engineered translation initiation element comprising an IRES-like polynucleotide sequence, wherein the IRES-like polynucleotide sequence comprises a nucleic acid sequence that is about or at least about 90%, 95%, 97%, 98%, 99% or 100% identical to SEQ ID NO: 1025-14161 or SEQ ID NO: 14412-15341;
[0159] Z1 A is the first part of the expression sequence encoding a therapeutic product;
[0160] Z1 B is the second part of the expression sequence encoding a therapeutic product;
[0161] Each L is independently a linker sequence;
[0162] A1 and B1 each independently comprise nucleotide derivatives capable of joining the 5' and 3' ends by a 3' to 5' phosphodiester linkage to cyclize the RNA polynucleotide; and
[0163] n is an integer selected from 0 to 2.
[0164] Embodiment 3
[0165] An RNA polynucleotide comprising a construct of formula III:
[0166] 5’-(3’ intron fragment)-(E2)-(L) n -TI-(L) n -Z1-(L) n -(E1)-(5’ intron fragment)-3’(III)
[0167] Wherein:
[0168] TI is an engineered translation initiation element comprising an IRES-like polynucleotide sequence, wherein the IRES-like polynucleotide sequence comprises a nucleic acid sequence that is about or at least about 90%, 95%, 97%, 98%, 99% or 100% identical to SEQ ID NO: 1025 - 14161 or SEQ ID NO: 14412 - 15341;
[0169] Z1 is an expression sequence encoding a therapeutic product;
[0170] Each L is independently a linker sequence;
[0171] The 5' intron fragment and the 3' intron fragment are each a fragment of a type II intron, wherein the 5' intron fragment is located 5' to the 3' intron fragment in the type II intron;
[0172] E1 is a 5'-adjacent exon fragment of the type II intron, having a length ≥ 0 nucleotides;
[0173] E2 is a 3'-adjacent exon fragment of the type II intron, having a length ≥ 0 nucleotides; and
[0174] n is an integer selected from 0 to 2.
[0175] Embodiment 4
[0176] An RNA polynucleotide comprising a construct of formula IV:
[0177] 5’-(3’ intron fragment)-(E2)-(L) n -Z1 B -(L) n -TI-(L) n -Z1 A- -(L) n -(E1)-(5’ intron fragment)-3’(IV)
[0178] Wherein:
[0179] TI is an engineered translation initiation element comprising an IRES-like polynucleotide sequence, wherein the IRES-like polynucleotide sequence comprises a nucleic acid sequence that is about or at least about 90%, 95%, 97%, 98%, 99% or 100% identical to SEQ ID NO: 1025-14161 or SEQ ID NO: 14412-15341;
[0180] Z1 A is the first part of an expression sequence encoding a therapeutic product;
[0181] Z1 B is the second part of an expression sequence encoding a therapeutic product;
[0182] Each L is independently a linker sequence;
[0183] The 5' intron fragment and the 3' intron fragment are each a fragment of a type II intron, wherein the 5' intron fragment is located 5' to the 3' intron fragment in the type II intron;
[0184] The E1 is the 5'-flanking exon fragment of the type II intron and has a length ≥ 0 nucleotides;
[0185] The E2 is the 3'-flanking exon fragment of the type II intron and has a length ≥ 0 nucleotides; and
[0186] n is an integer selected from 0 to 2.
[0187] Embodiment 5
[0188] The RNA polynucleotide according to Embodiment 3 or 4, wherein the RNA polynucleotide further comprises a 5' homologous arm at the 5' end of the 3' intron fragment.
[0189] Embodiment 6
[0190] The RNA polynucleotide according to Embodiment 3 or 4, wherein the RNA polynucleotide further comprises a 3' homologous arm at the 3' end of the 5' intron fragment.
[0191] Embodiment 7
[0192] The RNA polynucleotide according to Embodiment 3 or 4, wherein the RNA polynucleotide further comprises a 5' homologous arm at the 5' end of the 3' intron fragment and a 3' homologous arm at the 3' end of the 5' intron fragment.
[0193] Embodiment 8
[0194] An RNA polynucleotide according to any one of embodiments 3-7, wherein the lengths of said E1 and said E2 are each independently 0 to 20 nucleotides, preferably 0 to 10 nucleotides, such as 0, 1, 2, 3, 4, 5, 6, 7, 8, 9 or 10 nucleotides.
[0195] Embodiment 9
[0196] An RNA polynucleotide according to any one of embodiments 3-8, wherein said 5'-intron fragment and said 3'-intron fragment are obtained by splitting a group II intron into two fragments at an unpaired region, and wherein said unpaired region is preferably selected from the linear region between two adjacent domains of said group II intron and the loop region of the stem-loop structure of domain 4 of said group II intron.
[0197] Embodiment 10
[0198] An RNA polynucleotide comprising a construct of formula IV
[0199] 5'-TI-(L) n -Z1-3'(IV)
[0200] wherein:
[0201] TI is an engineered translation initiation element comprising an IRES-like polynucleotide sequence, wherein said IRES-like polynucleotide sequence comprises a nucleic acid sequence that is about or at least about 90%, 95%, 97%, 98%, 99% or 100% identical to SEQ ID NO: 1025-14161 or SEQ ID NO: 14412-15341;
[0202] Z1 is an expression sequence encoding a therapeutic product;
[0203] each L is independently a linker sequence; and
[0204] n is an integer selected from 0 to 2.
[0205] Embodiment 11
[0206] An RNA polynucleotide comprising an engineered translation initiation element (TI) comprising an IRES-like polynucleotide sequence, wherein said IRES-like polynucleotide sequence comprises a nucleic acid sequence that is about or at least about 90%, 95%, 97%, 98%, 99% or 100% identical to SEQ ID NO: 1025-14161 or SEQ ID NO: 14412-15341.
[0207] Embodiment 12
[0208] The RNA polynucleotide according to any one of embodiments 1-11, wherein the RNA polynucleotide is circularized by a ligation reaction.
[0209] Embodiment 13
[0210] The RNA polynucleotide according to embodiment 12, wherein the RNA polynucleotide is circularized in the presence of T4 ligase.
[0211] Embodiment 14
[0212] The RNA polynucleotide according to embodiment 12, wherein the RNA polynucleotide is circularized in the absence of T4 ligase.
[0213] Embodiment 15
[0214] The RNA polynucleotide according to any one of embodiments 1-11, wherein the RNA polynucleotide is circularized by a splicing reaction.
[0215] Embodiment 16
[0216] The RNA polynucleotide according to embodiment 15, wherein the RNA polynucleotide is circularized in the presence of a spliceosome.
[0217] Embodiment 17
[0218] The RNA polynucleotide according to embodiment 15, wherein the RNA polynucleotide is circularized in the absence of a spliceosome.
[0219] Embodiment 18
[0220] The RNA polynucleotide according to any one of embodiments 1-11, wherein the RNA polynucleotide is circularized by a self-splicing reaction.
[0221] Embodiment 19
[0222] The RNA polynucleotide according to any one of embodiments 1-18, wherein the length of the IRES-like polynucleotide sequence is between 6 and 12 residues.
[0223] Embodiment 20
[0224] The RNA polynucleotide according to any one of embodiments 1-19, wherein the TI further comprises a second IRES-like polynucleotide sequence.
[0225] Embodiment 20-1
[0226] The RNA polynucleotide according to embodiment 20, wherein the first IRES-like polynucleotide sequence and the second IRES-like polynucleotide sequence are the same.
[0227] Embodiment 20-2
[0228] The RNA polynucleotide according to Embodiment 20, wherein the first IRES-like polynucleotide sequence and the second IRES-like polynucleotide sequence are different.
[0229] Embodiment 21
[0230] The RNA polynucleotide according to Embodiment 20, 20-1 or 20-2, wherein the second IRES-like polynucleotide sequence comprises a nucleic acid sequence that is about or at least about 90%, 95%, 97%, 98%, 99% or 100% identical to SEQ ID NO: 1025-14161 or SEQ ID NO: 14412-15341.
[0231] Embodiment 22
[0232] The RNA polynucleotide according to any one of Embodiments 1-21, wherein the TI further comprises a native IRES sequence. The RNA polynucleotide according to any one of Embodiments 1-21, wherein the TI further comprises an IRES sequence isolated from or derived from a native IRES sequence.
[0233] Embodiment 23
[0234] The RNA polynucleotide according to any one of Embodiments 1-22, wherein the IRES sequence is an RNA sequence capable of engaging a eukaryotic ribosome.
[0235] In some embodiments, the native IRES sequence is an endogenous IRES sequence isolated from or derived from Homo sapiens. In some embodiments, the native IRES sequence is an endogenous IRES sequence isolated from or derived from a human tissue or a human sample.
[0236] In some embodiments, the native IRES sequence is isolated from or derived from IRES sequences of viruses, mammals and Drosophila.
[0237] In some embodiments, the native IRES sequence is isolated from or derived from picornavirus complementary DNA (cDNA), encephalomyocarditis virus (EMCV) cDNA, poliovirus cDNA, ABPV IGRpred, AEV, ALPV IGRpred, BQCV IGRpred, BVDV1 1-385, BVDV1 29-391, CrPV 5NCR, CrPV IGR, crTMV IREScp, crTMV_IRESmp75, crTMV_IRESmp228, crTMV IREScp, crTMV IREScp, CSFV, CVB3, DCV IGR, EMCV-R, EoPV_5NTR, ERAV_245-96l, ERBV_l62-920, EV7l_l-748, FeLV-Notch2, FMDV serotype C, GBV-A, GBV-B, GBV-C, gypsy_env, gypsyD5, gypsyD2, HAV HM175, HCV genotype la, HiPV JGRpred, HIV-1, HoCVl JGRpred, HRV-2, IAPV JGRpred, idefix, KBV IGRpred, LINE-1_ORF1_-101_to_-1, LINE-1_ORF1_-302_to_-202, LINE-1_ORF2_-138_to_-86, LINE-1_ORF1_-44_to_-1, PSIV IGR, PV type 1 Mahoney, PV_type3_Leon, REV-A, RhPV 5NCR, RhPV IGR, SINV l IGRpred, SV40 661-830, TMEV, TMV_UI_IRESmp228, TRV 5NTR, TrV IGR or TSV IGR sequence.
[0238] In some embodiments, the native IRES sequence of Drosophila is the antennapedia gene of Drosophila melanogaster.
[0239] In some embodiments, the native IRES sequence is isolated from or derived from AML1 / RUNX1, Antp-D, Antp-DE, Antp-CDE, Apaf-l, Apaf-l, AQP4, ATlR varl, ATlR_var2, ATlR_var3, ATlR_var4, BAGl_p36delta236nt, BAGl_p36, BCL2, BiP_-222_-3, C-IAP1 285-1399, c IAP1 13 13-1462, c-jun, c-myc, Cat-l_224, CCND1, DAP5, eIF4G, eIF4GI-ext, eIF4GII, eIF4GII-long, ELG1, ELH, FGF1A, FMR1, Gtx-l33-l4l, Gtx-l-l66, Gtx-l-l20, Gtx-l-l96, hairless, HAP4, HIFla, hSNMl, HsplOl, hsp70, hsp70, Hsp90, IGF2_leader2, Kvl.4_l.2, L-myc, LamB 1 -335 -1, LEF1, MNT 75-267, MNT 36-160, MTG8a, MYB, MYT2 997-1 152, n-MYC, NDST1, NDST2, NDST3, NDST4L, NDST4S, NRF_-653_-l7, NtHSFl, ODC1, p27kipl, p53_l28-269, PDGF2 / c-sis, Pim-l, PITSLRE_p58, Rbm3, reaper, Scamper, TFIID, TIF4631, Ubx_l-966, Ubx_373-96l, UNR, Ure2, UtrA, VEGF-A-133 -1, XIAP 5-464, XIAP 305-466 or YAP1 sequence.
[0240] In some embodiments, the native IRES sequence is isolated from or derived from synthetic (GAAA)16, (PPT19)4, KMI1, KMI1, KMI2, KMI2, KMIX, XI or X2 IRES sequence.
[0241] Embodiment 24
[0242] The RNA polynucleotide according to any one of embodiments 1-23, wherein the L comprises a 5'UTR, 3'UTR, poly-A sequence, polyA-C sequence, poly-C sequence, poly-U sequence, poly-G sequence, ribosome binding site, aptamer, riboswitch, ribozyme, small RNA binding site, translation regulatory element (e.g., Kozak sequence), protein binding site (e.g., PTBP1 or HUR), unnatural nucleotide or non-nucleotide chemical linker.
[0243] Embodiment 25
[0244] The RNA polynucleotide according to any one of embodiments 1-24, wherein the length of the L is about 3 to about 100 nucleotide residues.
[0245] Embodiment 26
[0246] The RNA polynucleotide according to any one of embodiments 1-25, wherein the L comprises the nucleic acid sequence of RCC, wherein R is guanine or adenine.
[0247] Embodiment 27
[0248] The RNA polynucleotide according to any one of embodiments 1-26, wherein the RNA polynucleotide is a single-stranded RNA polynucleotide.
[0249] Embodiment 28
[0250] The RNA polynucleotide according to any one of embodiments 1-27, wherein the RNA polynucleotide is a circular RNA polynucleotide.
[0251] Embodiment 29
[0252] The RNA polynucleotide according to any one of embodiments 1-27, wherein the RNA polynucleotide is a linear RNA polynucleotide.
[0253] Embodiment 30
[0254] The RNA polynucleotide according to any one of embodiments 1-29, wherein the RNA polynucleotide is capable of cyclizing in the absence of an enzyme.
[0255] Embodiment 31
[0256] The RNA polynucleotide according to any one of embodiments 1-30, wherein the RNA polynucleotide is an isolated RNA polynucleotide or a synthetic RNA polynucleotide.
[0257] Embodiment 32
[0258] A polypeptide expressed by an RNA polynucleotide according to any one of embodiments 1-31.
[0259] Embodiment 33
[0260] A DNA vector encoding an RNA polynucleotide according to any one of embodiments 1-31.
[0261] Embodiment 34
[0262] A DNA vector suitable for synthesizing an RNA polynucleotide according to any one of embodiments 1-31.
[0263] Embodiment 35
[0264] A cell comprising an RNA polynucleotide according to any one of embodiments 1-31, a polypeptide according to embodiment 32, or a DNA vector according to embodiment 33 or 34.
[0265] Embodiment 36
[0266] A composition comprising an RNA polynucleotide according to any one of embodiments 1-31, a polypeptide according to embodiment 32, a DNA vector according to embodiment 33 or 34, or a cell according to embodiment 35 and a pharmaceutically acceptable carrier.
[0267] Embodiment 37
[0268] A method for preparing a cell population, the method comprising contacting the cells of the population with an RNA polynucleotide according to any one of embodiments 1-31, a polypeptide according to embodiment 32, or a DNA vector according to embodiment 33 or 34.
[0269] Embodiment 38
[0270] A method for generating an internal ribosome entry site (IRES)-like polynucleotide sequence, the method comprising the steps of:
[0271] (a) generating a polynucleotide query sequence consisting of X nucleic acid residues in length, where X is an integer greater than or equal to 3;
[0272] (b) generating X - Y + 1 overlapping polynucleotide fragment sequences within the polynucleotide query sequence,
[0273] where each polynucleotide fragment sequence consists of Y nucleic acid residues in length,
[0274] where the first position of each polynucleotide fragment sequence is n and the last position of the same polynucleotide fragment sequence is Y + n - 1, and
[0275] where n represents each positive integer between 1 and X - Y + 1;
[0276] (c) determining an enrichment score for each polynucleotide fragment sequence of (b);
[0277] (d) determining a numerical score for the polynucleotide query sequence by summing the enrichment scores of each polynucleotide fragment sequence of (c); and
[0278] (e) identifying the polynucleotide query sequence as an engineered IRES-like polynucleotide sequence according to a reference value.
[0279] Embodiment 38-1
[0280] The method according to Embodiment 38, wherein the reference value is calculated according to the following formula: reference value = (23.75 * X) - 84.85 - Z, where Z represents any number from 1 to 10000, and X represents the length of the IRES-like polynucleotide sequence.
[0281] Embodiment 38-2
[0282] The method according to Embodiment 38-1, wherein Z represents any number from 50 to 5000.
[0283] Embodiment 38-3
[0284] The method according to Embodiment 38-1, wherein Z represents any number from 175 to 2500.
[0285] Embodiment 38-4
[0286] The method according to Embodiment 38-1, wherein Z represents any number from 150 to 200.
[0287] Embodiment 39
[0288] The method according to Embodiment 38, wherein the reference value is a characteristic in the absence of therapeutic product expression.
[0289] Embodiment 40
[0290] The method according to Embodiment 38, wherein the reference value is the mean score of all numerical scores of two or more IRES-like polynucleotide sequences or two or more natural IRES sequences or a combination thereof.
[0291] Embodiment 41
[0292] The method according to embodiment 38, wherein the reference value is the average score of all numerical scores of two or more IRES-like polynucleotide sequences or two or more native IRES sequences or a combination thereof.
[0293] Embodiment 42
[0294] The method according to embodiment 38, wherein the reference value is about 95%, about 90%, about 85%, about 80%, about 75%, about 70%, about 65%, about 60%, about 55%, about 50%, about 45%, about 40%, about 35%, about 30%, about 25%, about 20%, about 15%, about 10%, about 9%, about 8%, about 7%, about 6%, about 5%, about 4%, about 3%, about 2%, about 1%, about 0.9%, about 0.8%, about 0.7%, about 0.6%, about 0.5%, about 0.4%, about 0.3%, about 0.2% or about 0.1% of the highest numerical score of two or more IRES-like polynucleotide sequences or two or more native IRES sequences or a combination thereof.
[0295] Embodiment 43
[0296] The method according to embodiment 38, wherein the reference value is about 50%, about 45%, about 40%, about 35%, about 30%, about 25%, about 20%, about 15% or about 10% of the highest numerical score of two or more IRES-like polynucleotide sequences or two or more native IRES sequences or a combination thereof.
[0297] Embodiment 44
[0298] The method according to embodiment 38, wherein the reference value is at least about 95%, about 90%, about 85%, about 80%, about 75%, about 70%, about 65%, about 60%, about 55%, about 50%, about 45%, about 40%, about 35%, about 30%, about 25%, about 20%, about 15%, about 10%, about 9%, about 8%, about 7%, about 6%, about 5%, about 4%, about 3%, about 2%, about 1%, about 0.9%, about 0.8%, about 0.7%, about 0.6%, about 0.5%, about 0.4%, about 0.3%, about 0.2% or about 0.1% higher than the lowest numerical score of two or more IRES-like polynucleotide sequences or two or more native IRES sequences or a combination thereof.
[0299] Embodiment 45
[0300] The method according to embodiment 38, wherein the reference value is at least about 95%, about 90%, about 85%, about 80%, about 75%, about 70%, about 65%, about 60%, about 55%, about 50% higher than the lowest numerical score of two or more IRES-like polynucleotide sequences or two or more native IRES sequences or a combination thereof.
[0301] Embodiment 46
[0302] The method according to embodiment 38, wherein the reference value is greater than or equal to 0.
[0303] Embodiment 47
[0304] The method according to any one of embodiments 38-46, wherein X is an integer selected from 3-300.
[0305] Embodiment 48
[0306] The method according to any one of embodiments 38-47, wherein X is an integer selected from 3-100.
[0307] Embodiment 49
[0308] The method according to any one of embodiments 38-48, wherein X is an integer selected from 5-100.
[0309] Embodiment 50
[0310] The method according to any one of embodiments 38-49, wherein X is an integer selected from 6-100.
[0311] Embodiment 51
[0312] The method according to any one of embodiments 38-50, wherein X is an integer greater than or equal to 5.
[0313] Embodiment 52
[0314] The method according to any one of embodiments 38-51, wherein X is an integer greater than or equal to 6.
[0315] Embodiment 53
[0316] The method according to any one of embodiments 38-52, wherein the polynucleotide query sequence is generated in a DNA vector suitable for synthesizing the polynucleotide query sequence.
[0317] Embodiment 54
[0318] The method according to any one of embodiments 38-53, wherein the polynucleotide query sequence is chemically synthesized.
[0319] Embodiment 55
[0320] The method according to any one of embodiments 38 - 54, wherein the length of the overlapping polynucleotide fragment sequences within the polynucleotide query sequence is 5, 6, 7, 8, 9 or 10 nucleic acid residues.
[0321] Embodiment 56
[0322] The method according to any one of embodiments 38 - 55, wherein the enrichment score of the polynucleotide fragment sequence is determined by:
[0323] i) Generating a library of expression plasmids, wherein each expression plasmid in the library contains a different polynucleotide fragment sequence and a reporter gene;
[0324] ii) Contacting a cell population with the library of expression plasmids;
[0325] iii) Quantifying the expression level of the reporter gene corresponding to each expression plasmid of the library of expression plasmids;
[0326] iv) Dividing the total cell population into a first population and a second population based on the protein expression level of iii);
[0327] v) Determining the enrichment score of the polynucleotide fragment sequence using the following system of equations:
[0328]
[0329] where f 1 is the frequency of the polynucleotide fragment sequence in the first population,
[0330] where f 2 is the frequency of the same polynucleotide fragment sequence in the second population,
[0331] where N 1 is the size of the first population, and
[0332] where N 2 is the size of the second population.
[0333] Embodiment 57
[0334] The method according to embodiment 56, wherein the first population has a protein expression level in the top 0 - 10% of the protein expression level of the total population, and wherein the second population has a protein expression level in the bottom 10% - 90% of the protein expression level of the total population.
[0335] Embodiment 58
[0336] The method according to embodiment 56 or 57, wherein the first population has a protein expression level in the top 50.1% of the protein expression levels of the total population, and wherein the second population has a protein expression level in the bottom 49.9% of the protein expression levels of the total population.
[0337] Embodiment 59
[0338] The method according to embodiment 56 or 57, wherein the first population has a protein expression level in the top 10% of the protein expression levels of the total population, and wherein the second population has a protein expression level in the bottom 90% of the protein expression levels of the total population.
[0339] Embodiment 60
[0340] The method according to any one of embodiments 38 - 59, wherein the length of the polynucleotide fragment sequence is 5 nucleic acid residues.
[0341] Embodiment 60 - 1
[0342] The method according to embodiment 60, wherein the polynucleotide fragment sequence is selected from SEQ ID NO: 1 - 1024, and wherein the enrichment score of the polynucleotide fragment sequence is as shown in Table 1.
[0343] Embodiment 61
[0344] The method according to any one of embodiments 38 - 59, wherein the length of the polynucleotide fragment sequence is 6 nucleic acid residues.
[0345] Embodiment 62
[0346] The method according to any one of embodiments 38 - 59, wherein the length of the polynucleotide fragment sequence is 7 nucleic acid residues.
[0347] Embodiment 63
[0348] The method according to any one of embodiments 38 - 59, wherein the length of the polynucleotide fragment sequence is 8 nucleic acid residues.
[0349] Embodiment 64
[0350] The method according to any one of embodiments 38 - 59, wherein the length of the polynucleotide fragment sequence is 9 nucleic acid residues.
[0351] Embodiment 65
[0352] The method according to any one of embodiments 38 - 59, wherein the length of the polynucleotide fragment sequence is 10 nucleic acid residues.
[0353] Embodiment 66
[0354] A method according to any one of embodiments 38-59, wherein the polynucleotide fragment sequence has a length of 11 nucleic acid residues.
[0355] Embodiment 67
[0356] A method according to any one of embodiments 38-59, wherein the polynucleotide fragment sequence has a length of 12 nucleic acid residues.
[0357] Embodiment 68
[0358] A method according to any one of embodiments 38-59, wherein the polynucleotide fragment sequence has a length of 13-100 nucleic acid residues.
[0359] Embodiment 69
[0360] An RNA polynucleotide comprising an engineered translation initiation element (TI), wherein the TI comprises an internal ribosome entry site (IRES)-like polynucleotide sequence generated by a method according to any one of embodiments 38-68.
[0361] Embodiment 70
[0362] A polypeptide expressed by an RNA polynucleotide according to embodiment 69.
[0363] Embodiment 71
[0364] A DNA vector encoding an RNA polynucleotide according to embodiment 69.
[0365] Embodiment 72
[0366] A DNA vector suitable for synthesizing an RNA polynucleotide according to embodiment 69.
[0367] Embodiment 73
[0368] A cell comprising an RNA polynucleotide according to embodiment 69, a polypeptide according to embodiment 70, or a DNA vector according to embodiment 71 or 72.
[0369] Embodiment 74
[0370] A composition comprising an RNA polynucleotide according to embodiment 69, a polypeptide according to embodiment 70, a DNA vector according to embodiment 71 or 72, or a cell according to embodiment 73 and a pharmaceutically acceptable carrier.
[0371] Embodiment 75
[0372] A method for preparing a cell population, the method comprising contacting the cells of the population with the RNA polynucleotide according to Embodiment 69, the polypeptide according to Embodiment 70, or the DNA vector according to Embodiment 71 or 72.
[0373] Embodiment 76
[0374] A method for modulating protein expression in a subject in need thereof, the method comprising administering to the subject a therapeutically effective amount of the polynucleotide according to any one of Embodiments 1-31 and 69, the polypeptide according to Embodiment 32 or 70, the DNA vector according to Embodiments 33, 34, 71 or 72, or the cell according to Embodiment 35 or 73.
[0375] Embodiment 77
[0376] A method for treating or preventing a disease or disorder in a subject in need thereof, the method comprising administering to the subject a therapeutically effective amount of the polynucleotide according to any one of Embodiments 1-31 and 69, the polypeptide according to Embodiment 32 or 69, the DNA vector according to Embodiments 33, 34, 71 or 72, or the cell according to Embodiment 35 or 73.
[0377] Embodiment 78
[0378] Use of the polynucleotide according to any one of Embodiments 1-31 and 69, the polypeptide according to Embodiment 32 or 70, the DNA vector according to Embodiments 33, 34, 71 or 72, or the cell according to Embodiment 35 or 73 in the manufacture of a medicament for modulating protein expression or treating or preventing a disease or disorder in a subject in need thereof.
[0379] Embodiment 79
[0380] The polynucleotide according to any one of Embodiments 1-31 and 69, the polypeptide according to Embodiment 32 or 70, the DNA vector according to Embodiments 33, 34, 71 or 72, or the cell according to Embodiment 35 or 73 for modulating protein expression or treating or preventing a disease or disorder in a subject in need thereof. 7. BRIEF DESCRIPTION OF THE DRAWINGS
[0381] Figure 1Schematic diagram showing the method for generating IRES-like sequences of the present disclosure. An exemplary polynucleotide query sequence (e.g., 12-mer nucleotide sequence) was generated. Overlapping polynucleotide fragment sequences within the polynucleotide query sequence were generated (e.g., pentamers 1-8). An enrichment score for each polynucleotide fragment sequence was generated (e.g., z-score for pentamers 1-8). A numerical score for the polynucleotide query sequence was generated by adding the enrichment scores of each polynucleotide fragment sequence (e.g., the sum of the z-score values). When the numerical score of the polynucleotide query sequence is higher than a determined threshold, the polynucleotide query sequence is identified as an IRES-like sequence.
[0382] Figure 2 Western blot analysis showing protein expression of an exemplary circular RNA having an IRES-like sequence, the IRES-like sequence being 12 nucleotides in length. Also shown are the z-score (numerical value) and the percentile ranking of all 12-mer nucleotide sequences.
[0383] Figure 3 Schematic flowchart showing the IRES-like sequence optimization method of the present disclosure.
[0384] Figure 4 Schematic flowchart showing the IRES-like sequence optimization method of the present disclosure.
[0385] Figure 5 Flowchart showing the method for obtaining circular RNAs of the present disclosure.
[0386] Figure 6 Bar graph showing the luminescence signal of luciferase protein expression of an exemplary circular RNA having an IRES-like sequence from the present disclosure.
[0387] Figure 7 Image of an in vivo imaging system (IVIS) spectrum of protein expression after IV injection of an exemplary circular RNA having a combination of a native IRES and an IRES-like sequence of the present disclosure.
[0388] Figure 8 Diagram showing the self-cyclization of type I or type II introns.
[0389] Figure 9 Gel image showing RNA self-splicing based on two splicing type II introns and flanking exon sequences.
[0390] Figure 10A Gel image showing RNA splicing (upper inset) and bar graph indicating the percentage of RNA cyclization determined by electrophoresis (lower inset).
[0391] Figure 10BIt is a gel image showing the experimental results of successfully forming circular RNAs through different methods.
[0392] Figure 11 It is a gel image showing the results of the cyclization products of three scarless target sequences at different magnesium ion concentrations.
[0393] Figure 12A It is a flow cytometry plot showing the protein expression from cells expressing an exemplary circular RNA having a translation initiation element.
[0394] Figure 12B It shows a Western blot analysis of the protein expression from an exemplary circular RNA having a translation initiation element, the translation initiation element having a length of 6 nucleotides (i.e., an IRES-like element, 6-mer). The z-score (numerical value) and the percentile ranking of all 6-mer nucleotide sequences are shown.
[0395] Figure 12C - 12E It shows a series of Western blot analyses of the protein expression from an exemplary circular RNA having a translation initiation element, the translation initiation element having a length of 12 nucleotides (i.e., an IRES-like element, 12-mer). The z-score (numerical value) and the percentile ranking of all 12-mer nucleotide sequences are shown.
[0396] Figure 13A It shows the IRES-like sequences having the highest ranked scores at each polynucleotide query length. The numbers on the left represent the total length of each nucleotide. Nucleotide lengths (SEQ ID NO) - length 6 (SEQ ID NO: 1025); length 7 (SEQ ID NO: 1225); length 8 (SEQ ID NO: 1425); length 9 (SEQ ID NO: 1625); length 10 (SEQ ID NO: 1826); length 11 (SEQ ID NO: 2026); length 12 (SEQ ID NO: 2226); length 13 (SEQ ID NO: 2426); length 14 (SEQ ID NO: 2626); length 15 (SEQ ID NO: 2827); length 16 (SEQ ID NO: 3028); length 17 (SEQ ID NO: 3229); length 18 (SEQ ID NO: 3429); length 19 (SEQ ID NO: 3629); length 20 (SEQ ID NO: 3829); length 21 (SEQ ID NO: 4033).
[0397] Figure 13BShows the highest ranked numerical score for each polynucleotide query length. The y-axis shows the highest ranked numerical score (i.e., the highest Z-score). The x-axis shows the length of the polynucleotide query sequences tested. The equation of the line is, numerical score = 23.74869 x length of the polynucleotide query sequence – 84.85243.
[0398] Figure 14A Shows the predicted numerical scores for polynucleotide query sequence lengths of 12, 18, 50 or 100 nucleotides. The y-axis shows the numerical scores. The x-axis shows the percentile rank.
[0399] Figure 14B Shows the predicted numerical scores for a polynucleotide query sequence length of 12 nucleotides. This figure represents a magnified image of the highest percentile rank (99.99250% to 100%). The y-axis shows the numerical scores. The x-axis shows the percentile rank. For the 12-mer sequence, the 99.99533% percentile corresponds to a numerical score of 165.55; the 99.99600% percentile corresponds to a numerical score of 167.43; the 99.99700% percentile corresponds to a numerical score of 171.34; the 99.99800% percentile corresponds to a numerical score of 174.88; the 99.99900% percentile corresponds to a numerical score of 181.30; the 100% percentile corresponds to a numerical score of 198.32.
[0400] Figure 14C Shows the predicted numerical scores for a polynucleotide query sequence length of 18 nucleotides. This figure represents a magnified image of the highest percentile rank (99.99250% to 100%). The y-axis shows the numerical scores. The x-axis shows the percentile rank. For the 18-mer sequence, the 99.9950% percentile corresponds to a score of 210.98, the 99.9960% percentile corresponds to a score of 216.97, the 99.9970% percentile corresponds to a score of 222.21, the 99.9980% percentile corresponds to a score of 229.26, the 99.9990% percentile corresponds to a score of 240.88, and the 100% percentile corresponds to a score of 341.65.
[0401] Figure 14DShows the predicted numerical scores for the lengths of polynucleotide query sequences of 50 nucleotides. This figure represents a scaled image of the highest percentile rankings (99.99250% to 100%). The y-axis shows the numerical scores. The x-axis shows the percentile rankings. For 50-mer sequences, the 99.9950% percentile corresponds to a score of 309.20, the 99.9960% percentile corresponds to a score of 315.56, the 99.9970% percentile corresponds to a score of 323.64, the 99.9980% percentile corresponds to a score of 335.22, the 99.9990% percentile corresponds to a score of 354.05, and the 100% percentile corresponds to a score of 1101.43.
[0402] Figure 14E Shows the predicted numerical scores for the lengths of polynucleotide query sequences of 100 nucleotides. This figure represents a scaled image of the highest percentile rankings (99.99250% to 100%). The y-axis shows the numerical scores. The x-axis shows the percentile rankings. For 100-mer sequences, the 99.9950% percentile corresponds to a score of 384.85, the 99.9960% percentile corresponds to a score of 393.02, the 99.9970% percentile corresponds to a score of 403.11, the 99.9980% percentile corresponds to a score of 416.68, the 99.9990% percentile corresponds to a score of 439.34, and the 100% percentile corresponds to a score of 2289.57.
[0403] Figure 15 Shows the luminescence signals of luciferase protein expression of exemplary circular RNAs with IRES-like sequences of the present disclosure. The y-axis shows the luminescence values. The x-axis shows six exemplary circular RNAs (V1, V3, V4, V5, V6, and V7), where the total lengths (linker sequences + IRES-like sequences) of V1, V3, V4, V5, V6, and V7, the Z-scores corresponding to the SEQ ID, and the percentiles corresponding to the SEQ ID are shown in the table.
[0404] Figure 16 Shows the general structure of group II introns. As shown, a typical group II intron can have six stem-loop structures, called domains 1-6 or D1-D6. D4 contains an open reading frame. These 6 domains are arranged in sequence and contain multiple exon-binding sequences (EBSs), such as EBS1, EBS2, and EBS3. These EBS sequences interact with intron-binding sequences (IBSs) in the exon region (such as IBS1, IBS2, and IBS3), such as complementary pairing, to trigger self-splicing. Note that a single nucleotide, the δ nucleotide, upstream of EBS1 in domain 1 can also pair with IBS3, and the interaction between δ and IBS3 is called δ-IBS3 pairing. 8. Specific embodiments
[0405] The present disclosure overcomes problems associated with the current art by providing novel methods for identifying and generating synthetic internal ribosome entry site (IRES)-like sequences. The present disclosure is at least in part based on the development of algorithmic methods and high-throughput reporter assays that can systematically screen for and quantify the IRES activity of RNA sequences that can promote the translation of circular RNAs. The inventors have systematically screened libraries of random polynucleotide sequences that drive protein expression and have used systematic ranking methods to distinguish nucleic acid sequence motifs that enhance protein expression (i.e., IRES-like sequences) from nucleic acid sequence motifs that reduce protein expression (i.e., non-IRES sequences), or to distinguish nucleic acid sequence motifs that more potently enhance protein expression (i.e., IRES-like sequences) from other nucleic acid sequence motifs. This is useful for identifying novel sequences with IRES activity that function more effectively than comparable native IRES sequences. This is also useful for identifying IRES elements that are general and species-independent.
[0406] The discovery of these nucleic acid sequence motifs allows for the generation of synthetic IRES-like sequences for enhancing protein expression in eukaryotic cells. Increasing protein expression is desirable for improving gene expression and / or protein expression and production in therapeutic applications, including but not limited to protein replacement therapy and vaccination. In addition, the use of combinations of IRES-like sequences and non-IRES sequences provides opportunities for fine-tuning the control and regulation of protein expression. The use of IRES-like sequences in combination with non-IRES sequences is also beneficial for precisely controlling the ratio of peptide expression in the manufacture and production of multimeric proteins and / or polycistronic cassettes for therapeutic applications, such as antibody production.
[0407] Accordingly, the present disclosure also provides IRES-like sequences and polynucleotides comprising an IRES-like sequence and an expression sequence encoding a therapeutic product of interest, which polynucleotides may be suitable for use as a drug or vaccine for applications such as in gene therapy and / or gene vaccination. In some embodiments, the present disclosure provides polynucleotides comprising an IRES-like sequence and an expression sequence encoding a therapeutic product of interest for improving protein manufacture and production. In some embodiments, the polynucleotide is an RNA polynucleotide. In some embodiments, the RNA polynucleotide is a circular RNA polynucleotide. In summary, the present disclosure provides improved polynucleotides comprising IRES-like sequences that overcome the disadvantages in the prior art by providing a cost-effective, systematic, and direct method for regulating protein expression.
[0408] This article provides non-naturally occurring RNAs with Group II intron self-splicing activity, which can form circular RNAs during self-splicing. Circular RNAs (circRNAs) are single-stranded RNAs with their ends joined together. As is known in the art, regardless of the cyclization method employed, circular RNAs can be generated in vitro using chemical means or through the enzymatic activity of a precursor RNA, where the precursor RNA refers to the linear RNA molecule that directly gives rise to the circular RNA. For example, the 5'-end and 3'-end of a linear nucleic acid can be chemically ligated through the catalysis of cyanogen bromide and morpholino derivatives, or ligated head-to-tail through the activity of a nucleic acid ligase. CircRNAs can also be produced through splicing. When a linear precursor undergoes splicing, a portion of the molecule is excised, resulting in a circular RNA with fewer total nucleotides than the precursor RNA.
[0409] Circular RNAs can also be produced through ribozyme-catalyzed RNA splicing. As used herein and understood in the art, the term "ribozyme" refers to an RNA molecule with enzymatic activity. Some ribozymes can catalyze self-splicing independently of the spliceosome and are referred to as "ribozymes with self-splicing activity", "self-splicing ribozymes" or "self-splicing introns". Naturally occurring self-splicing ribozymes can be classified into Group I and Group II introns. Although the splicing products of these two types of ribozymes are similar, the structures and splicing mechanisms of the ribozymes themselves are quite different. Group I introns have a 9-helix structure, require the external hydroxyl group in guanosine monophosphate (pG-OH) to trigger the reaction during catalytic splicing, and are highly dependent on the exon sequences located at both ends of the Group I intron. Group II introns rely on their own hydroxyl groups in the nucleotide sequence to trigger splicing. (See Figure 8 ). This splicing mechanism is closer to the spliceosome-mediated splicing reaction and better mimics the splicing in higher organisms. The term "Group I intron self-splicing activity" or "Group I intron activity" refers to the self-splicing activity derived from Group I introns; the term "Group II intron self-splicing activity" or "Group II intron activity" refers to the self-splicing activity derived from Group II introns. The method for preparing circular RNAs based on Group II intron self-splicing activity has at least the following advantages: reducing the use of biological and chemical reagents (such as ligases and related reagents), simple operation, and simple design.
[0410] Engineered ribozyme RNAs with self-splicing activity that form circRNAs after self-splicing, also referred to herein as "cRNAzymes". In some embodiments, the cRNAzymes can have in vitro self-splicing activity. Novel non-natural RNAs with group II intron self-splicing activity are disclosed herein that form circular RNAs or "group II cRNAzymes" after self-splicing. Also provided herein are vectors comprising polynucleotides encoding these group II cRNAzymes, methods of preparing the group II cRNAzymes disclosed herein by transcribing these vectors, and uses of these group II cRNAzymes in the preparation of circRNAs.
[0411] Before further describing the disclosure, it is to be understood that the disclosure is not limited to the specific embodiments described herein, and it is also to be understood that the terminology used herein is for the purpose of describing particular embodiments and is not intended to be limiting.
[0412] 8.1. Definitions
[0413] As used herein, "substantially free of" with respect to a specified component is used herein to mean that no specified component is purposefully formulated into a composition and / or is present only as a contaminant or in trace amounts. Thus, the total amount of the specified component resulting from any accidental contamination of the composition is far less than 0.1%, preferably less than 0.05%, and more preferably less than 0.01%. Most preferably, it is a composition in which the amount of the specified component is not detectable by standard analytical methods.
[0414] As used in the specification, "a" or "an" can mean one or more. As used in the claims, when used in conjunction with the word "comprising", the word "a" or "an" can mean one or more than one.
[0415] As used herein, the term "or" in the claims is used to mean "and / or" unless expressly stated to refer only to alternatives or the alternatives are mutually exclusive, although the disclosure text supports definitions that refer only to alternatives and "and / or". As used herein, "another" or "additional" can mean at least a second or more.
[0416] As used herein, the term "about" is used to indicate that a value includes the inherent variations of that value, e.g., due to the error of the device, the method for determining the value, or the variations present in the subject matter being studied. In some embodiments, "about" means a variation of ±10%, ±9%, ±8%, ±7%, ±6%, ±5%, ±4%, ±3%, ±2%, ±1%, ±0.5%, ±0.2%, or ±0.1% of the value so recited. In some embodiments, "about" means a variation of ±10%, ±9%, ±8%, ±7%, ±6%, ±5%, ±4%, ±3%, or ±2% of the value so recited. In some embodiments, "about" means a variation of ±10% of the value so recited. In some embodiments, "about" means a variation of ±5%, ±4%, ±3%, ±2, or ±1% of the value so recited. In some embodiments, "about" means a variation of ±1%, ±0.5%, ±0.2%, or ±0.1% of the value so recited.
[0417] The terms "average" and "mean" in this document are used interchangeably unless there is a clear indication to the contrary.
[0418] As used herein, the term "portion" when used in reference to a polypeptide or peptide refers to a fragment of the polypeptide or peptide. In some embodiments, a "portion" of a polypeptide or peptide retains at least one function and / or activity of the full-length polypeptide or peptide from which the portion is derived. For example, in some embodiments, if the full-length polypeptide binds a given ligand, a portion of the full-length polypeptide also binds the same ligand.
[0419] Unless explicitly stated to the contrary, the terms "protein" and "polypeptide" are used interchangeably herein.
[0420] When used in connection with a protein, gene, nucleic acid, or polynucleotide in a cell or organism, the term "exogenous" refers to a protein, gene, nucleic acid, or polynucleotide that is introduced into the cell or organism by artificial or natural means; or when used in connection with a cell, refers to a cell that is isolated by artificial or natural means and subsequently introduced into a cell population or organism. An exogenous nucleic acid can be from a different organism or cell, or it can be one or more additional copies of a nucleic acid that is naturally present in the organism or cell. An exogenous cell can be from a different organism, or it can be from the same organism. By way of non-limiting example, an exogenous nucleic acid is a nucleic acid at a chromosomal location different from that in the natural cell, or otherwise, flanked by nucleic acid sequences different from those found in nature. The term "exogenous" is used interchangeably with the term "heterologous".
[0421] "Expression construct" or "expression cassette" is used to mean a nucleic acid molecule capable of directing transcription. An expression construct includes at least one or more transcriptional control elements (such as promoters, enhancers, or their functional equivalents) that direct gene expression in one or more desired cell types, tissues, or organs. Additional elements, such as transcriptional termination signals, may also be included.
[0422] "Vector" or "construct" (sometimes referred to as a gene delivery system or gene transfer "vehicle") refers to a macromolecule or molecular complex that contains a polynucleotide to be delivered to a host cell (either in vitro or in vivo) or a protein expressed by the polynucleotide.
[0423] "Plasmid", as a common type of vector, is an extrachromosomal DNA molecule that is separated from chromosomal DNA and capable of replicating independently of chromosomal DNA. In some cases, it is circular double-stranded.
[0424] The term "cRNAzyme" is used herein to refer to a linear ribonucleic acid (RNA) capable of generating circular RNA through an autocatalytic back-splicing reaction.
[0425] The term "cRNAzyme construct" is a linear RNA construct having cRNAzyme activity.
[0426] The term "EBS" is used herein to refer to an exon-binding sequence that interacts (e.g., forms a complementary pair) with an intron-binding sequence (IBS) in an exon region to trigger splicing by its own hydroxyl group within the EBS.
[0427] The term "IBS" is used herein to refer to an intron-binding sequence that interacts (e.g., forms a complementary pair) with an exon-binding sequence (EBS) to locate the splice site.
[0428] The term "EBS1" is used herein to refer to exon-binding sequence 1. In some embodiments, EBS1 comprises a nucleic acid sequence selected from: (a) UAGGGC; (b) UUAUGG; (c) UCAACG; and (d) UGUGGC.
[0429] The term "EBS2" is used herein to refer to exon-binding sequence 2.
[0430] The term "EBS3" is used herein to refer to exon-binding sequence 3.
[0431] The term "EBS1'" is used herein to refer to a modified EBS1 sequence that interacts with IBS1'. The interaction between EBS1' and IBS1' is similar to the interaction between EBS1 and IBS1. In some embodiments, EBS1' comprises a nucleic acid sequence selected from: (a) UAGGGC; (b) UUAUGG; (c) UCAACG; and (d) UGUGGC.
[0432] The term "EBS3'" is used herein to refer to a modified EBS3 sequence that interacts with IBS3'. The interaction between EBS3' and IBS3' is similar to the interaction between EBS3 and IBS3.
[0433] The term "IBS1" is used herein to refer to intron binding sequence 1, which interacts with exon binding sequence 1 (EBS1) to locate the splicing site. In some embodiments, IBS1 comprises a nucleic acid sequence selected from: (a) GCCCUG; (b) CCAUGG; (c) CGUUGA; and (d) GCCAUA.
[0434] The term "IBS1'" is used herein to refer to a region on the target sequence that has a similar function to IBS1.
[0435] The term "IBS2" is used herein to refer to intron binding sequence 2, which interacts with exon binding sequence 2 (EBS2) to locate the splicing site.
[0436] The term "IBS3" is used herein to refer to intron binding sequence 3, which interacts with exon binding sequence 3 (EBS3) to locate the splicing site. In some embodiments, IBS3 and its downstream sequence comprise a nucleic acid sequence selected from: (a) AGCAAA; (b) AGCAGU; (c) AGAGAA; and (d) AGCAAA.
[0437] The term "IBS3'" is used herein to refer to a region on the target sequence that has a similar function to IBS3.
[0438] The term "δ" (delta) is used herein to refer to a region in domain 1 of group II introns, which is a single nucleotide located upstream of EBS1. δ pairs with IBS3, and the interaction between δ and IBS3 is called δ-IBS3 pairing. In some embodiments, the δ sequence and its upstream comprise a nucleic acid sequence selected from: (a) UGUGCU; (b) AAUGCU; (c) UGCUCU; and (d) UGUGCU.
[0439] The term "δ" (delta) is used herein to refer to the region in domain 1 of a group II intron, which is a single nucleotide located just upstream of EBS1'. δ pairs with IBS3', and the interaction between δ and IBS3' is called the δ-IBS3' pairing.
[0440] As used herein, the terms "group II intron" and "Group II intron" are used interchangeably herein to refer to an RNA molecule encoded by a group II intron, having similar secondary and tertiary structures. A group II intron RNA molecule typically has six domains. Group II introns mainly include six stem-loop structures, called domains 1-6 (D1-D6), and the six domains are arranged in sequence, containing multiple exon-binding sequences (EBS), such as EBS1, EBS2, and EBS3.
[0441] Type II introns can have modifications of one or more nucleotides relative to their wild-type form, and the modifications are selected from one or more of deletions, substitutions, and additions. In some embodiments, the modifications include modifying one or more EBS sequences of the Type II intron, wherein the EBS sequences are complementary paired with one or more regions of corresponding length in the target sequence at at least 60% of the nucleotide positions. In some embodiments, the modifications are to modify two EBS sequences (such as EBS1 and EBS3) of the Type II intron, wherein the EBS sequences are complementary paired with two regions of corresponding length in the target sequence at at least 60% of the nucleotide positions; preferably, the two regions are located at both ends of the target sequence respectively. In some embodiments, the modifications are to modify two EBS sequences (such as EBS1' and EBS3') of the Type II intron, wherein the EBS sequences are complementary paired with two regions of corresponding length in the target sequence at at least 60% of the nucleotide positions; preferably, the two regions are located at both ends of the target sequence respectively. In some embodiments, the modifications are to modify EBS1 and / or the δ sequence of the Type II intron, or to modify EBS1' and / or the δ" sequence, wherein the EBS1 and / or the δ sequence is complementary paired with a region of corresponding length in the target sequence at at least 60% of the nucleotides, optionally the modifications are to modify EBS1 and / or the δ sequence and its upstream sequence, wherein the EBS1 and / or the δ sequence and its upstream are complementary paired with a region of corresponding length in the target sequence at at least 60% of the nucleotides. In some embodiments, the modifications are to modify EBS1 and / or the δ sequence of the Type II intron, or to modify EBS1' and / or the δ" sequence, wherein the EBS1 and / or the δ sequence is complementary paired with a region of corresponding length in the target sequence at at least 60%, 61%, 62%, 63%, 64%, 65%, 66%, 67%, 68%, 69%, 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100% of the nucleotides, optionally the modifications are to modify EBS1 and / or the δ sequence and its upstream sequence, wherein the EBS1 and / or the δ sequence and its upstream sequence are complementary paired with a region of corresponding length in the target sequence at at least 60% of the nucleotides. In some embodiments, the region of corresponding length in the target sequence is IBS3, IBS3', IBS3 and its downstream sequence, or IBS3' and its downstream sequence. In some embodiments, the modifications include deleting part or all of domain 4, such as deleting the intron-encoded protein (IEP) sequence in domain 4, preferably deleting all of domain 4.In some embodiments, the modification comprises deletion of an open reading frame (ORF).
[0442] The term "domain 1" or "D1" is used herein to refer to the stem-loop structure of domain 1 of a Group II intron. The term "domain 2" or "D2" is used herein to refer to the stem-loop structure of domain 2 of a Group II intron. The term "domain 3" or "D3" is used herein to refer to the stem-loop structure of domain 3 of a Group II intron. The term "domain 4" or "D4" is used herein to refer to the stem-loop structure of domain 4 of a Group II intron. The term "domain 5" or "D5" is used herein to refer to the stem-loop structure of domain 5 of a Group II intron. The term "domain 6" or "D6" is used herein to refer to the stem-loop structure of domain 6 of a Group II intron. A stem-loop structure is a type of RNA secondary structure that can be determined by any suitable polynucleotide folding algorithm. Some programs are based on the calculation of the minimum Gibbs free energy. An exemplary algorithm is mFold (Zuker and Stiegler, Nucleic Acids Res. 9, 133-148 (1981)). Other exemplary folding algorithms are known in the art, such as those described in: AR Gruber et al., Cell 106, 23-24 (2008); PA Carr and GM Church, Nature Biotechnology 27, 1151-62 (2009); PCT application No. WO 2014093709, which is incorporated herein by reference.
[0443] The terms "5' intron fragment" and "3' intron fragment" are used herein to refer to intron fragments obtained by splitting a Group II intron at the loop region of the stem-loop structure of domain 1, 2, 3, 4, 5, or 6.
[0444] Unless otherwise specified or the context indicates to the contrary, the terms "nucleic acid sequence", "polynucleotide", and "oligonucleotide" are used interchangeably herein and refer to polymers or oligomers of pyrimidine and / or purine bases (such as cytosine, thymine, and uracil, adenine and guanine, respectively) (see Albert L. Lehninger, Principles of Biochemistry, 793 - 800 (Worth Pub. 1982)). These terms encompass any deoxyribonucleotide, ribonucleotide, or peptide nucleic acid component, as well as any chemical variants thereof, such as methylated, hydroxymethylated, or glycosylated forms of these bases. The composition of the polymer or oligomer can be heterogeneous or homogeneous, can be isolated from naturally occurring sources, or can be produced in an artificial or synthetic manner. Additionally, the nucleic acid can be DNA or RNA or a mixture thereof, and can exist permanently or transiently in single-stranded or double-stranded form (including homoduplex, heteroduplex, and hybrid states). A nucleic acid or nucleic acid sequence can contain other types of nucleic acid structures, such as DNA / RNA helices, peptide nucleic acids (PNA), morpholino nucleic acids (see, for example, Braasch and Corey, Biochemistry, 41(14):4503 - 4510 (2002) and U.S. Patent 5,034,506), locked nucleic acids (LNA; see Wahlestedt et al., Proc. Natl. Acad. Sci. U.S.A., 97:5633 - 5638 (2000)), cyclohexenyl nucleic acids (see Wang, J. Am. Chem. Soc., 122:8595 - 8602 (2000)), and / or ribozymes. The terms "nucleic acid", "nucleic acid sequence", "polynucleotide", and "oligonucleotide" can also encompass chains containing unnatural nucleotides, modified nucleotides, and / or non-nucleotide building blocks that can exhibit the same functions as natural nucleotides (e.g., "nucleotide analogs"). The term "DNA sequence" is used herein to refer to a nucleic acid containing a series of DNA bases.
[0445] As is known in the art, nucleic acid strands have an inherent directionality since the carbon atoms in the sugar ring are numbered from 1' to 5', with the "5'-end" having a free hydroxyl (or phosphate) group on the 5'-carbon and the "3'-starting end" having a free hydroxyl (or phosphate) group on the 3'-carbon. As used herein and understood in the art, a nucleic acid having certain sequence elements in the "5'-to-3'" direction means that these sequence elements are linearly arranged from the 5'-end to the 3'-end of the nucleic acid.
[0446] As used herein, "complementary" with respect to an oligonucleotide means that when the nucleobase sequence of the oligonucleotide is aligned with another nucleic acid in the opposite direction, at least 70% of the nucleobases of the oligonucleotide or one or more regions thereof are capable of forming hydrogen bonds with the nucleobases of the other nucleic acid or one or more regions thereof. Complementary nucleobases mean nucleobases that are capable of forming hydrogen bonds with each other. Unless otherwise specified, complementary nucleobase pairs include, but are not limited to, adenine (A) and thymine (T); adenine (A) and uracil (U); cytosine (C) and guanine (G); and 5-methylcytosine ( m C) and guanine (G). Complementary oligonucleotides and / or nucleic acids need not have nucleobase complementarity at every nucleoside. Rather, some mismatches are tolerated. As used herein, "fully complementary" or "100% complementary" with respect to an oligonucleotide means that the oligonucleotide is complementary to another oligonucleotide or nucleic acid at every nucleoside of the oligonucleotide.
[0447] As used herein, the term "nucleobase" means an unmodified nucleobase or a modified nucleobase. As used herein, an "unmodified nucleobase" is adenine (A), thymine (T), cytosine (C), uracil (U), and guanine (G). As used herein, a "modified nucleobase" is a group of atoms other than unmodified A, T, C, U, or G that is capable of pairing with at least one unmodified nucleobase. "5-Methylcytosine" is a modified nucleobase. A universal base is a modified nucleobase that can pair with any one of the five unmodified nucleobases. As used herein, a "nucleobase sequence" means the order of consecutive nucleobases in a nucleic acid or oligonucleotide independent of any sugar or internucleoside linkage modification.
[0448] As used herein, the term "nucleoside" means a compound comprising a nucleobase and a sugar moiety. The nucleobase and the sugar moiety are each independently unmodified or modified. As used herein, a "modified nucleoside" means a nucleoside comprising a modified nucleobase and / or a modified sugar moiety. Modified nucleosides include abasic nucleosides that lack a nucleobase. "Linked nucleosides" are nucleosides that are linked in a continuous sequence (i.e., there are no additional nucleosides between the linked nucleosides).
[0449] Unless otherwise specified or indicated to the contrary in the context, the terms "polypeptide" and "protein" are used interchangeably herein and refer to polymeric forms of amino acids that include at least two or more contiguous amino acids that are chemically or biochemically modified or derivatized. As used herein, the term "peptide" refers to a class of short polypeptides. The term peptide can refer to a polymer of amino acids (naturally or non-naturally occurring) that is up to about 100 amino acids in length. For example, the length of a peptide can be about 1 to about 10, about 10 to about 25, about 25 to about 50, about 50 to about 75, about 75 to about 100 amino acid residues. In some embodiments, the length of a peptide can be about 100, about 200, about 300, about 400, about 500, about 600, about 700, about 800, about 900, about 1000, about 1250, about 1500, about 1750, about 2000, about 2250, about 2500, about 2750, about 3000, about 3250, about 3500, about 3750, about 4000, about 4250, about 4500, about 4750, about 5000 amino acid residues.
[0450] The nomenclature for nucleotides, nucleic acids, nucleosides, and amino acids used herein is consistent with the International Union of Pure and Applied Chemistry (IUPAC) standards (see, e.g., bioinformatics.org / smsylupac.html).
[0451] When referring to nucleic acid or protein sequences, the term "identity" is used to denote the similarity between two sequences. Sequence similarity or identity can be determined using standard techniques known in the art, including but not limited to the local sequence identity algorithm of Smith & Waterman, Adv. Appl. Math. 2, 482 (1981), the sequence identity alignment algorithm of Needleman & Wunsch, J Mol Biol. 48, 443 (1970), the similarity search method of Pearson & Lipman, Proc. Natl. Acad. Sci. USA 85, 2444 (1988), computerized implementations of these algorithms (GAP, BESTFIT, FASTA, and TFASTA in the Wisconsin Genetics Software Package (Genetics Computer Group, 575 Science Drive, Madison, Wisconsin)), the best-fit sequence program described by Devereux et al., Nucl. Acid Res. 12, 387-395 (1984), or by inspection. Another algorithm is the BLAST algorithm described in Altschul et al., J Mol Biol. 215, 403-410, (1990) and Karlin et al., Proc. Natl. Acad. Sci. USA 90, 5873-5787 (1993). A particularly useful BLAST program is the WU-BLAST-2 program, which is obtained from Altschul et al., Methods in Enzymology, 266, 460-480 (1996); blast.wustl / edu / blast / README.html. WU-BLAST-2 uses a number of search parameters, which are optionally set to default values. These parameters are dynamic values and are established by the program itself based on the composition of the particular sequence and the composition of the particular database against which the target sequence is being searched; however, the values can be adjusted to increase sensitivity. In addition, another useful algorithm is gapped BLAST, as reported in Altschul et al., (1997) Nucleic Acids Res. 25, 3389-3402. Unless otherwise specified, the percent identity herein is determined using the algorithm available at the Internet address blast.ncbi.nlm.nih.gov / Blast.cgi.
[0452] The terms "internal ribosome entry site", "internal ribosome entry site sequence", "IRES", and "IRES sequence region" are used interchangeably herein and refer to viral or human cellular RNA (e.g., messenger RNA (mRNA) and / or circRNA) cis - elements that allow bypassing of the classical eukaryotic cap - dependent translation initiation step. The classical cap - dependent mechanism used by the vast majority of eukaryotic mRNAs requires the m 7 G cap, the initiator Met - tRNAmet, more than a dozen initiation factor proteins, directed scanning, and GTP hydrolysis to place the translation - competent ribosome at the start codon. An IRES typically consists of: a long and highly structured 5 - UTR that mediates the binding of the translation initiation complex and catalyzes the formation of a functional ribosome.
[0453] The term "IRES - like sequence" or "internal ribosome entry site - like sequence" refers to a synthetic nucleotide sequence that exhibits native IRES function. In some embodiments, the IRES - like sequence can recruit ribosomal components to mediate cap - independent translation.
[0454] When referring to a nucleic acid sequence, the terms "coding sequence", "coding sequence region", "coding region", and "CDS" are used interchangeably herein to refer to, for example, the part of a DNA or RNA sequence that is translated or can be translated into a protein. The terms "reading frame", "open reading frame", and "ORF" are used interchangeably herein to refer to a nucleotide sequence that starts with a start codon (e.g., ATG) and, in some embodiments, ends with a stop codon (e.g., TAA, TAG, or TGA). An open reading frame can contain introns and exons; thus, all CDSs are ORFs, but not all ORFs are CDSs.
[0455] The term "X - mer" refers to having "X" nucleic acid residues / nucleotides in a polynucleotide sequence (e.g., an IRES - like sequence), where "X" is an integer. For example, "10 - mer" means that the polynucleotide sequence of the present disclosure has 10 nucleic acid residues / nucleotides.
[0456] The terms "complementary" and "complementarity" refer to the relationship between two nucleic acid sequences or nucleic acid monomers that can form one or more hydrogen bonds with each other through traditional Watson-Crick base pairing or other non-traditional types of pairing. The degree of complementarity between two nucleic acid sequences can be expressed as the percentage of nucleotides in the nucleic acid sequence that can form hydrogen bonds (e.g., Watson-Crick base pairing) with a second nucleic acid sequence (e.g., about 50%, about 60%, about 70%, about 80%, about 90%, and 100% complementary). Two nucleic acid sequences are "fully complementary" if all consecutive nucleotides of one nucleic acid sequence form hydrogen bonds with the same number of consecutive nucleotides in the second nucleic acid sequence. Two nucleic acid sequences are "substantially complementary" if the degree of complementarity between them is at least 60% (e.g., at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 97%, at least 98%, at least 99%, or 100%) over a region of at least 8 nucleotides (e.g., at least 9, at least 10, at least 11, at least 12, at least 13, at least 14, at least 15, at least 16, at least 17, at least 18, at least 19, at least 20, at least 21, at least 22, at least 23, at least 24, at least 25, at least 30, at least 35, at least 40, at least 45, at least 50, or more nucleotides), or if the two nucleic acid sequences hybridize under medium stringency conditions or, in some embodiments, under high stringency conditions. Exemplary medium stringency conditions include overnight incubation at 37°C in a solution containing 20% formamide, 5% SSC (150 mM NaCl, 15 mM sodium citrate), 50 mM sodium phosphate (pH 7.6), 5x Denhardt's solution, 10% dextran sulfate, and 20 mg / ml denatured, sheared salmon sperm DNA, followed by washing the filter in 1*SSC at about 37°C - 50°C, or substantially similar conditions, such as those described as medium stringency conditions in Sambrook, J., Molecular Cloning: A Laboratory Manual, Cold Spring Harbor Laboratory Press; 4th Edition (June 15, 2012).High-stringency conditions are used, e.g., conditions such as: (1) low ionic strength and high temperature washes, such as 0.015 M sodium chloride / 0.0015 M sodium citrate / 0.1% sodium dodecyl sulfate (SDS) at 50°C, (2) use of a denaturing agent, such as formamide, during hybridization, e.g., 50% (v / v) formamide and 0.1% bovine serum albumin (BSA) / 0.1% Ficoll / 0.1% polyvinylpyrrolidone (PVP) / 50 mM sodium phosphate buffer (pH 6.5), and 750 mM sodium chloride and 75 mM sodium citrate at 42°C, or (3) use of 50% formamide, 5x SSC (0.75 M NaCl, 0.075 M sodium citrate), 50 mM sodium phosphate (pH 6.8), 0.1% sodium pyrophosphate, 5x Denhardt's solution, sonicated salmon sperm DNA (50 μg / ml), 0.1% SDS, and 10% dextran sulfate at 42°C, and washing under the following conditions: (i) 42°C in 0.2x SSC, (ii) 55°C in 50% formamide, and (iii) 55°C in 0.1x SSC (optionally in combination with EDTA). Additional details and explanations of hybridization reaction stringency are provided, e.g., in Sambrook, supra, and Ausubel et al., eds., Short Protocols in Molecular Biology, 5th ed., John Wiley & Sons, Inc., Hoboken, N.J. (2002).
[0457] The term "animal" refers to both human and non-human animals, including non-human primates (e.g., bonobos, chimpanzees, gorillas, monkeys) and other animals such as birds, pigs, horses, dogs, cats, rabbits, mice, rats, cows, sheep, goats, and deer.
[0458] When referring to nucleic acid sequences, the term "hybridization" or "hybridized" refers to the association formed between and / or within sequences having complementarity.
[0459] The term "control element" collectively refers to promoter regions, polyadenylation signals, transcription termination sequences, upstream regulatory domains, origins of replication, internal ribosome entry sites (IRESs), enhancers, splice junctions, etc., which together provide for the replication, transcription, post-transcriptional processing, and translation of the coding sequence in a recipient cell. Not all of these control elements need to be present so long as the selected coding sequence is capable of being replicated, transcribed, and translated in an appropriate host cell.
[0460] The term "promoter" as used herein refers to a nucleotide region containing a DNA regulatory sequence, where the regulatory sequence is derived from a gene capable of binding to RNA polymerase and allowing transcription of a downstream (3' direction) coding sequence to be initiated. It may contain genetic elements to which regulatory proteins and molecules (such as RNA polymerase and other transcription factors) can bind to initiate the specific transcription of a nucleic acid sequence. The phrases "operably positioned", "operably linked", "under control", and "under transcriptional control" mean that the promoter is in the correct functional position and / or orientation relative to a nucleic acid sequence to control the initiation and / or expression of transcription of this sequence.
[0461] "Enhancer" means a nucleic acid sequence that, when located near a promoter, confers increased transcriptional activity relative to the transcriptional activity caused by the promoter in the absence of the enhancer domain.
[0462] "Operably linked" with respect to a nucleic acid molecule means that two or more nucleic acid molecules (e.g., the nucleic acid molecule to be transcribed, a promoter, and a functional effector element) are linked in such a way as to permit transcription of the nucleic acid molecule.
[0463] The term "homology" refers to the percentage of identity between the nucleic acid residues of two polynucleotides or the amino acid residues of two polypeptides. The correspondence between one sequence and another can be determined by techniques known in the art. For example, homology can be determined by aligning sequence information and directly comparing the sequence information between two polypeptides using readily available computer programs. As determined using the above methods, two polynucleotides (e.g., DNA) or two polypeptide sequences are "substantially homologous" to each other when the nucleotides or amino acids match at least about 80%, preferably at least about 90%, and most preferably at least about 95% over the defined molecular length.
[0464] "Treatment" or "treatment of a disease or disorder" means implementing a protocol or treatment plan that may include administering to a patient one or more drugs or active agents in an effort to alleviate the signs or symptoms of the disease or disease recurrence. Desired treatment effects include reducing the rate of disease progression, improving or alleviating the disease state, and regression, increased survival, improved quality of life, or improved prognosis. Remission or prevention can occur before and after the signs or symptoms of the disease or disorder appear. In addition, "treatment" ("treating" or "treatment") does not require complete alleviation of signs or symptoms, nor does it require a cure.
[0465] As used throughout this disclosure, the term "therapeutic benefit" or "therapeutically effective" refers to anything that promotes or enhances the health of a subject in the medical treatment of the disorder. This includes, but is not limited to, reducing the frequency, severity, or rate of progression of disease signs or symptoms. For example, the treatment of cancer can involve, for example, reducing the size of a tumor, reducing the invasiveness of a tumor, reducing the growth rate of cancer, or reducing the rate of metastasis or recurrence. The treatment of cancer can also refer to prolonging the survival period of a subject with cancer.
[0466] The phrase "pharmaceutically or pharmacologically acceptable" refers to molecular entities and compositions that do not produce adverse, allergic, or other untoward reactions when administered to an animal, such as a human, in an appropriate manner. For administration to an animal (e.g., a human), it is understood that the formulation should meet sterility, pyrogenicity, general safety, and purity standards as required by, for example, the Office of Biologics Standards of the FDA.
[0467] As used herein, "pharmaceutically acceptable carrier" includes any and all aqueous biocompatible solvents (e.g., saline solution, phosphate buffered saline, parenteral vehicles such as sodium chloride, Ringer's dextrose, etc.), antioxidants, preservatives (e.g., antibacterial or antifungal agents, antioxidants, chelating agents, and inert gases), isotonic agents, such like materials and combinations thereof, as known to those of ordinary skill in the art. The pH and exact concentration of the various components in the pharmaceutical composition are adjusted according to well-known parameters.
[0468] As used herein, the term "bulged adenosine" (also referred to as bulged A) refers to a residue in an intron sequence that serves as the initiating nucleophile. Conventional group II bulged adenosines are located in domain 6 (D6 or DVI) of group II introns. In some group II introns, the bulged adenosine is 7 or 8 nucleotides from the 3' splice site. The bulged adenosine is generally conserved and plays a central role in the splicing process. During this process, the 2'-hydroxyl of the bulged adenosine attacks the 5' splice site, followed by a nucleophilic attack of the 3'-OH of the upstream exon on the 3' splice site. This results in a branched intron lariat that is joined at the bulged adenosine by a 2'-phosphodiester bond (Van der Veen et al., The EMBO Journal, 6(12):3827-3831 (1987); Jacquier et al., Journal of molecular biology, 219(3):415-428 (1991); Daniels et al., Journal of molecular biology, 256(1):31-49 (1996)).
[0469] As used herein, the term "atypical bulged adenosine" (also referred to as atypical bulge A) refers to a region found within D6 of group IIB intron C.te I1 (Cte), which is found in the human pathogen Clostridium tetani. D6 of Cte does not have an obvious bulged adenosine. Instead, it has a loop region, see Figure 2 A-2E, which acts as a nucleophile in the first step of splicing (the branching pathway) (McNeil et al., RNA, 20(6):855-866 (2014)).
[0470] As used herein, the term "catalytic triad" refers to a highly conserved region (AGC) that forms a base triad with other nucleotides to form a triple helix called a "catalytic triplex". This catalytic triplex forms the binding pocket for two active site magnesium ions. (Chan et al. Nature Comm. 9.1 (2018):1-10)
[0471] As used herein, the term "scar" refers to a non-target sequence region in a circRNA splicing product. As used herein, the term "scarless splicing" refers to the self-splicing of a cRNAzyme that produces a circRNA that does not contain additional sequence elements outside of the target sequence. Thus, a "scarless" circRNA does not contain a scar, meaning it consists only of the target sequence. As used herein, the term "near-scarless splicing" refers to the self-splicing of a cRNAzyme that produces a circRNA that contains no more than 20 nucleotides outside of the target sequence. A "near-scarless circRNA" is a circRNA produced by "near-scarless" self-splicing of a cRNAzyme that, in addition to the target sequence, includes no more than 20 nucleotides.
[0472] The term "in vitro transcription" or "IVT" refers to a general method for producing RNA in vitro that uses RNA polymerase, ribonucleotides, and appropriate buffer conditions to synthesize RNA from a DNA template.
[0473] As used herein, the term "resulting target sequence" refers to the target sequence formed in a circRNA after self-splicing of the RNA (or cRNAzyme) provided herein.
[0474] 8.2. Cap-Independent Translation Initiation Elements
[0475] In some embodiments, the present disclosure provides an RNA polynucleotide comprising an engineered translation initiation element (TI) comprising an internal ribosome entry site (IRES)-like polynucleotide sequence. In some embodiments, the IRES-like polynucleotide sequence can mediate cap-independent translation initiation.
[0476] Translation initiation of mRNA in eukaryotic cells is a complex process that involves the coordinated interaction of many factors (Pain (1996) Eur. J. Biochem. 236, 747-771). For most mRNAs, the first step is the recruitment of the ribosomal 40S subunit to the mRNA at or near the capped 5' end. The cap-binding protein complex eIF4F greatly facilitates the association of the 40S with the mRNA. The eIF4F factor consists of three subunits: the RNA helicase eIF4A, the cap-binding protein eIF4E, and the multi-adaptor protein eIF4G, which serves as a scaffold for the proteins in the complex and has binding sites for eIF4E, eIF4a, eIF3, and the poly(A)-binding protein.
[0477] Circular RNAs are a type of single-stranded covalently closed-loop RNAs that do not contain the 5' cap that is typically known to be required for cap-dependent translation. Thus, circular RNA translation utilizes alternative mechanisms to initiate cap-independent translation, such as using internal ribosome entry site (IRES) sequences that are recognized by ribosomes. See, for example, Wesselhoeft, R. A. et al., Nat. Commun. 9, 2629 (2018) and Chinese Patent Application No. 202110594352.4.
[0478] 8.2.1. Natural internal ribosome entry site (IRES) sequences
[0479] In some embodiments, the RNA polynucleotides, precursor RNAs, and circular RNAs provided herein contain natural internal ribosome entry site (IRES) sequences. In some embodiments, the IRES sequence is an RNA sequence capable of engaging ribosomes (e.g., eukaryotic ribosomes). The IRES sequence allows for the translation of one or more open reading frames (e.g., open reading frames that form an expression sequence) from the circular RNA. The IRES attracts ribosomes (e.g., eukaryotic ribosomes) to the translation initiation complex and facilitates translation initiation.
[0480] A large number of natural IRES sequences are available and include sequences derived from or isolated from various viruses, such as the leader sequences of picornaviruses such as: encephalomyocarditis virus (EMCV) UTR, poliovirus leader sequence, hepatitis A virus leader sequence, hepatitis C virus IRES, human rhinovirus type 2 IRES, IRES element from foot-and-mouth disease virus, Giardia virus IRES, and the like.
[0481] In some embodiments, the native IRES sequence is isolated from or derived from the IRES sequences of the following: Taura syndrome virus, Tiiatoma virus, Theiler's encephalomyelitis virus, simian virus 40, Solenopsis invicta virus 1, Rhopalosiphum padi virus, reticuloendotheliosis virus, human poliovirus 1, Plautia stali intestine virus, Kashmir bee virus, human rhinovirus 2, Homalodisca coagulata virus-1, human immunodeficiency virus type 1, Himetobi P virus, hepatitis C virus, hepatitis A virus, hepatitis A virus HA16, hepatitis GB virus, foot-and-mouth disease virus, human enterovirus 71, equine rhinitis virus, Ectropis obliqua picorna-like virus, encephalomyocarditis virus, Drosophila C virus, human coxsackievirus B3, tobacco mosaic virus, cricket paralysis virus, bovine viral diarrhea virus 1, black queen cell virus, aphid lethal paralysis virus, avian encephalomyelitis virus, acute bee paralysis virus, Hibiscus chlorotic ringspot virus, classical swine fever virus, tobacco etch virus, turnip crinkle virus, EMCV-A, EMCV-B, EMCV-Bf, EMCV-Cf, EMCV pEC9, Picobimavirus, HCV QC64, human Cosavirus E / D, human Cosavirus F, human Cosavirus JMY, rhinovirus NAT001, HRV14, HRV89, HRVC-02, HRV-A21, Salivirus A SHI, Salivirus FHB, Salivirus NG-J1, human parechovirus 1, Crohivirus B, Yc-3, Rosavirus M-7, Shanbavirus A, Pasivirus A, Pasivirus A2, echovirus E14, human parechovirus 5, Aichi virus, Phopivirus, CVA10, enterovirus C, enterovirus D, enterovirus J, human Pegivirus 2, GBV-C GT110, GBV-C K1737, GBV-C Iowa, Pegivirus A1220, Pasivirus A3, Sapelovirus, Rosavirus B, Bakunsa Virus, Tremovirus A, porcine Pasivirus 1, PLV-CHN, Pasivirus A, Sicinivirus, HepacivirusK, Hepacivirus A, BVDV1, Border disease virus, BVDV2, CSFV-PK15C, SF573 Dicistravirus, Hubei picornavirus-like virus, CRPV, Salivirus ABN5, Salivirus ABN2, Salivirus A02394, Salivirus AGUT, Salivirus ACH, Salivirus ASZ1, Salivirus FHB, Coxsackievirus (e.g., CVA3, CVA12, CVB1, CVB3, CVB5), Echovirus 7, Enterovirus A71, and / or EV24.
[0482] In some embodiments, the native IRES sequence is isolated from or derived from a eukaryotic IRES element selected from: human FGF2, human SFTPA1, human AML1 / RUNX1, Drosophila Antennapedia, human AQP4, human AT1R, human BAG-1, human BCL2, human BiP, human c-IAP1, human c-myc, human eIF4G, mouse NDST4L, human LEF1, mouse HIF1α, human n.myc, mouse Gtx, human p27kip1, human PDGF2 / c-sis, human p53, human Pim-1, mouse Rbm3, Drosophila reaper, canine Scamper, Drosophila Ubx, human UNR, mouse UtrA, human VEGF-A, human XIAP, Drosophila hairless, Saccharomyces cerevisiae TFIID, and Saccharomyces cerevisiae YAP1.
[0483] In some embodiments, the native IRES sequence is derived from or isolated from an endogenous IRES sequence of Homo sapiens. In some embodiments, the native IRES sequence is derived from or isolated from an endogenous IRES sequence of a human tissue or a human sample.
[0484] In some embodiments, the IRES sequence is isolated from or derived from a cellular IRES element selected from the group consisting of: AML1 / RUNX1, Antp-D, Antp-DE, Antp-CDE, ATlR varl, ATlR_var2, ATlR_var3, ATlR_var4, BAGl_p36delta236nt, BAGl_p36, BiP_-222_-3, C-IAP1 285-1399, c IAP1 13 13-1462, c-jun, Cat-l_224, CCND1, eIF4GI-ext, eIF4GII, eIF4GII-long, FGF1A, FMR1, Gtx-l33-l4l, Gtx-l-l66, Gtx-l-l20, Gtx-l-l96, HAP4, HIFla, hSNMl, HsplOl, hsp70, hsp70, Hsp90, IGF2_leader2, L-myc, MNT 75-267, MNT 36-160, MTG8a, MYB, MYT2 997-1 152, NRF_-653_-l7, NtHSFl, ODC1, p27kipl, p53_l28-269, PDGF2 / c-sis, PITSLRE_p58, Rbm3, reaper, Scamper, TFIID, TIF4631, Ubx_l-966, Ubx_373-96l, UNR, Ure2, XIAP 5-464, XIAP 305-466, YAP1, (GAAA)l6, (PPTl9)4, and XI.
[0485] In some embodiments, the IRES sequence is isolated from or derived from a viral IRES element selected from: ABPVIGRpred, AEV, ALPV IGRpred, BQCV IGRpred, BVDV1 1-385, BVDV1 29-391, CrPV 5NCR, CrPVIGR, crTMV_IRESmp228, CSFV, DCV IGR, EoPV_5NTR, ERBV_162-920, EV71_1-748, FMDV serotype C, GBV-A, GBV-C, HAV HM175, HiPVJGRpred, HIV-1, HoCV1JGRpred, IAPVJGRpred, idefix, KBVIGRpred, PSIV IGR, PV type 1 Mahoney, PV_type3_Leon, REV-A, RhPV 5NCR, RhPV IGR, SINV 1IGRpred, SV40 661-830, TMEV, TMV_UI_IRESmp228, TRV 5NTR, TrV IGR, TSV, and IGR.
[0486] In some embodiments, the IRES sequence of the present disclosure comprises a sequence isolated from or derived from a native IRES sequence. Exemplary native IRES sequences include, but are not limited to, those listed in Tables 98-100. In some embodiments, the IRES sequence comprises a nucleic acid sequence that is at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 100% identical to the nucleic acid sequence or a functional portion thereof of the native IRES sequences of Tables 98-100.
[0487] 8.2.2. IRES-like sequences
[0488] In some embodiments, the RNA polynucleotides, precursor RNAs, and circular RNAs provided herein comprise IRES-like sequences. As used herein, the term "IRES-like sequence" refers to a synthetic or artificial IRES sequence that has the function of a native IRES sequence. In some embodiments, the IRES-like sequence is an RNA sequence capable of recruiting ribosomes (e.g., eukaryotic ribosomes). In some embodiments, the IRES-like sequence is an RNA sequence capable of mediating cap-independent translation initiation. IRES-like sequences can be identified by any of the methods disclosed herein.
[0489] In some embodiments, the length of the IRES-like sequence is greater than or equal to 3 nucleic acid residues. In some embodiments, the length of the IRES-like sequence is 3 - 300 nucleic acid residues. In some embodiments, the length of the IRES-like sequence is 3 - 200 nucleic acid residues. In some embodiments, the length of the IRES-like sequence is 4 - 200 nucleic acid residues. In some embodiments, the length of the IRES-like sequence is 5 - 200 nucleic acid residues. In some embodiments, the length of the IRES-like sequence is 6 - 200 nucleic acid residues. In some embodiments, the length of the IRES-like sequence is 7 - 200 nucleic acid residues. In some embodiments, the length of the IRES-like sequence is 3 - 100 nucleic acid residues. In some embodiments, the length of the IRES-like sequence is 4 - 100 nucleic acid residues. In some embodiments, the length of the IRES-like sequence is 5 - 100 nucleic acid residues. In some embodiments, the length of the IRES-like sequence is 6 - 100 nucleic acid residues. In some embodiments, the length of the IRES-like sequence is 7 - 100 nucleic acid residues. In some embodiments, the length of the IRES-like sequence is 3 - 50 nucleic acid residues. In some embodiments, the length of the IRES-like sequence is 4 - 50 nucleic acid residues. In some embodiments, the length of the IRES-like sequence is 5 - 50 nucleic acid residues. In some embodiments, the length of the IRES-like sequence is 6 - 50 nucleic acid residues. In some embodiments, the length of the IRES-like sequence is 7 - 50 nucleic acid residues. In some embodiments, the length of the IRES-like sequence is 3 - 40 nucleic acid residues. In some embodiments, the length of the IRES-like sequence is 4 - 40 nucleic acid residues. In some embodiments, the length of the IRES-like sequence is 5 - 40 nucleic acid residues. In some embodiments, the length of the IRES-like sequence is 6 - 40 nucleic acid residues. In some embodiments, the length of the IRES-like sequence is 7 - 40 nucleic acid residues. In some embodiments, the length of the IRES-like sequence is 3 - 30 nucleic acid residues. In some embodiments, the length of the IRES-like sequence is 4 - 30 nucleic acid residues. In some embodiments, the length of the IRES-like sequence is 5 - 30 nucleic acid residues. In some embodiments, the length of the IRES-like sequence is 6 - 30 nucleic acid residues. In some embodiments, the length of the IRES-like sequence is 7 - 30 nucleic acid residues. In some embodiments, the length of the IRES-like sequence is 3 - 20 nucleic acid residues. In some embodiments, the length of the IRES-like sequence is 4 - 20 nucleic acid residues. In some embodiments, the length of the IRES-like sequence is 5 - 20 nucleic acid residues. In some embodiments, the length of the IRES-like sequence is 6 - 20 nucleic acid residues. In some embodiments, the length of the IRES-like sequence is 7 - 20 nucleic acid residues.In some embodiments, the length of the IRES-like sequence is 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99 or 100 nucleic acid residues. In some embodiments, the length of the IRES-like sequence is 5 nucleic acid residues. In some embodiments, the length of the IRES-like sequence is 6 nucleic acid residues. In some embodiments, the length of the IRES-like sequence is 7 nucleic acid residues. In some embodiments, the length of the IRES-like sequence is 8 nucleic acid residues. In some embodiments, the length of the IRES-like sequence is 9 nucleic acid residues. In some embodiments, the length of the IRES-like sequence is 10 nucleic acid residues. In some embodiments, the length of the IRES-like sequence is 11 nucleic acid residues. In some embodiments, the length of the IRES-like sequence is 12 nucleic acid residues.
[0490] Exemplary IRES-like sequences of the present disclosure include, but are not limited to, those listed in Tables 2-96, 110, and 112. In some embodiments, the IRES-like sequence comprises a nucleic acid sequence that is at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99% or at least 100% identical to the nucleic acid sequence of the IRES-like sequences of Tables 2-96, 110, and 112 or a functional portion thereof.
[0491] In some embodiments, the IRES-like sequence or a portion thereof described herein can be combined with the native IRES sequence or a portion thereof described herein to form additional IRES-like sequences. In some embodiments, the complete native IRES sequence is combined with the complete IRES-like sequence. In some embodiments, a portion of the native IRES sequence is combined with the complete IRES-like sequence. In some embodiments, the complete native IRES sequence is combined with a portion of the IRES-like sequence. In some embodiments, a portion of the native IRES sequence is combined with a portion of the IRES-like sequence. Exemplary sequences comprising a combination of a portion of the native IRES sequence and a portion of the IRES-like sequence include, but are not limited to, those listed in Table 97.
[0492] 8.2.2.1. Method for generating IRES-like sequences
[0493] (a) The present disclosure provides a method for generating an IRES-like polynucleotide sequence, the method comprising the steps of: generating a polynucleotide query sequence consisting of X nucleic acid residues (i.e., X-mer), where X is an integer greater than or equal to 3;
[0494] (b) generating X - Y + 1 overlapping polynucleotide fragment sequences within the polynucleotide query sequence,
[0495] where each polynucleotide fragment sequence consists of Y nucleic acid residues,
[0496] where the first position of each polynucleotide fragment sequence is n and the last position of the same polynucleotide fragment sequence is Y + n - 1, and
[0497] where n represents each positive integer between 1 and X - Y + 1;
[0498] (c) determining an enrichment score for each polynucleotide fragment sequence of (b);
[0499] (d) determining a numerical score for the polynucleotide query sequence by summing the enrichment scores of each polynucleotide fragment sequence of (c); and
[0500] (e) identifying the polynucleotide query sequence as an engineered IRES-like polynucleotide sequence according to a reference value.
[0501] In some embodiments, the reference value is determined as described herein, for example, in: 8.2.2.1.a. Generating a polynucleotide query sequence, 6.2.2.1.b. Generating overlapping polynucleotide subsequences within the query sequence.
[0502] In some embodiments, the reference value is a characteristic of the absence of therapeutic product expression.
[0503] In some embodiments, the reference value is the average score of all numerical scores of two or more IRES-like polynucleotide sequences or two or more natural IRES sequences or a combination thereof.
[0504] In some embodiments, the reference value is about 95%, about 90%, about 85%, about 80%, about 75%, about 70%, about 65%, about 60%, about 55%, about 50%, about 45%, about 40%, about 35%, about 30%, about 25%, about 20%, about 15%, about 10%, about 9%, about 8%, about 7%, about 6%, about 5%, about 4%, about 3%, about 2%, about 1%, about 0.9%, about 0.8%, about 0.7%, about 0.6%, about 0.5%, about 0.4%, about 0.3%, about 0.2%, or about 0.1% of the highest numerical score of two or more IRES-like polynucleotide sequences or two or more native IRES sequences or a combination thereof.
[0505] In some embodiments, the reference value is about 50%, about 45%, about 40%, about 35%, about 30%, about 25%, about 20%, about 15%, or about 10% of the highest numerical score of two or more IRES-like polynucleotide sequences or two or more native IRES sequences or a combination thereof.
[0506] In some embodiments, the reference value is at least about 95%, about 90%, about 85%, about 80%, about 75%, about 70%, about 65%, about 60%, about 55%, about 50%, about 45%, about 40%, about 35%, about 30%, about 25%, about 20%, about 15%, about 10%, about 9%, about 8%, about 7%, about 6%, about 5%, about 4%, about 3%, about 2%, about 1%, about 0.9%, about 0.8%, about 0.7%, about 0.6%, about 0.5%, about 0.4%, about 0.3%, about 0.2%, or about 0.1% higher than the lowest numerical score of two or more IRES-like polynucleotide sequences or two or more native IRES sequences or a combination thereof.
[0507] In some embodiments, the reference value is at least about 95%, about 90%, about 85%, about 80%, about 75%, about 70%, about 65%, about 60%, about 55%, about 50% higher than the lowest numerical score of two or more IRES-like polynucleotide sequences or two or more native IRES sequences or a combination thereof.
[0508] In some embodiments, the reference value is greater than or equal to 0.
[0509] In some embodiments, the numerical score of the polynucleotide query sequence is calculated using a system of equations, wherein the numerical score of the polynucleotide query sequence is greater than or equal to 0.
[0510] The number of nucleic acid residues in the polynucleotide query sequence (i.e., X) is described in detail herein, for example, in 8.2.2.1.a. Generate a polynucleotide query sequence and 8.2.2.1.b. Generate overlapping polynucleotide subsequences within the query sequence.
[0511] In some embodiments, X is an integer selected from 3 - 300. In some embodiments, X is an integer selected from 3 - 100. In some embodiments, X is an integer selected from 5 - 100. In some embodiments, X is an integer selected from 6 - 100. In some embodiments, X is an integer greater than or equal to 5. In some embodiments, X is an integer greater than or equal to 6.
[0512] In some embodiments, the polynucleotide query sequence is generated within a DNA vector suitable for synthesizing the polynucleotide query sequence. In some embodiments, the polynucleotide query sequence is chemically synthesized.
[0513] In some embodiments, the enrichment score of a polynucleotide fragment sequence is determined by: i) generating an expression plasmid library, wherein each expression plasmid in the library contains a different polynucleotide fragment sequence and a reporter gene; ii) contacting a cell population with the expression plasmid library; iii) quantifying the expression level of the reporter gene corresponding to each expression plasmid in the expression plasmid library; iv) dividing the total cell population into a first population and a second population based on the protein expression level in iii); v) determining the enrichment score of the polynucleotide fragment sequence.
[0514] The enrichment score of a polynucleotide fragment sequence can be determined according to the methods described herein, for example, in 8.2.2.1.c. Determine the enrichment score.
[0515] The numerical score of a polynucleotide query sequence can be determined according to the methods described herein, for example, in 8.2.2.1.d. Determine the numerical score of the polynucleotide sequence and identify IRES - like sequences.
[0516] 8.2.2.1.a. Generate a polynucleotide query sequence
[0517] A library of random polynucleotide query sequences can be generated by any suitable method known in the art or described herein. In some embodiments, the random polynucleotide query sequences (e.g., RNA polynucleotides) are generated using any DNA vector suitable for synthesizing the polynucleotide query sequences. In some embodiments, the random polynucleotide query sequences (e.g., RNA polynucleotides) are generated by chemical synthesis, error-prone PCR, or reverse transcription PCR of a transcriptome. For example, to generate a library of random 10-mer sequences, a hairpin primer (ATTCCGTCTCAAGTAANNNNNNNNNNATCATGGAGACGCACTGTTTTTTTCAGTGCGTCTCCATGA (SEQ ID NO: 15342)) can be generated using the Klenow fragment (NEB), and the resulting DNA can be digested with BsmBI and ligated into pcircGFP-BsmBI digested with BsmBI. The ligation product can be transformed into ElectroMax DH-5α (Invitrogen) to generate E. coli clones to achieve approximately 2-fold coverage of all possible DNA decamers (a total of 2 million clones).
[0518] In some embodiments, the length of the random polynucleotide query sequence is greater than or equal to 3 residues. In some embodiments, the length of the random polynucleotide query sequence is 3 - 300 nucleic acid residues. In some embodiments, the length of the random polynucleotide query sequence is 3 - 200 nucleic acid residues. In some embodiments, the length of the random polynucleotide query sequence is 3 - 100 nucleic acid residues. In some embodiments, the length of the random polynucleotide query sequence is 5 - 100 nucleic acid residues. In some embodiments, the length of the random polynucleotide query sequence is 6 - 100 nucleic acid residues. In some embodiments, the length of the random polynucleotide query sequence is 7 - 100 nucleic acid residues. In some embodiments, the length of the random polynucleotide query sequence is 7 - 90 nucleic acid residues. In some embodiments, the length of the random polynucleotide query sequence is 7 - 80 nucleic acid residues. In some embodiments, the length of the random polynucleotide query sequence is 7 - 70 nucleic acid residues. In some embodiments, the length of the random polynucleotide query sequence is 7 - 60 nucleic acid residues. In some embodiments, the length of the random polynucleotide query sequence is 7 - 50 nucleic acid residues. In some embodiments, the length of the random polynucleotide query sequence is 7 - 40 nucleic acid residues. In some embodiments, the length of the random polynucleotide query sequence is 7 - 30 nucleic acid residues. In some embodiments, the length of the random polynucleotide query sequence is 7 - 20 nucleic acid residues. In some embodiments, the length of the random polynucleotide query sequence is 7 - 15 nucleic acid residues. In some embodiments, the length of the random polynucleotide query sequence is 6 - 90 nucleic acid residues. In some embodiments, the length of the random polynucleotide query sequence is 6 - 80 nucleic acid residues. In some embodiments, the length of the random polynucleotide query sequence is 6 - 70 nucleic acid residues. In some embodiments, the length of the random polynucleotide query sequence is 6 - 60 nucleic acid residues. In some embodiments, the length of the random polynucleotide query sequence is 6 - 50 nucleic acid residues. In some embodiments, the length of the random polynucleotide query sequence is 6 - 40 nucleic acid residues. In some embodiments, the length of the random polynucleotide query sequence is 6 - 30 nucleic acid residues. In some embodiments, the length of the random polynucleotide query sequence is 6 - 20 nucleic acid residues. In some embodiments, the length of the random polynucleotide query sequence is 6 - 15 nucleic acid residues.
[0519] In some embodiments, the length of the random polynucleotide query sequence is at least 5 nucleic acid residues. In some embodiments, the length of the random polynucleotide query sequence is at least 6 nucleic acid residues. In some embodiments, the length of the random polynucleotide query sequence is at least 7 nucleic acid residues. In some embodiments, the length of the random polynucleotide query sequence is at least 8 nucleic acid residues. In some embodiments, the length of the random polynucleotide query sequence is at least 9 nucleic acid residues. In some embodiments, the length of the random polynucleotide query sequence is at least 10 nucleic acid residues. In some embodiments, the length of the random polynucleotide query sequence is at least 11 nucleic acid residues. In some embodiments, the length of the random polynucleotide query sequence is at least 12 nucleic acid residues. In some embodiments, the length of the random polynucleotide query sequence is at least 13 nucleic acid residues. In some embodiments, the length of the random polynucleotide query sequence is at least 14 nucleic acid residues. In some embodiments, the length of the random polynucleotide query sequence is at least 15 nucleic acid residues. In some embodiments, the length of the random polynucleotide query sequence is at least 16 nucleic acid residues. In some embodiments, the length of the random polynucleotide query sequence is at least 17 nucleic acid residues. In some embodiments, the length of the random polynucleotide query sequence is at least 18 nucleic acid residues. In some embodiments, the length of the random polynucleotide query sequence is at least 19 nucleic acid residues. In some embodiments, the length of the random polynucleotide query sequence is at least 20 nucleic acid residues.
[0520] In some embodiments, the length of the random polynucleotide query sequence is 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99 or 100 nucleic acid residues. In some embodiments, the length of the random polynucleotide query sequence is 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99 or 100 nucleic acid residues. In some embodiments, the length of the random polynucleotide query sequence is 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99 or 100 nucleic acid residues.In some embodiments, the length of the random polynucleotide query sequence is 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99 or 100 nucleic acid residues. In some embodiments, the length of the random polynucleotide query sequence is 5 nucleic acid residues. In some embodiments, the length of the random polynucleotide query sequence is 6 nucleic acid residues. In some embodiments, the length of the random polynucleotide query sequence is 7 nucleic acid residues. In some embodiments, the length of the random polynucleotide query sequence is 8 nucleic acid residues. In some embodiments, the length of the random polynucleotide query sequence is 9 nucleic acid residues. In some embodiments, the length of the random polynucleotide query sequence is 10 nucleic acid residues. In some embodiments, the length of the random polynucleotide query sequence is 11 nucleic acid residues. In some embodiments, the length of the random polynucleotide query sequence is 12 nucleic acid residues. In some embodiments, the length of the random polynucleotide query sequence is 13 nucleic acid residues. In some embodiments, the length of the random polynucleotide query sequence is 14 nucleic acid residues. In some embodiments, the length of the random polynucleotide query sequence is 15 nucleic acid residues. In some embodiments, the length of the random polynucleotide query sequence is 16 nucleic acid residues. In some embodiments, the length of the random polynucleotide query sequence is 17 nucleic acid residues. In some embodiments, the length of the random polynucleotide query sequence is 18 nucleic acid residues. In some embodiments, the length of the random polynucleotide query sequence is 19 nucleic acid residues. In some embodiments, the length of the random polynucleotide query sequence is 20 nucleic acid residues.
[0521] In some embodiments, a library of random polynucleotide query sequences is constructed using primers with random regions. In some embodiments, the primers with random regions are hairpin primers. In some embodiments, the library of random polynucleotide query sequences is constructed using the method described by Fan et al. in 2020 (doi: https: / / doi.org / 10.1101 / 473207). In some embodiments, the library of random polynucleotide query sequences is constructed using hairpin primers containing random regions. In some embodiments, the hairpin primers are extended with a DNA polymerase or a DNA polymerase fragment. In some embodiments, the DNA polymerase fragment is the Klenow fragment. In some embodiments, the DNA polymerase is Taq polymerase.
[0522] In some embodiments, the length of the random region of the primer or the reverse primer is 5-mer, 6-mer, 7-mer, 8-mer, 9-mer, 10-mer, 11-mer, 12-mer, 13-mer, 14-mer, 15-mer, 16-mer, 17-mer, 18-mer, 19-mer, 20-mer, 21-mer, 22-mer, 23-mer, 24-mer, 25-mer, 26-mer, 27-mer, 28-mer, 29-mer, 30-mer, 31-mer, 32-mer, 33-mer, 34-mer, 35-mer, 36-mer, 37-mer, 38-mer, 39-mer, 40-mer, 41-mer, 42-mer, 43-mer, 44-mer, 45-mer, 46-mer, 47-mer, 48-mer, 49-mer, 50-mer, 51-mer, 52-mer, 54-mer, 55-mer, 56-mer, 57-mer, 58-mer, 59-mer, 60-mer, 61-mer, 62-mer, 63-mer, 64-mer, 65-mer, 66-mer, 67-mer, 68-mer, 69-mer, 70-mer, 81-mer, 82-mer, 83-mer, 84-mer, 85-mer, 86-mer, 87-mer, 88-mer, 89-mer, 90-mer, 91-mer, 92-mer, 93-mer, 94-mer, 95-mer, 96-mer, 97-mer, 98-mer, 99-mer or 100-mer. In some embodiments, the length of the random region of the reverse primer is 10-mer. In some embodiments, the length of the random region of the reverse primer is 11-mer. In some embodiments, the length of the random region of the reverse primer is 12-mer. In some embodiments, the length of the random region of the reverse primer is 13-mer. In some embodiments, the length of the random region of the reverse primer is 14-mer. In some embodiments, the length of the random region of the reverse primer is 15-mer.
[0523] In some embodiments, the reverse primer comprises the following sequence or consists of the following sequence: AT TCCGTCTCAAGTAANNNNNNNNNNATCATGGAGACGCACTGTTTTT TTCAGTGCGTCTCCATGA (SEQ ID NO:15342), where "N" represents any nucleic acid residue.
[0524] In some embodiments, the resulting PCR product is ligated upstream of an expression cassette (e.g., encoding a reporter protein). In some embodiments, the PCR product obtained from the reverse primer is digested with one or more restriction enzymes and ligated into an expression vector having an expression cassette (e.g., encoding a reporter protein).
[0525] Any suitable expression vector known in the art or described herein can be used. Those skilled in the art will be well able to construct expression vectors by standard recombinant techniques (see, e.g., Sambrook et al., 2001 (supra) and Ausubel et al., 1996 (supra), both incorporated herein by reference) to express the random polynucleotide query sequences of the present disclosure. Vectors include, but are not limited to, plasmids, cosmids, viruses (phages, animal viruses, and plant viruses), and artificial chromosomes (e.g., YACs), such as retroviral vectors (e.g., derived from Moloney murine leukemia virus vectors (MoMLV), MSCV, SFFV, MPSV, SNV, etc.), lentiviral vectors (e.g., derived from HIV-1, HIV-2, SIV, BIV, FIV, etc.), adenovirus (Ad) vectors (including their replication-competent, replication-defective, and empty capsid forms), adeno-associated virus (AAV) vectors, simian virus 40 (SV-40) vectors, bovine papillomavirus vectors, Epstein-Barr virus vectors, yeast-based vectors, bovine papillomavirus (BPV)-based vectors, herpesvirus vectors, vaccinia virus vectors, Harvey murine sarcoma virus vectors, mouse mammary tumor virus vectors, Rous sarcoma virus vectors, parvovirus vectors, poliovirus vectors, vesicular stomatitis virus vectors, Maraba virus vectors, and group B adenovirus enadenotucirev vectors.
[0526] In some embodiments, the expression vector is the pcircGFP-BsmBI vector (Fan et al. 2020 (doi:https: / / doi.org / 10.1101 / 473207). In some embodiments, the restriction enzyme is BsmBI.
[0527] Any suitable reporter protein known in the art or described herein (e.g., a reporter protein encoded by an expression cassette) can be used. In some embodiments, cells containing a construct of the present disclosure (e.g., an RNA polynucleotide having an IRES-like sequence and an expression cassette encoding a reporter protein) can be identified in vitro or in vivo by including a marker in the expression vector. Such a marker will confer an identifiable change on the cell, allowing easy identification of cells containing the expression vector. Generally, a selectable marker is one that confers a property that allows selection. A positive selectable marker is one in which the presence of the marker allows its selection, while a negative selectable marker is one in which its presence prevents its selection. An example of a positive selectable marker is a drug resistance marker (e.g., a gene conferring resistance to neomycin, puromycin, hygromycin, DHFR, GPT, bleomycin, and histidinol). Other types of markers are also contemplated, including screening markers such as GFP.
[0528] Those skilled in the art are also aware of how to use selectable markers in conjunction with FACS analysis. The marker used is not considered important as long as it can be co-expressed with the nucleic acid encoding the gene product. Other examples of selectable markers and screening markers are well known to those skilled in the art. In some embodiments, the reporter protein is a fluorescent protein. In some embodiments, the reporter protein is green fluorescent protein (GFP).
[0529] 8.2.2.1.b. Generate overlapping polynucleotide subsequences within the query sequence
[0530] In some embodiments, a cell population is contacted with an expression plasmid library. In some embodiments, a random polynucleotide query sequence expression plasmid library is transfected into a cell line to generate a cell population. In some embodiments, the transfected cell line is a human cell line. In some embodiments, the transfected cell line is a human HEK293T cell line.
[0531] In some embodiments, the expression level of the reporter gene corresponding to each expression plasmid of the expression plasmid library is quantified. In some embodiments, a random polynucleotide query sequence expression plasmid library is transfected into a cell line to generate a cell population, and the reporter protein expression levels of multiple cells within the population are determined.
[0532] In some embodiments, the cells transfected with the plasmid library as described above are collected and sorted by reporter signal intensity (e.g., fluorescence intensity). In some embodiments, the reporter signal intensity is determined by fluorescence-activated cell sorting (FACS).
[0533] In some embodiments, the reporter signal intensities of multiple cells within a population are ranked. In some embodiments, the multiple cells are divided into two or more populations. In some embodiments, the multiple cells are at least divided into a first population and a second population. In some embodiments, the multiple cells are divided into two populations (e.g., high and low reporter protein expression). In some embodiments, the multiple cells are divided into three populations (e.g., high, medium, and low reporter protein expression). In some embodiments, the multiple cells are divided into four populations (e.g., high, medium, low, and no reporter protein expression).
[0534] In some embodiments, the high-expression population is determined by the signal intensity of the reporter protein. In some embodiments, when ranked by reporter signal intensity, the high-expression population consists of cells having a reporter protein signal intensity in the top 0-0.1%, 0-0.2%, 0-0.3%, 0-0.4%, 0-0.5%, 0-0.6%, 0-0.7%, 0-0.8%, 0-0.9%, 0-1%, 0-1.1%, 0-1.2%, 0-1.3%, 0-1.4%, 0-1.5%, 0-1.6%, 0-1.7%, 0-1.8%, 0-1.9%, 0-2%, 0-2.1%, 0-2.2%, 0-2.3%, 0-2.4%, 0-2.5%, 0-2.6%, 0-2.7%, 0-2.8%, 0-2.9%, 0-3%, 0-3.1%, 0-3.2%, 0-3.3%, 0-3.4%, 0-3.5%, 0-3.6%, 0-3.7%, 0-3.8%, 0-3.9%, 0-4%, 0-4.1%, 0-4.2%, 0-4.3%, 0-4.4%, 0-4.5%, 0-4.6%, 0-4.7%, 0-4.8%, 0-4.9%, 0-5%, 0-5.1%, 0-5.2%, 0-5.3%, 0-5.4%, 0-5.5%, 0-5.6%, 0-5.7%, 0-5.8%, 0-5.9%, 0-6%, 0-6.1%, 0-6.2%, 0-6.3%, 0-6.4%, 0-6.5%, 0-6.6%, 0-6.7%, 0-6.8%, 0-6.9%, 0-7%, 0-7.1%, 0-7.2%, 0-7.3%, 0-7.4%, 0-7.5%, 0-7.6%, 0-7.7%, 0-7.8%, 0-7.9%, 0-8%, 0-8.1%, 0-8.2%, 0-8.3%, 0-8.4%, 0-8.5%, 0-8.6%, 0-8.7%, 0-8.8%, 0-8.9%, 0-9%, 0-9.1%, 0-9.2%, 0-9.3%, 0-9.4%, 0-9.5%, 0-9.6%, 0-9.7%, 0-9.8%, 0-9.9% or 0-10% of the total cell population. In some embodiments, when ranked by reporter signal intensity, the high-expression population consists of cells having a reporter protein signal intensity in the top approximately 0.1%, approximately 0.2%, approximately 0.3%, approximately 0.4%, approximately 0.5%, approximately 0.6%, approximately 0.7%, approximately 0.8%, approximately 0.9%, approximately 1%, approximately 2%, approximately 3%, approximately 4%, approximately 5%, approximately 6%, approximately 7%, approximately 8%, approximately 9% or approximately 10% of the total cell population. In some embodiments, when ranked by reporter signal intensity, the high-expression population consists of cells having a reporter protein signal intensity in the 0-10% of the total cell population.
[0535] In some embodiments, the low-expression population is determined by the signal intensity of a reporter protein. In some embodiments, as ranked by reporter signal intensity, the low-expression population consists of cells having reporter protein signal intensities in the bottom 0 - 0.1%, 0 - 0.2%, 0 - 0.3%, 0 - 0.4%, 0 - 0.5%, 0 - 0.6%, 0 - 0.7%, 0 - 0.8%, 0 - 0.9%, 0 - 1%, 0 - 1.1%, 0 - 1.2%, 0 - 1.3%, 0 - 1.4%, 0 - 1.5%, 0 - 1.6%, 0 - 1.7%, 0 - 1.8%, 0 - 1.9%, 0 - 2%, 0 - 2.1%, 0 - 2.2%, 0 - 2.3%, 0 - 2.4%, 0 - 2.5%, 0 - 2.6%, 0 - 2.7%, 0 - 2.8%, 0 - 2.9%, 0 - 3%, 0 - 3.1%, 0 - 3.2%, 0 - 3.3%, 0 - 3.4%, 0 - 3.5%, 0 - 3.6%, 0 - 3.7%, 0 - 3.8%, 0 - 3.9%, 0 - 4%, 0 - 4.1%, 0 - 4.2%, 0 - 4.3%, 0 - 4.4%, 0 - 4.5%, 0 - 4.6%, 0 - 4.7%, 0 - 4.8%, 0 - 4.9%, 0 - 5%, 0 - 5.1%, 0 - 5.2%, 0 - 5.3%, 0 - 5.4%, 0 - 5.5%, 0 - 5.6%, 0 - 5.7%, 0 - 5.8%, 0 - 5.9%, 0 - 6%, 0 - 6.1%, 0 - 6.2%, 0 - 6.3%, 0 - 6.4%, 0 - 6.5%, 0 - 6.6%, 0 - 6.7%, 0 - 6.8%, 0 - 6.9%, 0 - 7%, 0 - 7.1%, 0 - 7.2%, 0 - 7.3%, 0 - 7.4%, 0 - 7.5%, 0 - 7.6%, 0 - 7.7%, 0 - 7.8%, 0 - 7.9%, 0 - 8%, 0 - 8.1%, 0 - 8.2%, 0 - 8.3%, 0 - 8.4%, 0 - 8.5%, 0 - 8.6%, 0 - 8.7%, 0 - 8.8%, 0 - 8.9%, 0 - 9%, 0 - 9.1%, 0 - 9.2%, 0 - 9.3%, 0 - 9.4%, 0 - 9.5%, 0 - 9.6%, 0 - 9.7%, 0 - 9.8%, 0 - 9.Composed of 9%, 0 - 10%, 0 - 11%, 0 - 12%, 0 - 13%, 0 - 14%, 0 - 15%, 0 - 16%, 0 - 17%, 0 - 18%, 0 - 19%, 0 - 20%, 0 - 21%, 0 - 22%, 0 - 23%, 0 - 24%, 0 - 25%, 0 - 30%, 0 - 35%, 0 - 40%, 0 - 45%, 0 - 50%, 0 - 55%, 0 - 60%, 0 - 65%, 0 - 70%, 0 - 75%, 0 - 80%, 0 - 85%, 0 - 90%, 5% - 10%, 5% - 11%, 5% - 12%, 5% - 13%, 5% - 14%, 5% - 15%, 5% - 16%, 5% - 17%, 5% - 18%, 5% - 19%, 5% - 20%, 5% - 21%, 5% - 22%, 5% - 23%, 5% - 24%, 5% - 25%, 5% - 30%, 5% - 35%, 5% - 40%, 5% - 45%, 5% - 50%, 5% - 55%, 5% - 60%, 5% - 65%, 5% - 70%, 5% - 75%, 5% - 80%, 5% - 85%, 5% - 90%, 10% - 15%, 10% - 16%, 10% - 17%, 10% - 18%, 10% - 19%, 10% - 20%, 10% - 21%, 10% - 22%, 10% - 23%, 10% - 24%, 10% - 25%, 10% - 30%, 10% - 35%, 10% - 40%, 10% - 45%, 10% - 50%, 10% - 55%, 10% - 60%, 10% - 65%, 10% - 70%, 10% - 75%, 10% - 80%, 10% - 85%, or 10% - 90% of the cells. In some embodiments, as ranked by reporter signal intensity, the low - expressing population is composed of cells with reporter protein signal intensity in the bottom approximately 0.1%, approximately 0.2%, approximately 0.3%, approximately 0.4%, approximately 0.5%, approximately 0.6%, approximately 0.7%, approximately 0.8%, approximately 0.9%, approximately 1%, approximately 2%, approximately 3%, approximately 4%, approximately 5%, approximately 6%, approximately 7%, approximately 8%, approximately 9%, approximately 10%, approximately 15%, approximately 20%, approximately 25%, approximately 30%, approximately 35%, approximately 40%, approximately 45%, approximately 50%, approximately 55%, approximately 60%, approximately 65%, approximately 70%, approximately 75%, approximately 80%, approximately 85%, or approximately 90% of the total cell population.
[0536] In some embodiments, the medium-expression population is determined by the reporter protein signal intensity. In some embodiments, as ranked by reporter signal intensity, the medium-expression population consists of cells having a reporter protein signal intensity between about 49%-51%, 48%-52%, 47%-53%, 46%-54%, 45%-55%, 44%-56%, 43%-57%, 42%-58%, 41%-59%, 40%-60%, 39%-61%, 38%-62%, 37%-63%, 36%-64%, 35%-65%, 34%-66%, 33%-67%, 32%-68%, 31%-69%, 30%-70%, 29%-71%, 28%-72%, 27%-73%, 26%-74%, or 25%-75% of the total cell population. In some embodiments, as ranked by reporter signal intensity, the selected expression population consists of cells having a reporter protein signal intensity between about 49%-51%, 48%-52%, 47%-53%, 46%-54%, 45%-55%, 44%-56%, 43%-57%, 42%-58%, 41%-59%, 40%-60%, 39%-61%, 38%-62%, 37%-63%, 36%-64%, 35%-65%, 34%-66%, 33%-67%, 32%-68%, 31%-69%, 30%-70%, 29%-71%, 28%-72%, 27%-73%, 26%-74%, or 25%-75% of the total cell population.
[0537] In some embodiments, the non-expression population is determined by the reporter protein signal intensity. In some embodiments, as ranked by reporter signal intensity, the non-expression population consists of cells having a reporter protein signal intensity in the bottom 0-5%, 0-10%, 0-15%, 0-20%, 0-25%, 0-30%, 0-35%, 0-40%, 0-45%, or 0-50% of the total cell population. In some embodiments, as ranked by reporter signal intensity, the low-expression population consists of cells having a reporter protein signal intensity in the bottom about 0.1%, about 0.2%, about 0.3%, about 0.4%, about 0.5%, about 0.6%, about 0.7%, about 0.8%, about 0.9%, about 1%, about 2%, about 3%, about 4%, about 5%, about 6%, about 7%, about 8%, about 9%, about 10%, about 11%, about 12%, about 13%, about 14%, about 15%, about 16%, about 17%, about 18%, about 19%, about 20%, about 21%, about 22%, about 23%, about 24%, about 25%, about 30%, about 35%, about 40%, about 45%, or about 50% of the total cell population.
[0538] In some embodiments, the high-expression population, ranked by reporter signal intensity, consists of cells with reporter protein signal intensities in the top 0 - 10% of the total cell population, and the low-expression population, ranked by reporter signal intensity, consists of cells with reporter protein signal intensities in the bottom 0 - 90% of the total cell population. In some embodiments, the high-expression population, ranked by reporter signal intensity, consists of cells with reporter protein signal intensities in the top 0 - 5% of the total cell population, and the low-expression population, ranked by reporter signal intensity, consists of cells with reporter protein signal intensities in the bottom 0 - 80% of the total cell population. In some embodiments, the high-expression population, ranked by reporter signal intensity, consists of cells with reporter protein signal intensities in the top 0 - 1% of the total cell population, and the low-expression population, ranked by reporter signal intensity, consists of cells with reporter protein signal intensities in the bottom 0 - 75% of the total cell population. In some embodiments, the high-expression population, ranked by reporter signal intensity, consists of cells with reporter protein signal intensities in the top 0 - 0.5% of the total cell population, and the low-expression population, ranked by reporter signal intensity, consists of cells with reporter protein signal intensities in the bottom 0 - 75% of the total cell population. In some embodiments, the high-expression population, ranked by reporter signal intensity, consists of cells with reporter protein signal intensities in the top approximately 0.5% of the total cell population, and the low-expression population, ranked by reporter signal intensity, consists of cells with reporter protein signal intensities in the bottom approximately 75% of the total cell population.
[0539] In some embodiments, a reference value for identifying a polynucleotide query sequence as an engineered IRES-like polynucleotide sequence is determined based on reporter signal intensity, the high-expression population, and / or the low-expression population.
[0540] In some embodiments, the reference value for identifying a polynucleotide query sequence as an engineered IRES-like polynucleotide sequence is determined using the following formula: Reference value = (23.75 x the length of the polynucleotide query sequence) – 84.85 – Z, where Z represents any number from 1 - 10000, and X represents the length of the IRES-like polynucleotide sequence. In some embodiments, Z represents any number from 50 - 5000. In some embodiments, Z represents any number from 175 - 2500. In some embodiments, Z represents any number from 150 - 200.
[0541] In some embodiments, the number represented by Z is: at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 20, at least 30, at least 40, at least 50, at least 60, at least 70, at least 80, at least 90, at least 100, at least 200, at least 300, at least 400, at least 500, at least 600, at least 700, at least 800, at least 900, at least 1000, at least 1500, at least 2000, at least 2500, at least 3000, at least 3500, at least 4000, at least 4500, at least 5000, at least 5500, at least 6000, at least 6500, at least 7000, at least 7500, at least 8000, at least 8500, at least 9000, at least 9500 or at least 10000.
[0542] In some embodiments, the number represented by Z is: about 1, about 2, about 3, about 4, about 5, about 6, about 7, about 8, about 9, about 10, about 20, about 30, about 40, about 50, about 60, about 70, about 80, about 90, about 100, about 200, about 300, about 400, about 500, about 600, about 700, about 800, about 900, about 1000, about 1500, about 2000, about 2500, about 3000, about 3500, about 4000, about 4500, about 5000, about 5500, about 6000, about 6500, about 7000, about 7500, about 8000, about 8500, about 9000, about 9500 or about 10000.
[0543] In some embodiments, a random polynucleotide query sequence is recovered from a sorted cell population, and the enriched sequences are analyzed for the random polynucleotide query sequence. In some embodiments, the length of the random sequence recovered from the sorted cell population is 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99 or 100 nucleic acid residues.
[0544] In some embodiments, X-mer random polynucleotide query sequences (i.e., random polynucleotide query sequences of X nucleic acid residues) are recovered from high-expression populations and low-expression populations. In some embodiments, the X-mer random polynucleotide sequences are from high-expression populations and non-expression populations.
[0545] In some embodiments, 4-mers to 100-mers, 4-mers to 90-mers, 4-mers to 80-mers, 4-mers to 70-mers, 4-mers to 60-mers, 4-mers to 50-mers, 4-mers to 40-mers, 4-mers to 30-mers, 4-mers to 20-mers, 4-mers to 19-mers, 4-mers to 18-mers, 4-mers to 17-mers, 4-mers to 16-mers, 4-mers to 15-mers, 4-mers to 14-mers, 4-mers to 13-mers, 4-mers to 12-mers, 4-mers to 11-mers, 4-mers to 10-mers, 4-mers to 9-mers, 4-mers to 8-mers, 4-mers to 7-mers, 5-mers to 100-mers, 5-mers to 90-mers, 5-mers to 80-mers, 5-mers to 70-mers, 5-mers to 60-mers, 5-mers to 50-mers, 5-mers to 40-mers, 5-mers to 30-mers, 5-mers to 20-mers, 5-mers to 19-mers, 5-mers to 18-mers, 5-mers to 17-mers, 5-mers to 16-mers, 5-mers to 15-mers, 5-mers to 14-mers, 5-mers to 13-mers, 5-mers to 12-mers, 5-mers to 11-mers, 5-mers to 10-mers, 5-mers to 9-mers, 5-mers to 8-mers, 5-mers to 7-mers, 6-mers to 100-mers, 6-mers to 90-mers, 6-mers to 80-mers, 6-mers to 70-mers, 6-mers to 60-mers, 6-mers to 50-mers, 6-mers to 40-mers, 6-mers to 30-mers, 6-mers to 20-mers, 6-mers to 19-mers, 6-mers to 18-mers, 6-mers to 17-mers, 6-mers to 16-mers, 6-mers to 15-mers, 6-mers to 14-mers, 6-mers to 13-mers, 6-mers to 12-mers, 6-mers to 11-mers, 6-mers to 10-mers, 6-mers to 9-mers, 6-mers to 8-mers, 6-mers to 7-mers, 7-mers to 100-mers, 7-mers to 90-mers, 7-mers to 80-mers, 7-mers to 70-mers, 7-mers to 60-mers, 7-mers to 50-mers, 7-mers to 40-mers, 7-mers to 30-mers, 7-mers to 20-mers,7-mer to 19-mer, 7-mer to 18-mer, 7-mer to 17-mer, 7-mer to 16-mer, 7-mer to 15-mer, 7-mer to 14-mer, 7-mer to 13-mer, 7-mer to 12-mer, 7-mer to 11-mer, 7-mer to 10-mer, 7-mer to 9-mer, or 7-mer to 8-mer random polynucleotide query sequences (i.e., random polynucleotide query sequences of X nucleic acid residues). In some embodiments, 4-mer to 100-mer, 4-mer to 90-mer, 4-mer to 80-mer, 4-mer to 70-mer, 4-mer to 60-mer, 4-mer to 50-mer, 4-mer to 40-mer, 4-mer to 30-mer, 4-mer to 20-mer, 4-mer to 19-mer, 4-mer to 18-mer, 4-mer to 17-mer, 4-mer to 16-mer, 4-mer to 15-mer, 4-mer to 14-mer, 4-mer to 13-mer, 4-mer to 12-mer, 4-mer to 11-mer, 4-mer to 10-mer, 4-mer to 9-mer, 4-mer to 8-mer, 4-mer to 7-mer, 5-mer to 100-mer, 5-mer to 90-mer, 5-mer to 80-mer, 5-mer to 70-mer, 5-mer to 60-mer, 5-mer to 50-mer, 5-mer to 40-mer, 5-mer to 30-mer, 5-mer to 20-mer, 5-mer to 19-mer, 5-mer to 18-mer, 5-mer to 17-mer, 5-mer to 16-mer, 5-mer to 15-mer, 5-mer to 14-mer, 5-mer to 13-mer, 5-mer to 12-mer, 5-mer to 11-mer, 5-mer to 10-mer, 5-mer to 9-mer, 5-mer to 8-mer, 5-mer to 7-mer, 6-mer to 100-mer, 6-mer to 90-mer, 6-mer to 80-mer, 6-mer to 70-mer, 6-mer to 60-mer, 6-mer to 50-mer, 6-mer to 40-mer, 6-mer to 30-mer, 6-mer to 20-mer, 6-mer to 19-mer, 6-mer to 18-mer, 6-mer to 17-mer, 6-mer to 16-mer, 6-mer to 15-mer, 6-mer to 14-mer, 6-mer to 13-mer, 6-mer to 12-mer, 6-mer to 11-mer,6-mer to 10-mer, 6-mer to 9-mer, 6-mer to 8-mer, 6-mer to 7-mer, 7-mer to 100-mer, 7-mer to 90-mer, 7-mer to 80-mer, 7-mer to 70-mer, 7-mer to 60-mer, 7-mer to 50-mer, 7-mer to 40-mer, 7-mer to 30-mer, 7-mer to 20-mer, 7-mer to 19-mer, 7-mer to 18-mer, 7-mer to 17-mer, 7-mer to 16-mer, 7-mer to 15-mer, 7-mer to 14-mer, 7-mer to 13-mer, 7-mer to 12-mer, 7-mer to 11-mer, 7-mer to 10-mer, 7-mer to 9-mer, or 7-mer to 8-mer random polynucleotide sequences are from high-expression populations and non-expression populations.
[0546] In some embodiments, an X-mer random polynucleotide sequence (XN) is recovered from a population and extended by appending random nucleotides (r) of a vector (e.g., a DNA vector) sequence to the 5' and / or 3' end to generate a polynucleotide query sequence (e.g., 5’-r–XN–r-3’). In some embodiments, the length of the appended random nucleotides of the vector sequence is 1, 2, 3, 4, or 5 nucleic acid residues. In some embodiments, the length of the appended random nucleotides at each end is 1 nucleic acid residue (e.g., 5’-1r–XN–1r-3’). In some embodiments, the length of the appended random nucleotides is 2 nucleic acid residues (e.g., 5’-2r–XN–2r-3’).
[0547] By way of illustration, in some embodiments, 10-mer random polynucleotide sequences are recovered from high-expression populations and low-expression populations. In some embodiments, the 10-mer random polynucleotide sequences are from high-expression populations and non-expression populations.
[0548] In some embodiments, a 10-mer random polynucleotide sequence (10N) is recovered from a population and extended by appending random nucleotides (r) of a vector (e.g., a DNA vector) sequence to the 5' and / or 3' end to generate a polynucleotide query sequence (e.g., 5’-r–10N–r-3’). In some embodiments, the length of the appended random nucleotides of the vector sequence is 1, 2, 3, 4, or 5 nucleic acid residues. In some embodiments, the length of the appended random nucleotides at each end is 1 nucleic acid residue (e.g., 5’-1r–10N–1r-3’). In some embodiments, the length of the appended random nucleotides is 2 nucleic acid residues (e.g., 5’-2r–10N–2r-3’).
[0549] In some embodiments, a series of overlapping polynucleotide fragment sequences (e.g., overlapping subsequences) are generated from a polynucleotide query sequence. In some embodiments, there are X - Y + 1 overlapping polynucleotide fragment sequences within the polynucleotide query sequence, where X is the length of the polynucleotide query sequence, where each polynucleotide fragment sequence consists of Y nucleic acid residues in length, where the first position of each polynucleotide fragment sequence is n and the last position of the same polynucleotide fragment sequence is Y + n - 1, and where n represents each positive integer between 1 and X - Y + 1.
[0550] In some embodiments, the length of the overlapping polynucleotide fragment sequences is greater than or equal to 2 nucleic acid residues. In some embodiments, the length of the overlapping polynucleotide fragment sequences is 2 - 299 nucleic acid residues. In some embodiments, the length of the overlapping polynucleotide fragment sequences is 2 - 200 nucleic acid residues. In some embodiments, the length of the overlapping polynucleotide fragment sequences is 2 - 100 nucleic acid residues. In some embodiments, the length of the overlapping polynucleotide fragment sequences is 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99 or 100 nucleic acid residues. In some embodiments, the length of the overlapping polynucleotide fragment sequences is 5 nucleic acid residues. In some embodiments, the length of the overlapping polynucleotide fragment sequences is 6 nucleic acid residues. In some embodiments, the length of the overlapping polynucleotide fragment sequences is 7 nucleic acid residues. In some embodiments, the length of the overlapping polynucleotide fragment sequences is 8 nucleic acid residues. In some embodiments, the length of the overlapping polynucleotide fragment sequences is 9 nucleic acid residues. In some embodiments, the length of the overlapping polynucleotide fragment sequences is 10 nucleic acid residues. In some embodiments, the length of the overlapping polynucleotide fragment sequences is 11 nucleic acid residues. In some embodiments, the length of the overlapping polynucleotide fragment sequences is 12 nucleic acid residues.
[0551] 8.2.2.1.c. Determine the enrichment score
[0552] In some embodiments, the enrichment score of a polynucleotide fragment sequence is determined by: i) generating a library of expression plasmids, wherein each expression plasmid of the library contains a different polynucleotide fragment sequence and a reporter gene as described above; ii) contacting a cell population with the library of expression plasmids; iii) quantifying the expression level of the reporter gene corresponding to each expression plasmid of the library of expression plasmids; iv) dividing the total cell population into a first population and a second population based on the protein expression level of iii); and v) determining the enrichment score of the polynucleotide fragment sequence.
[0553] The enrichment score can be calculated using any suitable statistical analysis method known in the art or described herein. Exemplary methods for calculating the enrichment score include, but are not limited to, the Z-test, odds ratio, T-test, or Fisher's exact test. Wang, Z. et al. describe exemplary methods for calculating the enrichment score (Cell 119, 831 (2004)), which are incorporated herein by reference in their entirety for each method.
[0554] In some embodiments, the enrichment score of each overlapping polynucleotide fragment sequence within the two populations is calculated using the Z-test, resulting in a Z-score. In some embodiments, the enrichment score (e.g., Z-score) of each overlapping polynucleotide fragment sequence is calculated using the Z-test. In some embodiments, for calculating
[0555] where f 1 is the frequency of a certain sequence in population 1, f 2 is the frequency of the same sequence in population 2, N 1 is the size of population 1, and N 2 is the size of population 2.
[0556] In some embodiments, the overlapping polynucleotide fragment sequence is a hexamer (length of 6 nucleic acid residues or 6-mer). In some embodiments, f 1 is the frequency of each hexamer sequence in population 1, f 2 is the frequency of the same hexamer sequence in population 2, N 1 is the size of population 1 (i.e., the total number of hexamers in population 1), and N 2 is the size of population 2 (i.e., the total number of hexamers in population 2).
[0557] In some embodiments, the overlapping polynucleotide fragment sequence is a pentamer (length of 5 nucleic acid residues or 5-mer). In some embodiments, f 1 is the frequency of each pentamer sequence in population 1, f 2 is the frequency of the same pentamer sequence in population 2, N 1is the size of population 1 (i.e., the total number of pentamers in population 1), and N 2 is the size of population 2 (i.e., the total number of pentamers in population 2).
[0558] In some embodiments, the overlapping polynucleotide fragment sequences are pentamers (length 5 nucleic acid residues or 5-mer). In some embodiments, f 1 is the frequency of each pentamer sequence in population 1, f 2 is the frequency of the same pentamer sequence in population 2, N 1 is the size of population 1, and N 2 is the size of population 2.
[0559] In some embodiments, population 1 is the highly expressed population described herein. In some embodiments, population 1 is the moderately expressed population described herein. In some embodiments, population 1 is the lowly expressed population described herein. In some embodiments, population 1 is the non-expressed population described herein.
[0560] In some embodiments, population 2 is the highly expressed population described herein. In some embodiments, population 2 is the moderately expressed population described herein. In some embodiments, population 2 is the lowly expressed population described herein. In some embodiments, population 2 is the non-expressed population described herein.
[0561] In some embodiments, population 1 is the highly expressed population described herein, and population 2 is the moderately expressed population described herein. In some embodiments, population 1 is the highly expressed population described herein, and population 2 is the lowly expressed population described herein. In some embodiments, population 1 is the highly expressed population described herein, and population 2 is the non-expressed population described herein.
[0562] In some embodiments, population 1 is the moderately expressed population described herein, and population 2 is the lowly expressed population described herein. In some embodiments, population 1 is the moderately expressed population described herein, and population 2 is the non-expressed population described herein. In some embodiments, population 1 is the moderately expressed population described herein, and population 2 is the highly expressed population described herein.
[0563] In some embodiments, population 1 is the lowly expressed population described herein, and population 2 is the non-expressed population described herein. In some embodiments, population 1 is the lowly expressed population described herein, and population 2 is the moderately expressed population described herein. In some embodiments, population 1 is the lowly expressed population described herein, and population 2 is the highly expressed population described herein.
[0564] In some embodiments, population 1 is the non-expression population described herein, and population 2 is the low-expression population described herein. In some embodiments, population 1 is the non-expression population described herein, and population 2 is the medium-expression population described herein. In some embodiments, population 1 is the non-expression population described herein, and population 2 is the high-expression population described herein.
[0565] Exemplary overlapping polynucleotide fragment sequences (e.g., pentamers) and enrichment scores (e.g., Z-scores) are shown in Table 1.
[0566] As will be understood by those skilled in the art, the rank ordering of the overlapping polynucleotide fragments described herein provides a method by which functional synthetic IRES-like sequences can be designed or generated. Without wishing to be bound by theory, it is believed that including a high frequency of the top-ranked polynucleotide fragment motifs in a synthetic IRES-like sequence will result in a higher probability of IRES-like function. Top-ranked polynucleotide fragment motifs within an IRES-like sequence include, but are not limited to, SEQ ID NO: 1, 2, 3, 4, 5, 6, 8, 9, 10, 11, 12, 13, 14, 15, 17, 22, 23, 25, 58, 62.
[0567] Conversely, without wishing to be bound by theory, it is believed that including a high frequency of the bottom-ranked polynucleotide fragment motifs in a synthetic IRES-like sequence will result in a lower probability of IRES-like function.
[0568] 8.2.2.1.d. Determine the numerical score of a polynucleotide sequence and identify IRES-like sequences
[0569] In some embodiments, the numerical score of a polynucleotide query sequence is determined by summing the enrichment scores (e.g., Z-scores) of each overlapping polynucleotide fragment sequence. In some embodiments, the numerical score of a polynucleotide query sequence is determined by averaging the enrichment scores (e.g., Z-scores) of each overlapping polynucleotide fragment sequence.
[0570] In some embodiments, the numerical score is normally distributed according to the following system of equations: z = (X 数值得分 - μ) / σ = 2.06; and thus X = z*σ + μ. In some embodiments, a polynucleotide query sequence having greater than or equal to z standard deviations is identified as an IRES-like sequence.
[0571] In some embodiments, a polynucleotide query sequence having less than or equal to X standard deviations is identified as a non-IRES sequence.
[0572] In some embodiments, a polynucleotide query sequence having a numerical score greater than or equal to 0 is identified as an IRES-like sequence. In some embodiments, a polynucleotide query sequence having a numerical score greater than or equal to 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 1100, 1200, 1300, 1400, 1500, 1600, 1700, 1800, 1900, 2000, 2100, 2200, 2300, or 2500 is identified as an IRES-like sequence.
[0573] In some embodiments, a polynucleotide query sequence having a numerical score less than 0 is identified as a non-IRES sequence. In some embodiments, a polynucleotide query sequence having a numerical score less than 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 1100, 1200, 1300, 1400, 1500, 1600, 1700, 1800, 1900, 2000, 2100, 2200, 2300, or 2500 is identified as an IRES-like sequence.
[0574] In some embodiments, a polynucleotide query sequence whose ranked numerical score is in the top 5% of all polynucleotide query sequences of equal length is considered an IRES-like sequence. In some embodiments, a polynucleotide query sequence whose ranked numerical score is in the top 10% of all polynucleotide query sequences of equal length is considered an IRES-like sequence. In some embodiments, a polynucleotide query sequence whose ranked numerical score is in the top 15% of all polynucleotide query sequences of equal length is considered an IRES-like sequence. In some embodiments, a polynucleotide query sequence whose ranked numerical score is in the top 20% of all polynucleotide query sequences of equal length is considered an IRES-like sequence. In some embodiments, a polynucleotide query sequence whose ranked numerical score is in the top 25% of all polynucleotide query sequences of equal length is considered an IRES-like sequence.
[0575] In some embodiments, polynucleotide query sequences with rank scores in the bottom 5% of all polynucleotide query sequences of equal length are considered non-IRES sequences. In some embodiments, polynucleotide query sequences with rank scores in the bottom 10% of all polynucleotide query sequences of equal length are considered non-IRES sequences. In some embodiments, polynucleotide query sequences with rank scores in the bottom 15% of all polynucleotide query sequences of equal length are considered non-IRES sequences. In some embodiments, polynucleotide query sequences with rank scores in the bottom 20% of all polynucleotide query sequences of equal length are considered non-IRES sequences. In some embodiments, polynucleotide query sequences with rank scores in the bottom 25% of all polynucleotide query sequences of equal length are considered non-IRES sequences.
[0576] In some embodiments, polynucleotide query sequences with rank scores in the top 10% of all polynucleotide query sequences of equal length are considered IRES-like sequences. In some embodiments, polynucleotide query sequences with rank scores in the top 5% of all polynucleotide query sequences of equal length are considered IRES-like sequences. In some embodiments, polynucleotide query sequences with rank scores in the top 2% of all polynucleotide query sequences of equal length are considered IRES-like sequences. Exemplary IRES-like sequences and rank scores are shown in Tables 2-96, 110, and 112. Accordingly, exemplary IRES-like sequences of the present disclosure include, but are not limited to, those listed in Tables 2-96, 110, and 112. In some embodiments, an IRES-like sequence comprises a nucleic acid sequence that is at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 100% identical to the nucleic acid sequence or a functional portion thereof of the IRES-like sequences of Tables 2-96, 110, and 112.
[0577] 8.2.2.2. Methods for Optimizing IRES-Like Sequences
[0578] A genetic algorithm is a search algorithm that uses random perturbations of a list of random subsets to generate new motifs (see Schmitt, Lothar M (2001), Theory of Genetic Algorithms, Theoretical Computer Science (259), pp. 1-61). Genetic algorithms represent a branch of the research area known as evolutionary computation, as they mimic the biological processes of reproduction and natural selection to find the "fittest" solutions (Goldberg, D.E. (1989), Genetic Algorithms in Search, Optimization, and Machine Learning. Reading: Addison-Wesley). A genetic algorithm is an iterative optimization procedure that repeatedly applies operators (such as selection, crossover, and mutation) to a set of solutions until some convergence criterion is met. Each iterative step of obtaining a new population is called a generation.
[0579] In some embodiments, a genetic algorithm is used to generate a top sequence library ( Figure 3 (130) and Figure 4 (230)).
[0580] In some embodiments, the genetic algorithm program includes the following steps as Figure 3 shown: a) randomly generate an initial population of a defined length as the current parental population, where the size of the initial population = N(100); b) evaluate the objective function (110) of each individual sequence in the initial population using an evaluation function (also called a fitness function; sequences with higher scores are fitter) (e.g., Z-score); c) select the top sequences to form a top sequence library (size = D) (130); d) select some sequences in the top sequence library using a selection operator (the higher the fitness of a sequence, the more likely it is to be selected) (120); and e) generate an offspring population through a crossover operator (121) and / or a mutation operator (122); and repeat steps b)-e) in a loop.
[0581] In some embodiments, the genetic algorithm program includes as Figure 4The following steps are shown: a) randomly generate an initial population with a defined length as the current parental population, where the size of the initial population = N(200); b) evaluate the objective function (210) of each individual sequence in the initial population using an evaluation function (also known as a fitness function; a sequence with a higher score is fitter) (e.g., Z-score); c) select the top sequences to form a top sequence library (size = D) (230); d) select some sequences from the top sequence library using a selection operator (the higher the fitness of a sequence, the more likely it is to be selected) (220); and e) generate an offspring population through a crossover operator (221) and / or a mutation operator (222); and repeat steps b)-e) in a loop.
[0582] In some embodiments, the initial population has a defined length of 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99 or 100 nucleic acid residues. In some embodiments, the initial population has a defined length between 5 and 12 nucleic acid residues. In some embodiments, the initial population has a defined length between 12 and 20 nucleic acid residues. In some embodiments, the initial population has a defined length between 20 and 100 nucleic acid residues. In some embodiments, the initial population has a defined length between 100 and 200 nucleic acid residues. In some embodiments, the initial population has a defined length between 200 and 300 nucleic acid residues. In some embodiments, the initial population has a defined length between 300 and 400 nucleic acid residues. In some embodiments, the initial population has a defined length between 400 and 500 nucleic acid residues. In some embodiments, the initial population has a defined length between 500 and 600 nucleic acid residues. In some embodiments, the initial population has a defined length between 600 and 700 nucleic acid residues. In some embodiments, the initial population has a defined length between 500 and 600 nucleic acid residues. In some embodiments, the initial population has a defined length between 600 and 700 nucleic acid residues. In some embodiments, the initial population has a defined length between 700 and 800 nucleic acid residues. In some embodiments, the initial population has a defined length between 800 and 900 nucleic acid residues. In some embodiments, the initial population has a defined length between 900 and 1000 nucleic acid residues.
[0583] In some embodiments, the iterative genetic algorithm is iterated until the top sequence library is stable and does not change over many generations. In some embodiments, the iterative genetic algorithm is iterated until the top sequence library is stable and does not change within 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95 or 100 generations.
[0584] In some embodiments, the iterative genetic algorithm is run until the top sequence library stabilizes and does not change within 20 generations.
[0585] In some embodiments, the selection, hybridization, and mutation processes continue until the number of offspring is the same as the initial population, so that the offspring consist entirely of new offspring and the parents are completely replaced.
[0586] In some embodiments, the selection, hybridization, and mutation processes allow the high-scoring sequences of the parents to survive in the offspring.
[0587] In some embodiments, a genetic algorithm is used to select the top-ranked sequences.
[0588] In some embodiments, a genetic algorithm is used to select the bottom-ranked sequences.
[0589] In some embodiments, hybridization and mutation are applied to the sequence library of the current generation to produce the next generation. In some embodiments, random addition or deletion is applied to the sequence library of the current generation to produce the next generation. The hybridization step combines the features from a pair of subsets to form a new subset. In a preferred embodiment, hybridization is the same operator as crossover.
[0590] In a preferred embodiment, a t-test is used to determine the probability that a subset is better than the current best subset by at least one small user-specified threshold for a competitive search to be appropriate.
[0591] In some embodiments, hybridization and mutation are applied to the sequence library of the current generation to produce the next generation. In some embodiments, random addition or deletion is applied to the sequence library of the current generation to produce the next generation. The hybridization step combines the features from a pair of subsets to form a new subset. In a preferred embodiment, hybridization is the same operator as crossover.
[0592] In a preferred embodiment, a t-test is used to determine the probability that a subset is better than the current best subset by at least one small user-specified threshold for a competitive search to be appropriate.
[0593] 8.3. Therapeutic Products
[0594] In some embodiments, the polynucleotides of the present disclosure encode therapeutic products.
[0595] In some embodiments, the therapeutic product is a polypeptide, protein, enzyme, antibody, or a combination thereof. In some embodiments, the therapeutic product comprises one or more polypeptides, proteins, enzymes, antibodies, or a combination thereof.
[0596] In some embodiments, the therapeutic product is a protein or an enzyme. In some embodiments, the protein or enzyme is associated with a genetic disease (e.g., a disease in which a genetic alteration (e.g., a mutation) and / or abnormal protein regulation plays a role in the initiation, development, and / or manifestation of the disease).
[0597] In some embodiments, the therapeutic product is a polypeptide or a protein. In some embodiments, the polypeptide or protein is similar to a weakened or inactivated form of a disease-causing agent (e.g., an infectious agent such as a pathogen), which can be selected from microorganisms (such as bacteria, viruses, fungi, parasites) or one or more components of such microorganisms (such as toxins, proteins (e.g., surface proteins) and / or cell walls). In some embodiments, the therapeutic product is an antigen or an agent that can stimulate the body's immune system to recognize the antigen or agent, generate antibodies against the antigen or agent, destroy the antigen or agent, and / or develop immunological memory of the antigen or agent. In some embodiments, the therapeutic product is an antigen or an agent that can induce and / or enhance vaccine-induced memory and / or enable the immune system to respond rapidly and effectively to the antigen or agent upon subsequent encounter with the antigen or agent.
[0598] In some embodiments, the therapeutic product is derived from an infectious agent or a part, component, and / or product thereof (e.g., cell wall, genomic sequence, membrane, capsid, protein, lipid, glycan, toxin). In some embodiments, the infectious agent is selected from viruses, bacteria, fungi, protozoa, and worms. In some embodiments, the infectious agent is selected from viruses and bacteria.
[0599] In some embodiments, the infectious agent is a virus selected from the following: adenovirus; herpes simplex type 1; herpes simplex type 2; encephalitis virus, papillomavirus; varicella-zoster virus; Epstein-Barr virus; human cytomegalovirus; human herpesvirus 8; BK virus; JC virus; smallpox; poliovirus; human bocavirus; parvovirus B19; human astrovirus; norovirus; coxsackievirus; hepatitis A virus; hepatitis B virus; hepatitis C virus; hepatitis D virus; hepatitis E virus; rhinovirus; severe acute respiratory syndrome (SARS) virus; yellow fever virus; dengue virus; West Nile virus; rubella virus; human immunodeficiency virus (HIV); influenza virus; Guanarito virus; Junin virus; Lassa virus; Machupo virus; Sabia virus; Crimean-Congo hemorrhagic fever virus; Ebola virus; Marburg virus; measles virus; mumps virus; parainfluenza virus; respiratory syncytial virus (RSV); human metapneumovirus; Hendra virus; Nipah virus; rabies virus; rotavirus; orbivirus; Coltivirus; Banna virus; human enterovirus; hantavirus; West Nile virus; coronavirus, SARS-related coronavirus (SARS-CoV), SARS-CoV-2 virus (COVID-19-related); Middle East respiratory syndrome coronavirus; Japanese encephalitis virus; vesicular exanthernavirus; and eastern equine encephalitis.
[0600] In some embodiments, the infectious agent is a bacterium selected from the following: tuberculosis (Mycobacterium tuberculosis), Clostridium difficile resistant to clindamycin, Clostridium difficile resistant to fluoroquinolones, methicillin-resistant Staphylococcus aureus (MRSA), multidrug-resistant Enterococcus faecalis, multidrug-resistant Enterococcus faecium, multidrug-resistant Pseudomonas aeruginosa, multidrug-resistant Acinetobacter baumannii, and vancomycin-resistant Staphylococcus aureus (VRSA).
[0601] In some embodiments, the infectious agent is associated with humans, non-human primates, or other animals such as birds, pigs, horses, dogs, cats, rabbits, mice, rats, cows, sheep, goats, and deer.
[0602] In some embodiments, the therapeutic product is an antibody. In some embodiments, the antibodies include, but are not limited to, monoclonal antibodies, polyclonal antibodies, recombinantly produced antibodies, human antibodies, humanized antibodies, chimeric antibodies, synthetic antibodies, tetrameric antibodies comprising two heavy chain and two light chain molecules, antibody light chain monomers, antibody heavy chain monomers, antibody light chain dimers, antibody heavy chains, antibody heavy chain dimers, antibody light chain-heavy chain pairs, intracellular antibodies, heteroconjugate antibodies, monovalent antibodies, antigen-binding fragments of full-length antibodies, and fusion proteins of the foregoing. Antigen-binding fragments include, but are not limited to, single domain antibodies (heavy chain variable domain antibodies (VHH) or nanobodies), Fab, Fab', F(ab') 2, Fd, Fv, scFv (single-chain variable fragment).
[0603] 8.3.1. Method for producing a therapeutic product
[0604] The present disclosure also provides a method for producing a protein in a cell, the method comprising contacting the cell with an RNA polynucleotide (e.g., a circular RNA molecule) described herein, a DNA vector encoding the RNA polynucleotide or suitable for synthesizing the RNA polynucleotide, whereby the protein-coding nucleic acid sequence is translated and a protein is produced in the cell.
[0605] In some embodiments, the method for producing a protein in a cell comprises contacting the cell with an RNA polynucleotide (e.g., circular RNA) sequence described herein, or a vector comprising the RNA polynucleotide (e.g., circular RNA) sequence, under conditions under which the protein-coding nucleic acid sequence is translated and a protein is produced in the cell. Also provided is a protein produced by the disclosed method.
[0606] In some embodiments, the production of the protein is tissue-specific. For example, the protein can be selectively produced in one or more of the following tissues: muscle, liver, kidney, brain, lung, skin, pancreas, blood, or heart.
[0607] In some embodiments, the protein is expressed recursively in the cell.
[0608] In some embodiments, the protein is produced in the cell for at least about 10%, at least about 20%, or at least about 30% longer compared to the case where the protein-coding nucleic acid sequence is provided to the cell using a viral vector encoding a linear RNA or provided to the cell as a linear RNA.
[0609] In some embodiments, the level of the protein produced in the cell is at least about 10%, at least about 20%, or at least about 30% higher compared to the case where the protein-coding nucleic acid sequence is provided to the cell using a viral vector or provided to the cell as a linear RNA.
[0610] In some embodiments, the RNA polynucleotides (e.g., circular RNAs) described herein (including their components, such as IRES-like sequences) promote cap-independent translational activity from circRNAs. In some human diseases, classical translation via a cap-independent mechanism may be reduced. Thus, using circRNAs to express proteins can be particularly helpful for treating such diseases. In some embodiments, the cap-independent translational activity from circRNAs is promoted using the circRNAs described herein under conditions where cap-dependent translation in the cell is reduced or shut off.
[0611] As discussed above, when an IRES is in-frame with a protein-coding nucleic acid sequence and the protein-coding sequence lacks a stop codon, translation of the protein-coding nucleic acid sequence can occur in an infinite loop (i.e., recursively). Thus, in some embodiments, a method of producing a protein in a cell produces tandem proteins.
[0612] Any of the prokaryotic or eukaryotic host cells described herein can be contacted with a recombinant circRNA molecule or a vector comprising a circRNA molecule. The host cell can be a mammalian cell, such as a human cell. In some embodiments, the cell is in vivo. In some embodiments, the cell is in vitro. In some embodiments, the cell is ex vivo. In some embodiments, the cell is in a mammal such as a human.
[0613] In some embodiments, regardless of the cell type selected, 5'-cap-dependent translation is impaired (e.g., reduced, decreased, inhibited, or completely absent) in the cell. In some embodiments, there is substantially no 5'-cap-dependent translation in the cell.
[0614] A recombinant circular RNA molecule, a DNA molecule encoding them, or a vector comprising them can be introduced into a cell by any method (including, for example, by transfection, transformation, or transduction). The terms "transfection," "transformation," and "transduction" are used interchangeably herein and refer to the introduction of one or more exogenous polynucleotides into a host cell by using physical or chemical methods. Many transfection techniques are known in the art and include, for example, calcium phosphate DNA co-precipitation (see, e.g., Murray E.J. (ed.), Methods in Molecular Biology, Vol. 7, Gene Transfer and Expression Protocols, Humana Press (1991)); DEAE-dextran; electroporation; cationic lipid-mediated transfection; tungsten particle-facilitated particle bombardment (Johnston, Nature, 346:776-777 (1990)); strontium phosphate DNA co-precipitation (Brash et al., Mol. Cell. Biol., 7:2031-2034 (1987)); and magnetic nanoparticle-based gene delivery (Dobson, J., Gene Ther, 13(4):283-7 (2006)).
[0615] 8.3.2. Post-Translational Modification of Therapeutic Products
[0616] In some embodiments, the therapeutic product (e.g., a protein) is post-translationally modified. In some embodiments, the post-translational modification is specific to the cell type, tissue type, or organ in which the therapeutic product is produced or delivered.
[0617] In some embodiments, the post-translational modification is acetylation, SUMOylation, glycosylation, tyrosine sulfation, phosphorylation, ADP-ribosylation, isoprenylation, myristoylation, palmitoylation, ubiquitination, sentrinization, and / or ubiquitin-like protein modification.
[0618] In some embodiments, the therapeutic product is post-translationally modified in vitro, ex vivo, or in vivo.
[0619] In some embodiments, the therapeutic product is post-translationally modified after expression in a cell (e.g., a primary cell, a cell line derived from a primary cell, an immortalized cell). In some embodiments, the cell is an immortalized cell (e.g., a human immortalized cell).
[0620] In some embodiments, the therapeutic product is post-translationally modified in a living organism (e.g., a bacterium, an animal).
[0621] In some embodiments, the therapeutic product is post-translationally modified in an animal (e.g., an animal as described herein). In some embodiments, the therapeutic product is post-translationally modified in a human.
[0622] In some embodiments, post-translational modifications can be detected by techniques known in the art, including by gel electrophoresis, Western blotting, Eastern blotting, immunoprecipitation, mass spectrometry, chromatography, and flow cytometry. Analysis of modified proteins is typically performed by electrophoresis and autoradiography, with specificity enhanced by immunoprecipitation of the protein of interest prior to electrophoresis.
[0623] In some embodiments, detection of post-translational modification can include in vivo labeling of the cellular substrate pool with a radioactive substrate or substrate precursor molecule, which results in incorporation of a radiolabeled moiety into the therapeutic product, including but not limited to phosphate, fatty acyl (e.g., myristoyl or palmitoyl), sentrin, methyl, acetyl, hydroxyl, iodine, flavin, ubiquitin, or ADP-ribose. In some embodiments, detection of post-translational modification can include in vitro enzymatic incorporation of a labeled moiety into the therapeutic product to estimate the state of modification in vivo. In some embodiments, the labeled moiety includes but is not limited to a radioisotope, a luminescent enzyme (e.g., horseradish peroxidase, HRP), or a fluorescent label.
[0624] In some embodiments, post-translational modification can be detected by analyzing a change in the electrophoretic mobility of the modified therapeutic product compared to the unmodified therapeutic product.
[0625] In some embodiments, post-translational modifications can be detected by thin layer chromatography of radiolabeled fatty acids extracted from a therapeutic product.
[0626] In some embodiments, post-translational modifications can be detected by: partitioning a therapeutic product into a detergent-rich layer or a detergent layer by phase separation, and the effect of enzymatic treatment of the therapeutic product on partitioning between an aqueous environment and a detergent-rich environment.
[0627] 8.4. RNA polynucleotides
[0628] In some embodiments, the present disclosure provides RNA polynucleotides that include an engineered translation initiation element (TI) that contains an IRES-like polynucleotide sequence.
[0629] In some embodiments, the present disclosure provides RNA polynucleotides that include a construct of Formula I: 5’-(A1) 0-1 -(L) n -TI-(L) n -Z1-(L) n -(B1) 0-1 -3’(I).
[0630] In some embodiments, in Formula I, TI is an engineered translation initiation element that contains an IRES-like polynucleotide sequence; Z1 is an expression sequence encoding a therapeutic product; each L is independently a linker sequence; and n is a positive integer (e.g., an integer selected from 0 to 2).
[0631] In some embodiments, in Formula I: A1 and B1 are each independently a sequence capable of cyclizing the RNA polynucleotide. In some embodiments, in Formula I: A1 and B1 are a pair of homologous sequences capable of self-cleaving and cyclizing the RNA polynucleotide.
[0632] In some embodiments, in Formula I: A1 and B1 each independently include nucleotide derivatives capable of joining the 5’ and 3’ ends by a 3' to 5' phosphodiester linkage to cyclize the RNA polynucleotide.
[0633] In some embodiments, in Formula I: A1 and B1 each independently include nucleotide triphosphate derivatives capable of joining the 5’ and 3’ ends by a 3' to 5' phosphodiester linkage to cyclize the RNA polynucleotide.
[0634] In some embodiments, the present disclosure provides RNA polynucleotides that include a construct of Formula II: 5’-(A1) 0-1 -(L) n -Z1 B -(L)n -TI-(L) n -Z1 A- (L) n -(B1) 0-1 -3’(II).
[0635] In some embodiments, in Formula II, TI is an engineered translation initiation element comprising an IRES-like polynucleotide sequence; Z1 A is the first part of the expression sequence encoding the therapeutic product; Z1 B is the second part of the expression sequence encoding the therapeutic product; each L is independently a linker sequence; and n is a positive integer (e.g., an integer selected from 0 to 2).
[0636] In some embodiments, in Formula II: A1 and B1 are each independently a sequence capable of cyclizing the RNA polynucleotide. In some embodiments, in Formula I: A1 and B1 are a pair of homologous sequences capable of self-cleaving and cyclizing an RNA polynucleotide.
[0637] In some embodiments, in Formula II: A1 and B1 each independently comprise nucleotide derivatives capable of ligating the 5'-end and the 3'-end by a 3' to 5' phosphodiester linkage to cyclize the RNA polynucleotide.
[0638] In some embodiments, in Formula II: A1 and B1 each independently comprise nucleotide triphosphate derivatives capable of ligating the 5'-end and the 3'-end by a 3' to 5' phosphodiester linkage to cyclize the RNA polynucleotide.
[0639] In some embodiments, provided herein is an RNA polynucleotide comprising a construct of Formula III: 5'-(3' intron fragment)-(E2)-(L) n -TI-(L) n -Z1-(L) n -(E1)-(5' intron fragment)-3'(III).
[0640] In some embodiments, in Formula III, TI is an engineered translation initiation element comprising an IRES-like polynucleotide sequence; Z1 is the expression sequence encoding the therapeutic product; each L is independently a linker sequence; and n is a positive integer (e.g., an integer selected from 0 to 2).
[0641] In some embodiments, in Formula III, the 5' intron fragment and the 3' intron fragment are each a fragment of a type II intron, wherein the 5' intron fragment is on the 5'-side of the 3' intron fragment in the type II intron.
[0642] In some embodiments, in Formula III, E1 is a 5'-adjacent exon fragment of a Group II intron, and its length is ≥0 nucleotides.
[0643] In some embodiments, in Formula III, E2 is a 3'-adjacent exon fragment of a Group II intron, and its length is ≥0 nucleotides.
[0644] In some embodiments, in Formula III, E1 is a 5'-adjacent exon fragment of a Group II intron, and its length is ≥0 nucleotides; E2 is a 3'-adjacent exon fragment of a Group II intron, and its length is ≥0 nucleotides.
[0645] In some embodiments, provided herein is an RNA polynucleotide comprising a construct of Formula IV: 5'-(3' intron fragment)-(E2)-(L) n -Z1 B -(L) n -TI-(L) n -Z1 A- -(L) n -(E1)-(5' intron fragment)-3' (IV).
[0646] In some embodiments, in Formula IV, TI is an engineered translation initiation element comprising an IRES-like polynucleotide sequence; Z1 A is the first part of an expression sequence encoding a therapeutic product; Z1 B is the second part of an expression sequence encoding a therapeutic product; each L is independently a linker sequence; and n is a positive integer (e.g., an integer selected from 0 to 2).
[0647] In some embodiments, in Formula IV, the 5' intron fragment and the 3' intron fragment are each a fragment of a Group II intron, wherein the 5' intron fragment is on the 5'-side of the 3' intron fragment in the Group II intron.
[0648] In some embodiments, in Formula IV, E1 is a 5'-adjacent exon fragment of a Group II intron, and its length is ≥0 nucleotides.
[0649] In some embodiments, in Formula IV, E2 is a 3'-adjacent exon fragment of a Group II intron, and its length is ≥0 nucleotides.
[0650] In some embodiments, in Formula IV, E1 is a 5'-adjacent exon fragment of a Group II intron, and its length is ≥0 nucleotides; E2 is a 3'-adjacent exon fragment of a Group II intron, and its length is ≥0 nucleotides.
[0651] In some embodiments, provided herein is an RNA polynucleotide comprising a construct of Formula V: 5'-TI-(L)n -Z1-3’(V).
[0652] In some embodiments, in Formula V, TI is an engineered translation initiation element comprising an IRES-like polynucleotide sequence; Z1 is an expression sequence encoding a therapeutic product; each L is independently a linker sequence; and n is a positive integer (e.g., an integer selected from 0 to 2).
[0653] In some embodiments, in Formula III or IV, the RNA polynucleotide further comprises a 5' homologous arm at the 5' end of the 3' intron fragment.
[0654] In some embodiments, in Formula III or IV, the RNA polynucleotide further comprises a 3' homologous arm at the 3' end of the 5' intron fragment.
[0655] In some embodiments, in Formula III or IV, the RNA polynucleotide further comprises a 5' homologous arm at the 5' end of the 3' intron fragment and a 3' homologous arm at the 3' end of the 5' intron fragment.
[0656] In some embodiments, in Formula III or IV, E1 and E2 are each independently 0 to 20 nucleotides in length.
[0657] In some embodiments, in Formula III or IV, the 5' intron fragment and the 3' intron fragment are obtained by splitting a Group II intron into two fragments at an unpaired region, where the unpaired region is preferably selected from the linear region between two adjacent domains of the Group II intron and the loop region of the stem-loop structure of domain 4 of the Group II intron.
[0658] In some embodiments, the 5' homologous arms of the present disclosure include, but are not limited to, those listed in Table 109.
[0659] In some embodiments, the 5' homology arm of the present disclosure comprises a nucleic acid sequence that is at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99% or at least 100% identical to the 5' homology arm sequence of Table 109 or a functional portion thereof. In some embodiments, the 5' homology arm of the present disclosure comprises a combination of more than one (e.g., two, three or four) nucleic acid sequences that are at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99% or at least 100% identical to the 5' homology arm sequence of Table 109 or a functional portion thereof. In some embodiments, the 5' homology arm of the present disclosure comprises a nucleic acid sequence selected from the 5' homology arm sequences of Table 109 or a functional portion thereof. In some embodiments, the 5' homology arm of the present disclosure comprises a combination of more than one (e.g., two, three or four) nucleic acid sequences selected from the 5' homology arm sequences of Table 109 or a functional portion thereof.
[0660] In some embodiments, the 5' homology arm of the present disclosure is a nucleic acid sequence that is at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99% or at least 100% identical to the 5' homology arm sequence of Table 109 or a functional portion thereof. In some embodiments, the 5' homology arm of the present disclosure is a combination of more than one (e.g., two, three or four) nucleic acid sequences that are at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99% or at least 100% identical to the 5' homology arm sequence of Table 109 or a functional portion thereof. In some embodiments, the 5' homology arm of the present disclosure is a nucleic acid sequence selected from the 5' homology arm sequences of Table 109 or a functional portion thereof. In some embodiments, the 5' homology arm of the present disclosure is a combination of more than one (e.g., two, three or four) nucleic acid sequences selected from the 5' homology arm sequences of Table 109 or a functional portion thereof.
[0661] In some embodiments, the 3' homology arms of the present disclosure include, but are not limited to, those listed in Table 109.
[0662] In some embodiments, the 3' homology arm of the present disclosure comprises a nucleic acid sequence that is at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99% or at least 100% identical to the 3' homology arm sequence of Table 109 or a functional portion thereof. In some embodiments, the 3' homology arm of the present disclosure comprises a combination of more than one (e.g., two, three or four) nucleic acid sequences that are at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99% or at least 100% identical to the 3' homology arm sequence of Table 109 or a functional portion thereof. In some embodiments, the 3' homology arm of the present disclosure comprises a nucleic acid sequence selected from the 3' homology arm sequences of Table 109 or a functional portion thereof. In some embodiments, the 3' homology arm of the present disclosure comprises a combination of more than one (e.g., two, three or four) nucleic acid sequences selected from the 3' homology arm sequences of Table 109 or a functional portion thereof.
[0663] In some embodiments, the 3' homology arm of the present disclosure is a nucleic acid sequence that is at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99% or at least 100% identical to the 3' homology arm sequence of Table 109 or a functional portion thereof. In some embodiments, the 3' homology arm of the present disclosure is a combination of more than one (e.g., two, three or four) nucleic acid sequences that are at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99% or at least 100% identical to the 3' homology arm sequence of Table 109 or a functional portion thereof. In some embodiments, the 3' homology arm of the present disclosure is a nucleic acid sequence selected from the 3' homology arm sequences of Table 109 or a functional portion thereof. In some embodiments, the 3' homology arm of the present disclosure is a combination of more than one (e.g., two, three or four) nucleic acid sequences selected from the 3' homology arm sequences of Table 109 or a functional portion thereof.
[0664] In some embodiments, the E1 sequences of the present disclosure include, but are not limited to, those listed in Table 103. In some embodiments, the E1 sequence comprises a nucleic acid sequence that is at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 100% identical to the nucleic acid sequence of the E1 sequence of Table 103 or a functional portion thereof. In some embodiments, the E1 sequence is a nucleic acid sequence that is at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 100% identical to the nucleic acid sequence of the E1 sequence of Table 103 or a functional portion thereof.
[0665] In some embodiments, the E2 sequences of the present disclosure include, but are not limited to, those listed in Table 104. In some embodiments, the E2 sequence comprises a nucleic acid sequence that is at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 100% identical to the nucleic acid sequence of the E2 sequence of Table 104 or a functional portion thereof. In some embodiments, the E2 sequence is a nucleic acid sequence that is at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 100% identical to the nucleic acid sequence of the E2 sequence of Table 104 or a functional portion thereof.
[0666] In some embodiments, the Group II intron sequences of the present disclosure include, but are not limited to, those listed in Table 105. In some embodiments, the Group II intron sequence comprises a nucleic acid sequence that is at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 100% identical to the nucleic acid sequence of the Group II intron sequence of Table 105 or a functional portion thereof. In some embodiments, the Group II intron sequence is a nucleic acid sequence that is at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 100% identical to the nucleic acid sequence of the Group II intron sequence of Table 105 or a functional portion thereof.
[0667] In some embodiments, the 3' intron fragment sequences of the present disclosure include, but are not limited to, those listed in Table 106. In some embodiments, the 3' intron fragment sequence comprises a nucleic acid sequence that is at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99% or at least 100% identical to the nucleic acid sequence or a functional portion thereof of the 3' intron fragment sequence of Table 106. In some embodiments, the 3' intron fragment sequence is a nucleic acid sequence that is at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99% or at least 100% identical to the nucleic acid sequence or a functional portion thereof of the 3' intron fragment sequence of Table 106.
[0668] In some embodiments, the 5' intron fragment sequences of the present disclosure include, but are not limited to, those listed in Table 107. In some embodiments, the 5' intron fragment sequence comprises a nucleic acid sequence that is at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99% or at least 100% identical to the nucleic acid sequence or a functional portion thereof of the 5' intron fragment sequence of Table 107. In some embodiments, the 5' intron fragment sequence is a nucleic acid sequence that is at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99% or at least 100% identical to the nucleic acid sequence or a functional portion thereof of the 5' intron fragment sequence of Table 107.
[0669] In some embodiments, in any of the formulas described herein, n is an integer selected from 0 to 5. In some embodiments, n is an integer selected from 0 to 2. In some embodiments, n is 2. In some embodiments, n is 1. In some embodiments, n is 0.
[0670] In some embodiments, in any of the formulas described herein, TI further comprises a second IRES-like polynucleotide sequence. In some embodiments, the first IRES-like polynucleotide sequence and the second IRES-like polynucleotide sequence are the same. In some embodiments, the first IRES-like polynucleotide sequence and the second IRES-like polynucleotide sequence are different.
[0671] In some embodiments, in any of the formulas described herein, the IRES-like polynucleotide sequence is determined by any of the methods described herein. In some embodiments, the IRES-like polynucleotide sequence comprises a nucleic acid sequence that is about or at least about 90%, 95%, 97%, 98%, 99% or 100% identical to the nucleic acid sequence of SEQ ID NO: 1025-14161 or SEQ ID NO: 14412-15341. In embodiments, the IRES-like polynucleotide comprises a nucleic acid sequence that is about or at least about 90%, 95%, 97%, 98%, 99% or 100% identical to the nucleic acid sequence of SEQ ID NO: 1, 2, 3, 4, 5, 6, 8, 9, 10, 11, 12, 13, 14, 15, 17, 22, 23, 25, 58, 62.
[0672] In some embodiments, in any of the formulas described herein, the length of the IRES-like sequence is greater than or equal to 3 residues. In some embodiments, the length of the IRES-like sequence is 3-300 nucleic acid residues. In some embodiments, the length of the IRES-like sequence is 3-200 nucleic acid residues. In some embodiments, the length of the IRES-like sequence is 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99 or 100 nucleic acid residues. In some embodiments, the length of the IRES-like sequence is 5 nucleic acid residues. In some embodiments, the length of the IRES-like sequence is 6 nucleic acid residues. In some embodiments, the length of the IRES-like sequence is 7 nucleic acid residues. In some embodiments, the length of the IRES-like sequence is 8 nucleic acid residues. In some embodiments, the length of the IRES-like sequence is 9 nucleic acid residues. In some embodiments, the length of the IRES-like sequence is 10 nucleic acid residues. In some embodiments, the length of the IRES-like sequence is 11 nucleic acid residues. In some embodiments, the length of the IRES-like sequence is 12 nucleic acid residues.
[0673] In some embodiments, in any of the formulas described herein, the TI (i.e., an engineered translation initiation element comprising an IRES-like polynucleotide sequence) further comprises an IRES sequence isolated from or derived from a native IRES sequence. In some embodiments, the native IRES sequence comprises a nucleic acid sequence that is at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99% or 100% identical to the nucleic acid sequence of any of the native IRES sequences of Tables 98 - 100 or a functional portion thereof.
[0674] 8.4.1. Expression sequence
[0675] In some embodiments, the RNA polynucleotides provided herein comprise an expression sequence (Z1) encoding a therapeutic product. In some embodiments, the expression sequence encodes a reporter protein. In some embodiments, the expression sequence encodes a therapeutic protein. Exemplary expression sequences include, but are not limited to, those listed in Tables 101 and 102.
[0676] In some embodiments, the RNA polynucleotide comprises one expression sequence. In some embodiments, the polynucleotide comprises more than one expression sequence, such as 2, 3, 4 or 5 expression sequences.
[0677] In some embodiments, the RNA polynucleotide encodes a protein composed of subunits encoded by more than one gene. For example, the protein can be a heterodimer, where each chain or subunit of the protein is encoded by a separate gene. It is possible to deliver more than one RNA polynucleotide in a delivery vehicle, and each RNA polynucleotide encodes a separate subunit of the protein. Alternatively, a single RNA polynucleotide can be engineered to encode more than one subunit. In some embodiments, the individual RNA polynucleotide molecules encoding the respective subunits can be administered in separate delivery vehicles.
[0678] 8.4.2. Linker
[0679] In some embodiments, the RNA polynucleotides provided herein comprise a linker sequence (L). In some embodiments, the linker sequence is from 3 to 300 nucleic acid residues in length. In some embodiments, the linker sequence is about 3 - 10, 10 - 20, 20 - 30, 30 - 40, 40 - 50, 50 - 60, 60 - 70, 70 - 80, 80 - 90, 90 - 100, 100 - 125, 125 - 150, 150 - 175, 175 - 200, 200 - 225, 225 - 250, 250 - 275, or 275 - 300 nucleic acid sequences in length. In some embodiments, the linker sequence is about 3N nucleic acid residues in length, where N is an integer selected from 1 - 100. In some embodiments, the linker sequence is about 3N nucleic acid residues in length, where N is an integer selected from 1 - 50. In some embodiments, the linker sequence is about 3N nucleic acid residues in length, where N is an integer selected from 1 - 20. In some embodiments, the linker sequence is about 3N nucleic acid residues in length, where N is an integer selected from 1 - 10. In some embodiments, the linker sequence is about 3N nucleic acid residues in length, where N is 1, 2, 3, 4, or 5. In some embodiments, the linker sequence is about 3N nucleic acid residues in length, where N is 1, 2, or 3. In some embodiments, the linker sequence is about 3 nucleic acid residues in length.
[0680] In some embodiments, the linker sequence comprises the nucleic acid sequence of RCC, where R is guanine or adenine. In some embodiments, the linker comprises the nucleic acid sequence of RCCRCC, where R is guanine or adenine. In some embodiments, the linker comprises the nucleic acid sequence of RCCRCCRCC, where R is guanine or adenine.
[0681] In some embodiments, the linker comprises a nucleic acid sequence encoding: 5'UTR, 3'UTR, poly-A sequence, polyA-C sequence, poly-C sequence, poly-U sequence, poly-G sequence, ribosome binding site, aptamer, riboswitch, ribozyme, small RNA binding site, translation regulatory element (e.g., Kozak sequence), protein binding site (e.g., PTBP1 or HUR), unnatural nucleotide, or non-nucleotide chemical linker sequence.
[0682] In some embodiments, the RNA polynucleotides provided herein comprise a 3' UTR. In some embodiments, the 3' UTR is from human β-globin, human α-globin, Xenopus laevis β-globin, Xenopus laevis α-globin, human prolactin, human GAP-43, human eEF1α1, human tau protein, human TNFα, dengue virus, hantavirus small mRNA, bunyavirus small mRNA, turnip yellow mosaic virus, hepatitis C virus, rubella virus, tobacco mosaic virus, human IL-8, human actin, human GAPDH, human tubulin, hibiscus chlorotic ringspot virus, woodchuck hepatitis virus post-translational regulatory element, sindbis virus, turnip crinkle virus, tobacco etch virus, or Venezuelan equine encephalitis virus.
[0683] In some embodiments, the RNA polynucleotides provided herein comprise a 5' UTR. In some embodiments, the 5' UTR is from human β-globin, Xenopus laevis β-globin, human α-globin, Xenopus laevis α-globin, rubella virus, tobacco mosaic virus, mouse Gtx, dengue virus, heat shock protein 70 kDa protein 1A, tobacco alcohol dehydrogenase, tobacco etch virus, turnip crinkle virus, or adenovirus tripartite leader.
[0684] In some embodiments, the RNA polynucleotides provided herein comprise a polyA region. In some embodiments, the length of the polyA region is at least 30 nucleotides or at least 60 nucleotides.
[0685] In some embodiments, the RNA polynucleotides described herein are circularized by a ligation reaction. In some embodiments, the RNA polynucleotides described herein are circularized in the presence of T4 ligase. In some embodiments, the RNA polynucleotides described herein are circularized in the absence of T4 ligase.
[0686] In some embodiments, the RNA polynucleotides described herein are circularized by a splicing reaction. In some embodiments, the RNA polynucleotides described herein are circularized in the presence of a spliceosome. In some embodiments, the RNA polynucleotides described herein are circularized in the absence of a spliceosome.
[0687] In some embodiments, the RNA polynucleotides described herein are circularized by a self-splicing reaction.
[0688] 8.4.3. Vectors, Linear RNA, Precursor RNA, and Circular RNA
[0689] In some embodiments, the RNA polynucleotides provided herein are single-stranded RNA. In some embodiments, the polynucleotide is linear RNA. In some embodiments, precursor RNA is provided herein. In some embodiments, the provided RNA polynucleotides are encoded by a vector. In some embodiments, the precursor RNA is linear RNA produced by in vitro transcription of a vector provided herein.
[0690] In some embodiments, the RNA polynucleotide is circular RNA or can be used to prepare circular RNA polynucleotides. In some embodiments, circular RNA is provided herein. In some embodiments, the circular RNA is circular RNA produced by a vector provided herein. In some embodiments, the circular RNA is circular RNA produced by circularizing a precursor RNA provided herein.
[0691] Circular RNA
[0692] Circular RNA (also referred to as "circRNA" or "cRNA") is single-stranded RNA that is head-to-tail linked. circRNAs have been recognized as a class of non-coding RNAs that are ubiquitous in eukaryotic cells. Typically generated by back-splicing, circRNAs are found to be very stable.
[0693] Circular RNA can be generated by any suitable method known in the art or disclosed herein. Exemplary methods for generating circular RNA are described in WO 2021263124, Wesselhoeft, R.A. et al., Nat. Commun. 9, 2629 (2018), and Chinese Patent Application No. 202110594352.4, each of which is incorporated herein by reference in its entirety.
[0694] Circular RNA can be generated by any non-mammalian splicing method. For example, linear RNAs containing various types of introns can be circularized, and the various types of introns include self-splicing group I introns, self-splicing group II introns, spliceosomal introns, and tRNA introns. In particular, the advantage of group I and group II introns is that they can be easily used to generate circular RNA in vitro and in vivo because they are capable of self-splicing due to their self-catalytic ribozyme activity.
[0695] Alternatively, circular RNAs can be generated in vitro from linear RNAs by chemical or enzymatic ligation of the 5' and 3' ends of the RNA. Chemical ligation can be performed, for example, using cyanogen bromide (BrCN) or ethyl-3-(3-dimethylaminopropyl)carbodiimide (EDC) to activate the nucleotide monophosphate groups to allow formation of phosphodiester bonds (Sokolova, FEBS Lett, 232:153-155 (1988); Dolinnaya et al., Nucleic Acids Res., 19:3067-3072 (1991); Fedorova, Nucleosides Nucleotides Nucleic Acids, 15:1137-1147 (1996)). Alternatively, enzymatic ligation can be used to circularize RNAs. Exemplary ligases that can be used include T4 DNA ligase (T4 Dnl), T4 RNA ligase 1 (T4Rnl 1), and T4 RNA ligase 2 (T4 Rnl2).
[0696] In some embodiments, splint ligation can be used to generate circular RNAs. Splint ligation involves using an oligonucleotide splint that hybridizes to the two ends of a linear RNA to join the ends of the linear RNA together. Hybridization of the splint (which can be a deoxyribooligonucleotide or a ribooligonucleotide) orients the 5'-phosphate and 3'-OH of the RNA ends for ligation. As described above, subsequent ligation can be performed using chemical or enzymatic techniques. For example, enzymatic ligation can be performed with T4 DNA ligase (which requires a DNA splint), T4 RNA ligase 1 (which requires an RNA splint), or T4 RNA ligase 2 (DNA or RNA splints). Chemical ligation (such as with BrCN or EDC) can be more effective than enzymatic ligation in some cases if the structure of the hybridized splint-RNA complex interferes with enzyme activity (see, e.g., Dolinnaya et al Nucleic Acids Res, 2 / (23):5403-5407 (1993); Petkovic et al., Nucleic Acids Res, 43(4):2454-2465 (2015)).
[0697] Although circular RNAs are generally more stable than their linear counterparts mainly due to the absence of free ends required for exonucleolytic degradation, the circRNAs described herein can be additionally modified to further enhance stability. There are also other types of modifications that can improve the circularization efficiency, purification of circRNAs, and / or protein expression from circRNAs. For example, recombinant circRNAs can be engineered to include "homology arms" (i.e., 9-19 nucleotide lengths placed at the 5' and 3' ends of the precursor RNA with the aim of bringing the 5' and 3' splice sites closer to each other), spacer sequences (i.e., linker sequences), and / or phosphorothioate (PS) caps (Wesselhoeft et al., Nat. Commun., 9:2629 (2018)). Recombinant circRNAs can also be engineered to include 2'-O-methylfluor- or -O-methoxyethyl conjugates, phosphorothioate backbones, or 2',4'-cyclic 2'-O-ethyl modifications to increase stability (Holdt et al., Front Physiol., 9:1262 (2018); Kratzfeldt et al., Nature, 438(7068):685-9 (2005); and Crooke et al., Cell Metab., 27(4):714-739 (2018)). Recombinant circRNA molecules can also contain one or more modifications that reduce the innate immunogenicity of the circRNA molecule in the host, such as at least one N6-methyladenosine (m 6 A).
[0698] In some embodiments, the recombinant circular RNA molecule is encoded by a nucleic acid comprising at least two introns and at least one exon. In some embodiments, the DNA sequence encoding the circular RNA molecule comprises a sequence encoding at least two introns and at least one exon. As used herein, the term "exon" refers to a nucleic acid sequence present in a gene that is represented in the mature form of the RNA molecule after the introns are excised during transcription. Exons can be translated into proteins (e.g., in the case of messenger RNA (mRNA)). As used herein, the term "intron" refers to a nucleic acid sequence present in a given gene that is removed by RNA splicing during the maturation of the final RNA product. Introns are typically located between exons. During transcription, introns are removed from the precursor messenger RNA (pre-mRNA), and exons are joined by RNA splicing. In some embodiments, the recombinant circular RNA molecule comprises a nucleic acid sequence containing one or more exons and one or more introns.
[0699] Thus, circular RNAs can be generated by splicing endogenous or exogenous introns, as described in WO 2017 / 222911, which is hereby incorporated by reference in its entirety. As used herein, the term "endogenous intron" means an intron sequence that is native to the host cell in which the circRNA is produced. For example, when a circRNA is expressed in a human cell, a human intron is an endogenous intron. An "exogenous intron" is an intron that is heterologous to the host cell in which the circRNA is produced. For example, when a circRNA is expressed in a human cell, a bacterial intron is an exogenous intron. A large number of intron sequences from various organisms and viruses are known and include sequences derived from genes encoding proteins, ribosomal RNAs (rRNAs), or transfer RNAs (tRNAs). Representative intron sequences are available in various databases, including the Group I Intron Sequence and Structure Database (ma.whu.edu.cn / gissd / ), the Bacterial Group II Intron Database (webapps2.ucalgary.ca / ~groupii / index.html), the Mobile Group II Intron Database (fp.ucalgary.ca / group2introns), the Yeast Intron Database (emblS16heidelberg.de / Extemallnfo / seraphin / yidb.html), the Ares Lab Yeast Intron Database (compbio.soe.ucsc.edu / yeast_introns.html), the U12 Intron Database (genome.crg.es / cgibin / ul2db / ul2db.cgi), and the Exon-Intron Database (bpg.utoledo.edu / ~afedorov / lab / eid.html).
[0700] In some embodiments, the RNA polynucleotide (e.g., circular RNA) can have any length or size. In some embodiments, the length of the RNA polynucleotide is between 300 and 10,000, 400 and 9,000, 500 and 8,000, 600 and 7,000, 700 and 6,000, 800 and 5,000, 900 and 5,000, 1,000 and 5,000, 1,100 and 5,000, 1,200 and 5,000, 1,300 and 5,000, 1,400 and 5,000, and / or 1,500 and 5,000 nucleotides.
[0701] In some embodiments, the length of the RNA polynucleotide (e.g., circular RNA) is at least 300 nt, 400 nt, 500 nt, 600 nt, 700 nt, 800 nt, 900 nt, 1000 nt, 1100 nt, 1200 nt, 1300 nt, 1400 nt, 1500 nt, 2000 nt, 2500 nt, 3000 nt, 3500 nt, 4000 nt, 4500 nt, or 5000 nt. In some embodiments, the length of the RNA polynucleotide is no more than 3000 nt, 3500 nt, 4000 nt, 4500 nt, 5000 nt, 6000 nt, 7000 nt, 8000 nt, 9000 nt, or 10000 nt.
[0702] In some embodiments, the length of the RNA polynucleotide (e.g., circular RNA) is about 300 nt, 400 nt, 500 nt, 600 nt, 700 nt, 800 nt, 900 nt, 1000 nt, 1100 nt, 1200 nt, 1300 nt, 1400 nt, 1500 nt, 2000 nt, 2500 nt, 3000 nt, 3500 nt, 4000 nt, 4500 nt, 5000 nt, 6000 nt, 7000 nt, 8000 nt, 9000 nt, or 10000 nt.
[0703] In some embodiments, the length of the RNA polynucleotide (e.g., circular RNA) is at least 500, 600, 700, 800, 900, 1000, 1500, 2000, 2500, 3000, 3500, 4000, 4500, 5000, 5500, 6000, 6500, 7000, 7500, 8000, 8500, 9000, 9500, or 10000 nt. The RNA polynucleotide (e.g., circular RNA) can be unmodified, partially modified, or fully modified.
[0704] In some embodiments, the circular RNAs provided herein have higher functional stability than mRNAs containing the same expression sequence. In some embodiments, the circular RNAs provided herein have higher functional stability than mRNAs containing the same expression sequence, 5moU modification, optimized UTR, cap, and / or polyA tail.
[0705] In some embodiments, the circular RNA polynucleotides provided herein have a functional half-life of at least 5 hours, 10 hours, 15 hours, 20 hours, 30 hours, 40 hours, 50 hours, 60 hours, 70 hours, or 80 hours. In some embodiments, the circular RNA polynucleotides provided herein have a functional half-life of 5 - 80, 10 - 70, 15 - 60, and / or 20 - 50 hours. In some embodiments, the circular RNA polynucleotides provided herein have a longer (e.g., at least 1.5-fold longer, at least 2-fold longer) functional half-life than an equivalent linear RNA polynucleotide encoding the same protein. In some embodiments, the functional half-life can be evaluated by detecting functional protein synthesis.
[0706] In some embodiments, the circular RNA polynucleotides provided herein have a half-life of at least 5 hours, 10 hours, 15 hours, 20 hours, 30 hours, 40 hours, 50 hours, 60 hours, 70 hours, or 80 hours. In some embodiments, the circular RNA polynucleotides provided herein have a half-life of 5 - 80, 10 - 70, 15 - 60, and / or 20 - 50 hours. In some embodiments, the circular RNA polynucleotides provided herein have a longer (e.g., at least 1.5-fold longer, at least 2-fold longer) half-life than an equivalent linear RNA polynucleotide encoding the same protein.
[0707] In some embodiments, the circular RNAs provided herein can have a higher expression level than an equivalent linear mRNA, e.g., a higher expression level 24 hours after administering the RNA to cells. In some embodiments, the circular RNAs provided herein have a higher expression level than an mRNA comprising the same expression sequence, 5moU modification, optimized UTRs, cap, and / or polyA tail. In some embodiments, the circular RNAs provided herein can have a higher stability than an equivalent linear mRNA. In some embodiments, this can be demonstrated by measuring the presence and density of receptors at a time point 1 week after in vitro or in vivo electroporation. In some embodiments, this can be demonstrated by the presence of RNA measured via qPCR or ISH.
[0708] In some embodiments, the circular RNA polynucleotides provided herein comprise modified RNA nucleotides and / or modified nucleosides. In some embodiments, the modified nucleoside is m 5 C (5-methylcytidine). In some embodiments, the modified nucleoside is m 5 U (5-methyluridine). In some embodiments, the modified nucleoside is m 6 A (N 6 -methyladenosine). In some embodiments, the modified nucleoside is s 2U (2-thiouridine). In some embodiments, the modified nucleoside is Y (pseudouridine). In some embodiments, the modified nucleoside is Um (2'-O-methyluridine). In some embodiments, the modified nucleoside is m!A (1-methyladenosine); m 2 A (2-methyladenosine); Am (2'-0-methyladenosine); ms 2 m 6 A (2-methylthio-N 6 -methyladenosine); i 6 A (N 6 -isopentenyladenosine); ms2i6A (2-methylthio-N 6 -isopentenyladenosine); io 6 A (N 6 -(cis-hydroxyisopentenyl)adenosine); ms 2 io 6 A (2-methylthio-N 6 -(cis-hydroxyisopentenyl)adenosine); g 6 A (N 6 -glycylcarbamoyladenosine); t 6 A (N 6 -threonylcarbamoyladenosine); ms 2 t 6 A (2-methylthio-N 6 -threonylcarbamoyladenosine); m 6 t 6 A (N 6 -methyl-N 6 -threonylcarbamoyladenosine); hn 6 A (N 6 -hydroxy-norvalylcarbamoyladenosine); ms 2 hn 6 A (2-methylthio-N 6 -hydroxy-norvalylcarbamoyladenosine); Ar(p) (2'-0-ribosyladenosine (phosphate)); I (inosine); m 1 I (1-methylinosine); m 1 hn (1,2'-O-dimethylinosine); m 3 C (3-methylcytidine); Cm (2'-0-methylcytidine); s 2 C (2-thiocytidine); ac 4 C (N 4 -acetylcytidine); (5-formylcytidine); m 5 Cm (5,2'-O-dimethylcytidine); ac 4 Cm (N 4 -acetyl-2'-O-methylcytidine); k 2C (lysine); m!G (1-methylguanosine); m 2 G (N 2 -methylguanosine); m 7 G (7-methylguanosine); Gm (2'-0-methylguanosine); m 2 2G (N 2 ,N 2 -dimethylguanosine); m 2 Gm (N 2 ,2'-O-dimethylguanosine); m 2 aGm (N 2 ,N 2 ,2'-O-trimethylguanosine); Gr(p) (2'-0-ribosylguanosine (phosphate)); yW (wybutosine); oayW (oxywybutosine); OHyW (hydroxywybutosine); OHyW* (undermodified hydroxywybutosine); imG (wyosine); mimG (methylwyosine); Q (queuosine); oQ (epoxyqueuosine); galQ (galactosyl-queuosine); manQ (mannosyl-queuosine); preQo (7-cyano-7-deazaguanosine); preQi (7-aminomethyl-7-deazaguanosine); G + (archiguanosine); D (dihydrouridine); m 5 Um (5,2'-0-dimethyluridine); s 4 U (4-thiouridine); m 5 s 2 U (5-methyl-2-thiouridine); s 2 Um (2-thio-2'-0-methyluridine); acp 3 U (3-(3-amino-3-carboxypropyl)uridine); ho 5 U (5-hydroxyuridine); mo 5 U (5-methoxyuridine); cmo 5 U (uridine 5-oxyacetic acid); mcmo 5 U (uridine 5-oxyacetic acid methyl ester); chm 5 U (5-(carboxyhydroxymethyl)uridine); mchm 5 U (5-(carboxyhydroxymethyl)uridine methyl ester); mcm 5 U (5-methoxycarbonylmethyluridine); mcm 5 Um (5-methoxycarbonylmethyl-2'-0-methyluridine); mcm 5 s 2 U (5-methoxycarbonylmethyl-2-thiouridine); nm 5 S 2 U (5-aminomethyl-2-thiouridine); mnm 5 U (5-methylaminomethyluridine); mnm5 s 2 U (5-methylaminomethyl-2-thiouridine); mnm 5 se 2 U (5-methylaminomethyl-2-selenouridine); ncm 5 U (5-carbamoylmethyluridine); ncm 5 Um (5-carbamoylmethyl-2'-O-methyluridine); cmnm 5 U (5-carboxymethylaminomethyluridine); cmnm 5 Um (5-carboxymethylaminomethyl-2'-0-methyluridine); cmnm 5 s 2 U (5-carboxymethylaminomethyl-2-thiouridine); m 6 2 A (N 6 ,N 6 -dimethyladenosine); Im (2'-0-methylinosine); m 4 C (N 4 -methylcytidine); m 4 Cm (N 4 ,2'-0-dimethylcytidine); hm 5 C (5-hydroxymethylcytidine); m 3 U (3-methyluridine); cm 5 U (5-carboxymethyluridine); m 6 Am (N 6 ,2'-O-dimethyladenosine); m 6 2 Am (N 6 ,N 6 ,0-2'-trimethyladenosine); m 2,7 G (N 2 ,7-dimethylguanosine); m 2,2,7 G (N 2 ,N 2 ,7-trimethylguanosine); m 3 Um (3,2'-0-dimethyluridine); m 5 D (5-methyldihydrouridine); f 5 Cm (5-formyl-2'-0-methylcytidine); m'Gm (l,2'-0-dimethylguanosine); m'Am (l,2'-0-dimethyladenosine); rm 5 U (5-tauromethyluridine); τm5s2U (5-tauromethyl-2-thiouridine); imG-14 (4-demethylwyosine); imG2 (isowyosine); or ac 6 A (N 6 -acetyladenosine).
[0709] In some embodiments, the modified nucleosides can include compounds selected from the group consisting of: pyridin-4-one ribonucleoside, 5-aza-uridine, 2-thio-5-aza-uridine, 2-thio-uridine, 4-thio-pseudouridine, 2-thio-pseudouridine, 5-hydroxyuridine, 3-methyluridine, 5-carboxymethyl-uridine, 1-carboxymethyl-pseudouridine, 5-propynyl-uridine, 1-propynyl-pseudouridine, 5-tauromethyluridine, 1-tauromethyl-pseudouridine, 5-tauromethyl-2-thio-uridine, 1-tauromethyl-4-thio-uridine, 5-methyl-uridine, 1-methyl-pseudouridine, 4-thio-1-methyl-pseudouridine, 2-thio-1-methyl-pseudouridine, 1-methyl-1-deaza-pseudouridine, 2-thio-1-methyl-1-deaza-pseudouridine, dihydrouridine, dihydropseudouridine, 2-thio-dihydrouridine, 2-thio-dihydropseudouridine, 2-methoxyuridine, 2-methoxy-4-thiouridine, 4-methoxy-pseudouridine, 4-methoxy-2-thio-pseudouridine, 5-aza-cytidine, pseudoisocytidine, 3-methyl-cytidine, N4-acetylcytidine, 5-formylcytidine, N4-methylcytidine, 5-hydroxymethylcytidine, 1-methyl-pseudoisocytidine, pyrrolo-cytidine, pyrrolo-pseudoisocytidine, 2-thio-cytidine, 2-thio-5-methyl-cytidine, 4-thio-pseudoisocytidine, 4-thio-1-methyl-pseudoisocytidine, 4-thio-1-methyl-1-deaza-pseudoisocytidine, 1-methyl-1-deaza-pseudoisocytidine, zebularine, 5-aza-zebularine, 5-methyl-zebularine, 5-aza-2-thio-zebularine, 2-thio-zebularine, 2-methoxy-cytidine, 2-methoxy-5-methyl-cytidine, 4-methoxy-pseudoisocytidine, 4-methoxy-1-methyl-pseudoisocytidine, 2-aminopurine, 2,6-diaminopurine, 7-deaza-adenine, 7-deaza-8-aza-adenine, 7-deaza-2-aminopurine, 7-deaza-8-aza-2-aminopurine, 7-deaza-2,6-diaminopurine, 7-deaza-8-aza-2,6-diaminopurine, 1-methyladenosine, N6-methyladenosine, N6-isopentenyladenosine, N6-(cis-hydroxyisopentenyl)adenosine, 2-methylthio-N6-(cis-hydroxyisopentenyl)adenosine, N6-glycylcarbamoyladenosine, N6-threonylcarbamoyladenosine, 2-methylthio-N6-threonylcarbamoyladenosine, N6,N6-dimethyladenosine, 7-methyladenine, 2-methylthio-adenine, 2-methoxy-adenine, inosine, 1-methyl-inosine, wyosine, wybutosine, 7-deazaguanosine, 7-deaza-8-azaguanosine, 6-thioguanosine, 6-thio-7-deazaguanosine, 6-thio-7-deaza-8-azaguanosine, 7-methylguanosine, 6-thio-7-methylguanosine, 7-methylinosine, 6-methoxyguanosine, 1-methylguanosine, N2-methylguanosine, N2,N2-dimethylguanosine, 8-oxoguanosine, 7-methyl-8-oxoguanosine, 1-methyl-6-thioguanosine, N2-methyl-6-thioguanosine, and N2,N2-dimethyl-6-thioguanosine. In some embodiments, the modification is independently selected from 5-methylcytosine, pseudouridine, and 1-methylpseudouridine.,
[0710] In some embodiments, the polynucleotide can be codon-optimized. A codon-optimized sequence can be a sequence in which the codons in the polynucleotide encoding a therapeutic product have been replaced to increase the expression, stability, and / or activity of the therapeutic product. Factors affecting codon optimization include, but are not limited to, one or more of the following: (i) codon bias variations between two or more organisms or genes or a synthetically constructed bias table, (ii) variations in the degree of codon bias within an organism, gene, or gene set, (iii) systematic variations in codons (including context), (iv) codon variations according to the decoding tRNA, (v) codon variations according to GC% whether for the entire triplet or at one of its positions, (vi) variations in the degree of similarity to a reference sequence (e.g., a naturally occurring sequence), (vii) variations in codon frequency cutoffs, (viii) structural properties of the mRNA transcribed from the DNA sequence, (ix) prior knowledge of the DNA sequence function on which the design of the codon substitution set is based, and / or (x) systematic variations in the codon sets for each amino acid. In some embodiments, the codon-optimized polynucleotide can minimize ribozyme resistance and / or limit structural interference between the expression sequence and the IRES.
[0711] 8.4.4. Generation of RNA polynucleotides
[0712] The vectors provided herein can be prepared using standard techniques of molecular biology. For example, the various elements of the vectors provided herein can be obtained using recombinant methods, such as by screening cDNA and genomic libraries from cells or by deriving polynucleotides from vectors known to contain them.
[0713] The various elements of the vectors provided herein may also be produced synthetically based on known sequences rather than being cloned. The complete sequences may be assembled from overlapping oligonucleotides prepared by standard methods and assembled into the complete sequence. See, e.g., Edge, Nature (1981) 292:756; Nambair et al., Science (1984) 223:1299; and Jay et al., J. Biol. Chem. (1984) 259:6311.
[0714] Thus, a specific nucleotide sequence may be obtained from a vector containing the desired sequence or synthesized in whole or in part using a variety of oligonucleotide synthesis techniques known in the art (such as site-directed mutagenesis and polymerase chain reaction (PCR) techniques where appropriate). One method of obtaining a nucleotide sequence encoding a desired vector element is to anneal complementary sets of overlapping synthetic oligonucleotides produced in a conventional automated polynucleotide synthesizer, followed by ligation with an appropriate DNA ligase and amplification of the ligated nucleotide sequence by PCR. See, e.g., Jayaraman et al., Proc. Natl. Acad. Sci. USA (1991) 88:4084-4088. Additionally, the following may be used: oligonucleotide-directed synthesis (Jones et al., Nature (1986) 54:75-82), oligonucleotide-directed mutagenesis of pre-existing nucleotide regions (Riechmann et al., Nature (1988) 332:323-327 and Verhoeyen et al., Science (1988) 239:1534-1536), and enzymatic filling of nicked oligonucleotides with T4 DNA polymerase (Queen et al., Proc. Natl. Acad. Sci. USA (1989) 86:10029-10033).
[0715] The precursor RNA provided herein may be generated by incubating the vectors provided herein under conditions that permit transcription of the precursor RNA encoded by the vector. For example, in some embodiments, the precursor RNA is synthesized by incubating a vector provided herein that contains its 5' duplex-forming region and / or an RNA polymerase promoter upstream of the expression sequence with a compatible RNA polymerase under conditions that permit in vitro transcription. In some embodiments, the vector is incubated intracellularly with a bacteriophage RNA polymerase or in the nucleus with a host RNA polymerase P.
[0716] In some embodiments, provided herein is a method for generating precursor RNA by in vitro transcription using a vector provided herein (e.g., a vector provided herein having an RNA polymerase promoter upstream of the 5' homology region) as a template.
[0717] In some embodiments, the resulting precursor RNA can be used to generate circular RNA (e.g., the circular RNA polynucleotides provided herein).
[0718] Accordingly, in some embodiments, provided herein is a method of preparing circular RNA. In some embodiments, the method includes synthesizing a precursor RNA by transcription (e.g., run-off transcription) using a vector provided herein as a template, and incubating the resulting precursor RNA under conditions suitable for cyclization to form circular RNA.
[0719] In some embodiments, a composition comprising circular RNA has been purified. The circular RNA can be purified by any known method commonly used in the art, such as column chromatography, gel filtration chromatography, and size exclusion chromatography. In some embodiments, purification includes one or more of the following steps: phosphatase treatment, HPLC size exclusion purification, and RNase R digestion. In some embodiments, purification sequentially includes the following steps: RNase R digestion, phosphatase treatment, and HPLC size exclusion purification. In some embodiments, purification includes reverse phase HPLC. In some embodiments, the purified composition contains less double-stranded RNA, DNA splints, triphosphorylated RNA, phosphatase proteins, protein ligases, capping enzymes, and / or nicked RNA compared to unpurified RNA.
[0720] 8.5. Formulations and Delivery
[0721] In some embodiments, the polynucleotides or circular RNAs disclosed herein can be formulated using liposomes, liposome complexes, lipid nanoparticles, polymer-based delivery systems, and viral vectors. In some embodiments, the polynucleotides or circular RNAs can be formulated in lipid nanoparticles (such as those described in WO 2012170930, which is incorporated herein by reference in its entirety). In some embodiments, the lipids can be cleavable lipids, such as those described in WO2012170889, which is incorporated herein by reference in its entirety. In some embodiments, the pharmaceutical compositions of the polynucleotides or circular RNAs can include at least one PEGylated lipid described in WO 2012099755, which is incorporated herein by reference. In some embodiments, the lipid nanoparticle formulations can be prepared by the methods described in WO2011127255 or WO 2008103276, each of which is incorporated herein by reference in its entirety. The lipid nanoparticles can be coated or associated with copolymers, such as but not limited to block copolymers, such as the branched polyether-polyamide block copolymers described in WO 2013012476, which is incorporated herein by reference in its entirety. Liposomes, liposome complexes, or lipid nanoparticles can be used to enhance the efficacy of polynucleotide- or circular RNA-directed protein production, as these formulations may be able to increase the cellular transfection of polynucleotides and circular RNAs, increase the in vivo or in vitro half-life of polynucleotides and circular RNAs, and / or permit controlled release.
[0722] In some embodiments, the polynucleotides or circular RNAs disclosed herein encode a protein composed of subunits encoded by more than one gene. For example, the protein can be a heterodimer, where each chain or subunit of the protein is encoded by a separate gene. It is possible to deliver more than one polynucleotide or circular RNA molecule in a delivery vehicle, and each polynucleotide or circular RNA encodes a separate subunit of the protein. Alternatively, a single polynucleotide or circular RNA molecule can be engineered to encode more than one subunit. In some embodiments, the separate polynucleotide or circular RNA molecules encoding the individual subunits can be administered in separate delivery vehicles.
[0723] The present disclosure contemplates differential targeting of target cells and tissues by passive and active targeting means. The phenomenon of passive targeting exploits the natural distribution pattern of the delivery vehicle in the body, without relying on the use of additional excipients or means to enhance the recognition of the delivery vehicle by the target cells. For example, delivery vehicles that are phagocytosed by cells of the reticuloendothelial system may accumulate in the liver or spleen, thus providing a means of passively directing the delivery of the composition to such target cells.
[0724] Alternatively, the present disclosure contemplates active targeting, which involves using targeting moieties that can bind (covalently or non-covalently) to the transfer vehicle to facilitate the localization of such transfer vehicle at certain target cells or tissues. For example, targeting can be mediated by including one or more endogenous targeting moieties in or on the transfer vehicle to facilitate distribution to target cells or tissues. Recognition of the targeting moiety by the target tissue actively promotes the tissue distribution and cellular uptake of the transfer vehicle and / or its contents in the target cells and tissues (e.g., inclusion of an apolipoprotein-E targeting ligand in or on the transfer vehicle promotes recognition and binding by the endogenous low-density lipoprotein receptor expressed by hepatocytes). As provided herein, the compositions can include moieties capable of enhancing the affinity of the composition for target cells. The targeting moieties can be attached to the outer bilayer of the lipid particle during or after formulation. These methods are well known in the art. Additionally, some lipid particle formulations can use fusogenic polymers (such as PEAA), hemagglutinins, other lipopeptides (see U.S. Patent No. 6,417,326, which is incorporated herein by reference), and other features useful for in vivo and / or intracellular delivery. In some other embodiments, the compositions of the present disclosure exhibit improved transfection efficiency and / or exhibit enhanced selectivity for the target cells or tissues of interest. Accordingly, compositions are contemplated that include one or more moieties (such as peptides, aptamers, oligonucleotides, vitamins, or other molecules) capable of enhancing the affinity of the composition and its nucleic acid contents for target cells or tissues. Suitable moieties can optionally be bound or linked to the surface of the transfer vehicle. In some embodiments, the targeting moiety can span the surface of the transfer vehicle or be encapsulated within the transfer vehicle. Suitable moieties are selected based on their physical, chemical, or biological properties (e.g., selective affinity and / or recognition of target cell surface markers or features). Cell-specific target sites and their corresponding targeting ligands can vary widely. Appropriate targeting moieties are selected to take advantage of the unique features of the target cells, thereby allowing the composition to discriminate between target and non-target cells. For example, the compositions of the present disclosure can include surface markers (such as apolipoprotein-B or apolipoprotein-E) that selectively enhance the recognition or affinity for hepatocytes (e.g., through receptor-mediated recognition and binding to such surface markers). For example, it is contemplated that using galactose as a targeting moiety can direct the compositions of the present disclosure to parenchymal hepatocytes, or alternatively, it is contemplated that using sugar residues containing mannose as targeting ligands can direct the compositions of the present disclosure to liver endothelial cells (e.g., sugar residues containing mannose can preferentially bind to the asialoglycoprotein receptor present in hepatocytes).(See Hillery AM, et al., “Drug Delivery and Targeting: For Pharmacists and Pharmaceutical Scientists” (2002) Taylor & Francis, Inc.). Thus, the presentation of such targeting moieties conjugated to a moiety present in a delivery vehicle (e.g., a lipid nanoparticle) facilitates the recognition and uptake of the compositions of the present disclosure in target cells and tissues. Examples of suitable targeting moieties include one or more peptides, proteins, aptamers, vitamins, and oligonucleotides.
[0725] In some embodiments, the polynucleotides or circular RNAs disclosed herein are formulated according to the methods described in U.S. Publication No. 20180153822. In some embodiments, the present disclosure provides a method of encapsulating a polynucleotide or circular RNA in a lipid nanoparticle, the method comprising the steps of forming the lipid into a pre-formed lipid nanoparticle (i.e., formed in the absence of RNA) and then combining the pre-formed lipid nanoparticle with the RNA. In some embodiments, the formulated method produces an RNA formulation with potentially better tolerance, higher potency (peptide or protein expression), and higher efficacy (improvement in biologically relevant endpoints) in vitro and in vivo compared to the same RNA formulation prepared without the step of pre-forming the lipid nanoparticle (e.g., directly combining the lipid with the RNA).
[0726] In some embodiments, the delivery vehicle is formulated and / or targeted as described by Shobaki N et al., Int J Nanomedicine 2018; 13:8395 - 8410. In some embodiments, the delivery vehicle consists of 3 lipid types. In some embodiments, the delivery vehicle consists of 4 lipid types. In some embodiments, the delivery vehicle consists of 5 lipid types. In some embodiments, the delivery vehicle consists of 6 lipid types.
[0727] For certain cationic lipid nanoparticle formulations of RNA, heating of the RNA in a buffer (e.g., citrate buffer) is necessary to achieve high encapsulation of the RNA. In these processes or methods, heating is required prior to the formulation process (i.e., heating the individual components) because heating after formulation (after nanoparticle formation) does not increase the efficiency of encapsulating the RNA in the lipid nanoparticles. In contrast, in some embodiments, the order of heating of the RNA does not affect the RNA encapsulation percentage. In some embodiments, heating of one or more of a solution containing preformed lipid nanoparticles, a solution containing RNA, and a mixed solution containing RNA encapsulated in lipid nanoparticles is not required (i.e., kept at ambient temperature) before or after the formulation process.
[0728] The RNA can be provided in a solution to be mixed with the lipid solution such that the RNA can be encapsulated in the lipid nanoparticles. Suitable RNA solutions can be any aqueous solution containing the RNA to be encapsulated at various concentrations. For example, a suitable RNA solution can contain RNA at a concentration of or greater than about 0.01 mg / ml, 0.05 mg / ml, 0.06 mg / ml, 0.07 mg / ml, 0.08 mg / ml, 0.09 mg / ml, 0.1 mg / ml, 0.15 mg / ml, 0.2 mg / ml, 0.3 mg / ml, 0.4 mg / ml, 0.5 mg / ml, 0.6 mg / ml, 0.7 mg / ml, 0.8 mg / ml, 0.9 mg / ml, or 1.0 mg / ml. In some embodiments, a suitable RNA solution can contain RNA at a concentration range of about 0.01 - 1.0 mg / ml, 0.01 - 0.9 mg / ml, 0.01 - 0.8 mg / ml, 0.01 - 0.7 mg / ml, 0.01 - 0.6 mg / ml, 0.01 - 0.5 mg / ml, 0.01 - 0.4 mg / ml, 0.01 - 0.3 mg / ml, 0.01 - 0.2 mg / ml, 0.01 - 0.1 mg / ml, 0.05 - 1.0 mg / ml, 0.05 - 0.9 mg / ml, 0.05 - 0.8 mg / ml, 0.05 - 0.7 mg / ml, 0.05 - 0.6 mg / ml, 0.05 - 0.5 mg / ml, 0.05 - 0.4 mg / ml, 0.05 - 0.3 mg / ml, 0.05 - 0.2 mg / ml, 0.05 - 0.1 mg / ml, 0.1 - 1.0 mg / ml, 0.2 - 0.9 mg / ml, 0.3 - 0.8 mg / ml, 0.4 - 0.7 mg / ml, or 0.5 - 0.6 mg / ml.
[0729] Typically, a suitable RNA solution may also contain a buffer and / or a salt. Typically, the buffer may include HEPES, ammonium sulfate, Tris, sodium bicarbonate, sodium citrate, sodium acetate, potassium phosphate, or sodium phosphate. In some embodiments, the suitable concentration of the buffer may range from about 0.1 mM to 100 mM, 0.5 mM to 90 mM, 1.0 mM to 80 mM, 2 mM to 70 mM, 3 mM to 60 mM, 4 mM to 50 mM, 5 mM to 40 mM, 6 mM to 30 mM, 7 mM to 20 mM, 8 mM to 15 mM, or 9 to 12 mM.
[0730] Exemplary salts may include sodium chloride, magnesium chloride, and potassium chloride. In some embodiments, the suitable concentration of the salt in the RNA solution may range from about 1 mM to 500 mM, 5 mM to 400 mM, 10 mM to 350 mM, 15 mM to 300 mM, 20 mM to 250 mM, 30 mM to 200 mM, 40 mM to 190 mM, 50 mM to 180 mM, 50 mM to 170 mM, 50 mM to 160 mM, 50 mM to 150 mM, or 50 mM to 100 mM.
[0731] In some embodiments, a suitable RNA solution may have a pH in the range of about 3.5 - 6.5, 3.5 - 6.0, 3.5 - 5.5, 3.5 - 5.0, 3.5 - 4.5, 4.0 - 5.5, 4.0 - 5.0, 4.0 - 4.9, 4.0 - 4.8, 4.0 - 4.7, 4.0 - 4.6, or 4.0 - 4.5.
[0732] Various methods can be used to prepare an RNA solution suitable for the present disclosure. In some embodiments, the RNA can be directly dissolved in the buffer solution described herein. In some embodiments, an RNA solution can be produced by mixing an RNA stock solution with a buffer solution, and then with a lipid solution for encapsulation. In some embodiments, an RNA solution can be produced by mixing an RNA stock solution with a buffer solution and then immediately with a lipid solution for encapsulation.
[0733] According to the present disclosure, the lipid solution contains a lipid mixture suitable for forming a transfer vehicle for encapsulating the RNA. In some embodiments, the suitable lipid solution is ethanol-based. For example, a suitable lipid solution may contain a mixture of the desired lipids dissolved in pure ethanol (i.e., 100% ethanol). In some embodiments, the suitable lipid solution is isopropanol-based. In some embodiments, the suitable lipid solution is dimethyl sulfoxide-based. In some embodiments, the suitable lipid solution is a mixture of suitable solvents including, but not limited to, ethanol, isopropanol, and dimethyl sulfoxide.
[0734] Suitable lipid solutions can contain mixtures of the desired lipids at various concentrations. In some embodiments, suitable lipid solutions can contain mixtures of the desired lipids at a total concentration in the range of about 0.1 - 100 mg / ml, 0.5 - 90 mg / ml, 1.0 - 80 mg / ml, 1.0 - 70 mg / ml, 1.0 - 60 mg / ml, 1.0 - 50 mg / ml, 1.0 - 40 mg / ml, 1.0 - 30 mg / ml, 1.0 - 20 mg / ml, 1.0 - 15 mg / ml, 1.0 - 10 mg / ml, 1.0 - 9 mg / ml, 1.0 - 8 mg / ml, 1.0 - 7 mg / ml, 1.0 - 6 mg / ml, or 1.0 - 5 mg / ml.
[0735] Any desired lipids can be mixed in any ratio suitable for encapsulating RNA. In some embodiments, suitable lipid solutions contain mixtures of the desired lipids, which include cationic lipids, helper lipids (e.g., non-cationic lipids and / or cholesterol lipids), and / or PEGylated lipids. In some embodiments, suitable lipid solutions contain mixtures of the desired lipids, which include one or more cationic lipids, one or more helper lipids (e.g., non-cationic lipids and / or cholesterol lipids), and one or more PEGylated lipids.
[0736] In some embodiments, the polynucleotides or circular RNAs disclosed herein are formulated using viral vectors. Viral vectors can be derived from a variety of viruses, including adenoviruses, adeno-associated viruses, lentiviruses (e.g., HIV, FIV, and EIAV), and herpesviruses. Examples of commercially available viral vectors include pSilencer adeno (Ambion, Austin, Texas) and pLenti6 / BLOCK-iT TM -DEST (Invitrogen, Carlsbad, California). The selection of viral vectors, the method of expressing polynucleotides or circular RNAs from the vectors, and the method of delivering viral vectors are within the ordinary skill of those in the art. In some embodiments, the viral vector is a recombinant AAV (rAAV) vector, such as those known in the art (PMID: 30245471, PMID: 33614232).
[0737] The present disclosure also provides delivery systems comprising the polynucleotides or circular RNAs disclosed herein. In some embodiments, the delivery system is any one of a liposome, a nanoparticle, a polymer-based delivery system, or a ligand-conjugate delivery system. In some embodiments, the ligand-conjugate delivery system comprises one or more of an antibody, a peptide, a sugar moiety, or a combination thereof.
[0738] In some embodiments, the delivery systems of the present disclosure include nanoparticles containing the polynucleotides or circular RNAs disclosed herein.
[0739] In some embodiments, the nanoparticles include polymer-based nanoparticles, lipid-polymer-based nanoparticles, metal-based nanoparticles, carbon nanotube-based nanoparticles, nanocrystals, or polymeric micelles. In some embodiments, the polymer-based nanoparticles comprise multiblock copolymers, diblock copolymers, polymeric micelles, or hyperbranched macromolecules. In some embodiments, the polymer-based nanoparticles comprise multiblock copolymers and diblock copolymers. In some embodiments, the polymer-based nanoparticles are pH-responsive. In some embodiments, the polymer-based nanoparticles further comprise a buffering component.
[0740] In some embodiments, the delivery system comprises liposomes. Liposomes are spherical vesicles having at least one lipid bilayer and, in some embodiments, an aqueous core. In some embodiments, the lipid bilayer of the liposome may comprise phospholipids. Exemplary but non-limiting examples of phospholipids are phosphatidylcholine, but the lipid bilayer may comprise additional lipids such as phosphatidylethanolamine. Liposomes can be multilamellar, i.e., composed of several lamellar phase lipid bilayers; or unilamellar liposomes having a single lipid bilayer. Liposomes can be made in a specific size range to make them a viable target for phagocytosis. The size range of liposomes can be from 20 nm to 100 nm, 100 nm to 400 nm, 1 μM and greater, or 200 nm to 3 μM. Examples of lipids and lipid-based formulations are provided in U.S. Publication No. 20090023673. In some embodiments, one or more lipids are one or more cationic lipids.
[0741] In some embodiments, the liposomes or nanoparticles of the present invention include micelles. Micelles are aggregates of surfactant molecules. Exemplary micelles comprise aggregates of amphiphilic macromolecules, polymers, or copolymers in an aqueous solution, where the hydrophilic heads contact the surrounding solvent while the hydrophobic tail regions are sequestered at the center of the micelle.
[0742] In some embodiments, the nanoparticles include nanocrystals. Exemplary nanocrystals are crystalline particles having at least one dimension less than 1000 nanometers, preferably less than 100 nanometers.
[0743] In some embodiments, the nanoparticles include polymer-based nanoparticles. In some embodiments, the polymers include multiblock copolymers, diblock copolymers, polymeric micelles, or hyperbranched macromolecules. In some embodiments, the particles comprise one or more cationic polymers. In some embodiments, the cationic polymer is chitosan, protamine, polylysine, polyhistidine, polyarginine, or poly(ethylene)imine. In some embodiments, one or more of the polymers contain a buffering component, a degradable component, a hydrophilic component, a cleavable bond component, or some combination thereof.
[0744] In some embodiments, the nanoparticles or some portions thereof are degradable. In some embodiments, the lipids and / or polymers of the nanoparticles are degradable.
[0745] In some embodiments, any of these delivery systems of the present disclosure may include a buffering component. In some embodiments, any of the present disclosure may include a buffering component and a degradable component. In some embodiments, any of the present disclosure may include a buffering component and a hydrophilic component. In some embodiments, any of the present disclosure may include a buffering component and a cleavable bond component. In some embodiments, any of the present disclosure may include a buffering component, a degradable component, and a hydrophilic component. In some embodiments, any of the present disclosure may include a buffering component, a degradable component, and a cleavable bond component. In some embodiments, any of the present disclosure may include a buffering component, a hydrophilic component, and a cleavable bond component. In some embodiments, any of the present disclosure may include a buffering component, a degradable component, a hydrophilic component, and a cleavable bond component. In some embodiments, the particles are composed of one or more polymers containing any of the foregoing combinations of components.
[0746] In some embodiments, the delivery system includes a ligand-conjugate delivery system. In some embodiments, the ligand-conjugate delivery system comprises one or more of an antibody, a peptide, a sugar moiety, a lipid, or a combination thereof.
[0747] In some embodiments, the polynucleotides or circular RNAs disclosed herein are conjugated, complexed, or encapsulated by one or more lipids or polymers of the delivery system. In some embodiments, the polynucleotides or circular RNAs may be encapsulated in the hollow core of the nanoparticles. Alternatively or additionally, the polynucleotides or circular RNAs may be incorporated, for example, by embedding, into the lipid- or polymer-based shell of the delivery system. Alternatively or additionally, the polynucleotides or circular RNAs may be attached to the surface of the delivery system. In some embodiments, the polynucleotides or circular RNAs are conjugated to one or more lipids or polymers of the delivery system, for example, by covalent attachment.
[0748] In some embodiments, the ligand conjugate delivery system further comprises a targeting agent. In some embodiments, the targeting agent includes a peptide ligand, a nucleotide ligand, a polysaccharide ligand, a fatty acid ligand, a lipid ligand, a small molecule ligand, an antibody, an antibody fragment, an antibody mimetic, or an antibody mimetic fragment.
[0749] In some embodiments, the delivery system of the present disclosure includes a polymer-based delivery system. In some embodiments, the polymer-based delivery system comprises a blend polymer. In some embodiments, the blend polymer is a copolymer comprising a degradable component and a hydrophilic component. In some embodiments, the degradable component of the blend polymer is a polyester, a poly(orthoester), a poly(ethyleneimine), a poly(caprolactone), a polyanhydride, a poly(acrylic acid), a polyglycolide, or a polyurethane. In some embodiments, the degradable component of the blend polymer is poly(lactic acid) (PLA) or poly(lactic-co-glycolic acid) (PLGA). In some embodiments, the hydrophilic component of the blend polymer is a polyalkylene glycol or a polyalkylene oxide. In some embodiments, the polyalkylene glycol is polyethylene glycol (PEG). In some embodiments, the polyalkylene oxide is polyethylene oxide (PEO).
[0750] In some embodiments, the delivery system of the present disclosure is a polymer-based nanoparticle. The polymer-based nanoparticle comprises one or more polymers. In some embodiments, the one or more polymers include a polyester, a poly(orthoester), a poly(ethyleneimine), a poly(caprolactone), a polyanhydride, a poly(acrylic acid), a polyglycolide, or a polyurethane. In some embodiments, the one or more polymers include poly(lactic acid) (PLA) or poly(lactic-co-glycolic acid) (PLGA). In some embodiments, the one or more polymers include poly(lactic-co-glycolic acid) (PLGA). In some embodiments, the one or more polymers include poly(lactic acid) (PLA). In some embodiments, the one or more polymers include a polyalkylene glycol or a polyalkylene oxide. In some embodiments, the polyalkylene glycol is polyethylene glycol (PEG), or the polyalkylene oxide is polyethylene oxide (PEO).
[0751] In some embodiments, the polymer-based nanoparticle comprises a poly(lactic-co-glycolic acid) PLGA polymer. In some embodiments, the PLGA nanoparticles further comprise a targeting agent as described herein.
[0752] In some embodiments, the delivery system of the present disclosure is nanoparticles having an average feature size of less than about 500 nm, 400 nm, 300 nm, 250 nm, 200 nm, 180 nm, 150 nm, 120 nm, 100 nm, 90 nm, 80 nm, 70 nm, 60 nm, 50 nm, 40 nm, 30 nm, or 20 nm. In some embodiments, the nanoparticles have an average feature size of 10 nm, 20 nm, 30 nm, 40 nm, 50 nm, 60 nm, 70 nm, 80 nm, 90 nm, 100 nm, 120 nm, 150 nm, 180 nm, 200 nm, 250 nm, or 300 nm. In some embodiments, the nanoparticles have an average feature size of 10-500 nm, 10-400 nm, 10-300 nm, 10-250 nm, 10-200 nm, 10-150 nm, 10-100 nm, 10-75 nm, 10-50 nm, 50-500 nm, 50-400 nm, 50-300 nm, 50-200 nm, 50-150 nm, 50-100 nm, 50-75 nm, 100-500 nm, 100-400 nm, 100-300 nm, 100-250 nm, 100-200 nm, 100-150 nm, 150-500 nm, 150-400 nm, 150-300 nm, 150-250 nm, 150-200 nm, 200-500 nm, 200-400 nm, 200-300 nm, 200-250 nm, 200-500 nm, 200-400 nm, or 200-300 nm.
[0753] In some embodiments, the target cells lack a protein or enzyme of interest. For example, in the case where nucleic acids need to be delivered to hepatocytes, the hepatocytes represent the target cells. In some embodiments, the compositions of the present disclosure differentially transfect the target cells (i.e., do not transfect non-target cells). The compositions of the present disclosure can also be prepared to preferentially target a variety of target cells, including but not limited to hepatocytes, epithelial cells, hematopoietic cells, epithelial cells, endothelial cells, lung cells, bone cells, stem cells, mesenchymal cells, nerve cells (e.g., meningeal cells, astrocytes, motor neurons, dorsal root ganglion cells, and anterior horn motor neurons), photoreceptor cells (e.g., rod cells and cone cells), retinal pigment epithelial cells, secretory cells, cardiomyocytes, adipocytes, vascular smooth muscle cells, cardiomyocytes, skeletal muscle cells, beta cells, pituitary cells, synovial lining cells, ovarian cells, testicular cells, fibroblasts, B cells, T cells, NK cells, dendritic cells, macrophages, reticulocytes, white blood cells, granulocytes, tumor cells (including tumor cell lines such as Hela, MCF7, PC3, A549, NCI-H727, and HCT-116), immortalized cells (such as MCF10A, HEK293, HEK293T), and primary cell lines (such as hepatic stellate cells, HPrEC, FHC).
[0754] The compositions of the present disclosure can be prepared to preferentially distribute to target cells in specific organs such as the heart, lung, kidney, liver, and spleen as described herein. In some embodiments, the compositions of the present disclosure distribute to liver cells to facilitate delivery and subsequent expression of the circular RNA contained therein by the liver cells (e.g., hepatocytes). The targeted cells can function as a biological "reservoir" or "depot" capable of producing and systematically secreting a functional protein or enzyme. Thus, in some embodiments, the delivery vehicle can target hepatocytes and / or preferentially distribute to liver cells after delivery. In some embodiments, after transfection of the target hepatocytes, the circular RNA loaded in the vehicle is translated and a functional protein product is produced, secreted, and distributed throughout the body. In some embodiments, cells other than hepatocytes (e.g., cells of the lung, spleen, heart, eye, or central nervous system) can be used as reservoir sites for protein production.
[0755] In some embodiments, the compositions of the present disclosure promote the endogenous production of one or more functional proteins and / or enzymes in a subject. In some embodiments, the transfer vehicle comprises a circular RNA encoding the deficient protein or enzyme. After distributing such a composition to a target tissue and subsequently transfecting such target cells, the exogenous circular RNA loaded into the transfer vehicle (e.g., a lipid nanoparticle) can be translated in vivo to produce a functional protein or enzyme encoded by the exogenously administered circular RNA (e.g., a protein or enzyme deficient in the subject). Thus, the compositions of the present disclosure utilize the ability of a subject to translate exogenously or recombinantly produced circular RNAs to produce endogenously translated proteins or enzymes, thereby producing (and, where applicable, secreting) functional proteins or enzymes. The expressed or translated protein or enzyme can also be characterized by the inclusion of native post-translational modifications that are typically not present in recombinantly produced proteins or enzymes, thereby further reducing the immunogenicity of the translated protein or enzyme.
[0756] Administration of the circular RNA encoding the deficient protein or enzyme obviates the need to deliver the nucleic acid to a specific organelle within the target cell. Instead, after transfecting the target cell and delivering the nucleic acid to the cytoplasm of the target cell, the circular RNA content of the transfer vehicle can be translated and a functional protein or enzyme expressed.
[0757] In some embodiments, the circular RNA comprises one or more miRNA binding sites. In some embodiments, the circular RNA comprises one or more miRNA binding sites recognized by miRNAs present in one or more non-target cells or non-target cell types (e.g., Kupffer cells) and absent in one or more target cells or target cell types (e.g., hepatocytes). In some embodiments, the circular RNA comprises one or more miRNA binding sites recognized by miRNAs present in one or more non-target cells or non-target cell types (e.g., Kupffer cells) at a concentration increased compared to the concentration present in one or more target cells or target cell types (e.g., hepatocytes). miRNAs are thought to act by pairing with complementary sequences within an RNA molecule, thereby causing gene silencing.
[0758] 8.6. Pharmaceutical Compositions
[0759] In some embodiments, provided herein are compositions (e.g., pharmaceutical compositions) comprising a therapeutic agent provided herein. In some embodiments, the therapeutic agent is a circular RNA polynucleotide provided herein. In some embodiments, the therapeutic agent is a vector provided herein. In some embodiments, the therapeutic agent is a cell comprising a circular RNA or vector provided herein. In some embodiments, the composition further comprises a pharmaceutically acceptable carrier. In some embodiments, the compositions provided herein comprise a combination of a therapeutic agent provided herein with other pharmaceutically active agents or drugs. In a preferred embodiment, the pharmaceutical composition comprises a cell or a population thereof provided herein.
[0760] With respect to pharmaceutical compositions, the pharmaceutically acceptable carrier can be any conventionally used carrier and is limited only by chemical-physical factors such as solubility and lack of reactivity with one or more active agents, as well as the route of administration. The pharmaceutically acceptable carriers described herein, e.g., vehicles, adjuvants, excipients, and diluents, are well known to those skilled in the art and are readily available to the public. Preferably, the pharmaceutically acceptable carrier is a carrier that is chemically inert to one or more therapeutic agents and has no adverse side effects or toxicity under the conditions of use.
[0761] The choice of carrier will be determined in part by the particular therapeutic agent and by the particular method used to administer the therapeutic agent. Accordingly, there are a variety of suitable formulations for the pharmaceutical compositions provided herein.
[0762] In some embodiments, the pharmaceutical composition comprises a preservative. In some embodiments, suitable preservatives can include, for example, methylparaben, propylparaben, sodium benzoate, and benzalkonium chloride. Optionally, a mixture of two or more preservatives can be used. The preservative or mixture thereof is typically present in an amount of from about 0.0001% to about 2% by weight of the total composition.
[0763] In some embodiments, the pharmaceutical composition comprises a buffer. In some embodiments, suitable buffers can include, for example, citric acid, sodium citrate, phosphoric acid, potassium phosphate, and various other acids and salts. A mixture of two or more buffers can be used optionally. The buffer or mixture thereof is typically present in an amount of from about 0.001% to about 4% by weight of the total composition.
[0764] In some embodiments, the concentration of the therapeutic agent in the pharmaceutical composition can vary, e.g., be less than about 1% or at least about 1%, 2%, 3%, 4%, 5%, 6%, 7%, 8%, 9%, 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, or about 50% or more by weight, and can be selected primarily by fluid volume and viscosity depending on the particular mode of administration chosen.
[0765] The following formulations for oral, aerosol, parenteral (e.g., subcutaneous, intravenous, intra-arterial, intramuscular, intradermal, intraperitoneal, and intrathecal), and topical administration are merely exemplary and in no way limiting. More than one route can be used to administer the therapeutic agents provided herein, and in some cases, a particular route may provide a more immediate and effective response than another route.
[0766] Formulations suitable for oral administration may include or consist of the following: (a) liquid solutions, such as an effective amount of a therapeutic agent dissolved in a diluent (such as water, saline, or orange juice); (b) capsules, sachets, tablets, lozenges, and troches, each containing a predetermined amount of the active ingredient in solid or granular form; (c) powders; (d) suspensions in a suitable liquid; and (e) suitable emulsions. Liquid formulations may include diluents, such as water and alcohols, for example ethanol, benzyl alcohol, and polyvinyl alcohol, with or without pharmaceutically acceptable surfactants. Capsule forms may be of the ordinary hard or soft gelatin type, containing, for example, surfactants, lubricants, and inert fillers such as lactose, sucrose, calcium phosphate, and corn starch. Tablet forms may include one or more of the following: lactose, sucrose, mannitol, corn starch, potato starch, alginic acid, microcrystalline cellulose, acacia, gelatin, guar gum, colloidal silicon dioxide, croscarmellose sodium, talc, magnesium stearate, calcium stearate, zinc stearate, stearic acid, and other excipients, coloring agents, diluents, buffering agents, disintegrants, wetting agents, preservatives, flavoring agents, and other pharmacologically compatible excipients. Lozenge forms may contain the therapeutic agent and a flavoring agent, usually sucrose, acacia, or tragacanth. Pastille forms may contain the therapeutic agent and an inert matrix, such as gelatin and glycerin, or sucrose and acacia, emulsions, gels, etc., in addition to such excipients known in the art.
[0767] Formulations suitable for parenteral administration include aqueous and non-aqueous isotonic sterile injection solutions, which may contain antioxidants, buffers, bacteriostatic agents, and solutes that render the formulation isotonic with the blood of the intended recipient; and aqueous and non-aqueous sterile suspensions, which may contain suspending agents, solubilizers, thickening agents, stabilizers, and preservatives. In some embodiments, the therapeutic agents provided herein may be administered in a physiologically acceptable diluent in a pharmaceutical carrier, with or without the addition of pharmaceutically acceptable surfactants (such as soaps or detergents), suspending agents (such as pectin, carbomer, methylcellulose, hydroxypropylmethylcellulose, or carboxymethylcellulose), or emulsifying agents and other pharmaceutical adjuvants, the diluent being a sterile liquid or liquid mixture, including water, saline, dextrose aqueous solutions, and related sugar solutions, alcohols (such as ethanol or cetyl alcohol), diols (such as propylene glycol or polyethylene glycol), dimethyl sulfoxide, glycerol, ketals (such as 2,2-dimethyl-1,3-dioxolane-4-methanol), ethers, poly(ethylene glycol) 400, oils, fatty acids, fatty acid esters, or glycerol esters, or acetylated fatty acid glycerol esters.
[0768] Oils that can be used in parenteral formulations in some embodiments include petroleum, animal oils, vegetable oils, or synthetic oils. Specific examples of oils include peanut oil, soybean oil, sesame oil, cottonseed oil, corn oil, olive oil, petrolatum, and mineral oil. Suitable fatty acids for parenteral formulations include oleic acid, stearic acid, and isostearic acid. Ethyl oleate and isopropyl myristate are examples of suitable fatty acid esters.
[0769] Suitable soaps for some embodiments of parenteral formulations include alkali metal salts of fatty acids, ammonium salts, and triethanolamine salts, and suitable detergents include (a) cationic detergents, such as dimethyldialkylammonium halides and alkylpyridinium halides, (b) anionic detergents, such as alkyl, aryl, and olefin sulfonates, alkyl, olefin, ether, and glycerol monoester sulfates, and sulfosuccinates, (c) nonionic detergents, such as fatty amine oxides, fatty acid alkanolamides, and polyoxyethylene polypropylene copolymers, (d) amphoteric detergents, such as alkyl-β-aminopropionates and 2-alkyl-imidazoline quaternary salts, and (e) mixtures thereof.
[0770] In some embodiments, the parenteral formulation will contain, for example, from about 0.5% to about 25% by weight of the therapeutic agent in solution. Preservatives and buffers may be used. To minimize or eliminate irritation at the injection site, such compositions may contain one or more nonionic surfactants having a hydrophilic-lipophilic balance value (HLB) of, for example, from about 12 to about 17. The amount of surfactant in such formulations typically ranges from, for example, about 5% to about 15% by weight. Suitable surfactants include polyethylene glycols, sorbitan fatty acid esters (such as sorbitan monooleate), and high molecular weight adducts of ethylene oxide with hydrophobic bases formed by the condensation of propylene oxide with propylene glycol. The parenteral formulation can be presented in unit-dose or multi-dose sealed containers, such as ampoules or vials, and can be stored under lyophilized (freeze-dried) conditions, requiring only the immediate addition of a sterile liquid excipient (e.g., water) for injection prior to use. The temporary injection solutions and suspensions can be prepared from sterile powders, granules, and tablets of the foregoing types.
[0771] In some embodiments, injectable formulations are provided herein. The requirements for an effective pharmaceutical carrier for injectable compositions are well known to those of ordinary skill in the art (see, e.g., Pharmaceutics and Pharmacy Practice, J.B. Lippincott Company, Philadelphia, PA, Banker and Chalmers, eds., pages 238-250 (1982), and ASHP Handbook on Injectable Drugs, Toissel, 4th ed., pages 622-630 (1986)).
[0772] In some embodiments, topical formulations are provided herein. Topical formulations, including those useful for transdermal drug delivery, are suitable for use in the context of certain embodiments provided herein for application to the skin. In some embodiments, the therapeutic agent alone or in combination with other suitable components can be made into an aerosol formulation for administration by inhalation. These aerosol formulations can be placed in a pressurized acceptable propellant, such as dichlorodifluoromethane, propane, nitrogen, etc. They can also be formulated as drugs for non-pressurized formulations, such as in a nebulizer or atomizer. Such spray formulations can also be used for spraying on mucous membranes.
[0773] In some embodiments, the therapeutic agents provided herein can be formulated as inclusion complexes (such as cyclodextrin inclusion complexes) or liposomes. Liposomes can be used to target therapeutic agents to specific tissues. Liposomes can also be used to increase the half-life of therapeutic agents. Many methods can be used to prepare liposomes, such as those described in, for example, Szoka et al., Ann. Rev. Biophys. Bioeng., 9, 467 (1980) and U.S. Patents 4,235,871, 4,501,728, 4,837,028, and 5,019,369.
[0774] In some embodiments, the therapeutic agents provided herein are formulated in timed-release, delayed-release, or sustained-release delivery systems such that delivery of the composition occurs prior to sensitization at the site to be treated and for a time sufficient to cause sensitization. Such systems can avoid repeated administration of the therapeutic agent, thereby increasing convenience for the subject and the physician, and can be particularly suitable for certain composition embodiments provided herein. In some embodiments, the compositions of the present disclosure are formulated such that they are suitable for extended release of the circRNA contained therein. Such extended-release compositions can be conveniently administered to a subject at extended dosing intervals. For example, in some embodiments, the compositions of the present disclosure are administered to a subject twice daily, daily, or every other day. In some embodiments, the compositions of the present disclosure are administered to a subject twice weekly, once weekly, every ten days, every two weeks, every three weeks, every four weeks, once monthly, every six weeks, every eight weeks, every three months, every four months, every six months, every eight months, every nine months, or annually.
[0775] In some embodiments, the protein encoded by the polynucleotides described herein is produced by the target cells for a duration of time. For example, the protein can be produced for more than one hour, more than four hours, more than six hours, more than 12 hours, more than 24 hours, more than 48 hours, or more than 72 hours after administration. In some embodiments, the therapeutic product is expressed at peak levels approximately six hours after administration. In some embodiments, the expression of the therapeutic product is maintained at least at a therapeutic level. In some embodiments, the therapeutic product is expressed at least at a therapeutic level for more than one hour, more than four hours, more than six hours, more than 12 hours, more than 24 hours, more than 48 hours, or more than 72 hours after administration. In some embodiments, a therapeutic level of the therapeutic product can be detected in the patient's serum or tissue (e.g., liver or lung). In some embodiments, the level of the detectable therapeutic product results from continuous expression of the circRNA composition for the following time periods after administration: more than one hour, more than four hours, more than six hours, more than 12 hours, more than 24 hours, more than 48 hours, or more than 72 hours.
[0776] In some embodiments, the protein encoded by the polynucleotides described herein is produced at levels above normal physiological levels. The protein levels can be increased compared to a control. In some embodiments, the control is the baseline physiological level of the therapeutic product in a normal individual or a population of normal individuals. In some embodiments, the control is the baseline physiological level of the therapeutic product in an individual lacking the relevant protein or polypeptide or in a population of individuals lacking the relevant protein or polypeptide. In some embodiments, the control can be the normal level of the relevant protein or polypeptide in an individual to whom the composition is administered. In some embodiments, the control is the expression level of the therapeutic product at one or more comparable time points after other therapeutic interventions, such as after direct injection of the corresponding therapeutic product.
[0777] In some embodiments, the level of the protein encoded by the polynucleotides described herein can be detected 3 days, 4 days, 5 days, or 1 week or longer after administration. Increased levels of the secreted protein can be observed in serum and / or tissue (e.g., liver or lung).
[0778] In some embodiments, the method results in a sustained circulating half-life of the protein encoded by the polynucleotides described herein. For example, the protein can be detected for a longer number of hours or days compared to the half-life observed by subcutaneous injection of the protein or mRNA encoding the protein. In some embodiments, the half-life of the protein is 1 day, 2 days, 3 days, 4 days, 5 days, or 1 week or longer.
[0779] Many types of release delivery systems are available and known to those of ordinary skill in the art. They include polymer-based systems such as poly(lactide-co-glycolide), poly(ortho esters), polycaprolactone, poly(ester amides), poly(ortho esters), poly(hydroxybutyric acid), and poly(anhydrides). Microcapsules of the above drug-containing polymers are described, for example, in U.S. Patent No. 5,075,109. Delivery systems also include: non-polymer systems that are lipids, including sterols such as cholesterol, cholesterol esters, and fatty acids or neutral fats (such as monoglycerides and triglycerides); hydrogel release systems; sylastic systems; peptide-based systems; wax coatings; compressed tablets using conventional binders and excipients; partially fused implants; and the like. Specific examples include, but are not limited to: (a) erosion systems in which the active composition is contained in a matrix in the form of a matrix, such as those described in U.S. Patent Nos. 4,452,775, 4,667,014, 4,748,034, and 5,239,660, and (b) diffusion systems in which the active ingredient permeates from the polymer at a controlled rate, such as those described in U.S. Patent Nos. 3,832,253 and 3,854,480. In addition, pump-based hardware delivery systems can be used, some of which are adapted for implantation.
[0780] In some embodiments, a therapeutic agent can be conjugated directly or indirectly, via a linking moiety, to a targeting moiety. Methods for conjugating therapeutic agents to targeting moieties are known in the art. See, e.g., Wadwa et al., J. Drug Targeting 3:111 (1995) and U.S. Patent 5,087,616.
[0781] In some embodiments, the therapeutic agents provided herein are formulated in depot form such that the manner in which the therapeutic agent is released into the body to which it is administered is controlled with respect to time and location in the body (see, e.g., U.S. Patent 4,450,150). The depot form of the therapeutic agent can be, for example, an implantable composition that comprises the therapeutic agent and a porous or non-porous material (such as a polymer), wherein the therapeutic agent is encapsulated by the material or diffused throughout the material, and / or degradation of the non-porous material. The depot is then implanted at the desired location in the body, and the therapeutic agent is released from the implant at a predetermined rate.
[0782] 8.7. Methods of Use
[0783] In some aspects, provided herein is a method of treating and / or preventing a disorder such as cancer, the method comprising introducing a pharmaceutical composition provided herein into a subject in need thereof (e.g., a subject having cancer). In some embodiments, the pharmaceutical composition comprises a circular RNA polynucleotide provided herein. In some embodiments, the pharmaceutical composition comprises a vector provided herein. In some embodiments, the pharmaceutical composition comprises a cell (e.g., a human cell) comprising a polynucleotide provided herein (e.g., a circular RNA or a vector provided herein).
[0784] Thus, in some embodiments, provided herein are methods of treating and / or preventing a disease in a subject (e.g., a mammalian subject such as a human subject). Without being bound by a particular theory or mechanism, the circular RNAs provided herein can be used to express a therapeutic protein for protein replacement therapy.
[0785] In some embodiments, the therapeutic agents provided herein are co-administered with one or more additional therapeutic agents (e.g., in the same pharmaceutical composition or in separate pharmaceutical compositions). In some embodiments, the therapeutic agents provided herein can be administered first, and then one or more additional therapeutic agents can be administered, and vice versa. Alternatively, the therapeutic agents provided herein and one or more additional therapeutic agents can be administered simultaneously.
[0786] In some embodiments, the therapeutic agent is a cell or population of cells comprising a circular RNA or a vector provided herein that expresses a protein encoded by the circular RNA or vector. In some embodiments, the cells administered are allogeneic to the subject being treated. In some embodiments, the cells administered are autologous to the subject being treated.
[0787] In some embodiments, the subject is a mammal. In some embodiments, the mammals referred to herein can be any mammal, including but not limited to mammals of the order Rodentia (such as mice and hamsters), or mammals of the order Lagomorpha (such as rabbits). Mammals can be from the order Carnivora, including felines (cats) and canines (dogs). Mammals can be from the order Artiodactyla, including bovines (cows) and porcines (pigs); or from the order Perissodactyla, including equines (horses). Mammals can be primates, New World monkeys (Ceboid) or Simoids (monkeys), or anthropoid primates (humans and apes). Preferably, the mammal is a human.
[0788] A therapeutically effective amount of an RNA polynucleotide (e.g., circular RNA) can be administered by a variety of routes, including parenteral administration, such as intravenous, intraperitoneal, intramuscular, intrasternal, or intra-articular injection or infusion.
[0789] 8.7.1. Gene Therapy
[0790] Compositions and methods are described for administering a therapeutically effective amount of an RNA polynucleotide or circular RNA disclosed herein to a subject (e.g., a human) to treat a disease or disorder, such as a disease or disorder in which a therapeutic product is involved in the initiation, development, and / or manifestation of the disease or disorder, including diseases or disorders associated with protein misfolding and / or protein degenerative diseases and diseases or disorders associated with gene mutations. In some embodiments, the RNA polynucleotide or circular RNA described herein comprises an IRES-like sequence, an endogenous IRES sequence, or a variant thereof, or a combination thereof. In some embodiments, the IRES-like sequence, endogenous IRES sequence, or a variant thereof, or a combination thereof is optimized as disclosed herein for the expression of a therapeutic product. In some embodiments, the RNA polynucleotide or circular RNA provided herein can be used, for example, in gene therapy. In some embodiments, the RNA polynucleotide or circular RNA described herein can be used in combination with a CRISPR-Cas system for RNA editing, zinc finger nucleases (ZFNs), transcription activator-like effector nucleases (TALENs), meganucleases, and Puf.
[0791] 8.7.2. Vaccines
[0792] This disclosure describes compositions and methods for administering a therapeutically effective amount of an RNA polynucleotide or circular RNA disclosed herein to a subject (e.g., a human) to stimulate an immune response against an antigen or agent (e.g., an infective agent such as a pathogen), generate antibodies against the antigen or agent, destroy the antigen or agent, and / or develop immunological memory of the antigen or agent. In some embodiments, the RNA polynucleotides described herein comprise an IRES-like sequence, an endogenous IRES sequence, or a variant thereof, or a combination thereof. In some embodiments, the RNA polynucleotides described herein comprise an expression sequence encoding a therapeutic product. In some embodiments, the IRES-like sequence, endogenous IRES sequence, or variant thereof, or a combination thereof is optimized as disclosed herein for expressing a therapeutic product in a subject. In some embodiments, the RNA polynucleotides described herein can be used directly to prepare a vaccine (e.g., an RNA vaccine).
[0793] 8.7.3. Antibody-Based Therapies
[0794] In some embodiments, the therapeutic product is a polypeptide or protein. In some embodiments, the polypeptide or protein is similar to a weakened or inactivated form of a disease-causing agent (e.g., a pathogen), which can be selected from microorganisms (such as bacteria, viruses, fungi, parasites) or one or more components of such microorganisms (such as toxins, proteins (e.g., surface proteins) and / or cell walls). In some embodiments, the therapeutic product is an antigen or agent that can stimulate the body's immune system to recognize the antigen or agent, generate antibodies against the antigen or agent, destroy the antigen or agent, and / or develop immunological memory of the antigen or agent. In some embodiments, the therapeutic product is an antigen or agent that can induce and / or enhance vaccine-induced memory and / or enable the immune system to respond rapidly and effectively to the antigen or agent upon subsequent encounter with the antigen or agent.
[0795] In some embodiments, the therapeutic product is derived from an infective agent or a part, component, and / or product thereof (e.g., cell wall, genomic sequence, membrane, capsid, protein, lipid, glycan, toxin). In some embodiments, the infective agent is selected from viruses, bacteria, fungi, protozoa, and worms. In some embodiments, the infective agent is selected from viruses and bacteria.
[0796] In some embodiments, the infectious agent is a virus selected from the following: adenovirus; herpes simplex type 1; herpes simplex type 2; encephalitis virus, papillomavirus; varicella-zoster virus; Epstein-Barr virus; human cytomegalovirus; human herpesvirus 8; BK virus; JC virus; smallpox; poliovirus; human bocavirus; parvovirus B19; human astrovirus; norovirus; coxsackievirus; hepatitis A virus; hepatitis B virus; hepatitis C virus; hepatitis D virus; hepatitis E virus; rhinovirus; severe acute respiratory syndrome (SARS) virus; yellow fever virus; dengue virus; West Nile virus; rubella virus; human immunodeficiency virus (HIV); influenza virus; Guanarito virus; Junin virus; Lassa virus; Machupo virus; Sabia virus; Crimean-Congo hemorrhagic fever virus; Ebola virus; Marburg virus; measles virus; mumps virus; parainfluenza virus; respiratory syncytial virus (RSV); human metapneumovirus; Hendra virus; Nipah virus; rabies virus; rotavirus; orbivirus; Coltivirus; Banna virus; human enterovirus; hantavirus; West Nile virus; coronavirus, SARS-related coronavirus (SARS-CoV), SARS-CoV-2 virus (COVID-19-related); Middle East respiratory syndrome coronavirus; Japanese encephalitis virus; vesicular exanthema virus; and Eastern equine encephalitis.
[0797] In some embodiments, the infectious agent is a bacterium selected from the following: tuberculosis (Mycobacterium tuberculosis), Clostridium difficile resistant to clindamycin, Clostridium difficile resistant to fluoroquinolone, methicillin-resistant Staphylococcus aureus (MRSA), multidrug-resistant Enterococcus faecalis, multidrug-resistant Enterococcus faecium, multidrug-resistant Pseudomonas aeruginosa, multidrug-resistant Acinetobacter baumannii, and vancomycin-resistant Staphylococcus aureus (VRSA).
[0798] In some embodiments, the infectious agent is associated with a human, non-human primate, or other animal, such as a bird, pig, horse, dog, cat, rabbit, mouse, rat, cow, sheep, goat, and deer.
[0799] As will be understood by those skilled in the art, the terms "antibody therapy" or "antibody-based therapy" are used interchangeably herein to refer to a method of treating a disease or disorder in a subject, wherein the method comprises administering one or more antibodies.
[0800] In some embodiments, the RNA polynucleotides described herein comprise an expression sequence encoding a therapeutic product. In some embodiments, the therapeutic product is a polypeptide or protein. In some embodiments, the polypeptide or protein is a therapeutic antibody. In some embodiments, the RNA polynucleotides described herein and the compositions comprising the RNA polynucleotides can be used in the manufacture of antibodies for use in antibody therapy.
[0801] Exemplary therapeutic antibodies include antibodies for treating cancer. Non-limiting examples of therapeutic antibodies include 131I-tositumomab (follicular lymphoma, B-cell lymphoma, leukemia), 3F8 (neuroblastoma), 8H9, abagovomab (ovarian cancer), adecatumumab (prostate cancer and breast cancer), afutuzumab (lymphoma), pegaptanib sodium, alemtuzumab (B-cell chronic lymphocytic leukemia, T-cell lymphoma), amatuximab, AME-133v (follicular lymphoma, cancer), AMG 102 (advanced renal cell carcinoma), anatumomab mafenatox (non-small cell lung cancer), apolizumab (solid tumors, leukemia, non-Hodgkin lymphoma, lymphoma), batokizumab (cancer, viral infections), betomuzumab (non-Hodgkin lymphoma), belimumab (non-Hodgkin lymphoma), bevacizumab (colon cancer, breast cancer, brain and central nervous system tumors, lung cancer, hepatocellular carcinoma, renal cancer, breast cancer, pancreatic cancer, bladder cancer, sarcoma, melanoma, esophageal cancer; gastric cancer, metastatic renal cell carcinoma;Renal cell carcinoma, glioblastoma, liver cancer, proliferative diabetic retinopathy, macular degeneration), Bivatuzumab mertansine (squamous cell carcinoma), Blinatumomab, Brentuximab vedotin (blood cancer), Canertinib (colon cancer, gastric cancer, pancreatic cancer, NSCLC), Moxetumomab pasudotox (colorectal cancer), Lexatumumab (cancer), Caroatuximab pentetate (prostate cancer), Carlumab, Catumaxomab (ovarian cancer, fallopian tube tumor, peritoneal tumor), Cetuximab (metastatic colorectal cancer and head and neck cancer), Pertuzumab (ovarian cancer and other solid tumors), Sicutuximab (solid tumors), Ticlatuzumab (pancreatic cancer), CNTO 328 (B-cell non-Hodgkin lymphoma, multiple myeloma, Castleman's disease, ovarian cancer), CNTO 95 (melanoma), Conatumumab, Dacetuzumab (blood cancer), Daratumumab, Denosumab (myeloma, giant cell tumor of bone, breast cancer, prostate cancer, osteoporosis), Dimab (lymphoma), Zalutumumab, Elotuzumab (malignant melanoma), Edrecolomab (colorectal cancer), Elotuzumab (multiple myeloma), Asimadoline, Inotuzumab ozogamicin, Enfortumab vedotin, Ipratuzumab (autoimmune disease, systemic lupus erythematosus, non-Hodgkin lymphoma, leukemia), Eutrombopag (breast cancer), Eutrombopag (breast cancer), Ixazomib (melanoma, prostate cancer, ovarian cancer), Farletuzumab (ovarian cancer), FBTA05 (chronic lymphocytic leukemia), Figitumumab (cancer), Figitumumab (adrenocortical carcinoma, non-small cell lung cancer), Flanvotumab (melanoma), Galiximab (B-cell lymphoma), Galiximab (non-Hodgkin lymphoma), Ganitumab, GC1008 (advanced renal cell carcinoma;Malignant melanoma, pulmonary fibrosis), Zinotuzumab (leukemia), Inotuzumab Ozogamicin (acute myeloid leukemia), Gemtuzumab (clear cell renal cell carcinoma), Vitaxin-Gemtuzumab (melanoma, breast cancer), GS6624 (idiopathic pulmonary fibrosis and solid tumors), HuC242-DM4 (colon cancer, gastric cancer, pancreatic cancer), HuHMFG1 (breast cancer), HuN901-DM1 (myeloma), Ertumaxomab (recurrent or refractory low-grade follicular or transformed B-cell non-Hodgkin lymphoma (NHL)), Elotuzumab, ID09C3 (non-Hodgkin lymphoma), Removab, Obinutuzumab, Intetumumab (solid tumors (prostate cancer, melanoma)), Ipilimumab (sarcoma, melanoma, lung cancer, ovarian cancer leukemia, lymphoma, brain and central nervous system tumors, testicular cancer, prostate cancer, pancreatic cancer, breast cancer), Eritumaxomab (Hodgkin lymphoma), Rabatuzumab (colorectal cancer), Lesatuximab, Lintuzumab, Motexafin-Lovastatin (multiple myeloma, non-Hodgkin lymphoma, Hodgkin lymphoma), Lucatumumab (chronic lymphocytic leukemia), Mapatumumab (colon cancer, myeloma), Matuzumab (lung cancer, cervical cancer, esophageal cancer), MDX-060 (Hodgkin lymphoma, lymphoma), MEDI 522 (solid tumors, leukemia, lymphoma, small intestine cancer, melanoma), Mitumomab (small cell lung cancer), Mogamulizumab, MORab-003 (ovarian cancer, fallopian tube cancer, peritoneal cancer), MORab-009 (pancreatic cancer, mesothelioma, ovarian cancer, non-small cell lung cancer, fallopian tube cancer, peritoneal cavity cancer), Pasotuxizumab, MT103 (non-Hodgkin lymphoma), Talactoferrin-Nacotuzumab (colorectal cancer), Etaracizumab (non-small cell lung cancer, renal cell carcinoma), Natalizumab, Necitumumab (non-small cell lung cancer), Nimotuzumab (squamous cell carcinoma, head and neck cancer, nasopharyngeal cancer, glioma), Nimotuzumab (squamous cell carcinoma, glioma, solid tumors, lung cancer), Olaparib, Onartuzumab (cancer), Montelukast-Optumumab, Oregovomab (ovarian cancer), Oregovomab (ovarian cancer, fallopian tube cancer, peritoneal cavity cancer), PAM4 (pancreatic cancer), Panitumumab (colon cancer, lung cancer, breast cancer; bladder cancer;Ovarian cancer), Pertuzumab, Pemtumomab, Pertuzumab (breast cancer, ovarian cancer, lung cancer, prostate cancer), Pretomanid (brain cancer), Recotamab, Retuximab, Ramucirumab (solid tumors), Relotamab (solid tumors), Rituximab (urticaria, rheumatoid arthritis, ulcerative colitis, chronic focal encephalitis, non-Hodgkin lymphoma, lymphoma, chronic lymphocytic leukemia), Lobatamab, Samalizumab, SGN-30 (Hodgkin lymphoma, lymphoma), SGN-40 (non-Hodgkin lymphoma, myeloma, leukemia, chronic lymphocytic leukemia), Cilazumab, Siltuximab, Tabalumab (B-cell carcinoma), Tacatuzumab tetraxetan, Taplitumomab paptox, Tenatumomab, Teprotumumab (hematological malignancies), TGN1412 (chronic lymphocytic leukemia, rheumatoid arthritis), Ticilimumab (= Tremelimumab), Tigatuzumab, TNX-650 (Hodgkin lymphoma), Tositumomab (follicular lymphoma, B-cell lymphoma, leukemia, myeloma), Trastuzumab (breast cancer, endometrial cancer, solid tumors), TRBS07 (melanoma), Tremelimumab, TRU-016 (chronic lymphocytic leukemia), TRU-016 (non-Hodgkin lymphoma), Tucotuzumab celmoleukin, Ublituximab, Urelumab, Veltuzumab (non-Hodgkin lymphoma), Veltuzumab (IMMU-106) (non-Hodgkin lymphoma), Volociximab (renal cell carcinoma, pancreatic cancer, melanoma), Votumumab (colorectal tumors), WX-G250 (renal cell carcinoma), Zalutumumab (head and neck cancer, squamous cell carcinoma) and Zanolimumab (T-cell lymphoma).;
[0802] Exemplary therapeutic antibodies include antibodies for treating immune disorders. Non-limiting examples of therapeutic antibodies include zanilimumab (psoriasis), epratuzumab (autoimmune diseases, systemic lupus erythematosus, non-Hodgkin lymphoma, leukemia), etrolizumab (inflammatory bowel disease), fontolizumab (Crohn's disease), elsilimomab (autoimmune diseases), mepolizumab (hypereosinophilic syndrome, asthma, eosinophilic gastroenteritis, Churg-Strauss syndrome, eosinophilic esophagitis), mirvetuximab (multiple myeloma and other hematologic malignancies), immune globulin (primary immunodeficiency), priliximab (Crohn's disease, multiple sclerosis), rituximab (urticaria, rheumatoid arthritis, ulcerative colitis, chronic focal encephalitis, non-Hodgkin lymphoma, lymphoma, chronic lymphocytic leukemia), loncastuximab (systemic lupus erythematosus), lulizumab (rheumatic diseases), sarilumab (rheumatoid arthritis, ankylosing spondylitis), vedolizumab (Crohn's disease, ulcerative colitis), wiselizumab (Crohn's disease, ulcerative colitis), reslizumab (airway, skin, and gastrointestinal inflammation), adalimumab (rheumatoid arthritis, Crohn's disease, ankylosing spondylitis, psoriatic arthritis), aselizumab (severely injured patients), atinumab (treating the nervous system), atezolizumab (rheumatoid arthritis, systemic juvenile idiopathic arthritis), bapineuzumab (severe allergic disorders), besilesomab (inflammatory lesions and metastases), BMS-945429, ALD518 (cancer and rheumatoid arthritis), brodalumab (psoriasis, rheumatoid arthritis, inflammatory bowel disease, multiple sclerosis), blisibimod (inflammatory diseases), canakinumab (rheumatoid arthritis), canakinumab (cryopyrin-associated periodic syndromes (CAPS), rheumatoid arthritis, chronic obstructive pulmonary disease), pegcetacoplan (Crohn's disease), erlizumab (heart attack, stroke, traumatic shock), fezakinumab (rheumatoid arthritis, psoriasis), golimumab (rheumatoid arthritis, psoriatic arthritis, ankylosing spondylitis), gomiliximab (allergic asthma), infliximab (rheumatoid arthritis, Crohn's disease, ankylosing spondylitis, psoriatic arthritis, plaque psoriasis, Bykhterev's disease (MorbusBechterew), ulcerative colitis), Mavrilimumab (rheumatoid arthritis), Natalizumab (multiple sclerosis), Ocrelizumab (multiple sclerosis, rheumatoid arthritis, lupus erythematosus, blood cancer), Odulimomab (prevention of organ transplant rejection, immune diseases), Ofatumumab (chronic lymphocytic leukemia, follicular non-Hodgkin lymphoma, B-cell lymphoma, rheumatoid arthritis, relapsing-remitting multiple sclerosis, lymphoma, B-cell chronic lymphocytic leukemia), Ozzolizumab (inflammation), Certolizumab (reducing side effects of heart surgery), Lovelizumab (hemorrhagic shock), SBI-087 (rheumatoid arthritis), SBI-087 (systemic lupus erythematosus), Secukinumab (uveitis, rheumatoid arthritis psoriasis), Sirukumab (rheumatoid arthritis), Talizumab (allergic reaction), Tocilizumab (rheumatoid arthritis, systemic juvenile idiopathic arthritis, Castleman disease), Toripalimab (rheumatoid arthritis, lupus nephritis), TRU-015 (rheumatoid arthritis), TRU-016 (autoimmune diseases and inflammation), Ustekinumab (multiple sclerosis, psoriasis, psoriatic arthritis), Ustekinumab (IL-12 / IL-23 blocker) (plaque psoriasis, psoriatic arthritis, multiple sclerosis, sarcoidosis, latter comparison), Vepalimomab (inflammation), Ato-ozolizumab (systemic lupus erythematosus, graft-versus-host disease), Jislimumab (SLE, dermatomyositis, polymyositis), Lucatumumab (allergy) and Rho(D) immunoglobulin (rhesus disease); or an antibody selecte...
Claims
1. An RNA polynucleotide, the RNA polynucleotide comprising a construct of Formula I, Formula II, Formula III, Formula IV or Formula V: 5’-(A1) 0-1 -(L) n -TI-(L) n -Z1-(L) n -(B1) 0-1 -3’(I), 5’-(A1) 0-1 -(L) n -Z1 B -(L) n -TI-(L) n -Z1 A- (L) n -(B1) 0-1 -3’(II), 5’-(3’ intron fragment)-(E2)-(L) n -TI-(L) n -Z1-(L) n -(E1)-(5’ intron fragment)-3’ (III), 5’-(3’ intron fragment)-(E2)-(L) n -Z1 B -(L) n -TI-(L) n -Z1 A- -(L) n -(E1)-(5’ intron fragment)-3’(IV), or 5’-TI-(L) n -Z1-3’(V), wherein: TI is an engineered translation initiation element comprising an IRES-like polynucleotide sequence, wherein the IRES-like polynucleotide sequence comprises a nucleic acid sequence that is about or at least about 90%, 95%, 97%, 98%, 99% or 100% identical to SEQ ID NO: 1025 - 14161 or SEQ ID NO: 14412 - 15341; Z1 is an expression sequence encoding a therapeutic product; Z1 A is the first part of an expression sequence encoding a therapeutic product; Z1 B is the second part of the expression sequence encoding the therapeutic product; each L is independently a linker sequence; A1 and B1 are each independently a sequence capable of cyclizing the RNA polynucleotide, or A1 and B1 each independently comprise a nucleotide derivative capable of ligating the 5'-end and the 3'-end by a 3' to 5' phosphodiester linkage to cyclize the RNA polynucleotide; the 5' intron fragment and the 3' intron fragment are each a fragment of a Group II intron, wherein the 5' intron fragment is on the 5' side of the 3' intron fragment in the Group II intron; E1 is the 5'-adjacent exon fragment of the Group II intron, having a length ≥ 0 nucleotides; E2 is the 3'-adjacent exon fragment of the Group II intron, having a length ≥ 0 nucleotides; and n is an integer selected from 0 to 2.
2. The RNA polynucleotide according to claim 1, the RNA polynucleotide further comprising a 5' homologous arm at the 5' end of the 3' intron fragment.
3. The RNA polynucleotide according to claim 1, the RNA polynucleotide further comprising a 3' homologous arm at the 3' end of the 5' intron fragment.
4. The RNA polynucleotide according to claim 1, the RNA polynucleotide further comprising a 5' homologous arm at the 5' end of the 3' intron fragment and a 3' homologous arm at the 3' end of the 5' intron fragment.
5. The RNA polynucleotide according to any one of claims 1 - 4, wherein E1 and E2 each independently have a length of 0 to 20 nucleotides.
6. The RNA polynucleotide according to any one of claims 1 - 5, wherein the 5' intron fragment and the 3' intron fragment are obtained by splitting a Group II intron into two fragments at an unpaired region, wherein the unpaired region is preferably selected from the linear region between two adjacent domains of the Group II intron and the loop region of the stem-loop structure of domain 4 of the Group II intron.
7. An RNA polynucleotide, the RNA polynucleotide comprising: an engineered translation initiation element (TI) comprising an IRES-like polynucleotide sequence, wherein the IRES-like polynucleotide sequence comprises a nucleic acid sequence that is about or at least about 90%, 95%, 97%, 98%, 99% or 100% identical to SEQ ID NO: 1025 - 14161 or SEQ ID NO: 14412 - 15341.
8. The RNA polynucleotide according to any one of claims 1-7, wherein the length of the IRES-like polynucleotide sequence is 6-12 nucleotide residues.
9. The RNA polynucleotide according to any one of claims 1-8, wherein the TI further comprises a second IRES-like polynucleotide sequence.
10. The RNA polynucleotide according to claim 9, wherein the second IRES-like polynucleotide sequence comprises a nucleic acid sequence that is about or at least about 90%, 95%, 97%, 98%, 99% or 100% identical to SEQ ID NO: 1025-14161 or SEQ ID NO: 14412-15341.
11. The RNA polynucleotide according to any one of claims 1-6 and 8-10, wherein each L independently comprises a 5'UTR, 3'UTR, poly-A sequence, poly-A-C sequence, poly-C sequence, poly-U sequence, poly-G sequence, ribosome binding site, aptamer, riboswitch, ribozyme, small RNA binding site, translation regulatory element (e.g., Kozak sequence), protein binding site (e.g., PTBP1 or HUR), unnatural nucleotide or non-nucleotide chemical linker.
12. The RNA polynucleotide according to any one of claims 1-6 and 8-11, wherein each L independently has a length of about 3 to about 100 nucleotide residues.
13. The RNA polynucleotide according to any one of claims 1-6 and 8-12, wherein each L independently comprises the nucleic acid sequence of RCC, where R is guanine or adenine.
14. The RNA polynucleotide according to any one of claims 1-13, wherein the RNA polynucleotide is a single-stranded RNA polynucleotide.
15. The RNA polynucleotide according to any one of claims 1-14, wherein the RNA polynucleotide is a circular RNA polynucleotide.
16. The RNA polynucleotide according to any one of claims 1-14, wherein the RNA polynucleotide is a linear RNA polynucleotide.
17. The RNA polynucleotide according to any one of claims 1-14 and 16, wherein the RNA polynucleotide is capable of cyclizing in the absence of an enzyme.
18. A polypeptide expressed by the RNA polynucleotide according to any one of claims 1-17.
19. A DNA vector encoding the RNA polynucleotide according to any one of claims 1-17.
20. A cell comprising the RNA polynucleotide according to any one of claims 1-17, the polypeptide according to claim 18, or the DNA vector according to claim 19.
21. A composition comprising the RNA polynucleotide according to any one of claims 1-17, the polypeptide according to claim 18, the DNA vector according to claim 19, or the cell according to claim 20 and a pharmaceutically acceptable carrier.
22. A method for preparing a cell population, the method comprising contacting the cells of the population with the RNA polynucleotide according to any one of claims 1-17, the polypeptide according to claim 18, or the DNA vector according to claim 19.
23. A method for generating an internal ribosome entry site (IRES)-like polynucleotide sequence, the method comprising the steps of: (a) generating a polynucleotide query sequence consisting of X nucleic acid residues in length, where X is an integer greater than or equal to 3; (b) generating X-Y+1 overlapping polynucleotide fragment sequences within the polynucleotide query sequence, where each polynucleotide fragment sequence consists of Y nucleic acid residues in length, where the first position of each polynucleotide fragment sequence is n and the last position of the same polynucleotide fragment sequence is Y+n-1, and where n represents each positive integer between 1 and X-Y+1; (c) determining the enrichment score of each polynucleotide fragment sequence of (b); (d) determining the numerical score of the polynucleotide query sequence by summing the enrichment scores of each polynucleotide fragment sequence of (c); and (e) identifying the polynucleotide query sequence as an engineered IRES-like polynucleotide sequence according to a reference value.
24. The method according to claim 23, wherein the reference value is calculated according to the following formula: Reference value = (23.75*X) – 84.85 – Z, where Z represents any number from 1 to 10000, and X represents the length of the IRES-like polynucleotide sequence.
25. The method according to claim 24, wherein Z represents any number from 50 to 5000.
26. The method according to claim 24, wherein Z represents any number from 175 to 2500.
27. The method according to claim 24, wherein Z represents any number from 150 to 200.
28. The method according to claim 23, wherein the reference value is a characteristic in the absence of therapeutic product expression.
29. The method according to claim 23, wherein the reference value is the average score of all numerical scores of two or more IRES-like polynucleotide sequences or two or more natural IRES sequences or a combination thereof.
30. The method according to claim 23, wherein the reference value is greater than or equal to 0.
31. The method according to any one of claims 23-30, wherein X is an integer greater than or equal to 5.
32. The method according to any one of claims 23-30, wherein X is an integer greater than or equal to 6.
33. The method according to any one of claims 23-32, wherein the length of the overlapping polynucleotide fragment sequences within the polynucleotide query sequence is 5, 6, 7, 8, 9 or 10 nucleic acid residues.
34. The method according to any one of claims 23-33, wherein the enrichment score of the polynucleotide fragment sequence is determined by the following formula: i) generating an expression plasmid library, wherein each expression plasmid in the library contains a different polynucleotide fragment sequence and a reporter gene; ii) contacting the cell population with the expression plasmid library; iii) quantifying the expression level of the reporter gene corresponding to each expression plasmid of the expression plasmid library; iv) dividing the total cell population into a first population and a second population based on the protein expression level of iii); v) determining the enrichment score of the polynucleotide fragment sequence using the following system of equations: where f 1 is the frequency of the polynucleotide fragment sequence in the first population, where f 2 is the frequency of the same polynucleotide fragment sequence in the second population, where N 1 is the size of the first population, and Where N 2 is the size of the second population.
35. The method according to claim 34, wherein the first population has a protein expression level in the top 0-10% of the protein expression level of the total population, and wherein the second population has a protein expression level in the bottom 10%-90% of the protein expression level of the total population.
36. The method according to claim 34 or 35, wherein the first population has a protein expression level in the top 50.1% of the protein expression level of the total population, and wherein the second population has a protein expression level in the bottom 49.9% of the protein expression level of the total population.
37. The method according to claim 34 or 35, wherein the first population has a protein expression level in the top 10% of the protein expression level of the total population, and wherein the second population has a protein expression level in the bottom 90% of the protein expression level of the total population.
38. The method according to any one of claims 23-37, wherein the length of the polynucleotide fragment sequence is 5 nucleic acid residues.
39. The method according to claim 38, wherein the polynucleotide fragment sequence is selected from SEQ ID NO: 1-1024, and the enrichment score of the polynucleotide fragment sequence is shown in Table 1.
40. An RNA polynucleotide comprising an engineered translation initiation element (TI), wherein the TI comprises an internal ribosome entry site (IRES)-like polynucleotide sequence generated by the method according to any one of claims 23-39.
41. A polypeptide expressed by the RNA polynucleotide according to claim 40.
42. A DNA vector encoding the RNA polynucleotide according to claim 40.
43. A cell comprising the RNA polynucleotide according to claim 40, the polypeptide according to claim 41, or the DNA vector according to claim 42.
44. A composition comprising the RNA polynucleotide according to claim 40, the polypeptide according to claim 41, the DNA vector according to claim 42, or the cell according to claim 43 and a pharmaceutically acceptable carrier.
45. A method for preparing a cell population, the method comprising contacting the cells of the population with the RNA polynucleotide according to claim 40, the polypeptide according to claim 41, or the DNA vector according to claim 42.
46. A method for modulating protein expression in a subject in need thereof, the method comprising administering to the subject a therapeutically effective amount of the polynucleotide according to any one of claims 1-17 and 40, the polypeptide according to claim 18 or 41, the DNA vector according to claim 19 or 42, or the cell according to claim 20 or 43.
47. A method for treating or preventing a disease or disorder in a subject in need thereof, the method comprising administering to the subject a therapeutically effective amount of the polynucleotide according to any one of claims 1171 and 40, the polypeptide according to claim 18 or 41, the DNA vector according to claim 19 or 42, or the cell according to claim 20 or 43.
48. Use of the polynucleotide according to any one of claims 1-17 and 40, the polypeptide according to claim 18 or 41, the DNA vector according to claim 19 or 42, or the cell according to claim 20 or 43, in the manufacture of a medicament for modulating protein expression or treating or preventing a disease or disorder in a subject in need thereof.
49. The polynucleotide according to any one of claims 1-17 and 40, the polypeptide according to claim 18 or 41, the DNA vector according to claim 19 or 42, or the cell according to claim 20 or 43, for modulating protein expression or treating or preventing a disease or disorder in a subject in need thereof.
Citation Information
Patent Citations
Construction body for preparing circular RNA, method and application thereof
CN115404240A
Apparatus for obtaining combustible gas.
US1111995A
Lipid containing formulations
US20090023673A1
Process of Preparing mRNA-Loaded Lipid Nanoparticles
US20180153822A1
Method of making an inflatable balloon catheter
US3832253A
Cited By
MITE transposon for regulating and controlling content of sulfo metabolite in tea tree and application of MITE transposon
CN119824018A
Mite transposon for regulating content of sulfuryl metabolite of camellia sinensis and application thereof
CN119824018B