Ribozyme-mediated RNA assembly and expression
The system employs ribozymes to catalyze and ligate RNA molecules encoding protein fragments, overcoming vector size constraints and facilitating the expression of large proteins, including therapeutic ones, by using 3' and 5' ribozymes for efficient protein assembly.
Patent Information
- Application Number
- JP2025553908
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-10-10
- Filing Date
- 2024-03-18
- Publication Date
- 2026-03-19
AI Technical Summary
The expression of full-length proteins is limited due to plasmid and vector size constraints, particularly in gene therapy settings where nucleic acids encoding such proteins exceed the packaging size of AAV, and proteins with multiple repeat sequences pose challenges for efficient expression.
A system and method utilizing nucleic acid molecules encoding RNA molecules with coding regions and ribozymes, specifically 3' and 5' ribozymes, to catalyze self-processing and ligate termini, enabling the assembly of RNA molecules encoding proteins, including those with multiple domains, through ribozyme-mediated RNA assembly and expression.
Enables the efficient expression of proteins exceeding 1000 amino acids, including therapeutic proteins, by overcoming size limitations and facilitating seamless ligation of protein fragments, thereby addressing the challenges of vector packaging and repeat sequences.
Smart Images

Figure 2026509506000001_ABST
Abstract
Description
Technical Field
[0001] Cross - reference to Related Applications This application claims priority based on U.S. Provisional Application No. 63 / 490,898 filed on March 17, 2023, U.S. Provisional Application No. 63 / 491,372 filed on March 21, 2023, U.S. Provisional Application No. 63 / 504,923 filed on May 30, 2023, and U.S. Provisional Application No. 63 / 589,217 filed on October 10, 2023, and each of these applications is hereby incorporated by reference in its entirety.
Background Art
[0002] In certain situations, the expression of full - length proteins is limited due to plasmid and vector size constraints. For example, in a therapeutic setting, some nucleic acids encoding full - length proteins exceed the packaging size of AAV, thereby limiting their applicability in gene therapy settings. Additionally, certain proteins that are biologically and industrially relevant contain multiple repeat sequences that can make expression difficult.
[0003] Therefore, there is a need in the art for improved compositions and methods for efficient protein expression. The present invention meets this unmet need.
Summary of the Invention
[0004] In one embodiment, the present invention relates to a system for generating an RNA molecule encoding a protein of interest, the system comprising a nucleic acid molecule encoding a first RNA molecule comprising a coding region encoding a first portion of the protein of interest and a 3' ribozyme, and a nucleic acid molecule encoding a second RNA molecule comprising a coding region encoding a second portion of the protein of interest and a 5' ribozyme.
[0005] In one embodiment, the 3' ribozyme catalyzes itself from a first RNA molecule, thereby generating a 3'P or 2'3'cP terminus. In one embodiment, the 5' ribozyme catalyzes itself from a second RNA molecule, thereby generating a 5'OH terminus. In one embodiment, the 3'P or 2'3'cP terminus is ligated to the 5'OH terminus to form an RNA molecule containing the coding regions of the first RNA molecule and the second RNA molecule. In one embodiment, the 3' ribozyme is a member of the HDV family of ribozymes. In one embodiment, the 5' ribozyme is a member of the HH family of ribozymes.
[0006] In one embodiment, the system further comprises one or more additional nucleic acid molecules encoding one or more additional RNA molecules, each additional RNA molecule comprising a coding region encoding a domain of the protein of interest, a 5' ribozyme, and a 3' ribozyme.
[0007] In one embodiment, the system further comprises one or more additional nucleic acid molecules encoding one or more additional RNA molecules, each additional RNA molecule comprising a coding region encoding a domain of the protein of interest, a 5' ribozyme, and a 3' ribozyme recognition sequence. In one embodiment, the system further comprises a ribozyme that interacts with the 3' ribozyme recognition sequence and induces the removal of the 3' recognition sequence. In one embodiment, the 3' ribozyme recognition sequence comprises VS-S, and the ribozyme is VS-Rz.
[0008] In one embodiment, the present invention relates to a method for generating an RNA molecule encoding a target protein, the method comprising administering a nucleic acid molecule encoding a first RNA molecule containing a coding region and a 3' ribozyme that encode a first portion of the target protein to a cell or tissue, and administering a nucleic acid molecule encoding a second RNA molecule containing a coding region and a 5' ribozyme that encodes a second portion of the target protein to a cell or tissue.
[0009] In one embodiment, the 3' ribozyme catalyzes itself from a first RNA molecule, thereby generating a 3'P or 2'3'cP terminus. In one embodiment, the 5' ribozyme catalyzes itself from a second RNA molecule, thereby generating a 5'OH terminus. In one embodiment, the 3'P or 2'3'cP terminus is ligated to the 5'OH terminus to form an RNA molecule containing the coding regions of the first RNA molecule and the second RNA molecule. In one embodiment, the 3' ribozyme is a member of the HDV family of ribozymes. In one embodiment, the 5' ribozyme is a member of the HH family of ribozymes.
[0010] In one embodiment, the method further comprises administering to a cell or tissue one or more additional nucleic acid molecules encoding one or more additional RNA molecules, each additional RNA molecule comprising a coding region encoding a domain of the protein of interest, a 5' ribozyme, and a 3' ribozyme.
[0011] In one embodiment, the method further comprises administering to a cell or tissue one or more additional nucleic acid molecules encoding one or more additional RNA molecules, each additional RNA molecule comprising a coding region encoding a domain of the protein of interest, a 5' ribozyme, and a 3' ribozyme recognition sequence.
[0012] In one embodiment, the method further comprises administering to cells or tissue a ribozyme that interacts with a 3' ribozyme recognition sequence and induces the removal of the 3' recognition sequence. In one embodiment, the 3' ribozyme recognition sequence comprises VS-S, and the ribozyme is VS-Rz. In one embodiment, the method further comprises administering to cells or tissue a ligase to induce the assembly of an RNA molecule. In one embodiment, the ligase is an RNA 2',3'-cyclic phosphate and 5'-OH(RtcB) ligase.
[0013] In one embodiment, the present invention includes an in vitro method for generating an RNA molecule encoding a target protein, the method being To provide a first RNA molecule containing a coding region that encodes the first portion of the target protein and a 3' ribozyme, To provide a second RNA molecule containing a coding region that encodes a second portion of the target protein and a 5' ribozyme, The present invention includes providing a ligase that induces the assembly of RNA molecules from the coding regions of a first RNA molecule and a second RNA molecule.
[0014] In one embodiment, the present invention relates to an in vitro method for generating an RNA molecule encoding a target repeat domain protein, the method comprising: a) providing a first RNA molecule comprising a coding region encoding a first portion of the target protein and a 3' ribozyme; b) providing one or more additional RNA molecules comprising a coding region encoding a domain of the target protein, a 5' ribozyme, and a 3' ribozyme recognition sequence; c) providing a ligase to ligate the coding region of the first RNA molecule with the coding regions of one or more additional RNA molecules; d) providing a ribozyme that recognizes the 3' ribozyme recognition sequence and catalyzes its removal; e) repeating steps b) to d) one or more times to generate RNA molecules encoding multiple repeat domains; f) providing a final RNA molecule comprising a coding region encoding the final portion of the target protein and a 5' ribozyme; and g) providing a ligase to ligate the coding regions of one or more additional RNA molecules with the coding region of the final RNA molecule, thereby generating a complete RNA molecule encoding a repeat domain protein.
[0015] In one embodiment, the present invention includes a method for treating a disease or disorder in a subject caused by a mutation of a large protein of interest, the method comprising administering to the subject a first nucleic acid molecule comprising a coding region encoding a first portion of the protein of interest and a 3' ribozyme, and administering to the subject a second nucleic acid molecule comprising a coding region encoding a second portion of the protein of interest and a 5' ribozyme.
[0016] In one embodiment, the disease or disorder is Duchenne muscular dystrophy, Becker muscular dystrophy (BMD), autosomal recessive polycystic kidney disease, hemophilia A, Stargard macular degeneration, limb-girdle muscular dystrophy, autosomal recessive severe congenital deafness, autosomal recessive non-syndromic hearing loss (ARNSHL), sensorineural hearing loss, cystic fibrosis, Wilson's disease, Miyoshi type myopathy, autosomal recessive hearing loss type 9 (DFNB9), Usher syndrome type I, GJB2-associated autosomal recessive non-syndromic hearing loss (GJB2-AR NSHL), Autosomal recessive cerebellar parenchymal disorder type 3, Non-syndromic hearing loss, Autosomal recessive hearing loss type 16 (DFNB16), Meniere's disease, Autosomal dominant non-syndromic sensorineural hearing loss type 12 (DFNA12), Autosomal recessive spinocerebellar ataxia type 21 (SCAR21), Usher syndrome type 1F (USH1F), Autosomal recessive hearing loss type 23 (DFNB23), Autosomal recessive hearing loss type 30 (DFNB30), Oto-spine-megaly epiphysis dysplasia (OSMED), Autosomal recessive hearing loss type 77 (DFNB77), Autosomal recessive hearing loss type 84A (DFNB84A), Autosomal recessive hearing loss type 84B (DFNB84B), Peripheral neuropathy These include harm, myopathy, autosomal dominant non-syndromic hearing loss type 4A (DFNA4), congenital thrombocytopenia, sensorineural hearing loss, autosomal dominant non-syndromic hearing loss type 56 (DFNA56), epileptic encephalopathy, Timothy syndrome, long QT syndrome, X-linked retinal disease, aldosteronism, autosomal recessive hearing loss type 42 (DFNB42), primary aldosteronism (Conn syndrome), seizures, neurological abnormalities, sinoatrial node dysfunction, neurodevelopmental disorders, hypokalemic periodic paralysis, epilepsy, developmental epileptic encephalopathy, Brodymyopathy, Darier's disease, heart disease, von Willebrand disease, and Zellweger syndrome.
[0017] In one embodiment, the present invention includes a system for generating an RNA molecule encoding a target protein, and a circular RNA molecule comprising a synthetic intron including a first portion of the target protein, a 5' ribozyme, a cargo sequence, and a 3' ribozyme, as well as a nucleic acid encoding a second portion of the target protein.
[0018] In one embodiment, the target protein is one or more proteins selected from the group consisting of therapeutic proteins, reporter proteins, and Cas9 proteins.
[0019] In one embodiment, the cargo sequence is one or more selected from the group consisting of a sequence encoding a therapeutic protein of interest, a CRISPR guide RNA sequence, a small RNA sequence, and a trans-cleaved ribozyme sequence. In one embodiment, the small RNA sequence includes one or more selected from the group consisting of microRNA (miRNA), Piwi-interacting RNA (piRNA), small interfering RNA (siRNA), small nucleolar RNA (snoRNA), small tRNA-derived RNA (tsRNA), small rDNA-derived RNA (srRNA), and small nuclear RNA (snRNA).
[0020] In one embodiment, the 3' ribozyme of the synthetic intron is a member of the HH family of ribozymes. In one embodiment, the 5' ribozyme of the synthetic intron is one or more selected from the group consisting of members of the HDV family of ribozymes, members of the HDV family of ribozymes, and VS-S ribozyme recognition sequences. In one embodiment, the system further comprises one or more selected from the group consisting of RtcB ligase and nucleic acids encoding RtcB ligase.
[0021] In one embodiment, the present invention is a system for generating an RNA molecule encoding a target protein, wherein the system is a) A first RNA molecule containing a coding region that encodes the first portion of the target protein, which is directly linked at the 3' end to the nucleotide sequence of a ribozyme (3' ribozyme), b) A second RNA molecule comprising a coding region that encodes a second portion of the target protein, which is directly ligated at the 5' end to the nucleotide sequence of a ribozyme (5' ribozyme), Here, the autocleavage of the 3' ribozyme generates a 3' phosphate group at the 3' end of the first portion of the target protein, and the autocleavage of the 5' ribozyme generates a 5' hydroxyl group at the 5' end of the second portion of the target protein, enabling scarless ligation of the first and second RNA molecules and generating the RNA molecule encoding the target protein.
[0022] In one embodiment, at least one of the 3'-ribozyme and 5'-ribozyme is SEQ ID NO: 131, SEQ ID NO: 132, SEQ ID NO: 133, SEQ ID NO: 134, SEQ ID NO: 135, SEQ ID NO: 136, SEQ ID NO: 137, SEQ ID NO: 138, SEQ ID NO: 139, SEQ ID NO: 140, SEQ ID NO: 141, SEQ ID NO: 142, SEQ ID NO: 143, SEQ ID NO: 144, SEQ ID NO: 145, SEQ ID NO: 166, SEQ ID NO: 167, SEQ ID NO: 168, SEQ ID NO: 169, SEQ ID NO: 170, SEQ ID NO: 171, SEQ ID NO: 172, SEQ ID NO: 173, SEQ ID NO: 174, SEQ ID NO: 175, SEQ ID NO: 176, SEQ ID NO: 192, SEQ ID NO: 193, SEQ ID NO: 194, SEQ ID NO: 195, SEQ ID NO: 196, SEQ ID NO: 197, SEQ ID NO: 198, SEQ ID NO: 199, SEQ ID NO: 200, SEQ ID NO: 202, SEQ ID NO: 203, SEQ ID NO: 204, SEQ ID NO: 205, SEQ ID NO: 206, SEQ ID NO: 207, SEQ ID NO: 217, SEQ ID NO: 218, or SEQ ID NO: 219.
[0023] In one embodiment, the length of the target protein is longer than 1000 amino acid residues.
[0024] In one embodiment, the target protein is DMD, PKHD1, F8, ABCA4, DYSF, OTOF, CFTR, ATP7B, MYOF, MYO7A, MYO15A, CDH23, STRC, OTOG, TECTA, PCDH15, TRIOBP, MYO3A, COL11A2, LOXHD1, PTPRQ, OTOGL, MYH14, MYH9, TNC, CACNA1A, CACNA1C, CACNA1F, CACNA1H, CACNA1G, CACNA1D, CACNA1B, CACNA1S, CACNA1I, CACNA1E, ATP2A1, ATP2A2, VWF, PEX1, or CMYA5.
[0025] In one embodiment, the present invention provides a method for treating a disease or disorder of interest caused by a mutation of a target protein, comprising administering to a subject a first nucleic acid molecule comprising a coding region encoding a first portion of the target protein directly linked at its 3'-end to a nucleotide sequence of a ribozyme (3'-ribozyme), and administering to the subject a second nucleic acid molecule comprising a coding region encoding a second portion of the target protein directly linked at its 5'-end to a nucleotide sequence of a ribozyme (5'-ribozyme), wherein self-cleavage of the 3'-ribozyme generates a 3'-phosphate at the 3'-end of the first portion of the target protein, and self-cleavage of the 5'-ribozyme generates a 5'-hydroxy group at the 5'-end of the second portion of the target protein, thereby enabling seamless ligation of the target protein.
[0026] In one embodiment, at least one of the 3'-ribozyme and the 5'-ribozyme has a sequence selected from SEQ ID NO: 131, SEQ ID NO: 132, SEQ ID NO: 133, SEQ ID NO: 134, SEQ ID NO: 135, SEQ ID NO: 136, SEQ ID NO: 137, SEQ ID NO: 138, SEQ ID NO: 139, SEQ ID NO: 140, SEQ ID NO: 141, SEQ ID NO: 142, SEQ ID NO: 143, SEQ ID NO: 144, SEQ ID NO: 145, SEQ ID NO: 166, SEQ ID NO: 167, SEQ ID NO: 168, SEQ ID NO: 169, SEQ ID NO: 170, SEQ ID NO: 171, SEQ ID NO: 172, SEQ ID NO: 173, SEQ ID NO: 174, SEQ ID NO: 175, SEQ ID NO: 176, SEQ ID NO: 192, SEQ ID NO: 193, SEQ ID NO: 194, SEQ ID NO: 195, SEQ ID NO: 196, SEQ ID NO: 197, SEQ ID NO: 198, SEQ ID NO: 199, SEQ ID NO: 200, SEQ ID NO: 202, SEQ ID NO: 203, SEQ ID NO: 204, SEQ ID NO: 205, SEQ ID NO: 206, SEQ ID NO: 207, SEQ ID NO: 217, SEQ ID NO: 218 or SEQ ID NO: 219.
[0027] In one embodiment, the disease or disorder is Duchenne muscular dystrophy, Becker muscular dystrophy (BMD), autosomal recessive polycystic kidney disease, hemophilia A, Stargard macular degeneration, limb-girdle muscular dystrophy, autosomal recessive severe congenital deafness, autosomal recessive non-syndromic hearing loss (ARNSHL), sensorineural hearing loss, cystic fibrosis, Wilson's disease, Miyoshi type myopathy, autosomal recessive hearing loss type 9 (DFNB9), Usher syndrome type I, GJB2-associated autosomal recessive non-syndromic hearing loss (GJB2-AR NSHL), Autosomal recessive cerebellar parenchymal disorder type 3, Non-syndromic hearing loss, Autosomal recessive hearing loss type 16 (DFNB16), Meniere's disease, Autosomal dominant non-syndromic sensorineural hearing loss type 12 (DFNA12), Autosomal recessive spinocerebellar ataxia type 21 (SCAR21), Usher syndrome type 1F (USH1F), Autosomal recessive hearing loss type 23 (DFNB23), Autosomal recessive hearing loss type 30 (DFNB30), Oto-spine-megaly epiphysis dysplasia (OSMED), Autosomal recessive hearing loss type 77 (DFNB77), Autosomal recessive hearing loss type 84A (DFNB84A), Autosomal recessive hearing loss type 84B (DFNB84B), Peripheral neuropathy These include harm, myopathy, autosomal dominant non-syndromic hearing loss type 4A (DFNA4), congenital thrombocytopenia, sensorineural hearing loss, autosomal dominant non-syndromic hearing loss type 56 (DFNA56), epileptic encephalopathy, Timothy syndrome, long QT syndrome, X-linked retinal disease, aldosteronism, autosomal recessive hearing loss type 42 (DFNB42), primary aldosteronism (Conn syndrome), seizures, neurological abnormalities, sinoatrial node dysfunction, neurodevelopmental disorders, hypokalemic periodic paralysis, epilepsy, developmental epileptic encephalopathy, Brodymyopathy, Darier's disease, heart disease, von Willebrand disease, and Zellweger syndrome.
[0028] In one embodiment, the target protein is DMD, PKHD1, F8, ABCA4, DYSF, OTOF, CFTR, ATP7B, MYOF, MYO7A, MYO15A, CDH23, STRC, OTOG, TECTA, PCDH15, TRIOBP, MYO3A, COL11A2, LOXHD1, PTPRQ, OTOGL, MYH14, MYH9, TNC, CACNA1A, CACNA1C, CACNA1F, CACNA1H, CACNA1G, CACNA1D, CACNA1B, CACNA1S, CACNA1I, CACNA1E, ATP2A1, ATP2A2, VWF, PEX1, or CMYA5.
[0029] In one embodiment, the first nucleic acid molecule comprises a nucleic acid sequence encoding an N-terminal fragment of a dystrophin, minidystrophin or microdystrophin protein, the second nucleic acid molecule comprises a nucleic acid sequence encoding a C-terminal fragment of a dystrophin, minidystrophin or microdystrophin protein, and administration of the first and second nucleic acid molecules results in the production of a full-length dystrophin, minidystrophin or microdystrophin protein. In one embodiment, the first nucleic acid comprises the nucleic acid sequence of SEQ ID NO: 150 and the second nucleic acid comprises the nucleic acid sequence of SEQ ID NO: 151. In one embodiment, the first nucleic acid comprises the nucleic acid sequence of SEQ ID NO: 152 and the second nucleic acid comprises the nucleic acid sequence of SEQ ID NO: 153.
[0030] In one embodiment, the first nucleic acid molecule includes a nucleic acid sequence encoding the N-terminal fragment of dysferlin, and the second nucleic acid molecule includes a nucleic acid sequence encoding the C-terminal fragment of dysferlin, and administration of the first and second nucleic acid molecules produces full-length dysferlin protein. In one embodiment, the first nucleic acid includes the nucleic acid sequence of SEQ ID NO: 157, and the second nucleic acid includes the nucleic acid sequence of SEQ ID NO: 158. In one embodiment, the first nucleic acid includes the nucleic acid sequence of SEQ ID NO: 177, and the second nucleic acid includes the nucleic acid sequence of SEQ ID NO: 178. In one embodiment, the first nucleic acid includes the nucleic acid sequence of SEQ ID NO: 179, and the second nucleic acid includes the nucleic acid sequence of SEQ ID NO: 178. In one embodiment, the first nucleic acid includes the nucleic acid sequence of SEQ ID NO: 181, and the second nucleic acid includes the nucleic acid sequence of SEQ ID NO: 182. In one embodiment, the first nucleic acid includes the nucleic acid sequence of SEQ ID NO: 183, and the second nucleic acid includes the nucleic acid sequence of SEQ ID NO: 182.
[0031] In one embodiment, the first nucleic acid molecule contains a nucleic acid sequence encoding the N-terminal fragment of STRC, and the second nucleic acid molecule contains a nucleic acid sequence encoding the C-terminal fragment of STRC, and the administration of the first and second nucleic acid molecules produces the full-length STRC protein. In one embodiment, the first nucleic acid contains the nucleic acid sequence of SEQ ID NO: 160, and the second nucleic acid contains the nucleic acid sequence of SEQ ID NO: 161.
[0032] In one embodiment, the present invention is a method for generating an RNA molecule encoding a target protein, wherein the method is The procedure comprises administering to cells or tissue a first RNA molecule containing a coding region that encodes a first portion of the target protein, directly ligated at the 3' end to the nucleotide sequence of a ribozyme (3' ribozyme), and administering to cells or tissue a second RNA molecule containing a coding region that encodes a second portion of the target protein, directly ligated at the 5' end to the nucleotide sequence of a ribozyme (5' ribozyme), Here, the autocleavage of the 3' ribozyme generates a 3' phosphate group at the 3' end of the first portion of the target protein, and the autocleavage of the 5' ribozyme generates a 5' hydroxyl group at the 5' end of the second portion of the target protein, enabling scarless ligation of the first and second RNA molecules and generating the RNA molecule encoding the target protein.
[0033] In one embodiment, at least one of the 3'-ribozyme and 5'-ribozyme is SEQ ID NO: 131, SEQ ID NO: 132, SEQ ID NO: 133, SEQ ID NO: 134, SEQ ID NO: 135, SEQ ID NO: 136, SEQ ID NO: 137, SEQ ID NO: 138, SEQ ID NO: 139, SEQ ID NO: 140, SEQ ID NO: 141, SEQ ID NO: 142, SEQ ID NO: 143, SEQ ID NO: 144, SEQ ID NO: 145, SEQ ID NO: 166, SEQ ID NO: 167, SEQ ID NO: 168, SEQ ID NO: 169, SEQ ID NO: 170, SEQ ID NO: 171, SEQ ID NO: 172, SEQ ID NO: 173, SEQ ID NO: 174, SEQ ID NO: 175, SEQ ID NO: 176, SEQ ID NO: 192, SEQ ID NO: 193, SEQ ID NO: 194, SEQ ID NO: 195, SEQ ID NO: 196, SEQ ID NO: 197, SEQ ID NO: 198, SEQ ID NO: 199, SEQ ID NO: 200, SEQ ID NO: 202, SEQ ID NO: 203, SEQ ID NO: 204, SEQ ID NO: 205, SEQ ID NO: 206, SEQ ID NO: 207, SEQ ID NO: 217, SEQ ID NO: 218, or SEQ ID NO: 219.
[0034] In one embodiment, the total length of the target protein is longer than 1000 amino acid residues. In one embodiment, the target protein is selected from the group consisting of DMD, PKHD1, F8, ABCA4, DYSF, OTOF, CFTR, ATP7B, MYOF, MYO7A, MYO15A, CDH23, STRC, OTOG, TECTA, PCDH15, TRIOBP, MYO3A, COL11A2, LOXHD1, PTPRQ, OTOGL, MYH14, MYH9, TNC, CACNA1A, CACNA1C, CACNA1F, CACNA1H, CACNA1G, CACNA1D, CACNA1B, CACNA1S, CACNA1I, CACNA1E, ATP2A1, ATP2A2, VWF, PEX1, and CMYA5.
[0035] In one embodiment, the present invention includes an in vitro method for generating an RNA molecule encoding a target protein, the method comprising: providing a first RNA molecule comprising a coding region encoding a first portion of the target protein, which is directly ligated at the 3' end to the nucleotide sequence of a ribozyme (3' ribozyme); providing a second RNA molecule comprising a coding region encoding a second portion of the target protein, which is directly ligated at the 5' end to the nucleotide sequence of a ribozyme (5' ribozyme), wherein the 3' ribozyme self-cleavage generates a 3' phosphate group at the 3' end of the first portion of the target protein, and the 5' ribozyme self-cleavage generates a 5' hydroxyl group at the 5' end of the second portion of the target protein, thereby enabling scarless ligation of the first and second RNA molecules to generate an RNA molecule encoding the target protein; and providing a ligase that induces the assembly of RNA molecules from the coding regions of the first RNA molecule and the second RNA molecule.
[0036] In one embodiment, at least one of the 3'-ribozyme and 5'-ribozyme is SEQ ID NO: 131, SEQ ID NO: 132, SEQ ID NO: 133, SEQ ID NO: 134, SEQ ID NO: 135, SEQ ID NO: 136, SEQ ID NO: 137, SEQ ID NO: 138, SEQ ID NO: 139, SEQ ID NO: 140, SEQ ID NO: 141, SEQ ID NO: 142, SEQ ID NO: 143, SEQ ID NO: 144, SEQ ID NO: 145, SEQ ID NO: 166, SEQ ID NO: 167, SEQ ID NO: 168, SEQ ID NO: 169, SEQ ID NO: 170, SEQ ID NO: 171, SEQ ID NO: 172, SEQ ID NO: 173, SEQ ID NO: 174, SEQ ID NO: 175, SEQ ID NO: 176, SEQ ID NO: 192, SEQ ID NO: 193, SEQ ID NO: 194, SEQ ID NO: 195, SEQ ID NO: 196, SEQ ID NO: 197, SEQ ID NO: 198, SEQ ID NO: 199, SEQ ID NO: 200, SEQ ID NO: 202, SEQ ID NO: 203, SEQ ID NO: 204, SEQ ID NO: 205, SEQ ID NO: 206, SEQ ID NO: 207, SEQ ID NO: 217, SEQ ID NO: 218, or SEQ ID NO: 219.
[0037] In one embodiment, the present invention provides a system for generating a circular RNA molecule encoding a target protein, the system comprising a 5' ribozyme, an IRES sequence, and a nucleic acid encoding a second portion of the target protein, directly linked to a 3' ribozyme, wherein self-cleavage of the 3' ribozyme generates a 3' phosphate group at the 3' end of the second portion of the target protein, and self-cleavage of the 5' ribozyme generates a 5' hydroxyl group at the 5' end of the first portion of the target protein, thereby enabling scarless ligation of the 5' hydroxyl group and 3' phosphate group of the RNA molecule, and generating a circular RNA molecule encoding the target protein.
[0038] In one embodiment, the RNA molecule includes an in vitro transcription RNA molecule.
[0039] In one embodiment, the target protein is a therapeutic protein, a reporter protein, or a Cas9 protein.
[0040] In one embodiment, at least one of the 3'-ribozyme and 5'-ribozyme is SEQ ID NO: 131, SEQ ID NO: 132, SEQ ID NO: 133, SEQ ID NO: 134, SEQ ID NO: 135, SEQ ID NO: 136, SEQ ID NO: 137, SEQ ID NO: 138, SEQ ID NO: 139, SEQ ID NO: 140, SEQ ID NO: 141, SEQ ID NO: 142, SEQ ID NO: 143, SEQ ID NO: 144, SEQ ID NO: 145, SEQ ID NO: 166, SEQ ID NO: 167, SEQ ID NO: 168, SEQ ID NO: 169, SEQ ID NO: 170, SEQ ID NO: 171, SEQ ID NO: 172, SEQ ID NO: 173, SEQ ID NO: 174, SEQ ID NO: 175, SEQ ID NO: 176, SEQ ID NO: 192, SEQ ID NO: 193, SEQ ID NO: 194, SEQ ID NO: 195, SEQ ID NO: 196, SEQ ID NO: 197, SEQ ID NO: 198, SEQ ID NO: 199, SEQ ID NO: 200, SEQ ID NO: 202, SEQ ID NO: 203, SEQ ID NO: 204, SEQ ID NO: 205, SEQ ID NO: 206, SEQ ID NO: 207, SEQ ID NO: 217, SEQ ID NO: 218, or SEQ ID NO: 219.
[0041] In one embodiment, the present invention relates to a method for generating a circular RNA molecule in vivo, the method comprising administering a linear in vitro transcription RNA molecule or a DNA molecule encoding a linear RNA molecule to a cell or tissue, wherein the linear RNA molecule comprises a 5' ribozyme, an IRES sequence, and a second portion of the protein of interest, directly ligated to a 3' ribozyme, wherein autocleavage of the 3' ribozyme generates a 3' phosphate group at the 3' end of the second portion of the protein of interest, and autocleavage of the 5' ribozyme generates a 5' hydroxyl group at the 5' end of the first portion of the protein of interest, thereby enabling seamless ligation of the 5' hydroxyl group and 3' phosphate group of the RNA molecule, thereby generating a circular RNA molecule encoding the protein of interest.
[0042] In one embodiment, the target protein is a therapeutic protein, a reporter protein, or a Cas9 protein.
[0043] In one embodiment, at least one of the 3'-ribozyme and 5'-ribozyme is SEQ ID NO: 131, SEQ ID NO: 132, SEQ ID NO: 133, SEQ ID NO: 134, SEQ ID NO: 135, SEQ ID NO: 136, SEQ ID NO: 137, SEQ ID NO: 138, SEQ ID NO: 139, SEQ ID NO: 140, SEQ ID NO: 141, SEQ ID NO: 142, SEQ ID NO: 143, SEQ ID NO: 144, SEQ ID NO: 145, SEQ ID NO: 166, SEQ ID NO: 167, SEQ ID NO: 168, SEQ ID NO: 169, SEQ ID NO: 170, SEQ ID NO: 171, SEQ ID NO: 172, SEQ ID NO: 173, SEQ ID NO: 174, SEQ ID NO: 175, SEQ ID NO: 176, SEQ ID NO: 192, SEQ ID NO: 193, SEQ ID NO: 194, SEQ ID NO: 195, SEQ ID NO: 196, SEQ ID NO: 197, SEQ ID NO: 198, SEQ ID NO: 199, SEQ ID NO: 200, SEQ ID NO: 202, SEQ ID NO: 203, SEQ ID NO: 204, SEQ ID NO: 205, SEQ ID NO: 206, SEQ ID NO: 207, SEQ ID NO: 217, SEQ ID NO: 218, or SEQ ID NO: 219.
[0044] The following detailed description of embodiments of the present invention will be better understood when read in conjunction with the accompanying drawings. It should be understood that the present invention is not limited to the exact arrangement and means of the embodiments shown in the drawings. [Brief explanation of the drawing]
[0045] [Figure 1A] Figures 1A–1E show data demonstrating ribozyme-mediated transsplicing and expression in mammalian cells. The figures show vectors encoding the N-terminal (Nt) half of GFP containing the 3'HDV ribozyme and the C-terminal (Ct) half of GFP containing the 5' hammerhead (HH) ribozyme. [Figure 1B] Exemplary results show that co-expression of both Nt-GFP-HDV and HH-Ct-GFP in COS7 cells and HEK293T cells resulted in detectable GFP fluorescence, whereas transfection with each expressed separately did not. [Figure 1C-1D] Exemplary results of RT-PCR amplification (Figure 1C) and Sanger sequencing analysis (Figure 1D) using primers specific to each independent RNA (G1 and G2) are shown, demonstrating ribozyme removal, scarless splicing, and recovery of the GFP coding sequence. [Figure 1E] Exemplary Western blot results using a GFP-specific antibody, showing the predicted full-length protein size for GFP, are presented. [Figure 2A] Figures 2A-2E show data demonstrating the expression of a luciferase-based reporter to quantify the effect of ribozyme sequences on trans-splicing in mammalian cells. The figures show vectors encoding the N-terminal (Nt) half of luciferase containing the 3'HDV ribozyme and the C-terminal (Ct) half of luciferase containing the 5' hammerhead (HH) ribozyme. [Figure 2B-2C]Exemplary results of RT-PCR amplification (Figure 2B) and Sanger sequencing analysis (Figure 2C) using primers specific to each independent LucRNA (L1 and L2) are shown, demonstrating ribozyme removal and scarless splicing of the luciferase open reading frame. [Figure 2D-2E] We demonstrate the effects of different HDV (Figure 2D) and HH (Figure 2E) ribozyme sequences on transsplicing in mammalian cells. Furthermore, mutations in ribozyme-catalyzed nucleotides resulted in loss of luciferase activity (last column). [Figure 3A] Figures 3A-3D show data demonstrating the regulation of protein expression from Nt and Ct vectors. The diagrams show the arrangement of the C-terminal proteolytic sequence that inhibits the expression of proteins encoded by the Nt vector. [Figure 3B] This paper presents exemplary results demonstrating the efficiency of various proteolytic sequences in inhibiting GFP-HDV expression from Nt vectors encoding full-length GFP. [Figure 3C] The diagram shows the arrangement of the N-terminal translational regulatory sequence to prevent the translation of protein sequences in a Ct vector. [Figure 3D] This paper presents exemplary results demonstrating the efficiency of different GFP sequence modifications or translational regulatory sequences in preventing GFP fluorescence in mammalian cells. [Figure 4A] Figures 4A–4D show data demonstrating single and multiple ribozyme-mediated transsplicing in mammalian cells. The figures show vectors encoding 4xMTS and full-length GFP (without the start ATG codon) containing ribozymes that mediate transsplicing and expression of mitochondrial-targeted GFP protein. [Figure 4B] The co-expression of these vectors yields exemplary results demonstrating that it results in mitochondrial-localized green fluorescence that overlaps with the red fluorescence of the mitochondrial tracker CMXRos. [Figure 4C]The figure shows vectors for multiple transsplicing and expression of mitochondrial-targeted GFP protein (4xMTS-GFP) in reading frame 1 and myristoylated membrane-targeted red fluorescent protein (F2-Myr-RFP) in reading frame 2. [Figure 4D] This example demonstrates that the simultaneous expression of all four vectors in mammalian Cos7 cells results in specific green fluorescence in mitochondria and red fluorescence in the membrane. [Figure 5A] Figures 5A and 5B show data demonstrating the enhancement of ribozyme-mediated trans-splicing using optimized ribozyme sequences as well as cis-splicing splice acceptor and splice donor sequences. The diagrams show the arrangement of chimeric splice donor (SD) and splice acceptor (SA) sequences in typical Nt-GFP-3'Rz and 5'Rz-Ct-GFP trans-splicing GFP reporters, where Rz represents a cis-cleaved ribozyme. [Figure 5B] Exemplary results of GFP fluorescence in Cos7 cells 18 hours post-transfection (first three columns) or 36 hours post-transfection (last column) after single-vector transfection (first two columns) or after simultaneous transfection (last two columns) are shown. The first column shows the use of unoptimized HH and HDV ribozymes, the second column shows the use of optimized Twister and RzB ribozymes, and the last column shows Twister and RzB ribozymes as well as SD and SA sequence combinations. [Figure 6A] Figures 6A–6D show data demonstrating ribozyme-mediated trans-splicing of large protein-coding genes. A figure shows a vector encoding a split μ-dystrophin-GFP fusion protein for delivery using an AAV vector. [Figure 6B-6C]Exemplary results of RT-PCR (Figure 6B) and Sanger sequencing analysis (Figure 6C) of cells transfected with Nt-Dys and Ct-Dys vectors, demonstrating specific transsplicing, are shown. [Figure 6D] Exemplary results of GFP fluorescence from cells transfected with both Nt and Ct dystrophin vectors, imaged using confocal microscopy to demonstrate the predicted membrane localization of dystrophin, are shown. [Figure 7A] Figures 7A–7C show data demonstrating lentiviral delivery of ribozyme-containing RNA for trans-splicing in target cells. Figures also show negative sense orientation of Nt and Ct split GFP expression cassettes in lentiviral gene transvestments. [Figure 7B] This example shows that only cells co-transfected with lentiviruses encoding both the Nt-GFP and Ct-GFP genes exhibit GFP fluorescence. [Figure 7C] The figure shows the negative sense orientation of the Nt and Ct split Dys expression cassette in a lentiviral gene transfer vector. [Figure 8A] Figures 8A and 8B show data demonstrating ribozyme-mediated trans-splicing and toxic DTA gene expression. Figures also show vectors encoding the split Nt and Ct DTA genes. [Figure 8B] We present exemplary results demonstrating that cells co-transfected with both Nt-DTA and Ct-DTA result in decreased expression of the co-transfected GFP reporter, consistent with the translational repressor function of DTA in mammalian cells. [Figure 9] We present exemplary results demonstrating that co-expression of exogenous RNA regulatory enzymes can enhance or inhibit ribozyme-mediated transsplicing in mammalian cells. [Figure 10A]Figures 10A–10D show data demonstrating that RtcB is sufficient to catalyze ribozyme-mediated trans-splicing in vitro. The figures show a trans-splicing reporter for a split ciferase containing an upstream T7 RNA promoter that enables in vitro RNA transcription. [Figure 10B] Exemplary RT-PCR results demonstrating that in vitro transspliced luciferase RNA is dependent on the addition of RtcB protein (NEB) using the manufacturer's recommended reaction conditions are shown. [Figure 10C] The figure shows transsplicing vectors for the conserved N-terminal (N1L) and C-terminal (N3R) domains of spidoin. [Figure 10D] We present data demonstrating that RtcB is sufficient to catalyze ribozyme-mediated transsplicing in vitro. Exemplary Sanger sequencing results are shown demonstrating that RtcB ligase from *E. coli* was sufficient to catalyze the transligation of ribozyme-cleaved RNA encoding N1L and N3R. [Figure 11] This figure shows in vitro directed ligation of ribozyme-catalyzed RNA using RtcB, VS-S, and VS-Rz. [Figure 12A] Figures 12A-12D show data demonstrating the use of trans-cleaved ribozymes for RNA trans-splicing. The secondary structures of cis-cleaved ribozymes are shown. [Figure 12B] This shows a manipulated ribozyme that can be cleaved by a transformer. [Figures 12C-12D] Data demonstrating the use of trans-cleaved ribozymes for RNA trans-splicing are presented. Figures demonstrating the potential application of trans-cleaved ribozymes to delete disease-causing mutations such as frameshifts or immature stop codons in order to restore protein expression and function are also presented. [Figure 13A]Figures 13A and 13B show data illustrating the secondary structures of representative ribozymes that can be used for scarless 5' cleavage of RNA. Representative ribozymes that can be used for scarless 5' cleavage are shown. [Figure 13B] This diagram shows representative ribozymes that can be used for scarless 3' cleavage. N = any nucleotide. The red scissors define the cleavage site. Red nucleotides indicate catalytic mutations. Orange nucleotides represent the RNA sequence to be transspliced. Dark blue nucleotides indicate the ribozyme sequence required to form the stem. Light blue indicates the stem 1 tertiary stabilization motif (TSM) that interacts with the stem 2 loop. HH - Hammerhead, HDV - Hepatitis delta virus, Rz - Ribozyme. (HH: Gao and Zhao, 2013, JIPB; HDV: Schurer et al., 2002, Nuc Acids Res; RzB: Saksmerprome et al., 2004, RNA; Twister: Liu et al., 2014, Nat Chem Bio) [Figure 14A] Figures 14A–14C show data demonstrating scarless cleavage and inducible RNA trans-splicing, as well as expression by trans-activated ribozymes. The figures show that the VS ribozyme can be split into two components: a small VS-S stem-loop lacking autocatalytic activity and a larger VS-Rz that induces VS-S cleavage when delivered trans. The VS-S / VS-Rz ribozyme pair can be used to generate inducible scarless splicing. [Figure 14B] This figure illustrates a method utilizing a VS-S / VS-Rz trans-activated ribozyme pair to generate an inducible RNA trans-splicing system. Does the Nt-GFP-VS-S RNA generate a suitable RNA end that can participate in trans-splicing with co-expressed Ct-GFP RNA only upon delivery or expression of VS-Rz? [Figure 14C]This figure shows a method for generating RNA having an N-terminal sequence, a variable or non-variable repeat region, and a C-terminal sequence. The "repeat" RNA contains a 5' autocatalytic ribozyme and a 3' trans-activating ribozyme (such as VS-S), which enables controlled repeat addition that depends on the selective addition of a trans-activating VS-Rz and ligase (such as RtcB). [Figure 15A] Figures 15A–15E show data demonstrating ribozyme-mediated transsplicing with the generation of stable intronic RNA sequences. Figures also show the use of cis-cleaved ribozymes to mediate transsplicing of two independent RNAs. [Figure 15B] The diagram shows the use of internal cis-cleavage ribozymes to create synthetic introns. [Figure 15C] Exemplary results are shown demonstrating efficient cis-cleavage of synthetic introns and independent RNA trans-splicing to generate a functional protein (GFP). [Figures 15D-15E] The figure shows the use of internal cis-cleavage ribozymes to generate trans-spliced and translated reporter and intron sequence "cargos" that can be any useful RNA sequence or gene expression cassette. [Figure 16A] Figures 16A–16C show exemplary results for optimized ribozyme sequences for ribozyme-mediated trans-splicing in vivo. A comparison of relative ribozyme activity using a luciferase trans-splicing reporter is shown. RzB hammerhead ribozyme variants containing a tertiary stabilization motif active at low magnesium concentrations exhibited maximum luciferase activity in mammalian cells. [Figure 16B] This report compares HDV ribozymes (HDV68 and GenomicHDV) with the Twister ribozyme (Twst). The Twister ribozyme on the 3' end of Nt-Luc provided maximum luciferase activity, which was lost with the catalytic inactivation mutation (Twst mut). [Figure 16C]This shows a comparison of Twister ribozyme sequence modifications. Shortening of the P1 stem reduced reporter activity. Modification of the first residue revealed that Twister can tolerate the A nucleotide (U1A) at position 1. [Figure 17A] Figures 17A–17C show the identification of optimal ribozyme pairs for StitchR-mediated RNA transsplicing in mammalian cells. We show the relative activity of representative ribozyme sequences from each of the major ribozyme families, including Twister, Twister Sister, Hammerhead, HDV, Pistol, Varkud Satellite (VS), Hairpin, and Hovlinc (Hov), screened in mammalian cells using a StitchR-responsive luciferase-based reporter. [Figure 17B] This report describes the screening of the relative activity of representative ribozyme sequences from each of the major ribozyme families, including Twister, Twister Sister, Hammerhead, HDV, Pistol, Varkud Satellite (VS), Hairpin, and Hovlinc (Hov), in mammalian cells using a StitchR-compatible luciferase-based reporter. The inclusion of splice donor (SD) and splice acceptor (SA) sequences enabled the generation of functional introns in the sutured mRNA, allowing for the restoration of a complete luciferase open reading frame after processing. The highest luciferase expression was found in both the Luc Nt and Luc Ct vectors using the Twister ribozyme, which closely approximated the expression observed from vectors encoding luciferase in a single open reading frame (ORF). [Figure 17C] We demonstrate that we assayed different subtypes of Twister ribozymes to determine their relative activity in promoting mRNA stitchr activity and expression in mammalian cells. [Figure 18A]Figures 18A–18C show data demonstrating the application of StitchR for dual AAV gene delivery and expression in mammalian cells. The design of a StitchR-compatible split GFP reporter cassette, subcloned into a vector adjacent to the AAV2 ITR sequence, under the control of a human CMV (hCMV) promoter and a bovine growth hormone polyadenylated sequence (bGH pA), is shown. [Figures 18B-18C] Robust full-length GFP expression, detectable by epifluorescence (Figure 18B) and Western blotting (Figure 18C), was obtained only when AAV2 / 1 serotype viruses were generated using the construct and simultaneously transduced into human cells (HEK293T). [Figures 19A-19B] Figures 19A–19C show data demonstrating the application of StitchR for in vivo dual AAV gene delivery and expression. StitchR-compatible dual AAV GFP reporter virus (2E+ 12vg / ml) was injected into 10-day-old (P10) mice and imaged 2 months after injection using IVIS fluorescence-based whole-animal imaging (Figure 19A) and epifluorescence (Figure 19B). Strong GFP fluorescence was detected throughout the body and readily observed in hindlimb muscle tissue. [Figure 19C] The data shows that full-length GFP protein expression was observed only in mice injected with both GFPnt and GFPct viruses, as detected by Western blotting using an anti-GFP antibody. [Figure 20] This shows a subset of human loss-of-function monogenetic disorders resulting from large gene mutations. The numbers in parentheses represent the amino acid sequence lengths of adjacent disease genes. [Figure 21]The data shows that mutations in the large dystrophin (Dys) protein can cause severe Duchenne muscular dystrophy (DMD) or mild Becker muscular dystrophy (BMD). The figure shows characterized human dystrophin (Dys) protein domains, key protein interaction sites, sequence deletions found in mild Becker muscular dystrophy (BMD), and engineered therapeutic dystrophins that can be fitted into either a single AAV (microDys) or a dual AAV vector pair (Stitch® miniDys). [Figure 22] The figure shows the microDys AAV expression vector regulated by the cardiac and skeletal muscle-specific CK8e promoter and bGH pA. [Figure 23] The figure shows a dual StitchR-compatible N-terminal (StitchR Dys-Nt) and C-terminal (StitchR Dys-Ct) AAV expression vector controlled by a cardiac and skeletal muscle-specific CK8e promoter and bGH pA. [Figure 24] The figure shows the N-terminal (Dual AKDys-Nt) and C-terminal (Dual AKDys-Ct) expression vectors of a dual AAV vector pair, regulated by a cardiac and skeletal muscle-specific CK8e promoter and bGHpA, utilizing a recombinant AK sequence. [Figure 25]Data are shown demonstrating the detection of full-length mini-dystrophin protein expression in human HEK293T cells transfected with either AK, StitchR, or a single ORF expression vector under the control of the core EF1a promoter, using Western blotting with N-terminal or C-terminal anti-dystrophin-specific antibodies. The StitchR technique was efficient in generating full-length protein (lane 7). This activity was dependent on ribozyme-mediated RNA cleavage, as a mutation in a single catalytic residue resulted in loss of full-length expression (lane 8). StitchR was more efficient than sequence-based vectors with AK recombination ability, which did not yield detectable full-length mini-dystrophin expression at this exposure level (comparing lanes 4 and 7). Notably, StitchR was nearly as efficient in generating full-length mini-dystrophin protein in a single open reading frame as vectors encoding full-length mini-dystrophin (comparing lanes 7 and 9). [Figure 26] The data shown, detected using Western blotting, demonstrate in vivo expression of full-length mini-dystrophin protein transduced with StitchR-compatible AAV virus under the control of the core CK8e promoter. The StitchR technique was effective in generating full-length mini-dystrophin protein at levels approximated endogenous mouse dystrophin protein in skeletal muscle (quadriceps) and heart. Mini-dystrophin was not detected in the liver, demonstrating the muscle specificity of the CK8e promoter. [Figure 27] The diagram shows the human dysferlin (DYSF) protein, indicating the location of known domains and important protein interaction sites. Human mutations in DYSF can lead to limb-girdle muscular dystrophy type 2b (LGMD2B) and Miyoshi myopathy (MM), which are now commonly referred to as dysferlin disorders. [Figure 28] The figure shows a StitchR-compatible dual AAV vector pair expressing a full-length human codon-optimized dysferlin (hcoDYSF) vector under the control of cardiac and skeletal muscle-specific CK8e promoters and bGH pA. [Figure 29] The following data shows full-length human dysferlin protein expression in human HEK293T cells transfected with either a dual StitchR vector or a single ORF expression vector encoding dysferlin, under the control of the core EF1a promoter, as detected by Western blotting using an anti-dysferlin antibody. [Figure 30] The following data shows full-length human sterocillin (STRC) protein expression in human HEK293T cells transfected with a dual StitchR vector under the control of the core EF1a promoter, as detected using Western blotting with either an N-terminal or C-terminal anti-STRC antibody. [Figure 31] This is a schematic diagram illustrating ribozyme-mediated scarless cis-splicing for generating circular RNA that can be translated via IRES-mediated translation. [Figure 32] This paper presents experimental data demonstrating the generation of GFP-coding circular RNA and the expression of GFP from circular RNA. [Figure 33] This is a schematic diagram illustrating the functions of IRES and ribozymes for protein expression from CirculR. [Figure 34] This figure shows lentivirus-based screening for functional IRES sequences using CirculR. [Figure 35] The diagram shows "Stitch RNA" or StitchR, which provides scarless splicing of RNA molecules in the absence of sequence homology or splendor oligonucleotides. [Figure 36] This figure shows data illustrating ribozyme-activated mRNA assembly and expression. [Figure 37A] Figures 37A and 17B show data illustrating mRNA rearrangement and full-length GFP protein expression in cells. They demonstrate scarless splicing of two independent split GFP RNAs. [Figure 37B] This shows the expression of full-length GFP protein. [Figure 38]This figure shows the development of a StitchR reporter assay using luciferase. The reporter provides quantitative measurement of StitchR activity in mammalian cells. Ribozyme activity is essential for StitchR-mediated luciferase activity. [Figure 39] This paper presents data identifying the optimal ribozyme for StitchR activity in mammalian cells (mRzB mutant RzB, mTwst mutant Twister). [Figure 40] The diagrams show the conventional RNA splicing pathway and the non-conventional RNA splicing pathway. [Figure 41] This document shows the RNA enzyme that catalyzes StitchR activity in vitro. RtcB is sufficient and necessary to catalyze StitchR activity in vitro. [Figure 42] This shows RNA enzymes that regulate StitchR activity in mammalian cells. RtcB enhances StitchR activity in mammalian cells. T4 PNK inhibits StitchR activity in mammalian cells. [Figure 43] This shows data demonstrating the specific detection of "stitched" RNA in mammalian cells. [Figure 44] This shows data demonstrating that intron sequences can be "stitched" and spliced. [Figure 45] Data is presented demonstrating that intron sequences can be “stitched” and spliced. Including “stitched” introns significantly enhanced StitchR activity in cells. [Figure 46] This figure shows RNA enzymes that can be used to "stitch" intron sequences and splice them. [Figure 47] The data demonstrates the optimization of StitchR activity in mammalian cells. Optimized StitchR activity approaches single-vector efficiency. StitchR utilizing Twister Rz in both Nt and Ct cells achieved 89% of the activity of a single ORF vector. [Figure 48]This report presents data demonstrating the optimization of StitchR activity in mammalian cells. Rice's compact P1-type Twister ribozyme (Oryza sativa, Osa) provides the strongest StitchR activity. [Figure 49] This figure shows data demonstrating the reconstruction of endogenous full-length proteins. [Figure 50] This figure shows data demonstrating the reconstruction of endogenous full-length proteins. [Figure 51] This figure shows data demonstrating the application of StitchR for dual AAV gene therapy. [Figure 52] This figure shows data demonstrating StitchR-mediated AAV delivery and full-length GFP expression in mammalian cells. [Figure 53] This figure shows data demonstrating StitchR-mediated AAV delivery and in vivo expression of full-length GFP. [Figure 54] This figure shows data demonstrating StitchR-mediated AAV delivery and in vivo expression of full-length GFP. [Figure 55] This paper presents data demonstrating the transsplicing efficiency of StitchR RNA in vivo. [Figure 56] This demonstrates a significant health need to treat diseases resulting from genes that are too large for a single AAV vector. [Figure 57] This paper demonstrates StitchR AAV gene therapy for treating major muscle genetic disorders. Dysferlin (DYSF) mutations cause limb-girdle muscular dystrophy type 2B and Miyoshi myopathy (MM), which are now commonly referred to as dysferlin disorders. With over 3000 patient-specific mutations and no common mutation hotspots, gene editing approaches have been challenging. Gene therapy approaches could potentially treat all patients. [Figure 58] This figure shows data demonstrating StitchR AAV gene therapy for dysferlin (DYSF). [Figure 59]This figure shows data demonstrating StitchR AAV gene therapy for dysferlin (DYSF). CK8e promoter - a compact, robust cardiac and skeletal muscle-specific promoter in both mouse and human. AAV 9 - an AAV serotype effective for targeting cardiac and skeletal muscle in both mouse and human. Constructs designed and tested in mice can be directly translated to treat patients. [Figure 60] This figure shows data demonstrating StitchR-mediated AAV delivery and in vivo expression of large therapeutic muscle proteins. [Figure 61] StitchR has multiple research and therapeutic applications, one of which is to expand the packaging capabilities of therapeutic viral vectors. [Figure 62A] Figures 62A–62G show experimental results demonstrating that StitchR-mediated expression of human mididystrophin restores the disease phenotype in a mouse model of Duchenne muscular dystrophy. Loss-of-function mutations in the large dystrophin (DMD) gene cause Duchenne muscular dystrophy. Large sequence deletions in the DMD gene can result in mild muscular dystrophy (BMD). Figures show the functional domains of the human dystrophin (Dys) protein and known DMD-protein interactions. The full-length dystrophin open reading frame (approximately 11kb) is too large to be packaged into a single AAV particle or into a dual AAV approach. Fully functional mididystrophin (ΔH2-R15), which is even larger than the deletions seen in BMD, can be packaged using a dual AAV approach. Here, the inventors reconstituted mididystrophin in vivo using a dual AAV9 serotype virus and a CK8e muscle-specific promoter with StitchR-activated RNA transligation. D2-MDX (dystrophin KO) mice lack dystrophin expression and develop myopathy characterized by muscle atrophy, muscle fiber necrosis, and fibrosis by 1 month of age. [Figures 62B-62C]This report presents experimental results demonstrating that StitchR-mediated expression of human-type mididystrophin restores the disease phenotype in a mouse model of Duchenne muscular dystrophy. Injection of D2-MDX mice with dual AAV StitchR gene therapy resulted in mididystrophin expression to levels comparable to full-length dystrophin in both male and female mice. [Figure 62D] StitchR mizidystrophin is expressed in the muscle fiber membrane, similar to endogenous full-length dystrophin, improving muscle pathology and restoring the membrane localization of known dystrophin-interacting proteins nNOS, which are normally disrupted in D2-MDX mice. [Figure 62E] This paper presents experimental results demonstrating that StitchR-mediated expression of human midijidystrophin restores the disease phenotype in a mouse model of Duchenne muscular dystrophy. Western blot gel decitometry demonstrated that midijidystrophin was expressed at 94.9% of WT levels. [Figures 62F-62G] We present experimental results demonstrating that StitchR-mediated expression of human mididystrophin restores the disease phenotype in a mouse model of Duchenne muscular dystrophy. Mididystrophin expression in D2-MDX mice was sufficient to normalize serum creatine kinase (CK) levels back to WT levels (Figure 62F), and both significantly reduced the proportion of muscle fibers with central nuclei, a characteristic of DMD (Figure 62G). [Figure 63] This figure shows the quantification of StitchR-mediated expression of ΔH2-R15 compared to WT DMD Western blot using densitometry. Samples were normalized to a vinculin-loaded control. [Figure 64A]Figures 64A–64E show the results of experiments demonstrating StitchR-mediated expression of human full-length dysferlin (DYSF) in dysferlin knockout mice (AJ starin). The dysferlin gene encodes a large fascial protein in which loss-of-function mutations cause limb-girdle muscular dystrophy type 2B and Miyoshi myopathy, commonly referred to as dysferlin disorders. The dysferlin open reading frame exceeds the packaging capacity of a single AAV vector, but can be fully packaged using a dual AAV vector approach. Full-length dysferlin was reconstituted and expressed under the control of a muscle-specific CK8e promoter using a StitchR-mediated dual AAV9 vector approach. [Figures 64B-64C] Full-length human dysferlin protein was expressed at levels exceeding those observed in both wild-type male and wild-type female mice. [Figure 64D] Human dysferlin, like WT-type dysferlin, was localized to the membrane, but it also accumulated inside cells, a phenomenon that has been observed in other dysferlin transgenic approaches. [Figure 64E] This shows the relative intensity of dysferlin. [Figure 65A] Figures 65A and 65B show experimental results demonstrating the development of stitchR activity-dependent gene cassettes and cell lines. The figures show a stitchR-dependent ribotron-disrupted HSV-TK expression cassette encoded within a lentiviral vector in reverse orientation to enable lentiviral RNA delivery. Lentiviral delivery allows for non-lysable integration of the ribotron-encoded HSV-TK cassette into cells, although other methods can be used to stably express the ribotron-HSV-TK transgene. Including antibiotic selection genes such as blastosidine allows for the selection of stably integrated cell clones. [Figure 65B]Expression of ribozyme-activated synthetic introns (or ribotrons) enables stitchR-dependent expression of the HSV-TK cell death gene in cells that induce cell death in the presence of the non-toxic molecule ganciclovir (GCN). In control, mutations in catalytic nucleotides within the ribozyme sequence disrupted sensitization to ganciclovir. [Figure 66]This figure shows experimental results demonstrating that in vitro transcribed CirculRRNA is reliably translated without requiring further processing (cap and poly(A) tail) necessary to produce linear in vitro transcribed RNA suitable for translation. Linear mRNA undergoes innate processing similar to endogenous RNA when transcribed from a DNA plasmid within a cell, being capped, spliced, and polyadenylated—modifications necessary for translation into protein. This can be visualized by a DNA plasmid encoding a complete open reading frame of green fluorescent protein (GFP), which, under the control of a mammalian promoter (e.g., sCMVIE94), exhibits robust green fluorescence compared to untransfected cells. Interestingly, however, in vitro transcribed RNA containing the same GFP open reading frame using phage RNA polymerase (e.g., T7) is not translated when transfected into cells. Linear in vitro transcribed RNA requires the addition of both a 5' cap and a 3' poly(A) tail for this RNA to produce green fluorescence. Adding internal ribosome entry sequences (IRESs), such as those derived from coxsackievirus B3 virus (CVB3IRES), which enable cap-independent translation, dramatically reduces the translation of linear RNA transcribed from plasmid DNA, but is insufficient to enable protein translation from linear in vitro transcribed RNA. CirculR RNA encodes adjacent autocatalytic ribozyme sequences that generate unique 5' and 3' ends that catalyze the scarless circularization of RNA in eukaryotic cells. Notably, circularization occurs in ribozyme-cleaved RNA transcribed into cells from plasmid DNA or generated from in vitro transcription in linear form and then transfected into cells. Because circular RNA lacks a 5' or 3' end, CirculR RNA requires IRES addition for translation to be equivalent in activity to linear RNA translated from IRES when transfected from plasmid DNA.However, CirculR RNA containing the IRES sequence is the only in vitro transcription RNA that can be translated into protein, and its activity is equivalent to that of linear RNA processed with a 5' cap and 3' poly(A) tail. [Figure 67A] Figures 67A–67C illustrate CirculR-mediated scarless rearrangement of functional RNA motifs. Previously, we demonstrated the unique ability of CirculR (ribozyme-mediated RNA cyclization in eukaryotic cells) for scarless rearrangement of proteins encoding open reading frames (GFP, blastosidine, etc.). Here, we demonstrate that CirculR can be used for scarless rearrangement of functional RNA motifs in cyclized RNA. [Figure 67B] For example, small RNA motifs with known protein-binding partners (such as MS2, PRR1, Qβ, PP7, and S1) can be incorporated into the design of scarless self-cleaving ribozymes such as Twister and RzB Hammerhead. [Figure 67C] During autocleavage and ligation of two ends, either in vitro or intracellularly, circularization leads to the reconstruction of functional RNA motifs. This may be useful for the specific recognition of circular-paired-linear RNAs, which can then be used for the purification, intracellular targeting, or packaging of circular RNAs. [Figure 68] Figure 17B shows the average relative luciferase levels. [Figure 69] Figure 17C shows the average relative luciferase levels. [Figure 70]This paper describes sequence optimization and requirements for StitchR activity. In cell-based assays, StitchR-mediated transligation of full-length dysferlin has been previously demonstrated. Here, the data show that the polyadenylated sequence (PAS) in the N-terminal StitchR vector is not essential for StitchR activity. StitchR may be compatible with other dual-vector techniques, such as DNA transsplicing, which can be mediated via AAV ITR sequences or recombinant sequences (such as AK sequences). In cell-based assays using transfected plasmid DNA, the addition of AK sequences neither promoted nor inhibited StitchR activity. [Figure 71] This figure compares intein-mediated protein trans-splicing with StitchR-activated RNA transligation. Transligation approaches using intein or intein + extein proteins were compared to StitchR to express full-length dysferlin. Because intein requires an adjacent extein sequence for efficiency, the intein sequence alone was not sufficient to promote full-length dysferlin expression at this splitting site. Inteins with an optimal extein sequence (leaving a small protein scar) were not as efficient as StitchR-based dysferlin expression. [Modes for carrying out the invention]
[0046] definition Unless otherwise defined, all technical and / or scientific terms used herein have the same meaning as those generally understood by those skilled in the art to which this invention pertains.
[0047] In general, the nomenclature used herein, as well as the laboratory procedures in cell culture, molecular genetics, organic chemistry, nucleic acid chemistry, and hybridization, are well known and commonly used in the art.
[0048] Standard techniques are used for the synthesis of nucleic acids and peptides. The techniques and procedures are generally carried out in accordance with conventional methods in the art and various general references provided throughout this document (e.g., Sambrook and Russell, 2012, Molecular Cloning, A Laboratory Approach, Cold Spring Harbor Press, Cold Spring Harbor, NY, and Ausubel et al., 2012, Current Protocols in Molecular Biology, John Wiley & Sons, NY).
[0049] The nomenclature used herein, as well as the experimental procedures used in analytical chemistry and organic synthesis described below, are well known and commonly used in the art. Standard techniques or their modifications are used in chemical synthesis and chemical analysis.
[0050] In the context of this invention (particularly in the context of the claims), the terms "a," "an," "the," and similar terms should be interpreted as encompassing both singular and plural forms unless otherwise specifically indicated herein or unless otherwise clearly contradicted by the context.
[0051] As used herein, "about" when referring to measurable values such as quantity or duration means that it includes variations of ±20%, ±10%, ±5%, ±1%, or ±0.1% from the specified value, and such variations are suitable for performing the disclosed method.
[0052] "Antisense" specifically refers to the nucleic acid sequence of the non-coding strand of a double-stranded DNA molecule that codes for a protein, or a sequence that is substantially homologous to the non-coding strand. As defined herein, an antisense sequence is complementary to the sequence of a double-stranded DNA molecule that codes for a protein. An antisense sequence does not need to be complementary only to the coding portion of the coding strand of the DNA molecule. An antisense sequence may be complementary to a regulatory sequence specified on the coding strand of a protein-coding DNA molecule, which controls the expression of the coding sequence.
[0053] When referring to the immobilization of molecules (e.g., nucleic acid molecules) onto a solid support, the term “adhered” as used herein is intended to encompass direct or indirect, covalent or non-covalent adhesion, either explicitly or in context, unless otherwise specified.
[0054] As used interchangeably herein, “microspheres,” “beads,” or their grammatical equivalents refer to small, distinct particles that can act as solid supports for attaching biomolecules (e.g., nucleic acid molecules).
[0055] "Disease" is a state of animal health in which the animal is unable to maintain homeostasis and, if the disease does not improve, the animal's health continues to deteriorate.
[0056] In contrast, a “disorder” in animals is a state of health in which the animal can maintain homeostasis, but the animal’s health is less desirable than when there is no disorder. Even if left untreated, a disorder does not necessarily lead to a further decline in the animal’s health.
[0057] A disease or disorder is “relieved” if the severity of the signs or symptoms of the disease or disorder, the frequency with which the patient experiences such signs or symptoms, or both, decreases.
[0058] "Code" refers to the inherent properties of a specific nucleotide sequence contained within a polynucleotide such as a gene, cDNA, or mRNA. These properties allow such a sequence to function as a template for the synthesis of other polymers and macromolecules with defined nucleotide sequences (i.e., rRNA, tRNA, and mRNA) or defined amino acid sequences in biological processes, from which biological properties are derived. Therefore, if the transcription and translation of the mRNA corresponding to that gene produces a protein in a cell or other biological system, then that gene codes for a protein. Both the coding strand, whose nucleotide sequence is identical to the mRNA sequence and is typically provided in a sequence listing, and the non-coding strand, used as a template for the transcription of the gene or cDNA, can be said to code for the protein or other product of that gene or cDNA.
[0059] The terms “patient,” “subject,” and “individual” are used interchangeably herein and refer to any animal or cell suitable for the methods herein, whether in vitro or in vivo. In one embodiment, the subject includes vertebrates and invertebrates. Invertebrates include, but are not limited to, Drosophila melanogaster and Caenorhabditis elegans. Vertebrates include, but are not limited to, primates, rodents, domesticated animals, or play animals. Examples of primates include, but are not limited to, chimpanzees, crab-eating macaques, spider monkeys, and macaques (e.g., rhesus macaques). Examples of rodents include, but are not limited to, mice, rats, woodchucks, ferrets, rabbits, and hamsters. Examples of livestock and game animals include, but are not limited to, cattle, horses, pigs, deer, bison, buffalo, feline species (e.g., domestic cats), canine species (e.g., dogs, foxes, wolves), avian species (e.g., chickens, emus, ostriches), and fish (e.g., zebrafish, trout, catfish, and salmon). In some embodiments, the subject is a mammal, such as a primate, such as a human. In certain non-limiting embodiments, the patient, subject, or individual is a human.
[0060] As used herein with respect to antibodies, the term “specifically binding” means an antibody that recognizes a particular antigen but substantially does not recognize or bind to other molecules in the sample. For example, an antibody that specifically binds to an antigen from one species may also bind to that antigen from one or more species. However, such cross-species reactivity does not in itself alter the specific classification of the antibody. In another example, an antibody that specifically binds to an antigen may also bind to different alleles of the antigen. However, such cross-reactivity does not in itself alter the specific classification of the antibody.
[0061] In some cases, the terms "specific binding" or "specifically binding" can be used in reference to the interaction between an antibody, protein, or peptide and a second chemical species, meaning that the interaction depends on the presence of a specific structure on the chemical species (e.g., an antigenic determinant or epitope). For example, antibodies generally recognize and bind to specific protein structures rather than proteins. If an antibody is specific to epitope "A", then in a reaction involving label "A" and the antibody, the presence of a molecule containing epitope A (or free, unlabeled A) reduces the amount of labeled A bound to the antibody.
[0062] The "coding region" of a gene consists of nucleotide residues in the coding chain and nucleotides in the non-coding chain of the gene that are homologous or complementary to the coding region of the mRNA molecule produced by gene transcription.
[0063] The "coding region" of an mRNA molecule also consists of nucleotide residues of the mRNA molecule that either coincide with the anti-codon region of the transfer RNA molecule during translation of the mRNA molecule, or encode a stop codon. Therefore, the coding region may contain nucleotide residues that contain codons for amino acid residues that are not present in the mature protein encoded by the mRNA molecule (e.g., amino acid residues in the protein transport signal sequence).
[0064] As used herein to refer to nucleic acids, “complementary” refers to the broad concept of sequence complementarity between regions of two nucleic acid chains or between two regions of the same nucleic acid chain. It is known that an adenine residue in a first nucleic acid region can form a specific hydrogen bond ("base pairing") with a residue in a second nucleic acid region that is antiparallel to the first region, if the residue is thymine or uracil. Similarly, it is known that a cytosine residue in a first nucleic acid chain can base pair with a residue in a second nucleic acid chain that is antiparallel to the first chain, if the residue is guanine. When two regions are arranged antiparallel, a first region of a nucleic acid is complementary to a second region of the same or different nucleic acid if at least one nucleotide residue in the first region can base pair with a residue in the second region. In one embodiment, if the first region includes a first portion and the second region includes a second portion, and the first and second portions are arranged antiparallel, then at least about 50%, at least about 75%, at least about 90%, or at least about 95% of the nucleotide residues of the first portion can base pair with the nucleotide residues of the second portion. In one embodiment, all nucleotide residues of the first portion can base pair with the nucleotide residues of the second portion.
[0065] As used herein, the term "DNA" is defined as deoxyribonucleic acid.
[0066] As used herein, the term “expression” is defined as the transcription and / or translation of a particular nucleotide sequence driven by its promoter.
[0067] As used herein, the term “expression vector” refers to a vector containing a nucleic acid sequence that codes for at least a portion of a gene product that can be transcribed. In some cases, the RNA molecule is then translated into a protein, polypeptide, or peptide. In other cases, these sequences are not translated, for example, in the production of antisense molecules, siRNA, ribozymes, etc. Expression vectors may contain a variety of regulatory sequences that point to nucleic acid sequences necessary for the transcription and possibly translation of the operably linked coding sequence in a particular host organism. In addition to regulatory sequences that govern transcription and translation, vectors and expression vectors may also contain nucleic acid sequences that serve other functions.
[0068] As used herein, the term “wild type” is a term of the art as understood by those skilled in the art, and means a typical form of an organism, strain, gene or characteristic occurring in nature, distinct from a mutant or variant form.
[0069] The term "homology" refers to the degree of complementarity. Partial homology or complete homology (i.e., identity) can exist. Homologousity is often measured using sequence analysis software (e.g., the Sequence Analysis Software Package (Genetics Computer Group, University of Wisconsin Biotechnology Center, 1710 University Avenue, Madison, Wis. 53705)). Such software matches similar sequences by assigning a degree of homology to various substitutions, deletions, insertions, and other modifications. Conservative substitutions typically include substitutions within the following groups: glycine, alanine, valine, isoleucine, leucine, aspartic acid, glutamic acid, asparagine, glutamine, serine, threonine, lysine, arginine, and phenylalanine, tyrosine.
[0070] "Isolated" means altered or removed from its natural state. For example, a nucleic acid or peptide that is naturally present in a living animal under its normal conditions is not "isolated," but the same nucleic acid or peptide that has been partially or completely separated from its coexisting substances in its natural conditions is "isolated." Isolated nucleic acids or proteins may exist in a substantially purified form or in a non-natural environment, such as a host cell.
[0071] The term “isolated,” when used in relation to nucleic acids, as in “isolated oligonucleotide” or “isolated polynucleotide,” refers to a nucleic acid sequence that is identified and isolated from at least one contaminant that it normally associates with at its source. Therefore, isolated nucleic acids exist in a form or environment different from the form or environment in which they are found in nature. In contrast, non-isolated nucleic acids (e.g., DNA and RNA) are found in the state in which they exist in nature. For example, a given DNA sequence (e.g., a gene) is found on a host cell chromosome adjacent to neighboring genes. An RNA sequence (e.g., a specific mRNA sequence encoding a particular protein) is found in a cell as a mixture with numerous other mRNAs encoding numerous proteins. However, isolated nucleic acids include, for example, nucleic acids in cells that normally express nucleic acids that are located at a chromosomal position different from their natural chromosomal position in a cell, or that are adjacent to nucleic acid sequences different from those found naturally. Isolated nucleic acids or oligonucleotides can exist in single-stranded or double-stranded forms. When expressing a protein using isolated nucleic acids or oligonucleotides, the oligonucleotide may contain at least a sense strand or a coding strand (i.e., the oligonucleotide may be single-stranded), but may also contain both a sense strand and an antisense strand (i.e., the oligonucleotide may be double-stranded).
[0072] The term "isolated," when used in reference to polypeptides such as "isolated protein" or "isolated polypeptide," refers to a polypeptide identified and isolated from at least one contaminant that is typically associated with its source. Therefore, isolated polypeptides exist in a form or environment different from the form or environment in which they are found in nature. In contrast, non-isolated polypeptides (e.g., proteins and enzymes) are found in the state in which they exist in nature.
[0073] "Nucleic acid" means any nucleic acid, whether or not it is composed of deoxyribonucleosides or ribonucleosides, and whether or not it is composed of phosphodiester bonds or modified links, such as phosphotriesters, phosphoramidates, siloxanes, carbonates, carboxymethyl esters, acetamidates, carbamates, thioethers, cross-linked phosphoramidates, cross-linked methylenephosphonates, phosphorothioates, methylphosphonates, phosphorodithioates, cross-linked phosphorothioates, or sulfone bonds, and combinations of such bonds. Specifically, the term nucleic acid also includes nucleic acids composed of bases other than the five biologically present bases (adenine, guanine, thymine, cytosine, and uracil). The term "nucleic acid" typically refers to large polynucleotides.
[0074] Conventional notation is used in this specification to describe polynucleotide sequences. The left end of a single-stranded polynucleotide sequence is the 5' end. The leftward direction of a double-stranded polynucleotide sequence is called the 5' direction.
[0075] The direction in which nucleotides are added to a newly generated RNA transcript from 5' to 3' is called the transcription direction. The DNA strand that has the same sequence as the mRNA is called the "coding strand." The sequence on the DNA strand located 5' to the reference point on the DNA is called the "upstream sequence." The sequence on the DNA strand that is 3' relative to the reference point on the DNA is called the "downstream sequence."
[0076] An "expression cassette" refers to a nucleic acid molecule containing a coding sequence that is operably linked to a promoter / regulatory sequence necessary for the transcription and, if applicable, translation of the coding sequence.
[0077] As used herein, the term “operatably linked” refers to the linking of nucleic acid sequences in such a manner that a nucleic acid molecule is produced that can direct the transcription of a given gene and / or the synthesis of a desired protein molecule. The term also refers to the linking of amino acid-coding sequences in such a manner that a functional protein or polypeptide is produced (e.g., enzymatically active, capable of binding to a binding partner, capable of inhibiting, etc.).
[0078] As used herein, the term “promoter / regulatory sequence” means a nucleic acid sequence required for the expression of a gene product operably ligated to a promoter / regulatory sequence. In some examples, this sequence may be a core promoter sequence, while in others, this sequence may also include enhancer sequences and other regulatory elements required for the expression of the gene product. A promoter / regulatory sequence may, for example, inducibly express a gene product.
[0079] As used herein, “stringent conditions” for hybridization refer to conditions under which a nucleic acid complementary to the target sequence primarily hybridizes with the target sequence and substantially does not hybridize with the non-target sequence. Stringent conditions are generally sequence-dependent and vary depending on many factors. Generally, the longer the sequence, the higher the temperature at which the sequence specifically hybridizes with its target sequence. Non-limiting examples of stringent conditions are described in detail in Tijssen (1993), Laboratory Techniques In Biochemistry And Molecular Biology—Hybridization With Nucleic Acid Probes Part 1, Second Chapter, “Overview of principles of hybridization and the strategy of nucleic acid probe assay,” Elsevier, NY.
[0080] Hybridization refers to a reaction in which one or more polynucleotides react to form a complex that is stabilized via hydrogen bonds between the bases of nucleotide residues. Hydrogen bonding can occur in Watson-Crick base pairing, Hoogstein bonds, or any other sequence-specific manner. The complex can consist of two strands forming a double-stranded structure, three or more strands forming a multi-stranded complex, a single self-hybridized strand, or any combination thereof. Hybridization reactions can constitute a step in a broader process, such as the initiation of PCR or enzymatic cleavage of polynucleotides. A sequence that can hybridize with a given sequence is called a "complement" of the given sequence.
[0081] An "inducible" promoter is a nucleotide sequence that, when operably ligated to a polynucleotide encoding or designating a gene product, substantially produces the gene product only if the corresponding inducer is present.
[0082] A "constitutive" promoter is a nucleotide sequence that, when operably linked to a polynucleotide encoding or designating a gene product, causes the cell to produce that gene product under most or all physiological conditions.
[0083] As used herein, the term “polynucleotide” is defined as a chain of nucleotides. Furthermore, nucleic acids are polymers of nucleotides. Thus, nucleic acids and polynucleotides as used herein are interchangeable. Those skilled in the art have general knowledge that nucleic acids are polynucleotides that can be hydrolyzed to monomeric “nucleotides.” Monomeric nucleotides can be hydrolyzed to nucleosides. As used herein, polynucleotides include, but are not limited to, all nucleic acid sequences obtained by recombinant means, i.e., any means available in the art, including but not limited to the cloning of nucleic acid sequences from recombinant libraries or cell genomes using conventional cloning techniques and PCR, and by synthetic means.
[0084] In the context of this invention, the following abbreviations for commonly existing nucleic acid bases are used: "A" refers to adenosine, "C" to cytosine, "G" to guanosine, "T" to thymidine, and "U" to uridine.
[0085] As used herein, the terms “peptide,” “polypeptide,” and “protein” are interchangeable and refer to compounds composed of amino acid residues covalently linked by peptide bonds. A protein or peptide must contain at least two amino acids, and there is no limit to the maximum number of amino acids that can be contained in a protein or peptide sequence. A polypeptide includes any peptide or protein containing two or more amino acids linked to each other by peptide bonds. As used herein, this term refers to both short chains, also commonly called peptides, oligopeptides, and oligomers in the art, and long chains, also commonly called proteins in the art, of which there are many types. A “polypeptide” includes, for example, biologically active fragments, substantially homologous polypeptides, oligopeptides, homodimers, heterodimers, polypeptide variants, modified polypeptides, derivatives, analogs, and fusion proteins. Polypeptides include native peptides, recombinant peptides, synthetic peptides, or combinations thereof.
[0086] As used herein, the term "RNA" is defined as ribonucleic acid.
[0087] As used herein, the term "ribozyme" refers to an RNA molecule capable of acting as an enzyme. For example, some ribozymes can cleave RNA molecules. RNA-cleaving ribozymes typically consist of at least a catalytic domain and a recognition sequence recognized by the catalytic domain. The catalytic domain may be part of the same RNA molecule as the recognition sequence and thus mediate cis-cleavage. Alternatively, the catalytic domain may be a separate RNA molecule from the RNA molecule containing the recognition sequence and thus mediate trans-cleavage.
[0088] "Recombinant polynucleotides" refer to polynucleotides that have sequences that are not naturally linked together. Amplified or assembled recombinant polynucleotides can be included in suitable vectors, which can be used to transform suitable host cells.
[0089] Recombinant polynucleotides can also perform non-coding functions (e.g., promoters, origins of replication, ribosome binding sites, etc.).
[0090] As used herein, the term “recombinant polypeptide” is defined as a polypeptide produced by the use of recombinant DNA methods.
[0091] As used herein, the terms “solid surface,” “solid support,” and other grammatical equivalents refer to any material that is suitable for the attachment of biomolecules (e.g., nucleic acid molecules), or that can be modified to be suitable for the attachment of biomolecules.
[0092] As used herein, the term “tag” refers to any chemical modification of a biomolecule (e.g., a nucleic acid molecule) that provides additional functionality (e.g., adhesion to a solid support, fluorescence visualization, etc.).
[0093] Where used herein, a “variant” is a nucleic acid or peptide sequence that differs in sequence from a reference nucleic acid or peptide sequence, but retains the essential biological properties of the reference molecule. Sequence changes in nucleic acid variants may not alter the amino acid sequence of the peptide encoded by the reference nucleic acid, or they may result in amino acid substitutions, additions, deletions, fusions, and cleavages. Sequence changes in peptide variants are typically limited or conserved, so the sequences of the reference peptide and the variant are closely similar overall and identical in many regions. Variants and reference peptides may differ in amino acid sequence by one or more substitutions, additions, or deletions in any combination. Variants of nucleic acids or peptides may be naturally occurring variants, such as allele variants, or variants not known to exist naturally. Variants of nucleic acids and peptides not known to exist naturally may be produced by mutagenesis or direct synthesis.
[0094] A “vector” is a composition of substances containing isolated nucleic acids that can be used to deliver the isolated nucleic acids into the interior of a cell. Numerous vectors are known in the art, including, but not limited to, linear polynucleotides, polynucleotides associated with ionic or amphiphilic compounds, plasmids, and viruses. Therefore, the term “vector” includes autonomously replicating plasmids or viruses. This term should also be interpreted to include non-plasmid and non-viral compounds that facilitate the transfer of nucleic acids into cells, such as polylysine compounds and liposomes. Examples of viral vectors include, but are not limited to, adenovirus vectors, adeno-associated virus vectors, and retroviral vectors.
[0095] Scope: Throughout this disclosure, various aspects of the invention can be presented in scope form. It should be understood that scope form descriptions are merely for convenience and brevity and should not be interpreted as inflexible limitations on the scope of the invention. Therefore, scope descriptions should be considered to specifically disclose all possible sub-ranges and the individual numbers within those ranges. For example, a scope description such as 1-6 should be considered to specifically disclose sub-ranges such as 1-3, 1-4, 1-5, 2-4, 2-6, 3-6, and the individual numbers within those ranges, e.g., 1, 2, 2.7, 3, 4, 5, 5.3, and 6. This applies regardless of the width of the range.
[0096] Detailed explanation The present invention provides compositions and methods for efficiently and reliably linking two or more individual RNA molecules to produce a larger single RNA molecule encoding a protein and a fusion protein. The present invention utilizes ribozyme-mediated trans-splicing of multiple RNA molecules to assemble a single RNA molecule encoding a target protein or fusion protein. The present invention can be used to efficiently produce fusion proteins, chimeric proteins, and the like. Furthermore, the present invention is useful for producing large expression products (e.g., proteins or fusion proteins) whose coding sequences are too large to be packaged in a single vector. For example, in some embodiments, the full-length expression product is longer than 1000 amino acid residues. In some embodiments, the full-length expression product is longer than 2000 amino acid residues. In some embodiments, the full-length expression product is longer than 3000 amino acid residues. Furthermore, the techniques of the present invention also enable the rapid and easy combination of two different sequences, which can have synergistic effects for generating novel protein combinations or library sequences. This may be particularly useful, for example, for the production of synthetic antibodies (such as nanobodies) or the functional selection of enzymes.
[0097] The present invention also provides compositions and methods for efficiently delivering one or more RNA molecules having a synthetic intron adjacent to a ribozyme. The synthetic intron adjacent to the ribozyme can be positioned between a first RNA portion encoding the N-terminal portion of the target protein and a second RNA portion encoding the C-terminal portion of the target protein. The synthetic intron adjacent to the ribozyme may include a cargo sequence, such as a sequence encoding a therapeutic protein or a sequence containing functional RNA. The use of two ribozymes makes it possible to generate three RNA fragments by cis-splicing: 1) a first RNA portion encoding the N-terminal portion of the target protein, 2) a synthetic intron adjacent to the ribozyme, and 3) a second RNA portion encoding the C-terminal portion of the target protein. Cis-splicing generates ends that are suitable for ligation. Ligation of the suitable ends of the cis-spliced synthetic intron generates a circular RNA molecule that is more resistant to degradation than linear RNA molecules. The matching ends of a first RNA portion encoding the N-terminal portion of the target protein and a second RNA portion encoding the C-terminal portion of the target protein are ligated to generate an RNA molecule encoding the full-length target protein. The full-length target protein may be, for example, a therapeutic protein, a CRISPR-Cas protein, or a reporter protein to provide a surrogate indicator for the delivery and expression of cargo sequences within a circular RNA molecule containing synthetic introns adjacent to a ribozyme.
[0098] In one embodiment, the present invention provides one or more nucleic acid molecules encoding two or more RNA molecules. In one embodiment, one or more RNA molecules include a ribozyme. In one embodiment, one or more RNA molecules include a coding region and a ribozyme. In a particular embodiment, the ribozyme autocleaves from the RNA molecule, leaving the coding region. Exemplary ribozymes that may be used in the context of the present invention include, but are not limited to, Hammerhead (HH), Hepatitis Delta Virus (HDV), Barkood Satellite (VS), Sister, Twister-sister, Hairpin, Hatchet, Pistol, HOV Linc, or members of the lanthanum family of ribozymes. However, the present invention is not limited to any particular ribozyme, but rather encompasses all known members of the family of endogenous ribozymes and any potential artificial ribozymes. That is, as described elsewhere in this specification, the ribozymes described, all known endogenous ribozymes, and potential artificial ribozymes can be used for ligation, trans-splicing, and cyclization of multiple RNAs.
[0099] For example, in one embodiment, the composition comprises a nucleic acid molecule encoding a first RNA molecule, the first RNA molecule comprising a coding region and a 3' ribozyme, the 3' ribozyme being able to catalyze itself from the RNA molecule, leaving behind a coding region having a 3'P or 2'3' cyclic phosphate (cP) terminus. In one embodiment, the 3' ribozyme comprises an HDV ribozyme. Furthermore, in one embodiment, the composition comprises a nucleic acid molecule encoding a second RNA molecule, the second RNA molecule comprising a coding region and a 5' ribozyme, the 5' ribozyme being able to catalyze itself from the RNA molecule, leaving behind a coding region with a 5'OH terminus. In one embodiment, the 5' ribozyme comprises an HH ribozyme. In certain cases, the ligase ligates the coding region of the first RNA molecule together with the coding region of the second RNA molecule to form a longer RNA molecule encoding the protein of interest.
[0100] For example, in one embodiment, the composition comprises a first RNA molecule, the first RNA molecule comprising a coding region and a 3' ribozyme, the 3' ribozyme being able to catalyze itself from the RNA molecule, leaving behind a coding region having a 3'P or 2'3' cyclic phosphate (cP) terminus. In one embodiment, the 3' ribozyme comprises an HDV ribozyme. Furthermore, in one embodiment, the composition comprises a second RNA molecule, the second RNA molecule comprising a coding region and a 5' ribozyme, the 5' ribozyme being able to catalyze itself from the RNA molecule, leaving behind a coding region having a 5'OH terminus. In one embodiment, the 5' ribozyme comprises an HH ribozyme. In certain cases, the ligase ligates the coding region of the first RNA molecule together with the coding region of the second RNA molecule to form a longer RNA molecule encoding the protein of interest.
[0101] In certain embodiments, the first RNA contains a coding region encoding a first portion of the protein of interest, and the second RNA contains a coding region encoding a second portion of the protein of interest, and thus ribozyme-mediated cleavage and ligase-mediated assembly of the RNA molecules result in the production of an RNA molecule encoding a protein having both the first and second portions. The present invention can be used to produce a full-length protein from a plurality of RNAs, each containing a coding region encoding a portion of the full-length protein. Furthermore, the present invention can be used to produce a fusion protein containing multiple domains, where each RNA molecule contains a coding region encoding a domain of the fusion protein. For example, the present invention can be used to generate an RNA molecule encoding a protein having a leader sequence, an N-terminal tag, or a C-terminal tag by assembling RNA from a first RNA molecule containing a coding sequence encoding a leader sequence, an N-terminal tag, or a C-terminal tag, and a second RNA molecule containing a coding sequence encoding a protein.
[0102] In certain embodiments, the present invention relates to the formation of a single RNA molecule from three or more individual RNA molecules. For example, in certain embodiments, the composition comprises a nucleic acid molecule encoding a first RNA molecule and a nucleic acid molecule encoding a second RNA molecule, the first RNA molecule comprising a coding region encoding the N-terminal region of a protein, the second RNA molecule comprising a coding region encoding the C-terminal region of a protein, and further comprising one or more additional RNA molecules (each comprising a coding region encoding a protein domain [e.g., a repeat domain]) and one or more nucleic acid molecules. In one embodiment, the first RNA molecule comprises a coding region encoding the N-terminal region and a 3' ribozyme, the 3' ribozyme can catalyze itself from the RNA molecule, leaving a coding region having a 3'P or 2'3' cyclic phosphate (cP) terminus. In one embodiment, the 3' ribozyme comprises an HDV ribozyme. In one embodiment, the second RNA molecule comprises a coding region encoding the C-terminal region and a 5' ribozyme, the 5' ribozyme can catalyze itself from the RNA molecule, leaving a coding region having a 5'OH terminus. In one embodiment, the 5' ribozyme comprises an HH ribozyme. In one embodiment, each additional RNA molecule comprises a coding region encoding a protein domain, a 3' ribozyme, and a 5' ribozyme. In one embodiment, the 3' ribozyme is an HDV ribozyme. In one embodiment, the 5' ribozyme is an HH ribozyme. In certain embodiments, the 3' ribozyme can catalyze itself from the RNA molecule, and the 5' ribozyme can catalyze itself from the RNA molecule, leaving a coding region having a 5'OH and a 3'P or 2'3'cP terminus. In one embodiment, each additional RNA molecule comprises a coding region encoding a protein domain, a 5' ribozyme, and a 3' ribozyme recognition sequence. In certain embodiments, the 5' ribozyme can catalyze itself from the RNA molecule, leaving behind a coding region having a 5'OH terminus, while the 3' ribozyme recognition sequence interacts with the ribozyme to induce splicing of the 3' ribozyme recognition sequence from the RNA molecule, leaving behind a coding region having a 3'P or 2'3'cP terminus.In one embodiment, the 3' ribozyme recognition sequence includes a Vsv1 sequence that interacts with a VS ribozyme. This technique can be used to generate an RNA molecule encoding a protein having multiple repeat domains by sequentially adding coding regions that encode repeat domains, by sequentially providing ribozymes (e.g., VS ribozymes) that interact with the 3' ribozyme recognition sequence to generate a 3'P or 2'3'cP terminus, and ligating this coding region to the 5'OH terminus of another coding region encoding a repeat domain. In certain embodiments, the sequential addition of repeat domains can be performed on a solid substrate or support to which a first RNA molecule encoding an N-terminal region is bound.
[0103] In certain embodiments, multiple RNA molecules are ligated together after ribozyme-mediated generation of 5'OH and 3'P or 2'3'cP termini. In some examples, the RNA molecules are ligated together by endogenous ligases present in the native cells or tissues where RNA assembly is taking place. In some cases, the method of the present invention includes the step of adding an exogenous ligase to induce joint ligation of the processed RNA molecules. In one embodiment, the ligase is an RNA 2',3'-cyclic phosphate and 5'-OH(RtcB) ligase.
[0104] In certain embodiments, the present invention relates to the use of deoxyribozymes, enzymes, or DNAases for cleaving single-stranded DNA sequences for trans-splicing or editing in trans. For example, a deoxyribozyme (a self-cleaved DNA sequence leaving 3'-P (or 2'3'-cP) and 5'-OH ends), an enzyme, or DNAse (which may be fused with a DNA-targeting protein such as TALEN or CRISPR) that cleaves and leaves the same ends, can be used on a single-stranded DNA substrate for trans-splicing or editing in trans with the trans-cleaved deoxyribozyme sequence. A RTCB can be used to act on a single-stranded DNA substrate having 3'-P (or 2'3'-cP) and 5'-OH ends generated by the deoxyribozyme, enzyme, or DNAase.
[0105] composition In one embodiment, the present invention relates to a composition comprising one or more nucleic acid molecules encoding one or more ribozymes. In one embodiment, the present invention comprises one or more RNA molecules comprising one or more ribozymes. In some embodiments, the one or more RNA molecules comprise at least a first RNA molecule and a second RNA molecule.
[0106] In some embodiments, one or more ribozymes in the composition can spontaneously cis-cleave from one or more RNA molecules. In some embodiments, one or more ribozymes are 3' ribozymes. In some embodiments, the 3' ribozymes produce a 3'P or 2'3'cP terminus on the remaining one or more RNA molecules after spontaneous cis-cleavage. In some embodiments, one or more ribozymes are 5' ribozymes. In some embodiments, the 5' ribozymes produce a 5'OH terminus on the remaining one or more RNA molecules after spontaneous cis-cleavage. In some embodiments, the 3'P or 2'3'cP terminus and the 5'OH terminus can be ligated together.
[0107] In some embodiments, the first RNA molecule comprises a 3' ribozyme. In some embodiments, the 3' ribozyme is derived from one or more families selected from the group consisting of Hammerhead (HH), Hepatitis Delta Virus (HDV), Barkood Satellite (VS), Twister (Twst), Sister, Twister-sister (TS), Hairpin, Hatchet and Pistol, HOV Linc, or variants or fragments thereof that maintain cis-cleavage function. In one embodiment, the 3' ribozyme comprises a lanthanum ribozyme (available at Zhou et al. Human Lantern Ribozymes: Smallest Known Self-cleaving Ribozymes, 07 March 2023, PREPRINT (Version 1) available at Research Square [doi.org / 10.21203 / rs.3.rs-2567304 / v1]). However, the present invention is not limited to any particular ribozyme, but rather encompasses all known members of the family of endogenous ribozymes and any potentially artificial ribozymes. In one embodiment, the 3' ribozyme comprises a P1-type twister, a P3-type twister, or a P5-type twister. In one embodiment, the 3' ribozyme comprises a P1-type twister. In one embodiment, the 3' ribozyme comprises a P1-type twister derived from rice (Oryza sativa, Osa). In some embodiments, the 3' ribozyme comprises one or more nucleotide overhangs. In one embodiment, the overhang comprises a nucleotide sequence that hybridizes to an upstream sequence of the 3' ribozyme in a first RNA molecule. In some embodiments, the overhang improves the efficiency of spontaneous cis-cleavage.
[0108] In some embodiments, the second RNA molecule comprises a 5' ribozyme. In some embodiments, the 5' ribozyme is derived from one or more families selected from the group consisting of Hammerhead (HH), Hepatitis Delta Virus (HDV), Barkood Satellite (VS), Twister (Twst), Sister, Twister-sister (TS), Hairpin, Hatchet and Pistol, HOV Linc, or variants or fragments thereof that maintain cis-cleavage function. In one embodiment, the 5' ribozyme comprises a lanthanum ribozyme (available at Zhou et al. Human Lantern Ribozymes: Smallest Known Self-cleaving Ribozymes, 07 March 2023, PREPRINT (Version 1) available at Research Square [doi.org / 10.21203 / rs.3.rs-2567304 / v1]). However, the present invention is not limited to any particular ribozyme, but rather encompasses all known members of the family of endogenous ribozymes and any potentially artificial ribozymes. In one embodiment, the 5' ribozyme comprises a P1-type twister, a P3-type twister, or a P5-type twister. In one embodiment, the 5' ribozyme comprises a P1-type twister. In one embodiment, the 5' ribozyme comprises a P1-type twister derived from rice (Oryza sativa, Osa). In some embodiments, the 5' ribozyme comprises one or more nucleotide overhangs. In one embodiment, the overhang comprises a nucleotide sequence that hybridizes to a downstream sequence of the 5' ribozyme in a second RNA molecule. In some embodiments, the overhang improves the efficiency of spontaneous cis-cleavage.
[0109] In one embodiment, the HDV ribozyme of the composition comprises one or more selected from the group consisting of HDV, HDV68, HDV67, HDV56, genHDV, and anti-HDV, or a variant or fragment thereof. In one embodiment, HDV68 comprises the nucleic acid sequence of SEQ ID NO: 9. In one embodiment, HDV67 comprises the nucleic acid sequence of SEQ ID NO: 10. In one embodiment, HDV56 comprises the nucleic acid sequence of SEQ ID NO: 11. In one embodiment, genHDV comprises the nucleic acid sequence of SEQ ID NO: 12. In one embodiment, anti-HDV comprises the nucleic acid sequence of SEQ ID NO: 13.
[0110] In one embodiment, the HH ribozyme contains one or more nucleotides in a stem 1 overhang that hybridizes with nucleotides in an upstream or downstream sequence of the HH ribozyme. In one embodiment, the number of nucleotides in the stem 1 overhang may be one or more nucleotides, two or more nucleotides, four or more nucleotides, six or more nucleotides, eight or more nucleotides, ten or more nucleotides, twelve or more nucleotides, fourteen or more nucleotides, sixteen or more nucleotides, eighteen or more nucleotides, or twenty or more nucleotides. In one embodiment, the HH ribozyme containing one or more nucleotide stem 1 overhangs contains a nucleic acid sequence selected from the group consisting of SEQ ID NOs: 111, SEQ ID NOs: 112, SEQ ID NOs: 113, SEQ ID NOs: 114, SEQ ID NOs: 115, SEQ ID NOs: 116, SEQ ID NOs: 117, and SEQ ID NOs: 118, where the nucleotide indicated as N corresponds to a nucleotide that hybridizes with a nucleotide in a downstream sequence of the HH ribozyme. In one embodiment, the HH ribozyme has one or more nucleotides in a stem 3 overhang. In one embodiment, the HH ribozyme has a 5-nucleotide stem with a 3-overhang. In one embodiment, the HH ribozyme comprises the nucleic acid sequence of SEQ ID NO: 105, where the nucleotide indicated as N corresponds to a nucleotide that hybridizes with the nucleotide of the upstream sequence of the HH ribozyme. In one embodiment, the HH ribozyme is modified in the stem 2 loop. In one embodiment, the HH ribozyme having the modified stem 2 loop comprises a nucleic acid sequence selected from the group consisting of SEQ ID NO: 119, SEQ ID NO: 120, SEQ ID NO: 121, SEQ ID NO: 122, SEQ ID NO: 123, and SEQ ID NO: 124, where the nucleotide indicated as N corresponds to a nucleotide that hybridizes with the nucleotide of the downstream sequence of the HH ribozyme. In one embodiment, the HH ribozyme is modified in the stem 1 to include a tertiary stabilization motif (TSM). In one embodiment, the HH ribozyme is modified in the stem 2 loop and modified in the stem 1 to include a tertiary stabilization motif (TSM). In one embodiment, the modified HH ribozyme cleaves more efficiently in the cis-section than the HH ribozyme. In one embodiment, the modified HH ribozyme is RzB.In one embodiment, RzB comprises the nucleic acid sequence of sequence number 125, where the nucleotide indicated as N corresponds to a nucleotide that hybridizes with the nucleotide of the downstream sequence of the HH ribozyme.
[0111] In one embodiment, the Twister ribozyme includes the nucleic acid sequence of SEQ ID NO: 32. In one embodiment, the Twister ribozyme includes one or more nucleotides in the P1 stem overhang. In one embodiment, the number of nucleotides in the P1 stem overhang may be 1 or more, 2 or more, 3 or more, 4 or more, or 5 or more. In one embodiment, the Twister ribozyme including one or more nucleotide P1 stem overhangs includes a nucleic acid sequence selected from the group consisting of SEQ ID NO: 106, SEQ ID NO: 107, SEQ ID NO: 108, SEQ ID NO: 109, and SEQ ID NO: 110, where the nucleotide indicated as N corresponds to a nucleotide that hybridizes with a nucleotide in the downstream sequence of the Twister ribozyme.
[0112] In some embodiments, one or more ribozymes in the composition consist of a first portion and a second portion. In some embodiments, the first portion is incorporated into one or more RNA molecules. In some embodiments, the first portion is a ribozyme recognition sequence. In some embodiments, the second portion is introduced separately. In some embodiments, cis-cleavage of the first portion from one or more RNA molecules occurs only when the first and second portions are in contact with each other. In some embodiments, one or more ribozymes are VS ribozymes. In one embodiment, the VS ribozyme contains the nucleic acid sequence of SEQ ID NO: 14. In one embodiment, the first portion is a VS ribozyme stem-loop (VS-S). In one embodiment, VS-S contains the nucleic acid sequence of SEQ ID NO: 15. In one embodiment, the second portion is the remaining portion of VS without the stem-loop (VS-Rz). In one embodiment, VS-Rz contains the nucleic acid sequence of SEQ ID NO: 16.
[0113] Ribozymes are autocatalytic RNAs that cleave in cis to produce unique RNA 3' and 5' ends, as described herein. However, cis-cleaved ribozymes can be manipulated to cleave in trans, thereby enabling nucleotide-specific cleavage of target RNA and obtaining similar RNA ends. In some embodiments, the present invention includes a composition comprising a single nucleic acid molecule encoding a single RNA molecule containing a trans-cleaved ribozyme. In one embodiment, the trans-cleaved ribozyme can trans-cleave a separate RNA molecule. In one embodiment, the trans-cleaved ribozyme recognizes a specific nucleic acid sequence within a separate RNA molecule. In some embodiments, the trans-cleaved ribozyme is targeted to delete a disease-causing mutation. In some embodiments, the disease-causing mutation is located in an exon. In some embodiments, the disease-causing mutation is located in an intron. In some embodiments, the composition comprises two trans-cleaved ribozymes targeting upstream and downstream of a disease-causing mutation. In some embodiments, trans-cleavage upstream and downstream of the disease-causing mutation results in the removal of the disease-causing mutation. In some embodiments, the rest of the gene is trans-spliced together after trans-cleaving the disease-causing mutation. In some embodiments, the trans-spliced gene is expressed as a functional protein.
[0114] As described herein, the 3'P or 2'3'cP and 5'OH ends of an RNA molecule that has undergone ribozyme-mediated cleavage can be ligated together. Thus, isolated RNA sequences encoding distinct portions of a larger full-length protein can be transspliced together in a scarless manner to enable the expression of the full-length protein. In one embodiment, the present invention relates to a composition comprising one or more nucleic acid molecules encoding two or more portions of a protein of interest and encoding one or more ribozymes. In one embodiment, the present invention relates to a composition comprising one or more RNA molecules comprising two or more portions of a protein of interest and containing one or more ribozymes.
[0115] In one embodiment, one or more nucleic acid molecules encoding two or more portions of the target protein include a first nucleic acid molecule encoding a first portion of the target protein and a second nucleic acid molecule encoding a second portion of the target protein. In one embodiment, the first nucleic acid includes a first RNA molecule. In one embodiment, the second nucleic acid includes a second RNA molecule. In one embodiment, the first RNA molecule is ligated to a 3' ribozyme at its 3' end. In one embodiment, the second RNA molecule is ligated to a 5' ribozyme at its 5' end. In one embodiment, upon cis-cleavage of the 3' and 5' ribozyme sequences, the 3'P or 2'3'cP end of the first RNA molecule is ligated to the 5'OH end of the second RNA molecule, thereby generating a single RNA molecule encoding the target full-length protein. In one embodiment, the target full-length protein functions identically to an endogenously expressed full-length protein of the same sequence.
[0116] In one embodiment, the target full-length protein includes therapeutic proteins. In one embodiment, therapeutic proteins are not limited to these, but include eutrophin, dystrophin, dysferlin, myoferin, cystic fibrosis membrane conductance regulator (CFTR), coagulation factor VIII, fibrocystin, retina-specific phospholipid transport ATPase (ABCA4), otoferin, copper transport ATPase 2, MYO7A, MYO15A, CDH23, STRC, OTOG, TECTA, PCDH15, and TRIOB. It comprises one or more selected from the group consisting of P, MYO3A, COL11A2, LOXHD1, PTPRQ, OTOGL, MYH14, MYH9, TNC, CACNA1A, CACNA1C, CACNA1F, CACNA1H, CACNA1G, CACNA1D, CACNA1B, CACNA1S, CACNA1I, CACNA1E, ATP2A1, ATP2A2, Adcy6, FKBP12-rapamycin-binding domain, and Cas9. In one embodiment, the target full-length protein is a recombinase. In one embodiment, the recombinase is one or more selected from the group consisting of CRE recombinase and FLP recombinase, but is not limited to these. In one embodiment, the target full-length protein is a eukaryotic / prokaryotic antibiotic resistance gene product. In one embodiment, the eukaryotic / prokaryotic antibiotic resistance gene product is one or more selected from the group consisting of ampicillin, kanamycin, blastosidine, puromycin, neomycin, and hygromycin. In certain embodiments, the full-length protein of interest is an antibody. In one embodiment, the antibody can bind to the target protein of interest. In some embodiments, the antibody is an antibody fragment, synthetic antibody, nanobody, or a fragment or variant thereof that maintains the ability to bind to the target protein. In one embodiment, the full-length protein of interest includes, but is not limited to, synthetic repeat proteins that constitute hydrogels, synthetic spider silk, and collagen.In one embodiment, the synthetic repeat protein includes, but is not limited to, one or more selected from the group consisting of spidrin, silk, keratin, collagen, elastin, resilin, and squid trefoil proteins, beta-solenoid proteins, zinc finger nucleases (ZFNs), and Tal effector nucleases (TALENs). In one embodiment, the full-length protein of interest includes a toxic or antiviral protein that can inhibit the generation of lentiviral particles in mammalian packing cells. In one embodiment, the toxic protein is a cell suicide gene. In one embodiment, the cell suicide gene includes, but is not limited to, one or more selected from the group consisting of diphtheria toxin A, herpes simplex virus thymidine kinase, lysine, cholera toxin, prion protein, pertussis toxin, ectatomin, conopeptide, abrin, verotoxin, tetanospasmin, botulinum toxin, Pseudomonas aeruginosa exotoxin A, anthrax, saporin, and porkweed antiviral protein. In one embodiment, the antiviral protein includes, but is not limited to, one or more selected from the group consisting of interferon-inducible GTP-binding protein (MxA), myeloperoxidase (MPO), and interferon.
[0117] If the N-terminal or C-terminal RNA molecule encoding a portion of the target protein can be translated before ribozyme-mediated cleavage or expressed separately, it can potentially lead to undesirable or truncated protein expression. However, this undesirable expression can be limited by utilizing proteolytic translational regulatory sequences. In one embodiment, one or more RNA molecules in the composition include nucleic acid sequences encoding proteolytic translational regulatory sequences. In one embodiment, the first RNA molecule includes nucleic acid sequences encoding proteolytic translational regulatory sequences. In one embodiment, the second RNA molecule includes nucleic acid sequences encoding proteolytic translational regulatory sequences. In some embodiments, the proteolytic translational regulatory sequences prevent partial expression of the protein before ribozyme cleavage and splicing. In some embodiments, the translational control sequence for proteolysis includes one or more selected from the group consisting of hCL1-PEST sequences, E1A-PEST sequences, removal of the poly(A) sequence of a nucleic acid, pseudo-translation generating a polyK tail through a polyA tail, deletion of an ATG stop codon, silent mutations in the N-terminal NTG codon, the 5'UTR of a yeast GCN4 sequence encoding four small upstream ORFs that function as translation inhibitors, and small internal fragments of the 5'UTR of a yeast GCN4 sequence. In some embodiments, the translational control sequence for proteolysis includes one or more nucleic acid sequences selected from the group consisting of SEQ ID NOs: 43, 44, 45, 46, 47, 48, 49, 77, 79, and 104. In some embodiments, the translational control sequence for proteolysis includes one or more amino acid sequences selected from the group consisting of SEQ ID NOs: 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 76, 78, and 80.
[0118] In certain embodiments, RNA nuclear localization signals may be useful in preventing cytosolic transport and translation of unsplicing RNA molecules in order to further prevent undesirable or truncated protein expression. In one embodiment, one or more RNA molecules of the composition include nucleic acid sequences encoding an RNA nuclear localization sequence. In one embodiment, the first RNA molecule includes a nucleic acid sequence encoding an RNA nuclear localization sequence. In one embodiment, the second RNA molecule includes a nucleic acid sequence encoding an RNA nuclear localization sequence. In one embodiment, the RNA nuclear localization sequence prevents the transport of cytoplasmic RNA and the translation of partial proteins before cleavage and splicing of the ribozyme sequence. In one embodiment, the RNA nuclear localization sequence includes one or more nucleic acid sequences selected from the group consisting of SEQ ID NOs: 50 and SEQ ID NOs: 51.
[0119] In some embodiments, the composition further comprises one or more additional RNA molecules, each additional RNA molecule comprising a coding region encoding the protein of interest, a 5' ribozyme, and a 3' ribozyme domain. In some embodiments, the system further comprises one or more additional nucleic acid molecules encoding one or more additional RNA molecules, each additional RNA molecule comprising a coding region encoding the domain of the protein of interest, a 5' ribozyme, and a 3' ribozyme.
[0120] In some embodiments, the composition further comprises one or more additional RNA molecules, each additional RNA molecule comprising a coding region encoding the domain of the protein of interest, a 5' ribozyme, and a 3' ribozyme recognition sequence. In some embodiments, the system further comprises one or more additional nucleic acid molecules encoding one or more additional RNA molecules, each additional RNA molecule comprising a coding region encoding the domain of the protein of interest, a 5' ribozyme, and a 3' ribozyme recognition sequence.
[0121] Pre-mRNA splicing by spliceosomes has been shown to enhance mRNA translation by depositing factors that promote the first round of translation, or by promoting RNA processing and transport to the cytoplasm. The addition of chimeric cis-splicing introns within a transgene has also been shown to enhance transgene protein expression. Therefore, in certain embodiments, the addition of splice donor and splice acceptor sites recognized by spliceosomes and cis-spliced may enhance protein expression from a split precursor RNA molecule. In one embodiment, the composition comprises one or more RNA molecules containing a splice donor sequence or a splice acceptor sequence. In one embodiment, the first RNA molecule of the composition contains a splice donor sequence. In one embodiment, the splice donor sequence is ligated to the 3' end of the first RNA molecule following a ribozyme sequence. In one embodiment, the second RNA molecule of the composition contains a splice acceptor sequence. In one embodiment, the splice acceptor sequence is ligated to the 5' end of the second RNA molecule prior to a ribozyme sequence. In one embodiment, the inclusion of splice donor and splice acceptor sequences enhances protein expression after ribozyme-mediated transsplicing.
[0122] Due to the three open reading frames through which the protein is translated, ribozyme-mediated trans-splicing and the expression of multiple different functional proteins may also be possible. By utilizing this capability, functional proteins can be generated using trans-splicing of RNA in three different mismatched open reading frames. In one embodiment, the composition of the present invention comprises at least four nucleic acid molecules, each comprising at least two pairs of nucleic acid molecules. In one embodiment, each pair of nucleic acid molecules encodes at least two portions of the protein of interest and encodes at least two ribozymes. In one embodiment, the composition comprises at least four RNA molecules, each comprising at least two pairs of RNA molecules. In one embodiment, each pair of RNA molecules encodes at least two portions of the protein of interest and encodes at least two ribozymes.
[0123] In one embodiment, at least two pairs of RNA molecules include a first pair of RNA molecules and a second pair of RNA molecules. In one embodiment, the first pair of RNA molecules includes a first RNA molecule and a second RNA molecule. In one embodiment, the second pair of RNA molecules includes a third RNA molecule and a fourth RNA molecule. In some embodiments, the third and fourth RNA molecules have different open reading frames than the first and second RNA molecules, and upon spontaneous cis-cleavage, ligation between either the first or second RNA molecule and either the third or fourth RNA molecule is not able to translate the full-length functional protein product.
[0124] In one embodiment, at least two pairs of RNA molecules further comprise a third pair of RNA molecules. In one embodiment, the third pair of RNA molecules comprises a fifth RNA molecule and a sixth RNA molecule. In some embodiments, the fifth and sixth RNA molecules have different open reading frames than the first and second pairs of RNA molecules, and the full-length functional protein product can be translated only by ligation of the first, second, or third pair of RNA molecules upon spontaneous cis-cleavage.
[0125] Ribozyme-mediated trans-splicing between two independent RNAs can occur, as described herein, when one RNA contains a 3' ribozyme and the other contains a 5' ribozyme. However, when transcribed in cis within the same RNA molecule, the two ribozymes can mediate their own scarless removal. This approach similarly generates two independent RNAs with 3'-P and 5' OH ends, which can undergo trans-splicing and translation within the cell. Including a cargo sequence between the 3' and 5' ribozymes also gives rise to the possibility of generating a circularized RNA molecule during ligation.
[0126] In one embodiment, the present invention relates to a composition comprising a single nucleic acid molecule encoding two or more portions of a target protein and encoding one or more ribozymes. In one embodiment, the present invention relates to a composition comprising a single RNA molecule encoding two or more portions of a target protein and comprising one or more ribozymes.
[0127] In one embodiment, a single nucleic acid molecule encodes a first portion of RNA, a synthetic intron, and a second portion of RNA. In one embodiment, the synthetic intron includes a 5' ribozyme and a 3' ribozyme. In one embodiment, the first portion of RNA encodes a first portion of the protein of interest. In one embodiment, the second portion of RNA encodes a second portion of the protein of interest. In one embodiment, the single nucleic acid contains a sequence linked in the order (first portion of RNA encoding the first portion of the protein of interest) - (5' ribozyme of the synthetic intron) - (3' ribozyme of the synthetic intron) - (second portion of RNA encoding the second portion of the protein of interest). In one embodiment, the first portion of the protein of interest is the N-terminal portion of GFP. In one embodiment, the 5' ribozyme of the synthetic intron includes HDV. In one embodiment, the first portion of the RNA and the 5' ribozyme of the synthetic intron comprise the nucleic acid sequence of SEQ ID NO: 127, where lowercase letters indicate the 5' ribozyme sequence and uppercase letters indicate the sequence encoding the N-terminal portion of GFP (see Example 4, "GFP with internally synthesized ribozyme introns with and without cargo"). In one embodiment, the second portion of the protein of interest is the C-terminal portion of GFP. In one embodiment, the 3' ribozyme of the synthetic intron comprises HH. In one embodiment, the second portion of the RNA and the 3' ribozyme of the synthetic intron comprise the nucleic acid sequence of SEQ ID NO: 128, where lowercase letters indicate the 3' ribozyme sequence and uppercase letters indicate the sequence encoding the C-terminal portion of GFP (see Example 4, "GFP with internally synthesized ribozyme introns with and without cargo").
[0128] In one embodiment, the synthetic intron includes a cargo sequence positioned between the 5' ribozyme and the 3' ribozyme. In one embodiment, the single nucleic acid includes a sequence concatenated in the order of (first portion of RNA encoding the first portion of the protein of interest)-(5' ribozyme of the synthetic intron)-(cargo sequence)-(3' ribozyme of the synthetic intron)-(second portion of RNA encoding the second portion of the protein of interest).
[0129] In one embodiment, the 5' ribozyme sequence of the synthetic intron does not require a bilateral flanking sequence for activity. In one embodiment, the circular RNA generated from terminal ligation of a synthetic intron containing a 5' ribozyme sequence that does not require a bilateral flanking sequence for activity can exist in both circular and re-cleaved linear forms. In one embodiment, the ribozyme sequence is an HDV ribozyme.
[0130] In one embodiment, the 5' ribozyme sequence of the synthetic intron requires a bilateral flanking sequence for activity. In one embodiment, the circular RNA generated from terminal ligation of a synthetic intron containing a 5' ribozyme sequence requiring a bilateral flanking sequence for activity can only exist in a circular form. In one embodiment, the ribozyme sequence is an HH ribozyme.
[0131] In one embodiment, the 5' ribozyme sequence of the synthetic intron is a ribozyme recognition sequence. In one embodiment, the ribozyme recognition sequence requires the addition of a trans-cleaved ribozyme for inducible cleavage. In one embodiment, the ribozyme recognition sequence comprises VS-S. In some embodiments, VS-S is encoded by a nucleic acid sequence comprising SEQ ID NO: 15. In one embodiment, the trans-cleaved ribozyme comprises VS-Rz. In some embodiments, VS-Rz is encoded by a nucleic acid sequence comprising SEQ ID NO: 16.
[0132] In one embodiment, self-cleavage of the 5' and 3' ribozyme sequences generates three distinct RNA molecules: 1) a first fragment containing a first portion of RNA encoding a first portion of the target protein; 2) a second fragment containing a synthetic intron; and 3) a third fragment containing a second portion of RNA encoding a second portion of the target protein. In one embodiment, the matching ends of the second fragment are ligated to generate a circular RNA molecule containing a synthetic intron with a cargo sequence. In another embodiment, the first and third fragments are ligated together to generate a single full-length linear RNA molecule.
[0133] In one embodiment, the cargo sequence of the synthetic intron is one or more selected from the group consisting of a sequence encoding a therapeutic protein of interest, a CRISPR guide RNA sequence, a small RNA sequence, and a trans-cleaved ribozyme sequence. In one embodiment, the small RNA sequence includes one or more selected from the group consisting of microRNA (miRNA), Piwi-interacting RNA (piRNA), small interfering RNA (siRNA), small nucleolar RNA (snoRNA), small tRNA-derived RNA (tsRNA), small rDNA-derived RNA (srRNA), and small nuclear RNA (snRNA).
[0134] In one embodiment, a single full-length linear RNA molecule encodes the full-length protein of interest. In one embodiment, the full-length protein of interest is a therapeutic protein. In one embodiment, the therapeutic protein is not limited to, but may include, eutrophin, dystrophin, dysferlin, myoferin, cystic fibrosis membrane conductance regulator (CFTR), coagulation factor VIII, fibrocystin, retina-specific phospholipid transport ATPase (ABCA4), otoferin, copper transport ATPase 2, MYO7A, MYO15A, CDH23, STRC, OTOG, TECTA, PCDH15, and TRIOBP. It may be one or more selected from the group consisting of MYO3A, COL11A2, LOXHD1, PTPRQ, OTOGL, MYH14, MYH9, TNC, CACNA1A, CACNA1C, CACNA1F, CACNA1H, CACNA1G, CACNA1D, CACNA1B, CACNA1S, CACNA1I, CACNA1E, ATP2A1, ATP2A2, Adcy6, FKBP12-rapamycin-binding domain, and Cas9. In one embodiment, the target full-length protein is a recombinase. In one embodiment, the recombinase is one or more selected from the group consisting of CRE recombinase and FLP recombinase, but is not limited to these. In one embodiment, the target full-length protein is a eukaryotic / prokaryotic antibiotic resistance gene product. In one embodiment, the eukaryote / prokaryote antibiotic resistance gene product is one or more selected from the group consisting of ampicillin, kanamycin, blastosidine, puromycin, neomycin, and hygromycin. In one embodiment, the full-length protein of interest is a reporter protein. In one embodiment, the reporter protein is one or more selected from the group consisting of green fluorescent protein (GFP), red fluorescent protein (RFP), and luciferase (Luc). In one embodiment, the reporter protein is used as a proxy indicator for evaluating the delivery and expression of cargo sequences. In certain embodiments, the full-length protein of interest is an antibody. In one embodiment, the antibody can bind to the target protein of interest.In some embodiments, the antibody is an antibody fragment, synthetic antibody, nanobody, or a fragment or variant thereof that maintains the ability to bind to a target protein. In one embodiment, the full-length protein of interest comprises a toxic or antiviral protein that can inhibit the generation of lentiviral particles in mammalian packing cells. In one embodiment, the toxic protein is a cell suicide gene. In one embodiment, the cell suicide gene comprises, but is not limited to, one or more selected from the group consisting of diphtheria toxin A, herpes simplex virus thymidine kinase, lysine, cholera toxin, prion protein, pertussis toxin, ectatomin, conopeptide, abrin, verotoxin, tetanospasmin, botulinum toxin, Pseudomonas aeruginosa exotoxin A, anthrax, saporin, and porkweed antiviral protein. In one embodiment, the antiviral protein comprises, but is not limited to, one or more selected from the group consisting of interferon-inducible GTP-binding protein (MxA), myeloperoxidase (MPO), and interferon.
[0135] In certain embodiments, the techniques of the present invention can be used to assemble a full-length RNA viral genome. In one embodiment, one or more nucleic acid molecules encoding one or more ribozymes of the present invention encode one or more portions of an RNA viral genome. In one embodiment, one or more RNA molecules containing one or more ribozymes of the present invention contain one or more portions of an RNA viral genome.
[0136] In one embodiment, one or more nucleic acid molecules include a first nucleic acid molecule encoding a first portion of the RNA virus genome and encoding a 3' ribozyme. In one embodiment, one or more nucleic acid molecules include a second nucleic acid molecule encoding a second portion of the RNA virus genome and a 5' ribozyme. In one embodiment, one or more RNA molecules include a first RNA molecule containing a first portion of the RNA virus genome and a 3' ribozyme. In one embodiment, one or more RNA molecules include a second RNA molecule containing a second portion of the RNA virus genome and a 5' ribozyme. In one embodiment, the composition includes a ligase or a nucleic acid encoding a ligase. In one embodiment, upon cis-cleavage of the 3' and 5' ribozymes, the first portion of the RNA virus genome and the second portion of the RNA virus genome are ligated together, thereby producing a full-length RNA virus genome. Exemplary RNA viruses include, but are not limited to, coronaviruses, paramyxoviruses, orthomyxoviruses, retroviruses, lentiviruses, alphaviruses, flaviviruses, rhabdoviruses, measles viruses, Newcastle disease viruses, and picornaviruses.
[0137] In some embodiments, the present invention comprises a composition comprising a nucleic acid encoding a ligase. In some embodiments, the ligase mediates ligation of the 3'P or 2'3'cP terminus and the 5'OH terminus. In some embodiments, the ligase is an RNA 2',3'-cyclic phosphate and 5'-OH (RtcB) ligase. In some embodiments, the RtcB ligase is derived from one or more domains of an organism selected from the group consisting of eukaryotes, bacteria, and archaea. In some embodiments, the organism is selected from humans, Escherichia coli (E. coli), Deinococcus radiodurans, Pyrococcus horikoshii, Pyrococcus sp. ST04 strain, and Thermococcus sp. EP strain. In some embodiments, the nucleic acid sequence encoding the ligase is one or more sequences selected from the group consisting of SEQ ID NOs: 82, 84, 86, 88, 90, and 92. In some embodiments, the nucleic acid sequence encoding the ligase encodes one or more amino acid sequences selected from the group consisting of SEQ ID NOs: 81, 83, 85, 87, 89, and 91.
[0138] nucleic acid In some embodiments, one or more nucleic acids of the present invention include nucleic acid sequences substantially homologous to the nucleic acid sequences described herein. For example, in some embodiments, the nucleic acids have a degree of identity of at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% with respect to the original nucleic acid sequence.
[0139] In some embodiments, one or more nucleic acids of the present invention include nucleic acid sequences that are part of the nucleic acid sequences described herein. For example, in some embodiments, the nucleic acid has a length of at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% of the original nucleic acid sequence.
[0140] In some embodiments, one or more nucleic acids of the present invention include nucleic acid sequences that are part of a nucleic acid sequence described herein and are substantially homologous to a nucleic acid sequence described herein. For example, in some embodiments, the nucleic acid has a degree of identity of at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% with respect to the original nucleic acid sequence. and / or having a length of at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% of the original nucleic acid sequence.
[0141] The nucleic acids of the present invention may include, but are not limited to, DNA and RNA, any type of nucleic acid. For example, in one embodiment, the composition includes an isolated DNA molecule (e.g., an isolated cDNA molecule) encoding the fusion protein of the present invention. In one embodiment, the composition includes an isolated RNA molecule encoding the fusion protein of the present invention or a functional fragment thereof.
[0142] The nucleic acid molecules of the present invention can be modified to improve their stability in serum or growth media for cell culture. Modifications can be made to enhance stability, functionality, and / or specificity, and to minimize the immunostimulatory properties of the nucleic acid molecules of the present invention. For example, to enhance stability, 3' residues can be stabilized against degradation, and for example, they can be selected to consist of purine nucleotides, particularly adenosine or guanosine nucleotides. Alternatively, substitution of pyrimidine nucleotides with modification analogs, such as substitution of uridine with 2'-deoxythymidine, is acceptable and does not affect the function of the molecule.
[0143] In one embodiment of the present invention, a nucleic acid molecule may contain at least one modified nucleotide analog. For example, the terminal can be stabilized by incorporating a modified nucleotide analog.
[0144] Non-limiting examples of nucleotide analogs include sugar- and / or skeletal-modified ribonucleotides (i.e., modifications to the phosphate-sugar skeleton). For example, phosphodiester links in native RNA may be modified to include at least one nitrogen or sulfur heteroatom. In exemplary skeletal-modified ribonucleotides, the phosphoester group bound to an adjacent ribonucleotide is replaced by a modifying group, such as a phosphorothioate group. In exemplary sugar-modified ribonucleotides, the 2'OH group is replaced by a group selected from H, OR, R, halo, SH, SR, NH2, NHR, NR2, or ON, where R is a C1-C6 alkyl, alkenyl, or alkynyl, and halo is F, Cl, Br, or I.
[0145] Other examples of modification include ribonucleotides, i.e., ribonucleotides, which contain at least one non-naturally occurring nucleotide base in place of a naturally occurring nucleotide base. The bases may be modified to block the activity of adenosine deaminase. Exemplary modified nucleotide bases, though not limited to these, include uridine and / or cytidine modified at position 5, e.g., 5-(2-amino)propyluridine, 5-bromouridine; adenosine and / or guanosine modified at position 8, e.g., 8-bromoguanosine; deazanucleotides, e.g., 7-deaza-adenosine; and O- and N-alkylated nucleotides, e.g., N6-methyladenosine. Combinations of the above modifications may also be used.
[0146] In some examples, nucleic acid molecules include at least one of the following chemical modifications: 2'-H, 2'-O-methyl, or 2'-OH modifications of one or more nucleotides. In certain embodiments, nucleic acid molecules of the present invention can have enhanced resistance to nucleases. To enhance nuclease resistance, nucleic acid molecules may include, for example, 2'-modified ribose units and / or phosphorothioate linkages. For example, the 2'-hydroxyl group (OH) may be modified or substituted with multiple different "oxy" or "deoxy" substituents. To enhance nuclease resistance, nucleic acid molecules of the present invention may include 2'-O-methyl, 2'-fluorine, 2'-O-methoxyethyl, 2'-O-aminopropyl, 2'-amino, and / or phosphorothioate linkages. Binding affinity to targets can also be increased by including locked nucleic acids (LNA), ethylene nucleic acids (ENA), e.g., 2'-4'-ethylene-bridged nucleic acids, and specific nucleic acid base modifications such as 2-amino-A, 2-thio (e.g., 2-thio-U), and G-clamp modifications.
[0147] In one embodiment, the nucleic acid molecule may be a 2'-modified nucleotide, such as 2'-deoxy, 2'-deoxy-2'-fluoro, 2'-O-methyl, 2'-O-methoxyethyl (2'-O-MOE), 2'-O-aminopropyl (2'-O-AP), 2'-O-dimethylaminoethyl (2'-O-DMAOE), 2'-O-dimethylaminopropyl (2'-O-DMAP), 2'-O-dimethylaminoethyloxyethyl (2'-O-DMAEOE), or 2'-ON-methylacetamide (2'-O-NMA). In one embodiment, the nucleic acid molecule contains at least one 2'-O-methyl modified nucleotide, and in some embodiments, all nucleotides of the nucleic acid molecule contain 2'-O-methyl modification.
[0148] In certain embodiments, the nucleic acid molecule of the present invention has one or more of the following characteristics.
[0149] Nucleic acid agents discussed herein include unmodified RNA and DNA, as well as modified RNA and DNA, for example, to improve efficacy, and polymers of nucleoside substitutes. Unmodified RNA refers to molecules in which the components of nucleic acid, namely sugars, bases, and phosphate groups, are the same as or essentially the same as those found in nature, or the same as those found in nature in the human body. In the art, rare or abnormal but naturally occurring RNA is referred to as modified RNA. See, for example, Limbach et al. (Nucleic Acids Res., 1994, 22:2183-2196). Such rare or abnormal RNAs, often referred to as modified RNA, are typically the result of post-transcriptional modification and fall within the scope of the term unmodified RNA as used herein. Modified RNA, as used herein, refers to molecules in which one or more components of nucleic acid, namely sugars, bases, and phosphate groups, are different from those found in nature, or from those found in the human body. They are called “modified RNA,” but naturally include molecules that are not strictly RNA due to the modification. Nucleoside substitutes are molecules in which the ribophosphate skeleton is replaced with a non-ribophosphate construct that allows the bases to be presented in the correct spatial relationships, so that hybridization occurs substantially similarly to that seen in the ribophosphate skeleton, for example, an uncharged mimic of the ribophosphate skeleton.
[0150] The nucleic acid modification of the present invention may be present in one or more of the following: a phosphate group, a sugar group, a skeleton, the N-terminus, the C-terminus, or a nucleic acid base.
[0151] vector The present invention also includes a composition comprising one or more vectors into which one or more nucleic acid molecules of the present invention are inserted. In one embodiment, the vector encodes at least two RNA molecules. In one embodiment, the vector comprises at least two RNA molecules. In some embodiments, at least two RNA molecules are encoded by the same vector. In some embodiments, at least two RNA molecules are contained within the same vector. In one embodiment, the at least two RNA molecules comprise a first RNA molecule and a second RNA molecule.
[0152] In some embodiments, the present invention comprises at least two vectors encoding at least two RNA molecules. In some embodiments, the at least two vectors comprise at least two RNA molecules. In some embodiments, the at least two vectors encode distinct RNA molecules. In some embodiments, the at least two vectors comprise distinct RNA molecules. In some embodiments, the at least two distinct RNA molecules comprise a first RNA molecule and a second RNA molecule. In some embodiments, the first RNA molecule is encoded by a first vector, and the second RNA molecule is encoded by a second vector. In some embodiments, the first vector comprises a first RNA molecule, and the second vector comprises a second RNA molecule.
[0153] In some embodiments, the present invention further comprises a vector encoding one or more additional RNA molecules. In some embodiments, the present invention further comprises one or more vectors containing one or more additional RNA molecules. In some embodiments, each additional RNA molecule comprises a coding region encoding the protein of interest, the 5' ribozyme and the 3' ribozyme domains. In some embodiments, each additional RNA molecule comprises a coding region encoding the domain of the protein of interest, the 5' ribozyme and the 3' ribozyme recognition sequence.
[0154] The art is filled with suitable vectors useful in the present invention. In summary, the expression of native or synthetic nucleic acids encoding the fusion protein of the present invention is typically achieved by operably ligating the nucleic acid encoding the fusion protein or a portion thereof to a promoter and incorporating the construct into an expression vector. The vectors used are suitable for replication and optional incorporation in eukaryotic cells. Typical vectors include transcriptional and translational terminators, start sequences, and promoters useful for regulating the expression of the desired nucleic acid sequence.
[0155] The vectors of the present invention may also be used in nucleic acid immunization and gene therapy using standard gene delivery protocols. Methods for gene delivery are known in the art. For example, they are described in U.S. Patents 5,399,346, 5,580,859 and 5,589,466, which are incorporated herein by reference in their entirety. In another embodiment, the present invention provides a gene therapy vector.
[0156] The isolated nucleic acids of the present invention can be cloned into several types of vectors. For example, the nucleic acids can be cloned into vectors including, but not limited to, plasmids, phagemids, phage derivatives, animal viruses, and cosmids. Vectors for specific purposes include expression vectors, replication vectors, probe generation vectors, and sequencing vectors.
[0157] Furthermore, vectors can be delivered to cells in the form of viral vectors. Viral vector technology is well known in the art and is described, for example, in Sambrook et al. (2012, Molecular Cloning: A Laboratory Manual, Cold Spring Harbor Laboratory, New York), as well as in other virology and molecular biology manuals. Useful viruses as vectors include, but are not limited to, retroviruses, adenoviruses, adeno-associated viruses, herpesviruses, and lentiviruses. Generally, a suitable vector includes a functional origin of replication in at least one organism, a promoter sequence, a convenient restriction endonuclease site, and one or more selection markers (e.g., International Publication No. 01 / 96584, International Publication No. 01 / 29058, and U.S. Patent No. 6,326,193).
[0158] Furthermore, several additional virus-based systems have been developed for gene delivery into mammalian cells. For example, retroviruses provide a convenient platform for gene delivery systems. Selected genes can be inserted into vectors and packaged into retroviral particles using techniques known in the art. Recombinant viruses can then be isolated and delivered to target cells either in vivo or ex vivo. Numerous retroviral systems are known in the art. In some embodiments, adenovirus vectors are used. Several adenovirus vectors are known in the art.
[0159] In one embodiment, the composition comprises a vector derived from adeno-associated virus (AAV). The term “AAV vector” means a vector derived from adeno-associated virus serotypes, for example, but not limited to, AAV-1, AAV-2, AAV-3, AAV-4, AAV-5, AAV-6, AAV-7, AAV-8, and AAV-9. AAV vectors have become a potent means of gene delivery for treating a variety of disorders. AAV vectors possess several characteristics that make them ideally suited for gene therapy, including lack of pathogenicity, minimal immunogenicity, and the ability to transduce postmittal cells in a stable and efficient manner. The expression of specific genes contained within an AAV vector can specifically target one or more types of cells by selecting an appropriate combination of AAV serotype, promoter, and delivery method.
[0160] AAV vectors may have deletions in one or more AAV wild-type genes, preferably all or part of the rep and / or cap genes, but may retain functional facile ITR sequences. Despite high homology, different serotypes have tropisms to different tissues. The receptor for AAV1 is unknown, but AAV1 is known to transduce skeletal muscle and cardiac muscle more efficiently than AAV2. Since most studies have been performed using pseudotyping vectors in which vector DNA facile to the AAV2 ITR is packaged in a capsid of an alternative serotype, it is clear that the biological differences are related to the capsid rather than the genome. Recent evidence has shown that DNA expression cassettes packaged in an AAV1 capsid are at least 1 log 10 more efficient in transduction into cardiomyocytes than those packaged in an AAV2 capsid. In one embodiment, the viral delivery system is an adeno-associated virus delivery system. Adeno-associated viruses may be serotype 1 (AAV1), serotype 2 (AAV2), serotype 3 (AAV3), serotype 4 (AAV4), serotype 5 (AAV5), serotype 6 (AAV6), serotype 7 (AAV7), serotype 8 (AAV8), or serotype 9 (AAV9).
[0161] Desirable AAV fragments for assembly into vectors include cap proteins containing vp1, vp2, vp3 and the hypervariable region, rep proteins containing rep78, rep68, rep52, and rep40, and sequences encoding these proteins. These fragments can be readily utilized in a variety of vector systems and host cells. Such fragments can be used alone, in combination with other AAV serotype sequences or fragments, or in combination with elements from other AAV or non-AAV viral sequences. As used herein, artificial AAV serotypes include, but are not limited to, AAVs having non-natural capsid proteins. Such artificial capsids can be generated by any suitable technique using selected AAV sequences (e.g., fragments of the vp1 capsid protein) in combination with heterologous sequences that may be obtained from different selected AAV serotypes, discontinuous portions of the same AAV serotype, non-AAV viral sources, or non-viral sources. Artificial AAV serotypes may be, but are not limited to, chimeric AAV capsids, recombinant AAV capsids, or "humanized" AAV capsids. Therefore, exemplary AAVs or artificial AAVs suitable for the expression of one or more proteins include, in particular, AAV2 / 8 (U.S. Patent No. 7,282,199), AAV2 / 5 (available from the National Institutes of Health), AAV2 / 9 (International Patent Publication WO2005 / 033321), AAV2 / 6 (U.S. Patent No. 6,156,303), and AAVrh8 (International Patent Publication WO2003 / 042397).
[0162] In one embodiment, the composition comprises a lentiviral vector for delivering one or more nucleic acids of the present invention. In one embodiment, the present invention comprises a lentiviral vector comprising one or more RNA molecules encoding one or more proteins of interest. For example, retrovirus-derived vectors such as lentiviruses are suitable tools for achieving long-term gene transfer because they allow for the long-term and stable integration of the transgene and its proliferation in daughter cells. Lentiviral vectors have an additional advantage over vectors derived from onco-retroviruses such as mouse leukemia virus in that they can transduce non-proliferating cells such as hepatocytes. They also have the additional advantage of low immunogenicity.
[0163] In certain embodiments, the vector also includes conventional regulatory elements operably ligated to the transgene in a manner that enables transcription, translation, and / or expression in cells transfected with the plasmid vector or infected with a virus produced by the present invention. As used herein, “operably ligated” sequences include both expression regulatory sequences adjacent to the gene of interest and expression regulatory sequences acting trans or at any distance to control the gene of interest. Expression regulatory sequences include appropriate transcription start, termination, promoter, and enhancer sequences, efficient RNA processing signals such as splicing and polyadenylation (poly-A) signals, sequences that stabilize cytoplasmic mRNA, sequences that enhance translation efficiency (i.e., Kozak consensus sequences), sequences that enhance protein stability, and, if desired, sequences that enhance the secretion of the encoded product. Numerous expression regulatory sequences, including promoters that are innate, constitutive, inducible, and / or tissue-specific, are known and available in the art.
[0164] Further promoter elements, such as enhancers, regulate the frequency of transcription initiation. Typically, these are located 30–110 bp upstream of the initiation site, although some promoters have recently been shown to also contain functional elements downstream of the initiation site. The spacing between promoter elements is often flexible, resulting in conserved promoter function if elements are reversed or moved relative to one another. In thymidine kinase (TK) promoters, the spacing between promoter elements can be widened to 50 bp before activity begins to decline. Depending on the promoter, individual elements appear to function cooperatively or independently to activate transcription.
[0165] A suitable example of a suitable promoter is the pre-early cytomegalovirus (CMV) promoter sequence. This promoter sequence is a potent constitutive promoter sequence that can drive high levels of expression of any polynucleotide sequence operably ligated to it. Another example of a suitable promoter is elongation growth factor-1α (EF-1α). However, other constitutive promoter sequences may also be used, including but not limited to human gene promoters such as the monkey virus 40 (SV40) early promoter, mouse mammary tumor virus (MMTV), human immunodeficiency virus (HIV) long-term repeat (LTR) promoter, MoMuLV promoter, avian leukemia virus promoter, Epstein-Barr virus pre-early promoter, Roussarcoma virus promoter, as well as actin promoters, myosin promoters, hemoglobin promoters, and creatine kinase promoters. Furthermore, the present invention should not be limited to the use of constitutive promoters. Inducible promoters are also contemplated as part of the present invention. The use of inducible promoters provides a molecular switch that can turn on the expression of a polynucleotide sequence operably ligated to it when such expression is desired, or turn off the expression when expression is not desired. Examples of inductive promoters include, but are not limited to, the metallothione promoter, glucocorticoid promoter, progesterone promoter, and tetracycline promoter.
[0166] Enhancer sequences found on a vector also regulate the expression of the genes contained therein. Typically, enhancers bind to protein factors to enhance gene transcription. Enhancers may be located upstream or downstream of the gene they regulate. Enhancers may also be tissue-specific to enhance transcription in a particular cell or tissue type. In one embodiment, the vector of the present invention includes one or more enhancers to promote the transcription of genes present within the vector.
[0167] To determine the expression of the fusion protein of the present invention, the expression vector introduced into cells may also include either or both a selection marker gene and / or a reporter gene to facilitate the identification and selection of expression cells from a population of cells to be transfected or infected via a viral vector. In other embodiments, the selection marker may be supported on a separate DNA fragment and used in a co-transfection procedure. Both the selection marker and the reporter gene may be flanked by appropriate regulatory sequences to enable expression in host cells. Useful selection markers include, for example, antibiotic resistance genes such as neo.
[0168] Reporter genes are used to identify potentially transfected cells and evaluate the functionality of regulatory sequences. Generally, a reporter gene is a gene encoding a polypeptide that is not present in or expressed by the recipient organism or tissue, and whose expression is revealed by some readily detectable characteristic, such as enzymatic activity. Reporter gene expression is assayed at a suitable time after the DNA has been introduced into recipient cells. Suitable reporter genes may include those encoding luciferase, β-galactosidase, chloramphenicol acetyltransferase, secreted alkaline phosphatase, or the green fluorescent protein gene (e.g., Ui-Tei et al., 2000 FEBS Letters 479:79-82). Suitable expression systems are well known and can be prepared using known techniques or are commercially available. Generally, a construct having a minimum 5' flanking region that exhibits the highest level of reporter gene expression is identified as a promoter. Such a promoter region can be ligated to the reporter gene and used to evaluate the agonist's ability to regulate promoter-driven transcription.
[0169] protein In some embodiments, the present invention comprises a composition comprising a ligase. In some embodiments, the ligase mediates ligation between the 3'P or 2'3'cP end of an RNA molecule and the 5'OH end of an RNA molecule. In some embodiments, the ligase is an RNA 2',3'-cyclic phosphate and 5'-OH (RtcB) ligase. In some embodiments, the RtcB ligase is derived from one or more domains of an organism selected from the group consisting of eukaryotes, bacteria, and archaea. In some embodiments, the organism is selected from humans, Escherichia coli (E. coli), Deinococcus radiodurans, Pyrococcus horikoshii, Pyrococcus sp. ST04 strain, and Thermococcus sp. EP strain. In some embodiments, the ligase comprises one or more amino acid sequences selected from the group consisting of SEQ ID NOs: 81, 83, 85, 87, 89, and 91.
[0170] In some embodiments, one or more proteins of the present invention include an amino acid sequence substantially homologous to the amino acid sequences described herein. For example, in some embodiments, the protein has at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% identity with respect to the original amino acid sequence.
[0171] In some embodiments, one or more proteins of the present invention include an amino acid sequence that is part of an amino acid sequence described herein. For example, in some embodiments, the protein has a length of at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% of the original amino acid sequence.
[0172] In some embodiments, one or more proteins of the present invention include an amino acid sequence that is part of an amino acid sequence described herein and is substantially homologous to an amino acid sequence described herein. For example, in some embodiments, the protein includes at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% homologous to the original amino acid sequence. Having a degree of uniformity and / or having a length of at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% of the original amino acid sequence.
[0173] Pharmaceutical composition The present invention also includes the use of a pharmaceutical composition or a salt thereof of the present invention for carrying out the method of the present invention. Such a pharmaceutical composition may consist of at least one nucleic acid or a salt thereof of the present invention in a form suitable for administration to a subject, or the pharmaceutical composition may comprise at least one nucleic acid or a salt thereof of the present invention, and one or more pharmaceutically acceptable carriers, one or more additional components, or any combination thereof. The nucleic acids of the present invention may be present in the pharmaceutical composition in the form of a physiologically acceptable salt, for example, in combination with a physiologically acceptable cation or anion, as is well known in the art.
[0174] In one embodiment, a pharmaceutical composition useful for carrying out the method of the present invention may be administered to deliver a dose of 1 ng / kg / day to 100 mg / kg / day. In another embodiment, a pharmaceutical composition useful for carrying out the present invention may be administered to deliver a dose of 1 ng / kg / day to 500 mg / kg / day.
[0175] The relative amounts of the active ingredient, pharmaceutically acceptable carrier, and any additional ingredients in the pharmaceutical composition of the present invention vary depending on the characteristics, physique, and condition of the subject being treated, and further depending on the route through which the composition is administered. For example, the composition may contain 0.1% to 100% (w / w) of the active ingredient.
[0176] Pharmaceutical compositions useful in the methods of the present invention can be suitably developed for oral, rectal, vaginal, parenteral, topical, pulmonary, intranasal, buccal, ocular, or other routes of administration. Compositions useful in the methods of the present invention can be administered directly to the skin or any other tissue of a mammal. Other intended formulations include liposomal preparations, resealed red blood cells containing active ingredients, and immunological formulations. The route(s) of administration are readily apparent to those skilled in the art and depend on any number of factors, including the type and severity of the disease being treated, and the type and age of the veterinary or human subject being treated.
[0177] Formulations of the pharmaceutical compositions described herein can be prepared by any method known or hereafter developed in the field of pharmacology. Generally, such preparation methods include associating an active ingredient with a carrier or one or more other auxiliary components, and then, if necessary or desirable, forming or packaging the product into desired single-dose or multi-dose units.
[0178] As used herein, “unit dose” is a discrete amount of a pharmaceutical composition containing a predetermined amount of active ingredient. The amount of active ingredient is generally equal to the amount of active ingredient to be administered to the subject, or a convenient proportion of such an amount, for example, 1 / 2 or 1 / 3 of such an amount. A unit dosage form may be for one of two doses: once daily or multiple daily doses (e.g., about 1 to 4 times / day or more). If multiple daily doses are used, the unit dosage form may be the same or different for each dose.
[0179] In one embodiment, the composition of the present invention is formulated using one or more pharmaceutically acceptable excipients or carriers. In one embodiment, the pharmaceutical composition of the present invention comprises a therapeutically effective amount of the nucleic acid of the present invention and a pharmaceutically acceptable carrier. Useful pharmaceutically acceptable carriers include, but are not limited to, glycerol, water, saline, ethanol, and other pharmaceutically acceptable salt solutions such as phosphates and organic acid salts. Examples of these and other pharmaceutically acceptable carriers are described in Remington's Pharmaceutical Sciences (1991, Mack Publication Co., New Jersey).
[0180] The carrier may be a solvent or dispersion medium containing, for example, water, ethanol, polyols (e.g., glycerol, propylene glycol, liquid polyethylene glycol, etc.), suitable mixtures thereof, and vegetable oils. Adequate fluidity can be maintained, for example, by the use of coatings such as lecithin, maintaining the required particle size in the case of dispersions, and the use of surfactants. Prevention of microbial action can be achieved by various antimicrobial and antifungal agents, such as parabens, chlorobutanol, phenol, ascorbic acid, thimerosal, etc. Often, isotonic agents, such as sugars, sodium chloride, or polyalcohols such as mannitol and sorbitol, are included in the composition. Sustained absorption of the injectable composition can be achieved by including absorption-delaying agents, such as aluminum monostearate or gelatin, in the composition. In one embodiment, the pharmaceutically acceptable carrier is not DMSO alone.
[0181] The formulations may be used in additive mixtures with conventional excipients, i.e., pharmaceutically acceptable organic or inorganic carrier substances suitable for oral, vaginal, parenteral, nasal, intravenous, subcutaneous, enteral, or any other suitable mode of administration known in the art. The pharmaceutical preparations may be sterilized and mixed, if desired, with adjuvants such as salts, colorants, flavorings, and / or aromatics to affect lubrication, preservatives, stabilizers, wetting agents, emulsifiers, and osmotic buffers. They may also be combined with other activators, such as other analgesics, as needed.
[0182] As used herein, “additional components” include, but are not limited to, the following excipients: surfactants, dispersants, inert diluents, granulators and disintegrants, binders, lubricants, sweeteners, flavorings, colorants, preservatives, physiologically degradable compositions such as gelatin, aqueous vehicles and solvents, oily vehicles and solvents, suspending agents, dispersants or wetting agents, emulsifiers, lubricants, buffers, salts, thickeners, fillers, emulsifiers, antioxidants, antibiotics, antifungals, stabilizers, and one or more pharmaceutically acceptable polymer or hydrophobic materials. Other “additional components” that may be included in the pharmaceutical compositions of the present invention are known in the art, for example, as described in Genaro, ed. (1985, Remington's Pharmaceutical Sciences, Mack Publishing Co., Easton, PA), which is incorporated herein by reference.
[0183] The compositions of the present invention may contain a preservative in an amount of about 0.005% to 2.0% of the total weight of the composition. The preservative is used to prevent spoilage when exposed to environmental pollutants. Examples of preservatives useful according to the present invention include, but are not limited to, those selected from the group consisting of benzyl alcohol, sorbic acid, parabens, imidoureas, and combinations thereof. An exemplary preservative is a combination of about 0.5% to 2.0% benzyl alcohol and 0.05% to 0.5% sorbic acid.
[0184] In one embodiment, the composition contains an antioxidant and a chelating agent that inhibits the degradation of nucleic acids. Exemplary antioxidants for some compounds include BHT, BHA, alpha-tocopherol, and ascorbic acid in amounts ranging from about 0.01% to 0.3% by weight of the total weight of the composition, and BHT in amounts ranging from 0.03% to 0.1% by weight. In one embodiment, the chelating agent is present in an amount ranging from 0.01% to 0.5% by weight of the total weight of the composition. Exemplary chelating agents include EDTA salts (e.g., disodium EDTA) and citric acid in amounts ranging from about 0.01% to 0.20% by weight. In some embodiments, the chelating agent is in the range of 0.02% to 0.10% by weight of the total weight of the composition. The chelating agent is useful for chelating metal ions in the composition that may be detrimental to the shelf life of the formulation. While BHT and disodium EDTA are exemplary antioxidants and chelating agents for some compounds, respectively, other suitable equivalent antioxidants and chelating agents may be substituted, as is known to those skilled in the art.
[0185] Liquid suspensions can be prepared using conventional methods for achieving suspension of active ingredients in aqueous or oily vehicles. Examples of aqueous vehicles include water and isotonic saline. Examples of oily vehicles include almond oil, oily esters, ethyl alcohol, vegetable oils such as peanut, olive, sesame, and coconut oil, fractionated vegetable oils, and mineral oils such as liquid paraffin. Liquid suspensions may further contain one or more additional components, but are not limited to, suspending agents, dispersants or wetting agents, emulsifiers, lubricants, preservatives, buffers, salts, flavorings, colorants, and sweeteners. Oily suspensions may further contain thickeners. Known suspending agents, but are not limited to, sorbitol syrup, hydrogenated edible fats, sodium alginate, polyvinylpyrrolidone, tragacanth gum, acacia gum, and cellulose derivatives such as sodium carboxymethylcellulose, methylcellulose, and hydroxypropyl methylcellulose. Known dispersants or wetting agents include, but are not limited to, natural phosphatides, e.g., lecithin, condensation products of alkylene oxides and fatty acids, condensation products of long-chain aliphatic alcohols, condensation products of fatty acids and partial esters derived from hexitol, or condensation products of fatty acids and partial esters derived from hexitol anhydrides (e.g., polyoxyethylene stearate, heptadecaethyleneoxycetanol, polyoxyethylene sorbitol monooleate, and polyoxyethylene sorbitan monooleate, respectively). Known emulsifiers include, but are not limited to, lecithin and acacia. Known preservatives include, but are not limited to, methyl, ethyl, or n-propyl-para-hydroxybenzoate, ascorbic acid, and sorbic acid. Known sweeteners include, for example, glycerol, propylene glycol, sorbitol, sucrose, and saccharin. Known thickeners for oily suspensions include, for example, beeswax, hard paraffin, and cetyl alcohol.
[0186] Liquid solutions of active ingredients in aqueous or oily solvents can be prepared in substantially the same manner as liquid suspensions, the main difference being that the active ingredient is dissolved rather than suspended in the solvent. As used herein, an "oily" liquid is a liquid containing carbon-containing liquid molecules and exhibiting properties of lower polarity than water. Liquid solutions of the pharmaceutical compositions of the present invention may contain each of the components described with respect to liquid suspensions, and it is understood that the suspending agent does not necessarily aid in the dissolution of the active ingredient in the solvent. Examples of aqueous solvents include water and isotonic saline. Examples of oily solvents include almond oil, oily esters, ethyl alcohol, vegetable oils such as peanut, olive, sesame, and coconut oil, fractionated vegetable oils, and mineral oils such as liquid paraffin.
[0187] Powder and granular formulations of the pharmaceutical preparations of the present invention can be prepared using known methods. Such formulations may be administered directly to a subject and may be used, for example, to form tablets, to fill capsules, or to prepare aqueous or oily suspensions or solutions by adding an aqueous or oily vehicle. Each of these formulations may further contain one or more of the following: dispersants or wetting agents, suspending agents, and preservatives. Fillers and additional excipients such as sweeteners, flavoring agents, or coloring agents may also be included in these formulations.
[0188] The pharmaceutical compositions of the present invention may also be prepared, packaged, or sold in the form of oil-in-water emulsions or water-in-oil emulsions. The oily phase may be a vegetable oil such as olive oil or peanut oil, a mineral oil such as liquid paraffin, or a combination thereof. Such compositions may further contain one or more emulsifiers, for example, naturally occurring gums such as gum arabic or tragacanth gum, naturally occurring phosphatides such as soy or lecithin phosphatides, esters or partial esters derived from combinations of fatty acids and hexitol anhydrides, for example, sorbitan monooleate, and condensation products of such partial esters with ethylene oxide, for example, polyoxyethylene sorbitan monooleate. These emulsions may also contain additional components, such as sweeteners or flavorings.
[0189] Methods for impregnating or coating materials with chemical compositions are known in the art, but are not limited to these, and include methods for depositing or bonding chemical compositions to a surface, methods for incorporating chemical compositions into the structure of a material during its synthesis (i.e., using, for example, physiologically degradable materials), and methods for absorbing aqueous or oily solutions or suspensions into an absorbent material, with or without subsequent drying.
[0190] The administration regimen may affect what constitutes the effective dose. The therapeutic formulation may be administered to the subject either before or after the diagnosis of the disease. Furthermore, some divided doses, as well as staggered dosages, may be administered daily or sequentially, or the dose may be infused continuously or administered as a bolus injection. In addition, the dosage of the therapeutic formulation may be increased or decreased proportionally, as indicated by the urgency of the therapeutic or preventive situation.
[0191] The compositions of the present invention may be administered to subjects (mammals, including humans) using known procedures in doses and durations effective for preventing or treating a disease. The effective amount of nucleic acid required to achieve a therapeutic effect may vary depending on factors such as the activity of the specific nucleic acid used, the time of administration, the rate of nucleic acid efflux, the duration of treatment, other drugs, compounds, or materials used in combination with the nucleic acid, the state of the disease or disorder, age, sex, weight, condition, the overall health status and prior medical history of the subject being treated, and similar factors well known in the medical field. The administration regimen may be adjusted to produce an optimal therapeutic response. For example, several divided doses may be administered daily, or the dose may be proportionally reduced as indicated by the urgency of the treatment situation. Non-limiting examples of the effective dose range of the nucleic acid of the present invention range from about 1 mg / kg body weight / day to 5,000 mg / kg body weight / day. Those skilled in the art will be able to study the relevant factors and make a determination regarding the effective amount of therapeutic nucleic acid without excessive experimentation.
[0192] Nucleic acids may be administered to subjects several times a day, or less frequently, for example, once a day, once a week, once every two weeks, once a month, or even less frequently, for example, once every few months, or even once a year or less. The amount of nucleic acid administered per day may be, in non-limiting examples, daily, every other day, every two days, every three days, every four days, or every five days. For example, in every-other-day administration, a dose of 5 mg / day may be started on Monday, the first subsequent dose of 5 mg / day may be administered on Wednesday, the second subsequent dose of 5 mg / day may be administered on Friday, and so on. The frequency of these doses will be readily apparent to those skilled in the art and will not be limited, but will depend on any number of factors, such as the type and severity of the disease being treated, the type and age of the animal, etc.
[0193] The actual dosage level of the active ingredient in the pharmaceutical composition of the present invention can be varied to obtain an amount of the active ingredient effective in achieving a desired therapeutic response for a particular subject, composition, and mode of administration, without being toxic to the subject.
[0194] A physician skilled in the art, such as an internist or veterinarian, can easily determine and prescribe the effective amount of the required pharmaceutical composition. For example, an internist or veterinarian can start with a dose of the nucleic acid of the present invention used in the pharmaceutical composition at a level lower than the level required to achieve the desired therapeutic effect, and gradually increase the dosage until the desired effect is achieved.
[0195] In certain embodiments, it is particularly advantageous to formulate nucleic acids in dosage units for ease of administration and uniformity of dosage. As used herein, a dosage unit refers to a physically distinct unit suitable as a unit dosage for the subject being treated. Each unit contains a predetermined amount of therapeutic nucleic acid calculated to produce a desired therapeutic effect when combined with the required pharmaceutical vehicle. The dosage unit forms of the present invention are determined by and directly depend on (a) the inherent characteristics of the nucleic acid and the specific therapeutic effect to be achieved, and (b) the limitations inherent in the art of formulating / manufacturing such nucleic acids for the treatment of diseases in the subject.
[0196] In one embodiment, the compositions of the present invention are administered to a subject in dosages ranging from one to five times or more per day. In another embodiment, the compositions of the present invention are administered to a subject in dosages ranging from once a day, once every two days, once a day, once every three days to once a week, and once every two weeks, but are not limited to these. It will be readily apparent to those skilled in the art that the frequency of administration of the various combination compositions of the present invention will vary from subject to subject depending on many factors, including but not limited to the age, disease or disorder being treated, sex, overall health condition, and other factors. Therefore, the present invention should not be construed as being limited to any particular drug regimen, and the exact dosage and composition to be administered to any subject will be determined by the attending physician, taking into account all other factors relating to the subject.
[0197] The compositions of the present invention for administration are approximately 1 mg to approximately 10,000 mg, approximately 20 mg to approximately 9,500 mg, approximately 40 mg to approximately 9,000 mg, approximately 75 mg to approximately 8,500 mg, approximately 150 mg to approximately 7,500 mg, approximately 200 mg to approximately 7,000 mg, approximately 3,050 mg to approximately 6,000 mg, approximately 500 mg to approximately 5,000 mg, approximately 750 mg to approximately 4,000 mg, and approximately 1 mg to approximately 3,000 mg. This can range from mg, approximately 10 mg to approximately 2,500 mg, approximately 20 mg to approximately 2,000 mg, approximately 25 mg to approximately 1,500 mg, approximately 50 mg to approximately 1,000 mg, approximately 75 mg to approximately 900 mg, approximately 100 mg to approximately 800 mg, approximately 250 mg to approximately 750 mg, approximately 300 mg to approximately 600 mg, approximately 400 mg to approximately 500 mg, and any whole or partial increments in between.
[0198] In some embodiments, the dose of the composition of the present invention is about 1 mg to about 2,500 mg. In some embodiments, the dose of the composition of the present invention used in the compositions described herein is less than about 10,000 mg, or less than about 8,000 mg, or less than about 6,000 mg, or less than about 5,000 mg, or less than about 3,000 mg, or less than about 2,000 mg, or less than about 1,000 mg, or less than about 500 mg, or less than about 200 mg, or less than about 50 mg. Similarly, in some embodiments, the dose of the second composition described herein (i.e., the drug used to treat the same or a different disease as treated by the composition of the present invention) is less than about 1,000 mg, or less than about 800 mg, or less than about 600 mg, or less than about 500 mg, or less than about 400 mg, or less than about 300 mg, or less than about 200 mg, or less than about 100 mg, or less than about 50 mg, or less than about 40 mg, or less than about 30 mg, or less than about 25 mg, or less than about 20 mg, or less than about 15 mg, or less than about 10 mg, or less than about 5 mg, or less than about 2 mg, or less than about 1 mg, or less than about 0.5 mg, and any whole or partial increase thereof.
[0199] In one embodiment, the present invention relates to a packaged pharmaceutical composition comprising a container holding a therapeutically effective amount of the nucleic acid of the present invention, either alone or in combination with a second pharmaceutical agent, and instructions for using the nucleic acid to treat, prevent or alleviate one or more symptoms of a disease in a subject.
[0200] The term “container” includes any container for holding a pharmaceutical composition. For example, in one embodiment, the container is packaging containing the pharmaceutical composition. In other embodiments, the container is not packaging containing the pharmaceutical composition; that is, the container is a container such as a box or vial containing the packaged or unpackaged pharmaceutical composition and instructions for use of the pharmaceutical composition. Furthermore, packaging techniques are well known in the art. Instructions for use of the pharmaceutical composition may be included in the packaging containing the pharmaceutical composition, and it should be understood that the instructions form an increased functional relationship with the packaged product. However, it should be understood that the instructions may include information about its intended function, for example, the ability of the nucleic acid to treat or prevent a disease in a subject, or to deliver an imaging agent or diagnostic agent to a subject.
[0201] Any route of administration of the composition of the present invention includes oral, nasal, parenteral, sublingual, transdermal, and transmucosal administration (e.g., sublingual, tongue, (trans)buccal mucosa, and (trans)nasal cavity, intrabladder, intraduodenum, intrastomy, intrarectum, intraperitoneal, subcutaneous, intramuscular, intradermal, intra-arterial, and intravenous administration).
[0202] Suitable compositions and dosage forms include, for example, tablets, capsules, caplets, pills, gel caps, lozenges, dispersions, suspensions, solutions, syrups, granules, beads, transdermal patches, gels, powders, pellets, magmas, lozenges, creams, pastes, plasters, lotions, discs, suppositories, liquid sprays for nasal or oral administration, dry powders or aerosolized formulations for inhalation, and compositions and formulations for intravesical administration. It should be understood that formulations and compositions that may be useful in the present invention are not limited to the specific formulations and compositions described herein.
[0203] system In some embodiments, the present invention relates to a system for cis-cleavage and trans-splicing of independent RNA molecules. In some embodiments, the present invention relates to a system for cis-cleavage and trans-splicing of a single RNA molecule. In some embodiments, cis-cleavage and trans-splicing of independent RNA molecules or fragments of a single RNA molecule yields a single RNA molecule encoding the desired full-length protein, as described herein. In some embodiments, the system includes a ligase or nucleic acid encoding a ligase, such as RtcB, as described herein.
[0204] In one embodiment, the present invention relates to an inducible system for generating a single RNA encoding a full-length protein from two distinct RNA molecules encoding a first and second portion of a full-length protein via cis-cleavage of a ribozyme and trans-splicing of two independent RNA molecules. In some embodiments, the system includes a ribozyme recognition sequence and a ribozyme as described herein. In some embodiments, the system includes a ligase or nucleic acid encoding a ligase as described herein.
[0205] In one embodiment, the present invention relates to a system for assembling a full-length RNA virus genome. Exemplary RNA viruses include, but are not limited to, coronaviruses, paramyxoviruses, orthomyxoviruses, retroviruses, lentiviruses, alphaviruses, flaviviruses, rhabdoviruses, measles viruses, Newcastle disease viruses, and picornaviruses. In one embodiment, the system comprises a first nucleic acid encoding a first portion of the RNA virus genome and encoding a 3' ribozyme. In one embodiment, the system comprises a second nucleic acid encoding a second portion of the RNA virus genome and encoding a 5' ribozyme. In one embodiment, the system comprises the first portion of the RNA virus genome and a 3' ribozyme. In one embodiment, the system comprises the second portion of the RNA virus genome and a 5' ribozyme. In one embodiment, the system comprises a nucleic acid encoding a ligase or a ligase. In one embodiment, upon cis-cleavage of the 3' and 5' ribozymes, the first portion of the RNA virus genome and the second portion of the RNA virus genome are ligated together to produce a full-length RNA virus genome.
[0206] In Vivo In one embodiment, the present invention relates to a system for the delivery and expression of one or more full-length proteins via cis-cleavage and trans-splicing of an independent RNA molecule encoding a portion of a full-length protein. In some embodiments, the system enables the delivery and expression of large proteins exceeding the package size of conventional vectors (e.g., dystrophin exceeding the packaging size of AAV vectors), synthetic repeat domain proteins that are difficult to synthesize in vitro (e.g., synthetic spider silk), or toxic / antiviral proteins (e.g., DTA). In one embodiment, the present invention comprises an AAV system for the delivery and expression of one or more full-length proteins of interest. In some embodiments, the system comprises a ligase or nucleic acid encoding a ligase as described herein.
[0207] In one embodiment, the present invention includes a lentiviral delivery system for delivering one or more nucleic acid molecules encoding one or more target proteins. In one embodiment, the lentiviral delivery system includes (1) a packaging plasmid, (2) an envelope plasmid, and (3) an import plasmid. In one embodiment, the import plasmid encodes a first RNA molecule and a second RNA molecule.
[0208] In one embodiment, the present invention includes a dual lentiviral delivery system comprising a first lentiviral vector and a second lentiviral vector. In one embodiment, the first lentiviral vector system comprises (1) a packaging plasmid, (2) an envelope plasmid, and (3) a first import plasmid. In one embodiment, the second lentiviral vector system comprises (1) a packaging plasmid, (2) an envelope plasmid, and (3) a second import plasmid. In one embodiment, the first import plasmid encodes a first RNA molecule. In one embodiment, the second import plasmid encodes a second RNA molecule.
[0209] In one embodiment, the packaging plasmid contains a nucleic acid sequence encoding a gag-pol polyprotein. In one embodiment, the gag-pol polyprotein contains a catalytically inactive integrase. In one embodiment, the gag-pol polyprotein contains a D116N integrase mutation.
[0210] In one embodiment, the envelope plasmid includes a nucleic acid sequence encoding an envelope protein. In one embodiment, the envelope plasmid includes a nucleic acid sequence encoding an HIV envelope protein. In one embodiment, the envelope plasmid includes a nucleic acid sequence encoding a vesicular stomatitis virus g protein (VSV-g) envelope protein. In one embodiment, the envelope protein can be selected based on a desired cell type.
[0211] In one embodiment, the first RNA molecule of one transfer plasmid contains a protein-coding region encoding a first portion of the protein of interest and a 3' ribozyme. In one embodiment, the second RNA molecule of one transfer plasmid contains a protein-coding region encoding a second portion of the protein of interest and a 5' ribozyme. In one embodiment, the transfer plasmid contains a 5' long-term repeat (LTR) sequence and a 3' LTR sequence. In one embodiment, the 3' LTR is a self-inactivating (SIN) LTR. Thus, in one embodiment, the 5' LTR contains a U3 sequence, an R sequence, and a U5 sequence, and the 3' LTR contains an R sequence and a U5 sequence, but does not contain a U3 sequence. In one embodiment, the 5' LTR and 3' LTR are adjacent to the sequences encoding the first portion and the second portion of the protein of interest.
[0212] In one embodiment, the first RNA molecule of the first import plasmid includes a protein-coding region encoding a first portion of the protein of interest and a 3' ribozyme. In one embodiment, the second RNA molecule of the second import plasmid includes a protein-coding region encoding a second portion of the protein of interest and a 5' ribozyme. In one embodiment, the first and second import plasmids include a 5' long-term repeat (LTR) sequence and a 3' LTR sequence. In one embodiment, the 3' LTR is a self-inactivating (SIN) LTR. Thus, in one embodiment, the 5' LTR includes a U3 sequence, an R sequence, and a U5 sequence, and the 3' LTR includes an R sequence and a U5 sequence, but does not include a U3 sequence. In one embodiment, the 5' LTR and 3' LTR of the first import plasmid are adjacent to the sequences encoding the first portion of the protein of interest and the 3' ribozyme. In one embodiment, the 5'LTR and 3'LTR of the second transfer plasmid are adjacent to the sequences encoding the second portion of the target protein and the 5' ribozyme.
[0213] In one embodiment, a packaging plasmid, an envelope plasmid, and a transfer plasmid are introduced into cells. In one embodiment, the cells transcribe and translate a nucleic acid sequence encoding the gag-pol protein to produce the gag-pol polyprotein. In one embodiment, the cells transcribe and translate a nucleic acid sequence encoding the envelope protein to produce the envelope protein. In one embodiment, the cells transcribe one transfer plasmid to provide the first RNA molecule and the second RNA molecule. In one embodiment, the cells transcribe the first transfer plasmid to provide the first RNA molecule and transcribe the second transfer plasmid to provide the second RNA molecule. In one embodiment, the gag-pol protein, the envelope polyprotein, the first RNA molecule, and the second RNA molecule are packaged into virus particles. In one embodiment, the virus particles are recovered from the cell culture medium. In one embodiment, the virus particles transduce target cells, the 3' ribozyme catalyzes itself from the first RNA molecule, thereby generating a 3'P or 2'3'cP terminus, the 5' ribozyme catalyzes itself from the second RNA molecule, thereby generating a 5'OH terminus, and the endogenous RNA 2',3'-cyclic phosphate and 5'-OH (RtcB) ligase ligates the 3'P or 2'3'cP terminus to the 5'OH terminus, thereby generating a complete RNA molecule encoding the protein of interest, and the cells translate the protein of interest.
[0214] In one embodiment, a packaging plasmid, an envelope plasmid, and a first transfer plasmid are introduced into a cell. In one embodiment, the cell transcribes and translates a nucleic acid sequence encoding a gag-pol protein to produce a gag-pol polyprotein. In one embodiment, the cell transcribes and translates a nucleic acid sequence encoding an envelope protein to produce an envelope protein. In one embodiment, the cell transcribes the first transfer plasmid to provide a first RNA molecule. In one embodiment, the gag-pol protein, the envelope polyprotein, and the first RNA molecule are packaged into a first viral particle. In one embodiment, the first viral particle is recovered from the cell culture medium.
[0215] In one embodiment, a packaging plasmid, an envelope plasmid, and a second transfer plasmid are introduced into a cell. In one embodiment, the cell transcribes and translates a nucleic acid sequence encoding a gag-pol protein to produce a gag-pol polyprotein. In one embodiment, the cell transcribes and translates a nucleic acid sequence encoding an envelope protein to produce an envelope protein. In one embodiment, the cell transcribes the second transfer plasmid to provide a second RNA molecule. In one embodiment, the gag-pol protein, the envelope polyprotein, and the second RNA molecule are packaged into a second viral particle. In one embodiment, the second viral particle is recovered from the cell culture medium.
[0216] In one embodiment, first and second viral particles transduce target cells, a 3' ribozyme catalyzes itself from the first RNA molecule, thereby generating a 3'P or 2'3'cP terminus, a 5' ribozyme catalyzes itself from the second RNA molecule, thereby generating a 5'OH terminus, and endogenous RNA 2',3'-cyclic phosphate and 5'-OH(RtcB) ligase ligates the 3'P or 2'3'cP terminus to the 5'OH terminus, thereby generating a complete RNA molecule encoding the desired protein, which the cell then translates. In one embodiment, the present invention relates to a system for preventing undesirable partial protein expression from a split precursor RNA molecule. In one embodiment, the system includes incorporating a translational control sequence for proteolysis into the split precursor RNA molecule, as described herein.
[0217] In one embodiment, the present invention relates to a system for expressing two or more target proteins from two or more pairs of independent RNA molecules encoding portions of a target protein via cis-cleavage of ribozymes and trans-splicing of pairs of independent RNA molecules. In one embodiment, each individual pair of independent RNA molecules has a separate reading frame so as described herein that undesirable pair trans-splicing does not result in translation of a full-length functional protein. In some embodiments, the system includes a ligase or nucleic acid encoding a ligase as described herein.
[0218] In one embodiment, the present invention includes a system for the delivery and expression of a full-length protein of interest and a cargo sequence. In one embodiment, the system includes a first portion of RNA encoding a first portion of the protein of interest, ligated to a synthetic intron at its 3' end, and a second portion of RNA encoding a second portion of the protein of interest, ligated to a synthetic intron at its 5' end. In one embodiment, the synthetic intron has a 5' ribozyme sequence and a 3' ribozyme sequence adjacent to each other on both sides. In one embodiment, the synthetic intron includes a cargo sequence positioned between the 5' ribozyme sequence and the 3' ribozyme sequence. In one embodiment, self-cleavage of the 5' ribozyme sequence and the 3' ribozyme sequence produces three distinct RNA molecules: 1) a first fragment containing the first portion of RNA encoding the first portion of the protein of interest, 2) a second fragment containing the synthetic intron, and 3) a third fragment containing the second portion of RNA encoding the second portion of the protein of interest. In one embodiment, the matching ends of the second fragment are ligated to produce a circular RNA molecule containing the synthetic intron containing the cargo sequence. In some embodiments, the first and third fragments are ligated together to produce a single full-length linear RNA molecule. In one embodiment, the full-length protein of interest comprises a therapeutic protein, a reporter protein, a recombinase, an antibiotic resistance gene product, an antibody, or a Cas9 protein. In one embodiment, the cargo sequence comprises a therapeutic nucleic acid sequence (e.g., a miRNA sequence or a CRISPR guide RNA sequence) or encodes a therapeutic protein. In some embodiments, the full-length protein of interest comprises Cas9, and the cargo sequence comprises a guide RNA sequence, thereby targeting Cas9 to a specific genomic sequence for editing. In some embodiments, the system comprises a ligase or nucleic acid encoding a ligase as described herein.
[0219] In one embodiment, the present invention includes a system for gene editing comprising one or more trans-cleaved ribozymes. In some embodiments, the system comprises two trans-cleaved ribozymes targeting upstream and downstream of a disease-causing mutation. In some embodiments, trans-cleaving upstream and downstream of the disease-causing mutation results in the removal of the disease-causing mutation. In some embodiments, the remainder of the gene is trans-spliced together after trans-cleaving of the disease-causing mutation. In some embodiments, the trans-spliced gene is expressed as a functional protein. In some embodiments, the system comprises a ligase or nucleic acid encoding a ligase as described herein.
[0220] In Vitro In one embodiment, the present invention includes an in vitro system for generating an RNA molecule encoding a target protein. In one embodiment, the system includes at least two RNA molecules. In one embodiment, the at least two RNA molecules include a first RNA molecule and a second RNA molecule.
[0221] In one embodiment, the first RNA molecule includes a coding region that encodes a first portion of the protein of interest. In one embodiment, the first RNA molecule includes a 3' ribozyme. In one embodiment, the first RNA molecule includes a coding region that encodes a first portion of the protein of interest and a 3' ribozyme, as described herein.
[0222] In one embodiment, the second RNA molecule includes a coding region that encodes a second portion of the protein of interest. In one embodiment, the second RNA molecule includes a 5' ribozyme. In one embodiment, the second RNA molecule includes a coding region that encodes a second portion of the protein of interest and a 5' ribozyme, as described herein.
[0223] In one embodiment, the in vitro system for generating an RNA molecule encoding a protein of interest further comprises a ligase. In one embodiment, the ligase induces the assembly of RNA molecules from the coding regions of a first RNA molecule and a second RNA molecule. In one embodiment, the ligase is an RNA 2',3'-cyclic phosphate and 5'-OH(RtcB) ligase as described herein.
[0224] In one embodiment, the present invention includes an in vitro system for generating an RNA molecule encoding a target repeat domain protein. In one embodiment, the system includes a first RNA molecule, one or more additional RNA molecules, and a final RNA molecule.
[0225] In one embodiment, the first RNA molecule includes a coding region encoding a first portion of the protein of interest. In one embodiment, the first RNA molecule includes a 3' ribozyme. In one embodiment, the first RNA molecule includes a coding region encoding a first portion of the protein of interest and a 3' ribozyme. In one embodiment, the 3' ribozyme catalyzes itself from the first RNA molecule, thereby generating a 3'P or 2'3'cP terminus. In one embodiment, the first RNA molecule further includes a 5' tag. In one embodiment, the 5' tag mediates the attachment of the first RNA molecule to a solid support.
[0226] In one embodiment, one or more additional RNA molecules include a coding region encoding the domain of the protein of interest, a 5' ribozyme, and a 3' ribozyme recognition sequence. In one embodiment, the 5' ribozyme cleaves itself to produce a 5' OH end. In one embodiment, the 3' ribozyme recognition sequence includes the VS-S sequence described herein.
[0227] In one embodiment, the last RNA molecule includes a coding region that encodes the last portion of the protein of interest. In one embodiment, the last RNA molecule includes a 5' ribozyme. In one embodiment, the last RNA molecule includes a coding region that encodes the last portion of the protein of interest and a 5' ribozyme. In one embodiment, the 5' ribozyme cleaves itself to produce a 5' OH end.
[0228] In one embodiment, the system further comprises a ribozyme. In one embodiment, the ribozyme comprises VS-Rz as described herein. In one embodiment, VS-Rz recognizes VS-S as described herein and mediates its cleavage from one or more additional RNA molecules. In one embodiment, the cleavage produces a 3'P or 2'3'cP terminus.
[0229] In one embodiment, the system includes a ligase. In some embodiments, the ligase ligates the 3'P or 2'3'cP end of a first RNA molecule to the 5'OH end of one or more additional RNA molecules. In some embodiments, the ligase ligates the 3'P or 2'3'cP end of one or more additional RNA molecules to the 5'OH end of the last RNA molecule. In some embodiments, the ligase ligates the 3'P or 2'3'cP end of a first RNA molecule to the 5'OH end of one or more additional RNA molecules, and ligates the 3'P or 2'3'cP end of one or more additional RNA molecules to the 5'OH end of the last RNA molecule, thereby producing a complete RNA molecule encoding an N-terminal domain, one or more additional domains, and a C-terminal domain. In some embodiments, the ligase is the RNA 2',3'-cyclic phosphate and 5'-OH(RtcB) ligase described herein.
[0230] method In some embodiments, the present invention relates to a method for cis-cleaving and trans-splicing or scarless slicing of independent RNA molecules. In some embodiments, the present invention relates to a method for cis-cleaving and trans-splicing or scarless slicing of a single RNA molecule. In some embodiments, cis-cleaving and trans-splicing of an independent RNA molecule or a fragment of a single RNA molecule yields a single RNA molecule encoding the desired full-length protein, as described herein. In some embodiments, the method comprises administering a ligase or nucleic acid encoding a ligase as described herein.
[0231] In one embodiment, the present invention relates to an inducible method for generating a single RNA encoding a full-length protein from two separate RNA molecules encoding a first and a second portion of the full-length protein, via cis-cleavage of a ribozyme contained in each of the two separate RNA molecules and trans-splicing of the two independent RNA molecules. In some embodiments, scarless ligation of the two independent RNA molecules generates a single RNA molecule encoding the full-length protein of interest. In some embodiments, the method includes a ribozyme recognition sequence and a ribozyme as described herein. In some embodiments, the method includes administering a ligase or nucleic acid encoding a ligase as described herein.
[0232] In some embodiments, the present invention relates to a method for circularizing an RNA molecule. In some embodiments, the method comprises cis-cleavage and cis-ligation or scarless cis-splicing of a single RNA molecule. In some embodiments, the linear RNA molecule (transcribed in vitro or in vivo from a DNA plasmid) comprises a 5' ribozyme and a 3' ribozyme. In one embodiment, the 5' ribozyme is directly ligated to the N-terminal coding sequence of the protein of interest, and the 3' ribozyme is directly ligated to the C-terminal coding sequence of the protein of interest. In one embodiment, the 3' ribozyme is directly ligated to the N-terminal coding sequence of the protein of interest, and the 5' ribozyme is directly ligated to the C-terminal coding sequence of the protein of interest. In some embodiments, cis-cleavage of both the 3' and 5' ribozymes, as well as cis-ligation or scarless cis-splicing of a fragment of the RNA molecule containing the coding sequence, yields a single circular RNA molecule encoding the full-length protein of interest, as described herein. In some embodiments, the method includes administering a ligase or nucleic acid encoding a ligase as described herein.
[0233] In Vivo In one embodiment, the present invention includes a method for generating an RNA molecule encoding a target protein. In some embodiments, the method includes administering at least two nucleic acid molecules to a cell or tissue. In one embodiment, the at least two nucleic acid molecules include a first RNA molecule and a second RNA molecule. In some embodiments, the at least two nucleic acid molecules encode a first RNA molecule and a second RNA molecule.
[0234] In one embodiment, the first RNA molecule comprises a coding region encoding the first portion of the protein of interest. In one embodiment, the first RNA molecule comprises a 3' ribozyme. In one embodiment, the first RNA molecule comprises a coding region encoding the first portion of the protein of interest and a 3' ribozyme. In one embodiment, the 3' ribozyme catalyzes itself from the first RNA molecule, thereby generating a 3'P or 2'3'cP terminus. In one embodiment, the 3' ribozyme is a member of the ribozymes of the HDV family.
[0235] In one embodiment, the second RNA molecule comprises a coding region encoding the second portion of the protein of interest. In one embodiment, the second RNA molecule comprises a 5' ribozyme. In one embodiment, the second RNA molecule comprises a coding region encoding the second portion of the protein of interest and a 5' ribozyme. In one embodiment, the 5' ribozyme catalyzes itself from the second RNA molecule, thereby generating a 5'OH terminus. In one embodiment, the 5' ribozyme is a member of the ribozymes of the HH family.
[0236] In one embodiment, the 3'P or 2'3'cP terminus is ligated to the 5'OH terminus to form an RNA molecule comprising the coding region of the first RNA molecule and the coding region of the second RNA molecule. In some embodiments, the trans-ligation of the coding sequences of the first and second portions of the protein of interest is performed in a scarless manner such that there is no intervening sequence between the first and second portions of the protein of interest after translation.
[0237] In one embodiment, the method comprises administering to the cell or tissue one or more additional nucleic acid molecules encoding one or more additional RNA molecules, each additional RNA molecule comprising a coding region encoding a domain of the protein of interest, a 5' ribozyme, and a 3' ribozyme.
[0238] <In one embodiment, the method comprises administering one or more additional nucleic acid molecules encoding one or more additional RNA molecules to a cell or tissue, each additional RNA molecule comprising a coding region encoding a domain of the protein of interest, a 5' ribozyme, and a 3' ribozyme recognition sequence. In one embodiment, the 3' ribozyme recognition sequence comprises VS-S. In one embodiment, the ribozyme is VS.
[0239] In one embodiment, the method involves administering one or more nucleic acid molecules encoding ligases and ligases selected from the group to a cell or tissue. In one embodiment, the ligase induces the assembly of RNA molecules from the coding regions of a first RNA molecule and a second RNA molecule. In some embodiments, the transligation of the coding sequences of the first and second portions of the protein of interest is performed in a scarless manner such that no intervening sequence is present between the first and second portions of the protein of interest after translation. In one embodiment, the ligase is an RNA 2',3'-cyclic phosphate and 5'-OH(RtcB) ligase.
[0240] In some embodiments, the method involves administering to cells or tissues at least one AAV vector encoding a first RNA molecule containing a protein-coding region and a 3' ribozyme encoding a first portion of the protein of interest, and a second RNA molecule containing a protein-coding region and a 5' ribozyme encoding a second portion of the protein of interest. In some embodiments, the method involves administering a ligase or nucleic acid encoding a ligase as described herein.
[0241] In some embodiments, the method includes administering at least two AAV vectors, including a first AAV vector and a second AAV vector. In one embodiment, the first AAV vector encodes a first RNA molecule comprising a protein-coding region encoding a first portion of the protein of interest and a 3' ribozyme. In one embodiment, the second AAV vector encodes a second RNA molecule comprising a protein-coding region encoding a second portion of the protein of interest and a 5' ribozyme to a cell or tissue. In some embodiments, the method includes administering a ligase or nucleic acid encoding a ligase as described herein. In some embodiments, the transligation of the coding sequences of the first and second portions of the protein of interest is performed in a scarless manner such that no intervening sequence is present between the first and second portions of the protein of interest after translation.
[0242] In some embodiments, the method involves administering to cells or tissues a lentiviral vector encoding at least one lentiviral vector comprising a first RNA molecule containing a protein-coding region and a 3' ribozyme encoding a first portion of the protein of interest, and a second RNA molecule containing a protein-coding region and a 5' ribozyme encoding a second portion of the protein of interest. In some embodiments, the method involves administering a ligase or nucleic acid encoding a ligase as described herein.
[0243] In some embodiments, the method includes administering at least two lentiviral vectors, including a first lentiviral vector and a second lentiviral vector. In one embodiment, the first lentiviral vector encodes a first RNA molecule comprising a protein-coding region encoding a first portion of the protein of interest and a 3' ribozyme. In one embodiment, the second lentiviral vector encodes a second RNA molecule comprising a protein-coding region encoding a second portion of the protein of interest to be administered to a cell or tissue and a 5' ribozyme. In some embodiments, the method includes administering a ligase or nucleic acid encoding a ligase as described herein.
[0244] In some embodiments, the method includes administering at least one lentiviral vector delivery system to deliver a first RNA molecule containing a protein-coding region encoding a first portion of the protein of interest and a 3' ribozyme, and a second RNA molecule containing a protein-coding region encoding a second portion of the protein of interest and a 5' ribozyme, to cells or tissues. In some embodiments, the method includes administering a ligase or nucleic acid encoding a ligase as described herein.
[0245] In some embodiments, the method includes administering at least two lentiviral vector delivery systems, including a first lentiviral vector delivery system and a second lentiviral vector delivery system. In one embodiment, the first lentiviral vector delivery system provides a first RNA molecule comprising a protein-coding region encoding a first portion of the protein of interest and a 3' ribozyme. In one embodiment, the second lentiviral vector delivery system provides a second RNA molecule comprising a protein-coding region encoding a second portion of the protein of interest and a 5' ribozyme to cells or tissues. In some embodiments, the method includes administering a ligase or nucleic acid encoding a ligase as described herein.
[0246] In some embodiments, the method involves administering two or more delivery vehicles selected from the group consisting of AAV vectors, lentiviral vectors, lentiviral vector delivery systems, or combinations thereof. In one embodiment, the two or more delivery vehicles include a first delivery vehicle and a second delivery vehicle. In one embodiment, the first delivery vehicle provides a first RNA molecule comprising a protein-coding region encoding a first portion of the protein of interest and a 3' ribozyme. In one embodiment, the second delivery vehicle provides a second RNA molecule comprising a protein-coding region encoding a second portion of the protein of interest and a 5' ribozyme to a cell or tissue. In some embodiments, the method involves administering a ligase or nucleic acid encoding a ligase as described herein.
[0247] In one embodiment, the present invention includes a method for generating a circular RNA molecule encoding a protein of interest. In some embodiments, the method includes administering at least one nucleic acid molecule to a cell or tissue, the nucleic acid molecule comprising at least a 5' ribozyme directly ligated to a sequence encoding the C-terminal or N-terminal portion of the protein of interest and a 3' ribozyme directly ligated to a sequence encoding the C-terminal or N-terminal portion of the protein of interest. In one embodiment, the 3' ribozyme catalyzes itself from the RNA molecule, thereby generating a 3'P or 2'3'cP end. In one embodiment, the 3' ribozyme is a member of the HDV family of ribozymes. In one embodiment, the 5' ribozyme catalyzes itself from the RNA molecule, thereby generating a 5'OH end. In one embodiment, the 5' ribozyme is a member of the HH family of ribozymes. In one embodiment, the 3'P or 2'3'cP end is ligated to the 5'OH end to form a circular RNA molecule in which the N-terminal and C-terminal coding sequences of the protein of interest are operably ligated. In some embodiments, cyclicization occurs in a scarless manner, such that no intervening sequence exists between the N-terminus and C-terminus after translation.
[0248] Methods for introducing and expressing genes in cells are well known in the art. In relation to expression vectors, vectors can be readily introduced into host cells, such as mammalian, bacterial, yeast, or insect cells, by any method in the art. For example, expression vectors can be introduced into host cells by physical, chemical, or biological means.
[0249] Physical methods for introducing polynucleotides into host cells include calcium phosphate precipitation, lipofection, particle bombardment, microinjection, and electroporation. Methods for producing cells containing vectors and / or exogenous nucleic acids are well known in the art. See, for example, Sambrook et al. (2012, Molecular Cloning: A Laboratory Manual, Cold Spring Harbor Laboratory, New York). An exemplary method for introducing polynucleotides into host cells is calcium phosphate transfection.
[0250] Biological methods for introducing target polynucleotides into host cells include the use of DNA vectors and RNA vectors. Viral vectors, particularly retroviral vectors, are the most widely used method for inserting genes into mammalian cells, such as human cells. Other viral vectors may be derived from lentiviruses, poxviruses, herpes simplex virus I, adenoviruses, and adeno-associated viruses, among others. See, for example, U.S. Patents 5,350,674 and 5,585,362.
[0251] Chemical means for introducing polynucleotides into host cells include colloidal dispersion systems such as macromolecular complexes, nanocapsules, microspheres, and beads, as well as lipid-based systems including oil-in-water emulsions, micelles, mixed micelles, and liposomes. An exemplary colloidal system for use as a delivery vehicle in vitro and in vivo is liposomes (e.g., artificial membrane vesicles).
[0252] When a nonviral delivery system is used, an exemplary delivery vehicle is a liposome. Lipid formulations are intended for the delivery of nucleic acids to host cells (in vitro, ex vivo, or in vivo). In another embodiment, the nucleic acid may be associated with a lipid. Lipid-associated nucleic acids may be encapsulated within the aqueous interior of a liposome, dispersed within the lipid bilayer of a liposome, attached to a liposome via a linking molecule associated with both the liposome and the oligonucleotide, encapsulated within a liposome, complexed with a liposome, dispersed in a lipid-containing solution, mixed with a lipid, combined with a lipid, contained as a suspension in a lipid, containing micelles or complexed with micelles, or otherwise associated with a lipid. Lipids, lipid / DNA, or lipid / expression vector-related compositions are not limited to any particular structure in solution. For example, they may exist as a bilayer structure, as micelles, or in a "broken-down" structure. They may also simply be scattered in solution, or in some cases form aggregates that are not uniform in size or shape. Lipids are fatty substances that can be naturally occurring or synthetic. For example, lipids include naturally occurring lipid droplets in the cytoplasm, as well as a class of compounds including fatty acids, alcohols, amines, amino alcohols, and long-chain aliphatic hydrocarbons such as aldehydes and their derivatives.
[0253] Lipids suitable for use can be obtained from commercial sources. For example, dimyristylphosphatidylcholine ("DMPC") can be obtained from Sigma, St. Louis, MO. Dicetyl phosphate ("DCP") can be obtained from K&K Laboratories (Plainview, NY). Cholesterol ("Choi") can be obtained from Calbiochem-Behring. Dimyristylphosphatidylglycerol ("DMPG") and other lipids can be obtained from AvantiPolarLipids, Inc. (Birmingham, AL). Storage solutions of lipids in chloroform or chloroform / methanol can be stored at approximately -20°C. Chloroform is used as the sole solvent because it evaporates more readily than methanol. "Liposomes" is a general term encompassing various monolayer and multilayer lipid vehicles formed by the formation of encapsulated lipid bilayers or aggregates. Liposomes can be characterized as having a vesicular structure with a phospholipid bilayer membrane and an internal aqueous medium. Multilayer liposomes have multiple lipid layers separated by an aqueous medium. They form spontaneously when phospholipids are suspended in an excess aqueous solution. The lipid components undergo self-rearrangement before the formation of a closed structure, trapping water and dissolved solute between the lipid bilayers (Ghosh et al., 1991 Glycobiology 5:505-10). However, compositions that have structures different from the usual vesicular structure in solution are also included. For example, lipids may take the form of micellar structures or simply exist as heterogeneous aggregates of lipid molecules. Lipofectamine-nucleic acid complexes are also considered.
[0254] Regardless of the method used to introduce exogenous nucleic acids into host cells, various assays can be performed to confirm the presence of recombinant DNA sequences in host cells. Such assays include, for example, “molecular biological” assays well known to those skilled in the art, such as Southern blotting and Northern blotting, RT-PCR and PCR; “biochemical” assays, such as immunological means (ELISA and Western blotting); or assays described herein for identifying activators that fall within the scope of the present invention, to detect the presence or absence of specific peptides.
[0255] In one embodiment, the present invention relates to a method for expressing two or more target proteins from two or more pairs of independent RNA molecules encoding portions of a target protein via cis-cleavage of ribozymes and trans-splicing of pairs of independent RNA molecules. In one embodiment, the method comprises administering one, two, or three pairs of nucleic acid molecules that encode or contain RNA molecules, such that each individual pair of independent RNA molecules has a separate reading frame, so that even if an undesirable pair is trans-spliced, the full-length functional protein will not be translated. In one embodiment, the method further comprises administering one or more selected from the group consisting of nucleic acid molecules encoding ligases and ligases to cells or tissues. In one embodiment, the ligase is an RNA 2',3'-cyclic phosphate and 5'-OH(RtcB) ligase as described herein.
[0256] In one embodiment, the present invention includes a method for delivering and expressing a full-length protein and cargo sequence of interest. In one embodiment, the method includes administering to a cell or tissue a first portion of RNA encoding a first portion of the protein of interest, ligated to a synthetic intron at its 3' end, and a second portion of RNA encoding a second portion of the protein of interest, ligated to a synthetic intron at its 5' end. In one embodiment, the synthetic intron has a 5' ribozyme sequence and a 3' ribozyme sequence adjacent to each other on both sides. In one embodiment, the synthetic intron includes a cargo sequence positioned between the 5' ribozyme sequence and the 3' ribozyme sequence. In one embodiment, autocleavage of the 5' ribozyme sequence and the 3' ribozyme sequence generates three distinct RNA molecules: 1) a first fragment containing the first portion of RNA encoding the first portion of the protein of interest, 2) a second fragment containing a synthetic intron, and 3) a third fragment containing the second portion of RNA encoding the second portion of the protein of interest. In one embodiment, the matching ends of a second fragment are ligated to produce a circular RNA molecule containing a synthetic intron with a cargo sequence. In another embodiment, the first and third fragments are ligated together to produce a single full-length linear RNA molecule. In one embodiment, the full-length protein of interest comprises a therapeutic protein, a reporter protein, a recombinase, an antibiotic resistance gene product, an antibody, or a Cas9 protein. In one embodiment, the cargo sequence comprises a therapeutic nucleic acid sequence (e.g., a miRNA sequence or a CRISPR guide RNA sequence) or encodes a therapeutic protein. In some embodiments, the full-length protein of interest comprises Cas9, and the cargo sequence comprises a guide RNA sequence, thereby targeting Cas9 to a specific genomic sequence for editing. In some embodiments, the method comprises administering a ligase or nucleic acid encoding a ligase as described herein to cells or tissues.
[0257] In one embodiment, the present invention includes a gene editing method comprising one or more trans-cleaved ribozymes. In some embodiments, the method comprises administering a first trans-cleaved ribozyme and a second trans-cleaved ribozyme, wherein the first trans-cleaved ribozyme targets upstream of the disease-causing mutation, and the second trans-cleaved ribozyme targets downstream of the disease-causing mutation. In some embodiments, trans-cleaving upstream and downstream of the disease-causing mutation results in the removal of the disease-causing mutation. In some embodiments, the remainder of the gene is trans-spliced together after trans-cleaving of the disease-causing mutation. In some embodiments, the trans-spliced gene is expressed as a functional protein.
[0258] In one embodiment, the present invention relates to an in vivo method for assembling a full-length RNA virus genome. Exemplary RNA viruses, but not limited to, include coronaviruses, paramyxoviruses, orthomyxoviruses, retroviruses, lentiviruses, alphaviruses, flaviviruses, rhabdoviruses, measles viruses, Newcastle disease viruses, and picornaviruses. In one embodiment, the method comprises administering a first nucleic acid encoding a first portion of the RNA virus genome and encoding a 3' ribozyme to a cell or tissue. In one embodiment, the method comprises administering a second nucleic acid encoding a second portion of the RNA virus genome and encoding a 5' ribozyme to a cell or tissue. In one embodiment, the method comprises administering a first RNA molecule containing the first portion of the RNA virus genome and the 3' ribozyme to a cell or tissue. In one embodiment, the method comprises administering a second RNA molecule containing the second portion of the RNA virus genome and the 5' ribozyme to a cell or tissue. In one embodiment, the method comprises administering a ligase described herein, or a nucleic acid encoding a ligase, to a cell or tissue. In one embodiment, during cis-cleavage of the 3' and 5' ribozymes, the first and second portions of the RNA virus genome are ligated together, thereby generating the full-length RNA virus genome.
[0259] In one embodiment, the present invention relates to an in vivo method for generating a circular RNA molecule encoding a target protein. In one embodiment, the method includes administering a DNA molecule or an in vitro transcribed linear RNA molecule containing a coding region encoding a first portion of the target protein (N-terminal coding sequence), a 5' ribozyme, an intervening sequence to be removed, and a 3' ribozyme directly ligated to a coding region encoding a second portion of the target protein (C-terminal coding sequence). In one embodiment, the N-terminal coding sequence is directly ligated to the 5' ribozyme sequence, and the 3' ribozyme is directly ligated to the C-terminal coding sequence. In some embodiments, the N-terminal and C-terminal coding sequences are oriented in opposite directions in the linear RNA molecule. In some embodiments, each of the 3' ribozyme and 5' ribozyme includes the sequence of SEQ ID NOs: 131, 132, 133, 134, 135, 136, 137, 138, 139, 140, 141, 142, 143, 144, 145, 166, 167, 168, 169, 170, 171, 172, 173, 174, 175, 176, 192, 193, 194, 195, 196, 197, 198, 199, 200, 202, 203, 204, 205, 206, 207, 217, 218, or 219.In some embodiments, the 5' ribozyme is SEQ ID NO: 131, SEQ ID NO: 132, SEQ ID NO: 133, SEQ ID NO: 134, SEQ ID NO: 135, SEQ ID NO: 136, SEQ ID NO: 137, SEQ ID NO: 138, SEQ ID NO: 139, SEQ ID NO: 140, SEQ ID NO: 141, SEQ ID NO: 142, SEQ ID NO: 143, SEQ ID NO: 144, SEQ ID NO: 145, SEQ ID NO: 166, SEQ ID NO: 167, SEQ ID NO: 168, SEQ ID NO: 169, SEQ ID NO: 170, SEQ ID NO: 171, SEQ ID NO: 172, SEQ ID NO: 173, SEQ ID NO: 174, SEQ ID NO: 175, SEQ ID NO: 176, SEQ ID NO: 192, SEQ ID NO: 193, SEQ ID NO: 194, SEQ ID NO: 195, SEQ ID NO: 196, SEQ ID NO: 197, SEQ ID NO: 198, SEQ ID NO: 199, SEQ ID NO: 200, SEQ ID NO: 202, SEQ ID NO: 203, SEQ ID NO: 204, SEQ ID NO: 205, SEQ ID NO: 206, SEQ ID NO: 207, SEQ ID NO: 217, SEQ ID NO: 218, or SEQ ID NO: 219. In some embodiments, the 3'-ribozyme is SEQ ID NO: 131, SEQ ID NO: 132, SEQ ID NO: 133, SEQ ID NO: 134, SEQ ID NO: 135, SEQ ID NO: 136, SEQ ID NO: 137, SEQ ID NO: 138, SEQ ID NO: 139, SEQ ID NO: 140, SEQ ID NO: 141, SEQ ID NO: 142, SEQ ID NO: 143, SEQ ID NO: 144, SEQ ID NO: 145, SEQ ID NO: 166, SEQ ID NO: 167, SEQ ID NO: 168, SEQ ID NO: 169, SEQ ID NO: 170, SEQ ID NO: 171, SEQ ID NO: 172, SEQ ID NO: 173, SEQ ID NO: 174, SEQ ID NO: 175, SEQ ID NO: 176, SEQ ID NO: 192, SEQ ID NO: 193, SEQ ID NO: 194, SEQ ID NO: 195, SEQ ID NO: 196, SEQ ID NO: 197, SEQ ID NO: 198, SEQ ID NO: 199, SEQ ID NO: 200, SEQ ID NO: 202, SEQ ID NO: 203, SEQ ID NO: 204, SEQ ID NO: 205, SEQ ID NO: 206, SEQ ID NO: 207, SEQ ID NO: 217, SEQ ID NO: 218, or SEQ ID NO: 219. In some embodiments, the 5' ribozyme comprises SEQ ID NO: 174, and the 3' ribozyme comprises SEQ ID NO: 175. In some embodiments, the 5' ribozyme comprises SEQ ID NO: 176, and the 3' ribozyme comprises SEQ ID NO: 172. In some embodiments, the coding sequence for the protein of interest is reversed orientation in the linear RNA molecule. In some embodiments, the linear RNA molecule further comprises IRES, a polyadenylated sequence, a sequence with AK recombination ability (SEQ ID NO: 180), or any combination thereof.
[0260] In one embodiment, the method includes the step of providing a linear RNA molecule comprising a 3' ribozyme, a coding region encoding a second portion of the protein of interest (C-terminal coding sequence), a coding region encoding a first portion of the protein of interest (N-terminal coding sequence), and a 5' ribozyme. In one embodiment, the N-terminal coding sequence is directly ligated to the 5' ribozyme sequence, and the 3' ribozyme is directly ligated to the C-terminal coding sequence. In some embodiments, the N-terminal and C-terminal coding sequences are oriented in opposite directions in the linear RNA molecule. In some embodiments, each of the 3' ribozyme and 5' ribozyme includes the sequence of SEQ ID NOs: 131, 132, 133, 134, 135, 136, 137, 138, 139, 140, 141, 142, 143, 144, 145, 166, 167, 168, 169, 170, 171, 172, 173, 174, 175, 176, 192, 193, 194, 195, 196, 197, 198, 199, 200, 202, 203, 204, 205, 206, 207, 217, 218, or 219. In some embodiments, the 3' ribozyme comprises SEQ ID NO: 171, and the 5' ribozyme comprises SEQ ID NO: 170. In some embodiments, the 3' ribozyme comprises SEQ ID NO: 169, and the 5' ribozyme comprises SEQ ID NO: 168. In some embodiments, the coding sequence for the protein of interest is reversed orientation in the linear RNA molecule. In some embodiments, the linear RNA molecule further comprises IRES, a polyadenylated sequence, a sequence with AK recombination ability (SEQ ID NO: 180), or any combination thereof.
[0261] In Vitro In one embodiment, the present invention includes an in vitro method for generating an RNA molecule encoding a target protein. In one embodiment, the method includes a step of providing at least two RNA molecules. In one embodiment, the step includes providing a first RNA molecule and a second RNA molecule.
[0262] In one embodiment, the first RNA molecule includes a coding region that encodes a first portion of the protein of interest. In one embodiment, the first RNA molecule includes a 3' ribozyme. In one embodiment, the first RNA molecule includes a coding region that encodes a first portion of the protein of interest and a 3' ribozyme.
[0263] In one embodiment, the second RNA molecule includes a coding region that encodes a second portion of the protein of interest. In one embodiment, the second RNA molecule includes a 5' ribozyme. In one embodiment, the second RNA molecule includes a coding region that encodes a second portion of the protein of interest and a 5' ribozyme.
[0264] In one embodiment, an in vitro method for generating an RNA molecule encoding a protein of interest further includes providing a ligase. In one embodiment, the ligase induces the assembly of RNA molecules from the coding regions of a first RNA molecule and a second RNA molecule. In one embodiment, the ligase is an RNA 2',3'-cyclic phosphate and 5'-OH(RtcB) ligase as described herein.
[0265] In one embodiment, the present invention includes an in vitro method for generating an RNA molecule encoding a target multi-domain protein. In one embodiment, the method includes a) providing a first RNA molecule, b) providing one or more additional RNA molecules, c) providing a ribozyme, and d) providing a final RNA molecule.
[0266] In one embodiment, the first RNA molecule in step a) includes a coding region encoding a first portion of the protein of interest. In one embodiment, the first RNA molecule includes a 3' ribozyme. In one embodiment, the first RNA molecule includes a coding region encoding a first portion of the protein of interest and a 3' ribozyme. In one embodiment, the 3' ribozyme catalyzes itself from the first RNA molecule, thereby generating a 3'P or 2'3'cP terminus. In one embodiment, the first RNA molecule further includes a 5' tag. In one embodiment, the 5' tag mediates the attachment of the first RNA molecule to a solid support.
[0267] In one embodiment, at least one additional RNA molecule in step b) comprises a coding region encoding the domain of the protein of interest, a 5' ribozyme, and a 3' ribozyme recognition sequence. In one embodiment, the 5' ribozyme cleaves itself to produce a 5' OH end. In one embodiment, a ligase is provided to catalyze the ligation of the first RNA molecule to one or more additional RNA molecules. In one embodiment, the ligase is an RNA 2',3'-cyclic phosphate and 5'-OH(RtcB) ligase as described herein. In one embodiment, the 3' ribozyme recognition sequence comprises a VS-S sequence as described herein.
[0268] In one embodiment, the ribozyme in step c) comprises VS-Rz as described herein. In one embodiment, VS-Rz recognizes VS-S and mediates its cleavage from one or more additional RNA molecules. In one embodiment, the cleavage produces a 3'P or 2'3'cP terminus. In one embodiment, steps b) to c) are repeated at least once to produce an RNA molecule encoding multiple domains. In one embodiment, VS-Rz is removed before repeating step b).
[0269] In one embodiment, the RNA molecule of step d) includes a coding region encoding the last portion of the protein of interest. In one embodiment, the RNA molecule encoding the last portion of the protein of interest further includes a 5' ribozyme. In one embodiment, the RNA molecule includes a coding region encoding the last portion of the protein of interest directly ligated to the 5' ribozyme. In one embodiment, the 5' ribozyme catalyzes itself from the last RNA molecule, thereby generating a 5' OH terminus. In one embodiment, a ligase is provided that catalyzes the ligation of one or more additional RNA molecules to the RNA molecule encoding the last portion of the protein of interest, thereby generating a complete RNA molecule encoding an N-terminal domain, one or more additional domains, and a C-terminal domain. In one embodiment, the ligase is an RNA 2',3'-cyclic phosphate and 5'-OH(RtcB) ligase as described herein.
[0270] In one embodiment, the present invention includes an in vitro method for generating a circular RNA molecule encoding a protein of interest. In one embodiment, the method includes the step of providing a plasmid or vector comprising a coding region (N-terminal coding sequence) encoding a first portion of the protein of interest, a 5' ribozyme, an intervening sequence to be removed, and a 3' ribozyme directly ligated to the coding region encoding a second portion of the protein of interest (C-terminal coding sequence). In one embodiment, the N-terminal coding sequence is directly ligated to the 5' ribozyme sequence, and the 3' ribozyme is directly ligated to the C-terminal coding sequence. In some embodiments, the N-terminal and C-terminal coding sequences are oriented in opposite directions in the linear RNA molecule. In some embodiments, the 3' ribozyme and the 5' ribozyme each include the sequence of SEQ ID NOs: 131, 132, 133, 134, 135, 136, 137, 138, 139, 140, 141, 142, 143, 144, 145, 166, 167, 168, 169, 170, 171, 172, 173, 174, 175, or 176. In some embodiments, the 5' ribozyme is SEQ ID NOs: 168, 170, 174, or 176. In some embodiments, the 3' ribozyme is SEQ ID NOs: 169, 171, 172, or 174. In some embodiments, the 5' ribozyme comprises SEQ ID NO: 174, and the 3' ribozyme comprises SEQ ID NO: 175. In some embodiments, the 5' ribozyme comprises SEQ ID NO: 176, and the 3' ribozyme comprises SEQ ID NO: 172. In some embodiments, the coding sequence for the protein of interest is reversed orientation in the linear RNA molecule. In some embodiments, the linear RNA molecule further comprises IRES, a polyadenylated sequence, a sequence with AK recombination ability (SEQ ID NO: 180), or any combination thereof.
[0271] In one embodiment, the method includes the step of providing a linear RNA molecule comprising a 3' ribozyme, a coding region encoding a second portion of the protein of interest (C-terminal coding sequence), a coding region encoding a first portion of the protein of interest (N-terminal coding sequence), and a 5' ribozyme. In one embodiment, the N-terminal coding sequence is directly ligated to the 5' ribozyme sequence, and the 3' ribozyme is directly ligated to the C-terminal coding sequence. In some embodiments, the N-terminal and C-terminal coding sequences are oriented in opposite directions in the linear RNA molecule. In some embodiments, each of the 3' ribozyme and 5' ribozyme includes the sequence of SEQ ID NOs: 131, 132, 133, 134, 135, 136, 137, 138, 139, 140, 141, 142, 143, 144, 145, 166, 167, 168, 169, 170, 171, 172, 173, 174, 175, 176, 192, 193, 194, 195, 196, 197, 198, 199, 200, 202, 203, 204, 205, 206, 207, 217, 218, or 219.In some embodiments, the 5' ribozyme is SEQ ID NO: 131, SEQ ID NO: 132, SEQ ID NO: 133, SEQ ID NO: 134, SEQ ID NO: 135, SEQ ID NO: 136, SEQ ID NO: 137, SEQ ID NO: 138, SEQ ID NO: 139, SEQ ID NO: 140, SEQ ID NO: 141, SEQ ID NO: 142, SEQ ID NO: 143, SEQ ID NO: 144, SEQ ID NO: 145, SEQ ID NO: 166, SEQ ID NO: 167, SEQ ID NO: 168, SEQ ID NO: 169, SEQ ID NO: 170, SEQ ID NO: 171, SEQ ID NO: 172, SEQ ID NO: 173, SEQ ID NO: 174, SEQ ID NO: 175, SEQ ID NO: 176, SEQ ID NO: 192, SEQ ID NO: 193, SEQ ID NO: 194, SEQ ID NO: 195, SEQ ID NO: 196, SEQ ID NO: 197, SEQ ID NO: 198, SEQ ID NO: 199, SEQ ID NO: 200, SEQ ID NO: 202, SEQ ID NO: 203, SEQ ID NO: 204, SEQ ID NO: 205, SEQ ID NO: 206, SEQ ID NO: 207, SEQ ID NO: 217, SEQ ID NO: 218, or SEQ ID NO: 219. In some embodiments, the 3'-ribozyme is SEQ ID NO: 131, SEQ ID NO: 132, SEQ ID NO: 133, SEQ ID NO: 134, SEQ ID NO: 135, SEQ ID NO: 136, SEQ ID NO: 137, SEQ ID NO: 138, SEQ ID NO: 139, SEQ ID NO: 140, SEQ ID NO: 141, SEQ ID NO: 142, SEQ ID NO: 143, SEQ ID NO: 144, SEQ ID NO: 145, SEQ ID NO: 166, SEQ ID NO: 167, SEQ ID NO: 168, SEQ ID NO: 169, SEQ ID NO: 170, SEQ ID NO: 171, SEQ ID NO: 172, SEQ ID NO: 173, SEQ ID NO: 174, SEQ ID NO: 175, SEQ ID NO: 176, SEQ ID NO: 192, SEQ ID NO: 193, SEQ ID NO: 194, SEQ ID NO: 195, SEQ ID NO: 196, SEQ ID NO: 197, SEQ ID NO: 198, SEQ ID NO: 199, SEQ ID NO: 200, SEQ ID NO: 202, SEQ ID NO: 203, SEQ ID NO: 204, SEQ ID NO: 205, SEQ ID NO: 206, SEQ ID NO: 207, SEQ ID NO: 217, SEQ ID NO: 218, or SEQ ID NO: 219. In some embodiments, the 3' ribozyme comprises SEQ ID NO: 171, and the 5' ribozyme comprises SEQ ID NO: 170. In some embodiments, the 3' ribozyme comprises SEQ ID NO: 169, and the 5' ribozyme comprises SEQ ID NO: 168. In some embodiments, the coding sequence for the protein of interest is reversed orientation in the linear RNA molecule. In some embodiments, the linear RNA molecule further comprises IRES, a polyadenylated sequence, a sequence with AK recombination ability (SEQ ID NO: 180), or any combination thereof.
[0272] Any RNA molecule of this disclosure can be transcribed in vitro from a template DNA called an “in vitro transcription template.” The source of DNA may be, for example, genomic DNA, plasmid DNA, phage DNA, cDNA, synthetic DNA sequences, or any other suitable DNA source. In some embodiments, the in vitro transcription template includes a 5' untranslated (UTR) region, an open reading frame, and has a 3' UTR and a poly-A tail. In some embodiments, the in vitro transcription template lacks a poly-A tail. The specific nucleic acid sequence composition and length of the in vitro transcription template depend on the mRNA or fragment thereof (e.g., an N-terminal or C-terminal fragment) encoded by the template.
[0273] In one embodiment, the 5'UTR is 0 to 3000 nucleotides long. The lengths of the 5'UTR and 3'UTR sequences appended to the coding region can be modified by different methods, including but not limited to designing PCR primers that anneal to different regions of the UTR. Using this approach, those skilled in the art can modify the lengths of the 5'UTR and 3'UTR to achieve optimal translation efficiency after transfection of the transfected RNA.
[0274] The 5'UTR and 3'UTR may be the naturally occurring endogenous 5'UTR and 3'UTR of the gene of interest. Alternatively, a non-endogenous UTR sequence for the gene of interest can be added by incorporating the UTR sequence into forward and reverse primers, or by any other modification of the template. The use of a non-endogenous UTR sequence for the gene of interest may be useful for modifying RNA stability and / or translation efficiency. For example, AU-rich elements in a 3'UTR sequence are known to reduce mRNA stability. Therefore, the 3'UTR may be selected or designed to increase the stability of the transcription RNA based on the UTR properties known in the art.
[0275] In one embodiment, the 5'UTR may contain a Kozak sequence of an endogenous gene. Alternatively, if a non-endogenous 5'UTR has been added to the gene of interest by PCR as described above, the consensus Kozak sequence can be redesigned by adding a 5'UTR sequence. While Kozak sequences can enhance the translation efficiency of some RNA transcripts, they are not considered necessary for all RNAs to enable efficient translation. The need for Kozak sequences for many mRNAs is well known in the art. In other embodiments, the 5'UTR may be derived from an RNA virus whose RNA genome is stable in the cell. In other embodiments, various nucleotide analogs can be used in the 3'UTR or 5'UTR to prevent exonuclease degradation of mRNA.
[0276] To enable RNA synthesis from a DNA template, a transcription promoter needs to be attached to the DNA template upstream of the sequence to be transcribed. When a sequence that functions as an RNA polymerase promoter is attached to the 5' end of a forward primer, the RNA polymerase promoter is incorporated into the PCR product upstream of the open reading frame to be transcribed. In one embodiment, the promoter is the T7 RNA polymerase promoter, as described elsewhere in this specification. Other useful promoters include, but are not limited to, the T3 and SP6 RNA polymerase promoters. Consensus nucleotide sequences for the T7, T3, and SP6 promoters are known in the art.
[0277] In one embodiment, mRNA has a cap at the 5' end and a poly(A) tail at the 3' end, which determine ribosome binding, translation initiation, and intracellular mRNA stability. With circular DNA templates, such as plasmid DNA, RNA polymerase produces long concatemer products, which are unsuitable for expression in eukaryotic cells. Transcription of plasmid DNA linearized at the 3'UTR end yields mRNA of normal size, which is effective in eukaryotic transfection when polyadenylated after transcription.
[0278] In linear DNA templates, phage T7 RNA polymerase can extend the 3' end of the transcript beyond the last base of the template (Schenborn and Mierendorf, Nuc Acids Res., 13:6223-36 (1985); Nacheva and Berzal-Herranz, Eur.J.Biochem., 270:1485-65 (2003)).
[0279] The conventional method for incorporating polyA / T stretches into DNA templates is molecular cloning. However, polyA / T sequences incorporated into plasmid DNA can cause plasmid instability, which can be mitigated by using recombinant-unsuitable bacterial cells for plasmid proliferation.
[0280] The RNA poly(A) tail can be further elongated after in vitro transcription using a poly(A) polymerase such as E. coli (E. coli) poly(A) polymerase (E-PAP) or yeast poly(A) polymerase. In one embodiment, increasing the length of the poly(A) tail from 100 nucleotides to 300-400 nucleotides approximately doubles the RNA translation efficiency. Furthermore, the stability of mRNA can be increased by attaching different chemical groups to the 3' end. Such attachments can include modified / artificial nucleotides, aptamers, and other compounds. For example, ATP analogs can be incorporated into the poly(A) tail using a poly(A) polymerase. ATP analogs can further increase the stability of RNA.
[0281] The 5' cap also provides stability to the mRNA molecule. In one embodiment, the RNA produced by the above method contains a 5'cap1 structure. Such a cap1 structure can be generated using a vaccinia capping enzyme and a 2'-O-methyltransferase enzyme (CellScript, Madison, WI). Alternatively, the 5' cap is known in the art and is provided using the techniques described herein (Cougot, et al., Trends in Biochem. Sci., 29:436-444 (2001); Stepinski, et al., RNA, 7:1468-95 (2001); Elango, et al., Biochim. Biophys. Res. Commun., 330:958-966 (2005)).
[0282] Certain embodiments of the present invention may utilize a solid support comprising an inert substrate or matrix (e.g., a glass slide, polymer beads, etc.) functionalized by the application of a layer or coating of an intermediate material containing reactive groups that enable covalent attachment to biomolecules such as polynucleotides. Examples of such supports, but not limited to these, include polyacrylamide hydrogels supported on an inert substrate such as glass, particularly the polyacrylamide hydrogels described in International Publication No. 2005 / 065814 and U.S. Patent Application Publication No. 2008 / 0280773, the contents of which are incorporated herein by reference in their entirety. In such embodiments, biomolecules (e.g., polynucleotides) may be directly covalently attached to the intermediate material (e.g., hydrogel), or the intermediate material itself may be noncovalently attached to the substrate or matrix (e.g., a glass substrate). The term “covalent attachment to a solid support” should be interpreted to encompass this type of arrangement accordingly.
[0283] As those skilled in the art will understand, the number of possible substrates is very large. Possible substrates include, but are not limited to, glass and modified or functionalized glass, plastics (including acrylic, polystyrene and copolymers of styrene with other materials, polypropylene, polyethylene, polybutylene, polyurethane, Teflon®, etc.), polysaccharides, nylon or nitrocellulose, ceramics, resins, silica or silica-based materials including silicon and modified silicon, carbon, metals, inorganic glass, plastics, optical fiber bundles, and various other polymers.
[0284] In some embodiments, the solid support comprises microspheres or beads. Suitable bead compositions include, but are not limited to, plastics, ceramics, glass, polystyrene, methylstyrene, acrylic polymers, paramagnetic materials, triazoles, carbon graphite, titanium dioxide, latex or crosslinked dextran, e.g., Sepharose, cellulose, nylon, crosslinked micelles and Teflon, and any other materials outlined herein for the solid support. The "Microsphere Detection Guide" (Bangs Laboratories, Fishers Ind.) is a useful guide. In certain embodiments, the microspheres are magnetic microspheres or beads.
[0285] The beads do not need to be spherical, and irregular particles may be used. Alternatively or additionally, the beads may be porous. The bead size ranges from nanometers, i.e., 100 nm, to millimeters, i.e., 1 mm, with beads of about 0.2 microns to about 200 microns being preferred, and about 0.5 to about 5 microns being particularly preferred, although in some embodiments smaller or larger beads may be used.
[0286] In one embodiment, the present invention relates to an in vitro method for assembling a full-length RNA virus genome. Exemplary RNA viruses include, but are not limited to, coronaviruses, paramyxoviruses, orthomyxoviruses, retroviruses, lentiviruses, alphaviruses, flaviviruses, rhabdoviruses, measles viruses, Newcastle disease viruses, and picornaviruses. In one embodiment, the method comprises providing a first RNA molecule comprising a first portion of the RNA virus genome and a 3' ribozyme. In one embodiment, the method comprises providing a second RNA molecule comprising a second portion of the RNA virus genome and a 5' ribozyme. In one embodiment, upon cis-cleavage of the 3' and 5' ribozymes, the first portion of the RNA virus genome and the second portion of the RNA virus genome have compatible ends for ligation, as described herein. In one embodiment, the method comprises contacting the first RNA molecule and the second RNA molecule with a ligase described herein, thereby generating a full-length RNA virus genome.
[0287] Treatment and Use The present invention provides methods for treating, alleviating, and / or reducing the risk of developing a disease or disorder in a subject. For example, in one embodiment, the method of the present invention treats, alleviates, and / or reduces the risk of developing a disease or disorder in a mammal. In one embodiment, the method of the present invention treats, alleviates, and / or reduces the risk of developing a disease or disorder in a plant. In one embodiment, the method of the present invention treats, alleviates, and / or reduces the risk of developing a disease or disorder in a yeast organism.
[0288] In one embodiment, the subject is a cell. In one embodiment, the cell is a prokaryotic cell or a eukaryotic cell. In one embodiment, the cell is a eukaryotic cell. In one embodiment, the cell is a plant, animal, or fungal cell. In one embodiment, the cell is a plant cell. In one embodiment, the cell is an animal cell. In one embodiment, the cell is a yeast cell.
[0289] In one embodiment, the subject is a mammal. For example, in one embodiment, the subject is a human, a non-human primate, a dog, a cat, a horse, a cattle, a goat, a sheep, a rabbit, a pig, a rat, or a mouse. In one embodiment, the subject is a non-mammalian subject. For example, in one embodiment, the subject is a zebrafish, a fruit fly, or a nematode.
[0290] In one embodiment, the disease or disorder is caused by a missing or defective protein whose nucleic acid sequence exceeds the packaging size of the viral vector. Therefore, in one embodiment, the disease or disorder can be treated, mitigated, or the risk reduced using the compositions, systems, and methods of the present invention. Accordingly, in one embodiment, the method comprises administering one or more compositions of the present invention to a subject. Furthermore, in one embodiment, the method comprises utilizing one or more systems of the present invention to treat, mitigate, and / or reduce the risk of developing the disease or disorder in a subject.
[0291] In one embodiment, the disease or disorder is Duchenne muscular dystrophy, Becker muscular dystrophy (BMD), autosomal recessive polycystic kidney disease, hemophilia A, Stargard macular degeneration, limb-girdle muscular dystrophy, autosomal recessive severe congenital deafness, autosomal recessive non-syndromic hearing loss (ARNSHL), sensorineural hearing loss, cystic fibrosis, Wilson's disease, Miyoshi type myopathy, autosomal recessive hearing loss type 9 (DFNB9), Usher syndrome type I, GJB2-associated autosomal recessive non-syndromic hearing loss (GJB2-AR NSHL), Autosomal recessive cerebellar parenchymal disorder type 3, Non-syndromic hearing loss, Autosomal recessive hearing loss type 16 (DFNB16), Meniere's disease, Autosomal dominant non-syndromic sensorineural hearing loss type 12 (DFNA12), Autosomal recessive spinocerebellar ataxia type 21 (SCAR21), Usher syndrome type 1F (USH1F), Autosomal recessive hearing loss type 23 (DFNB23), Autosomal recessive hearing loss type 30 (DFNB30), Oto-spine-megaly epiphysis dysplasia (OSMED), Autosomal recessive hearing loss type 77 (DFNB77), Autosomal recessive hearing loss type 84A (DFNB84A), Autosomal recessive hearing loss type 84B (DFNB84B), Peripheral neuropathy These include disorders, myopathy, autosomal dominant non-syndromic hearing loss type 4A (DFNA4), congenital thrombocytopenia, sensorineural hearing loss, autosomal dominant non-syndromic hearing loss type 56 (DFNA56), epileptic encephalopathy, Timothy syndrome, long QT syndrome, X-linked retinal disease, aldosteronism, autosomal recessive hearing loss type 42 (DFNB42), primary aldosteronism (Conn syndrome), seizures, neurological abnormalities, sinoatrial node dysfunction, neurodevelopmental disorders, hypokalemic periodic paralysis, epilepsy, developmental epileptic encephalopathy, Brodymyopathy, Darier's disease, heart disease, von Willebrand disease, and Zellweger syndrome. In one embodiment, the disease or disorder is caused by any of the genetic mutations that accept CRISPR-Cas9-mediated editing.
[0292] In one embodiment, the method of the present invention involves administering to a subject having Duchenne muscular dystrophy a composition comprising a first nucleic acid containing a coding region encoding a first portion of dystrophin and a 3' ribozyme, and a second nucleic acid containing a coding region encoding a second portion of dystrophin and a 5' ribozyme, wherein the first nucleic acid transcribes a first RNA molecule, the second nucleic acid transcribes a second RNA molecule, cis-cleavage of the 3' and 5' ribozymes and trans-splicing of the coding regions encoding the first and second portions of dystrophin occur to produce a single RNA molecule encoding the full-length dystrophin protein.
[0293] In one embodiment, the method of the present invention involves administering to a subject having Duchenne muscular dystrophy a composition comprising a first nucleic acid encoding the nucleic acid sequence of SEQ ID NO: 129 and a second nucleic acid encoding the nucleic acid sequence of SEQ ID NO: 130, wherein transcription of the first nucleic acid yields a first RNA molecule, transcription of the second nucleic acid molecule yields a second RNA molecule, and cis-cleavage of 3' and 5' ribozymes and trans-splicing of the first and second RNA molecules generate a single RNA molecule encoding dystrophin protein.
[0294] In one embodiment, the method of the present invention comprises administering to a subject having Duchenne muscular dystrophy a composition comprising a first nucleic acid encoding the nucleic acid sequence of SEQ ID NO: 22 and a second nucleic acid encoding the nucleic acid sequence of SEQ ID NO: 23, wherein the first nucleic acid transcribes a first RNA molecule, the second nucleic acid transcribes a second RNA molecule, and cis-cleavage of the 3' and 5' ribozymes and trans-splicing of the first and second RNA molecules generate a single RNA molecule encoding a full-length dystrophin protein having a C-terminal GFP reporter. In one embodiment, the second nucleic acid encodes a fragment of SEQ ID NO: 23, wherein the fragment does not contain a coding sequence for the C-terminal GFP reporter.
[0295] In one embodiment, the method involves administering to a subject having Duchenne muscular dystrophy a composition comprising a first RNA molecule encoding a first portion of dystrophin and containing a 3' ribozyme, and a second RNA molecule encoding a second portion of dystrophin and containing a 5' ribozyme, wherein a single RNA molecule encoding the dystrophin protein is generated by cis-cleavage of the 3' and 5' ribozymes and trans-splicing of the first and second RNA molecules.
[0296] In one embodiment, the method involves administering to a subject having Duchenne muscular dystrophy a composition comprising a first RNA molecule containing the nucleic acid sequence of SEQ ID NO: 129 and a second RNA molecule containing the nucleic acid sequence of SEQ ID NO: 130, wherein a single RNA molecule encoding a minidystrophin protein is generated by cis-cleavage and trans-splicing of the 3' and 5' ribozymes of the first and second RNA molecules.
[0297] In one embodiment, the method comprises administering to a subject having Duchenne muscular dystrophy a composition comprising a first RNA molecule containing the nucleic acid sequence of SEQ ID NO: 22 and a second RNA molecule containing the nucleic acid sequence of SEQ ID NO: 23, thereby generating a single RNA molecule encoding a full-length dystrophin protein having a C-terminal GFP reporter by cis-cleavage of the 3' and 5' ribozymes and trans-splicing of the first and second RNA molecules. In one embodiment, the second nucleic acid encodes the fragment of SEQ ID NO: 23, wherein the fragment does not contain the coding sequence for the C-terminal GFP reporter.
[0298] In one embodiment, the method of the present invention involves administering to a subject having one or more diseases selected from Table 1 a composition comprising a first nucleic acid containing a coding region encoding a first portion of a therapeutic protein corresponding to the related disease in Table 1 and a 3' ribozyme, and a second nucleic acid containing a coding region encoding a second portion of a therapeutic protein corresponding to the related disease in Table 1 and a 5' ribozyme, wherein the first nucleic acid transcribes a first RNA molecule, the second nucleic acid transcribes a second RNA molecule, and performs cis-cleavage of the 3' and 5' ribozymes, as well as trans-splicing of the coding regions encoding the first portion of the therapeutic protein and the second portion of the therapeutic protein, thereby generating a single RNA molecule encoding the full-length therapeutic protein.
[0299] In one embodiment, the method involves administering to a subject having one or more diseases selected from Table 1 a composition comprising a first RNA molecule encoding a first portion of a therapeutic protein corresponding to an associated disease in Table 1 and containing a 3' ribozyme, and a second RNA molecule encoding a second portion of a therapeutic protein corresponding to an associated disease in Table 1 and containing a 5' ribozyme, wherein a single RNA molecule encoding the full-length therapeutic protein is generated by cis-cleavage of the 3' and 5' ribozymes and trans-splicing of the first and second RNA molecules. [Table 1] JPEG2026509506000003.jpg188159
[0300] In one embodiment, the method of the present invention involves administering to a subject having Duchenne muscular dystrophy a composition comprising a first nucleic acid comprising the nucleic acid sequence of SEQ ID NO: 150 and a second nucleic acid comprising the nucleic acid sequence of SEQ ID NO: 151, wherein the transcription of the first nucleic acid yields a first RNA molecule, the transcription of the second nucleic acid molecule yields a second RNA molecule, and cis-cleavage of 3' and 5' ribozymes and trans-splicing of the first and second RNA molecules generate a single RNA molecule encoding a minidystrophin protein.
[0301] In one embodiment, the method of the present invention involves administering to a subject having Duchenne muscular dystrophy a composition comprising a first nucleic acid containing the nucleic acid sequence of SEQ ID NO: 152 and a second nucleic acid containing the nucleic acid sequence of SEQ ID NO: 153, wherein the transcription of the first nucleic acid yields a first RNA molecule, the transcription of the second nucleic acid molecule yields a second RNA molecule, and cis-cleavage of 3' and 5' ribozymes and trans-splicing of the first and second RNA molecules generate a single RNA molecule encoding a minidystrophin protein.
[0302] In one embodiment, the method of the present invention involves administering a composition comprising a nucleic acid molecule containing the nucleic acid sequence of Sequence ID No. 162 to a subject having Duchenne muscular dystrophy, wherein transcription of the nucleic acid yields a linear RNA molecule, and cis-cleavage and cis-ligation of the 3' and 5' ribozymes of the RNA molecule generate a single circular RNA molecule encoding a dystrophin protein.
[0303] In one embodiment, the method of the present invention involves administering a composition comprising a nucleic acid molecule containing the nucleic acid sequence of Sequence ID No. 163 to a subject having Duchenne muscular dystrophy, wherein transcription of the nucleic acid yields a linear RNA molecule, and cis-cleavage and cis-ligation of the 3' and 5' ribozymes of the RNA molecule generate a single circular RNA molecule encoding a dystrophin protein.
[0304] In one embodiment, the method of the present invention involves administering a composition comprising a nucleic acid molecule containing the nucleic acid sequence of Sequence ID No. 164 to a subject having Duchenne muscular dystrophy, wherein transcription of the nucleic acid yields a linear RNA molecule, and cis-cleavage and cis-ligation of the 3' and 5' ribozymes of the RNA molecule generate a single circular RNA molecule encoding a dystrophin protein.
[0305] In one embodiment, the method of the present invention involves administering a composition comprising a nucleic acid molecule containing the nucleic acid sequence of Sequence ID No. 165 to a subject having Duchenne muscular dystrophy, wherein transcription of the nucleic acid yields a linear RNA molecule, and cis-cleavage and cis-ligation of the 3' and 5' ribozymes of the RNA molecule generate a single circular RNA molecule encoding a dystrophin protein.
[0306] In one embodiment, the method of the present invention involves administering a composition comprising a first nucleic acid containing the nucleic acid sequence of SEQ ID NO: 157 and a second nucleic acid containing the nucleic acid sequence of SEQ ID NO: 158 to a subject having dysferlin disorder or myopathy or mitochondrial dysfunction due to a mutation in dysferlin, wherein a first RNA molecule is produced by transcription of the first nucleic acid, a second RNA molecule is produced by transcription of the second nucleic acid molecule, and a single RNA molecule encoding the dysferlin protein is produced by cis-cleavage of the 3' and 5' ribozymes and trans-splicing of the first and second RNA molecules.
[0307] In one embodiment, the method of the present invention involves administering a composition comprising a first nucleic acid containing the nucleic acid sequence of SEQ ID NO: 177 and a second nucleic acid containing the nucleic acid sequence of SEQ ID NO: 178 to a subject having dysferlin disorder or myopathy or mitochondrial dysfunction due to a mutation in dysferlin, wherein a first RNA molecule is produced by transcription of the first nucleic acid, a second RNA molecule is produced by transcription of the second nucleic acid molecule, and a single RNA molecule encoding the dysferlin protein is produced by cis-cleavage of the 3' and 5' ribozymes and trans-splicing of the first and second RNA molecules.
[0308] In one embodiment, the method of the present invention involves administering a composition comprising a first nucleic acid containing the nucleic acid sequence of SEQ ID NO: 179 and a second nucleic acid containing the nucleic acid sequence of SEQ ID NO: 178 to a subject having dysferlin disorder or myopathy or mitochondrial dysfunction due to a mutation in dysferlin, wherein a first RNA molecule is produced by transcription of the first nucleic acid, a second RNA molecule is produced by transcription of the second nucleic acid molecule, and a single RNA molecule encoding the dysferlin protein is produced by cis-cleavage of the 3' and 5' ribozymes and trans-splicing of the first and second RNA molecules.
[0309] In one embodiment, the method of the present invention involves administering a composition comprising a first nucleic acid containing the nucleic acid sequence of SEQ ID NO: 181 and a second nucleic acid containing the nucleic acid sequence of SEQ ID NO: 182 to a subject having dysferlin disorder or myopathy or mitochondrial dysfunction due to a mutation in dysferlin, wherein a first RNA molecule is produced by transcription of the first nucleic acid, a second RNA molecule is produced by transcription of the second nucleic acid molecule, and a single RNA molecule encoding the dysferlin protein is produced by cis-cleavage of the 3' and 5' ribozymes and trans-splicing of the first and second RNA molecules.
[0310] In one embodiment, the method of the present invention involves administering a composition comprising a first nucleic acid comprising the nucleic acid sequence of SEQ ID NO: 183 and a second nucleic acid comprising the nucleic acid sequence of SEQ ID NO: 182 to a subject having dysferlin disorder or myopathy or mitochondrial dysfunction due to a mutation in dysferlin, wherein a first RNA molecule is produced by transcription of the first nucleic acid, a second RNA molecule is produced by transcription of the second nucleic acid molecule, and a single RNA molecule encoding the dysferlin protein is produced by cis-cleavage of the 3' and 5' ribozymes and trans-splicing of the first and second RNA molecules.
[0311] In one embodiment, the method of the present invention involves administering a composition comprising a first nucleic acid containing the nucleic acid sequence of SEQ ID NO: 160 and a second nucleic acid containing the nucleic acid sequence of SEQ ID NO: 161 to a subject having autosomal recessive non-syndromic sensorineural hearing loss-16 or hearing loss due to a mutation in sterocillin (STRC), wherein the transcription of the first nucleic acid yields a first RNA molecule, the transcription of the second nucleic acid molecule yields a second RNA molecule, and cis-cleavage of 3' and 5' ribozymes and trans-splicing of the first and second RNA molecules generate a single RNA molecule encoding a sterocillin protein.
[0312] Experimental example The present invention will be described in more detail with reference to the following experimental examples. These examples are provided for illustrative purposes only and are not intended to limit the invention unless otherwise specified. Therefore, the present invention should not be construed as being limited in any way to the following examples, but rather as encompassing all possible modifications that become apparent as a result of the teachings provided herein.
[0313] Unless further explanation is provided, it is likely that those skilled in the art can construct and utilize the present invention and implement the claimed method by using the foregoing description and the following exemplary embodiments. Therefore, the following embodiments should not be construed as limiting the remainder of this disclosure.
[0314] Example 1: Assembly and expression of ribozyme-mediated RNA in mammalian cells Ribozymes (Rzs) are small catalytic RNA sequences capable of nucleotide-specific autocleavage (Doherty and Doudna 2000). Ribozyme-mediated RNA cleavage generates unique 3'-phosphate and 5'-hydroxy termini, which are analogous to substrates of the ubiquitous RNA repair pathway present in all three kingdoms of life. As shown herein, ribozyme-mediated cis-cleavage can be utilized for trans-splicing of independent RNA transcripts in mammalian cells, an approach called stitchR (stitch RNA). Notably, messenger RNA reconstitution by stitchR has enabled efficient translation and expression of full-length proteins in mammalian cells. As demonstrated, stitchR can be utilized for the delivery and expression of protein-coding functional domains or large protein-coding sequences via viral vectors. Furthermore, overexpression of RNA 2',3'-cyclic phosphate and 5'-OH(RtcB) ligase enhances stitchR activity in mammalian cells and is sufficient to catalyze stitchR activity in vitro. These data characterize a novel approach that utilizes ribozymes for scarless splicing of functional RNA in cells, which could be useful for countless research and therapeutic applications.
[0315] Autocatalytic RNA sequences are widespread in nature and catalyze diverse biological processes, including intron splicing, rolling-circle viral genome replication, and peptide bond formation (Weinberg et al. 2019). At least seven major ribozyme families, including hammerhead (HH), hepatitis delta virus (HDV), barkood satellite (VS), Sister, Twister-sister, hairpin, Hatchet, and Pistol, have been identified with distinct sequence and structural features. The most widely studied are the HH, HDV, and Twister family members, which have been utilized in vitro and in vivo due to their small size and cleavage properties to generate RNA with precise ends lacking ribozyme sequences (Figure 13) (Ferre-D'Amare and Doudna 1996; Avis et al. 2012; Zhang et al. 2017).
[0316] In prokaryotes and eukaryotes, most cellular RNA is synthesized and spliced with 5'-phosphate (P) and 3'-hydroxyl (OH) ends, including messenger RNA and long non-coding RNAs. In contrast, atypical cis-splicing of many tRNAs and mRNA encoding the ER stress-responsive protein XBP1 is catalyzed by an enzymatic pathway, resulting in either a unique 5'-OH and 3'-P or 2'3' cyclic phosphate (cP) end. Recent findings suggest that atypical cis-splicing of RNA is catalyzed in mammals by ubiquitous RNA 2',3'-cyclic phosphate and 5'-OH (RtcB) ligases. Furthermore, RtcB and several other enzyme families may function to repair host cellular RNA damaged by stress or exogenous ribotoxins. Ribozyme-mediated cleavage results in similar ends, so ribozyme-cleaved RNA can undergo trans-splicing by the endogenous RNA repair pathway.
[0317] Ribozyme-cleaved mRNA is transspliced and translated in mammalian cells. To determine whether ribozymes can be used for scarless splicing of RNA in mammalian cells, we designed two expression plasmids containing non-duplication N-terminal (Nt) and C-terminal (Ct) fragments of the fluorescent reporter GFP (Nt-GFP and Ct-GFP, respectively). The ribozymes were designed to catalyze the removal of themselves from adjacent nucleotides of GFP fragments containing a 3'HDV ribozyme on Nt-GFP and a 5'HH ribozyme on Ct-GFP (Figure 1A). Expression of either Nt or Ct encoding GFP-ribozyme RNA alone did not result in detectable GFP fluorescence when transfected into mammalian COS-7 or HEK293T cells (Figure 1B). Notably, co-expression of Nt-GFP encoding RNA and Ct-GFP encoding RNA emitted green fluorescence after 48 hours (Figure 1B). RT-PCR analysis and Sanger sequencing revealed that trans-splicing of distinct Nt-GFP RNA and Ct-GFP RNA occurred between predicted ribozyme-catalyzed cleavage sites (Figure 1C and Figure 1D). Furthermore, full-length GFP protein was detected by Western blotting in co-transfected cells (Figure 1E). These data demonstrate that the endogenous mammalian cellular RNA repair pathway was sufficient to catalyze trans-splicing of independent ribozyme-processing RNAs efficiently translated into full-length proteins. This RNA trans-splicing approach was named stitchR.
[0318] Influence of ribozyme sequence and type on ribozyme-mediated transsplicing To accurately quantify the relative amount of functional full-length protein produced by ribozyme-mediated trans-splicing in cells, reporters were generated using two non-overlapping halves of firefly luciferase (Figure 2A). Consistent with our previous findings, only simultaneous transfection of both RNA encoding Nt- and Ct-luciferase ribozymes resulted in trans-splicing and luciferase activity in cells (Figures 2B and 2C). Using this assay, we further characterized the effects of different HH and HDV ribozyme sequences on trans-splicing activity in mammalian cells. A 6-base pair (bp) duplication in the stem 1 HH ribozyme resulted in maximum luciferase activity, while mutations in the HH catalytic residue led to loss of activity, consistent with previous reports on HH ribozyme activity characterized in vitro (Figure 2D). Furthermore, genomic HDV ribozyme sequences and anti-genomic HDV ribozyme sequences were equivalent in luciferase activity, with the exception of the minimum 56-nucleotide HDV ribozyme (HDV56), which showed significantly reduced activity (Figure 2E). Also, consistent with previous reports, mutation from C to U in the nucleotides required for HDV catalysis resulted in complete loss of luciferase activity (Figure 2E). These findings demonstrate that ribozyme-mediated trans-splicing activity is dependent on ribozyme cleavage in mammalian cells.
[0319] Prevention of undesirable or truncated protein expression from Nt or Ct vectors using translational control and / or proteolytic sequences. Nt RNA or Ct RNA, if translated before ribozyme-mediated cleavage or expressed separately, can potentially lead to the expression of undesirable or truncated proteins. To limit the expression of unspliced Nt or Ct vectors, the effectiveness of previously characterized translational control of proteolytic sequences on the stability of full-length GFP-encoding vectors was tested. Addition of HDV ribozymes to the 3' end of GFP was not thought to alter GFP fluorescence (Figures 3A and 3B). To selectively prevent GFP expression, the effects of proteolytic sequences hCL1-PEST, E1A-PEST, removal of the poly(A) sequence from the vector, or generation of a polyK tail by simulated translation via the polyA tail were tested (Figures 3A and 3B). All degradation sequences were cloned in GFP open reading frame and in frame so that translation occurred via the HDV ribozyme sequence. Including hCL1-PEST showed a strong decrease in GFP fluorescence, while EF1a PEST did not. Deletion of the vector poly(A) sequence from the expression vector resulted in complete loss of GFP expression, and translation via the poly(A) sequence to generate the poly(K) tail also led to a decrease in fluorescence.
[0320] In the case of a Ct-encoded GFP reporter, inclusion of the 5'HH ribozyme and deletion of the GFP start codon (ATG) still resulted in weak but detectable GFP expression despite the absence of a predicted upstream alternative ATG (Figures 3C and 3D). Further silent mutations to the N-terminal NTG codon of GFP (GFPcdn) further reduced GFP detection, but weak fluorescence was still observed. Including the 5'UTR of the yeast GCN4 gene encoding four small upstream ORFs that function as translation inhibitors eliminated detectable GFP fluorescence. Smaller internal fragments of the GCN4 5'UTR containing only the four uORFs were similarly effective in preventing GFP expression. These data demonstrate that undesirable protein expression from individual Nt or Ct vectors can be prevented by utilizing the translational regulatory sequences of proteolysis.
[0321] These translational control or proteolytic sequences can be used in other dual-vector applications where it is desirable to restrict undesirable or truncated protein expression, such as dual-AAV vector strategies that rely on homologous recombination to generate large protein-coded open reading frames.
[0322] Single and multiple transsplicing of functional protein-coding RNAs To determine whether ribozyme-mediated trans-splicing can be used to combine protein-coding functional domains within cells, we generated RNA encoding four copies of a mitochondrial targeting sequence (NT-4xMTS) and an open reading frame encoding full-length GFP lacking the ATG start codon (Ct-GFP) (Figure 4A). Co-expression of these two independent RNAs resulted in strong expression of mitochondrially localized GFP overlapping with the red fluorescent mitochondrial marker MitoTracker Red CMXRos (Figure 4B). These findings demonstrate that ribozyme-mediated trans-splicing can be used to rapidly combine two independent RNAs and express a specific functional fusion protein in cells.
[0323] Due to the three open reading frames from which the protein is translated, ribozyme-mediated transsplicing and the expression of multiple different functional proteins may also be possible. By leveraging this feature, functional proteins can be generated using RNA transsplicing adapted to three different open reading frames. To demonstrate this functionality, we designed additional ribozyme pairs in reading frame 2 (F2) encoding a myristoylation membrane-targeting sequence (Nt-F2-Myr) and a red fluorescent protein (Ct-F2-RFP) (Figure 4C). These Nt and Ct vector pairs also contained hCL1-PEST proteolytic sequences and GCN4 translation inhibitor sequences to restrict the expression of truncated proteins from the individual Nt and Ct vectors, respectively. In co-transfected cells, GFP fluorescence was highly specific to mitochondria and RFP fluorescence was highly specific to the membrane (Figure 4D), demonstrating the ability of this approach for RNA transsplicing to generate different functional proteins within the cell.
[0324] Optimized ribozymes enhance protein expression in ribozyme-mediated trans-splicing. Slight sequence modifications can significantly affect ribozyme catalytic activity by altering secondary structure, stability, or binding to metal ion cofactors. Using our trans-splice single-ciferase reporter assay, we identified improved ribozyme types and sequence modifications that enhance trans-splice single-ciferase reporter activity in mammalian cells (Figure 16). The RzB hammerhead variant ribozyme containing a tertiary stabilizing motif (TSM) showed higher activity than the ribozyme without the TSM (Figure 16A). Furthermore, the Twister (twst) ribozyme showed greater activity than the HDV ribozyme when the 3' was cloned to Nt-Luc. Catalytic mutations within the Twister ribozyme could similarly eliminate luciferase activity (Figure 16B), and these depend on P1 stem formation (Figure 16C). Since the Twister ribozyme requires U at position 1, this requirement may limit the design of scarless splicing to sequences ending in U. Therefore, the inventors tested whether nucleotide substitutions could be tolerated at position 1 and found that U1A did not show significantly different activity, but U1C or U1G substitutions retained activity, albeit with some reduction (Figure 16C).
[0325] Optimized splice donor and acceptor sequences enhance protein expression in ribozyme-mediated trans-splicing. Pre-mRNA splicing by spliceosomes has been shown to enhance mRNA translation by depositing factors that promote the first round of translation, or by enhancing RNA processing and transport to the cytoplasm. It has also been shown to enhance transgene protein expression by adding chimeric cis-splicing introns to transgenes. We then investigated whether trans-spliced RNA can undergo cis-splicing by spliceosomes and whether this affects the translation and expression of trans-spliced mRNA. To test this, splice donor (SD) and splice acceptor (SA) sequences were incorporated into a trans-splicing GFP reporter so that the trans-spliced RNA would reconstitute chimeric introns (Figure 5A). Notably, the addition of SD and SA sequences resulted in a robust enhancement of GFP fluorescence compared to trans-splicing GFP reporters without SD or SA sequences (Figure 5B). RT-PCR and Sanger sequencing demonstrated that trans- and cis-splicing occurred for both Nt-GFP RNA and Ct-GFP RNA, including SD and SA sequences, resulting in the restoration of a normal GFP open reading frame (data not shown). These data suggest that trans-splicing can occur in the nucleus, and that subsequent cis-splicing is a useful strategy for enhancing expression from trans-spliced RNA.
[0326] Ribozyme-mediated transsplicing and large gene sequence expression for delivery using viral therapeutic vectors Ribozyme-mediated transsplicing can be used to deliver and express large protein-coding mRNAs that exceed the packaging size limits of therapeutic viral gene therapy vectors such as AAV (Figure 6A). This may be useful for restoring the expression of large, mutated genes in numerous human monogenetic diseases, such as dystrophin (Dys) in Duchenne muscular dystrophy (DMD), CFTR in cystic fibrosis (CF), and factor VIII (F8) in hemophilia A. In cell-based transfection assays, co-expression of vectors encoding Nt and Ct spirit μdystrophin with a C-terminal GFP tag was transspliced in mammalian cells (Figures 6B and 6C) and localized to the membrane (Figure 6D). These data demonstrate the feasibility of using ribozyme-mediated transsplicing to reconstitute and express large protein-coding genes.
[0327] Lentiviral delivery of ribozyme-responsive RNA for intracellular transsplicing The autocatalytic cleavage of ribozymes can interfere with the packaging of ribozyme-coding RNA by positive-strand RNA viruses, such as commonly used gamma-retroviruses and lentiviral vectors. To circumvent this potential problem, Nt and Ct split GFP expression cassettes were encoded on the negative-sense strand of a third-generation lentiviral vector backbone (Figure 7A). Lentiviral particles were prepared separately for each Nt and Ct vector and subsequently used for transduction into HEK293T cells. Cells transduced with both Nt-GFP and Ct-GFP showed green fluorescence expression, while cells transduced with either Nt-GFP or Ct-GFP alone showed no detectable fluorescence (Figure 7B). These data demonstrate that lentiviral vectors are capable of delivering and expressing RNA encoding ribozymes for trans-splicing.
[0328] This approach may also be useful for delivering large gene sequences that exceed the packaging size of these viral vectors, such as Dys (Figure 7C). Furthermore, ribozyme-mediated transsplicing may enable the safe handling or reconfiguration of viral genomes, such as lentiviral or large coronavirus RNA genomes.
[0329] Safe handling, delivery, and expression of toxic or antiviral genes using viral vectors. Furthermore, ribozyme-mediated trans-splicing may enable the safe handling or reconstitution of toxic or antiviral proteins that can inhibit the generation of lentiviral particles in mammalian packaging cells. These include several cell suicide genes, such as translation-inhibitory diphtheria toxin A (DTA) (Figure 8A). We show that a vector encoding a split DTA sequence inhibits the co-expression of a CS2GFP reporter construct during trans-splicing and expression, consistent with DTA's translation-inhibiting role in mammalian cells (Figure 8B).
[0330] Enzymes that enhance or inhibit ribozyme-mediated transsplicing Several enzyme families have been suggested to ligate the 5'-OH group to either a 3'-P or 2'3' cyclic phosphate (cP) terminus, and most notably, RtcB has been found to be conserved across all three domains of life. Human codon-optimized RtcB orthologues (from eukaryotes (H. sapiens), bacteria (E. coli), and archaea (P. horikoshii)) were cloned and co-expressed, and their effects on the activity of transspliced single luciferase reporters were measured. Interestingly, co-expression of RtcB from P. horikoshii resulted in a 4.5-fold enhancement of luciferase activity, while human and bacterial orthologues showed moderate enhancement or no enhancement, respectively (Figure 9).
[0331] Other enzyme families have been shown to regulate these RNA ends. Interestingly, expression of T4 polynucleotide kinase (T4PNK), which acts as a 5'-hydroxyl kinase, 3'-phosphatase, and 2',3'-cyclic phosphodiesterase, significantly inhibited luciferase activity (Figure 9). These data suggest that co-expression of exogenous enzymes can enhance or inhibit ribozyme-mediated transsplicing in mammalian cells.
[0332] RtcB is sufficient to catalyze ribozyme-mediated RNA transsplicing in vitro. Due to their nucleotide-specific cleavage, ribozymes are widely used in vitro to generate precise RNA ends. Next, we investigated whether directional trans-splicing of independently synthesized RNA could be performed in vitro using ribozymes. Using in vitro RNA transcription of Nt- and Ct-luciferase ribozyme reporter constructs with T7 RNA polymerase, it was found that the addition of recombinant E. coli (E. coli) RtcB was necessary and sufficient to catalyze trans-splicing detected by RT-PCR (Figures 10A and 10B). Similarly, we designed RNA encoding the domain of the spider protein spidoin (Figure 10C). Spidoin is the main component of spider silk and is a material highly valued for its tensile properties, but its highly repeatable protein nature has made it difficult to synthesize in heterologous systems. Spidoin is naturally composed of multiple A-repeats and Q-repeats and has conserved N-terminal (N1L) and C-terminal (N3R) domains at both ends. Following in vitro synthesis of spidoin RNA using T7 polymerase, it was found that the addition of recombinant RtcB ligase derived from Escherichia coli (E. coli) was sufficient to catalyze the transligation of ribozyme-cleaved N1L and N3R coding RNAs, as detected by RT-PCR and Sanger sequencing (Figure 10D).
[0333] Regulated tandem trans splicing of RNA encoding multidomain proteins Next, we investigated whether adding a third RNA encoding an AQ fusion domain with ribozymes at both ends would result in an uncontrolled tandem repeat assembly (Figure 11A). While directional trans-splicing between each of the separate RNAs could be detected, the assembly of three or more independent RNA fragments could not be detected (data not shown). This may be due to the rapid circularization of the RNA, both containing ends suitable for ligation by RtcB. As an alternative approach, utilizing a trans-activated VS ribozyme has the potential to enable continuous and controlled assembly of RNA sequences in vitro (Figures 11B and 11C). In this approach, the 3'-terminal RNA ribozyme is suitable for ligation by RtcB only when addition and trans-cleavage by VS-Rz are performed. Since VS-Rz trans-activated ribozyme RNA is not covalently attached, stepwise addition of stitchR-compatible RNA, VS-Rz, and RtcB ligase may enable controlled tandem assembly of RNA sequences, which could be useful for assembling repeat RNAs encoding biologically or industrially important proteins such as synthetic spider silk, elastin, and collagen.
[0334] Trans-splicing of endogenous RNA using trans-cleaved ribozymes - therapeutic application to correct disease-causing mutations Ribozymes are autocatalytic RNAs that cleave in cis to produce the unique RNA ends shown by the inventors, which are then trans-spliced and expressed in mammalian cells (Figure 12A). Notably, cis-cleaved ribozymes can be manipulated to cleave in trans so that the target RNA is cleaved in a nucleotide-specific manner, resulting in similar RNA ends (Figure 12B) (Carbonell et al. 2011; Webb and Luptak 2018). Thus, trans-cleaved ribozymes can be used to catalyze scarless splicing of RNA in cells or in vitro. This approach could be useful for a multitude of applications, one of which is the deletion of disease-causing mutations in gene transcripts by targeting mutant flanking sequences in either exon or intron sequences (Figures 12C and 12D).
[0335] In conclusion, this specification demonstrates that ribozyme-mediated cleavage of independent RNAs expressed in cells can be efficiently assembled and translated in mammalian cells. This approach, referred to herein as stitchR, has the potential to function as a novel method for combinatorial assembly of functional RNAs and proteins for both basic and therapeutic applications. Due to the autocatalytic nature of ribozymes and the endogenous RNA repair pathway present in cells, stitchR requires the expression of only separate RNAs for trans-splicing and translation to occur in cells. In vitro, RtcB ligase has been demonstrated to be sufficient for trans-splicing, and furthermore, given that RtcB is ubiquitous and widespread across all three kingdoms of life, stitchR has the potential to be a useful approach in a diverse range of organisms.
[0336] The robust nature of this system depends on the efficient and precise properties of ribozyme-mediated RNA cleavage, which produces reliable and accurate nucleotide-specific ends essential for the recovery of the open reading frame encoding proteins. Furthermore, the ability to generate RNA using ribozymes that completely catalyze their own removal enables scarless assembly, resulting in RNA that is essentially indistinguishable from its native counterpart.
[0337] While ribozyme cleavage has been extensively studied in vitro, in vivo ribozyme cleavage is not well understood and is thought to be influenced by the availability of metal ions necessary for folding and catalysis through interaction with RNA-binding proteins. StitchR functions as an indirect readout of ribozyme-mediated cleavage and, interestingly, is found herein to be significantly influenced by changes in ribozyme sequence and structure. This suggests that optimizing ribozyme cleavage may be a useful approach to enhance stitchR activity in vivo. Further analysis of the effects of RNA repair pathway components such as RtcB, RtcA, and artiases may also serve as important factors in regulating stitchR activity.
[0338] While ribozymes naturally evolved to function in cis to facilitate their self-cleavage, some ribozyme families (particularly HDV and HH) have been engineered to cleave target RNA in trans. It is noteworthy that combining trans-cleaving ribozymes with stitchR could enable even more robust RNA cleavage and repair methods in cells or in vitro. This approach could serve as a nucleotide-specific "cut and paste" approach to RNA, potentially useful for generating RNA diversity or removing specific harmful mutations in disease-causing RNA.
[0339] Example 2: Inducible trans-splicing and RNA expression using trans-activated ribozymes Most ribozymes are autocatalytic, requiring only metal ions as cofactors, readily found in biological environments, and assist in folding and chemocatalysis. When donor RNA terminates at a G nucleotide, barkood satellite (VS) ribozymes can be utilized for scarless splicing. Interestingly, VS ribozymes can be modified to enable ribozyme transactivation and induce catalytic activity (Guo and Collins 1995; Ouellet et al. 2009). When split into two components, the smaller VS stem-loop (VS-S) alone is not sufficient to induce cis-cleavage, but the addition of the remaining sequence VS-Rz promotes efficient cleavage of VS-S (Figure 14A). This transactivation feature can enable inducible ribozyme-mediated trans-cleavage, and VS-S cleavage on Nt donor RNA requires the addition of the VS-Rz sequence, which may then be favorable for trans-splicing with Ct acceptor RNA containing the 5'-OH terminus (Figure 14B). VS-Rz sequences containing typical 5'-P- and 3'-OH RNA ends cannot participate in transsplicing and therefore can function as multi-turnover catalysts for the reaction.
[0340] For example, the ability to control ribozyme-mediated cleavage by the necessary addition of transactivating sequences such as VS-Rz can enable controlled addition of variable or non-variable RNA sequences, thereby generating synthetic repeat RNA (Figure 14C). One approach is to generate RNA with a unique N-terminal domain, a unique C-terminal domain, and an internal variable or non-variable "repeat" domain. This approach requires that both the N-terminal and C-terminal RNAs contain a single ribozyme at their 3' and 5' ends, respectively. The internal repeat RNA requires ribozymes at both the 5' and 3' ends to function as both acceptor and donor during trans-splicing. However, adding ribozymes to both ends of RNA, or RNA with both 3'-P and 5'-OH, can lead to cyclization by ligases such as RtcB (Desai et al. 2015), preventing involvement in the elongating linear chain. However, by utilizing inducible transactivation ribozymes, stepwise ligation of the 5' and 3' ends becomes possible through the addition and removal of both VS-Rz and RtcB ligases, potentially leading to the synthesis of controlled RNA domains (Figure 14C). This approach may be useful for generating highly repeatable RNA sequences, which can then be translated to create synthetic repeat proteins such as those constituting hydrogels, synthetic spider silk, or collagen, which may be difficult to recombinate to generate and encode DNA. These approaches may be useful for drug delivery, biomaterials, or industrial material production (Chambre et al. 2020).
[0341] Example 3: Generation of stable synthetic intron sequences using ribozymes Ribozyme-mediated trans-splicing between two independent RNAs can occur when one RNA contains a 3' ribozyme and the other contains a 5' ribozyme (Figure 15A). However, it has been shown that when transcribed in cis within the same RNA, the two ribozymes can mediate their own scarless removal (Figure 15B). This approach similarly generates two independent RNAs with 3'-P and 5'OH ends, which can undergo trans-splicing and translation in the cell (Figure 15B). This can also be achieved in vitro by adding a ligase such as RtcB.
[0342] Ribozyme-generating intron sequences, including compatible 5'-OH and 3'-P terminals, can be cis-spliced or circularized, which is a common readout for RtcB ligase activity in vitro. In contrast to lariat RNA, which is generated by spliceosomes during exon splicing and rapidly degraded, the RNA ring is considered highly stable because it no longer contains 5' or 3' ends and therefore cannot be degraded by RNA exonucleases. A cargo sequence, which may contain any number of functional or useful RNAs (e.g., microRNAs, CRISPR-guide RNAs, etc.) or gene expression sequences, can be inserted as "cargo" between two ribozymes (Figure 15C). This approach may be useful for ribozyme-mediated trans-splicing and simultaneous delivery and expression of useful RNA sequences during expression. If one of the internal ribozymes does not require a bilateral flanking sequence for activity, for example, in the case of a 5'HDV ribozyme, the RNA ring can exist in both circular and re-cleaved linear forms (Figure 15C). When VS-S is used instead of HDV, the system can be made induceable, in which case delivery or expression of VS-Rz is required. By using a ribozyme that requires a bilateral flanking sequence for cleavage, such as an HH ribozyme, the cleavage can be designed so that the RNA circularization of cargo RNA is unidirectional (Figure 15D). Example 4: Array Transsplicing protein-coding nucleic acid sequences Nt-GFP (SEQ ID NO: 1) AUGGUGAGCAAGGGCGAGGAGCUGUUCACCGGGGUGGUGCCCAUCCUGGUCGAGCUGGACGGCGACGUAAACGGCCACAAGUUCAGCGUGUCCGGCGAGGGCGAGGGCGAUGCCACCUACGGCAAGCUGACCCUGAAGUUCAUCUGCACCACCGGCAAGCUGCCCGUGCCCUGGCCCACCCUCGUGACCACCCUGACCUACGGCGUGCAGUGCUUCAGCCGCUACCCCGACCACAUGAAGCAGCACGACUUCUUCAAGUCCGCCAUGCCCGAAGGCUACGUCCAGGAGCGCACCAUCUUCUU Ct-GFP (SEQ ID NO: 2) CAAGGACGACGGCAACUACAAGACCCGCGCCGAGGUGAAGUUCGAGGGCGACACCCUGGUGAACCGCAUCGAGCUGAAGGGCAUCGACUUCAAGGAGGACGGCAACAUCCUGGGGCACAAGCUGGAGUACAACUACAACAGCCACAACGUCUAUAUCAUGGCCGACAAGCAGAAGAACGGCAUCAAGGUGAACUUCAAGAUCCGCCACAACAUCGAGGACGGCAGCGUGCAGCUCGCCGACCACUACCAGCAGAACACCCCCAUCGGCGACGGCCCCGUGCUGCUGCCCGACAACCACUACCUGAGCACCCAGUCCGCCCUGAGCAAAGACCCCAACGAGAAGCGCGAUCACAUGGUCCUGCUGGAGUUCGUGACCGCCGCCGGGAUCACUCUCGGCAUGGACGAGCUGUACAAGUAGUAA Nt-Luciferase (SEQ ID NO: 3) AUGGAAGACGCCAAAAACAUAAAGAAAGGCCCGGCGCCAUUCUAUCCGCUGGAAGAUGGAACCGCUGGAGAGCAACUGCAUAAGGCUAUGAAGAGAUACGCCCUGGUUCCUGGAACAAUUGCUUUUACAGAUGCACAUAUCGAGGUGGACAUCACUUACGCUGAGUACUUCGAAAUGUCCGUUCGGUUGGCAGAAGCUAUGAAACGAUAUGGGCUGAAUACAAAUCACAGAAUCGUCGUAUGCAGUGAAAACUCUCUUCAAUUCUUUAUGCCGGUGUUGGGCGCGUUAUUUAUCGGAGUUGCAGUUGCGCCCGCGAACGACAUUUAUAAUGAACGUGAAUUGCUCAACAGUAUGGGCAUUUCGCAGCCUACCGUGGUGUUCGUUUCCAAAAAGGGGUUGCAAAAAAUUUUGAACGUGCAAAAAAAGCUCCCAAUCAUCCAAAAAAUUAUUAUCAUGGAUUCUAAAACGGAUUACCAGGGAUUUCAGUCGAUGUACACGUUCGUCACAUCUCAUCUACCUCCCGGUUUUAAUGAAUACGAUUUUGUGCCAGAGUCCUUCGAUAGGGACAAGACAAUUGCACUGAUCAUGAACUCCUCUGGAUCUACUGGUCUGCCUAAAGGUGUCGCUCUGCCUCAUAGAACUGCCUGCGUGAGAUUCUCGCAUGCCAGAGAUCCUAUUUUUGGCAAUCAAAUCAUUCCGGAUACUGCGAUUUUAAGUGUUGUUCCAUUCCAUCACGGUUUUGGAAUGUUUACUACACUCGGAUAUUUGAUAUGUGGAUUUCGAGUCGUCUUAAUGUAUAGAUUUGAAGAAGAGCUGUUUCUGAGGAGCCUU Ct-luciferase (SEQ ID NO: 4) CAGGAUUACAAGAUUCAAAGUGCGCUGCUGGUGCCAACCCUAUUCUCCUUCUUCGCCAAAAGCACUCUGAUUGACAAAUACGAUUUAUCUAAUUUACACGAAAUUGCUUCUGGUGGCGCUCCCCUCUCUAAGGAAGUCGGGGAAGCGGUUGCCAAGAGGUUCCAUCUGCCAGGUAUCAGGCAAGGAUAUGGGCUCACUGAGACUACAUCAGCUAUUCUGAUUACACCCGAGGGGGAUGAUAAACCGGGCGCGGUCGGUAAAGUUGUUCCAUUUUUUGAAGCGAAGGUUGUGGAUCUGGAUACCGGGAAAACGCUGGGCGUUAAUCAAAGAGGCGAACUGUGUGUGAGAGGUCCUAUGAUUAUGUCCGGUUAUGUAAACAAUCCGGAAGCGACCAACGCCUUGAUUGACAAGGAUGGAUGGCUACAUUCUGGAGACAUAGCUUACUGGGACGAAGACGAACACUUCUUCAUCGUUGACCGCCUGAAGUCUCUGAUUAAGUACAAAGGCUAUCAGGUGGCUCCCGCUGAAUUGGAAUCCAUCUUGCUCCAACACCCCAACAUCUUCGACGCAGGUGUCGCAGGUCUUCCCGACGAUGACGCCGGUGAACUUCCCGCCGCCGUUGUUGUUUUGGAGCACGGAAAGACGAUGACGGAAAAAGAGAUCGUGGAUUACGUCGCCAGUCAAGUAACAACCGCGAAAAAGUUGCGCGGAGGAGUUGUGUUUGUGGACGAAGUACCGAAAGGUCUUACCGGAAAACUCGACGCAAGAAAAAUCAGAGAGAUCCUCAUAAAGGCCAAGAAGGGCGGAAAGAUCGCCGUGUAGUAA N1L (SEQ ID NO: 5) ATGGGTCAGGCCAATACGCCCTGGAGCAGTAAGGCAAACGCGGATGCCTTTATAAATTCATTCATCAGTGCAGCATCCAATACTGGTTCCTTCTCTCAAGACCAAATGGAGGACATGTCACTCATCGGCAATACTCTGATGGCTGCCATGGACAATATGGGAGGCCGCATAACACCATCTAAGTTGCAGGCGTTGGATATGGCCTTCGCATCATCAGTGGCCGAGATCGCGGCTAGTGAGGGCGGCGACTTGGGAGTCACTACCAACGCGATCGCGGATGCCCTCACTTCTGCTTTTTATCAAACGACCGGGGTTGTCAATTCACGATTCATATCTGAGATCAGGAGCCTCATAGGAATGTTCGCGCAGGCTTCCGCAAATGACGTTTATGCATCTGCTGGCTCTGGCAGCGGGGGTGGTGGGTATGGAGCCAGCTCAGCATCTGCGGCTTCTGCAAGTGCTGCTGCCCCGAGTGGCGTAGCTTATCAGGCTCCTGCTCAGGCTCAAATCAGTTTTACGTTGCGAGGGCAACAACCTGTTTCC AQ (Accession No. 6) GGTCCTTATGGACCCGGTGCTAGCGCTGCGGCAGCAGCCGCTGGCGGTTATGGCCCAGGTTCAGGGCAACAGGGGCCTGGGCAACAAGGACCTGGCCAACAAGGTCCTGGTCAGCAGGGTCCAGGGCAGCAG NR3 (Accession No. 7) GGCGCTGCTTCCGCTGCAGTATCAGTAGGTGGCTATGGACCTCAATCTAGTAGCGCCCCTGTTGCCTCTGCCGCCGCATCTCGACTTTCAAGTCCCGCCGCTAGTTCCAGGGTCAGTTCCGCGGTATCTAGCTTGGTAAGTAGCGGACCCACTAATCAAGCGGCACTTTCAAACACAATATCCTCAGTAGTCAGTCAAGTAAGCGCATCAAACCCTGGCTTGTCAGGGTGTGACGTTCTGGTTCAGGCACTTCTGGAAGTTGTCTCAGCGTTGGTAAGCATCCTGGGTAGCTCCTCCATAGGTCAAATTAATTATGGCGCGAGCGCCCAATACACACAAATGGTGGGTCAGAGTGTGGCGCAGGCACTCGCAGGCGACTACAAGGATCATGACGGAGACTATAAGGATCATGATATAGATTACAAGGACGATGATGACAAGGCCTAGTAA Nt-4xMTS (SEQ ID NO: 8) AUGAGUGUGUUGACGCCGUUGCUUCUGCGAGGGCUUACCGGGUCUGCUAGAAGACUUCCGGUCCCCAGGGCCAAGAUACAUAGCCUCGGAGACCCGAUGUCUGUGCUCACUCCUCUGCUUUUGCGAGGACUGACUGGGUCCGCCAGACGACUCCCGGUGCCGAGAGCUAAAAUCCAUAGCCUGGGAAAAUUGGCAACUAUGUCAGUCCUGACGCCGCUUCUUCUCCGGGGUCUUACAGGGUCUGCAAGAAGGCUGCCUGUACCUCGGGCGAAAAUUCAUAGCUUGGGCGACCCGAUGAGUGUAUUGACGCCCCUGUUGCUGAGAGGAUUGACUGGGUCAGCGCGCCGGCUCCCUGUCCCCCGAGCUAAGAUUCACUCCCUUGGUAAGCUGAGAAUCCUCCAAUCAACGGUUCCGAGAGCAAGAGAUCCGCCGGUCGCCACGAGGCCUCUCGAG Nt-DTA (SEQ ID NO: 17) AUGGACCCCGACGACGUGGUGGACAGCAGCAAGAGCUUCGUGAUGGAGAACUUCAGCAGCUACCACGGCACCAAGCCCGGCUACGUGGACAGCAUCCAGAAGGGCAUCCAGAAGCCCAAGAGCGGCACCCAGGGCAACUACGACGACGACUGGAAGGGCUUCUACAGCACCGACAACAAGUACGACGCUGCCGGCUACAGCGUGGACAACGAGAACCCCCUGAGCGGCAAGGCCGGCGGCGUGGUGAAGGUGACCUACCCCGGCCUGACCAAGGUGCUGGCCCUGAAGGUG Ct-DTA (SEQ ID NO: 18) GACAAUGCCGAGACCAUCAAGAAGGAGCUGGGCCUGAGCCUGACCGAGCCCCUGAUGGAGCAGGUGGGCACCGAGGAGUUCAUCAAGAGAUUCGGCGACGGCGCCAGCAGAGUGGUGCUGAGCCUGCCCUUCGCCGAGGGCAGCAGCAGCGUGGAGUACAUCAACAACUGGGAGCAGGCCAAGGCCCUGAGCGUGGAGCUGGAGAUCAACUUCGAGACCAGAGGCAAGAGAGGCCAGGACGCCAUGUACGAGUACAUGGCCCAGGCUUGCGCCGGCAACAGAGUGAGAAGAUAGUAA GFPcdn (without start ATG codon) (SEQ ID NO: 19) GUUAGCAAGGGCGAGGAGCUCUUCACCGGGGUCGUCCCCAUCCUCGUCGAGCUCGACGGCGACGUAAACGGCCACAAGUUCAGCGUCUCCGGCGAGGGCGAGGGCGAUGCCACCUACGGCAAGCUCACCCUGAAGUUCAUCUGCACCACCGGCAAGCUGCCCGUGCCCUGGCCCACCCUCGUGACCACCCUGACCUACGGCGUGCAGUGCUUCAGCCGCUACCCCGACCACAUGAAGCAGCACGACUUCUUCAAGUCCGCCAUGCCCGAAGGCUACGUCCAGGAGCGCACCAUCUUCUUCAAGGACGACGGCAACUACAAGACCCGCGCCGAGGUGAAGUUCGAGGGCGACACCCUGGUGAACCGCAUCGAGCUGAAGGGCAUCGACUUCAAGGAGGACGGCAACAUCCUGGGGCACAAGCUGGAGUACAACUACAACAGCCACAACGUCUAUAUCAUGGCCGACAAGCAGAAGAACGGCAUCAAGGUGAACUUCAAGAUCCGCCACAACAUCGAGGACGGCAGCGUGCAGCUCGCCGACCACUACCAGCAGAACACCCCCAUCGGCGACGGCCCCGUGCUGCUGCCCGACAACCACUACCUGAGCACCCAGUCCGCCCUGAGCAAAGACCCCAACGAGAAGCGCGAUCACAUGGUCCUGCUGGAGUUCGUGACCGCCGCCGGGAUCACUCUCGGCAUGGACGAGCUGUACAAGUAG F2-Myr (SEQ ID NO: 20) AUGGGUUGUUGUUUCAGCAAGACAGCGGCGAAAGGUGAAGCAGCAGCAGAAAGACCAGGCGAGGCUGCGGUAGCAUCAAGUCCCUCCAAGGCUAAUGGGCAGGAAAACGGACACGUCAAAGUUGGAAGCGU F2-RFP (SEQ ID NO: 21) AGCCAUCAUCAAGGAGUUCAUGCGCUUCAAGGUGCACAUGGAGGGCUCCGUGAACGGCCACGAGUUCGAGAUCGAGGGCGAGGGCGAGGGCCGCCCCUACGAGGGCACCCAGACCGCCAAGCUGAAGGUGACCAAGGGUGGCCCCCUGCCCUUCGCCUGGGACAUCCUGUCCCCUCAGUUCAUGUACGGCUCCAAGGCCUACGUGAAGCACCCCGCCGACAUCCCCGACUACUUGAAGCUGUCCUUCCCCGAGGGCUUCAAGUGGGAGCGCGUGAUGAACUUCGAGGACGGCGGCGUGGUGACCGUGACCCAGGACUCCUCCCUGCAGGACGGCGAGUUCAUCUACAAGGUGAAGCUGCGCGGCACCAACUUCCCCUCCGACGGCCCCGUAAUGCAGAAGAAGACCAUGGGCUGGGAGGCCUCCUCCGAGCGGAUGUACCCCGAGGACGGCGCCCUGAAGGGCGAGAUCAAGCAGAGGCUGAAGCUGAAGGACGGCGGCCACUACGACGCUGAGGUCAAGACCACCUACAAGGCCAAGAAGCCCGUGCAGCUGCCCGGCGCCUACAACGUCAACAUCAAGUUGGACAUCACCUCCCACAACGAGGACUACACCAUCGUGGAACAGUACGAACGCGCCGAGGGCCGCCACUCCACCGGCGGCAUGGACGAGCUGUACAAGUAGUAA Nt-uDys (SEQ ID NO: 22) Ct-uDys-GFP (SEQ ID NO: 23) Nt-miniDys(ΔH2-R15)(Sequence ID 129) Ct-miniDys(ΔH2-R15)(Sequence ID 130) Ribozyme nucleic acid sequences for Scarless 3' RNA cleavage HDV68 (Sequence ID 9) GGCCGGCAUGGUCCCAGCCUCCUCGCUGGCGCCGGCUGGGCAACAUGCUUCGGCAUGGCGAAUGGGAC HDV68 catalytic variant (SEQ ID NO: 24) 5'-GGCCGGCAUGGUCCCAGCCUCCUCGCUGGCGCCGGCUGGGCAACAUGCUUCGGCAUGGUGAAUGGGAC-3' HDV67 (Sequence ID 10) GGGUCGGCAUGGCAUCUCCACCUCCUCGCGGUCCGACCUGGGCUACUUCGGUAGGCUAAGGGAGAAG HDV56 (Sequence ID 11) GAGGGAUAGUACAGAGCCUCCCCGUGGCUCCCUUGGAUAACCAACUGAUACUGUAC Genome HDV (genHDV) (Sequence ID 12) GGCCGGCAUGGUCCCAGCCUCCUCGCUGGCGCCGGCUGGGCAACAUUCCGAGGGGACCGUCCCCUCGGUAAUGGCGAAUGGGACCCA Antigenome HDV (antiHDV) (SEQ ID NO: 13) GGGUCGGCAUGGCAUCUCCACCUCCUCGCGGUCCGACCUGGGCAUCCGAAGGAGGACGCACGUCCACUCGGAUGGCUAAGGGAGAGCCACU VS Ribozyme (SEQ ID NO: 14) GCGGUAGUAAGCAGGGAACUCACCUCCAAUUUCAGUACUGAAAUUGUCGUAGCAGUUGACUACUGUUAUGUGAUUGGUAGAGGCUAAGUGACGGUAUUGGCGUAAGUCAGUAUUGCAGCACAGCACAAGCCCGCUUGCGAGAAU VS-S (Sequence ID 15) GAAGGGCGUCGUCGCCCCGAG VS-Rz (Sequence ID 16) GCGGUAGUAAGCAGGGAACUCACCUCCAAUUUCAGUACUGAAAUUGUCGUAGCAGUUGACUACUGUUAUGUGAUUGGUAGAGGCUAAGUGACGGUAUUGGCGUAAGUCAGUAUUGCAGCACAGCACAAGCCCGCUUGCGAGAAU Hammerhead with a stem 3 overhang specific to Nt-Luc (SEQ ID NO: 25) 5'-GAGCCUUACCGGAUGUGUUUUCCGGUCUGAUGAGUC...
Claims
1. A method for treating a disease or disorder caused by a mutation in a target protein in a subject, Administering a first nucleic acid molecule to the subject, which includes a coding region that encodes a first portion of the target protein directly linked at the 3' end to the nucleotide sequence of a ribozyme (3' ribozyme), Administering to the subject a second nucleic acid molecule containing a coding region that encodes a second portion of the target protein, directly linked at the 5' end to the nucleotide sequence of a ribozyme (5' ribozyme), and Includes, A method comprising the following steps: self-cleavage of the 3' ribozyme generates a 3' phosphate group at the 3' end of the first portion of the target protein; self-cleavage of the 5' ribozyme generates a 5' hydroxyl group at the 5' end of the second portion of the target protein; thereby enabling scarless ligation of the target protein.
2. At least one of the 3'-ribozyme and the 5'-ribozyme is sequence number 131, 132, 133, 134, 135, 136, 137, 138, 139, 140, 141, 142, 143, 144, 145, 166, 167, 168, 169, 170, 171, sequence number The method according to claim 1, selected from the group consisting of No. 172, SEQ ID NO: 173, SEQ ID NO: 174, SEQ ID NO: 175, SEQ ID NO: 176, SEQ ID NO: 192, SEQ ID NO: 193, SEQ ID NO: 194, SEQ ID NO: 195, SEQ ID NO: 196, SEQ ID NO: 197, SEQ ID NO: 198, SEQ ID NO: 199, SEQ ID NO: 200, SEQ ID NO: 202, SEQ ID NO: 203, SEQ ID NO: 204, SEQ ID NO: 205, SEQ ID NO: 206, SEQ ID NO: 207, SEQ ID NO: 217, SEQ ID NO: 218, or SEQ ID NO:
219.
3. The aforementioned diseases or disorders include Duchenne muscular dystrophy, Becker muscular dystrophy (BMD), autosomal recessive polycystic kidney disease, hemophilia A, Stargard macular degeneration, limb-girdle muscular dystrophy, autosomal recessive severe congenital deafness, autosomal recessive non-syndromic hearing loss (ARNSHL), sensorineural hearing loss, cystic fibrosis, Wilson's disease, Miyoshi type myopathy, autosomal recessive hearing loss type 9 (DFNB9), Usher syndrome type I, and GJB2-associated autosomal recessive non-syndromic hearing loss (GJB2-AR). NSHL), Autosomal recessive cerebellar parenchymal disorder type 3, Non-syndromic hearing loss, Autosomal recessive hearing loss type 16 (DFNB16), Meniere's disease, Autosomal dominant non-syndromic sensorineural hearing loss type 12 (DFNA12), Autosomal recessive spinocerebellar ataxia type 21 (SCAR21), Usher syndrome type 1F (USH1F), Autosomal recessive hearing loss type 23 (DFNB23), Autosomal recessive hearing loss type 30 (DFNB30), Oto-spine-megaly epiphysis dysplasia (OSMED), Autosomal recessive hearing loss type 77 (DFNB77), Autosomal recessive hearing loss type 84A (DFNB84A), Autosomal recessive hearing loss type 84B (DFNB84B), Peripheral neuropathy, Myopathy, Autosomal dominant The method according to claim 1, wherein the patient is one or more selected from the group consisting of non-syndromic hearing loss type 4A (DFNA4), congenital thrombocytopenia, sensorineural hearing loss, autosomal dominant non-syndromic hearing loss type 56 (DFNA56), epileptic encephalopathy, Timothy syndrome, long QT syndrome, X-linked retinal disease, aldosteronism, autosomal recessive hearing loss type 42 (DFNB42), primary aldosteronism (Conn syndrome), seizures, neurological abnormalities, sinoatrial node dysfunction, neurodevelopmental disorders, hypokalemic periodic paralysis, epilepsy, developmental epileptic encephalopathy, brodymyopathy, Darier's disease, heart disease, von Willebrand disease, and Zellweger syndrome.
4. The method according to claim 1, wherein the target protein is selected from the group consisting of DMD, PKHD1, F8, ABCA4, DYSF, OTOF, CFTR, ATP7B, MYOF, MYO7A, MYO15A, CDH23, STRC, OTOG, TECTA, PCDH15, TRIOBP, MYO3A, COL11A2, LOXHD1, PTPRQ, OTOGL, MYH14, MYH9, TNC, CACNA1A, CACNA1C, CACNA1F, CACNA1H, CACNA1G, CACNA1D, CACNA1B, CACNA1S, CACNA1I, CACNA1E, ATP2A1, ATP2A2, VWF, PEX1, and CMYA5.
5. The method according to claim 1, wherein the first nucleic acid molecule comprises a nucleic acid sequence encoding the N-terminal fragment of a dystrophin, mini-dystrophin, or micro-dystrophin protein, and the second nucleic acid molecule comprises a nucleic acid sequence encoding the C-terminal fragment of a dystrophin, mini-dystrophin, or micro-dystrophin protein, and administration of the first and second nucleic acid molecules results in the production of a full-length dystrophin, mini-dystrophin, or micro-dystrophin protein.
6. The first and second nucleic acid molecules described above, a) A first nucleic acid containing the nucleic acid sequence of SEQ ID NO: 150, and a second nucleic acid containing the nucleic acid sequence of SEQ ID NO: 151, and b) A first nucleic acid containing the nucleic acid sequence of SEQ ID NO: 152 and a second nucleic acid containing the nucleic acid sequence of SEQ ID NO: 153 The method according to claim 5, selected from the group consisting of the following.
7. The method according to claim 1, wherein the first nucleic acid molecule comprises a nucleic acid sequence encoding the N-terminal fragment of dysferlin, and the second nucleic acid molecule comprises a nucleic acid sequence encoding the C-terminal fragment of dysferlin, and a full-length dysferlin protein is produced by administering the first and second nucleic acid molecules.
8. The first and second nucleic acid molecules described above, a) A first nucleic acid containing the nucleic acid sequence of sequence number 157 and a second nucleic acid containing the nucleic acid sequence of sequence number 158, b) A first nucleic acid containing the nucleic acid sequence of sequence number 177, and a second nucleic acid containing the nucleic acid sequence of sequence number 178, c) A first nucleic acid containing the nucleic acid sequence of sequence number 179, and a second nucleic acid containing the nucleic acid sequence of sequence number 178, d) A first nucleic acid containing the nucleic acid sequence of SEQ ID NO: 181, a second nucleic acid containing the nucleic acid sequence of SEQ ID NO: 182, and e) A first nucleic acid containing the nucleic acid sequence of SEQ ID NO: 183, and a second nucleic acid containing the nucleic acid sequence of SEQ ID NO: 182 The method according to claim 7, selected from the group consisting of the following.
9. The method according to claim 1, wherein the first nucleic acid molecule comprises a nucleic acid sequence encoding the N-terminal fragment of STRC, and the second nucleic acid molecule comprises a nucleic acid sequence encoding the C-terminal fragment of STRC, and a full-length STRC protein is produced by administering the first and second nucleic acid molecules.
10. The method according to claim 9, wherein the first nucleic acid molecule comprises the nucleic acid sequence of SEQ ID NO: 160, and the second nucleic acid molecule comprises the nucleic acid sequence of SEQ ID NO:
161.
11. A system for generating RNA molecules that encode a target protein, A first RNA molecule comprising a coding region that encodes a first portion of the target protein, which is directly linked at the 3' end to the nucleotide sequence of a ribozyme (3' ribozyme), A second RNA molecule comprising a coding region that encodes a second portion of the target protein, which is directly linked at the 5' end to the nucleotide sequence of a ribozyme (5' ribozyme), Includes, A system in which the self-cleavage of the 3' ribozyme generates a 3' phosphate group at the 3' end of the first portion of the target protein, and the self-cleavage of the 5' ribozyme generates a 5' hydroxyl group at the 5' end of the second portion of the target protein, enabling scarless ligation of the first and second RNA molecules, thereby generating an RNA molecule encoding the target protein.
12. At least one of the 3'-ribozyme and the 5'-ribozyme is sequence number 131, 132, 133, 134, 135, 136, 137, 138, 139, 140, 141, 142, 143, 144, 145, 166, 167, 168, 169, 170, 171, sequence number The system according to claim 11, selected from number 172, sequence number 173, sequence number 174, sequence number 175, sequence number 176, sequence number 192, sequence number 193, sequence number 194, sequence number 195, sequence number 196, sequence number 197, sequence number 198, sequence number 199, sequence number 200, sequence number 202, sequence number 203, sequence number 204, sequence number 205, sequence number 206, sequence number 207, sequence number 217, sequence number 218, or sequence number 219.
13. The system according to claim 11, wherein the total length of the target protein is longer than 1,000 amino acid residues.
14. The system according to claim 12, wherein the target protein is selected from the group consisting of DMD, PKHD1, F8, ABCA4, DYSF, OTOF, CFTR, ATP7B, MYOF, MYO7A, MYO15A, CDH23, STRC, OTOG, TECTA, PCDH15, TRIOBP, MYO3A, COL11A2, LOXHD1, PTPRQ, OTOGL, MYH14, MYH9, TNC, CACNA1A, CACNA1C, CACNA1F, CACNA1H, CACNA1G, CACNA1D, CACNA1B, CACNA1S, CACNA1I, CACNA1E, ATP2A1, ATP2A2, VWF, PEX1, and CMYA5.
15. A method for generating an RNA molecule encoding a target protein, Administering a first RNA molecule to a cell or tissue, which includes a coding region that encodes a first portion of the target protein, directly ligated at the 3' end to the nucleotide sequence of a ribozyme (3' ribozyme); Administering a second RNA molecule to a cell or tissue, which includes a coding region that encodes a second portion of the target protein, directly ligated at the 5' end to the nucleotide sequence of a ribozyme (5' ribozyme); Includes, A method comprising the following steps: self-cleavage of the 3' ribozyme generates a 3' phosphate group at the 3' end of the first portion of the target protein; self-cleavage of the 5' ribozyme generates a 5' hydroxyl group at the 5' end of the second portion of the target protein; thereby enabling scarless ligation of the first and second RNA molecules, and generating an RNA molecule encoding the target protein.
16. At least one of the 3'-ribozyme and the 5'-ribozyme is sequence number 131, 132, 133, 134, 135, 136, 137, 138, 139, 140, 141, 142, 143, 144, 145, 166, 167, 168, 169, 170, 171, sequence number The method according to claim 15, selected from the group consisting of 172, SEQ ID NO: 173, SEQ ID NO: 174, SEQ ID NO: 175, SEQ ID NO: 176, SEQ ID NO: 192, SEQ ID NO: 193, SEQ ID NO: 194, SEQ ID NO: 195, SEQ ID NO: 196, SEQ ID NO: 197, SEQ ID NO: 198, SEQ ID NO: 199, SEQ ID NO: 200, SEQ ID NO: 202, SEQ ID NO: 203, SEQ ID NO: 204, SEQ ID NO: 205, SEQ ID NO: 206, SEQ ID NO: 207, SEQ ID NO: 217, SEQ ID NO: 218, or SEQ ID NO:
219.
17. The method according to claim 14, wherein the total length of the target protein is longer than 1,000 amino acid residues.
18. The method according to claim 17, wherein the target protein is selected from the group consisting of DMD, PKHD1, F8, ABCA4, DYSF, OTOF, CFTR, ATP7B, MYOF, MYO7A, MYO15A, CDH23, STRC, OTOG, TECTA, PCDH15, TRIOBP, MYO3A, COL11A2, LOXHD1, PTPRQ, OTOGL, MYH14, MYH9, TNC, CACNA1A, CACNA1C, CACNA1F, CACNA1H, CACNA1G, CACNA1D, CACNA1B, CACNA1S, CACNA1I, CACNA1E, ATP2A1, ATP2A2, VWF, PEX1, and CMYA5.
19. An in vitro method for generating an RNA molecule encoding a target protein, To provide a first RNA molecule comprising a coding region that encodes a first portion of the target protein, which is directly linked at the 3' end to the nucleotide sequence of a ribozyme (3' ribozyme), To provide a second RNA molecule comprising a coding region that encodes a second portion of the target protein, which is directly ligated at the 5' end to the nucleotide sequence of a ribozyme (5' ribozyme), Here, the self-cleavage of the 3' ribozyme generates a 3' phosphate group at the 3' end of the first portion of the target protein, and the self-cleavage of the 5' ribozyme generates a 5' hydroxyl group at the 5' end of the second portion of the target protein, thereby enabling scarless ligation of the first and second RNA molecules, and generating the RNA molecule encoding the target protein. To provide a ligase that induces the assembly of the RNA molecule from the coding region of the first RNA molecule and the coding region of the second RNA molecule. Methods that include...
20. At least one of the 3'-ribozyme and the 5'-ribozyme is sequence number 131, 132, 133, 134, 135, 136, 137, 138, 139, 140, 141, 142, 143, 144, 145, 166, 167, 168, 169, 170, 171, sequence number The method according to claim 19, selected from the group consisting of 172, SEQ ID NO: 173, SEQ ID NO: 174, SEQ ID NO: 175, SEQ ID NO: 176, SEQ ID NO: 192, SEQ ID NO: 193, SEQ ID NO: 194, SEQ ID NO: 195, SEQ ID NO: 196, SEQ ID NO: 197, SEQ ID NO: 198, SEQ ID NO: 199, SEQ ID NO: 200, SEQ ID NO: 202, SEQ ID NO: 203, SEQ ID NO: 204, SEQ ID NO: 205, SEQ ID NO: 206, SEQ ID NO: 207, SEQ ID NO: 217, SEQ ID NO: 218, or SEQ ID NO:
219.
21. A system for generating circular RNA molecules encoding a target protein, It comprises a nucleic acid molecule encoding a 5' ribozyme, an IRES sequence, and a 3' ribozyme, directly linked to the first portion of the target protein. A system in which the self-cleavage of the 3' ribozyme generates a 3' phosphate group at the 3' end of the second portion of the target protein, and the self-cleavage of the 5' ribozyme generates a 5' hydroxyl group at the 5' end of the first portion of the target protein, thereby enabling scarless ligation of the 5' hydroxyl group and 3' phosphate group of the RNA molecule, and generating a circular RNA molecule encoding the target protein.
22. The system according to claim 21, wherein the RNA molecule includes an in vitro transcription RNA molecule.
23. The system according to claim 21, wherein the target protein is one or more selected from the group consisting of therapeutic proteins, reporter proteins, and Cas9 proteins.
24. At least one of the 3'-ribozyme and the 5'-ribozyme is sequence number 131, 132, 133, 134, 135, 136, 137, 138, 139, 140, 141, 142, 143, 144, 145, 166, 167, 168, 169, 170, 171, sequence number The system according to claim 23, selected from number 172, sequence number 173, sequence number 174, sequence number 175, sequence number 176, sequence number 192, sequence number 193, sequence number 194, sequence number 195, sequence number 196, sequence number 197, sequence number 198, sequence number 199, sequence number 200, sequence number 202, sequence number 203, sequence number 204, sequence number 205, sequence number 206, sequence number 207, sequence number 217, sequence number 218, or sequence number 219.
25. A method for generating a circular RNA molecule in vivo, comprising administering a linear in vitro transcription RNA molecule or a DNA molecule encoding a linear RNA molecule to a cell or tissue, wherein the linear RNA molecule is It comprises a 5' ribozyme, an IRES sequence, directly linked to the first portion of the target protein, and a second portion of the target protein directly linked to the 3' ribozyme. A method comprising the following steps: self-cleavage of the 3' ribozyme generates a 3' phosphate group at the 3' end of the second portion of the target protein; self-cleavage of the 5' ribozyme generates a 5' hydroxyl group at the 5' end of the first portion of the target protein; thereby enabling seamless ligation of the 5' hydroxyl group and 3' phosphate group of the RNA molecule, and generating a circular RNA molecule encoding the target protein.
26. The method according to claim 25, wherein the target protein is one or more selected from the group consisting of therapeutic proteins, reporter proteins, and Cas9 proteins.
27. At least one of the 3'-ribozyme and the 5'-ribozyme is sequence number 131, 132, 133, 134, 135, 136, 137, 138, 139, 140, 141, 142, 143, 144, 145, 166, 167, 168, 169, 170, 171, sequence number The method according to claim 25, selected from the group consisting of 172, SEQ ID NO: 173, SEQ ID NO: 174, SEQ ID NO: 175, SEQ ID NO: 176, SEQ ID NO: 192, SEQ ID NO: 193, SEQ ID NO: 194, SEQ ID NO: 195, SEQ ID NO: 196, SEQ ID NO: 197, SEQ ID NO: 198, SEQ ID NO: 199, SEQ ID NO: 200, SEQ ID NO: 202, SEQ ID NO: 203, SEQ ID NO: 204, SEQ ID NO: 205, SEQ ID NO: 206, SEQ ID NO: 207, SEQ ID NO: 217, SEQ ID NO: 218, or SEQ ID NO: 219.