Ribozyme mediated RNA assembly and expression

By using ribozyme-mediated RNA molecule self-splicing and ligation technology, scarless RNA molecules are generated, solving the problem of limited full-length protein expression and achieving efficient expression of proteins with more than 1,000 amino acid residues, which is suitable for the treatment of a variety of diseases.

CN120882872APending Publication Date: 2025-10-31UNIVERSITY OF ROCHESTER
View PDF 14 Cites 0 Cited by

Patent Information

Application Number
CN202480019417.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-10-10
Filing Date
2024-03-18
Publication Date
2025-10-31

AI Technical Summary

Technical Problem

In existing technologies, the expression of full-length proteins is limited by the size of plasmids and vectors. In particular, in therapeutic settings, some nucleic acids encoding full-length proteins exceed the packaging size of AAVs, limiting their applicability in gene therapy. At the same time, some biological and industrially relevant proteins contain a large number of repetitive sequences, making expression difficult.

Method used

Nucleic acid molecules encoding a first RNA molecule and a second RNA molecule, each containing different parts of the target protein and a ribozyme, are used. Specific ends are generated through self-cleavage of the 3' and 5' ribozymes, and then ligated to form a complete RNA molecule. Scarless ligation is achieved using ligase to generate the RNA molecule encoding the target protein.

Benefits of technology

It achieves efficient expression of full-length proteins with more than 1,000 amino acid residues, solving the size limitation problem, and generates scarless complete RNA molecules through ribozyme self-cleavage and ligation technology, which are suitable for the treatment of a variety of diseases.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120882872A_ABST
    Figure CN120882872A_ABST
Patent Text Reader

Abstract

The present invention provides compositions, systems and methods for expressing a protein of interest or a fusion protein using ribozyme mediated cis-cleavage and trans-cleavage or cis-ligation of RNA molecules.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross-references to related applications

[0002] This application claims priority to U.S. Provisional Application No. 63 / 490,898, filed March 17, 2023; U.S. Provisional Application No. 63 / 491,372, filed March 21, 2023; U.S. Provisional Application No. 63 / 504,923, filed May 30, 2023; and U.S. Provisional Application No. 63 / 589,217, filed October 10, 2023, the entire contents of which are incorporated herein by reference. Background Technology

[0003] In some cases, the expression of full-length proteins is limited by the size constraints of plasmids and vectors. For example, in therapeutic settings, some nucleic acids encoding full-length proteins exceed the packaging size of AAVs, thus limiting their applicability in gene therapy environments. Furthermore, some biologically and industrially relevant proteins contain numerous repetitive sequences, which can make expression difficult.

[0004] Therefore, there is a need in the art for improved compositions and methods to achieve efficient protein expression. This invention addresses this unmet need. Summary of the Invention

[0005] In one embodiment, the present invention includes a system for generating an RNA molecule encoding a target protein, comprising: a nucleic acid molecule encoding a first RNA molecule, the first RNA molecule comprising a coding region encoding a first portion of the target protein and a 3' ribozyme; and a nucleic acid molecule encoding a second RNA molecule, the second RNA molecule comprising a coding region encoding a second portion of the target protein and a 5' ribozyme.

[0006] In one embodiment, the 3' ribozyme catalyzes its own detachment from the first RNA molecule, thereby generating a 3'P or 2'3' cP terminus. In one embodiment, the 5' ribozyme catalyzes its own detachment from the second RNA molecule, thereby generating a 5'OH terminus. In one embodiment, the 3'P or 2'3' cP terminus is linked to the 5'OH terminus to form an RNA molecule containing the coding regions of the first and second RNA molecules. In one embodiment, the 3' ribozyme is a member of the HDV ribozyme family. In one embodiment, the 5' ribozyme is a member of the HH ribozyme family.

[0007] In one embodiment, the system further includes one or more additional nucleic acid molecules encoding one or more additional RNA molecules, each additional RNA molecule containing a coding region encoding a target protein domain; a 5' ribozyme and a 3' ribozyme.

[0008] In one embodiment, the system further comprises one or more additional nucleic acid molecules encoding one or more additional RNA molecules, each additional RNA molecule containing a coding region encoding a target protein domain; and 5' and 3' ribozyme recognition sequences. In one embodiment, the system further comprises a ribozyme that interacts with the 3' ribozyme recognition sequence, inducing the removal of the 3' recognition sequence. In one embodiment, the 3' ribozyme recognition sequence comprises VS-S, and wherein the ribozyme is VS-Rz.

[0009] In one embodiment, the present invention relates to a method for generating an RNA molecule encoding a target protein, the method comprising: administering a nucleic acid molecule encoding a first RNA molecule to a cell or tissue, the first RNA molecule comprising a coding region encoding a first portion of the target protein and a 3' ribozyme; and administering a nucleic acid molecule encoding a second RNA molecule to a cell or tissue, the second RNA molecule comprising a coding region encoding a second portion of the target protein and a 5' ribozyme.

[0010] In one embodiment, the 3' ribozyme catalyzes its own detachment from the first RNA molecule, thereby generating a 3'P or 2'3' cP terminus. In one embodiment, the 5' ribozyme catalyzes its own detachment from the second RNA molecule, thereby generating a 5'OH terminus. In one embodiment, the 3'P or 2'3' cP terminus is joined to the 5'OH terminus to form an RNA molecule containing the coding regions of the first and second RNA molecules. In one embodiment, the 3' ribozyme is a member of the HDV ribozyme family. In one embodiment, the 5' ribozyme is a member of the HH ribozyme family.

[0011] In one embodiment, the method further includes administering to cells or tissue one or more additional nucleic acid molecules encoding one or more additional RNA molecules, each additional RNA molecule containing a coding region encoding a target protein domain; a 5' ribozyme; and a 3' ribozyme.

[0012] In one embodiment, the method further includes administering to cells or tissue one or more additional nucleic acid molecules encoding one or more additional RNA molecules, each additional RNA molecule containing a coding region encoding a target protein domain; a 5' ribozyme; and a 3' ribozyme recognition sequence.

[0013] In one embodiment, the method further includes administering a ribozyme that interacts with a 3' ribozyme recognition sequence to cells or tissues, the ribozyme inducing the removal of the 3' recognition sequence. In one embodiment, the 3' ribozyme recognition sequence comprises VS-S, and wherein the ribozyme is VS-Rz. In one embodiment, the method further includes administering a ligase to cells or tissues to induce the assembly of RNA molecules. In one embodiment, the ligase is an RNA 2',3'-cyclic phosphate and 5'-OH (RtcB) ligase.

[0014] In one embodiment, the present invention includes a method for generating RNA molecules encoding a target protein in vitro, comprising: providing a first RNA molecule, the first RNA molecule comprising a coding region of a first portion encoding the target protein and a 3' ribozyme; providing a second RNA molecule, the second RNA molecule comprising a coding region of a second portion encoding the target protein and a 5' ribozyme; and providing a ligase to induce the assembly of RNA molecules from the coding regions of the first RNA molecule and the coding regions of the second RNA molecule.

[0015] In one embodiment, the present invention includes a method for generating an RNA molecule encoding a target protein with repeating domains in vitro, the method comprising the steps of: a) providing a first RNA molecule comprising a coding region encoding a first portion of the target protein and a 3' ribozyme; b) providing one or more additional RNA molecules comprising a coding region encoding a domain of the target protein, a 5' ribozyme, and a 3' ribozyme recognition sequence; c) providing a ligase to ligate the coding region of the first RNA molecule to the coding region of one or more additional RNA molecules; d) providing a ribozyme that recognizes and catalyzes the removal of the 3' ribozyme recognition sequence; e) repeating steps b)-d) once or more to generate an RNA molecule encoding multiple repeating domains; f) providing a last RNA molecule comprising a coding region encoding a final portion of the target protein and a 5' ribozyme; and g) providing a ligase to ligate the coding regions of one or more additional RNA molecules to the coding region of the last RNA molecule, thereby generating a complete RNA molecule encoding a protein with repeating domains.

[0016] In one embodiment, the present invention includes a method for treating a disease or disorder in a subject caused by a mutation in a large target protein, comprising: administering to the subject a first nucleic acid molecule, the first nucleic acid molecule comprising a coding region encoding a first portion of the target protein and a 3' ribozyme; and administering to the subject a second nucleic acid, the second nucleic acid comprising a coding region encoding a second portion of the target protein and a 5' ribozyme.

[0017] In one embodiment, the disease or disorder is Duchenne Muscular Dystrophy, Becker Muscular Dystrophy (BMD), autosomal recessive polycystic kidney disease, hemophilia A, Stargardt macular degeneration, limb-girdle muscular dystrophies, autosomal recessive profound prelingual deafness, autosomal recessive nonsyndromic hearing loss (ARNSHL), sensorineural deafness, cystic fibrosis, Wilson's disease, Miyoshi myopathy, or autosomal recessive deafness-9. Deafness-9 (DFNB9), Usher Syndrome Type I, GJB2-related autosomal recessive non-syndromic hearing loss (GJB2-AR NSHL), Autosomal recessive cerebelloparenchymal disorder type 3, non-syndromic hearing loss, autosomal recessive deafness-16 (DFNB16), Meniere's disease (MD), and autosomal dominant non-syndromic sensorineural deafness-12.DFNA12), autosomal recessive spinocerebellar ataxia-21 (SCAR21), Usher syndrome type 1F (USH1F), autosomal recessive deafness-23 (DFNB23), autosomal recessive deafness-30 (DFNB30), otospondylomegaepiphyseal dysplasia (OSMED), autosomal recessive deafness-77 (DFNB77), autosomal recessive deafness-84A (DFNB84A), autosomal recessive deafness-84B (DFNB84A). Deafness-84B (DFNB84B), peripheral neuropathy, myopathy, autosomal dominant non-syndromic deafness-4A (DFNA4), congenital thrombocytopenia, sensorineural hearing loss, autosomal dominant non-syndromic hearing loss-56 (DFNA56), epileptic encephalopathy, Timothy syndrome, long QT syndrome, X-linked retinal disorder, hyperaldosteronism, autosomal recessive deafness-42 (DFNA42).DFNB42), primary aldosteronism (Conn's syndrome), seizures, neurologic abnormalities, sinoatrial node dysfunction, neurodevelopmental disorders, hypokalemic periodic paralysis, epilepsy, developmental and epileptic encephalopathies, Brody's myopathy, Darier's disease, heart disease, von Willebrand disease, or Zellweger syndrome.

[0018] In one embodiment, the present invention includes a system for generating RNA molecules and circular RNA molecules encoding a target protein, the system comprising nucleic acids encoding: a first portion of the target protein; a synthetic intron comprising a 5' ribozyme, a cargo sequence, and a 3' ribozyme; and a second portion of the target protein.

[0019] In one embodiment, the target protein is selected from one or more of the following: therapeutic proteins, reporter proteins, and Cas9 proteins.

[0020] In one embodiment, the cargo sequence is selected from one or more of the following: sequences encoding a target therapeutic protein, CRISPR guide RNA sequences, small RNA sequences, and trans-cleaving ribozyme sequences. In one embodiment, the small RNA sequence includes one or more of the following: microRNA (miRNA), Piwi-interacting RNA (piRNA), small interfering RNA (siRNA), small nucleolar RNA (snoRNA), small tRNA-derived RNA (tsRNA), small rDNA-derived RNA (srRNA), and small nuclear RNA (snRNA).

[0021] In one embodiment, the 3' ribozyme that synthesizes the intron is a member of the HH ribozyme family. In one embodiment, the 5' ribozyme that synthesizes the intron is selected from one or more of the following: members of the HDV ribozyme family, members of the HDV ribozyme family, and VS-S ribozyme recognition sequences. In one embodiment, the system further comprises one or more of the following: an RtcB ligase and a nucleic acid encoding the RtcB ligase.

[0022] In one embodiment, the present invention includes a system for generating an RNA molecule encoding a target protein, the system comprising:

[0023] a) A first RNA molecule containing a coding region encoding a first part of a target protein, wherein the coding region is directly linked at its 3' end to a nucleotide sequence of a ribozyme (3' ribozyme);

[0024] b) A second RNA molecule containing a coding region encoding a second part of the target protein, wherein the coding region is directly linked at its 5' end to the nucleotide sequence of a ribozyme (5' ribozyme);

[0025] Specifically, the self-cleavage of the 3' ribozyme generates a 3' phosphate at the 3' end of the first part of the target protein, and the self-cleavage of the 5' ribozyme generates a 5' hydroxyl group at the 5' end of the second part of the target protein, thereby allowing the first and second RNA molecules to join without scarring to generate an RNA molecule encoding the target protein.

[0026] In one embodiment, at least one of the 3' ribozyme and the 5' ribozyme is SEQ ID NO:131, SEQ ID NO:132, SEQ ID NO:133, SEQ ID NO:134, SEQ ID NO:135, SEQ ID NO:136, SEQ ID NO:137, SEQ ID NO:138, SEQ ID NO:139, SEQ ID NO:140, SEQ ID NO:141, SEQ ID NO:142, SEQ ID NO:143, SEQ ID NO:144, SEQ ID NO:145, SEQ ID NO:166, SEQ ID NO:167, SEQ ID NO:168, SEQ ID NO:169, SEQ ID NO:170, SEQ ID NO:171, SEQ ID NO:172, SEQ ID NO:173, SEQ ID NO:174, SEQ ID NO:175, SEQ ID NO:176, SEQ ID NO:192, SEQ ID NO:193, SEQ ID NO:194, SEQ ID NO:195, SEQ ID NO:196, SEQ ID NO:192, SEQ ID NO:193, SEQ ID NO:194, SEQ ID NO:195, SEQ ID NO:196, SEQ ID NO:195, SEQ ID NO:196, SEQ ID NO:197, SEQ ID NO:198, SEQ ID NO:19 ... NO:194, SEQ ID NO:195, SEQ ID NO:196, SEQ ID NO:197, SEQ ID NO:198, SEQ ID NO:199, SEQ ID NO:200, SEQ ID NO:202, SEQ ID NO:203, SEQ ID NO:204, SEQ ID NO:205, SEQ ID NO:206, SEQ ID NO:207, SEQ ID NO:217, SEQ ID NO:218 or SEQ ID NO:219.

[0027] In one embodiment, the target protein is longer than 1000 amino acid residues.

[0028] In one embodiment, the target protein is DMD, PKHD1, F8, ABCA4, DYSF, OTOF, CFTR, ATP7B, MYOF, MYO7A, MYO15A, CDH23, STRC, OTOG, TECTA, PCDH15, TRIOBP, MYO3A, COL11A2, LOXHD1, PTPRQ, OTOGL, MYH14, MYH9, TNC, CACNA1A, CACNA1C, CACNA1F, CACNA1H, CACNA1G, CACNA1D, CACNA1B, CACNA1S, CACNA1I, CACNA1E, ATP2A1, ATP2A2, VWF, PEX1, or CMYA5.

[0029] In one embodiment, the present invention includes a method for treating a disease or disorder in a subject caused by a mutation in a target protein, comprising administering to the subject a first nucleic acid molecule comprising a coding region encoding a first portion of the target protein, the coding region being directly linked at its 3' end to a nucleotide sequence of a ribozyme (3' ribozyme); and administering to the subject a second nucleic acid molecule comprising a coding region encoding a second portion of the target protein, the coding region being directly linked at its 5' end to a nucleotide sequence of a ribozyme (5' ribozyme); wherein self-cleavage of the 3' ribozyme generates a 3' phosphate at the 3' end of the first portion of the target protein, and self-cleavage of the 5' ribozyme generates a 5' hydroxyl group at the 5' end of the second portion of the target protein, thereby allowing scarless attachment of the target protein.

[0030] In one embodiment, at least one of the 3' ribozyme and the 5' ribozyme is SEQ ID NO:131, SEQ ID NO:132, SEQ ID NO:133, SEQ ID NO:134, SEQ ID NO:135, SEQ ID NO:136, SEQ ID NO:137, SEQ ID NO:138, SEQ ID NO:139, SEQ ID NO:140, SEQ ID NO:141, SEQ ID NO:142, SEQ ID NO:143, SEQ ID NO:144, SEQ ID NO:145, SEQ ID NO:166, SEQ ID NO:167, SEQ ID NO:168, SEQ ID NO:169, SEQ ID NO:170, SEQ ID NO:171, SEQ ID NO:172, SEQ ID NO:173, SEQ ID NO:174, SEQ ID NO:175, SEQ ID NO:176, SEQ ID NO:192, SEQ ID NO:193, SEQ ID NO:194, SEQ ID NO:195, SEQ ID NO:196, SEQ ID NO:192, SEQ ID NO:193, SEQ ID NO:194, SEQ ID NO:195, SEQ ID NO:196, SEQ ID NO:195, SEQ ID NO:196, SEQ ID NO:197, SEQ ID NO:198, SEQ ID NO:19 ... NO:194, SEQ ID NO:195, SEQ ID NO:196, SEQ ID NO:197, SEQ ID NO:198, SEQ ID NO:199, SEQ ID NO:200, SEQ ID NO:202, SEQ ID NO:203, SEQ ID NO:204, SEQ ID NO:205, SEQ ID NO:206, SEQ ID NO:207, SEQ ID NO:217, SEQ ID NO:218 or SEQ ID NO:219.

[0031] In one embodiment, the disease or disorder is Duchenne muscular dystrophy, Becker's muscular dystrophy (BMD), autosomal recessive polycystic kidney disease, hemophilia A, Stargardt's macular degeneration, limb-girdle muscular dystrophy, autosomal recessive deep prelingual hearing loss, autosomal recessive nonsyndromic hearing loss (ARNSHL), sensorineural hearing loss, cystic fibrosis, Wilson's disease, Trihyo's myopathy, autosomal recessive deafness-9 (DFNB9), Usher syndrome type I, GJB2-related autosomal recessive nonsyndromic hearing loss (GJB2-AR). NSHL), type 3 autosomal recessive cerebellar cortical disorder, nonsyndromic hearing loss, autosomal recessive deafness-16 (DFNB16), Meniere's disease (MD), autosomal dominant nonsyndromic sensorineural hearing loss-12 (DFNA12), autosomal recessive spinocerebellar ataxia-21 (SCAR21), Usher syndrome type 1F (USH1F), autosomal recessive deafness-23 (DFNB23), autosomal recessive deafness-30 (DFNB30), macrophage dysplasia of the vertebral body (OSMED), autosomal recessive deafness-77 (DFNB77), autosomal recessive deafness-84A (DFNB84A), autosomal recessive deafness-84B (DFNB77). NB84B), peripheral neuropathy, myopathy, autosomal dominant nonsyndromic hearing loss-4A (DFNA4), congenital thrombocytopenia, sensorineural hearing loss, autosomal dominant nonsyndromic hearing loss-56 (DFNA56), epileptic encephalopathy, Timothy syndrome, long QT syndrome, X-linked retinal disease, aldosteronism, autosomal recessive hearing loss-42 (DFNB42), primary aldosteronism (Conn syndrome), seizures, neurological abnormalities, sinoatrial node dysfunction, neurodevelopmental disorders, hypokalemic periodic paralysis, epilepsy, developmental and epileptic encephalopathy, Brody's myopathy, Darier's disease, heart disease, von Willebrand disease, or brain-hepatorenal syndrome.

[0032] In one embodiment, the target protein is DMD, PKHD1, F8, ABCA4, DYSF, OTOF, CFTR, ATP7B, MYOF, MYO7A, MYO15A, CDH23, STRC, OTOG, TECTA, PCDH15, TRIOBP, MYO3A, COL11A2, LOXHD1, PTPRQ, OTOGL, MYH14, MYH9, TNC, CACNA1A, CACNA1C, CACNA1F, CACNA1H, CACNA1G, CACNA1D, CACNA1B, CACNA1S, CACNA1I, CACNA1E, ATP2A1, ATP2A2, VWF, PEX1, or CMYA5.

[0033] In one embodiment, the first nucleic acid molecule comprises a nucleic acid sequence encoding an N-terminal fragment of dystrophin, mini-dystrophin, or micro-dystrophin; and wherein the second nucleic acid molecule comprises a nucleic acid sequence encoding a C-terminal fragment of dystrophin, mini-dystrophin, or micro-dystrophin; and wherein administration of the first and second nucleic acid molecules results in the production of full-length dystrophin, mini-dystrophin, or micro-dystrophin. In one embodiment, the first nucleic acid comprises the nucleic acid sequence of SEQ ID NO:150, and the second nucleic acid comprises the nucleic acid sequence of SEQ ID NO:151. In one embodiment, the first nucleic acid comprises the nucleic acid sequence of SEQ ID NO:152, and the second nucleic acid comprises the nucleic acid sequence of SEQ ID NO:153.

[0034] In one embodiment, the first nucleic acid molecule comprises a nucleic acid sequence encoding an N-terminal fragment of Dysferlin; and the second nucleic acid molecule comprises a nucleic acid sequence encoding a C-terminal fragment of Dysferlin; and administration of the first and second nucleic acid molecules results in the production of a full-length Dysferlin protein. In one embodiment, the first nucleic acid comprises the nucleic acid sequence of SEQ ID NO:157, and the second nucleic acid comprises the nucleic acid sequence of SEQ ID NO:158. In one embodiment, the first nucleic acid comprises the nucleic acid sequence of SEQ ID NO:177, and the second nucleic acid comprises the nucleic acid sequence of SEQ ID NO:178. In one embodiment, the first nucleic acid comprises the nucleic acid sequence of SEQ ID NO:179, and the second nucleic acid comprises the nucleic acid sequence of SEQ ID NO:178. In one embodiment, the first nucleic acid comprises the nucleic acid sequence of SEQ ID NO:181, and the second nucleic acid comprises the nucleic acid sequence of SEQ ID NO:182. In one embodiment, the first nucleic acid comprises the nucleic acid sequence of SEQ ID NO:183, and the second nucleic acid comprises the nucleic acid sequence of SEQ ID NO:182.

[0035] In one embodiment, the first nucleic acid molecule comprises a nucleic acid sequence encoding an N-terminal fragment of STRC; wherein the second nucleic acid molecule comprises a nucleic acid sequence encoding a C-terminal fragment of STRC; and wherein administration of the first and second nucleic acid molecules results in the production of a full-length STRC protein. In one embodiment, the first nucleic acid comprises the nucleic acid sequence of SEQ ID NO:160, and the second nucleic acid comprises the nucleic acid sequence of SEQ ID NO:161.

[0036] In one embodiment, the present invention includes a method for generating an RNA molecule encoding a target protein, the method comprising: administering a first RNA molecule to a cell or tissue, the first RNA molecule comprising a coding region encoding a first portion of the target protein, the coding region being directly linked at its 3' end to a nucleotide sequence of a ribozyme (3' ribozyme); and administering a second RNA molecule to the cell or tissue, the second RNA molecule comprising a coding region encoding a second portion of the target protein, the coding region being directly linked at its 5' end to a nucleotide sequence of a ribozyme (5' ribozyme); wherein self-cleavage of the 3' ribozyme generates a 3' phosphate at the 3' end of the first portion of the target protein, and self-cleavage of the 5' ribozyme generates a 5' hydroxyl group at the 5' end of the second portion of the target protein, thereby allowing the first and second RNA molecules to be joined without scarring to generate an RNA molecule encoding the target protein.

[0037] In one embodiment, at least one of the 3' ribozyme and the 5' ribozyme is SEQ ID NO:131, SEQ ID NO:132, SEQ ID NO:133, SEQ ID NO:134, SEQ ID NO:135, SEQ ID NO:136, SEQ ID NO:137, SEQ ID NO:138, SEQ ID NO:139, SEQ ID NO:140, SEQ ID NO:141, SEQ ID NO:142, SEQ ID NO:143, SEQ ID NO:144, SEQ ID NO:145, SEQ ID NO:166, SEQ ID NO:167, SEQ ID NO:168, SEQ ID NO:169, SEQ ID NO:170, SEQ ID NO:171, SEQ ID NO:172, SEQ ID NO:173, SEQ ID NO:174, SEQ ID NO:175, SEQ ID NO:176, SEQ ID NO:192, SEQ ID NO:193, SEQ ID NO:194, SEQ ID NO:195, SEQ ID NO:196, SEQ ID NO:192, SEQ ID NO:193, SEQ ID NO:194, SEQ ID NO:195, SEQ ID NO:196, SEQ ID NO:195, SEQ ID NO:196, SEQ ID NO:197, SEQ ID NO:198, SEQ ID NO:19 ... NO:194, SEQ ID NO:195, SEQ ID NO:196, SEQ ID NO:197, SEQ ID NO:198, SEQ ID NO:199, SEQ ID NO:200, SEQ ID NO:202, SEQ ID NO:203, SEQ ID NO:204, SEQ ID NO:205, SEQ ID NO:206, SEQ ID NO:207, SEQ ID NO:217, SEQ ID NO:218 or SEQ ID NO:219.

[0038] In one embodiment, the full-length of the target protein is greater than 1000 amino acid residues. In one embodiment, the target protein is selected from DMD, PKHD1, F8, ABCA4, DYSF, OTOF, CFTR, ATP7B, MYOF, MYO7A, MYO15A, CDH23, STRC, OTOG, TECTA, PCDH15, TRIOBP, MYO3A, COL11A2, LOXHD1, PTPRQ, OTOGL, MYH14, MYH9, TNC, CACNA1A, CACNA1C, CACNA1F, CACNA1H, CACNA1G, CACNA1D, CACNA1B, CACNA1S, CACNA1I, CACNA1E, ATP2A1, ATP2A2, VWF, PEX1, and CMYA5.

[0039] In one embodiment, the present invention includes a method for generating an RNA molecule encoding a target protein in vitro, comprising: providing a first RNA molecule comprising a coding region of a first portion encoding the target protein, the coding region being directly linked at its 3' end to a nucleotide sequence of a ribozyme (3' ribozyme); and providing a second RNA molecule comprising a coding region of a second portion encoding the target protein, the coding region being directly linked at its 5' end to a nucleotide sequence of a ribozyme (5' ribozyme); wherein self-cleavage of the 3' ribozyme produces a 3' phosphate at the 3' end of the first portion of the target protein, and self-cleavage of the 5' ribozyme produces a 5' hydroxyl group at the 5' end of the second portion of the target protein, thereby allowing the first and second RNA molecules to be joined without scarring to generate an RNA molecule encoding the target protein; and providing a ligase to induce the assembly of RNA molecules from the coding regions of the first RNA molecule and the coding regions of the second RNA molecule.

[0040] In one embodiment, at least one of the 3' ribozyme and 5' ribozyme is SEQ ID NO:131, SEQ ID NO:132, SEQ ID NO:133, SEQ ID NO:134, SEQ ID NO:135, SEQ ID NO:136, SEQ ID NO:137, SEQ ID NO:138, SEQ ID NO:139, SEQ ID NO:140, SEQ ID NO:141, SEQ ID NO:142, SEQ ID NO:143, SEQ ID NO:144, SEQ ID NO:145, SEQ ID NO:166, SEQ ID NO:167, SEQ ID NO:168, SEQ ID NO:169, SEQ ID NO:170, SEQ ID NO:171, SEQ ID NO:172, SEQ ID NO:173, SEQ ID NO:174, SEQ ID NO:175, SEQ ID NO:176, SEQ ID NO:192, SEQ ID NO:193, SEQ ID NO:194, SEQ ID NO:195, SEQ ID NO:19 ...7, SEQ ID NO:168, SEQ ID NO:169, SEQ ID NO:199, SEQ ID NO:194, SEQ ID NO:195, SEQ ID NO:196, SEQ ID NO:197, SEQ ID NO:198, SEQ ID NO:199, SEQ ID NO:199, SEQ ID NO:199, SEQ ID NO:199, SEQ ID NO:199, SEQ ID NO:199, SEQ ID NO:199, SEQ ID NO:199, SEQ ID NO:19 NO:194, SEQ ID NO:195, SEQ ID NO:196, SEQ ID NO:197, SEQ ID NO:198, SEQ ID NO:199, SEQ ID NO:200, SEQ ID NO:202, SEQ ID NO:203, SEQ ID NO:204, SEQ ID NO:205, SEQ ID NO:206, SEQ ID NO:207, SEQ ID NO:217, SEQ ID NO:218 or SEQ ID NO:219.

[0041] In one embodiment, the present invention includes a system for generating a circular RNA molecule encoding a target protein, the system comprising a nucleic acid encoding: a 5' ribozyme directly linked to a first portion of the target protein, an IRES sequence, and a second portion of the target protein directly linked to a 3' ribozyme, wherein self-cleavage of the 3' ribozyme produces a 3' phosphate at the 3' end of the second portion of the target protein, and self-cleavage of the 5' ribozyme produces a 5' hydroxyl group at the 5' end of the first portion of the target protein, thereby allowing scarless ligation of the 5' hydroxyl group and the 3' phosphate group of the RNA molecule to generate a circular RNA molecule encoding the target protein.

[0042] In one implementation, the RNA molecule includes an in vitro transcribed RNA molecule.

[0043] In one implementation, the target protein is a therapeutic protein, a reporter protein, or a Cas9 protein.

[0044] In one embodiment, at least one of the 3' ribozyme and the 5' ribozyme is SEQ ID NO:131, SEQ ID NO:132, SEQ ID NO:133, SEQ ID NO:134, SEQ ID NO:135, SEQ ID NO:136, SEQ ID NO:137, SEQ ID NO:138, SEQ ID NO:139, SEQ ID NO:140, SEQ ID NO:141, SEQ ID NO:142, SEQ ID NO:143, SEQ ID NO:144, SEQ ID NO:145, SEQ ID NO:166, SEQ ID NO:167, SEQ ID NO:168, SEQ ID NO:169, SEQ ID NO:170, SEQ ID NO:171, SEQ ID NO:172, SEQ ID NO:173, SEQ ID NO:174, SEQ ID NO:175, SEQ ID NO:176, SEQ ID NO:192, SEQ ID NO:193, SEQ ID NO:194, SEQ ID NO:195, SEQ ID NO:196, SEQ ID NO:192, SEQ ID NO:193, SEQ ID NO:194, SEQ ID NO:195, SEQ ID NO:196, SEQ ID NO:195, SEQ ID NO:196, SEQ ID NO:197, SEQ ID NO:198, SEQ ID NO:19 ... NO:194, SEQ ID NO:195, SEQ ID NO:196, SEQ ID NO:197, SEQ ID NO:198, SEQ ID NO:199, SEQ ID NO:200, SEQ ID NO:202, SEQ ID NO:203, SEQ ID NO:204, SEQ ID NO:205, SEQ ID NO:206, SEQ ID NO:207, SEQ ID NO:217, SEQ ID NO:218 or SEQ ID NO:219.

[0045] In one embodiment, the present invention includes a method for generating a circular RNA molecule in vivo, the method comprising administering a linearly transcribed RNA molecule or a DNA molecule encoding a linear RNA molecule to cells or tissues, wherein the linear RNA molecule comprises: a 5' ribozyme directly linked to a first portion of a target protein, an IRES sequence, and a second portion of the target protein directly linked to a 3' ribozyme, wherein self-cleavage of the 3' ribozyme generates a 3' phosphate at the 3' end of the second portion of the target protein, and self-cleavage of the 5' ribozyme generates a 5' hydroxyl group at the 5' end of the first portion of the target protein, thereby allowing scarless ligation of the 5' hydroxyl group and the 3' phosphate group of the RNA molecule to generate a circular RNA molecule encoding the target protein.

[0046] In one implementation, the target protein is a therapeutic protein, a reporter protein, or a Cas9 protein.

[0047] In one embodiment, at least one of the 3' ribozyme and the 5' ribozyme is SEQ ID NO:131, SEQ ID NO:132, SEQ ID NO:133, SEQ ID NO:134, SEQ ID NO:135, SEQ ID NO:136, SEQ ID NO:137, SEQ ID NO:138, SEQ ID NO:139, SEQ ID NO:140, SEQ ID NO:141, SEQ ID NO:142, SEQ ID NO:143, SEQ ID NO:144, SEQ ID NO:145, SEQ ID NO:166, SEQ ID NO:167, SEQ ID NO:168, SEQ ID NO:169, SEQ ID NO:170, SEQ ID NO:171, SEQ ID NO:172, SEQ ID NO:173, SEQ ID NO:174, SEQ ID NO:175, SEQ ID NO:176, SEQ ID NO:192, SEQ ID NO:193, SEQ ID NO:194, SEQ ID NO:195, SEQ ID NO:196, SEQ ID NO:192, SEQ ID NO:193, SEQ ID NO:194, SEQ ID NO:195, SEQ ID NO:196, SEQ ID NO:195, SEQ ID NO:196, SEQ ID NO:197, SEQ ID NO:198, SEQ ID NO:19 ... NO:194, SEQ ID NO:195, SEQ ID NO:196, SEQ ID NO:197, SEQ ID NO:198, SEQ ID NO:199, SEQ ID NO:200, SEQ ID NO:202, SEQ ID NO:203, SEQ ID NO:204, SEQ ID NO:205, SEQ ID NO:206, SEQ ID NO:207, SEQ ID NO:217, SEQ ID NO:218 or SEQ ID NO:219. Attached Figure Description

[0048] The following detailed description of embodiments of the invention will be more readily understood when read in conjunction with the accompanying drawings. It should be understood that the invention is not limited to the precise arrangement and means of the embodiments shown in the drawings.

[0049] Figures 1A through 1E depict data illustrating ribozyme-mediated transcription and expression in mammalian cells. Figure 1A shows a schematic diagram of vectors encoding the N-terminal (Nt) half of GFP with a 3' HDV ribozyme and the C-terminal (Ct) half of GFP with a 5' hammerhead (HH) ribozyme. Figure 1B depicts exemplary results showing that GFP fluorescence was detectable with co-expression of Nt-GFP-HDV and HH-Ct-GFP in COS7 and HEK293T cells, but not with individual transfection. Figures 1C-1D depict exemplary results of RT-PCR amplification (Figure 1C) and Sanger sequence analysis (Figure 1D) using specific primers for each individual RNA (G1 and G2), showing ribozyme removal and scarless trans-splicing and restoration of the GFP coding sequence. Figure 1E depicts exemplary Western blot analysis results using a GFP-specific antibody, showing the predicted full-length GFP protein size.

[0050] Figures 2A through 2E depict data illustrating the development of luciferase-based reporter genes to quantify the effects of ribozyme sequences on trans-splicing in mammalian cells. A schematic diagram in Figure 2A depicts vectors encoding the N-terminal (Nt) half of luciferase with a 3' HDV ribozyme and the C-terminal (Ct) half of luciferase with a 5' hammerhead (HH) ribozyme. Figures 2B-2C depict exemplary results of RT-PCR amplification (Figure 2B) and Sanger sequence analysis (Figure 2C) using specific primers for each individual Luc RNA (L1 and L2), showing ribozyme removal and scarless trans-splicing of the luciferase open reading frame. Figures 2D-2E demonstrate the effects of different HDV (Figure 2D) and HH (Figure 2E) ribozyme sequences on trans-splicing in mammalian cells. Furthermore, mutations in the ribozyme-catalyzed nucleotides lead to loss of luciferase activity (last column in Figure 2D and last column in Figure 2E).

[0051] Figures 3A through 3D depict data illustrating the regulation of protein expression in Nt and Ct vectors. Figure 3A shows the placement of a C-terminal protein degradation sequence, which prevents the expression of the protein encoded by the Nt vector. Figure 3B illustrates the effectiveness of different protein degradation sequences in preventing the expression of GFP-HDV in Nt vectors encoding full-length GFP. Figure 3C shows the placement of an N-terminal translation control sequence in the Ct vector to prevent the translation of the protein sequence. Figure 3D illustrates the effectiveness of different GFP sequence modifications or translation control sequences in preventing GFP fluorescence in mammalian cells.

[0052] Figures 4A through 4D depict data illustrating single and multiple ribozyme-mediated trans-splicing in mammalian cells. Figure 4A shows vectors encoding 4xMTS with a ribozyme and full-length GFP (without the start ATG codon) to mediate trans-splicing and expression of GFP proteins targeting mitochondria. Exemplary results depicted in Figure 4B show that co-expression of these vectors produced mitochondrial-localized green fluorescence that overlapped with the red fluorescence of the mitochondrial tracer CMXRos. Figure 4C shows vectors used for multiple trans-splicing and expression of a mitochondrial-targeting GFP protein (4xMTS-GFP) in reading frame 1 and a red fluorescent protein (F2-Myr-RFP) targeting the myristylated membrane in reading frame 2. Exemplary results depicted in Figure 4D demonstrate that co-expression of all four vectors in mammalian Cos7 cells resulted in specific green fluorescence in mitochondria and red fluorescence in the membrane.

[0053] Figures 5A and 5B depict data showing how optimized ribozyme sequences, along with cis-splicing splice acceptor and splice donor sequences, enhance ribozyme-mediated trans-splicing. Figure 5A shows the positions of the chimeric splice donor (SD) and splice acceptor (SA) sequences in the universal Nt-GFP-3'Rz and 5'Rz-Ct-GFP trans-splicing GFP reporter genes, where Rz represents the cis-splicing ribozyme. Figure 5B depicts exemplary results of GFP fluorescence in Cos7 cells after single-vector transfection (first two columns), co-transfection (last two columns), 18 hours post-transfection (first three columns), or 36 hours post-transfection (last column). The first row describes the use of unoptimized HH and HDV ribozymes, the second row describes the use of optimized Twister and RzB ribozymes, and the last row describes combinations of Twister and RzB ribozymes with SD and SA sequences.

[0054] Figures 6A through 6D depict data illustrating ribozyme-mediated trans-splicing of large protein-coding genes. Figure 6A shows a diagram illustrating a vector used for delivery of a split μ dystrophin-GFP fusion protein using an AAV vector. Figures 6B-6C depict exemplary results of RT-PCR (Figure 6B) and Sanger sequencing (Figure 6C) analyses of cells transfected with Nt-Dys and Ct-Dys vectors, showing specific trans-splicing. Figure 6D depicts exemplary results of GFP fluorescence imaging of cells transfected with Nt and Ct dystrophin vectors using confocal microscopy, showing predicted membrane localization of dystrophin.

[0055] Figures 7A through 7C depict data illustrating lentiviral delivery of ribozyme-containing RNA for trans-splicing in target cells. Figure 7A shows the negative sense orientation of the Nt and Ct split GFP expression cassettes in the lentiviral gene transfer vector. Figure 7B illustrates exemplary results demonstrating that only cells co-transduced with lentiviruses encoding both Nt-GFP and Ct-GFP genes exhibit GFP fluorescence. Figure 7C shows the negative sense orientation of the Nt and Ct split Dys expression cassettes in the lentiviral gene transfer vector.

[0056] Figures 8A and 8B depict data illustrating ribozyme-mediated trans-splicing and expression of the toxic DTA gene. Figure 8A shows a diagram of the vector encoding the splitting Nt and Ct DTA genes. Figure 8B illustrates exemplary results showing that co-transfection of Nt-DTA and Ct-DTA cells resulted in reduced expression of the co-transfected GFP reporter gene, consistent with the translational repressive function of DTA in mammalian cells.

[0057] Figure 9 The exemplary results described demonstrate that co-expression of exogenous RNA regulatory enzymes can enhance or inhibit ribozyme-mediated trans-splicing in mammalian cells.

[0058] Figures 10A through 10D depict data showing that RtcB is sufficient to catalyze ribozyme-mediated trans-splicing in vitro. Figure 10A shows a diagram depicting the splitting of a luciferase trans-splicing reporter gene containing an upstream T7 RNA promoter to allow in vitro RNA transcription. Figure 10B shows exemplary RT-PCR results demonstrating that in vitro trans-spliced ​​luciferase RNA depends on the addition of the RtcB protein (NEB) using manufacturer-recommended reaction conditions. Figure 10C shows a diagram depicting a trans-splicing vector for the conserved N-terminal (N1L) and C-terminal (N3R) domains of spider silk protein. Figure 10D depicts exemplary Sanger sequencing results demonstrating that the RtcB ligase from *E. coli* is sufficient to catalyze the trans-ligation of ribozyme-cleaved N1L and N3R-encoded RNAs.

[0059] Figure 11 The in vitro directional ligation of ribozyme-catalyzed RNA was described using RtcB, VS-S, and VS-Rz.

[0060] Figures 12A through 12D depict data illustrating RNA trans-splicing using trans-splicing ribozymes. Figure 12A depicts the secondary structure of a cis-splicing ribozyme. Figure 12B depicts an engineered ribozyme capable of trans-splicing. The graphs shown in Figures 12C and 12D illustrate the potential applications of trans-splicing ribozymes for deleting pathogenic mutations (such as frame-shifting or premature stop codons) to restore protein expression and function.

[0061] Figure 13A and Figure 13B The data depicted show the secondary structures of representative ribozymes that can be used for scarless trans-splicing of RNA. Figure 13A Representative ribozymes that can be used for scarless 5' shearing are described. Figure 13B Representative ribozymes suitable for scarless 3' splicing are depicted. N = any nucleotide. Red scissors indicate splicing sites. Red nucleotides represent catalytic mutations. Orange nucleotides represent RNA sequences to be trans-spliced. Dark blue nucleotides represent ribozyme sequences required to form the stem region. Light blue represents the tertiary stable motif (TSM) in stem region 1, which interacts with the loop in stem region 2. HH – Hammerhead, HDV – Hepatitis D virus, Rz – Ribozyme. (HH: Gao and Zhao, 2013, JIPB; HDV: Schürer et al., 2002, Nuc Acids Res; RzB: Saksmerprome et al., 2004, RNA; Twister: Liu et al., 2014, NatChem Bio).

[0062] Figure 14 A to Figure 14 The data depicted in C show scarless cleavage and induced RNA trans-splicing and expression using trans-activated ribozymes. Figure 14 The diagram depicted in Figure A shows that the VS ribozyme can be broken down into two components: a small VS-S stem-loop (lacking autocatalytic activity) and a larger VS-Rz (inducing VS-S cleavage upon trans delivery). The VS-S / VS-Rz ribozyme pair can be used to generate induced, scarless trans-splicing. Figure 14 The diagram in B illustrates a method for generating an inducible RNA trans-splicing system using a ribozyme pair that is trans-activated by VS-S / VS-Rz. Only upon delivery or expression of VS-Rz does the Nt-GFP-VS-S RNA generate the appropriate RNA terminus, which is capable of participating in trans-splicing with the co-expressed Ct-GFP RNA. Figure 14The diagram shown in C depicts a method for generating RNA with an N-terminal sequence, a variable or immutable repeat region, and a C-terminal sequence. This 'repetitive' RNA contains a 5' autocatalytic ribozyme and a 3' trans-activating ribozyme, such as VS-S, which allows for controlled repeat addition based on the selective addition of trans-activating VS-Rz and a ligase (e.g., RtcB).

[0063] Figure 15 A to Figure 15 The data depicted by E show ribozyme-mediated trans-splicing that generates stable intron RNA sequences. Figure 15 The diagram shown in Figure A illustrates the use of cis-cleaving ribozymes to mediate the trans-splicing of two independent RNAs. Figure 15 The diagram shown in B depicts the use of internal cis-cleavage ribozymes to generate synthetic introns. Figure 15 The exemplary results depicted by C demonstrate the efficient cis-splicing of synthetic introns and trans-splicing of independent RNA to produce a functional protein (GFP). Figure 15 D and Figure 15 The diagram shown in E depicts the generation of a reporter gene and intron sequence (“cargo”) through trans-splicing and translation using an internal cis-cutting ribozyme. This sequence can be any useful RNA sequence or gene expression cassette.

[0064] Figure 16 A to Figure 16 C depicts exemplary results of optimized ribozyme sequences for in vivo ribozyme-mediated trans-splicing. Figure 16 A depicts a comparison of relative ribozyme activities using a reporter gene with luciferase trans-splicing. The RzB hammerhead ribozyme variant contains a tertiary stable motif, is active at low magnesium concentrations, and exhibits the highest luciferase activity in mammalian cells. Figure 16 B describes a comparison between HDV ribozymes (HDV68 and genomic HDV) and Twister ribozymes (Twst). The Twister ribozyme at the Nt-Luc 3' end provides the greatest luciferase activity, while the catalytic inactivation mutation (Twst mut) eliminates this activity. Figure 16 C describes a comparison of Twister ribozyme sequence modifications. Shortening the P1 stem region reduces reporter gene activity. Modification of the first residue indicates that Twister can tolerate the A nucleotide at position 1 (U1A).

[0065] Figure 17 A to Figure 17 C describes the identification of the optimal ribozyme pair for StitchR-mediated RNA transsplicing in mammalian cells. Figure 17 A and Figure 17B depicts the relative activities of representative ribozyme sequences from major ribozyme families (including Twister, Twister Sister, Hammerhead, HDV, Pistol, Varkud Satellite (VS), Hairpin, and Hovlinc (Hov) ribozymes) in mammalian cells using a luciferase-based reporter gene screening method activated by StitchR. The addition of splice donor (SD) and splice acceptor (SA) sequences allows for the generation of functional introns in the spliced ​​mRNA, which can restore the complete luciferase open reading frame after processing. The highest luciferase expression was achieved using Twister ribozymes in both LucNt and LucCt vectors, which is very close to the expression results of vectors encoding luciferase in a single open reading frame (ORF). Figure 17 The data depicted in C show that different isoforms of Twister ribozymes were measured to determine their relative activities in promoting mRNA stitchr activity and expression in mammalian cells.

[0066] Figure 18 A to Figure 18 The data depicted by C show that StitchR is applied in mammalian cells for the delivery and expression of dual AAV genes. Figure 18 A describes the design of a StitchR-activated splitting GFP reporter cassette, which, under the control of the human cytomegalovirus (hCMV) promoter and the bovine growth hormone polyadenylated sequence (bGH pA), was subcloned into a vector flanked by the AAV2ITR sequence. The construct was used to generate the AAV2 / 1 serotype virus, which only exhibits epifluorescence upon co-transduction into human cells (HEK293T). Figure 18 B) and protein blotting ( Figure 18 C) Robust full-length GFP expression was detected.

[0067] Figure 19 A to Figure 19 Data depicted in C show that StitchR was administered in vivo for the delivery and expression of dual AAV genes. StitchR-activated dual AAV GFP reporter virus (2E+12 vg / ml) was injected into mice 10 days after birth (P10), and used 2 months post-injection. Figure 19 A) Whole-animal imaging based on IVIS fluorescence or ( Figure 19 B) Imaging with epifluorescence. Robust GFP fluorescence was detected throughout the mouse body and was also readily observed in the hind limb muscle tissue. Figure 19Data depicted by C showed that full-length GFP protein expression was observed only in mice simultaneously injected with GFPnt and GFPct viruses, as detected by Western blotting using an anti-GFP antibody.

[0068] Figure 20 This describes a subset of human loss-of-function monogenic diseases caused by mutations in large genes. The numbers in parentheses represent the amino acid sequence lengths of adjacent disease genes.

[0069] Figure 21 The data depicted show that mutations in large dystrophin (Dys) can lead to severe Duchenne muscular dystrophy (DMD) or less severe Becker muscular dystrophy (BMD). The figure depicts the characterized human dystrophin (Dys) domains, the locations of important protein-protein interactions, sequence deletions found in mild Becker muscular dystrophy (BMD), and engineered therapeutic dystrophins that can be loaded into single AAV (microDys) or dual AAV vector pairs (StitchR miniDys).

[0070] Figure 22 A schematic diagram depicts a miniature dystrophin AAV expression vector controlled by the heart and skeletal muscle-specific CK8e promoter and bGH pA.

[0071] Figure 23 A schematic diagram depicts a dual-StitchR activated N-terminal (StitchR Dys-Nt) and C-terminal (StitchR Dys-Ct) AAV expression vector controlled by a heart- and skeletal muscle-specific CK8e promoter and bGH pA.

[0072] Figure 24 A schematic diagram of a dual AAV vector expressing N-terminal (Dual AK Dys-Nt) and C-terminal (Dual AK Dys-Ct) sequences is depicted, which utilizes recombinant AK sequences and is controlled by a heart- and skeletal muscle-specific CK8e promoter and bGH pA.

[0073] Figure 25The data depicted show the expression of full-length mini-dystrophin in human HEK293T cells transfected with AK, StitchR, or a single ORF expression vector, under the control of the core EF1a promoter, as detected by Western blotting using N-terminal or C-terminal dystrophin-specific antibodies. StitchR technology efficiently generates the full-length protein (lane 7). This activity depends on ribozyme-mediated RNA splicing, as mutations in a single catalytic residue eliminate full-length expression (lane 8). StitchR is more efficient than vectors based on AK recombinant sequences, which did not produce detectable full-length mini-dystrophin expression at this exposure level (compare lanes 4 and 7). Notably, StitchR is almost as efficient at generating full-length mini-dystrophin as vectors encoding it in a single open reading frame (compare lanes 7 and 9).

[0074] Figure 26 The data depicted demonstrate, using Western blotting, the in vivo expression of full-length mini dystrophin via AAV transduction activated by StitchR, under the control of the core CK8e promoter. The StitchR technology efficiently generates full-length mini dystrophin, with levels in skeletal muscle (quadriceps) and the heart approaching those of endogenous mouse dystrophin. Dystrophin was not detected in the liver, demonstrating the muscle specificity of the CK8e promoter.

[0075] Figure 27 The diagram depicts the human Dysferlin (DYSF) protein, along with its known domains and the locations of important protein interactions. Mutations in DYSF in humans can cause limb-girdle muscular dystrophy type 2b (LGMD2B) and triyoshi myopathy (MM), now commonly referred to as Dysferlinopathies.

[0076] Figure 28 A schematic diagram of a pair of dual AAV vectors for StitchR activation of the full-length human codon-optimized Dysferlin (hcoDYSF) vector under the control of the cardiac and skeletal muscle-specific CK8e promoter and bGH pA is depicted.

[0077] Figure 29 The data depicted show the expression of full-length human Dysferlin protein under the control of the core EF1a promoter in human HEK293T cells transfected with a vector encoding dysferlin in a dual-StitchR vector or a single ORF expression vector, as detected by Western blotting using an anti-Dysferlin antibody.

[0078] Figure 30 The data depicted show the expression of full-length human Sterocilin (STRC) protein under the control of the core EF1a promoter in human HEK293T cells transfected with dual StitchR vectors, as detected by Western blotting using N-terminal or C-terminal anti-STRC antibodies.

[0079] Figure 31 This is a schematic diagram illustrating ribozyme-mediated scarless cis-splicing to produce circular RNA (which can be translated via IRES-mediated translation).

[0080] Figure 32 The experimental data depicted showed the generation of a circular RNA encoding GFP and the expression of GFP from that circular RNA.

[0081] Figure 33 This is a schematic diagram depicting the roles of IRES and ribozymes in proteins expressed by CirculR.

[0082] Figure 34 The use of CirculR for lentivirus-based screening of functional IRES sequences is described.

[0083] Figure 35 A schematic diagram depicts 'Stitch RNA' or StitchR, which provides scarless splicing of RNA molecules in the absence of sequence homology or splinting oligonucleotides.

[0084] Figure 36 The data depicted show the assembly and expression of ribozyme-activated mRNA.

[0085] The data in Figures 37A and 37B show the reconstruction of mRNA in cells and the expression of full-length GFP protein. Figure 37A depicts the scarless trans-splicing of two independently splitting GFP RNAs. Figure 37B depicts the expression of full-length GFP protein.

[0086] Figure 38 The development of a luciferase-based StitchR reporter gene assay was described. This reporter gene can be used to quantitatively measure StitchR activity in mammalian cells. Ribozyme activity is crucial for StitchR-mediated luciferase activity.

[0087] Figure 39 Data were drawn for the optimal ribozymes (mRzB-mutant RzB; mTwst-mutant Twister) for identifying StitchR activity in mammalian cells.

[0088] Figure 40 A schematic diagram illustrating conventional and unconventional RNA splicing pathways is provided.

[0089] Figure 41 An RNase that catalyzes the activity of StitchR in vitro was described. RtcB is a sufficient and necessary condition for the in vitro catalysis of StitchR activity.

[0090] Figure 42 The RNases that regulate StitchR activity in mammalian cells were described. RtcB enhances StitchR activity in mammalian cells. T4 PNK eliminates StitchR activity in mammalian cells.

[0091] Figure 43 The data depicted show the specific detection of 'stitched' RNA in mammalian cells.

[0092] Figure 44 The data depicted shows that intron sequences can be 'stitched' and spliced.

[0093] Figure 45 The data depicted show that intron sequences can be 'stitched' and spliced. The addition of 'stitched' introns significantly enhances StitchR activity in cells.

[0094] Figure 46 The RNases that can be used for intron sequences that will be 'stitched' and spliced ​​are described.

[0095] Figure 47 The data depicted demonstrate the optimization of StitchR activity in mammalian cells. The optimized StitchR activity approaches single-vector efficiency. StitchR using Twister Rz in Nt and Ct achieves 89% of the activity of a single ORF vector.

[0096] Figure 48 The data depicted demonstrate the optimization of StitchR activity in mammalian cells. The compact P1-type Twister ribozyme from rice (Oryzasativa, Osa) provides the strongest StitchR activity.

[0097] Figure 49 The data depicted show the reconstruction of endogenous full-length protein.

[0098] Figure 50 The data depicted show the reconstruction of endogenous full-length protein.

[0099] Figure 51 The data depicted demonstrates the application of StitchR in dual AAV gene therapy.

[0100] Figure 52 The data depicted show AAV delivery and expression of full-length GFP mediated by StitchR in mammalian cells.

[0101] Figure 53 The data depicted show AAV delivery and expression of full-length GFP mediated by StitchR in vivo.

[0102] Figure 54 The data depicted show AAV delivery and expression of full-length GFP mediated by StitchR in vivo.

[0103] Figure 55 The data depicted show the efficiency of trans-splicing of Stitch RNA in vivo.

[0104] Figure 56 This demonstrates the urgent health need to treat diseases caused by genes that are too large for a single AAV vector.

[0105] Figure 57 The StitchR AAV gene therapy for treating large muscle gene diseases is shown. Dysferlin (DYSF) – Dysferlin mutations cause limb-girdle muscular dystrophy type 2B and trigeminal myopathy (MM), now commonly referred to as Dysferlin proteinopathy. With over ~3000 patient-specific mutations and no common mutation hotspots, gene editing approaches face challenges. Gene therapy methods can be used to treat all patients.

[0106] Figure 58 The data depicted show the StitchR AAV gene therapy treatment for Dysferlin (DYSF).

[0107] Figure 59 The data depicted demonstrates the use of the StitchR AAV gene for Dysferlin (DYSF). The CK8e promoter is a compact, robust cardiac and skeletal muscle-specific promoter in mice and humans. AAV9 is an AAV serotype that effectively targets cardiac and skeletal muscle in mice and humans. Constructs designed and tested in mice can be directly translated for patient treatment.

[0108] Figure 60 The data depicted show the AAV delivery and expression of large therapeutic muscle proteins mediated by StitchR in vivo.

[0109] Figure 61 The diagram shows that StitchR has a variety of research and therapeutic applications, one of which is expanding the packaging capacity of therapeutic viral vectors.

[0110] Figures 62A through 62G depict experimental results demonstrating that StitchR-mediated expression of human midiDystrophin rescued the disease phenotype in a mouse model of Duchenne muscular dystrophy. Loss-of-function mutations in the large dystrophin (DMD) gene lead to Duchenne muscular dystrophy. Large sequence deletions in the DMD gene can lead to a milder form of Becker muscular dystrophy (BMD). (Figure 62A) This graph depicts the functional domains of human dystrophin (Dys) and known DMD-protein interactions. The full-length dystrophin open reading frame (~11 kb) is too large to be packaged into a single AAV particle or used in a dual AAV approach. A fully functional midiDystrophin (ΔH2-R15) can be packaged using a dual AAV approach, and its size is even larger than the deletion found in BMD. Here we reconstructed midiDystrophin in vivo using StitchR-activated RNA trans-ligation with a dual AAV9 serotype virus and the CK8e muscle-specific promoter. D2-MDX (dystrophin knockout (KO)) mice lack dystrophin expression and exhibit muscle pathology characterized by muscle atrophy, myofibril necrosis, and fibrosis before one month of age. (Figs. 62B and 62C) Injection of D2-MDX mice with dual AAV StitchR gene therapy resulted in midiDystrophin expression levels in both male and female mice being similar to those of full-length dystrophin. (Fig. 62D) StitchR midiDystrophin expression levels on the myofibril membrane were similar to those of endogenous full-length dystrophin and rescued the muscle pathology and membrane targeting of the known dystrophin-interacting protein nNOS, whose localization is normally disrupted in D2-MDX mice. (Fig. 62E) Western blot densitometric analysis showed that midiDystrophin expression was 94.9% of the wild-type level. In D2-MDX mice, expression of MidiDystrophin was sufficient to restore serum creatine kinase (CK) levels to wild-type levels (Fig. 62F) and significantly reduced the percentage of myofibrils with central nuclei (Fig. 62G), both of which are hallmarks of DMD.

[0111] Figure 63 The quantitative analysis of StitchR-mediated ΔH2-R15 expression was depicted using densitometrics and compared with WTDMD Western blot. Samples were normalized according to the Vinculin loading control.

[0112] Figures 64A to 64E depict the experimental results showing the expression of full-length human Dysferlin (DYSF) mediated by StitchR in Dysferlin knockout mice (AJ strain). (Figure 64A) The Dysferlin gene encodes a large muscle membrane protein, in which loss-of-function mutations lead to limb-girdle muscular dystrophy type 2B and Triyoshi myopathy, now commonly referred to as Dysferlinopathies. The Dysferlin open reading frame exceeds the packaging capacity of a single AAV vector, but can be fully packaged using a dual AAV vector approach. Using a StitchR-mediated dual AAV9 vector approach, full-length Dysferlin was reconstructed and expressed under the control of the muscle-specific CK8e promoter. (Figures 64B and 64C) The expression level of full-length human Dysferlin protein exceeded that observed in wild-type male and female mice. (Fig. 64D) Human Dysferlin is similar to wild-type Dysferlin, located on the cell membrane, but it also accumulates inside the cell, as observed in other Dysferlin transgenic methods. (Fig. 64E) Relative strength of Dysferlin.

[0113] Figures 65A and 65B depict experimental results showing the development of the stitchR activity-dependent gene cassette and cell lines. (Figure 65A) This figure depicts the stitchR-dependent ribotron-blocking HSV-TK expression cassette encoded within a lentiviral vector, reverse-oriented to enable lentiviral RNA delivery. Lentiviral delivery allows the ribotron-encoded HSV-TK cassette to be non-indelibly integrated into cells; however, other methods can also be used to stably express the ribotron HSV-TK transgene. Adding an antibiotic selection gene (e.g., blasticidin) can screen for stably integrated cell clones. (Figure 65B) Expression of the ribotron-activated synthetic intron (or ribotron) leads to the stitchR-dependent expression of the HSV-TK cell death gene in cells, which induces cell death in the presence of the non-toxic molecule ganciclovir (GCN). As a control, mutations in the catalytic nucleotides of the ribozyme sequence disrupt sensitization to ganciclovir.

[0114] Figure 66The experimental results depicted show that in vitro transcribed Circul RNA can be robustly translated without the additional processing (capping and poly(A) tail) required to make linear in vitro transcribed RNA suitable for translation. Linear mRNA undergoes natural processing similar to that of endogenous RNA, such as capping, splicing, and polyadenylation, when transcribed from DNA plasmids in cells—modifications essential for protein translation. This can be visualized using DNA plasmids encoding the green fluorescent protein (GFP) open reading frame, which, when controlled by mammalian promoters (e.g., sCMV IE94), produce robust green fluorescence compared to untransfected cells. Interestingly, however, RNA containing the same GFP open reading frame, transcribed in vitro using phage RNA polymerases (e.g., T7), is not translated upon transfection into cells. Linear in vitro transcribed RNA requires the addition of a 5' cap and a 3' poly(A) tail to produce green fluorescence. Adding an internal ribosome entry sequence (IRES), such as the IRES of Coxsackievirus B3 (CVB3 IRES), allows cap-independent translation, but this significantly reduces the translation of linear RNA transcribed from plasmid DNA and is insufficient to allow protein translation from linearly transcribed RNA. CirculR RNA encodes flanking autocatalytic ribozyme sequences that produce unique 5' and 3' ends, catalyzing scarless circularization of RNA in eukaryotic cells. Notably, ribozyme-cleaved RNA undergoes circularization; these RNAs are either transcribed intracellularly from plasmid DNA or generated in vitro in a linear form and then transfected into cells. Because circular RNA lacks a 5' or 3' end, CirculR RNA requires the addition of an IRES for translation, and its activity upon transfection from plasmid DNA is comparable to that of linear RNA translated from IRES. However, CirculR RNA containing IRES sequences is the only in vitro transcribed RNA capable of being translated into protein, and its activity is comparable to that of linear RNA processed with a 5' cap and a 3' poly(A) tail.

[0115] Figures 67A to 67C illustrate the scarless reconstruction of functional RNA motifs mediated by CirculR. Previously, we demonstrated the unique ability of CirculR (ribozyme-mediated RNA circularization in eukaryotic cells) for scarless reconstruction of open reading frames encoding proteins (GFP, blastcinon, etc.). This paper demonstrates that CirculR can be used for scarless reconstruction of functional RNA motifs in circular RNA form (Figure 67A). For example, small RNA motifs with known protein-binding chaperones (MS2, PRR1, Qβ, PP7, S1, etc.) can be integrated into the design of scarless self-cleaving ribozymes (e.g., Twister, RzB hammerhead ribozymes, etc.) (Figure 67B). Self-cleavage and ligation of these two ends in vitro or in cells result in the reconstruction of the functional RNA motif upon circularization (Figure 67C). This can be used to specifically recognize both circular and linear RNA forms, thus enabling purification, intracellular targeting, or packaging of circular RNA.

[0116] Figure 68 Depicting Figure 17 The average relative luciferase value of B.

[0117] Figure 69 Depicting Figure 17 The average relative luciferase value of C.

[0118] Figure 70 Sequence optimization and requirements for StitchR activity are described. StitchR-mediated trans-ligation of full-length Dysferlin has previously been confirmed in cell-based assays. The data here show that the polyadenylated sequence (PAS) in the N-terminal StitchR vector is not essential for stitchR activity. StitchR is compatible with other dual-vector technologies, such as DNA trans-splicing mediated by AAV ITR sequences or recombinant sequences (e.g., AK sequences). In cell-based assays using transfected plasmid DNA, the addition of the AK sequence neither promotes nor interferes with StitchR activity.

[0119] Figure 71 This study compares intronipeptide-mediated protein trans-splicing with StitchR-activated RNA trans-ligation. The INTEIN (intronipeptide) or INTEIN+EXTEIN (intronipeptide + exopeptide) protein trans-ligation methods are compared with StitchR in terms of full-length Dysferlin expression. A single INTEIN sequence is insufficient to promote full-length Dysferlin expression at this cleavage site because INTEIN requires flanking EXTEIN sequences to improve efficiency. INTEIN with the optimal EXTEIN sequence (leaving a small protein scar) is less efficient than StitchR-based Dysferlin expression. Detailed Implementation

[0120] definition

[0121] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains.

[0122] Generally, the nomenclature used in this article, as well as the laboratory procedures in cell culture, molecular genetics, organic chemistry, nucleic acid chemistry, and hybridization, are well-known and commonly used in the field.

[0123] Nucleic acids and peptides are synthesized using standard techniques. These techniques and procedures are generally performed according to conventional methods in the field and various general references (e.g., Sambrook and Russell, 2012, Molecular Cloning, A Laboratory Approach, Cold Spring Harbor Press, Cold Spring Harbor, NY; and Ausubel et al., 2012, Current Protocols in Molecular Biology, John Wiley & Sons, NY), which are provided throughout this article.

[0124] The nomenclature used in this article, as well as the laboratory procedures used in analytical chemistry and organic synthesis described below, are well-known and commonly used in the art. Standard techniques, or their modifications, can be used in chemical synthesis and chemical analysis.

[0125] The terms “a,” “an,” “the,” and similar terms used in the context of this invention (especially in the claims) should be understood to cover both the singular and the plural, unless otherwise stated herein or clearly contradicted by the context.

[0126] As used herein, the word “about” when referring to a measurable value (e.g., quantity, duration of time, etc.) is intended to cover deviations from a specified value of ±20%, or ±10%, or ±5%, or ±1%, or ±0.1%, provided that such deviations are suitable for implementing the disclosed method.

[0127] "Antense" specifically refers to the nucleic acid sequence of the non-coding strand of a double-stranded DNA molecule encoding a protein, or a sequence substantially homologous to that non-coding strand. As defined herein, an antisense sequence is complementary to the sequence of the double-stranded DNA molecule encoding the protein. An antisense sequence does not necessarily have to be complementary only to the coding portion of the DNA molecule. An antisense sequence can be complementary to a specified regulatory sequence on the coding strand of the DNA molecule encoding the protein, which controls the expression of the coding sequence.

[0128] When referring to the attachment of molecules (such as nucleic acid molecules) to a solid support, the term “linkage” as used herein is intended to encompass direct or indirect, covalent or non-covalent linkages, unless explicitly stated or indicated by context otherwise.

[0129] As used interchangeably in this paper, “microspheres,” “beads,” or their grammatical equivalents describe small, discrete particles that can serve as solid supports for connecting biomolecules (e.g., nucleic acid molecules).

[0130] "Disease" refers to a state of health in which an animal is unable to maintain homeostasis. If the disease is not treated, the animal's health will continue to deteriorate.

[0131] In contrast, an animal's "disorder" refers to a condition where the animal is able to maintain homeostasis, but its health is not as good as it would be without a disorder. Without treatment, a disorder does not necessarily lead to a further decline in the animal's health.

[0132] The disease or disorder is considered to be "relieved" if the severity of the signs or symptoms of the disease or disorder, the frequency of the patient's occurrence of such signs or symptoms, or both decrease.

[0133] "Encoding" refers to the inherent characteristics of a specific nucleotide sequence in a polynucleotide (such as a gene, cDNA, or mRNA) that serves as a template for the synthesis of other polymers and macromolecules in biological processes. These polymers and macromolecules have defined nucleotide sequences (i.e., rRNA, tRNA, and mRNA) or defined amino acid sequences, and the resulting biological properties. Therefore, if the transcription and translation of the mRNA corresponding to a gene produces a protein in a cell or other biological system, then the gene encodes that protein. Both the coding strand (whose nucleotide sequence is identical to the mRNA sequence and is usually provided in the sequence listing) and the non-coding strand (which serves as a template for gene or cDNA transcription) can be called the gene or cDNA that encodes that protein or other product.

[0134] The terms “patient,” “subject,” and “individual,” etc., are used interchangeably herein to refer to any animal or cell, whether in vitro or in vivo, suitable for the methods described herein. In one embodiment, the subject includes vertebrates and invertebrates. Invertebrates include, but are not limited to, *Drosophila melanogaster* and *Caenorhabditis elegans*. Vertebrates include, but are not limited to, primates, rodents, livestock, or game animals. Primates include, but are not limited to, chimpanzees, cynomolgus monkeys, spider monkeys, and macaques (e.g., rhesus monkeys). Rodents include, but are not limited to, mice, rats, marmots, ferrets, rabbits, and hamsters. Livestock and game animals include, but are not limited to, cattle, horses, pigs, deer, bison, buffalo, felines (e.g., domestic cats), canines (e.g., dogs, foxes, wolves), birds (e.g., chickens, emus, ostriches), and fish (e.g., zebrafish, trout, catfish, and salmon). In some embodiments, the subject is a mammal, such as a primate, such as a human. In some non-limiting implementations, the patient, subject, or individual is a human.

[0135] The term "specific binding" used for antibodies in this article refers to antibodies that recognize a specific antigen but substantially do not recognize or bind to other molecules in a sample. For example, an antibody that specifically binds to an antigen of one species may also bind to that antigen of one or more species. However, this cross-species reactivity itself does not change the antibody's specific classification. In another example, an antibody that specifically binds to an antigen may also bind to different allelic forms of that antigen. However, this cross-reactivity itself does not change the antibody's specific classification.

[0136] In some cases, the terms "specific binding" or "specifically binding" can be used to refer to the interaction of an antibody, protein, or peptide with a second chemical substance in a way that depends on the presence of a specific structure (e.g., an antigenic determinant or epitope) on that chemical substance; for example, the antibody recognizes and binds to a specific protein structure rather than the usual protein. If an antibody is specific for epitope "A," then in a reaction containing labeled "A" and an antibody, the presence of a molecule containing epitope A (or free, unlabeled A) will reduce the amount of labeled A that binds to the antibody.

[0137] The "coding region" of a gene consists of nucleotide residues in the coding strand and nucleotides in the non-coding strand, which are homologous to or complementary to the coding region of the mRNA molecule transcribed from the gene.

[0138] The coding region of an mRNA molecule is also composed of nucleotide residues that match the anticodon region of the transfer RNA molecule or encode a stop codon during the translation of the mRNA molecule. Therefore, the coding region can contain nucleotide residues containing codons for amino acid residues that are not present in the mature protein encoded by the mRNA molecule (e.g., amino acid residues in the protein output signal sequence).

[0139] The term "complementarity" used herein to refer to nucleic acids is a broad concept of sequence complementarity between regions of two nucleic acid strands or between two regions of the same nucleic acid strand. It is known that adenine residues in a first nucleic acid region can form specific hydrogen bonds ("base pairing") with residues in a second nucleic acid region antiparallel to the first region (if the residue is thymine or uracil). Similarly, it is known that cytosine residues in a first nucleic acid strand can base pair with residues in a second nucleic acid strand antiparallel to the first strand (if the residue is guanine). If, when the two regions are arranged in an antiparallel manner, at least one nucleotide residue in the first region can base pair with a residue in the second region, then the first region of the nucleic acid is complementary to the second region of the same or different nucleic acids. In one embodiment, the first region comprises a first portion, and the second region comprises a second portion, such that, when the first and second portions are arranged in an antiparallel manner, at least about 50%, at least about 75%, at least about 90%, or at least about 95% of the nucleotide residues in the first portion can base pair with nucleotide residues in the second portion. In one embodiment, all nucleotide residues in the first portion can base pair with nucleotide residues in the second portion.

[0140] The term "DNA" as used in this article is defined as deoxyribonucleic acid.

[0141] As used in this article, the term “expression” is defined as the transcription and / or translation of a specific nucleotide sequence driven by its promoter.

[0142] As used herein, the term "expression vector" refers to a vector containing at least a portion of a nucleic acid sequence encoding a gene product that can be transcribed. In some cases, the RNA molecule is subsequently translated into a protein, polypeptide, or peptide. In other cases, these sequences are not translated, such as in the production of antisense molecules, siRNA, ribozymes, etc. Expression vectors can contain a variety of control sequences, which are nucleic acid sequences essential for transcription and possible translation of coding sequences that are operablely linked in a particular host organism. In addition to control sequences that regulate transcription and translation, vectors and expression vectors may also contain nucleic acid sequences that perform other functions.

[0143] As used herein, the term "wild type" is a term in the art as understood by those skilled in the art, referring to the typical form of an organism, strain, gene, or trait that exists in nature, in order to distinguish it from mutant or variant forms.

[0144] The term "homology" refers to the degree of complementarity. Homology can be partial or complete (i.e., identical). Homology is typically measured using sequence analysis software (e.g., the sequence analysis software package from the Genetics Computing Group at the University of Wisconsin-Madison Biotechnology Center, 1710 University Avenue, Madison, Wisconsin 53705). This software matches similar sequences by assigning degrees of homology to various substitutions, deletions, insertions, and other modifications. Conserved substitutions typically include substitutions within the following groups: glycine, alanine; valine, isoleucine, leucine; aspartic acid, glutamic acid, asparagine, glutamine; serine, threonine; lysine, arginine; and phenylalanine, tyrosine.

[0145] "Separated" means altered or removed from its natural state. For example, nucleic acids or peptides that are naturally present in the normal environment of a living animal are not "separated," but the same nucleic acids or peptides that are partially or completely separated from their natural counterparts are "separated." Separated nucleic acids or proteins can exist in a substantially purified form or in non-natural environments, such as host cells.

[0146] When the term "isolated" is used for nucleic acids, such as "isolated oligonucleotides" or "isolated polynucleotides," it refers to a nucleic acid sequence that has been identified and isolated from at least one contaminant that typically accompanies its source. Thus, isolated nucleic acids exist in a form or environment different from their natural occurrence. Conversely, non-isolated nucleic acids (e.g., DNA and RNA) exist in their natural state. For example, a given DNA sequence (e.g., a gene) is located on the chromosome of a host cell, adjacent to neighboring genes; an RNA sequence (e.g., a specific mRNA sequence encoding a particular protein) exists in the cell in a mixed form with many other mRNAs encoding multiple proteins. However, isolated nucleic acids include, for example, nucleic acids in cells that normally express the nucleic acid, where the nucleic acid is located at a different chromosomal location than in natural cells, or flanked by nucleic acid sequences different from those found in nature. Isolated nucleic acids or oligonucleotides can exist in single-stranded or double-stranded form. When isolated nucleic acids or oligonucleotides are used to express proteins, the oligonucleotide must contain at least a sense strand or a coding strand (i.e., the oligonucleotide can be single-stranded), but it can also contain both sense and antisense strands (i.e., the oligonucleotide can be double-stranded).

[0147] When the term "isolated" is used to refer to polypeptides, such as "isolated protein" or "isolated polypeptide," it refers to a polypeptide that has been identified and isolated from at least one contaminant that typically accompanies its source. Therefore, isolated polypeptides exist in a form or environment different from their natural occurrence. In contrast, non-isolated polypeptides (such as proteins and enzymes) exist in their natural state.

[0148] "Nucleic acid" refers to any nucleic acid, whether it is composed of deoxyribonucleosides or ribonucleosides, and whether it is composed of phosphodiester linkages or modified linkages, such as phosphotriester linkages, aminophosphate linkages, siloxane linkages, carbonate linkages, carboxymethyl ester linkages, glycine linkages, carbamate linkages, thioether linkages, bridged aminophosphate linkages, bridged methylene phosphonate linkages, thiophosphate linkages, methylphosphonate linkages, dithiophosphate linkages, bridged thiophosphate linkages, or sulfone linkages, as well as combinations of these linkages. The term nucleic acid also specifically includes nucleic acids composed of bases other than those found in the five biologically known bases (adenine, guanine, thymine, cytosine, and uracil). The term "nucleic acid" generally refers to large polynucleotides.

[0149] This article uses conventional symbols to describe polynucleotide sequences: the left-hand end of a single-stranded polynucleotide sequence is called the 5' end; the left-hand direction of a double-stranded polynucleotide sequence is called the 5' direction.

[0150] The direction in which nucleotides are added to nascent RNA transcripts from the 5' end to the 3' end is called the transcription direction. The DNA strand with the same sequence as the mRNA is called the "coding strand"; the sequence on the DNA strand located at the 5' end of the reference point on the DNA strand is called the "upstream sequence"; the sequence on the DNA strand located at the 3' end of the reference point on the DNA strand is called the "downstream sequence".

[0151] An "expression cassette" is a nucleic acid molecule that contains a coding sequence that is operatively linked to a promoter / regulatory sequence necessary for transcription and, optionally, translation of that coding sequence.

[0152] As used herein, the term "operably ligated" refers to the ligation of nucleic acid sequences in a manner that produces a nucleic acid molecule capable of directing the transcription of a given gene and / or the synthesis of a desired protein molecule. The term also refers to the ligation of sequences encoding amino acids in a manner that produces a functional protein or polypeptide (e.g., possessing enzymatic activity, capable of binding to binding partners, capable of inhibition, etc.).

[0153] As used herein, the term "promoter / regulatory sequence" refers to a nucleic acid sequence necessary for the expression of a gene product operatively linked to a promoter / regulatory sequence. In some cases, this sequence may be a core promoter sequence; in others, it may also contain enhancer sequences and other regulatory elements necessary for gene product expression. For example, a promoter / regulatory sequence may be a sequence that expresses a gene product in an inducible manner.

[0154] The “strict hybridization conditions” used in this article refer to conditions under which nucleic acids complementary to the target sequence primarily hybridize with the target sequence and minimally hybridize with non-target sequences. Strict conditions are typically sequence-dependent and vary depending on various factors. Generally, the longer the sequence, the higher the temperature required for specific hybridization with its target sequence. Non-restrictive examples of strict conditions are described in detail in Tijssen (1993), Laboratory Techniques In Biochemistry And Molecular Biology—Hybridization With Nucleic Acid Probes Part 1, Second Chapter “Overview of principles of hybridization and the strategy of nucleic acid probe assay”, Elsevier, NY.

[0155] "Hybridization" refers to a reaction in which one or more polynucleotides react to form a complex that is stabilized by hydrogen bonds between the bases of the nucleotide residues. Hydrogen bonds can occur through Watson-Crick base pairing, Hoogstein binding, or any other sequence-specific mechanism. The complex can consist of two strands forming a double-stranded structure, three or more strands forming a multi-stranded complex, a single self-hybridizing strand, or any combination of these structures. Hybridization can constitute a step in a broader process, such as the initiation of PCR or the cleavage of polynucleotides by an enzyme. A sequence capable of hybridizing with a given sequence is called the "complementary sequence" of that given sequence.

[0156] An "inducible" promoter is a nucleotide sequence that, when operatively linked to a polynucleotide encoding or specifying a gene product, will essentially only produce the gene product in the presence of an inducer corresponding to that promoter.

[0157] A "constitutive" promoter is a nucleotide sequence that, when operatively linked to a polynucleotide that encodes or specifies a gene product, results in the production of that gene product in the cell under most or all physiological conditions.

[0158] As used herein, the term "polynucleotide" is defined as a nucleotide chain. Furthermore, nucleic acids are polymers of nucleotides. Therefore, the terms nucleic acid and polynucleotide are used interchangeably as used herein. Those skilled in the art will generally understand that nucleic acids are polynucleotides, which can be hydrolyzed into monomeric "nucleotides." Monomeric nucleotides can be hydrolyzed into nucleosides. Polynucleotides as used herein include, but are not limited to, all nucleic acid sequences obtained by any method available in the art (including, but not limited to, recombinant methods (i.e., cloning nucleic acid sequences from recombinant libraries or cell genomes using common cloning techniques and PCR, etc.) and synthetic methods).

[0159] In the context of this invention, the following abbreviations for common nucleic acid bases are used: “A” for adenosine, “C” for cytosine, “G” for guanosine, “T” for thymidine, and “U” for uridine.

[0160] As used herein, the terms “peptide,” “polypeptide,” and “protein” are used interchangeably to refer to compounds composed of amino acid residues covalently linked by peptide bonds. A protein or peptide must contain at least two amino acids, and there is no limit to the maximum number of amino acids constituting the protein or peptide sequence. A polypeptide includes any peptide or protein containing two or more amino acids linked together by peptide bonds. The term as used herein refers both to short chains, such as peptides, oligopeptides, and oligomers commonly referred to in the art, and to long chains, such as proteins, which come in various types. “Polypeptide” includes, for example, biologically active fragments, substantially homologous polypeptides, oligopeptides, homodimers, heterodimers, polypeptide variants, modified polypeptides, derivatives, analogs, fusion proteins, etc. Polypeptides include natural peptides, recombinant peptides, synthetic peptides, or combinations thereof.

[0161] The term “RNA” used in this article is defined as ribonucleic acid.

[0162] As used in this article, the term "ribozyme" refers to an RNA molecule capable of acting as an enzyme. For example, some ribozymes can cleave RNA molecules. RNA-cleaving ribozymes typically consist of at least a catalytic domain and a recognition sequence recognized by that catalytic domain. The catalytic domain may be part of an RNA molecule that is identical to the recognition sequence, thereby mediating cis-cleavage. Alternatively, the catalytic domain may be a different RNA molecule from the one containing the recognition sequence, thereby mediating trans-cleavage.

[0163] "Recombinant polynucleotides" are polynucleotides whose sequences are not naturally linked together. Amplified or assembled recombinant polynucleotides can be contained in a suitable vector, and this vector can be used to transform suitable host cells.

[0164] Recombinant polynucleotides can also perform non-coding functions (such as promoters, origins of replication, ribosome binding sites, etc.).

[0165] The term “recombinant polypeptide” as used in this article is defined as a polypeptide produced using recombinant DNA methods.

[0166] As used herein, the terms “solid surface,” “solid support,” and their grammatical equivalents refer to any material suitable or modifiable for connecting biomolecules (e.g., nucleic acid molecules).

[0167] As used in this article, the term "tag" refers to any chemical modification of a biomolecule (e.g., a nucleic acid molecule) to provide additional functionality (e.g., attachment to a solid support, fluorescence visualization, etc.).

[0168] As used in this paper, the term "variant" refers to a nucleic acid or peptide sequence that differs from a reference nucleic acid or peptide sequence, but retains the essential biological characteristics of the reference molecule. Changes in the nucleic acid variant sequence may not alter the amino acid sequence of the peptide encoded by the reference nucleic acid, but may result in amino acid substitutions, additions, deletions, fusions, and truncations. Changes in the peptide variant sequence are typically limited or conserved; therefore, the reference peptide and variant sequences are generally very similar and identical in many regions. The amino acid sequence differences between the variant and the reference peptide can be due to any combination of one or more substitutions, additions, or deletions. Nucleic acid or peptide variants can be naturally occurring, such as allelic variants, or unknown naturally occurring variants. Non-naturally occurring nucleic acid and peptide variants can be prepared by mutagenesis or direct synthesis.

[0169] "Vector" refers to a composition containing isolated nucleic acids that can be used to deliver the isolated nucleic acids into cells. Various vectors are known in the art, including, but not limited to, linear polynucleotides, polynucleotides bound to ionic or amphiphilic compounds, plasmids, and viruses. Therefore, the term "vector" includes autonomously replicating plasmids or viruses. The term should also be understood to include non-plasmid and non-viral compounds that facilitate the transfer of nucleic acids into cells, such as polylysine compounds, liposomes, etc. Examples of viral vectors include, but are not limited to, adenovirus vectors, adeno-associated virus vectors, retroviral vectors, etc.

[0170] Scope: In this disclosure, various aspects of the invention may be presented in the form of scope. It should be understood that the use of scope is merely for convenience and brevity and should not be construed as a rigid limitation on the scope of the invention. Therefore, the scope description should be considered as having specifically disclosed all possible sub-scopes and the various values ​​within those scopes. For example, the description of a scope such as 1 to 6 should be considered as having specifically disclosed sub-scopes such as 1 to 3, 1 to 4, 1 to 5, 2 to 4, 2 to 6, 3 to 6, etc., and the various values ​​within those scopes such as 1, 2, 2.7, 3, 4, 5, 5.3, and 6. This applies regardless of the width of the scope.

[0171] describe

[0172] This invention provides compositions and methods for efficiently and reliably linking two or more individual RNA molecules to produce larger single RNA molecules encoding proteins and fusion proteins. The invention utilizes ribozyme-mediated trans-splicing of multiple RNA molecules to assemble single RNA molecules encoding a target protein or fusion protein. This invention can be used for the efficient production of fusion proteins, chimeric proteins, etc. Furthermore, this invention can be used to produce large expression products (e.g., proteins or fusion proteins) whose coding sequences may be too large to be packaged into a single vector. For example, in some embodiments, the full-length expression product is longer than 1000 amino acid residues. In some embodiments, the full-length expression product is longer than 2000 amino acid residues. In some embodiments, the full-length expression product is longer than 3000 amino acid residues. Moreover, the technique of this invention allows for the rapid and convenient combination of two different sequences, which has a multiplier effect for generating new protein combinations or library sequences. This may be particularly suitable for, for example, generating synthetic antibodies (such as nanobodies) or functional screening of enzymes.

[0173] This invention also provides compositions and methods for the efficient delivery of one or more RNA molecules with ribozyme-flanked synthetic introns. The ribozyme-flanked synthetic intron may be located between a first RNA portion encoding the N-terminal portion of a target protein and a second RNA portion encoding the C-terminal portion of the target protein. The ribozyme-flanked synthetic intron may contain a cargo sequence, for example, a sequence encoding a therapeutic protein or containing functional RNA. Using two ribozymes allows cis-splicing to produce three RNA fragments: 1) a first RNA portion encoding the N-terminal portion of the target protein, 2) the ribozyme-flanked synthetic intron, and 3) a second RNA portion encoding the C-terminal portion of the target protein. The cis-splicing produces compatible ends for ligation. Ligating the compatible ends of the cis-spliced ​​synthetic introns generates a circular RNA molecule, which is more resistant to degradation than linear RNA molecules. Ligating the first RNA portion encoding the N-terminal portion of the target protein with the compatible ends of the second RNA portion encoding the C-terminal portion of the target protein generates an RNA molecule encoding the full-length target protein. The full-length target protein can be, for example, a therapeutic protein, a CRISPR-Cas protein, or a reporter protein, to provide an alternative indicator for the delivery and expression of cargo sequences in circular RNA molecules containing synthetic introns flanking ribozymes.

[0174] In one aspect, the present invention provides one or more nucleic acid molecules encoding two or more RNA molecules. In some embodiments, the one or more RNA molecules contain a ribozyme. In one embodiment, the one or more RNA molecules contain a coding region and a ribozyme. In some embodiments, the ribozyme self-cleaves from the RNA molecule, leaving the coding region. Exemplary ribozymes that can be used in the context of the present invention include, but are not limited to, hammerhead ribozymes (HH), hepatitis D virus (HDV) ribozymes, Varkud Satellite (VS) ribozymes, Sister ribozymes, Twister-sister ribozymes, Hairpin ribozymes, Hatchet ribozymes, Pistol ribozymes, HOV Linc ribozymes, or members of the lantern family. However, the present invention is not limited to any particular ribozyme, but covers all known endogenous ribozyme family members and any potential artificial ribozymes. That is, the described ribozymes, all known endogenous ribozymes, and potential artificial ribozymes can be used for the ligation, trans-splicing, and circularization of multiple RNAs, as described elsewhere herein.

[0175] For example, in one embodiment, the composition comprises a nucleic acid molecule encoding a first RNA molecule, wherein the first RNA molecule includes a coding region and a 3' ribozyme, wherein the 3' ribozyme is capable of catalyzing its own separation from the RNA molecule, leaving behind a coding region with a 3'P or 2'3' cyclic phosphate (cP) terminus. In one embodiment, the 3' ribozyme comprises HDV ribozyme. Furthermore, in one embodiment, the composition comprises a nucleic acid molecule encoding a second RNA molecule, wherein the second RNA molecule includes a coding region and a 5' ribozyme, wherein the 5' ribozyme is capable of catalyzing its own separation from the RNA molecule, leaving behind a coding region with a 5'OH terminus. In one embodiment, the 5' ribozyme comprises HH ribozyme. In some cases, a ligase joins the coding region of the first RNA molecule with the coding region of the second RNA molecule to form a longer RNA molecule encoding the target protein.

[0176] For example, in one embodiment, the composition comprises a first RNA molecule containing a coding region and a 3' ribozyme, wherein the 3' ribozyme is capable of catalyzing its own separation from the RNA molecule, leaving behind a coding region with a 3'P or 2'3' cyclic phosphate (cP) terminus. In one embodiment, the 3' ribozyme comprises HDV ribozyme. Furthermore, in one embodiment, the composition comprises a second RNA molecule containing a coding region and a 5' ribozyme, wherein the 5' ribozyme is capable of catalyzing its own separation from the RNA molecule, leaving behind a coding region with a 5'OH terminus. In one embodiment, the 5' ribozyme comprises HH ribozyme. In some cases, a ligase links the coding region of the first RNA molecule to the coding region of the second RNA molecule, forming a longer RNA molecule encoding the target protein.

[0177] In some embodiments, the first RNA contains a coding region encoding a first portion of the target protein, and the second RNA contains a coding region encoding a second portion of the target protein. Therefore, ribozyme-mediated cleavage and ligase-mediated assembly of the RNA molecules result in the production of an RNA molecule encoding a protein having both the first and second portions. This invention can be used to generate full-length proteins from multiple RNAs, each RNA containing a coding region encoding a portion of the full-length protein. Furthermore, this invention can be used to generate fusion proteins comprising multiple domains, wherein each RNA molecule contains a coding region encoding a fusion protein domain. For example, this invention can be used to generate RNA molecules encoding proteins having a leader sequence, an N-terminal tag, a C-terminal tag, etc., by assembling RNA from a first RNA molecule containing a coding sequence encoding a leader sequence, an N-terminal tag, or a C-terminal tag, and a second RNA molecule containing a coding sequence encoding the protein.

[0178] In some embodiments, the present invention relates to a single RNA molecule formed from three or more separate RNA molecules. For example, in some aspects, the composition comprises a nucleic acid molecule encoding a first RNA molecule, wherein the first RNA molecule includes a coding region encoding an N-terminal region of a protein; a nucleic acid molecule encoding a second RNA molecule, wherein the second RNA molecule includes a coding region encoding a C-terminal region of a protein; and one or more nucleic acid molecules encoding one or more additional RNA molecules, each additional RNA molecule including a coding region encoding a protein domain (e.g., a repeating domain). In one embodiment, the first RNA molecule includes a coding region encoding the N-terminal region and a 3' ribozyme, wherein the 3' ribozyme is capable of catalyzing its own separation from the RNA molecule, leaving a coding region with a 3'P or 2'3' cyclic phosphate (cP) end. In one embodiment, the 3' ribozyme comprises HDV ribozyme. In one embodiment, the second RNA molecule includes a coding region encoding the C-terminal region and a 5' ribozyme, wherein the 5' ribozyme is capable of catalyzing its own separation from the RNA molecule, leaving a coding region with a 5'OH end. In one embodiment, the 5' ribozyme comprises HH ribozyme. In one embodiment, each additional RNA molecule comprises a coding region encoding a protein domain, a 3' ribozyme, and a 5' ribozyme. In one embodiment, the 3' ribozyme is an HDV ribozyme. In one embodiment, the 5' ribozyme is an HH ribozyme. In some aspects, the 3' ribozyme is capable of catalyzing its own separation from the RNA molecule, and the 5' ribozyme is capable of catalyzing its own separation from the RNA molecule, leaving a coding region with a 5'OH and a 3'P or 2'3' cP terminus. In one embodiment, each additional RNA molecule comprises a coding region encoding a protein domain, a 5' ribozyme, and a 3' ribozyme recognition sequence. In some aspects, the 5' ribozyme is capable of catalyzing its own separation from the RNA molecule, leaving a coding region with a 5'OH terminus; and the 3' ribozyme recognition sequence interacts with the ribozyme, inducing the 3' ribozyme recognition sequence to splice out of the RNA molecule, leaving a coding region with a 3'P or 2'3' cP terminus. In one embodiment, the 3' ribozyme recognition sequence comprises a Vsv1 sequence that interacts with the VS ribozyme. This technique can be used to generate RNA molecules encoding proteins with multiple repeating domains by sequentially adding a coding region encoding a repeating domain by sequentially providing a ribozyme (e.g., VS ribozyme) to interact with the 3' ribozyme recognition sequence to generate a 3'P or 2'3' cP terminus and then linking the coding region to the 5' OH terminus of another coding region encoding a repeating domain. In some aspects, the sequential addition of repeating domains can be performed on a solid matrix or support, wherein the first RNA molecule encoding the N-terminal region is bound to the matrix or support.

[0179] In some respects, multiple RNA molecules are linked together after ribozyme-mediated generation of 5'OH and 3'P or 2'3'cP ends. In other cases, RNA molecules are linked together by endogenous ligases present in the natural cells or tissues where RNA assembly takes place. In some cases, the method of the present invention includes the step of adding an exogenous ligase to induce the processed RNA molecules to link together. In one embodiment, the ligase is an RNA 2',3'-cyclic phosphate and 5'-hydroxyl (RtcB) ligase.

[0180] In some aspects, the present invention relates to the use of deoxyribozymes, enzymes, or DNases to cleave single-stranded DNA sequences for trans splicing or trans editing. For example, deoxyribozymes (which self-cleave DNA sequences leaving 3'-P (or 2'3'-cP) and 5'-OH ends), enzymes that cleave and leave the same ends, or DNases (which can be fused with DNA-targeting proteins such as TALEN or CRISPR) can be used for trans splicing of single-stranded DNA substrates or for trans editing using trans-cleavage of deoxyribozyme sequences. RTCBs can be used to act on single-stranded DNA substrates with 3'-P (or 2'3'-cP) and 5'-OH ends produced by deoxyribozymes, enzymes, or DNases.

[0181] Composition

[0182] In one embodiment, the present invention relates to a composition comprising one or more nucleic acid molecules encoding one or more ribozymes. In one embodiment, the invention comprises one or more RNA molecules comprising one or more ribozymes. In some embodiments, the one or more RNA molecules comprise at least a first RNA molecule and a second RNA molecule.

[0183] In some embodiments, the one or more ribozymes in the composition are capable of spontaneously cis-cleaving the one or more RNA molecules. In some embodiments, the one or more ribozymes are 3' ribozymes. In some embodiments, the 3' ribozyme generates a 3'P or 2'3' cP terminus on the remaining one or more RNA molecules after spontaneous cis-cleavage. In some embodiments, the one or more ribozymes are 5' ribozymes. In some embodiments, the 5' ribozyme generates a 5'OH terminus on the remaining one or more RNA molecules after spontaneous cis-cleavage. In some embodiments, the 3'P or 2'3' cP terminus and the 5'OH terminus may be linked together.

[0184] In some embodiments, the first RNA molecule comprises a 3' ribozyme. In some embodiments, the 3' ribozyme is selected from one or more families of: hammerhead (HH) ribozymes, hepatitis D virus (HDV) ribozymes, Varkud Satellite (VS) ribozymes, Sister ribozymes, Twister-sister ribozymes, Hairpin ribozymes, Hatchet ribozymes, Pistol ribozymes, HOV Linc ribozymes, or variants or fragments thereof that retain cis-cleaving function. In one embodiment, the 3' ribozyme includes lantern ribozymes (Zhou et al. Human Lantern Ribozymes: Smallest Known Self-cleaving Ribozymes, 07 March 2023, PREPRINT (Version 1), available at Research Square [doi.org / 10.21203 / rs.3.rs-2567304 / v1]). However, the invention is not limited to any particular ribozyme, but covers members of all known endogenous ribozyme families and any potential artificial ribozymes. In one embodiment, the 3' ribozyme includes a P1 type Twister, a P3 type Twister, or a P5 type Twister. In one embodiment, the 3' ribozyme includes a P1 type Twister. In one embodiment, the 3' ribozyme includes a P1 type Twister from rice (Oryza sativa, Osa). In some embodiments, the 3' ribozyme comprises an overhang of one or more nucleotides. In one embodiment, the overhang comprises a nucleotide sequence that hybridizes to the upstream sequence of the 3' ribozyme in the first RNA molecule. In some embodiments, the overhang improves the efficiency of spontaneous cis-cleavage.

[0185] In some embodiments, the second RNA molecule comprises a 5' ribozyme. In some embodiments, the 5' ribozyme is selected from one or more families of: hammerhead (HH) ribozymes, hepatitis D virus (HDV) ribozymes, VarkudSatellite (VS) ribozymes, Sister ribozymes, Twister-sister ribozymes, Hairpin ribozymes, Hatchede ribozymes, Pistol ribozymes, HOV Linc ribozymes, or variants or fragments thereof that retain cis-cleaving function. In one embodiment, the 5' ribozyme includes lantern ribozymes (Zhou et al. Human Lantern Ribozymes: Smallest Known Self-cleaving Ribozymes, 07 March 2023, PREPRINT (Version 1), available at Research Square [doi.org / 10.21203 / rs.3.rs-2567304 / v1]). However, the invention is not limited to any particular ribozyme, but covers members of all known endogenous ribozyme families and any potential artificial ribozymes. In one embodiment, the 5' ribozyme includes a P1-type Twister, a P3-type Twister, or a P5-type Twister. In one embodiment, the 5' ribozyme includes a P1-type Twister. In one embodiment, the 5' ribozyme includes a P1-type Twister from rice (Oryza sativa, Osa). In some embodiments, the 5' ribozyme comprises a pendant of one or more nucleotides. In one embodiment, the pendant comprises a nucleotide sequence that hybridizes to a downstream sequence of the 5' ribozyme in a second RNA molecule. In some embodiments, the pendant improves the efficiency of spontaneous cis-cleavage.

[0186] In one embodiment, the HDV ribozyme in the composition comprises one or more selected from the group consisting of HDV, HDV68, HDV67, HDV56, genHDV, and antiHDV, or variants or fragments thereof. In one embodiment, HDV68 comprises the nucleic acid sequence of SEQ ID NO:9. In one embodiment, HDV67 comprises the nucleic acid sequence of SEQ ID NO:10. In one embodiment, HDV56 comprises the nucleic acid sequence of SEQ ID NO:11. In one embodiment, genHDV comprises the nucleic acid sequence of SEQ ID NO:12. In one embodiment, antiHDV comprises the nucleic acid sequence of SEQ ID NO:13.

[0187] In one embodiment, the HH ribozyme comprises one or more nucleotides in the stem region 1 pendant, the nucleotides hybridizing with nucleotides upstream or downstream of the sequence of the HH ribozyme. In one embodiment, the number of nucleotides in the stem region 1 pendant can be one or more nucleotides, two or more nucleotides, four or more nucleotides, six or more nucleotides, eight or more nucleotides, ten or more nucleotides, twelve or more nucleotides, fourteen or more nucleotides, sixteen or more nucleotides, eighteen or more nucleotides, or twenty or more nucleotides. In one embodiment, the HH ribozyme in the stem region 1 pendant containing one or more nucleotides comprises a nucleic acid sequence selected from the following: SEQ ID NO:111, SEQ ID NO:112, SEQ ID NO:113, SEQ ID NO:114, SEQ ID NO:115, SEQ ID NO:116, SEQ ID NO:117, and SEQ ID NO:118, wherein the nucleotide labeled N corresponds to the nucleotide hybridizing with the nucleotide downstream of the sequence of the HH ribozyme. In one embodiment, the HH ribozyme has one or more nucleotides in the stem region 3 pendant. In one embodiment, the HH ribozyme has five nucleotides in the stem region 3 pendant. In one embodiment, the HH ribozyme comprises the nucleic acid sequence of SEQ ID NO: 105, wherein the nucleotide labeled N corresponds to a nucleotide that hybridizes with a nucleotide upstream of the sequence of the HH ribozyme. In one embodiment, the HH ribozyme is modified in the stem region 2 loop. In one embodiment, the HH ribozyme having the modified stem region 2 loop comprises a nucleic acid sequence selected from: SEQ ID NO: 119, SEQ ID NO: 120, SEQ ID NO: 121, SEQ ID NO: 122, SEQ ID NO: 123 and SEQ ID NO: 124, wherein the nucleotide labeled N corresponds to a nucleotide that hybridizes with a nucleotide downstream of the sequence of the HH ribozyme. In one embodiment, the HH ribozyme is modified in the stem region 1 to include a tertiary stable motif (TSM). In one embodiment, the HH ribozyme is modified in the stem region 2 loop and in the stem region 1 to include a tertiary stable motif (TSM). In one embodiment, the modified HH ribozyme performs cis-cleavage more efficiently than the HH ribozyme. In one embodiment, the modified HH ribozyme is RzB. In one embodiment, RzB comprises the nucleic acid sequence of SEQ ID NO:125, wherein the nucleotide marked N corresponds to a nucleotide that hybridizes to a nucleotide downstream of the sequence of the HH ribozyme.

[0188] In one embodiment, the Twister ribozyme comprises the nucleic acid sequence of SEQ ID NO:32. In one embodiment, the Twister ribozyme comprises one or more nucleotides in the P1 stem pendant. In one embodiment, the number of nucleotides in the P1 stem pendant can be one or more, two or more, three or more, four or more, or five or more. In one embodiment, the Twister ribozyme of the P1 stem pendant comprising one or more nucleotides comprises a nucleic acid sequence selected from the following: SEQ ID NO:106, SEQ ID NO:107, SEQ ID NO:108, SEQ ID NO:109, and SEQ ID NO:110, wherein the nucleotide labeled N corresponds to a nucleotide that hybridizes to a nucleotide downstream of the sequence of the Twister ribozyme.

[0189] In some embodiments, one or more ribozymes in the composition consist of a first portion and a second portion. In some embodiments, the first portion is integrated into the one or more RNA molecules. In some embodiments, the first portion is a ribozyme recognition sequence. In some embodiments, the second portion is introduced separately. In some embodiments, cis-cleavage of the first portion occurs from the one or more RNA molecules only when the first and second portions come into contact with each other. In some embodiments, the one or more ribozymes are VS ribozymes. In one embodiment, the VS ribozyme comprises the nucleic acid sequence of SEQ ID NO:14. In one embodiment, the first portion is a VS ribozyme stem-loop (VS-S). In one embodiment, VS-S comprises the nucleic acid sequence of SEQ ID NO:15. In one embodiment, the second portion is the remaining portion of the VS without the stem-loop (VS-Rz). In one embodiment, VS-Rz comprises the nucleic acid sequence of SEQ ID NO:16.

[0190] Ribozymes are autocatalytic RNAs that cis-cleave, as described herein, to produce unique 3' and 5' ends of RNA. However, cis-cleaving ribozymes can be engineered to trans-cleave, allowing target RNA to be cleaved in a nucleotide-specific manner, producing similar RNA ends. In some embodiments, the invention comprises a composition comprising a single nucleic acid molecule encoding a single RNA molecule containing a trans-cleaving engineered ribozyme. In one embodiment, the trans-cleaving engineered ribozyme is capable of trans-cleaving an individual RNA molecule. In one embodiment, the trans-cleaving engineered ribozyme recognizes a specific nucleic acid sequence in the individual RNA molecule. In some embodiments, the trans-cleaving engineered ribozyme targets a pathogenic mutation for deletion. In some embodiments, the pathogenic mutation is located in an exon. In some embodiments, the pathogenic mutation is located in an intron. In some embodiments, the composition comprises two trans-cleaving engineered ribozymes targeting upstream and downstream of a pathogenic mutation. In some embodiments, trans-cleavage upstream and downstream of the pathogenic mutation results in the removal of the pathogenic mutation. In some embodiments, after the pathogenic mutation is trans-spliced, the remaining portion of the gene is trans-spliced ​​together. In some embodiments, the trans-spliced ​​gene is expressed as a functional protein.

[0191] As described herein, the 3'P or 2'3'cP ends and the 5'OH ends of ribozyme-mediated cleavage RNA molecules can be joined together. This allows separate RNA sequences encoding different portions of a larger full-length protein to be trans-spliced ​​together without scarring, thereby enabling the expression of the full-length protein. In one embodiment, the present invention relates to a composition comprising two or more portions encoding a target protein and one or more nucleic acid molecules encoding one or more ribozymes. In another embodiment, the present invention relates to a composition comprising two or more portions encoding a target protein and one or more RNA molecules containing one or more ribozymes.

[0192] In one embodiment, the nucleic acid molecule encoding two or more portions of the target protein comprises a first nucleic acid molecule encoding a first portion of the target protein and a second nucleic acid molecule encoding a second portion of the target protein. In one embodiment, the first nucleic acid comprises a first RNA molecule. In one embodiment, the second nucleic acid comprises a second RNA molecule. In one embodiment, the first RNA molecule is linked to a 3' ribozyme at its 3' end. In one embodiment, the second RNA molecule is linked to a 5' ribozyme at its 5' end. In one embodiment, after cis-cleavage of the 3' and 5' ribozyme sequences, the 3'P or 2'3'cP end of the first RNA molecule is linked to the 5'OH end of the second RNA molecule, thereby generating a single RNA molecule encoding the full-length target protein. In one embodiment, the full-length target protein functions identically to a full-length protein of the same sequence expressed endogenously.

[0193] In one embodiment, the full-length target protein includes a therapeutic protein. In one embodiment, the therapeutic protein comprises one or more selected from the following: dystrophin-associated protein (Utrophin), dystrophin, Dysferlin, Myoferlin, cystic fibrosis transmembrane transport regulator (CFTR), coagulation factor VIII, Fibrocystin, retinal-specific phospholipid transporter ATPase (ABCA4), Otoferlin, copper ATP transporter 2, MYO7A, MYO15A, CDH23, STRC, OTOG, TEC. The target protein may contain, but is not limited to, TA, PCDH15, TRIOBP, MYO3A, COL11A2, LOXHD1, PTPRQ, OTOGL, MYH14, MYH9, TNC, CACNA1A, CACNA1C, CACNA1F, CACNA1H, CACNA1G, CACNA1D, CACNA1B, CACNA1S, CACNA1I, CACNA1E, ATP2A1, ATP2A2, Adcy6, FKBP12-rapamycin binding domain, and Cas9. In one embodiment, the full-length target protein is a recombinase. In one embodiment, the recombinase is selected from one or more of the following: CRE recombinase, FLP recombinase, but is not limited to. In one embodiment, the full-length target protein is a product of a eukaryotic / prokaryotic antibiotic resistance gene. In one embodiment, the eukaryotic / prokaryotic antibiotic resistance gene product is selected from one or more of the following: ampicillin, kanamycin, blasticidin, puromycin, neomycin, and hygromycin, but is not limited thereto. In some embodiments, the full-length target protein is an antibody. In one embodiment, the antibody is capable of binding to the target protein. In some embodiments, the antibody is an antibody fragment, a synthetic antibody, a nanobody, or a fragment or variant thereof that retains the ability to bind to the target protein. In one embodiment, the full-length target protein comprises, but is not limited to, synthetic repeating proteins, including, proteins constituting a hydrogel, synthetic spider silk, and collagen. In one embodiment, the synthetic repeating protein comprises, but is not limited to, one or more of the following: spider silk protein, silk fibroin, keratin, collagen, elastin, arthropod elastin (Resilin), squid ring teeth, beta-solenoid protein, zinc finger nuclease (ZFN), and Tal effector nuclease (TALEN). In one embodiment, the full-length target protein includes a toxic or antiviral protein that inhibits the production of lentiviral particles in mammalian packaging cells. In one embodiment, the toxic protein is a cell suicide gene.In one embodiment, the cell suicide gene includes, but is not limited to, one or more of the following: diphtheria toxin A (DTA), HSV-tk, ricin, cholera toxin, major prion protein, pertussis toxin, ectatomin, conopeptides, aburin, verotoxin, tetanospasmin, botulinum toxin, Pseudomonas exotoxin A, anthraxin, saporin, and pokeweed antiviral protein (PAP). In one embodiment, the antiviral protein includes, but is not limited to, one or more of the following: interferon-induced GTP-binding protein (MxA), myeloperoxidase (MPO), and interferon.

[0194] N-terminal or C-terminal RNA molecules encoding a portion of a target protein may be translated prior to ribozyme-mediated splicing, or, when expressed alone, may result in unwanted or truncated protein expression. However, such unwanted expression can be limited by translation control of translation control protein degradation sequences. In one embodiment, the one or more RNA molecules in the composition comprise a nucleic acid sequence encoding a translation control protein degradation sequence. In one embodiment, the first RNA molecule comprises a nucleic acid sequence encoding a translation control protein degradation sequence. In one embodiment, the second RNA molecule comprises a nucleic acid sequence encoding a translation control protein degradation sequence. In some embodiments, the translation control protein degradation sequence prevents partial protein expression prior to ribozyme sequence splicing and cleavage. In some embodiments, the translation control protein degradation sequence comprises one or more selected from: hCL1-PEST sequences, E1A-PEST sequences, poly(A) sequences with nucleic acid removed, polyK-tails generated by polyA-tailed translation, deletion of an ATG stop codon, silencing mutations within the N-terminal NTG codon, the 5'UTR of a yeast GCN4 sequence encoding four small upstream ORFs that function as translation inhibitors, and small internal fragments of the 5'UTR of the yeast GCN4 sequence. In some embodiments, the translation control protein degradation sequence comprises one or more nucleic acid sequences selected from the following: SEQ ID NO:43, SEQ ID NO:44, SEQ ID NO:45, SEQ ID NO:46, SEQ ID NO:47, SEQ ID NO:48, SEQ ID NO:49, SEQ ID NO:77, SEQ ID NO:79 and SEQ ID NO:104. In some embodiments, the translation control protein degradation sequence comprises one or more amino acid sequences selected from the following: SEQ ID NO:52, SEQ ID NO:53, SEQ ID NO:54, SEQ ID NO:55, SEQ ID NO:56, SEQ ID NO:57, SEQ ID NO:58, SEQ ID NO:59, SEQ ID NO:60, SEQ ID NO:61, SEQ ID NO:62, SEQ ID NO:63, SEQ ID NO:64, SEQ ID NO:65, SEQ ID NO:66, SEQ ID NO:67, SEQ ID NO:68, SEQ ID NO:69, SEQ ID NO:70, SEQ ID NO:71, SEQ ID NO:72, SEQ ID NO:73, SEQ ID NO:74, SEQ ID NO:76, SEQ ID NO:78, and SEQ ID NO:80.

[0195] In some aspects, to further prevent unwanted or truncated protein expression, RNA nuclear localization signals can be used to block the cytoplasmic export and translation of unspliced ​​RNA molecules. In one embodiment, the one or more RNA molecules in the composition comprise a nucleic acid sequence encoding an RNA nuclear localization sequence. In one embodiment, the first RNA molecule comprises a nucleic acid sequence encoding an RNA nuclear localization sequence. In one embodiment, the second RNA molecule comprises a nucleic acid sequence encoding an RNA nuclear localization sequence. In one embodiment, the RNA nuclear localization sequence blocks the cytoplasmic RNA export and translation of a portion of the protein prior to ribozyme cleavage and splicing. In one embodiment, the RNA nuclear localization sequence comprises one or more nucleic acid sequences selected from SEQ ID NO:50 and SEQ ID NO:51.

[0196] In some embodiments, the composition further comprises one or more additional RNA molecules, each additional RNA molecule containing a coding region encoding a target protein domain; a 5' ribozyme; and a 3' ribozyme. In some embodiments, the system further comprises one or more additional nucleic acid molecules encoding one or more additional RNA molecules, each additional RNA molecule containing a coding region encoding a target protein domain; a 5' ribozyme; and a 3' ribozyme.

[0197] In some embodiments, the composition further comprises one or more additional RNA molecules, each additional RNA molecule containing a coding region encoding a target protein domain; a 5' ribozyme; and a 3' ribozyme recognition sequence. In some embodiments, the system further comprises one or more additional nucleic acid molecules encoding one or more additional RNA molecules, each additional RNA molecule containing a coding region encoding a target protein domain; a 5' ribozyme; and a 3' ribozyme recognition sequence.

[0198] Pre-mRNA splicing by the spliceosome has been shown to enhance mRNA translation by promoting the deposition of translation factors in the leader ring or by promoting RNA processing and export to the cytoplasm. The addition of chimeric cis-splicing introns to transgenes has also been shown to promote transgenic protein expression. Therefore, in some embodiments, the addition of spliceosome-recognizing and cis-splicing splice donor and splice acceptor sites can enhance protein expression of pre-splicing RNA molecules. In one embodiment, the composition comprises one or more RNA molecules containing a splice donor or splice acceptor sequence. In one embodiment, the first RNA molecule in the composition contains a splice donor sequence. In one embodiment, the splice donor sequence is attached to the 3' end of the first RNA molecule, following the ribozyme sequence. In one embodiment, the second RNA molecule in the composition contains a splice acceptor sequence. In one embodiment, the splice acceptor sequence is attached to the 5' end of the second RNA molecule, preceding the ribozyme sequence. In one embodiment, the addition of splice donor and splice acceptor sequences enhances ribozyme-mediated trans-splicing protein expression.

[0199] Because there are three open reading frames (ORFs) for protein translation, it is possible to simultaneously perform ribozyme-mediated trans-splicing and expression of multiple different functional proteins. Utilizing this characteristic, functional proteins can be generated by back-splicing RNA from three different incompatible ORFs. In one embodiment, the composition of the present invention comprises at least four nucleic acid molecules, which comprise at least two pairs of nucleic acid molecules. In one embodiment, each pair of nucleic acid molecules encodes at least two portions of the target protein and encodes at least two ribozymes. In one embodiment, the composition comprises at least four RNA molecules, which comprise at least two pairs of RNA molecules. In one embodiment, each pair of RNA molecules encodes at least two portions of the target protein and contains at least two ribozymes.

[0200] In one embodiment, the at least two pairs of RNA molecules include a first pair of RNA molecules and a second pair of RNA molecules. In one embodiment, the first pair of RNA molecules includes a first RNA molecule and a second RNA molecule. In one embodiment, the second pair of RNA molecules includes a third RNA molecule and a fourth RNA molecule. In some embodiments, the third RNA molecule and the fourth RNA molecule have different open reading frames than the first RNA molecule and the second RNA molecule, such that during spontaneous cis-secretion, the linking of the first RNA molecule or the second RNA molecule with the third RNA molecule or the fourth RNA molecule cannot be translated into a full-length functional protein product.

[0201] In one embodiment, the at least two pairs of RNA molecules further include a third pair of RNA molecules. In one embodiment, the third pair of RNA molecules includes a fifth RNA molecule and a sixth RNA molecule. In some embodiments, the fifth RNA molecule and the sixth RNA molecule have open reading frames different from those of the first and second pairs of RNA molecules, such that during spontaneous cis-cleavage, only the ligation of the first, second, or third pair of RNA molecules can translate into a full-length functional protein product.

[0202] As described herein, when one RNA contains a 3' ribozyme and the other a 5' ribozyme, ribozyme-mediated trans splicing can occur between two independent RNAs. However, during cis transcription within the same RNA molecule, both ribozymes can mediate their own scarless removal. This approach similarly produces two independent RNAs with 3'-P and 5'-OH ends, which can undergo trans splicing and translation in the cell. Inserting a cargo sequence between the 3' and 5' ribozymes also makes it possible to generate circular RNA molecules upon ligation.

[0203] In one embodiment, the present invention relates to a composition comprising two or more portions encoding a target protein and a single nucleic acid molecule encoding one or more ribozymes. In another embodiment, the present invention relates to a composition comprising two or more portions encoding a target protein and a single RNA molecule comprising one or more ribozymes.

[0204] In one embodiment, the single nucleic acid molecule encodes a first-part RNA, a synthetic intron, and a second-part RNA. In one embodiment, the synthetic intron comprises a 5' ribozyme and a 3' ribozyme. In one embodiment, the first-part RNA encodes a first portion of the target protein. In one embodiment, the second-part RNA encodes a second portion of the target protein. In one embodiment, the single nucleic acid comprises a sequence linked in the following order: (first-part RNA encoding the first portion of the target protein) - (5' ribozyme of the synthetic intron) - (3' ribozyme of the synthetic intron) - (second-part RNA encoding the second portion of the target protein). In one embodiment, the first portion of the target protein is the N-terminal portion of GFP. In one embodiment, the 5' ribozyme of the synthetic intron comprises HDV. In one embodiment, the first-part RNA and the 5' ribozyme of the synthetic intron comprise the nucleic acid sequence of SEQ ID NO:127, wherein lowercase letters represent the 5' ribozyme sequence and uppercase letters represent the sequence encoding the N-terminal portion of GFP (see Example 4, "GFP with or without internal synthetic ribozyme introns"). In one embodiment, the second portion of the target protein is the C-terminal portion of GFP. In one embodiment, the 3' ribozyme of the synthetic intron comprises HH. In one embodiment, the second RNA portion and the 3' ribozyme of the synthetic intron comprise the nucleic acid sequence of SEQ ID NO:128, wherein lowercase letters represent the 3' ribozyme sequence and uppercase letters represent the sequence encoding the C-terminal portion of GFP. (See Example 4, "GFP with or without internal synthetic ribozyme introns").

[0205] In one embodiment, the synthetic intron comprises a cargo sequence located between the 5' ribozyme and the 3' ribozyme. In one embodiment, the single nucleic acid comprises sequences linked in the following order: (first part RNA encoding the first part of the target protein) - (5' ribozyme of the synthetic intron) - (cargo sequence) - (3' ribozyme of the synthetic intron) - (second part RNA encoding the second part of the target protein).

[0206] In one embodiment, the 5' ribozyme sequence of the synthetic intron is active without the need for flanking sequences. In one embodiment, the circular RNA generated by ligating the ends of a synthetic intron containing a 5' ribozyme sequence that is active without the need for flanking sequences can exist in a circular or re-spliced ​​linear form. In one embodiment, the ribozyme sequence is an HDV ribozyme.

[0207] In one embodiment, the 5' ribozyme sequence of the synthetic intron does indeed require bilateral flanking sequences to be active. In one embodiment, the circular RNA generated by ligating the ends of a synthetic intron containing a 5' ribozyme sequence that does indeed require bilateral flanking sequences to be active exists only in a circular form. In one embodiment, the ribozyme sequence is an HH ribozyme.

[0208] In one embodiment, the 5' ribozyme sequence of the synthesized intron is a ribozyme recognition sequence. In one embodiment, this ribozyme recognition sequence requires the addition of a trans-cleaving ribozyme for inducible cleavage. In one embodiment, the ribozyme recognition sequence comprises VS-S. In some embodiments, VS-S is encoded by a nucleic acid sequence comprising SEQ ID NO:15. In one embodiment, the trans-cleaving ribozyme comprises VS-Rz. In some embodiments, VS-Rz is encoded by a nucleic acid sequence comprising SEQ ID NO:16.

[0209] In one embodiment, self-cleavage of the 5' and 3' ribozyme sequences yields three separate RNA molecules: 1) a first fragment containing a first part of RNA encoding a first portion of the target protein; 2) a second fragment containing a synthetic intron; and 3) a third fragment containing a second part of RNA encoding a second portion of the target protein. In one embodiment, the compatible ends of the second fragment are joined together to produce a circular RNA molecule containing a synthetic intron that contains a cargo sequence. In one embodiment, the first and third fragments are joined together to produce a single full-length linear RNA molecule.

[0210] In one embodiment, the cargo sequence of the synthesized intron is selected from one or more of the following: sequences encoding a target therapeutic protein, CRISPR guide RNA sequences, small RNA sequences, and trans-cleavage ribozyme sequences. In one embodiment, the small RNA sequence includes one or more of the following: microRNA (miRNA), Piwi-interacting RNA (piRNA), small interfering RNA (siRNA), small nucleolar RNA (snoRNA), small tRNA-derived RNA (tsRNA), small rDNA-derived RNA (srRNA), and small nuclear RNA (snRNA).

[0211] In one embodiment, the single full-length linear RNA molecule encodes a full-length target protein. In one embodiment, the full-length target protein is a therapeutic protein. In one embodiment, the therapeutic protein may be selected from one or more of the following: dystrophin, dystrophin, dysferlin, myoferlin, cystic fibrosis transmembrane transport regulator (CFTR), coagulation factor VIII, Fibrocystin, retinal-specific phospholipid transporter ATPase (ABCA4), Otoferlin, copper ATP transporter 2, MYO7A, MYO15A, CDH23, STRC, OTOG, TE. CTA, PCDH15, TRIOBP, MYO3A, COL11A2, LOXHD1, PTPRQ, OTOGL, MYH14, MYH9, TNC, CACNA1A, CACNA1C, CACNA1F, CACNA1H, CACNA1G, CACNA1D, CACNA1B, CACNA1S, CACNA1I, CACNA1E, ATP2A1, ATP2A2, Adcy6, FKBP12-rapamycin binding domain, and Cas9, but not limited thereto. In one embodiment, the full-length target protein is a recombinase. In one embodiment, the recombinase is selected from one or more of the following: CRE recombinase, FLP recombinase, but not limited thereto. In one embodiment, the full-length target protein is a eukaryotic / prokaryotic antibiotic resistance gene product. In one embodiment, the eukaryotic / prokaryotic antibiotic resistance gene product is selected from one or more of the following: ampicillin, kanamycin, blasticidin, puromycin, neomycin, and hygromycin. In one embodiment, the full-length target protein is a reporter protein, but is not limited thereto. In one embodiment, the reporter protein is selected from one or more of the following: green fluorescent protein (GFP), red fluorescent protein (RFP), and luciferase (Luc), but is not limited thereto. In one embodiment, the reporter protein is used as an alternative indicator for assessing cargo sequence delivery and expression. In some embodiments, the full-length target protein is an antibody. In one embodiment, the antibody is capable of binding to the target protein. In some embodiments, the antibody is an antibody fragment, a synthetic antibody, a nanobody, or a fragment or variant thereof that retains the ability to bind to the target protein. In one embodiment, the full-length target protein comprises a toxic or antiviral protein that can inhibit the production of lentiviral particles in mammalian packaging cells. In one embodiment, the toxic protein is a cell suicide gene.In one embodiment, the cell suicide gene comprises one or more of the following: diphtheria toxin A (DTA), HSV-tk, ricin, cholera toxin, major prion protein, pertussis toxin, exotoxin, conopodyl peptide, abrinogen toxin, shiga toxin, tetanus toxin, botulinum toxin, Pseudomonas exotoxin A, anthrax toxin, saponins, and pokeweed antiviral protein (PAP), but is not limited thereto. In one embodiment, the antiviral protein comprises one or more of the following: interferon-induced GTP-binding protein (MxA), myeloperoxidase (MPO), and interferon, but is not limited thereto.

[0212] In some aspects, the techniques of the present invention can be used to assemble full-length RNA virus genomes. In one embodiment, the one or more nucleic acid molecules encoding one or more ribozymes of the present invention encode one or more portions of an RNA virus genome. In one embodiment, the RNA molecule containing one or more ribozymes of the present invention contains one or more portions of an RNA virus genome.

[0213] In one embodiment, the one or more nucleic acid molecules comprise a first nucleic acid molecule encoding a first portion of an RNA virus genome and encoding a 3' ribozyme. In one embodiment, the one or more nucleic acid molecules comprise a second nucleic acid encoding a second portion of an RNA virus genome and encoding a 5' ribozyme. In one embodiment, the one or more RNA molecules comprise a first RNA molecule containing a first portion of an RNA virus genome and a 3' ribozyme. In one embodiment, the one or more RNA molecules comprise a second RNA molecule containing a second portion of an RNA virus genome and a 5' ribozyme. In one embodiment, the composition comprises a nucleic acid or a ligase encoding a ligase. In one embodiment, upon cis-cleavage of the 3' and 5' ribozymes, the first portion and the second portion of the RNA virus genome are joined together to produce a full-length RNA virus genome. Exemplary RNA viruses include, but are not limited to: coronaviruses, paramyxoviruses, orthomyxoviruses, retroviruses, lentiviruses, alphaviruses, flaviviruses, rhabdoviruses, measlesviruses, Newcastle disease viruses, and picornaviruses.

[0214] In some embodiments, the present invention comprises a composition containing a nucleic acid encoding a ligase. In some embodiments, the ligase mediates the ligation of a 3'P or 2'3' cP terminus to a 5'OH terminus. In some embodiments, the ligase is an RNA 2',3'-cyclic phosphate and 5'-OH (RtcB) ligase. In some embodiments, the RtcB ligase is derived from one or more domains of an organism selected from eukaryotes, bacteria, and archaea. In some embodiments, the organism is selected from: humans, *Escherichia coli*, *Deinococcus radiodurans*, *Pyrococcus horikoshii*, *Pyrococcus sp.* ST04, and *Thermococcus sp.* EP. In some embodiments, the nucleic acid sequence encoding the ligase is selected from one or more of the following: SEQ ID NO:82, SEQ ID NO:84, SEQ ID NO:86, SEQ ID NO:88, SEQ ID NO:90, SEQ ID NO:92. In some embodiments, the nucleic acid sequence encoding the ligase encodes one or more of the following amino acid sequences: SEQ ID NO:81, SEQ ID NO:83, SEQ ID NO:85, SEQ ID NO:87, SEQ ID NO:89, SEQ ID NO:91.

[0215] Nucleic acid

[0216] In some embodiments, one or more nucleic acids of the present invention comprise nucleic acid sequences substantially homologous to the nucleic acid sequences described herein. For example, in some embodiments, the degree of identity between the nucleic acid and the original nucleic acid sequence is at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5%.

[0217] In some embodiments, one or more nucleic acids of the present invention comprise a nucleic acid sequence that is part of the nucleic acid sequence described herein. For example, in some embodiments, the nucleic acid is at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% of the length of the original nucleic acid sequence.

[0218] In some embodiments, one or more nucleic acids of the present invention comprise a nucleic acid sequence that is part of the nucleic acid sequence described herein and is substantially homologous to the nucleic acid sequence described herein. For example, in some embodiments, the nucleic acid has an identity degree with the original nucleic acid sequence of at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99%. Or at least 99.5%, and / or at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% relative to the length of the original nucleic acid sequence.

[0219] The nucleic acids of the present invention can comprise any type of nucleic acid, including but not limited to DNA and RNA. For example, in one embodiment, the composition comprises a separated DNA molecule encoding the fusion protein of the present invention, such as a separated cDNA molecule. In one embodiment, the composition comprises a separated RNA molecule encoding the fusion protein of the present invention or a functional fragment thereof.

[0220] The nucleic acid molecules of the present invention can be modified to improve their stability in serum or cell culture growth media. Modifications can be added to enhance the stability, functionality, and / or specificity of the nucleic acid molecules of the present invention and to minimize their immunostimulatory properties. For example, to enhance stability, the 3'-residues can be stabilized to prevent degradation; for instance, they can be selected to consist of purine nucleotides, particularly adenosine or guanosine nucleotides. Alternatively, replacing pyrimidine nucleotides with modified analogs, such as replacing uridine with 2'-deoxythymidine, is tolerable and does not affect the function of the molecule.

[0221] In one embodiment of the invention, the nucleic acid molecule may contain at least one modified nucleotide analog. For example, the ends can be stabilized by incorporating a modified nucleotide analog.

[0222] Non-limiting examples of nucleotide analogs include sugar and / or backbone-modified ribonucleotides (i.e., including modifications to the phosphate-sugar backbone). For example, the phosphodiester bonds of native RNA can be modified to include at least one nitrogen or sulfur heteroatom. In exemplary backbone-modified ribonucleotides, the phosphate ester groups linking adjacent ribonucleotides are replaced by modifying groups (e.g., thiophosphate groups). In exemplary sugar-modified ribonucleotides, the 2'OH group is replaced by a group selected from H, OR, R, halogen, SH, SR, NH2, NHR, NR2, or ON, wherein R is a C1-C6 alkyl, alkenyl, or alkynyl group, and the halogen is F, Cl, Br, or I.

[0223] Other examples of modifications are nucleobase-modified ribonucleotides, i.e., ribonucleotides containing at least one non-naturally occurring nucleobase in place of a naturally occurring nucleobase. Bases can be modified to block the activity of adenosine deaminase. Exemplary modified nucleobases include, but are not limited to, uridine and / or cytidine modified at the 5-position, such as 5-(2-amino)propyluridine, 5-bromouridine; adenosine and / or guanosine modified at the 8-position, such as 8-bromoguanosine; denitronucleotides, such as 7-denitroadenosine; and O- and N-alkylated nucleotides, such as N6-methyladenosine. It should be noted that the above modifications can be used in combination.

[0224] In some cases, nucleic acid molecules contain at least one of the following chemical modifications: 2'-H, 2'-O-methyl, or 2'-OH modification of one or more nucleotides. In some embodiments, the nucleic acid molecules of the present invention may have enhanced nuclease resistance. To improve nuclease resistance, nucleic acid molecules may contain, for example, 2'-modified ribose units and / or thiophosphate bonds. For example, the 2' hydroxyl (OH) may be modified or substituted with a variety of different "oxygen" or "deoxy" substituents. To improve nuclease resistance, nucleic acid molecules of the present invention may contain 2'-O-methyl, 2'-fluorine, 2'-O-methoxyethyl, 2'-O-aminopropyl, 2'-amino, and / or thiophosphate bonds. Containing locked nucleic acids (LNAs), ethylene nucleic acids (ENAs) (e.g., 2'-4'-ethylene-bridging nucleic acids), and certain nucleobase modifications (e.g., 2-amino-A, 2-thio (e.g., 2-thio-U), G-clamp modifications) may also increase binding affinity to the target.

[0225] In one embodiment, the nucleic acid molecule comprises a 2'-modified nucleotide, such as 2'-deoxy, 2'-deoxy-2'-fluorine, 2'-O-methyl, 2'-O-methoxyethyl (2'-O-MOE), 2'-O-aminopropyl (2'-O-AP), 2'-O-dimethylaminoethyl (2'-O-DMAOE), 2'-O-dimethylaminopropyl (2'-O-DMAP), 2'-O-dimethylaminoethoxyethyl (2'-O-DMAEOE), or 2'-ON-methylacetamido (2'-O-NMA). In one embodiment, the nucleic acid molecule comprises at least one 2'-O-methyl modified nucleotide; in some embodiments, all nucleotides of the nucleic acid molecule comprise 2'-O-methyl modification.

[0226] In some embodiments, the nucleic acid molecule of the present invention has one or more of the following properties:

[0227] The nucleic acid reagents discussed in this article include unmodified RNA and DNA, as well as modified RNA and DNA (e.g., for improved efficacy), and polymers of nucleoside substitutes. Unmodified RNA refers to molecules whose nucleic acid components (i.e., the sugar, base, and phosphate moieties) are the same or substantially the same as those found in nature or naturally occurring in the human body. Rare or unusual but naturally occurring RNA is referred to in this art as modified RNA, see, for example, Limbach et al. (Nucleic Acids Res., 1994, 22:2183-2196). Such rare or unusual RNAs, generally referred to as modified RNA, are usually the result of post-transcriptional modification and fall within the scope of the term "unmodified RNA" as used herein. Modified RNA, as used herein, refers to molecules whose one or more components (i.e., the sugar, base, and phosphate moieties) differ from those found in nature or in the human body. Although they are called "modified RNA," they certainly also include molecules that are not strictly RNA due to modification. Nucleoside substitutes are molecules in which the ribose phosphate backbone is replaced by a non-ribose phosphate structure. This non-ribose phosphate structure allows the bases to be presented in the correct spatial positions, thus making the hybridization essentially similar to the hybridization of the ribose phosphate backbone. For example, uncharged ribose phosphate backbone analogs.

[0228] The nucleic acid modification of the present invention can exist in one or more of the phosphate group, sugar group, backbone, N-terminus, C-terminus or nucleobase.

[0229] carrier

[0230] The present invention also includes a composition comprising one or more vectors in which one or more nucleic acid molecules of the present invention are inserted. In one embodiment, the vector encodes at least two RNA molecules. In one embodiment, the vector comprises at least two RNA molecules. In some embodiments, the at least two RNA molecules are encoded by the same vector. In some embodiments, the at least two RNA molecules are contained within the same vector. In one embodiment, the at least two RNA molecules comprise a first RNA molecule and a second RNA molecule.

[0231] In some embodiments, the present invention comprises at least two vectors encoding at least two RNA molecules. In some embodiments, the at least two vectors comprise at least two RNA molecules. In some embodiments, the at least two vectors encode individual RNA molecules. In some embodiments, the at least two vectors comprise individual RNA molecules. In some embodiments, the at least two individual RNA molecules comprise a first RNA molecule and a second RNA molecule. In some embodiments, the first RNA molecule is encoded by a first vector, and the second RNA molecule is encoded by a second vector. In some embodiments, the first RNA molecule comprises a first vector, and the second RNA molecule comprises a second vector.

[0232] In some embodiments, the present invention further includes a vector encoding one or more additional RNA molecules. In some embodiments, the present invention further includes one or more vectors containing one or more additional RNA molecules. In some embodiments, each additional RNA molecule includes a coding region encoding a target protein domain; a 5' ribozyme; and a 3' ribozyme. In some embodiments, each additional RNA molecule includes a coding region encoding a target protein domain; a 5' ribozyme; and a 3' ribozyme recognition sequence.

[0233] Numerous suitable vectors are available in the art for use with this invention. In short, expression of a natural or synthetic nucleic acid encoding the fusion protein of this invention is typically achieved by operatively linking a nucleic acid encoding the fusion protein of this invention or a portion thereof to a promoter and introducing this construct into an expression vector. The vector to be used is suitable for replication in eukaryotic cells and optionally for integration. Typical vectors contain transcription and translation terminators, a start sequence, and a promoter for regulating the expression of the desired nucleic acid sequence.

[0234] Using standard gene delivery protocols, the vector of the present invention can also be used for nucleic acid immunotherapy and gene therapy. Gene delivery methods are known in the art. See, for example, U.S. Patents 5,399,346, 5,580,859, and 5,589,466, the entire contents of which are incorporated herein by reference. In another embodiment, the present invention provides a gene therapy vector.

[0235] The isolated nucleic acid of this invention can be cloned into various types of vectors. For example, the nucleic acid can be cloned into vectors, including but not limited to plasmids, phage particles, phage derivatives, animal viruses, and granules. Vectors of particular interest include expression vectors, replication vectors, probe generation vectors, and sequencing vectors.

[0236] Furthermore, vectors can be provided to cells in the form of viral vectors. Viral vector technology is well known in the art, for example, as described in Sambrook et al. (2012, Molecular Cloning: A Laboratory Manual, Cold Spring Harbor Laboratory, New York) and other virology and molecular biology manuals. Viruses that can be used as vectors include, but are not limited to, retroviruses, adenoviruses, adeno-associated viruses, herpesviruses, and lentiviruses. Generally, a suitable vector contains at least one origin of replication that is functional in at least one organism, a promoter sequence, a convenient restriction endonuclease site, and one or more selection markers (e.g., WO 01 / 96584; WO 01 / 29058; and U.S. Patent No. 6,326,193).

[0237] In addition, many other virus-based systems have been developed for transferring genes into mammalian cells. For example, retroviruses provide a convenient platform for gene delivery systems. Selected genes can be inserted into vectors and packaged into retroviral particles using techniques known in the art. The recombinant virus can then be isolated and delivered to the cells of a subject (either in vivo or in vitro). Various retroviral systems are known in the art. In some embodiments, adenovirus vectors are used. Various adenovirus vectors are known in the art.

[0238] In one embodiment, the composition comprises a vector derived from adeno-associated virus (AAV). The term "AAV vector" refers to a vector derived from an adeno-associated virus serotype, including but not limited to AAV-1, AAV-2, AAV-3, AAV-4, AAV-5, AAV-6, AAV-7, AAV-8, and AAV-9. AAV vectors have become a powerful gene delivery tool for treating a wide range of diseases. AAV vectors possess many properties that make them well-suited for gene therapy, including non-pathogenicity, extremely low immunogenicity, and the ability to transduce post-mitotic cells in a stable and efficient manner. By selecting a suitable combination of AAV serotype, promoter, and delivery method, the expression of a specific gene contained in the AAV vector can be specifically targeted to one or more cell types.

[0239] AAV vectors may have one or more AAV wild-type genes, preferably rep and / or cap genes, deleted entirely or partially, but retain functional flanking ITR sequences. Despite high homology, different serotypes exhibit tissue tropism. The receptor for AAV1 is unclear; however, AAV1 is known to transduce skeletal and cardiac muscle more efficiently than AAV2. Since most studies have been conducted using pseudotyped vectors that package vector DNA with flanking AAV2 ITRs into capsids of different serotypes, it is clear that biological differences are related to the capsid rather than the genome. Recent evidence suggests that DNA expression cassettes packaged in AAV1 capsids are at least 1 log10 more efficient at transducing cardiomyocytes than those packaged in AAV2 capsids. In one embodiment, the viral delivery system is an adeno-associated virus (AAV) delivery system. Adeno-associated viruses can be serotype 1 (AAV1), serotype 2 (AAV2), serotype 3 (AAV3), serotype 4 (AAV4), serotype 5 (AAV5), serotype 6 (AAV6), serotype 7 (AAV7), serotype 8 (AAV8), or serotype 9 (AAV9).

[0240] Suitable AAV fragments for assembly into vectors include cap proteins (including vp1, vp2, vp3, and hypervariable regions), rep proteins (including rep 78, rep 68, rep 52, and rep 40), and sequences encoding these proteins. These fragments are readily available for use in a variety of vector systems and host cells. Such fragments can be used alone, in combination with other AAV serotype sequences or fragments, or in combination with elements of other AAV or non-AAV viral sequences. Artificial AAV serotypes used herein include, but are not limited to, AAVs with non-naturally occurring capsid proteins. Such artificial capsids can be generated using any suitable technique, combining selected AAV sequences (e.g., fragments of the vp1 capsid protein) with heterologous sequences that can be derived from different selected AAV serotypes, discontinuous portions of the same AAV serotype, from non-AAV viral sources, or from non-viral sources. Artificial AAV serotypes can be chimeric AAV capsids, recombinant AAV capsids, or “humanized” AAV capsids, but are not limited to these. Therefore, exemplary AAVs or artificial AAVs suitable for expressing one or more proteins include AAV2 / 8 (see U.S. Patent No. 7,282,199), AAV2 / 5 (available from the National Institutes of Health), AAV2 / 9 (International Patent Publication No. WO2005 / 033321), AAV2 / 6 (U.S. Patent No. 6,156,303), and AAVrh8 (International Patent Publication No. WO2003 / 042397), etc.

[0241] In one embodiment, the composition comprises a lentiviral vector for delivering one or more nucleic acids of the present invention. In one embodiment, the present invention comprises a lentiviral vector containing one or more RNA molecules encoding one or more target proteins. For example, vectors derived from retroviruses (e.g., lentiviruses) are suitable tools for achieving long-term gene transfer because they allow for the long-term stable integration of transgenes and their proliferation in daughter cells. Lentiviral vectors have additional advantages over vectors derived from tumor retroviruses (e.g., murine leukemia virus) because they can transduce non-proliferating cells (e.g., hepatocytes). They also have the additional advantage of low immunogenicity.

[0242] In some embodiments, the vector also includes conventional control elements that are operatively linked to the transgene in a manner that allows transcription, translation, and / or expression of the transgene in cells transfected with the plasmid vector produced by this invention or infected with a virus. The “operatively linked” sequence as used herein includes both expression control sequences adjacent to the target gene and expression control sequences that control the target gene by acting trans or at a distance. Expression control sequences include: suitable transcription initiation, termination, promoter, and enhancer sequences; effective RNA processing signals, such as splicing and polyadenylation (polyA) signals; sequences stabilizing cytoplasmic mRNA; sequences that enhance translation efficiency (e.g., Kozak concordant sequences); sequences that enhance protein stability; and sequences that, if necessary, enhance the secretion of the encoded product. A variety of expression control sequences are known in the art and can be utilized, including natural promoters, constitutive promoters, inducible promoters, and / or tissue-specific promoters.

[0243] Other promoter elements, such as enhancers, regulate the frequency of transcription initiation. These elements are typically located in a region 30–110 bp upstream of the start site, although recent studies have shown that many promoters also contain functional elements downstream of the start site. The spacing between promoter elements is generally flexible, so promoter function can be preserved even when elements are inverted or moved relative to each other. In the thymidine kinase (TK) promoter, the spacing between promoter elements can increase to 50 bp before activity begins to decline. Depending on the promoter, the individual elements appear to function synergistically or independently to activate transcription.

[0244] An example of a suitable promoter is the immediate early cytomegalovirus (CMV) promoter sequence. This promoter sequence is a strongly constitutive promoter sequence capable of driving high-level expression of any polynucleotide sequence operatively linked to it. Another example of a suitable promoter is elongation growth factor-1α (EF-1α). However, other constitutive promoter sequences may also be used, including but not limited to the simian virus 40 (SV40) early promoter, mouse mammary tumor virus (MMTV), human immunodeficiency virus (HIV) long terminal repeat (LTR) promoter, MoMuLV promoter, avian leukosis virus promoter, Epstein-Barr virus immediate early promoter, Rous sarcoma virus promoter, and human gene promoters, such as, but not limited to, actin promoter, myosin promoter, hemoglobin promoter, and creatine kinase promoter. Furthermore, the invention should not be limited to the use of constitutive promoters. Inducible promoters are also covered in this invention. Using an inducible promoter provides a molecular switch that can turn on the expression of a polynucleotide sequence operatively linked to it when such expression is needed, or turn off expression when it is not needed. Examples of inducible promoters include, but are not limited to, metallothionein promoters, glucocorticoid promoters, progesterone promoters, and tetracycline promoters.

[0245] Enhancer sequences found on vectors also regulate the expression of genes contained therein. Typically, enhancers bind to protein factors to enhance gene transcription. Enhancers can be located upstream or downstream of the gene they regulate. Enhancers can also be tissue-specific, enhancing transcription in specific cell or tissue types. In one embodiment, the vector of the present invention contains one or more enhancers to enhance the transcription of genes present within the vector.

[0246] To evaluate the expression of the fusion protein of the present invention, the expression vector to be introduced into cells may further contain a selectable marker gene or a reporter gene, or both, to facilitate the identification and selection of expressing cells from a population of cells to be transfected or infected via a viral vector. In other aspects, the selectable marker may be carried on a separate DNA fragment and used in a co-transfection procedure. Both the selectable marker and the reporter gene may be side-linked with appropriate regulatory sequences to enable their expression in host cells. Useful selectable markers include, for example, antibiotic resistance genes, such as neo, etc.

[0247] Reporter genes are used to identify potentially transfected cells and assess the functionality of regulatory sequences. Generally, a reporter gene is a gene that is absent or not expressed in the recipient organism or tissue and whose encoded polypeptide expression is indicated by easily detectable properties such as enzyme activity. The expression of this DNA is measured at an appropriate time after the reporter gene is introduced into the recipient cells. Suitable reporter genes may include genes encoding luciferase, β-galactosidase, chloramphenicol acetyltransferase, secretory alkaline phosphatase, or green fluorescent protein (e.g., Ui-Tei et al., 2000 FEBS Letters 479:79-82). Suitable expression systems are well-known and can be prepared or commercially available using known techniques. Typically, constructs with minimal 5' flanking regions that exhibit the highest reporter gene expression levels are identified as promoters. Such promoter regions can be linked to reporter genes to assess the ability of drugs to regulate promoter-driven transcription.

[0248] protein

[0249] In some embodiments, the present invention includes a composition containing a ligase. In some embodiments, the ligase mediates the ligation of the 3'P or 2'3'cP terminus of an RNA molecule to the 5'OH terminus of an RNA molecule. In some embodiments, the ligase is an RNA 2',3'-cyclic phosphate and 5'-OH (RtcB) ligase. In some embodiments, the RtcB ligase is derived from one or more domains selected from organisms including eukaryotes, bacteria, and archaea. In some embodiments, the organism is selected from: humans, *Escherichia coli*, *Radiataetomyces*, *Porcine horikosa*, *Porcine horikosa* ST04, and *Thermococcus sp. EP*. In some embodiments, the ligase comprises one or more amino acid sequences selected from: SEQ ID NO:81, SEQ ID NO:83, SEQ ID NO:85, SEQ ID NO:87, SEQ ID NO:89, SEQ ID NO:91.

[0250] In some embodiments, one or more proteins of the present invention comprise an amino acid sequence substantially homologous to the amino acid sequence described herein. For example, in some embodiments, the degree of identity of the protein with respect to the original amino acid sequence is at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5%.

[0251] In some embodiments, one or more proteins of the present invention comprise an amino acid sequence that is part of the amino acid sequence described herein. For example, in some embodiments, the length of the protein relative to the original amino acid sequence is at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5%.

[0252] In some embodiments, one or more proteins of the present invention comprise an amino acid sequence that is part of the amino acid sequence described herein and is substantially homologous to the amino acid sequence described herein. For example, in some embodiments, the degree of identity of the protein with respect to the original amino acid sequence is at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or... At least 99.5%, and / or at least 60%, 65%, 70%, 75%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 99.5% relative to the length of the original amino acid sequence.

[0253] Pharmaceutical Composition

[0254] This invention also covers methods for practicing the invention using pharmaceutical compositions of the invention or salts thereof. Such pharmaceutical compositions may consist of at least one nucleic acid of the invention or a salt thereof in a form suitable for administration to a subject, or the pharmaceutical composition may contain at least one nucleic acid of the invention or a salt thereof, and one or more pharmaceutically acceptable carriers, one or more additional ingredients, or combinations thereof. The nucleic acids of the invention may be present in the pharmaceutical composition in the form of physiologically acceptable salts, for example, in combination with physiologically acceptable cations or anions, as is well known in the art.

[0255] In one embodiment, the pharmaceutical composition for practicing the method of the present invention may be administered at a dose of 1 ng / kg / day to 100 mg / kg / day. In another embodiment, the pharmaceutical composition for practicing the present invention may be administered at a dose of 1 ng / kg / day to 500 mg / kg / day.

[0256] The relative amounts of the active ingredient, pharmaceutically acceptable carrier, and any additional ingredients in the pharmaceutical compositions of the present invention will vary depending on the identity, body type, and condition of the treated subject, and further depend on the route of administration of the composition. For example, the composition may contain 0.1% to 100% (w / w) of the active ingredient.

[0257] The pharmaceutical compositions used in the methods of the present invention can be suitably developed for oral, rectal, vaginal, parenteral, topical, pulmonary, intranasal, oral, ocular, or other routes of administration. The compositions used in the methods of the present invention can be applied directly to the skin or any other tissue of a mammal. Other formulations considered include liposomal formulations, resealed red blood cells containing the active ingredient, and immunologically based formulations. Routes of administration will be apparent to those skilled in the art and depend on many factors, including the type and severity of the disease being treated, the type and age of the veterinary or human subject being treated, etc.

[0258] The formulations of the pharmaceutical compositions described herein can be prepared by any method known or to be developed in the field of pharmacology. Typically, such preparation methods include the following steps: mixing the active ingredient with a carrier or one or more other auxiliary ingredients, and then, if desired or anticipated, shaping or packaging the product into the desired single-dose or multi-dose units.

[0259] As used herein, “unit dose” refers to a discrete amount of a pharmaceutical composition containing a predetermined amount of the active ingredient. The amount of the active ingredient is typically equal to the dose of the active ingredient to be administered to the subject, or a convenient fraction of that dose, such as half or one-third of the dose. The unit dose may be in the form of a single daily dose or one of several daily doses (e.g., about 1 to 4 or more times per day). When using multiple daily doses, the unit dose form for each dose may be the same or different.

[0260] In one embodiment, the compositions of the present invention are formulated using one or more pharmaceutically acceptable excipients or carriers. In one embodiment, the pharmaceutical compositions of the present invention comprise a therapeutically effective amount of the nucleic acid of the present invention and a pharmaceutically acceptable carrier. Available pharmaceutically acceptable carriers include, but are not limited to, glycerol, water, saline, ethanol, and other pharmaceutically acceptable salt solutions, such as phosphates and organic acid salts. Examples of these and other pharmaceutically acceptable carriers are described in Remington's Pharmaceutical Sciences (1991, Mack Publication Co., New Jersey).

[0261] The carrier can be a solvent or dispersion medium, such as comprising water, ethanol, polyols (e.g., glycerol, propylene glycol, and liquid polyethylene glycol, etc.), suitable mixtures thereof, and vegetable oils. Suitable flowability can be maintained, for example, by using coating (e.g., lecithin), by maintaining the desired particle size in the dispersed state, and by using surfactants. Microbial action can be prevented by various antibacterial and antifungal agents (e.g., parabens, chlorobutanol, phenol, ascorbic acid, thimerosal, etc.). In many cases, the composition contains isotonic agents, such as sugars, sodium chloride, or polyols (e.g., mannitol and sorbitol). The absorption of the injectable composition can be prolonged by including substances that delay absorption (e.g., aluminum monostearate or gelatin) in the composition. In one embodiment, pharmaceutically acceptable carriers are not limited to DMSO.

[0262] The formulations can be mixed with conventional excipients, namely pharmaceutically acceptable organic or inorganic carrier substances, suitable for oral, vaginal, parenteral, nasal, intravenous, subcutaneous, enteral, or any other suitable route of administration known in the art. The pharmaceutical formulations can be sterilized and can be mixed with adjuvants (e.g., lubricants, preservatives, stabilizers, wetting agents, emulsifiers, salts for influencing osmotic pressure, buffers, colorants, flavoring agents, and / or aromatizers, etc.) as needed. They can also be combined with other active agents (e.g., other analgesics) as needed.

[0263] As used herein, “additional ingredients” include, but are not limited to, one or more of the following: excipients; surfactants; dispersants; inert diluents; granulators and disintegrants; binders; lubricants; sweeteners; flavorings; colorants; preservatives; physiologically degradable compositions, such as gelatin; aqueous carriers and solvents; oily carriers and solvents; suspending agents; dispersants or wetting agents; emulsifiers, modifiers; buffers; salts; thickeners; fillers; emulsifiers; antioxidants; antibiotics; antifungals; stabilizers; and pharmaceutically acceptable polymers or hydrophobic materials. Other “additional ingredients” that may be included in the pharmaceutical compositions of the present invention are known in the art and described, for example, in Genaro’s (1985, Remington’s Pharmaceutical Sciences, Mack Publishing Co., Easton, PA), which is incorporated herein by reference.

[0264] The compositions of the present invention may contain a preservative comprising about 0.005% to 2.0% by weight of the total composition. This preservative is used to prevent spoilage of the product when exposed to contaminants in the environment. Examples of preservatives usable according to the present invention include, but are not limited to, preservatives selected from benzyl alcohol, sorbic acid, parabens, imidazolidinyl urea, and combinations thereof. An exemplary preservative is a combination of about 0.5% to 2.0% benzyl alcohol and 0.05% to 0.5% sorbic acid.

[0265] In one embodiment, the composition comprises an antioxidant and a chelating agent that inhibit nucleic acid degradation. Exemplary antioxidants include BHT, BHA, α-tocopherol, and ascorbic acid, in amounts ranging from about 0.01 wt% to 0.3 wt% of the total weight of the composition, and BHT in amounts ranging from 0.03 wt% to 0.1 wt% of the total weight of the composition. In one embodiment, the chelating agent is present in amounts ranging from 0.01 wt% to 0.5 wt% of the total weight of the composition. Exemplary chelating agents include ethylenediaminetetraacetic acid (e.g., disodium ethylenediaminetetraacetate) and citric acid, in amounts ranging from about 0.01 wt% to 0.20% by weight. In some embodiments, the chelating agent is present in amounts ranging from 0.02 wt% to 0.10 wt% of the total weight of the composition. The chelating agent can be used to chelate metal ions in the composition that may affect the shelf life of the formulation. Although BHT and disodium EDTA are typical antioxidants and chelators for certain compounds, as is known to those skilled in the art, they can be replaced by other suitable and equivalent antioxidants and chelators.

[0266] Liquid suspensions can be prepared using conventional methods to suspend the active ingredient in an aqueous or oily carrier. Aqueous carriers include, for example, water and isotonic saline. Oily carriers include, for example, almond oil, oily esters, ethanol, vegetable oils (e.g., peanut oil, olive oil, sesame oil, or coconut oil), fractionated vegetable oils, and mineral oils (e.g., liquid paraffin). Liquid suspensions may also contain one or more additional ingredients, including but not limited to suspending agents, dispersants or wetting agents, emulsifiers, modifiers, preservatives, buffers, salts, flavoring agents, coloring agents, and sweeteners. Oily suspensions may also contain thickeners. Known suspending agents include, but are not limited to, sorbitol syrup, hydrogenated edible fats, sodium alginate, polyvinylpyrrolidone, tragacanth gum, gum arabic, and cellulose derivatives such as sodium carboxymethyl cellulose, methylcellulose, and hydroxypropyl methylcellulose. Known dispersants or wetting agents include, but are not limited to, naturally occurring phospholipids (e.g., lecithin), alkylene oxides with fatty acids, with long-chain fatty alcohols, with esters derived from fatty acids and hexitols, or with esters derived from fatty acids and hexitol anhydrides (e.g., polyoxyethylene stearate, heptadecaethyleneoxycetanol, polyoxyethylene sorbitan monooleate, and polyoxyethylene dehydrated sorbitan monooleate). Known emulsifiers include, but are not limited to, lecithin and gum arabic. Known preservatives include, but are not limited to, methylparaben, ethylparaben, or n-propylparaben, ascorbic acid, and sorbic acid. Known sweeteners include, for example, glycerin, propylene glycol, sorbitol, sucrose, and saccharin. Known thickeners for oily suspensions include beeswax, paraffin wax, and cetyl alcohol.

[0267] The preparation of liquid solutions of active ingredients in aqueous or oily solvents is essentially the same as that of liquid suspensions, the main difference being that the active ingredient is dissolved in the solvent, rather than suspended in it. As used herein, "oily" liquid refers to a liquid containing carbon-containing liquid molecules and having a lower polarity than water. Liquid solutions of the pharmaceutical compositions of this invention may contain each of the components described regarding liquid suspensions; however, it should be understood that suspending agents do not necessarily contribute to the dissolution of the active ingredient in the solvent. Aqueous solvents include, for example, water and isotonic saline. Oily solvents include, for example, almond oil, oily esters, ethanol, vegetable oils (e.g., peanut oil, olive oil, sesame oil, or coconut oil), fractionated vegetable oils, and mineral oils (e.g., liquid paraffin).

[0268] The powder and granule formulations of the pharmaceutical preparations of this invention can be prepared using known methods. Such formulations can be administered directly to a subject, for example, for making tablets, for filling capsules, or for preparing aqueous or oily suspensions or solutions by adding an aqueous or oily carrier. Each of these formulations may also contain one or more dispersants or wetting agents, suspending agents, and preservatives. Furthermore, these formulations may also contain other excipients, such as fillers and sweeteners, flavoring agents, or coloring agents.

[0269] The pharmaceutical compositions of the present invention can also be prepared, packaged, or marketed as oil-in-water emulsions or water-in-oil emulsions. The oil phase can be a vegetable oil, such as olive oil or peanut oil; a mineral oil, such as liquid paraffin; or a combination thereof. Such compositions may also contain one or more emulsifiers, such as naturally occurring gums, such as gum arabic or tragacanth; naturally occurring phospholipids, such as soybean lecithin or soybean lecithin; esters or metaesters derived from a combination of fatty acids and hexitanic anhydrides, such as sorbitan monooleate; and condensation products of such metaesters with ethylene oxide, such as polyoxyethylene sorbitan monooleate. These emulsions may also contain other ingredients, such as sweeteners or flavoring agents.

[0270] Methods of impregnating or coating materials with chemical compositions are known in the art, and include, but are not limited to, methods of depositing or incorporating chemical compositions onto surfaces, methods of incorporating chemical compositions into the structure of materials during material synthesis (i.e., using physiologically degradable materials), and methods of absorbing aqueous or oily solutions or suspensions into absorbent materials (with or without subsequent drying).

[0271] Dosing regimens may affect the composition of the effective dose. Therapeutic agents can be administered to subjects before or after disease diagnosis. Furthermore, multiple fractional or staggered doses can be given daily or continuously, or the dose can be administered via continuous infusion or bolus injection. Additionally, the dosage of the therapeutic agent can be increased or decreased proportionally depending on the urgency of the treatment or prevention situation.

[0272] The compositions of the present invention can be administered to subjects (including mammals, such as humans) using known procedures at doses and durations appropriate for the effective prevention or treatment of disease. The effective amount of nucleic acid required to achieve a therapeutic effect can vary depending on a variety of factors, such as the activity of the specific nucleic acid used; the timing of administration; the rate of nucleic acid excretion; the duration of treatment; other drugs, compounds, or materials used in combination with said nucleic acid; the state of the disease or disorder of the treated subject; age, sex, weight, condition, general health status, and medical history; and similar factors well known in the medical field. Dosing regimens can be adjusted to provide an optimal therapeutic response. For example, multiple daily doses can be administered, or the dose can be reduced proportionally according to the urgency of the treatment situation. A non-limiting example of the effective dose range of the nucleic acids of the present invention is about 1 to 5,000 mg / kg body weight / day. Those skilled in the art can investigate the relevant factors and determine the effective dose of therapeutic nucleic acids without extensive experimentation.

[0273] Nucleic acid can be administered to subjects multiple times daily, or at lower frequencies, such as once daily, once weekly, once every two weeks, once monthly, or even lower frequencies, such as once every few months, or even once a year or less. It should be understood that the daily nucleic acid dosage can be administered daily, every other day, every 2 days, every 3 days, every 4 days, or every 5 days, as in non-limiting examples. For example, with every other day administration, a dose of 5 mg daily could be given starting on Monday, the first subsequent daily dose of 5 mg on Wednesday, the second subsequent daily dose of 5 mg on Friday, and so on. The frequency of administration will be apparent to those skilled in the art and depends on many factors, such as (but not limited to) the type and severity of the disease being treated, the type and age of the animal, etc.

[0274] The actual dose level of the active ingredient in the pharmaceutical composition of the present invention can be varied to obtain an amount of active ingredient that effectively achieves the therapeutic response required for a specific subject, composition, and route of administration, without causing toxicity to the subject.

[0275] A physician or veterinarian with ordinary skills in the art can easily determine and prescribe the effective amount of the desired pharmaceutical composition. For example, a physician or veterinarian can start with a dose of the nucleic acid of the present invention used in the pharmaceutical composition that is below the dose required to achieve the desired therapeutic effect and gradually increase the dose until the desired effect is achieved.

[0276] In certain embodiments, it is particularly advantageous to formulate nucleic acids into dosage units to facilitate administration and maintain dosage uniformity. As used herein, dosage units refer to physically discrete units suitable as unit doses for use in a subject to be treated; each unit contains a predetermined amount of therapeutic nucleic acid, calculated to bind with a desired drug carrier to produce the intended therapeutic effect. The dosage unit form of the present invention depends on and is directly dependent on (a) the unique properties of the nucleic acid and the specific therapeutic effect to be achieved, and (b) the inherent limitations of techniques for compounding / formulating such nucleic acids for treating a subject's disease.

[0277] In one embodiment, the compositions of the present invention are administered to the subject at a dose range of one to five or more times daily. In another embodiment, the compositions of the present invention are administered to the subject at a dose range including, but not limited to, once daily, once every two days, once every three days, once weekly, and once every two weeks. Those skilled in the art will readily understand that the frequency of administration of the various combinations of the present invention will vary from subject to subject and will depend on many factors, including but not limited to age, disease or disorder to be treated, sex, overall health condition, and other factors. Therefore, the present invention should not be construed as limited to any particular dosage regimen, and the precise dose given to any subject and the composition to be administered to the subject will be determined by the attending physician taking into account all other factors of the subject.

[0278] The compositions of the present invention for application can be in the range of about 1 mg to about 10,000 mg, about 20 mg to about 9,500 mg, about 40 mg to about 9,000 mg, about 75 mg to about 8,500 mg, about 150 mg to about 7,500 mg, about 200 mg to about 7,000 mg, about 3,050 mg to about 6,000 mg, about 500 mg to about 5,000 mg, about 750 mg to about 4,000 mg, about 1 mg to about 3,000 mg, about 10 mg to about 2,500 mg, about 20 mg to about 2,000 mg, about 25 mg to about 1,500 mg, about 50 mg to about 1,000 mg, about 75 mg to about 900 mg, about 100 mg to about 800 mg, about 250 mg to about 750 mg, about 300 mg to about 600 mg, about 400 mg to about 500 mg, and any and all integer or partial increments therebetween.

[0279] In some embodiments, the dosage of the composition of the present invention is from about 1 mg to about 2,500 mg. In some embodiments, the dosage of the composition of the present invention used in the compositions described herein is less than about 10,000 mg, or less than about 8,000 mg, or less than about 6,000 mg, or less than about 5,000 mg, or less than about 3,000 mg, or less than about 2,000 mg, or less than about 1,000 mg, or less than about 500 mg, or less than about 200 mg, or less than about 50 mg. Similarly, in some embodiments, the dose of the second composition as described herein (i.e., a medicament for treating the same disease as or a different disease treated by the composition of the present invention) is less than about 1,000 mg, or less than about 800 mg, or less than about 600 mg, or less than about 500 mg, or less than about 400 mg, or less than about 300 mg, or less than about 200 mg, or less than about 100 mg, or less than about 50 mg, or less than about 40 mg, or less than about 30 mg, or less than about 25 mg, or less than about 20 mg, or less than about 15 mg, or less than about 10 mg, or less than about 5 mg, or less than about 2 mg, or less than about 1 mg, or less than about 0.5 mg, and any and all of these increments.

[0280] In one embodiment, the present invention relates to a packaged pharmaceutical composition comprising a container holding a therapeutically effective amount of the nucleic acid of the present invention, alone or in combination with a second pharmaceutical agent; and instructions for using the nucleic acid to treat, prevent, or alleviate one or more symptoms of a disease in a subject.

[0281] The term "container" includes any container used to contain a pharmaceutical composition. For example, in one embodiment, the container is packaging containing the pharmaceutical composition. In other embodiments, the container is not packaging containing the pharmaceutical composition; that is, the container is a receptacle, such as a box or vial, containing either a packaged or unpackaged pharmaceutical composition and instructions for use of the pharmaceutical composition. Furthermore, packaging techniques are well known in the art. It should be understood that instructions for use of the pharmaceutical composition may be included on packaging containing the pharmaceutical composition, thus creating an enhanced functional relationship between the instructions and the packaged product. However, it should be understood that the instructions may contain information related to the ability of the nucleic acid to perform its intended function, such as treating or preventing a disease in a subject, or delivering an imaging or diagnostic agent to a subject.

[0282] The routes of administration for any composition of the present invention include oral, nasal, parenteral, sublingual, transdermal, transmucosal (e.g., sublingual, tongue, buccal, and nasal), intravesical, intraduodenal, gastric, rectal, intraperitoneal, subcutaneous, intramuscular, intradermal, intra-arterial, and intravenous administration.

[0283] Suitable compositions and dosage forms include, for example, tablets, capsules, pouches, pills, gel capsules, troches, dispersions, suspensions, solutions, syrups, granules, beads, transdermal patches, gels, powders, pills, creams, lozenges, ointments, pastes, plasters, lotions, discs, suppositories, liquid sprays for nasal or oral administration, dry powders or aerosols for inhalation, and compositions and formulations for intravesical administration. It should be understood that the formulations and compositions used in this invention are not limited to the specific formulations and compositions described herein.

[0284] system

[0285] In some embodiments, the present invention relates to systems for cis- and trans-splicing of independent RNA molecules. In some embodiments, the present invention relates to systems for cis- and trans-splicing of single RNA molecules. In some embodiments, as described herein, cis- and trans-splicing of independent RNA molecules or single RNA molecule fragments produces single RNA molecules encoding full-length target proteins. In some embodiments, as described herein, the system comprises a ligase or a nucleic acid encoding a ligase, such as RtcB.

[0286] In one embodiment, the present invention relates to an inducible system for generating a single RNA encoding a full-length protein from two separate RNA molecules encoding a first and a second portion via cis-cleavage of a ribozyme and trans-splicing of two separate RNA molecules. In some embodiments, as described herein, the system comprises a ribozyme recognition sequence and a ribozyme. In some embodiments, as described herein, the system comprises a ligase or a nucleic acid encoding a ligase.

[0287] In one embodiment, the present invention relates to a system for assembling a full-length RNA virus genome. Exemplary RNA viruses include, but are not limited to: coronaviruses, paramyxoviruses, orthomyxoviruses, retroviruses, lentiviruses, alphaviruses, flaviviruses, rhabdoviruses, measles viruses, Newcastle disease viruses, and picornaviruses. In one embodiment, the system comprises a first nucleic acid encoding a first portion of the RNA virus genome and a first nucleic acid encoding a 3' ribozyme. In one embodiment, the system comprises a second nucleic acid encoding a second portion of the RNA virus genome and a 5' ribozyme. In one embodiment, the system comprises a first portion of the RNA virus genome and a 3' ribozyme. In one embodiment, the system comprises a second portion of the RNA virus genome and a 5' ribozyme. In one embodiment, the system comprises a nucleic acid or ligase encoding a ligase. In one embodiment, the first portion and the second portion of the RNA virus genome are ligated together during cis-cleavage of the 3' and 5' ribozymes, thereby producing a full-length RNA virus genome.

[0288] in vivo

[0289] In one embodiment, the present invention relates to a system for delivering and expressing one or more full-length proteins by cis- and trans-splicing a separate RNA molecule encoding a portion of a full-length protein. In some embodiments, the system allows for the delivery and expression of large proteins exceeding the packaging size of conventional vectors (e.g., dystrophin exceeding the packaging size of an AAV vector), synthetic repeating domain proteins that are difficult to synthesize in vitro using nucleic acid constructs (e.g., synthetic spider silk), or toxic / antiviral proteins (e.g., DTA). In one embodiment, the invention comprises an AAV system for delivering and expressing one or more full-length target proteins. In some embodiments, as described herein, the system comprises a ligase or a nucleic acid encoding a ligase.

[0290] In one embodiment, the present invention includes a lentiviral delivery system for delivering one or more nucleic acid molecules encoding one or more target proteins. In one aspect, the lentiviral delivery system includes (1) a packaging plasmid, (2) an envelope plasmid, and (3) a transfer plasmid. In one embodiment, the transfer plasmid encodes a first RNA molecule and a second RNA molecule.

[0291] In one embodiment, the present invention includes a dual lentiviral delivery system comprising a first lentiviral vector and a second lentiviral vector. In one embodiment, the first lentiviral vector system comprises (1) a packaging plasmid, (2) an envelope plasmid, and (3) a first transfer plasmid. In one embodiment, the second lentiviral vector system comprises (1) a packaging plasmid, (2) an envelope plasmid, and (3) a second transfer plasmid. In one embodiment, the first transfer plasmid encodes a first RNA molecule. In one embodiment, the second transfer plasmid encodes a second RNA molecule.

[0292] In one embodiment, the packaging plasmid contains a nucleic acid sequence encoding the gag-pol polyprotein. In one embodiment, the gag-pol polyprotein contains a catalytically inactivating integrase. In one embodiment, the gag-pol polyprotein contains a D116N integrase mutant.

[0293] In one embodiment, the envelope plasmid contains a nucleic acid sequence encoding an envelope protein. In one embodiment, the envelope plasmid contains a nucleic acid sequence encoding an HIV envelope protein. In one embodiment, the envelope plasmid contains a nucleic acid sequence encoding a vesicular stomatitis virus g protein (VSV-g) envelope protein. In one embodiment, the envelope protein can be selected according to the desired cell type.

[0294] In one embodiment, the first RNA molecule of the single transfer plasmid contains a protein-coding region encoding a first part of the target protein and a 3' ribozyme. In one embodiment, the second RNA molecule of the single transfer plasmid contains a protein-coding region encoding a second part of the target protein and a 5' ribozyme. In one embodiment, the transfer plasmid contains a 5' long terminal repeat (LTR) sequence and a 3' LTR sequence. In one embodiment, the 3' LTR is a self-inactivating (SIN) LTR. Therefore, in one embodiment, the 5' LTR contains the U3 sequence, the R sequence, and the U5 sequence, while the 3' LTR contains the R sequence and the U5 sequence, but not the U3 sequence. In one embodiment, the 5' LTR and the 3' LTR are located flanking the sequences encoding the first and second parts of the target protein, respectively.

[0295] In one embodiment, the first RNA molecule of the first transfer plasmid contains a protein-coding region encoding a first portion of the target protein and a 3' ribozyme. In one embodiment, the second RNA molecule of the second transfer plasmid contains a protein-coding region encoding a second portion of the target protein and a 5' ribozyme. In one embodiment, the first and second transfer plasmids contain a 5' long terminal repeat (LTR) sequence and a 3' LTR sequence. In one embodiment, the 3' LTR is a self-inactivating (SIN) LTR. Therefore, in one embodiment, the 5' LTR contains the U3 sequence, the R sequence, and the U5 sequence, while the 3' LTR contains the R sequence and the U5 sequence, but not the U3 sequence. In one embodiment, the 5' LTR and 3' LTR of the first transfer plasmid are located flanking the sequence encoding the first portion of the target protein and the 3' ribozyme. In one embodiment, the 5' LTR and 3' LTR of the second transfer plasmid are located flanking the sequence encoding the second portion of the target protein and the 5' ribozyme.

[0296] In one embodiment, a packaging plasmid, an envelope plasmid, and a transfer plasmid are introduced into a cell. In one embodiment, the cell transcribes and translates a nucleic acid sequence encoding a gag-pol protein to produce a gag-pol polyprotein. In one embodiment, the cell transcribes and translates a nucleic acid sequence encoding an envelope protein to produce an envelope protein. In one embodiment, the cell transcribes a single transfer plasmid to provide a first RNA molecule and a second RNA molecule. In one embodiment, the cell transcribes a first transfer plasmid to provide a first RNA molecule and a second transfer plasmid to provide a second RNA molecule. In one embodiment, the gag-pol protein, the envelope polyprotein, the first RNA molecule, and the second RNA molecule are packaged into viral particles. In one embodiment, viral particles are collected from a cell culture medium. In one embodiment, viral particles transduce target cells, wherein a 3' ribozyme catalyzes its detachment from a first RNA molecule to generate a 3'P or 2'3'cP terminus, and a 5' ribozyme catalyzes its detachment from a second RNA molecule to generate a 5'OH terminus. An endogenous RNA 2',3'-cyclic phosphate and 5'-OH (RtcB) ligase connects the 3'P or 2'3'cP terminus to the 5'OH terminus, thereby generating a complete RNA molecule encoding the target protein, and the cell translates the target protein.

[0297] In one embodiment, a packaging plasmid, an envelope plasmid, and a first transfer plasmid are introduced into a cell. In one embodiment, the cell transcribes and translates a nucleic acid sequence encoding a gag-pol protein to produce a gag-pol polyprotein. In one embodiment, the cell transcribes and translates a nucleic acid sequence encoding an envelope protein to produce an envelope protein. In one embodiment, the cell transcribes a first transfer plasmid to provide a first RNA molecule. In one embodiment, the gag-pol protein, the envelope polyprotein, and the first RNA molecule are packaged into a first viral particle. In one embodiment, the first viral particle is collected from a cell culture medium.

[0298] In one embodiment, a packaging plasmid, an envelope plasmid, and a second transfer plasmid are introduced into cells. In one embodiment, the cells transcribe and translate a nucleic acid sequence encoding a gag-pol protein to produce a gag-pol polyprotein. In one embodiment, the cells transcribe and translate a nucleic acid sequence encoding an envelope protein to produce an envelope protein. In one embodiment, the cells transcribe a second transfer plasmid to provide a second RNA molecule. In one embodiment, the gag-pol protein, the envelope polyprotein, and the second RNA molecule are packaged into a second viral particle. In one embodiment, the second viral particle is collected from a cell culture medium.

[0299] In one embodiment, a first viral particle and a second viral particle transduce a target cell, wherein a 3' ribozyme catalyzes its detachment from a first RNA molecule to generate a 3'P or 2'3'cP terminus; a 5' ribozyme catalyzes its detachment from a second RNA molecule to generate a 5'OH terminus; an endogenous RNA 2',3'-cyclic phosphate and 5'-OH (RtcB) ligase connects the 3'P or 2'3'cP terminus to the 5'OH terminus, thereby generating a complete RNA molecule encoding the target protein, and the cell translates the target protein. In one embodiment, the present invention relates to a system for preventing the expression of unwanted partial proteins in a pre-dividing RNA molecule. In one embodiment, as described herein, the system includes translational control of introducing protein degradation sequences into the pre-dividing RNA molecule.

[0300] In one embodiment, the present invention relates to a system for expressing two or more target proteins from two or more independent RNA molecule pairs encoding a portion of a target protein via ribozymatic cis- and trans-splicing of independent RNA molecule pairs. In one embodiment, as described herein, each individual independent RNA molecule pair has a separate reading frame such that trans-splicing of unwanted RNA molecule pairs does not result in the translation of a full-length functional protein. In some embodiments, as described herein, the system comprises a ligase or a nucleic acid encoding a ligase.

[0301] In one embodiment, the present invention includes a system for delivering and expressing a full-length target protein and a cargo sequence. In one embodiment, the system includes: a first portion of RNA encoding a first portion of the target protein, with its 3' end linked to a synthetic intron; and a second portion of RNA encoding a second portion of the target protein, with its 5' end linked to the synthetic intron. In one embodiment, a 5' ribozyme sequence and a 3' ribozyme sequence are respectively attached to both sides of the synthetic intron. In one embodiment, the synthetic intron contains a cargo sequence located between the 5' and 3' ribozyme sequences. In one embodiment, self-cleaving of the 5' and 3' ribozyme sequences produces three separate RNA molecules: 1) a first fragment containing the first portion of the first portion of RNA encoding the first portion of the target protein, 2) a second fragment containing the synthetic intron, and 3) a third fragment containing the second portion of the second portion of RNA encoding the second portion of the target protein. In one embodiment, the compatible ends of the second fragment are joined together to generate a circular RNA molecule containing the synthetic intron, which contains the cargo sequence. In one embodiment, the first and third fragments are joined together to generate a full-length linear RNA molecule. In one embodiment, the full-length target protein comprises a therapeutic protein, a reporter protein, a recombinase, an antibiotic resistance gene product, an antibody, or a Cas9 protein. In one embodiment, the cargo sequence comprises a therapeutic nucleic acid sequence (e.g., a miRNA sequence or a CRISPR guide RNA sequence) or encodes a therapeutic protein. In some embodiments, the full-length target protein comprises Cas9, and the cargo sequence comprises a guide RNA sequence, thereby targeting Cas9 to a specific genomic sequence for editing. In some embodiments, as described herein, the system comprises a ligase or a nucleic acid encoding a ligase.

[0302] In one embodiment, the present invention includes a gene editing system comprising one or more trans-splicing engineered ribozymes. In some embodiments, the system comprises two trans-splicing engineered ribozymes targeting upstream and downstream of a pathogenic mutation. In some embodiments, trans-splicing upstream and downstream of the pathogenic mutation removes the pathogenic mutation. In some embodiments, after the pathogenic mutation is trans-spliced, the remaining portion of the gene is trans-spliced ​​together. In some embodiments, the trans-spliced ​​gene is expressed as a functional protein. In some embodiments, as described herein, the system comprises a ligase or a nucleic acid encoding a ligase.

[0303] in vitro

[0304] In one embodiment, the present invention includes an in vitro system for generating RNA molecules encoding a target protein. In one embodiment, the system includes at least two RNA molecules. In one embodiment, the at least two RNA molecules include a first RNA molecule and a second RNA molecule.

[0305] In one embodiment, the first RNA molecule includes a coding region encoding a first portion of the target protein. In one embodiment, the first RNA molecule includes a 3' ribozyme. In one embodiment, as described herein, the first RNA molecule includes a coding region encoding a first portion of the target protein and a 3' ribozyme.

[0306] In one embodiment, the second RNA molecule includes a coding region encoding a second portion of the target protein. In one embodiment, the second RNA molecule includes a 5' ribozyme. In one embodiment, as described herein, the second RNA molecule includes a coding region encoding a second portion of the target protein and a 5' ribozyme.

[0307] In one embodiment, the in vitro system for generating RNA molecules encoding a target protein further comprises a ligase. In one embodiment, the ligase induces the assembly of RNA molecules from the coding regions of a first RNA molecule and a second RNA molecule. In one embodiment, as described herein, the ligase is an RNA 2',3'-cyclic phosphate and 5'-OH (RtcB) ligase.

[0308] In one embodiment, the present invention includes an in vitro system for generating an RNA molecule encoding a target protein with a repeating domain. In one embodiment, the system comprises a first RNA molecule, one or more additional RNA molecules, and a final RNA molecule.

[0309] In one embodiment, the first RNA molecule includes a coding region encoding a first portion of the target protein. In one embodiment, the first RNA molecule includes a 3' ribozyme. In one embodiment, the first RNA molecule includes a coding region encoding a first portion of the target protein and a 3' ribozyme. In one embodiment, the 3' ribozyme catalyzes its own separation from the first RNA molecule, thereby generating a 3'P or 2'3' cP terminus. In one embodiment, the first RNA molecule further includes a 5' tag. In one embodiment, the 5' tag mediates the connection of the first RNA molecule to a solid support.

[0310] In one embodiment, the one or more additional RNA molecules comprise a coding region for a domain encoding a target protein; a 5' ribozyme; and a 3' ribozyme recognition sequence. In one embodiment, the 5' ribozyme self-cleaves to generate a 5' OH terminus. In one embodiment, the 3' ribozyme recognition sequence comprises the VS-S sequence as described herein.

[0311] In one embodiment, the last RNA molecule contains a coding region encoding the final portion of the target protein. In one embodiment, the last RNA molecule contains a 5' ribozyme. In one embodiment, the last RNA molecule contains a coding region encoding the final portion of the target protein and a 5' ribozyme. In one embodiment, the 5' ribozyme self-cleaves to generate a 5' OH terminus.

[0312] In one embodiment, the system further comprises a ribozyme. In one embodiment, the ribozyme comprises VS-Rz as described herein. In one embodiment, the VS-Rz recognizes VS-S as described herein and mediates its cleavage from one or more additional RNA molecules. In one embodiment, the cleavage produces a 3'P or 2'3'cP terminus.

[0313] In one embodiment, the system comprises a ligase. In some embodiments, the ligase ligates the 3'P or 2'3'cP end of a first RNA molecule to the 5'OH end of one or more additional RNA molecules. In some embodiments, the ligase ligates the 3'P or 2'3'cP end of one or more additional RNA molecules to the 5'OH end of a last RNA molecule. In some embodiments, the ligase ligates the 3'P or 2'3'cP end of the first RNA molecule to the 5'OH end of one or more additional RNA molecules and ligates the 3'P or 2'3'cP end of one or more additional RNA molecules to the 5'OH end of the last RNA molecule, thereby generating a complete RNA molecule encoding an N-terminal domain, one or more additional domains, and a C-terminal domain. In some embodiments, as described herein, the ligase is an RNA 2',3'-cyclic phosphate and 5'-OH (RtcB) ligase.

[0314] method

[0315] In some embodiments, the present invention relates to methods for cis-cleaving, trans-splicing, or scarless trans-ligation of independent RNA molecules. In some embodiments, the present invention relates to methods for cis-cleaving, trans-splicing, or scarless trans-ligation of single RNA molecules. In some embodiments, as described herein, cis-cleaving and trans-splicing of independent RNA molecules or fragments of single RNA molecules can produce single RNA molecules encoding full-length target proteins. In some embodiments, as described herein, the method includes administering a ligase or a nucleic acid encoding a ligase.

[0316] In one embodiment, the present invention relates to an inducible method for generating a single RNA encoding a full-length protein from two separate RNA molecules by cis-splicing of ribozymes contained on two separate RNA molecules (the two separate RNA molecules encoding a first and a second portion of a full-length protein, respectively) and trans-splicing of the two separate RNA molecules. In some embodiments, scarless ligation of the two separate RNA molecules generates a single RNA molecule encoding a full-length target protein. In some embodiments, as described herein, the method includes a ribozyme recognition sequence and a ribozyme. In some embodiments, as described herein, the method includes administering a ligase or a nucleic acid encoding a ligase.

[0317] In some embodiments, the present invention relates to a method for circularizing RNA molecules. In some embodiments, the method includes cis-cleavage and cis-ligation or scarless cis-splicing of a single RNA molecule. In some embodiments, the linear RNA molecule (transcribed in vitro or in vivo from a DNA plasmid) comprises a 5' ribozyme and a 3' ribozyme. In one embodiment, the 5' ribozyme is directly ligated to the N-terminal coding sequence of the target protein, and the 3' ribozyme is directly ligated to the C-terminal coding sequence of the target protein. In one embodiment, the 3' ribozyme is directly ligated to the N-terminal coding sequence of the target protein, and the 5' ribozyme is directly ligated to the C-terminal coding sequence of the target protein. In some embodiments, as described herein, cis-cleavage of the 3' and 5' ribozymes, and cis-ligation or scarless cis-splicing of the RNA molecule fragment containing the coding sequence, produce a single circular RNA molecule encoding a full-length target protein. In some embodiments, as described herein, the method includes administering a ligase or a nucleic acid encoding a ligase.

[0318] in vivo

[0319] In one embodiment, the present invention includes a method for generating an RNA molecule encoding a target protein. In some embodiments, the method includes administering at least two nucleic acid molecules to cells or tissues. In one embodiment, the at least two nucleic acid molecules include a first RNA molecule and a second RNA molecule. In some embodiments, the at least two nucleic acid molecules encode the first RNA molecule and the second RNA molecule.

[0320] In one embodiment, the first RNA molecule contains a coding region encoding a first portion of the target protein. In one embodiment, the first RNA molecule contains a 3' ribozyme. In one embodiment, the first RNA molecule contains a coding region encoding a first portion of the target protein and a 3' ribozyme. In one embodiment, the 3' ribozyme catalyzes its own detachment from the first RNA molecule, thereby generating a 3'P or 2'3' cP terminus. In one embodiment, the 3' ribozyme is a member of the HDV ribozyme family.

[0321] In one embodiment, the second RNA molecule contains a coding region encoding a second portion of the target protein. In one embodiment, the second RNA molecule contains a 5' ribozyme. In one embodiment, the second RNA molecule contains a coding region encoding a second portion of the target protein and a 5' ribozyme. In one embodiment, the 5' ribozyme catalyzes its own separation from the second RNA molecule, thereby generating a 5' OH terminus. In one embodiment, the 5' ribozyme is a member of the HH ribozyme family.

[0322] In one embodiment, the 3'P or 2'3'cP end is joined to the 5'OH end to form an RNA molecule containing the coding regions of a first RNA molecule and a second RNA molecule. In some embodiments, the trans-joining of the coding sequences of the first and second portions of the target protein is performed in a scarless manner, such that there are no interpolated sequences between the first and second portions of the translated target protein.

[0323] In one embodiment, the method includes administering to a cell or tissue one or more additional nucleic acid molecules encoding one or more additional RNA molecules, each additional RNA molecule containing a coding region encoding a target protein domain; a 5' ribozyme; and a 3' ribozyme.

[0324] In one embodiment, the method includes administering to cells or tissue one or more additional nucleic acid molecules encoding one or more additional RNA molecules, each additional RNA molecule containing a coding region encoding a target protein domain; a 5' ribozyme; and a 3' ribozyme recognition sequence. In one embodiment, the 3' ribozyme recognition sequence comprises VS-S. In one embodiment, the ribozyme is VS.

[0325] In one embodiment, the method includes administering to cells or tissue a nucleic acid molecule selected from one or more of the following: encoding a ligase and a ligase. In one embodiment, the ligase induces the assembly of RNA molecules from the coding regions of a first RNA molecule and a second RNA molecule. In some embodiments, the trans-ligation of the coding sequences of the first and second portions of the target protein is performed in a scarless manner, such that there are no intercalation sequences between the first and second portions of the translated target protein. In one embodiment, the ligase is an RNA 2',3'-cyclic phosphate and 5'-OH (RtcB) ligase.

[0326] In some embodiments, the method includes administering at least one AAV vector to cells or tissues, the AAV vector encoding a first RNA molecule and a second RNA molecule, the first RNA molecule comprising a protein-coding region encoding a first portion of a target protein and a 3' ribozyme, and the second RNA molecule comprising a protein-coding region encoding a second portion of the target protein and a 5' ribozyme. In some embodiments, as described herein, the method includes administering a ligase or a nucleic acid encoding a ligase.

[0327] In some embodiments, the method includes administering at least two AAV vectors, including a first AAV vector and a second AAV vector. In one embodiment, the first AAV vector encodes a first RNA molecule containing a protein-coding region encoding a first portion of the target protein and a 3' ribozyme. In one embodiment, the second AAV vector encodes a second RNA molecule containing a protein-coding region encoding a second portion of the target protein and a 5' ribozyme. In some embodiments, as described herein, the method includes administering a ligase or a nucleic acid encoding a ligase. In some embodiments, the trans-ligation of the coding sequences of the first and second portions of the target protein is performed in a scarless manner, such that there are no intercalation sequences between the first and second portions of the translated target protein.

[0328] In some embodiments, the method includes administering at least one lentiviral vector to cells or tissues, the lentiviral vector encoding a first RNA molecule and a second RNA molecule, the first RNA molecule comprising a protein-coding region encoding a first portion of a target protein and a 3' ribozyme, and the second RNA molecule comprising a protein-coding region encoding a second portion of the target protein and a 5' ribozyme. In some embodiments, as described herein, the method includes administering a ligase or a nucleic acid encoding a ligase.

[0329] In some embodiments, the method includes administering at least two lentiviral vectors, including a first lentiviral vector and a second lentiviral vector. In one embodiment, the first lentiviral vector encodes a first RNA molecule containing a protein-coding region encoding a first portion of a target protein and a 3' ribozyme. In one embodiment, the second lentiviral vector encodes a second RNA molecule containing a protein-coding region encoding a second portion of the target protein and a 5' ribozyme. In some embodiments, as described herein, the method includes administering a ligase or a nucleic acid encoding a ligase.

[0330] In some embodiments, the method includes administering at least one lentiviral vector delivery system to cells or tissues to provide a first RNA molecule comprising a protein-coding region encoding a first portion of a target protein and a 3' ribozyme, and a second RNA molecule comprising a protein-coding region encoding a second portion of the target protein and a 5' ribozyme. In some embodiments, as described herein, the method includes administering a ligase or nucleic acid encoding a ligase.

[0331] In some embodiments, the method includes administering at least two lentiviral vector delivery systems, including a first lentiviral vector delivery system and a second lentiviral vector delivery system. In one embodiment, the first lentiviral vector delivery system provides a first RNA molecule containing a protein-coding region encoding a first portion of a target protein and a 3' ribozyme. In one embodiment, the second lentiviral vector delivery system provides a second RNA molecule containing a protein-coding region encoding a second portion of a target protein and a 5' ribozyme to a cell or tissue. In some embodiments, as described herein, the method includes administering a ligase or a nucleic acid encoding a ligase.

[0332] In some embodiments, the method includes administering two or more delivery vectors selected from AAV vectors, lentiviral vectors, lentiviral vector delivery systems, or combinations thereof. In one embodiment, the two or more delivery vectors include a first delivery vector and a second delivery vector. In one embodiment, the first delivery vector provides a first RNA molecule containing a protein-coding region encoding a first portion of a target protein and a 3' ribozyme. In one embodiment, the second delivery vector provides a second RNA molecule to a cell or tissue, the second RNA molecule containing a protein-coding region encoding a second portion of a target protein and a 5' ribozyme. In some embodiments, as described herein, the method includes administering a ligase or a nucleic acid encoding a ligase.

[0333] In one embodiment, the present invention includes a method for generating a circular RNA molecule encoding a target protein. In some embodiments, the method includes administering at least one nucleic acid molecule to a cell or tissue, wherein the nucleic acid molecule comprises at least a 5' ribozyme with a sequence directly linked to the C-terminal or N-terminal portion encoding the target protein and a 3' ribozyme with a sequence directly linked to the C-terminal or N-terminal portion encoding the target protein. In one embodiment, the 3' ribozyme catalyzes its own detachment from the RNA molecule, thereby generating a 3'P or 2'3'cP terminus. In one embodiment, the 3' ribozyme is a member of the HDV ribozyme family. In one embodiment, the 5' ribozyme catalyzes its own detachment from the RNA molecule, thereby generating a 5'OH terminus. In one embodiment, the 5' ribozyme is a member of the HH ribozyme family. In one embodiment, the 3'P or 2'3'cP terminus is linked to the 5'OH terminus to form a circular RNA molecule in which the coding sequences of the N-terminus and C-terminus of the target protein are operatively linked. In some embodiments, the circularization is performed in a scarless manner, such that there are no intercalation sequences between the post-translational N-terminus and C-terminus.

[0334] Methods for introducing and expressing genes into cells are known in the art. As for expression vectors, they can be readily introduced into host cells, such as mammalian, bacterial, yeast, or insect cells, by any method in the art. For example, expression vectors can be transferred into host cells by physical, chemical, or biological means.

[0335] Physical methods for introducing polynucleotides into host cells include calcium phosphate precipitation, liposome transfection, particle bombardment, microinjection, and electroporation. Methods for preparing cells containing vectors and / or exogenous nucleic acids are well known in the art. See, for example, Sambrook et al. (2012, Molecular Cloning: A Laboratory Manual, Cold Spring Harbor Laboratory, New York). An exemplary method for introducing polynucleotides into host cells is calcium phosphate transfection.

[0336] Biological methods for introducing target polynucleotides into host cells include the use of DNA and RNA vectors. Viral vectors, especially retroviral vectors, have become the most widely used method for inserting genes into mammalian (e.g., human) cells. Other viral vectors can be derived from lentiviruses, poxviruses, herpes simplex virus type I, adenoviruses, and adeno-associated viruses, etc. See, for example, U.S. Patents 5,350,674 and 5,585,362.

[0337] Chemical methods for introducing polynucleotides into host cells include colloidal dispersion systems, such as macromolecular complexes, nanocapsules, microspheres, beads, and lipid-based systems, including oil-in-water emulsions, micelles, mixed micelles, and liposomes. An exemplary colloidal system used as a delivery carrier in vitro and in vivo is a liposome (e.g., an artificial membrane vesicle).

[0338] In the case of using non-viral delivery systems, an exemplary delivery vector is the liposome. Consider using lipid formulations to introduce nucleic acids into host cells (in vitro, ex vivo, or in vivo). Alternatively, nucleic acids can bind to lipids. Lipid-bound nucleic acids can be encapsulated within the aqueous interior of liposomes, dispersed within the lipid bilayer of liposomes, attached to liposomes by linkers that bind to both liposomes and oligonucleotides, embedded in liposomes, complexed with liposomes, dispersed in a solution containing lipids, mixed with lipids, bound to lipids, contained in lipids in suspension form, contained or complexed with micelles, or otherwise bound to lipids. Combinations of lipids, lipid / DNA, or lipid / expression vectors in solution are not limited to any particular structure. For example, they can exist in bilayer structures, micelles, or “collapsed” structures. They can also simply be dispersed in solution, possibly forming aggregates of non-uniform size or shape. Lipids are fatty substances and can be naturally occurring or synthetic. For example, lipids include fat droplets naturally present in the cytoplasm, as well as compounds containing long-chain aliphatic hydrocarbons and their derivatives, such as fatty acids, alcohols, amines, amino alcohols, and aldehydes.

[0339] Suitable lipids are available from commercial sources. For example, dimyristyl phosphatidylcholine (“DMPC”) is available from Sigma, Inc. (St. Louis, Missouri); diceryl phosphate (“DCP”) is available from K&K Laboratories, Inc. (Plainview, NY); cholesterol (“Choi”) is available from Calbiochem-Behring, Inc.; dimyristyl phosphatidylglycerol (“DMPG”) and other lipids are available from Avanti Polar Lipids, Inc. (Birmingham, Alabama, NY). The lipids can be stored in chloroform or chloroform / methanol stock solutions at approximately -20°C. Chloroform is used as the sole solvent because it evaporates more readily than methanol. “Liposome” is a general term encompassing various monolayer and multilayer lipid carriers formed by closed lipid bilayers or aggregates. Liposomes are characterized by a vesicular structure with a phospholipid bilayer membrane and an internal aqueous medium. Multilayer liposomes have multiple lipid layers separated by an aqueous medium. When phospholipids are suspended in excess aqueous solution, they spontaneously form. The lipid components rearrange themselves before forming a closed structure, encapsulating water and dissolved solutes between the lipid bilayers (Ghosh et al., 1991 Glycobiology 5:505-10). However, compositions exhibiting structures in solution that differ from normal vesicle structures are also included. For example, lipids can present as micelle structures or exist solely as heterogeneous aggregates of lipid molecules. Lipofectamine-nucleic acid complexes have also been envisioned.

[0340] Regardless of the method used to introduce exogenous nucleic acids into host cells, various assays can be performed to confirm the presence of recombinant DNA sequences in the host cells. Such assays include, for example, "molecular biology" assays well known to those skilled in the art, such as DNA blotting (Southern blotting) and RNA blotting (Northern blotting), RT-PCR, and PCR; and "biochemical" assays, such as detecting the presence or absence of specific peptides, for example, identifying agents within the scope of this invention by immunological methods (ELISA and Western blotting) or by the assays described herein.

[0341] In one embodiment, the present invention relates to a method for expressing two or more target proteins from two or more independent RNA molecule pairs encoding a target protein via ribozymatic cis- and trans-splicing of independent RNA molecule pairs. In one embodiment, the method includes administering one, two, or three pairs of nucleic acid molecules encoding or containing RNA molecules, wherein each individual independent RNA molecule pair has a separate reading frame such that trans-splicing of unwanted RNA molecule pairs does not result in the translation of a full-length functional protein. In one embodiment, the method further includes administering to a cell or tissue one or more substances selected from: nucleic acid molecules encoding a ligase and a ligase. In one embodiment, as described herein, the ligase is an RNA 2',3'-cyclic phosphate and 5'-OH (RtcB) ligase.

[0342] In one embodiment, the present invention includes a method for delivering and expressing a full-length target protein and a cargo sequence. In one embodiment, the method includes administering a first portion of RNA and a second portion of RNA to cells or tissues, the first portion of RNA encoding a first portion of the target protein, ligated to a synthetic intron at its 3' end, and the second portion of RNA encoding a second portion of the target protein, ligated to a synthetic intron at its 5' end. In one embodiment, a 5' ribozyme sequence and a 3' ribozyme sequence are respectively flanked by the synthetic intron. In one embodiment, the synthetic intron contains a cargo sequence located between the 5' and 3' ribozyme sequences. In one embodiment, self-cleaving of the 5' and 3' ribozyme sequences produces three independent RNA molecules: 1) a first fragment containing the first portion of the first RNA encoding the first portion of the target protein; 2) a second fragment containing the synthetic intron; and 3) a third fragment containing the second portion of the second RNA encoding the second portion of the target protein. In one embodiment, the compatible ends of the second fragment are joined together to generate a circular RNA molecule containing a synthetic intron, the synthetic intron containing the cargo sequence. In one embodiment, the first and third fragments are joined together to generate a single full-length linear RNA molecule. In one embodiment, the full-length target protein includes a therapeutic protein, a reporter protein, a recombinase, an antibiotic resistance gene product, an antibody, or a Cas9 protein. In one embodiment, the cargo sequence comprises a therapeutic nucleic acid sequence (e.g., a miRNA sequence or a CRISPR guide RNA sequence) or encodes a therapeutic protein. In some embodiments, the full-length target protein comprises Cas9, while the cargo sequence comprises a guide RNA sequence, thereby targeting Cas9 to a specific genomic sequence for editing. In some embodiments, as described herein, the method includes administering a ligase or a nucleic acid encoding a ligase to cells or tissues.

[0343] In one embodiment, the present invention includes a gene editing method comprising one or more trans-splicing engineered ribozymes. In some embodiments, the method includes administering a first trans-splicing engineered ribozyme and a second trans-splicing engineered ribozyme, wherein the first trans-splicing engineered ribozyme targets upstream of a pathogenic mutation, and the second trans-splicing engineered ribozyme targets downstream of the pathogenic mutation. In some embodiments, trans-splicing upstream and downstream of the pathogenic mutation results in the removal of the pathogenic mutation. In some embodiments, following trans-splicing of the pathogenic mutation, the remaining portion of the gene is trans-spliced ​​together. In some embodiments, the trans-spliced ​​gene is expressed as a functional protein.

[0344] In one embodiment, the present invention relates to a method for assembling a full-length RNA virus genome in vivo. Exemplary RNA viruses include, but are not limited to, coronaviruses, paramyxoviruses, orthomyxoviruses, retroviruses, lentiviruses, alphaviruses, flaviviruses, rhabdoviruses, measles viruses, Newcastle disease viruses, and picornaviruses. In one embodiment, the method includes administering to cells or tissue a first nucleic acid encoding a first portion of the RNA virus genome and encoding a 3' ribozyme. In one embodiment, the method includes administering to cells or tissue a second nucleic acid encoding a second portion of the RNA virus genome and encoding a 5' ribozyme. In one embodiment, the method includes administering to cells or tissue a first RNA molecule comprising the first portion of the RNA virus genome and a 3' ribozyme. In one embodiment, the method includes administering to cells or tissue a second RNA molecule comprising the second portion of the RNA virus genome and a 5' ribozyme. In one embodiment, as described herein, the method includes administering to cells or tissue a nucleic acid or ligase encoding a ligase. In one embodiment, upon cis-cleavage of the 3' and 5' ribozymes, the first portion and the second portion of the RNA virus genome are ligated together to produce a full-length RNA virus genome.

[0345] In one embodiment, the present invention relates to a method for generating a circular RNA molecule encoding a target protein in vivo. In one embodiment, the method includes the step of administering a DNA molecule or an in vitro transcribed linear RNA molecule, the linear RNA molecule comprising a coding region (N-terminal coding sequence) of a first portion encoding the target protein, a 5' ribozyme, an intercalation sequence to be removed, and a 3' ribozyme directly linked to a coding region encoding a second portion (C-terminal coding sequence) of the target protein. In one embodiment, the N-terminal coding sequence is directly linked to the 5' ribozyme sequence, and the 3' ribozyme is directly linked to the C-terminal coding sequence. In some embodiments, the N-terminal and C-terminal coding sequences in the linear RNA molecule are in opposite directions. In some embodiments, the 3' ribozyme and the 5' ribozyme each comprise SEQ ID NO:131, SEQ ID NO:132, SEQ ID NO:133, SEQ ID NO:134, SEQ ID NO:135, SEQ ID NO:136, SEQ ID NO:137, SEQ ID NO:138, SEQ ID NO:139, SEQ ID NO:140, SEQ ID NO:141, SEQ ID NO:142, SEQ ID NO:143, SEQ ID NO:144, SEQ ID NO:145, SEQ ID NO:166, SEQ ID NO:167, SEQ ID NO:168, SEQ ID NO:169, SEQ ID NO:170, SEQ ID NO:171, SEQ ID NO:172, SEQ ID NO:173, SEQ ID NO:174, SEQ ID NO:175, SEQ ID NO:176, SEQ ID NO:192, SEQ ID NO:193, SEQ ID NO:194, SEQ ID NO:195, SEQ ID NO:196, SEQ ID NO:192, SEQ ID NO:193, SEQ ID NO:194, SEQ ID NO:195, SEQ ID NO:196, SEQ ID NO:197, SEQ ID NO:198, SEQ ID NO:19 ... NO: 194, SEQ ID NO: 195, SEQ ID NO: 196, SEQ ID NO: 197, SEQ ID NO: 198, SEQ ID NO: 199, SEQ ID NO: 200, SEQ ID NO: 202, SEQ ID NO: 203, SEQ ID NO: 204, SEQ ID NO: 205, SEQ ID NO: 206, SEQ ID NO: 207, SEQ ID The sequence in NO:217, SEQ ID NO:218 or SEQ ID NO:219.In some embodiments, the 5' ribozyme is SEQ ID NO:131, SEQ ID NO:132, SEQ ID NO:133, SEQ ID NO:134, SEQ ID NO:135, SEQ ID NO:136, SEQ ID NO:137, SEQ ID NO:138, SEQ ID NO:139, SEQ ID NO:140, SEQ ID NO:141, SEQ ID NO:142, SEQ ID NO:143, SEQ ID NO:144, SEQ ID NO:145, SEQ ID NO:166, SEQ ID NO:167, SEQ ID NO:168, SEQ ID NO:169, SEQ ID NO:170, SEQ ID NO:171, SEQ ID NO:172, SEQ ID NO:173, SEQ ID NO:174, SEQ ID NO:175, SEQ ID NO:176, SEQ ID NO:192, SEQ ID NO:193, SEQ ID NO:194, SEQ ID NO:195, SEQ ID NO:196, SEQ ID NO:197, SEQ ID NO:198, SEQ ID NO:199, SEQ ID NO:200, SEQ ID NO:202, SEQ ID NO:203, SEQ ID NO:204, SEQ ID NO:205, SEQ ID NO:206, SEQ ID NO:207, SEQ ID NO:217, SEQ ID NO:218 or SEQ ID NO:219.In some embodiments, the 3' ribozyme is SEQ ID NO:131, SEQ ID NO:132, SEQ ID NO:133, SEQ ID NO:134, SEQ ID NO:135, SEQ ID NO:136, SEQ ID NO:137, SEQ ID NO:138, SEQ ID NO:139, SEQ ID NO:140, SEQ ID NO:141, SEQ ID NO:142, SEQ ID NO:143, SEQ ID NO:144, SEQ ID NO:145, SEQ ID NO:166, SEQ ID NO:167, SEQ ID NO:168, SEQ ID NO:169, SEQ ID NO:170, SEQ ID NO:171, SEQ IDNO:172, SEQ ID NO:173, SEQ ID NO:174, SEQ ID NO:175, SEQ ID NO:176, SEQ ID NO:192, SEQ ID NO:193, SEQ ID NO:194, SEQ ID SEQ ID NO:195, SEQ ID NO:196, SEQ ID NO:197, SEQ ID NO:198, SEQ ID NO:199, SEQ ID NO:200, SEQ ID NO:202, SEQ ID NO:203, SEQ ID NO:204, SEQ ID NO:205, SEQ ID NO:206, SEQ ID NO:207, SEQ ID NO:217, SEQ ID NO:218, or SEQ ID NO:219. In some embodiments, the 5' ribozyme comprises SEQ ID NO:174, and the 3' ribozyme comprises SEQ ID NO:175. In some embodiments, the 5' ribozyme comprises SEQ ID NO:176, and the 3' ribozyme comprises SEQ ID NO:172. In some embodiments, the coding sequence of the target protein is in the opposite direction in the linear RNA molecule. In some embodiments, the linear RNA molecule further comprises IRES, a polyadenylated sequence, an AK recombinant sequence (SEQ ID NO:180), or any combination thereof.

[0346] In one embodiment, the method includes the following steps: providing a linear RNA molecule comprising a 3' ribozyme, a coding region encoding a second part of a target protein (C-terminal coding sequence), a coding region encoding a first part of the target protein (N-terminal coding sequence), and a 5' ribozyme. In one embodiment, the N-terminal coding sequence is directly linked to the 5' ribozyme sequence, and the 3' ribozyme is directly linked to the C-terminal coding sequence. In some embodiments, the N-terminal and C-terminal coding sequences in the linear RNA molecule are in opposite directions. In some embodiments, the 3' ribozyme and the 5' ribozyme each comprise SEQ ID NO:131, SEQ ID NO:132, SEQ ID NO:133, SEQ ID NO:134, SEQ ID NO:135, SEQ ID NO:136, SEQ ID NO:137, SEQ ID NO:138, SEQ ID NO:139, SEQ ID NO:140, SEQ ID NO:141, SEQ ID NO:142, SEQ ID NO:143, SEQ ID NO:144, SEQ ID NO:145, SEQ ID NO:166, SEQ ID NO:167, SEQ ID NO:168, SEQ ID NO:169, SEQ ID NO:170, SEQ ID NO:171, SEQ ID NO:172, SEQ ID NO:173, SEQ ID NO:174, SEQ ID NO:175, SEQ ID NO:176, SEQ ID NO:192, SEQ ID NO:193, SEQ ID NO:194, SEQ ID NO:195, SEQ ID NO:196, SEQ ID NO:192, SEQ ID NO:193, SEQ ID NO:194, SEQ ID NO:195, SEQ ID NO:196, SEQ ID NO:197, SEQ ID NO:168, SEQ ID NO:169, SEQ ID NO:194, SEQ ID NO:195, SEQ ID NO:196, SEQ ID NO:197, SEQ ID NO:198, SEQ ID NO:19 ... The sequences are SEQ ID NO:194, SEQ ID NO:195, SEQ ID NO:196, SEQ ID NO:197, SEQ ID NO:198, SEQ ID NO:199, SEQ ID NO:200, SEQ ID NO:202, SEQ ID NO:203, SEQ ID NO:204, SEQ ID NO:205, SEQ ID NO:206, SEQ ID NO:207, SEQ ID NO:217, SEQ ID NO:218, or SEQ ID NO:219. In some embodiments, the 3' ribozyme comprises SEQ ID NO:171, and the 5' ribozyme comprises SEQ ID NO:170. In some embodiments, the 3' ribozyme comprises SEQ ID NO:169, and the 5' ribozyme comprises SEQ ID NO:168. In some embodiments, the coding sequence of the target protein is in the opposite direction in the linear RNA molecule.In some embodiments, the linear RNA molecule also includes IRES, a polyadenylated sequence, an AK recombinant sequence (SEQ ID NO:180), or any combination thereof.

[0347] in vitro

[0348] In one embodiment, the present invention includes a method for generating RNA molecules encoding a target protein in vitro. In one embodiment, the method includes the step of providing at least two RNA molecules. In one embodiment, the step includes providing a first RNA molecule and a second RNA molecule.

[0349] In one embodiment, the first RNA molecule includes a coding region encoding a first portion of the target protein. In one embodiment, the first RNA molecule includes a 3' ribozyme. In one embodiment, the first RNA molecule includes a coding region encoding a first portion of the target protein and a 3' ribozyme.

[0350] In one embodiment, the second RNA molecule includes a coding region encoding a second portion of the target protein. In one embodiment, the second RNA molecule includes a 5' ribozyme. In one embodiment, the second RNA molecule includes a coding region encoding a second portion of the target protein and a 5' ribozyme.

[0351] In one embodiment, the method for generating RNA molecules encoding a target protein in vitro further includes providing a ligase. In one embodiment, the ligase induces the assembly of RNA molecules from the coding regions of a first RNA molecule and a second RNA molecule. In one embodiment, as described herein, the ligase is an RNA 2',3'-cyclic phosphate and 5'-OH (RtcB) ligase.

[0352] In one embodiment, the present invention includes a method for generating an RNA molecule encoding a multi-domain target protein in vitro. In one embodiment, the method includes the steps of: a) providing a first RNA molecule, b) providing one or more additional RNA molecules, c) providing a ribozyme, and d) providing a final RNA molecule.

[0353] In one embodiment, the first RNA molecule in step a) comprises a coding region encoding a first portion of the target protein. In one embodiment, the first RNA molecule comprises a 3' ribozyme. In one embodiment, the first RNA molecule comprises a coding region encoding a first portion of the target protein and a 3' ribozyme. In one embodiment, the 3' ribozyme catalyzes its own separation from the first RNA molecule, thereby generating a 3'P or 2'3' cP terminus. In one embodiment, the first RNA molecule further comprises a 5' tag. In one embodiment, the 5' tag mediates the attachment of the first RNA molecule to a solid support.

[0354] In one embodiment, at least one additional RNA molecule in step b) comprises a coding region encoding a target protein domain; a 5' ribozyme; and a 3' ribozyme recognition sequence. In one embodiment, the 5' ribozyme self-cleaves to generate a 5' OH terminus. In one embodiment, a ligase is provided to catalyze the ligation of the first RNA molecule to one or more additional RNA molecules. In one embodiment, as described herein, the ligase is an RNA 2',3'-cyclic phosphate and 5'-OH (RtcB) ligase. In one embodiment, as described herein, the 3' ribozyme recognition sequence comprises a VS-S sequence.

[0355] In one embodiment, as described herein, the ribozyme in step c) comprises VS-Rz. In one embodiment, the VS-Rz recognizes VS-S and mediates its cleavage from one or more additional RNA molecules. In one embodiment, the cleavage produces a 3'P or 2'3'cP end. In one embodiment, steps b) through c) are repeated at least once to generate an RNA molecule encoding multiple domains. In one embodiment, the VS-Rz is removed prior to repeating step b).

[0356] In one embodiment, the RNA molecule in step d) contains a coding region encoding the final portion of the target protein. In one embodiment, the RNA molecule encoding the final portion of the target protein also contains a 5' ribozyme. In one embodiment, the RNA molecule contains a coding region encoding the final portion of the target protein directly linked to the 5' ribozyme. In one embodiment, the 5' ribozyme catalyzes its detachment from the final RNA molecule, thereby generating a 5' OH terminus. In one embodiment, a ligase is provided to catalyze the ligation of one or more additional RNA molecules to the RNA molecule encoding the final portion of the target protein, thereby generating a complete RNA molecule encoding an N-terminal domain, one or more additional domains, and a C-terminal domain. In one embodiment, as described herein, the ligase is an RNA 2',3'-cyclic phosphate and 5'-OH (RtcB) ligase.

[0357] In one embodiment, the present invention includes a method for generating a circular RNA molecule encoding a target protein in vitro. In one embodiment, the method includes the steps of: providing a plasmid or vector comprising a coding region of a first portion (N-terminal coding sequence) encoding the target protein, a 5' ribozyme, an intercalation sequence to be removed, and a 3' ribozyme directly linked to a coding region of a second portion (C-terminal coding sequence) encoding the target protein. In one embodiment, the N-terminal coding sequence is directly linked to the 5' ribozyme sequence, and the 3' ribozyme is directly linked to the C-terminal coding sequence. In some embodiments, the N-terminal and C-terminal coding sequences in the linear RNA molecule are in opposite directions. In some embodiments, the 3' ribozyme and the 5' ribozyme each comprise the sequence from SEQ ID NO:131, SEQ ID NO:132, SEQ ID NO:133, SEQ ID NO:134, SEQ ID NO:135, SEQ ID NO:136, SEQ ID NO:137, SEQ ID NO:138, SEQ ID NO:139, SEQ ID NO:140, SEQ ID NO:141, SEQ ID NO:142, SEQ ID NO:143, SEQ ID NO:144, SEQ ID NO:145, SEQ ID NO:166, SEQ ID NO:167, SEQ ID NO:168, SEQ ID NO:169, SEQ ID NO:170, SEQ ID NO:171, SEQ ID NO:172, SEQ ID NO:173, SEQ ID NO:174, SEQ ID NO:175, or SEQ ID NO:176. In some embodiments, the 5' ribozyme is SEQ ID NO:168, SEQ ID NO:170, SEQ ID NO:174, or SEQ ID NO:176. In some embodiments, the 3' ribozyme is SEQ ID NO:169, SEQ ID NO:171, SEQ ID NO:172, or SEQ ID NO:174. In some embodiments, the 5' ribozyme comprises SEQ ID NO:174, and the 3' ribozyme comprises SEQ ID NO:175. In some embodiments, the 5' ribozyme comprises SEQ ID NO:176, and the 3' ribozyme comprises SEQ ID NO:172. In some embodiments, the coding sequence of the target protein is in the opposite direction in the linear RNA molecule. In some embodiments, the linear RNA molecule further comprises IRES, a polyadenylated sequence, an AK recombinant sequence (SEQ ID NO:180), or any combination thereof.

[0358] In one embodiment, the method includes providing a linear RNA molecule comprising a 3' ribozyme, a coding region encoding a second portion of a target protein (C-terminal coding sequence), a coding region encoding a first portion of the target protein (N-terminal coding sequence), and a 5' ribozyme. In one embodiment, the N-terminal coding sequence is directly linked to the 5' ribozyme sequence, and the 3' ribozyme is directly linked to the C-terminal coding sequence. In some embodiments, the N-terminal and C-terminal coding sequences in the linear RNA molecule are in opposite directions. In some embodiments, the 3' ribozyme and the 5' ribozyme each comprise SEQ ID NO:131, SEQ ID NO:132, SEQ ID NO:133, SEQ ID NO:134, SEQ ID NO:135, SEQ ID NO:136, SEQ ID NO:137, SEQ ID NO:138, SEQ ID NO:139, SEQ ID NO:140, SEQ ID NO:141, SEQ ID NO:142, SEQ ID NO:143, SEQ ID NO:144, SEQ ID NO:145, SEQ ID NO:166, SEQ ID NO:167, SEQ ID NO:168, SEQ ID NO:169, SEQ ID NO:170, SEQ ID NO:171, SEQ ID NO:172, SEQ ID NO:173, SEQ ID NO:174, SEQ ID NO:175, SEQ ID NO:176, SEQ ID NO:192, SEQ ID NO:193, SEQ ID NO:194, SEQ ID NO:195, SEQ ID NO:196, SEQ ID NO:192, SEQ ID NO:193, SEQ ID NO:194, SEQ ID NO:195, SEQ ID NO:196, SEQ ID NO:197, SEQ ID NO:198, SEQ ID NO:19 ... NO:194, SEQ ID NO:195, SEQ ID NO:196, SEQ ID NO:197, SEQ ID NO:198, SEQ ID NO:199, SEQ ID NO:200, SEQ ID NO:202, SEQ ID NO:203, SEQ ID NO:204, SEQ ID NO:205, SEQ ID NO:206, SEQ ID NO:207, SEQ ID The sequence in NO:217, SEQ ID NO:218 or SEQ ID NO:219.In some embodiments, the 5' ribozyme is SEQ ID NO:131, SEQ ID NO:132, SEQ ID NO:133, SEQ ID NO:134, SEQ ID NO:135, SEQ ID NO:136, SEQ ID NO:137, SEQ ID NO:138, SEQ ID NO:139, SEQ ID NO:140, SEQ ID NO:141, SEQ ID NO:142, SEQ ID NO:143, SEQ ID NO:144, SEQ ID NO:145, SEQ ID NO:166, SEQ ID NO:167, SEQ ID NO:168, SEQ ID NO:169, SEQ ID NO:170, SEQ ID NO:171, SEQ ID NO:172, SEQ ID NO:173, SEQ ID NO:174, SEQ ID NO:175, SEQ ID NO:176, SEQ ID NO:192, SEQ ID NO:193, SEQ ID NO:194, SEQ ID NO:195, SEQ ID NO:19; SEQ ID NO:197, SEQ ID NO:198, SEQ ID NO:199, SEQ ID NO:200, SEQ ID NO:202, SEQ ID NO:203, SEQ ID NO:204, SEQ ID NO:205, SEQ ID NO:206, SEQ ID NO:207, SEQ ID NO:217, SEQ ID NO:218 or SEQ ID NO:219.In some embodiments, the 3' ribozyme is SEQ ID NO:131, SEQ ID NO:132, SEQ ID NO:133, SEQ ID NO:134, SEQ ID NO:135, SEQ ID NO:136, SEQ ID NO:137, SEQ ID NO:138, SEQ ID NO:139, SEQ ID NO:140, SEQ ID NO:141, SEQ ID NO:142, SEQ ID NO:143, SEQ ID NO:144, SEQ ID NO:145, SEQ ID NO:166, SEQ ID NO:167, SEQ ID NO:168, SEQ ID NO:169, SEQ ID NO:170, SEQ ID NO:171, SEQ ID NO:172, SEQ ID NO:173, SEQ ID NO:174, SEQ ID NO:175, SEQ ID NO:176, SEQ ID NO:192, SEQ ID NO:193, SEQ ID NO:194, SEQ ID SEQ ID NO:195, SEQ ID NO:196, SEQ ID NO:197, SEQ ID NO:198, SEQ ID NO:199, SEQ ID NO:200, SEQ ID NO:202, SEQ ID NO:203, SEQ ID NO:204, SEQ ID NO:205, SEQ ID NO:206, SEQ ID NO:207, SEQ ID NO:217, SEQ ID NO:218, or SEQ ID NO:219. In some embodiments, the 3' ribozyme comprises SEQ ID NO:171, and the 5' ribozyme comprises SEQ ID NO:170. In some embodiments, the 3' ribozyme comprises SEQ ID NO:169, and the 5' ribozyme comprises SEQ ID NO:168. In some embodiments, the coding sequence of the target protein is reversed in the linear RNA molecule. In some embodiments, the linear RNA molecule further comprises IRES, a polyadenylated sequence, an AK recombinant sequence (SEQ ID NO:180), or any combination thereof.

[0359] Any RNA molecule disclosed herein can be transcribed in vitro from template DNA (referred to as the "in vitro transcription template"). The DNA source can be, for example, genomic DNA, plasmid DNA, phage DNA, cDNA, synthetic DNA sequences, or any other suitable DNA source. In some embodiments, the in vitro transcription template encodes a 5' untranslated (UTR) region containing an open reading frame and encodes a 3' UTR and a polyA tail. In some embodiments, the in vitro transcription template lacks a polyA tail. The specific nucleic acid sequence composition and length of the in vitro transcription template will depend on the mRNA or a fragment thereof encoded by the template (e.g., an N-terminal or C-terminal fragment).

[0360] In one implementation, the length of the 5' UTR is between 0 and 3000 nucleotides. The lengths of the 5' and 3' UTR sequences to be added to the coding region can be varied by various methods, including but not limited to designing PCR primers that anneal to regions different from the UTR. In this way, those skilled in the art can modify the desired 5' and 3' UTR lengths to achieve optimal translation efficiency after transfection of transcribed RNA.

[0361] The 5' and 3' UTRs can be naturally occurring endogenous 5' and 3' UTRs of the target gene. Alternatively, non-endogenous UTR sequences of the target gene can be added by incorporating the UTR sequences into the forward and reverse primers, or by making any other modifications to the template. Using non-endogenous UTR sequences can be used to alter RNA stability and / or translation efficiency. For example, it is known that AU-rich elements in the 3' UTR sequence reduce mRNA stability. Therefore, 3' UTRs can be selected or designed to improve the stability of transcribed RNA based on UTR characteristics known in the art.

[0362] In one embodiment, the 5'UTR may contain the Kozak sequence of an endogenous gene. Alternatively, when a non-target gene endogenous 5'UTR is added via PCR as described above, a shared Kozak sequence can be redesigned by adding this 5'UTR sequence. Kozak sequences can improve the translation efficiency of some RNA transcripts, but not all RNAs require a Kozak sequence for efficient translation. Many mRNAs are known in the art to require a Kozak sequence. In other embodiments, the 5'UTR may be derived from an RNA virus whose RNA genome is stably present in the cell. In other embodiments, various nucleotide analogs may be used in the 3' or 5'UTR to prevent mRNA degradation by exonucleases.

[0363] To enable RNA synthesis from a DNA template, a transcription promoter should be ligated upstream of the DNA template containing the sequence to be transcribed. When a sequence serving as an RNA polymerase promoter is added to the 5' end of a forward primer, the RNA polymerase promoter will be incorporated upstream of the PCR product of the open reading frame to be transcribed. In one embodiment, the promoter is the T7 RNA polymerase promoter, as described elsewhere herein. Other useful promoters include, but are not limited to, the T3 and SP6 RNA polymerase promoters. The common nucleotide sequences of the T7, T3, and SP6 promoters are known in the art.

[0364] In one implementation, the mRNA possesses both a 5' cap and a 3' poly(A) tail, which determines ribosome binding, translation initiation, and stability within the cell. On circular DNA templates (e.g., plasmid DNA), RNA polymerase produces a long concatameric product unsuitable for expression in eukaryotic cells. Plasmid DNA linearized at the 3' UTR end transcribes to produce normal-sized mRNA, which, when polyadenylated post-transcriptionally, allows for efficient transfection into eukaryotic cells.

[0365] On a linear DNA template, phage T7 RNA polymerase can extend the 3' end of the transcript beyond the last base of the template (Schenborn and Mierendorf, Nuc Acids Res., 13:6223-36 (1985); Nacheva and Berzal-Herranz, Eur. J. Biochem., 270:1485-65 (2003)).

[0366] The conventional method for integrating polyA / T fragments into a DNA template is molecular cloning. However, polyA / T sequences integrated into plasmid DNA can lead to plasmid instability, which can be mitigated by using recombinant agonistant bacterial cells for plasmid proliferation.

[0367] The poly(A) tail of RNA can be further extended post-transcriptionally using poly(A) polymerases (e.g., E. coli poly(A) polymerase (E-PAP) or yeast poly(A) polymerase). In one embodiment, increasing the length of the poly(A) tail from 100 nucleotides to 300-400 nucleotides can improve RNA translation efficiency by approximately two-fold. Furthermore, linking different chemical groups to the 3' end can improve mRNA stability. Such links can include modified / artificial nucleotides, aptamers, and other compounds. For example, ATP analogs can be incorporated into the poly(A) tail using poly(A) polymerase. ATP analogs can further improve RNA stability.

[0368] The 5' cap also provides stability to the mRNA molecule. In one embodiment, the RNA produced by said method contains a 5' cap1 structure. Such a cap1 structure can be generated using a vaccinia capping enzyme and a 2'-O-methyltransferase (CellScript, Madison, WI, USA). Alternatively, the 5' cap can be provided using techniques known in the art and as described herein (Cougot, et al., Trends in Biochem. Sci., 29:436-444 (2001); Stepinski, et al., RNA, 7:1468-95 (2001); Elango, et al., Biochim. Biophys. Res. Commun., 330:958-966 (2005)).

[0369] Some embodiments of the present invention may utilize a solid support consisting of an inert substrate or matrix (e.g., glass slides, polymer beads, etc.) that has been functionalized, for example by applying an intermediate material layer or coating containing reactive groups that allow covalent linkage with biomolecules (e.g., polynucleotides). Examples of such supports include, but are not limited to, polyacrylamide hydrogels loaded on an inert substrate (e.g., glass), particularly the polyacrylamide hydrogels described in WO2005 / 065814 and US2008 / 0280773, the entire contents of which are incorporated herein by reference. In such embodiments, biomolecules (e.g., polynucleotides) may be directly covalently linked to the intermediate material (e.g., hydrogel), but the intermediate material itself may be non-covalently linked to the substrate or matrix (e.g., glass substrate). Accordingly, the term "covalently linked to a solid support" should be interpreted to encompass this type of arrangement.

[0370] Those skilled in the art will understand that the number of possible substrates is enormous. Possible substrates include, but are not limited to: glass and modified or functionalized glass, plastics (including acrylic resins, polystyrene and copolymers of styrene and other materials, polypropylene, polyethylene, polybutene, polyurethane, Teflon) TM (etc.), polysaccharides, nylon or nitrocellulose, ceramics, resins, silica or silica-based materials (including silicon and modified silicon), carbon, metals, inorganic glass, plastics, fiber bundles and various other polymers.

[0371] In some embodiments, the solid support comprises microspheres or beads. Suitable bead compositions include, but are not limited to, plastics, ceramics, glass, polystyrene, methylstyrene, acrylic polymers, paramagnetic materials, thorium oxide sol, carbon graphite, titanium dioxide, latex or cross-linked dextran (e.g., agarose gel), cellulose, nylon, cross-linked micelles, and Teflon, as well as any other materials for solid supports outlined herein. The Microsphere Detection Guide from Bangs Laboratories, Fishers Inc., Indiana, USA, is a useful guide. In some embodiments, the microspheres are magnetic microspheres or beads.

[0372] The beads do not have to be spherical; irregular particles can be used. Alternatively, the beads can be porous. The size of the beads ranges from nanometers (i.e., 100 nanometers) to millimeters (i.e., 1 millimeter), preferably beads from about 0.2 micrometers to about 200 micrometers, particularly preferably beads from about 0.5 micrometers to about 5 micrometers, although smaller or larger beads can be used in some embodiments.

[0373] In one embodiment, the present invention relates to a method for assembling a full-length RNA virus genome in vitro. Exemplary RNA viruses include, but are not limited to, coronaviruses, paramyxoviruses, orthomyxoviruses, retroviruses, lentiviruses, alphaviruses, flaviviruses, rhabdoviruses, measles viruses, Newcastle disease viruses, and picornaviruses. In one embodiment, the method includes providing a first RNA molecule comprising a first portion of an RNA virus genome and a 3' ribozyme. In one embodiment, the method includes providing a second RNA molecule comprising a second portion of an RNA virus genome and a 5' ribozyme. In one embodiment, as described herein, the first and second portions of the RNA virus genome have compatible ends for ligation upon cis-cleavage of the 3' and 5' ribozymes. In one embodiment, as described herein, the method includes contacting the first and second RNA molecules with a ligase to generate a full-length RNA virus genome.

[0374] Treatment and Uses

[0375] This invention provides methods for treating diseases or disorders in subjects, alleviating their symptoms, and / or reducing their risk of disease. For example, in one embodiment, the method of this invention treats diseases or disorders in mammals, alleviating their symptoms, and / or reducing their risk of disease. In one embodiment, the method of this invention treats diseases or disorders in plants, alleviating their symptoms, and / or reducing their risk of disease. In one embodiment, the method of this invention treats diseases or disorders in yeast organisms, alleviating their symptoms, and / or reducing their risk of disease.

[0376] In one embodiment, the subject is a cell. In one embodiment, the cell is a prokaryotic or eukaryotic cell. In one embodiment, the cell is a eukaryotic cell. In one embodiment, the cell is a plant, animal, or fungal cell. In one embodiment, the cell is a plant cell. In one embodiment, the cell is an animal cell. In one embodiment, the cell is a yeast cell.

[0377] In one embodiment, the subject is a mammal. For example, in one embodiment, the subject is a human, a non-human primate, a dog, a cat, a horse, a cow, a goat, a sheep, a rabbit, a pig, a rat, or a mouse. In one embodiment, the subject is a non-mammal subject. For example, in one embodiment, the subject is a zebrafish, a fruit fly, or a roundworm.

[0378] In one embodiment, the disease or disorder is caused by a missing or defective protein whose nucleic acid sequence exceeds the packaging size of a viral vector. Therefore, in one embodiment, the compositions, systems, and methods of the present invention can be used to treat, alleviate, or reduce the risk of a disease or disorder. Thus, in one embodiment, the method includes administering one or more compositions of the present invention to a subject. Furthermore, in one embodiment, the method includes using one or more systems of the present invention to treat a subject's disease or disorder, alleviate its symptoms, and / or reduce its risk of developing the disease.

[0379] In one embodiment, the disease or disorder is Duchenne muscular dystrophy, Becker's muscular dystrophy (BMD), autosomal recessive polycystic kidney disease, hemophilia A, Stargardt's macular degeneration, limb-girdle muscular dystrophy, autosomal recessive deep prelingual hearing loss, autosomal recessive nonsyndromic hearing loss (ARNSHL), sensorineural hearing loss, cystic fibrosis, Wilson's disease, Trihyo's myopathy, autosomal recessive deafness-9 (DFNB9), Usher syndrome type I, GJB2-related autosomal recessive nonsyndromic hearing loss (GJB2-AR). NSHL), type 3 autosomal recessive cerebellar cortical disorder, nonsyndromic hearing loss, autosomal recessive deafness-16 (DFNB16), Meniere's disease (MD), autosomal dominant nonsyndromic sensorineural hearing loss-12 (DFNA12), autosomal recessive spinocerebellar ataxia-21 (SCAR21), Usher syndrome type 1F (USH1F), autosomal recessive deafness-23 (DFNB23), autosomal recessive deafness-30 (DFNB30), macrophage dysplasia of the vertebral body (OSMED), autosomal recessive deafness-77 (DFNB77), autosomal recessive deafness-84A (DFNB84A), autosomal recessive deafness-84B (DFN B84B), peripheral neuropathy, myopathy, autosomal dominant nonsyndromic deafness-4A (DFNA4), congenital thrombocytopenia, sensorineural hearing loss, autosomal dominant nonsyndromic hearing loss-56 (DFNA56), epileptic encephalopathy, Timothy syndrome, long QT syndrome, X-linked retinal disease, aldosteronism, autosomal recessive deafness-42 (DFNB42), primary aldosteronism (Conn syndrome), seizures, neurological abnormalities, sinoatrial node dysfunction, neurodevelopmental disorders, hypokalemic periodic paralysis, epilepsy, developmental and epileptic encephalopathy, Brody's myopathy, Darier's disease, heart disease, von Willebrand disease, or brain-hepatorenal syndrome. In one embodiment, the disease or disorder is caused by a gene mutation suitable for CRISPR-Cas9-mediated editing.

[0380] In one embodiment, the method of the present invention includes administering a composition to a subject suffering from Duchenne muscular dystrophy, the composition comprising a first nucleic acid and a second nucleic acid, the first nucleic acid comprising a coding region encoding a first portion of dystrophin and a 3' ribozyme, the second nucleic acid comprising a coding region encoding a second portion of dystrophin and a 5' ribozyme, wherein the first nucleic acid transcribes a first RNA molecule, and the second nucleic acid transcribes a second RNA molecule, and wherein cis-splicing of the 3' and 5' ribozymes and trans-splicing of the coding regions encoding the first and second portions of dystrophin produce a single RNA molecule encoding the full-length dystrophin.

[0381] In one embodiment, the method of the present invention includes administering a composition to a subject suffering from Duchenne muscular dystrophy, the composition comprising a first nucleic acid encoding a nucleic acid sequence of SEQ ID NO:129 and a second nucleic acid encoding a nucleic acid sequence of SEQ ID NO:130, wherein transcription of the first nucleic acid produces a first RNA molecule, and transcription of the second nucleic acid molecule produces a second RNA molecule, and wherein cis-cleavage of 3' and 5' ribozymes and trans-splicing of the first and second RNA molecules produce a single RNA molecule encoding dystrophin.

[0382] In one embodiment, the method of the present invention includes administering a composition to a subject suffering from Duchenne muscular dystrophy, the composition comprising a first nucleic acid encoding a nucleic acid sequence of SEQ ID NO:22 and a second nucleic acid encoding a nucleic acid sequence of SEQ ID NO:23, wherein the first nucleic acid is transcribed into a first RNA molecule, and the second nucleic acid is transcribed into a second RNA molecule, and wherein cis-splicing of 3' and 5' ribozymes and trans-splicing of the first and second RNA molecules produce a single RNA molecule encoding a full-length dystrophin protein having a C-terminal GFP reporter gene. In one embodiment, the second nucleic acid encodes a fragment of SEQ ID NO:23, wherein the fragment does not contain the coding sequence of the C-terminal GFP reporter gene.

[0383] In one embodiment, the method includes administering a composition to a subject suffering from Duchenne muscular dystrophy, the composition comprising a first RNA molecule encoding a first portion of dystrophin and containing a 3' ribozyme, and a second RNA molecule encoding a second portion of dystrophin and containing a 5' ribozyme, wherein cis-splicing of the 3' and 5' ribozymes and trans-splicing of the first and second RNA molecules produce a single RNA molecule encoding dystrophin.

[0384] In one embodiment, the method includes administering a composition to a subject suffering from Duchenne muscular dystrophy, the composition comprising a first RNA molecule and a second RNA molecule, the first RNA molecule comprising the nucleic acid sequence of SEQ ID NO:129 and the second RNA molecule comprising the nucleic acid sequence of SEQ ID NO:130, wherein cis-cleavage of 3' and 5' ribozymes and trans-splicing of the first and second RNA molecules produce a single RNA molecule encoding miniDystrophin.

[0385] In one embodiment, the method includes administering a composition to a subject suffering from Duchenne muscular dystrophy, the composition comprising: a first RNA molecule containing the nucleic acid sequence of SEQ ID NO:22; and a second RNA molecule containing the nucleic acid sequence of SEQ ID NO:23, wherein cis-cleavage of the 3' and 5' ribozymes and trans-splicing of the first and second RNA molecules generate a single RNA molecule encoding a full-length dystrophin protein having a C-terminal GFP reporter gene. In one embodiment, the second nucleic acid encodes a fragment of SEQ ID NO:23, wherein the fragment does not contain the coding sequence of the C-terminal GFP reporter gene.

[0386] In one embodiment, the method of the present invention includes administering a composition to a subject suffering from one or more diseases selected from Table 1, the composition comprising a first nucleic acid and a second nucleic acid, the first nucleic acid comprising a coding region encoding a first portion of a therapeutic protein corresponding to the relevant disease in Table 1 and a 3' ribozyme, the second nucleic acid comprising a coding region encoding a second portion of a therapeutic protein corresponding to the relevant disease in Table 1 and a 5' ribozyme, wherein the first nucleic acid transcribes a first RNA molecule, and the second nucleic acid transcribes a second RNA molecule, and wherein cis-splicing of the 3' and 5' ribozymes and trans-splicing of the coding regions encoding the first and second portions of the therapeutic protein generate a single RNA molecule encoding a full-length therapeutic protein.

[0387] In one embodiment, the method includes administering a composition to a subject suffering from one or more diseases selected from Table 1, the composition comprising a first RNA molecule and a second RNA molecule, the first RNA molecule encoding a first portion of a therapeutic protein corresponding to the relevant disease in Table 1 and containing a 3' ribozyme, the second RNA molecule encoding a second portion of a therapeutic protein corresponding to the relevant disease in Table 1 and containing a 5' ribozyme, wherein cis-splicing of the 3' and 5' ribozymes and trans-splicing of the first and second RNA molecules generate a single RNA molecule encoding a full-length therapeutic protein.

[0388]

[0389]

[0390] In one embodiment, the method of the present invention includes administering a composition to a subject suffering from Duchenne muscular dystrophy, the composition comprising a first nucleic acid and a second nucleic acid, the first nucleic acid comprising the nucleic acid sequence of SEQ ID NO:150, the second nucleic acid comprising the nucleic acid sequence of SEQ ID NO:151, wherein transcription of the first nucleic acid produces a first RNA molecule, and transcription of the second nucleic acid molecule produces a second RNA molecule, and wherein cis-cleavage of 3' and 5' ribozymes and trans-splicing of the first and second RNA molecules generate a single RNA molecule encoding mini dystrophin.

[0391] In one embodiment, the method of the present invention includes administering a composition to a subject suffering from Duchenne muscular dystrophy, the composition comprising a first nucleic acid and a second nucleic acid, the first nucleic acid comprising the nucleic acid sequence of SEQ ID NO:152, the second nucleic acid comprising the nucleic acid sequence of SEQ ID NO:153, wherein transcription of the first nucleic acid produces a first RNA molecule, and transcription of the second nucleic acid molecule produces a second RNA molecule, and wherein cis-cleavage of 3' and 5' ribozymes and trans-splicing of the first and second RNA molecules produce a single RNA molecule encoding mini dystrophin.

[0392] In one embodiment, the method of the present invention includes administering a composition to a subject suffering from Duchenne muscular dystrophy, the composition comprising a nucleic acid molecule containing the nucleic acid sequence of SEQ ID NO:162, wherein transcription of said nucleic acid produces a linear RNA molecule, and cis-cleavage of said 3' and 5' ribozymes and cis-ligation of said RNA molecule produce a single circular RNA molecule encoding dystrophin.

[0393] In one embodiment, the method of the present invention includes administering a composition to a subject suffering from Duchenne muscular dystrophy, the composition comprising a nucleic acid molecule containing the nucleic acid sequence of SEQ ID NO:163, wherein transcription of the nucleic acid produces a linear RNA molecule, and wherein cis-cleavage of 3' and 5' ribozymes and cis-ligation of the RNA molecule generate a single circular RNA molecule encoding dystrophin.

[0394] In one embodiment, the method of the present invention includes administering a composition to a subject suffering from Duchenne muscular dystrophy, the composition comprising a nucleic acid molecule containing the nucleic acid sequence of SEQ ID NO:164, wherein transcription of said nucleic acid produces a linear RNA molecule, and cis-cleavage of said 3' and 5' ribozymes and cis-ligation of said RNA molecule produce a single circular RNA molecule encoding dystrophin.

[0395] In one embodiment, the method of the present invention includes administering a composition to a subject suffering from Duchenne muscular dystrophy, the composition comprising a nucleic acid molecule containing the nucleic acid sequence of SEQ ID NO:165, wherein transcription of the nucleic acid produces a linear RNA molecule, and wherein cis-cleavage of 3' and 5' ribozymes and cis-ligation of the RNA molecule generate a single circular RNA molecule encoding dystrophin.

[0396] In one embodiment, the method of the present invention includes administering a composition to a subject suffering from Dysferlinopathy or myopathy or mitochondrial dysfunction due to dysferlin mutation, the composition comprising: a first nucleic acid comprising the nucleic acid sequence of SEQ ID NO:157; and a second nucleic acid comprising the nucleic acid sequence of SEQ ID NO:158, wherein transcription of the first nucleic acid produces a first RNA molecule, and transcription of the second nucleic acid molecule produces a second RNA molecule, and wherein cis-cleavage of 3' and 5' ribozymes and trans-splicing of the first and second RNA molecules generate a single RNA molecule encoding the dysferlin protein.

[0397] In one embodiment, the method of the present invention includes administering a composition to a subject suffering from Dysferlin proteinopathy or myopathy or mitochondrial dysfunction due to dysferlin mutation, the composition comprising a first nucleic acid and a second nucleic acid, the first nucleic acid comprising the nucleic acid sequence of SEQ ID NO:177, the second nucleic acid comprising the nucleic acid sequence of SEQ ID NO:178, wherein transcription of the first nucleic acid produces a first RNA molecule, and transcription of the second nucleic acid molecule produces a second RNA molecule, and wherein cis-cleavage of 3' and 5' ribozymes and trans-splicing of the first and second RNA molecules generate a single RNA molecule encoding the dysferlin protein.

[0398] In one embodiment, the method of the present invention includes administering a composition to a subject suffering from Dysferlin proteinopathy or myopathy or mitochondrial dysfunction due to dysferlin mutation, the composition comprising: a first nucleic acid comprising the nucleic acid sequence of SEQ ID NO: 179; and a second nucleic acid comprising the nucleic acid sequence of SEQ ID NO: 178, wherein transcription of the first nucleic acid produces a first RNA molecule, and transcription of the second nucleic acid molecule produces a second RNA molecule, and wherein cis-cleavage of 3' and 5' ribozymes and trans-splicing of the first and second RNA molecules generate a single RNA molecule encoding the dysferlin protein.

[0399] In one embodiment, the method of the present invention includes administering a composition to a subject suffering from Dysferlin proteinopathy or myopathy or mitochondrial dysfunction due to dysferlin mutation, the composition comprising a first nucleic acid and a second nucleic acid, the first nucleic acid comprising the nucleic acid sequence of SEQ ID NO:181, the second nucleic acid comprising the nucleic acid sequence of SEQ ID NO:182, wherein transcription of the first nucleic acid produces a first RNA molecule, and transcription of the second nucleic acid molecule produces a second RNA molecule, and wherein cis-cleavage of 3' and 5' ribozymes and trans-splicing of the first and second RNA molecules generate a single RNA molecule encoding the dysferlin protein.

[0400] In one embodiment, the method of the present invention includes administering a composition to a subject suffering from Dysferlin proteinopathy or myopathy or mitochondrial dysfunction due to dysferlin mutation, the composition comprising: a first nucleic acid comprising the nucleic acid sequence of SEQ ID NO: 183; and a second nucleic acid comprising the nucleic acid sequence of SEQ ID NO: 182, wherein transcription of the first nucleic acid produces a first RNA molecule, and transcription of the second nucleic acid molecule produces a second RNA molecule, and wherein cis-cleavage of 3' and 5' ribozymes and trans-splicing of the first and second RNA molecules generate a single RNA molecule encoding the dysferlin protein.

[0401] In one embodiment, the method of the present invention includes administering a composition to a subject suffering from autosomal recessive nonsyndromic sensorineural hearing loss-16 or deafness due to a sterocilin (STRC) mutation, the composition comprising a first nucleic acid containing the nucleic acid sequence of SEQ ID NO:160 and a second nucleic acid containing the nucleic acid sequence of SEQ ID NO:161, wherein transcription of the first nucleic acid produces a first RNA molecule, and transcription of the second nucleic acid molecule produces a second RNA molecule, and wherein cis-cleavage of the 3' and 5' ribozymes and trans-splicing of the first and second RNA molecules generate a single RNA molecule encoding a sterocilin protein.

[0402] Experimental Examples

[0403] The present invention will be further described in detail through the following experimental embodiments. These embodiments are for illustrative purposes only and are not intended to be limiting unless otherwise stated. Therefore, the present invention should not be construed as limited to the following embodiments, but should be understood to cover any and all variations that are obvious from the teachings provided herein.

[0404] Without further explanation, those skilled in the art will believe that the invention can be made and utilized, and the claimed methods practiced, by using the foregoing description and the following exemplary embodiments. Therefore, the following working examples should not be construed as limiting the remainder of this disclosure in any way.

[0405] Example 1: Ribozyme-mediated RNA assembly and expression in mammalian cells

[0406] Ribozymes (Rzs) are small catalytic RNA sequences capable of nucleotide-specific self-splicing (Doherty and Doudna 2000). Ribozyme-mediated RNA scission produces unique 3' phosphate and 5'-hydroxyl ends, which are analogous to substrates of RNA repair pathways ubiquitous in all three kingdoms of life. As shown in this paper, ribozyme-mediated cis-splicing can be used for trans-splicing of independent RNA transcripts in mammalian cells, a method known as stitchR (stitchRNA). Notably, reconstructing messenger RNA via stitchR allows for efficient translation and expression of full-length proteins in mammalian cells. As demonstrated, stitchR can be used to assemble protein-coding functional domains or for the delivery and expression of large protein-coding sequences via viral vectors. Furthermore, overexpression of RNA 2',3'-cyclic phosphate and 5'-hydroxyl (RtcB) ligases enhances stitchR activity in mammalian cells and is sufficient to catalyze stitchR activity in vitro. These data describe a novel approach to scarless trans-splicing of functional RNA within cells using ribozymes, which holds promise for numerous research and therapeutic applications.

[0407] Autocatalytic RNA sequences are widely distributed in nature and catalyze a variety of biological processes, including intron splicing, rolling circle virus genome replication, and peptide bond formation (Weinberg et al. 2019). At least seven major ribozyme families have been identified with distinct sequence and structural features, including hammerhead (HH) ribozymes, hepatitis D virus (HDV) ribozymes, Varkud Satellite (VS) ribozymes, Sister ribozymes, Twister-sister ribozymes, Hairpin ribozymes, Hatchet ribozymes, and Pistol ribozymes. The HH, HDV, and Twister family members are the most extensively studied. Due to their small size and splicing properties, they have been used in vitro and in vivo to generate RNA with precise ends that do not contain ribozyme sequences (Figure 13) (Ferre-D'Amare and Doudna 1996; Avis et al. 2012; Zhang et al. 2017).

[0408] In both prokaryotes and eukaryotes, most cellular RNAs (including messenger RNA and long non-coding RNA) are synthesized and spliced ​​with 5'-phosphate (P) and 3'-hydroxyl (OH) ends. Conversely, unconventional cis-splicing of many tRNAs and mRNAs encoding the endoplasmic reticulum (ER) stress response protein XBP1 is catalyzed by enzymatic pathways, resulting in distinctive 5'-OH ends and either 3'-P or 2'3' cyclic phosphate (cP) ends. Recent findings suggest that unconventional cis-splicing of RNA is catalyzed by RNA ligases with 2',3'-cyclic phosphate and 5'-OH (RtcB), which are ubiquitous in mammals. Furthermore, RtcB and several other enzyme families may have the function of repairing host cellular RNA that may be damaged by stress or exogenous ribotoxins. Because ribozyme-mediated cleavage produces similar ends, ribozyme-cleaved RNA can be trans-spliced ​​via endogenous RNA repair pathways.

[0409] In mammalian cells, mRNA cleaved by ribozymes undergoes trans-splicing and translation.

[0410] To determine whether ribozymes could be used for scarless RNA trans-splicing in mammalian cells, two expression plasmids were designed, each containing non-overlapping N-terminal (Nt) and C-terminal (Ct) fragments of the fluorescent reporter gene GFP (Nt-GFP and Ct-GFP, respectively). The ribozymes were designed to catalyze the removal of themselves from adjacent nucleotides of the GFP fragments, including the 3'HDV ribozyme on Nt-GFP and the 5'HH ribozyme on Ct-GFP (Fig. 1A). When transfected into mammalian COS-7 or HEK293T cells, no GFP fluorescence was detected when either the GFP-ribozyme RNA encoding Nt or Ct was expressed alone (Fig. 1B). Notably, co-expression of the Nt-GFP and Ct-GFP RNAs together resulted in green fluorescence after 48 hours (Fig. 1B). RT-PCR analysis and Sanger sequencing indicated that trans-splicing of the Nt-GFP and Ct-GFP RNAs occurred separately between the predicted ribozyme-catalyzed cleavage sites (Fig. 1C and Fig. 1D). Furthermore, full-length GFP protein was detected by Western blot in co-transfected cells (Figure 1E). These data indicate that the endogenous mammalian cell RNA repair pathway is sufficient to catalyze the independent ribozyme processing of RNA for trans-splicing and its efficient translation into full-length protein. This RNA trans-splicing method was named stitchR.

[0411] The Influence of Ribozyme Sequence and Type on Ribozyme-Mediated Trans-Splicing

[0412] To accurately quantify the relative amounts of functional full-length proteins produced in cells via ribozyme-mediated trans-splicing, a reporter gene was constructed using two non-overlapping half-fragments of a firefly luciferase fragment (Fig. 2A). Consistent with our previous findings, only simultaneous transfection of RNA encoding Nt- and Ct-luciferase ribozymes induced both trans-splicing and luciferase activity in cells (Figs. 2B and 2C). Using this assay, the effects of different HH and HDV ribozyme sequences on trans-splicing activity in mammalian cells were further characterized. Six-base-pair overlap in the stem 1HH ribozyme provided the highest luciferase activity, while mutations in the HH catalytic residues eliminated the activity, consistent with previous reports on the in vitro characterization of HH ribozyme activity (Fig. 2D). Furthermore, the luciferase activities of genomic and antigenomic HDV ribozyme sequences were comparable, except for the smallest 56-nucleotide HDV ribozyme (HDV56), which exhibited significantly reduced activity (Fig. 2E). Consistent with previous reports, the C-to-U mutation of the nucleotide required for HDV catalysis resulted in a complete loss of luciferase activity (Figure 2E). These findings suggest that ribozyme-mediated trans-splicing activity depends on ribozyme cleavage in mammalian cells.

[0413] Use translation control and / or protein degradation sequences to prevent Nt or Ct vectors from expressing unwanted or truncated proteins.

[0414] Nt or Ct RNA may be translated prior to ribozyme-mediated splicing, or when expressed alone, may lead to unwanted or truncated protein expression. To limit the expression of unspliced ​​Nt or Ct vectors, the effect of translational control of previously characterized protein degradation sequences on the stability of vectors encoding full-length GFP was tested. Adding an HDV ribozyme to the 3' end of GFP did not appear to alter GFP fluorescence (Fig. 3A and B). To selectively block GFP expression, the effects of protein degradation sequences hCL1-PEST, E1A-PEST, removal of the vector's poly(A) sequence, or translation via a polyA tail to generate a polyK tail were tested (Fig. 3A and B). All degradation sequences were cloned in-frame with the GFP open reading frame for translation via the HDV ribozyme sequence. Addition of hCL1-PEST resulted in a significant decrease in GFP fluorescence, while EF1a PEST did not. Removal of the vector's poly(A) sequence from the expression vector resulted in complete loss of GFP expression, while translation via the polyA sequence to generate a polyK tail also resulted in decreased fluorescence.

[0415] For Ct-encoded GFP reporter genes, the addition of the 5'HH ribozyme and deletion of the GFP start codon (ATG) still resulted in weak but detectable GFP expression, despite the lack of the predicted upstream substitution ATG (Fig. 3C and D). Further silencing mutations within the N-terminal NTG codon (GFPcdn) of GFP further reduced GFP detection; however, weak fluorescence remained. Addition of the 5'UTR of the yeast GCN4 gene (which encodes four small upstream ORFs acting as translation inhibitors) eliminated detectable GFP fluorescence. A smaller internal fragment of the GCN4 5'UTR, encoding only four uORFs, also effectively blocked GFP expression. These data suggest that translational control of protein degradation sequences can be used to prevent unwanted protein expression in single Nt or Ct vectors.

[0416] These translation control or protein degradation sequences can be used in other dual-vector applications where it is necessary to restrict the expression of unwanted or truncated proteins, such as dual AAV vector strategies that rely on homologous recombination to generate large protein-coding open reading frames.

[0417] Single and multiple trans-splicing of functional protein-coding RNA

[0418] To determine whether ribozyme-mediated transsplicing could be used for the combination of protein-coding domains within the cell, RNA encoding a four-copy mitochondrial targeting sequence (Nt-4xMTS) and an open reading frame (Ct-GFP) encoding full-length GFP (with the ATG start codon missing) were generated (Fig. 4A). Co-expression of these two independent RNAs resulted in robust expression of mitochondrial-targeted GFP, which overlapped with the red fluorescent mitochondrial marker MitoTracker Red CMXRos (Fig. 4B). These findings demonstrate that ribozyme-mediated transsplicing can be used to rapidly combine two independent RNAs to express specific functional fusion proteins within the cell.

[0419] Because protein translation involves three open reading frames (ORFs), ribozyme-mediated trans-splicing and expression of multiple functional proteins can occur simultaneously. Utilizing this property, multiple functional proteins can be generated by trans-splicing multiple RNAs within compatible ORFs. To demonstrate this functionality, an additional ribozyme pair located in reading frame 2 (F2) was designed, encoding a myristylated membrane-targeting sequence (Nt-F2-Myr) and a red fluorescent protein (Ct-F2-RFP) (Figure 4C). These Nt and Ct vector pairs also contain an hCL1-PEST protein degradation sequence and a GCN4 translation repressor sequence to restrict the expression of truncated proteins from individual Nt and Ct vectors, respectively. In co-transfected cells, GFP fluorescence showed high specificity for mitochondria, while RFP fluorescence showed high specificity for the membrane (Figure 4D), demonstrating that this RNA trans-splicing method can generate diverse functional proteins in cells.

[0420] Optimized ribozymes enhance protein expression in ribozyme-mediated trans-splicing.

[0421] Small sequence modifications can significantly affect the catalytic activity of ribozymes by altering secondary structure, stability, or binding to metal ion cofactors. Using our trans-splicing luciferase reporter gene assay, we identified improved ribozyme types and sequence modifications that enhanced trans-splicing luciferase reporter gene activity in mammalian cells. Figure 16 RzB hammerhead variant ribozymes containing tertiary stable motifs (TSMs) exhibited higher activity than ribozymes without TSMs. Figure 16 A). Furthermore, when the Twister(twst) ribozyme is cloned to the 3' end of Nt-Luc, its activity is higher than that of the HDV ribozyme. Catalytic mutations within the Twister ribozyme can also eliminate luciferase activity. Figure 16 B), and depends on the formation of P1 stem ( Figure 16 C). Since the Twister ribozyme requires U at position 1, this requirement may limit the design of scarless trans-splicing to U-terminated sequences. Therefore, we tested whether nucleotide substitutions at position 1 were acceptable, and found that U1A substitutions did not show significantly different activity, while U1C or U1G substitutions retained activity, albeit with a slight decrease. Figure 16 C).

[0422] Optimized splice donor and acceptor sequences enhance protein expression in ribozyme-mediated transsplicing.

[0423] The splicing of pre-mRNA by the spliceosome has been shown to enhance mRNA translation, either by depositing factors that promote pioneer-round translation or by facilitating RNA processing and export to the cytoplasm. The addition of chimeric cis-splicing introns to transgenic proteins has also been shown to promote transgenic protein expression. Subsequently, it was investigated whether trans-spliced ​​RNA could undergo cis-splicing via the spliceosome, and whether this would affect the translation and expression of trans-spliced ​​mRNA. To verify this, splice donor (SD) and splice acceptor (SA) sequences were integrated into a trans-spliced ​​GFP reporter gene, allowing the trans-spliced ​​RNA to reconstruct a chimeric intron (Figure 5A). Notably, the addition of SD and SA sequences significantly enhanced GFP fluorescence compared to a trans-spliced ​​GFP reporter gene without SD or SA sequences (Figure 5B). RT-PCR and Sanger sequencing revealed that both Nt-GFP and Ct-GFP RNAs containing SD and SA sequences underwent trans-splicing and cis-splicing, resulting in the restoration of normal GFP open reading frames (data not shown). These data suggest that trans-splicing may occur in the cell nucleus, and that subsequent cis-splicing is an effective strategy for enhancing trans-spliced ​​RNA expression.

[0424] Ribozyme-mediated trans-splicing and large gene sequence expression for delivery using viral therapeutic vectors

[0425] Ribozyme-mediated transsplicing can be used to deliver and express large protein-coding mRNAs that exceed the packaging size limitations of therapeutic viral gene therapy vectors (e.g., AAV) (Fig. 6A). This could potentially facilitate the restoration of expression of large, mutated genes in many human monogenic diseases, such as dystrophin (Dys) in Duchenne muscular dystrophy (DMD), CFTR in cystic fibrosis (CF), and factor VIII (F8) in hemophilia A. In cell-based transfection assays, co-expression of a vector encoding Nt and Ct cleavage μ-dystrophin (μDystrophin) with a C-terminal GFP tag was trans-spliced ​​in mammalian cells (Figs. 6B and 6C) and localized to the membrane (Fig. 6D). These data demonstrate the feasibility of reconstructing and expressing large protein-coding genes using ribozyme-mediated transsplicing.

[0426] Lentiviral delivery of ribozyme-mediated RNA for intracellular trans-splicing

[0427] The autocatalytic cleavage of ribozymes can hinder the packaging of ribozyme-encoded RNA by positive-sense RNA viruses (such as commonly used gamma retroviruses and lentiviral vectors). To avoid this potential problem, Nt and Ct splitting GFP expression cassettes were encoded on the negative strand of the third-generation lentiviral vector backbone (Figure 7A). Lentiviral particles of Nt and Ct vectors were constructed separately and then used to transduce HEK293T cells. Cells transduced with both Nt-GFP and Ct-GFP simultaneously showed green fluorescence expression, while no fluorescence was detected in cells transduced with Nt-GFP or Ct-GFP alone (Figure 7B). These data demonstrate that lentiviral vectors can deliver and express ribozyme-encoded RNA for trans-splicing.

[0428] This method can also be used to deliver large gene sequences that exceed the packaging size of these viral vectors, such as Dys (Figure 7C). Ribozyme-mediated trans-splicing can also safely process or reconstruct viral genomes, such as lentiviruses or large coronavirus RNA genomes.

[0429] Use viral vectors to safely process, deliver, and express virulent or antiviral genes.

[0430] Ribozyme-mediated trans-splicing can also safely process or reconstruct virulence or antiviral proteins that may suppress the production of lentiviral particles in mammalian packaging cells. These proteins include many cell suicide genes, such as translation-repressive diphtheria toxin A (DTA) (Fig. 8A). We show that vectors encoding splitting DTA sequences, upon trans-splicing and expression, suppress the co-expression of the CS2GFP reporter gene construct, consistent with the translational repression of DTA in mammalian cells (Fig. 8B).

[0431] Enzymes that enhance or inhibit ribozyme-mediated transsplicing

[0432] Many enzyme families have been proposed to link 5'-OH and 3'-P or 2'3' cyclic phosphate (cP) ends, the most notable being RtcB, which is found to be conserved across all three biological domains. Human codon-optimized RtcB orthologs from eukaryotes (H. sapiens), bacteria (E. coli), and archaea (P. horikoshii) were cloned and co-expressed to measure their effects on the activity of trans-splicing luciferase reporter genes. Interestingly, co-expression of RtcB from *Thermococcus horikoshii* led to a 4.5-fold increase in luciferase activity, while the human and bacterial orthologs showed moderate or no increase, respectively. Figure 9 ).

[0433] Other enzyme families have been shown to regulate these RNA ends. Interestingly, the expression of T4 polynucleotide kinase (T4 PNK) (which functions as both a 5'-hydroxykinase and a 3'-phosphatase, as well as a 2',3'-cyclic phosphodiesterase) significantly inhibited luciferase activity. Figure 9 These data suggest that co-expression of exogenous enzymes can enhance or inhibit ribozyme-mediated trans-splicing in mammalian cells.

[0434] RtcB is sufficient to catalyze ribozyme-mediated RNA transsplicing in vitro.

[0435] Ribozymes have been widely used in vitro to generate precise RNA ends due to their nucleotide-specific cleavage. Next, we attempted to determine whether ribozymes could be used for in vitro directed trans-splicing of independently synthesized RNA. In vitro RNA transcription of Nt and Ct-luciferase-ribozyme reporter gene constructs was performed using T7 RNA polymerase. The results showed that, according to RT-PCR, the addition of recombinant E. coli RtcB was a necessary and sufficient condition for catalyzing trans-splicing (Figures 10A and 10B). Similarly, RNA encoding the spider protein Spidroin domain was designed (Figure 10C). Spidroin is a major component of spider silk, a highly valued material due to its tensile strength, but its highly repetitive nature makes it difficult to synthesize in heterologous systems. Spidroin naturally consists of multiple A and Q repeat sequences flanked by conserved N-terminal (N1L) and C-terminal (N3R) domains. After synthesizing Spidroin RNA in vitro using T7 polymerase, RT-PCR and Sanger sequencing revealed that the addition of recombinant RtcB ligase from E. coli was sufficient to catalyze the trans ligation of ribozyme-cleaved N1L and N3R-encoded RNAs (Figure 10D).

[0436] Controlled tandem trans-splicing of RNA encoding multi-domain proteins

[0437] The next step investigated whether adding a third RNA encoding an AQ fusion domain with a flanking ribozyme would lead to tandem repeat assembly (although uncontrolled). Figure 11 A). While directed trans-splicing between each individual RNA segment could be detected, the assembly of three or more independent RNA segments could not be detected (data not shown). This is likely due to rapid circularization of RNA segments containing RtcB-compatible ends. As an alternative, it is possible to achieve sequential and controlled assembly of RNA sequences in vitro using trans-activated VS ribozymes. Figure 11 B and Figure 11(C) In this method, the 3' end RNA ribozyme is only suitable for RtcB ligation after VS-Rz addition and trans-cleavage. Since the VS-Rz trans-activated ribozyme RNA is not covalently ligated, the stepwise addition of stitchR-compatible RNA, VS-Rz, and RtcB ligases can achieve controlled tandem assembly of RNA sequences. This can be used to assemble repetitive RNAs encoding biologically or industrially important proteins (e.g., for synthesizing spider silk, elastin, collagen, etc.).

[0438] Trans-splicing of endogenous RNA using trans-cleavage ribozymes—therapeutic applications for correcting pathogenic mutations.

[0439] Ribozymes are autocatalytic RNA cleavage processes that produce unique RNA ends through cis-splicing. We have demonstrated that these ends undergo trans-splicing and are subsequently expressed in mammalian cells (Fig. 12A). Notably, cis-cleaving ribozymes can be engineered to trans-splice, allowing target RNA to be cleaved in a nucleotide-specific manner, resulting in similar RNA ends (Fig. 12B) (Carbonell et al. 2011; Webb and Luptak 2018). Therefore, trans-cleaving ribozymes can be used to catalyze scarless trans-splicing of RNA, either intracellularly or in vitro. This approach has potential applications, one major one being the deletion of pathogenic mutations in gene transcripts by targeting mutant flanking sequences in exons or introns (Figs. 12C and 12D).

[0440] In summary, this paper demonstrates that ribozyme-mediated splicing of independently expressed RNA in cells can efficiently assemble and translate into mammalian cells. This method, termed `stitchR`, offers a novel approach for the combined assembly of functional RNA and proteins for basic research and therapeutic applications. Due to the autocatalytic properties of ribozymes and the presence of endogenous RNA repair pathways in cells, `stitchR` requires only the expression of a single RNA to undergo trans-splicing and translation in cells. In vitro experiments have shown that the RtcB ligase is sufficient for trans-splicing, and given the ubiquitous and widespread expression of RtcB in all three biological kingdoms, `stitchR` holds the potential to be a useful method applicable to a wide range of organisms.

[0441] The robustness of this system relies on the efficiency and precision of ribozyme-mediated RNA cleavage, which produces reliable and precise nucleotide-specific ends crucial for the restoration of protein-coding open reading frames. Furthermore, the ability to generate RNA using ribozymes capable of completely catalyzing their own removal enables scarless assembly, producing RNA that is virtually indistinguishable from its natural counterpart.

[0442] Although the in vitro cleavage of ribozymes has been extensively studied, their in vivo cleavage mechanism remains unclear. Currently, it is believed to be influenced by folding due to interactions with RNA-binding proteins and the availability of metal ions required for catalysis. StitchR can serve as an indirect readout for ribozyme-mediated cleavage. Interestingly, this study found that ribozyme-mediated cleavage is significantly affected by changes in ribozyme sequence and structure. This suggests that optimizing ribozyme cleavage may be an effective method to enhance in vivo StitchR activity. Further analysis of the effects of RNA repair pathway components (e.g., RtcB, RtcA, and Archease) may also be important factors regulating StitchR activity.

[0443] Ribozymes naturally evolved to cis-cleave to facilitate their own cleavage; however, some ribozyme families (especially HDV and HH) have been engineered to trans-cleave target RNAs. This article discusses how binding trans-cleaving ribozymes to stitchR could further enable highly efficient RNA cleavage and repair methods, both intracellularly and in vitro. This approach could serve as a nucleotide-specific RNA "cut-and-paste" method for generating RNA diversity or for removing certain harmful mutations from pathogenic RNAs.

[0444] Example 2: RNA activation using trans-activated ribozymes Inducible Re-splicing and expression

[0445] Most ribozymes are autocatalytic, requiring only metal ions as cofactors, which are readily available in the biological environment and contribute to folding and chemical catalysis. Varkud Satellite (VS) ribozymes can be used for scarless trans-splicing if the donor RNA ends in a G nucleotide. Interestingly, VS ribozymes can be modified to allow trans-activation to induce catalysis (Guo and Collins 1995; Ouellet et al. 2009). When split into two parts, the small VS stem-loop (VS-S) alone is insufficient to induce cis-splicing; however, adding the remaining sequence VS-Rz promotes efficient splicing of VS-S. Figure 14 A). This transactivation property enables inducible ribozyme-mediated trans scission, in which the VS-Rz sequence required for VS-S scission is added to the Nt donor RNA, which is then adapted for trans splicing with Ct acceptor RNA containing a 5'-OH terminus. Figure 14 B). The VS-Rz sequence contains typical 5'-P- and 3'-OH RNA ends, which prevent it from participating in trans-splicing, thus it can serve as a multiple conversion catalyst for this reaction.

[0446] The ability to control ribozyme-mediated cleavage (e.g., by adding a desired trans-activation sequence, such as VS-Rz) allows for the controlled addition of variable or immutable RNA sequences to generate synthetic repetitive RNA. Figure 14C). One approach is to generate RNA with a unique N-terminal domain, a unique C-terminal domain, and an internal variable or immutable "repeated" domain. This approach requires the N-terminal and C-terminal RNA to contain a single ribozyme at the 3' and 5' ends, respectively. Internally repeating RNA requires ribozymes at both the 5' and 3' ends to allow it to act as both acceptor and donor during trans-splicing. However, adding ribozymes to both ends of RNA, or RNA with both 3'-P and 5'-OH, leads to cyclization by ligases (e.g., RtcB) (Desai et al. 2015), preventing its participation in linear strand growth. However, using inducible trans-activated ribozymes allows for stepwise ligation of the 5' and 3' ends by adding and removing VS-Rz and RtcB ligases, enabling controlled RNA domain synthesis. Figure 14 (C) This method can be used to generate highly repetitive RNA sequences, which can then be translated into synthetic repetitive proteins, such as those that make up hydrogels, synthesize spider silk, or form collagen, proteins that are difficult to generate and encode as DNA due to recombination. These methods can be used for drug delivery, the generation of biomaterials, or industrial materials (Chambre et al. 2020).

[0447] Example 3: Generating stable synthetic intron sequences using ribozymes

[0448] When one RNA contains a 3' ribozyme and another RNA contains a 5' ribozyme, ribozyme-mediated trans-splicing can occur between the two independent RNAs. Figure 15 A). However, when cis-transcription was performed within the same RNA, results showed that both ribozymes could mediate self-scarless removal ( Figure 15 B). This method can also generate two independent RNAs, one with a 3'-P end and the other with a 5'-OH end, which can be trans-spliced ​​and translated within the cell. Figure 15 B). This can also be achieved in vitro by adding a ligase (e.g., RtcB).

[0449] The intron sequences produced by ribozymes also contain compatible 5'-OH and 3'-P ends, allowing for cis-splicing or circularization, a common indicator of RtcB ligase activity in vitro. Unlike the lasso RNA produced by spliceosomes during exon splicing (which degrades rapidly), RNA loops are considered highly stable because they no longer contain 5' or 3' ends and are therefore not degraded by RNA exonucleases. The cargo sequence can contain any number of functional or useful RNAs (e.g., microRNAs, CRISPR guide RNAs, etc.) or gene expression sequences, and can be inserted as "cargo" between two ribozymes. Figure 15C). This method can be used to co-deliver and express useful RNA sequences during ribozyme-mediated trans-splicing and expression. If an internal ribozyme does not require bilateral flanking sequences to function, such as the 5'HDV ribozyme, the RNA loop can exist in a circular or re-spliced ​​linear form. Figure 15 C). When VS-S is used instead of HDV, the system can be made inducible to deliver or express VS-Rz. The use of ribozymes requiring cleavage of bilateral flanking sequences (e.g., HH ribozymes) can be designed so that RNA circularization of the cargo RNA is unidirectional. Figure 15 D).

[0450] Example 4: Sequence

[0451] Trans-splicing protein-coding nucleic acid sequences

[0452] Nt-GFP (SEQ ID NO:1)

[0453] AUGGUGAGCAAGGGCGAGGAGCUGUUCACCGGGGUGGUGCCCAUCCUGGUCGAGCUGGACGGCG

[0454] ACGUAAACGGCCACAAGUUCAGCGUGUCCGGCGAGGGCGAGGGCGAUGCCACCUACGGCAAGCUG

[0455] ACCCUGAAGUUCAUCUGCACCACCGGCAAGCUGCCCGUGCCCUGGCCCACCCUCGUGACCACCCU

[0456] GACCUACGGCGUGCAGUGCUUCAGCCGCUACCCGACCACAUGAAGCAGCACGACUUCUUCAAGU

[0457] CCGCCAUGCCCGAAGGCUACGUCCAGGAGCGCACCAUCUUCUU

[0458] Ct-GFP (SEQ ID NO:2)

[0459] CAAGGACGGCAACUACAAGACCCGCGCCGAGGUGAAGUUCGAGGGCGACACCCUGGUGAACC

[0460] GCAUCGAGCUGAAGGGCAUCGACUUCAAGGAGGACGGCAACAUCCUGGGGCACAAGCUGGAGUA

[0461] CAACUACAACAGCCACAACGUCUAUAUCAUGGCCGACAAGCAGAAGAACGGCAUCAAGGUGAAC

[0462] UUCAAGAUCCGCCACAACAUCGAGGACGGCAGCGUGCAGCUCGCCGACCACUACCAGCAGAACAC

[0463] CCCCAUCGGCGACGGCCCCGUGCUGCUGCCCGACAACCACUACCUGAGCACCCAGUCCGCCCUGA

[0464] GCAAAGACCCCAACGAGAAGCGCGAUCACAUGGUCCUGCUGGAGUUCGUGACCGCCGCCGGGAUC

[0465] ACUCUCGGCAUGGACGAGCUGUACAAGUAGUAA

[0466] Nt - Luciferase (SEQ ID NO:3)

[0467] AUGGAAGACGCCAAAAACAUAAAGAAAGGCCCGGCGCCAUUCUAUCCGCUGGAAGAUGGAACCG

[0468] CUGGAGAGCAACUGCAUAAGGCUAUGAAGAGAUACGCCCUGGUUCCUGGAACAAUUGCUUUUAC

[0469] AGAUGCACAUAUCGAGGUGGACAUCACUUACGCUGAGUACUUCGAAAUGUCCGUUCGGUUGGCA

[0470] GAAGCUAUGAAACGAUAUGGGCUGAAUACAAAUCACAGAAUCGUCGUAUGCAGUGAAAACUCUC

[0471] UUCAAUUCUUUAUGCCGGUGUUGGGCGCGUUAUUUAUCGGAGUUGCAGUUGCGCCCGCGAACGA

[0472] CAUUUAUAAUGAACGUGAAUUGCUCAACAGUAUGGGCAUUUCGCAGCCUACCGUGGUGUUCGUU

[0473] UCCAAAAAGGGGUUGCAAAAAAUUUUGAACGUGCAAAAAAAGCUCCCAAUCAUCCAAAAAAUUA

[0474] UUAUCAUGGAUUCUAAAACGGAUUACCAGGGAUUUCAGUCGAUGUACACGUUCGUCACAUCUCA

[0475] UCUACCUCCCGGUUUUAAUGAAUACGAUUUUGUGCCAGAGUCCUUCGAUAGGGACAAGACAAUU

[0476] GCACUGAUCAUGAACUCCUCUGGAUCUACUGGUCUGCCUAAAGGUGUCGCUCUGCCUCAUAGAAC

[0477] UGCCUGCGUGAGAUUCUCGCAUGCCAGAGAUCCUAUUUUUGGCAAUCAAAUCAUUCCGGAUACU

[0478] GCGAUUUUAAGUGUUGUUCCAUUCCAUCACGGUUUUGGAAUGUUUACUACACUCGGAUAUUUGA

[0479] UAUGUGGAUUUCGAGUCGUCUUAAUGUAUAGAUUUGAAGAAGAGCUGUUUCUGAGGAGCCUU

[0480] Ct - Luciferase (SEQ ID NO:4)

[0481] CAGGAUUACAAGAUUCAAAGUGCGCUGCUGGUGCCAACCCUAUUCUCCUUCUUCGCCAAAAGCAC

[0482] UCUGAUUGACAAAUACGAUUUAUCUAAUUUACACGAAAUUGCUUCUGGUGGCGCUCCCCUCUCU

[0483] AAGGAAGUCGGGGAAGCGGGUUGCCAAGAGGUUCCAUCUGCCAGGUAUCAGGCAAGGAUUGGGC

[0484] UCACUGAGACUACAUCAGCUAUUCUGAUAUACACCCGAGGGGGAUGAUAAACCGGGCGCGGUCGG

[0485] UAAAGUUGUUCCAUUUUUGAAGCGAAGGUUGUGGAUCUGGAUACCGGGAAAACGCUGGGCGUU

[0486] AAUCAAAGGGCGAACUGUGUGUGAGAGGUCCUAUGAUUAUGUCCGGUUAUGUAAACAAUCCGG

[0487] AAGCGACCAACGCCUUGAUUGACAAGGAUGGAUGGCUACAUUCUGGAGACAUAGCUUACUGGGA

[0488] CGAAGACGAACAUUCUCUUCAUCGUUGACCGCCUGAGUCUCUGAUUAAGUACAAAGGCUAUCAG

[0489] GUGGCUCCCGCUGAAUUGGAAUCCAUCUUGCUCCAACACCCCCAACAUCUUCGACGCAGGUGUCGC

[0490] AGGUCUUCCCGACGAUGACGCCGGUGAACUUCCCGCCGCCGUUGUUGUUUGGAGCACGGAAAG

[0491] ACGAUGACGGAAAAGAGAUCGUGGAUAUACGUCGCCCAGUCAAGUAACAACCGCGAAAAAGUUGC

[0492] GCGGAGGAGUUGUGUUGUGGACGAGAGUACCGAAAGGUCUUACCGGAAACUCGACGCAAGAAAAUCAGAGAGAUGUCCUCUAUAAAGCCAAGAAGGGCGGAAAGAUCGCCGUGUAGUA

[0493] N1L(SEQ ID NO:5)

[0494] ATGGGTCAGGCCAATACGCCCTGGAGCAGTAAGGCAAACGCGGATGCCTTTATAAATTCATTCATCAG

[0495] TGCAGCATCCAATACTGGTTCCTTCTCTCAAGACCAAATGGAGGACATGTCACTCATCGGCAATACTC

[0496] TGATGGCTGCCATGGACAATATGGGAGGCCGCATAACACCATCTAAGTTGCAGGCGTTGGATATGGCC

[0497] TTCGCATCATCAGTGGCCGAGATCGCGGCTAGTGAGGGCGGCGACTTGGGAGTCACTACCAACGCGA

[0498] TCGCGGATGCCCTCACTTCTGCTTTTTATCAAACGACCGGGGTTGTCAATTCACGATTCATATCTGAGA

[0499] TCAGGAGCCTCATAGGAATGTTCGCGCAGGCTTCCGCAAATGACGTTTATGCATCTGCTGGCTCTGGC

[0500] AGCGGGGGTGGTGGGTATGGAGCCAGCTCAGCATCTGCGGCTTCTGCAAGTGCTGCTGCCCCGAGTG

[0501] GCGTAGCTTATCAGGCTCCTGCTCAGGCTCAAATCAGTTTTACGTTGCGAGGGCAACAACCTGTTTCC

[0502] AQ(SEQ ID NO:6)

[0503] GGTCCTTATGGACCCGGTGCTAGCGCTGCGGCAGCAGCCGCTGGCGGTTATGGCCCAGGTTCAGGGC

[0504] AACAGGGGCCTGGGCAACAAGGACCTGGCCAACAAGGTCCTGGTCAGCAGGGTCCAGGGCAGCAG

[0505] NR3(SEQ ID NO:7)

[0506] GGCGCTGCTTCCGCTGCAGTATCAGTAGGTGGCTATGGACCTCAATCTAGTAGCGCCCCTGTTGCCTCT

[0507] GCCGCCGCATCTCGACTTTCAAGTCCCGCCGCTAGTTCCAGGGTCAGTTCCGCGGTATCTAGCTTGGT

[0508] AAGTAGCGGACCCACTAATCAAGCGGCACTTTCAAACACAATATCCTCAGTAGTCAGTCAAGTAAGC

[0509] GCATCAAACCCTGGCTTGTCAGGGTGTGACGTTCTGGTTCAGGCACTTCTGGAAGTTGTCTCAGCGTT

[0510] GGTAAGCATCCTGGGTAGCTCCTCCATAGGTCAAATTAATTATGGCGCGAGCGCCCAATACACACAAA

[0511] TGGTGGGTCAGAGTGTGGCGCAGGCACTCGCAGGCGACTACAAGGATCATGACGGAGACTATAAGGA

[0512] TCATGATATAGATTACAAGGACGATGATGACAAGGCCTAGTAA

[0513] Nt-4xMTS(SEQ ID NO:8)

[0514] AUGAGUGUGUUGACGCCGUUGCUUCUGCGAGGGCUUACCGGGUCUGCUAGAAGACUUCCGGUCC

[0515] CCAGGGCCAAGAUACAUAGCCUCGGAGACCCGAUGUCUGUGCUCACUCCUCUGCUUUUGCGAGGA

[0516] CUGACUGGGUCCGCCAGACGACUCCCGGUGCCGAGAGCUAAAAUCCAUAGCCUGGGAAAAUUGGC

[0517] AACUAUGUCAGUCCUGACGCCGCUUCUUCUCCGGGGUCUUACAGGGUCUGCAAGAAGGCUGCCUG

[0518] UACCUCGGGCGAAAAUUCAUAGCUUGGGCGACCCGAUGAGUGUAUUGACGCCCCUGUUGCUGAG

[0519] AGGAUUGACUGGGUCAGCGCGCCGGCUCCCUGUCCCCCGAGCUAAGAUUCACUCCCUUGGUAAGC

[0520] UGAGAAUCCUCCAAUCAACGGUUCCGAGAGCAAGAGAUCCGCCGGUCGCCACGAGGCCUCUCGAG

[0521] Nt-DTA(SEQ ID NO:17)

[0522] AUGGACCCCGACGACGUGGUGGACAGCAGCAAGAGCUUCGUGAUGGAGAACUUCAGCAGCUACC

[0523] ACGGCACCAAGCCCGGCUACGUGGACAGCAUCCAGAAGGGCAUCCAGAAGCCCAAGAGCGGCACC

[0524] CAGGGCAACUACGACGACGACUGGAAGGGCUUCUACAGCACCGACAACAAGUACGACGCUGCCGG

[0525] CUACAGCGUGGACAACGAGAACCCCCUGAGCGGCAAGGCCGGCGGCGUGGUGAAGGUGACCUACC

[0526] CCGGCCUGACCAAGGUGCUGGCCCUGAAGGUG

[0527] Ct-DTA(SEQ ID NO:18)

[0528] GACAAUGCCGAGACCAUCAAGAAGGAGCUGGGCCUGAGCCUGACCGAGCCCCUGAUGGAGCAGG

[0529] UGGGCACCGAGGAGUUCAUCAAGAGAUUCGGCGACGGCGCCAGCAGAGUGGUGCUGAGCCUGCC

[0530] CUUCGCCGAGGGCAGCAGCAGCGUGGAGUACAUCAACAACUGGGAGCAGGCCAAGGCCCUGAGCG

[0531] UGGAGCUGGAGAUCAACUUCGAGACCAGAGGCAAGAGAGGCCAGGACGCCAUGUACGAGUACAU

[0532] GGCCCAGGCUUGCGCCGGCAACAGAGUGAGAAGAUAGUAA

[0533] GFPcdn (without start ATG codon) (SEQ ID NO:19)

[0534] GUUAGCAAGGGCGAGGAGCUCUUCACCGGGGUCGUCCCCAUCCUCGUCGAGCUCGACGGCGACGU

[0535] AAACGGCCACAAGUUCAGCGUCUCCGGCGAGGGCGAGGGCGAUGCCACCUACGGCAAGCUCACCC

[0536] UGAAGUUCAUCUGCACCACCGGCAAGCUGCCCGUGCCCUGGCCCACCCUCGUGACCACCCUGACC

[0537] UACGGCGUGCAGUGCUUCAGCCGCUACCCCGACCACAUGAAGCAGCACGACUUCUUCAAGUCCGC

[0538] CAUGCCCGAAGGCUACGUCCAGGAGCGCACCAUCUUCUUCAAGGACGACGGCAACUACAAGACCC

[0539] GCGCCGAGGUGAAGUUCGAGGGCGACACCCUGGUGAACCGCAUCGAGCUGAAGGGCAUCGACUU

[0540] CAAGGAGGACGGCAACAUCCUGGGGCACAAGCUGGAGUACAACUACAACAGCCACAACGUCUAU

[0541] AUCAUGGCCGACAAGCAGAAGAACGGCAUCAAGGUGACUUCAAGAUCCGCCACAACAUCGAGG

[0542] ACGGCAGCGUGCAGCUCGCCGAACCACUACCAGCAGAACACCCCAUCGGCGACGGCCCCGUGCUG

[0543] CUGCCCGACAACCACUACCUGAGCACCCAGUCCGCCCUGAGCAAAGACCCCAACGAGAAGCGCGA

[0544] UCACAUGGUCCUGCUGGAGUUCGUGACCGCCGCGGGAUCACUCUCGGCAUGGACGAGCUGUACA

[0545] AGUAG

[0546] F2-Myr(SEQ ID NO:20)

[0547] AUGGGGUUGUGUUUCAGCAAGACAGCGGCGAAGGUGAAGCAGCAGCAGAAAGACCAGGCGAGG

[0548] CUGCGGUAGCAUCAAGUCCCUCCAAGGCUAAUGGGCAGGAAAAACGGACACGUCAAAGUUGGAAG

[0549] General Terms and Conditions

[0550] F2-RFP(SEQ ID NO:21)

[0551] AGCCAUCAUCAAGGAGUUCAUGCGCUUCAAGGUGCACAUGGAGGCUCCGUGAACGGCCACGAG

[0552] UUCGAGAUCGAGGGCGAGGGCGAGGCCGCCCCUACGAGGGCACCCAGACCGCCAAGCUGAAGGU ...

Claims

1. A method for treating a disease or disorder caused by a mutation in a target protein in a subject, comprising: The subject is given a first nucleic acid molecule containing a coding region encoding a first portion of the target protein, the coding region being directly linked at its 3' end to a nucleotide sequence of a ribozyme (3' ribozyme); as well as The subject is given a second nucleic acid molecule containing a coding region encoding a second part of the target protein, the coding region being directly linked at its 5' end to a nucleotide sequence of a ribozyme (5' ribozyme); Specifically, the self-cleavage of the 3' ribozyme generates a 3' phosphate at the 3' end of the first part of the target protein, and the self-cleavage of the 5' ribozyme generates a 5' hydroxyl group at the 5' end of the second part of the target protein, thereby achieving scarless connection of the target protein.

2. The method according to claim 1, wherein at least one of the 3' ribozyme and the 5' ribozyme is selected from SEQ ID NO:131, SEQ ID NO:132, SEQ ID NO:133, SEQ ID NO:134, SEQ ID NO:135, SEQ ID NO:136, SEQ ID NO:137, SEQ ID NO:138, SEQ ID NO:139, SEQ ID NO:140, SEQ ID NO:141, SEQ ID NO:142, SEQ ID NO:143, SEQ ID NO:144, SEQ ID NO:145, SEQ ID NO:166, SEQ ID NO:167, SEQ ID NO:168, SEQ ID NO:169, SEQ ID NO:170, SEQ ID NO:171, SEQ ID NO:172, SEQ ID NO:173, SEQ ID NO:174, SEQ ID NO:175, SEQ ID NO:176, SEQ ID NO:192, SEQ ID NO:193, SEQ ID NO:194, SEQ ID NO:195, SEQ ID NO:19 ...7, SEQ ID NO:168, SEQ ID NO:169, SEQ ID NO:170, SEQ ID NO:171, SEQ ID NO:172, SEQ ID NO:173, SEQ ID NO:174, SEQ ID NO:175, SEQ ID NO:176, SEQ ID NO:192, SEQ ID NO:193, SEQ ID NO:194, SEQ ID NO:195, SEQ ID NO:196, SEQ ID NO:197, SEQ ID NO: ID NO:194, SEQ ID NO:195, SEQ ID NO:196, SEQ ID NO:197, SEQ ID NO:198, SEQ ID NO:199, SEQ ID NO:200, SEQ ID NO:202, SEQ ID NO:203, SEQ ID NO:204, SEQ ID NO:205, SEQ ID NO:206, SEQ ID NO:207, SEQ ID NO:217, SEQ ID NO:218 or SEQ ID NO:

219.

3. The method according to claim 1, wherein the disease or disorder is selected from one or more of the following: Duchenne muscular dystrophy, Becker muscular dystrophy (BMD), autosomal recessive polycystic kidney disease, hemophilia A, Stargardt macular degeneration, limb-girdle muscular dystrophy, autosomal recessive profound prelingual hearing loss, autosomal recessive nonsyndromic hearing loss (ARNSHL), sensorineural hearing loss, cystic fibrosis, Wilson's disease, Trihyomyopathy, and autosomal recessive deafness-9 (DFNB9). Type I Usher syndrome, GJB2-associated autosomal recessive nonsyndromic hearing loss (GJB2-ARNSHL), Type 3 autosomal recessive cerebellar cortical disorder, nonsyndromic hearing loss, autosomal recessive deafness-16 (DFNB16), Meniere's disease (MD), autosomal dominant nonsyndromic sensorineural hearing loss-12 (DFNA12), autosomal recessive spinocerebellar ataxia-21 (SCAR21), Type 1F Usher syndrome (USH1F), autosomal Acute hearing loss-23 (DFNB23), autosomal recessive hearing loss-30 (DFNB30), macrophage dysplasia of the vertebrae (OSMED), autosomal recessive hearing loss-77 (DFNB77), autosomal recessive hearing loss-84A (DFNB84A), autosomal recessive hearing loss-84B (DFNB84B), peripheral neuropathy, myopathy, autosomal dominant nonsyndromic hearing loss-4A (DFNA4), congenital thrombocytopenia, sensorineural hearing loss, autosomal dominant nonsyndromic hearing loss Hearing loss-56 (DFNA56), epileptic encephalopathy, Timothy syndrome, long QT syndrome, X-linked retinopathy, aldosteronism, autosomal recessive deafness-42 (DFNB42), primary aldosteronism (Conn syndrome), seizures, neurological abnormalities, sinoatrial node dysfunction, neurodevelopmental disorders, hypokalemic periodic paralysis, epilepsy, developmental and epileptic encephalopathy, Brody's myopathy, Darier's disease, heart disease, von Willebrand disease, or brain-hepatorenal syndrome.

4. The method according to claim 1, wherein the target protein is selected from DMD, PKHD1, F8, ABCA4, DYSF, OTOF, CFTR, ATP7B, MYOF, MYO7A, MYO15A, CDH23, STRC, OTOG, TECTA, PCDH15, TRIOBP, MYO3A, COL11A2, LOXHD1, PTPRQ, OTOGL, MYH14, MYH9, TNC, CACNA1A, CACNA1C, CACNA1F, CACNA1H, CACNA1G, CACNA1D, CACNA1B, CACNA1S, CACNA1I, CACNA1E, ATP2A1, ATP2A2, VWF, PEX1, and CMYA5.

5. The method of claim 1, wherein the first nucleic acid molecule comprises a nucleic acid sequence encoding an N-terminal fragment encoding dystrophin, mini dystrophin, or micro dystrophin; and wherein the second nucleic acid molecule comprises a nucleic acid sequence encoding a C-terminal fragment encoding dystrophin, mini dystrophin, or micro dystrophin; and wherein administration of the first and second nucleic acid molecules results in the production of full-length dystrophin, mini dystrophin, or micro dystrophin.

6. The method of claim 5, wherein the first and second nucleic acid molecules are selected from: a) A first nucleic acid molecule containing the nucleic acid sequence of SEQ ID NO:150 and a second nucleic acid molecule containing the nucleic acid sequence of SEQ ID NO:151; and b) A first nucleic acid molecule containing the nucleic acid sequence of SEQ ID NO:152 and a second nucleic acid molecule containing the nucleic acid sequence of SEQ ID NO:

153.

7. The method of claim 1, wherein the first nucleic acid molecule comprises a nucleic acid sequence encoding an N-terminal fragment of Dysferlin; and wherein the second nucleic acid molecule comprises a nucleic acid sequence encoding a C-terminal fragment of Dysferlin; and wherein administration of the first and second nucleic acid molecules results in the production of a full-length Dysferlin protein.

8. The method of claim 7, wherein the first and second nucleic acid molecules are selected from: a) A first nucleic acid containing the nucleic acid sequence of SEQ ID NO:157 and a second nucleic acid containing the nucleic acid sequence of SEQ ID NO:158; b) A first nucleic acid containing the nucleic acid sequence of SEQ ID NO:177 and a second nucleic acid containing the nucleic acid sequence of SEQ ID NO:178; c) A first nucleic acid containing the nucleic acid sequence of SEQ ID NO:179 and a second nucleic acid containing the nucleic acid sequence of SEQ ID NO:178; d) A first nucleic acid containing the nucleic acid sequence of SEQ ID NO:181 and a second nucleic acid containing the nucleic acid sequence of SEQ ID NO:182; e) A first nucleic acid containing the nucleic acid sequence shown in SEQ ID NO:183 and a second nucleic acid containing the nucleic acid sequence shown in SEQ ID NO:

182.

9. The method of claim 1, wherein the first nucleic acid molecule comprises a nucleic acid sequence encoding an N-terminal segment of STRC; and wherein the second nucleic acid molecule comprises a nucleic acid sequence encoding a C-terminal segment of STRC; and wherein administration of the first and second nucleic acid molecules results in the production of a full-length STRC protein.

10. The method according to claim 9, wherein the first nucleic acid molecule comprises the nucleic acid sequence of SEQ ID NO:160, and the second nucleic acid molecule comprises the nucleic acid sequence of SEQ ID NO:

161.

11. A system for generating an RNA molecule encoding a target protein, comprising: A first RNA molecule comprising a coding region encoding a first portion of a target protein, wherein the coding region is directly linked at its 3' end to a nucleotide sequence of a ribozyme (3' ribozyme); and The second RNA molecule contains a coding region encoding the second part of the target protein, the coding region being directly linked at the 5' end to the nucleotide sequence of the ribozyme (5' ribozyme); The self-cleavage of the 3' ribozyme generates a 3' phosphate at the 3' end of the first part of the target protein, and the self-cleavage of the 5' ribozyme generates a 5' hydroxyl group at the 5' end of the second part of the target protein, thereby allowing the first RNA molecule and the second RNA molecule to be joined without scarring to generate an RNA molecule encoding the target protein.

12. The system according to claim 11, wherein at least one of the 3' ribozyme and the 5' ribozyme is selected from SEQ ID NO:131, SEQ ID NO:132, SEQ ID NO:133, SEQ ID NO:134, SEQ ID NO:135, SEQ ID NO:136, SEQ ID NO:137, SEQ ID NO:138, SEQ ID NO:139, SEQ ID NO:140, SEQ ID NO:141, SEQ ID NO:142, SEQ ID NO:143, SEQ ID NO:144, SEQ ID NO:145, SEQ ID NO:166, SEQ ID NO:167, SEQ ID NO:168, SEQ ID NO:169, SEQ ID NO:170, SEQ ID NO:171, SEQ ID NO:172, SEQ ID NO:173, SEQ ID NO:174, SEQ ID NO:175, SEQ ID NO:176, SEQ ID NO:192, SEQ ID NO:19 ... NO:193, SEQ ID NO:194, SEQ ID NO:195, SEQ ID NO:196, SEQ ID NO:197, SEQ ID NO:198, SEQ ID NO:199, SEQ ID NO:200, SEQ ID NO:202, SEQ ID NO:203, SEQ ID NO:204, SEQ ID NO:205, SEQ ID NO:206, SEQ ID NO:207, SEQ ID NO:217, SEQ ID NO:218 or SEQ ID NO:

219.

13. The system according to claim 11, wherein the full length of the target protein is greater than 1,000 amino acid residues.

14. The system according to claim 12, wherein the target protein is selected from DMD, PKHD1, F8, ABCA4, DYSF, OTOF, CFTR, ATP7B, MYOF, MYO7A, MYO15A, CDH23, STRC, OTOG, TECTA, PCDH15, TRIOBP, MYO3A, COL11A2, LOXHD1, PTPRQ, OTOGL, MYH14, MYH9, TNC, CACNA1A, CACNA1C, CACNA1F, CACNA1H, CACNA1G, CACNA1D, CACNA1B, CACNA1S, CACNA1I, CACNA1E, ATP2A1, ATP2A2, VWF, PEX1, and CMYA5.

15. A method for generating an RNA molecule encoding a target protein, comprising: The first RNA molecule is administered to cells or tissues, the first RNA molecule containing a coding region encoding a first part of a target protein, the coding region being directly linked at the 3' end to a nucleotide sequence of a ribozyme (3' ribozyme); as well as The second RNA molecule is administered to cells or tissues. The second RNA molecule contains a coding region encoding a second part of a target protein, which is directly linked at the 5' end to the nucleotide sequence of a ribozyme (3' ribozyme). The 3' ribozyme self-cleavage produces a 3' phosphate at the 3' end of the first part of the target protein, and the 5' ribozyme self-cleavage produces a 5' hydroxyl group at the 5' end of the second part of the target protein, thereby allowing the first RNA molecule and the second RNA molecule to join without scarring to generate an RNA molecule encoding the target protein.

16. The method according to claim 15, wherein at least one of the 3' ribozyme and the 5' ribozyme is selected from SEQ ID NO:131, SEQ ID NO:132, SEQ ID NO:133, SEQ ID NO:134, SEQ ID NO:135, SEQ ID NO:136, SEQ ID NO:137, SEQ ID NO:138, SEQ ID NO:139, SEQ ID NO:140, SEQ ID NO:141, SEQ ID NO:142, SEQ ID NO:143, SEQ ID NO:144, SEQ ID NO:145, SEQ ID NO:166, SEQ ID NO:167, SEQ ID NO:168, SEQ ID NO:169, SEQ ID NO:170, SEQ ID NO:171, SEQ ID NO:172, SEQ ID NO:173, SEQ ID NO:174, SEQ ID NO:175, SEQ ID NO:176, SEQ ID NO:192, SEQ ID NO:16 ... NO:193, SEQ ID NO:194, SEQ ID NO:195, SEQ ID NO:196, SEQ ID NO:197, SEQ ID NO:198, SEQ ID NO:199, SEQ ID NO:200, SEQ ID NO:202, SEQ ID NO:203, SEQ ID NO:204, SEQ ID NO:205, SEQ ID NO:206, SEQ ID NO:207, SEQ ID NO:217, SEQ ID NO:218 or SEQ ID NO:

219.

17. The method of claim 14, wherein the full length of the target protein is greater than 1,000 amino acid residues.

18. The method according to claim 17, wherein the target protein is selected from DMD, PKHD1, F8, ABCA4, DYSF, OTOF, CFTR, ATP7B, MYOF, MYO7A, MYO15A, CDH23, STRC, OTOG, TECTA, PCDH15, TRIOBP, MYO3A, COL11A2, LOXHD1, PTPRQ, OTOGL, MYH14, MYH9, TNC, CACNA1A, CACNA1C, CACNA1F, CACNA1H, CACNA1G, CACNA1D, CACNA1B, CACNA1S, CACNA1I, CACNA1E, ATP2A1, ATP2A2, VWF, PEX1, and CMYA5.

19. A method for generating an RNA molecule encoding a target protein in vitro, comprising: Provided is a first RNA molecule containing a coding region encoding a first portion of a target protein, the coding region being directly linked at its 3' end to a nucleotide sequence of a ribozyme (3' ribozyme); and Provide a second RNA molecule containing a coding region encoding a second part of a target protein, wherein the coding region is directly linked at its 5' end to a nucleotide sequence of a ribozyme (5' ribozyme); The 3' ribozyme self-cleavage produces a 3' phosphate at the 3' end of the first part of the target protein, and the 5' ribozyme self-cleavage produces a 5' hydroxyl group at the 5' end of the second part of the target protein, thereby allowing the first and second RNA molecules to join without scarring to generate an RNA molecule encoding the target protein; and A ligase is provided to induce the assembly of the coding regions of the first RNA molecule and the second RNA molecule into an RNA molecule.

20. The method according to claim 19, wherein at least one of the 3' ribozyme and the 5' ribozyme is selected from SEQ ID NO:131, SEQ ID NO:132, SEQ ID NO:133, SEQ ID NO:134, SEQ ID NO:135, SEQ ID NO:136, SEQ ID NO:137, SEQ ID NO:138, SEQ ID NO:139, SEQ ID NO:140, SEQ ID NO:141, SEQ ID NO:142, SEQ ID NO:143, SEQ ID NO:144, SEQ ID NO:145, SEQ ID NO:166, SEQ ID NO:167, SEQ ID NO:168, SEQ ID NO:169, SEQ ID NO:170, SEQ ID NO:171, SEQ ID NO:172, SEQ ID NO:173, SEQ ID NO:174, SEQ ID NO:175, SEQ ID NO:176, SEQ ID NO:192, SEQ ID NO:193, SEQ ID NO:174, SEQ ID NO:175, SEQ ID NO:176, SEQ ID NO:192, SEQ ID NO:193, SEQ ID NO:194, SEQ ID NO:195, SEQ ID NO:196, SEQ ID NO:197, SEQ ID NO:198, SEQ ID NO:19 ... NO:193, SEQ ID NO:194, SEQ ID NO:195, SEQ ID NO:196, SEQ ID NO:197, SEQ ID NO:198, SEQ ID NO:199, SEQ ID NO:200, SEQ ID NO:202, SEQ ID NO:203, SEQ ID NO:204, SEQ ID NO:205, SEQ ID NO:206, SEQ ID NO:207, SEQ ID NO:217, SEQ ID NO:218 or SEQ ID NO:

219.

21. A system for generating a circular RNA molecule encoding a target protein, comprising a nucleic acid molecule encoding: The 5' ribozyme and IRES sequence are directly linked to the first part of the target protein, and the second part of the target protein is directly linked to the 3' ribozyme. The self-cleavage of the 3' ribozyme generates a 3' phosphate at the 3' end of the second part of the target protein, and the self-cleavage of the 5' ribozyme generates a 5' hydroxyl group at the 5' end of the first part of the target protein, thereby allowing the 5' hydroxyl group of the RNA molecule to be linked to the 3' phosphate without scarring, thereby generating a circular RNA molecule encoding the target protein.

22. The system of claim 21, wherein the RNA molecule comprises an in vitro transcribed RNA molecule.

23. The system of claim 21, wherein the target protein is one or more selected from therapeutic proteins, reporter proteins, and Cas9 proteins.

24. The system according to claim 23, wherein at least one of the 3' ribozyme and the 5' ribozyme is selected from SEQ ID NO:131, SEQ ID NO:132, SEQ ID NO:133, SEQ ID NO:134, SEQ ID NO:135, SEQ ID NO:136, SEQ ID NO:137, SEQ ID NO:138, SEQ ID NO:139, SEQ ID NO:140, SEQ ID NO:141, SEQ ID NO:142, SEQ ID NO:143, SEQ ID NO:144, SEQ ID NO:145, SEQ ID NO:166, SEQ ID NO:167, SEQ ID NO:168, SEQ ID NO:169, SEQ ID NO:170, SEQ ID NO:171, SEQ ID NO:172, SEQ ID NO:173, SEQ ID NO:174, SEQ ID NO:175, SEQ ID NO:176, SEQ ID NO:192, SEQ ID NO:19 ... NO:193, SEQ ID NO:194, SEQ ID NO:195, SEQ ID NO:196, SEQ ID NO:197, SEQ ID NO:198, SEQ ID NO:199, SEQ ID NO:200, SEQ ID NO:202, SEQ ID NO:203, SEQ ID NO:204, SEQ ID NO:205, SEQ ID NO:206, SEQ ID NO:207, SEQ ID NO:217, SEQ ID NO:218 or SEQ ID NO:

219.

25. A method for generating a circular RNA molecule in vivo, the method comprising administering a linear in vitro transcribed RNA molecule or a DNA molecule encoding a linear RNA molecule to a cell or tissue, wherein the linear RNA molecule comprises: The 5' ribozyme and IRES sequence are directly linked to the first part of the target protein, and the second part of the target protein is directly linked to the 3' ribozyme. The 3' ribozyme self-cleavage produces a 3' phosphate at the 3' end of the second part of the target protein, and the 5' ribozyme self-cleavage produces a 5' hydroxyl group at the 5' end of the first part of the target protein, thereby allowing the 5' hydroxyl group of the RNA molecule to be linked to the 3' phosphate without scarring, thereby generating a circular RNA molecule encoding the target protein.

26. The method of claim 25, wherein the target protein is one or more selected from therapeutic proteins, reporter proteins, and Cas9 proteins.

27. The method according to claim 25, wherein at least one of the 3' ribozyme and the 5' ribozyme is selected from SEQ ID NO:131, SEQ ID NO:132, SEQ ID NO:133, SEQ ID NO:134, SEQ ID NO:135, SEQ ID NO:136, SEQ ID NO:137, SEQ ID NO:138, SEQ ID NO:139, SEQ ID NO:140, SEQ ID NO:141, SEQ ID NO:142, SEQ ID NO:143, SEQ ID NO:144, SEQ ID NO:145, SEQ ID NO:166, SEQ ID NO:167, SEQ ID NO:168, SEQ ID NO:169, SEQ ID NO:170, SEQ ID NO:171, SEQ ID NO:172, SEQ ID NO:173, SEQ ID NO:174, SEQ ID NO:175, SEQ ID NO:176, SEQ ID NO:192, SEQ ID NO:16 ... NO:193, SEQ ID NO:194, SEQ ID NO:195, SEQ ID NO:196, SEQ ID NO:197, SEQ ID NO:198, SEQ ID NO:199, SEQ ID NO:200, SEQ ID NO:202, SEQ ID NO:203, SEQ ID NO:204, SEQ ID NO:205, SEQ ID NO:206, SEQ ID NO:207, SEQ ID NO:217, SEQ ID NO:218 or SEQ ID NO:219.

Citation Information

Patent Citations

  • Method of Nucleotide Detection

    US20080280773A1

  • Intrinsic factor - horse peroxidase conjugates and a method for increasing the stability thereof

    US5350674A

  • Gene therapy

    US5399346A

  • Delivery of exogenous DNA sequences in a mammal

    US5580859A

  • Adenovirus vectors for gene therapy

    US5585362A