Compositions and methods for high efficiency recombination of RNA molecules

The use of RNA molecules with dimerization domains for splicing and recombination addresses the AAV packaging limitations, enabling efficient delivery and expression of large proteins for genetic disease treatment.

JP7759106B2Active Publication Date: 2025-10-23SALK INST FOR BIOLOGICAL STUDIES
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
JP2022526727
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2020-03-27
Filing Date
2020-09-30
Publication Date
2025-10-23
Estimated Expiration
2040-09-30

AI Technical Summary

Technical Problem

Existing gene therapy methods using AAV vectors struggle to deliver large proteins due to packaging constraints, often resulting in truncated proteins or inefficiencies, leaving many genetic diseases untreatable.

Method used

A composition comprising two RNA molecules, each encoding a portion of the target protein, linked by dimerization domains that facilitate splicing and recombination within cells to express full-length proteins, using systems and methods that include synthetic DNA molecules with promoters for expression.

Benefits of technology

Efficient and safe delivery of large proteins, enabling effective treatment of genetic diseases by achieving high-level expression of full-length proteins in target cells.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007759106000008
    Figure 0007759106000008
  • Figure 0007759106000009
    Figure 0007759106000009
  • Figure 0007759106000010
    Figure 0007759106000010
Patent Text Reader

Abstract

Provided herein are compositions and systems for reconstructing RNA molecules, including the method for using RNA molecules.For example, such molecules can be used to deliver protein coding sequences through two or more viral vectors (such as AAV), thereby causing the reconstructing of full-length protein in cells.Such methods can be used to deliver therapeutic protein, for example, to treat genetic disease or cancer.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application is a continuation-in-part of PCT / US2020 / 025430, filed March 27, 2020, which claims priority to U.S. Provisional Patent Application No. 62 / 933,714, filed November 11, 2019, all of which are incorporated herein by reference in their entireties.

[0002] Field The present disclosure provides systems, kits, compositions, and methods that allow for the recombination of two or more RNA molecules, thereby allowing for the expression of a full-length protein. [Background technology]

[0003] background Gene therapy is a promising method for treating genetic diseases caused by loss-of-function mutations. Replacement genes are typically reintroduced into target cells using vectors such as AAV, because this virus is generally safe and efficient at entering cells. However, in the case of AAV, it is difficult to encapsulate more than about 5,000 nucleotides using conventional capsids. The length of genes encoding large proteins often exceeds the packaging constraints of AAV, and many genetic diseases remain untreatable. Previously, strategies to overcome this limitation have been explored, but they have proven inefficient, resulted in high-level expression of potentially toxic truncated proteins, or both. A safe and highly efficient strategy for delivering large proteins to treat diseases is needed. Summary of the Invention [Means for solving the problem]

[0004] overview Compositions for expressing a target protein are provided herein. In one example, the composition includes: (a) a first RNA molecule, the first RNA molecule including, from 5' to 3', (i) a coding sequence for the N-terminal portion of the target protein, (ii) a splice donor, and (iii) a first dimerization domain; and (b) a second RNA molecule, the second RNA molecule including, from 5' to 3', (i) a second dimerization domain that binds to the first dimerization domain, (ii) a branch point sequence, (iii) a polypyrimidine tract, (iv) a splice acceptor, and (v) a coding sequence for the C-terminal portion of the target protein.

[0005] In some examples, the first dimerization domain and the second dimerization domain are linked by a direct bond, an indirect bond, or both.

[0006] In some examples, the dimerization domain is a kissing loop domain or a hypodiverse domain.

[0007] In some examples, the first RNA molecule and / or the second RNA molecule comprises at least one splice enhancer.

[0008] Also provided is a composition for expressing a target protein, the composition comprising: (a) a first synthetic DNA molecule encoding a first RNA molecule of any one of claims 1 to 16, the first synthetic DNA molecule comprising (i) a first promoter operably linked to a sequence encoding the first RNA molecule; and (b) a second synthetic DNA molecule encoding a second RNA molecule of any one of claims 1 to 16, the second synthetic DNA molecule comprising (i) a second promoter operably linked to a sequence encoding the second RNA molecule.

[0009] Also provided are systems for expressing target proteins comprising the compositions described.

[0010] Also provided is a method for expressing a protein in a cell using the system disclosed herein or the RNA encoded by the system.Such a method can include introducing the system into a cell and expressing a first synthetic RNA molecule and a second synthetic RNA molecule in the same cell.In some embodiments, the cell is present in a subject, and the method treats a disease in the subject, such as a genetic disease caused by a mutation in the gene encoding the target protein.In some embodiments, the genetic disease is Duchenne muscular dystrophy, hemophilia A, Stargardt's disease, or Usher syndrome.

[0011] The foregoing and other objects and features disclosed herein will become more apparent from the following detailed description which proceeds with reference to the accompanying drawings.

[0012] The patent or application file contains at least one drawing executed in color. Copies of this patent or patent application publication with color drawing(s) will be provided by the Office upon request and payment of the necessary fee. [Brief explanation of the drawings]

[0013] [Figure 1A]Figure 1A shows a schematic diagram of the vector design (left) and RNA interaction and splicing (right). Left: 5' trans-splice (trsp) DNA vector; open arrows indicate two opposing promoters. The 3' UTR with the RFP-encoding domain and polyadenylation element is expressed in opposite directions from the N-terminal portion of YFP (n-yfp), followed by a splice donor sequence (SD), a downstream intronic splicing enhancer (DISE), two intronic splicing enhancers (2xISE), a binding domain (BD, also known as the dimerization domain), and a stable stem-loop box B element (box B), a self-cleaving hammerhead ribozyme (HHrz), and finally a polyadenylation element. A small intron has been inserted into the n-yfp segment (white segment within n-yfp). 3' trsp DNA vector; open arrows indicate two opposing promoters. The 3'UTR with the YFP-encoding domain and polyadenylation element is expressed in opposite directions from the 3'UTR, which contains a complementary binding domain (anti-BD, also called the dimerization domain), followed by three intronic splicing enhancer sequences (3xISE), a branch point (BP), a polypyrimidine tract (PPT), a splice acceptor sequence (SA), the C-terminal portion of the YFP-encoding sequence (c-terminal proton), and finally a polyadenylation element. Right: Pre-mRNA interactions (5'trsp-RNA + 3'trsp-RNA) and trans-splicing are shown, resulting in the generation of the YFP-encoding mRNA. [Figure 1B] FIG. 1B shows that transfection of the N-terminal expression plasmid alone does not result in YFP fluorescence. [Figure 1C] FIG. 1C shows that transfection of the C-terminal expression plasmid alone does not result in YFP fluorescence. [Figure 1D] FIG. 1D shows that expression of the N- and C-terminal fragments without the binding domain showed low levels of YFP induction. [Figure 1E]Figure 1E shows rationally designed dimerization / binding domains in a looped configuration (all-pyrimidine or all-purine low-diversity sequences interrupted by complementary sequences that form a double-stranded stem structure). [Figure 1F] FIG. 1F shows a 3D rendering of the "looped" dimerization domain configuration. [Figure 1G] FIG. 1G shows a negative control lacking the binding domain in the C-terminal half. [Figure 1H] FIG. 1H shows a negative control lacking the binding domain in the N-terminal half. [Figure 1I] FIG. 1I shows that matching binding domains in loop configurations in both the N-terminal and C-terminal halves exhibited strong YFP induction in 90% of cells. [Figure 1J] Figures 1J-1N represent equivalent data to those in Figures 1E-1I for the construction of binding domains with 150-nucleotide low-diversity sequences composed exclusively of pyrimidines (or alternatively exclusively of purines) containing sequences that result in a fully open configuration. Figure 1J shows a 150-nucleotide low-diversity pyrimidine sequence that results in a fully open configuration for complimentary base pairing. [Figure 1K] FIG. 1K shows a 3D rendering of the 150-nucleotide low-diversity pyrimidine sequence of (1J). [Figure 1L] Figure 1L shows transfection of control HEK293T cells with a construct encoding a C-terminal YFP lacking the complementary low-diversity binding domain. Few transfected cells express YFP. [Figure 1M] Figure 1M shows transfection of control HEK293T cells with a construct encoding an N-terminal YFP lacking the complementary low-diversity binding domain. Few transfected cells express YFP. [Figure 1N]Figure 1N shows the transfection of HEK293T cells with N- and C-terminal YFP constructs, both of which have complementary low-divergence dimerization binding domains. Many cells express high levels of YFP. [Figure 1O] Figure 10 shows a representative fluorescence image for the cells shown in Figure 1G. Positive markers for transfection (RFP+BFP) are expressed, but YFP protein is not efficiently reconstituted. [Figure 1P] Figure 1P shows a representative fluorescence image of the cells shown in Figure 1L. The positive markers for transfection (RFP + BFP) are expressed, and YFP protein is reconstituted at high levels in RFP and BFP double-positive cells. [Figure 1Q] Figure 1Q shows a comparison of the conditions shown in Figure 1D, Figures 1G-1I, and Figures 1L-1N: N: no binding domain, Loop: looped low diversity binding domain configuration, Lin: linear low diversity configuration. [Figure 2A] Figure 2A shows a schematic diagram of the vector design. The protein coding sequence for yellow fluorescent protein (YFP) is divided into an N-terminus, a middle fragment (m-yfp), and a C-terminal fragment. The junction between the RNA encoding the n fragment and the RNA encoding the m fragment is connected by a loop-shaped binding domain (BD1), and the junction between the m fragment and the c fragment is connected by a loop-shaped binding domain (BD2). The pyrimidine (Y) and purine (R) sequences are positioned to prevent self-circularization of the m fragment and direct recombination between the N and C fragments. The N-terminal fragment is coexpressed with red fluorescent protein as a transfection control, and the C-terminal fragment is coexpressed with blue fluorescent protein as a transfection control. The promoter sequences are indicated by open arrows. The splice donor (SD) and splice acceptor (SA) sites are indicated. Intron splicing elements including a splice enhancer, polypyrimidine tract, and branch point similar to those used upstream (5') of the SA and downstream (3') of the SD in Figure 1A are included. [Figure 2B] Figure 2B shows that transfection of plasmid I+II+III (see Figure 2A) into a human cell line efficiently reconstitutes high-level YFP expression in 80% of the transfected cells. [Figure 2C] FIG. 2C shows a representative fluorescence image of expression of the n and m fragments (plasmids I+II, see FIG. 2A) in which no yfp fluorescence is shown (negative control). [Figure 2D] FIG. 2D shows a representative fluorescence image of expression of the m and c fragments (plasmids II+III, see FIG. 2A) (negative control), showing no yfp fluorescence. [Figure 2E] FIG. 2E shows a representative fluorescence image demonstrating that co-transfection of all three fragments (plasmids I+II+III, see FIG. 2A) induces strong YFP fluorescence. [Figure 3]Figures 3A-3D show efficient reconstitution of yellow fluorescent protein (YFP) from two fragments (SEQ ID NOs: 1 and 2) expressed from two AAV2 / 8 vectors after systemic administration to neonatal mice (P3). (A) AAV1 encodes the N-terminal half fragment of YFP, and AAV2 encodes the C-terminal half fragment. AAV1 and AAV2 were mixed at equal titers and intravenously injected into mice. Tissue samples were collected 3 weeks after injection. (B) YFP fluorescence in the liver of a young mouse at the time of sacrifice (green). An uninjected liver is shown for comparison (control: no YFP detected). DRAQ5 nuclear staining is shown in magenta for context. (C) Strong YFP fluorescence (green) in the myocardium at the time of sacrifice. The top panel shows the macroscopic image with red autofluorescence (magenta) for context. The bottom panel shows a cross-section (magenta) with DRAQ5 nuclear staining for context. An uninjected mouse heart lacking YFP is shown as a control. (D) Strong YFP fluorescence in skeletal muscle of the limb at the time of sacrifice. An uninjected mouse limb is shown for comparison (negative control, no YFP detected). The top panel shows a macroscopic image with magenta red autofluorescence. The bottom panel shows a microscopic image of a cross section through the limb. The bottom panel shows magenta DRAQ5 nuclear staining for context. [Figure 4] Figures 4A-4B show efficient reconstitution of yellow fluorescent protein (YFP) from three fragments (SEQ ID NOs: 145, 146, and 2, respectively) in the tibialis anterior muscle of a mouse neonatal (P3) pup after intramuscular injection of three AAV2 / 8 vectors. (A) A schematic diagram of three AAV particles carrying the N-, M-, and C-terminal fragments of YFP is shown (similar to Figure 2A). (B) Strong YFP fluorescence in a longitudinal section of the tibialis anterior muscle of a mouse injected with all three viral particles is shown. DRAQ5 nuclear staining is shown in magenta for context. [Figure 5]Figures 5A-5F show efficient reconstitution of yellow fluorescent protein (YFP) from two and three fragments in adult mouse tibialis anterior muscle. (A) The N- and C-terminal halves of the YFP coding sequence with synthetic RNA dimerization and recombination domains are shown. (B) Two AAV transfer plasmids expressing these two fragments were electroporated percutaneously into adult mouse tibialis anterior (TA) muscle, and strong fluorescence was detected 5 days after electroporation. (C) No fluorescence was detectable in the contralateral, uninjected TA. (D) The N-, middle, and C-termini of the YFP coding sequence with synthetic RNA dimerization and recombination domains are shown, with each fragment ligated to its adjacent fragment. (E) Transcutaneous electroporation of three AAV transfer plasmids expressing these three fragments is shown. Strong YFP fluorescence is detected, indicating efficient reconstitution of YFP from the three fragments. (F) Fluorescence in the contralateral, uninjected TA is shown. The fluorescence channel is overlaid on the grayscale image for context. [Figure 6A] 6A is a schematic diagram presenting an exemplary system for the RNA recombination methods disclosed herein, using two nucleic acid molecules 110, 150, in which a target protein is split into two portions, each portion encoded by a different nucleic acid molecule. In some examples, the nucleic acid molecules 110, 150 of the system are DNA and include a promoter 112, 152. In some examples, the nucleic acid molecules 110, 150 of the system are RNA and therefore lack a promoter 112, 152. The drawings are not to scale. [Figure 6B] Figure 6B is a schematic diagram presenting exemplary dimerization domains (e.g., 122, 154 in Figure 6A) that contain low diversity sequences interspersed with sequences that can form stems, resulting in local RNA loops that are open and available for base pairing in the absence of pseudoknot formation. The drawing is not to scale. [Figure 6C]Figure 6C is a schematic diagram showing that the interaction and hybridization (base pairing) of pre-mRNA dimerization domain 122 of molecule 110 (Figure 6A) with pre-mRNA dimerization domain 154 of molecule 150 (Figure 6A) allows recombination of N-terminal coding sequence 114 and C-terminal coding sequence 164 by spliceosome components. This results in a seamless junction between the 3' end of the fused N-terminal protein coding sequence 114 and the 5' end of the C-terminal protein sequence 164, and between the N-terminal and C-terminal portions. The drawing is not to scale. [Figure 6D] Figure 6D is a schematic diagram presenting an exemplary system for the RNA recombination methods disclosed herein using three nucleic acid molecules 110, 200, 150, in which the target protein is divided into three portions (N-terminal, middle, and C-terminal), each encoded by a different nucleic acid molecule. Before transcription, the nucleic acid molecules 110, 150, 200 of the system are DNA and include promoters 112, 152, 202. After transcription, the nucleic acid molecules 110, 150, 200 of the system are RNA and therefore lack promoters 112, 152, 202. The drawing is not to scale. [Figure 6E] 6E is a schematic diagram illustrating that the interaction and hybridization (base pairing) between dimerization domain 122 of molecule 110 (FIG. 6D) and dimerization domain 204 of molecule 200 (FIG. 6D), as well as the interaction and hybridization (base pairing) between dimerization domain 226 of molecule 200 (FIG. 6D) and dimerization domain 154 of molecule 150 (FIG. 6D), allows spliceosome components to recombine N-terminal coding sequence 114, intermediate coding sequence 216, and C-terminal coding sequence 164. This results in seamless junctions between the fused 3' end of N-terminal coding sequence 114 and the 5' end of intermediate protein sequence 216, and between the fused 3' end of intermediate coding sequence 216 and the 5' end of C-terminal sequence 216, as well as between the N-terminal, intermediate, and C-terminal portions. In some examples, for example, after transcription, the elements shown are RNA. The drawings are not to scale. [Figure 6F]Figure 6F is a schematic diagram presenting an exemplary system for the RNA recombination methods disclosed herein, using two nucleic acid molecules 110, 150, in which the target protein is divided into two parts, each part encoded by a different nucleic acid molecule. In this example, DNA is transcribed into RNA, and therefore the nucleic acid molecules 110, 150 of the system are RNA and therefore lack the promoters 112, 152 present in DNA (see Figure 6A). The drawing is not to scale. [Figure 7A] Figure 7A is a schematic diagram presenting an exemplary system for the RNA recombination methods disclosed herein, which uses two nucleic acid molecules 500, 600 as in Figure 6A, but in which the dimerization domains are aptamers 512, 602 that recognize the same target molecule 700. In some examples, for example, after transcription, the elements shown are RNA. The drawing is not to scale. [Figure 7B] Figure 7B is a schematic diagram presenting an exemplary system for the RNA recombination methods disclosed herein, related to Figure 7A, that uses dimerization domains that recognize the same target molecule. Here, the target recognized by the dimerization domains is a specific RNA molecule (instead of molecule 700 in Figure 7A, e.g., a protein or small molecule). Each domain recognizes a different portion of an mRNA molecule that is expressed only in target cells (i.e., cells in which expression of the target protein is desired), such as, for example, a cancer-specific transcript. In some examples, for example, after transcription, the depicted element is RNA. The drawing is not to scale. [Figure 7C] Figure 7C is a schematic diagram presenting an exemplary system for the RNA recombination methods disclosed herein, using two nucleic acid molecules 800, 900 similar to Figures 6A and 7A, showing that dimerization domains 812, 902 are hybridized to an oligonucleotide 1000 that prevents the dimerization domains from interacting with each other, thus preventing or reducing recombination of the N-terminal coding sequence 802 and the C-terminal coding sequence 914. In some examples, for example, after transcription, the elements shown are RNA. The drawings are not to scale. [Figure 8] Figure 8 is a bar graph comparing the reconstitution of YFP protein expression in the presence (w / ) or absence (w / o) of the WPRE3 sequence in the 3' untranslated region. N=3 replicates per sample are shown. [Figure 9A] Figure 9A is a schematic diagram presenting an example of the use of a dimerization domain (e.g., 122, 154 in Figure 6A) containing a kissing loop interaction for high-affinity dimerization. It will be understood that, using the teachings presented herein, any of the coding portions disclosed herein (e.g., YFP) can be replaced with other target protein coding sequences. The drawings are not to scale. [Figure 9B] Figure 9B shows RFP, BFP, and YFP signals in HEK293T cells transfected with both split YFP halves, either a linear dimerization domain attached using low-diversity design principles or a structured dimerization domain designed for kissing loop-loop interactions. Efficient reconstitution is indicated by a strong yellow fluorescent signal. [Figure 10A]10A-10Z are exemplary synthetic nucleic acid molecules that can be used with the systems and methods. In some examples, the synthetic nucleic acid molecule has at least 80%, at least 85%, at least 90%, at least 95%, at least 98%, at least 99%, or 100% sequence identity to any one of SEQ ID NOs: 1 (Figures 10A-10B), 2 (Figures 10C-10E), 7 (Figure 10E), 8 (Figure 10F), 9 (Figure 10G), 10 (Figure 10H), 11 (Figure 10I), 12 (Figure 10J), 13 (Figure 10K), 14 (Figure 10L), 15 (Figure 10M), 16 (Figure 10N), 17 (Figure 10O), 18 (Figure 10P), 19 (Figure 10Q), 20 (Figures 10R-10U), and 21 (Figures 10V-10Z), but has a different target protein coding sequence. Thus, an intron region used with any of the systems or methods presented herein can have at least 80%, at least 85%, at least 90%, at least 95%, at least 98%, at least 99%, or 100% sequence identity to any of the intron sequences of SEQ ID NOs: 1, 2, 3, 4, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, or 21. For example, Figures 10A-D show exemplary (A, B) first synthetic molecule (SEQ ID NO: 1) and (C, D) second synthetic molecule (SEQ ID NO: 2) that can be used to express full-length YFP, while SEQ ID NOs: 3 and 4 present the corresponding synthetic intron portions without the YFP-encoding portion. In some examples, the synthetic intron sequence has at least 80%, at least 85%, at least 90%, at least 95%, at least 98%, at least 99%, or 100% sequence identity to SEQ ID NO: 3 or 4. Thus, portions of the coding sequence of any of the synthetic molecules presented herein (eg, nt 544-1032 of SEQ ID NO: 1 and nt 905-1141 of SEQ ID NO: 2) can be replaced with alternative portions of the coding sequence. [Figure 10B] Same as above. [Figure 10C] Same as above. [Figure 10D] Same as above. [Figure 10E] Same as above. [Figure 10F] Same as above. [Figure 10G] Same as above. [Figure 10H] Same as above. [Figure 10I] Same as above. [Figure 10J] Same as above. [Figure 10K] Same as above. [Figure 10L] Same as above. [Figure 10M] Same as above. [Figure 10N] Same as above. [Figure 10O] Same as above. [Figure 10P] Same as above. [Figure 10Q] Same as above. [Figure 10R] Same as above. [Figure 10S] Same as above. [Figure 10T] Same as above. [Figure 10U] Same as above. [Figure 10V] Same as above. [Figure 10W] Same as above. [Figure 10X] Same as above. [Figure 10Y] Same as above. [Figure 10Z] Same as above. [Figure 11] Figure 11 is a bar graph showing the reconstitution efficiency of random, complementary base-pairing binding domains of different lengths (50 bp, 100 bp, 150 bp, 200 bp, 300 bp, 400 bp, and 500 bp). Median YFP fluorescence intensity is compared between cells with comparable RFP and BFP transfection levels. n = 3 samples per condition. n = 3 samples per condition. [Figure 12]Figures 12A-12B show that including a splice enhancer in the synthetic intron increases reconstitution efficiency. Figure 12A is a schematic diagram of the 5'-N-terminal and 3'-C-terminal constructs (SEQ ID NOS: 1 and 2) used. (See Figure 1A for abbreviations.) Figure 12B is a bar graph showing the YFP fluorescence obtained after transfection of cells with SEQ ID NOS: 1 and 2 or their various truncations, indicated by Δ. n=3 samples per condition. [Figure 13] Figures 13A-13D show tracing of midline-crossing cortical neurons by reconstituting full-length flp recombinase (Flpo) from two fragments (SEQ ID NOs: 147 and 148). (A) Schematic representation of the 5' and 3' sequences used to reconstitute flpo (similar to the construct in Figure 12A). (B) Schematic representation of an injected flp-reporter mouse line injected with AAV viruses encoding N-flpo and C-flpo into the left and right cortical regions, respectively. (C and D) Show neuronal cell body and axonal labeling of cortical neurons projecting to the opposite hemisphere of the brain and therefore infected with both N-flpo and C-flpo viruses. Hoechst staining (nuclei) is shown for context. [Figure 14-1]Figures 14A-14D show the expression of oversized cargo (i.e., proteins encoded by long RNAs) in cell culture and in vivo in the mouse primary motor cortex. (A) Schematic representation of the 5' and 3' sequences used to reconstitute YFP, including a long stuffer sequence (uninterrupted open reading frame; SEQ ID NOs: 22 and 23, respectively). (B) Quantitative real-time PCR analysis of the reconstitution efficiency of oversized YFP constructs in HEK 293T cells. N=3 per condition. (C) Expression of reconstituted YFP proteins from full-length oversized YFP expression and split-REJ expression assessed by flow cytometry in transiently transfected HEK 293T cells. Median yellow fluorescence intensity is compared between cell populations with equal transfection control (blue and red) fluorescence for different conditions. The Y-axis indicates median yellow fluorescence intensity [au]. N=3 per condition. (D) Schematic of injection into mouse primary motor cortex and image of brain tissue 10 days after injection showing successful reconstitution of the long (2401 amino acid) YFP protein in vivo. [Figure 14-2] Same as above. [Figure 15-1]Figures 15A-15C show efficient reconstitution of full-length human coagulation factor VIII (FVIII) (2317 amino acids) with an N-terminal HA tag (in place of the N-terminal signal peptide). (A) Schematic representation of the 5' and 3' sequences used to reconstitute FVIII (SEQ ID NOs: 24 and 25, respectively). (B) PCR amplification of the junction. (C) Western blot showing FVIII expression. Lanes 1-3: Expression of full-length FVIII (the 290 kDa band indicates full-length, unprocessed FVIII). Lanes 4-6: Expression of reconstituted FVIII (the 290 kDa band indicates successfully reconstituted FVIII). Lanes 7 and 8: Expression of the N-terminus only indicates the absence of the 290 kDa full-length FVIII band. For all lanes: the expected protein processing products are observed ranging from approximately 75 kDa to approximately 210 kDa. FVIII is probed using a mouse anti-HA primary antibody. All lanes were loaded with 5 micrograms of clarified cellular protein extract and probed with GAPDH (rabbit anti-GAPDH) as a loading control. [Figure 15-2] Same as above. [Figure 16-1]Figures 16A-16F show efficient reconstitution of full-length human Abca4 (2300 amino acids) with a C-terminal FLAG tag. (A) Schematic representation of the 5' and 3' sequences used to reconstitute Abca4 (SEQ ID NOs: 20 and 21, respectively), along with a Sanger sequencing trace across the junction. (B) PCR amplification of the junction. (C) Schematic representation of the probes used to assay recombination of the 5' and 3' fragments. (D) Quantification of reconstitution efficiency by PCR after 2 days of expression in HEK 293T cells. N=2 per condition. (E) Western blot showing Abca4 expression. Lanes 1-3: Expression of full-length Abca4 (the approximately 260 kDa band represents full-length Abca4). Lanes 4-6: Expression of reconstituted Abca4 (the 260 kDa band represents successfully reconstituted Abca4). Lanes 7 and 8: No-transfection control (i.e., HEK 293t lysate only) shows the absence of any signal. Abca4 is probed using a mouse anti-FLAG primary antibody. All lanes were loaded with 5 micrograms of clarified cell protein extract. GAPDH (rabbit anti-GAPDH) is probed as a loading control. (F) Quantification of the Western blot in (E) normalized for differential BFP concentration. Data are shown normalized to the mean of the full-length expression control. [Figure 16-2] Same as above. [Figure 17] Figures 17A and 17B show (A) the kissing loop dimerization domain based on HIV-1 (N fragment, SEQ ID NO: 139, C fragment, SEQ ID NO: 140); and (B) the kissing loop dimerization domain based on HIV-2 (N fragment, SEQ ID NO: 141, C fragment, SEQ ID NO: 142). [Figure 18]Figures 18A-18C show efficient reconstitution of full-length mouse Otof (2,019 amino acids) with a C-terminal FLAG tag. The DNA sequences of the 5' and 3' molecules used are shown in SEQ ID NOs: 155 and 156. (A) Western blot showing Otof expression. Lanes 1-3: Expression of full-length Otof (the approximately 250 kDa band indicates full-length Otof). Lanes 4-6: Expression of reconstituted Otof (the 250 kDa band indicates successfully reconstituted Otof). Lane 7: No-transfection control (i.e., HEK 293t lysate only) shows the absence of any signal. Otof is probed using a mouse anti-FLAG primary antibody. All lanes were loaded with 5 micrograms of clarified cell protein extract. GAPDH (rabbit anti-GAPDH) is probed as a loading control. (B) Raw quantification of Western blots and (C) normalization for differential BFP concentration. Data are shown normalized to the mean of full-length expression controls. [Figure 19] Figures 19A-19C show efficient reconstitution of full-length human Myo7a (2243 amino acids) with a C-terminal FLAG tag. The DNA sequences of the 5' and 3' molecules used are shown in SEQ ID NOs: 157 and 158. (A) Western blot showing Myo7a expression. Lanes 1-3: Expression of full-length Myo7a (the approximately 270 kDa band indicates full-length Myo7a). Lanes 4-6: Expression of reconstituted Myo7a (the 270 kDa band indicates successfully reconstituted Myo7a). Lane 7: No-transfection control (i.e., HEK 293t lysate only) shows the absence of any signal. Myo7a is probed using a mouse anti-FLAG primary antibody. All lanes were loaded with 5 micrograms of clarified cell protein extract. GAPDH (rabbit anti-GAPDH) is probed as a loading control. (B) Raw quantification of the Western blot and (C) normalization for differential BFP concentration. Data are shown normalized to the mean of full-length expression controls. [Figure 20]Figures 20A-20D show efficient reconstitution of full-length DCas9-VPR (1951 amino acids). The DNA sequences of the 5' and 3' molecules used are shown in SEQ ID NOs: 159 and 160. (A) Western blot showing DCas9-VPR expression. Lanes 1-3: Expression of full-length DCas9-VPR (the approximately 250 kDa band indicates full-length DCas9-VPR). Lanes 4-6: Expression of reconstituted DCas9-VPR (the 250 kDa band indicates successfully reconstituted DCas9-VPR). Lane 7: No-transfection control (i.e., HEK 293t lysate only) shows the absence of any signal. DCas9-VPR is probed using a mouse anti-Cas9 primary antibody. All lanes were loaded with 5 micrograms of clarified cell protein extract. GAPDH (rabbit anti-GAPDH) is probed as a loading control. (B) Raw quantification of Western blots and (C) normalization for differential BFP concentration. Data are shown normalized to the mean of the full-length expression control. (D) Example of transcriptional activation of YFP expression plasmids in HEK 293T cells. Full-length dCas9-VPR (upper panel) or two-way split-REJ dual dCas9-VPR (lower panel) are transiently transfected with a non-targeting guide RNA expression plasmid (left panel) or a UAS-targeting guide RNA expression plasmid (right panel). All cells are also transfected with a UAS-YFP plasmid, which is transcriptionally inactive until dCas9-VPR is targeted to the upstream region of a minimal promoter that drives yellow fluorescent protein expression. Red fluorescent protein (RFP) is expressed with the N-terminal fragment of dCas9-VPR, and blue fluorescent protein (BFP) is expressed with full-length dCas9-VPR or the C-terminal fragment of dCas9-VPR, respectively. RFP and BFP serve as transfection controls. When both the full-length dCas9-VPR paired with the UAS-targeting guide RNA and the bidirectional split dCas9-VPR were expressed, yellow fluorescent protein expression was observed, confirming the functionality of the reconstituted full-length protein. [Figure 21]Figures 21A-21D show efficient reconstitution of a full-length humanized prime editor (2118 amino acids). The DNA sequences of the 5' and 3' molecules used are shown in SEQ ID NOs: 161 and 162. (A) Western blot showing expression of the prime editor. Lanes 1-3: Expression of the full-length prime editor (the approximately 260 kDa band indicates the full-length prime editor). Lanes 4-6: Expression of the reconstituted prime editor (the 260 kDa band indicates the successfully reconstituted prime editor). Lane 7: No transfection control (i.e., HEK 293t lysate only) shows the absence of any signal. The prime editor is probed using a mouse anti-Cas9 primary antibody. All lanes were loaded with 5 micrograms of clarified cell protein extract. GAPDH (rabbit anti-GAPDH) is probed as a loading control. (B) Raw quantification of the Western blot and (C) normalization for differential BFP concentration. Data are shown normalized to the mean of the full-length expression control. (D) Prime editor-induced G-to-T transversion mutations were induced at the FANCF and VEGFA3 loci in HEK293T cells. The top panel shows the sequence context of the FANCF and VEGFA3 loci, respectively. The gray arrow indicates the sequence targeted by the prime editor guide RNA (pegRNA). The protospacer adjacent motif (PAM) is indicated by a gray box. The G targeted for transversion to T is highlighted in the sequence. The genomic locus was sequenced using Sanger sequencing in three conditions. The top panel shows a representative Sanger trace for the unedited wild-type condition. The second panel from the top shows a representative Sanger trace representing the full-length expressed prime editor construct. The region highlighted by the black box indicates the appearance of a T band in Sanger sequencing, indicating successful integration of the edit into a portion of the cells. The bottom panel shows a representative Sanger trace for cells edited using the two-way split-recombinant prime editor.The appearance of a T trace (black box) demonstrates the functionality of the prime editor when reassembled from the two fragments. [Figure 22]Figures 22A-22C show efficient reconstitution of a full-length humanized cytosine base editor (AncBE4) (1854 amino acids). The DNA sequences of the 5' and 3' molecules used are shown in SEQ ID NOs: 163 and 164. (A) Western blot showing AncBE4 expression. Lanes 1-3: Expression of full-length AncBE4 (the approximately 230 kDa band indicates full-length AncBE4). Lanes 4-6: Expression of reconstituted AncBE4 (the 230 kDa band indicates successfully reconstituted AncBE4). Lane 7: No-transfection control (i.e., HEK 293t lysate only) shows the absence of any signal. AncBE4 is probed using a mouse anti-Cas9 primary antibody. All lanes were loaded with 5 micrograms of clarified cell protein extract. GAPDH (rabbit anti-GAPDH) is probed as a loading control. (B) Raw quantification of the Western blot. Data are shown normalized to the mean of the full-length expression control. (C) AncBE4-induced C-to-T transposition mutations were induced at the EMX1 and HEK site 3 loci in HEK293T cells. The top panel shows the sequence context of the EMX1 and HEK site 3 loci, respectively. The gray arrow indicates the sequence targeted by the AncBE4 guide RNA (sgRNA). The protospacer adjacent motif (PAM) is indicated by a gray box. The C targeted for transposition to a T is highlighted in the sequence. The genomic locus was sequenced using Sanger sequencing in three conditions. The top panel shows a representative Sanger trace for the unedited wild-type condition. The second panel from the top shows a representative Sanger trace representing the full-length expressed AncBE4 construct. The region highlighted by the black box indicates the appearance of a T band in Sanger sequencing, indicating successful integration of the edit into a portion of the cells. The bottom panel shows a representative Sanger trace for cells edited using the two-way split reconstituted AncBE4. The appearance of the T trace (black box) demonstrates the functionality of AncBE4 when reconstituted from the two fragments. [Figure 23]Figures 23A-23C show efficient reconstitution of the full-length humanized adenine base editor (Abe8e) (1606 amino acids). The DNA sequences of the 5' and 3' molecules used are shown in SEQ ID NOs: 165 and 166. (A) Western blot showing Abe8e expression. Lanes 1-3: Expression of full-length Abe8e (the approximately 230 kDa band indicates full-length Abe8e). Lanes 4-6: Expression of reconstituted Abe8e (the 230 kDa band indicates successfully reconstituted Abe8e). Lane 7: No-transfection control (i.e., HEK 293t lysate only) shows the absence of any signal. Abe8e is probed using a mouse anti-Cas9 primary antibody. All lanes were loaded with 5 micrograms of clarified cell protein extract. GAPDH (rabbit anti-GAPDH) is probed as a loading control. (B) Raw quantification of the Western blot. Data are shown normalized to the mean of the full-length expression control. (C) Abe8e-induced A-to-G transition mutations were induced at the BCL11A and HGB1 / 2 loci in HEK293T cells. The top panel shows the sequence context of the BCL11A and HGB1 / 2 loci, respectively. The gray arrow indicates the sequence targeted by the Abe8e guide RNA (sgRNA). The protospacer adjacent motif (PAM) is indicated by a gray box. The A targeted for G transition is highlighted within the sequence. The genomic locus was sequenced using Sanger sequencing in three conditions. The top panel shows a representative Sanger trace for the unedited wild-type condition. The second panel from the top shows a representative Sanger trace representing the full-length expressed Abe8e construct. The region highlighted by the black box indicates the appearance of a G band in Sanger sequencing, indicating successful integration of the edit into a portion of the cells. The bottom panel shows a representative Sanger trace for cells edited using the two-way split reconstituted Abe8e. The appearance of the G trace (black box) demonstrates the functionality of Abe8e when reconstituted from the two fragments. [Figure 24A]The effects of downstream intron splicing enhancer (DISE) and intron splicing enhancer (ISE) and acceptor sequences on the efficiency of RNA end-joining. (A) Schematic diagram of the screening setup. The 5' fragment is an RNA molecule transcribed from a DNA construct using the human CMV promoter and enhancer. The resulting RNA molecule contains a long stuffer open reading frame to simulate a large cargo size. This stuffer sequence ends with a 2A self-cleaving peptide sequence, followed by the coding region for the 5' fragment of yellow fluorescent protein (n-yfp). The 5' fragment of yfp ends with a splice donor site (SD). This splice donor site is followed by the 5' intron portion of the RNA end-joining module. To determine the effect of DISE and ISE sequences on the efficiency of the RNA end-joining reaction, the 5' intron portion is subdivided into three fragments: from 5' to 3': ds: downstream segment; m: middle intron segment; and dd: donor distal segment. The 5' intronic portion is followed by a trimodal kissing loop RNA dimerization domain. The message is terminated by a short polyadenylation signal. The overall length of this 5' RNA molecule is approximately 4 kb to simulate the reassembly scenario of a large cargo. The 3' fragment is an RNA molecule transcribed from a DNA construct using the human CMV promoter and enhancer. The 3' fragment begins with a trimodal kissing loop RNA dimerization domain complementary to that of the RNA molecule encoding the 5' fragment. The dimerization domain is followed by the 3' intronic portion of the RNA end-binding module. This 3' intronic portion is subdivided into three segments: ad: acceptor distal segment; m: middle intron segment; and ap: acceptor proximal segment. The acceptor proximal segment contains a branch point and a polypyrimidine tract variation, both of which are essential for spliceosome-mediated RNA binding reactions. The splice acceptor (SA) site is followed by the 3'yfp coding sequence, followed by a self-cleaving 2A sequence, followed by a long stuffer open reading frame. The message is terminated by an SV40 polyadenylation signal.The overall length of the 3' RNA molecule is approximately 4 kb to simulate a reconstitution scenario for a large cargo. Association of the two RNA molecules (5' and 3' fragments) is mediated by a trimodal kissing-loop RNA dimerization domain, while spliceosome recruitment and RNA end-joining are mediated by an intron segment. Successful RNA end-joining results in the reconstitution of the YFP open reading frame and subsequent YFP translation. (B) Median YFP fluorescence intensity determined by flow cytometry is shown for multiple intron configurations. In the first grouping (bars 1–9), a selection of potential downstream intron splicing enhancer sequences was paired with the consensus splice donor site (GTAAGTATT in the DNA construct and GUAAGUAUU in the RNA sequence) shown in bars 1–8. These are compared to the consensus splice donor (ds9) followed by a scrambled sequence consisting of equal portions of all four bases. In the second grouping, m1–m16, the selection of potential intron splicing enhancers was compared with a scrambled sequence (m16). The final grouping compared the selection of potential strong branch points, polypyrimidine tracts, and splice acceptors. The reference construct consisted of a scrambled sequence and consensus donor at all nonvariable positions, followed by a scrambled sequence and consensus splice acceptor at the ds position (where the entire polypyrimidine tract was composed of Ts in the DNA construct and Us in the RNA fragment, respectively). (C) List of the different DISE, ISE, and splice acceptor elements used. [Figure 24B] Same as above. [Figure 24C] Same as above. DETAILED DESCRIPTION OF THE INVENTION

[0014] Sequence Listing The nucleic acid and amino acid sequences listed in the accompanying sequence listing are shown using standard letter abbreviations for nucleotide bases and three-letter codes for amino acids, as defined in 37 C.F.R. 1.822. Only one strand of each nucleic acid sequence is shown, but it should be understood that any reference to the shown strand includes the complementary strand. The sequence listing was submitted as a 157KB ASCII text file created on September 30, 2020, and is incorporated herein by reference. In the accompanying sequence listing:

[0015] SEQ ID NOs: 1 and 2 are the N- and C-terminal sequences, respectively, used to express full-length YFP. SEQ ID NO: 1 contains the CMV promoter nt 1 to 543, the YFP coding sequence nt 544 to 1032, the synthetic intron nt 1033 to 1436, and the untranslated poly(A) region nt 1437 to 1491. SEQ ID NO: 2 contains the CMV promoter nt 1 to 522, the synthetic intron nt 523 to 904, the YFP coding sequence nt 905 to 1141, and the untranslated poly(A) region nt 1142 to 1302.

[0016] SEQ ID NOs: 3 and 4 are 5' and 3' intron sequences, respectively, that can be used to express a desired full-length protein, where the N-terminal portion of the full-length protein can be added at nt 1 of SEQ ID NO: 3, and the C-terminal portion of the full-length protein can be added at nt 382 of SEQ ID NO: 4.

[0017] SEQ ID NOs: 5 and 6 are the N- and C-terminal coding sequences, respectively, used to express full-length YFP.

[0018] SEQ ID NO: 7 is an exemplary synthetic intron dimerization domain (FIG. 10E).

[0019] SEQ ID NO:8 is an exemplary synthetic intron without an intronic splicing enhancer (FIG. 10F).

[0020] SEQ ID NO: 9 is an exemplary synthetic intron without an intronic splicing enhancer (FIG. 10G).

[0021] SEQ ID NO: 10 is an exemplary synthetic intron without an intronic splicing enhancer (FIG. 10H).

[0022] SEQ ID NO: 11 is an exemplary synthetic intron without a binding domain (FIG. 10I).

[0023] SEQ ID NO: 12 is an exemplary synthetic intron with a dimerization domain (FIG. 10J).

[0024] SEQ ID NO: 13 is an exemplary synthetic intron with a dimerization domain (FIG. 10K).

[0025] SEQ ID NO: 14 is an exemplary synthetic intron without an intronic splicing enhancer (FIG. 10L).

[0026] SEQ ID NO: 15 is an exemplary synthetic intron with only DISE (Figure 10M).

[0027] SEQ ID NO: 16 is an exemplary synthetic intron without HHrz (FIG. 10N).

[0028] SEQ ID NO: 17 is an exemplary synthetic intron without an intronic splicing enhancer (FIG. 10O).

[0029] SEQ ID NO: 18 is an exemplary U12-dependent intron with a binding domain (FIG. 10P).

[0030] SEQ ID NO: 19 is an exemplary U12-dependent intron with a binding domain (FIG. 10Q).

[0031] SEQ ID NOs: 20 and 21 are the N- and C-terminal DNA sequences, respectively, used to express RNA (pre-mRNA) resulting in full-length Abca4. In SEQ ID NO: 20, the sequence corresponding to the N-terminal Abca4 coding region is located between nt 22 and nt 3702, a synthetic intron is located between nt 3703 and nt 3912, and an untranslated poly(A) region is located between nt 3921 and nt 3969. SEQ ID NO: 20 also contains a splice donor between nt 3703 and nt 3711, a rat FGFR2 DISE between nt 3714 and nt 3737, a cTNT intron splicing enhancer between nt 3747 and nt 3770, an M2 intron splicing enhancer between nt 3782 and nt 3794, and a kissing loop dimerization domain between nt 3801 and nt 3975. In SEQ ID NO:21, nts 1 to 228 are a synthetic intron, nts 229 to 3366 are the C-terminal Abca4 coding region, nts 3367 to 3447 are a FLAG epitope tag, and nts 3476 to 3607 are an untranslated poly(A) region (signal). SEQ ID NO:21 also contains a kissing loop dimerization domain at nts 3 to 114, an M2 intron splicing enhancer at nts 121 to 133, a cTNT intron splicing enhancer at nts 140 to 163, an M2 intron splicing enhancer at nts 175 to 187, a branchpoint motif at nts 194 to 201, a polypyrimidine tract at nts 207 to 226, and a splice acceptor at nt 228.

[0032] SEQ ID NOs: 22 and 23 are the N-terminal and C-terminal DNA sequences, respectively, used to express the RNA (pre-mRNA) that yields full-length YFP, each containing a splice enhancer. In SEQ ID NO: 22, the N-terminal YFP coding region is from nt 22 to 3702, nt 3703 to 3912 is a synthetic intron, and nt 3921 to 3969 is an untranslated poly(A) region. SEQ ID NO: 22 also contains a splice donor from nt 3703 to 3711, a rat FGFR2 DISE from nt 3714 to 3737, a cTNT intron splicing enhancer from nt 3747 to 3770, an M2 intron splicing enhancer from nt 3782 to 3794, and a kissing loop dimerization domain from nt 3801 to 3975. In SEQ ID NO: 23, nts 1 to 225 are a synthetic intron, nts 226 to 3747 are a C-terminal YFP coding region, and nts 3748 to 3912 are an untranslated poly(A) region. SEQ ID NO: 23 also contains a kissing loop dimerization domain from nts 3 to 114, an M2 intron splicing enhancer from nts 118 to 130, a cTNT intron splicing enhancer from nts 137 to 160, an M2 intron splicing enhancer from nts 172 to 184, a branchpoint motif from nts 191 to 198, a polypyrimidine tract from nts 204 to 223, and a splice acceptor at nt 225.

[0033] SEQ ID NOs: 24 and 25 are the N- and C-terminal sequences, respectively, used to express RNA (pre-mRNA) resulting in full-length human factor VIII. In SEQ ID NO: 24, the N-terminal FVIII coding region with an N-terminal HA epitope tag is located between nt 22 and nt 3561, nt 3562 and nt 3771 is a synthetic intron, and nt 3780 and nt 3828 is an untranslated poly(A) region. SEQ ID NO: 24 also contains a splice donor between nt 3562 and nt 3570, a rat FGFR2 DISE between nt 3573 and nt 3596, a cTNT intron splicing enhancer between nt 3606 and nt 3629, an M2 intron splicing enhancer between nt 3641 and nt 3653, and a kissing loop dimerization domain between nt 3660 and nt 3834. In SEQ ID NO: 25, nt 1 to 225 are a synthetic intron, nt 226 to 3636 are a C-terminal FVIII coding region, and nt 3665 to 3797 are an untranslated poly(A) region. SEQ ID NO: 25 also contains a splice donor at nt 3703 to 3711, a rat FGFR2 DISE at nt 3714 to 3737, a cTNT intron splicing enhancer at nt 3747 to 3770, an M2 intron splicing enhancer at nt 3782 to 3794, and a kissing loop dimerization domain at nt 3801 to 3975.

[0034] SEQ ID NOs: 26-136 are exemplary splicing enhancers that can be used in the systems presented herein (e.g., 118, 120, 156 in Figure 6A).

[0035] SEQ ID NOs: 137 and 138 are exemplary splice donor sequences.

[0036] SEQ ID NOs: 139 and 140 are the N and C fragments, respectively, of the HIV-1 based kissing loop dimerization domain.

[0037] SEQ ID NOs: 141 and 142 are the N and C fragments, respectively, of the kissing loop dimerization domain based on HIV-2.

[0038] SEQ ID NO: 143 is an exemplary cryptic splice acceptor sequence.

[0039] SEQ ID NO: 144 is an exemplary branch point consensus sequence.

[0040] SEQ ID NOs: 145 and 146 are the N- and middle sequences, respectively, used together with SEQ ID NO: 2 (C-terminal fragment) to express full-length YFP. In SEQ ID NO: 145, nt 1 to 543 is the CMV promoter sequence, nt 544 to 849 is the N-terminal YFP coding region, and nt 850 to 1305 is a synthetic intron. In SEQ ID NO: 146, nt 1 to 522 is the CMV promoter sequence, nt 523 to 901 is a synthetic intron, nt 902 to 1084 is the middle YFP coding region, and nt 1085 to 1543 is an untranslated poly(A) region.

[0041] SEQ ID NOs: 147 and 148 are the 5' and 3' synthetic sequences, respectively, used to express full-length Flpo. In SEQ ID NO: 147, nt 1 to 540 is the CMV promoter sequence, nt 541 to 1112 is the N-terminal Flpo coding region, and nt 1113 to 1571 is a synthetic intron. In SEQ ID NO: 148, nt 1 to 522 is the CMV promoter sequence, nt 523 to 904 is a synthetic intron, nt 905 to 1604 is the C-terminal Flpo coding region, and nt 1605 to 1765 is an untranslated poly(A) region.

[0042] SEQ ID NOs: 149 and 150 are exemplary low diversity sequences.

[0043] SEQ ID NOs: 151 and 152 are exemplary splice donor consensus sequences.

[0044] SEQ ID NO: 153 is an exemplary kissing loop based on the HIV-2 kissing loop dimerization domain (SEQ ID NOs: 141 and 142, Figure 17B).

[0045] SEQ ID NO: 154 is an exemplary Kozak-enhanced start codon.

[0046] SEQ ID NOs: 155 and 156 are exemplary constructs that can be used to express the mouse Otof coding sequence in vivo. SEQ ID NO: 155 is used to generate the N-terminal Otof RNA. SEQ ID NO: 155 contains the human CMV enhancer and promoter from nt 1 to 522, the putative transcription start site from nt 523, and the polyadenylation signal from nt 4263 to 4311. SEQ ID NO: 155 encodes the N-terminal Otof RNA elements as follows: a 5' untranslated region containing a Kozak sequence from nt 523 to 546; a 5' otoferlin coding sequence from nt 547 to 4044; a 5' synthetic intron sequence from nt 4045 to 4142; a 5' trimodal kissing loop dimerization domain from nt 4143 to 4254; and a linker from nt 4255 to 4262. SEQ ID NO: 155 is used to generate the C-terminal Otof RNA. It contains the human CMV enhancer and promoter from nt 1 to 522, a putative transcription start site from nt 523, and a polyadenylation signal from nt 3335 to 3467. It encodes the C-terminal Otof RNA element as follows: a 3' trimodal kissing-loop dimerization domain from nt 525 to 636; a 3' synthetic intron sequence from nt 637 to 747; a 3' otoferlin coding sequence from nt 748 to 3225; a C-terminal 3xFlag tag from nt 3226 to 3306; and a linker from nt 3307 to 3334.

[0047] SEQ ID NOs:157 and 158 are exemplary constructs that can be used to express the human myosin VIIA (Myo7a) coding sequence in vivo. SEQ ID NO:157 is used to generate N-terminal Myo7a RNA. SEQ ID NO:157 contains the human CMV enhancer and promoter from nt 1 to 522, the putative transcription start site from nt 523, and the polyadenylation signal from nt 4344 to 4392. SEQ ID NO:157 encodes the N-terminal Myo7A RNA elements as follows: 5' untranslated region containing a Kozak sequence from nt 523 to 543; 5' Myo7a coding sequence from nt 544 to 4125; 5' synthetic intron sequence from nt 4126 to 4223; 5' trimodal kissing loop dimerization domain from nt 4224 to 4335; and a linker from nt 4336 to 4343. SEQ ID NO:158 is used to generate C-terminal Myo7a RNA. SEQ ID NO:158 contains the human CMV enhancer and promoter from nt 1 to 522, a putative transcription start site at nt 523, and a polyadenylation signal from nt 3923 to 4055. SEQ ID NO:158 encodes the C-terminal Myo7a RNA elements as follows: 3' trimodal kissing loop dimerization domain from nt 525 to 636; 3' synthetic intron sequence from nt 637 to 747; 3' Myo7a coding sequence from nt 748 to 3813; C-terminal 3xFlag tag from nt 3814 to 3894; and a linker from nt 3895 to 3922.

[0048] SEQ ID NOs: 159 and 160 are exemplary constructs that can be used to express in vivo the full-length, enzymatically inactive Cas9 (dCas9-VPR) coding sequence fused to the VPR transcriptional activator domain. SEQ ID NO: 159 is used to generate N-terminal DCas9-VPR RNA. SEQ ID NO: 159 contains the human CMV enhancer and promoter from nt 1 to 522, the putative transcription start site from nt 523, and the polyadenylation signal from nt 4112 to 4161. SEQ ID NO: 159 encodes the N-terminal DCas9-VPR RNA elements as follows: 5' untranslated region containing a Kozak sequence from nt 523 to 543; 5' DCas9-VPR coding sequence from nt 544 to 3894; 5' synthetic intron sequence from nt 3895 to 3992; 5' trimodal kissing loop dimerization domain from nt 3993 to 4104; and linker from nt 4105 to 4112. SEQ ID NO: 160 is used to generate the C-terminal DCas9-VPR RNA. SEQ ID NO: 160 includes the human CMV enhancer and promoter from nt 1 to 522, the putative transcription start site from nt 523, and the polyadenylation signal from nt 3278 to 3410. SEQ ID NO: 160 encodes the C-terminal DCas9-VPR RNA elements as follows: a 3' trimodal kissing loop dimerization domain from nt 525 to 636; a 3' synthetic intron sequence from nt 637 to 747; a 3' DCas9-VPR coding sequence from nt 748 to 3249; and a linker from nt 3250 to 3277.

[0049] SEQ ID NOs: 161 and 162 are exemplary constructs that can be used to express the full-length humanized Cas9 prime editor (prime editor) coding sequence in vivo. SEQ ID NO: 161 encodes the N-terminal prime editor sequence as follows: human CMV enhancer and promoter nt 1 to 522; putative transcription start site nt 523; 5' untranslated region containing a Kozak sequence nt 523 to 543; 5' prime editor coding sequence nt 544 to 3894; 5' synthetic intron sequence nt 3895 to 3992; 5' trimodal kissing loop dimerization domain nt 3993 to 4104; linker nt 4105 to 4112; and polyadenylation signal nt 4112 to 4161. SEQ ID NO: 162 encodes the C-terminal prime editor sequence as follows: human CMV enhancer and promoter nt 1-522; putative transcription start site nt 523; 3' trimodal kissing loop dimerization domain nt 525-636; 3' synthetic intron sequence nt 637-747; 3' prime editor coding sequence nt 748-3750; linker nt 3751-3778; polyadenylation signal nt 3779-3911.

[0050] SEQ ID NOs: 163 and 164 are exemplary constructs that can be used to express the full-length humanized cytosine base editor (AncBE4) coding sequence in vivo. SEQ ID NO: 163 encodes the N-terminal AncBE4 sequence as follows: human CMV enhancer and promoter nt 1 to 522; putative transcription start site nt 523; 5' untranslated region containing a Kozak sequence nt 523 to 540; 5' AncBE4 coding sequence nt 541 to 2892; 5' synthetic intron sequence nt 2893 to 2990; 5' trimodal kissing loop dimerization domain nt 2991 to 3102; linker nt 3103 to 3110; and polyadenylation signal nt 3111 to 3159. SEQ ID NO: 164 encodes the C-terminal AncBE4 sequence as follows: human CMV enhancer and promoter nt 1 to 522; putative transcription start site nt 523; 3' trimodal kissing loop dimerization domain nt 525 to 636; 3' synthetic intron sequence nt 637 to 747; 3' AncBE4 coding sequence nt 748 to 3957; linker nt 3958 to 3982; polyadenylation signal nt 3983 to 4115.

[0051] SEQ ID NOs: 165 and 166 are exemplary constructs that can be used to express the full-length humanized adenine base editor (Abe8e) coding sequence in vivo. SEQ ID NO: 165 encodes the N-terminal Abe8e sequence as follows: human CMV enhancer and promoter nt 1 to 522; putative transcription start site nt 523; 5' untranslated region containing a Kozak sequence nt 523 to 540; 5' Abe8e coding sequence nt 541 to 2706; 5' synthetic intron sequence nt 2707 to 2804; 5' trimodal kissing loop dimerization domain nt 2805 to 2916; linker nt 2917 to 2924; and polyadenylation signal nt 2925 to 2973. SEQ ID NO: 166 encodes the C-terminal Abe8e sequence as follows: human CMV enhancer and promoter nt 1 to 522; putative transcription start site nt 523; 3' trimodal kissing loop dimerization domain nt 525 to 636; 3' synthetic intron sequence nt 637 to 747; 3' Abe8e coding sequence nt 748 to 3399; linker nt 3400 to 3427; polyadenylation signal nt 3428 to 3560.

[0052] SEQ ID NO: 167 is an exemplary kissing loop domain (GATTTTTGACCTGCTCGATTGTCCACTGCGAGCAGGTCTTTTGGAGTCGGGCGAGGCGGAAGCCCGACTCCTTTTGGCATGCACGCTAGCCGCGTCGTGCATGCCTTTTATC).

[0053] SEQ ID NO: 168 is an exemplary ISE, M2(GGGTTATGGGACC).

[0054] SEQ ID NO: 169 is an exemplary ISE, cTNT(GGCTGAGGGAAGGACTGTCCTGGG).

[0055] SEQ ID NO: 170 is an exemplary DISE, rat FGFR2 (CTCTTTCTTTCCATGGGTTGGCCT).

[0056] SEQ ID NOs: 171 and 172 are exemplary constructs that can be used to express the full-length YFP coding sequence. SEQ ID NO: 171 encodes the N-terminal YFP sequence as follows: human CMV enhancer and promoter nt 1 to 522; putative transcription start site nt 523; 5' untranslated region containing the Kozak sequence nt 523 to 543; 5' stuffer open reading frame nt 544 to 3654; self-cleaving 2A sequence nt 3655 to 3729; 5' yellow fluorescent protein segment nt 3730 to 4224; 5' synthetic intron sequence (variable) nt 4225 to 4294; 5' trimodal kissing loop dimerization domain (uppercase): nt 4295 to 4406; linker nt 4407 to 4414; and polyadenylation signal nt 4415 to 4463. SEQ ID NO: 172 encodes the C-terminal YFP sequence as follows: Name: 3' intron screening split YFP; human CMV enhancer and promoter nt 1-522; putative transcription start site nt 523; 3' trimodal kissing loop dimerization domain nt 525-636; 3' synthetic intron sequence (variable) nt 637-706; 3' yfp coding sequence nt 707-940; self-cleaving 2A sequence nt 941-1006; 3' stuffer open reading frame nt 1007-4228; linker nt 4229-4265; polyadenylation signal nt 4257-4388.

[0057] SEQ ID NOs: 173 to 180 are exemplary intron splicing enhancer sequences.

[0058] SEQ ID NO: 181 is a scrambled sequence.

[0059] SEQ ID NOs: 182 to 196 are exemplary intron splicing enhancer sequences.

[0060] SEQ ID NOs: 197 to 198 are scrambled sequences.

[0061] SEQ ID NOs: 199-203 are exemplary intron splicing enhancer sequences.

[0062] SEQ ID NO: 204 is a scrambled sequence.

[0063] SEQ ID NO: 205 is an exemplary branch point sequence (TACTAACA).

[0064] SEQ ID NO: 206 is an exemplary polyadenylation signal AATAAAAATATCTTTATTTTCATTACATCTGTGTGTTGGTTTTTTGTGTG.

[0065] Detailed Description Unless otherwise specified, technical terms are used according to conventional usage.The definitions of common terms in molecular biology can be found in Benjamin Lewin, Genes VII, published by Oxford University Press, 1999;Kendrew et al. (eds.), The Encyclopedia of Molecular Biology, published by Blackwell Science Ltd., 1994;And Robert A. Meyers (ed.), Molecular Biology and Biotechnology: a Comprehensive Desk Reference, published by VCH Publishers, Inc., 1995;And other similar references.

[0066] As used herein, the singular forms "a," "an," and "the" refer to both the singular and the plural unless the context clearly indicates otherwise. As used herein, the term "comprises" means "includes." Thus, "comprising a nucleic acid molecule" does not exclude other elements and means "including a nucleic acid molecule." It should further be understood that any and all base sizes given for nucleic acids are approximate unless otherwise specified and are provided for illustrative purposes. While many methods and materials similar or equivalent to those described herein can be used, certain suitable methods and materials are described below. In case of conflict, the present specification, including explanations of terms, will control. Furthermore, the materials, methods, and examples are merely illustrative and not limiting. All references, including patent applications and patents, and GenBank accession numbers, are incorporated herein by reference in their entirety.

[0067] In order to facilitate review of the various embodiments disclosed herein, the following explanations of specific terms are provided:

[0068] Administration: Providing or administering to a subject an agent, such as a therapeutic nucleic acid molecule or other therapeutic agent, provided herein by any effective route. Exemplary administration routes include, but are not limited to, injection (e.g., subcutaneous, intramuscular, intradermal, intraperitoneal, intrathecal, intratumoral, intraosseous, and intravenous), transdermal, intranasal, and inhalation routes. Administration can be systemic or local.

[0069] Aptamer: A nucleic acid molecule (e.g., DNA or RNA) that binds with high affinity and specificity to a specific target agent or molecule. Aptamers can be used as dimerization domains in the nucleic acid molecules disclosed herein. In one example, two aptamers can bind to each other, for example, by canonical base pairing, non-canonical base pairing interactions, non-base pairing interactions, or a combination thereof, to mediate dimer formation. In one example, an aptamer allows RNA dimer formation (and subsequent recombination) only in the presence of one or more targets recognized by the aptamers. Aptamers have been obtained by a combinatorial selection process called systematic evolution of ligands by exponential enrichment (SELEX) (e.g., Ellington et al., Nature 1990, 346, 818-822; Tuerk and Gold Science 1990, 249, 505-510; Liu et al., Chem. Rev. 2009, 109, 1948-1998; Shamah et al., Acc. Chem. Res. 2008, 41, 130-138; Famulok, et al., Chem. Rev. 2007, 107, 3715-3743; Manimala et al., Recent Dev. Nucleic Acids Res. 2004, 1, 207-231; Famulok et al., Acc. Chem. Res. 2000, 33, 591-599; Hesselberth, et al., Rev. Mol. Biotech. 2000, 74, 15-25; Wilson et al., Annu. Rev. Biochem. 1999, 68, 611-647; Morris et al., Proc. Natl. Acad. Sci. USA 1998, 95, 2902-2907). In such a process, DNA or RNA molecules capable of binding to a desired target molecule are synthesized. 14 ~10 15Aptamers are selected from nucleic acid libraries consisting of diverse sequences through repeated steps of selection, amplification, and mutation. The affinity of aptamers for their targets can compete with that of antibodies, with dissociation constants in the low picomolar range (Morris et al., Proc. Natl. Acad. Sci. USA 1998, 95, 2902-2907; Green et al., Biochemistry 1996, 35, 14413-14424).

[0070] Specific aptamers have been identified for a wide range of targets, from small organic molecules such as adenosine to proteins such as thrombin, and even viruses and cells (Liu et al., Chem. Rev. 2009, 109, 1948-1998; Lee et al., Nucleic Acids Res. 2004, 32, D95-D100; Navani and Li, Curr. Opin. Chem. Biol. 2006, 10, 272-281; ​​Song et al., TrAC, Trends Anal. Chem. 2008, 27, 108-117).For example, metal ions such as Zn(II) (Ciesiolka et al., RNA 1: 538-550, 1995) and Ni(II) (Hofmann et al., RNA, 3: 1289-1300, 1997); nucleotides such as adenosine triphosphate (ATP) (Huizenga and Szostak, Biochemistry, 34: 656-665, 1995); and guanine (Kiga et al., Nucleic Acids Res., 26: 1755-60, 1998); cofactors such as NAD (Kiga et al., Nucleic Acids Res., 26: 1755-60, 1998) and flavins (Lauhon and Szostak, J. Am. Chem. Soc., 117: 1246-57, 1995); antibiotics such as viomycin (Wallis et al., Chem. Biol. 4: 357-366, 1997) and streptomycin (Wallace and Schroeder, RNA 4: 112-123, 1998); proteins such as HIV reverse transcriptase (Chaloin et al., Nucleic Acids Res., 30: 4001-8, 2002) and hepatitis C virus RNA-dependent RNA polymerase (Biroccio et al., J. Virol. 76: 3688-96, 2002); toxins such as cholera toxin and staphylococcal enterotoxin B (Bruno and Kiel, BioTechniques, 32: pp. 178-180 and 182-183, 2002); and bacterial spores such as Bacillus anthracis (Bruno and Aptamers that recognize the nucleotide sequence of the nucleotide sequence of interest (Kiel, Biosensors & Bioelectronics, 14: 457-464, 1999) are available.

[0071] Binding: An association between two substances or molecules, such as hybridization between one nucleic acid molecule and another nucleic acid molecule (or itself), or binding between an aptamer and its target, such as between two dimerization domains. An oligonucleotide molecule and another nucleic acid molecule bind or are stably bound when there are a sufficient number of complementary base pairs between the oligonucleotide molecules and the target nucleic acid to allow for detection of binding. In some examples, binding between nucleic acid molecules can occur directly. In some examples, binding between nucleic acid molecules can occur indirectly, for example, through an intermediate molecule. Either direct or indirect binding can occur by standard base pairing, non-standard base pairing interactions, non-base pairing interactions, or combinations thereof. Non-standard base pairing interactions can occur by any stabilization means known to those skilled in the art, including, but not limited to, Hoogsteen base pairs and wobble base pairs. Non-base pairing interactions can include binding through an intermediate molecule. In some examples, direct binding is between kissing loop dimerization domains. In some examples, direct binding is between low diversity dimerization domains. In some embodiments, the direct binding is between aptamer regions. In some embodiments, the direct binding between aptamer regions involves non-canonical base pairing interactions. In some embodiments, the direct binding between aptamer regions involves canonical base pairing and non-canonical base pairing interactions. In some embodiments, the indirect binding occurs through a nucleic acid bridge. In some embodiments, the nucleic acid bridge is mRNA. A non-limiting example of a nucleic acid bridge is shown in FIG. 7B. In some embodiments, the indirect binding occurs through an aptamer molecule. A non-limiting example of indirect binding through an aptamer molecule is shown in FIG. 7A. In some embodiments, the indirect binding through an aptamer molecule involves non-base pairing interactions between the aptamer molecule and the binding region. In some embodiments, the indirect binding through an aptamer molecule involves non-base pairing interactions between the aptamer molecule and the binding region, and base pairing interactions between the binding regions.

[0072] C-terminal portion: A region of a protein sequence comprising a contiguous stretch of amino acids beginning at or near the C-terminal residue of the protein. The C-terminal portion of a protein can be defined by a contiguous stretch of amino acids (e.g., several amino acid residues).

[0073] Cancer: A malignant tumor characterized by abnormal or uncontrolled cell growth. Other features often associated with cancer include metastasis, interference with the normal function of nearby cells, release of abnormal levels of cytokines or other secretory products, and suppressed or exacerbated inflammatory or immunological responses, and infiltration of surrounding or distant tissues or organs, such as lymph nodes. "Metastatic disease" refers to cancer cells that leave the original tumor site and travel to other parts of the body, for example, via the bloodstream or lymphatic system.

[0074] Complementarity: The ability of a nucleic acid to form hydrogen bond(s) with another nucleic acid sequence, either by conventional Watson-Crick base pairing or other non-conventional types. Percent complementarity indicates the percentage of residues in a nucleic acid molecule that can form hydrogen bonds (e.g., Watson-Crick base pairing) with a second nucleic acid sequence (e.g., 5, 6, 7, 8, 9, 10 out of 10 would be 50%, 60%, 70%, 80%, 90%, and 100% complementary). "Fully complementary" means that all contiguous residues of a nucleic acid sequence will hydrogen bond with the same number of contiguous residues in a second nucleic acid sequence. "Substantially complementary," as used herein, refers to a degree of complementarity of at least 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 97%, 98%, 99%, or 100% over a region of 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 30, 35, 40, 45, 50, or more nucleotides, or that two nucleic acids hybridize under stringent conditions. Thus, in some embodiments, the first dimerization domain and the second dimerization domain are fully complementary to each other (e.g., 100%). In other examples, the first dimerization domain and the second dimerization domain are substantially complementary (eg, at least 80%) to each other.

[0075] Contacting: To be placed in direct physical association, including in solid or liquid form. Contacting can occur in vitro or ex vivo, for example, by adding a reagent to a sample (such as one containing cells), or in vivo, by administration to a subject.

[0076] Downregulated or knocked down: When used in reference to the expression of a molecule such as a target nucleic acid or protein, refers to any process that results in a reduction in the production of the target RNA or protein, but in some examples does not result in the complete elimination of the target RNA product or target RNA function. In one example, downregulation or knockdown does not result in the complete elimination of the detectable expression or activity of the target nucleic acid / protein. In some examples, downregulation or knockdown of a target nucleic acid includes processes that may reduce the translation of the target RNA and therefore the presence of the corresponding protein. The systems disclosed herein can be used to downregulate any target nucleic acid / protein of interest.

[0077] Downregulation or knockdown includes any detectable reduction in a target nucleic acid / protein. In certain examples, detectable target nucleic acid / protein in a cell or cell-free system is reduced by at least 10%, at least 20%, at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, at least 75%, at least 80%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% (e.g., 40%-90%, 40%-80%, or 50%-95% reduction) compared to a control (e.g., the amount of target nucleic acid / protein detected in a corresponding untreated cell or sample). In one example, the control is the relative amount of expression in a normal cell (e.g., a non-recombinant cell that does not contain a nucleic acid molecule for RNA recombination as provided herein).

[0078] Effective amount: An amount of an agent (such as a system providing multiple vectors, each encoding a different portion of a therapeutic protein, such as dystrophin) that is sufficient to produce a beneficial or desired result. An effective amount can also refer to the amount of correctly bound RNA or therapeutic protein produced that is sufficient to produce a beneficial or desired result.

[0079] The effective amount (also referred to as a therapeutically effective amount) may vary depending on one or more of the subject and disease state being treated, the subject's weight and age, the severity of the disease state, the mode of administration, etc., and can be determined by one skilled in the art. Beneficial therapeutic effects can include the feasibility of a diagnostic determination; amelioration of a disease, symptom, disorder, or pathological condition; reduction or prevention of the onset of a disease, symptom, disorder, or condition; and generally negating a disease, symptom, disorder, or pathological condition.

[0080] In one embodiment, an "effective amount" of two or more synthetic nucleic acid molecules provided herein is sufficient to treat a disease, such as a genetic disease or cancer. In one embodiment, an "effective amount" of two or more synthetic nucleic acid molecules provided herein is an amount sufficient to increase the survival time of a treated patient by, for example, at least 10%, at least 20%, at least 25%, at least 50%, at least 70%, at least 75%, at least 80%, at least 90%, at least 95%, at least 99%, at least 100%, at least 200%, at least 300%, at least 400%, at least 500%, or at least 600% (compared to not administering the two or more synthetic nucleic acid molecules provided herein). In one embodiment, an "effective amount" of two or more synthetic nucleic acid molecules provided herein is an amount sufficient to increase the survival time of a treated patient by, for example, at least 6 months, at least 9 months, at least 1 year, at least 1.5 years, at least 2 years, at least 2.5 years, at least 3 years, at least 4 years, at least 5 years, at least 10 years, at least 12 years, at least 15 years, or at least 20 years (compared to not administering the two or more synthetic nucleic acid molecules provided herein). In one embodiment, an "effective amount" of two or more synthetic nucleic acid molecules provided herein is an amount sufficient to increase mobility in a treated patient (such as a DMD patient) by, for example, at least 10%, at least 20%, at least 25%, at least 50%, at least 70%, at least 75%, at least 80%, at least 90%, at least 95%, at least 99%, at least 100%, at least 200%, at least 300%, at least 400%, at least 500%, or at least 600% (compared to not administering two or more synthetic nucleic acid molecules provided herein).In one embodiment, an "effective amount" of two or more synthetic nucleic acid molecules provided herein is an amount sufficient to increase mobility in a treated patient (such as a DMD patient) by, for example, at least 10%, at least 20%, at least 25%, at least 50%, at least 70%, at least 75%, at least 80%, at least 90%, at least 95%, at least 99%, at least 100%, at least 200%, at least 300%, at least 400%, at least 500%, or at least 600% (compared to not administering two or more synthetic nucleic acid molecules provided herein). In one embodiment, an "effective amount" of two or more synthetic nucleic acid molecules provided herein is an amount sufficient to increase the cognitive ability of a treated patient (such as a DMD patient) by, for example, at least 10%, at least 20%, at least 25%, at least 50%, at least 70%, at least 75%, at least 80%, at least 90%, at least 95%, at least 99%, at least 100%, at least 200%, at least 300%, at least 400%, at least 500%, or at least 600% (compared to not administering the two or more synthetic nucleic acid molecules provided herein). In one embodiment, an "effective amount" of two or more synthetic nucleic acid molecules provided herein is an amount sufficient to increase respiratory function in a treated patient (such as a DMD patient) by, for example, at least 10%, at least 20%, at least 25%, at least 50%, at least 70%, at least 75%, at least 80%, at least 90%, at least 95%, at least 99%, at least 100%, at least 200%, at least 300%, at least 400%, at least 500%, or at least 600% (compared to not administering the two or more synthetic nucleic acid molecules provided herein).In one embodiment, an "effective amount" of two or more synthetic nucleic acid molecules presented herein is an amount sufficient to increase blood clotting in a treated patient (such as a hemophilia patient) by, for example, at least 10%, at least 20%, at least 25%, at least 50%, at least 70%, at least 75%, at least 80%, at least 90%, at least 95%, at least 99%, at least 100%, at least 200%, at least 300%, at least 400%, at least 500%, or at least 600% (compared to not administering two or more synthetic nucleic acid molecules presented herein). In one embodiment, an "effective amount" of two or more synthetic nucleic acid molecules provided herein is an amount sufficient to increase the vision of a treated patient (such as a patient with Usher syndrome or Stargardt disease) by, for example, at least 10%, at least 20%, at least 25%, at least 50%, at least 70%, at least 75%, at least 80%, at least 90%, at least 95%, at least 99%, at least 100%, at least 200%, at least 300%, at least 400%, at least 500%, or at least 600% (compared to not administering the two or more synthetic nucleic acid molecules provided herein). In one embodiment, an "effective amount" of two or more synthetic nucleic acid molecules provided herein is an amount sufficient to increase the hearing of a treated patient (such as a patient with Usher syndrome) by, for example, at least 10%, at least 20%, at least 25%, at least 50%, at least 70%, at least 75%, at least 80%, at least 90%, at least 95%, at least 99%, at least 100%, at least 200%, at least 300%, at least 400%, at least 500%, or at least 600% (compared to not administering the two or more synthetic nucleic acid molecules provided herein).

[0081] In one embodiment, an "effective amount" of two or more synthetic nucleic acid molecules provided herein is an amount sufficient to reduce the size of the gastrocnemius muscle in a treated DMD patient by, for example, at least 10%, at least 20%, at least 25%, at least 50%, at least 70%, at least 75%, at least 80%, at least 90%, or at least 95% (compared to not administering the two or more synthetic nucleic acid molecules provided herein). In one embodiment, an "effective amount" of two or more synthetic nucleic acid molecules provided herein is an amount sufficient to reduce the size of the cardiomyopathic muscle in a treated DMD patient by, for example, at least 10%, at least 20%, at least 25%, at least 50%, at least 70%, at least 75%, at least 80%, at least 90%, or at least 95% (compared to not administering the two or more synthetic nucleic acid molecules provided herein). In some examples, a combination of these effects is achieved.

[0082] Increase / increase or decrease / decrease: A statistically significant positive or negative change in quantity, respectively, from a control value (a value representing the absence of a therapeutic agent, such as the absence of administration of two or more synthetic nucleic acid molecules provided herein). An increase / increase is a positive change, e.g., an increase / decrease of at least 50%, at least 100%, at least 200%, at least 300%, at least 400%, or at least 500% compared to the control value. A decrease / decrease is a negative change, e.g., a decrease / decrease of at least 20%, at least 25%, at least 50%, at least 75%, at least 80%, at least 90%, at least 95%, at least 98%, at least 99%, or at least 100% compared to the control value. In some examples, the decrease / decrease is less than 100%, e.g., a decrease / decrease of 90% or less, 95% or less, or 99% or less.

[0083] Hybridization: nucleic acid hybridization occurs when two nucleic acid molecules form a certain amount of hydrogen bonds with each other.The stringency of hybridization can vary depending on the environmental conditions surrounding nucleic acid, the nature of hybridization method, and the composition and length of nucleic acid used.The calculation of the hybridization conditions required to achieve a certain degree of stringency is discussed in Sambrook et al., Molecular Cloning: A Laboratory Manual (Cold Spring Harbor Laboratory Press, Cold Spring Harbor, NY, 2001); and Tijssen, Laboratory Techniques in Biochemistry and Molecular Biology-Hybridization with Nucleic Acid Probes Part I, Chapter 2 (Elsevier, New York, 1993). m is the temperature at which 50% of a given strand of nucleic acid hybridizes to its complementary strand.

[0084] Isolated: An "isolated" biological component (e.g., a nucleic acid molecule or protein) is one that has been substantially separated from, made separately from, or purified away from other biological components, e.g., other cells (e.g., RBCs), chromosomal and extrachromosomal DNA and RNA, and proteins, within the cells or tissues of the organism in which it resides. "Isolated" nucleic acids and proteins include nucleic acids and proteins purified by standard purification methods. The term also encompasses nucleic acids and proteins prepared by recombinant expression in a host cell as well as chemically synthesized nucleic acids and proteins.

[0085] Kissing loop / kissing stem loop: An RNA structure formed when the bases between two hairpin loops form paired interactions. These intermolecular "kissing interactions" occur when an unpaired nucleotide in one hairpin loop base pairs with an unpaired nucleotide in another hairpin loop, forming a stable interaction complex. See Figure 9A for an example.

[0086] N-terminal portion: A region of a protein sequence comprising a contiguous stretch of amino acids beginning with the N-terminal residue of the protein. The N-terminal portion of a protein can be defined by a contiguous stretch of amino acids (e.g., several amino acid residues).

[0087] Non-naturally occurring, synthetic, or engineered: Terms used interchangeably herein and indicating the involvement of the hand of man. These terms, when referring to a nucleic acid molecule or polypeptide, indicate that the nucleic acid molecule or polypeptide is at least substantially free from at least one other component with which it is naturally associated and found in nature. Additionally, these terms can indicate that the nucleic acid molecule or polypeptide has a sequence not found in nature.

[0088] Nucleic acid molecule: A deoxyribonucleotide (DNA) or ribonucleotide (RNA) polymer that may contain natural nucleotides / ribonucleotides and / or analogs of natural nucleotides / ribonucleotides that hybridize to nucleic acid molecules in the same way as naturally occurring nucleotides. A nucleic acid molecule may be a single-stranded (ss) DNA or RNA molecule or a double-stranded (ds) nucleic acid molecule. RNA or mRNA, as used herein, may refer to a pre-mRNA molecule or a mature RNA transcript. A pre-mRNA molecule contains sequences that are removed by processing, such as intron sequences that are removed by splicing after the binding of the dimerization domain described herein. A nucleic acid molecule described herein may be, for example, a DNA molecule from which RNA is transcribed by a promoter on the DNA in the case of a DNA expression vector.

[0089] Operably linked: A first nucleic acid sequence is operably linked to a second nucleic acid sequence when the first nucleic acid sequence is placed into a functional relationship with the second nucleic acid sequence. For example, a promoter sequence is operably linked to a nucleic acid sequence when the promoter affects expression of the nucleic acid sequence, e.g., when the promoter affects transcription of a pre-mRNA that, when spliced, can result in expression of a protein (such as, for example, a portion of the DMD, Factor VIII, Factor IX, or ABCA4 coding sequence).

[0090] Pharmaceutically acceptable carriers: Pharmaceutically acceptable carriers useful in the present invention are conventional. Remington's Pharmaceutical Sciences, by E. W. Martin, Mack Publishing Co., Easton, PA, 15th Edition (1975), describes compositions and formulations suitable for pharmaceutical delivery of therapeutic agents, such as the nucleic acid molecules disclosed herein.

[0091] Generally, the nature of the carrier depends on the specific administration mode used.For example, parenteral preparations usually contain an injectable fluid as a vehicle, which includes pharmaceutically and physiologically acceptable fluids such as water, physiological saline, balanced salt solution, aqueous dextrose, glycerol, etc. In addition to biologically neutral carriers, the pharmaceutical composition to be administered may contain minor amounts of non-toxic auxiliary substances, such as wetting agents or emulsifying agents, preservatives, and pH buffering agents, for example, sodium acetate or sorbitan monolaurate.

[0092] Polypeptide, Peptide, and Protein: Refers to polymers of amino acids of any length. The polymers may be linear or branched, may contain modified amino acids, and may be interrupted by non-amino acids. The terms also encompass modified amino acid polymers; for example, amino acid polymers that have undergone any other manipulation, such as disulfide bond formation, glycosylation, lipidation, acetylation, phosphorylation, or conjugation with a labeling component. As used herein, the term "amino acid" includes natural and / or unnatural or synthetic amino acids, including glycine and both the D- and L-optical isomers, as well as amino acid analogs and peptidomimetics. In one example, the protein is one that is associated with a disease, such as a genetic disease (see, e.g., Table 1). In one example, the protein is a therapeutic protein, such as one used to treat a disease, such as cancer. In one example, the protein is at least 50 amino acids in length, at least 100 amino acids in length, at least 500 amino acids in length, at least 1000 amino acids in length, at least 1500 amino acids in length, e.g., at least 2000 amino acids, at least 2500 amino acids, at least 3000 amino acids, or at least 5000 amino acids in length.

[0093] Polypyrimidine tract: A region of pre-messenger RNA (mRNA) that facilitates the assembly of the spliceosome, a specialized protein complex for RNA splicing during the process of post-transcriptional modification. This tract can be primarily pyrimidine nucleotides such as uracil, and in some examples is 15-20 base pairs in length, located approximately 5-40 base pairs before the 3' end of the intron to be spliced.

[0094] Promoter / Enhancer: A group of nucleic acid control sequences that direct the transcription of a nucleic acid sequence. A promoter includes necessary nucleic acid sequences near the start site of transcription, such as a TATA element in the case of a polymerase II type promoter. A promoter also includes distal enhancer or repressor elements, which may be located as far as several thousand base pairs from the start site of transcription, as needed. In some embodiments, the promoter sequence plus its corresponding coding sequence is larger than the capacity of an AAV. In some embodiments, the promoter sequence of a target protein is at least 3500 nt, at least 4000 nt, at least 5000 nt, or even at least 6000 nt.

[0095] A "constitutive promoter" is a promoter that is constantly active and is not regulated by external signals or molecules. In contrast, the activity of an "inducible promoter" is regulated by external signals or molecules (e.g., transcription factors). Both constitutive and inducible promoters can be used in the methods and systems provided herein (see, e.g., Bitter et al., Methods in Enzymology 153: 516-544, 1987). Tissue-specific promoters can be used in the methods and systems provided herein, for example, to direct expression primarily in a desired tissue or cell of interest, such as muscle, neurons, bone, skin, blood, a specific organ (e.g., liver, pancreas), or a specific cell type (e.g., lymphocytes). In some examples, the promoters used herein are endogenous to the target protein to be expressed. In some examples, the promoters used herein are exogenous to the target protein to be expressed.

[0096] Also included are promoter elements that are sufficient to render promoter-dependent gene expression cell-type-specific, tissue-specific, or inducible by external signals or agents, and such elements can be located in the 5' or 3' region of the gene. Promoters produced by recombinant DNA or synthetic techniques can also be used to effect transcription of nucleic acid sequences.

[0097] Exemplary promoters that can be used with the methods and systems presented herein include, but are not limited to, the SV40 promoter, the cytomegalovirus (CMV) promoter (optionally with a CMV enhancer), pol III promoters (e.g., the U6 promoter and the H1 promoter), pol II promoters (e.g., the retroviral Rous sarcoma virus (RSV) LTR promoter (optionally with an RSV enhancer), the dihydrofolate reductase promoter, the β-actin promoter, the phosphoglycerol kinase (PGK) promoter, and the EF1α promoter).

[0098] Recombinant: A recombinant nucleic acid molecule or protein sequence has a sequence that does not occur in nature or has a sequence that is created by the artificial combination of two otherwise separate sequence segments (e.g., a viral vector containing a portion of the dystrophin coding sequence, e.g., about one-third, half, or two-thirds of the coding sequence). This artificial combination can be achieved, for example, by chemical synthesis or artificial manipulation of isolated nucleic acid segments, for example, by genetic engineering techniques. Similarly, a recombinant or transgenic cell contains a recombinant nucleic acid molecule.

[0099] Sequence identity: The similarity between amino acid (or nucleotide) sequences is expressed in terms of the similarity between the sequences, otherwise referred to as sequence identity. Sequence identity is frequently measured in terms of percentage identity (or similarity or homology); the higher the percentage, the more similar the two sequences are.

[0100] Methods for aligning sequences for comparison are known. Various programs and alignment algorithms are described in Smith and Waterman, Adv. Appl. Math. 2:482, 1981; Needleman and Wunsch, J. Mol. Biol. 48:443, 1970; Pearson and Lipman, Proc. Natl. Acad. Sci. USA 85:2444, 1988; Higgins and Sharp, Gene 73:237, 1988; Higgins and Sharp, CABIOS 5:151, 1989; Corpet et al., Nucleic Acids Research 16:10881, 1988; and Pearson and Lipman, Proc. Natl. Acad. Sci. USA 85:2444, 1988. Altschul et al., Nature Genet. 6:119, 1994, presents a detailed discussion of sequence alignment methods and homology calculations.

[0101] The NCBI Basic Local Alignment Search Tool (BLAST) (Altschul et al., J. Mol. Biol. 215: 403, 1990) for use in conjunction with the sequence analysis programs blastp, blastn, blastx, tblastn, and tblastx is available from several sources, including the National Center for Biotechnology Information (NCBI, Bethesda, MD) and the Internet. A description of how to determine sequence identity using this program is available on the Internet at the NCBI website.

[0102] Variants of a native protein or coding sequence (e.g., DMD, Factor VIII, Factor IX, or ABCA4 sequences) are generally characterized by having at least about 80%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity, as counted over the entire length of the alignment with the amino acid sequence, using NCBI Blast 2.0 with gap insertion blastp set to default parameters. For comparisons of amino acid sequences of more than about 30 amino acids, the Blast 2 alignment function is used with the default BLOSUM62 matrix set to default parameters (gap existence cost of 11, and per residue gap cost of 1). When aligning short peptides (fewer than approximately 30 amino acids), alignments should be performed using the Blast 2 alignment function with the PAM30 matrix set to default parameters (open gap penalty of 9, extension gap penalty of 1). Proteins with greater similarity to the reference sequence, when assessed by this method, will exhibit increasingly greater percentage identity, e.g., at least 95%, at least 98%, or at least 99% sequence identity. When comparing less than the entire sequence for sequence identity, homologs and variants generally have at least 80% sequence identity over a short window of 10-20 amino acids, and may have at least 85%, at least 90%, or at least 95% sequence identity depending on their similarity to the reference sequence. Methods for determining sequence identity over such short windows are available on the Internet at the NCBI website. These sequence identity ranges are provided merely as guidance; highly significant homologs may be obtained outside the provided ranges.

[0103] Variants of the nucleic acid sequences disclosed herein (e.g., synthetic intron sequences and coding sequences) are generally characterized as having at least about 80%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity, counted over the entire length of the alignment with the nucleic acid sequence, using NCBI Blast 2.0, gap-insertion blastn, set to default parameters. It will be understood by those skilled in the art that these sequence identity ranges are provided merely as guidance, and that functional sequences outside the provided ranges may be obtainable.

[0104] Subject: A mammal, e.g., a human. Mammals include, but are not limited to, mice, monkeys, humans, farm animals, sport animals, and pet animals. In one embodiment, the subject is a non-human mammalian subject, e.g., a monkey or other non-human primate, mouse, rat, rabbit, pig, goat, sheep, dolphin, dog, cat, horse, or cow. In some examples, the subject is a laboratory animal / organism such as a mouse, rabbit, or rat. In some examples, the subject treated using the methods disclosed herein is a human.

[0105] In some examples, the subject has a genetic disease, such as those listed in Table 1, that can be treated using the methods disclosed herein. In some examples, the subject treated using the methods disclosed herein is a human subject with a genetic disease. In some examples, the subject treated using the methods disclosed herein is a human subject with cancer.

[0106] Therapeutic agent: refers to one or more molecules or compounds that produce some beneficial effect when administered to a subject. The synthetic nucleic acid molecules disclosed herein and the systems presented herein are therapeutic agents. Beneficial therapeutic effects can include enabling diagnostic determinations to be made; improving a disease, symptom, disorder, or pathological condition; reducing or preventing the onset of a disease, symptom, disorder, or condition; and generally negating a disease, symptom, disorder, or pathological condition.

[0107] Transduced, transformed, and transfected: A cell has been "transduced" by a virus or vector when the virus or vector has transferred the nucleic acid molecule to the cell. A cell has been "transformed" or "transfected" by a nucleic acid introduced into the cell when the nucleic acid is stably replicated by the cell, either by integration of the nucleic acid into the cell's genome or by episomal replication.

[0108] These terms encompass all techniques by which nucleic acid molecules can be introduced into such cells, including transfection with viral vectors, transformation with plasmid vectors, and introduction of naked DNA by electroporation, lipofection, particle gun acceleration, and other methods known in the art. In some examples, the methods are chemical methods (e.g., calcium phosphate transfection), physical methods (e.g., electroporation, microinjection, particle bombardment), fusion (e.g., liposomes), receptor-mediated endocytosis (e.g., DNA-protein complexes, viral envelope / capsid-DNA complexes), and biological infection with viruses such as recombinant viruses (Wolff, JA, ed., Gene Therapeutics, Birkhauser, Boston, USA, 1994). Methods for introducing nucleic acid molecules into cells are known (see, e.g., U.S. Patent No. 6,110,743). These methods can be used to transduce cells with the nucleic acid molecules disclosed herein.

[0109] Transgene: An exogenous gene delivered by a vector, e.g., AAV. In one example, a transgene encodes a portion of a target protein, e.g., about one-third, half, or two-thirds of a target protein, e.g., operably linked to a promoter sequence. In one example, a transgene includes a portion of a dystrophin coding sequence, e.g., about one-third, half, or two-thirds of a dystrophin coding sequence (or other therapeutic drug coding sequence, e.g., encoding a protein listed in Table 1), e.g., operably linked to a promoter sequence.

[0110] Treating, treatment, and therapy: Any objective or subjective parameter, such as relief, remission, or reduction of symptoms, or any success or indication of success regarding the attenuation or reversal of an injury, lesion, or condition, including making the condition more tolerable to the patient, slowing the rate of degeneration or decline, making the end point of degeneration less debilitating, improving the physical or mental well-being of the subject, or extending the length of survival. Treatment can be evaluated by objective or subjective parameters, including the results of physical exams, blood tests, and other clinical tests. In some examples, treatment using the methods disclosed herein results in a reduction in the number or severity of symptoms associated with the genetic disease, e.g., an increase in the survival time of the patient with the genetic disease being treated.

[0111] In some examples, treatment with the methods disclosed herein results in a reduction in the number or severity of symptoms associated with DMD or other genetic diseases, such as increased survival, increased mobility (e.g., walking, climbing), improved cognitive ability, reduced gastrocnemius muscle size, reduced cardiomyopathy, improved vision, improved hearing, improved blood clotting, or improved respiratory function. In some examples, a combination of these effects is achieved.

[0112] Tumor, neoplasia, malignancy, or cancer: A neoplasm is an abnormal growth of tissue or cells resulting from excessive cell division. Neoplastic growth can result in a tumor. The amount of tumor in an individual is the "tumor burden," which can be measured as the number, volume, or weight of tumors. Tumors that do not metastasize are called "benign." Tumors that have the potential to invade surrounding tissues and / or metastasize are called "malignant." "Non-cancerous tissue" is tissue that originates from the same organ in which a malignant neoplasm forms, but does not have the characteristic lesions of a neoplasm. Generally, non-cancerous tissue appears histologically normal. "Normal tissue" is tissue that originates from an organ, where the organ is not affected by cancer or another disease or disorder of that organ. A "cancer-free" subject has not been diagnosed with cancer of that organ and does not have detectable cancer.

[0113] Exemplary tumors, such as cancers, that can be treated using the methods and systems disclosed herein include solid tumors, such as breast cancer (e.g., lobular carcinoma and ductal carcinoma), sarcoma, lung cancer (e.g., non-small cell carcinoma, large cell carcinoma, squamous cell carcinoma, and adenocarcinoma), pulmonary mesothelioma, colorectal adenocarcinoma, gastric cancer, prostate adenocarcinoma, ovarian cancer (e.g., serous cystadenocarcinoma and mucinous cystadenocarcinoma), ovarian germ cell tumor, testicular cancer and germ cell tumor, pancreatic adenocarcinoma, and pancreatic adenocarcinoma. cancer, bile duct adenocarcinoma, hepatocellular carcinoma, bladder cancer (including, for example, transitional cell carcinoma, adenocarcinoma, and squamous cell carcinoma), renal cell adenocarcinoma, endometrial cancer (including, for example, adenocarcinoma and mixed Müllerian tumor (carcinosarcoma)), endocervical cancer, epicervical cancer, and vaginal cancer (e.g., adenocarcinoma and squamous cell carcinoma of the endocervix, epicervix, and vagina, respectively), tumors of the skin (e.g., squamous cell carcinoma, basal cell carcinoma, malignant melanoma, skin adnexal tumors These tumors include esophageal cancer, nasopharyngeal and oropharyngeal cancer (including squamous cell carcinoma and adenocarcinoma of the nasopharynx and oropharynx), salivary gland cancer, tumors of the brain and central nervous system (including, for example, tumors of glial origin, tumors of neuronal origin, and tumors of meningeal origin), tumors of the peripheral nerves, soft tissue sarcomas, and sarcomas of bone and cartilage, and lymphoid tumors (including B-cell and T-cell malignant lymphomas). In one embodiment, the tumor is an adenocarcinoma.

[0114] The methods and systems can also be used to treat liquid tumors, such as lymphocytic, leukemia, or other types of leukemia. In certain examples, the tumors treated are blood tumors, such as leukemias (e.g., acute lymphoblastic leukemia (ALL), chronic lymphocytic leukemia (CLL), acute myeloid leukemia (AML), chronic myeloid leukemia (CML), hairy cell leukemia (HCL), T-cell prolymphocytic leukemia (T-PLL), large granular lymphocytic leukemia, and adult T-cell leukemia), lymphomas (e.g., Hodgkin's lymphoma and non-Hodgkin's lymphoma), and myelomas.

[0115] Upregulated: When used in reference to the expression of a molecule such as a target nucleic acid / protein, refers to any process that results in increased production of the target nucleic acid / protein. In some examples, upregulation or activation of a target RNA includes processes that can increase translation of the target RNA and therefore increase the presence of the corresponding protein.

[0116] Upregulation includes any detectable increase in target nucleic acid / protein. In certain examples, detectable expression of a target nucleic acid / protein in a cell or cell-free system is increased by at least 20%, at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, at least 75%, at least 80%, at least 90%, at least 95%, at least 100%, at least 200%, at least 400%, or at least 500% compared to a control (e.g., the amount of target nucleic acid / protein detected in a corresponding sample not treated with a nucleic acid molecule provided herein). In one example, the control is the relative amount of expression in a normal cell (e.g., a non-recombinant cell not comprising a system provided herein).

[0117] Under conditions sufficient for: A phrase used to describe any environment that allows for a desired activity. In one example, the desired activity is an increase in expression or activity of a protein necessary for the treatment of a disease. In one example, the desired activity is the treatment or slowing of progression of a genetic disease, such as DMD (or other genetic diseases listed in Table 1), in vivo, using, for example, the methods and systems disclosed herein.

[0118] Vector: A nucleic acid molecule into which a foreign nucleic acid molecule can be introduced without interfering with the vector's ability to replicate in and / or integrate into a host cell. Vectors include, but are not limited to, single-stranded, double-stranded, or partially double-stranded nucleic acid molecules; nucleic acid molecules containing one or more free ends, nucleic acid molecules with no free ends (e.g., circular); nucleic acid molecules comprising DNA, RNA, or both; and other types of polynucleotides.

[0119] A vector may contain a nucleic acid sequence that allows it to replicate in a host cell, such as an origin of replication. A vector may also contain one or more selectable marker genes and other genetic elements. An integrating vector is capable of integrating itself into a host nucleic acid. An expression vector is a vector that contains the necessary regulatory sequences to allow the transcription and translation of the inserted gene(s).

[0120] One type of vector is a "plasmid," which refers to a circular double-stranded DNA loop into which additional DNA segments can be inserted, such as by standard molecular cloning techniques. Another type of vector is a viral vector, in which viral-derived DNA or RNA sequences are present within the vector for packaging into a virus (e.g., retrovirus, replication-deficient retrovirus, adenovirus, replication-deficient adenovirus, and adeno-associated virus). Viral vectors also contain polynucleotides carried by the virus for transfection into host cells. In some embodiments, the vector is a lentiviral vector (e.g., an integration-deficient lentiviral vector) or an adeno-associated viral (AAV) vector.

[0121] In some embodiments, the vector is an AAV, such as AAV serotype AAV9 or AAVrh.10. In some embodiments, the vector is capable of penetrating the blood-brain barrier, for example, after intravenous administration. Adeno-associated virus serotype rh.10 (AAV.rh10) vectors partially penetrate the blood-brain barrier, resulting in high-level and widespread transgene expression.

[0122] II. Overview of Some Embodiments One approach to cure patients suffering from genetic diseases is gene replacement therapy (commonly referred to as gene therapy). In this approach, defective genes are replaced with intact versions of the genes, delivered, for example, by viral vectors, thereby achieving sustained expression for months to years. Adeno-associated viruses (AAVs) have been used in clinical gene replacement therapy, but their packaging capacity is limited (e.g., less than about 5 kb). Therefore, strategies to overcome this packaging limitation are needed to achieve gene replacement for genes that exceed the size limit of about 5 kb. For example, some promoters alone, coding sequences alone, or a combination of promoters and coding sequences exceed the size limit of about 5 kb of AAV. Therefore, such proteins encoded by such promoters and coding sequences can be expressed using the system disclosed herein.

[0123] Previous methods for overcoming AAV cargo limitations do not appear to achieve the efficiency required to produce sufficient levels of target protein in sufficient numbers of cells for disease treatment. For example, dystrophin is approximately 11 kb and must be delivered in at least three fragments to fit the packaging limitations of AAV.

[0124] Splicing-mediated recombination of two RNA molecules using naturally occurring intron sequences in one or both of these RNA fragments is inefficient. First, these natural intron sequences are derived from naturally occurring introns and are composed of a mixture of all four RNA nucleotides. Such sequences tend to fold into structures that can interfere with trans-interactions by forming strong intramolecular base pairs rather than becoming available for intermolecular interactions. Second, because exon definition in higher eukaryotes is driven by exons rather than introns, these naturally occurring intron sequences have not evolved to strongly attract spliceosome components. These two limitations of previous strategies are addressed in the present invention by designing synthetic intron sequences not found in nature. These synthetic sequences contain elements that strongly attract and stimulate spliceosome recruitment while minimizing secondary structures (and in some instances, other structures, such as tertiary structures) that interfere with the assembly of the two RNA fragments.

[0125] The present inventors have developed novel nucleic acid-based elements that can be used to efficiently reconstruct the coding sequence of large genes from multiple sequential fragments. The methods and systems disclosed herein differ from previous methods. The highly efficient synthetic introns disclosed herein utilize optimal placement of RNA elements (or DNA encoding these elements) that efficiently drive the RNA splicing reaction between non-covalently linked RNAs (pre-mRNAs). The method / system represents a significant advance over previous attempts to utilize trans-splicing, as it produces high levels of functional protein that more closely approach therapeutic levels of protein for treating genetic diseases. This innovation is based on selecting a non-natural RNA domain that is essentially unable to form strong cis-binding interactions that interfere with trans interactions with a second RNA with a complementary strand (which also has inherently low cis-binding ability). These optimized dimerization domains and / or synthetic introns can contain non-natural sequences (e.g., sequences not found in human cells and / or in other biological systems) used in combination with optimized motifs (including splice donors, splice acceptors, splice enhancers, and splice branch point sequences) that facilitate RNA splicing. Synthetic nucleic acids can be non-natural nucleic acid sequences, such as sequences not found in human cells and / or in other biological systems. The present invention demonstrates for the first time that by optimizing trans-dimerization of RNA strands with appropriate RNA motifs that mediate efficient splicing, two or three different RNAs can be precisely and efficiently covalently linked in the same cell in vivo and in vitro to produce high levels of functional protein.Unlike "hybrid" approaches, where DNA recombination ultimately results in inefficient combination at the DNA level, followed by RNA splicing in cis, which removes the DNA recombination site from the mature transcript, in the methods / systems disclosed herein, two protein-coding RNA fragments join at the pre-mRNA level, facilitating a more efficient reaction with a lower risk of producing a recombinant product that encodes a non-functional and / or harmful product.

[0126] The data demonstrate that by using an efficient synthetic RNA dimerization and recombination domain (sRdR domain, also referred to as RNA end joining (REJ) domain), a gene of interest can be efficiently reconstituted from two or three separate gene fragments expressed in the same cell. These results demonstrate the ability of the methods and systems disclosed herein to reconstitute large genes such as dystrophin, blood coagulation factor VIII, or ATP-binding cassette subfamily A member 4 (Abca4) using AAV to treat Duchenne muscular dystrophy and hemophilia A, or Stargardt's disease, respectively. Based on these findings, other genetic diseases, such as those that benefit from the expression of large proteins (see, for example, the disorders listed in Table 1), can be similarly treated. Other applications include research and biotechnology applications.

[0127] To address some of the limitations of existing strategies for reconstructing fragmented genes from multiple AAVs, a system is provided herein for sequentially aligning and recombining two or more individual synthetic RNA molecules in a target cell. Each of the individual synthetic RNA molecules contains a dimerization domain and a synthetic intron sequence containing elements necessary for RNA splicing, thereby mediating efficient RNA recombination of the individual fragments when the dimerization domains bind to each other in the correct order. In one example, reconstitution of the coding sequence from the two fragments is achieved by adding a first synthetic intron (A) to the 3' end of the N-terminal coding fragment and a complementary second synthetic domain (A') to the 5' end of the C-terminal coding fragment. The two RNAs are recombined by the cell's endogenous RNA splicing machinery (i.e., the spliceosome machinery). The synthetic intron domain contains bifunctional elements: (1) a dimerization domain to mediate base pairing between the two halves to be recombined, and (2) a domain optimized to efficiently recruit the splicing machinery to mediate efficient reconstitution of the two RNA molecules. In some examples, the synthetic intron comprises a sequence having at least 50%, at least 60%, at least 70%, at least 75%, 80%, at least 85%, at least 90%, at least 95%, at least 98%, at least 99%, or 100% sequence identity to any of the synthetic introns set forth in SEQ ID NOs: 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 145, 146, 147, 148, 155, 156, 157, 158, 159, 160, 161, 162, 163, 164, 165, and 166 (see, e.g., Figures 10A-10Z).In some examples, a synthetic intron is an RNA molecule encoded by a sequence having at least 50%, at least 60%, at least 70%, at least 75%, 80%, at least 85%, at least 90%, at least 95%, at least 98%, at least 99%, or 100% sequence identity to any of the synthetic introns set forth in SEQ ID NOs: 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 145, 146, 147, 148, 155, 156, 157, 158, 159, 160, 161, 162, 163, 164, 165, and 166, but without the promoter sequence set forth. Those skilled in the art will understand that any of the molecules set forth in SEQ ID NOs: 1, 2, 20, 21, 22, 23, 24, 25, 145, 146, 147, 148, 155, 156, 157, 158, 159, 160, 161, 162, 163, 164, 165, and 166 can be modified to replace the protein-encoding portions (e.g., 114 and 164 in Figure 6A) with another protein-encoding sequence of interest (e.g., the YFP-encoding sequence of SEQ ID NO: 1, 2, 22, or 23 can be replaced with a therapeutic protein-encoding sequence). Thus, also provided herein are synthetic intron molecules having at least 80%, at least 85%, at least 90%, at least 95%, at least 98%, at least 99%, or 100% sequence identity to any of the synthetic intron portions provided in SEQ ID NOs: 1, 2, 20, 21, 22, 23, 24, 25, 145, 146, 147, 148, 155, 156, 157, 158, 159, 160, 161, 162, 163, 164, 165, and 166 (e.g., nt 3703-3975 of SEQ ID NO: 22 and nt 1-225 of SEQ ID NO: 23).Also provided are synthetic intron RNA molecules encoded by sequences having at least 50%, at least 60%, at least 70%, at least 75%, 80%, at least 85%, at least 90%, at least 95%, at least 98%, at least 99%, or 100% sequence identity to any of the synthetic introns set forth in SEQ ID NOs: 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 145, 146, 147, 148, 155, 156, 157, 158, 159, 160, 161, 162, 163, 164, 165, and 166, but without the promoter sequences set forth).

[0128] Exemplary dimerization domains were selected bioinformatically to minimize / optimize their internal secondary / tertiary structures. The dimerization domains tested contained long stretches of low-diversity nucleotide sequences to avoid intramolecular annealing. By avoiding intramolecular annealing, these dimerization domains exist in an open configuration, and are therefore available for pairing with corresponding complementary dimerization domain sequences. The synthetic intron domain contains an intron splice enhancer element, which leads to efficient recruitment of the splicing machinery.

[0129] The RNA molecules disclosed herein are designed to have at least an open, accessible single-stranded region that can bind to complementary dimerization domains to allow efficient splicing and recombination of the RNA. In some embodiments, this is achieved by using only purines or only pyrimidines in the binding domain. Due to the inability of purines (and pyrimidines) to pair with themselves, these stretches of RNA have an open, predicted structure.

[0130] RNA molecules exist as single strands in cells. Being single-stranded, RNA molecules inherently tend to hybridize with themselves, thereby forming strong secondary and tertiary structures. The most stable base pairs are G-C, A-U, and G-U wobble pairs. Thermodynamically, two-base pairing is favored over open configurations. To design efficient synthetic nucleic acid molecules, two complementary dimerization domains are present in an open configuration, thus making the dimerization domains available for intermolecular base pairing. To avoid intramolecular base pairing between other parts of the synthetic nucleic acid molecule, long stretches of non-variable sequences containing incompatible bases can be included. For example, long stretches of pyrimidines (i.e., C and T) or purines (i.e., A and G) can be present in the synthetic nucleic acid molecule. Pyrimidines cannot form standard base pairs with other pyrimidines, and purines cannot form standard base pairs with other purines. Such stretches of purines or pyrimidines can range from 2-3 bases to 200-300 bases. These stretches cannot bind intramolecularly and are therefore available for intermolecular base pairing with complementary fragments. For example, synthetic nucleic acid molecules A and A' can be constructed such that A contains a pyrimidine stretch (e.g., 5'-CCUU(...)CCUU-3') and A' contains a complementary purine sequence (e.g., 5'-AAGG(...)AAGG-3').

[0131] The synthetic nucleic acid molecule disclosed herein (for example, RNA or DNA encoding RNA) is designed to minimize any off-target binding to the incorrect site in genome.Off-target binding can be reduced by changing the sequence of nucleic acid molecule.

[0132] The same design principle of using stretches of low diversity RNA bases to achieve open synthetic nucleic acid architectures can be extended to the use of stretches of single bases in the dimerization domain, for example, a stretch of Gs base-pairing with a stretch of Cs, and a stretch of As base-pairing with a stretch of Us.

[0133] The following methods can be used to increase recombination of two or more synthetic nucleic acid molecules. RNA splicing relies on the recruitment of spliceosome components to the 5' end of the intron (splice donor site) and the 3' end of the intron (splice acceptor site, along with its associated branchpoint sequence and polypyrimidine tract). Different ribonucleoproteins are recruited to introns through base pairing of protein-associated small nuclear RNAs (snRNAs) with intron sequences. Placing perfectly matched consensus sequences within the RNA dimerization and recombination domain can facilitate the recruitment of spliceosome components, which in turn enhances the efficiency of spliceosome-mediated recombination. Previously characterized intron splice enhancer sequences can recruit additional splicing-enhancing factors, termed intron splice enhancers.

[0134] In some embodiments, instead of using naturally occurring RNA sequences for RNA splicing sequences, consensus sequences are used. For example, consensus sequences can be used for any sequence involved in splicing, including splice donors, splice acceptors, splice enhancers, and splice branchpoint sequences. These synthetic nucleic acid molecules can be used to sequentially join two (or more) RNA molecules ex vivo, in vitro, or in vivo in cells. The synthetic nucleic acid molecule can include any promoter and coding sequence outside of the encoded synthetic intron domain. For example, two synthetic nucleic acid molecules can carry both halves of a single gene. This was tested in vitro and in vivo by reconstituting both halves of yellow fluorescent protein (YFP) and was shown to be efficient (see Figures 3A-3D).

[0135] The modular nature of synthetic nucleic acid molecules allowed us to test the efficiency of achieving serial recombination of multiple RNA fragments (i.e., >2) using a combinatorial set of optimized complementary dimerization domains (Figure 4A-4B). A tripartite yellow fluorescent protein was efficiently reconstituted and expressed at high levels in >80% of transfected cells.

[0136] These results demonstrate that a single RNA molecule can be reconstituted from at least three different synthetic nucleic acid molecules, such as when expressing disease-causing genes (or therapeutic proteins) with promoters and / or coding sequences that are too long to fit into a single gene therapy vector, e.g., AAV.

[0137] In some embodiments, the synthetic nucleic acid molecules, eg, synthetic DNA molecules, of the compositions, systems, kits, and methods of the invention are generated by transcription of an RNA viral genome by reverse transcriptase.

[0138] The systems disclosed herein allow for efficient RNA recombination between the individual fragments. In some examples, the efficiency of reconstitution (i.e., splicing or recombination) achieved using the compositions, systems, or methods of the present disclosure is determined using any suitable method known to those of skill in the art. In some examples, the efficiency of reconstitution is represented by a measure of correctly ligated RNA compared to a control RNA, or a measure of full-length protein or protein activity compared to a control protein. In some examples, the control RNA is unligated RNA, in which case the efficiency of reconstitution is represented by a measure of bound RNA compared to unligated RNA. This measurement can be made by detecting and comparing the junction RNA and the unligated 3' RNA species 3' (e.g., junction RNA:3' RNA). In some examples, in which more than two RNAs are ligated, binding at any or all of the junctions is assessed. In some examples, the efficiency of reconstitution is represented by a measure of full-length protein or active protein compared to a protein fragment or inactive protein.

[0139] In some embodiments, the efficiency of rearrangement, recombination, or splicing (a measure of the correct joining of two or more different coding sequences present on different RNA molecules and / or the production of a desired full-length protein) is between about 10% and about 100%. In some embodiments, the rearrangement efficiency is between about 10% and about 15%, between about 10% and about 20%, between about 10% and about 25%, between about 10% and about 30%, between about 10% and about 40%, between about 10% and about 50%, between about 10% and about 60%, between about 10% and about 70%, between about 10% and about 80%, between about 10% and about 90%, between about 10% and about 100%, between about 15% and about 20%, between about 15% and about 25%, between about 15% and about 30%, between about 15% and about 40%, between about 15% and about 50%, or between about 15% and about 50%. 0%, about 15% to about 60%, about 15% to about 70%, about 15% to about 80%, about 15% to about 90%, about 15% to about 100%, about 20% to about 25%, about 20% to about 30%, about 20% to about 40%, about 20% to about 50%, about 20% to about 60%, about 20% to about 70%, about 20% to about 80%, about 20% to about 90%, about 20% to about 100%, about 25% to about 30%, about 25% to about 40%, about 25% to about 50 %, about 25% to about 60%, about 25% to about 70%, about 25% to about 80%, about 25% to about 90%, about 25% to about 100%, about 30% to about 40%, about 30% to about 50%, about 30% to about 60%, about 30% to about 70%, about 30% to about 80%, about 30% to about 90%, about 30% to about 100%, about 40% to about 50%, about 40% to about 60%, about 40% to about 70%, about 40% to about 80%, about 40% to about 90% , about 40% to about 100%, about 50% to about 60%, about 50% to about 70%, about 50% to about 80%, about 50% to about 90%, about 50% to about 100%, about 60% to about 70%, about 60% to about 80%, about 60% to about 90%, about 60% to about 100%, about 70% to about 80%, about 70% to about 90%, about 70% to about 100%, about 80% to about 90%, about 80% to about 100%, or about 90% to about 100%. In some embodiments, the reconstitution efficiency is about 10%, about 15%, about 20%, about 25%, about 30%, about 40%, about 50%, about 60%, about 70%, about 80%, about 90%, or about 100%. In some embodiments, the reconstitution efficiency is at least about 10%, about 15%, about 20%, about 25%, about 30%, about 40%, about 50%, about 60%, about 70%, about 80%, or about 90%.In some embodiments, the reconstitution efficiency is at most about 15%, about 20%, about 25%, about 30%, about 40%, about 50%, about 60%, about 70%, about 80%, about 90%, or about 100%.

[0140] In some examples, the rearrangement, recombination, or splicing efficiency (in this example, a measure of the correct joining of two different coding sequences present on different RNA molecules and / or the production of a desired full-length protein, where the two different coding sequences encode transcripts of about 3200 nt to 9000 nt, e.g., about 4000 to 9000 nt, about 4400 to 9000 nt, about 3200 to 4000 nt, about 3200 to 3600 nt, e.g., about 4500 nt, about 4000 nt, about 3800 nt, about 3600 nt, or about 3200 nt) is about 10% to about 100%. In some embodiments, the reconstitution efficiency using the two-part system is between about 10% and about 15%, between about 10% and about 20%, between about 10% and about 25%, between about 10% and about 30%, between about 10% and about 40%, between about 10% and about 50%, between about 10% and about 60%, between about 10% and about 70%, between about 10% and about 80%, between about 10% and about 90%, between about 10% and about 100%, between about 15% and about 20%, between about 15% and about 25%, between about 15% and about 30%, between about 15 ... About 40%, about 15% to about 50%, about 15% to about 60%, about 15% to about 70%, about 15% to about 80%, about 15% to about 90%, about 15% to about 100%, about 20% to about 25%, about 20% to about 30%, about 20% to about 40%, about 20% to about 50%, about 20% to about 60%, about 20% to about 70%, about 20% to about 80%, about 20% to about 90%, about 20% to about 100%, about 25% to about 30%, about 25% to about 40%, Approximately 25% to approximately 50%, approximately 25% to approximately 60%, approximately 25% to approximately 70%, approximately 25% to approximately 80%, approximately 25% to approximately 90%, approximately 25% to approximately 100%, approximately 30% to approximately 40%, approximately 30% to approximately 50%, approximately 30% to approximately 60%, approximately 30% to approximately 70%, approximately 30% to approximately 80%, approximately 30% to approximately 90%, approximately 30% to approximately 100%, approximately 40% to approximately 50%, approximately 40% to approximately 60%, approximately 40% to approximately 70%, approximately 40% to approximately 80%, approximately 40% to about 90%, about 40% to about 100%, about 50% to about 60%, about 50% to about 70%, about 50% to about 80%, about 50% to about 90%, about 50% to about 100%, about 60% to about 70%, about 60% to about 80%, about 60% to about 90%, about 60% to about 100%, about 70% to about 80%, about 70% to about 90%, about 70% to about 100%, about 80% to about 90%, about 80% to about 100%, or about 90% to about 100%.In some embodiments, the reconstitution efficiency is about 10%, about 15%, about 20%, about 25%, about 30%, about 40%, about 50%, about 60%, about 70%, about 80%, about 90%, or about 100%. In some embodiments, the reconstitution efficiency is at least about 10%, about 15%, about 20%, about 25%, about 30%, about 40%, about 50%, about 60%, about 70%, about 80%, or about 90%. In some embodiments, the reconstitution efficiency is at most about 15%, about 20%, about 25%, about 30%, about 40%, about 50%, about 60%, about 70%, about 80%, about 90%, or about 100%.

[0141] In some examples, the efficiency of rearrangement, recombination, or splicing (in this example, a measure of the correct joining of two different coding sequences present on different RNA molecules and / or the production of a desired full-length protein, where the two different coding sequences encode transcripts of approximately 4000 nt) is between about 40% and about 60%, e.g., between about 40% and about 50%, between about 42% and about 47%, e.g., about 45%.

[0142] In some examples, the efficiency of rearrangement, recombination, or splicing (in this example, a measure of the correct joining of two different coding sequences present on different RNA molecules and / or the production of a desired full-length protein, where the two different coding sequences encode a transcript of approximately 3800 nt) is between about 40% and about 60%, e.g., between about 40% and about 50%, between about 42% and about 47%, e.g., about 45%.

[0143] In some examples, the efficiency of rearrangement, recombination, or splicing (in this example, a measure of the correct joining of two different coding sequences present on different RNA molecules and / or the production of a desired full-length protein, where the two different coding sequences encode a transcript of approximately 3600 nt) is about 25% to about 50%, e.g., about 30% to about 40%, e.g., about 35%.

[0144] In some examples, the efficiency of rearrangement, recombination, or splicing (in this example, a measure of the correct joining of two different coding sequences present on different RNA molecules and / or the production of a desired full-length protein, where the two different coding sequences encode a transcript of approximately 3200 nt) is about 25% to about 50%, e.g., about 30% to about 40%, e.g., about 35%.

[0145] In some examples, the rearrangement, recombination, or splicing efficiency (in this example, a measure of the correct joining of three different coding sequences present on different RNA molecules and / or the production of a desired full-length protein, where the three different coding sequences encode transcripts of about 3,200 nt to about 13,500 nt, e.g., about 4,000 nt to about 5,000 nt, about 4,000 nt to about 13,500 nt, about 6,000 nt to about 12,000 nt, about 6,000 nt to about 10,000 nt, or about 8,000 nt to about 12,000 nt, e.g., up to about 13,500 nt) is about 10% to about 100%. In some embodiments, the reconstitution efficiency using the three-part system is between about 10% and about 15%, between about 10% and about 20%, between about 10% and about 25%, between about 10% and about 30%, between about 10% and about 40%, between about 10% and about 50%, between about 10% and about 60%, between about 10% and about 70%, between about 10% and about 80%, between about 10% and about 90%, between about 10% and about 100%, between about 15% and about 20%, between about 15% and about 25%, between about 15% and about 30%, between about 15 ... About 40%, about 15% to about 50%, about 15% to about 60%, about 15% to about 70%, about 15% to about 80%, about 15% to about 90%, about 15% to about 100%, about 20% to about 25%, about 20% to about 30%, about 20% to about 40%, about 20% to about 50%, about 20% to about 60%, about 20% to about 70%, about 20% to about 80%, about 20% to about 90%, about 20% to about 100%, about 25% to about 30%, about 25% to about 40%, Approximately 25% to approximately 50%, approximately 25% to approximately 60%, approximately 25% to approximately 70%, approximately 25% to approximately 80%, approximately 25% to approximately 90%, approximately 25% to approximately 100%, approximately 30% to approximately 40%, approximately 30% to approximately 50%, approximately 30% to approximately 60%, approximately 30% to approximately 70%, approximately 30% to approximately 80%, approximately 30% to approximately 90%, approximately 30% to approximately 100%, approximately 40% to approximately 50%, approximately 40% to approximately 60%, approximately 40% to approximately 70%, approximately 40% to approximately 80%, approximately 40% to about 90%, about 40% to about 100%, about 50% to about 60%, about 50% to about 70%, about 50% to about 80%, about 50% to about 90%, about 50% to about 100%, about 60% to about 70%, about 60% to about 80%, about 60% to about 90%, about 60% to about 100%, about 70% to about 80%, about 70% to about 90%, about 70% to about 100%, about 80% to about 90%, about 80% to about 100%, or about 90% to about 100%.In some embodiments, the reconstitution efficiency is about 10%, about 15%, about 20%, about 25%, about 30%, about 40%, about 50%, about 60%, about 70%, about 80%, about 90%, or about 100%. In some embodiments, the reconstitution efficiency is at least about 10%, about 15%, about 20%, about 25%, about 30%, about 40%, about 50%, about 60%, about 70%, about 80%, or about 90%. In some embodiments, the reconstitution efficiency is at most about 15%, about 20%, about 25%, about 30%, about 40%, about 50%, about 60%, about 70%, about 80%, about 90%, or about 100%.

[0146] In some examples, the rearrangement, recombination, or splicing efficiency (in this example, a measure of the correct joining of four different coding sequences present on different RNA molecules and / or the production of a desired full-length protein, where the four different coding sequences encode transcripts of about 3,200 nt to about 18,000 nt, e.g., about 4,000 nt to about 18,000 nt, about 4,000 nt to about 5,000 nt, about 10,000 nt to about 18,000 nt, about 15,000 nt to about 18,000 nt, or about 12,000 nt to about 15,000 nt, e.g., up to about 18,000 nt) is about 10% to about 100%. In some embodiments, the reconstitution efficiency using the four-part system is between about 10% and about 15%, between about 10% and about 20%, between about 10% and about 25%, between about 10% and about 30%, between about 10% and about 40%, between about 10% and about 50%, between about 10% and about 60%, between about 10% and about 70%, between about 10% and about 80%, between about 10% and about 90%, between about 10% and about 100%, between about 15% and about 20%, between about 15% and about 25%, between about 15% and about 30%, between about 15 ... About 40%, about 15% to about 50%, about 15% to about 60%, about 15% to about 70%, about 15% to about 80%, about 15% to about 90%, about 15% to about 100%, about 20% to about 25%, about 20% to about 30%, about 20% to about 40%, about 20% to about 50%, about 20% to about 60%, about 20% to about 70%, about 20% to about 80%, about 20% to about 90%, about 20% to about 100%, about 25% to about 30%, about 25% to about 40%, Approximately 25% to approximately 50%, approximately 25% to approximately 60%, approximately 25% to approximately 70%, approximately 25% to approximately 80%, approximately 25% to approximately 90%, approximately 25% to approximately 100%, approximately 30% to approximately 40%, approximately 30% to approximately 50%, approximately 30% to approximately 60%, approximately 30% to approximately 70%, approximately 30% to approximately 80%, approximately 30% to approximately 90%, approximately 30% to approximately 100%, approximately 40% to approximately 50%, approximately 40% to approximately 60%, approximately 40% to approximately 70%, approximately 40% to approximately 80%, approximately 40% to about 90%, about 40% to about 100%, about 50% to about 60%, about 50% to about 70%, about 50% to about 80%, about 50% to about 90%, about 50% to about 100%, about 60% to about 70%, about 60% to about 80%, about 60% to about 90%, about 60% to about 100%, about 70% to about 80%, about 70% to about 90%, about 70% to about 100%, about 80% to about 90%, about 80% to about 100%, or about 90% to about 100%.In some examples, the reconstitution efficiency is about 10%, about 15%, about 20%, about 25%, about 30%, about 40%, about 50%, about 60%, about 70%, about 80%, about 90%, or about 100%. In some examples, the reconstitution efficiency is at least about 10%, about 15%, about 20%, about 25%, about 30%, about 40%, about 50%, about 60%, about 70%, about 80%, or about 90%. In some examples, the reconstitution efficiency is at most about 15%, about 20%, about 25%, about 30%, about 40%, about 50%, about 60%, about 70%, about 80%, about 90%, or about 100%. In some examples, the compositions, systems, or methods of the present disclosure are evaluated by determining the level of RNA or protein production using any suitable method known to one of skill in the art. In some embodiments, the RNA production level is expressed by the measure of correctly bound RNA compared to a control RNA, or the measure of full-length protein compared to a control. In some embodiments, the control RNA is a corresponding mutant RNA or endogenous RNA. For example, the ratio of the amount of bound RNA produced in transfected cells to the amount of mutant or endogenous RNA is compared with the same ratio in untransfected cells to determine the production level of correctly bound RNA. In some embodiments, the ratio of the amount of correctly bound RNA, full-length protein, or protein activity to the amount of control RNA, or the amount or activity of a control protein is compared.

[0147] In some embodiments, the achieved RNA production level is 5% to 100%. In some embodiments, the achieved RNA production level is about 5% to about 100%. In some embodiments, the achieved RNA production level is about 5% to about 10%, about 5% to about 20%, about 5% to about 25%, about 5% to about 30%, about 5% to about 40%, about 5% to about 50%, about 5% to about 60%, about 5% to about 70%, about 5% to about 80%, about 5% to about 90%, about 5% to about 100%, about 10% to about 20%, about 10% to about 25%, about 10% to about 30%, about 10% to about 40%, or about 10% to about 50%. , about 10% to about 60%, about 10% to about 70%, about 10% to about 80%, about 10% to about 90%, about 10% to about 100%, about 20% to about 25%, about 20% to about 30%, about 20% to about 40%, about 20% to about 50%, about 20% to about 60%, about 20% to about 70%, about 20% to about 80%, about 20% to about 90%, about 20% to about 100%, about 25% to about 30%, about 25% to about 40%, about 25% to about 50% , about 25% to about 60%, about 25% to about 70%, about 25% to about 80%, about 25% to about 90%, about 25% to about 100%, about 30% to about 40%, about 30% to about 50%, about 30% to about 60%, about 30% to about 70%, about 30% to about 80%, about 30% to about 90%, about 30% to about 100%, about 40% to about 50%, about 40% to about 60%, about 40% to about 70%, about 40% to about 80%, about 40% to about 90% , about 40% to about 100%, about 50% to about 60%, about 50% to about 70%, about 50% to about 80%, about 50% to about 90%, about 50% to about 100%, about 60% to about 70%, about 60% to about 80%, about 60% to about 90%, about 60% to about 100%, about 70% to about 80%, about 70% to about 90%, about 70% to about 100%, about 80% to about 90%, about 80% to about 100%, or about 90% to about 100%. In some examples, the achieved RNA production level is about 5%, about 10%, about 20%, about 25%, about 30%, about 40%, about 50%, about 60%, about 70%, about 80%, about 90%, or about 100%. In some examples, the level of RNA production achieved is at least about 5%, about 10%, about 20%, about 25%, about 30%, about 40%, about 50%, about 60%, about 70%, about 80%, or about 90%.In some examples, the level of RNA production achieved is at most about 10%, about 20%, about 25%, about 30%, about 40%, about 50%, about 60%, about 70%, about 80%, about 90%, or about 100%.

[0148] In some examples, the protein production level is expressed by a measure of the amount of full-length protein or protein activity compared to the amount of full-length protein or protein activity of a control protein. In some examples, the control protein is a corresponding mutant protein or endogenous protein. For example, the ratio of the amount of full-length protein or protein activity produced in transfected cells to the amount of mutant or endogenous protein is compared to the same ratio in untransfected cells. In some examples, the control protein is, for example, a full-length protein produced in cells engineered to express the control full-length protein (where the cells are not transfected with a construct of the invention) or in untransfected cells from a normal subject expressing the control full-length protein, and the protein production level is determined by measuring the amount or activity of the protein in the transfected cells and comparing it to that of the control protein. In some examples, the control protein is a mutant form of the protein produced in cells transfected with the construct or in untransfected cells, and the protein production level is determined by comparing the amount of full-length protein or protein activity to that of the control protein. In some examples, the protein production level is determined by comparing the amount of full-length protein or protein activity to that of an endogenous protein or housekeeping protein.

[0149] In some embodiments, the protein production level achieved is about 1% to about 100%. In some embodiments, the protein production level achieved is about 10% to about 100%. In some embodiments, the protein production level achieved is about 10% to about 20%, about 10% to about 30%, about 10% to about 40%, about 10% to about 50%, about 10% to about 60%, about 10% to about 70%, about 10% to about 75%, about 10% to about 80%, about 10% to about 85%, about 10% to about 90%, about 10% to about 100%, about 20% to about 30%, about 20% to about 40%, about 20% to about 50%, about 20% to about 6 ... %, about 20% to about 70%, about 20% to about 75%, about 20% to about 80%, about 20% to about 85%, about 20% to about 90%, about 20% to about 100%, about 30% to about 40%, about 30% to about 50%, about 30% to about 60%, about 30% to about 70%, about 30% to about 75%, about 30% to about 80%, about 30% to about 85%, about 30% to about 90%, about 30% to about 100%, about 40% to about 50%, about 40% to about 60%, about 4 0% to about 70%, about 40% to about 75%, about 40% to about 80%, about 40% to about 85%, about 40% to about 90%, about 40% to about 100%, about 50% to about 60%, about 50% to about 70%, about 50% to about 75%, about 50% to about 80%, about 50% to about 85%, about 50% to about 90%, about 50% to about 100%, about 60% to about 70%, about 60% to about 75%, about 60% to about 80%, about 60% to about 85%, about 60% to about 90%, about 60% to about 100%, about 70% to about 75%, about 70% to about 80%, about 70% to about 85%, about 70% to about 90%, about 70% to about 100%, about 75% to about 80%, about 75% to about 85%, about 75% to about 90%, about 75% to about 100%, about 80% to about 85%, about 80% to about 90%, about 80% to about 100%, about 85% to about 90%, about 85% to about 100%, or about 90% to about 100%. In some embodiments, the protein production level achieved is about 10%, about 20%, about 30%, about 40%, about 50%, about 60%, about 70%, about 75%, about 80%, about 85%, about 90%, or about 100%. In some examples, the protein production level achieved is at least about 10%, about 20%, about 30%, about 40%, about 50%, about 60%, about 70%, about 75%, about 80%, about 85%, or about 90%.In some examples, the level of protein production achieved is at most about 20%, about 30%, about 40%, about 50%, about 60%, about 70%, about 75%, about 80%, about 85%, about 90%, or about 100%.

[0150] In some embodiments, the protein activity level achieved is about 50% to about 100%. In some embodiments, the protein activity level achieved is about 50% to about 100%. In some embodiments, the protein activity level achieved is about 50% to about 55%, about 50% to about 60%, about 50% to about 65%, about 50% to about 70%, about 50% to about 75%, about 50% to about 80%, about 50% to about 85%, about 50% to about 90%, about 50% to about 95%, about 50% to about 100%, about 55% to about 60%, about 55% to about 65%, about 55% to about 70%, about 55% to about 75%, about 55% to about 80%, about 55% to about 85%, about 55% to about 90%, about 55% to about 95%, about 55% to about 100%, about 60% to about 65%, about 60% to about 70%, about 60% to about 75%, about 60% to about 80%, about 60% to about 85%, about 60% to about 90%, about 60% to about 95%, about 60% to about 10 0%, approximately 65% ​​to approximately 70%, approximately 65% ​​to approximately 75%, approximately 65% ​​to approximately 80%, approximately 65% ​​to approximately 85%, approximately 65% ​​to approximately 90%, approximately 65% ​​to approximately 95%, approximately 65% ​​to approximately 100%, approximately 70% to approximately 75%, approximately 70% to approximately 80%, approximately 70% to approximately 85%, approximately 70% to approximately 90%, approximately 70% to approximately 95%, approximately 70% to approximately 100%, approximately 75% to approximately 80%, approximately 75 In some embodiments, the protein activity level achieved is about 50%, about 55%, about 60%, about 65%, about 70%, about 75%, about 80%, about 85%, about 90%, about 95%, about 95%, or about 100%. In some embodiments, the level of protein activity achieved is at least about 50%, about 55%, about 60%, about 65%, about 70%, about 75%, about 80%, about 85%, about 90%, or about 95%. In some embodiments, the level of protein activity achieved is at most about 55%, about 60%, about 65%, about 70%, about 75%, about 80%, about 85%, about 90%, about 95%, or about 100%.

[0151] In some embodiments, the amount of correctly linked RNA or full-length protein produced in the cell is sufficient to ameliorate or cure a condition or disease in a subject, as understood by one of skill in the art for a particular condition or disease. In some embodiments, the amount of correctly linked RNA or full-length protein produced in the cell is an effective amount. In some embodiments, this amount corresponds to about 50% to 100% of the amount of RNA or protein produced in a normal cell. In some embodiments, this amount corresponds to about 40% to 100% of the amount of RNA or protein produced in a normal cell. In some embodiments, this amount is about 40% to about 45%, about 40% to about 50%, about 40% to about 55%, about 40% to about 60%, about 40% to about 65%, about 40% to about 70%, about 40% to about 75%, about 40% to about 80%, about 40% to about 85%, about 40% to about 90%, about 40% to about 100%, about 45% to about 50%, about 45% to about 55%, about 45 ... 0%, approximately 45% to approximately 65%, approximately 45% to approximately 70%, approximately 45% to approximately 75%, approximately 45% to approximately 80%, approximately 45% to approximately 85%, approximately 45% to approximately 90%, approximately 45% to approximately 100%, approximately 50% to approximately 55%, approximately 50% to approximately 60%, approximately 50% to approximately 65%, approximately 50% to approximately 70%, approximately 50% to approximately 75%, approximately 50% to approximately 80%, approximately 50% to approximately 85%, approximately 50% to approximately 90%, approximately 50% to approximately 100%, approximately 55% to approximately 60%, approximately 55% to About 65%, about 55% to about 70%, about 55% to about 75%, about 55% to about 80%, about 55% to about 85%, about 55% to about 90%, about 55% to about 100%, about 60% to about 65%, about 60% to about 70%, about 60% to about 75%, about 60% to about 80%, about 60% to about 85%, about 60% to about 90%, about 60% to about 100%, about 65% to about 70%, about 65% to about 75%, about 65% to about 80%, about 65% to about 85%, about 65 % to about 90%, about 65% to about 100%, about 70% to about 75%, about 70% to about 80%, about 70% to about 85%, about 70% to about 90%, about 70% to about 100%, about 75% to about 80%, about 75% to about 85%, about 75% to about 90%, about 75% to about 100%, about 80% to about 85%, about 80% to about 90%, about 80% to about 100%, about 85% to about 90%, about 85% to about 100%, or about 90% to about 100%.In some examples, this amount corresponds to about 40%, about 45%, about 50%, about 55%, about 60%, about 65%, about 70%, about 75%, about 80%, about 85%, about 90%, or about 100% of the amount of RNA or protein produced in a normal cell. In some examples, this amount corresponds to about at least about 40%, about 45%, about 50%, about 55%, about 60%, about 65%, about 70%, about 75%, about 80%, about 85%, or about 90% of the amount of RNA or protein produced in a normal cell. In some examples, this amount corresponds to about at most about 45%, about 50%, about 55%, about 60%, about 65%, about 70%, about 75%, about 80%, about 85%, about 90%, or about 100% of the amount of RNA or protein produced in a normal cell.

[0152] The RNA or protein measurements used to determine recombination efficiency or production level can be performed by any suitable method known to those skilled in the art. In some examples, recombination efficiency or production level is determined by measuring the amount of expressed functional protein, for example, by Western blotting. In some examples, recombination efficiency or production level is determined by measuring RNA transcripts, for example, using quantitative real-time PCR based on two probes. For example, the first assay spans the sequence completely contained in the 3' exon coding sequence (labeled with the 3' probe). The second assay spans the junction between the 5' exon coding sequence and the 3' exon coding sequence (labeled with the junction probe). Recombination efficiency can be calculated as the ratio of (junction probe count) / (3' probe count). The terms "recombination efficiency," "recombination efficiency," and "splicing efficiency" are used interchangeably herein.

[0153] In some embodiments, the dimerization domain is about 20 to about 1000 nt, or about 50 to about 160 nt, or about 50 to about 500 nt, or about 50 to 1000 nt, wherein the reconstitution efficiency is such that an effective amount of correctly linked RNA or full-length protein is produced. In some embodiments, the dimerization domain is about 50 to about 160 nt, wherein the reconstitution efficiency is such that an effective amount of correctly linked RNA or full-length protein is produced.

[0154] Efficient recombination between multiple RNA molecules allows for the packaging and delivery of transgenes into AAVs that exceed the packaging limitations of a single AAV. AAV packaging limitations are a major obstacle to gene therapy approaches for diseases caused by the absence or defect of large genes. One application of this system is to express large disease-causing genes using viral vectors with limited packaging capacity. Diseases and genes include, but are not limited to, (Disease (Gene, OMIM Gene Identifier)): 1) Duchenne muscular dystrophy and Becker muscular dystrophy (dystrophin, OMIM:300377); 2) dysferlinopathy (dysferlin, OMIM:603009); 3) cystic fibrosis (CFTR, OMIM:602421); 4) Usher syndrome 1B (myosin VIIA, OMIM:276903); 5) stromal cell carcinoma (STD) syndrome (STD). These include Talgart disease 1 (ABCA4, OMIM:601691); 6) hemophilia A (clotting factor VIII, OMIM:300841); 7) von Willebrand disease (von Willebrand factor, OMIM:613160); 8) Marfan syndrome (fibrillin 1, OMIM:134797); 9) von Recklinghausen disease (neurofibromatosis-1, OMIM:162200), and hearing loss (OTOF, OMIM:603681). Others are listed in Table 1. Additionally, Cas9 proteins (e.g., those exemplified in Examples 20-23) can be expressed using the disclosed systems presented herein, for example, to treat genomic point mutations or to activate or overexpress genes. Transgene delivery can be achieved by splitting the gene into multiple fragments using the techniques presented herein.

[0155] Additional applications of the methods and systems disclosed herein include intersectional gene delivery for targeted gene expression. The differential infection / expression patterns of two viruses encoding fragmented genes can be used. The reconstituted protein is expressed in overlapping cell populations, representing a crossover where both viruses express themselves. Examples of such applications include: (1) delivering both halves (or three thirds, or other portions) of a protein from two (or more) projection targets using a retrogradely transported viral vector to label bifurcated dual-projection neurons; (2) delivering one fragment under the control of a promoter active in population A and a second fragment from a promoter active in population B to specifically tag / engineer the A∪B population; (3) delivering the first half of a protein using a viral vector with tropism for population A and the second half using a viral vector with tropism for population B to specifically tag / engineer the A∪B population; or combinations of these approaches.

[0156] In one example, the dimerization domain is, for example, (a) a small molecule trigger recognized by the aptamer, or (b) an aptamer sequence to facilitate dimerization in the presence of a protein present in the cell that binds to both halves and thus stimulates dimerization.

[0157] In some embodiments, the RNA-RNA interaction required for end-joining can be positively or negatively regulated by other nucleotides, such as (a) an antisense oligonucleotide sequence with homology to both halves (ssDNA-induced dimerization). In such an example, an antisense oligonucleotide with complementary sequences to both halves bridges the two molecules together, thus facilitating spliceosome-mediated recombination of the two molecules. (b) an antisense oligonucleotide sequence with homology to one of the two linked RNAs can prevent RNA dimerization of the two molecules and act as an off switch for gene expression, or (c) an endogenous cellular RNA with homology to both halves (RNA-induced dimerization). In such an example, a cellular RNA (e.g., mRNA or retroelement) with complementary sequences to both halves bridges the two molecules together, thus facilitating spliceosome-mediated recombination of the two molecules.

[0158] These molecular, protein, or RNA-mediated interactions allow for controllable / fine-tuned gene expression levels: by titrating molecules that interact with the binding domain (e.g., antisense oligonucleotides, small molecules, endogenous cellular RNA), the efficiency of dimer formation between the two halves can be modulated, regulating expression levels independently of promoter activity. Such installments can be used when a narrow range of protein expression levels is required.

[0159] III. Series This paper presents a system that can be used to recombine two or more RNA molecules, for example, at least two, at least three, at least four, or at least five different RNA molecules (for example, two, three, four, five, six, seven, eight, nine, or ten different RNA molecules) using a synthetic intron containing a dimerization sequence.Unlike the fragmentation and recombination of two fragments at the protein level, the method disclosed herein does not require extensive protein manipulation to find the appropriate division point.Recombination at the RNA level allows the seamless joining of two fragments of the protein. The methods and systems disclosed herein allow for large genes (and corresponding proteins), e.g., those greater than about 4.5 kb, at least greater than 5 kb, at least greater than 5.5 kb, at least greater than 6 kb, at least greater than 10 kb, at least greater than 13.5 kb, or at least greater than 18 kb, to be separated into two or more fragments or portions, each of which can be introduced into a cell or subject by a separate vector, e.g., multiple AAVs. In one example, the system includes two portions for recombining two RNA molecules, where, for example, a target protein is encoded by at least about 4500 nt to about 9000 nt, e.g., 4000 nt to 5000 nt. In one embodiment, the system includes three portions for recombining three RNA molecules, where, for example, the target protein is encoded by a maximum of about 13,500 nt, e.g., about 4,500 nt to about 13,500 nt, or 4,000 nt to 5,000 nt. In one embodiment, the system includes four portions for recombining four RNA molecules, where, for example, the target protein is encoded by a maximum of about 18,000 nt, e.g., about 4,500 nt to about 18,000 nt, or 4,000 nt to 5,000 nt. This helps overcome limitations on the space available in the vector. In some embodiments, the length of the endogenous promoter limits the capacity of its corresponding gene to be expressed in the AAV.In some embodiments, the length of a coding sequence limits its capacity to be expressed in an AAV. In some embodiments, the length of an endogenous promoter and the length of its coding sequence limit their capacity to be expressed together in an AAV. The systems disclosed herein can be used to express such long sequences that were previously difficult to express in an AAV.

[0160] In some examples, the target protein to be rearranged is a protein associated with a disease such as a monogenic disease, a recessive genetic disease, a disease caused by mutations in large genes (e.g., larger than about 4500 nt, e.g., at least 5 kb, at least 5.5 kb, at least 6 kb, at least 10 kb, at least 13.5 kb, or at least 18 kb), and / or a disease caused by a gene (e.g., promoter + coding sequence) that exceeds the capacity of an AAV (e.g., larger than 5000 nt). Examples of such diseases include, but are not limited to, hemophilia A (caused by mutations in the 7 kb coding region of the F8 gene), hemophilia B (caused by mutations in the F9 gene), Duchenne muscular dystrophy (caused by mutations in the 11 kb coding region of the dystrophin gene), sickle cell anemia (caused by mutations in the beta globin domain of hemoglobin, which has a promoter of approximately 3.5 kb), Stargardt disease (caused by mutations in the 6.9 kb coding region of the ABCA4 gene), and Usher syndrome (caused by mutations in the 7 kb coding region of MYO7A, resulting in hearing loss and visual impairment).

[0161] In one example, the reconstituted target protein is one that can treat a disease such as cancer, e.g., breast cancer, lung cancer, prostate cancer, liver cancer, kidney cancer, brain cancer, bone cancer, ovarian cancer, uterine cancer, skin cancer, or colon cancer. In one example, the reconstituted therapeutic target protein is a toxin, e.g., an AB toxin such as diphtheria toxin A or Pseudomonas exotoxin A, or a form lacking receptor binding activity (e.g., diphtheria toxin DAB389, DAB486, DT388, DT390, or Pseudomonas exotoxin A PE38 or PE40).

[0162] In some examples, RNA sequences encoding target proteins and used in the methods and systems disclosed herein are codon-optimized for expression in target organisms or cells, e.g., human, dog, pig, cat, mouse, or rat cells. Thus, in some examples, the RNA coding sequence contains preferred codons (e.g., does not contain rare codons with low utilization). Codon optimization can be performed by identifying abundant tRNA levels in the target organism or cell. In some examples, protein-encoding RNA sequences are de-enriched for cryptic splice donor and acceptor sites to maximize RNA recombination reactions.

[0163] In some examples, the protein is divided into two parts, for example, into approximately two equal halves (or other proportions, e.g., part A expressing about 1 / 3 and part B expressing about 2 / 3, or part A expressing about 1 / 4 and part B expressing about 3 / 4, etc.). However, each part need not have the same number of nucleotides (or encode the same number of amino acids). In such examples, the method can use two synthetic nucleic acid molecules (e.g., RNA or DNA encoding such RNA), one of which contains the coding sequence for the N-terminal part of the protein and the other of which contains the coding sequence for the C-terminal part of the protein. On this basis, those skilled in the art will understand that in addition to dividing a protein into two fragments or parts, a protein of interest can be divided or split into more than two fragments, for example, three fragments. The design principles of the intron sequences of the three RNA molecules are similar to those of the two RNA molecules, but instead utilize a different pair of dimerization domains for one of the two binding sites. Thus, for example, an N-terminal protein coding sequence may be followed by an intron sequence having a specific binding domain (e.g., a first dimerization sequence), and the intermediate coding sequence may include an intron sequence having a complementary sequence to the first dimerization sequence (a second dimerization sequence). The intermediate coding fragment may be followed by another intron fragment having another dimerization sequence (a third dimerization sequence, different from the second dimerization sequence). The third fragment may include a C-terminal coding sequence for the protein and may also include an intron region having a dimerization sequence complementary to the third dimerization sequence (a fourth dimerization sequence). When more than one intermediate portion is used, the two intermediate portions may be referred to, for example, as an intermediate portion and a first intermediate portion, or a first intermediate portion and a second intermediate portion, or a first intermediate portion, a second intermediate portion, and a third intermediate portion, so that it is understood that each portion is distinct.

[0164] In one example, a desired protein can be split into N-terminal and C-terminal portions (e.g., roughly in half, or unequal portions, e.g., 1 / 3 and 2 / 3, or 1 / 4 and 3 / 4), which can then be reconstituted using the systems and methods disclosed herein. Referring to Figure 6A, in such an example, the system includes at least two synthetic nucleic acid molecules 110, 150. Each nucleic acid molecule 110, 150 can be composed of DNA or RNA (in the case of RNA, the promoter 112, 152 is not present). In some embodiments, molecules 110, 150 are each approximately at least 100 nucleotides / ribonucleotides (nt) in length, e.g., at least 200 nt, at least 300 nt, at least 500 nt, at least 1000 nt, at least 2000 nt, at least 3000 nt, at least 4000 nt, at least 5000 nt, at least 6000 nt, at least 7000 nt, at least 8000 nt, at least 10,000 nt, e.g., 200-10,000 nt, 200-8000 nt, 500-5000 nt, or 200-1000 nt, etc. Molecules 110, 150 can comprise natural and / or unnatural nucleotides or ribonucleotides.

[0165] Molecule 110 is a molecule located 5' of the system because it includes splice donor 116. In embodiments in which molecule 110 is DNA, molecule 110 includes promoter 112 operably linked to a sequence encoding an RNA molecule, which includes, from 5' to 3': a coding sequence for the N-terminal portion of the target protein 114, including a splice junction at the 3' end of the target protein coding sequence, SD 116, optional DISE 118, optional ISE 120, dimerization domain 122, and optional polyadenylation sequence 124. Any promoter 112 (or enhancer) can be used, such as one that utilizes RNA polymerase II, e.g., a constitutive promoter or an inducible promoter. In some examples, promoter 112 is a tissue-specific promoter, such as one that is constitutively active in muscle tissue (e.g., skeletal muscle tissue or cardiac muscle tissue), eye tissue (e.g., retina tissue), inner ear tissue, liver tissue, pancreatic tissue, lung tissue, skin tissue, bone, or kidney tissue. In some embodiments, promoter 112 is a cell-specific promoter, such as one that is constitutively active in cancer cells or normal cells. In some embodiments, promoter 112 is the endogenous promoter of the target protein to be expressed, and in some embodiments, is long (e.g., at least 2500 nt, at least 3000 nt, at least 4000 nt, at least 5000 nt, or at least 7500 nt). In some embodiments, the promoter 112 is at least about 50 nucleotides (nt) in length, e.g., at least 100 nt, at least 200 nt, at least 300 nt, at least 500 nt, at least 1000 nt, at least 2000 nt, at least 3000 nt, at least 4000 nt, at least 5000 nt, at least 6000 nt, at least 7000 nt, at least 8000 nt, at least 9000 nt, or at least 10,000 nt, e.g., 50-10,000 nt, 100-5000 nt, 500-5000 nt, or 50-1000 nt in length.In some embodiments, molecule 110 is DNA and is at least 200 nt, at least 300 nt, at least 500 nt, at least 1000 nt, at least 2000 nt, at least 3000 nt, at least 4000 nt, at least 5000 nt, at least 6000 nt, at least 7000 nt, or at least 8000 nt in length, e.g., 200-10,000 nt, 200-8000 nt, 500-5000 nt, or 200-1000 nt. As shown in Figure 6F, in embodiments where molecule 110 is RNA, e.g., after transcription of DNA into RNA, molecule 110 does not include promoter 112, and 114 is RNA encoded by a coding sequence for the N-terminal portion of a target protein. In some examples, molecule 110 is RNA, does not include a promoter 112, and is at least 200 nt, at least 300 nt, at least 500 nt, at least 1000 nt, at least 2000 nt, at least 3000 nt, at least 4000 nt, at least 5000 nt, at least 6000 nt, at least 7000 nt, or at least 8000 nt, e.g., 200-10,000 nt, 200-8000 nt, 500-5000 nt, or 200-1000 nt in length. Molecule 110 (with or without promoter 112) can include natural and / or non-natural nucleotides or ribonucleotides.

[0166] The splice junction around the 3' end of the N-terminal coding sequence (or RNA sequence encoded thereby) 114 can match a consensus sequence found in the target cell or organism into which the molecule 110, 150 is introduced. In humans, the splice junction sequence is AG (adenine-guanine) or UG (uracil-guanine) at positions -1 and -2 of the 5' splice site for U2-dependent introns, or AG, UG, CU (cytosine-uracil), or UU for U12-dependent introns. Thus, in some embodiments, the splice junction is 2 nt in length, and the 3' end of the N-terminal coding portion 114 is AG, UG, CU, or UU. In some embodiments, a DNA molecule encoding a portion of a target protein includes sequences encoding portions of multiple splice junctions, for example, at the 3' end of a DNA molecule encoding the N-terminal portion of the target protein and at the 5' end of a DNA molecule encoding the C-terminal portion of the target protein.

[0167] The remaining 3'-terminal portion of molecule 110 is an intron 130. In some embodiments, intron sequence 130 is approximately at least 10 nt in length, e.g., at least 20 nt, at least 50 nt, at least 100 nt, at least 250 nt, at least 250 nt, at least 300 nt, at least 400 nt, or at least 500 nt, e.g., 20-500 nt, 20-250 nt, 20-100 nt, 50-100 nt, or 50-200 nt in length. Immediately following the N-terminal coding sequence (or the RNA encoded thereby) 114 is a splice donor (SD) 116 (e.g., an SD consensus sequence, such as the SD human consensus sequence). Thus, the SD 116 of intron sequence 130 is 3' to the N-terminal coding sequence 114. The SD 116 forms a recognition sequence for spliceosome components to bind to the RNA molecule. The sequence of SD116 can be an SD consensus sequence found in a target cell or organism into which the molecule 110, 150 is introduced. In some examples, SD116 is at least 2 nt, e.g., at least 5 nt, or at least 10 nt in length, e.g., 2-10 nt, 2-8 nt, 2-5, or 5-10 nt. SD116 can be used to recruit U2- or U12-dependent splicing machinery. In one example, U2-dependent splicing is used in human cells, and the SD116 sequence includes or is GUAAGUAUU. In one example, U12-dependent splicing is used in human cells, and the SD116 sequence includes or is AUAUCCUUUUUA (SEQ ID NO: 137) or GUAUCCUUUUUA (SEQ ID NO: 138). It is understood throughout that RNA sequences can be written using the nucleotides A, G, T, and C, and DNA sequences can be written using the nucleotides A, G, U, and C.

[0168] The intron sequence 130 optionally includes one or both of a set of splicing enhancer sequences, referred to as downstream intron splice enhancers (DISEs) 118 and intron splice enhancers (ISEs) 120, which stimulate the action (e.g., increase activity) of the spliceosome. In some examples, the intron sequence 130 includes at least two splicing enhancer sequences, e.g., at least three, at least four, or at least five splicing enhancer sequences. Exemplary splicing enhancer sequences include DISEs 118 and ISEs 120. In some examples, including one or more splicing enhancer sequences 118, 120 in the intron sequence 130 increases splicing efficiency by at least 20%, at least 30%, at least 40%, at least 50%, at least 75%, at least 80%, at least 90%, or at least 95%. Exemplary splicing enhancer sequences that can be used are SEQ ID NOs: 26-136, 151, and 152, as well as GGGTTT, GGTGGT, TTTGGG, GAGGGG, GGTATT, GTAACG, GGGGGTAGG, GGAGGGTTT, GGGTGGTGT TTCAT, CCATTT, TTTTAAA, TGCAT, TGCATG, TGTGTT, CTAAC, TCTCT, TCTGT, TCTTT, TGCATG, CTAAC, CTGCT, TAACC, AGCTT, TTCATTA, GTTAG, TTTTGC, ACTAAT, ATGTTT, CTCTG, GGG, GGG(N)2-4GGG, TGGG, YCAY, UGCAUG, or 3x(G 3~6 N 1~7In some examples, DISE118, if present, may be at least 3 nt, at least 4 nt, at least 5 nt, at least 10 nt, at least 25 nt, at least 50 nt, at least 75 nt, or at least 100 nt in length, e.g., 3-10 nt, 3-11 nt, 4-11 nt, 5-11 nt, 10-50 nt, 5-100 nt, 10-25 nt, 10-20 nt, or 20-75 nt, and the sequence of DISE118 is or includes CUCUUUCUUUTCCAUGGGUUGGCU (SEQ ID NO: 134), TGCATG, CTAAC, CTGCT, TAACC, AGCTT, TTCATTA, GTTAG, TTTTGC, ACTAAT, ATGTTT, or CTCTG. In some embodiments, ISE120, if present, can be approximately at least 3 nt, at least 4 nt, at least 5 nt, at least 10 nt, e.g., at least 20 nt, at least 25 nt, at least 30 nt, at least 40 nt, or at least 50 nt in length, e.g., 3-10 nt, 3-11 nt, 4-11 nt, 5-11 nt, 10-50 nt, 20-25 nt, 10-25 nt, 10-20 nt, or 20-40 nt in length. In one embodiment, the sequence of ISE120 is or includes GGCUGAGGGAAGGACUGUCCUGGG (SEQ ID NO: 135), GGGUUAUGGGACC (SEQ ID NO: 136), TTCAT, CCATTT, TTTTAAA, TGCAT, TGCATG, TGTGTT, CTAAC, TCTCT, TCTGT, or TCTTT. In some embodiments, intron sequence 130 includes at least two, at least three, or at least four ISE120s.In some embodiments, ISE120 is or comprises at least one sequence, e.g., at least two, at least three, e.g., one, two, three, four, or five, having at least 80%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity to SEQ ID NO: 173, 174, 175, 176, 177, 178, 179, 180, 182, 183, 184, 185, 186, 187, 188, 189, 190, 191, 192, 193, 194, 195, 196, 199, 200, 201, 202, or 203. In some examples, DISE118 is or comprises at least one sequence, e.g., at least two, at least three such sequences, e.g., one, two, three, four or five such sequences, having at least 80%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99% or 100% sequence identity to SEQ ID NO: 173, 174, 175, 176, 177, 178, 179, 180, 182, 183, 184, 185, 186, 187, 188, 189, 190, 191, 192, 193, 194, 195, 196, 199, 200, 201, 202, or 203.

[0169] 3' to SD 116 (and enhancer sequences 118, 120, if present) is a dimerization domain 122 used to bring together the N-terminal coding sequence (or RNA encoded thereby) 114 and the C-terminal coding sequence 154. The intron sequence 130 portion of molecule 110 can optionally include a polyadenylation site 124 at its 3' end to terminate transcription of the fragment. In some examples, polyadenylation sequence 124 is a polyA sequence of at least 15 A's, e.g., 15-30 or 15-20 A's.

[0170] In some embodiments, the first dimerization domain 122 (and the second dimerization domain 154 of molecule 150) includes multiple unpaired nucleotides (i.e., unpaired within the structure of molecule 110 itself). Having unpaired nucleotides within the dimerization domains allows the 5' (or first) dimerization domain 122 and the 3' (or second) dimerization domain 154 to interact through base pairing. Through this interaction, molecules 110 and 150 are held in close proximity, thereby prompting the spliceosome to recombine the two molecules by joining the N-terminal coding region (or the RNA encoded thereby) 114 and the C-terminal coding region (or the RNA encoded thereby) 164.

[0171] In one example, the dimerization domains 122 (and 154) comprise "low-diversity sequences" containing limited nucleotide diversity, and therefore are unlikely to form stem-loops themselves in the secondary structure of each molecule 110, 150. Such low-diversity dimerization domains 122 (and 154) can be in a relatively open configuration, regardless of the sequences of the DNA (or RNA encoded thereby) 114, 164 encoding the N- and C-termini of the protein. This allows the nucleotides of the first dimerization domain 122 to be available for base pairing with the second dimerization domain 154 of the corresponding molecule 150, thereby enabling subsequent binding of the N-terminal coding sequence (or RNA encoded thereby) 114 and the C-terminal coding sequence (or RNA encoded thereby) 164. In some examples, the first dimerization domain 122 and the second dimerization domain 154 comprise low-diversity sequences interspersed with sequences that can form stems, resulting in local RNA loops that are open and available for base pairing in the absence of pseudoknot formation (FIG. 6B). Exemplary low-diversity sequences include a repeated run of U (e.g., 30-500 U), a repeated run of A (e.g., 30-500 A), a repeated run of G (e.g., 30-500 G), a repeated run of C (e.g., 30-500 C), a mixture containing only A and G (e.g., 30-500 A and G, e.g., AAAGAAGGAA(...) (SEQ ID NO: 149), or a mixture containing only C and U (e.g., 30-500 C and U, e.g., CUUUCUUUUCUU(...) (SEQ ID NO: 150)). Other exemplary low diversity sequences include complementary sequences that form a helix flanked by low diversity sequences.

[0172] In some examples, the first dimerization domain 122 and the second dimerization domain 154 contain only purines or only pyrimidines. In one example, the first dimerization domain 122 contains only purines, while the second dimerization domain 154 contains only pyrimidines. In another example, the first dimerization domain 122 contains only pyrimidines, while the second dimerization domain 154 contains only purines. Because purines cannot pair with themselves (nor can pyrimidines), these stretches of RNA have a predicted open structure.

[0173] In some examples, the first dimerization domain and the second dimerization domain 122, 154 do not contain cryptic splice acceptors that may compete with RNA recombination, such as sequences similar to the splice donor consensus sequence NNNAGGUNNNN (SEQ ID NO: 151) or NNNUGGUNNNN (SEQ ID NO: 152), where N refers to any nucleotide. In some examples, the first dimerization domain 122 is 1000 nt or less, e.g., 750 nt or less, or more than 500 nt, e.g., 6 to 1000 nt, 10 to 1000 nt, 20 to 1000 nt, 30 to 1000 nt, 30 to 750 nt, 30 to 500 nt, 50 to 500 nt, 50 to 100 nt, or 100 to 250 nt. In some embodiments, the first dimerization domain 122 is greater than 50 nt, e.g., at least 51 nt, at least 100 nt, at least 150 nt, at least 161 nt, or at least 170 nt, e.g., 51-159 nt, 51-150 nt, 51-120 nt, 51-100 nt, or 51-70 nt. In some embodiments, the first dimerization domain 122 is greater than 160 nt, e.g., at least 161 nt, at least 170 nt, at least 180 nt, at least 200 nt, at least 300 nt, at least 400 nt, at least 500 nt, at least 600 nt, at least 700 nt, at least 800 nt, at least 900 nt, or at least 1000 nt, e.g., 161-100 nt, 161-500 nt, 161-300 nt, 161-200 nt, or 161-170 nt. In some examples, the first dimerization domain 122 is less than 50 nt, for example, 6 to 49 nt, 6 to 45 nt, 6 to 40 nt, 6 to 30 nt, 6 to 20 nt, or 6 to 10 nt.

[0174] In some embodiments, the dimerization domain is 20 to 160 nt, 50 to 500 nt, or 500 to 1000 nt. In some embodiments, the dimerization domain is about 20 nt to about 160 nt. In some embodiments, the dimerization domain is about 20 nt to about 40 nt, about 20 nt to about 50 nt, about 20 nt to about 70 nt, about 20 nt to about 90 nt, about 20 nt to about 100 nt, about 20 nt to about 110 nt, about 20 nt to about 120 nt, about 20 nt to about 130 nt, about 20 nt to about 140 nt, about 20 nt to about 150 nt, about 20 nt to about 160 nt, about 40 nt to about 50 nt, about 40 nt to about 70 nt, about 40 nt to about 90 nt, about 40 nt to about 100 nt, about 40 nt to about 110 nt, about 4 ... 0nt to about 120nt, about 40nt to about 130nt, about 40nt to about 140nt, about 40nt to about 150nt, about 40nt to about 160nt, about 50nt to about 70nt, about 50nt to about 90nt, about 50nt to about 100nt, about 50nt to about 110nt , about 50nt to about 120nt, about 50nt to about 130nt, about 50nt to about 140nt, about 50nt to about 150nt, about 50nt to about 160nt, about 70nt to about 90nt, about 70nt to about 100nt, about 70nt to about 110nt, about 70nt to about 1 20nt, about 70nt to about 130nt, about 70nt to about 140nt, about 70nt to about 150nt, about 70nt to about 160nt, about 90nt to about 100nt, about 90nt to about 110nt, about 90nt to about 120nt, about 90nt to about 130nt, about 90 nt ~ about 140nt, about 90nt to about 150nt, about 90nt to about 160nt, about 100nt to about 110nt, about 100nt to about 120nt, about 100nt to about 130nt, about 100nt to about 140nt, about 100nt to about 150nt, about 100nt about 160nt, about 110nt to about 120nt, about 110nt to about 130nt, about 110nt to about 140nt, about 110nt to about 150nt, about 110nt to about 160nt, about 120nt to about 130nt, about 120nt to about 140nt, about 120nt to about 150nt, about 120nt to about 160nt, about 130nt to about 140nt, about 130nt to about 150nt, about 130nt to about 160nt, about 140nt to about 150nt, about 140nt to about 160nt, or about 150nt to about 160nt.In some embodiments, the dimerization domain is about 20 nt, about 40 nt, about 50 nt, about 70 nt, about 90 nt, about 100 nt, about 110 nt, about 120 nt, about 130 nt, about 140 nt, about 150 nt, or about 160 nt. In some embodiments, the dimerization domain is at least about 20 nt, about 40 nt, about 50 nt, about 70 nt, about 90 nt, about 100 nt, about 110 nt, about 120 nt, about 130 nt, about 140 nt, or about 150 nt. In some embodiments, the dimerization domain is at most about 40 nt, about 50 nt, about 70 nt, about 90 nt, about 100 nt, about 110 nt, about 120 nt, about 130 nt, about 140 nt, about 150 nt, or about 160 nt.

[0175] In some embodiments, the dimerization domain is about 50 nt to about 500 nt. In some embodiments, the dimerization domain is about 50 nt to about 100 nt, about 50 nt to about 150 nt, about 50 nt to about 200 nt, about 50 nt to about 250 nt, about 50 nt to about 300 nt, about 50 nt to about 350 nt, about 50 nt to about 400 nt, about 50 nt to about 500 nt, about 100 nt to about 150 nt, about 100 nt to about 200 nt, about 100 nt to about 250 nt, about 100 nt to about 300 nt, about 100 nt to about 350 nt, about 100 nt to about 400 nt, about 100 nt to about 500 nt, about 150 nt to about 200 nt, about 150 nt to about 250 nt, about 150 nt to about 300 nt, about 150nt to about 350nt, about 150nt to about 400nt, about 150nt to about 500nt, about 200nt to about 250nt, about 200nt to about 300nt, about 200nt to about 350nt, about 200nt to about 400nt, about 200nt to about 500nt, about 250nt to about 300nt, about 250nt to about 350nt, about 250nt to about 400nt, about 250nt to about 500nt, about 300nt to about 350nt, about 300nt to about 400nt, about 300nt to about 500nt, about 350nt to about 400nt, about 350nt to about 500nt, or about 400nt to about 500nt. In some embodiments, the dimerization domain is about 50 nt, about 100 nt, about 150 nt, about 200 nt, about 250 nt, about 300 nt, about 350 nt, about 400 nt, or about 500 nt. In some embodiments, the dimerization domain is at least about 50 nt, about 100 nt, about 150 nt, about 200 nt, about 250 nt, about 300 nt, about 350 nt, or about 400 nt. In some embodiments, the dimerization domain is at most about 100 nt, about 150 nt, about 200 nt, about 250 nt, about 300 nt, about 350 nt, about 400 nt, or about 500 nt.

[0176] In some examples, the sequences of the first and second dimerization domains 122 and 154 are determined by in silico structure prediction screening (e.g., using RNA folding structure prediction to screen a library of potential dimerization domain sequences; selecting sequences with a high proportion of unpaired nucleotides in both the dimerization domain and the corresponding anti-dimerization domain), low diversity nucleotide design (e.g., designing the dimerization domain to contain a stretch of low diversity sequence, such as a repeat sequence of only U, only A, only C, only G, only R (G and A), or only Y (U and C), a sequence that cannot fold with itself), or empirical screening (e.g., synthesizing a library of dimerization domains and corresponding anti-dimerization domains and screening for maximum recombination efficiency).

[0177] In some embodiments, the sequences of the first and second dimerization domains 122, 154 are designed to contain complementary RNA hairpin structures (also referred to as stem loops) that can form strong kissing loop interactions with their counterparts. In some embodiments, kissing loops are used when three or more dimerization domains, e.g., four or more or five or more dimerization domains, e.g., three, four, five, six, seven, eight, nine, or ten dimerization domains, are used to connect three or more portions of a coding sequence (e.g., Figure 6E). Each hairpin loop (or stem loop) of a kissing loop is composed of at least two complementary sequences (e.g., forming a stem) separated by a region of non-complementary sequence (e.g., forming a loop). In some examples, the dimerization domain may be comprised of one or more (e.g., at least two, at least three, at least four, or at least five, e.g., 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20) loops. In some examples using multiple loops, all or some of the loops may be repeated. In some examples using multiple loops, all or some of the loops may be different. In some examples, each complementary sequence is about 4-100 nt and is separated by a loop of about 3-20 nt. Base pairing between the two complementary sequences results in a helix (or stem) of, for example, at least 4 bp, at least 5 bp, at least 10 bp, at least 20 bp, at least 30 bp, at least 40 bp, at least 50 bp, at least 75 bp, at least 90 bp, or at least 100 bp, e.g., 4 to 100 bp, 5 to 75 bp, or 10 to 50 bp. In some embodiments, the loop portion is at least 3 nt, at least 5 nt, at least 10 nt, at least 15 nt, or at least 20 nt, e.g., 3 to 20 nt, 5 to 15 nt, or 5 to 10 nt, wherein the loop is not base paired.The complementary sequence between the two hairpin loops results in base pairing and the creation of a kissing loop / kissing stem-loop interaction. In some embodiments, the complementary sequence between the two hairpin loops is between at least 3 nucleotides of one loop and at least 3 nucleotides of the second loop, e.g., at least 4 nt, at least 5 nt, at least 6 nt, at least 7 nt, at least 8 nt, at least 9 nt, at least 10 nt, at least 11 nt, at least 12 nt, at least 13 nt, at least 14 nt, at least 15 nt, at least 16 nt, at least 17 nt, at least 19 nt, or at least 20 nt (e.g., 3 nt, 4 nt, 5 nt, 6 nt, 7 nt, 8 nt, 9 nt, 10 nt, 11 nt, 12 nt, 13 nt, 14 nt, The amino acid sequence may be between at least one of the first and second loops (e.g., 15nt, 16nt, 17nt, 18nt, 19nt, or 20nt) and at least 4 nt, at least 5 nt, at least 6 nt, at least 7 nt, at least 8 nt, at least 9 nt, at least 10 nt, at least 11 nt, at least 12 nt, at least 13 nt, at least 14 nt, at least 15 nt, at least 16 nt, at least 17 nt, at least 19 nt, or at least 20 nt (e.g., 3 nt, 4 nt, 5 nt, 6 nt, 7 nt, 8 nt, 9 nt, 10 nt, 11 nt, 12 nt, 13 nt, 14 nt, 15 nt, 16 nt, 17 nt, 18 nt, 19 nt, or 20 nt) of the amino acid sequence. In some embodiments, the complementary sequence between two hairpin loops is present in at least 15%, at least 20%, at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, or 100% of the total loop sequence.

[0178] In some cases, the stems of the kissing loops are selected to base pair in trans between the two RNA molecules. In such examples, after the formation of a kissing loop interaction between one hairpin loop on one molecule and another hairpin loop on the second molecule, the stem (or helix) region of each of the initial hairpin loops can base pair in trans between the two RNA molecules through strand displacement / invasion and extension of duplex formation. In some examples, up to about 85% of the nucleotides within the initial loop sequence can remain unpaired after extension of duplex formation (e.g., about 15% of the nucleotides are paired between the two loops). In some examples, the kissing loop is based on the HIV-1 DIS loop (SEQ ID NOs: 139 and 140, FIG. 17A ) and includes two A nucleotides on the 5′ side of a 6-nucleotide complementary sequence followed by one A nucleotide on the 3′ side (e.g., AANNNNNNA, where N can be A, U, G, or C). In some examples, the kissing loop is based on the HIV-2 kissing loop dimerization domain (SEQ ID NOs: 141 and 142, FIG. 17B) and includes a G nucleotide and an A nucleotide on the 5' side of a six-nucleotide complementary sequence, followed by three A nucleotides on the 3' side (e.g., GANNNNNNAAA (SEQ ID NO: 153), where N can be A, U, G, or C).

[0179] In one configuration, extended duplex formation is favored by including mismatches in the initial stem, resulting in a high percentage of matches in the extended duplex. Thus, in some embodiments, the helix or stem region of a hairpin loop may contain up to 30% of initially unpaired base pairs (e.g., 30% or less, 20% or less, 15% or less, 10% or less, 5% or less, or 1% or less of the base pairs, e.g., 1-30%, 5-30%, 10-30%, or 25-30% of the base pairs are initially unpaired). These unpaired regions may form bulges, mismatches, or internal loops.

[0180] In addition to the interaction of two hairpin loops (kissing loop interactions), other forms of loop interactions can be utilized for the first and second dimerization domains 122, 154. In one example, the loop is a bulge, and one strand of the base-paired helix contains one or more nucleotides that protrude from the stem structure. Exemplary bulges are at least 1 nt, at least 2 nt, at least 3 nt, at least 4 nt, at least 5 nt, at least 10 nt, or at least 20 nt, e.g., 1-20 nt, 1-15 nt, 1-10 nt, or 5-10 nt, or 1 nt, 2 nt, 3 nt, 4 nt, 5 nt, 6 nt, 7 nt, 8 nt, 9 nt, 10 nt, 11 nt, 12 nt, 13 nt, 14 nt, 15 nt, 16 nt, 17 nt, 18 nt, 19 nt, or 20 nt. In one example, the loop is an internal loop, eg, one or more nucleotides within a helix are mismatched, resulting in a helix interrupted by an internal loop at the position of the mismatch. In some examples, the helix is ​​at least 4 nt on each strand (e.g., at least 5 nt, at least 10 nt, at least 20 nt, at least 30 nt, at least 40 nt, at least 50 nt, at least 75 nt, at least 90 nt, or at least 100 nt, e.g., between 4 and 100 nt, between 5 and 75 nt, or between 10 and 50 nt, e.g., between 4 and 100 nt) and at least 1 nt on each side of the internal loop (e.g., at least 2 nt, at least 3 nt, at least 4 nt, at least 5 nt, at least 10 nt, or at least 20 nt on each strand, e.g., between 1 and 20 nt, between 1 and 15 nt, between 1 and 10 nt, or between 5 and 10 nt, or 1 nt, 2 nt, 3 nt, 4 nt, 5 nt, 6 nt, 7 nt, 8 nt, 9 nt, 10 nt, 11 nt, 12 nt, 13 nt, 14 nt, 15 nt, 16 nt, 17 nt, 18 nt, 19 nt, or 20 nt). In one embodiment, the loop is a multi-branched loop, from which three helices or stems form a triangle with the three helices connected by one or more unpaired nucleotides.In some examples, each of the helices is at least 4 bp (e.g., at least 5 bp, at least 10 bp, at least 20 bp, at least 30 bp, at least 40 bp, at least 50 bp, at least 75 bp, at least 90 bp, or at least 100 bp, e.g., 4 to 100 bp, 5 to 75 bp, or 10 to 50 bp), and the unpaired nucleotides forming the triangle are at least 3 nt (e.g., at least 4 nt, at least 5 nt, at least 10 nt, at least 20, at least 15, at least 30, at least 40, at least 50, or at least 60 nt, e.g., 3 to 60 nt, 3 to 30 nt, 3 to 25 nt, or 5 to 20 nt, e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 2, 25, 30, 35, 40, 45, 50, 55, or 60 nucleotides). Kissing interactions can occur between any two of these types of loops (e.g., between two or more binding domains each containing one or more helices). In some examples, to allow for extension of duplex formation after the initial loop kissing interaction, a helix in one dimerization domain (e.g., first dimerization domain 122) has a direct counterpart in the other binding domain (e.g., second dimerization domain 154). In some examples, dimerization domains containing helices to generate loops form a single kissing stem-loop when two or more dimerization domains (e.g., 122, 154 in Figure 6A) interact. In some examples, dimerization domains containing helices form multiple loops for kissing loop interactions when two or more dimerization domains (e.g., 122, 154 in Figure 6A) interact. In some examples, one or more dimerization domains (e.g., 122 in Figure 6A) contain a bulge, a single-base bulge, a mismatch or internal loop, or a helix that is destabilized by the inclusion of a GU wobble pair, but is matched with the other binding domain (e.g., 154 in Figure 6A) to favor extended duplex formation after initial kissing / pairing.In some examples, one or more dimerization domains (e.g., 122 in Figure 6A) contain destabilized helices that, when stabilized (e.g., theophylline-switched kissing loops), expose loops that can interact with a second dimerization domain (e.g., 122 in Figure 6A) through loop-loop interactions (e.g., kissing / pairing).

[0181] In some examples, these stem loops contain at least 10 nt, e.g., at least 20 nt, at least 25 nt, at least 50 nt, at least 75 nt, or at least 100 nt in length, e.g., 10-50 nt, 20-25 nt, 10-100 nt, 10-20 nt, or 20-40 nt in length. Each dimerization domain can contain at least one individual stem loop, e.g., at least 2, at least 5, at least 10, at least 15, or at least 20 individual stem loops, e.g., 1-20, 2-5, or 1-10.

[0182] In some embodiments, 3 to 10 portions of the coding sequence are connected by 2 to 9 kissing loops, e.g., 3 portions are connected by 2 kissing loops, 4 portions are connected by 3 kissing loops, etc., where each of the 2 to 9 kissing loops is different. In some embodiments, the kissing loop comprises a plurality of stem loops, e.g., 2 to 20 stem loops. In some embodiments, each of the plurality of stem loops within a kissing loop is the same. In some embodiments, each of the plurality of stem loops within a kissing loop is different. In some embodiments, the dimerization domain comprises 1 to 20 stem loops. In some embodiments, the dimerization domain comprises 1 stem loop to 20 stem loops. In some embodiments, the dimerization domain is one stem loop to two stem loops, one stem loop to three stem loops, one stem loop to four stem loops, one stem loop to five stem loops, one stem loop to six stem loops, one stem loop to seven stem loops, one stem loop to eight stem loops, one stem loop to nine stem loops, one stem loop to ten stem loops, one stem loop to fifteen stem loops, one stem loop to twenty stem loops, two stem loops to three stem loops, two stem loops to four stem loops, two stem loops to five stem loops, two stem loops to six stem loops, stem loops, 2 stem loops to 7 stem loops, 2 stem loops to 8 stem loops, 2 stem loops to 9 stem loops, 2 stem loops to 10 stem loops, 2 stem loops to 15 stem loops, 2 stem loops to 20 stem loops, 3 stem loops to 4 stem loops, 3 stem loops to 5 stem loops, 3 stem loops to 6 stem loops, 3 stem loops to 7 stem loops, 3 stem loops to 8 stem loops, 3 stem loops to 9 stem loops, 3 stem loops to 10 stem loops, 3 stem loops to 15 stem loops, 3 stem loops to 20 stem loops,4 stem loops to 5 stem loops, 4 stem loops to 6 stem loops, 4 stem loops to 7 stem loops, 4 stem loops to 8 stem loops, 4 stem loops to 9 stem loops, 4 stem loops to 10 stem loops, 4 stem loops to 15 stem loops, 4 stem loops to 20 stem loops, 5 stem loops to 6 stem loops, 5 stem loops to 7 stem loops, 5 stem loops to 8 stem loops, 5 stem loops to 9 stem loops, 5 stem loops to 10 stem loops, 5 stem loops to 15 stem loops, 5 stem loops to 20 stem loops, 6 stem loops to 7 stem loops, 6 stem loops to 8 stem loops, 6 stem loops to 9 stem loops, 6 stem loops The stem loops may include up to 10 stem loops, 6 stem loops to 15 stem loops, 6 stem loops to 20 stem loops, 7 stem loops to 8 stem loops, 7 stem loops to 9 stem loops, 7 stem loops to 10 stem loops, 7 stem loops to 15 stem loops, 7 stem loops to 20 stem loops, 8 stem loops to 9 stem loops, 8 stem loops to 10 stem loops, 8 stem loops to 15 stem loops, 8 stem loops to 20 stem loops, 9 stem loops to 10 stem loops, 9 stem loops to 15 stem loops, 9 stem loops to 20 stem loops, 10 stem loops to 15 stem loops, 10 stem loops to 20 stem loops, or 15 stem loops to 20 stem loops. In some embodiments, the dimerization domain comprises 1 stem loop, 2 stem loops, 3 stem loops, 4 stem loops, 5 stem loops, 6 stem loops, 7 stem loops, 8 stem loops, 9 stem loops, 10 stem loops, 15 stem loops, or 20 stem loops. In some embodiments, the dimerization domain comprises at least 1 stem loop, 2 stem loops, 3 stem loops, 4 stem loops, 5 stem loops, 6 stem loops, 7 stem loops,In some embodiments, the dimerization domain comprises at most 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, or 20 stem loops.

[0183] Other mechanisms can be used that allow two or more dimerization domains (e.g., 122, 154 in FIG. 6A ) to bind or interact with each other sufficiently to cause recombination of the coding sequences. In some examples, the two or more dimerization domains (e.g., 122, 154 in FIG. 6A ) can interact with each other, e.g., through non-base pairing interactions, or are nucleic acid aptamers (e.g., RNA aptamers) that can bind to a common molecule (e.g., a protein, ATP, a metal ion, a cofactor, or a synthetic ligand). In some examples, the two or more dimerization domains (e.g., 122, 154 in FIG. 6A ) do not hybridize with each other, but both (or all) hybridize with the same bridging nucleic acid molecule. In some examples, such bridging nucleic acid molecules can be exogenously provided to a cell, tissue, or organism. In some examples, such bridging nucleic acid molecules can be DNA or RNA sequences within a cell, e.g., a transcript or genomic locus. In some embodiments, two or more dimerization domains (eg, 122, 154 in Figure 6A) are sequences that can interact with each other, for example, through non-base pairing interactions.

[0184] Molecule 150 is a 3'-located molecule and includes a splice acceptor (SA) 162 and a second dimerization domain 154. In embodiments in which molecule 150 is DNA, molecule 150 includes a second promoter 152 followed by an intron sequence 170. Promoter 152 can be operably linked to intron sequence 170. Any promoter 152 can be used, whether constitutive or inducible. In some examples, promoter 152 is a tissue-specific promoter, such as one constitutively active in muscle tissue (e.g., skeletal muscle tissue or cardiac muscle tissue), eye tissue (e.g., retinal tissue), inner ear tissue, liver tissue, pancreatic tissue, lung tissue, skin tissue, bone, or kidney tissue. In some examples, promoter 152 is a cell-specific promoter, such as one constitutively active in cancer cells or normal cells. In some embodiments, promoter 112 is the endogenous promoter of the target protein to be expressed, and in some embodiments, is long (e.g., at least 2500 nt, at least 3000 nt, at least 4000 nt, at least 5000 nt, or at least 7500 nt). In some embodiments, promoter 112 is at least about 50 nucleotides (nt) in length, e.g., at least 100 nt, at least 200 nt, at least 300 nt, at least 500 nt, at least 1000 nt, at least 2000 nt, at least 3000 nt, at least 4000 nt, at least 5000 nt, at least 6000 nt, at least 7000 nt, at least 8000 nt, at least 9000 nt, or at least 10,000 nt, e.g., 50-10,000 nt, 100-5000 nt, 500-5000 nt, or 50-1000 nt in length. In some embodiments, promoter 112 and promoter 152 are the same promoter. In other examples, promoter 112 and promoter 152 are different promoters.In some embodiments, molecule 150 is DNA and is at least 200 nt, at least 300 nt, at least 500 nt, at least 1000 nt, at least 2000 nt, at least 3000 nt, at least 4000 nt, at least 5000 nt, at least 6000 nt, at least 7000 nt, or at least 8000 nt, e.g., 200-10,000 nt, 200-8000 nt, 500-5000 nt, or 200-1000 nt in length. As shown in Figure 6F, in embodiments where molecule 150 is RNA, e.g., after expression of DNA into RNA, molecule 150 no longer includes promoter 152, and 164 is RNA encoded by a coding sequence for the C-terminal portion of the target protein. In some examples, molecule 150 is RNA, does not include promoter 152, and is at least 200 nt, at least 300 nt, at least 500 nt, at least 1000 nt, at least 2000 nt, at least 3000 nt, at least 4000 nt, at least 5000 nt, at least 6000 nt, at least 7000 nt, or at least 8000 nt, e.g., 200-10,000 nt, 200-8000 nt, 500-5000 nt, or 200-1000 nt in length. Molecule 150 (with or without promoter 152) can include natural and / or non-natural nucleotides or ribonucleotides.

[0185] Intron sequence 170 includes second dimerization domain 154, optional ISE 156, branch point 158, polypyrimidine tract 160, followed by splice acceptor sequence 162. In some examples, intron sequence 130 is approximately at least 10 nt, e.g., at least 20 nt, at least 30 nt, at least 50 nt, at least 100 nt, at least 250 nt, at least 250 nt, at least 300 nt, at least 400 nt, or at least 500 nt in length, e.g., 20-500, 20-250, 20-100, 50-100, 30-500, or 50-200 nt in length.

[0186] The second dimerization domain 154 has a sequence that is the reverse complement of the sequence of the first dimerization domain 122 of the molecule 110. Accordingly, the same design features and considerations as those for the first dimerization domain 122 described above also apply to the second dimerization domain 154. For example, in some embodiments, the second dimerization domain 154 contains a stem-loop that can form a kissing-loop interaction with the first dimerization domain 122. In some embodiments, the second dimerization domain 154 does not contain a cryptic splice acceptor (e.g., NNNAGGUNNN; SEQ ID NO: 143) that may compete with RNA recombination. In some embodiments, the second dimerization domain 154 has a low diversity sequence. In some embodiments, second dimerization domain 154 is 1000 nt or less, e.g., 750 nt or less, or greater than 500 nt, e.g., 30-1000 nt, 30-750 nt, 30-500 nt, 50-500 nt, 50-100 nt, or 100-250 nt. In some embodiments, second dimerization domain 154 is greater than 50 nt, e.g., at least 51 nt, at least 100 nt, at least 150 nt, at least 161 nt, or at least 170 nt, e.g., 51-159 nt, 51-150 nt, 51-120 nt, 51-100 nt, or 51-70 nt. In some embodiments, second dimerization domain 154 is greater than 160 nt, e.g., at least 161 nt, at least 170 nt, at least 180 nt, at least 200 nt, at least 300 nt, at least 400 nt, at least 500 nt, at least 600 nt, at least 700 nt, at least 800 nt, at least 900 nt, or at least 1000 nt, e.g., 161-100 nt, 161-500 nt, 161-300 nt, 161-200 nt, or 161-170 nt. In some embodiments, second dimerization domain 154 is less than 50 nt, e.g., 6-49 nt, 6-45 nt, 6-40 nt, 6-30 nt, 6-20 nt, or 6-10 nt.

[0187] 3' to second dimerization domain 154 is an optional ISE 156, a branchpoint sequence 158 (e.g., a branchpoint consensus sequence), a polypyrimidine tract 160, followed by a splice acceptor sequence 162. ISE 156, like ISE 120 and DISE 118 of molecule 110, stimulates the spliceosome to catalyze the recombination reaction. In some examples, intron sequence 150 includes at least two ISEs 156, e.g., at least three, at least four, or at least five ISEs 156. Exemplary splicing enhancer sequences include ISE 156. In some examples, inclusion of one or more splicing enhancer sequences 156 in intron sequence 150 increases the efficiency of recombination or splicing by at least 10%, at least 20%, at least 30%, at least 40%, or at least 50%. Exemplary splicing enhancer sequences that can be used are SEQ ID NOs: 26-136, 151, and 152, as well as GGGTTT, GGTGGT, TTTGGG, GAGGGG, GGTATT, GTAACG, GGGGGTAGG, GGAGGGTTT, GGGTGGTGT TTCAT, CCATTT, TTTTAAA, TGCAT, TGCATG, TGTGTT, CTAAC, TCTCT, TCTGT, TCTTT, TGCATG, CTAAC, CTGCT, TAACC, AGCTT, TTCATTA, GTTAG, TTTTGC, ACTAAT, ATGTTT, CTCTG, GGG, GGG(N)2-4GGG, TGGG, YCAY, UGCAUG, or 3x(G 3~6 N 1~7). In some examples, ISE156, if present, can be approximately at least 3 nt, at least 4 nt, at least 5 nt, at least 10 nt, e.g., at least 20 nt, at least 25 nt, at least 30 nt, at least 40 nt, or at least 50 nt in length, e.g., 3-10 nt, 3-11 nt, 4-11 nt, 5-11 nt, 10-50 nt, 20-25 nt, 10-25 nt, 10-20 nt, or 20-40 nt in length. In one example, the sequence of ISE156 is or includes GGCUGAGGGAAGGACUGUCCUGGG (SEQ ID NO: 135), GGGUUAUGGGACC (SEQ ID NO: 136), TTCAT, CCATTT, TTTTAAA, TGCAT, TGCATG, TGTGTT, CTAAC, TCTCT, TCTGT, or TCTTT. In some examples, ISE120 and ISE156 are the same sequence. In another example, ISE120 and ISE156 are different sequences.

[0188] 3' to the second dimerization domain 154 (and ISE 156, if present) is a branchpoint sequence 158 (e.g., a branchpoint consensus sequence), a polypyrimidine tract 160, followed by a splice acceptor sequence 162 (e.g., a splice acceptor consensus sequence). The sequence of the branchpoint 158 ​​is based on the consensus sequence of the species of the target cell or organism. For example, for human splicing, the consensus sequence can include or be YUNAY. Thus, the sequence used can be CUAAC for a U2-dependent intron, or UUUUCCUUAACU (SEQ ID NO: 144) for a U12-dependent intron.

[0189] The polypyrimidine tract 160 can contain C, U, or both C and U nucleotides, e.g., CnUy (where n+y is greater than or equal to 10 nucleotides), and can also include nucleotides -3 to -22 relative to the 3' splice junction. In some embodiments, the polypyrimidine tract 160 contains at least 80% Y nucleotides (i.e., U, C, or both U and C). In some embodiments, the polypyrimidine tract 160 is a poly-C or poly-U sequence. In some embodiments, the polypyrimidine tract 160 is a poly-U sequence with at least 15 Us, e.g., 15-30 or 15-20 Us. The branch point 158 ​​and the polypyrimidine tract 160 are essential splicing components. The sequence of SA 162 can be based on a consensus sequence for the species of the target cell or organism. For example, in humans, the SA sequence is AG at positions -1 and -2 relative to the 3' splice site for U2-dependent introns, and can be AC ​​or AG for U12-dependent introns. Thus, in some examples, SA162 can be 2 nt in length, e.g., AG or AC.

[0190] Immediately following SA162 is an exon sequence containing a DNA sequence encoding a C-terminal portion of target protein 164 with a splice junction at its 5' end. The splice junction at the 5' end of the DNA sequence encoding the C-terminal portion of target protein 164 can match a consensus sequence found in the target cell or organism into which molecule 110, 150 is introduced. In some embodiments, the splice junction can be GA or GU at positions +1 and +2 relative to the 3' splice site for a U2-dependent intron, or GU or AU for a U12-dependent intron. Thus, in some embodiments, the splice junction is 2 nt in length, and the 5' end of C-terminal coding portion 164 is GA, GU, or AU.

[0191] The exon sequence following intron portion 170 of molecule 150 includes a second coding portion (e.g., half) of the target protein, e.g., a C-terminal fragment 164, and an optional polyadenylation sequence 166. Thus, molecule 150 includes sequence 164 encoding the C-terminal portion of the target protein. The 3' end of molecule 150 optionally includes polyadenylation sequence 166, which promotes spliceosome assembly. In some examples, polyadenylation sequence 166 is a polyA sequence of at least 15 A, e.g., 15-30 or 15-20 A. In some examples, polyadenylation sequence 166 and polyadenylation sequence 124 are the same sequence. In other examples, polyadenylation sequence 166 and polyadenylation sequence 124 are different sequences.

[0192] In some examples, N-terminal coding region 114 and / or C-terminal coding region 164 are native coding sequences. For example, the coding sequences are those found in the cell or organism into which the systems disclosed herein are introduced (e.g., human coding sequences in the case of introduction into a human cell or subject). In some examples, N-terminal coding region 114 and / or C-terminal coding region 164 are codon-optimized relative to the native coding sequence, e.g., to maximize tRNA availability or to de-enrich cryptic splice sites (e.g., to reduce or avoid incorrect splicing and promote correct junction formation). In some examples, portions of N-terminal coding region 114 and / or C-terminal coding region 164 are codon-optimized relative to the native coding sequence, e.g., approximately 200 nt adjacent to each junction (e.g., the 3' end of 114 and the 5' end of 164) can be codon-optimized or modified to contain an exonic splice enhancer site (ESE) (which binds an SR protein). For example, the coding sequence can be one that is not found in the cell or organism into which the system disclosed herein is introduced (e.g., a human coding sequence in the case of introduction into a mouse cell or subject).

[0193] In some examples, N-terminal coding region 114 and / or C-terminal coding region 164 include introns, either natural or synthetic in nature, containing both splice donor and acceptor sites. For example, an intron embedded within the coding sequence to be expressed can be included upstream of sequence 116 (e.g., about 200 nt upstream), within N-terminal coding region 114, or an intron embedded within the coding sequence to be expressed downstream of sequence 162 (e.g., about 200 nt downstream) and within C-terminal coding region 164, or both. The inclusion of such introns can be used to stimulate attachment of the splicing machinery to the trans-splicing intron donor and acceptor. In some examples, such (stimulatory) introns can be derived from the host expressing 110 and 150. In some examples, such (stimulatory) introns can be derived from other organisms, or from viral or synthetic sources.

[0194] In some examples, inclusion of a sequence to stabilize molecule 150 (e.g., located between 164 and 166 in the 3' untranslated region of 150 in Figure 6A) can increase the efficiency of expression of the recombinant product by at least 25%, at least 30%, at least 40%, at least 50%, at least 60%, or at least 75%, e.g., 25-95%, 25-75%, 25-60%, 25-50%, 40-95%, 40-60%, or 50-60%. In some examples, a woodchuck posttranscriptional regulatory element (WPRE) or a truncated version thereof (e.g., WPRE3) is included in the 3'-UTR as a stabilizing element to enhance the efficiency of expression of the recombinant product. In some examples, the WPRE sequence has at least 80%, at least 85%, at least 90%, at least 95%, or 100% sequence identity to nt 1093-1684 of GenBank Accession No. J04514 or the 247 bp sequence of WPRE3.

[0195] As shown in Figure 6C, the interaction and hybridization (base pairing) of first dimerization domain 122 of molecule 110 and second dimerization domain 154 of molecule 150 allows the spliceosome components to recombine N-terminal coding sequence 114 and C-terminal coding sequence 164. Specifically, the 3' end of N-terminal protein coding sequence 114 fuses with the 5' end of C-terminal protein sequence 164 as a seamless junction between the two moieties.

[0196] Figure 6D shows a schematic diagram of a system in which a target protein is divided into three portions: an N-terminal portion, a middle portion, and a C-terminal portion (each portion may be of similar or different size). Thus, one skilled in the art will understand that a protein can be divided into any number of desired segments or portions, and a reasonable number of molecules can be designed using the information provided herein. In such an example, the system includes at least three synthetic nucleic acid molecules 110, 200, and 150, where molecule 110 includes molecule 114 encoding the N-terminal portion of the protein, molecule 200 includes molecule 216 encoding the middle portion of the protein, and molecule 150 includes molecule 164 encoding the C-terminal portion of the protein. Each nucleic acid molecule 110, 200, 150 may be composed of DNA and, after translation, may be RNA without the promoters 112, 202, 152 present. In some examples, molecules 110, 200, 150 (with or without promoters 112, 202, 152) are each at least about 100 nucleotides / ribonucleotides (nt) in length, e.g., at least 200 nt, at least 300 nt, at least 500 nt, at least 1000 nt, at least 2000 nt, at least 3000 nt, at least 4000 nt, at least 5000 nt, at least 6000 nt, at least 7000 nt, or at least 8000 nt, e.g., 200-10,000 nt, 200-8000 nt, 500-5000 nt, or 200-1000 nt. Molecules 110, 150, 200 (with or without promoters 112, 202, 152) can include natural and / or non-natural nucleotides or ribonucleotides. In addition to using two (or more) orthogonal dimerization domains, one of the two introns can be a U2-type intron and the second intron can be a U12-type intron. The splice donors and acceptors of U2- and U12-dependent introns show minimal cross-reactivity because the consensus recognition sequences differ between the two types of introns.Both strategies (i.e., orthogonal dimerization domains and U2-type introns versus U12-type introns) promote the recombination of the three fragments in the correct order (e.g., preventing the first fragment from joining directly with the last fragment and preventing the middle fragment from circularizing on itself).

[0197] 6D includes the same features as disclosed above with respect to FIG. 1A, i.e., a promoter 112 operably linked to a sequence encoding an RNA molecule, the RNA molecule including, from 5' to 3': a coding sequence for the N-terminal portion of the target protein 114 including a splice junction at the 3' end of the target protein coding sequence, SD 116, optional DISE 118, optional ISE 120, a dimerization domain 122, and an optional polyadenylation sequence 124, wherein first dimerization domain 122 has reverse complementarity to third dimerization domain 204 of molecule 200. In embodiments where molecule 110 is RNA, e.g., after expression of DNA into RNA, as shown in FIG. 6F, molecule 110 does not include promoter 112, and 114 is RNA encoded by a coding sequence for the N-terminal portion of the target protein. The molecule 110 (with or without the promoter 112) can include natural and / or non-natural nucleotides or ribonucleotides.

[0198] Molecule 150 of Figure 6D includes the same features as those disclosed above with respect to Figure 1A, i.e., a promoter 152 operably linked to a sequence encoding an RNA molecule, which includes, from 5' to 3': a second dimerization domain 154, an optional ISE 156, a branch point sequence 158, a polypyrimidine tract 160, a splice acceptor (SA) 162; and a coding sequence for a C-terminal portion of a target protein 164, including a splice junction at the 5' end of the target protein coding sequence, and an optional polyadenylation sequence 166. Second dimerization domain 154 has reverse complementarity to fourth dimerization domain 226 of molecule 200. Molecule 150 (with or without promoter 152) can include natural and / or non-natural nucleotides or ribonucleotides.

[0199] Molecule 200 allows for the joining of N-terminal coding region 114 and C-terminal coding region 164 by providing a dimerization domain that is reverse complementary to dimerization domains 122, 154 of molecules 110 and 150, respectively. Molecule 200 includes features of both molecules 110 and 150, including two intron sequences 230, 240. Specifically, in embodiments in which molecule 200 is DNA, molecule 220 includes promoter 210 (which may be the same as or different from promoters 112 and / or 152) operably linked to a sequence encoding an RNA molecule, the RNA molecule including, from 5' to 3': third dimerization domain 204 (which is the reverse complement of first dimerization domain 122 of molecule 110 in FIG. 6D ), optional ISE 206, branch point 208, polypyrimidine tract 210, SA 212, targeting target nucleotide sequence 214, and / or targeting target sequence 216. 6D , molecule 150 includes a target protein intermediate portion coding sequence 216, which includes a splice junction at the 5' end of the target protein coding sequence and a splice junction at the 3' end of the target protein coding sequence, SD 220, optional DISE 222, optional ISE 224, fourth dimerization domain 226 (which is the reverse complement of fourth dimerization domain 154 of molecule 150 of FIG. 6D ), and optional polyadenylation sequence 228. In some embodiments, molecule 220 is DNA and is at least 200 nt, at least 300 nt, at least 500 nt, at least 1000 nt, at least 2000 nt, at least 3000 nt, at least 4000 nt, at least 5000 nt, at least 6000 nt, at least 7000 nt, or at least 8000 nt, e.g., 200-10,000 nt, 200-8000 nt, 500-5000 nt, or 200-1000 nt in length. In embodiments where molecule 200 is RNA, e.g., after expression of DNA into RNA, molecule 200 no longer comprises promoter 202, and 216 is RNA encoded by a coding sequence for the middle portion of the target protein.In some examples, molecule 200 is RNA, does not include promoter 202, and is at least 200 nt, at least 300 nt, at least 500 nt, at least 1000 nt, at least 2000 nt, at least 3000 nt, at least 4000 nt, at least 5000 nt, at least 6000 nt, at least 7000 nt, or at least 8000 nt, e.g., 200-10,000 nt, 200-8000 nt, 500-5000 nt, or 200-1000 nt in length. Molecule 200 (with or without promoter 202) can include natural and / or non-natural nucleotides or ribonucleotides.

[0200] 6E, the interaction and hybridization (base pairing) between first dimerization domain 122 of molecule 110 and third dimerization domain 204 of molecule 200, and the interaction and hybridization (base pairing) between fourth dimerization domain 226 of molecule 200 and second dimerization domain 154 of molecule 150, allow the spliceosome components to recombine N-terminal coding sequence 114, intermediate coding sequence 216, and C-terminal coding sequence 164. Specifically, the 3' end of N-terminal protein coding sequence 114 is fused to the 5' end of intermediate protein sequence 216, and the 3' end of intermediate protein sequence 216 is fused to the 5' end of C-terminal protein sequence 164, forming a seamless junction between the three portions.

[0201] Alternative dimerization domains are shown in Figures 7A-7B and 9A. That is, as an alternative to using hybridizable dimerization domains (e.g., 112-204, 226-154, Figures 6D and 6E), in one embodiment, an aptamer sequence is used. As shown in Figure 7A, in both synthetic nucleic acid molecules 500, 600, aptamer sequences 512, 602 are used in place of dimerization domains, and the aptamers join together through their interaction with a target (e.g., adenosine, dopamine, or caffeine). In such an embodiment, the aptamer sequences 512, 602 in each molecule 500, 600 can be the same sequence, or even different sequences. Molecule 500 of Figure 7A includes the same features as disclosed above with respect to molecule 110 of Figure 6A, and, if DNA, includes a promoter operably linked to a sequence encoding an RNA molecule, the RNA molecule including, from 5' to 3': a coding sequence for the N-terminal portion of the target protein 502 including a splice junction at the 3' end of the target protein coding sequence, SD 506, optional DISE 508, optional ISE 510, a first aptamer 512 in place of the first dimerization domain, and an optional polyadenylation sequence. In embodiments where molecule 500 is RNA, for example, when transcribed from a DNA molecule, molecule 500 does not include a promoter (e.g., as shown in Figure 7A). Similarly, molecule 600 of Figure 7A includes the same features as disclosed above with respect to molecule 150 of Figure 6A, and if it is DNA, includes a promoter operably linked to a sequence encoding an RNA molecule, which includes, from 5' to 3': an aptamer 602 in place of second dimerization domain 154, an optional ISE 604, a branch point 606, a polypyrimidine tract 608, an SA 610, DNA encoding a C-terminal portion 614 of a target protein having a splice junction at the 5' end, and an optional polyadenylation sequence 616. In embodiments where molecule 600 is RNA, for example, when transcribed from a DNA molecule, molecule 500 does not include a promoter (e.g., as shown in Figure 7A).Interaction of the two aptamers 512, 602 with each other or with molecule 700 allows spliceosome components to recombine the N-terminal coding sequence 502 and the C-terminal coding sequence 614. Specifically, the 3' end of the N-terminal protein coding sequence 502 is fused to the 5' end of the C-terminal protein sequence 614 as a seamless junction between the two moieties. Molecules 500 and 600 can contain natural and / or unnatural nucleotides or ribonucleotides.

[0202] In some examples, the aptamer sequences 512, 602 may recognize (e.g., specifically bind to) the same target 700 (FIG. 7A), or even recognize different targets (in which case each molecule specifically recognized by each aptamer, such as a caffeine / dopamine hybrid molecule, or a synthetic molecule containing a portion of a molecule recognized by an aptamer, is also administered with the systems presented herein). Exemplary targets recognized by aptamers include cellular proteins, small molecules, exogenous proteins, or RNA molecules.

[0203] Figure 7B shows an example similar to Figure 7A. The dimerization domains (512, 602 in Figure 7A) recognize RNA molecules. In the example shown in Figure 7B, each domain recognizes a different portion of an mRNA molecule, e.g., a cancer-specific transcript, that is expressed only in target cells (cells in which expression of the target protein is desired). In such an example, the RNA coding sequence (502, 614 in Figure 7A) recombines only in the presence of the specific RNA molecule recognized by the dimerization domain. Here, the target protein is expressed only in cancer cells and not in normal cells. Such a system allows for controlled expression of a target protein (e.g., a cancer therapeutic protein, e.g., a toxin or a cytotoxic enzyme such as thymidine kinase and ganciclovir; thus, in some examples, the target protein is a toxin or thymidine kinase) in cancer cells while reducing undesirable side effects of target protein expression in normal, non-cancer cells.

[0204] 7C presents an illustrative "off switch" example. Here, hybridization / binding of dimerization domains 812, 902 (which are reverse complements to each other) of synthetic nucleic acid molecules 800, 900 can be reduced by providing anti-binding domain oligonucleotides (e.g., RNA or DNA) 1000 (which can be two different anti-binding domain oligonucleotides 1000, one the reverse complement of 812 and one the reverse complement of 912) that compete for binding / hybridization. Thus, anti-binding domain oligonucleotides 1000 can act as an "off switch" for the reconstitution of proteins encoded by N-terminal and C-terminal encoding portions 802 and 914, respectively. Molecule 800 of Figure 7C includes the same features as disclosed above with respect to molecule 110 of Figure 6A, which is an RNA molecule (and thus lacks a promoter), comprising, from 5' to 3': a coding sequence for the N-terminal portion of the target protein 802, including a splice junction at the 3' end of the target protein coding sequence, SD 806, optional DISE 808, optional ISE 810, dimerization domain 812, and optional polyadenylation sequence 814. Similarly, molecule 900 of Figure 7C includes the same features as disclosed above with respect to molecule 150 of Figure 6A, which is an RNA molecule (and thus lacks a promoter), comprising, from 5' to 3': an anti-dimerization domain 902, optional ISE 904, branch point 906, polypyrimidine tract 908, SA 910, RNA 914 encoding the C-terminal portion of the target protein, and optional polyadenylation sequence 916. The two dimerization domains 812, 902 are unable to interact / hybridize with each other in the presence of the anti-binding domain oligonucleotide 1000, thus preventing or reducing recombination of the N-terminal coding sequence 802 and the C-terminal coding sequence 914. Such application can be used to reduce or eliminate expression of the protein encoded by the system. Molecules 800 and 900 can contain natural and / or non-natural nucleotides or ribonucleotides.

[0205] Figure 9A shows an exemplary dimerization domain that uses kissing loop interactions instead of reverse-complementary hybridization for dimerization. Kissing loop interactions are formed when bases within the loops of two RNA hairpins form an interacting pair between two RNA molecules. The molecule on the left, labeled n-yfp, represents an RNA molecule encoding the N-terminal fragment of yfp linked to a synthetic intron containing a splice donor site, a downstream intron splicing enhancer element, and two intron splicing enhancer elements. The dimerization domain of this molecule contains three RNA hairpin loops, each consisting of a stem (where the RNA hybridizes to itself) and a loop (where the RNA does not hybridize to itself). In this example, the dimerization domain contains three stem and loop elements (also referred to as hairpin loops) and is referred to as a trimodal kissing loop dimerization domain. The molecule on the right, labeled c-yfp, represents an RNA molecule encoding the C-terminal portion of yfp. This molecule consists of a trimodal kissing loop dimerization domain containing a set of three hairpin loops from 5' to 3'. The loops can form kissing loop interactions with corresponding loops on a complementary n-yfp molecule. The trimodal kissing loop dimerization domain is followed by a synthetic intron sequence containing three intron splicing enhancer sequences, a branch point sequence, a polypyrimidine tract, and a splice acceptor site. The synthetic intron sequence is followed by a C-terminal yfp coding sequence, followed by a 3' untranslated region containing a polyadenylation signal. A representative three-dimensional rendering of the kissing loop interaction is shown at the top of the figure. This rendering illustrates how the twisted conformation of the hairpin loop exposes the loop residues on the outside, making them available for kissing loop interactions.

[0206] Once the two molecules bind, a trans-splicing reaction is mediated by the spliceosome, resulting in the joining of the N-terminal and C-terminal ypf coding sequences, which then allows for the expression of the full-length fluorescent protein.

[0207] Although Figures 6A-7C and 9A illustrate an embodiment in which the system uses two synthetic nucleic acid molecules (i.e., the target protein coding sequence is split between two synthetic nucleic acid molecules), one of skill in the art will understand that such an embodiment can similarly be used with more than two synthetic nucleic acid molecules, e.g., 3, 4, 5, 6, 7, 8, 9, or 10 synthetic nucleic acid molecules, using the teachings herein.

[0208] In some embodiments, the system includes a nucleic acid molecule that suppresses expression of the unassembled / unrecombined fragment. In such embodiments, if two or more portions of the full-length coding sequence (e.g., 110 and 114, and 150 and 164, respectively, in FIG. 6A) are unrecombined, the nucleic acid molecule suppresses full-length protein expression of each portion of the full-length coding sequence that was not recombined. For example, such an inhibitory nucleic acid molecule can destabilize the RNA once it leaves the nucleus, prevent translation, stimulate translation from a shifted start codon, contain a microRNA target site, or contain a protein degron or destabilization domain that, once translated, suppresses protein activity or flags the protein for degradation.

[0209] In one example, destabilization of the unmodified RNA molecule is achieved by including a self-cleaving RNA sequence (e.g., a hammerhead ribozyme or HDV ribozyme) into a synthetic intron, for example, at any position within intron sequence 130 of Figure 6A or 6F. In one example, cleavage of the RNA molecule results in loss of the polyA tail that stabilizes the RNA, which can prevent expression of the unmodified protein from open reading frame 114 of Figure 6A or 6F. In one example, a self-cleaving RNA sequence is included at any position within intron sequence 170 of Figure 6A or 6F to cleave the 5'-terminal CAP, which can result in reduced expression of an open reading frame containing part or all of coding sequence 164 of Figure 6A or 6F. In one example, the self-cleaving RNA sequence is replaced with an RNA cleavage enzyme target site, such as a Csy4 target site.

[0210] In some examples, the inhibitory nucleic acid molecule contains a start codon (ATG) or a Kozak-enhanced start codon (GCCGCCACCATG (SEQ ID NO: 154) or GCCACCATG or ACCATG) anywhere within intron sequence 170 of Figure 6A or 6F that directs translation of an open reading frame shifted -1, -2, +1, or +2 nucleotides relative to open reading frame sequence 164 of Figure 6A or 6F. In one example, this decoy start codon strategy is used to reduce or suppress expression of unassembled fragments by directing translation away from the open reading frame of sequence 164 of Figure 6A or 6F to be suppressed.

[0211] In some examples, the inhibitory nucleic acid molecule includes one or more microRNA target sites anywhere within intron sequence 130 of FIG. 6A or 6F and / or anywhere within intron sequence 170 of FIG. 6A or 6F. When a particular molecule (e.g., 110 or 150 of FIG. 6A or 6F) is exported from the nucleus, it undergoes microRNA / short hairpin RNA-dependent degradation, thereby degrading / silencing the unbound RNA exported from the nucleus and thereby suppressing expression of unintended, unbound fragments. In one example, such microRNA target sequences can be complementary to microRNAs known to be expressed in cells, tissues, or animals into which molecules 110 and 150 of FIG. 6A or 6F are introduced. In one example, the microRNA target sequences are complementary to sequences introduced into the cells, tissues, or animals. In one example, such microRNAs can be expressed in the form of short hairpin RNAs from an RNA polymerase III-dependent promoter. In one example, such microRNAs can be expressed from an RNA polymerase II-dependent promoter and embedded in a microRNA processing loop (eg, a mir30 scaffold).

[0212] In some examples, destabilization of unrecombined protein products from an open reading frame (e.g., 114 in Figure 6) can be achieved by eliminating occurrences of stop codons in intron sequence 130 of Figure 6A or 6F and additionally including an RNA sequence (e.g., a degron sequence) located anywhere within intron sequence 130 of Figure 6A or 6F and in frame with the open reading frame extending outward from sequence 114 of Figure 6A or 6F that encodes an in-frame protein signal capable of flagging the protein for degradation. In one example, the degron sequence can be a PEST sequence or a CL1 degron sequence. The degron sequence used can use proteasome-dependent, proteasome-independent, ubiquitin-dependent, or ubiquitin-independent pathways. In one example, destabilization of unrecombined proteins is enhanced by including several of the same or different degron sequences.

[0213] In some embodiments, destabilization of the unrecombined protein product from open reading frame sequence 164 of Figure 6A is achieved by introducing an in-frame start codon (ATG) with the open reading frame in sequence 164 of Figure 6 followed by a degron sequence anywhere in intron sequence 170 of Figure 6A. In this example, the degron sequence binds to the unrecombined protein fragment at its N-terminus, flagging the unrecombined protein fragment for degradation, thereby inhibiting it.

[0214] IV. Compositions and Kits Compositions and kits are provided that include two or more synthetic nucleic acid molecules provided herein, wherein the synthetic nucleic acid molecules encode a full-length protein when recombined. In some examples, the two or more synthetic nucleic acid molecules provided herein are DNA. In some examples, the two or more synthetic nucleic acid molecules provided herein are RNA and do not include a promoter sequence. In one example, a composition or kit includes two synthetic nucleic acid molecules provided herein, wherein each of the two synthetic nucleic acid molecules encodes a different portion (i.e., N-terminus and C-terminus) of a target protein, such as those listed in Table 1 (or a therapeutic protein such as a toxin or thymidine kinase), where recombination between the two molecules generates the entire coding sequence. In one example, a composition or kit includes three synthetic nucleic acid molecules provided herein, wherein each of the three synthetic nucleic acid molecules encodes a different portion (i.e., N-terminus, middle, and C-terminus) of a target protein, such as those listed in Table 1 (or a therapeutic protein such as a toxin or thymidine kinase), where recombination between the three molecules generates the entire coding sequence. In one example, a composition or kit includes four or more synthetic nucleic acid molecules provided herein, each encoding a different portion of a target protein (i.e., the N-terminus, first middle, second middle (and optionally additional middle), and C-terminus, where recombination between the four or more synthetic nucleic acid molecules generates the entire coding sequence), such as those listed in Table 1 (or therapeutic proteins such as toxins or thymidine kinase). In one example, a composition or kit includes two or more sets of two or more synthetic nucleic acid molecules provided herein, where each set of synthetic nucleic acid molecules encodes a different target protein, e.g., two or more of those listed in Table 1 (and / or therapeutic proteins such as toxins or thymidine kinase).

[0215] In one example, each synthetic nucleic acid molecule in the composition or kit is part of a vector, such as an AAV or other gene therapy vector. In one example, the composition or kit includes cells, e.g., bacterial or eukaryotic cells, containing two or more of the disclosed synthetic nucleic acid molecules, where the synthetic nucleic acid molecules, when recombined, encode a full-length target protein.

[0216] Such compositions may include a pharmaceutically acceptable carrier (e.g., saline, water, glycerol, DMSO, or PBS). In some examples, the composition is a liquid, a lyophilized powder, or is frozen.

[0217] In some examples, the kit includes a delivery system (e.g., liposomes, particles, exosomes, or microvesicles), e.g., to direct cell-type specific uptake / enhance endosomal escape / enable crossing of the blood-brain barrier, etc. In some examples, the kit further includes a cell culture or growth medium, such as a suitable medium for growing bacterial cells, plant cells, insect cells, or mammalian cells. In some examples, such parts of the kit are contained in separate containers. Exemplary containers include plastic or glass vials or tubes.

[0218] In some examples, two or more synthetic nucleic acid molecules provided herein are each contained in a separate container, hi some examples, two or more sets of two or more synthetic nucleic acid molecules provided herein are contained in separate containers.

[0219] V. Treatment Method The methods and systems disclosed herein can be used to express any protein of interest, for example, when the protein is too large to be expressed by a therapeutic virus (e.g., AAV) or when the complete gene sequence (e.g., endogenous promoter + coding sequence) is too large to be expressed by a therapeutic virus (e.g., AAV). In such cases, the systems disclosed herein can be used to separate the coding sequence for the target protein into two or more parts and recombine them in the correct order, thereby allowing the protein to be expressed when and where desired.

[0220] The subject to be treated can be any mammal, such as one with a monogenetic disorder, such as those listed in Table 1. In one example, the subject has cancer. Thus, humans, cats, pigs, rats, mice, cows, goats, and dogs can be treated with the disclosed methods. In some examples, the subject is a human infant under 6 months of age. In some examples, the subject is a human infant under 1 year of age. In some examples, the subject is a young human. In some examples, the subject is a human adult at least 18 years of age. In some examples, the subject is a female. In some examples, the subject is a male.

[0221] Two or more synthetic nucleic acid molecules presented herein can be used to treat subject and adapted to the subject to be treated.Therefore, for example, when the subject to be treated is dog, can use the dog coding sequence of target protein, and intron sequence can be optimized for the expression in dog cell; when the subject to be treated is human, can use the human coding sequence of target protein, and intron sequence can be optimized for the expression in human cell.

[0222] Two or more synthetic nucleic acid molecules presented herein can be administered as part of a vector, such as an adeno-associated vector (AAV), e.g., AAV serotype rh.10. In some examples, a vector (e.g., an AAV) containing one of two or more synthetic nucleic acid molecules presented herein is administered systemically, such as intravenously. Thus, when the coding sequence is separated into two synthetic nucleic acid molecules presented herein, two AAVs are administered, each containing one of the two synthetic nucleic acid molecules presented herein.

[0223] Two or more synthetic nucleic acid molecules provided herein are administered in a therapeutically effective amount, e.g., as an AAV. In some examples, two or more synthetic nucleic acid molecules provided herein, when part of a viral vector (e.g., an AAV), are administered in a therapeutically effective amount of at least 1 x 10 per subject. 11 Genome copies (gc), at least 1 × 10 12 gc, at least 2 × 10 12 gc, at least 1 × 10 13 gc, at least 2 × 10 13 gc, or at least 1 × 10 per subject 14 gc, e.g., 2 x 10 per subject 11 gc, 2 × 10 per subject 12 gc, 2 × 10 per subject 13 gc, or 2 x 10 per subject 14 In some examples, two or more synthetic nucleic acid molecules provided herein, when part of a viral vector (e.g., AAV), are administered at a dose of at least 1 x 10 g. 11 gc / kg, at least 5 × 10 11 gc / kg, at least 1 × 10 12 gc / kg, at least 5 × 10 12 gc / kg, at least 1 × 10 13 gc / kg, or at least 4 × 10 13 gc / kg, e.g., 4 x 10 11 gc / kg, 4 × 10 12 gc / kg, or 4 × 10 13Administer at a dose of gc / kg.

[0224] If adverse symptoms occur, such as AAV capsid-specific T cells in the blood, corticosteroids can be administered (see, e.g., Nathwani et al., N Engl J Med. 365 (25): 2357-65, 2011).

[0225] Diseases that can be treated using the methods disclosed herein include any blood genetic disease (e.g., sickle cell disease, primary immunodeficiency disease), HIV (e.g., HIV-1), and hematological malignancies or cancers. Examples of primary immunodeficiency diseases and their corresponding mutations include those listed in Al-Herz et al. (Frontiers in Immunology, volume 5, article 162, April 22, 2014, incorporated herein by reference in its entirety). Hematological malignancies or cancers are tumors that affect the blood, bone marrow, and lymph nodes. Examples include leukemia (e.g., acute lymphoblastic leukemia, acute myeloid leukemia, chronic lymphocytic leukemia, chronic myeloid leukemia, acute monocytic leukemia), lymphoma (e.g., Hodgkin's lymphoma and non-Hodgkin's lymphoma), and myeloma. In some examples, the disease is a monogenic disease. Table 1 provides a list of exemplary disorders and genes that can be targeted by the systems and methods disclosed herein. Further examples are provided at rarediseases.info.nih.gov / diseases / diseases-by-category / 5 / congenital-and-genetic-diseases (the list is incorporated herein by reference). The systems and methods disclosed herein can be beneficial for any genetic disease caused by protein deficiency (e.g., recessive mutation) or protein insufficiency. When the coding region of a gene is relatively small, the systems and methods disclosed herein are useful for adding regulatory sequences, such as tissue-specific promoters or specific non-coding RNA segments, to direct gene expression to the appropriate cell type at the appropriate level. [Table 1-1] [Table 1-2] [Table 1-3] [Table 1-4] [Table 1-5]

[0226] The methods and systems disclosed herein can be used to treat any of the disorders listed in Table 1 or other known genetic disorders. The methods disclosed herein can also be used to treat other disorders, such as cancer, in which expressing a toxin or a therapeutic protein, such as thymidine kinase, in cancer cells can be beneficial. When administering two or more synthetic molecules presented herein that express full-length thymidine kinase to a subject, the subject is also administered ganciclovir. Treatment need not 100% eliminate all features of the disorder, but can reduce them. Specific examples are presented below, but it will be understood that based on this teaching, symptoms of other disorders can be similarly affected. For example, the methods disclosed herein can be used to increase the expression of a protein that is not expressed or has reduced expression by a subject, or to reduce the expression of a protein that is undesirably expressed or has reduced expression by a subject. For example, the methods disclosed herein can be used to treat or reduce the undesirable effects of a genetic disease.

[0227] For example, the methods and systems disclosed herein can treat or reduce the undesirable effects of sickle cell disease by expressing a full-length wild-type β-globin chain of hemoglobin. In one example, the methods disclosed herein reduce symptoms of sickle cell disease (e.g., one or more of the following: presence of sickle red blood cells in the blood, pain, ischemia, necrosis, anemia, vaso-occlusive crisis, aplastic crisis, splenic sequestration crisis, and hemolytic crisis) in a recipient subject, e.g., by at least 10%, at least 20%, at least 50%, at least 70%, or at least 90% (compared to not administering a therapeutic nucleic acid molecule). In one example, the methods disclosed herein reduce the number of sickle red blood cells in a recipient subject, e.g., by at least 10%, at least 20%, at least 50%, at least 70%, at least 90%, or at least 95% (compared to not administering a therapeutic nucleic acid molecule).

[0228] For example, the methods and systems disclosed herein can treat or reduce the undesirable effects of thrombophilia by expressing full-length wild-type Factor V Leiden or prothrombin genes. In one example, the methods disclosed herein reduce symptoms of thrombophilia (e.g., one or more of thrombosis, such as deep vein thrombosis, pulmonary embolism, venous thromboembolism, swelling, chest pain, palpitations) in a recipient subject, e.g., by at least 10%, at least 20%, at least 50%, at least 70%, or at least 90% (compared to not administering a therapeutic nucleic acid molecule). In one example, the methods disclosed herein reduce the activity of a coagulation factor in a recipient subject, e.g., by at least 10%, at least 20%, at least 50%, at least 70%, at least 90%, or at least 95% (compared to not administering a therapeutic nucleic acid molecule).

[0229] For example, the methods and systems disclosed herein can treat or reduce the undesirable effects of CD40 ligand deficiency by expressing a full-length wild-type CD40 ligand gene. In one example, the methods disclosed herein reduce symptoms of CD40 ligand deficiency (e.g., one or more of elevated serum IgM, low serum levels of other immunoglobulins, opportunistic infections, autoimmunity, and malignancies) in a recipient subject, e.g., by at least 10%, at least 20%, at least 50%, at least 70%, or at least 90% (compared to when the therapeutic nucleic acid molecule is not administered). In one example, the methods disclosed herein increase the amount or activity of CD40 ligand deficiency in a recipient subject, e.g., by at least 10%, at least 20%, at least 50%, at least 70%, at least 90%, at least 100%, at least 200%, or at least 500% (compared to when the therapeutic nucleic acid molecule is not administered).

[0230] For example, the methods disclosed herein can be used to treat or reduce the undesirable effects of primary immunodeficiency diseases caused by genetic defects. For example, the methods and systems disclosed herein (e.g., using AAV, two or more synthetic nucleic acid molecules can be used to express missing or defective functional proteins in a subject) can treat or reduce the undesirable effects of primary immunodeficiency diseases. In one example, the methods disclosed herein reduce the symptoms of primary immunodeficiency diseases in recipient subjects (e.g., one or more of bacterial infection, fungal infection, viral infection, parasitic infection, lymphadenopathy, splenomegaly, wounds, and weight loss), for example, by at least 10%, at least 20%, at least 50%, at least 70%, or at least 90% (compared to when not administering a therapeutic nucleic acid molecule). In one example, the methods disclosed herein result in an increase in the number of immune cells (e.g., T cells, such as CD8 cells) in a recipient subject having a primary immunodeficiency disorder (compared to not administering the therapeutic nucleic acid molecule), e.g., an increase of at least 10%, at least 20%, at least 50%, at least 70%, at least 90%, at least 95%, at least 100%, at least 200%, at least 300%, at least 400%, or at least 500%. In one example, the methods disclosed herein result in a decrease in the number of infections (e.g., bacterial, viral, fungal, or a combination thereof) in a recipient subject having a primary immunodeficiency disorder over a period of time (e.g., more than one year), e.g., a decrease of at least 10%, at least 20%, at least 50%, at least 70%, at least 90%, or at least 95% (compared to not administering the therapeutic nucleic acid molecule).

[0231] For example, the methods disclosed herein can be used to treat or reduce the undesirable effects of single gene disorders. For example, the methods disclosed herein (e.g., using AAV to express a functional protein that is missing or defective in a subject using two or more synthetic nucleic acid molecules) can treat or reduce the undesirable effects of single gene disorders. In one example, the methods disclosed herein reduce the symptoms of the single gene disorder in a recipient subject, for example, by at least 10%, at least 20%, at least 50%, at least 70%, or at least 90% (compared to not administering a therapeutic nucleic acid molecule). In one example, the methods disclosed herein increase the amount of a normal protein that is not normally expressed in a recipient subject with a single gene disorder, for example, by at least 10%, at least 20%, at least 50%, at least 70%, at least 90%, at least 95%, at least 100%, at least 200%, at least 300%, at least 400%, or at least 500% (compared to not administering a therapeutic nucleic acid molecule).

[0232] For example, the methods disclosed herein can be used to treat or reduce the undesirable effects of a hematological malignancy in a recipient subject. In one example, the methods disclosed herein reduce the number of abnormal white blood cells (e.g., B cells) in a recipient subject (e.g., a subject with leukemia), e.g., by at least 10%, at least 20%, at least 50%, at least 70%, or at least 90% (compared to not administering a treatment disclosed herein). In one example, administration of a treatment disclosed herein can be used to treat or reduce the undesirable effects of lymphoma, e.g., reduce lymphoma size, lymphoma volume, lymphoma growth rate, lymphoma metastasis, e.g., by at least 10%, at least 20%, at least 50%, at least 70%, or at least 90% (compared to not administering a treatment disclosed herein). In one example, administration of a therapy disclosed herein can be used to treat or reduce the undesirable effects of multiple myeloma, e.g., to reduce the number of abnormal plasma cells in a recipient subject, e.g., by at least 10%, at least 20%, at least 50%, at least 70%, or at least 90% (compared to not administering a therapy disclosed herein).

[0233] For example, the methods disclosed herein can be used to treat or reduce the undesirable effects of malignant tumors, such as those caused by genetic defects, in recipient subjects. In one example, the methods disclosed herein reduce the number of cancer cells, tumor size, tumor volume, or number of metastases in a recipient subject (e.g., a subject with a cancer listed herein), for example, by at least 10%, at least 20%, at least 50%, at least 70%, or at least 90% (compared to not administering a treatment disclosed herein). In one example, administering a treatment disclosed herein can be used to treat or reduce the undesirable effects of lymphoma, for example, by reducing tumor size, tumor volume, cancer growth rate, or cancer metastasis, for example, by at least 10%, at least 20%, at least 50%, at least 70%, or at least 90% (compared to not administering a treatment disclosed herein).

[0234] For example, the methods disclosed herein can be used to treat or reduce the undesirable effects of a neurological disease caused by a genetic defect in a recipient subject. In one example, the methods disclosed herein increase neurological function in a recipient subject (e.g., a subject with a neurological disease listed above), e.g., by at least 10%, at least 20%, at least 50%, at least 70%, at least 90%, at least 100%, at least 200%, at least 300%, at least 400%, or at least 500% (compared to without the treatment disclosed herein).

[0235] Treatment of Duchenne Muscular Dystrophy (DMD) Duchenne muscular dystrophy (DMD, MIM:310200) is a fatal genetic disease characterized by progressive muscle weakness and degeneration. As the disease progresses, degenerating muscle fibers are replaced by fatty and fibrotic tissue. DMD originates from a deficiency in the gene dystrophin (MIM:300377). The dystrophin gene spans a 22-kbp region and is prone to mutations. Therefore, DMD can manifest sporadically in some patients without a family history of disease-causing mutations. DMD is one of four conditions known as dystrophinopathies. The other three diseases in this group are Becker muscular dystrophy (BMD, a milder form of DMD); clinical findings intermediate between DMD and BMD; and DMD-associated dilated cardiomyopathy (heart disease) with little or no clinical skeletal or voluntary muscle disease. Thus, in some embodiments, patients with DMD, BMD, clinical findings intermediate between DMD and BMD; or DMD-associated dilated cardiomyopathy (heart disease) with little or no clinical skeletal or voluntary muscle disease are treated using the systems and methods disclosed herein.

[0236] The methods and systems disclosed herein can be used to treat monogenic causes of DMD by expressing dystrophin.Dystrophin, such as dystrophin, has a long coding region.Current methods for expressing dystrophin from a single AAV utilize shortened / truncated versions of dystrophin (microdystrophin and minidystrophin).Some of these shortened dystrophin delivery therapies have been tested in phase I / II clinical trials (NCT03362502, NCT00428935, NCT03368742, NCT03375164).Although these shortened versions of dystrophin can improve the worst outcomes of dystrophin deficiency in DMD, they are expected to lack complete functionality compared to full-length dystrophin, since the shortened versions lack important domains in the rod and hinge regions of the full-length protein. The methods and systems disclosed herein use "multiplexed" AAV combinations, allowing multiple AAV viruses to efficiently infect the same cells when introduced at a high multiplicity of infection (MOI, i.e., high titer), thereby mitigating the size limitation of AAV transgenic payloads.

[0237] Thus, in some examples, a therapeutically effective amount of a composition comprising two or more AAVs, each containing one of a set of disclosed synthetic molecules, e.g., a set comprising two, three, four, or five different synthetic RNA molecules (each as a different AAV) that recombine to generate a full-length dystrophin coding sequence, is administered to a DMD subject (e.g., iv).

[0238] VI. Illustrative Embodiments 1. A system for expressing a target protein, comprising: (a) a first synthetic nucleic acid molecule comprising, 5' to 3', a first promoter operably linked to a sequence encoding an RNA molecule comprising: a coding sequence for an N-terminal portion of the target protein; a splice donor; and a first dimerization domain; and (b) a second synthetic nucleic acid molecule comprising, 5' to 3', a second promoter operably linked to a sequence encoding an RNA molecule comprising: a second dimerization domain that binds to the first dimerization domain; a branch point sequence; a polypyrimidine tract; a splice acceptor; and a coding sequence for the C-terminal portion of the target protein.

[0239] 2. A system for expressing a target protein, comprising: (a) a first synthetic nucleic acid molecule comprising, from 5' to 3': a coding sequence for an N-terminal portion of the target protein; a splice donor; and a coding sequence for an RNA molecule comprising a first dimerization domain; and (b) from 5' to 3': a coding sequence for a second dimerization domain linked to the first dimerization domain; a branch point sequence; a polypyrimidine tract; a splice acceptor; and a coding sequence for an intermediate portion of the target protein; a second synthetic nucleic acid molecule comprising a second promoter operably linked to a sequence encoding an RNA molecule comprising a second splice donor; and a third dimerization domain; and (c) a third synthetic nucleic acid molecule comprising, 5' to 3': a fourth dimerization domain linked to the third dimerization domain; a branch point sequence; a polypyrimidine tract; a splice acceptor; and a coding sequence for the C-terminal portion of a target protein.

[0240] 3. A system for expressing a target protein, comprising: (a) from 5' to 3': a first synthetic nucleic acid molecule comprising a first promoter operably linked to a sequence encoding an RNA molecule comprising a coding sequence for an N-terminal portion of a target protein, a splice donor, and a first dimerization domain; (b) from 5' to 3': a second synthetic nucleic acid molecule comprising a second promoter operably linked to a sequence encoding an RNA molecule comprising a second dimerization domain that binds to the first dimerization domain; a branch point sequence; a polypyrimidine tract; a splice acceptor; and a coding sequence for an intermediate portion of the target protein; a second splice donor; and a third dimerization domain; and (c) from 5' to 3': a third dimerization domain. a fourth dimerization domain that binds to the fifth dimerization domain; a branch point sequence; a polypyrimidine tract; a splice acceptor; and a coding sequence for a first intermediate portion of the target protein; a second splice donor; and a third promoter operably linked to a sequence encoding an RNA molecule that comprises: a fourth dimerization domain that binds to the fifth dimerization domain; a branch point sequence; a polypyrimidine tract; a splice acceptor; and a coding sequence for a C-terminal portion of the target protein; and (c) a fourth synthetic nucleic acid molecule that comprises, 5' to 3': a fourth promoter operably linked to a sequence encoding an RNA molecule that comprises: a sixth dimerization domain that binds to the fifth dimerization domain; a branch point sequence; a polypyrimidine tract; a splice acceptor; and a coding sequence for a C-terminal portion of the target protein.

[0241] 4. The system of any one of embodiments 1 to 3, wherein each promoter is independently selected.

[0242] 5. The system of any one of embodiments 1 to 4, wherein the first promoter and the second promoter are the same promoter; the first promoter and the second promoter are different promoters; the first promoter, the second promoter, and the third promoter are the same promoter; the first promoter, the second promoter, and the third promoter are different promoters; the first promoter, the second promoter, the third promoter, and the fourth promoter are the same promoter; or the first promoter, the second promoter, the third promoter, and the fourth promoter are different promoters.

[0243] 6. The system of any one of embodiments 1 to 5, wherein each of the first promoter, the second promoter, the third promoter, and the fourth promoter is independently selected from a constitutive promoter; a tissue-specific promoter; and a promoter endogenous to the target protein.

[0244] 7. The system of any one of embodiments 1 to 6, wherein the first dimerization domain and the second dimerization domain, the third dimerization domain and the fourth dimerization domain, and / or the fifth dimerization domain and the sixth dimerization domain are linked by a direct bond, an indirect bond, or a combination thereof.

[0245] 8. The composition of claim 7, wherein the direct or indirect bond comprises a base-pairing interaction, a non-canonical base-pairing interaction, a non-base-pairing interaction, or a combination thereof.

[0246] 9. The composition of claim 7 or 8, wherein the direct binding comprises base-pairing interactions between kissing loops or low diversity regions.

[0247] 10. The composition of claim 7 or 8, wherein the direct binding comprises non-canonical base pairing interactions, non-canonical base pairing interactions, non-base pairing interactions, or a combination thereof between the aptamer regions.

[0248] 11. The composition of claim 7 or 8, wherein the indirect linkage comprises a base-pairing interaction through a nucleic acid bridge.

[0249] 12. The composition of claim 7 or 8, wherein the indirect binding comprises a non-base pairing interaction between the aptamer and the target of the aptamer, or between two aptamers.

[0250] 13. The system of any one of embodiments 1 to 12, wherein the first dimerization domain, the second dimerization domain, the third dimerization domain, the fourth dimerization domain, the fifth dimerization domain and / or the sixth dimerization domain does not comprise a cryptic splice acceptor.

[0251] 14. The system of any one of embodiments 1 to 13, comprising at least one pair of directly or indirectly linked aptamer sequence-dimerization domains.

[0252] 15. The system of any one of embodiments 1 to 14, comprising at least one pair of kissing loop interacting dimerization domains.

[0253] 16. The system of any one of embodiments 1 to 15, wherein the target protein is a disease-associated protein or a therapeutic protein.

[0254] 17. The system of embodiment 16, wherein the disease is a monogenic disease.

[0255] 18. The system of embodiment 17, wherein the therapeutic protein is a toxin.

[0256] 19. The system of any one of embodiments 16 to 18, wherein the disease and target protein are those listed in Table 1.

[0257] 20. The system of any one of embodiments 1 to 19, wherein the first synthetic nucleic acid molecule, the second synthetic nucleic acid molecule, the third synthetic nucleic acid molecule, and / or the fourth synthetic nucleic acid molecule further comprises a polyadenylation sequence at the 3' end of the first synthetic nucleic acid molecule, the second synthetic nucleic acid molecule, the third synthetic nucleic acid molecule, or the fourth synthetic nucleic acid molecule.

[0258] 21. The system of any one of embodiments 1 or 4 to 20, wherein the first synthetic nucleic acid molecule further comprises one or both of a downstream intron splice enhancer (DISE) 3' to the splice donor and 5' to the first dimerization domain, an intron splice enhancer (ISE) 3' to the splice donor and 5' to the first dimerization domain; and / or the second synthetic nucleic acid molecule further comprises one or both of an ISE 3' to the second dimerization domain and 5' to the branchpoint sequence, and a DISE 3' to the splice donor and 5' to the dimerization domain; and any combination thereof.

[0259] 22. The system of any one of embodiments 2 or 4 to 20, wherein the first synthetic nucleic acid molecule further comprises an ISE 3' to the first splice donor and 5' to the first dimerization domain, an ISE 3' to the first splice donor and 5' to the first dimerization domain, or both a DISE and an ISE; the second synthetic nucleic acid molecule further comprises an ISE 3' to the second dimerization domain and 5' to the first branch point sequence, a DISE 3' to the second splice donor and 5' to the second dimerization domain, an ISE 3' to the second splice donor and 5' to the third dimerization domain, or a combination thereof; and / or the third synthetic nucleic acid molecule further comprises an ISE 3' to the fourth dimerization domain and 5' to the second branch point sequence; and any combination thereof.

[0260] 23. The first synthetic nucleic acid molecule further comprises a DISE 3' to the first splice donor and 5' to the first dimerization domain, an ISE 3' to the first splice donor and 5' to the first dimerization domain, or both a DISE and an ISE; the second synthetic nucleic acid molecule further comprises an ISE 3' to the second dimerization domain and 5' to the first branch point sequence, a DISE 3' to the second splice donor and 5' to the second dimerization domain, an ISE 3' to the second splice donor and 5' to the third dimerization domain, or any of these. 21. The system of any one of embodiments 3 to 20, further comprising a combination; wherein the third synthetic nucleic acid molecule further comprises an ISE 3' to the fourth dimerization domain and 5' to the second branch point sequence; and / or the fourth synthetic nucleic acid molecule further comprises an ISE 3' to the fifth dimerization domain and 5' to the third branch point sequence, a DISE 3' to the third splice donor and 5' to the fifth dimerization domain, an ISE 3' to the third splice donor and 5' to the sixth dimerization domain, or any combination thereof.

[0261] 24. The system of any one of embodiments 1 to 23, which, when introduced into a cell, causes RNA molecules to be produced in the proper order and recombine, thereby resulting in a full-length coding sequence for the target protein.

[0262] 25. The system of any one of embodiments 1 to 24, wherein each of the first synthetic nucleic acid molecule, the second synthetic nucleic acid molecule, the third synthetic nucleic acid molecule, and the fourth synthetic nucleic acid molecule is part of a separate viral vector.

[0263] 26. The system of embodiment 25, wherein the viral vector is AAV.

[0264] 27. the first synthetic nucleic acid molecule and / or the third synthetic nucleic acid molecule further comprises a self-cleaving RNA sequence or an RNA cleavage enzyme target sequence located anywhere 3′ to the splice donor, such that the sequence cleaves the 3′-located polyadenylation tail, thereby reducing or suppressing expression of the protein fragment from the unrecombined RNA molecule; the second synthetic nucleic acid molecule and / or the fourth synthetic nucleic acid molecule further comprises a self-cleaving RNA sequence or an RNA cleavage enzyme target sequence located anywhere 5′ to the branch point sequence, such that the sequence cleaves the 5′-located RNA cap, thereby reducing or suppressing expression of the protein fragment from the unrecombined RNA molecule; the second synthetic nucleic acid molecule and / or the fourth synthetic nucleic acid molecule further comprises a start codon anywhere 5' to the branchpoint sequence that is shifted relative to the open reading frame that is 3' to the splice acceptor, thereby reducing or preventing translation of the target protein fragment from the unrecombined RNA molecule; the first synthetic nucleic acid molecule and / or the third synthetic nucleic acid molecule further comprises a microRNA target site anywhere 3′ to the splice donor, such that the unbound RNA fragment undergoes microRNA-dependent degradation upon exiting the nucleus; the second synthetic nucleic acid molecule and / or the fourth synthetic nucleic acid molecule further comprises a microRNA target site anywhere 3′ to the coding sequence, such that the unbound RNA fragment undergoes microRNA-dependent degradation upon exiting the nucleus; the first synthetic nucleic acid molecule and / or the third synthetic nucleic acid molecule further comprises a sequence encoding a degron proteolytic tag anywhere 3′ to the splice donor and in-frame with the open reading frame of the target protein 5′ to the splice donor site, such that unbound protein fragments are tagged for degradation; the second synthetic nucleic acid molecule and / or the fourth synthetic nucleic acid molecule further comprises a start codon and an in-frame degron proteolytic tag anywhere 5′ to the branchpoint sequence and in-frame with the open reading frame of the target protein 3′ to the splice acceptor site, such that unbound protein fragments are tagged for degradation; or a combination thereof, The system of any one of embodiments 1 to 26.

[0265] 28. Any one, two, three, or four synthetic nucleic acid molecules of the system each have a length of about 2,500 nt to about 5,000 nt, 2,500 nt to about 2,750 nt, about 2,500 nt to about 3,000 nt, about 2,500 nt to about 3,250 nt, about 2,500 nt to about 3,500 nt, about 2,500 nt to about 3,750 nt, about 2,500 nt to about 4,000 nt, about 2,500 nt to about 4,250 nt, about 2,500 nt to about 4,500 nt, about 2,500 nt to about 4,750 nt, about 2,500 nt to about 5,000 nt, about 2,750 nt to about 3,000 nt, Approximately 2,750nt to approximately 3,250nt, approximately 2,750nt to approximately 3,500nt, approximately 2,750nt to approximately 3,750nt, approximately 2,750nt to approximately 4,000nt, approximately 2,750nt to approximately 4,250nt, approximately 2,750nt to approximately 4,500nt, approximately 2,750nt to approximately 4,750nt , about 2,750nt to about 5,000nt, about 3,000nt to about 3,250nt, about 3,000nt to about 3,500nt, about 3,000nt to about 3,750nt, about 3,000nt to about 4,000nt, about 3,000nt to about 4,250nt, about 3,000nt to about 4,500nt t, about 3,000nt to about 4,750nt, about 3,000nt to about 5,000nt, about 3,250nt to about 3,500nt, about 3,250nt to about 3,750nt, about 3,250nt to about 4,000nt, about 3,250nt to about 4,250nt, about 3,250nt to about 4,500 nt, approximately 3,250nt to approximately 4,750nt, approximately 3,250nt to approximately 5,000nt, approximately 3,500nt to approximately 3,750nt, approximately 3,500nt to approximately 4,000nt, approximately 3,500nt to approximately 4,250nt, approximately 3,500nt to approximately 4,500nt, approximately 3,500nt to approximately 4,75 0nt, about 3,500nt to about 5,000nt, about 3,750nt to about 4,000nt, about 3,750nt to about 4,250nt, about 3,750nt to about 4,500nt, about 3,750nt to about 4,750nt, about 3,750nt to about 5,000nt, about 4,000nt to about 4,2 50nt, approximately 4,000nt to approximately 4,500nt, approximately 4,000nt to approximately 4,750nt, approximately 4,000nt to approximately 5,000nt, approximately 4,250nt to approximately 4,500nt, approximately 4,250nt to approximately 4,750nt, approximately 4,250nt to approximately 5,000nt, approximately 4,500nt to approximately 4,28. The system of any one of embodiments 1 to 27, having a size independently selected from 750 nt, about 4,500 nt to about 5,000 nt, about 4,750 nt to about 5,000 nt, about 2,500 nt, about 2,750 nt, about 3,000 nt, about 3,250 nt, about 3,500 nt, about 3,750 nt, about 4,000 nt, about 4,250 nt, about 4,500 nt, about 4,750 nt, and about 5,000 nt.

[0266] 29. The coding sequence of the N-terminal portion of the target protein, the coding sequence of the middle portion of the target protein, or the coding sequence of the C-terminal portion of the target protein encoded by the synthetic nucleic acid molecule of the system is each about 1,000 nt to about 4,000 nt, about 1,000 nt to about 1,500 nt, about 1,000 nt to about 2,000 nt, about 1,000 nt to about 2,500 nt, about 1,000 nt to about 3,000 nt, about 1,000 nt to about 3,500 nt, about 1,000 nt to about 4,000 nt, about 1,500 nt to about 2,000 nt, about 1,500 nt to about 2,500 nt, about 1,500 nt to about 3,000 nt, about 1,500 nt to about 3,500 nt, about 1,500 nt to about 1,500 nt 29. The system of any one of embodiments 1 to 28, having a size independently selected from about 4,000 nt, about 2,000 nt to about 2,500 nt, about 2,000 nt to about 3,000 nt, about 2,000 nt to about 3,500 nt, about 2,000 nt to about 4,000 nt, about 2,500 nt to about 3,000 nt, about 2,500 nt to about 3,500 nt, about 2,500 nt to about 4,000 nt, about 3,000 nt to about 3,500 nt, about 3,000 nt to about 4,000 nt, about 3,500 nt to about 4,000 nt, about 1,000 nt, about 1,500 nt, about 2,000 nt, about 2,500 nt, about 3,000 nt, about 3,500 nt, and about 4,000 nt.

[0267] 30. Each of the one, two, three, or four RNAs encoded by one, two, three, or four synthetic nucleic acid molecules of the system is about 2,500 to 4,500 nt, about 2,500 nt to about 2,750 nt, about 2,500 nt to about 3,000 nt, about 2,500 nt to about 3,250 nt, about 2,500 nt to about 3,500 nt, about 2,500 nt to about 3,750 nt, about 2,500 nt to about 4,000 nt, about 2,500 nt to about 4,250 nt, about 2,500nt~about 4,500nt, about 2,750nt~about 3,000nt, about 2,750nt~about 3,250nt, about 2,750nt~about 3,500nt, about 2,750nt~about 3,750nt, about 2,750nt~about 4,000nt, Approximately 2,750nt to approximately 4,250nt, approximately 2,750nt to approximately 4,500nt, approximately 3,000nt to approximately 3,250nt, approximately 3,000nt to approximately 3,500nt, approximately 3,000nt to approximately 3,750nt, approximately 3,000nt to approximately 4,000nt , about 3,000nt to about 4,250nt, about 3,000nt to about 4,500nt, about 3,250nt to about 3,500nt, about 3,250nt to about 3,750nt, about 3,250nt to about 4,000nt, about 3,250nt to about 4,250nt t, about 3,250nt to about 4,500nt, about 3,500nt to about 3,750nt, about 3,500nt to about 4,000nt, about 3,500nt to about 4,250nt, about 3,500nt to about 4,500nt, about 3,750nt to about 4,000 30. The system of any one of embodiments 1 to 29, wherein the system has a size independently selected from about 3,750 nt to about 4,250 nt, about 3,750 nt to about 4,500 nt, about 4,000 nt to about 4,250 nt, about 4,000 nt to about 4,500 nt, about 4,250 nt to about 4,500 nt, about 2,500 nt, about 2,750 nt, about 3,000 nt, about 3,250 nt, about 3,500 nt, about 3,750 nt, about 4,000 nt, about 4,250 nt, and about 4,500 nt.

[0268] 31. The synthetic nucleic acid molecule is selected from the group consisting of about 5,000 nt to about 10,000 nt, about 5,000 nt to about 5,500 nt, about 5,000 nt to about 6,000 nt, about 5,000 nt to about 6,500 nt, about 5,000 nt to about 7,000 nt, about 5,000 nt to about 7,500 nt, about 5,000 nt to about 8,000 nt, about 5,000 nt to about 8,500 nt, about 5,000 nt to about 9,000 nt, about 5,000 nt to about 9,500 nt, about 5,000 nt to about 10,000 nt, about 5,500 nt to about 6,000 nt, about 5,500 nt to about 6,500 nt, and about 5 , 500nt ~ approx. 7,000nt, approx. 5,500nt ~ approx. 7,500nt, approx. 5,500nt ~ approx. 8,000nt, approx. 5,500nt ~ approx. 8,500nt, approx. , about 6,000nt to about 6,500nt, about 6,000nt to about 7,000nt, about 6,000nt to about 7,500nt, about 6,000nt to about 8,000nt, about 6,000nt to about 8,500nt, about 6,000nt to about 9,000nt, about 6,000nt to about 9,500nt t, about 6,000nt to about 10,000nt, about 6,500nt to about 7,000nt, about 6,500nt to about 7,500nt, about 6,500nt to about 8,000nt, about 6,500nt to about 8,500nt, about 6,500nt to about 9,000nt, about 6,500nt to about 9,5 00nt, approximately 6,500nt to approximately 10,000nt, approximately 7,000nt to approximately 7,500nt, approximately 7,000nt to approximately 8,000nt, approximately 7,000nt to approximately 8,500nt, approximately 7,000nt to approximately 9,000nt, approximately 7,000nt to approximately 9,500nt, approximately 7,000nt to approximately 10,000nt, approx. 7,500nt ~ approx. 8,000nt, approx. 7,500nt ~ approx. 8,500nt, approx. 7,500nt ~ approx. 9,000nt, approx. 7,500nt ~ approx. 9,500nt, approx. nt ~ about 9,000nt, about 8,000nt - about 9,500nt, about 8,000nt - about 10,000nt, about 8,500nt - about 9,000nt, about 8,500nt - about 9,500nt, about 8,500nt - about 10,000nt, about 9,000nt - about 9,500nt, about 9,having a total size selected from about 1,000 nt to about 10,000 nt, about 9,500 nt to about 10,000 nt, about 5,000 nt, about 5,500 nt, about 6,000 nt, about 6,500 nt, about 7,000 nt, about 7,500 nt, about 8,000 nt, about 8,500 nt, about 9,000 nt, about 9,500 nt, and about 10,000 nt; The total target protein coding sequence is from about 2,000 nt to about 8,000 nt, from about 2,000 nt to about 3,000 nt, from about 2,000 nt to about 3,500 nt, from about 2,000 nt to about 4,000 nt, from about 2,000 nt to about 4,500 nt, from about 2,000 nt to about 5,000 nt, from about 2,000 nt to about 5,500 nt, from about 2,000 nt to about 6,000 nt, from about 2,000 nt to about 6,500 nt, from about 2,000 nt to about 7,000 nt, from about 2,000 nt to about 7,500 nt, from about 2,000 nt to about 8,000 nt, from about 3,000 nt to about 3,500 nt, from about 3,000 nt to about 4,500 nt, ,000nt~about 4,000nt, about 3,000nt~about 4,500nt, about 3,000nt~about 5,000nt, about 3,000nt~about 5,500nt, about 3,000nt~about 6,000nt, about 3,000nt~about 6,500nt, about 3,000nt~about 7,000nt, Approximately 3,000nt to approximately 7,500nt, approximately 3,000nt to approximately 8,000nt, approximately 3,500nt to approximately 4,000nt, approximately 3,500nt to approximately 4,500nt, approximately 3,500nt to approximately 5,000nt, approximately 3,500nt to approximately 5,500nt, approximately 3,500nt to approximately 6,000nt , about 3,500nt to about 6,500nt, about 3,500nt to about 7,000nt, about 3,500nt to about 7,500nt, about 3,500nt to about 8,000nt, about 4,000nt to about 4,500nt, about 4,000nt to about 5,000nt, about 4,000nt to about 5,500nt nt, approximately 4,000nt to approximately 6,000nt, approximately 4,000nt to approximately 6,500nt, approximately 4,000nt to approximately 7,000nt, approximately 4,000nt to approximately 7,500nt, approximately 4,000nt to approximately 8,000nt, approximately 4,500nt to approximately 5,000nt, approximately 4,500nt to approximately 5,50 0nt, about 4,500nt to about 6,000nt, about 4,500nt to about 6,500nt, about 4,500nt to about 7,000nt, about 4,500nt to about 7,500nt, about 4,500nt to about 8,000nt, about 5,000nt to about 5,500nt, about 5,000nt to about 6,0 00nt, about 5,000nt to about 6,500nt, about 5,000nt to about 7,000nt, about 5,000nt to about 7,500nt, about 5,000nt to about 8,000nt, about 5,500nt to about 6,000nt, about 5,500nt to about 6,500nt, about 5,500nt to about 7,000nt, about 5,500nt to about 7,500nt, about 5,500nt to about 8,000nt, about 6,000nt to about 6,500nt, about 6,000nt to about 7,000nt, about 6,000nt to about 7,50 0nt, approximately 6,000nt to approximately 8,000nt, approximately 6,500nt to approximately 7,000nt, approximately 6,500nt to approximately 7,500nt, approximately 6,500nt to approximately 8,000nt, approximately 7,000nt to approximately 7,500n nt, about 7,000 nt to about 8,000 nt, or about 7,500 nt to about 8,000 nt, with the total target protein coding sequence being about 2,000 nt, about 3,000 nt, about 3,500 nt, about 4,000 nt, about 4,500 nt, about 5,000 nt, about 5,500 nt, about 6,000 nt, about 6,500 nt, about 7,000 nt, about 7,500 nt, and about 8,000 nt; and / or The RNA encoded by the two synthetic nucleic acid molecules may be about 5,000nt to about 9,000nt, about 5,000nt to about 5,500nt, about 5,000nt to about 6,000nt, about 5,000nt to about 6,500nt, about 5,000nt to about 7,000nt, about 5,000nt to about 7,500nt, about 5,000nt to about 8,000nt, about 5,000nt to about 8,500nt, about 5,000nt to about 9,000nt, about 5,500nt to about 6,000nt, about 5, 500nt ~ approx. 6,500nt, approx. 5,500nt ~ approx. 7,000nt, approx. 5,500nt ~ approx. 7,500nt, approx. 5,500nt ~ approx. 8,000nt, approx. t, about 6,000nt to about 6,500nt, about 6,000nt to about 7,000nt, about 6,000nt to about 7,500nt, about 6,000nt to about 8,000nt, about 6,000nt to about 8,500nt, about 6,000nt to about 9 ,000nt, about 6,500nt to about 7,000nt, about 6,500nt to about 7,500nt, about 6,500nt to about 8,000nt, about 6,500nt to about 8,500nt, about 6,500nt to about 9,000nt, about 7,000nt nt~about 7,500nt, about 7,000nt~about 8,000nt, about 7,000nt~about 8,500nt, about 7,000nt~about 9,000nt, about 7,500nt~about 8,000nt, about 7,500nt~about 8,500nt, about and the RNAs encoded by the two synthetic nucleic acid molecules have a total size of about 5,000 nt, about 5,500 nt, about 6,000 nt, about 6,500 nt, about 7,000 nt, about 7,500 nt, about 8,000 nt, about 8,500 nt, and about 9,000 nt. The system of any one of embodiments 1 and 4 to 30.

[0269] 32. The synthetic nucleic acid molecule is from about 7,500 nt to about 15,000 nt, from about 7,500 nt to about 8,500 nt, from about 7,500 nt to about 9,500 nt, from about 7,500 nt to about 10,000 nt, from about 7,500 nt to about 10,500 nt, from about 7,500 nt to about 11,000 nt, from about 7,500 nt to about 11,500 nt, from about 7,500 nt to about 12,000 nt, from about 7,500 nt to about 12,500 nt, from about 7,500 nt to about 13,000 nt, from about 7,500 nt to about 14,000 nt, from about 7,500 nt to about 15,000 nt, or from about 8,500 nt to about 9,500 nt. 0nt, about 8,500nt to about 10,000nt, about 8,500nt to about 10,500nt, about 8,500nt to about 11,000nt, about 8,500nt to about 11,500nt, about 8,500nt to about 12,000nt, about 8,500nt to about 12,500nt, about 8,500nt t ~ about 13,000nt, about 8,500nt - about 14,000nt, about 8,500nt - about 15,000nt, about 9,500nt - about 10,000nt, about 9,500nt - about 10,500nt, about 9,500nt - about 11,000nt, about 9,500nt - about 11,500nt , about 9,500nt to about 12,000nt, about 9,500nt to about 12,500nt, about 9,500nt to about 13,000nt, about 9,500nt to about 14,000nt, about 9,500nt to about 15,000nt, about 10,000nt to about 10,500nt, about 10,000nt ~11,000nt, 10,000nt~11,500nt, 10,000nt~12,000nt, 10,000nt~12,500nt, 10,000nt~13,000nt, 10,000nt~14,000nt, 10,000nt~15, 000nt, approx. 10,500nt ~ approx. 11,000nt, approx. 10,500nt ~ approx. 11,500nt, approx. 10,500nt ~ approx. 12,000nt, approx. 10,500nt ~ approx. 12,500nt, approx. , approx. 10,500nt ~ approx. 15,000nt, approx. 11,000nt ~ approx. 11,500nt, approx. 11,000nt ~ approx. 12,000nt, approx. 11,000nt ~ approx. 12,500nt, approx.000nt ~ approx. 15,000nt, approx. 11,500nt ~ approx. 12,000nt, approx. 11,500nt ~ approx. 12,500nt, approx. 11,500nt ~ approx. 13,000nt, approx. 12,000nt ~ approx. 12,500nt, approx. 12,000nt ~ approx. 13,000nt, approx. 12,000nt ~ approx. 14,000nt, approx. 12,000nt ~ approx. 15,000nt, approx. 12,500nt ~ approx. 13,000nt, approx. , has a total size selected from about 12,500 nt to about 15,000 nt, about 13,000 nt to about 14,000 nt, about 13,000 nt to about 15,000 nt, or about 14,000 nt to about 15,000 nt, and the synthetic nucleic acid molecule has a total size of about 7,500 nt, about 8,500 nt, about 9,500 nt, about 10,000 nt, about 10,500 nt, about 11,000 nt, about 11,500 nt, about 12,000 nt, about 12,500 nt, about 13,000 nt, about 14,000 nt, and about 15,000 nt; The total target protein coding sequence is about 3,000 nt to about 12,000 nt, about 3,000 nt to about 4,000 nt, about 3,000 nt to about 5,000 nt, about 3,000 nt to about 6,000 nt, about 3,000 nt to about 7,000 nt, about 3,000 nt to about 7,500 nt, about 3,000 nt to about 8,000 nt, about 3,000 nt to about 8,500 nt, about 3,000 nt to about 9,000 nt, about 3,000 nt to about 1,000 nt, about 3,000 nt to about 11,000 nt, about 3,000 nt to about 12,000 nt, about 4,000 nt to about 5,000 nt, nt, approximately 4,000nt to approximately 6,000nt, approximately 4,000nt to approximately 7,000nt, approximately 4,000nt to approximately 7,500nt, approximately 4,000nt to approximately 8,000nt, approximately 4,000nt to approximately 8,500nt, approximately 4,000nt to approximately 9,000nt, approximately 4,000nt to approximately 1,0 00nt, about 4,000nt to about 11,000nt, about 4,000nt to about 12,000nt, about 5,000nt to about 6,000nt, about 5,000nt to about 7,000nt, about 5,000nt to about 7,500nt, about 5,000nt to about 8,000nt, about 5,000nt Approx. 8,500nt, Approx. 5,000nt~Approx. 9,000nt, Approx. 5,000nt~Approx. 1,000nt, Approx. 5,000nt~Approx. 11,000nt, Approx. 0nt~about 8,000nt, about 6,000nt~about 8,500nt, about 6,000nt~about 9,000nt, about 6,000nt~about 1,000nt, about 6,000nt~about 11,000nt, about 6,000nt~about 12,000nt, about 7,000nt~about 7,500nt, about 7,000nt to about 8,000nt, about 7,000nt to about 8,500nt, about 7,000nt to about 9,000nt, about 7,000nt to about 1,000nt, about 7,000nt to about 11,000nt, about 7,000nt to about 12,000nt, about 7,500nt to about 8,000 nt, about 7,500nt to about 8,500nt, about 7,500nt to about 9,000nt, about 7,500nt to about 1,000nt, about 7,500nt to about 11,000nt, about 7,500nt to about 12,000nt, about 8,000nt to about 8,500nt, about 8,000nt to about 9,000nt, about 8,000nt to about 1,000nt, about 8,000nt to about 11,000nt, about 8,000nt to about 12,000nt, about 8,500nt to about 9,000nt, about 8,500nt to about 1,000nt, about 8 , 500nt ~ approx. 11,000nt, approx. 8,500nt ~ approx. 12,000nt, approx. 9,000nt ~ approx. 1,000nt, approx. 9,000nt ~ approx. 11,000nt, approx. 9,000nt ~ approx. 12,000nt, approx. 1,000nt and / or, wherein the total target protein coding sequence is selected from about 3,000 nt, about 4,000 nt, about 5,000 nt, about 6,000 nt, about 7,000 nt, about 7,500 nt, about 8,000 nt, about 8,500 nt, about 9,000 nt, about 1,000 nt, about 11,000 nt, and about 12,000 nt; and / or The RNAs encoded by the three synthetic nucleic acid molecules are selected from the group consisting of about 7,500 nt to about 13,500 nt, about 7,500 nt to about 8,500 nt, about 7,500 nt to about 9,000 nt, about 7,500 nt to about 9,500 nt, about 7,500 nt to about 10,000 nt, about 7,500 nt to about 10,500 nt, about 7,500 nt to about 11,000 nt, about 7,500 nt to about 11,500 nt, about 7,500 nt to about 12,000 nt, about 7,500 nt to about 12,500 nt, about 7,500 nt to about 13,000 nt, and about 7,500 nt to about 13,500 nt. Approximately 8,500nt to approximately 9,000nt, approximately 8,500nt to approximately 9,500nt, approximately 8,500nt to approximately 10,000nt, approximately 8,500nt to approximately 10,500nt, approximately 8,500nt to approximately 11,000nt, approximately 8,500nt to approximately 11,500nt, approximately 12 ,000nt, about 8,500nt to about 12,500nt, about 8,500nt to about 13,000nt, about 8,500nt to about 13,500nt, about 9,000nt to about 9,500nt, about 9,000nt to about 10,000nt, about 9,000nt to about 10,500nt, about 9,0 00nt ~ approx. 11,000nt, approx. 9,000nt ~ approx. 11,500nt, approx. 9,000nt ~ approx. 12,000nt, approx. 9,000nt ~ approx. 12,500nt, approx. 9,000nt ~ approx. 13,000nt, approx. 00nt, about 9,500nt to about 10,500nt, about 9,500nt to about 11,000nt, about 9,500nt to about 11,500nt, about 9,500nt to about 12,000nt, about 9,500nt to about 12,500nt, about 9,500nt to about 13,000nt, about 9,50 0nt ~ approx. 13,500nt, approx. 10,000nt ~ approx. 10,500nt, approx. 10,000nt ~ approx. 11,000nt, approx. 10,000nt ~ approx. 11,500nt, approx. 10,000nt ~ approx. 12,000nt, approx. Approx. 13,000nt, Approx. 10,000nt~Approx. 13,500nt, Approx. 10,500nt~Approx. 11,000nt, Approx. 10,500nt~Approx. 11,500nt, Approx. 10,500nt~Approx. 12,000nt, Approx.000nt, approx. 10,500nt ~ approx. 13,500nt, approx. 11,000nt ~ approx. 11,500nt, approx. 11,000nt ~ approx. 12,000nt, approx. 11,000nt ~ approx. 12,500nt, approx. t, approx. 11,500nt ~ approx. 12,000nt, approx. 11,500nt ~ approx. 12,500nt, approx. 11,500nt ~ approx. 13,000nt, approx. and the RNAs encoded by the two synthetic nucleic acid molecules have a total size of about 7,500 nt, about 8,500 nt, about 9,000 nt, about 9,500 nt, about 10,000 nt, about 10,500 nt, about 11,000 nt, about 11,500 nt, about 12,000 nt, about 12,500 nt, about 13,000 nt, and about 13,500 nt. The system of any one of embodiments 2 and 4 to 30.

[0270] 33. The synthetic nucleic acid molecule has a length of about 10,000 nt to about 20,000 nt, about 10,000 nt to about 11,000 nt, about 10,000 nt to about 12,000 nt, about 10,000 nt to about 13,000 nt, about 10,000 nt to about 14,000 nt, about 10,000 nt to about 15,000 nt, about 10,000 nt to about 16,000 nt, about 10,000 nt to about 17,000 nt, about 10,000 nt to about 18,000 nt, about 10,000 nt to about 19,000 nt, about 10,000 nt to about 20,000 nt, or about 11,000 nt to about 12,000 nt. nt, approx. 11,000nt ~ approx. 13,000nt, approx. 11,000nt ~ approx. 14,000nt, approx. 11,000nt ~ approx. 15,000nt, approx. 11,000nt ~ approx. 16,000nt, approx. 11,000nt ~ approx. 19,000nt, approx. 11,000nt ~ approx. 20,000nt, approx. 12,000nt ~ approx. 13,000nt, approx. 12,000nt ~ approx. 14,000nt, approx. 12,000nt ~ approx. 15,000nt, approx. 0nt ~ approx. 17,000nt, approx. 12,000nt ~ approx. 18,000nt, approx. 12,000nt ~ approx. 19,000nt, approx. 12,000nt ~ approx. 20,000nt, approx. 13,000nt ~ approx. 14,000nt, approx. Approx. 16,000nt, Approx. 13,000nt~Approx. 17,000nt, Approx. 13,000nt~Approx. 18,000nt, Approx. 13,000nt ~ Approx. 19,000nt, Approx. 13,000nt~Approx. 20,000nt, Approx. 00nt, about 14,000nt to about 17,000nt, about 14,000nt to about 18,000nt, about 14,000nt to about 19,000nt, about 14,000nt to about 20,000nt, about 15,000nt to about 16,000nt, about 15,000nt to about 17,000nt , about 15,000nt to about 18,000nt, about 15,000nt to about 19,000nt, about 15,000nt to about 20,000nt, about 16,000nt to about 17,000nt, about 16,000nt to about 18,000nt, about 16,000nt to about 19,000nt, about 16,and the synthetic nucleic acid molecule has a total size selected from about 10,000 nt, about 11,000 nt, about 12,000 nt, about 13,000 nt, about 14,000 nt, about 15,000 nt, about 16,000 nt, about 17,000 nt, about 18,000 nt, about 19,000 nt, about 17,000 nt, about 20,000 nt, about 18,000 nt, about 19,000 nt, and about 20,000 nt; The total target protein coding sequence is from about 4,000 nt to about 16,000 nt, from about 5,000 nt to about 6,000 nt, from about 5,000 nt to about 7,000 nt, from about 5,000 nt to about 8,000 nt, from about 5,000 nt to about 9,000 nt, from about 5,000 nt to about 10,000 nt, from about 5,000 nt to about 11,000 nt, from about 5,000 nt to about 12,000 nt, from about 5,000 nt to about 13,000 nt, from about 5,000 nt to about 14,000 nt, from about 5,000 nt to about 15,000 nt, from about 5,000 nt to about 16,000 nt, from about 6,000 nt to about 7,000nt, about 6,000nt to about 8,000nt, about 6,000nt to about 9,000nt, about 6,000nt to about 10,000nt, about 6,000nt to about 11,000nt, about 6,000nt to about 12,000nt, about 6,000nt to about 13,000nt, about 6,0 00nt ~ approx. 14,000nt, approx. 6,000nt ~ approx. 15,000nt, approx. 6,000nt ~ approx. 16,000nt, approx. 7,000nt ~ approx. 8,000nt, approx. 7,000nt ~ approx. 9,000nt, approx. 7,000nt ~ approx. 10,000nt, approx. 7,000nt ~ approx. nt, about 7,000nt to about 12,000nt, about 7,000nt to about 13,000nt, about 7,000nt to about 14,000nt, about 7,000nt to about 15,000nt, about 7,000nt to about 16,000nt, about 8,000nt to about 9,000nt, about 8,000nt ~10,000nt, 8,000nt~11,000nt, 8,000nt~12,000nt, 8,000~13,000nt, 8,000~14,000nt, 8,000~15,000nt, 8,000~16,000nt , approximately 9,000nt ~ approximately 10,000nt, approximately 9,000nt ~ approximately 11,000nt, approximately 9,000nt ~ approximately 12,000nt, approximately 9,000nt ~ approximately 13,000nt, approximately 9,000nt ~ approximately 14,000nt, approximately 9,000nt ~ approximately 15,000nt, approximately 9,000nt ~ Approx. 16,000nt, Approx. 10,000nt~Approx. 11,000nt, Approx. 10,000nt~Approx. 12,000nt, Approx. 10,000nt~Approx. 13,000nt, Approx. 10,000nt~Approx. 14,000nt, Approx.000nt, approx. 11,000nt ~ approx. 12,000nt, approx. 11,000nt ~ approx. 13,000nt, approx. 11,000nt ~ approx. 14,000nt, approx. 11,000nt t ~ approx. 15,000nt, approx. 11,000nt ~ approx. 16,000nt, approx. 12,000nt ~ approx. 13,000nt, approx. 12,000nt ~ approx. 14,000nt, approx. 1 2,000nt to about 15,000nt, about 12,000nt to about 16,000nt, about 13,000nt to about 14,000nt, about 13,000nt to about 15,000nt, about 13,000nt to about 16,000nt, about 14,000nt to about 15,000nt, about 14,000nt to about 16,000nt, or about 15,000 nt to about 16,000 nt, wherein the total target protein coding sequence is about 5,000 nt, about 6,000 nt, about 7,000 nt, about 8,000 nt, about 9,000 nt, about 10,000 nt, about 11,000 nt, about 12,000 nt, about 13,000 nt, about 14,000 nt, about 15,000 nt, or about 16,000 nt, wherein the total target protein coding sequence is at least about 5,000 nt, about 6,000 nt, about 7,000 nt, about 8,000 nt, about 9,000 nt, about 10,000 nt, about 11,000 nt, about 12,000 nt, about 13,000 nt, about 14,000 nt, and about 15,000 nt; and / or The RNAs encoded by the two synthetic nucleic acid molecules are about 10,000 nt to about 18,000 nt, about 10,000 nt to about 11,000 nt, about 10,000 nt to about 12,000 nt, about 10,000 nt to about 13,000 nt, about 10,000 nt to about 14,000 nt, about 10,000 nt to about 15,000 nt, about 10,000 nt to about 16,000 nt, about 10,000 nt to about 17,000 nt, about 10,000 nt to about 18,000 nt, about 11,000 nt to about 12,000 nt, about 11 ,000nt~about 13,000nt, about 11,000nt~about 14,000nt, about 11,000nt~about 15,000nt, about 11,000nt~about 16,000nt, about 11,000nt~about 17,000nt, about 11,000nt~about 18,00 0nt, about 12,000nt~about 13,000nt, about 12,000nt~about 14,000nt, about 12,000nt~about 15,000nt, about 12,000nt~about 16,000nt, about 12,000nt~about 17,000nt, about 12,000nt~ Approx. 18,000nt, Approx. 13,000nt~Approx. 14,000nt, Approx. 13,000nt~Approx. 15,000nt, Approx. 13,000nt~Approx. 16,000nt, Approx. 13,000nt~Approx. 17,000nt, Approx. ,000nt ~ approx. 15,000nt, approx. 14,000nt ~ approx. 16,000nt, approx. 14,000nt ~ approx. 17,000nt, approx. 14,000nt ~ approx. 18,000nt, approx. nt, about 15,000nt to about 18,000nt, about 16,000nt to about 17,000nt, about 16,000nt to about 18,000nt, or about 17,000nt to about 18,000nt, and the RNAs encoded by the two synthetic nucleic acid molecules have a total size of about 10,000nt, about 11,000nt, about 12,000nt, about 13,000nt, about 14,000nt, about 15,000nt, about 16,000nt, about 17,000nt, and about 18,000nt. The system of any one of embodiments 3 and 4 to 30.

[0271] 34. RNA recombination efficiency is about 10% to about 95%, about 10% to about 20%, about 10% to about 30%, about 10% to about 35%, about 10% to about 40%, about 10% to about 45%, about 10% to about 50%, about 10% to about 55%, about 10% to about 60%, about 10% to about 70%, about 10% to about 80%, about 10% to about 90%, about 20% to about 30%, about 20% to about 35%, about 20% to about 40%, about 20% to about 45%, about 20% to about 50%, about 20% to about 55%, about 20% to about 60%, about 20% to about 70%, about 20% to about 80%, about 20% to about 90%, about 30% to about 35%, about 30% to about 40%, about 30% to about 45%, about 30% to about 50%, about 30% to about 55%, about 30% to about 60%, about 30% to about 70%, about 30% to about 80%, about 30% to about 90%, about 35% to about 40%, about 35% to about 45%, about 35% to about 50%, about 35% to about 55%, about 35% to about 60%, about 35% to about 70%, about 35% to about 80%, about 35% to about 90%, about 40% to about 45%, about 40% to about 50%, about 40% to about 55%, about 40% to about 60%, about 40% to about 70%, about 40% to about 80%, about 40% to about 90%, about 45% to about 50%, about 45% to about 55%, about 45% to about 60%, about 45% to about 70%, about 45% to about 80%, about 45% to about 90%, about 50% to about 55%, about 50% to about 60%, about 50% to about 70%, about 50% to about 80%, about 50% to about 90%, 34. The system of any one of embodiments 1 to 33, wherein the chromatographic efficiency is about 55% to about 60%, about 55% to about 70%, about 55% to about 80%, about 55% to about 90%, about 60% to about 70%, about 60% to about 80%, about 60% to about 90%, about 70% to about 80%, about 70% to about 90%, about 80% to about 90%, about 10%, about 20%, about 30%, about 35%, about 40%, about 45%, about 50%, about 55%, about 60%, about 70%, about 80%, or about 90%, or about 95%.

[0272] 35. The first and second dimerization domains, the third and fourth dimerization domains, and / or the fifth and sixth dimerization domains are each 1000 nt or less, for example, at least 50 nt, at least 100 nt, at least 150 nt, at least 200 nt, at least 300 nt, at least 400 nt, at least 500 nt, 50 to 1000 nt, 50 to 500 nt, 50 to 150 nt, 50 nt, 100 nt, 150 nt t, 200nt, 250nt, 300nt, 400nt, or 500nt; and the recombination efficiency of the system is at least about 20%, at least about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, or at least about 95%.

[0273] 36. The system of any one of embodiments 1 to 35, wherein each dimerization domain is 1000 nt or less, e.g., at least 50 nt, at least 100 nt, at least 150 nt, at least 200 nt, at least 300 nt, at least 400 nt, at least 500 nt, 50-1000 nt, 50-500 nt, 50-150 nt, 50 nt, 100 nt, 150 nt, 200 nt, 250 nt, 300 nt, 400 nt, or 500 nt; and the recombination efficiency of the system is at least 20%, at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, at least 75%, at least 80%, or at least 90%.

[0274] 37. A composition comprising the system of any one of embodiments 1 to 36.

[0275] 38. A composition comprising an RNA molecule according to any one of embodiments 1 to 37.

[0276] 39. A composition comprising one, two, three, or four RNA molecules according to any one of embodiments 1 to 37.

[0277] 40. The composition of any one of embodiments 37 to 39, comprising a first synthetic nucleic acid or RNA molecule, a second synthetic nucleic acid or RNA molecule, a third synthetic nucleic acid or RNA molecule, and optionally a fourth synthetic nucleic acid or RNA molecule, each encoding at least a portion of dystrophin, Factor VIII, ABCA4, or MYO7A.

[0278] 41. An RNA molecule according to any one of embodiments 1 to 36.

[0279] 42. A kit comprising the system of any one of embodiments 1 to 41, or the composition of any one of embodiments 37 to 40, wherein any of the first, second, third and fourth synthetic nucleic acid molecules may be in separate containers, and optionally further comprising a buffer, such as a pharmaceutically acceptable carrier.

[0280] 43. A method for expressing a target protein in a cell, comprising introducing into the cell the system of any one of embodiments 1 to 36 or the composition of any one of embodiments 35 to 37; and expressing in the cell a first synthetic RNA molecule and a second synthetic RNA molecule, or a first synthetic RNA molecule, a second synthetic RNA molecule, and a third synthetic RNA molecule, or a first synthetic RNA molecule, a second synthetic RNA molecule, a third synthetic RNA molecule, and a fourth synthetic RNA molecule, wherein the target protein is produced in the cell.

[0281] 44. The method of embodiment 43, wherein the cells are present in a subject and the introducing step comprises administering the system to the subject in a therapeutically effective amount.

[0282] 45. The method of embodiment 44, wherein a genetic disease caused by a mutation in a gene encoding a target protein in a subject is treated, resulting in expression of a functional target protein in the subject.

[0283] 46. The genetic disease is Duchenne muscular dystrophy and the target protein is dystrophin; The genetic disease is hemophilia A and the target protein is F8; the genetic disease is Stargardt disease and the target protein is ABCA4; or The genetic disease is Usher syndrome and the target protein is MYO7A. The method of embodiment 45.

[0284] 47. A nucleic acid molecule comprising at least 80%, at least 85%, at least 90%, at least 95%, at least 98%, at least 99%, or 100% sequence identity to the synthetic intron set forth in any one of SEQ ID NOs: 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 145, 146, 147, 148, 155, 156, 157, 158, 159, 160, 161, 162, 163, 164, 165, and 166.

[0285] 48. The nucleic acid molecule of embodiment 47, wherein the synthetic intron is nt 3703 to 3975 of SEQ ID NO: 20, nt 1 to 228 of SEQ ID NO: 21, nt 3703 to 3975 of SEQ ID NO: 22, nt 1 to 225 of SEQ ID NO: 23, nt 3560 to 3828 of SEQ ID NO: 24, or nt 1 to 225 of SEQ ID NO: 25.

[0286] 49. The synthetic nucleic acid molecule of embodiment 47 or 48, further comprising a portion of a protein-coding sequence.

[0287] 50. The synthetic nucleic acid molecule of embodiment 49, wherein the portion of the protein-coding sequence comprises the N-terminal half, the N-terminal third, the middle portion, the C-terminal half, or the C-terminal third of the protein-coding sequence.

[0288] 51. The system of any one of embodiments 1 to 36 or the composition of any one of embodiments 37 to 40, wherein at least one synthetic nucleic acid molecule comprises a synthetic intron comprising a nucleic acid molecule according to any one of embodiments 47 to 50.

[0289] 52. The composition, system, method, or kit of any of the preceding embodiments, wherein the synthetic nucleic acid is DNA generated by transcription of an RNA viral genome by reverse transcriptase.

[0290] VII. Additional Exemplary Embodiments 1. A composition for expressing a target protein, comprising: (a) a first RNA molecule comprising, from 5' to 3': (i) a coding sequence for the N-terminal portion of the target protein; (ii) a splice donor; and (iii) a first dimerization domain; and (b) a second RNA molecule comprising, from 5' to 3': (i) a second dimerization domain that binds to the first dimerization domain; (ii) a branch point sequence; (iii) a polypyrimidine tract; (iv) a splice acceptor; and (v) a coding sequence for the C-terminal portion of the target protein.

[0291] 2. A composition for expressing a target protein, comprising: (a) a first RNA molecule, the first RNA molecule comprising, from 5' to 3': (i) a coding sequence for an N-terminal portion of the target protein; (ii) a splice donor; and (iii) a first dimerization domain; and (b) a second RNA molecule, the second RNA molecule comprising, from 5' to 3': (i) a second dimerization domain that binds to the first dimerization domain; (ii) a branch point sequence; (iii) a polypyrimidine tract; and (iv) a splice acceptor. (v) a coding sequence for the middle portion of the target protein; (vi) a second splice donor; and (vii) a third dimerization domain; and (c) a third RNA molecule, the third RNA molecule comprising, from 5' to 3': (i) a fourth dimerization domain that binds to the third dimerization domain; (ii) a branch point sequence; (iii) a polypyrimidine tract; (iv) a splice acceptor; and (v) a coding sequence for the C-terminal portion of the target protein.

[0292] 3. A composition for expressing a target protein, comprising: (a) a first RNA molecule, the first RNA molecule comprising, from 5' to 3': (i) a coding sequence for an N-terminal portion of the target protein, (ii) a splice donor; and (iii) a first dimerization domain; (b) a second RNA molecule, the second RNA molecule comprising, from 5' to 3': (i) a second dimerization domain that binds to the first dimerization domain; (ii) a branch point sequence; (iii) a polypyrimidine tract; (iv) a splice acceptor; (v) a coding sequence for an intermediate portion of the target protein; (vi) a second splice donor; and (vii) a third dimerization domain; and (c) a third RNA molecule. (iv) a splice acceptor; (v) a coding sequence for a first intermediate portion of the target protein; (vi) a second splice donor; and (vii) a fifth dimerization domain; and (d) a fourth RNA molecule, the fourth RNA molecule comprising, from 5' to 3': (i) a sixth dimerization domain linked to the fifth dimerization domain; (ii) a branch point sequence; (iii) a polypyrimidine tract; (iv) a splice acceptor; and (v) a coding sequence for a C-terminal portion of the target protein.

[0293] 4. The composition of any one of embodiments 1 to 3, wherein the first dimerization domain and the second dimerization domain, the third dimerization domain and the fourth dimerization domain, and / or the fifth dimerization domain and the sixth dimerization domain are linked by a direct bond, an indirect bond, or a combination thereof.

[0294] 5. The composition of claim 4, wherein the direct or indirect bond comprises a base pairing interaction, a non-canonical base pairing interaction, a non-base pairing interaction, or a combination thereof.

[0295] 6. The composition of claim 4 or 5, wherein the direct binding comprises base-pairing interactions between kissing loops or low diversity regions.

[0296] 7. The composition of claim 4 or 5, wherein the direct binding comprises non-canonical base pairing interactions, non-canonical base pairing interactions, non-base pairing interactions, or a combination thereof between the aptamer regions.

[0297] 8. The composition of claim 4 or 5, wherein the indirect binding comprises a base-pairing interaction through a nucleic acid bridge.

[0298] 9. The composition of claim 4 or 5, wherein the indirect binding comprises a non-base pairing interaction between the aptamer and the aptamer targeting agent, or between two aptamers.

[0299] 10. The composition of any one of embodiments 1 to 9, wherein the first dimerization domain, the second dimerization domain, the third dimerization domain, the fourth dimerization domain, the fifth dimerization domain and / or the sixth dimerization domain does not comprise a cryptic splice acceptor.

[0300] 11. The composition of any one of embodiments 1 to 10, comprising at least one pair of directly or indirectly linked aptamer sequence dimerization domains.

[0301] 12. The composition of any one of embodiments 1 to 11, comprising at least one pair of kissing loop interacting dimerization domains.

[0302] 13. The composition of any one of embodiments 1 to 12, wherein the target protein is a disease-associated protein or a therapeutic protein.

[0303] 14. The composition of embodiment 13, wherein the disease is a monogenic disease.

[0304] 15. The composition of embodiment 14, wherein the therapeutic protein is a toxin.

[0305] 16. The composition of any one of embodiments 13 to 15, wherein the diseases and target proteins are those listed in Table 1.

[0306] 17. The composition of any one of embodiments 1 to 16, wherein the first RNA molecule, the second RNA molecule, the third RNA molecule, and / or the fourth RNA molecule further comprises a poly-A tail at the 3' end of the first RNA molecule, the second RNA molecule, the third RNA molecule, or the fourth RNA molecule.

[0307] 18. the first RNA molecule further comprises one or both of a downstream intron splice enhancer (DISE) 3' to the splice donor and 5' to the first dimerization domain, an intron splice enhancer (ISE) 3' to the splice donor and 5' to the first dimerization domain; and / or the secon...

Claims

1. A composition for expressing a target protein, said composition comprising: (a) a first DNA molecule encoding a first RNA molecule, the first DNA molecule comprising a first promoter operably linked to a sequence encoding the first RNA molecule, the first RNA molecule comprising, from 5' to 3': (i) a coding sequence for an N-terminal portion of the target protein; (ii) a splice donor; and (iii) a first dimerization domain; and (b) a second DNA molecule encoding a second RNA molecule, the second DNA molecule comprising a second promoter operably linked to a sequence encoding the second RNA molecule, the second RNA molecule comprising, from 5' to 3': (i) a second dimerization domain that binds to the first dimerization domain; (ii) a branch point sequence; (iii) a polypyrimidine tract; (iv) a splice acceptor; and (v) a coding sequence for the C-terminal portion of the target protein; the first dimerization domain comprises a first RNA hairpin, the first RNA hairpin comprising RNA complementary sequences separated by regions of non-complementary RNA sequence, wherein base pairing between the complementary sequences forms a first stem, and the regions of non-complementary RNA sequence form a first loop; the second dimerization domain comprises a second RNA hairpin, the second RNA hairpin comprising an RNA complementary sequence separated by a region of non-complementary RNA sequence, wherein base pairing between the complementary sequences forms a second stem, and the region of non-complementary RNA sequence of the second RNA hairpin forms a second loop; The composition, wherein the first loop hybridizes to the second loop.

2. The composition of claim 1 , wherein the first dimerization domain and the second dimerization domain are linked by a direct bond, an indirect bond, or a combination thereof.

3. The composition of claim 2 , wherein the direct or indirect bond comprises a base-pairing interaction, a non-canonical base-pairing interaction, a non-base-pairing interaction, or a combination thereof.

4. 4. The composition of claim 2 or 3, wherein the direct binding comprises a non-canonical base pairing interaction, a non-canonical base pairing interaction, a non-base pairing interaction, or a combination thereof between the aptamer regions.

5. The composition of claim 2 , wherein the indirect binding comprises a non-base pairing interaction between the aptamer and the target of the aptamer or between two aptamers.

6. The composition of any one of claims 1 to 5, wherein the first dimerization domain or the second dimerization domain does not comprise a cryptic splice acceptor.

7. 7. The composition of claim 1, wherein the dimerization domain is directly or indirectly bound to an aptamer sequence dimerization domain.

8. 8. The composition of claim 1, wherein the target protein is a disease-associated protein or a therapeutic protein.

9. The composition of claim 8 , wherein the disease is a monogenic disease.

10. The composition of claim 9 , wherein the therapeutic protein is a toxin.

11. 11. The composition of any one of claims 8 to 10, wherein the disease and the target protein are listed in Table 1.

12. the first RNA molecule further comprises one or both of a downstream intron splice enhancer (DISE) 3′ to the splice donor and 5′ to the first dimerization domain, and an intron splice enhancer (ISE) 3′ to the splice donor and 5′ to the first dimerization domain; and / or the second RNA molecule further comprises one or both of an ISE 3′ to the second dimerization domain and 5′ to the branch point sequence, and a DISE 3′ to the splice donor and 5′ to the second dimerization domain; or any combination thereof.

13. The first RNA molecule further comprises (ii-2) a DISE, an ISE, or both, wherein (ii-2) is 5' to the first dimerization domain and 3' to the splice donor; 12. The composition of any one of claims 1 to 11, wherein the second RNA molecule further comprises (i-2) at least one ISE sequence, wherein (i-2) is 5' to the branch point sequence and 3' to the second dimerization domain.

14. The first RNA molecule further comprises (ii-2) a DISE, a first ISE, and a second ISE, wherein (ii-2) is 5' to the first dimerization domain and 3' to the splice donor; 12. The composition of any one of claims 1 to 11, wherein the second RNA molecule further comprises (i-2) three ISE sequences, the (i-2) being 5' to the branch point sequence and 3' to the second dimerization domain.

15. the first RNA molecule further comprises a self-cleaving RNA sequence or an RNA cleavage enzyme target sequence located anywhere 3′ to the splice donor, such that the sequence cleaves the 3′-located polyadenylation tail, thereby reducing or suppressing expression of the protein fragment from the unrecombined RNA molecule; the second RNA molecule further comprises a self-cleaving RNA sequence or an RNA cleavage enzyme target sequence located anywhere 5′ to the branch point sequence, thereby cleaving the 5′-located RNA cap and reducing or suppressing expression of the protein fragment from the unrecombined RNA molecule; the second RNA molecule further comprises a start codon anywhere 5' to the branchpoint sequence that is shifted relative to the open reading frame that is 3' to the splice acceptor, thereby reducing or preventing translation of the target protein fragment from the unrecombined RNA molecule; the first RNA molecule further comprises a microRNA target site anywhere 3′ to the splice donor, such that the unbound RNA fragment undergoes microRNA-dependent degradation upon exiting the nucleus; the second RNA molecule further comprises a microRNA target site anywhere 3′ to the coding sequence, such that the unbound RNA fragment undergoes microRNA-dependent degradation upon exiting the nucleus; the first RNA molecule further comprises a sequence encoding a degron proteolytic tag anywhere 3′ to the splice donor and in frame with the open reading frame of the target protein 5′ to the splice donor site, such that unbound protein fragments are tagged for degradation; the second RNA molecule further comprises a start codon and an in-frame degron proteolytic tag anywhere 5′ to the branchpoint sequence and in-frame with the open reading frame of the target protein 3′ to the splice acceptor site, such that unbound protein fragments are tagged for degradation; or any combination thereof.

16. 16. The composition of any one of claims 1 to 15, wherein each promoter is independently selected.

17. 17. The composition of claim 16, wherein the first promoter and the second promoter are the same promoter or the first promoter and the second promoter are different promoters.

18. 18. The composition of any one of claims 1 to 17, wherein each of the first promoter and the second promoter is independently selected from a constitutive promoter; a tissue-specific promoter; and a promoter endogenous to the target protein.

19. 20. A system for expressing a target protein comprising the composition of any one of claims 1 to 18.

20. 20. The system of claim 19, wherein when introduced into a cell, the RNA molecules are produced and recombined in the appropriate order, thereby resulting in a full-length coding sequence for the target protein.

21. A system described in claim 19 or 20, wherein each of the first DNA molecule and the second DNA molecule is present in a separate viral vector.

22. 22. The system of claim 21, wherein the viral vector is AAV.

23. each of the DNA molecules is selected from the group consisting of 2,500nt to 5,000nt, 2,500nt to 2,750nt, 2,500nt to 3,000nt, 2,500nt to 3,250nt, 2,500nt to 3,500nt, 2,500nt to 3,750nt, 2,500nt to 4,000nt, 2,500nt to 4,250nt, 2,500nt to 4,500nt, 2,500nt to 4,750nt, 2,500nt to 5,000nt, 2,750nt to 3,000nt, 2,750nt to 3,250nt, 2,750nt to 3,500nt, 2,750nt to 3,750nt, 0nt, 2,750nt to 4,000nt, 2,750nt to 4,250nt, 2,750nt to 4,500nt, 2,750nt to 4, 750nt, 2,750nt to 5,000nt, 3,000nt to 3,250nt, 3,000nt to 3,500nt, 3,000nt to 3,750nt, 3,000nt to 4,000nt, 3,000nt to 4,250nt, 3,000nt to 4,500nt, 3,000n t~4,750nt, 3,000nt~5,000nt, 3,250nt~3,500nt, 3,250nt~3,750nt, 3,25 0nt to 4,000nt, 3,250nt to 4,250nt, 3,250nt to 4,500nt, 3,250nt to 4,750nt, 3, 250nt to 5,000nt, 3,500nt to 3,750nt, 3,500nt to 4,000nt, 3,500nt to 4,250nt, 3,500nt to 4,500nt, 3,500nt to 4,750nt, 3,500nt to 5,000nt, 3,750nt to 4,000n t, 3,750nt to 4,250nt, 3,750nt to 4,500nt, 3,750nt to 4,750nt, 3,750nt to 5,00 0nt, 4,000nt to 4,250nt, 4,000nt to 4,500nt, 4,000nt to 4,750nt, 4,000nt to 5, 000nt, 4,250nt to 4,500nt, 4,250nt to 4,750nt, 4,250nt to 5,000nt, 4,500nt to 4,750nt, 4,500nt to 5,000nt, 4,750nt to 5,000nt, 2,500nt, 2,750nt, 3,000nt, 3,250nt, 3,500nt, 3,750nt, 4,000nt, 4,250nt, 4,500nt, 4,750nt, and 5,23. The system according to any one of claims 19 to 22, having a size independently selected from 10,000 nt.

24. The coding sequence of the N-terminal portion of the target protein or the C-terminal portion of the target protein is 2,500 to 4,500 nt, 2,500 nt to 2,750 nt, 2,500 nt to 3,000 nt, 2,500 nt to 3,250 nt, 2,500 nt to 3,500 nt, 2,500 nt to 3,750 nt, 2,500 nt to 4,000 nt, 2,500 nt to 4,250 nt, 2,500 nt to 4,500 nt, nt, 2,750nt to 3,000nt, 2,750nt to 3,250nt, 2,750nt to 3,500nt, 2,750nt to 3,750nt, 2,750nt to 4,000nt, 2,750nt to 4,2 50nt, 2,750nt to 4,500nt, 3,000nt to 3,250nt, 3,000nt to 3,500nt, 3,000nt to 3,750nt, 3,000nt to 4,000nt, 3,000nt to 4, 250nt, 3,000nt to 4,500nt, 3,250nt to 3,500nt, 3,250nt to 3,750nt, 3,250nt to 4,000nt, 3,250nt to 4,250nt, 3,250nt to 4,500nt, 3,500nt to 3,750nt, 3,500nt to 4,000nt, 3,500nt to 4,250nt, 3,500nt to 4,500nt, 3,750nt to 4,000nt, 3,750nt 24. The system of any one of claims 19 to 23, having a size independently selected from: to 4,250 nt, 3,750 nt to 4,500 nt, 4,000 nt to 4,250 nt, 4,000 nt to 4,500 nt, 4,250 nt to 4,500 nt, 2,500 nt, 2,750 nt, 3,000 nt, 3,250 nt, 3,500 nt, 3,750 nt, 4,000 nt, 4,250 nt, and 4,500 nt.

25. Any one or both of the RNA molecules encoded by the DNA molecules of the system may each be 2,500 to 4,500 nt, 2,500 nt to 2,750 nt, 2,500 nt to 3,000 nt, 2,500 nt to 3,250 nt, 2,500 nt to 3,500 nt, 2,500 nt to 3,750 nt, 2,500 nt to 4,000 nt, 2,500 nt to 4,250 nt, 2,500 nt to 4,500 nt , 2,750nt to 3,000nt, 2,750nt to 3,250nt, 2,750nt to 3,500nt, 2,750nt to 3,750nt, 2,750nt to 4,000nt, 2,750nt to 4,250 nt, 2,750nt to 4,500nt, 3,000nt to 3,250nt, 3,000nt to 3,500nt, 3,000nt to 3,750nt, 3,000nt to 4,000nt, 3,000nt to 4,2 50nt, 3,000nt to 4,500nt, 3,250nt to 3,500nt, 3,250nt to 3,750nt, 3,250nt to 4,000nt, 3,250nt to 4,250nt, 3,250nt to 4 , 500nt, 3,500nt to 3,750nt, 3,500nt to 4,000nt, 3,500nt to 4,250nt, 3,500nt to 4,500nt, 3,750nt to 4,000nt, 3,750nt 25. The system of any one of claims 19 to 24, having a size independently selected from: to 4,250 nt, 3,750 nt to 4,500 nt, 4,000 nt to 4,250 nt, 4,000 nt to 4,500 nt, 4,250 nt to 4,500 nt, 2,500 nt, 2,750 nt, 3,000 nt, 3,250 nt, 3,500 nt, 3,750 nt, 4,000 nt, 4,250 nt, and 4,500 nt.

26. The DNA molecule may be selected from the group consisting of 5,000nt to 10,000nt, 5,000nt to 5,500nt, 5,000nt to 6,000nt, 5,000nt to 6,500nt, 5,000nt to 7,000nt, 5,000nt to 7,500nt, 5,000nt to 8,000nt, 5,000nt to 8,500nt, 5,000nt to 9,000nt, 5,000nt to 9,500nt, 5,000nt to 10,000nt, 5,500nt to 6,000nt, 5,500nt to 6,500nt, 5,500nt to 7,000nt, 5,500nt to 7,500nt. nt, 5,500nt to 8,000nt, 5,500nt to 8,500nt, 5,500nt to 9,000nt, 5,500nt to 9, 500nt, 5,500nt to 10,000nt, 6,000nt to 6,500nt, 6,000nt to 7,000nt, 6,000nt ~7,500nt, 6,000nt~8,000nt, 6,000nt~8,500nt, 6,000nt~9,000nt, 6,000nt nt ~ 9,500nt, 6,000nt ~ 10,000nt, 6,500nt ~ 7,000nt, 6,500nt ~ 7,500nt, 6, 500nt ~ 8,000nt, 6,500nt ~ 8,500nt, 6,500nt ~ 9,000nt, 6,500nt ~ 9,500nt , 6,500nt to 10,000nt, 7,000nt to 7,500nt, 7,000nt to 8,000nt, 7,000nt to 8,50 0nt, 7,000nt to 9,000nt, 7,000nt to 9,500nt, 7,000nt to 10,000nt, 7,500nt to 8 ,000nt, 7,500nt to 8,500nt, 7,500nt to 9,000nt, 7,500nt to 9,500nt, 7,500nt ~10,000nt, 8,000nt~8,500nt, 8,000nt~9,000nt, 8,000nt~9,500nt, 8,00nt 0nt to 10,000nt, 8,500nt to 9,000nt, 8,500nt to 9,500nt, 8,500nt to 10,000nt, 9,000nt to 9,500nt, 9,000nt to 10,000nt, 9,500nt to 10,000nt, 5,000nt, 5,50 0nt, 6,000nt, 6,500nt, 7,000nt, 7,500nt, 8,000nt, 8,500nt, 9,000nt, 9,500 nt, and 10,000 nt, The total target protein coding sequence is 2,000nt to 8,000nt, 2,000nt to 3,000nt, 2,000nt to 3,500nt, 2,000nt to 4,000nt, 2,000nt to 4,500nt, 2,000nt to 5,000nt, 2,000nt to 5,500nt, 2,000nt to 6,000nt, 2,000nt to 6,500nt, 2,000nt to 7,000nt, 2,000nt to 7,500nt, 2,000nt to 8,000nt, 3,000nt to 3,500nt, 3,000nt to 4,000nt, 3,000nt to 4,500nt, 0nt, 3,000nt to 5,000nt, 3,000nt to 5,500nt, 3,000nt to 6,000nt, 3,000nt to 6, 500nt, 3,000nt to 7,000nt, 3,000nt to 7,500nt, 3,000nt to 8,000nt, 3,500nt to 4 ,000nt, 3,500nt to 4,500nt, 3,500nt to 5,000nt, 3,500nt to 5,500nt, 3,500nt ~6,000nt, 3,500nt~6,500nt, 3,500nt~7,000nt, 3,500nt~7,500nt, 3,500nt t ~ 8,000nt, 4,000nt ~ 4,500nt, 4,000nt ~ 5,000nt, 4,000nt ~ 5,500nt, 4,00 0nt ~ 6,000nt, 4,000nt ~ 6,500nt, 4,000nt ~ 7,000nt, 4,000nt ~ 7,500nt, 4,0 00nt to 8,000nt, 4,500nt to 5,000nt, 4,500nt to 5,500nt, 4,500nt to 6,000nt, 4 , 500nt to 6,500nt, 4,500nt to 7,000nt, 4,500nt to 7,500nt, 4,500nt to 8,000nt, 5,000nt to 5,500nt, 5,000nt to 6,000nt, 5,000nt to 6,500nt, 5,000nt to 7,000n t, 5,000nt to 7,500nt, 5,000nt to 8,000nt, 5,500nt to 6,000nt, 5,500nt to 6,500 nt, 5,500nt to 7,000nt, 5,500nt to 7,500nt, 5,500nt to 8,000nt, 6,000nt to 6,5 00nt, 6,000nt to 7,000nt, 6,000nt to 7,500nt, 6,000nt to 8,000nt, 6,500nt to 7,000nt, 6,500nt to 7,500nt, 6,500nt to 8,000nt, 7,000nt to 7,500nt, 7,000nt to 8,000nt, or 7,500nt to 8,000nt, the total target protein coding sequence is 2,000nt, 3,000nt, 3,500nt, 4,000nt, 4,500nt, 5,000nt, 5,500nt, 6,000nt, 6,500nt, 7,000nt, 7,500nt, and 8,000nt; and / or The total size of the RNA molecule encoded by the first DNA molecule and the second DNA molecule is 5,000nt to 9,000nt, 5,000nt to 5,500nt, 5,000nt to 6,000nt, 5,000nt to 6,500nt, 5,000nt to 7,000nt, 5,000nt to 7,500nt, 5,000nt to 8,000nt, 5,000nt to 8,500nt, 5,000nt to 9,000nt. 00nt, 5,500nt to 6,000nt, 5,500nt to 6,500nt, 5,500nt to 7,000nt, 5,500nt to 7,500nt, 5,500nt to 8,000nt, 5,500nt to 8,500nt, 5,500nt to 9,000nt, 6,000nt to 6,500nt, 6,000nt to 7,000nt, 6,000nt to 7,500nt, 6,000nt to 8,000nt, 6,000nt nt ~ 8,500nt, 6,000nt ~ 9,000nt, 6,500nt ~ 7,000nt, 6,500nt ~ 7,500nt, 6,500nt ~ 8,000nt, 6,500nt ~ 8,500nt, 6, 500nt to 9,000nt, 7,000nt to 7,500nt, 7,000nt to 8,000nt, 7,000nt to 8,500nt, 7,000nt to 9,000nt, 7,500nt to 8,000nt , 7,500nt to 8,500nt, 7,500nt to 9,000nt, 8,000nt to 8,500nt, 8,000nt to 9,000nt, 8,500nt to 9,000nt, 5,000nt, 5,500nt, 6,000nt, 6,500nt, 7,000nt, 7,500nt, 8,000nt, 8,500nt, and 9,000nt.

27. Each of the first dimerization domain and the second dimerization domain is 1000 nt or less, for example, at least 50 nt, at least 100 nt, at least 150 nt, at least 200 nt, at least 300 nt, at least 400 nt, at least 500 nt, 50-1000 nt, 50-500 nt, 50-150 nt, 50 nt, 100 nt, 150 nt, 200 nt, 250 nt, 300 nt, 400 nt, or 27. The system of any one of claims 19 to 26, wherein the recombination efficiency of the system is at least 20%, at least 25%, at least 30%, at least 35%, at least 40%, at least 45%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, or 100%.

28. 28. The system of any one of claims 19 to 27, wherein each dimerization domain is 1000nt or less, for example, at least 50nt, at least 100nt, at least 150nt, at least 200nt, at least 300nt, at least 400nt, at least 500nt, 50-1000nt, 50-500nt, 50-150nt, 50nt, 100nt, 150nt, 200nt, 250nt, 300nt, 400nt, or 500nt; and wherein the recombination efficiency of the system is at least 20%, at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, at least 75%, at least 80%, at least 90%, or 100%.

29. RNA recombination efficiency is 10% to 100%, 10% to 20%, 10% to 30%, 10% to 35%, 10% to 40%, 10% to 45%, 10% to 50%, 10% to 55%, 10% to 60%, 10% to 70%, 10% to 80%, 10% to 90%, 20% to 30%, 20% to 35%, 20% to 40%, 20% to 45%, 20% to 50%, 20% to 55%, 20% ~60%, 20%~70%, 20%~80%, 20%~90%, 30%~35%, 30%~40%, 30%~45%, 30%~50%, 30%~55%, 30%~60%, 30%~70%, 30%~80%, 30%~90%, 35%~40%, 35%~45%, 35%~50%, 35%~55%, 35%~60%, 35%~70%, 35%~80%, 35 % to 90%, 40% to 45%, 40% to 50%, 40% to 55%, 40% to 60%, 40% to 70%, 40% to 80%, 40% to 90%, 45% to 50%, 45% to 55%, 45% to 60%, 45% to 70%, 45% to 80%, 45% to 90%, 50% to 55%, 50% to 60%, 50% to 70%, 50% to 80%, 50% to 90%, 55% to 60%, 29. The system of any one of claims 19 to 28, wherein the solubility of ...

30. 30. A composition comprising a system according to any one of claims 19 to 29.

31. The composition of claim 30, wherein each of the first RNA molecule and the second RNA molecule encodes at least a portion of dystrophin, factor VIII, ABCA4, or MYO7A.

32. 32. A kit comprising the system of any one of claims 19 to 29 or the composition of claim 30 or 31, wherein either the first DNA molecule or the second DNA molecule may be contained in separate containers, and optionally further comprising a buffer such as a pharmaceutically acceptable carrier.

33. 32. A system according to any one of claims 19 to 29 or a composition according to claim 30 or 31 for use in a method for expressing a target protein in a cell, said method comprising: A system or composition comprising the steps of introducing the system or composition into a cell and expressing the first RNA molecule and the second RNA molecule in the cell, wherein the target protein is produced in the cell.

34. 34. The system or composition of claim 33, wherein the cells are present in a subject and the introducing step comprises administering the system to the subject in a therapeutically effective amount.

35. 35. The system or composition of claim 34, which treats a genetic disease caused by a mutation in a gene encoding said target protein in said subject, resulting in expression of a functional target protein in said subject.

36. the genetic disease is Duchenne muscular dystrophy and the target protein is dystrophin; the genetic disease is hemophilia A and the target protein is F8; the genetic disease is Stargardt disease and the target protein is ABCA4; or 36. The system or composition of claim 35, wherein the genetic disease is Usher syndrome and the target protein is MYO7A.

37. 37. The system of any one of claims 19 to 29 and 33 to 36, or the composition of any one of claims 1 to 18, 30, 31 and 33 to 36, wherein one or both of the first and second RNA molecules comprise at least 80%, at least 85%, at least 90%, at least 95%, at least 98%, at least 99%, or 100% sequence identity to the synthetic intron presented in any one of SEQ ID NOs: 1, 2, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 20, 21, 22, 23, 24, 25, 145, 146, 147, 148, 155, 156, 157, 158, 159, 160, 161, 162, 163, 164, 165, and 166.

38. 38. The system of any one of claims 19 to 29 and 33 to 37, or the composition of any one of claims 1 to 18, 30, 31 and 33 to 37, wherein one or both of the first and second RNA molecules comprises a synthetic intron selected from nt 3703 to 3975 of SEQ ID NO:20, nt 1 to 228 of SEQ ID NO:21, nt 3703 to 3975 of SEQ ID NO:22, nt 1 to 225 of SEQ ID NO:23, nt 3560 to 3828 of SEQ ID NO:24, and nt 1 to 225 of SEQ ID NO:

25.

39. 39. The system of any one of claims 19 to 29 and 33 to 38, or the composition of any one of claims 1 to 18, 30, 31 and 33 to 38, wherein one or both of the first RNA molecule and the second RNA molecule further comprises a portion of a protein coding sequence.

40. 40. The system or composition of claim 39, wherein the portion of the protein coding sequence comprises the N-terminal half, N-terminal portion, C-terminal half, or C-terminal portion of the protein coding sequence.

41. Either one or two of the first RNA molecule and the second RNA molecule are each selected from the group consisting of 2,500 to 4,500 nt, 2,500 nt to 2,750 nt, 2,500 nt to 3,000 nt, 2,500 nt to 3,250 nt, 2,500 nt to 3,500 nt, 2,500 nt to 3,750 nt, 2,500 nt to 4,000 nt, 2,500 nt to 4,250 nt, 2,500 nt to 4,500 nt, 2 , 750nt to 3,000nt, 2,750nt to 3,250nt, 2,750nt to 3,500nt, 2,750nt to 3,750nt, 2,750nt to 4,000nt, 2,750nt to 4,250nt , 2,750nt to 4,500nt, 3,000nt to 3,250nt, 3,000nt to 3,500nt, 3,000nt to 3,750nt, 3,000nt to 4,000nt, 3,000nt to 4,250 nt, 3,000nt to 4,500nt, 3,250nt to 3,500nt, 3,250nt to 3,750nt, 3,250nt to 4,000nt, 3,250nt to 4,250nt, 3,250nt to 4,5 00nt, 3,500nt to 3,750nt, 3,500nt to 4,000nt, 3,500nt to 4,250nt, 3,500nt to 4,500nt, 3,750nt to 4,000nt, 3,750nt to 4 19. The composition of claim 1, wherein the nucleotides have a size independently selected from the following: 2,250 nt, 3,750 nt to 4,500 nt, 4,000 nt to 4,250 nt, 4,000 nt to 4,500 nt, 4,250 nt to 4,500 nt, 2,500 nt, 2,750 nt, 3,000 nt, 3,250 nt, 3,500 nt, 3,750 nt, 4,000 nt, 4,250 nt, and 4,500 nt.

42. The size of the total target protein coding sequence is 2,000nt to 8,000nt, 2,000nt to 3,000nt, 2,000nt to 3,500nt, 2,000nt to 4,000nt, 2,000nt to 4,500nt, 2,000nt to 5,000nt, 2,000nt to 5,500nt, 2,000nt to 6,000nt, 2,000nt to 6,500nt, 2,000nt to 7,000nt, 2,000nt to 7,500nt, 2,000nt to 8,000nt, 3,000nt to 3,500nt, 3,000nt to 4,000nt, 3,000nt ~4,500nt, 3,000nt~5,000nt, 3,000nt~5,500nt, 3,000nt~6,000nt, 3,000nt nt~6,500nt, 3,000nt~7,000nt, 3,000nt~7,500nt, 3,000nt~8,000nt, 3,5 00nt to 4,000nt, 3,500nt to 4,500nt, 3,500nt to 5,000nt, 3,500nt to 5,500nt, 3 , 500nt to 6,000nt, 3,500nt to 6,500nt, 3,500nt to 7,000nt, 3,500nt to 7,500nt, 3,500nt to 8,000nt, 4,000nt to 4,500nt, 4,000nt to 5,000nt, 4,000nt to 5,500n t, 4,000nt to 6,000nt, 4,000nt to 6,500nt, 4,000nt to 7,000nt, 4,000nt to 7,50 0nt, 4,000nt to 8,000nt, 4,500nt to 5,000nt, 4,500nt to 5,500nt, 4,500nt to 6, 000nt, 4,500nt to 6,500nt, 4,500nt to 7,000nt, 4,500nt to 7,500nt, 4,500nt to 8 ,000nt, 5,000nt to 5,500nt, 5,000nt to 6,000nt, 5,000nt to 6,500nt, 5,000nt ~7,000nt, 5,000nt~7,500nt, 5,000nt~8,000nt, 5,500nt~6,000nt, 5,500 nt ~ 6,500nt, 5,500nt ~ 7,000nt, 5,500nt ~ 7,500nt, 5,500nt ~ 8,000nt, 6,0 00nt ~ 6,500nt, 6,000nt ~ 7,000nt, 6,000nt ~ 7,500nt, 6,000nt ~ 8,000nt, 6,500nt to 7,000nt, 6,500nt to 7,500nt, 6,500nt to 8,000nt, 7,000nt to 7,500nt, 7,000nt to 8,000nt, 7,500nt to 8,000nt, 2,000nt, 3,000nt, 3,500nt, 4,000nt, 4,500nt, 5,000nt, 5,500nt, 6,000nt, 6,500nt, 7,000nt, 7,500nt, or 8,000nt; and / or The total size of the two RNA molecules is 5,000nt to 9,000nt, 5,000nt to 5,500nt, 5,000nt to 6,000nt, 5,000nt to 6,500nt, 5,000nt to 7,000nt, 5,000nt to 7,500nt, 5,000nt to 8,000nt, 5,000nt to 8,500nt, 5,000nt to 9,000nt, 5,500nt to 6,000nt, 5 , 500nt to 6,500nt, 5,500nt to 7,000nt, 5,500nt to 7,500nt, 5,500nt to 8,000nt, 5,500nt to 8,500nt, 5,500nt to 9, 000nt, 6,000nt to 6,500nt, 6,000nt to 7,000nt, 6,000nt to 7,500nt, 6,000nt to 8,000nt, 6,000nt to 8,500nt, 6,0 00nt ~ 9,000nt, 6,500nt ~ 7,000nt, 6,500nt ~ 7,500nt, 6,500nt ~ 8,000nt, 6,500nt ~ 8,500nt, 6,500nt ~ 9,00 0nt, 7,000nt to 7,500nt, 7,000nt to 8,000nt, 7,000nt to 8,500nt, 7,000nt to 9,000nt, 7,500nt to 8,000nt, 7,500 19. The composition of any one of claims 1 to 18, wherein the amino acid sequence of ...

43. each of the first dimerization domain and the second dimerization domain is 1000 nt or less, for example, at least 50 nt, at least 100 nt, at least 150 nt, at least 200 nt, at least 300 nt, at least 400 nt, at least 500 nt, 50-1000 nt, 50-500 nt, 50-150 nt, 50 nt, 100 nt, 150 nt, 200 nt, 250 nt, 300 nt, 400 nt, or 500 nt; and the recombination efficiency of the first RNA molecule and the second RNA molecule is 19. The composition of any one of claims 1 to 18, wherein the solubility of ...

44. 19. The composition of any one of claims 1 to 18, wherein each dimerization domain is 1000nt or less, e.g., at least 50nt, at least 100nt, at least 150nt, at least 200nt, at least 300nt, at least 400nt, at least 500nt, 50-1000nt, 50-500nt, 50-150nt, 50nt, 100nt, 150nt, 200nt, 250nt, 300nt, 400nt, or 500nt; and wherein the recombination efficiency of the first RNA molecule and the second RNA molecule is at least 20%, at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, at least 75%, at least 80%, or at least 90%.

45. RNA recombination efficiency is 10% to 100%, 10% to 20%, 10% to 30%, 10% to 35%, 10% to 40%, 10% to 45%, 10% to 50%, 10% to 55%, 10% to 60%, 10% to 70%, 10% to 80%, 10% to 90%, 20% to 30%, 20% to 35%, 20% to 40%, 20% to 45%, 20% to 50%, 20% to 55%, 20% ~60%, 20%~70%, 20%~80%, 20%~90%, 30%~35%, 30%~40%, 30%~45%, 30%~50%, 30%~55%, 30%~60%, 30%~70%, 30%~80%, 30%~90%, 35%~40%, 35%~45%, 35%~50%, 35%~55%, 35%~60%, 35%~70%, 35%~80%, 35 % to 90%, 40% to 45%, 40% to 50%, 40% to 55%, 40% to 60%, 40% to 70%, 40% to 80%, 40% to 90%, 45% to 50%, 45% to 55%, 45% to 60%, 45% to 70%, 45% to 80%, 45% to 90%, 50% to 55%, 50% to 60%, 50% to 70%, 50% to 80%, 50% to 90%, 55% to 60%, 5 19. The composition of any one of claims 1 to 18, wherein the saturation of ...

46. (a) each of the first RNA molecule and the second RNA molecule is between 2500 nt and 4500 nt; (b) the size of the total target protein coding sequence is between 2000 nt and 8000 nt; and / or (c) the total size of the two RNA molecules is between 5,000 nt and 9,000 nt; 19. The composition of any one of claims 1 to 18, wherein the RNA recombination efficiency is greater than 10%.

Citation Information

Patent Citations

  • Polypeptide expression system

    JP2017520255A

  • Use of RNA trans-splicing for generation of interfering RNA molecules

    US20060134658A1