Methods and compositions for expression of edited proteins

JP2024517939A5Pending Publication Date: 2025-05-22SALK INST FOR BIOLOGICAL STUDIES
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2023569888
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2021-05-14
Filing Date
2022-05-16
Publication Date
2025-05-22

AI Technical Summary

Technical Problem

Existing gene therapy methods face challenges in delivering large proteins due to packaging constraints of vectors like AAV, leading to inefficient delivery of truncated proteins or toxicity, limiting treatment of genetic diseases.

Method used

A method involving the expression of full-length nucleic acid editing proteins using two or more synthetic RNA molecules, each encoding a portion of the protein, which are introduced into the same cell to facilitate targeted editing and recombination, allowing for the expression of large proteins.

Benefits of technology

This approach enables safe and efficient expression of full-length proteins, overcoming packaging limitations and addressing genetic diseases by enabling targeted nucleic acid editing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

Provided herein are compositions and systems for reconstituting RNA molecules, including methods for using RNA molecules.For example, such molecules can be used to deliver protein coding sequences through two or more viral vectors (e.g., AAV), thereby resulting in the reconstitution of full-length proteins in cells.Such methods can be used to deliver proteins involved in editing nucleic acid molecules, for example, to treat genetic disease or cancer.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims priority to U.S. Provisional Patent Application No. 63 / 189,048, filed May 14, 2021, which is incorporated herein by reference in its entirety.

[0002] Field The present disclosure provides systems, kits, compositions, and methods that allow for the joining of two or more RNA molecules, thereby allowing for the expression of a full-length protein, such as a protein involved in gene editing, e.g., a Cas nuclease or a catalytically inactive form of a Cas nuclease. [Background technology]

[0003] background Gene therapy is a promising method for treating genetic diseases caused by loss-of-function mutations. Replacement genes are typically reintroduced into target cells using vectors such as AAV, because this virus is generally safe and efficient at entering cells. However, in the case of AAV, it is difficult to encapsulate more than about 5,000 nucleotides using conventional capsids. The length of genes encoding large proteins often exceeds the packaging constraints of AAV, and many genetic diseases remain untreatable. Previously, strategies to overcome this limitation have been explored, but they have proven inefficient, resulted in high-level expression of potentially toxic truncated proteins, or both. A safe and highly efficient strategy for delivering large proteins to treat diseases is needed. Summary of the Invention [Means for solving the problem]

[0004] overview Provided herein is a composition for expressing target proteins, such as proteins used to edit nucleic acid sequences (e.g., target DNA or RNA, such as genes).Included are compositions and methods for expressing nucleic acid editing proteins produced from two or more synthetic nucleic acid molecules that are introduced into the same cell separately.Using this strategy, full-length nucleic acid editing proteins and one or more guide RNAs can be provided to the same cell, thereby resulting in targeted nucleic acid editing.Cells can be those that require targeted nucleic acid editing, for example, to repair mutations in essential genes. In one example, the composition comprises: (a) a first RNA molecule comprising, from 5' to 3', (i) a coding sequence for the N-terminal portion of a nucleic acid editing protein; (ii) a splice donor; and (iii) a first dimerization domain; and (b) a second RNA molecule comprising, from 5' to 3', (i) a second dimerization domain that binds to the first dimerization domain; (ii) a branch point sequence; (iii) a polypyrimidine tract; (iv) a splice acceptor; and (v) a coding sequence for the C-terminal portion of a nucleic acid editing protein.In some examples, such compositions comprise: (c) a third RNA molecule comprising at least one first guide RNA (gRNA) specific to the first target nucleic acid molecule, wherein the at least one first gRNA directs a nucleic acid editing protein to a target editing site on the first target nucleic acid molecule; (d) (i) at least one second gRNA specific to the first target nucleic acid molecule that directs the nucleic acid editing protein to the same or a different target editing site on the first nucleic acid molecule, or (ii) a third RNA molecule comprising at least one second gRNA specific to the first target nucleic acid molecule that directs the nucleic acid editing protein to the same or a different target editing site on the first nucleic acid molecule. a fourth RNA molecule comprising at least one second gRNA specific to a second target nucleic acid molecule that directs the nucleic acid editing protein to a target editing site on the first nucleic acid molecule that is the same or different from the first gRNA and the second gRNA; (e) (i) at least one third gRNA specific to the first target nucleic acid molecule that directs the nucleic acid editing protein to a target editing site on the first nucleic acid molecule that is the same or different from the first gRNA and the second gRNA; (ii) at least one third gRNA specific to the second target nucleic acid molecule that directs the nucleic acid editing protein to a target editing site on the second nucleic acid molecule that is the same or different from the second gRNA; or (iii) a fifth RNA molecule comprising at least one third gRNA specific to a third target nucleic acid molecule that directs the nucleic acid editing protein to a target editing site on the third target nucleic acid molecule; and / or (f)(i) at least one fourth gRNA specific to the first target nucleic acid molecule that directs the nucleic acid editing protein to a target editing site that is the same or different from the first gRNA, second gRNA, and third gRNA on the first nucleic acid molecule, (ii) a fifth RNA molecule comprising at least one fourth gRNA specific to the first target nucleic acid molecule that directs the nucleic acid editing protein to a target editing site that is the same or different from the first gRNA, second gRNA, and third gRNA on the second nucleic acid molecule. (iii) at least one fourth gRNA specific to a second target nucleic acid molecule that directs the nucleic acid editing protein to the same or a different target editing site as the third gRNA on the third target nucleic acid molecule; or (iv) a sixth RNA molecule comprising at least one fourth gRNA specific to a fourth target nucleic acid molecule that directs the nucleic acid editing protein to a target editing site on the fourth target nucleic acid molecule.

[0005] In one example, the composition comprises: (a) a first RNA molecule comprising, from 5' to 3', (i) a coding sequence for the N-terminal portion of a nucleic acid editing protein; (ii) a splice donor; and (iii) a first dimerization domain; and (b) a second RNA molecule comprising, from 5' to 3', (i) a second dimerization domain that binds to the first dimerization domain; (ii) a branch point sequence; (iii) a polypyrimidine tract; (iv) a splice acceptor; and (v) a coding sequence for the C-terminal portion of a nucleic acid editing protein. In some examples, such compositions include: (c) a third RNA molecule and a fourth RNA molecule comprising at least one first crisprRNA (crRNA) and at least one first tracrRNA, each specific to a first target nucleic acid molecule, wherein the at least one first crRNA and the at least one first tracrRNA direct a nucleic acid editing protein to a target editing site on the first target nucleic acid molecule; (d)(i) a third RNA molecule and a fourth RNA molecule comprising at least one second crRNA and at least one second tracrRNA, each specific to the first target nucleic acid molecule, wherein the at least one (ii) a fifth RNA molecule and a sixth RNA molecule, each comprising at least one second crRNA specific to a second target nucleic acid molecule, and wherein the at least one second crRNA and the at least one second tracrRNA direct the nucleic acid editing protein to a target editing site on the second target nucleic acid molecule that is the same as or different from the first crRNA and the first tracrRNA and the second crRNA and the second tracrRNA; or (iii) a fifth RNA molecule and a sixth RNA molecule, each comprising at least one second crRNA specific to a second target nucleic acid molecule, and wherein the at least one second crRNA and the at least one second tracrRNA direct the nucleic acid editing protein to a target editing site on the second target nucleic acid molecule;(e) (i) comprising at least one third cRNA and at least one third tracrRNA, each specific to a first target nucleic acid molecule, wherein the at least one third crRNA and at least one third tracrRNA directs a nucleic acid editing protein to the same or different target editing site on the first nucleic acid molecule as the first crRNA and first tracrRNA and second crRNA and second tracrRNA, or (ii) comprising at least one third cRNA specific to a second target nucleic acid molecule, and at least (iii) a seventh RNA molecule and an eighth RNA molecule, each comprising at least one third crRNA and at least one third tracrRNA that direct a nucleic acid editing protein to the same or different target editing site on the second nucleic acid molecule as the second crRNA and second tracrRNA, or (iv) at least one third cRNA specific to a third target nucleic acid molecule, and at least one third crRNA and at least one third tracrRNA that direct a nucleic acid editing protein to the target editing site on the third target nucleic acid molecule;and / or (f) (i) comprising at least one fourth crRNA and at least one fourth tracrRNA specific to a first target nucleic acid molecule, respectively, wherein the at least one fourth crRNA and at least one fourth tracrRNA directs a nucleic acid editing protein to the same or different target editing site on the first nucleic acid molecule as the first crRNA and first tracrRNA, the second crRNA and second tracrRNA, and the third crRNA and third tracrRNA; or (ii) comprising at least one fourth crRNA specific to a second target nucleic acid molecule, wherein the at least one fourth crRNA and at least one fourth tracrRNA directs a nucleic acid editing protein to the second crRNA and second tracrRNA on the second nucleic acid molecule; and a third crRNA and a third tracrRNA that direct the nucleic acid editing protein to the same or a different target editing site as the third crRNA and the third tracrRNA, or (iii) at least one fourth crRNA specific to a third target nucleic acid molecule, wherein the at least one fourth crRNA and the at least one fourth tracrRNA direct the nucleic acid editing protein to the same or a different target editing site on the third target nucleic acid molecule as the third crRNA and the third tracrRNA, or (iv) at least one fourth crRNA specific to a fourth target nucleic acid molecule, wherein the at least one fourth crRNA and the at least one fourth tracrRNA direct the nucleic acid editing protein to the target editing site on the fourth target nucleic acid molecule;

[0006] In some examples, the first dimerization domain and the second dimerization domain are linked by a direct bond, an indirect bond, or both.

[0007] In some examples, the dimerization domain is a kissing loop domain or a hypodiverse domain.

[0008] In some examples, the first RNA molecule and / or the second RNA molecule comprises at least one splice enhancer.

[0009] Compositions for expressing a target protein are also provided. Such compositions may include (i) a first synthetic DNA molecule encoding the RNA molecule of (a) and the RNA molecule of (c), or the RNA molecule of (a), the RNA molecule of (c), and the RNA molecule of (d); and (ii) a second synthetic DNA molecule encoding the RNA molecule of (b) and the RNA molecule of (e), or the RNA molecule of (b), the RNA molecule of (e), and the RNA molecule of (f). In some embodiments, the first synthetic DNA molecule includes (i) a first promoter operably linked to a sequence encoding the first RNA molecule; and (b) a second promoter operably linked to a sequence encoding the second RNA molecule.

[0010] Also provided are systems for expressing nucleic acid editing proteins comprising the described compositions.

[0011] Also provided is a method for expressing a nucleic acid editing protein in a cell, for example, by using the system disclosed herein or the RNA encoded by the system in combination with an appropriate guide nucleic acid molecule that hybridizes with a target nucleic acid molecule. Such a method can include introducing the system into a cell and expressing a first synthetic RNA molecule and a second synthetic RNA molecule in the same cell. In some embodiments, the cell is present in a subject, and the method treats a disease in the subject, such as a genetic disease caused by a mutation in a target DNA or RNA (e.g., a gene). In some embodiments, the genetic disease is Duchenne muscular dystrophy, hemophilia A, Stargardt's disease, or Usher syndrome.

[0012] The foregoing and other objects and features disclosed herein will become more apparent from the following detailed description which proceeds with reference to the accompanying drawings.

[0013] The patent or application file contains at least one drawing executed in color. Copies of this patent or patent application publication with color drawing(s) will be provided by the Office upon request and payment of the necessary fee. [Brief explanation of the drawings]

[0014] [Figure 1-1] Figure 1A shows a schematic diagram of the vector design (left) and RNA interaction and splicing (right). Left: 5' trans-splice (trsp) DNA vector; open arrows indicate two opposing promoters. The 3' UTR with the RFP-encoding domain and polyadenylation element is expressed in opposite directions from the N-terminal portion of YFP (n-yfp), followed by a splice donor sequence (SD), a downstream intronic splicing enhancer (DISE), two intronic splicing enhancers (2xISE), a binding domain (BD, also known as the dimerization domain), and a stable stem-loop box B element (box B), a self-cleaving hammerhead ribozyme (HHrz), and finally a polyadenylation element. A small intron has been inserted into the n-yfp segment (white segment within n-yfp). 3' trsp DNA vector; open arrows indicate two opposing promoters. The 3'UTR with the YFP-encoding domain and polyadenylation element is expressed in opposite directions from the 3'UTR, which contains a complementary binding domain (anti-BD, also called the dimerization domain), followed by three intronic splicing enhancer sequences (3xISE), a branch point (BP), a polypyrimidine tract (PPT), a splice acceptor sequence (SA), the C-terminal portion of the YFP-encoding sequence (c-terminal proton), and finally a polyadenylation element. Right: Pre-mRNA interactions (5'trsp-RNA + 3'trsp-RNA) and trans-splicing are shown, resulting in the generation of the YFP-encoding mRNA.

[0015] FIG. 1B shows that transfection of the N-terminal expression plasmid alone does not result in YFP fluorescence.

[0016] [Figure 1-2]FIG. 1C shows that transfection of the C-terminal expression plasmid alone does not result in YFP fluorescence.

[0017] FIG. 1D shows that expression of the N- and C-terminal fragments without the binding domain showed low levels of YFP induction.

[0018] [Figure 1-3] Figure 1E shows rationally designed dimerization / binding domains in a looped configuration (all-pyrimidine or all-purine low-diversity sequences interrupted by complementary sequences that form a double-stranded stem structure).

[0019] FIG. 1F shows a 3D rendering of the "looped" dimerization domain configuration.

[0020] [Figure 1-4] FIG. 1G shows a negative control lacking the binding domain in the C-terminal half.

[0021] FIG. 1H shows a negative control lacking the binding domain in the N-terminal half.

[0022] [Figure 1-5] FIG. 1I shows that matching binding domains in loop configurations in both the N-terminal and C-terminal halves exhibited strong YFP induction in 90% of cells.

[0023] Figures 1J-1N represent data equivalent to those in Figures 1E-1I for the construction of binding domains with a 150-nucleotide low-diversity sequence composed exclusively of pyrimidines (or alternatively exclusively of purines), containing sequences that result in a completely open configuration.

[0024] Figure 1J shows the fully open configuration for complimentary base pairing. The resulting 150 nucleotide low diversity pyrimidine sequence is shown.

[0025] [Figure 1-6] FIG. 1K shows a 3D rendering of the 150-nucleotide low-diversity pyrimidine sequence of (1J).

[0026] Figure 1L shows transfection of control HEK293T cells with a construct encoding a C-terminal YFP lacking the complementary low-diversity binding domain. Few transfected cells express YFP.

[0027] [Figure 1-7] Figure 1M shows transfection of control HEK293T cells with a construct encoding an N-terminal YFP lacking the complementary low diversity binding domain. Few transfected cells express YFP.

[0028] Figure 1N shows the transfection of HEK293T cells with N- and C-terminal YFP constructs, both of which have complementary low-divergence dimerization binding domains. Many cells express high levels of YFP.

[0029] [Figure 1-8] Figure 10 shows a representative fluorescence image for the cells shown in Figure 1G. Positive markers for transfection (RFP+BFP) are expressed, but YFP protein is not efficiently reconstituted.

[0030] Figure 1P shows a representative fluorescence image of the cells shown in Figure 1L. The positive markers for transfection (RFP + BFP) are expressed, and YFP protein is reconstituted at high levels in RFP and BFP double-positive cells.

[0031] [Figure 1-9] Figure 1Q shows a comparison of the conditions shown in Figure 1D, Figures 1G-1I, and Figures 1L-1N: N: no binding domain, Loop: looped low diversity binding domain configuration, Lin: linear low diversity configuration.

[0032] [Figure 2-1] Figure 2A shows a schematic diagram of the vector design. The protein coding sequence for yellow fluorescent protein (YFP) is divided into an N-terminus, a middle fragment (m-yfp), and a C-terminal fragment. The junction between the RNA encoding the n fragment and the RNA encoding the m fragment is connected by a loop-shaped binding domain (BD1), and the junction between the m fragment and the c fragment is connected by a loop-shaped binding domain (BD2). The pyrimidine (Y) and purine (R) sequences are positioned to prevent self-circularization of the m fragment and direct recombination between the N and C fragments. The N-terminal fragment is coexpressed with red fluorescent protein as a transfection control, and the C-terminal fragment is coexpressed with blue fluorescent protein as a transfection control. The promoter sequences are indicated by open arrows. The splice donor (SD) and splice acceptor (SA) sites are indicated. Intron splicing elements including a splice enhancer, polypyrimidine tract, and branch point similar to those used upstream (5') of the SA and downstream (3') of the SD in Figure 1A are included.

[0033] Figure 2B shows that transfection of plasmid I+II+III (see Figure 2A) into a human cell line efficiently reconstitutes high-level YFP expression in 80% of the transfected cells.

[0034] [Figure 2-2] FIG. 2C shows a representative fluorescence image of expression of the n and m fragments (plasmids I+II, see FIG. 2A) in which no yfp fluorescence is shown (negative control).

[0035] FIG. 2D shows a representative fluorescence image of expression of the m and c fragments (plasmids II+III, see FIG. 2A) in which no yfp fluorescence is shown (negative control).

[0036] FIG. 2E shows a representative fluorescence image demonstrating that co-transfection of all three fragments (plasmids I+II+III, see FIG. 2A) induces strong YFP fluorescence.

[0037] [Figure 3] Figures 3A-3D show efficient reconstitution of yellow fluorescent protein (YFP) from two fragments (SEQ ID NOs: 1 and 2) expressed from two AAV2 / 8 vectors after systemic administration to neonatal mice (P3). (A) AAV1 encodes the N-terminal half fragment of YFP, and AAV2 encodes the C-terminal half fragment. AAV1 and AAV2 were mixed at equal titers and intravenously injected into mice. Tissue samples were collected 3 weeks after injection. (B) YFP fluorescence in the liver of a young mouse at the time of sacrifice (green). An uninjected liver is shown for comparison (control: no YFP detected). DRAQ5 nuclear staining is shown in magenta for context. (C) Strong YFP fluorescence (green) in the myocardium at the time of sacrifice. The top panel shows the macroscopic image with red autofluorescence (magenta) for context. The bottom panel shows a cross-section (magenta) with DRAQ5 nuclear staining for context. An uninjected mouse heart lacking YFP is shown as a control. (D) Strong YFP fluorescence in skeletal muscle of the limb at the time of sacrifice. An uninjected mouse limb is shown for comparison (negative control, no YFP detected). The top panel shows a macroscopic image with magenta red autofluorescence. The bottom panel shows a microscopic image of a cross section through the limb. The bottom panel shows magenta DRAQ5 nuclear staining for context.

[0038] [Figure 4]Figures 4A-4B show efficient reconstitution of yellow fluorescent protein (YFP) from three fragments (SEQ ID NOs: 145, 146, and 2, respectively) in the tibialis anterior muscle of a mouse neonatal (P3) pup after intramuscular injection of three AAV2 / 8 vectors. (A) A schematic diagram of three AAV particles carrying the N-, M-, and C-terminal fragments of YFP is shown (similar to Figure 2A). (B) Strong YFP fluorescence in a longitudinal section of the tibialis anterior muscle of a mouse injected with all three viral particles is shown. DRAQ5 nuclear staining is shown in magenta for context.

[0039] [Figure 5] Figures 5A-5F show efficient reconstitution of yellow fluorescent protein (YFP) from two and three fragments in adult mouse tibialis anterior muscle. (A) The N- and C-terminal halves of the YFP coding sequence with synthetic RNA dimerization and recombination domains are shown. (B) Two AAV transfer plasmids expressing these two fragments were electroporated percutaneously into adult mouse tibialis anterior (TA) muscle, and strong fluorescence was detected 5 days after electroporation. (C) No fluorescence was detectable in the contralateral, uninjected TA. (D) The N-, middle, and C-termini of the YFP coding sequence with synthetic RNA dimerization and recombination domains are shown, with each fragment ligated to its adjacent fragment. (E) Transcutaneous electroporation of three AAV transfer plasmids expressing these three fragments is shown. Strong YFP fluorescence is detected, indicating efficient reconstitution of YFP from the three fragments. (F) Fluorescence in the contralateral, uninjected TA is shown. The fluorescence channel is overlaid on the grayscale image for context.

[0040] [Figure 6-1]6A is a schematic diagram presenting an exemplary system for the RNA recombination methods disclosed herein, using two nucleic acid molecules 110, 150, in which a target protein is split into two portions, each portion encoded by a different nucleic acid molecule. In some examples, the nucleic acid molecules 110, 150 of the system are DNA and include a promoter 112, 152. In some examples, the nucleic acid molecules 110, 150 of the system are RNA and therefore lack a promoter 112, 152. The drawings are not to scale.

[0041] [Figure 6-2] Figure 6B is a schematic diagram presenting exemplary dimerization domains (e.g., 122, 154 in Figure 6A) that contain low diversity sequences interspersed with sequences that can form stems, resulting in local RNA loops that are open and available for base pairing in the absence of pseudoknot formation. The drawing is not to scale.

[0042] [Figure 6-3] Figure 6C is a schematic diagram showing that the interaction and hybridization (base pairing) of pre-mRNA dimerization domain 122 of molecule 110 (Figure 6A) with pre-mRNA dimerization domain 154 of molecule 150 (Figure 6A) allows recombination of N-terminal coding sequence 114 and C-terminal coding sequence 164 by spliceosome components. This results in a seamless junction between the 3' end of the fused N-terminal protein coding sequence 114 and the 5' end of the C-terminal protein sequence 164, and between the N-terminal and C-terminal portions. The drawing is not to scale.

[0043] [Figure 6-4]Figure 6D is a schematic diagram presenting an exemplary system for the RNA recombination methods disclosed herein using three nucleic acid molecules 110, 200, 150, in which the target protein is divided into three portions (N-terminal, middle, and C-terminal), each encoded by a different nucleic acid molecule. Before transcription, the nucleic acid molecules 110, 150, 200 of the system are DNA and include promoters 112, 152, 202. After transcription, the nucleic acid molecules 110, 150, 200 of the system are RNA and therefore lack promoters 112, 152, 202. The drawing is not to scale.

[0044] [Figure 6-5] 6E is a schematic diagram illustrating that the interaction and hybridization (base pairing) between dimerization domain 122 of molecule 110 (FIG. 6D) and dimerization domain 204 of molecule 200 (FIG. 6D), as well as the interaction and hybridization (base pairing) between dimerization domain 226 of molecule 200 (FIG. 6D) and dimerization domain 154 of molecule 150 (FIG. 6D), allows spliceosome components to recombine N-terminal coding sequence 114, intermediate coding sequence 216, and C-terminal coding sequence 164. This results in seamless junctions between the fused 3' end of N-terminal coding sequence 114 and the 5' end of intermediate protein sequence 216, and between the fused 3' end of intermediate coding sequence 216 and the 5' end of C-terminal sequence 216, as well as between the N-terminal, intermediate, and C-terminal portions. In some examples, for example, after transcription, the elements shown are RNA. The drawings are not to scale.

[0045] [Figure 6-6]Figure 6F is a schematic diagram presenting an exemplary system for the RNA recombination methods disclosed herein, using two nucleic acid molecules 110, 150, in which the target protein is divided into two parts, each part encoded by a different nucleic acid molecule. In this example, DNA is transcribed into RNA, and therefore the nucleic acid molecules 110, 150 of the system are RNA and therefore lack the promoters 112, 152 present in DNA (see Figure 6A). The drawing is not to scale.

[0046] [Figure 6-7]6G is a schematic diagram presenting an exemplary composition or system for the RNA recombination methods disclosed herein using two DNA molecules 110, 150, where a nucleic acid editing protein is split into two portions 114, 164, each portion encoded by a different DNA molecule 110, 150. In some examples, a promoter 112, 152 drives expression of each coding sequence 114, 164. This example system includes one or more guide nucleic acid molecules (e.g., gRNAs or gRNA-coding sequences) 140, 141, 171, 172 specific for one or more target nucleic acid molecules. However, such guide nucleic acid molecules 140, 141, 171, 172 are optional and, in some examples, are provided separately, e.g., as part of separate vectors, instead of being provided as part of molecules 110, 150. In this example, optional guide nucleic acid molecules 140, 141, 171, 172 are shown near the 5' and 3' ends of nucleic acid molecules 110, 150, although the disclosure is not limited to such locations. When one or more guide nucleic acid molecules 140, 141, 171, 172 are present, their expression can be driven by promoters 142, 143, 173, 174. The system can optionally include parvoviral inverted terminal repeats (ITRs) 176, 177, 178, 179 at the 5' and 3' ends of molecules 110, 150, respectively. Although this figure illustrates an embodiment in which the molecules 110, 150 are DNA, in some examples (e.g., after transcription), the nucleic acid molecules 110, 150 of the system are RNA and therefore lack guide nucleic acid molecules 140, 141, 171, 172, promoters 112, 152, 142, 143, 173, 174, and parvoviral ITRs 176, 177, 178, 179. The drawing is not to scale.

[0047] [Figure 6-8]Figure 6H is a schematic diagram presenting an exemplary system or composition for the RNA recombination methods disclosed herein using three DNA molecules 110, 200, 150, where a nucleic acid editing protein is divided into three portions (N-terminal, middle, and C-terminal, 114, 216, 154, respectively), each portion encoded by a different nucleic acid molecule 110, 200, 150. In some examples, promoters 112, 202, 152 drive expression of each coding sequence 114, 216, 164. This example system includes one or more guide nucleic acid molecules (e.g., gRNAs or gRNA-coding sequences) 140, 141, 231, 232, 171, 172 specific for one or more target nucleic acid molecules. However, such guide nucleic acid molecules 140, 141, 231, 232, 171, 172 are optional, and in some examples, instead of being provided as part of molecule 110, 200, 150, they are provided separately, for example, as part of a separate vector. This example shows optional guide nucleic acid molecules 140, 141, 231, 232, 171, 172 near the 5' and 3' ends of nucleic acid molecule 110, 200, 150, although the disclosure is not limited to such locations. When one or more guide nucleic acid molecules 140, 141, 231, 232, 171, 172 are present, their expression can be driven by promoters 142, 143, 233, 234, 173, 174. The system may optionally include parvoviral inverted terminal repeats (ITRs) 176, 177, 235, 236, 178, 179 at the respective 5' and 3' ends of the molecules 110, 200, 150. The diagram shows an embodiment in which the molecules 110, 150, 200 of the system are DNA, however, after transcription, the nucleic acid molecules 110, 150, 200 of the system are RNA and therefore lack the guide nucleic acid molecule 140, 141, 171, 172, 321, 232, promoter 112, 152, 202, 142, 143, 233, 234, 173, 174 and parvoviral ITRs 176, 177, 235, 236, 178, 179. The drawing is not to scale.

[0048] [Figure 7-1]Figure 7A is a schematic diagram presenting an exemplary system for the RNA recombination methods disclosed herein, which uses two nucleic acid molecules 500, 600 as in Figure 6A, but in which the dimerization domains are aptamers 512, 602 that recognize the same target molecule 700. In some examples, for example, after transcription, the elements shown are RNA. The drawing is not to scale.

[0049] [Figure 7-2] Figure 7B is a schematic diagram presenting an exemplary system for the RNA recombination methods disclosed herein, related to Figure 7A, that uses dimerization domains that recognize the same target molecule. Here, the target recognized by the dimerization domains is a specific RNA molecule (instead of molecule 700 in Figure 7A, e.g., a protein or small molecule). Each domain recognizes a different portion of an mRNA molecule that is expressed only in target cells (i.e., cells in which expression of the target protein is desired), such as, for example, a cancer-specific transcript. In some examples, for example, after transcription, the depicted element is RNA. The drawing is not to scale.

[0050] [Figure 7-3] Figure 7C is a schematic diagram presenting an exemplary system for the RNA recombination methods disclosed herein, using two nucleic acid molecules 800, 900 similar to Figures 6A and 7A, showing that dimerization domains 812, 902 are hybridized to an oligonucleotide 1000 that prevents the dimerization domains from interacting with each other, thus preventing or reducing recombination of the N-terminal coding sequence 802 and the C-terminal coding sequence 914. In some examples, for example, after transcription, the elements shown are RNA. The drawings are not to scale.

[0051] [Figure 8] Figure 8 is a bar graph comparing the reconstitution of YFP protein expression in the presence (w / ) or absence (w / o) of the WPRE3 sequence in the 3' untranslated region. N=3 replicates per sample are shown.

[0052] [Figure 9-1] Figure 9A is a schematic diagram presenting an example of the use of a dimerization domain (e.g., 122, 154 in Figure 6A) containing a kissing loop interaction for high-affinity dimerization. It will be understood that, using the teachings presented herein, any of the coding portions disclosed herein (e.g., YFP) can be replaced with other target protein coding sequences. The drawings are not to scale.

[0053] [Figure 9-2] Figure 9B shows RFP, BFP, and YFP signals in HEK293T cells transfected with both split YFP halves, either a linear dimerization domain attached using low-diversity design principles or a structured dimerization domain designed for kissing loop-loop interactions. Efficient reconstitution is indicated by a strong yellow fluorescent signal.

[0054] [Figure 10-1]10A-10Z are exemplary synthetic nucleic acid molecules that can be used with the systems and methods. In some examples, the synthetic nucleic acid molecule has at least 80%, at least 85%, at least 90%, at least 95%, at least 98%, at least 99%, or 100% sequence identity to any one of SEQ ID NOs: 1 (Figures 10A-10B), 2 (Figures 10C-10E), 7 (Figure 10E), 8 (Figure 10F), 9 (Figure 10G), 10 (Figure 10H), 11 (Figure 10I), 12 (Figure 10J), 13 (Figure 10K), 14 (Figure 10L), 15 (Figure 10M), 16 (Figure 10N), 17 (Figure 10O), 18 (Figure 10P), 19 (Figure 10Q), 20 (Figures 10R-10U), and 21 (Figures 10V-10Z), but has a different target protein coding sequence. Thus, an intron region used with any of the systems or methods presented herein can have at least 80%, at least 85%, at least 90%, at least 95%, at least 98%, at least 99%, or 100% sequence identity to any of the intron sequences of SEQ ID NOs: 1, 2, 3, 4, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, or 21. For example, Figures 10A-D show exemplary (A, B) first synthetic molecule (SEQ ID NO: 1) and (C, D) second synthetic molecule (SEQ ID NO: 2) that can be used to express full-length YFP, while SEQ ID NOs: 3 and 4 present the corresponding synthetic intron portions without the YFP-encoding portion. In some examples, the synthetic intron sequence has at least 80%, at least 85%, at least 90%, at least 95%, at least 98%, at least 99%, or 100% sequence identity to SEQ ID NO: 3 or 4. Thus, portions of the coding sequence of any of the synthetic molecules presented herein (eg, nt 544-1032 of SEQ ID NO: 1 and nt 905-1141 of SEQ ID NO: 2) can be replaced with alternative portions of the coding sequence. [Figure 10-2] Same as above. [Figure 10-3] Same as above. [Figure 10-4] Same as above. [Figure 10-5] Same as above. [Figure 10-6] Same as above. [Figure 10-7] Same as above. [Figure 10-8] Same as above. [Figure 10-9] Same as above. [Figure 10-10] Same as above. [Figure 10-11] Same as above. [Figure 10-12] Same as above. [Figure 10-13] Same as above. [Figure 10-14] Same as above. [Figure 10-15] Same as above. [Figure 10-16] Same as above. [Figure 10-17] Same as above. [Figure 10-18] Same as above. [Figure 10-19] Same as above. [Figure 10-20] Same as above. [Figure 10-21] Same as above. [Figure 10-22] Same as above. [Figure 10-23] Same as above. [Figure 10-24] Same as above. [Figure 10-25] Same as above. [Figure 10-26] Same as above.

[0055] [Figure 11] Figure 11 is a bar graph showing the reconstitution efficiency of random, complementary base-pairing binding domains of different lengths (50 bp, 100 bp, 150 bp, 200 bp, 300 bp, 400 bp, and 500 bp). Median YFP fluorescence intensity is compared between cells with comparable RFP and BFP transfection levels. n = 3 samples per condition. n = 3 samples per condition.

[0056] [Figure 12]Figures 12A-12B show that including a splice enhancer in the synthetic intron increases reconstitution efficiency. Figure 12A is a schematic diagram of the 5'-N-terminal and 3'-C-terminal constructs (SEQ ID NOS: 1 and 2) used. (See Figure 1A for abbreviations.) Figure 12B is a bar graph showing the YFP fluorescence obtained after transfection of cells with SEQ ID NOS: 1 and 2 or their various truncations, indicated by Δ. n=3 samples per condition.

[0057] [Figure 13] Figures 13A-13D show tracing of midline-crossing cortical neurons by reconstituting full-length flp recombinase (Flpo) from two fragments (SEQ ID NOs: 147 and 148). (A) Schematic representation of the 5' and 3' sequences used to reconstitute flpo (similar to the construct in Figure 12A). (B) Schematic representation of an injected flp-reporter mouse line injected with AAV viruses encoding N-flpo and C-flpo into the left and right cortical regions, respectively. (C and D) Show neuronal cell body and axonal labeling of cortical neurons projecting to the opposite hemisphere of the brain and therefore infected with both N-flpo and C-flpo viruses. Hoechst staining (nuclei) is shown for context.

[0058] [Figure 14-1]Figures 14A-14D show the expression of oversized cargo (i.e., proteins encoded by long RNAs) in cell culture and in vivo in the mouse primary motor cortex. (A) Schematic representation of the 5' and 3' sequences used to reconstitute YFP, including a long stuffer sequence (uninterrupted open reading frame; SEQ ID NOs: 22 and 23, respectively). (B) Quantitative real-time PCR analysis of the reconstitution efficiency of oversized YFP constructs in HEK 293T cells. N=3 per condition. (C) Expression of reconstituted YFP proteins from full-length oversized YFP expression and split-REJ expression assessed by flow cytometry in transiently transfected HEK 293T cells. Median yellow fluorescence intensity is compared between cell populations with equal transfection control (blue and red) fluorescence for different conditions. The Y-axis indicates median yellow fluorescence intensity [au]. N=3 per condition. (D) Schematic of injection into mouse primary motor cortex and image of brain tissue 10 days after injection showing successful reconstitution of the long (2401 amino acid) YFP protein in vivo. [Figure 14-2] Same as above.

[0059] [Figure 15-1]Figures 15A-15C show efficient reconstitution of full-length human coagulation factor VIII (FVIII) (2317 amino acids) with an N-terminal HA tag (in place of the N-terminal signal peptide). (A) Schematic representation of the 5' and 3' sequences used to reconstitute FVIII (SEQ ID NOs: 24 and 25, respectively). (B) PCR amplification of the junction. (C) Western blot showing FVIII expression. Lanes 1-3: Expression of full-length FVIII (the 290 kDa band indicates full-length, unprocessed FVIII). Lanes 4-6: Expression of reconstituted FVIII (the 290 kDa band indicates successfully reconstituted FVIII). Lanes 7 and 8: Expression of the N-terminus only indicates the absence of the 290 kDa full-length FVIII band. For all lanes: the expected protein processing products are observed ranging from approximately 75 kDa to approximately 210 kDa. FVIII is probed using a mouse anti-HA primary antibody. All lanes were loaded with 5 micrograms of clarified cellular protein extract and probed with GAPDH (rabbit anti-GAPDH) as a loading control. [Figure 15-2] Same as above.

[0060] [Figure 16-1]Figures 16A-16F show efficient reconstitution of full-length human Abca4 (2300 amino acids) with a C-terminal FLAG tag. (A) Schematic representation of the 5' and 3' sequences used to reconstitute Abca4 (SEQ ID NOs: 20 and 21, respectively), along with a Sanger sequencing trace across the junction. (B) PCR amplification of the junction. (C) Schematic representation of the probes used to assay recombination of the 5' and 3' fragments. (D) Quantification of reconstitution efficiency by PCR after 2 days of expression in HEK 293T cells. N=2 per condition. (E) Western blot showing Abca4 expression. Lanes 1-3: Expression of full-length Abca4 (the approximately 260 kDa band represents full-length Abca4). Lanes 4-6: Expression of reconstituted Abca4 (the 260 kDa band represents successfully reconstituted Abca4). Lanes 7 and 8: No-transfection control (i.e., HEK 293t lysate only) shows the absence of any signal. Abca4 is probed using a mouse anti-FLAG primary antibody. All lanes were loaded with 5 micrograms of clarified cell protein extract. GAPDH (rabbit anti-GAPDH) is probed as a loading control. (F) Quantification of the Western blot in (E) normalized for differential BFP concentration. Data are shown normalized to the mean of the full-length expression control. [Figure 16-2] Same as above.

[0061] [Figure 17] Figures 17A and 17B show (A) the kissing loop dimerization domain based on HIV-1 (N fragment, SEQ ID NO: 139, C fragment, SEQ ID NO: 140); and (B) the kissing loop dimerization domain based on HIV-2 (N fragment, SEQ ID NO: 141, C fragment, SEQ ID NO: 142).

[0062] [Figure 18]Figures 18A-18C show efficient reconstitution of full-length mouse Otof (2,019 amino acids) with a C-terminal FLAG tag. The DNA sequences of the 5' and 3' molecules used are shown in SEQ ID NOs: 155 and 156. (A) Western blot showing Otof expression. Lanes 1-3: Expression of full-length Otof (the approximately 250 kDa band indicates full-length Otof). Lanes 4-6: Expression of reconstituted Otof (the 250 kDa band indicates successfully reconstituted Otof). Lane 7: No-transfection control (i.e., HEK 293t lysate only) shows the absence of any signal. Otof is probed using a mouse anti-FLAG primary antibody. All lanes were loaded with 5 micrograms of clarified cell protein extract. GAPDH (rabbit anti-GAPDH) is probed as a loading control. (B) Raw quantification of Western blots and (C) normalization for differential BFP concentration. Data are shown normalized to the mean of full-length expression controls.

[0063] [Figure 19] Figures 19A-19C show efficient reconstitution of full-length human Myo7a (2243 amino acids) with a C-terminal FLAG tag. The DNA sequences of the 5' and 3' molecules used are shown in SEQ ID NOs: 157 and 158. (A) Western blot showing Myo7a expression. Lanes 1-3: Expression of full-length Myo7a (the approximately 270 kDa band indicates full-length Myo7a). Lanes 4-6: Expression of reconstituted Myo7a (the 270 kDa band indicates successfully reconstituted Myo7a). Lane 7: No-transfection control (i.e., HEK 293t lysate only) shows the absence of any signal. Myo7a is probed using a mouse anti-FLAG primary antibody. All lanes were loaded with 5 micrograms of clarified cell protein extract. GAPDH (rabbit anti-GAPDH) is probed as a loading control. (B) Raw quantification of the Western blot and (C) normalization for differential BFP concentration. Data are shown normalized to the mean of full-length expression controls.

[0064] [Figure 20] Figures 20A-20D show efficient reconstitution of full-length DCas9-VPR (1951 amino acids). The DNA sequences of the 5' and 3' molecules used are shown in SEQ ID NOs: 159 and 160. (A) Western blot showing DCas9-VPR expression. Lanes 1-3: Expression of full-length DCas9-VPR (the approximately 250 kDa band indicates full-length DCas9-VPR). Lanes 4-6: Expression of reconstituted DCas9-VPR (the 250 kDa band indicates successfully reconstituted DCas9-VPR). Lane 7: No-transfection control (i.e., HEK 293t lysate only) shows the absence of any signal. DCas9-VPR is probed using a mouse anti-Cas9 primary antibody. All lanes were loaded with 5 micrograms of clarified cell protein extract. GAPDH (rabbit anti-GAPDH) is probed as a loading control. (B) Raw quantification of Western blots and (C) normalization for differential BFP concentration. Data are shown normalized to the mean of the full-length expression control. (D) Example of transcriptional activation of YFP expression plasmids in HEK 293T cells. Full-length dCas9-VPR (upper panel) or two-way split-REJ dual dCas9-VPR (lower panel) are transiently transfected with a non-targeting guide RNA expression plasmid (left panel) or a UAS-targeting guide RNA expression plasmid (right panel). All cells are also transfected with a UAS-YFP plasmid, which is transcriptionally inactive until dCas9-VPR is targeted to the upstream region of a minimal promoter that drives yellow fluorescent protein expression. Red fluorescent protein (RFP) is expressed with the N-terminal fragment of dCas9-VPR, and blue fluorescent protein (BFP) is expressed with full-length dCas9-VPR or the C-terminal fragment of dCas9-VPR, respectively. RFP and BFP serve as transfection controls. When both the full-length dCas9-VPR paired with the UAS-targeting guide RNA and the bidirectional split dCas9-VPR were expressed, yellow fluorescent protein expression was observed, confirming the functionality of the reconstituted full-length protein.

[0065] [Figure 21]Figures 21A-21D show efficient reconstitution of a full-length humanized prime editor (2118 amino acids). The DNA sequences of the 5' and 3' molecules used are shown in SEQ ID NOs: 161 and 162. (A) Western blot showing expression of the prime editor. Lanes 1-3: Expression of the full-length prime editor (the approximately 260 kDa band indicates the full-length prime editor). Lanes 4-6: Expression of the reconstituted prime editor (the 260 kDa band indicates the successfully reconstituted prime editor). Lane 7: No transfection control (i.e., HEK 293t lysate only) shows the absence of any signal. The prime editor is probed using a mouse anti-Cas9 primary antibody. All lanes were loaded with 5 micrograms of clarified cell protein extract. GAPDH (rabbit anti-GAPDH) is probed as a loading control. (B) Raw quantification of the Western blot and (C) normalization for differential BFP concentration. Data are shown normalized to the mean of the full-length expression control. (D) Prime editor-induced G-to-T transversion mutations were induced at the FANCF and VEGFA3 loci in HEK293T cells. The top panel shows the sequence context of the FANCF and VEGFA3 loci, respectively. The gray arrow indicates the sequence targeted by the prime editor guide RNA (pegRNA). The protospacer adjacent motif (PAM) is indicated by a gray box. The G targeted for transversion to T is highlighted in the sequence. The genomic locus was sequenced using Sanger sequencing in three conditions. The top panel shows a representative Sanger trace for the unedited wild-type condition. The second panel from the top shows a representative Sanger trace representing the full-length expressed prime editor construct. The region highlighted by the black box indicates the appearance of a T band in Sanger sequencing, indicating successful integration of the edit into a portion of the cells. The bottom panel shows a representative Sanger trace for cells edited using the two-way split-recombinant prime editor.The appearance of a T trace (black box) demonstrates the functionality of the prime editor when reassembled from the two fragments.

[0066] [Figure 22]Figures 22A-22C show efficient reconstitution of a full-length humanized cytosine base editor (AncBE4) (1854 amino acids). The DNA sequences of the 5' and 3' molecules used are shown in SEQ ID NOs: 163 and 164. (A) Western blot showing AncBE4 expression. Lanes 1-3: Expression of full-length AncBE4 (the approximately 230 kDa band indicates full-length AncBE4). Lanes 4-6: Expression of reconstituted AncBE4 (the 230 kDa band indicates successfully reconstituted AncBE4). Lane 7: No-transfection control (i.e., HEK 293t lysate only) shows the absence of any signal. AncBE4 is probed using a mouse anti-Cas9 primary antibody. All lanes were loaded with 5 micrograms of clarified cell protein extract. GAPDH (rabbit anti-GAPDH) is probed as a loading control. (B) Raw quantification of the Western blot. Data are shown normalized to the mean of the full-length expression control. (C) AncBE4-induced C-to-T transposition mutations were induced at the EMX1 and HEK site 3 loci in HEK293T cells. The top panel shows the sequence context of the EMX1 and HEK site 3 loci, respectively. The gray arrow indicates the sequence targeted by the AncBE4 gRNA. The protospacer adjacent motif (PAM) is indicated by a gray box. The C targeted for transposition to a T is highlighted in the sequence. The genomic locus was sequenced using Sanger sequencing in three conditions. The top panel shows a representative Sanger trace for the unedited wild-type condition. The second panel from the top shows a representative Sanger trace representing the full-length expressed AncBE4 construct. The region highlighted by the black box indicates the appearance of a T band in the Sanger sequencing, indicating successful integration of the edit into a portion of the cells. The bottom panel shows a representative Sanger trace for cells edited using the two-way split reconstituted AncBE4. The appearance of the T trace (black box) demonstrates the functionality of AncBE4 when reconstituted from the two fragments.

[0067] [Figure 23-1] Figures 23A-23C show efficient reconstitution of full-length humanized adenine base editor (ABE8e) (1606 amino acids) (e.g., SEQ ID NOs: 225, 226). The DNA sequences of the 5' and 3' molecules used are shown in SEQ ID NOs: 165 and 166. (A) Western blot showing expression of ABE8e. Lanes 1-3: Expression of full-length ABE8e (the approximately 230 kDa band indicates full-length ABE8e). Lanes 4-6: Expression of reconstituted ABE8e (the 230 kDa band indicates successfully reconstituted ABE8e). Lane 7: No-transfection control (i.e., HEK293t lysate only) shows the absence of any signal. ABE8e is probed using a mouse anti-Cas9 primary antibody. All lanes were loaded with 5 micrograms of clarified cell protein extract. GAPDH (rabbit anti-GAPDH) is probed as a loading control. (B) Raw quantification of the Western blot. Data are shown normalized to the mean of the full-length expression control. (C) ABE8e-induced A-to-G transition mutations were induced at the BCL11A and HGB1 / 2 loci in HEK293t cells. The top panels show the sequence context of the BCL11A and HGB1 / 2 loci, respectively. Gray arrows indicate the sequence targeted by the ABE8e guide RNA (gRNA). The protospacer adjacent motif (PAM) is indicated by a gray box. The A targeted for G transition is highlighted in the sequence. Genomic loci were sequenced using Sanger sequencing in three conditions. The top panel shows a representative Sanger trace for the unedited wild-type condition. The second panel from the top shows a representative Sanger trace representing the full-length expressed ABE8e construct. The region highlighted by a black box indicates the appearance of a G band in Sanger sequencing, indicating successful integration of the edit into a subset of cells. The bottom panel shows a representative Sanger trace for cells edited with the two-way split reconstituted ABE8e. The appearance of the G trace (black square) demonstrates the functionality of ABE8e when reconstituted from two fragments.

[0068] [Figure 23-2]Figures 23D-G show efficient reconstitution of the full-length humanized adenine base editor (ABE8e) (1606 amino acids) to correct a premature stop codon in the mdx mouse model of Duchenne muscular dystrophy. (D) Schematic of the experiment. The ABE8e base editor is split into two fragments whose coding DNA is packaged separately into two separate adeno-associated virus capsids. A CRISPR gRNA is designed to target the locus of the premature stop codon, such that an A to G transition converts the TAA stop codon to a CAA codon. (E) Yellow fluorescent protein (YFP) is split into two coding fragments connected by a short stretch of mdx coding sequence surrounding the premature stop codon. The first half of YFP is non-fluorescent when the open reading frame is terminated by the premature TAA stop codon. Correction of the stop codon results in translation of the entire fluorescent YFP sequence (this construct is referred to as the YFP-editing-reporter). The left panel shows the absence of YFP fluorescence when HEK293T cells are transfected with the YFP-editing-reporter (co-expressing a red fluorescent transfection control), the N-terminal ABE8e vector, the C-terminal ABE8e vector, and a non-targeting gRNA. The right panel shows that a high percentage of cells express YFP fluorescence when the YFP-editing-reporter (co-expressing a red fluorescent transfection control), the N-terminal ABE8e vector, the C-terminal ABE8e vector, and an mdx gene-targeting gRNA are co-transfected. (F) In vivo editing of the mdx premature stop codon results in dystrophin expression in the treated muscle. Male mdx mutant mice were injected with a mixture of N-terminal ABE8e and C-terminal ABE8e packaged in an adeno-associated virus 8 vector (5 x 10 viral genomes per vector, for a total of 1 x 10 viral genomes per muscle). The viral genomes contained two gRNA expression cassettes (composed of an RNA polymerase III promoter and gRNA sequence) in each of the two genomes. The viral mix was injected intramuscularly into the tibialis anterior muscle.The top right panel shows dystrophin staining in a wild-type tibialis anterior muscle cross section as a reference. The bottom left panel shows untreated tibialis anterior muscle tissue. The bottom right panel shows dystrophin expression in tibialis anterior muscle injected with two adeno-associated viruses. (G) Dystrophin expression (top panel), ABE8e expression (middle panel), and GAPDH loading control (bottom panel) in tibialis anterior muscle treated with an adeno-associated virus mix to express ABE8e are shown. This illustrates the rescue of dystrophin expression using the ABE8e base editor. [Figure 23-3] Same as above. [Figure 23-4] Same as above. [Figure 23-5] Same as above.

[0069] [Figure 24-1]The effects of downstream intron splicing enhancer (DISE) and intron splicing enhancer (ISE) and acceptor sequences on the efficiency of RNA end-joining. (A) Schematic diagram of the screening setup. The 5' fragment is an RNA molecule transcribed from a DNA construct using the human CMV promoter and enhancer. The resulting RNA molecule contains a long stuffer open reading frame to simulate a large cargo size. This stuffer sequence ends with a 2A self-cleaving peptide sequence, followed by the coding region for the 5' fragment of yellow fluorescent protein (n-yfp). The 5' fragment of yfp ends with a splice donor site (SD). This splice donor site is followed by the 5' intron portion of the RNA end-joining module. To determine the effect of DISE and ISE sequences on the efficiency of the RNA end-joining reaction, the 5' intron portion is subdivided into three fragments: from 5' to 3': ds: downstream segment; m: middle intron segment; and dd: donor distal segment. The 5' intronic portion is followed by a trimodal kissing loop RNA dimerization domain. The message is terminated by a short polyadenylation signal. The overall length of this 5' RNA molecule is approximately 4 kb to simulate the reassembly scenario of a large cargo. The 3' fragment is an RNA molecule transcribed from a DNA construct using the human CMV promoter and enhancer. The 3' fragment begins with a trimodal kissing loop RNA dimerization domain complementary to that of the RNA molecule encoding the 5' fragment. The dimerization domain is followed by the 3' intronic portion of the RNA end-binding module. This 3' intronic portion is subdivided into three segments: ad: acceptor distal segment; m: middle intron segment; and ap: acceptor proximal segment. The acceptor proximal segment contains a branch point and a polypyrimidine tract variation, both of which are essential for spliceosome-mediated RNA binding reactions. The splice acceptor (SA) site is followed by the 3'yfp coding sequence, followed by a self-cleaving 2A sequence, followed by a long stuffer open reading frame. The message is terminated by an SV40 polyadenylation signal.The overall length of the 3' RNA molecule is approximately 4 kb to simulate a reconstitution scenario for a large cargo. Association of the two RNA molecules (5' and 3' fragments) is mediated by a trimodal kissing-loop RNA dimerization domain, while spliceosome recruitment and RNA end-joining are mediated by an intron segment. Successful RNA end-joining results in the reconstitution of the YFP open reading frame and subsequent YFP translation. (B) Median YFP fluorescence intensity determined by flow cytometry is shown for multiple intron configurations. In the first grouping (bars 1–9), a selection of potential downstream intron splicing enhancer sequences was paired with the consensus splice donor site (GTAAGTATT in the DNA construct and GUAAGUAUU in the RNA sequence) shown in bars 1–8. These are compared to the consensus splice donor (ds9) followed by a scrambled sequence consisting of equal portions of all four bases. In the second grouping, m1–m16, the selection of potential intron splicing enhancers was compared with a scrambled sequence (m16). The final grouping compared the selection of potential strong branch points, polypyrimidine tracts, and splice acceptors. The reference construct consisted of a scrambled sequence and consensus donor at all nonvariable positions, followed by a scrambled sequence and consensus splice acceptor at the ds position (where the entire polypyrimidine tract was composed of Ts in the DNA construct and Us in the RNA fragment, respectively). (C) List of the different DISE, ISE, and splice acceptor elements used. [Figure 24-2] Same as above. [Figure 24-3] Same as above.

[0070] [Figure 25]25A-25B show the in vivo expression of base editors in mouse muscle and correction of dystrophin gene mutations using the methods disclosed herein. (A) Western blot showing the expression of dystrophin, Cas9-ABE, and GAPDH in muscle extracts from wild-type (wt), untreated mdx-4cv, and treated mdx-4cv mice. (B) Immunohistochemical analysis of dystrophin expression in muscle from wild-type (wt), untreated mdx-4cv, and treated mdx-4cv mice. DETAILED DESCRIPTION OF THE INVENTION

[0071] Sequence Listing The nucleic acid and amino acid sequences listed in the accompanying sequence listing are shown using standard letter abbreviations for nucleotide bases and three-letter codes for amino acids, as defined in 37 C.F.R. 1.822. Only one strand of each nucleic acid sequence is shown; however, it is understood that any reference to the shown strand includes the complementary strand. The sequence listing was submitted as a 299 KB ASCII text file created on May 16, 2022, and is incorporated herein by reference. In the accompanying sequence listing:

[0072] SEQ ID NOs: 1 and 2 are the N- and C-terminal sequences, respectively, used to express full-length YFP. SEQ ID NO: 1 contains the CMV promoter nt 1 to 543, the YFP coding sequence nt 544 to 1032, the synthetic intron nt 1033 to 1436, and the untranslated poly(A) region nt 1437 to 1491. SEQ ID NO: 2 contains the CMV promoter nt 1 to 522, the synthetic intron nt 523 to 904, the YFP coding sequence nt 905 to 1141, and the untranslated poly(A) region nt 1142 to 1302.

[0073] SEQ ID NOs: 3 and 4 are 5' and 3' intron sequences, respectively, that can be used to express a desired full-length protein, where the N-terminal portion of the full-length protein can be added at nt 1 of SEQ ID NO: 3, and the C-terminal portion of the full-length protein can be added at nt 382 of SEQ ID NO: 4.

[0074] SEQ ID NOs: 5 and 6 are the N- and C-terminal coding sequences, respectively, used to express full-length YFP.

[0075] SEQ ID NO: 7 is an exemplary synthetic intron dimerization domain (FIG. 10E).

[0076] SEQ ID NO:8 is an exemplary synthetic intron without an intronic splicing enhancer (FIG. 10F).

[0077] SEQ ID NO: 9 is an exemplary synthetic intron without an intronic splicing enhancer (FIG. 10G).

[0078] SEQ ID NO: 10 is an exemplary synthetic intron without an intronic splicing enhancer (FIG. 10H).

[0079] SEQ ID NO: 11 is an exemplary synthetic intron without a binding domain (FIG. 10I).

[0080] SEQ ID NO: 12 is an exemplary synthetic intron with a dimerization domain (FIG. 10J).

[0081] SEQ ID NO: 13 is an exemplary synthetic intron with a dimerization domain (FIG. 10K).

[0082] SEQ ID NO: 14 is an exemplary synthetic intron without an intronic splicing enhancer (FIG. 10L).

[0083] SEQ ID NO: 15 is an exemplary synthetic intron with only DISE (Figure 10M).

[0084] SEQ ID NO: 16 is an exemplary synthetic intron without HHrz (FIG. 10N).

[0085] SEQ ID NO: 17 is an exemplary synthetic intron without an intronic splicing enhancer (FIG. 10O).

[0086] SEQ ID NO: 18 is an exemplary U12-dependent intron with a binding domain (FIG. 10P).

[0087] SEQ ID NO: 19 is an exemplary U12-dependent intron with a binding domain (FIG. 10Q).

[0088] SEQ ID NOs: 20 and 21 are the N- and C-terminal DNA sequences, respectively, used to express RNA (pre-mRNA) resulting in full-length Abca4. In SEQ ID NO: 20, the sequence corresponding to the N-terminal Abca4 coding region is located between nt 22 and nt 3702, a synthetic intron is located between nt 3703 and nt 3912, and an untranslated poly(A) region is located between nt 3921 and nt 3969. SEQ ID NO: 20 also contains a splice donor between nt 3703 and nt 3711, a rat FGFR2 DISE between nt 3714 and nt 3737, a cTNT intron splicing enhancer between nt 3747 and nt 3770, an M2 intron splicing enhancer between nt 3782 and nt 3794, and a kissing loop dimerization domain between nt 3801 and nt 3975. In SEQ ID NO:21, nts 1 to 228 are a synthetic intron, nts 229 to 3366 are the C-terminal Abca4 coding region, nts 3367 to 3447 are a FLAG epitope tag, and nts 3476 to 3607 are an untranslated poly(A) region (signal). SEQ ID NO:21 also contains a kissing loop dimerization domain at nts 3 to 114, an M2 intron splicing enhancer at nts 121 to 133, a cTNT intron splicing enhancer at nts 140 to 163, an M2 intron splicing enhancer at nts 175 to 187, a branchpoint motif at nts 194 to 201, a polypyrimidine tract at nts 207 to 226, and a splice acceptor at nt 228.

[0089] SEQ ID NOs: 22 and 23 are the N-terminal and C-terminal DNA sequences, respectively, used to express the RNA (pre-mRNA) that yields full-length YFP, each containing a splice enhancer. In SEQ ID NO: 22, the N-terminal YFP coding region is from nt 22 to 3702, nt 3703 to 3912 is a synthetic intron, and nt 3921 to 3969 is an untranslated poly(A) region. SEQ ID NO: 22 also contains a splice donor from nt 3703 to 3711, a rat FGFR2 DISE from nt 3714 to 3737, a cTNT intron splicing enhancer from nt 3747 to 3770, an M2 intron splicing enhancer from nt 3782 to 3794, and a kissing loop dimerization domain from nt 3801 to 3975. In SEQ ID NO: 23, nts 1 to 225 are a synthetic intron, nts 226 to 3747 are a C-terminal YFP coding region, and nts 3748 to 3912 are an untranslated poly(A) region. SEQ ID NO: 23 also contains a kissing loop dimerization domain from nts 3 to 114, an M2 intron splicing enhancer from nts 118 to 130, a cTNT intron splicing enhancer from nts 137 to 160, an M2 intron splicing enhancer from nts 172 to 184, a branchpoint motif from nts 191 to 198, a polypyrimidine tract from nts 204 to 223, and a splice acceptor at nt 225.

[0090] SEQ ID NOs: 24 and 25 are the N- and C-terminal sequences, respectively, used to express RNA (pre-mRNA) resulting in full-length human factor VIII. In SEQ ID NO: 24, the N-terminal FVIII coding region with an N-terminal HA epitope tag is located between nt 22 and nt 3561, nt 3562 and nt 3771 is a synthetic intron, and nt 3780 and nt 3828 is an untranslated poly(A) region. SEQ ID NO: 24 also contains a splice donor between nt 3562 and nt 3570, a rat FGFR2 DISE between nt 3573 and nt 3596, a cTNT intron splicing enhancer between nt 3606 and nt 3629, an M2 intron splicing enhancer between nt 3641 and nt 3653, and a kissing loop dimerization domain between nt 3660 and nt 3834. In SEQ ID NO: 25, nt 1 to 225 are a synthetic intron, nt 226 to 3636 are a C-terminal FVIII coding region, and nt 3665 to 3797 are an untranslated poly(A) region. SEQ ID NO: 25 also contains a splice donor at nt 3703 to 3711, a rat FGFR2 DISE at nt 3714 to 3737, a cTNT intron splicing enhancer at nt 3747 to 3770, an M2 intron splicing enhancer at nt 3782 to 3794, and a kissing loop dimerization domain at nt 3801 to 3975.

[0091] SEQ ID NOs: 26-136 are exemplary splicing enhancers that can be used in the systems presented herein (e.g., 118, 120, 156 in Figure 6A).

[0092] SEQ ID NOs: 137 and 138 are exemplary splice donor sequences.

[0093] SEQ ID NOs: 139 and 140 are the N and C fragments, respectively, of the HIV-1 based kissing loop dimerization domain.

[0094] SEQ ID NOs: 141 and 142 are the N and C fragments, respectively, of the kissing loop dimerization domain based on HIV-2.

[0095] SEQ ID NO: 143 is an exemplary cryptic splice acceptor sequence.

[0096] SEQ ID NO: 144 is an exemplary branch point consensus sequence.

[0097] SEQ ID NOs: 145 and 146 are the N- and middle sequences, respectively, used together with SEQ ID NO: 2 (C-terminal fragment) to express full-length YFP. In SEQ ID NO: 145, nt 1 to 543 is the CMV promoter sequence, nt 544 to 849 is the N-terminal YFP coding region, and nt 850 to 1305 is a synthetic intron. In SEQ ID NO: 146, nt 1 to 522 is the CMV promoter sequence, nt 523 to 901 is a synthetic intron, nt 902 to 1084 is the middle YFP coding region, and nt 1085 to 1543 is an untranslated poly(A) region.

[0098] SEQ ID NOs: 147 and 148 are the 5' and 3' synthetic sequences, respectively, used to express full-length Flpo. In SEQ ID NO: 147, nt 1 to 540 is the CMV promoter sequence, nt 541 to 1112 is the N-terminal Flpo coding region, and nt 1113 to 1571 is a synthetic intron. In SEQ ID NO: 148, nt 1 to 522 is the CMV promoter sequence, nt 523 to 904 is a synthetic intron, nt 905 to 1604 is the C-terminal Flpo coding region, and nt 1605 to 1765 is an untranslated poly(A) region.

[0099] SEQ ID NOs: 149 and 150 are exemplary low diversity sequences.

[0100] SEQ ID NOs: 151 and 152 are exemplary splice donor consensus sequences.

[0101] SEQ ID NO: 153 is an exemplary kissing loop based on the HIV-2 kissing loop dimerization domain (SEQ ID NOs: 141 and 142, Figure 17B).

[0102] SEQ ID NO: 154 is an exemplary Kozak-enhanced start codon.

[0103] SEQ ID NOs: 155 and 156 are exemplary constructs that can be used to express the mouse Otof coding sequence in vivo. SEQ ID NO: 155 is used to generate the N-terminal Otof RNA. SEQ ID NO: 155 contains the human CMV enhancer and promoter from nt 1 to 522, the putative transcription start site from nt 523, and the polyadenylation signal from nt 4263 to 4311. SEQ ID NO: 155 encodes the N-terminal Otof RNA elements as follows: a 5' untranslated region containing a Kozak sequence from nt 523 to 546; a 5' otoferlin coding sequence from nt 547 to 4044; a 5' synthetic intron sequence from nt 4045 to 4142; a 5' trimodal kissing loop dimerization domain from nt 4143 to 4254; and a linker from nt 4255 to 4262. SEQ ID NO: 155 is used to generate the C-terminal Otof RNA. It contains the human CMV enhancer and promoter from nt 1 to 522, a putative transcription start site from nt 523, and a polyadenylation signal from nt 3335 to 3467. It encodes the C-terminal Otof RNA element as follows: a 3' trimodal kissing-loop dimerization domain from nt 525 to 636; a 3' synthetic intron sequence from nt 637 to 747; a 3' otoferlin coding sequence from nt 748 to 3225; a C-terminal 3xFlag tag from nt 3226 to 3306; and a linker from nt 3307 to 3334.

[0104] SEQ ID NOs:157 and 158 are exemplary constructs that can be used to express the human myosin VIIA (Myo7a) coding sequence in vivo. SEQ ID NO:157 is used to generate N-terminal Myo7a RNA. SEQ ID NO:157 contains the human CMV enhancer and promoter from nt 1 to 522, the putative transcription start site from nt 523, and the polyadenylation signal from nt 4344 to 4392. SEQ ID NO:157 encodes the N-terminal Myo7A RNA elements as follows: 5' untranslated region containing a Kozak sequence from nt 523 to 543; 5' Myo7a coding sequence from nt 544 to 4125; 5' synthetic intron sequence from nt 4126 to 4223; 5' trimodal kissing loop dimerization domain from nt 4224 to 4335; and a linker from nt 4336 to 4343. SEQ ID NO:158 is used to generate C-terminal Myo7a RNA. SEQ ID NO:158 contains the human CMV enhancer and promoter from nt 1 to 522, a putative transcription start site at nt 523, and a polyadenylation signal from nt 3923 to 4055. SEQ ID NO:158 encodes the C-terminal Myo7a RNA elements as follows: 3' trimodal kissing loop dimerization domain from nt 525 to 636; 3' synthetic intron sequence from nt 637 to 747; 3' Myo7a coding sequence from nt 748 to 3813; C-terminal 3xFlag tag from nt 3814 to 3894; and a linker from nt 3895 to 3922.

[0105] SEQ ID NOs: 159 and 160 are exemplary constructs that can be used to express in vivo the full-length, enzymatically inactive Cas9 (dCas9-VPR) coding sequence fused to the VPR transcriptional activator domain. SEQ ID NO: 159 is used to generate N-terminal DCas9-VPR RNA. SEQ ID NO: 159 contains the human CMV enhancer and promoter from nt 1 to 522, the putative transcription start site from nt 523, and the polyadenylation signal from nt 4112 to 4161. SEQ ID NO: 159 encodes the N-terminal DCas9-VPR RNA elements as follows: 5' untranslated region containing a Kozak sequence from nt 523 to 543; 5' DCas9-VPR coding sequence from nt 544 to 3894; 5' synthetic intron sequence from nt 3895 to 3992; 5' trimodal kissing loop dimerization domain from nt 3993 to 4104; and linker from nt 4105 to 4112. SEQ ID NO: 160 is used to generate the C-terminal DCas9-VPR RNA. SEQ ID NO: 160 includes the human CMV enhancer and promoter from nt 1 to 522, the putative transcription start site from nt 523, and the polyadenylation signal from nt 3278 to 3410. SEQ ID NO: 160 encodes the C-terminal DCas9-VPR RNA elements as follows: a 3' trimodal kissing loop dimerization domain from nt 525 to 636; a 3' synthetic intron sequence from nt 637 to 747; a 3' DCas9-VPR coding sequence from nt 748 to 3249; and a linker from nt 3250 to 3277.

[0106] SEQ ID NOs: 161 and 162 are exemplary constructs that can be used to express the full-length humanized Cas9 prime editor (prime editor) coding sequence in vivo. SEQ ID NO: 161 encodes the N-terminal prime editor sequence as follows: human CMV enhancer and promoter nt 1 to 522; putative transcription start site nt 523; 5' untranslated region containing a Kozak sequence nt 523 to 543; 5' prime editor coding sequence nt 544 to 3894; 5' synthetic intron sequence nt 3895 to 3992; 5' trimodal kissing loop dimerization domain nt 3993 to 4104; linker nt 4105 to 4112; and polyadenylation signal nt 4112 to 4161. SEQ ID NO: 162 encodes the C-terminal prime editor sequence as follows: human CMV enhancer and promoter nt 1-522; putative transcription start site nt 523; 3' trimodal kissing loop dimerization domain nt 525-636; 3' synthetic intron sequence nt 637-747; 3' prime editor coding sequence nt 748-3750; linker nt 3751-3778; polyadenylation signal nt 3779-3911.

[0107] SEQ ID NOs: 163 and 164 are exemplary constructs that can be used to express the full-length humanized cytosine base editor (AncBE4) coding sequence in vivo. SEQ ID NO: 163 encodes the N-terminal AncBE4 sequence as follows: human CMV enhancer and promoter nt 1 to 522; putative transcription start site nt 523; 5' untranslated region containing a Kozak sequence nt 523 to 540; 5' AncBE4 coding sequence nt 541 to 2892; 5' synthetic intron sequence nt 2893 to 2990; 5' trimodal kissing loop dimerization domain nt 2991 to 3102; linker nt 3103 to 3110; and polyadenylation signal nt 3111 to 3159. SEQ ID NO: 164 encodes the C-terminal AncBE4 sequence as follows: human CMV enhancer and promoter nt 1 to 522; putative transcription start site nt 523; 3' trimodal kissing loop dimerization domain nt 525 to 636; 3' synthetic intron sequence nt 637 to 747; 3' AncBE4 coding sequence nt 748 to 3957; linker nt 3958 to 3982; polyadenylation signal nt 3983 to 4115.

[0108] SEQ ID NOs: 165 and 166 are exemplary constructs that can be used to express the full-length humanized adenine base editor (ABE8e) coding sequence in vivo. SEQ ID NO: 165 encodes the N-terminal ABE8e sequence as follows: human CMV enhancer and promoter nt 1 to 522; putative transcription start site nt 523; 5' untranslated region containing a Kozak sequence nt 523 to 540; 5' ABE8e coding sequence nt 541 to 2706; 5' synthetic intron sequence nt 2707 to 2804; 5' trimodal kissing loop dimerization domain nt 2805 to 2916; linker nt 2917 to 2924; and polyadenylation signal nt 2925 to 2973. SEQ ID NO: 166 encodes the C-terminal Abe8e sequence as follows: human CMV enhancer and promoter nt 1 to 522; putative transcription start site nt 523; 3' trimodal kissing loop dimerization domain nt 525 to 636; 3' synthetic intron sequence nt 637 to 747; 3' ABE8e coding sequence nt 748 to 3399; linker nt 3400 to 3427; polyadenylation signal nt 3428 to 3560.

[0109] SEQ ID NO: 167 is an exemplary kissing loop domain (GATTTTTGACCTGCTCGATTGTCCACTGCGAGCAGGTCTTTTGGAGTCGGGCGAGGCGGAAGCCCGACTCCTTTTGGCATGCACGCTAGCCGCGTCGTGCATGCCTTTTATC).

[0110] SEQ ID NO: 168 is an exemplary ISE, M2(GGGTTATGGGACC).

[0111] SEQ ID NO: 169 is an exemplary ISE, cTNT(GGCTGAGGGAAGGACTGTCCTGGG).

[0112] SEQ ID NO: 170 is an exemplary DISE, rat FGFR2 (CTCTTTCTTTCCATGGGTTGGCCT).

[0113] SEQ ID NOs: 171 and 172 are exemplary constructs that can be used to express the full-length YFP coding sequence. SEQ ID NO: 171 encodes the N-terminal YFP sequence as follows: human CMV enhancer and promoter nt 1 to 522; putative transcription start site nt 523; 5' untranslated region containing the Kozak sequence nt 523 to 543; 5' stuffer open reading frame nt 544 to 3654; self-cleaving 2A sequence nt 3655 to 3729; 5' yellow fluorescent protein segment nt 3730 to 4224; 5' synthetic intron sequence (variable) nt 4225 to 4294; 5' trimodal kissing loop dimerization domain (uppercase): nt 4295 to 4406; linker nt 4407 to 4414; and polyadenylation signal nt 4415 to 4463. SEQ ID NO: 172 encodes the C-terminal YFP sequence as follows: Name: 3' intron screening split YFP; human CMV enhancer and promoter nt 1-522; putative transcription start site nt 523; 3' trimodal kissing loop dimerization domain nt 525-636; 3' synthetic intron sequence (variable) nt 637-706; 3' yfp coding sequence nt 707-940; self-cleaving 2A sequence nt 941-1006; 3' stuffer open reading frame nt 1007-4228; linker nt 4229-4265; polyadenylation signal nt 4257-4388.

[0114] SEQ ID NOs: 173 to 180 are exemplary intron splicing enhancer sequences.

[0115] SEQ ID NO: 181 is a scrambled sequence.

[0116] SEQ ID NOs: 182 to 196 are exemplary intron splicing enhancer sequences.

[0117] SEQ ID NOs: 197 to 198 are scrambled sequences.

[0118] SEQ ID NOs: 199-203 are exemplary intron splicing enhancer sequences.

[0119] SEQ ID NO: 204 is a scrambled sequence.

[0120] SEQ ID NO: 205 is an exemplary branch point sequence (TACTAACA).

[0121] SEQ ID NO: 206 is an exemplary polyadenylation signal AATAAAAATATCTTTATTTTCATTACATCTGTGTGTTGGTTTTTTGTGTG.

[0122] SEQ ID NOs:207 and 208 are exemplary Cas9 coding and protein sequences, respectively.

[0123] SEQ ID NOs:209 and 210 are exemplary dCas9 coding and protein sequences, respectively.

[0124] SEQ ID NOs: 211 and 212 are the nucleic acid and amino acid sequences, respectively, of an exemplary Cas13d.

[0125] SEQ ID NOs: 213 and 214 are the nucleic acid and amino acid sequences, respectively, of an exemplary Cas13d.

[0126] SEQ ID NOs: 215 and 216 are exemplary dead Cas13d (e.g., catalytically inactive) amino acid sequences.

[0127] SEQ ID NO: 217 is an exemplary native HEPN domain RXXXXH.

[0128] SEQ ID NOs:218-221 are exemplary nuclear localization signal coding sequences and protein sequences.

[0129] SEQ ID NO: 222 is an exemplary Cas13d protein sequence.

[0130] SEQ ID NO: 223 is an exemplary DR sequence that can be included in a gRNA sequence and used with the Cas13d protein of SEQ ID NO: 212 for RNA editing.

[0131] SEQ ID NO: 224 is an exemplary DR sequence that can be included in a gRNA sequence and used with the Cas13d protein of SEQ ID NO: 224 for RNA editing.

[0132] SEQ ID NOs: 225 and 226 are exemplary constructs that can be used to express the full-length humanized adenine base editor (ABE8e) coding sequence in vivo. SEQ ID NO:225 encodes the N-terminal ABE8e sequence and contains two gRNA expression cassettes as follows: nt 1-141 AAV2 inverted repeat; nt 159-255 CRISPR gRNA (reverse), nt 256-504 human U6 RNA polymerase III promoter (reverse), nt 512-1019 CMV promoter, nt 1034-1051 5' untranslated region, nt 1052-3217 N-terminal ABE8e editor, nt 3218-3315 synthetic intron sequence, nt 3316-3427 dimerization domain, nt 3436-3485 polyadenylation signal, nt 3492-3706 H1 polymerase III promoter, nt 3707-3803 CRISPR gRNA, and nt 3825-3965 AAV2 inverted repeat. SEQ ID NO:225 contains SEQ ID NO:165 and an additional gRNA expression cassette. SEQ ID NO:226 encodes the C-terminal ABE8e sequence and contains two gRNA expression cassettes as follows: nt 1-141 AAV2 inverted repeat, nt 159-255 CRISPR gRNA (reverse), nt 256-504 human U6 RNA polymerase III promoter (reverse), nt 512-1019 CMV promoter, nt 1036-1147 dimerization domain, nt 1259-3910 3' ABE8e coding sequence, nt 3939-4069 polyadenylation signal, nt 4078-4292 H1 polymerase III promoter, nt 4293-4389 CRISPR gRNA, and nt 4411-4551 AAV2 inverted repeat. SEQ ID NO:226 contains SEQ ID NO:166 and an additional gRNA expression cassette.

[0133] Detailed Description Unless otherwise specified, technical terms are used according to conventional usage.The definitions of common terms in molecular biology can be found in Benjamin Lewin, Genes VII, published by Oxford University Press, 1999;Kendrew et al. (eds.), The Encyclopedia of Molecular Biology, published by Blackwell Science Ltd., 1994;And Robert A. Meyers (ed.), Molecular Biology and Biotechnology: a Comprehensive Desk Reference, published by VCH Publishers, Inc., 1995;And other similar references.

[0134] As used herein, the singular forms "a," "an," and "the" refer to both the singular and the plural unless the context clearly indicates otherwise. As used herein, the term "comprises" means "includes." Thus, "comprising a nucleic acid molecule" does not exclude other elements and means "including a nucleic acid molecule." It should further be understood that any and all base sizes given for nucleic acids are approximate unless otherwise specified and are provided for illustrative purposes. While many methods and materials similar or equivalent to those described herein can be used, certain suitable methods and materials are described below. In case of conflict, the present specification, including explanations of terms, will control. Furthermore, the materials, methods, and examples are merely illustrative and not limiting. All references, including patent applications and patents, and GenBank accession numbers, are incorporated herein by reference in their entirety. In order to facilitate review of the various embodiments disclosed herein, the following explanations of specific terms are provided:

[0135] Administration: Providing or administering to a subject an agent such as a therapeutic nucleic acid molecule (e.g., encoding one or more portions of a nucleic acid editor protein, a gRNA, or both) or other therapeutic agent presented herein by any effective route. Exemplary administration routes include, but are not limited to, injection (e.g., subcutaneous, intramuscular, intradermal, intraperitoneal, intrathecal, intratumoral, intraosseous, and intravenous), transdermal, intranasal, and inhalation routes. Administration can be systemic or local.

[0136] Aptamer: A nucleic acid molecule (e.g., DNA or RNA) that binds with high affinity and specificity to a specific target agent or molecule. Aptamers can be used as dimerization domains in the nucleic acid molecules disclosed herein. In one example, two aptamers can bind to each other, for example, by canonical base pairing, non-canonical base pairing interactions, non-base pairing interactions, or a combination thereof, to mediate dimer formation. In one example, an aptamer allows RNA dimer formation (and subsequent recombination) only in the presence of one or more targets recognized by the aptamers. Aptamers have been obtained by a combinatorial selection process called systematic evolution of ligands by exponential enrichment (SELEX) (e.g., Ellington et al., Nature 1990, 346, 818-822; Tuerk and Gold Science 1990, 249, 505-510; Liu et al., Chem. Rev. 2009, 109, 1948-1998; Shamah et al., Acc. Chem. Res. 2008, 41, 130-138; Famulok, et al., Chem. Rev. 2007, 107, 3715-3743; Manimala et al., Recent Dev. Nucleic Acids Res. 2004, 1, 207-231; Famulok et al., Acc. Chem. Res. 2000, 33, 591-599; Hesselberth, et al., Rev. Mol. Biotech. 2000, 74, 15-25; Wilson et al., Annu. Rev. Biochem. 1999, 68, 611-647; Morris et al., Proc. Natl. Acad. Sci. USA 1998, 95, 2902-2907). , DNA or RNA molecules capable of binding to a target molecule of interest are 14 ~10 15Aptamers are selected from nucleic acid libraries consisting of diverse sequences by repeated steps of selection, amplification, and mutation. The affinity of aptamers for their targets can compete with that of antibodies, with dissociation constants in the low picomolar range (Morris et al., Proc. Natl. Acad. Sci. USA 1998, 95, 2902-2907; Green et al., Biochemistry 1996, 35, 14413-14424).

[0137] Specific aptamers have been identified for a wide range of targets, from small organic molecules such as adenosine to proteins such as thrombin, and even viruses and cells (Liu et al., Chem. Rev. 2009, 109, 1948-1998; Lee et al., Nucleic Acids Res. 2004, 32, D95-D100; Navani and Li, Curr. Opin. Chem. Biol. 2006, 10, 272-281; ​​Song et al., TrAC, Trends Anal. Chem. 2008, 27, 108-117).For example, metal ions such as Zn(II) (Ciesiolka et al., RNA 1: 538-550, 1995) and Ni(II) (Hofmann et al., RNA, 3: 1289-1300, 1997); nucleotides such as adenosine triphosphate (ATP) (Huizenga and Szostak, Biochemistry, 34: 656-665, 1995); and guanine (Kiga et al., Nucleic Acids Res., 26: 1755-60, 1998); cofactors such as NAD (Kiga et al., Nucleic Acids Res., 26: 1755-60, 1998) and flavins (Lauhon and Szostak, J. Am. Chem. Soc., 117: 1246-57, 1995); antibiotics such as viomycin (Wallis et al., Chem. Biol. 4: 357-366, 1997) and streptomycin (Wallace and Schroeder, RNA 4: 112-123, 1998); proteins such as HIV reverse transcriptase (Chaloin et al., Nucleic Acids Res., 30: 4001-8, 2002) and hepatitis C virus RNA-dependent RNA polymerase (Biroccio et al., J. Virol. 76: 3688-96, 2002); toxins such as cholera toxin and staphylococcal enterotoxin B (Bruno and Kiel, BioTechniques, 32: pp. 178-180 and 182-183, 2002); and bacterial spores such as Bacillus anthracis (Bruno and Aptamers that recognize the nucleotide sequence of the nucleotide sequence of interest (Kiel, Biosensors & Bioelectronics, 14: 457-464, 1999) are available.

[0138] Binding: An association between two substances or molecules, such as hybridization between one nucleic acid molecule and another nucleic acid molecule (or itself), or binding between an aptamer and its target, such as between two dimerization domains. An oligonucleotide molecule and another nucleic acid molecule bind or are stably bound when there are a sufficient number of complementary base pairs between the oligonucleotide molecules and the target nucleic acid to allow for detection of binding. In some examples, binding between nucleic acid molecules can occur directly. In some examples, binding between nucleic acid molecules can occur indirectly, for example, through an intermediate molecule. Either direct or indirect binding can occur by standard base pairing, non-standard base pairing interactions, non-base pairing interactions, or combinations thereof. Non-standard base pairing interactions can occur by any stabilization means known to those skilled in the art, including, but not limited to, Hoogsteen base pairs and wobble base pairs. Non-base pairing interactions can include binding through an intermediate molecule. In some examples, direct binding is between kissing loop dimerization domains. In some examples, direct binding is between low diversity dimerization domains. In some embodiments, the direct binding is between aptamer regions. In some embodiments, the direct binding between aptamer regions involves non-canonical base pairing interactions. In some embodiments, the direct binding between aptamer regions involves canonical base pairing and non-canonical base pairing interactions. In some embodiments, the indirect binding occurs through a nucleic acid bridge. In some embodiments, the nucleic acid bridge is mRNA. A non-limiting example of a nucleic acid bridge is shown in FIG. 7B. In some embodiments, the indirect binding occurs through an aptamer molecule. A non-limiting example of indirect binding through an aptamer molecule is shown in FIG. 7A. In some embodiments, the indirect binding through an aptamer molecule involves non-base pairing interactions between the aptamer molecule and the binding region. In some embodiments, the indirect binding through an aptamer molecule involves non-base pairing interactions between the aptamer molecule and the binding region, and base pairing interactions between the binding regions.

[0139] C-terminal portion: A region of a protein sequence comprising a contiguous stretch of amino acids beginning at or near the C-terminal residue of the protein. The C-terminal portion of a protein can be defined by a contiguous stretch of amino acids (e.g., several amino acid residues).

[0140] Cancer: Malignant tumor characterized by abnormal or uncontrolled cell growth Other features often associated with cancer include metastasis, interference with the normal function of nearby cells, release of abnormal levels of cytokines or other secretory products, and suppressed or exacerbated inflammatory or immunological responses, and invasion of surrounding or distant tissues or organs, such as lymph nodes. "Metastatic disease" refers to cancer cells that leave the original tumor site and travel to other parts of the body, for example, via the bloodstream or lymphatic system.

[0141] Cas9: An RNA-guided DNA endonuclease enzyme involved in CRISPR-Cas immune defense against prokaryotic viruses. Cas9 has two active cutting sites (HNH and RuvC), one on each strand of the double helix. An exemplary native Cas9 sequence from S. pyogenes is shown in SEQ ID NO: 208.

[0142] The present disclosure also encompasses catalytically inactive (deactivated or dead) Cas9 (dCas9) proteins that have reduced or eliminated endonuclease activity but still bind to dsDNA. In some examples, dCas9 contains one or more mutations in the RuvC and HNH nuclease domains, such as one or more of the following point mutations: D10A, E762A, D839A, H840A, N854A, N863A, and D986A (e.g., based on the numbering in SEQ ID NO: 208). An exemplary dCas9 sequence with D10A and H840A substitutions is shown in SEQ ID NO: 210. In one example, the dCas9 protein has the mutations D10A, H840A, D839A, and N863A (see, e.g., Esvelt et al., Nat. Meth. 10: 1116-21, 2013).

[0143] Cas9 and dCas9 sequences are publicly available. For example, Cas9 nucleic acids are disclosed by GenBank® Accession Nos. CP012045.1, nucleotides 796693..800799, and CP014139.1, nucleotides 1100046..1104152, and Cas9 proteins are disclosed by GenBank® Accession Nos. NP_269215.1, AMA70685.1, and AKP81606.1. In some embodiments, inactivated forms of Cas9 (dCas9) are nuclease-deficient (e.g., those shown in GenBank® Accession Nos. AKA60242.1 and KR011748.1). Activatable Cas9 proteins are provided in U.S. Patent Application Publication No. 2018-0073002-A1. In certain examples, the Cas9 or dCas9 used in the compositions or methods disclosed herein have at least 80% sequence identity, e.g., at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity, to such sequences (e.g., SEQ ID NOs: 207, 208, 209, and 210), and retain their usability in the compositions and methods disclosed herein (e.g., can be encoded by two or more separate molecules disclosed herein and then recombined via the REJ methodology provided herein).

[0144] Cas13d: An RNA-guided RNA endonuclease enzyme capable of cleaving or binding to RNA. Cas13d protein specifically recognizes direct repeat (DR) sequences present in gRNAs with specific secondary structures. Cas13d protein contains one or two HEPN domains. The native HEPN domain contains the sequence RXXXXH (SEQ ID NO: 217), where X is any amino acid. Catalytically inactive or "dead" Cas13ds containing mutated HEPN domain(s) and thus unable to cleave RNA but capable of processing gRNA are also encompassed by the present disclosure (see, for example, SEQ ID NOs: 215 and 216). Such dead Cas13ds (dCas13ds) can be targeted to cis-elements of pre-mRNAs to manipulate alternative splicing.

[0145] Exemplary native and variant Cas13d protein sequences are provided in WO2019 / 040664, US 10,876,101 and US 10,392,616 (all incorporated by reference in their entirety), and herein as SEQ ID NOs: 212, 214, 215, 216, and 222.

[0146] In one example, a full-length (non-truncated) Cas13d protein is between 870 and 1080 amino acids in length. In one example, the Cas13d protein is derived from a genomic or metagenomic sequence of a bacterium of the order Clostridiales. In one example, the corresponding DR sequence of the Cas13d protein is located at the 5' end of a spacer sequence within a molecule comprising the Cas13d gRNA. In one example, the DR sequence within the Cas13d gRNA is truncated at the 5' end (e.g., by at least 1 nt, at least 2 nt, at least 3 nt, at least 4 nt, at least 5 nt, e.g., by 1-3 nt, 3-6 nt, 5-7 nt, or 5-10 nt) compared to the DR sequence of an unprocessed Cas13d guide array transcript. In one example, the DR sequence within the Cas13d gRNA is truncated at the 5' end by the Cas13d protein by 5-7 nt. In one example, the Cas13d protein can cleave the target RNA adjacent to the 3' end of the spacer-target duplex at either a, U, G, or C ribonucleotide, and adjacent to the 5' end at either a, U, G, or C ribonucleotide.

[0147] In one example, the Cas13d protein has at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to SEQ ID NO: 212, 214, 215, 216, or 222. In one example, the Cas13d coding sequence encodes a Cas13d protein that has at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to SEQ ID NO: 212, 214, 215, 216, or 222. In one example, the Cas13d coding sequence has at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to SEQ ID NO: 211 or 213.

[0148] Complementarity: The ability of a nucleic acid to form hydrogen bond(s) with another nucleic acid sequence, either by conventional Watson-Crick base pairing or other non-conventional types. Percent complementarity refers to the percentage of residues in a nucleic acid molecule (e.g., target DNA or RNA) that can form hydrogen bonds (e.g., Watson-Crick base pairing) with a second nucleic acid sequence, e.g., a gRNA (e.g., 5, 6, 7, 8, 9, 10 out of 10, which represent 50%, 60%, 70%, 80%, 90%, and 100% complementarity). "Fully complementary" means that all contiguous residues of a nucleic acid sequence will hydrogen bond with the same number of contiguous residues in a second nucleic acid sequence. "Substantially complementary," as used herein, refers to a degree of complementarity of at least 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 97%, 98%, 99%, or 100% over a region of 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 30, 35, 40, 45, 50, or more nucleotides, or that two nucleic acids hybridize under stringent conditions. Thus, in some embodiments, the first dimerization domain and the second dimerization domain are fully complementary to each other (e.g., 100%). In other examples, the first dimerization domain and the second dimerization domain are substantially complementary (eg, at least 80%) to each other.

[0149] Contacting: To be placed in direct physical association, including in solid or liquid form. Contacting can occur in vitro or ex vivo, for example, by adding a reagent to a sample (such as one containing cells), or in vivo, by administration to a subject.

[0150] CRISPR / Cas system: A prokaryotic immune system that confers resistance to foreign genetic elements such as plasmids and phages, providing a form of adaptive immunity. The system includes a Cas nuclease (e.g., Cas9, Cas13d) and a guide RNA (gRNA) that specifically binds to a target RNA or DNA and directs the Cas nuclease to the target site. The compositions, systems, and methods disclosed herein can be used to express Cas nucleases from two or more different DNA molecules, which in some examples further encode one or more gRNAs to regulate gene expression, e.g., increase or decrease expression of a target nucleic acid molecule, and / or edit the sequence of a target nucleic acid molecule (e.g., to repair one or more mutations associated with a disease, e.g., substitutions, insertions, or deletions).

[0151] Dead guide RNA (dgRNA): A guide RNA (gRNA) that can guide a wild-type Cas nuclease (e.g., Cas9) to a target nucleic acid but does not induce double-stranded DNA cleavage. A shortened gRNA contains a shortened targeting sequence of about 14-15 nucleotides, while a non-dead gRNA contains a targeting sequence of about 20 nucleotides. dgRNAs are further described, for example, in Dahlman et al. (2015) Nat. Biotechnol. 33: 1159-1161; Kiani et al. (2015) Nat. Methods, 12: 1051-1054; and Hsin-Kai Liao et al. (2017) Cell, 171: 1495-1507, all of which are incorporated herein by reference in their entirety. In some examples, the dgRNA is an RNA molecule (e.g., when expressed in a cell). In some examples, the dgRNA is encoded by a DNA molecule (e.g., in a vector, such as a viral vector).

[0152] DNA editing: A type of genetic engineering that uses nucleases (e.g., Cas9, Cas13d, or their dead versions) to create site-specific strand breaks at desired locations within DNA to insert, delete, or replace DNA molecules (or DNA nucleotides) in cells or organisms. The induced breaks are repaired, resulting in targeted mutation or repair. For example, the CRISPR / Cas method, which uses the REJ system presented herein to express Cas nucleases, can be used to edit one or more target DNA sequences, such as those associated with cancer (e.g., breast cancer, colon cancer, lung cancer, prostate cancer, melanoma), infectious diseases (e.g., HIV, hepatitis, HPV, and West Nile virus), or neurodegenerative disorders (e.g., Huntington's disease or ALS). For example, DNA editing can be used to treat disease or viral infection.

[0153] DNA insertion site: A site in DNA that is targeted for or has received insertion of an exogenous polynucleotide. The methods disclosed herein encompass the use of nucleic acid editors expressed from two or more nucleic acid molecules presented herein that can be used to target DNA for manipulation at the DNA insertion site.

[0154] Downregulated or knocked down: When used in reference to the expression of a molecule such as a target nucleic acid or protein, refers to any process that results in a reduction in the production of the target nucleic acid or protein, but in some examples does not result in the complete elimination of the target RNA product or target nucleic acid function. In one example, downregulation or knockdown does not result in the complete elimination of the detectable expression or activity of the target nucleic acid / protein. In some examples, downregulation or knockdown of a target nucleic acid includes processes that may reduce the translation of the target RNA and therefore the presence of the corresponding protein. The systems disclosed herein can be used to downregulate any target nucleic acid / protein of interest.

[0155] Downregulation or knockdown includes any detectable reduction in a target nucleic acid / protein. In certain examples, detectable target nucleic acid / protein in a cell or cell-free system is reduced by at least 10%, at least 20%, at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, at least 75%, at least 80%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% (e.g., 40%-90%, 40%-80%, or 50%-95% reduction) compared to a control (e.g., the amount of target nucleic acid / protein detected in a corresponding untreated cell or sample). In one example, the control is the relative amount of expression in a normal cell (e.g., a non-recombinant cell that does not contain a nucleic acid molecule for RNA recombination as provided herein).

[0156] Effective amount: An amount of an agent (e.g., a system providing multiple vectors, each encoding a different portion of a nucleic acid editing protein, such as a Cas9 or Cas13d protein, in combination with, e.g., an effective amount of, one or more gRNAs that can hybridize to a nucleic acid target) that is sufficient to produce a beneficial or desired result. An effective amount can also refer to the amount of correctly bound RNA or nucleic acid editing protein produced that is sufficient, e.g., in combination with, one or more gRNAs that can hybridize to a nucleic acid target, to produce a beneficial or desired result.

[0157] The effective amount (also referred to as a therapeutically effective amount) may vary depending on one or more of the subject and disease state being treated, the subject's weight and age, the severity of the disease state, the mode of administration, etc., and can be determined by one skilled in the art. Beneficial therapeutic effects can include the feasibility of a diagnostic determination; amelioration of a disease, symptom, disorder, or pathological condition; reduction or prevention of the onset of a disease, symptom, disorder, or condition; and generally negating a disease, symptom, disorder, or pathological condition.

[0158] In one embodiment, an "effective amount" of two or more synthetic nucleic acid molecules provided herein is sufficient to treat a disease, such as a genetic disease or cancer. In one embodiment, an "effective amount" of two or more synthetic nucleic acid molecules provided herein is an amount sufficient to increase the survival time of a treated patient by, for example, at least 10%, at least 20%, at least 25%, at least 50%, at least 70%, at least 75%, at least 80%, at least 90%, at least 95%, at least 99%, at least 100%, at least 200%, at least 300%, at least 400%, at least 500%, or at least 600% (compared to not administering the two or more synthetic nucleic acid molecules provided herein). In one embodiment, an "effective amount" of two or more synthetic nucleic acid molecules provided herein is an amount sufficient to increase the survival time of a treated patient by, for example, at least 6 months, at least 9 months, at least 1 year, at least 1.5 years, at least 2 years, at least 2.5 years, at least 3 years, at least 4 years, at least 5 years, at least 10 years, at least 12 years, at least 15 years, or at least 20 years (compared to not administering the two or more synthetic nucleic acid molecules provided herein). In one embodiment, an "effective amount" of two or more synthetic nucleic acid molecules provided herein is an amount sufficient to increase mobility in a treated patient (such as a DMD patient) by, for example, at least 10%, at least 20%, at least 25%, at least 50%, at least 70%, at least 75%, at least 80%, at least 90%, at least 95%, at least 99%, at least 100%, at least 200%, at least 300%, at least 400%, at least 500%, or at least 600% (compared to not administering two or more synthetic nucleic acid molecules provided herein).In one embodiment, an "effective amount" of two or more synthetic nucleic acid molecules provided herein is an amount sufficient to increase mobility in a treated patient (such as a DMD patient) by, for example, at least 10%, at least 20%, at least 25%, at least 50%, at least 70%, at least 75%, at least 80%, at least 90%, at least 95%, at least 99%, at least 100%, at least 200%, at least 300%, at least 400%, at least 500%, or at least 600% (compared to not administering two or more synthetic nucleic acid molecules provided herein). In one embodiment, an "effective amount" of two or more synthetic nucleic acid molecules provided herein is an amount sufficient to increase the cognitive ability of a treated patient (such as a DMD patient) by, for example, at least 10%, at least 20%, at least 25%, at least 50%, at least 70%, at least 75%, at least 80%, at least 90%, at least 95%, at least 99%, at least 100%, at least 200%, at least 300%, at least 400%, at least 500%, or at least 600% (compared to not administering the two or more synthetic nucleic acid molecules provided herein). In one embodiment, an "effective amount" of two or more synthetic nucleic acid molecules provided herein is an amount sufficient to increase respiratory function in a treated patient (such as a DMD patient) by, for example, at least 10%, at least 20%, at least 25%, at least 50%, at least 70%, at least 75%, at least 80%, at least 90%, at least 95%, at least 99%, at least 100%, at least 200%, at least 300%, at least 400%, at least 500%, or at least 600% (compared to not administering the two or more synthetic nucleic acid molecules provided herein).In one embodiment, an "effective amount" of two or more synthetic nucleic acid molecules presented herein is an amount sufficient to increase blood clotting in a treated patient (such as a hemophilia patient) by, for example, at least 10%, at least 20%, at least 25%, at least 50%, at least 70%, at least 75%, at least 80%, at least 90%, at least 95%, at least 99%, at least 100%, at least 200%, at least 300%, at least 400%, at least 500%, or at least 600% (compared to not administering two or more synthetic nucleic acid molecules presented herein). In one embodiment, an "effective amount" of two or more synthetic nucleic acid molecules provided herein is an amount sufficient to increase the vision of a treated patient (such as a patient with Usher syndrome or Stargardt disease) by, for example, at least 10%, at least 20%, at least 25%, at least 50%, at least 70%, at least 75%, at least 80%, at least 90%, at least 95%, at least 99%, at least 100%, at least 200%, at least 300%, at least 400%, at least 500%, or at least 600% (compared to not administering the two or more synthetic nucleic acid molecules provided herein). In one embodiment, an "effective amount" of two or more synthetic nucleic acid molecules provided herein is an amount sufficient to increase the hearing of a treated patient (such as a patient with Usher syndrome) by, for example, at least 10%, at least 20%, at least 25%, at least 50%, at least 70%, at least 75%, at least 80%, at least 90%, at least 95%, at least 99%, at least 100%, at least 200%, at least 300%, at least 400%, at least 500%, or at least 600% (compared to not administering the two or more synthetic nucleic acid molecules provided herein).

[0159] In one embodiment, an "effective amount" of two or more synthetic nucleic acid molecules provided herein is an amount sufficient to reduce the size of the gastrocnemius muscle in a treated DMD patient by, for example, at least 10%, at least 20%, at least 25%, at least 50%, at least 70%, at least 75%, at least 80%, at least 90%, or at least 95% (compared to not administering the two or more synthetic nucleic acid molecules provided herein). In one embodiment, an "effective amount" of two or more synthetic nucleic acid molecules provided herein is an amount sufficient to reduce the size of the cardiomyopathic muscle in a treated DMD patient by, for example, at least 10%, at least 20%, at least 25%, at least 50%, at least 70%, at least 75%, at least 80%, at least 90%, or at least 95% (compared to not administering the two or more synthetic nucleic acid molecules provided herein). In some examples, a combination of these effects is achieved.

[0160] Guide RNA (gRNA): A synthetic nucleic acid sequence used to direct a Cas nuclease (or dead Cas nuclease) protein to a target nucleic acid sequence, such as a target DNA (e.g., a genomic sequence) or a target RNA sequence. A gRNA molecule, whether as part of a single nucleic acid molecule or split across two or more nucleic acid molecules, comprises: (1) a portion having sequence complementarity to a target nucleic acid (e.g., at least 80%, at least 90%, at least 95%, or 100% sequence complementarity); and (2) a portion having a secondary structure that binds to a Cas nuclease. Thus, the target nucleic acid of a Cas protein can be altered simply by changing the target sequence present within the gRNA (see CRISPR-Cas9 Structures and Mechanisms. Fuguo Jiang and Jennifer A. Doudna, Annual Review of Biophysics, 46:1, 505-529 (2017)). In some examples, the gRNA is an RNA molecule (e.g., when expressed in a cell). In some examples, the gRNA is encoded by a DNA molecule (e.g., when part of a vector, such as a viral vector). The gRNA may contain modified bases or chemical modifications (e.g., Latorre et al., Angewandte Chemie 55: 3548-50, 2016).

[0161] In some examples, the gRNA comprises two or more MS2-binding loop sequences, which can be modified from the native MS2-binding loop sequence to have increased GC content and / or shortened repeat content. In some examples, the gRNA is modified to have increased GC content and / or shortened repeat content.

[0162] In some examples, the gRNA is a dead guide RNA (dgRNA). Increasing the GC content and / or shortening the repeat content of the gRNA can be used to convert the gRNA into a dgRNA, i.e., a guide nucleic acid molecule that can direct a Cas nuclease to a target nucleic acid sequence but does not induce DNA double-strand breaks (or RNA single-strand breaks).

[0163] In one example, the gRNA directs a Cas DNA nuclease (e.g., Cas9) to a target DNA. In one such example, the gRNA includes, from 5' to 3', (1) a CRISPR RNA (crRNA) region containing sequences designed to hybridize to (and, in some examples, edit) the target DNA sequence, and a region that hybridizes to a trans-activating crRNA (tracrRNA), and (2) a scaffolding sequence (tracrRNA) required for binding to Cas. Thus, the gRNA may combine the crRNA and tracrRNA into a single RNA transcript (referred to in the art as a single guide RNA, sgRNA; for simplicity, it is encompassed herein by the term "gRNA"). In this example, a region of the crRNA hybridizes to the tracrRNA to form a unique double-RNA hybrid structure that binds to and guides the Cas endonuclease protein to the target DNA molecule. In some examples, the cRNA and tracrRNA are two separate RNA molecules (e.g., a two-piece gRNA; for simplicity, this is also encompassed by the term "gRNA" herein). In some examples, a protospacer adjacent motif (PAM) immediately follows the site of the target DNA to be edited (e.g., the Cas9 cleavage site approximately 3 nt upstream of the PAM).

[0164] In another example, the gRNA directs a Cas RNA nuclease (e.g., Cas13d) to a target RNA. In one such example, the gRNA comprises, from 5' to 3', (1) a crRNA containing a direct repeat (DR) region and (2) a spacer region, e.g., for Cas13a, Cas13c, and Cas13d nucleases. In one example, the gRNA comprises a DR of approximately 36 nt followed by a spacer sequence of approximately 28-32 nt. In another such example, the gRNA comprises, from 5' to 3', a crRNA containing (1) a spacer and (2) a DR region, e.g., for Cas13b nuclease. In some examples, the gRNA is processed (shortened / modified) by a Cas RNA nuclease or other RNase to produce a shorter "mature" form. The DR is a constant portion of the gRNA that contains secondary structure that facilitates interaction of the Cas RNA nuclease protein with the gRNA. The spacer portion is a variable portion of the gRNA that contains a sequence designed to hybridize to (and in some embodiments, edit) the target RNA sequence. In some embodiments, the full-length spacer is about 28-32 nt (e.g., 30-32 nt) in length, while the mature (processed) spacer is about 14-30 nt.

[0165] Hybridization: Nucleic acid hybridization occurs when two nucleic acid molecules form a certain amount of hydrogen bonds with each other.The stringency of hybridization can vary depending on the environmental conditions surrounding nucleic acid, the nature of hybridization method, and the composition and length of nucleic acid used.Calculations of the hybridization conditions required to achieve a certain degree of stringency can be found in Sambrook et al., Molecular Cloning: A Laboratory Manual (Cold Spring Harbor Laboratory Press, Cold Spring Harbor, NY, 2001); and Tijssen, Laboratory Techniques in Biochemistry and Molecular Biology—Hybridization with Nucleic Acid Probes Part I, Chapter 2 (Elsevier, New York, 1993). m is the temperature at which 50% of a given strand of nucleic acid hybridizes to its complementary strand.

[0166] Increase or decrease: A statistically significant positive or negative change in quantity, respectively, from a control value (e.g., a value representing no therapeutic agent, such as no administration of two or more synthetic nucleic acid molecules provided herein). An increase is a positive change, e.g., an increase of at least 50%, at least 100%, at least 200%, at least 300%, at least 400%, or at least 500% compared to the control value. A decrease is a negative change, e.g., a decrease of at least 20%, at least 25%, at least 50%, at least 75%, at least 80%, at least 90%, at least 95%, at least 98%, at least 99%, or at least 100% compared to the control value. In some examples, the decrease is less than 100%, e.g., a decrease of 90% or less, 95% or less, or 99% or less.

[0167] Isolated: An "isolated" biological component (e.g., a nucleic acid molecule or protein) is one that has been substantially separated from, made separately from, or purified away from other biological components, e.g., other cells (e.g., RBCs), chromosomal and extrachromosomal DNA and RNA, and proteins, within the cells or tissues of the organism in which it resides. "Isolated" nucleic acids and proteins include nucleic acids and proteins purified by standard purification methods. The term also encompasses nucleic acids and proteins prepared by recombinant expression in a host cell as well as chemically synthesized nucleic acids and proteins.

[0168] Kissing loop / kissing stem loop: An RNA structure formed when the bases between two hairpin loops form paired interactions. These intermolecular "kissing interactions" occur when an unpaired nucleotide in one hairpin loop base pairs with an unpaired nucleotide in another hairpin loop, forming a stable interaction complex. See Figure 9A for an example.

[0169] N-terminal portion: A region of a protein sequence comprising a contiguous stretch of amino acids beginning with the N-terminal residue of the protein. The N-terminal portion of a protein can be defined by a contiguous stretch of amino acids (e.g., several amino acid residues).

[0170] Non-naturally occurring, synthetic, or engineered: Terms used interchangeably herein and indicating the involvement of the hand of man. These terms, when referring to a nucleic acid molecule or polypeptide, indicate that the nucleic acid molecule or polypeptide is at least substantially free from at least one other component with which it is naturally associated and found in nature. Additionally, these terms can indicate that the nucleic acid molecule or polypeptide has a sequence not found in nature.

[0171] Nucleic acid molecule: A deoxyribonucleotide (DNA) or ribonucleotide (RNA) polymer that may contain natural nucleotides / ribonucleotides and / or analogs of natural nucleotides / ribonucleotides that hybridize to nucleic acid molecules in the same way as naturally occurring nucleotides. A nucleic acid molecule may be a single-stranded (ss) DNA or RNA molecule or a double-stranded (ds) nucleic acid molecule. RNA or mRNA, as used herein, may refer to a pre-mRNA molecule or a mature RNA transcript. A pre-mRNA molecule contains sequences that are removed by processing, such as intron sequences that are removed by splicing after the binding of the dimerization domain described herein. A nucleic acid molecule described herein may be, for example, a DNA molecule from which RNA is transcribed by a promoter on the DNA in the case of a DNA expression vector.

[0172] Operably linked: A first nucleic acid sequence is operably linked to a second nucleic acid sequence when the first nucleic acid sequence is placed in a functional relationship with the second nucleic acid sequence. For example, a promoter is operably linked to a nucleic acid sequence when the promoter affects expression of the nucleic acid sequence, e.g., when the promoter affects transcription of a pre-mRNA that, when spliced, can result in expression of a protein (e.g., a portion of a nucleic acid editing protein coding sequence).

[0173] Pharmaceutically acceptable carriers: Pharmaceutically acceptable carriers useful in this disclosure are conventional. Remington's Pharmaceutical Sciences, by E. W. Martin, Mack Publishing Co., Easton, PA, 15th Edition (1975), describes compositions and formulations suitable for pharmaceutical delivery of therapeutic agents, such as the nucleic acid molecules disclosed herein.

[0174] Generally, the nature of the carrier depends on the specific administration mode used.For example, parenteral preparations usually contain an injectable fluid as a vehicle, which includes pharmaceutically and physiologically acceptable fluids such as water, physiological saline, balanced salt solution, aqueous dextrose, glycerol, etc. In addition to biologically neutral carriers, the pharmaceutical composition to be administered may contain minor amounts of non-toxic auxiliary substances, such as wetting agents or emulsifying agents, preservatives, and pH buffering agents, for example, sodium acetate or sorbitan monolaurate.

[0175] Polypeptide, peptide, and protein: These refer to polymers of amino acids of any length. The polymers may be linear or branched, may contain modified amino acids, and may be interrupted by non-amino acids. The terms also encompass modified amino acid polymers; for example, amino acid polymers that have undergone any other manipulation, such as disulfide bond formation, glycosylation, lipidation, acetylation, phosphorylation, or conjugation with a labeling component. As used herein, the term "amino acid" includes natural and / or unnatural or synthetic amino acids, including glycine and both the D- and L-optical isomers, as well as amino acid analogs and peptidomimetics. In one example, the protein is a nucleic acid-editing protein, such as Cas9, Cas13d, or a zinc finger nuclease. In one example, the protein is associated with a disease, such as a genetic disorder (see, e.g., Tables 1-4). In one example, the protein is a therapeutic protein, such as one used to treat a disease, such as cancer. In one example, the protein is at least 50 amino acids in length, at least 100 amino acids in length, at least 500 amino acids in length, at least 1000 amino acids in length, at least 1500 amino acids in length, e.g., at least 2000 amino acids, at least 2500 amino acids, at least 3000 amino acids, or at least 5000 amino acids in length.

[0176] Polypyrimidine tract: A region of pre-messenger RNA (mRNA) that facilitates the assembly of the spliceosome, a specialized protein complex for RNA splicing during the process of post-transcriptional modification. This tract can be primarily pyrimidine nucleotides such as uracil, and in some examples is 15-20 base pairs in length, located approximately 5-40 base pairs before the 3' end of the intron to be spliced.

[0177] Promoter / Enhancer: A group of nucleic acid control sequences that direct the transcription of a nucleic acid sequence. A promoter includes necessary nucleic acid sequences near the start site of transcription, such as a TATA element in the case of a polymerase II type promoter. A promoter also includes distal enhancer or repressor elements, which may be located as far as several thousand base pairs from the start site of transcription, as needed. In some embodiments, the promoter sequence plus its corresponding coding sequence is larger than the capacity of an AAV. In some embodiments, the promoter sequence of a target protein is at least 3500 nt, at least 4000 nt, at least 5000 nt, or even at least 6000 nt.

[0178] A "constitutive promoter" is a promoter that is constantly active and is not regulated by external signals or molecules. In contrast, the activity of an "inducible promoter" is regulated by external signals or molecules (e.g., transcription factors). Both constitutive and inducible promoters can be used in the methods and systems provided herein (see, e.g., Bitter et al., Methods in Enzymology 153: 516-544, 1987). Tissue-specific promoters can be used in the methods and systems provided herein, for example, to direct expression primarily in a desired tissue or cell of interest, such as muscle, neurons, bone, skin, blood, a specific organ (e.g., liver, pancreas), or a specific cell type (e.g., lymphocytes). In some examples, the promoters used herein are endogenous to the target protein to be expressed. In some examples, the promoters used herein are exogenous to the target protein to be expressed.

[0179] Also included are promoter elements that are sufficient to render promoter-dependent gene expression cell-type-specific, tissue-specific, or inducible by external signals or agents, and such elements can be located in the 5' or 3' region of the gene. Promoters produced by recombinant DNA or synthetic techniques can also be used to effect transcription of nucleic acid sequences.

[0180] Exemplary promoters that can be used with the methods and systems presented herein include, but are not limited to, the SV40 promoter, the cytomegalovirus (CMV) promoter (optionally with a CMV enhancer), pol III promoters (e.g., the U6 promoter and the H1 promoter), pol II promoters (e.g., the retroviral Rous sarcoma virus (RSV) LTR promoter (optionally with an RSV enhancer), the dihydrofolate reductase promoter, the β-actin promoter, the phosphoglycerol kinase (PGK) promoter, and the EF1α promoter).

[0181] Recombinant: A recombinant nucleic acid molecule or protein sequence has a sequence that does not occur in nature or has a sequence that is created by the artificial combination of two otherwise separate sequence segments (e.g., a viral vector that contains a portion of a nucleic acid editing protein coding sequence, e.g., about one-third, half, or two-thirds of the coding sequence). This artificial combination can be achieved, for example, by chemical synthesis or artificial manipulation of isolated nucleic acid segments, for example, by genetic engineering techniques. Similarly, a recombinant or transgenic cell contains a recombinant nucleic acid molecule.

[0182] RNA editing: A type of genetic engineering that uses engineered nucleases (e.g., Cas13d and dCas13d proteins) to create site-specific strand breaks at desired locations within the RNA to insert, delete, or replace RNA molecules (or ribonucleotides of RNA) in cells or organisms. The induced breaks are repaired, resulting in targeted mutation or repair. For example, the CRISPR / Cas method, which uses the REJ system presented herein to express Cas nucleases, can be used to edit the sequence of one or more target RNAs, such as those associated with cancer (e.g., breast cancer, colon cancer, lung cancer, prostate cancer, melanoma), infectious diseases (e.g., HIV, hepatitis, HPV, and West Nile virus), or neurodegenerative disorders (e.g., Huntington's disease or ALS). For example, RNA editing can be used to treat disease or viral infection.

[0183] RNA insertion site: A site in an RNA that is targeted for or has received insertion of an exogenous polynucleotide or polyribonucleotide. The methods disclosed herein encompass the use of nucleic acid editors expressed from two or more nucleic acid molecules presented herein that can be used to target RNA for manipulation at the RNA insertion site.

[0184] Sequence identity: The similarity between amino acid (or nucleotide) sequences is expressed in terms of the similarity between the sequences, otherwise referred to as sequence identity. Sequence identity is frequently measured in terms of percentage identity (or similarity or homology); the higher the percentage, the more similar the two sequences are.

[0185] Methods for aligning sequences for comparison are known. Various programs and alignment algorithms are described in Smith and Waterman, Adv. Appl. Math. 2:482, 1981; Needleman and Wunsch, J. Mol. Biol. 48:443, 1970; Pearson and Lipman, Proc. Natl. Acad. Sci. USA 85:2444, 1988; Higgins and Sharp, Gene 73:237, 1988; Higgins and Sharp, CABIOS 5:151, 1989; Corpet et al., Nucleic Acids Research 16:10881, 1988; and Pearson and Lipman, Proc. Natl. Acad. Sci. USA 85:2444, 1988. Altschul et al., Nature Genet. 6:119, 1994, presents a detailed discussion of sequence alignment methods and homology calculations.

[0186] The NCBI Basic Local Alignment Search Tool (BLAST) (Altschul et al., J. Mol. Biol. 215: 403, 1990) for use in conjunction with the sequence analysis programs blastp, blastn, blastx, tblastn, and tblastx is available from several sources, including the National Center for Biotechnology Information (NCBI, Bethesda, MD) and the Internet. A description of how to determine sequence identity using this program is available on the Internet at the NCBI website.

[0187] Variants of a native protein or coding sequence (e.g., DMD, Factor VIII, Factor IX, or ABCA4 sequences) are generally characterized by having at least about 80%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity, as counted over the entire length of the alignment with the amino acid sequence, using NCBI Blast 2.0 with gap insertion blastp set to default parameters. For comparisons of amino acid sequences of more than about 30 amino acids, the Blast 2 alignment function is used with the default BLOSUM62 matrix set to default parameters (gap existence cost of 11, and per residue gap cost of 1). When aligning short peptides (fewer than approximately 30 amino acids), alignments should be performed using the Blast 2 alignment function with the PAM30 matrix set to default parameters (open gap penalty of 9, extension gap penalty of 1). Proteins with greater similarity to the reference sequence, when assessed by this method, will exhibit increasingly greater percentage identity, e.g., at least 95%, at least 98%, or at least 99% sequence identity. When comparing less than the entire sequence for sequence identity, homologs and variants generally have at least 80% sequence identity over a short window of 10-20 amino acids, and may have at least 85%, at least 90%, or at least 95% sequence identity depending on their similarity to the reference sequence. Methods for determining sequence identity over such short windows are available on the Internet at the NCBI website. These sequence identity ranges are provided merely as guidance; highly significant homologs may be obtained outside the provided ranges.

[0188] Variants of the nucleic acid sequences disclosed herein (e.g., synthetic intron sequences and coding sequences) are generally characterized as having at least about 80%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity, counted over the entire length of the alignment with the nucleic acid sequence, using NCBI Blast 2.0, gap-insertion blastn, set to default parameters. It will be understood by those skilled in the art that these sequence identity ranges are provided merely as guidance, and that functional sequences outside the provided ranges may be obtainable.

[0189] Subject: A mammal, e.g., a human. Mammals include, but are not limited to, mice, monkeys, humans, farm animals, sport animals, and pet animals. In one embodiment, the subject is a non-human mammalian subject, e.g., a monkey or other non-human primate, mouse, rat, rabbit, pig, goat, sheep, dolphin, dog, cat, horse, or cow. In some examples, the subject is a laboratory animal / organism such as a mouse, rabbit, or rat. In some examples, the subject treated using the methods disclosed herein is a human.

[0190] In some embodiments, the subject has a genetic disease, such as those listed in Tables 1-4, that can be treated using the methods disclosed herein. In some embodiments, the subject treated using the methods disclosed herein is a human subject with a genetic disease. In some embodiments, the subject treated using the methods disclosed herein is a human subject with cancer. In some embodiments, the subject treated using the methods disclosed herein is a human subject with an infection, such as a bacterial or viral infection.

[0191] Target nucleic acid: A nucleic acid molecule, such as a gene, having a sequence to be altered and / or its expression modulated, e.g., a DNA or RNA sequence. In some examples, the target nucleic acid molecule is one nucleic acid molecule, two or more portions of the same nucleic acid molecule (e.g., the same RNA or gene), two or more different nucleic acid molecules (e.g., two different genes or two different RNAs), or two or more portions of two or more different nucleic acid molecules. The target nucleic acid molecule may contain one or more target editing sites, which are regions of the target nucleic acid molecule to be altered, e.g., substituted, deleted, or inserted (e.g., one or more targeted nucleotides or ribonucleotides, e.g., at least 10, at least 15, at least 20, or at least 30 consecutive nucleotides or ribonucleotides). In some examples, the target nucleic acid molecule is one whose expression is modulated, e.g., expression of a gene product (e.g., a protein) is increased or decreased. In one example, the target nucleic acid molecule is a DNA, RNA, or gene whose expression activation is desired. In one example, the target nucleic acid molecule is DNA, RNA, or a gene whose expression is desired to be reduced or eliminated. In one example, the target nucleic acid molecule is DNA, RNA, or a gene with one or more point mutations that result in a disease (e.g., those listed in Tables 1-4). In some examples, the targeting sequence (e.g., of the gRNA) is complementary to the target gene / nucleic acid. In other examples, the targeting sequence (e.g., of the gRNA) is complementary to the promoter and / or regulatory elements of the target nucleic acid molecule. In some examples, the target nucleic acid sequence is DNA and is immediately adjacent to a protospacer adjacent motif (PAM). In some examples, the target nucleic acid sequence is unique compared to other nucleic acid sequences in a cell.

[0192] Targeting sequence: The portion of a gRNA that has complementarity to a target nucleic acid sequence. In some embodiments, the targeting sequence is complementary to a promoter or regulatory element of a target nucleic acid whose expression is desired to be activated or repressed. In some embodiments, the targeting sequence of a gRNA is about 14-30 nt and has sufficient complementarity to the target nucleic acid sequence to hybridize with the target sequence and direct sequence-specific binding of a Cas nuclease to the target nucleic acid sequence. In some embodiments, the degree of complementarity between the targeting sequence of a gRNA and its corresponding target nucleic acid, when optimally aligned using a suitable alignment algorithm, is about 50%, about 60%, about 75%, about 80%, about 85%, about 90%, about 95%, about 97.5%, about 98%, about 99%, or more, or greater than about 50%, greater than about 60%, greater than about 75%, greater than about 80%, greater than about 85%, greater than about 90%, greater than about 95%, greater than about 97.5%, greater than about 98%, greater than about 99%, or more. In some embodiments, the degree of complementarity is 100%. Optimal alignment can be determined using any suitable algorithm for aligning sequences, non-limiting examples of which include the Smith-Waterman algorithm, the Needleman-Wunsch algorithm, algorithms based on the Burrows-Wheeler Transform (e.g., Burrows Wheeler Aligner), ClustalW, Clustal X, BLAT, Novoalign (Novocraft Technologies, ELAND (Illumina, San Diego, Calif.), SOAP (available at soap.genomics.org.cn), and Maq (available at maq.sourceforge.net).

[0193] Therapeutic agent: refers to one or more molecules or compounds that produce some beneficial effect when administered to a subject. The synthetic nucleic acid molecules disclosed herein and the systems presented herein are therapeutic agents. Beneficial therapeutic effects can include enabling diagnostic determinations to be made; improving a disease, symptom, disorder, or pathological condition; reducing or preventing the onset of a disease, symptom, disorder, or condition; and generally negating a disease, symptom, disorder, or pathological condition.

[0194] Transcriptional activator: A protein or protein domain that increases the transcription of nucleic acid molecules such as genes. Such proteins can be used in the compositions, systems, and methods presented herein. Such proteins and protein domains can have a DNA binding domain and a domain for activating transcription. These activators can be introduced into the system by binding to Cas nuclease or gRNA. Examples of such activators include VP64, p65, myogenic differentiation 1 (MyoD1), heat shock transcription factor (HSF) 1, RTA, CBP, SET7 / 9, or any combination thereof (e.g., p65 and HSF1).

[0195] Transduced, transformed, and transfected: A cell has been "transduced" by a virus or vector when the virus or vector has transferred the nucleic acid molecule to the cell. A cell has been "transformed" or "transfected" by a nucleic acid introduced into the cell when the nucleic acid is stably replicated by the cell, either by integration of the nucleic acid into the cell's genome or by episomal replication.

[0196] These terms encompass all techniques by which nucleic acid molecules can be introduced into such cells, including transfection with viral vectors, transformation with plasmid vectors, and introduction of naked DNA by electroporation, lipofection, particle gun acceleration, and other methods known in the art. In some examples, the methods are chemical methods (e.g., calcium phosphate transfection), physical methods (e.g., electroporation, microinjection, particle bombardment), fusion (e.g., liposomes), receptor-mediated endocytosis (e.g., DNA-protein complexes, viral envelope / capsid-DNA complexes), and biological infection with viruses such as recombinant viruses (Wolff, JA, ed., Gene Therapeutics, Birkhauser, Boston, USA, 1994). Methods for introducing nucleic acid molecules into cells are known (see, e.g., U.S. Patent No. 6,110,743). These methods can be used to transduce cells with the nucleic acid molecules disclosed herein.

[0197] Transgene: An exogenous gene delivered by a vector, e.g., AAV. In one example, a transgene encodes a portion of a nucleic acid editing protein, e.g., about one-third, half, or two-thirds of a nucleic acid editing protein, e.g., operably linked to a promoter sequence. In one example, a transgene includes a portion of a Cas nuclease coding sequence, e.g., about one-third, half, or two-thirds of a Cas nuclease coding sequence, e.g., operably linked to a promoter sequence.

[0198] Treating, treatment, and therapy: Any objective or subjective parameter, such as relief, remission, or reduction of symptoms, or any success or indication of success regarding the attenuation or reversal of an injury, lesion, or condition, including making the condition more tolerable to the patient, slowing the rate of degeneration or decline, making the end point of degeneration less debilitating, improving the physical or mental well-being of the subject, or extending the length of survival. Treatment can be evaluated by objective or subjective parameters, including the results of physical exams, blood tests, and other clinical tests. In some examples, treatment using the methods disclosed herein results in a reduction in the number or severity of symptoms associated with the genetic disease, e.g., an increase in the survival time of the patient with the genetic disease being treated.

[0199] In some examples, treatment with the methods disclosed herein results in a reduction in the number or severity of symptoms associated with DMD or other genetic diseases, such as increased survival, increased mobility (e.g., walking, climbing), improved cognitive ability, reduced gastrocnemius muscle size, reduced cardiomyopathy, improved vision, improved hearing, improved blood clotting, or improved respiratory function. In some examples, a combination of these effects is achieved.

[0200] Tumor, neoplasia, malignancy, or cancer: A neoplasm is an abnormal growth of tissue or cells resulting from excessive cell division. Neoplastic growth can result in a tumor. The amount of tumor in an individual is the "tumor burden," which can be measured as the number, volume, or weight of tumors. Tumors that do not metastasize are called "benign." Tumors that have the potential to invade surrounding tissues and / or metastasize are called "malignant." "Non-cancerous tissue" is tissue that originates from the same organ in which a malignant neoplasm forms, but does not have the characteristic lesions of a neoplasm. Generally, non-cancerous tissue appears histologically normal. "Normal tissue" is tissue that originates from an organ, where the organ is not affected by cancer or another disease or disorder of that organ. A "cancer-free" subject has not been diagnosed with cancer of that organ and does not have detectable cancer.

[0201] Exemplary tumors, such as cancers, that can be treated using the methods and systems disclosed herein include solid tumors, such as breast cancer (e.g., lobular carcinoma and ductal carcinoma), sarcoma, lung cancer (e.g., non-small cell carcinoma, large cell carcinoma, squamous cell carcinoma, and adenocarcinoma), pulmonary mesothelioma, colorectal adenocarcinoma, gastric cancer, prostate adenocarcinoma, ovarian cancer (e.g., serous cystadenocarcinoma and mucinous cystadenocarcinoma), ovarian germ cell tumor, testicular cancer and germ cell tumor, pancreatic adenocarcinoma, and pancreatic adenocarcinoma. cancer, bile duct adenocarcinoma, hepatocellular carcinoma, bladder cancer (including, for example, transitional cell carcinoma, adenocarcinoma, and squamous cell carcinoma), renal cell adenocarcinoma, endometrial cancer (including, for example, adenocarcinoma and mixed Müllerian tumor (carcinosarcoma)), endocervical cancer, epicervical cancer, and vaginal cancer (e.g., adenocarcinoma and squamous cell carcinoma of the endocervix, epicervix, and vagina, respectively), tumors of the skin (e.g., squamous cell carcinoma, basal cell carcinoma, malignant melanoma, skin adnexal tumors These tumors include esophageal cancer, nasopharyngeal and oropharyngeal cancer (including squamous cell carcinoma and adenocarcinoma of the nasopharynx and oropharynx), salivary gland cancer, tumors of the brain and central nervous system (including, for example, tumors of glial origin, tumors of neuronal origin, and tumors of meningeal origin), tumors of the peripheral nerves, soft tissue sarcomas, and sarcomas of bone and cartilage, and lymphoid tumors (including B-cell and T-cell malignant lymphomas). In one embodiment, the tumor is an adenocarcinoma.

[0202] The methods and systems can also be used to treat liquid tumors, such as lymphocytic, leukemia, or other types of leukemia. In certain examples, the tumors treated are blood tumors, such as leukemias (e.g., acute lymphoblastic leukemia (ALL), chronic lymphocytic leukemia (CLL), acute myeloid leukemia (AML), chronic myeloid leukemia (CML), hairy cell leukemia (HCL), T-cell prolymphocytic leukemia (T-PLL), large granular lymphocytic leukemia, and adult T-cell leukemia), lymphomas (e.g., Hodgkin's lymphoma and non-Hodgkin's lymphoma), and myelomas.

[0203] Upregulated: When used in reference to expression of a molecule such as a target nucleic acid / protein, refers to any process that results in increased production of the target nucleic acid / protein. In some examples, upregulation or activation of a target RNA includes processes that may increase translation of the target RNA and therefore the presence of the corresponding protein. An upregulated molecule can be a target nucleic acid of the compositions and methods described herein or a protein expressed from that nucleic acid molecule, e.g., a nucleic acid editor protein produced from a recombined transcript in an REJ division system; an edited target nucleic acid and / or resulting protein produced from that nucleic acid; or a representative marker, surrogate, or functional indicator of the target nucleic acid or protein or edited target nucleic acid and / or resulting protein.

[0204] Upregulation includes any detectable increase in target nucleic acid / protein. In certain examples, detectable expression of a target nucleic acid / protein in a cell or cell-free system is increased by at least 20%, at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, at least 75%, at least 80%, at least 90%, at least 95%, at least 100%, at least 200%, at least 400%, or at least 500% compared to a control. The control can be the amount of target nucleic acid / protein detected in a corresponding sample not treated with a nucleic acid molecule provided herein. In one example, the control is the relative amount of expression in normal cells (e.g., non-recombinant cells not containing the system provided herein). The control can be compared to a target nucleic acid or protein expressed from a nucleic acid molecule of the compositions and methods described herein, such as a nucleic acid editor protein produced from a recombined transcript in an REJ partitioning system; an edited target nucleic acid and / or the resulting protein produced from that nucleic acid; or a representative marker, surrogate, or functional indicator of the target nucleic acid or protein or the edited target nucleic acid and / or the resulting protein. The control used for comparison can be any reasonable control determined by one of skill in the art. A positive control can be the amount of protein produced from the corresponding mRNA and / or full-length construct. A negative control can be the amount of protein produced from the corresponding mRNA and / or an empty or otherwise defective construct. The control can be compared to a target nucleic acid or a protein produced in the presence of a nucleic acid editor protein expressed in a cell from a nucleic acid molecule using the compositions and methods provided herein. As described herein, when a nucleic acid editing protein is expressed in a cell from a nucleic acid molecule using the present compositions and methods, the engineered nucleic acid editing protein can edit its target nucleic acid.The edited nucleic acid and / or protein produced from the nucleic acid can be compared to a positive control, which can be the amount of corresponding nucleic acid, e.g., mRNA, and / or protein produced from the target nucleic acid, in a normal cell (e.g., a non-recombinant "normal" cell that does not have the target nucleic acid in need of editing and does not contain the system provided herein). The edited nucleic acid and / or protein produced from the nucleic acid can be compared to a negative control, which can be the amount of corresponding nucleic acid, e.g., mRNA, and / or protein produced from the target nucleic acid, in an untreated cell (e.g., a non-recombinant "mutant" cell that has the target nucleic acid in need of editing and does not contain the system provided herein). The detectable target nucleic acid and / or protein expression compared to the detectable positive control target nucleic acid and / or protein expression (e.g., a cell or cell-free system) can be 20% to 500% of the expression in the positive control, respectively. The detectable target nucleic acid and / or protein expression in a cell or cell-free system compared to the expression in the positive control can be about 20% to about 500%.Detectable expression of target nucleic acid and / or protein in cells or cell-free systems compared to expression in a positive control may be about 20% to about 30%, about 20% to about 40%, about 20% to about 50%, about 20% to about 60%, about 20% to about 70%, about 20% to about 90%, about 20% to about 95%, about 20% to about 100%, about 20% to about 150%, about 20% to about 200%, about 20% to about 500%, about 30% to about 40%, about 30% to about 50%, about 30% to about 60%, about 30% to about 70%, about 30% to about 9 ...100%, about 20% to about 150%, about 20% to about 200%, about 20% to about 500%, about 30% to about 100%, about 20% to about 150%, about 2 ~approximately 60%, ~approximately 30% to ~approximately 70%, ~approximately 30% to ~approximately 90%, ~approximately 30% to ~approximately 95%, ~approximately 30% to ~approximately 100%, ~approximately 30% to ~approximately 150%, ~approximately 30% to ~approximately 200%, ~approximately 30% to ~approximately 500%, ~approximately 40% to ~approximately 50%, ~approximately 40% to ~approximately 60%, ~approximately 40% to ~approximately 70%, ~approximately 40% to ~approximately 90%, ~approximately 40% to ~approximately 95%, ~approximately 40% to ~approximately 100%, ~approximately 40% to ~approximately 150%, ~approximately 40% to ~approximately 200%, ~approximately 40% to ~approximately 500%, ~approximately 50% to ~approximately 60%, ~approximately 50% to ~approximately 7% 0%, approximately 50% to approximately 90%, approximately 50% to approximately 95%, approximately 50% to approximately 100%, approximately 50% to approximately 150%, approximately 50% to approximately 200%, approximately 50% to approximately 500%, approximately 60% to approximately 70%, approximately 60% to approximately 90%, approximately 60% to approximately 95%, approximately 60% to approximately 100%, approximately 60% to approximately 150%, approximately 60% to approximately 200%, approximately 60% to approximately 500%, approximately 70% to approximately 90%, approximately 70% to approximately 95%, approximately 70% to approximately 100%, approximately 70% to approximately 150%, approximately 70% to approximately 20 The concentration may be 0%, about 70% to about 500%, about 90% to about 95%, about 90% to about 100%, about 90% to about 150%, about 90% to about 200%, about 90% to about 500%, about 95% to about 100%, about 95% to about 150%, about 95% to about 200%, about 95% to about 500%, about 100% to about 150%, about 100% to about 200%, about 100% to about 500%, about 150% to about 200%, about 150% to about 500%, or about 200% to about 500%. Detectable target nucleic acid and / or protein expression in a cell or cell-free system compared to expression in a positive control can be about 20%, about 30%, about 40%, about 50%, about 60%, about 70%, about 90%, about 95%, about 100%, about 150%, about 200%, or about 500%.Expression of the detectable target nucleic acid and / or protein in a cell or cell-free system relative to expression in a positive control can be at least about 20%, 30%, 40%, 50%, 60%, 70%, 90%, 95%, 100%, 150%, or 200%. Expression of the detectable target nucleic acid and / or protein in a cell or cell-free system relative to expression in a positive control can be at most about 30%, 40%, 50%, 60%, 70%, 90%, 95%, 100%, 150%, 200%, or 500%.

[0205] The edited nucleic acid and / or protein produced from the nucleic acid can be compared to a negative control, which can be the amount of the corresponding nucleic acid, e.g., mRNA, and / or protein produced from the target nucleic acid, in an untreated cell (e.g., a non-recombinant "mutant" cell that has the target nucleic acid to be edited and does not contain the system provided herein). Expression of the detectable target nucleic acid and / or protein can be increased by 20% to 500%, respectively, compared to expression of the detectable negative control target nucleic acid and / or protein. The detectable expression of the target nucleic acid / protein in a cell or cell-free system is about 20% to about 30%, about 20% to about 40%, about 20% to about 50%, about 20% to about 60%, about 20% to about 70%, about 20% to about 90%, about 20% to about 95%, about 20% to about 100%, about 20% to about 150%, about 20% to about 200%, about 20% to about 500%, about 30% to about 40%, about 30% to about 50%, about 30% to about 60%, about 30% to about 70%, about 30% to about 80%, about 30% to about 90%, about 30% to about 100%, about 20% to about 150%, about 20% to about 200%, about 20% to about 500%, about 30% to about 40%, about 30% to about 50%, about 30% to about 60%, about 30% to about 100%, about 30% to about 150%, about 20% to about 200%, about 20% to about 500%, about 30% to about 150%, about 2 ...300%, about 30% to about 40%, about 30% to about 50%, about 30% to about 60%, about 30% to about 150%, about 20% to about 150%, about 20% to about 150%, about 20% to 0% to about 70%, about 30% to about 90%, about 30% to about 95%, about 30% to about 100%, about 30% to about 150%, about 30% to about 200%, about 30% to about 500%, about 40% to about 50%, about 40% to about 60%, about 40% to about 70%, about 40% to about 90%, about 40% to about 95%, about 40% to about 100%, about 40% to about 150%, about 40% to about 200%, about 40% to about 500%, about 50% to about 60%, about 50% to about 70%, about 50 % to about 90%, about 50% to about 95%, about 50% to about 100%, about 50% to about 150%, about 50% to about 200%, about 50% to about 500%, about 60% to about 70%, about 60% to about 90%, about 60% to about 95%, about 60% to about 100%, about 60% to about 150%, about 60% to about 200%, about 60% to about 500%, about 70% to about 90%, about 70% to about 95%, about 70% to about 100%, about 70% to about 150%, about 70% to about 200%, The increase may be about 70% to about 500%, about 90% to about 95%, about 90% to about 100%, about 90% to about 150%, about 90% to about 200%, about 90% to about 500%, about 95% to about 100%, about 95% to about 150%, about 95% to about 200%, about 95% to about 500%, about 100% to about 150%, about 100% to about 200%, about 100% to about 500%, about 150% to about 200%, about 150% to about 500%, or about 200% to about 500%.Expression of the detectable target nucleic acid / protein in a cell or cell-free system may be increased by about 20%, about 30%, about 40%, about 50%, about 60%, about 70%, about 90%, about 95%, about 100%, about 150%, about 200%, or about 500% compared to a negative control. Expression of the detectable target nucleic acid / protein in a cell or cell-free system may be increased by at least about 20%, about 30%, about 40%, about 50%, about 60%, about 70%, about 90%, about 95%, about 100%, about 150%, or about 200% compared to a negative control. Detectable target nucleic acid / protein expression in a cell or cell-free system may be increased by at most about 30%, about 40%, about 50%, about 60%, about 70%, about 90%, about 95%, about 100%, about 150%, about 200%, or about 500% compared to a negative control.

[0206] Under conditions sufficient for: A phrase used to describe any environment that allows for a desired activity. In one example, the desired activity is an increase in protein expression or activity required to treat a disease. In one example, the desired activity is a decrease in protein expression or activity required to treat a disease. In one example, the desired activity is expression of a corrected protein sequence required to treat a disease. In one example, the desired activity is treating or slowing the progression of a genetic disease, such as DMD (or other genetic diseases listed in Tables 1-4) in vivo, for example, using the methods and systems disclosed herein.

[0207] Vector: A nucleic acid molecule into which a foreign nucleic acid molecule can be introduced without interfering with the vector's ability to replicate in and / or integrate into a host cell. Vectors include, but are not limited to, single-stranded, double-stranded, or partially double-stranded nucleic acid molecules; nucleic acid molecules containing one or more free ends, nucleic acid molecules with no free ends (e.g., circular); nucleic acid molecules comprising DNA, RNA, or both; and other types of polynucleotides.

[0208] A vector may contain a nucleic acid sequence that allows it to replicate in a host cell, such as an origin of replication. A vector may also contain one or more selectable marker genes and other genetic elements. An integrating vector is capable of integrating itself into a host nucleic acid. An expression vector is a vector that contains the necessary regulatory sequences to allow the transcription and translation of the inserted gene(s).

[0209] One type of vector is a "plasmid," which refers to a circular double-stranded DNA loop into which additional DNA segments can be inserted, such as by standard molecular cloning techniques. Another type of vector is a viral vector, in which viral-derived DNA or RNA sequences are present within the vector for packaging into a virus (e.g., retrovirus, replication-deficient retrovirus, adenovirus, replication-deficient adenovirus, and adeno-associated virus). Viral vectors also contain polynucleotides carried by the virus for transfection into host cells. In some embodiments, the vector is a lentiviral vector (e.g., an integration-deficient lentiviral vector) or an adeno-associated viral (AAV) vector.

[0210] In some embodiments, the vector is an AAV, such as AAV serotype AAV9 or AAVrh.10. In some embodiments, the vector is capable of penetrating the blood-brain barrier, for example, after intravenous administration. Adeno-associated virus serotype rh.10 (AAV.rh10) vectors partially penetrate the blood-brain barrier, resulting in high-level and widespread transgene expression. II. Overview of Some Embodiments

[0211] One approach to cure patients suffering from genetic diseases is gene editing therapy (commonly referred to as gene therapy). In this approach, defective genes are replaced with intact versions of the genes, delivered, for example, by viral vectors, thereby achieving sustained expression for months to years. Adeno-associated viruses (AAVs) have been used in clinical gene replacement therapy, but their packaging capacity is limited (e.g., less than about 5 kb). Therefore, in order to achieve gene replacement of genes that exceed the size limit of about 5 kb, strategies to overcome this packaging limit are required. For example, some promoters alone, coding sequences alone, or a combination of promoters and coding sequences exceed the size limit of about 5 kb of AAV. Therefore, such proteins encoded by such promoters and coding sequences can be expressed using the system disclosed herein.

[0212] Nucleic acid editing, for example, can be used to edit target DNA or RNA molecules, for example, gene editing, to increase or decrease the expression of target, and correct the mutation in target.Examples of such methods include the CRISPR / Cas method of editing DNA (for example, using Cas9 or dCas9 DNA endonuclease), the CRISPR / Cas method of editing RNA (for example, using Cas13d or dCas13d RNA endonuclease), the zinc finger nuclease method of genome editing (for example, using zinc finger nuclease that comprises zinc finger DNA binding domain and DNA cutting domain), and the transcription activator-like effector nuclease (TALEN)-based method of genome editing (for example, using transcription activator-like effector nuclease (TALEN) protein).All of these methods rely on the use of nucleic acid editing protein, which is the nuclease that can insert, delete and / or cross target nucleic acid sequence (for example, target DNA or RNA sequence) in cell. Thus, nucleic acid editing proteins can result in the insertion, deletion, and / or substitution of one or more selected nucleotides or ribonucleotides in a target DNA or RNA sequence. However, as noted above, cargo limitations of vectors such as AAV can make it difficult to produce sufficient levels of nucleic acid editing proteins (and in some examples, the corresponding gRNA) to treat disease.

[0213] Splicing-mediated recombination of two RNA molecules using naturally occurring intron sequences in one or both of these RNA fragments is inefficient. First, these natural intron sequences are derived from naturally occurring introns and are composed of a mixture of all four RNA nucleotides. Such sequences tend to fold into structures that can interfere with trans-interactions by forming strong intramolecular base pairs rather than becoming available for intermolecular interactions. Second, because exon definition in higher eukaryotes is driven by exons rather than introns, these naturally occurring intron sequences have not evolved to strongly attract spliceosome components. These two limitations of previous strategies are addressed in the present invention by designing synthetic intron sequences not found in nature. These synthetic sequences contain elements that strongly attract and stimulate spliceosome recruitment while minimizing secondary structures (and in some instances, other structures, such as tertiary structures) that interfere with the assembly of the two RNA fragments.

[0214] The present inventors have developed novel nucleic acid-based elements that can be used to efficiently reconstitute the coding sequence of large genes from multiple sequential fragments. The compositions, systems, and methods presented herein enable the reconstitution of full-length RNA from multiple sequential fragments. Reconstitution of full-length RNA, e.g., messenger RNA transcripts, in turn results in the production of full-length "reconstituted" proteins. The methods and systems disclosed herein differ from previous methods. The highly efficient synthetic introns disclosed herein utilize optimal placement of RNA elements (or DNA encoding these elements) that efficiently drive the RNA splicing reaction between non-covalently linked RNAs (pre-mRNAs). The methods / systems represent a significant advance over previous attempts to utilize trans-splicing because they generate high levels of functional nucleic acid editing protein that more closely approach therapeutic levels of nucleic acid editing protein for editing target nucleic acid molecules to treat genetic diseases. This innovation is based on selecting a non-natural RNA domain that is essentially unable to form strong cis-binding interactions that interfere with trans interactions with a second RNA with a complementary strand (which also has inherently low cis-binding ability). The intermolecular interaction between the binding regions of partnered dimerization domains, such as the first and second dimerization domains, the second and third dimerization domains, the third and fourth dimerization domains, the fourth and fifth dimerization domains, or the fifth and sixth dimerization domains, may be stronger than the intramolecular interaction between the binding region on a single dimerization domain and other sequences within the same dimerization domain. For example, a single-stranded kissing loop structure within a first kissing loop dimerization domain may bind, hybridize, or associate more strongly with its complementary kissing loop structure within a second kissing loop dimerization domain than with other sequences within the same dimerization domain. It is understood that within the same dimerization domain, there may be regions that bind more strongly intramolecularly than intermolecularly, such as the stem of a stem-loop structure.The resulting dimerization domains contain stable, accessible single-stranded binding regions that selectively and efficiently bind to their intended targets, i.e., complementary binding regions on partner dimerization domains. This strategy allows for unprecedented reconstitution efficiency of transcripts synthesized from separate templates, thereby resulting in high levels of target protein or therapeutic nucleic acid production. Numerous examples of such optimized synthetic nucleic acid molecules and methods for their use are presented herein. These optimized dimerization domains and / or synthetic introns can contain non-natural sequences (e.g., sequences not found in human cells and / or other biological systems) used in combination with optimized motifs that facilitate RNA splicing (including splice donors, splice acceptors, splice enhancers, and splice branchpoint sequences). Synthetic nucleic acids can be non-natural nucleic acid sequences, such as sequences not found in human cells and / or other biological systems. By optimizing trans-dimerization of RNA strands for relevant RNA motifs that mediate efficient splicing, it is demonstrated for the first time herein that two or three different RNAs can be precisely and efficiently covalently linked in vivo and in vitro in the same cell to produce high levels of functional nucleic acid editing proteins. Unlike "hybrid" approaches, which result in inefficient combinations at the DNA level due to DNA recombination, ultimately followed by RNA splicing in cis, which removes the DNA recombination site from the mature transcript, the methods / systems disclosed herein combine two protein-coding RNA fragments at the pre-mRNA level, facilitating a more efficient reaction with a lower risk of producing a recombinant product that encodes a non-functional and / or harmful product.

[0215] The data demonstrate that the use of efficient synthetic RNA dimerization and recombination domains (sRdR domains, also referred to as RNA end-joining (REJ) domains) allows for efficient production of nucleic acid editing proteins by reconstituting their full-length mRNA from two separate gene fragments expressed from two separate nucleic acid constructs in the same cell. The desired guide RNA, e.g., gRNA, can be expressed from one of the two constructs or from a different construct. The methods and systems disclosed herein can be used to reconstitute transcripts encoding large genes, such as ABE8e, to edit any target nucleic acid to treat any genetic disease. Based on these findings, any genetic disease, such as those benefiting from the expression of a nucleic acid editing protein (see, for example, the disorders listed in Tables 1-4), can be treated. Other diseases that can be treated using nucleic acid editing proteins include cancer and infectious diseases (e.g., bacterial or viral infections). Other applications include research and biotechnology applications.

[0216] In some embodiments, the present disclosure provides methods for using the nucleic acid editing compositions and systems of the present disclosure to treat a subject in need thereof by modifying a target nucleic acid sequence in the subject's cells. Uses of nucleic acid editing are described in the literature, for example, U.S. Patent Application No. 2020 / 392473, "Novel CRISPR enzymes and systems," which is incorporated herein by reference in its entirety.

[0217] In some embodiments, the compositions, systems, and methods described herein are used to treat a subject with a disease or disorder by repairing a nucleic acid mutation that causes defective or reduced production of a protein or another gene product, such as RNA. In some embodiments, the defective protein or gene product has reduced activity, is non-functional, or is toxic. In some embodiments, the mutation is present in a coding or non-coding region of a gene. In some embodiments, the disease or disorder is caused by a mutation in the coding region of a gene that results in a non-functional or otherwise defective gene product, and repairing the mutation restores expression of a functional gene product. In some embodiments, the disease or disorder is caused by a mutation in a regulatory region of a gene, such as a promoter region or splicing element, and repairing the mutation restores normal or desirable levels of expression of the gene product by upregulation or downregulation. In some embodiments, a mutation can be introduced such that splicing is altered to skip an exon, thereby restoring normal or desirable levels of a functional gene product. In some embodiments, the level of the functional gene product is modulated compared to a control, such as an untreated control.

[0218] In some embodiments, the disease gene with a loss-of-function mutation is a disease gene listed in Table 1, which is reproduced from Table 2 of Chen and Altman, 2017, "Opportunities for developing therapies for rare genetic diseases: focus on gain of function and allostery," Orphanet Journal of Rare Diseases 12: 61, which is incorporated herein by reference. In some embodiments, the nucleic acid editing compositions, systems, and methods described herein are used to treat a disease listed in Table 1. Table 1. Genetic diseases caused by loss-of-function mutations [Table 1] The corresponding disease name, OMIM identifier, mutated gene, and known allosteric activators are listed. The allosteric activators were searched in the allosteric database. The remaining potential candidates can be found in Additional file 1: Table S8.

[0219] In some embodiments, the compositions and methods described herein are useful for treating a subject with a disease or disorder by introducing a nucleic acid mutation to disrupt an undesired genomic sequence and / or downregulate the production of an undesired gene product in the subject's cells. In some embodiments, the undesired genomic sequence and / or the production of an undesired gene product is downregulated by modifying a coding or non-coding region. In some embodiments, one or more regulatory sequences (e.g., promoters or enhancers) are modified. In some embodiments, sequences that modulate splicing or other aspects of RNA processing (e.g., polyadenylation, subcellular localization, or half-life) are modified to disrupt an undesired genomic sequence and / or downregulate the level of an undesired gene product. In some embodiments, the coding region of an undesired genomic sequence is modified to introduce a mutation, such as a deletion, missense mutation, frameshift mutation, or stop codon, thereby downregulating the activity of the undesired genomic sequence and / or downregulating the production of an undesired gene product. In some embodiments, the activity and / or gene product level is downregulated compared to a control, for example, an untreated control. In some embodiments, the undesirable genomic sequence for targeting using the nucleic acid editing compositions, systems, and methods described herein can include any described in the literature, for example, cancer genes or disease genes with gain-of-function mutations. In some embodiments, the cancer gene is any listed in Table 2. In some embodiments, the nucleic acid editing compositions, systems, and methods described herein are used to treat the cancers listed in Table 2. Table 2. Oncogenes [Table 2-1] [Table 2-2] [Table 2-3] [Table 2-4] [Table 2-5]

[0220] In some embodiments, the disease gene having a gain-of-function mutation is a neurodegenerative disease gene. In some embodiments, the neurodegenerative disease gene is any of those listed in Table 3, which is reproduced from Table 1 of Chen and Altman, 2017. In some embodiments, the nucleic acid editing compositions, systems, and methods described herein are used to treat the neurodegenerative diseases listed in Table 3. Table 3. Genetic diseases caused by gain-of-function mutations [Table 3] The corresponding disease name, OMIM identifier, and mutated gene are listed. Diseases are grouped by molecular mechanism as indicated in the Type column. The remaining potential candidates can be found in Additional file 1: Table S7.

[0221] To address some of the limitations of existing strategies for reconstructing fragmented genes from multiple AAVs, a system is provided herein for sequentially aligning and recombining two or more individual synthetic RNA molecules in a target cell. Each of the individual synthetic RNA molecules contains a dimerization domain and a synthetic intron sequence containing elements necessary for RNA splicing, thereby mediating efficient RNA recombination of the individual fragments when the dimerization domains bind to each other in the correct order. In one example, reconstitution of the coding sequence from the two fragments is achieved by adding a first synthetic intron (A) to the 3' end of the N-terminal coding fragment and a complementary second synthetic domain (A') to the 5' end of the C-terminal coding fragment. The two RNAs are recombined by the cell's endogenous RNA splicing machinery (i.e., the spliceosome machinery). The synthetic intron domain contains bifunctional elements: (1) a dimerization domain to mediate base pairing between the two halves to be recombined, and (2) a domain optimized to efficiently recruit the splicing machinery to mediate efficient reconstitution of the two RNA molecules. The synthetic intron domain may contain elements to prevent proteins from being encoded by unspliced ​​RNA. In some embodiments, the synthetic intron comprises a sequence having at least 50%, at least 60%, at least 70%, at least 75%, 80%, at least 85%, at least 90%, at least 95%, at least 98%, at least 99%, or 100% sequence identity to any of the synthetic introns set forth in SEQ ID NOs: 159, 160, 161, 162, 163, 164, 165, and 166 (see, e.g., Figures 10A-10Z). In some examples, the synthetic intron is an RNA molecule encoded by a sequence having at least 50%, at least 60%, at least 70%, at least 75%, 80%, at least 85%, at least 90%, at least 95%, at least 98%, at least 99%, or 100% sequence identity to any of the synthetic introns set forth in SEQ ID NOs: 159, 160, 161, 162, 163, 164, 165, and 166, but without the promoter sequence set forth. Those skilled in the art will appreciate that any of the molecules set forth in SEQ ID NOs: 159, 160, 161, 162, 163, 164, 165, and 166 can be modified such that the protein-encoding portion (e.g., 114 and 164 in Figure 6A) is replaced with another nucleic acid-editing protein-encoding sequence of interest (e.g., the YFP-encoding sequence of SEQ ID NO: 1, 2, 22, or 23 can be replaced with a Cas9 or Cas13d protein-encoding sequence). Thus, synthetic intron molecules having at least 80%, at least 85%, at least 90%, at least 95%, at least 98%, at least 99%, or 100% sequence identity to any of the synthetic intron portions set forth in SEQ ID NOs: 159, 160, 161, 162, 163, 164, 165, and 166 are also provided herein.Also provided are synthetic intron RNA molecules encoded by sequences having at least 50%, at least 60%, at least 70%, at least 75%, 80%, at least 85%, at least 90%, at least 95%, at least 98%, at least 99%, or 100% sequence identity to any of the synthetic introns set forth in SEQ ID NOs: 159, 160, 161, 162, 163, 164, 165, and 166, but without the promoter sequences set forth.

[0222] Exemplary dimerization domains were selected bioinformatically to minimize / optimize their internal secondary / tertiary structures. The dimerization domains tested contained long stretches of low-diversity nucleotide sequences to avoid intramolecular annealing. By avoiding intramolecular annealing, these dimerization domains exist in an open configuration, and are therefore available for pairing with corresponding complementary dimerization domain sequences. The synthetic intron domain contains an intron splice enhancer element, which leads to efficient recruitment of the splicing machinery.

[0223] The RNA molecules disclosed herein are designed to have at least an open, accessible single-stranded region that can bind to complementary dimerization domains to allow efficient splicing and recombination of the RNA. In some embodiments, this is achieved by using only purines or only pyrimidines in the binding domain. Due to the inability of purines (and pyrimidines) to pair with themselves, these stretches of RNA have an open, predicted structure.

[0224] RNA molecules exist as single strands in cells. Being single-stranded, RNA molecules inherently tend to hybridize with themselves, thereby forming strong secondary and tertiary structures. The most stable base pairs are G-C, A-U, and G-U wobble pairs. Thermodynamically, two-base pairing is favored over open configurations. To design efficient synthetic nucleic acid molecules, two complementary dimerization domains are present in an open configuration, thus making the dimerization domains available for intermolecular base pairing. To avoid intramolecular base pairing between other parts of the synthetic nucleic acid molecule, long stretches of non-variable sequences containing incompatible bases can be included. For example, long stretches of pyrimidines (i.e., C and T) or purines (i.e., A and G) can be present in the synthetic nucleic acid molecule. Pyrimidines cannot form standard base pairs with other pyrimidines, and purines cannot form standard base pairs with other purines. Such stretches of purines or pyrimidines can range from 2-3 bases to 200-300 bases. These stretches cannot bind intramolecularly and are therefore available for intermolecular base pairing with complementary fragments. For example, synthetic nucleic acid molecules A and A' can be constructed such that A contains a pyrimidine stretch (e.g., 5'-CCUU(...)CCUU-3') and A' contains a complementary purine sequence (e.g., 5'-AAGG(...)AAGG-3').

[0225] The synthetic nucleic acid molecule disclosed herein (for example, RNA or DNA encoding RNA) is designed to minimize any off-target binding to the incorrect site in genome.Off-target binding can be reduced by changing the sequence of nucleic acid molecule.

[0226] The same design principle of using stretches of low diversity RNA bases to achieve open synthetic nucleic acid architectures can be extended to the use of stretches of single bases in the dimerization domain, for example, a stretch of Gs base-pairing with a stretch of Cs, and a stretch of As base-pairing with a stretch of Us.

[0227] The following methods can be used to increase recombination of two or more synthetic nucleic acid molecules. RNA splicing relies on the recruitment of spliceosome components to the 5' end of the intron (splice donor site) and the 3' end of the intron (splice acceptor site, along with its associated branchpoint sequence and polypyrimidine tract). Different ribonucleoproteins are recruited to introns through base pairing of protein-associated small nuclear RNAs (snRNAs) with intron sequences. Placing perfectly matched consensus sequences within the RNA dimerization and recombination domain can facilitate the recruitment of spliceosome components, which in turn enhances the efficiency of spliceosome-mediated recombination. Previously characterized intron splice enhancer sequences can recruit additional splicing-enhancing factors, termed intron splice enhancers.

[0228] In some embodiments, instead of using naturally occurring RNA sequences for the RNA splicing sequence, consensus sequences are used. For example, consensus sequences can be used for any sequence involved in splicing, including splice donors, splice acceptors, splice enhancers, and splice branchpoint sequences. These synthetic nucleic acid molecules can be used to sequentially link two (or more) RNA molecules in cells ex vivo, in vitro, or in vivo. The synthetic nucleic acid molecule can include any promoter and coding sequence outside of the encoded synthetic intron domain. For example, two synthetic nucleic acid molecules can carry both halves of a single nucleic acid editing gene. This has been tested in vitro and in vivo by reconstituting both halves of yellow fluorescent protein (YFP) and shown to be efficient (see Figures 3A-3D).

[0229] The modular nature of synthetic nucleic acid molecules allowed us to test the efficiency of achieving serial recombination of multiple RNA fragments (i.e., >2) using a combinatorial set of optimized complementary dimerization domains (Figure 4A-4B). A tripartite yellow fluorescent protein was efficiently reconstituted and expressed at high levels in >80% of transfected cells.

[0230] These results demonstrate that a single RNA molecule can be reconstituted from at least three different synthetic nucleic acid molecules, such as when expressing nucleic acid-editing proteins with promoters and / or coding sequences that are too long to fit into a single gene therapy vector, e.g., AAV.

[0231] In some embodiments, the synthetic nucleic acid molecules, eg, synthetic DNA molecules, of the compositions, systems, kits, and methods of the invention are generated by transcription of an RNA viral genome by reverse transcriptase.

[0232] The systems disclosed herein enable efficient RNA recombination between individual fragments. In some examples, the recombination (i.e., splicing or recombination) efficiency achieved using the compositions, systems, or methods of the present disclosure is determined using any suitable method known to those of skill in the art. In some examples, the recombination efficiency is represented by a measure of correctly ligated RNA compared to a control RNA, or a measure of full-length nucleic acid editing protein or protein activity compared to a control protein. In some examples, the control RNA is unligated RNA, in which case the recombination efficiency is represented by a measure of bound RNA compared to unligated RNA. This measurement can be made by detecting and comparing the junction RNA and the unligated 3' RNA species 3' (e.g., junction RNA:3' RNA). In some examples where more than two RNAs are ligated, ligation at any or all of the junctions is assessed. In some examples, the recombination efficiency is represented by a measure of full-length or active nucleic acid editing protein compared to a protein fragment or inactive protein.

[0233] In some examples, the efficiency of rearrangement, recombination, or splicing (a measure of the correct joining of two or more different coding sequences present on different RNA molecules and / or the production of a desired full-length protein) is between about 10% and about 100%. Synthetic intron domains can include elements that prevent proteins from being encoded by unspliced ​​RNA. In some examples, the rearrangement efficiency is between about 10% and about 15%, about 10% and about 20%, about 10% and about 25%, about 10% and about 30%, about 10% and about 40%, about 10% and about 50%, about 10% and about 60%, about 10% and about 70%, about 10% and about 80%, about 10% and about 90%, about 10% and about 100%, about 15% and about 20%, about 15% and about 25%, about 15% and about 30%, about 15% and about 40%, about 15% and about 50%, approximately 15% to approximately 60%, approximately 15% to approximately 70%, approximately 15% to approximately 80%, approximately 15% to approximately 90%, approximately 15% to approximately 100%, approximately 20% to approximately 25%, approximately 20% to approximately 30%, approximately 20% to approximately 40%, approximately 20% to approximately 50%, approximately 20% to approximately 60%, approximately 20% to approximately 70%, approximately 20% to approximately 80%, approximately 20% to approximately 90%, approximately 20% to approximately 100%, approximately 25% to approximately 30%, approximately 25% to approximately 40%, approximately 25% to approximately 5 0%, approximately 25% to approximately 60%, approximately 25% to approximately 70%, approximately 25% to approximately 80%, approximately 25% to approximately 90%, approximately 25% to approximately 100%, approximately 30% to approximately 40%, approximately 30% to approximately 50%, approximately 30% to approximately 60%, approximately 30% to approximately 70%, approximately 30% to approximately 80%, approximately 30% to approximately 90%, approximately 30% to approximately 100%, approximately 40% to approximately 50%, approximately 40% to approximately 60%, approximately 40% to approximately 70%, approximately 40% to approximately 80%, approximately 40% to approximately 90 %, about 40% to about 100%, about 50% to about 60%, about 50% to about 70%, about 50% to about 80%, about 50% to about 90%, about 50% to about 100%, about 60% to about 70%, about 60% to about 80%, about 60% to about 90%, about 60% to about 100%, about 70% to about 80%, about 70% to about 90%, about 70% to about 100%, about 80% to about 90%, about 80% to about 100%, or about 90% to about 100%. In some embodiments, the reconstitution efficiency is about 10%, about 15%, about 20%, about 25%, about 30%, about 40%, about 50%, about 60%, about 70%, about 80%, about 90%, or about 100%.In some embodiments, the reconstitution efficiency is at least about 10%, about 15%, about 20%, about 25%, about 30%, about 40%, about 50%, about 60%, about 70%, about 80%, or about 90%. In some embodiments, the reconstitution efficiency is at most about 15%, about 20%, about 25%, about 30%, about 40%, about 50%, about 60%, about 70%, about 80%, about 90%, or about 100%.

[0234] In some examples, the efficiency of rearrangement, recombination, or splicing (in this example, a measure of the correct joining of two different nucleic acid editing protein coding sequences present on different RNA molecules and / or the production of a desired nucleic acid editing full-length protein, where the two different nucleic acid editing protein coding sequences encode transcripts of about 3200 nt to 9000 nt, e.g., about 4000 to 9000 nt, about 4400 to 9000 nt, about 3200 to 4000 nt, about 3200 to 3600 nt, e.g., about 4500 nt, about 4000 nt, about 3800 nt, about 3600 nt, or about 3200 nt) is about 10% to about 100%. In some embodiments, the reconstitution efficiency using the two-part system is between about 10% and about 15%, between about 10% and about 20%, between about 10% and about 25%, between about 10% and about 30%, between about 10% and about 40%, between about 10% and about 50%, between about 10% and about 60%, between about 10% and about 70%, between about 10% and about 80%, between about 10% and about 90%, between about 10% and about 100%, between about 15% and about 20%, between about 15% and about 25%, between about 15% and about 30%, between about 15 ... About 40%, about 15% to about 50%, about 15% to about 60%, about 15% to about 70%, about 15% to about 80%, about 15% to about 90%, about 15% to about 100%, about 20% to about 25%, about 20% to about 30%, about 20% to about 40%, about 20% to about 50%, about 20% to about 60%, about 20% to about 70%, about 20% to about 80%, about 20% to about 90%, about 20% to about 100%, about 25% to about 30%, about 25% to about 40%, Approximately 25% to approximately 50%, approximately 25% to approximately 60%, approximately 25% to approximately 70%, approximately 25% to approximately 80%, approximately 25% to approximately 90%, approximately 25% to approximately 100%, approximately 30% to approximately 40%, approximately 30% to approximately 50%, approximately 30% to approximately 60%, approximately 30% to approximately 70%, approximately 30% to approximately 80%, approximately 30% to approximately 90%, approximately 30% to approximately 100%, approximately 40% to approximately 50%, approximately 40% to approximately 60%, approximately 40% to approximately 70%, approximately 40% to approximately 80%, approximately 40% to about 90%, about 40% to about 100%, about 50% to about 60%, about 50% to about 70%, about 50% to about 80%, about 50% to about 90%, about 50% to about 100%, about 60% to about 70%, about 60% to about 80%, about 60% to about 90%, about 60% to about 100%, about 70% to about 80%, about 70% to about 90%, about 70% to about 100%, about 80% to about 90%, about 80% to about 100%, or about 90% to about 100%.In some embodiments, the reconstitution efficiency is about 10%, about 15%, about 20%, about 25%, about 30%, about 40%, about 50%, about 60%, about 70%, about 80%, about 90%, or about 100%. In some embodiments, the reconstitution efficiency is at least about 10%, about 15%, about 20%, about 25%, about 30%, about 40%, about 50%, about 60%, about 70%, about 80%, or about 90%. In some embodiments, the reconstitution efficiency is at most about 15%, about 20%, about 25%, about 30%, about 40%, about 50%, about 60%, about 70%, about 80%, about 90%, or about 100%.

[0235] In some examples, the efficiency of rearrangement, recombination, or splicing (in this example, a measure of the correct joining of two different nucleic acid editing protein coding sequences present on different RNA molecules and / or the production of a desired full-length nucleic acid editing protein, where the two different nucleic acid editing protein coding sequences encode transcripts of approximately 4000 nt) is between about 40% and about 60%, e.g., between about 40% and about 50%, between about 42% and about 47%, e.g., about 45%.

[0236] In some examples, the efficiency of rearrangement, recombination, or splicing (in this example, a measure of the correct joining of two different nucleic acid editing protein coding sequences present on different RNA molecules and / or the production of a desired nucleic acid editing full-length protein, where the two different nucleic acid editing protein coding sequences encode transcripts of approximately 3800 nt) is about 40% to about 60%, e.g., about 40% to about 50%, about 42% to about 47%, e.g., about 45%.

[0237] In some examples, the efficiency of rearrangement, recombination, or splicing (in this example, a measure of the correct joining of two different nucleic acid editing protein coding sequences present on different RNA molecules and / or the production of a desired nucleic acid editing full-length protein, where the two different nucleic acid editing protein coding sequences encode transcripts of approximately 3600 nt) is about 25% to about 50%, for example, about 30% to about 40%, for example, about 35%.

[0238] In some examples, the efficiency of rearrangement, recombination, or splicing (in this example, a measure of the correct joining of two different nucleic acid editing protein coding sequences present on different RNA molecules and / or the production of a desired nucleic acid editing full-length protein, where the two different nucleic acid editing protein coding sequences encode transcripts of approximately 3200 nt) is about 25% to about 50%, for example, about 30% to about 40%, for example, about 35%.

[0239] In some examples, the efficiency of rearrangement, recombination, or splicing (in this example, a measure of the correct joining of three different nucleic acid editing protein coding sequences present on different RNA molecules and / or the production of a desired full-length nucleic acid editing protein, where the three different nucleic acid editing protein coding sequences encode transcripts of about 3,200 nt to about 13,500 nt, e.g., about 4,000 nt to about 5,000 nt, about 4,000 nt to about 13,500 nt, about 6,000 nt to about 12,000 nt, about 6,000 nt to about 10,000 nt, or about 8,000 nt to about 12,000 nt, e.g., up to about 13,500 nt) is about 10% to about 100%. In some embodiments, the reconstitution efficiency using the three-part system is between about 10% and about 15%, between about 10% and about 20%, between about 10% and about 25%, between about 10% and about 30%, between about 10% and about 40%, between about 10% and about 50%, between about 10% and about 60%, between about 10% and about 70%, between about 10% and about 80%, between about 10% and about 90%, between about 10% and about 100%, between about 15% and about 20%, between about 15% and about 25%, between about 15% and about 30%, between about 15 ... About 40%, about 15% to about 50%, about 15% to about 60%, about 15% to about 70%, about 15% to about 80%, about 15% to about 90%, about 15% to about 100%, about 20% to about 25%, about 20% to about 30%, about 20% to about 40%, about 20% to about 50%, about 20% to about 60%, about 20% to about 70%, about 20% to about 80%, about 20% to about 90%, about 20% to about 100%, about 25% to about 30%, about 25% to about 40%, Approximately 25% to approximately 50%, approximately 25% to approximately 60%, approximately 25% to approximately 70%, approximately 25% to approximately 80%, approximately 25% to approximately 90%, approximately 25% to approximately 100%, approximately 30% to approximately 40%, approximately 30% to approximately 50%, approximately 30% to approximately 60%, approximately 30% to approximately 70%, approximately 30% to approximately 80%, approximately 30% to approximately 90%, approximately 30% to approximately 100%, approximately 40% to approximately 50%, approximately 40% to approximately 60%, approximately 40% to approximately 70%, approximately 40% to approximately 80%, approximately 40% to about 90%, about 40% to about 100%, about 50% to about 60%, about 50% to about 70%, about 50% to about 80%, about 50% to about 90%, about 50% to about 100%, about 60% to about 70%, about 60% to about 80%, about 60% to about 90%, about 60% to about 100%, about 70% to about 80%, about 70% to about 90%, about 70% to about 100%, about 80% to about 90%, about 80% to about 100%, or about 90% to about 100%.In some embodiments, the reconstitution efficiency is about 10%, about 15%, about 20%, about 25%, about 30%, about 40%, about 50%, about 60%, about 70%, about 80%, about 90%, or about 100%. In some embodiments, the reconstitution efficiency is at least about 10%, about 15%, about 20%, about 25%, about 30%, about 40%, about 50%, about 60%, about 70%, about 80%, or about 90%. In some embodiments, the reconstitution efficiency is at most about 15%, about 20%, about 25%, about 30%, about 40%, about 50%, about 60%, about 70%, about 80%, about 90%, or about 100%.

[0240] In some examples, the efficiency of rearrangement, recombination, or splicing (in this example, a measure of the correct joining of four different nucleic acid editing protein coding sequences present on different RNA molecules and / or the production of a desired full-length nucleic acid editing protein, where the four different nucleic acid editing protein coding sequences encode transcripts of about 3,200 nt to about 18,000 nt, e.g., about 4,000 nt to about 18,000 nt, about 4,000 nt to about 5,000 nt, about 10,000 nt to about 18,000 nt, about 15,000 nt to about 18,000 nt, or about 12,000 nt to about 15,000 nt, e.g., up to about 18,000 nt) is about 10% to about 100%.In some embodiments, the reconstitution efficiency using the four-part system is between about 10% and about 15%, between about 10% and about 20%, between about 10% and about 25%, between about 10% and about 30%, between about 10% and about 40%, between about 10% and about 50%, between about 10% and about 60%, between about 10% and about 70%, between about 10% and about 80%, between about 10% and about 90%, between about 10% and about 100%, between about 15% and about 20%, between about 15% and about 25%, between about 15% and about 30%, between about 15 ... About 40%, about 15% to about 50%, about 15% to about 60%, about 15% to about 70%, about 15% to about 80%, about 15% to about 90%, about 15% to about 100%, about 20% to about 25%, about 20% to about 30%, about 20% to about 40%, about 20% to about 50%, about 20% to about 60%, about 20% to about 70%, about 20% to about 80%, about 20% to about 90%, about 20% to about 100%, about 25% to about 30%, about 25% to about 40%, Approximately 25% to approximately 50%, approximately 25% to approximately 60%, approximately 25% to approximately 70%, approximately 25% to approximately 80%, approximately 25% to approximately 90%, approximately 25% to approximately 100%, approximately 30% to approximately 40%, approximately 30% to approximately 50%, approximately 30% to approximately 60%, approximately 30% to approximately 70%, approximately 30% to approximately 80%, approximately 30% to approximately 90%, approximately 30% to approximately 100%, approximately 40% to approximately 50%, approximately 40% to approximately 60%, approximately 40% to approximately 70%, approximately 40% to approximately 80%, approximately 40% to In some embodiments, the reconstitution efficiency is about 10%, 15%, 20%, 25%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, or 100%. In some embodiments, the reconstitution efficiency is at least about 10%, about 15%, about 20%, about 25%, about 30%, about 40%, about 50%, about 60%, about 70%, about 80%, or about 90%. In some embodiments, the reconstitution efficiency is at most about 15%, about 20%, about 25%, about 30%, about 40%, about 50%, about 60%, about 70%, about 80%, about 90%, or about 100%.In some embodiments, the compositions, systems, or methods of the present disclosure are evaluated by determining the level of RNA or protein production using any suitable method known to those skilled in the art. In some embodiments, the level of RNA production is represented by the measure of correctly bound RNA compared to a control RNA, or the measure of full-length protein compared to a control. In some embodiments, the control RNA is a corresponding mutant RNA or endogenous RNA. For example, the ratio of the amount of bound RNA produced in transfected cells to the amount of mutant or endogenous RNA is compared to the same ratio in untransfected cells to determine the level of correctly bound RNA production. In some embodiments, the ratio of the amount of correctly bound RNA, full-length protein, or protein activity to the amount of control RNA, or the amount or activity of control protein, is compared.

[0241] In some embodiments, the achieved RNA production level is 5% to 100%. In some embodiments, the achieved RNA production level is about 5% to about 100%. In some embodiments, the achieved RNA production level is about 5% to about 10%, about 5% to about 20%, about 5% to about 25%, about 5% to about 30%, about 5% to about 40%, about 5% to about 50%, about 5% to about 60%, about 5% to about 70%, about 5% to about 80%, about 5% to about 90%, about 5% to about 100%, about 10% to about 20%, about 10% to about 25%, about 10% to about 30%, about 10% to about 40%, about 10% to about 50%, or %, about 10% to about 60%, about 10% to about 70%, about 10% to about 80%, about 10% to about 90%, about 10% to about 100%, about 20% to about 25%, about 20% to about 30%, about 20% to about 40%, about 20% to about 50%, about 20% to about 60%, about 20% to about 70%, about 20% to about 80%, about 20% to about 90%, about 20% to about 100%, about 25% to about 30%, about 25% to about 40%, about 25% to about 50% , about 25% to about 60%, about 25% to about 70%, about 25% to about 80%, about 25% to about 90%, about 25% to about 100%, about 30% to about 40%, about 30% to about 50%, about 30% to about 60%, about 30% to about 70%, about 30% to about 80%, about 30% to about 90%, about 30% to about 100%, about 40% to about 50%, about 40% to about 60%, about 40% to about 70%, about 40% to about 80%, about 40% to about 90% , about 40% to about 100%, about 50% to about 60%, about 50% to about 70%, about 50% to about 80%, about 50% to about 90%, about 50% to about 100%, about 60% to about 70%, about 60% to about 80%, about 60% to about 90%, about 60% to about 100%, about 70% to about 80%, about 70% to about 90%, about 70% to about 100%, about 80% to about 90%, about 80% to about 100%, or about 90% to about 100%. In some examples, the achieved RNA production level is about 5%, about 10%, about 20%, about 25%, about 30%, about 40%, about 50%, about 60%, about 70%, about 80%, about 90%, or about 100%. In some examples, the level of RNA production achieved is at least about 5%, about 10%, about 20%, about 25%, about 30%, about 40%, about 50%, about 60%, about 70%, about 80%, or about 90%.In some examples, the level of RNA production achieved is at most about 10%, about 20%, about 25%, about 30%, about 40%, about 50%, about 60%, about 70%, about 80%, about 90%, or about 100%.

[0242] In some examples, the protein production level is represented by a measure of the amount of full-length nucleic acid editing protein or nucleic acid editing protein activity compared to the amount of full-length nucleic acid editing protein or nucleic acid editing protein activity of a control protein. In some examples, the control protein is a corresponding mutant protein or endogenous protein. For example, the ratio of the amount of full-length nucleic acid editing protein or protein activity produced in a transfected cell to the amount of the mutant protein or endogenous protein is compared to the same ratio in an untransfected cell. In some examples, the control protein is, for example, a full-length nucleic acid editing protein produced in a cell engineered to express a control full-length protein (where the cell is not transfected with a construct of the invention) or an untransfected cell from a normal subject expressing the control full-length protein, and the protein production level is determined by measuring the amount or activity of the nucleic acid editing protein in the transfected cell and comparing it to that of the control protein. In some examples, the control protein is a mutant form of the protein produced in a cell transfected with the construct or an untransfected cell, and the amount of full-length protein or protein activity is compared to that of the control protein to determine the protein production level. In some examples, the amount of full-length protein or protein activity is compared to that of an endogenous or housekeeping protein to determine the level of protein production.

[0243] In some embodiments, the protein production level achieved is about 1% to about 100%. In some embodiments, the protein production level achieved is about 10% to about 100%. In some embodiments, the protein production level achieved is about 10% to about 20%, about 10% to about 30%, about 10% to about 40%, about 10% to about 50%, about 10% to about 60%, about 10% to about 70%, about 10% to about 75%, about 10% to about 80%, about 10% to about 85%, about 10% to about 90%, about 10% to about 100%, about 20% to about 30%, about 20% to about 40%, about 20% to about 50%, about 20% to about 6 ... %, about 20% to about 70%, about 20% to about 75%, about 20% to about 80%, about 20% to about 85%, about 20% to about 90%, about 20% to about 100%, about 30% to about 40%, about 30% to about 50%, about 30% to about 60%, about 30% to about 70%, about 30% to about 75%, about 30% to about 80%, about 30% to about 85%, about 30% to about 90%, about 30% to about 100%, about 40% to about 50%, about 40% to about 60%, about 4 0% to about 70%, about 40% to about 75%, about 40% to about 80%, about 40% to about 85%, about 40% to about 90%, about 40% to about 100%, about 50% to about 60%, about 50% to about 70%, about 50% to about 75%, about 50% to about 80%, about 50% to about 85%, about 50% to about 90%, about 50% to about 100%, about 60% to about 70%, about 60% to about 75%, about 60% to about 80%, about 60% to about 85%, about 60% to about 90%, about 60% to about 100%, about 70% to about 75%, about 70% to about 80%, about 70% to about 85%, about 70% to about 90%, about 70% to about 100%, about 75% to about 80%, about 75% to about 85%, about 75% to about 90%, about 75% to about 100%, about 80% to about 85%, about 80% to about 90%, about 80% to about 100%, about 85% to about 90%, about 85% to about 100%, or about 90% to about 100%. In some embodiments, the protein production level achieved is about 10%, about 20%, about 30%, about 40%, about 50%, about 60%, about 70%, about 75%, about 80%, about 85%, about 90%, or about 100%. In some examples, the protein production level achieved is at least about 10%, about 20%, about 30%, about 40%, about 50%, about 60%, about 70%, about 75%, about 80%, about 85%, or about 90%.In some examples, the level of protein production achieved is at most about 20%, about 30%, about 40%, about 50%, about 60%, about 70%, about 75%, about 80%, about 85%, about 90%, or about 100%.

[0244] In some embodiments, the protein activity level achieved is about 50% to about 100%. In some embodiments, the protein activity level achieved is about 50% to about 100%. In some embodiments, the protein activity level achieved is about 50% to about 55%, about 50% to about 60%, about 50% to about 65%, about 50% to about 70%, about 50% to about 75%, about 50% to about 80%, about 50% to about 85%, about 50% to about 90%, about 50% to about 95%, about 50% to about 100%, about 55% to about 60%, about 55% to about 65%, about 55% to about 70%, about 55% to about 75%, about 55% to about 80%, about 55% to about 85%, about 55% to about 90%, about 55% to about 95%, about 55% to about 100%, about 60% to about 65%, about 60% to about 70%, about 60% to about 75%, about 60% to about 80%, about 60% to about 85%, about 60% to about 90%, about 60% to about 95%, about 60% to about 10 0%, approximately 65% ​​to approximately 70%, approximately 65% ​​to approximately 75%, approximately 65% ​​to approximately 80%, approximately 65% ​​to approximately 85%, approximately 65% ​​to approximately 90%, approximately 65% ​​to approximately 95%, approximately 65% ​​to approximately 100%, approximately 70% to approximately 75%, approximately 70% to approximately 80%, approximately 70% to approximately 85%, approximately 70% to approximately 90%, approximately 70% to approximately 95%, approximately 70% to approximately 100%, approximately 75% to approximately 80%, approximately 75 In some embodiments, the protein activity level achieved is about 50%, about 55%, about 60%, about 65%, about 70%, about 75%, about 80%, about 85%, about 90%, about 95%, about 95%, or about 100%. In some embodiments, the level of protein activity achieved is at least about 50%, about 55%, about 60%, about 65%, about 70%, about 75%, about 80%, about 85%, about 90%, or about 95%. In some embodiments, the level of protein activity achieved is at most about 55%, about 60%, about 65%, about 70%, about 75%, about 80%, about 85%, about 90%, about 95%, or about 100%.

[0245] In some examples, the amount of correctly linked RNA or full-length nucleic acid editing protein produced in the cell, e.g., in combination with expression of one or more gRNAs, is sufficient to ameliorate or cure a condition or disease in a subject, as understood by those of skill in the art for a particular condition or disease. In some examples, the amount of correctly linked RNA or full-length nucleic acid editing protein produced in the cell (e.g., in combination with expression of one or more gRNAs) is an effective amount. In some examples, this amount corresponds to about 50% to 100% of the amount of RNA or protein produced in a normal cell. In some examples, this amount corresponds to about 40% to about 100% of the amount of RNA or protein produced in a normal cell.In some embodiments, this amount is about 40% to about 45%, about 40% to about 50%, about 40% to about 55%, about 40% to about 60%, about 40% to about 65%, about 40% to about 70%, about 40% to about 75%, about 40% to about 80%, about 40% to about 85%, about 40% to about 90%, about 40% to about 100%, about 45% to about 50%, about 45% to about 55%, about 45 ... 0%, approximately 45% to approximately 65%, approximately 45% to approximately 70%, approximately 45% to approximately 75%, approximately 45% to approximately 80%, approximately 45% to approximately 85%, approximately 45% to approximately 90%, approximately 45% to approximately 100%, approximately 50% to approximately 55%, approximately 50% to approximately 60%, approximately 50% to approximately 65%, approximately 50% to approximately 70%, approximately 50% to approximately 75%, approximately 50% to approximately 80%, approximately 50% to approximately 85%, approximately 50% to approximately 90%, approximately 50% to approximately 100%, approximately 55% to approximately 60%, approximately 55% to About 65%, about 55% to about 70%, about 55% to about 75%, about 55% to about 80%, about 55% to about 85%, about 55% to about 90%, about 55% to about 100%, about 60% to about 65%, about 60% to about 70%, about 60% to about 75%, about 60% to about 80%, about 60% to about 85%, about 60% to about 90%, about 60% to about 100%, about 65% to about 70%, about 65% to about 75%, about 65% to about 80%, about 65% to about 85%, about 65 % to about 90%, about 65% to about 100%, about 70% to about 75%, about 70% to about 80%, about 70% to about 85%, about 70% to about 90%, about 70% to about 100%, about 75% to about 80%, about 75% to about 85%, about 75% to about 90%, about 75% to about 100%, about 80% to about 85%, about 80% to about 90%, about 80% to about 100%, about 85% to about 90%, about 85% to about 100%, or about 90% to about 100%. In some examples, this amount corresponds to about 40%, about 45%, about 50%, about 55%, about 60%, about 65%, about 70%, about 75%, about 80%, about 85%, about 90%, or about 100% of the amount of RNA or protein produced in a normal cell, hi some examples, this amount corresponds to at least about 40%, about 45%, about 50%, about 55%, about 60%, about 65%, about 70%, about 75%, about 80%, about 85%, or about 90% of the amount of RNA or protein produced in a normal cell.In some examples, this amount corresponds to at most about 45%, about 50%, about 55%, about 60%, about 65%, about 70%, about 75%, about 80%, about 85%, about 90%, or about 100% of the amount of RNA or protein produced in a normal cell.

[0246] The RNA or protein measurements used to determine recombination efficiency or production level can be performed by any suitable method known to those skilled in the art. In some examples, recombination efficiency or production level is determined by measuring the amount of expressed functional protein, for example, by Western blotting. In some examples, recombination efficiency or production level is determined by measuring RNA transcripts, for example, using quantitative real-time PCR based on two probes. For example, the first assay spans the sequence completely contained in the 3' exon coding sequence (labeled with the 3' probe). The second assay spans the junction between the 5' exon coding sequence and the 3' exon coding sequence (labeled with the junction probe). Recombination efficiency can be calculated as the ratio of (junction probe count) / (3' probe count). The terms "recombination efficiency," "recombination efficiency," and "splicing efficiency" are used interchangeably herein.

[0247] The expression level of the rearranged protein or the successful editing of the nucleic acid sequence (and, for example, the production of the encoded protein) achieved using the compositions and methods presented herein can be assessed based on indirect measurements, such as the level of a representative marker, surrogate, or functional indicator. Any such marker known to those skilled in the art can be used to assess the level of the rearranged or edited protein according to the present disclosure and compared with a control based thereon. For example, the formation of the dystrophin-glycoprotein complex (DGC) or its subcomplexes can indicate the restoration of dystrophin function. See, for example, Gao and McNally, 2015, "The Dystrophin Complex: Structure, Function, and Implications for Therapy," Compr Physiol. 2015 5(3): 1223-39, and Omairi, et al., 2019, "Regulation of the dystrophinassociated glycoprotein complex composition by the metabolic properties of muscle fibers," Scientific Reports (2019) 9: 2770, each of which is incorporated herein by reference in its entirety.

[0248] In some embodiments, the dimerization domain is about 20 to about 1000 nt, or about 50 to about 160 nt, or about 50 to about 500 nt, or about 50 to 1000 nt, wherein the reconstitution efficiency is such that an effective amount of correctly linked RNA or full-length nucleic acid editing protein is produced. In some embodiments, the dimerization domain is about 50 to about 160 nt, wherein the reconstitution efficiency is such that an effective amount of correctly linked RNA or full-length nucleic acid editing protein is produced.

[0249] Efficient recombination between multiple RNA molecules allows for the packaging and delivery of transgenes into AAVs, exceeding the packaging limitations of a single AAV. AAV packaging limitations are a major obstacle to gene therapy approaches for diseases caused by the absence or defect of large genes. One application of this system is to express a nucleic acid editing protein and one or more gRNAs specific to a target gene using a viral vector with limited packaging capacity. Diseases and genes include, but are not limited to, (Disease (Gene, OMIM Gene Identifier)): 1) Duchenne muscular dystrophy and Becker muscular dystrophy (dystrophin, OMIM:300377); 2) dysferlinopathy (dysferlin, OMIM:603009); 3) cystic fibrosis (CFTR, OMIM:602421); 4) Usher syndrome 1B (myosin VIIA, OMIM:276903); 5) stromal fibrosis (stromal fibrosis, OMIM:602421); 6) stromal fibrosis (stromal fibrosis, OMIM:602422); 7) stromal fibrosis (stromal fibrosis, OMIM:602421); 8) stromal fibrosis (stromal fibrosis, OMIM:602422); 9) stromal fibrosis (stromal fibrosis, OMIM:602422); 10) stromal fibrosis (stromal fibrosis, OMIM:602422); 11) stromal fibrosis (stromal fibrosis, OMIM:602422); 12) stromal fibrosis (stromal fibrosis, OMIM:602422); 13) stromal fibrosis (stromal fibrosis, OMIM:602422); 14) stromal fibrosis (stromal fibrosis, These include Talgardt disease 1 (ABCA4, OMIM:601691); 6) hemophilia A (coagulation factor VIII, OMIM:300841); 7) von Willebrand disease (von Willebrand factor, OMIM:613160); 8) Marfan syndrome (fibrillin 1, OMIM:134797); 9) von Recklinghausen disease (neurofibromatosis-1, OMIM:162200), and hearing loss (OTOF, OMIM:603681).In some embodiments, the target nucleic acid is dystrophin; dysferlin; myosin VIIA; fibrillin 1; neurofibromatosis-1; the beta-globin chain of hemoglobin; clotting factor I; clotting factor II; clotting factor III; clotting factor IV; clotting factor V; clotting factor VI; clotting factor VII; clotting factor VIII; clotting factor IX; clotting factor X; clotting factor XI; clotting factor XII; clotting factor XIII; HBA1; HBA2; HBB; HBD; von Willebrand factor; MTHFR; FANCA; F ANCC; FANCD2; FANCG; FANCJ; ADAMTS13; Factor V Leiden prothrombin; IL-2RG, JAK3, IL-2 receptor gamma chain; IL-4 receptor gamma chain; IL-7 receptor gamma chain; IL-9 receptor gamma chain; IL-15 receptor gamma chain; IL-21 receptor gamma chain; RAG1; RAG2; CXCR4; IL7 receptor; ADA; PNP; WAS; CYBA, CYBB, NCF1, NCF2, NCF4; beta2 integrin; CC chemokine receptor Type 5 (CCR5), MSRB1;CSCR4;P17;PSIP1;CCR5;DMD;G6Pase;CEP290;ABCA4;MAGT1;Arylsulfatase A (ARSA);ABCD1;IDS;IDUA;IDUA; SGSH;NAGLU;HGSNAT;GNS;GALNS;GLB1;ARSB;GUSB;HYAL1;MAN2B1;SMPD1;NPC1;NPC2;CFTR;PKD-1;PDK-2;PDK-3;HEXA;GBA;HTT These genes are present in genes encoding proteins selected from (if wild-type) NF-1; NF2; APOB; LDLR; LDLRAP1; PCSK9; BCR-ABL; ASXL1; RUNX2; EPHA1; PD-1; androgen receptor; E6; E7; CD; NGF; ARSA; MBP; WASP; AADC; CLN2; ASPA; GAN; MT-ND4; SGSH; SUMF1; GAD; NTRN; TH; CH1; GDNF; GAA; SMN; and thymidine kinase. Others are presented in Tables 1-4.For example, to treat genomic point mutations or activate or overexpress genes, nucleic acid editing proteins and, for example, expression of one or more target nucleic acid-specific gRNAs can be expressed using the disclosed systems presented herein. Delivery of nucleic acid editing proteins can be achieved by splitting the protein into multiple fragments using the techniques presented herein.

[0250] Additional applications of the methods and systems disclosed herein include intersectional gene delivery for targeted gene expression. The differential infection / expression patterns of two viruses encoding fragmented genes can be used. The reconstituted protein is expressed in overlapping cell populations, representing a crossover where both viruses express themselves. Examples of such applications include: (1) delivering both halves (or three thirds, or other portions) of a protein from two (or more) projection targets using a retrogradely transported viral vector to label bifurcated dual-projection neurons; (2) delivering one fragment under the control of a promoter active in population A and a second fragment from a promoter active in population B to specifically tag / engineer the A∪B population; (3) delivering the first half of a protein using a viral vector with tropism for population A and the second half using a viral vector with tropism for population B to specifically tag / engineer the A∪B population; or combinations of these approaches.

[0251] In one example, the dimerization domain is, for example, (a) a small molecule trigger recognized by the aptamer, or (b) an aptamer sequence to facilitate dimerization in the presence of a protein present in the cell that binds to both halves and thus stimulates dimerization.

[0252] In some embodiments, the RNA-RNA interaction required for end-joining can be positively or negatively regulated by other nucleotides, such as (a) an antisense oligonucleotide sequence with homology to both halves (ssDNA-induced dimerization). In such an example, an antisense oligonucleotide with complementary sequences to both halves bridges the two molecules together, thus facilitating spliceosome-mediated recombination of the two molecules. (b) an antisense oligonucleotide sequence with homology to one of the two linked RNAs can prevent RNA dimerization of the two molecules and act as an off switch for gene expression, or (c) an endogenous cellular RNA with homology to both halves (RNA-induced dimerization). In such an example, a cellular RNA (e.g., mRNA or retroelement) with complementary sequences to both halves bridges the two molecules together, thus facilitating spliceosome-mediated recombination of the two molecules.

[0253] These molecular, protein, or RNA-mediated interactions allow for controllable / fine-tuned gene expression levels: by titrating molecules that interact with the binding domain (e.g., antisense oligonucleotides, small molecules, endogenous cellular RNA), the efficiency of dimer formation between the two halves can be modulated, regulating expression levels independently of promoter activity. Such installments can be used when a narrow range of protein expression levels is required.

[0254] III. Series This paper presents a system that can be used to recombine two or more RNA molecules, for example, at least two, at least three, at least four, or at least five different RNA molecules (for example, two, three, four, five, six, seven, eight, nine, or ten different RNA molecules) using a synthetic intron containing a dimerization sequence.Unlike the fragmentation and recombination of two fragments at the protein level, the method disclosed herein does not require extensive protein manipulation to find the appropriate division point.Recombination at the RNA level allows the seamless joining of two fragments of the protein. The methods and systems disclosed herein allow large genes (and corresponding proteins), e.g., those greater than about 4.5 kb, at least greater than 5 kb, at least greater than 5.5 kb, at least greater than 6 kb, at least greater than 1 kb, at least greater than 8 kb, at least greater than 8 kb, at least greater than 10 kb, at least greater than 13.5 kb, or at least greater than 18 kb, to be divided into two or more fragments or portions, each of which can be introduced into a cell or subject by a separate vector, e.g., multiple AAVs. In some examples, the methods and systems disclosed herein allow a nucleic acid editing gene (and corresponding protein, e.g., a Cas nuclease), e.g., one at least greater than about 2 kb, at least 2.5 kb, at least 3 kb, at least 3.5 kb, or at least 4 kb, to be divided into two or more fragments or portions, each of which can be introduced into a cell or subject by a separate vector, e.g., multiple AAVs, where such vectors may further comprise one or more gRNA coding sequences specific for one or more target nucleic acid molecules (e.g., specific for one or more target sites to be edited, e.g., by deletion, insertion, or substitution of one or more nucleotides or ribonucleotides). In some examples, multiple copies of the same gRNA coding sequence or sequences are present, e.g., to increase the number of gRNA molecules expressed from the vector.In one example, the system includes two portions for recombining two RNA molecules, where, for example, the target protein is encoded by at least about 4,500 nt to about 9,000 nt, e.g., 4,000 nt to 5,000 nt. In one example, the system includes three portions for recombining three RNA molecules, where, for example, the nucleic acid editing protein is encoded by a maximum of about 13,500 nt, e.g., about 2,000 nt to about 13,500 nt or 3,000 nt to 5,000 nt. In one example, the system includes four portions for recombining four RNA molecules, where, for example, the nucleic acid editing protein is encoded by a maximum of about 18,000 nt, e.g., about 2,000 nt to about 18,000 nt or 2,000 nt to 5,000 nt. This helps overcome limitations on the space available in the vector. In some examples, the length of the endogenous promoter limits the capacity of its corresponding gene to be expressed in the AAV. In some embodiments, the length of a coding sequence limits its capacity to be expressed in an AAV. In some embodiments, the length of an endogenous promoter and the length of its coding sequence limit their capacity to be expressed together in an AAV. The system disclosed herein can be used to express such long sequences that were previously difficult to express in an AAV. The system disclosed herein can also be used to express multiple copies of one or more gRNAs, for example, in combination with a nucleic acid editing protein, since the amount of gRNA can be rate-limiting.In some examples, the DNAs and systems disclosed herein express at least two gRNAs, at least three gRNAs, at least four gRNAs, at least five gRNAs, at least 10 gRNAs, at least 20 gRNAs, at least 50 gRNAs, at least 75 gRNAs, at least 100 gRNAs, at least 200 gRNAs, at least 500 gRNAs, or at least 1000 gRNAs, where the gRNAs can target the same nucleic acid molecule and the same target site, can target the same nucleic acid molecule and two or more target sites within the same nucleic acid molecule, can target two or more different nucleic acid molecules (e.g., one or more target sites within each different target nucleic acid molecule), or a combination thereof.

[0255] In some embodiments, one or more gRNAs target nucleic acid molecules such as genes associated with diseases such as single gene diseases, recessive gene diseases, and diseases caused by gene mutations. Examples of such diseases include, but are not limited to, hemophilia A (caused by mutations in the 7 kb coding region of the F8 gene, also known as coagulation factor VIII), hemophilia B (caused by mutations in the F9 gene), Duchenne muscular dystrophy (caused by mutations in the 11 kb coding region of the dystrophin gene), sickle cell anemia (caused by mutations in the beta-globin domain of hemoglobin, which has a promoter of approximately 3.5 kb), Stargardt disease (caused by mutations in the 6.9 kb coding region of the ABCA4 gene), and Usher syndrome (caused by mutations in the 7 kb coding region of MYO7A, resulting in hearing loss and visual impairment). In one embodiment, the gene is caused by a point mutation in the gene.

[0256] In one example, the one or more gRNAs target a nucleic acid molecule, such as a gene, to treat a disease such as cancer, e.g., breast cancer, lung cancer, prostate cancer, liver cancer, kidney cancer, brain cancer, bone cancer, ovarian cancer, uterine cancer, skin cancer, or colon cancer.

[0257] In some examples, RNA sequences encoding nucleic acid editing proteins and used in the methods and systems disclosed herein are codon-optimized for expression in target organisms or cells, e.g., human, dog, pig, cat, mouse, or rat cells. Thus, in some examples, the RNA coding sequence contains preferred codons (e.g., does not contain rare codons with low utilization). Codon optimization can be performed by identifying abundant tRNA levels in the target organism or cell. In some examples, the protein-encoding RNA sequence is de-enriched for cryptic splice donor and acceptor sites to maximize the RNA recombination reaction.

[0258] In some examples, the nucleic acid editing protein is divided into two parts, for example, approximately two equal halves (or other proportions, e.g., part A expressing about 1 / 3 and part B expressing about 2 / 3, or part A expressing about 1 / 4 and part B expressing about 3 / 4, etc.). However, each part need not have the same number of nucleotides (or encode the same number of amino acids). In such examples, the method can use two synthetic nucleic acid molecules (e.g., RNA or DNA encoding such RNA), one of which contains the coding sequence for the N-terminal part of the protein and the other of which contains the coding sequence for the C-terminal part of the protein. On this basis, those skilled in the art will understand that in addition to dividing the protein into two fragments or parts, the nucleic acid editing protein can be divided or split into more than two fragments, for example, three fragments. The design principles of the intron sequences of the three RNA molecules are similar to those of the two RNA molecules, but instead utilize different pairs of dimerization domains for one of the two binding sites. Thus, for example, an N-terminal protein coding sequence may be followed by an intron sequence having a specific binding domain (e.g., a first dimerization sequence), and the intermediate coding sequence may include an intron sequence having a complementary sequence to the first dimerization sequence (a second dimerization sequence). The intermediate coding fragment may be followed by another intron fragment having another dimerization sequence (a third dimerization sequence, different from the second dimerization sequence). The third fragment may include a C-terminal coding sequence for the protein and may also include an intron region having a dimerization sequence complementary to the third dimerization sequence (a fourth dimerization sequence). When more than one intermediate portion is used, the two intermediate portions may be referred to, for example, as an intermediate portion and a first intermediate portion, or a first intermediate portion and a second intermediate portion, or a first intermediate portion, a second intermediate portion, and a third intermediate portion, so that it is understood that each portion is distinct.

[0259] In one example, a nucleic acid editing protein can be split into an N-terminal portion and a C-terminal portion (e.g., roughly half, or unequal portions, e.g., 1 / 3 and 2 / 3, or 1 / 4 and 3 / 4), which can be reconstituted using the systems and methods disclosed herein. Referring to Figures 6A and 6G, in such an example, the system includes at least two synthetic nucleic acid molecules 110, 150. Each nucleic acid molecule 110, 150 can be composed of DNA or RNA (in the case of RNA, there is no corresponding promoter 112, 152). In some embodiments, molecules 110, 150 are each approximately at least 100 nucleotides / ribonucleotides (nt) in length, e.g., at least 200 nt, at least 300 nt, at least 500 nt, at least 1000 nt, at least 2000 nt, at least 3000 nt, at least 4000 nt, at least 5000 nt, at least 6000 nt, at least 7000 nt, at least 8000 nt, or at least 10,000 nt, e.g., 200-10,000 nt, 200-8000 nt, 500-5000 nt, 800-3000 nt, 1000-300 nt, or 200-1000 nt. Molecules 110, 150 can comprise natural and / or unnatural nucleotides or ribonucleotides.

[0260] Molecule 110 is a molecule located 5' of the system because it includes splice donor 116. In embodiments in which molecule 110 is DNA, molecule 110 includes promoter 112 operably linked to a sequence encoding an RNA molecule, which includes, from 5' to 3': a coding sequence for the N-terminal portion of the target protein 114, including a splice junction at the 3' end of the target protein coding sequence, SD 116, optional DISE 118, optional ISE 120, dimerization domain 122, and optional polyadenylation sequence 124. Any promoter 112 (or enhancer) can be used, such as one that utilizes RNA polymerase II, e.g., a constitutive promoter or an inducible promoter. In some examples, promoter 112 is a tissue-specific promoter, such as one that is constitutively active in muscle tissue (e.g., skeletal muscle tissue or cardiac muscle tissue), eye tissue (e.g., retina tissue), inner ear tissue, liver tissue, pancreatic tissue, lung tissue, skin tissue, bone, or kidney tissue. In some embodiments, promoter 112 is a cell-specific promoter, such as one that is constitutively active in cancer cells or normal cells. In some embodiments, promoter 112 is the endogenous promoter of the protein to be expressed, and in some embodiments, is long (e.g., at least 2500 nt, at least 3000 nt, at least 4000 nt, at least 5000 nt, or at least 7500 nt). In some embodiments, the promoter 112 is at least about 50 nucleotides (nt) in length, e.g., at least 100, at least 200, at least 300, at least 500, at least 1000, at least 2000, at least 3000, at least 4000, at least 5000, at least 6000, at least 7000, at least 8000 nt, at least 9000 nt, or at least 10,000 nt, e.g., 50-10,000 nt, 100-5000 nt, 500-5000 nt, or 50-1000 nt.In some embodiments, molecule 110 is DNA and is at least 200 nt, at least 300 nt, at least 500 nt, at least 800 nt, at least 1000 nt, at least 2000 nt, at least 3000 nt, at least 4000 nt, at least 5000 nt, at least 6000 nt, at least 7000 nt, or at least 8000 nt in length, e.g., 200-10,000 nt, 200-8000 nt, 500-5000 nt, 800-3000 nt, 1000-300 nt, or 200-1000 nt. In embodiments where molecule 110 is RNA, as shown in Figure 6F, for example, after transcription of DNA into RNA, molecule 110 does not include promoter 112, and 114 is RNA encoded by a coding sequence for the N-terminal portion of a nucleic acid editing protein. In some examples, molecule 110 is RNA, does not include a promoter 112, and is at least 200 nt, at least 300 nt, at least 500 nt, at least 1000 nt, at least 2000 nt, at least 3000 nt, at least 4000 nt, at least 5000 nt, at least 6000 nt, at least 7000 nt, or at least 8000 nt, e.g., 200-10,000 nt, 200-8000 nt, 500-5000 nt, 800-3000 nt, 1000-300 nt, or 200-1000 nt in length. Molecule 110 (with or without promoter 112) can include natural and / or non-natural nucleotides or ribonucleotides.

[0261] The splice junction around the 3' end of the N-terminal coding sequence (or RNA sequence encoded thereby) 114 can match a consensus sequence found in the target cell or organism into which the molecule 110, 150 is introduced. In humans, the splice junction sequence is AG (adenine-guanine) or UG (uracil-guanine) at positions -1 and -2 of the 5' splice site for U2-dependent introns, or AG, UG, CU (cytosine-uracil), or UU for U12-dependent introns. Thus, in some examples, the splice junction is 2 nt in length, and the 3' end of the N-terminal coding portion 114 is AG, UG, CU, or UU. In some examples, a DNA molecule encoding a portion of a nucleic acid editing protein includes sequences encoding portions of multiple splice junctions, for example, at the 3' end of a DNA molecule encoding the N-terminal portion of a nucleic acid editing protein and at the 5' end of a DNA molecule encoding the C-terminal portion of a nucleic acid editing protein.

[0262] The remaining 3'-terminal portion of molecule 110 is an intron 130. In some embodiments, intron sequence 130 is approximately at least 10 nt in length, e.g., at least 20 nt, at least 50 nt, at least 100 nt, at least 250 nt, at least 250 nt, at least 300 nt, at least 400 nt, or at least 500 nt, e.g., 20-500 nt, 20-250 nt, 20-100 nt, 50-100 nt, or 50-200 nt in length. Immediately following the N-terminal coding sequence (or the RNA encoded thereby) 114 is a splice donor (SD) 116 (e.g., an SD consensus sequence, such as the SD human consensus sequence). Thus, the SD 116 of intron sequence 130 is 3' to the N-terminal coding sequence 114. The SD 116 forms a recognition sequence for spliceosome components to bind to the RNA molecule. The sequence of SD116 can be an SD consensus sequence found in a target cell or organism into which the molecule 110, 150 is introduced. In some examples, SD116 is at least 2 nt, e.g., at least 5 nt, or at least 10 nt in length, e.g., 2-10 nt, 2-8 nt, 2-5, or 5-10 nt. SD116 can be used to recruit U2- or U12-dependent splicing machinery. In one example, U2-dependent splicing is used in human cells, and the SD116 sequence includes or is GUAAGUAUU. In one example, U12-dependent splicing is used in human cells, and the SD116 sequence includes or is AUAUCCUUUUUA (SEQ ID NO: 137) or GUAUCCUUUUUA (SEQ ID NO: 138). It is understood throughout that RNA sequences can be written using the nucleotides A, G, U, and C, and DNA sequences can be written using the nucleotides A, G, T, and C.It is also understood that sequences described herein as comprised by DNA molecules that are transcribed to form one or more RNA molecules, and sequences described as comprised by RNA molecules that may or may not be translated (i.e., protein-coding sequences) (e.g., gRNA), are sequences that have an intended function that can be recognized by one of skill in the art. For example, a gRNA sequence present in a DNA molecule of the present disclosure has a sequence that becomes functional after its transcription. As another example, a target protein-coding sequence present in a DNA molecule of the present disclosure has a sequence that can be translated into a target protein from an mRNA transcribed from the DNA. A sequence that is referred to as being encoded by or comprised by a DNA or RNA molecule is one that results in an intended functional product, as will be understood by one of skill in the art upon reading this disclosure.

[0263] The intron sequence 130 optionally includes one or both of a set of splicing enhancer sequences, referred to as downstream intron splice enhancers (DISEs) 118 and intron splice enhancers (ISEs) 120, which stimulate the action (e.g., increase activity) of the spliceosome. In some examples, the intron sequence 130 includes at least two splicing enhancer sequences, e.g., at least three, at least four, or at least five splicing enhancer sequences. Exemplary splicing enhancer sequences include DISEs 118 and ISEs 120. In some examples, including one or more splicing enhancer sequences 118, 120 in the intron sequence 130 increases splicing efficiency by at least 20%, at least 30%, at least 40%, at least 50%, at least 75%, at least 80%, at least 90%, or at least 95%. Exemplary splicing enhancer sequences that can be used are SEQ ID NOs: 26-136, 151, and 152, as well as GGGTTT, GGTGGT, TTTGGG, GAGGGG, GGTATT, GTAACG, GGGGGTAGG, GGAGGGTTT, GGGTGGTGT TTCAT, CCATTT, TTTTAAA, TGCAT, TGCATG, TGTGTT, CTAAC, TCTCT, TCTGT, TCTTT, TGCATG, CTAAC, CTGCT, TAACC, AGCTT, TTCATTA, GTTAG, TTTTGC, ACTAAT, ATGTTT, CTCTG, GGG, GGG(N)2-4GGG, TGGG, YCAY, UGCAUG, or 3x(G 3~6 N 1~7In some examples, DISE118, if present, may be at least 3 nt, at least 4 nt, at least 5 nt, at least 10 nt, at least 25 nt, at least 50 nt, at least 75 nt, or at least 100 nt in length, e.g., 3-10 nt, 3-11 nt, 4-11 nt, 5-11 nt, 10-50 nt, 5-100 nt, 10-25 nt, 10-20 nt, or 20-75 nt, and the sequence of DISE118 is or includes CUCUUUCUUUTCCAUGGGUUGGCU (SEQ ID NO: 134), TGCATG, CTAAC, CTGCT, TAACC, AGCTT, TTCATTA, GTTAG, TTTTGC, ACTAAT, ATGTTT, or CTCTG. In some embodiments, ISE120, if present, can be approximately at least 3 nt, at least 4 nt, at least 5 nt, at least 10 nt, e.g., at least 20 nt, at least 25 nt, at least 30 nt, at least 40 nt, or at least 50 nt in length, e.g., 3-10 nt, 3-11 nt, 4-11 nt, 5-11 nt, 10-50 nt, 20-25 nt, 10-25 nt, 10-20 nt, or 20-40 nt in length. In one embodiment, the sequence of ISE120 is or includes GGCUGAGGGAAGGACUGUCCUGGG (SEQ ID NO: 135), GGGUUAUGGGACC (SEQ ID NO: 136), TTCAT, CCATTT, TTTTAAA, TGCAT, TGCATG, TGTGTT, CTAAC, TCTCT, TCTGT, or TCTTT. In some embodiments, intron sequence 130 includes at least two, at least three, or at least four ISE120s.In some embodiments, ISE120 is or comprises at least one sequence, e.g., at least two, at least three, e.g., one, two, three, four, or five, having at least 80%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity to SEQ ID NO: 173, 174, 175, 176, 177, 178, 179, 180, 182, 183, 184, 185, 186, 187, 188, 189, 190, 191, 192, 193, 194, 195, 196, 199, 200, 201, 202, or 203. In some examples, DISE118 is or comprises at least one sequence, e.g., at least two, at least three such sequences, e.g., one, two, three, four or five such sequences, having at least 80%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99% or 100% sequence identity to SEQ ID NO: 173, 174, 175, 176, 177, 178, 179, 180, 182, 183, 184, 185, 186, 187, 188, 189, 190, 191, 192, 193, 194, 195, 196, 199, 200, 201, 202, or 203.

[0264] 3' to SD 116 (and enhancer sequences 118, 120, if present) is a dimerization domain 122 used to bring together the N-terminal coding sequence (or RNA encoded thereby) 114 and the C-terminal coding sequence 154. The intron sequence 130 portion of molecule 110 can optionally include a polyadenylation site 124 at its 3' end to terminate transcription of the fragment. In some examples, polyadenylation sequence 124 is a polyA sequence of at least 15 A's, e.g., 15-30 or 15-20 A's.

[0265] In some embodiments, the first dimerization domain 122 (and the second dimerization domain 154 of molecule 150) includes multiple unpaired nucleotides (i.e., unpaired within the structure of molecule 110 itself). Having unpaired nucleotides within the dimerization domains allows the 5' (or first) dimerization domain 122 and the 3' (or second) dimerization domain 154 to interact through base pairing. Through this interaction, molecules 110 and 150 are held in close proximity, thereby prompting the spliceosome to recombine the two molecules by joining the N-terminal coding region (or the RNA encoded thereby) 114 and the C-terminal coding region (or the RNA encoded thereby) 164.

[0266] In one example, the dimerization domains 122 (and 154) comprise "low-diversity sequences" containing limited nucleotide diversity, and therefore are unlikely to form stem-loops themselves in the secondary structure of each molecule 110, 150. Such low-diversity dimerization domains 122 (and 154) can be in a relatively open configuration, regardless of the sequences of the DNA (or RNA encoded thereby) 114, 164 encoding the N- and C-termini of the protein. This allows the nucleotides of the first dimerization domain 122 to be available for base pairing with the second dimerization domain 154 of the corresponding molecule 150, thereby enabling subsequent binding of the N-terminal coding sequence (or RNA encoded thereby) 114 and the C-terminal coding sequence (or RNA encoded thereby) 164. In some examples, the first dimerization domain 122 and the second dimerization domain 154 comprise low-diversity sequences interspersed with sequences that can form stems, resulting in local RNA loops that are open and available for base pairing in the absence of pseudoknot formation (FIG. 6B). Exemplary low-diversity sequences include a repeated run of U (e.g., 30-500 U), a repeated run of A (e.g., 30-500 A), a repeated run of G (e.g., 30-500 G), a repeated run of C (e.g., 30-500 C), a mixture containing only A and G (e.g., 30-500 A and G, e.g., AAAGAAGGAA(...) (SEQ ID NO: 149), or a mixture containing only C and U (e.g., 30-500 C and U, e.g., CUUUCUUUUCUU(...) (SEQ ID NO: 150)). Other exemplary low diversity sequences include complementary sequences that form a helix flanked by low diversity sequences.

[0267] In some examples, the first dimerization domain 122 and the second dimerization domain 154 contain only purines or only pyrimidines. In one example, the first dimerization domain 122 contains only purines, while the second dimerization domain 154 contains only pyrimidines. In another example, the first dimerization domain 122 contains only pyrimidines, while the second dimerization domain 154 contains only purines. Because purines cannot pair with themselves (nor can pyrimidines), these stretches of RNA have a predicted open structure.

[0268] In some examples, the first dimerization domain and the second dimerization domain 122, 154 do not contain cryptic splice acceptors that may compete with RNA recombination, such as sequences similar to the splice donor consensus sequence NNNAGGUNNNN (SEQ ID NO: 151) or NNNUGGUNNNN (SEQ ID NO: 152), where N refers to any nucleotide. In some examples, the first dimerization domain 122 is 1000 nt or less, e.g., 750 nt or less, or more than 500 nt, e.g., 6 to 1000 nt, 10 to 1000 nt, 20 to 1000 nt, 30 to 1000 nt, 30 to 750 nt, 30 to 500 nt, 50 to 500 nt, 50 to 100 nt, or 100 to 250 nt. In some embodiments, the first dimerization domain 122 is greater than 50 nt, e.g., at least 51 nt, at least 100 nt, at least 150 nt, at least 161 nt, or at least 170 nt, e.g., 51-159 nt, 51-150 nt, 51-120 nt, 51-100 nt, or 51-70 nt. In some embodiments, the first dimerization domain 122 is greater than 160 nt, e.g., at least 161 nt, at least 170 nt, at least 180 nt, at least 200 nt, at least 300 nt, at least 400 nt, at least 500 nt, at least 600 nt, at least 700 nt, at least 800 nt, at least 900 nt, or at least 1000 nt, e.g., 161-100 nt, 161-500 nt, 161-300 nt, 161-200 nt, or 161-170 nt. In some examples, the first dimerization domain 122 is less than 50 nt, for example, 6 to 49 nt, 6 to 45 nt, 6 to 40 nt, 6 to 30 nt, 6 to 20 nt, or 6 to 10 nt.

[0269] In some embodiments, the dimerization domain is 20 to 160 nt, 50 to 500 nt, or 500 to 1000 nt. In some embodiments, the dimerization domain is about 20 nt to about 160 nt. In some embodiments, the dimerization domain is about 20 nt to about 40 nt, about 20 nt to about 50 nt, about 20 nt to about 70 nt, about 20 nt to about 90 nt, about 20 nt to about 100 nt, about 20 nt to about 110 nt, about 20 nt to about 120 nt, about 20 nt to about 130 nt, about 20 nt to about 140 nt, about 20 nt to about 150 nt, about 20 nt to about 160 nt, about 40 nt to about 50 nt, about 40 nt to about 70 nt, about 40 nt to about 90 nt, about 40 nt to about 100 nt, about 40 nt to about 110 nt, about 4 ... 0nt to about 120nt, about 40nt to about 130nt, about 40nt to about 140nt, about 40nt to about 150nt, about 40nt to about 160nt, about 50nt to about 70nt, about 50nt to about 90nt, about 50nt to about 100nt, about 50nt to about 110nt , about 50nt to about 120nt, about 50nt to about 130nt, about 50nt to about 140nt, about 50nt to about 150nt, about 50nt to about 160nt, about 70nt to about 90nt, about 70nt to about 100nt, about 70nt to about 110nt, about 70nt to about 1 20nt, about 70nt to about 130nt, about 70nt to about 140nt, about 70nt to about 150nt, about 70nt to about 160nt, about 90nt to about 100nt, about 90nt to about 110nt, about 90nt to about 120nt, about 90nt to about 130nt, about 90 nt ~ about 140nt, about 90nt to about 150nt, about 90nt to about 160nt, about 100nt to about 110nt, about 100nt to about 120nt, about 100nt to about 130nt, about 100nt to about 140nt, about 100nt to about 150nt, about 100nt about 160nt, about 110nt to about 120nt, about 110nt to about 130nt, about 110nt to about 140nt, about 110nt to about 150nt, about 110nt to about 160nt, about 120nt to about 130nt, about 120nt to about 140nt, about 120nt to about 150nt, about 120nt to about 160nt, about 130nt to about 140nt, about 130nt to about 150nt, about 130nt to about 160nt, about 140nt to about 150nt, about 140nt to about 160nt, or about 150nt to about 160nt.In some embodiments, the dimerization domain is about 20 nt, about 40 nt, about 50 nt, about 70 nt, about 90 nt, about 100 nt, about 110 nt, about 120 nt, about 130 nt, about 140 nt, about 150 nt, or about 160 nt. In some embodiments, the dimerization domain is at least about 20 nt, about 40 nt, about 50 nt, about 70 nt, about 90 nt, about 100 nt, about 110 nt, about 120 nt, about 130 nt, about 140 nt, or about 150 nt. In some embodiments, the dimerization domain is at most about 40 nt, about 50 nt, about 70 nt, about 90 nt, about 100 nt, about 110 nt, about 120 nt, about 130 nt, about 140 nt, about 150 nt, or about 160 nt.

[0270] In some embodiments, the dimerization domain is about 50 nt to about 500 nt. In some embodiments, the dimerization domain is about 50 nt to about 100 nt, about 50 nt to about 150 nt, about 50 nt to about 200 nt, about 50 nt to about 250 nt, about 50 nt to about 300 nt, about 50 nt to about 350 nt, about 50 nt to about 400 nt, about 50 nt to about 500 nt, about 100 nt to about 150 nt, about 100 nt to about 200 nt, about 100 nt to about 250 nt, about 100 nt to about 300 nt, about 100 nt to about 350 nt, about 100 nt to about 400 nt, about 100 nt to about 500 nt, about 150 nt to about 200 nt, about 150 nt to about 250 nt, about 150 nt to about 300 nt, about 150nt to about 350nt, about 150nt to about 400nt, about 150nt to about 500nt, about 200nt to about 250nt, about 200nt to about 300nt, about 200nt to about 350nt, about 200nt to about 400nt, about 200nt to about 500nt, about 250nt to about 300nt, about 250nt to about 350nt, about 250nt to about 400nt, about 250nt to about 500nt, about 300nt to about 350nt, about 300nt to about 400nt, about 300nt to about 500nt, about 350nt to about 400nt, about 350nt to about 500nt, or about 400nt to about 500nt. In some embodiments, the dimerization domain is about 50 nt, about 100 nt, about 150 nt, about 200 nt, about 250 nt, about 300 nt, about 350 nt, about 400 nt, or about 500 nt. In some embodiments, the dimerization domain is at least about 50 nt, about 100 nt, about 150 nt, about 200 nt, about 250 nt, about 300 nt, about 350 nt, or about 400 nt. In some embodiments, the dimerization domain is at most about 100 nt, about 150 nt, about 200 nt, about 250 nt, about 300 nt, about 350 nt, about 400 nt, or about 500 nt.

[0271] In some examples, the sequences of the first and second dimerization domains 122 and 154 are determined by in silico structure prediction screening (e.g., using RNA folding structure prediction to screen a library of potential dimerization domain sequences; selecting sequences with a high proportion of unpaired nucleotides in both the dimerization domain and the corresponding anti-dimerization domain), low diversity nucleotide design (e.g., designing the dimerization domain to contain a stretch of low diversity sequence, such as a repeat sequence of only U, only A, only C, only G, only R (G and A), or only Y (U and C), a sequence that cannot fold with itself), or empirical screening (e.g., synthesizing a library of dimerization domains and corresponding anti-dimerization domains and screening for maximum recombination efficiency).

[0272] In some embodiments, the sequences of the first and second dimerization domains 122, 154 are designed to contain complementary RNA hairpin structures (also referred to as stem loops) that can form strong kissing loop interactions with their counterparts. In some embodiments, kissing loops are used when three or more dimerization domains, e.g., four or more or five or more dimerization domains, e.g., three, four, five, six, seven, eight, nine, or ten dimerization domains, are used to connect three or more portions of a coding sequence (e.g., Figure 6E). Each hairpin loop (or stem loop) of a kissing loop is composed of at least two complementary sequences (e.g., forming a stem) separated by a region of non-complementary sequence (e.g., forming a loop). In some examples, the dimerization domain may be comprised of one or more (e.g., at least two, at least three, at least four, or at least five, e.g., 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20) loops. In some examples using multiple loops, all or some of the loops may be repeated. In some examples using multiple loops, all or some of the loops may be different. In some examples, each complementary sequence is about 4-100 nt and is separated by a loop of about 3-20 nt. Base pairing between the two complementary sequences results in a helix (or stem) of, for example, at least 4 bp, at least 5 bp, at least 10 bp, at least 20 bp, at least 30 bp, at least 40 bp, at least 50 bp, at least 75 bp, at least 90 bp, or at least 100 bp, e.g., 4 to 100 bp, 5 to 75 bp, or 10 to 50 bp. In some embodiments, the loop portion is at least 3 nt, at least 5 nt, at least 10 nt, at least 15 nt, or at least 20 nt, e.g., 3 to 20 nt, 5 to 15 nt, or 5 to 10 nt, wherein the loop is not base paired.The complementary sequence between the two hairpin loops results in base pairing and the creation of a kissing loop / kissing stem-loop interaction. In some embodiments, the complementary sequence between the two hairpin loops is between at least 3 nucleotides of one loop and at least 3 nucleotides of the second loop, e.g., at least 4 nt, at least 5 nt, at least 6 nt, at least 7 nt, at least 8 nt, at least 9 nt, at least 10 nt, at least 11 nt, at least 12 nt, at least 13 nt, at least 14 nt, at least 15 nt, at least 16 nt, at least 17 nt, at least 19 nt, or at least 20 nt (e.g., 3 nt, 4 nt, 5 nt, 6 nt, 7 nt, 8 nt, 9 nt, 10 nt, 11 nt, 12 nt, 13 nt, 14 nt, The amino acid sequence may be between at least one of the first and second loops (e.g., 15nt, 16nt, 17nt, 18nt, 19nt, or 20nt) and at least 4 nt, at least 5 nt, at least 6 nt, at least 7 nt, at least 8 nt, at least 9 nt, at least 10 nt, at least 11 nt, at least 12 nt, at least 13 nt, at least 14 nt, at least 15 nt, at least 16 nt, at least 17 nt, at least 19 nt, or at least 20 nt (e.g., 3 nt, 4 nt, 5 nt, 6 nt, 7 nt, 8 nt, 9 nt, 10 nt, 11 nt, 12 nt, 13 nt, 14 nt, 15 nt, 16 nt, 17 nt, 18 nt, 19 nt, or 20 nt) of the amino acid sequence. In some embodiments, the complementary sequence between two hairpin loops is present in at least 15%, at least 20%, at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, or 100% of the total loop sequence.

[0273] In some cases, the stems of the kissing loops are selected to base pair in trans between the two RNA molecules. In such examples, after the formation of a kissing loop interaction between one hairpin loop on one molecule and another hairpin loop on the second molecule, the stem (or helix) region of each of the initial hairpin loops can base pair in trans between the two RNA molecules through strand displacement / invasion and extension of duplex formation. In some examples, up to about 85% of the nucleotides within the initial loop sequence can remain unpaired after extension of duplex formation (e.g., about 15% of the nucleotides are paired between the two loops). In some examples, the kissing loop is based on the HIV-1 DIS loop (SEQ ID NOs: 139 and 140, FIG. 17A ) and includes two A nucleotides on the 5′ side of a 6-nucleotide complementary sequence followed by one A nucleotide on the 3′ side (e.g., AANNNNNNA, where N can be A, U, G, or C). In some examples, the kissing loop is based on the HIV-2 kissing loop dimerization domain (SEQ ID NOs: 141 and 142, FIG. 17B) and includes a G nucleotide and an A nucleotide on the 5' side of a six-nucleotide complementary sequence, followed by three A nucleotides on the 3' side (e.g., GANNNNNNAAA (SEQ ID NO: 153), where N can be A, U, G, or C).

[0274] In one configuration, extended duplex formation is favored by including mismatches in the initial stem, resulting in a high percentage of matches in the extended duplex. Thus, in some embodiments, the helix or stem region of a hairpin loop may contain up to 30% of initially unpaired base pairs (e.g., 30% or less, 20% or less, 15% or less, 10% or less, 5% or less, or 1% or less of the base pairs, e.g., 1-30%, 5-30%, 10-30%, or 25-30% of the base pairs are initially unpaired). These unpaired regions may form bulges, mismatches, or internal loops.

[0275] In addition to the interaction of two hairpin loops (kissing loop interactions), other forms of loop interactions can be utilized for the first and second dimerization domains 122, 154. In one example, the loop is a bulge, and one strand of the base-paired helix contains one or more nucleotides that protrude from the stem structure. Exemplary bulges are at least 1 nt, at least 2 nt, at least 3 nt, at least 4 nt, at least 5 nt, at least 10 nt, or at least 20 nt, e.g., 1-20 nt, 1-15 nt, 1-10 nt, or 5-10 nt, or 1 nt, 2 nt, 3 nt, 4 nt, 5 nt, 6 nt, 7 nt, 8 nt, 9 nt, 10 nt, 11 nt, 12 nt, 13 nt, 14 nt, 15 nt, 16 nt, 17 nt, 18 nt, 19 nt, or 20 nt. In one example, the loop is an internal loop, eg, one or more nucleotides within a helix are mismatched, resulting in a helix interrupted by an internal loop at the position of the mismatch. In some examples, the helix is ​​at least 4 nt on each strand (e.g., at least 5 nt, at least 10 nt, at least 20 nt, at least 30 nt, at least 40 nt, at least 50 nt, at least 75 nt, at least 90 nt, or at least 100 nt, e.g., between 4 and 100 nt, between 5 and 75 nt, or between 10 and 50 nt, e.g., between 4 and 100 nt) and at least 1 nt on each side of the internal loop (e.g., at least 2 nt, at least 3 nt, at least 4 nt, at least 5 nt, at least 10 nt, or at least 20 nt on each strand, e.g., between 1 and 20 nt, between 1 and 15 nt, between 1 and 10 nt, or between 5 and 10 nt, or 1 nt, 2 nt, 3 nt, 4 nt, 5 nt, 6 nt, 7 nt, 8 nt, 9 nt, 10 nt, 11 nt, 12 nt, 13 nt, 14 nt, 15 nt, 16 nt, 17 nt, 18 nt, 19 nt, or 20 nt). In one embodiment, the loop is a multi-branched loop, from which three helices or stems form a triangle with the three helices connected by one or more unpaired nucleotides.In some examples, each of the helices is at least 4 bp (e.g., at least 5 bp, at least 10 bp, at least 20 bp, at least 30 bp, at least 40 bp, at least 50 bp, at least 75 bp, at least 90 bp, or at least 100 bp, e.g., 4 to 100 bp, 5 to 75 bp, or 10 to 50 bp), and the unpaired nucleotides forming the triangle are at least 3 nt (e.g., at least 4 nt, at least 5 nt, at least 10 nt, at least 20, at least 15, at least 30, at least 40, at least 50, or at least 60 nt, e.g., 3 to 60 nt, 3 to 30 nt, 3 to 25 nt, or 5 to 20 nt, e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 2, 25, 30, 35, 40, 45, 50, 55, or 60 nucleotides). Kissing interactions can occur between any two of these types of loops (e.g., between two or more binding domains each containing one or more helices). In some examples, to allow for extension of duplex formation after the initial loop kissing interaction, a helix in one dimerization domain (e.g., first dimerization domain 122) has a direct counterpart in the other binding domain (e.g., second dimerization domain 154). In some examples, dimerization domains containing helices to generate loops form a single kissing stem-loop when two or more dimerization domains (e.g., 122, 154 in Figure 6A) interact. In some examples, dimerization domains containing helices form multiple loops for kissing loop interactions when two or more dimerization domains (e.g., 122, 154 in Figure 6A) interact. In some examples, one or more dimerization domains (e.g., 122 in Figure 6A) contain a bulge, a single-base bulge, a mismatch or internal loop, or a helix that is destabilized by the inclusion of a GU wobble pair, but is matched with the other binding domain (e.g., 154 in Figure 6A) to favor extended duplex formation after initial kissing / pairing.In some examples, one or more dimerization domains (e.g., 122 in Figures 6A and 6G) ​​contain a destabilized helix that, when stabilized (e.g., a theophylline-switched kissing loop), exposes a loop that can interact with a second dimerization domain (e.g., 122 in Figures 6A and 6G) ​​through loop-loop interactions (e.g., kissing / pairing).

[0276] In some examples, these stem loops contain at least 10 nt, e.g., at least 20 nt, at least 25 nt, at least 50 nt, at least 75 nt, or at least 100 nt in length, e.g., 10-50 nt, 20-25 nt, 10-100 nt, 10-20 nt, or 20-40 nt in length. Each dimerization domain can contain at least one individual stem loop, e.g., at least 2, at least 5, at least 10, at least 15, or at least 20 individual stem loops, e.g., 1-20, 2-5, or 1-10.

[0277] In some embodiments, 3 to 10 portions of the coding sequence are connected by 2 to 9 kissing loops, e.g., 3 portions are connected by 2 kissing loops, 4 portions are connected by 3 kissing loops, etc., where each of the 2 to 9 kissing loops is different. In some embodiments, the kissing loop comprises a plurality of stem loops, e.g., 2 to 20 stem loops. In some embodiments, each of the plurality of stem loops within a kissing loop is the same. In some embodiments, each of the plurality of stem loops within a kissing loop is different. In some embodiments, the dimerization domain comprises 1 to 20 stem loops. In some embodiments, the dimerization domain comprises 1 stem loop to 20 stem loops. In some embodiments, the dimerization domain is one stem loop to two stem loops, one stem loop to three stem loops, one stem loop to four stem loops, one stem loop to five stem loops, one stem loop to six stem loops, one stem loop to seven stem loops, one stem loop to eight stem loops, one stem loop to nine stem loops, one stem loop to ten stem loops, one stem loop to fifteen stem loops, one stem loop to twenty stem loops, two stem loops to three stem loops, two stem loops to four stem loops, two stem loops to five stem loops, two stem loops to six stem loops, stem loops, 2 stem loops to 7 stem loops, 2 stem loops to 8 stem loops, 2 stem loops to 9 stem loops, 2 stem loops to 10 stem loops, 2 stem loops to 15 stem loops, 2 stem loops to 20 stem loops, 3 stem loops to 4 stem loops, 3 stem loops to 5 stem loops, 3 stem loops to 6 stem loops, 3 stem loops to 7 stem loops, 3 stem loops to 8 stem loops, 3 stem loops to 9 stem loops, 3 stem loops to 10 stem loops, 3 stem loops to 15 stem loops, 3 stem loops to 20 stem loops,4 stem loops to 5 stem loops, 4 stem loops to 6 stem loops, 4 stem loops to 7 stem loops, 4 stem loops to 8 stem loops, 4 stem loops to 9 stem loops, 4 stem loops to 10 stem loops, 4 stem loops to 15 stem loops, 4 stem loops to 20 stem loops, 5 stem loops to 6 stem loops, 5 stem loops to 7 stem loops, 5 stem loops to 8 stem loops, 5 stem loops to 9 stem loops, 5 stem loops to 10 stem loops, 5 stem loops to 15 stem loops, 5 stem loops to 20 stem loops, 6 stem loops to 7 stem loops, 6 stem loops to 8 stem loops, 6 stem loops to 9 stem loops, 6 stem loops The stem loops may include up to 10 stem loops, 6 stem loops to 15 stem loops, 6 stem loops to 20 stem loops, 7 stem loops to 8 stem loops, 7 stem loops to 9 stem loops, 7 stem loops to 10 stem loops, 7 stem loops to 15 stem loops, 7 stem loops to 20 stem loops, 8 stem loops to 9 stem loops, 8 stem loops to 10 stem loops, 8 stem loops to 15 stem loops, 8 stem loops to 20 stem loops, 9 stem loops to 10 stem loops, 9 stem loops to 15 stem loops, 9 stem loops to 20 stem loops, 10 stem loops to 15 stem loops, 10 stem loops to 20 stem loops, or 15 stem loops to 20 stem loops. In some embodiments, the dimerization domain comprises 1 stem loop, 2 stem loops, 3 stem loops, 4 stem loops, 5 stem loops, 6 stem loops, 7 stem loops, 8 stem loops, 9 stem loops, 10 stem loops, 15 stem loops, or 20 stem loops. In some embodiments, the dimerization domain comprises at least 1 stem loop, 2 stem loops, 3 stem loops, 4 stem loops, 5 stem loops, 6 stem loops, 7 stem loops,In some embodiments, the dimerization domain comprises at most 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, or 20 stem loops.

[0278] Other mechanisms can be used that allow two or more dimerization domains (e.g., 122, 154 in Figures 6A and 6G) ​​to bind or interact with each other sufficiently to cause recombination of the coding sequences. In some examples, the two or more dimerization domains (e.g., 122, 154 in Figures 6A and 6G) ​​can interact with each other, e.g., through non-base pairing interactions, or are nucleic acid aptamers (e.g., RNA aptamers) that can bind to a common molecule (e.g., a protein, ATP, a metal ion, a cofactor, or a synthetic ligand). In some examples, the two or more dimerization domains (e.g., 122, 154 in Figures 6A and 6G) ​​do not hybridize with each other, but both (or all) hybridize with the same bridged nucleic acid molecule. In some examples, such bridged nucleic acid molecules can be exogenously provided to a cell, tissue, or organism. In some examples, such bridged nucleic acid molecules can be DNA or RNA sequences within a cell, e.g., a transcript or genomic locus. In some embodiments, two or more dimerization domains (eg, 122, 154 in Figures 6A, 6G) are sequences that can interact with each other, for example, through non-base pairing interactions.

[0279] Molecule 150 is a 3'-located molecule and includes a splice acceptor (SA) 162 and a second dimerization domain 154. In embodiments in which molecule 150 is DNA, molecule 150 includes a second promoter 152 followed by an intron sequence 170. Promoter 152 can be operably linked to intron sequence 170. Any promoter 152 can be used, whether constitutive or inducible. In some examples, promoter 152 is a tissue-specific promoter, such as one constitutively active in muscle tissue (e.g., skeletal muscle tissue or cardiac muscle tissue), eye tissue (e.g., retinal tissue), inner ear tissue, liver tissue, pancreatic tissue, lung tissue, skin tissue, bone, or kidney tissue. In some examples, promoter 152 is a cell-specific promoter, such as one constitutively active in cancer cells or normal cells. In some embodiments, promoter 112 is the endogenous promoter of the target protein to be expressed, and in some embodiments, is long (e.g., at least 2500 nt, at least 3000 nt, at least 4000 nt, at least 5000 nt, or at least 7500 nt). In some embodiments, promoter 112 is at least about 50 nucleotides (nt) in length, e.g., at least 100 nt, at least 200 nt, at least 300 nt, at least 500 nt, at least 1000 nt, at least 2000 nt, at least 3000 nt, at least 4000 nt, at least 5000 nt, at least 6000 nt, at least 7000 nt, at least 8000 nt, at least 9000 nt, or at least 10,000 nt, e.g., 50-10,000 nt, 100-5000 nt, 500-5000 nt, or 50-1000 nt in length. In some embodiments, promoter 112 and promoter 152 are the same promoter. In other examples, promoter 112 and promoter 152 are different promoters.In some embodiments, molecule 150 is DNA and is at least 200 nt, at least 300 nt, at least 500 nt, at least 1000 nt, at least 2000 nt, at least 3000 nt, at least 4000 nt, at least 5000 nt, at least 6000 nt, at least 7000 nt, or at least 8000 nt, e.g., 200-10,000 nt, 200-8000 nt, 500-5000 nt, or 200-1000 nt in length. As shown in Figure 6F, in embodiments where molecule 150 is RNA, e.g., after expression of DNA into RNA, molecule 150 no longer comprises promoter 152, and 164 is RNA encoded by a coding sequence for the C-terminal portion of a nucleic acid editing protein. In some examples, molecule 150 is RNA, does not include promoter 152, and is at least 200 nt, at least 300 nt, at least 500 nt, at least 1000 nt, at least 2000 nt, at least 3000 nt, at least 4000 nt, at least 5000 nt, at least 6000 nt, at least 7000 nt, or at least 8000 nt, e.g., 200-10,000 nt, 200-8000 nt, 500-5000 nt, or 200-1000 nt in length. Molecule 150 (with or without promoter 152) can include natural and / or non-natural nucleotides or ribonucleotides.

[0280] Intron sequence 170 includes second dimerization domain 154, optional ISE 156, branch point 158, polypyrimidine tract 160, followed by splice acceptor sequence 162. In some examples, intron sequence 130 is approximately at least 10 nt, e.g., at least 20 nt, at least 30 nt, at least 50 nt, at least 100 nt, at least 250 nt, at least 250 nt, at least 300 nt, at least 400 nt, or at least 500 nt in length, e.g., 20-500, 20-250, 20-100, 50-100, 30-500, or 50-200 nt in length.

[0281] The second dimerization domain 154 has a sequence that is the reverse complement of the sequence of the first dimerization domain 122 of the molecule 110. Accordingly, the same design features and considerations as those for the first dimerization domain 122 described above also apply to the second dimerization domain 154. For example, in some embodiments, the second dimerization domain 154 contains a stem-loop that can form a kissing-loop interaction with the first dimerization domain 122. In some embodiments, the second dimerization domain 154 does not contain a cryptic splice acceptor (e.g., NNNAGGUNNN; SEQ ID NO: 143) that may compete with RNA recombination. In some embodiments, the second dimerization domain 154 has a low diversity sequence. In some embodiments, second dimerization domain 154 is 1000 nt or less, e.g., 750 nt or less, or greater than 500 nt, e.g., 30-1000 nt, 30-750 nt, 30-500 nt, 50-500 nt, 50-100 nt, or 100-250 nt. In some embodiments, second dimerization domain 154 is greater than 50 nt, e.g., at least 51 nt, at least 100 nt, at least 150 nt, at least 161 nt, or at least 170 nt, e.g., 51-159 nt, 51-150 nt, 51-120 nt, 51-100 nt, or 51-70 nt. In some embodiments, second dimerization domain 154 is greater than 160 nt, e.g., at least 161 nt, at least 170 nt, at least 180 nt, at least 200 nt, at least 300 nt, at least 400 nt, at least 500 nt, at least 600 nt, at least 700 nt, at least 800 nt, at least 900 nt, or at least 1000 nt, e.g., 161-100 nt, 161-500 nt, 161-300 nt, 161-200 nt, or 161-170 nt. In some embodiments, second dimerization domain 154 is less than 50 nt, e.g., 6-49 nt, 6-45 nt, 6-40 nt, 6-30 nt, 6-20 nt, or 6-10 nt.

[0282] 3' to second dimerization domain 154 is an optional ISE 156, a branchpoint sequence 158 (e.g., a branchpoint consensus sequence), a polypyrimidine tract 160, followed by a splice acceptor sequence 162. ISE 156, like ISE 120 and DISE 118 of molecule 110, stimulates the spliceosome to catalyze the recombination reaction. In some examples, intron sequence 150 includes at least two ISEs 156, e.g., at least three, at least four, or at least five ISEs 156. Exemplary splicing enhancer sequences include ISE 156. In some examples, inclusion of one or more splicing enhancer sequences 156 in intron sequence 150 increases the efficiency of recombination or splicing by at least 10%, at least 20%, at least 30%, at least 40%, or at least 50%. Exemplary splicing enhancer sequences that can be used are SEQ ID NOs: 26-136, 151, and 152, as well as GGGTTT, GGTGGT, TTTGGG, GAGGGG, GGTATT, GTAACG, GGGGGTAGG, GGAGGGTTT, GGGTGGTGT TTCAT, CCATTT, TTTTAAA, TGCAT, TGCATG, TGTGTT, CTAAC, TCTCT, TCTGT, TCTTT, TGCATG, CTAAC, CTGCT, TAACC, AGCTT, TTCATTA, GTTAG, TTTTGC, ACTAAT, ATGTTT, CTCTG, GGG, GGG(N)2-4GGG, TGGG, YCAY, UGCAUG, or 3x(G 3~6 N 1~7). In some examples, ISE156, if present, can be approximately at least 3 nt, at least 4 nt, at least 5 nt, at least 10 nt, e.g., at least 20 nt, at least 25 nt, at least 30 nt, at least 40 nt, or at least 50 nt in length, e.g., 3-10 nt, 3-11 nt, 4-11 nt, 5-11 nt, 10-50 nt, 20-25 nt, 10-25 nt, 10-20 nt, or 20-40 nt in length. In one example, the sequence of ISE156 is or includes GGCUGAGGGAAGGACUGUCCUGGG (SEQ ID NO: 135), GGGUUAUGGGACC (SEQ ID NO: 136), TTCAT, CCATTT, TTTTAAA, TGCAT, TGCATG, TGTGTT, CTAAC, TCTCT, TCTGT, or TCTTT. In some examples, ISE120 and ISE156 are the same sequence. In another example, ISE120 and ISE156 are different sequences.

[0283] 3' to the second dimerization domain 154 (and ISE 156, if present) is a branchpoint sequence 158 (e.g., a branchpoint consensus sequence), a polypyrimidine tract 160, followed by a splice acceptor sequence 162 (e.g., a splice acceptor consensus sequence). The sequence of the branchpoint 158 ​​is based on the consensus sequence of the species of the target cell or organism. For example, for human splicing, the consensus sequence can include or be YUNAY. Thus, the sequence used can be CUAAC for a U2-dependent intron, or UUUUCCUUAACU (SEQ ID NO: 144) for a U12-dependent intron.

[0284] The polypyrimidine tract 160 can contain C, U, or both C and U nucleotides, e.g., CnUy (where n+y is greater than or equal to 10 nucleotides), and can also include nucleotides -3 to -22 relative to the 3' splice junction. In some embodiments, the polypyrimidine tract 160 contains at least 80% Y nucleotides (i.e., U, C, or both U and C). In some embodiments, the polypyrimidine tract 160 is a poly-C or poly-U sequence. In some embodiments, the polypyrimidine tract 160 is a poly-U sequence with at least 15 Us, e.g., 15-30 or 15-20 Us. The branch point 158 ​​and the polypyrimidine tract 160 are essential splicing components. The sequence of SA 162 can be based on a consensus sequence for the species of the target cell or organism. For example, in humans, the SA sequence is AG at positions -1 and -2 relative to the 3' splice site for U2-dependent introns, and can be AC ​​or AG for U12-dependent introns. Thus, in some examples, SA162 can be 2 nt in length, e.g., AG or AC.

[0285] Immediately following SA162 is an exon sequence containing a DNA sequence encoding a C-terminal portion of target protein 164 with a splice junction at its 5' end. The splice junction at the 5' end of the DNA sequence encoding the C-terminal portion of nucleic acid editing protein 164 can match a consensus sequence found in the target cell or organism into which molecule 110, 150 is introduced. In some embodiments, the splice junction can be GA or GU at positions +1 and +2 relative to the 3' splice site for a U2-dependent intron, or GU or AU for a U12-dependent intron. Thus, in some embodiments, the splice junction is 2 nt in length, and the 5' end of C-terminal coding portion 164 is GA, GU, or AU.

[0286] The exon sequence following intron portion 170 of molecule 150 includes a second coding portion (e.g., half) of a nucleic acid editing protein, e.g., a C-terminal fragment 164, and an optional polyadenylation sequence 166. Thus, molecule 150 includes sequence 164 encoding the C-terminal portion of a nucleic acid editing protein. The 3' end of molecule 150 optionally includes polyadenylation sequence 166 that promotes spliceosome assembly. In some examples, polyadenylation sequence 166 is a polyA sequence of at least 15 A, e.g., 15-30 or 15-20 A. In some examples, polyadenylation sequence 166 and polyadenylation sequence 124 are the same sequence. In other examples, polyadenylation sequence 166 and polyadenylation sequence 124 are different sequences.

[0287] In some examples, N-terminal coding region 114 and / or C-terminal coding region 164 are native coding sequences. For example, the coding sequences are those found in the cell or organism into which the systems disclosed herein are introduced (e.g., human coding sequences in the case of introduction into a human cell or subject). In some examples, N-terminal coding region 114 and / or C-terminal coding region 164 are codon-optimized relative to the native coding sequence, e.g., to maximize tRNA availability or to de-enrich cryptic splice sites (e.g., to reduce or avoid incorrect splicing and promote correct junction formation). In some examples, portions of N-terminal coding region 114 and / or C-terminal coding region 164 are codon-optimized relative to the native coding sequence, e.g., approximately 200 nt adjacent to each junction (e.g., the 3' end of 114 and the 5' end of 164) can be codon-optimized or modified to contain an exonic splice enhancer site (ESE) (which binds an SR protein). For example, the coding sequence can be one that is not found in the cell or organism into which the system disclosed herein is introduced (e.g., a human coding sequence in the case of introduction into a mouse cell or subject).

[0288] In some examples, N-terminal coding region 114 and / or C-terminal coding region 164 include introns, either natural or synthetic in nature, containing both splice donor and acceptor sites. For example, an intron embedded within the coding sequence to be expressed can be included upstream of sequence 116 (e.g., about 200 nt upstream), within N-terminal coding region 114, or an intron embedded within the coding sequence to be expressed downstream of sequence 162 (e.g., about 200 nt downstream) and within C-terminal coding region 164, or both. The inclusion of such introns can be used to stimulate attachment of the splicing machinery to the trans-splicing intron donor and acceptor. In some examples, such (stimulatory) introns can be derived from the host expressing 110 and 150. In some examples, such (stimulatory) introns can be derived from other organisms, or from viral or synthetic sources.

[0289] In some examples, inclusion of a sequence to stabilize molecule 150 (e.g., located between 164 and 166 in the 3' untranslated region of 150 in Figure 6A) can increase the efficiency of expression of the recombinant product by at least 25%, at least 30%, at least 40%, at least 50%, at least 60%, or at least 75%, e.g., 25-95%, 25-75%, 25-60%, 25-50%, 40-95%, 40-60%, or 50-60%. In some examples, a woodchuck posttranscriptional regulatory element (WPRE) or a truncated version thereof (e.g., WPRE3) is included in the 3'-UTR as a stabilizing element to enhance the efficiency of expression of the recombinant product. In some examples, the WPRE sequence has at least 80%, at least 85%, at least 90%, at least 95%, or 100% sequence identity to nt 1093-1684 of GenBank Accession No. J04514 or the 247 bp sequence of WPRE3.

[0290] As shown in Figure 6G, in some embodiments, the system or DNA used to express a nucleic acid editing protein may further include one or more gRNA coding sequences 140, 141, 171, 172. Each gRNA comprises a first portion (in some embodiments, at least 15 nt, at least 16 nt, at least 17 nt, at least 18 nt, at least 19 nt, at least 20 nt, at least 25 nt, at least 30 nt, at least 35 nt, at least 40 nt, e.g., 15-50 nt, 15-40 nt, 15-30 nt, 15-25 nt, 28-32 nt, 25-35 nt, 17-24 nt, or 17-2 ...) that specifically hybridizes to a target nucleic acid molecule. a first portion (e.g., about 20 nt) that binds to a nucleic acid editing protein (in some examples, at least 20 nt, at least 30 nt, at least 40 nt, at least 50 nt, at least 60 nt, at least 70 nt, at least 75 nt, at least 80 nt, at least 90 nt, or at least 100 nt, e.g., 20-200 nt, 25-150 nt, 30-100 nt, 15-30 nt, 30-75 nt, or 75-100 nt). In some examples, the GC content of the portion of the gRNA that specifically hybridizes to the target nucleic acid molecule is about 40-80%. In some examples, one gRNA coding sequence 140, 141, 171, and 172 is at least about 60 nt, at least about 75 nt, at least about 80 nt, at least about 90 nt, at least about 100 nt, at least about 110 nt, or at least about 120 nt, e.g., 60-300 nt, 60-200 nt, 80-200 nt, 90-200 nt, 100-150 nt, or 100-120 nt. Expression of each gRNA 140, 141, 171, and 172 can be driven by a promoter 142, 143, 173, and 174, respectively, operably linked to the gRNA. In one example, each promoter 142, 143, 173, and 174 is the same. However, the promoters for each guide nucleic acid molecule 140, 141, 171, and 172 can be different. In some examples, the promoter is a polymerase III promoter, such as a human or mouse U6 or H1 promoter.The resulting gRNA forms a complex with the expressed nucleic acid editor, thereby enabling editing of the nucleic acid molecule to which the gRNA hybridizes. Thus, in some examples, the system or RNA composition is as shown in Figure 6C, showing the combination of N-terminal coding sequence 114 and C-terminal coding sequence 164, where the system or RNA may include one or more additional RNA molecules, i.e., one or more gRNA molecules (expressed from each gRNA 140, 141, 171, 172). In some examples, the systems or RNA compositions in Figure 6C are shown hybridized to one another. another) two RNA molecules, and the third RNA molecule comprises: (a) at least one first gRNA specific to a first target nucleic acid molecule, wherein the at least one first gRNA directs a nucleic acid editing protein to a target editing site on the first target nucleic acid molecule; (b) (i) at least one second gRNA specific to the first target nucleic acid molecule that directs the nucleic acid editing protein to the same or a different target editing site on the first nucleic acid molecule, or (i) at least one second gRNA specific to a second target nucleic acid molecule that directs the nucleic acid editing protein to a target editing site on the second target nucleic acid molecule; (c) (i) a nucleic acid editing protein directed by the first gRNA and second gRNA on the first nucleic acid molecule a fifth RNA molecule comprising: (i) at least one third gRNA specific to the first target nucleic acid molecule that directs the nucleic acid editing protein to the same or a different target editing site as the first gRNA, (ii) at least one third gRNA specific to the second target nucleic acid molecule that directs the nucleic acid editing protein to the same or a different target editing site as the second gRNA on the second nucleic acid molecule, or (iii) at least one third gRNA specific to the third target nucleic acid molecule that directs the nucleic acid editing protein to the target editing site on the third target nucleic acid molecule; and (d)(i) at least one fourth gRNA specific to the first target nucleic acid molecule that directs the nucleic acid editing protein to the same or a different target editing site as the first gRNA, the second gRNA, and the third gRNA on the first nucleic acid molecule;(ii) at least one fourth gRNA specific to the second target nucleic acid molecule that directs the nucleic acid editing protein to the same or a different target editing site as the second gRNA and the third gRNA on the second nucleic acid molecule; (iii) at least one fourth gRNA specific to the third target nucleic acid molecule that directs the nucleic acid editing protein to the same or a different target editing site as the third gRNA on the third target nucleic acid molecule; or (iv) at least one fourth gRNA specific to the fourth target nucleic acid molecule that directs the nucleic acid editing protein to the target editing site on the fourth target nucleic acid molecule.

[0291] Although one or more gRNA coding sequences 140, 141, 171, 172 are shown near the 5' and 3' ends of each molecule 110, 150 in Figure 6G, the gRNA coding sequences 140, 141, 171, 172 may be located elsewhere within each molecule 110, 150. Additionally, the promoter-gRNA coding sequences may be oriented forward or reverse relative to the direction of expression of the N-terminal and C-terminal coding sequences 114, 164. Thus, in some embodiments, molecule 110 includes (a) one or more promoter / gRNA coding sequences (e.g., 142 / 140) upstream of N-terminal coding sequence 114, (b) one or more promoter / gRNA coding sequences (e.g., 143 / 141) downstream of N-terminal coding sequence 114, or (c) both (a) and (b), and in some embodiments, molecule 140 includes (d) one or more promoter / gRNA coding sequences (e.g., 173 / 171) upstream of C-terminal coding sequence 164, or (e) one or more promoter / gRNA coding sequences (e.g., 174 / 172) downstream of C-terminal coding sequence 164, or both (d) and (e).

[0292] Figure 6G illustrates the presence of four gRNAs 140, 141, 171, and 172. However, the system or DNA may include fewer or more than four gRNAs. In some examples, the system, DNA, or RNA includes at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 20, at least 25, at least 50, at least 100, at least 500, or at least 500 gRNA or coding sequences. The gRNAs of the system, DNA, or RNA composition can (1) target the same target editing site of the same nucleic acid molecule (e.g., gene) and target nucleic acid; (2) target different target editing sites of the same nucleic acid molecule (e.g., gene) and target nucleic acid; (3) target different nucleic acid molecules (e.g., genes); or (4) any combination thereof. In some embodiments, one or more of gRNA coding sequences 140, 141, 171, 172 is a cassette that includes two or more gRNAs, allowing for expression of greater amounts of gRNAs. In some embodiments, the cassette encodes at least two gRNAs, at least three gRNAs, at least four gRNAs, at least five gRNAs, at least ten gRNAs, at least 20 gRNAs, at least 50 gRNAs, at least 100 gRNAs, or at least 500 gRNAs, where the sequence of each gRNA coding sequence within the cassette can be the same or different.

[0293] 6G illustrates four gRNA-encoding sequences 140, 141, 171, 172 as part of molecules 110, 150. However, in one example, one or more of the gRNA-encoding sequences 140, 141, 171, 172 are not part of molecules 110, 150, but instead are expressed from one or more separate DNA synthetic molecules, e.g., one or more other vectors. Thus, in some system and DNA embodiments, additional synthetic DNA molecules encoding one or more gRNAs are provided (e.g., in addition to molecules 110, 150). In one example, the system or DNA further includes at least one additional synthetic DNA encoding a first gRNA operably linked to a promoter, e.g., a synthetic DNA encoding one or more gRNAs operably linked to a promoter.

[0294] Figure 6G illustrates the presence of parvoviral inverted terminal repeats (ITRs) 176, 177, 178, and 179 at the 5' and 3' ends of each molecule 110, 150. Parvoviral ITRs contain origins of replication and are used to support AAV replication. Such ITRs can be used with parvoviral packaging plasmids, such as AAV. However, sequences 176, 177, 178, and 179 are optional. In one example, the parvoviral ITRs are adeno-associated virus (AAV) ITRs.

[0295] As shown in Figure 6C, the interaction and hybridization (base pairing) between the first dimerization domain 122 of molecule 110 and the second dimerization domain 154 of molecule 150 allows the spliceosome components to recombine the N-terminal coding sequence 114 and the C-terminal coding sequence 164. Specifically, the 3' end of the N-terminal protein coding sequence 114 is fused to the 5' end of the C-terminal protein sequence 164, creating a seamless junction between the two moieties. In examples where the system or DNA composition includes a gRNA coding sequence, such as that illustrated in Figure 6G, the same interaction and hybridization occurs to recombine the N-terminal coding sequence 114 and the C-terminal coding sequence 164, allowing for expression of a functional nucleic acid editing protein. The gRNAs expressed from 140, 141, 171, and 172 (whether as part of molecules 110, 150 or expressed from one or more different DNA molecules) are present in the cell in which expression occurs and form complexes with the expressed nucleic acid editing proteins (e.g., by interaction with the DR sequences of the tracrRNA and gRNA) to enable nucleic acid editing of the target nucleic acid molecule.

[0296] 6D and 6H show schematic diagrams of a system in which a nucleic acid editing protein is divided into three portions: an N-terminal portion, a middle portion, and a C-terminal portion (each portion may be of similar or different size). Thus, one skilled in the art will understand that a nucleic acid editing protein can be divided into any number of desired segments or portions, and a reasonable number of molecules can be designed using the information provided herein. In such an example, the system includes at least three synthetic nucleic acid molecules 110, 200, and 150, where molecule 110 includes molecule 114 encoding the N-terminal portion of the protein, molecule 200 includes molecule 216 encoding the middle portion of the protein, and molecule 150 includes molecule 164 encoding the C-terminal portion of the protein. Each nucleic acid molecule 110, 200, 150 may be composed of DNA and, after translation, may be RNA without the promoters 112, 202, 152 present. In some examples, molecules 110, 200, 150 (with or without promoters 112, 202, 152) are each at least about 100 nucleotides / ribonucleotides (nt) in length, e.g., at least 200 nt, at least 300 nt, at least 500 nt, at least 1000 nt, at least 2000 nt, at least 3000 nt, at least 4000 nt, at least 5000 nt, at least 6000 nt, at least 7000 nt, or at least 8000 nt, e.g., 200-10,000 nt, 200-8000 nt, 500-5000 nt, or 200-1000 nt. Molecules 110, 150, 200 (with or without promoters 112, 202, 152) can include natural and / or non-natural nucleotides or ribonucleotides. In addition to using two (or more) orthogonal dimerization domains, one of the two introns can be a U2-type intron and the second intron can be a U12-type intron. The splice donors and acceptors of U2- and U12-dependent introns show minimal cross-reactivity because the consensus recognition sequences differ between the two types of introns.Both strategies (i.e., orthogonal dimerization domains and U2-type introns versus U12-type introns) promote the recombination of the three fragments in the correct order (e.g., preventing the first fragment from joining directly with the last fragment and preventing the middle fragment from circularizing on itself).

[0297] 6D and 6H includes the same features as disclosed above with respect to FIGS. 1A and 6G, i.e., a promoter 112 operably linked to a sequence encoding an RNA molecule, the RNA molecule including, from 5' to 3': a coding sequence for the N-terminal portion of the target protein 114 including a splice junction at the 3' end of the target protein coding sequence, SD 116, optional DISE 118, optional ISE 120, a dimerization domain 122, and an optional polyadenylation sequence 124, wherein first dimerization domain 122 has reverse complementarity to third dimerization domain 204 of molecule 200. In embodiments where molecule 110 is RNA, e.g., after expression of DNA into RNA, as shown in FIG. 6F, molecule 110 does not include promoter 112, and 114 is RNA encoded by a coding sequence for the N-terminal portion of the target protein. The molecule 110 (with or without the promoter 112) can include natural and / or non-natural nucleotides or ribonucleotides.

[0298] Molecule 150 in Figures 6D and 6H includes the same features as those disclosed above with respect to Figures 1A and 6G, i.e., a promoter 152 operably linked to a sequence encoding an RNA molecule, which includes, from 5' to 3': a second dimerization domain 154, an optional ISE 156, a branch point sequence 158, a polypyrimidine tract 160, a splice acceptor (SA) 162; and a coding sequence for a C-terminal portion of a target protein 164, including a splice junction at the 5' end of the target protein coding sequence, and an optional polyadenylation sequence 166. Second dimerization domain 154 has reverse complementarity to fourth dimerization domain 226 of molecule 200. Molecule 150 (with or without promoter 152) can include natural and / or non-natural nucleotides or ribonucleotides.

[0299] Molecule 200 allows for the joining of N-terminal coding region 114 and C-terminal coding region 164 by providing a dimerization domain that is reverse complementary to dimerization domains 122, 154 of molecules 110 and 150, respectively. Molecule 200 includes features of both molecules 110 and 150, including two intron sequences 230, 240. Specifically, in embodiments in which molecule 200 is DNA, molecule 220 includes promoter 210 (which may be the same as or different from promoters 112 and / or 152) operably linked to a sequence encoding an RNA molecule, the RNA molecule including, from 5' to 3': third dimerization domain 204 (which is the reverse complement of first dimerization domain 122 of molecule 110 in FIG. 6D ), optional ISE 206, branch point 208, polypyrimidine tract 210, SA 212, targeting target nucleotide sequence 214, and / or targeting target sequence 216. 6D , molecule 150 includes a target protein intermediate portion coding sequence 216, which includes a splice junction at the 5' end of the target protein coding sequence and a splice junction at the 3' end of the target protein coding sequence, SD 220, optional DISE 222, optional ISE 224, fourth dimerization domain 226 (which is the reverse complement of fourth dimerization domain 154 of molecule 150 of FIG. 6D ), and optional polyadenylation sequence 228. In some embodiments, molecule 220 is DNA and is at least 200 nt, at least 300 nt, at least 500 nt, at least 1000 nt, at least 2000 nt, at least 3000 nt, at least 4000 nt, at least 5000 nt, at least 6000 nt, at least 7000 nt, or at least 8000 nt, e.g., 200-10,000 nt, 200-8000 nt, 500-5000 nt, or 200-1000 nt in length. In embodiments where molecule 200 is RNA, e.g., after expression of DNA into RNA, molecule 200 no longer comprises promoter 202, and 216 is RNA encoded by a coding sequence for the middle portion of the target protein.In some examples, molecule 200 is RNA, does not include promoter 202, and is at least 200 nt, at least 300 nt, at least 500 nt, at least 1000 nt, at least 2000 nt, at least 3000 nt, at least 4000 nt, at least 5000 nt, at least 6000 nt, at least 7000 nt, or at least 8000 nt, e.g., 200-10,000 nt, 200-8000 nt, 500-5000 nt, or 200-1000 nt in length. Molecule 200 (with or without promoter 202) can include natural and / or non-natural nucleotides or ribonucleotides.

[0300] As shown in Figure 6H, in some embodiments, the system or DNA used to express a nucleic acid editing protein may further comprise one or more gRNA coding sequences 140, 141, 171, 172, 231, 232. Each gRNA comprises a first portion (in some embodiments, at least 15 nt, at least 16 nt, at least 17 nt, at least 18 nt, at least 19 nt, at least 20 nt, at least 25 nt, at least 30 nt, at least 35 nt, at least 40 nt, e.g., 15-50 nt, 15-40 nt, 15-30 nt, 15-25 nt, 28-32 nt, 25-35 nt, 17-24 nt, or 17-2 ...) that specifically hybridizes to a target nucleic acid molecule. a first portion (e.g., about 20 nt) that binds to a nucleic acid editing protein (in some examples, at least 20 nt, at least 30 nt, at least 40 nt, at least 50 nt, at least 60 nt, at least 70 nt, at least 75 nt, at least 80 nt, at least 90 nt, or at least 100 nt, e.g., 20-200 nt, 25-150 nt, 30-100 nt, 15-30 nt, 30-75 nt, or 75-100 nt). In some examples, the GC content of the portion of the gRNA that specifically hybridizes to the target nucleic acid molecule is about 40-80%. In some embodiments, a gRNA coding sequence 140, 141, 171, 172, 231, 232 is at least about 60 nt, at least about 75 nt, at least about 80 nt, at least about 90 nt, at least about 100 nt, at least about 110 nt, or at least about 120 nt, e.g., 60-300 nt, 60-200 nt, 80-200 nt, 90-200 nt, 100-150 nt, or 100-120 nt. Expression of each gRNA 140, 141, 171, 172, 231, 232 can be driven by a promoter 142, 143, 173, 174, 233, 234, respectively, operably linked to the gRNA. In one embodiment, each promoter 142, 143, 173, 174, 233, 234 is the same. However, the promoter for each guide nucleic acid molecule 140, 141, 171, 172, 233, 234 may be different.173, 174, 233, 234 are polymerase III promoters, e.g., human or mouse U6 or H1 promoters. Upon expression, the resulting gRNA forms a complex with the expressed nucleic acid editor, thereby enabling editing of the nucleic acid molecule to which the gRNA hybridizes. Thus, in some examples, a system or RNA composition is as shown in Figure 6E, which shows the combination of N-terminal coding sequence 114, intermediate coding sequence 216, and C-terminal coding sequence 164, where the system or RNA may include one or more additional RNA molecules, i.e., one or more gRNA molecules (expressed from each gRNA 140, 141, 171, 172, 231, 232). In some examples, the system or RNA composition of Figure 6E comprises three RNA molecules shown hybridized to one another: (a) a fourth RNA molecule comprising at least one first gRNA specific to a first target nucleic acid molecule, wherein the at least one first gRNA directs a nucleic acid editing protein to a target editing site on the first target nucleic acid molecule; (b) a fifth RNA molecule comprising: (i) at least one second gRNA specific to the first target nucleic acid molecule that directs the nucleic acid editing protein to the same or a different target editing site on the first nucleic acid molecule, or (i) at least one second gRNA specific to a second target nucleic acid molecule that directs the nucleic acid editing protein to a target editing site on the second target nucleic acid molecule; and (c) a fifth RNA molecule comprising: (i) a nucleic acid editing protein specific to a first target nucleic acid molecule that directs the nucleic acid editing protein to a target editing site on the second target nucleic acid molecule. a sixth RNA molecule comprising: (i) at least one third gRNA specific to a first target nucleic acid molecule that directs a nucleic acid editing protein to the same or different target editing site as the first gRNA and the second gRNA on the first nucleic acid molecule; (ii) at least one third gRNA specific to a second target nucleic acid molecule that directs a nucleic acid editing protein to the same or different target editing site as the second gRNA on the second nucleic acid molecule; or (iii) at least one third gRNA specific to a third target nucleic acid molecule that directs a nucleic acid editing protein to the target editing site on the third target nucleic acid molecule; (d)(i) directs a nucleic acid editing protein to the same or different target editing site as the first gRNA, the second gRNA, and the third gRNA on the first nucleic acid molecule;(ii) at least one fourth gRNA specific to a first target nucleic acid molecule that directs a nucleic acid editing protein to the same or a different target editing site on the second nucleic acid molecule as the second gRNA and the third gRNA, (iii) at least one fourth gRNA specific to a third target nucleic acid molecule that directs a nucleic acid editing protein to the same or a different target editing site on the third target nucleic acid molecule as the third gRNA, or (iv) at least one fourth gRNA specific to a fourth target nucleic acid molecule that directs a nucleic acid editing protein to a target editing site on the fourth target nucleic acid molecule; (e) a seventh RNA molecule that includes at least one fifth gRNA, and (f) an eighth RNA molecule that includes at least one sixth gRNA.

[0301] Although one or more gRNA coding sequences 140, 141, 171, 172, 231, 232 are shown in Figure 6H near the 5' and 3' ends of each molecule 110, 150, the gRNA coding sequences 140, 141, 171, 172 may be located elsewhere within each molecule 110, 150. Furthermore, the promoter-gRNA coding sequences may be in either a forward or reverse orientation relative to the direction of expression of the N-terminal, intermediate, and C-terminal coding sequences 114, 216, 164. Furthermore, the promoter-gRNA coding sequences may be in either a forward or reverse orientation relative to the direction of expression of the N-terminal, intermediate, and C-terminal coding sequences 114, 216, 164. Thus, in some embodiments, molecule 110 includes (a) one or more promoter / gRNA coding sequences (e.g., 142 / 140) upstream of N-terminal coding sequence 114, (b) one or more promoter / gRNA coding sequences (e.g., 143 / 141) downstream of N-terminal coding sequence 114, or (c) both (a) and (b), and in some embodiments, molecule 140 includes (d) one or more promoter / gRNA coding sequences (e.g., 173 / 171) upstream of C-terminal coding sequence 164. ) (e) one or more promoter / gRNA coding sequences (e.g., 174 / 172) downstream of C-terminal coding sequence 164, or (f) both (d) and (e), and in some examples, molecule 200 includes (g) one or more promoter / gRNA coding sequences (e.g., 233 / 231) upstream of intermediate coding sequence 216, (h) one or more promoter / gRNA coding sequences (e.g., 234 / 232) downstream of intermediate coding sequence 216, or (i) both (g) and (h).

[0302] Figure 6H illustrates the presence of six gRNAs 140, 141, 171, 172, 231, and 232. However, the system or DNA may include fewer or more than six gRNAs. In some examples, the system, DNA, or RNA includes at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 20, at least 25, at least 50, at least 100, at least 500, or at least 500 gRNA or coding sequences. The gRNAs of the system, DNA, or RNA composition may (1) target the same target editing site of the same nucleic acid molecule (e.g., gene) and target nucleic acid; (2) target different target editing sites of the same nucleic acid molecule (e.g., gene) and target nucleic acid; (3) target different nucleic acid molecules (e.g., genes); or (4) any combination thereof. In some embodiments, one or more of gRNA coding sequences 140, 141, 171, 172, 231, 232 is a cassette that includes two or more gRNAs, allowing for expression of greater amounts of gRNAs. In some embodiments, the cassette encodes at least two gRNAs, at least three gRNAs, at least four gRNAs, at least five gRNAs, at least ten gRNAs, at least 20 gRNAs, at least 50 gRNAs, at least 100 gRNAs, or at least 500 gRNAs, where the sequence of each gRNA coding sequence within the cassette can be the same or different.

[0303] 6H illustrates four gRNA-encoding sequences 140, 141, 171, 172, 231, 232 as part of molecules 110, 150, 200. However, in one example, one or more of the gRNA-encoding sequences 140, 141, 171, 172, 231, 232 are not part of molecules 110, 150, 200, but instead are expressed from one or more separate DNA synthetic molecules, e.g., one or more other vectors. Thus, in some system and DNA embodiments, additional synthetic DNA molecules encoding one or more gRNAs are provided (e.g., in addition to molecules 110, 150, 200). In one example, the system or DNA further includes at least one additional synthetic DNA encoding a first gRNA operably linked to a promoter, e.g., a synthetic DNA encoding one or more gRNAs operably linked to a promoter.

[0304] Figure 6H illustrates the presence of parvoviral inverted terminal repeats (ITRs) 176, 177, 178, 179, 235, and 236 at the 5' and 3' ends of each molecule 110, 200, and 150. Such ITRs can be used with parvoviral packaging plasmids, such as AAV. However, the presence of sequences 176, 177, 178, 179, 235, and 236 is optional.

[0305] 6E, the interaction and hybridization (base pairing) between first dimerization domain 122 of molecule 110 and third dimerization domain 204 of molecule 200, and the interaction and hybridization (base pairing) between fourth dimerization domain 226 of molecule 200 and second dimerization domain 154 of molecule 150, allow the spliceosome components to recombine N-terminal coding sequence 114, intermediate coding sequence 216, and C-terminal coding sequence 164. Specifically, the 3' end of N-terminal protein coding sequence 114 is fused to the 5' end of intermediate protein sequence 216, and the 3' end of intermediate protein sequence 216 is fused to the 5' end of C-terminal protein sequence 164, forming a seamless junction between the three portions. In examples where the system or DNA composition includes a gRNA coding sequence, such as that illustrated in Figure 6H, the same interactions and hybridization occur to recombine the N-terminal coding sequence 114, the intermediate coding sequence 216, and the C-terminal coding sequence 164, thereby allowing expression of a functional nucleic acid editing protein. The gRNAs expressed from 140, 141, 171, 172, 231, 232 (whether as part of molecules 110, 150, 200 or expressed from one or more different DNA molecules) are present in the cell in which expression occurred and form a complex with the expressed nucleic acid editing protein (e.g., through interaction with the DR sequences of the tracrRNA and gRNA) to allow nucleic acid editing of the target nucleic acid molecule.

[0306] Alternative dimerization domains are shown in Figures 7A-7B and 9A. That is, as an alternative to using dimerization domains that hybridize to each other (e.g., 112-204, 226-154, Figures 6C, 6E), in one example, an aptamer sequence is used. As shown in Figure 7A, in both synthetic nucleic acid molecules 500, 600, aptamer sequences 512, 602 are used in place of dimerization domains, and the aptamers join together through their interaction with a target (e.g., adenosine, dopamine, or caffeine). In such an example, the aptamer sequences 512, 602 in each molecule 500, 600 can be the same sequence, or even different sequences. 6A and 6G, and if it is DNA, it includes a promoter operably linked to a sequence encoding an RNA molecule, the RNA molecule including, from 5' to 3': a coding sequence 502 for an N-terminal portion of a nucleic acid editing protein including a splice junction at the 3' end of the nucleic acid editing protein coding sequence, SD 506, optional DISE 508, optional ISE 510, a first aptamer 512 in place of a first dimerization domain, and an optional polyadenylation sequence. In embodiments where molecule 500 is RNA, for example, when transcribed from a DNA molecule, molecule 500 does not include a promoter (e.g., as shown in FIG. 7A). 6A and 6G, and, if DNA, includes a promoter operably linked to a sequence encoding an RNA molecule, the RNA molecule including, from 5' to 3': an aptamer 602 in place of second dimerization domain 154, an optional ISE 604, a branch point 606, a polypyrimidine tract 608, an SA 610, DNA encoding a C-terminal portion 614 of a target protein having a splice junction at the 5' end, and an optional polyadenylation sequence 616. In embodiments where molecule 600 is RNA, for example, when transcribed from a DNA molecule, molecule 500 does not include a promoter (e.g., as shown in FIG. 7A).Interaction of the two aptamers 512, 602 with each other or with molecule 700 allows spliceosome components to recombine the N-terminal coding sequence 502 and the C-terminal coding sequence 614. Specifically, the 3' end of the N-terminal protein coding sequence 502 is fused to the 5' end of the C-terminal protein sequence 614 as a seamless junction between the two moieties. Molecules 500 and 600 can contain natural and / or unnatural nucleotides or ribonucleotides.

[0307] In some examples, the aptamer sequences 512, 602 may recognize (e.g., specifically bind to) the same target 700 (FIG. 7A), or even recognize different targets (in which case each molecule specifically recognized by each aptamer, such as a caffeine / dopamine hybrid molecule, or a synthetic molecule containing a portion of a molecule recognized by an aptamer, is also administered with the systems presented herein). Exemplary targets recognized by aptamers include cellular proteins, small molecules, exogenous proteins, or RNA molecules.

[0308] Figure 7B shows an example similar to Figure 7A. The dimerization domains (512, 602 in Figure 7A) recognize RNA molecules. In the example shown in Figure 7B, each domain recognizes a different portion of an mRNA molecule, e.g., a cancer-specific transcript, that is expressed only in target cells (cells in which expression of the nucleic acid editing protein is desired). In such an example, the coding sequence (502, 614 in Figure 7A) composed of RNA is recombined only in the presence of the specific RNA molecule recognized by the dimerization domain. Here, the nucleic acid editing protein is expressed only in cancer cells and not in normal cells. Such a system makes it possible to control the expression of the nucleic acid editing protein.

[0309] 7C presents an illustrative "off switch" example. Here, hybridization / binding of dimerization domains 812, 902 (which are reverse complements to each other) of synthetic nucleic acid molecules 800, 900 can be reduced by providing anti-binding domain oligonucleotides (e.g., RNA or DNA) 1000 (which can be two different anti-binding domain oligonucleotides 1000, one the reverse complement of 812 and one the reverse complement of 912) that compete for binding / hybridization. Thus, anti-binding domain oligonucleotides 1000 can act as an "off switch" for the reconstitution of proteins encoded by N-terminal and C-terminal encoding portions 802 and 914, respectively. Molecule 800 of Figure 7C includes the same features as disclosed above with respect to molecule 110 of Figures 6A and 6G, which is an RNA molecule (and therefore lacks a promoter), and the RNA molecule includes, from 5' to 3': a coding sequence 802 for the N-terminal portion of a nucleic acid editing protein, including a splice junction at the 3' end of the nucleic acid editing protein coding sequence, SD 806, optional DISE 808, optional ISE 810, a dimerization domain 812, and an optional polyadenylation sequence 814. Similarly, molecule 900 in Figure 7C includes the same features as disclosed above with respect to molecule 150 in Figures 6A and 6G, which is an RNA molecule (and thus lacks a promoter), including from 5' to 3': an anti-dimerization domain 902, an optional ISE 904, a branch point 906, a polypyrimidine tract 908, an SA 910, an RNA 914 encoding the C-terminal portion of a nucleic acid editing protein, and an optional polyadenylation sequence 916. The two dimerization domains 812, 902 cannot interact / hybridize with each other in the presence of anti-binding domain oligonucleotide 1000, thus preventing or reducing recombination of the N-terminal coding sequence 802 and the C-terminal coding sequence 914. Molecules 800 and 900 may include natural and / or non-natural nucleotides or ribonucleotides.

[0310] Figure 9A shows an exemplary dimerization domain that uses kissing loop interactions instead of reverse-complementary hybridization for dimerization. Kissing loop interactions are formed when bases within the loops of two RNA hairpins form an interacting pair between two RNA molecules. The molecule on the left, labeled n-yfp, represents an RNA molecule encoding the N-terminal fragment of yfp linked to a synthetic intron containing a splice donor site, a downstream intron splicing enhancer element, and two intron splicing enhancer elements. The dimerization domain of this molecule contains three RNA hairpin loops, each consisting of a stem (where the RNA hybridizes to itself) and a loop (where the RNA does not hybridize to itself). In this example, the dimerization domain contains three stem an...

Claims

1. A dual vector composition for expressing a nucleic acid editing protein, comprising: (a) a first adeno-associated virus (AAV) vector packaging a first transgene, the first transgene being 5' to 3': (i) a first promoter; (ii) a first coding sequence encoding an N-terminal portion of the nucleic acid editing protein; (iii) a splice donor; and (iv) a first dimerization domain a first AAV vector comprising: (b) a second AAV vector packaging a second transgene, the second transgene being 5' to 3': (i) a second promoter; (ii) a second dimerization domain; (iii) branch point sequence; (iv) a polypyrimidine tract; (v) a splice acceptor; and (vi) a second coding sequence encoding a C-terminal portion of the nucleic acid editing protein; A second AAV vector comprising: Including, A dual vector composition, wherein the first promoter is operably linked to the first coding sequence and the second promoter is operably linked to the second coding sequence.

2. The first transgene further comprises one or more first guide RNA (gRNA) coding sequences and a third promoter operably linked to the one or more first gRNA coding sequences; the one or more first gRNA coding sequences encode one or more first gRNAs specific to a first target nucleic acid molecule; 2. The dual vector composition of Claim 1, wherein the one or more first gRNAs direct the nucleic acid editing protein to a target editing site on the first target nucleic acid molecule.

3. The first transgene further comprises one or more second gRNA coding sequences and a fourth promoter operably linked to the one or more second gRNA coding sequences; The one or more second gRNA coding sequences (i) the first target nucleic acid molecule, wherein the one or more second gRNAs direct the nucleic acid editing protein to the same or a different target editing site on the first nucleic acid molecule; or (ii) a second target nucleic acid molecule, wherein the one or more second gRNAs direct the nucleic acid editing protein to a target editing site on the second target nucleic acid molecule; 3. The dual vector composition of claim 2, further comprising one or more second gRNAs specific for the p53 gene.

4. The second transgene further comprises one or more third gRNA coding sequences and a fifth promoter operably linked to the one or more third gRNA coding sequences; The one or more third gRNA coding sequences (i) the first target nucleic acid molecule, wherein the one or more third gRNAs direct the nucleic acid editing protein to the same or different target editing site on the first nucleic acid molecule as the first gRNA and the second gRNA; (ii) the second target nucleic acid molecule, wherein the one or more third gRNAs direct the nucleic acid editing protein to the same or a different target editing site on the second nucleic acid molecule as the second gRNA; or (iii) a third target nucleic acid molecule, wherein the one or more third gRNAs direct the nucleic acid editing protein to a target editing site on the third target nucleic acid molecule; The dual vector composition of claim 3, further comprising one or more third gRNAs specific for the p53 gene.

5. The second transgene further comprises one or more fourth gRNA coding sequences and a sixth promoter operably linked to the one or more fourth gRNA coding sequences; The one or more fourth gRNA coding sequences (i) the first target nucleic acid molecule, wherein the one or more fourth gRNAs direct the nucleic acid editing protein to the same or different target editing site on the first nucleic acid molecule as the first gRNA, the second gRNA, and the third gRNA; (ii) the second target nucleic acid molecule, wherein the one or more fourth gRNAs direct the nucleic acid editing protein to the same or a different target editing site on the second nucleic acid molecule as the second gRNA and the third gRNA; (iii) the third target nucleic acid molecule, wherein the one or more fourth gRNAs direct the nucleic acid editing protein to the same or a different target editing site on the third target nucleic acid molecule as the third gRNA; or (iv) a fourth target nucleic acid molecule, wherein the one or more fourth gRNAs direct the nucleic acid editing protein to a target editing site on the fourth target nucleic acid molecule; The dual vector composition of claim 4, further comprising one or more fourth gRNAs specific for the p53A gene.

6. A dual vector composition described in any one of claims 1 to 5, wherein when the first introduced gene and the second introduced gene are transcribed, transcripts of the first dimerization domain and the second dimerization domain bind to each other by direct binding, indirect binding, or a combination thereof.

7. The dual vector composition of claim 6, wherein the indirect bond comprises a non-base pairing interaction between an aptamer and an aptamer target, or between two aptamers.

8. The dual vector composition of claim 6, wherein the transcript of the first dimerization domain and / or the transcript of the second dimerization domain does not contain a latent splice acceptor.

9. The dual vector composition of claim 6, wherein the transcripts of the first dimerization domain and the second dimerization domain are directly or indirectly bound to an aptamer sequence dimerization domain.

10. The dual vector composition of claim 6, wherein the transcript of the first dimerization domain comprises a first kissing loop interaction domain and the transcript of the second dimerization domain comprises a second kissing loop interaction domain.

11. The method of claim 10, wherein the first kissing loop interaction domain comprises a first RNA hairpin, the first RNA hairpin comprising an RNA complementary sequence separated by a region of non-complementary RNA sequence, wherein base pairing between the complementary sequences forms a first stem, and wherein the region of non-complementary RNA sequence forms a first loop; the second kissing loop interaction domain comprises a second RNA hairpin, the second RNA hairpin comprising an RNA complementary sequence separated by a region of non-complementary RNA sequence, wherein base pairing between the complementary sequences forms a second stem, and the region of non-complementary RNA sequence of the second RNA hairpin forms a second loop; The dual vector composition of claim 10 , wherein the first loop hybridizes to the second loop.

12. A dual vector composition described in any one of claims 2 to 5, wherein the target editing site is a part of a target nucleic acid molecule or a regulatory region of a target nucleic acid molecule associated with a disease.

13. The dual vector composition of claim 12, wherein the disease is a single-gene disease.

14. The dual vector composition of claim 13, wherein the target nucleic acid molecule contains one or more point mutations that result in the disease.

15. The first introduced gene further comprising one or both of a downstream intron splice enhancer (DISE) 3' to the splice donor and 5' to the first dimerization domain and an intron splice enhancer (ISE) 3' to the splice donor and 5' to the first dimerization domain; the second transgene further comprises one or both of an ISE 3' to the second dimerization domain and 5' to the branch point sequence, and a DISE 3' to the splice donor and 5' to the dimerization domain; or A dual vector composition according to any one of claims 1 to 5, which is any combination thereof.

16. The first introduced gene further comprises a first synthetic intron, the first synthetic intron encoding a self-cleaving RNA sequence or an RNA cleavage enzyme target sequence located anywhere 3' to the splice donor, such that the 3'-located polyadenylation tail is cleaved to reduce or inhibit expression of a protein fragment from an unmodified RNA molecule; the second transgene further comprises a second synthetic intron, which encodes a self-cleaving RNA sequence or an RNA cleavage enzyme target sequence located anywhere 5' to the branch point sequence, thereby cleaving the 5'-located RNA cap to reduce or suppress expression of a protein fragment from the unrecombined RNA molecule; the second transgene further comprises a start codon anywhere 5' to the branch point sequence that is shifted relative to the open reading frame 3' to the splice acceptor to reduce or prevent translation of the nucleic acid editing protein fragment from the unrecombined RNA molecule; said first transgene further comprises a first microRNA target site anywhere 3' to said splice donor, such that unbound RNA fragments are subject to microRNA-dependent degradation upon exiting the nucleus; said second transgene further comprises a second microRNA target site anywhere 3' to said coding sequence, such that unbound RNA fragments are subject to microRNA-dependent degradation once outside the nucleus; the first transgene further comprises a sequence encoding a degron proteolytic tag anywhere 3′ to the splice donor and in frame with the open reading frame of the nucleic acid editing protein 5′ to the splice donor site, such that unbound protein fragments are tagged for degradation; the second transgene further comprises a start codon and an in-frame degron proteolytic tag anywhere 5' to the branchpoint sequence and in-frame with the open reading frame of the nucleic acid editing protein 3' to the splice acceptor site, such that unbound protein fragments are tagged for degradation; or A dual vector composition according to any one of claims 1 to 5, which is any combination thereof.

17. A dual vector composition described in any one of claims 1 to 5, wherein the nucleic acid editing protein comprises a Cas nuclease, a zinc finger nuclease, or a transcription activator-like effector nuclease.

18. The dual vector composition of claim 5, wherein the first target nucleic acid molecule, the second target nucleic acid molecule, the third target nucleic acid molecule, and / or the fourth target nucleic acid molecule are target DNA molecules, and the one or more first gRNA, second gRNA, third gRNA, and / or fourth gRNA comprise crRNA and tracrRNA.

19. The dual vector composition of claim 18, wherein the nucleic acid editing protein comprises Cas9 or dead Cas9 (dCas9).

20. The Cas9 protein comprises at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity to SEQ ID NO:208 and is capable of functioning as an RNA-guided DNA endonuclease; the Cas9 protein is encoded by a sequence comprising at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity to SEQ ID NO:207 and encodes an RNA-guided DNA endonuclease; the dCas9 protein comprises at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity to SEQ ID NO:210 and is catalytically inactive; or 20. The dual vector composition of claim 19, wherein the dCas9 protein is encoded by a sequence that comprises at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity to SEQ ID NO:209 and encodes a catalytically inactive protein.

21. The dual vector composition of claim 19 or 20, wherein the Cas9 or dCas9 is part of a fusion protein.

22. The fusion protein comprising Cas9 or dCas9 and a transcription activation domain, e.g., VP64, P65, MyoD1, HSF1, RTA, CBP, SET7 / 9, or any combination thereof; Cytosine base editors (CBEs), such as those from sea lamprey [AID], CDA1, or APOBEC3G; The bacteriophage protein Gam; and 22. The dual vector composition of claim 21 , comprising one or more of an adenine base editor (ABE), e.g., ABE 6.3, 7.8, 7.9, 7.10, ABEmax, ABE8s, such as ABE8e (TadA-8e V106W).

23. The dual vector composition of claim 5, wherein the first target nucleic acid molecule, the second target nucleic acid molecule, the third target nucleic acid molecule, and / or the fourth target nucleic acid molecule is a target RNA, and the one or more first gRNAs, second gRNAs, third gRNAs, and / or fourth gRNAs comprise one or more direct repeats and one or more spacers.

24. The dual vector composition of any one of claims 1 to 5, wherein the nucleic acid editing protein comprises Cas13a, Cas13b, Cas13c, Cas13d or dead Cas13d (dCas13d).

25. The Cas13d protein comprises at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity to SEQ ID NO:212, 215, or 222 and is capable of functioning as an RNA-guided RNA endonuclease; the Cas13d protein is encoded by a sequence comprising at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity to SEQ ID NO: 211 or 213 and encodes an RNA-guided RNA endonuclease; the dCas13d protein comprises at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity to SEQ ID NO:215 or 216 and is catalytically inactive; or 25. The dual vector composition of Claim 24, wherein the dCas13d protein is encoded by a sequence encoding a protein that comprises at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity to SEQ ID NO: 215 or 216, and that encodes a protein that is catalytically inactive.

26. The dual vector composition of claim 24, wherein the Cas13a, Cas13b, Cas13c, Cas13d or dCas13d is part of a fusion protein.

27. ​​The fusion protein comprises Cas13a, Cas13b, Cas13c, Cas13d or dCas13d, A transcription activation domain, e.g., VP64, P65, MyoD1, HSF1, RTA, CBP, SET7 / 9, or any combination thereof; Cytosine base editors (CBEs), such as those from sea lamprey [AID], CDA1, or APOBEC3G; The bacteriophage protein Gam; and 27. The dual vector composition of claim 26, comprising one or more of an adenine base editor (ABE), e.g., ABE 6.3, 7.8, 7.9, 7.10, ABEmax, ABE8s, such as ABE8e (TadA-8e V106W).

28. A dual vector composition described in any one of claims 1 to 5, further comprising a parvovirus inverted terminal repeat (ITR) at the 5' and 3' ends of each of the first introduced gene and the second introduced gene.

29. The dual vector composition described in claim 28, wherein the serotype of the first AAV particle and the second AAV particle is AAV9.

30. The dual vector composition of claim 28, wherein the serotype of the first AAV particle and the second AAV particle is AAVrh.

10.

31. A dual vector composition described in any one of claims 1 to 5, wherein each promoter is independently selected.

32. The method of claim 31, wherein the first promoter and the second promoter are the same promoter. The first promoter and the second promoter are different promoters. The third promoter, the fourth promoter, the fifth promoter and the sixth promoter are the same promoter; the third promoter, the fourth promoter, the fifth promoter and the sixth promoter are different promoters; or A combination of these: The dual vector composition of claim 5.

33. The dual vector composition of any one of claims 1 to 5, wherein each of the first promoter and the second promoter is independently selected from a constitutive promoter; a tissue-specific promoter; and a promoter endogenous to the nucleic acid editing protein.

34. The dual vector composition of claim 32, wherein each of the third promoter, the fourth promoter, the fifth promoter and the sixth promoter is a polymerase III promoter, such as a U6 or H1 promoter.

35. A dual vector composition described in any one of claims 1 to 5, which, when introduced into a cell, produces and recombines RNA molecules in the appropriate order, thereby resulting in a full-length coding sequence for the nucleic acid editing protein.

36. Each of the first transgene and the second transgene is about 2500nt to about 5000nt, 2,500nt to about 2,750nt, about 2,500nt to about 3,000nt, about 2,500nt to about 3,250nt, about 2,500nt to about 3,500nt, about 2,500nt to about 3,750nt, about 2,500nt to about 4,000nt, about 2,500nt to about 4,250nt, about 2,500nt to about 4,500nt, about 2,500nt to about 4,750nt, about 2,500nt to about 5,000nt, about 2,750nt to about 3,000nt, about 2,75 0 nt to about 3,250 nt, about 2,750 nt to about 3,500 nt, about 2,750 nt to about 3,750 nt, about 2,750 nt to about 4.0 nt 00nt, about 2,750nt to about 4,250nt, about 2,750nt to about 4,500nt, about 2,750nt to about 4,750nt, about 2, 750 nt to about 5,000 nt, about 3,000 nt to about 3,250 nt, about 3,000 nt to about 3,500 nt, about 3,000 nt to about 3 , 750nt, about 3,000nt to about 4,000nt, about 3,000nt to about 4,250nt, about 3,000nt to about 4,500nt, about 3,000 nt to about 4,750 nt, about 3,000 nt to about 5,000 nt, about 3,250 nt to about 3,500 nt, about 3,250 nt to Approximately 3,750 nt, approximately 3,250 nt to approximately 4,000 nt, approximately 3,250 nt to approximately 4,250 nt, approximately 3,250 nt to approximately 4,500 nt , about 3,250 nt to about 4,750 nt, about 3,250 nt to about 5,000 nt, about 3,500 nt to about 3,750 nt, about 3,500 nt t to about 4,000 nt, about 3,500 nt to about 4,250 nt, about 3,500 nt to about 4,500 nt, about 3,500 nt to about 4,750 nt nt, about 3,500 nt to about 5,000 nt, about 3,750 nt to about 4,000 nt, about 3,750 nt to about 4,250 nt, about 3,75 0 nt to about 4,500 nt, about 3,750 nt to about 4,750 nt, about 3,750 nt to about 5,000 nt, about 4,000 nt to about 4,2 50 nt, about 4,000 nt to about 4,500 nt, about 4,000 nt to about 4,750 nt, about 4,000 nt to about 5,000 nt, about 4, 250 nt to about 4,500 nt, about 4,250 nt to about 4,750 nt, about 4,250 nt to about 5,000 nt, about 4,500 nt to about 4,The dual vector composition of any one of claims 1 to 5, having a size independently selected from about 750nt, about 4,500nt to about 5,000nt, about 4,750nt to about 5,000nt, about 2,500nt, about 2,750nt, about 3,000nt, about 3,250nt, about 3,500nt, about 3,750nt, about 4,000nt, about 4,250nt, about 4,500nt, about 4,750nt, and about 5,000nt.

37. The coding sequence of the N-terminal portion of the nucleic acid editing protein, or the coding sequence of the C-terminal portion of the nucleic acid editing protein, is about 2500 to 4500 nt, about 2,500 nt to about 2,750 nt, about 2,500 nt to about 3,000 nt, about 2,500 nt to about 3,250 nt, about 2,500 nt to about 3,500 nt, about 2,500 nt to about 3,750 nt, about 2,500 nt to about 4,000 nt, about 2,500 nt to about 4,250 nt, about 2,500 nt to about 4,5 00nt, about 2,750nt to about 3,000nt, about 2,750nt to about 3,250nt, about 2,750nt to about 3,500nt, about 2,750nt to about 3,750nt, about 2,750nt to about 4,000nt, about 2,750nt to about 4, 250 nt, about 2,750 nt to about 4,500 nt, about 3,000 nt to about 3,250 nt, about 3,000 nt to about 3,500 nt, about 3,000 nt to about 3,750 nt, about 3,000 nt to about 4,000 nt, about 3,000 nt to about 4 , 250 nt, about 3,000 nt to about 4,500 nt, about 3,250 nt to about 3,500 nt, about 3,250 nt to about 3,750 nt, about 3,250 nt to about 4,000 nt, about 3,250 nt to about 4,250 nt, about 3,250 nt to about 4,500 nt, about 3,500 nt to about 3,750 nt, about 3,500 nt to about 4,000 nt, about 3,500 nt to about 4,250 nt, about 3,500 nt to about 4,500 nt, about 3,750 nt to about 4,000 nt, about 3,750 nt to 6. The dual vector composition of claim 1, having a size independently selected from about 4,250 nt, about 3,750 nt to about 4,500 nt, about 4,000 nt to about 4,250 nt, about 4,000 nt to about 4,500 nt, about 4,250 nt to about 4,500 nt, about 2,500 nt, about 2,750 nt, about 3,000 nt, about 3,250 nt, about 3,500 nt, about 3,750 nt, about 4,000 nt, about 4,250 nt, and about 4,500 nt.

38. The dual vector composition of claim 6, wherein the first dimerization domain and the second dimerization domain are each 1000 nt or less; and the recombination efficiency between the transcript of the first introduced gene and the transcript of the second introduced gene is at least about 20%, at least about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, or about 100%.

39. The dual vector composition of any one of claims 1 to 5, wherein one or both of the first introduced gene and the second introduced gene further comprise a synthetic intron, the synthetic intron having at least 80%, at least 85%, at least 90%, at least 95%, at least 98%, at least 99%, or 100% sequence identity to any one of SEQ ID NOs: 159, 160, 161, 162, 163, 164, 165, 166, 225, and 226.

40. An N-terminal plasmid comprising: a first AAV inverted terminal repeat (ITR) comprising at least 90% sequence identity to nt 1-141 of SEQ ID NO:225; one or more first gRNA coding sequences; A coding sequence for the N-terminal portion of the nucleic acid editing protein; A synthetic intron comprising a DISE, an ISE, or both, having at least 90% sequence identity to nt 3218-3315 of SEQ ID NO:225; a dimerization domain encoding an RNA dimerization domain comprising at least 95% sequence identity to nt 3316-3427 of SEQ ID NO:225; One or more second gRNA coding sequences; and A second AAV ITR comprising at least 90% sequence identity to nt 3825-3965 of SEQ ID NO:

225. An N-terminal plasmid comprising:

41. A C-terminal plasmid comprising: a first AAV ITR comprising at least 90% sequence identity to nt 1-141 of SEQ ID NO:226; one or more first gRNA coding sequences; a dimerization domain encoding an RNA dimerization domain, comprising at least 95% sequence identity to nt 1259-3910 of SEQ ID NO:226; One or more ISE sequences comprising at least 90% sequence identity to SEQ ID NOs: 199, 200, 201, 202, or 203; Branch point sequence; Polypyrimidine tract; Splice acceptor; A coding sequence for the C-terminal portion of the nucleic acid editing protein; One or more second gRNA coding sequences; and A second AAV ITR comprising at least 90% sequence identity to nt 4411-4551 of SEQ ID NO:

226. A C-terminal plasmid comprising:

42. A dual AAV particle composition comprising: A first AAV particle packaging the N-terminal plasmid of claim 40; and A second AAV particle packaging the C-terminal plasmid of claim 41.

1. A dual AAV particle composition comprising:

43. The dual AAV particle composition described in claim 42, wherein the serotype of the first AAV particle and the second AAV particle is AAV9.

44. The dual AAV particle composition of claim 42, wherein the serotype of the first AAV particle and the second AAV particle is AAVrh.

10.

45. A kit comprising a dual vector composition according to any one of claims 1 to 5, an N-terminal plasmid according to claim 40 and a C-terminal plasmid according to claim 41, or a first AAV particle and a second AAV particle according to claim 43 or 44, wherein any of the first AAV vector, the second AAV vector, the N-terminal plasmid, the C-terminal plasmid, the first AAV particle, and the second AAV particle may be contained in separate containers, and the kit further comprises a buffer such as a pharma- ceutically acceptable carrier as necessary.