Methods and compositions for reducing unspliced RNA
By employing multiple synthetic RNA molecules to encode and join protein portions within cells, the method addresses the challenge of expressing large proteins in gene therapy, achieving efficient and safe full-length protein expression for genetic disease treatment.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- SALK INST FOR BIOLOGICAL STUDIES
- Filing Date
- 2025-05-19
- Publication Date
- 2026-07-30
AI Technical Summary
Existing gene therapy methods face challenges in delivering and expressing large proteins due to packaging constraints of vectors like AAV, leading to inefficient expression of truncated proteins or toxic fragments, limiting treatment of genetic diseases.
A strategy involving the use of two or more synthetic RNA molecules, each encoding a portion of a target protein, which are designed to join within a cell, allowing for the expression of a full-length protein through dimerization domains and splice sites, avoiding coding sequences for secondary structures that recruit nuclear RNA binding proteins.
This approach enables efficient and safe expression of full-length proteins, such as nucleic acid editing enzymes, in target cells, effectively treating genetic diseases by ensuring seamless protein reconstitution without toxic truncated forms.
Smart Images

Figure US2025028487_30072026_PF_FP_ABST
Abstract
Description
[0001] SLR 7158-102574-27 05 / 08 / 25 S2024-016A
[0002] METHODS AND COMPOSITIONS FOR REDUCING UNSPLICED RNA CROSS-REFERENCE TO RELATED APPLICATIONS
[0003] This application claims priority to US 63 / 644,407 filed May 8, 2024, herein incorporated by reference in its entirety.
[0004] FIELD
[0005] The present disclosure provides systems, kits, compositions, and methods that allow for joining of two or more RNA molecules, allowing expression of a full-length protein, such as a protein involved in gene editing such as a Cas nuclease, or catalytically inactive forms of a Cas nuclease. Also provided are systems, kits, compositions, and methods that reduce truncated protein expression from unspliced RNAs (unjoined fragments) in the cytosol.
[0006] INCORPORATION OF SEQUENCE LISTING
[0007] The Sequence Listing is submitted as an XML file titled “SequenceListing.xml”, created on May 8, 2025, approximately 614,959 bytes in size, and is incorporated by reference herein.
[0008] ACKNOWLEDGEMENT OF GOVERNMENT SUPPORT
[0009] This invention was made with government support under contract No. NICHD 5F30HD106732 awarded by the National Institutes of Health and under contract No. W81XWH-20- 1-0423 awarded by the Department of Defense. The government has certain rights in the invention.
[0010] BACKGROUND
[0011] Gene therapy is a promising method for treating genetic diseases caused by loss-of-function mutations. Replacement genes are typically reintroduced into target cells using vectors such as AAV because the virus is generally safe and efficient at entering cells. However, in the case of AAV it is difficult to encapsulate more than about 5000 nucleotides using conventional capsids. Since the length of genes that encode large proteins often exceed the packaging constraints of AAV, many genetic diseases remain untreatable. Strategies to overcome this limitation have been explored in the past, but proved inefficient, led to expression of high levels of potentially toxic truncated protein, or both. Safe, high efficiency strategies for delivery of large proteins to treat disease are needed.
[0012] SUMMARY
[0013] Provided herein are compositions for expressing a target protein, such as a protein used to edit a nucleic acid sequence (such as target DNA or RNA, such as a gene). Included are compositions and methods for expressing a nucleic acid editing protein produced from two or more synthetic nucleic acid molecules introduced individually to the same cell. Using this strategy, a full-length nucleic acid editingSLR 7158-102574-27 05 / 08 / 25 S2024-016A
[0014] protein and one or more guide RNAs can be provided to the same cell, resulting in targeted nucleic acid editing. The cell may be in need of targeted nucleic acid editing to repair a mutation, e.g., in an essential gene. In one example, the composition for expressing a target protein includes (a) a first RNA molecule, the RNA molecule comprising from 5’ to 3’: (i) a coding sequence for an N-terminal portion of the target protein; (ii) a splice donor; and (iii) a first dimerization domain; (b) a second RNA molecule, the RNA molecule comprising from 5’ to 3’: (i) a second dimerization domain, wherein the second dimerization domain binds to the first dimerization domain; (ii) a branch point sequence; (iii) a polypyrimidine tract; (iv) a splice acceptor; and (v) a coding sequence for a C-terminal portion of the target protein; and (c) one or more secondary structures in the first RNA molecule, one or more secondary structures in the second RNA molecule, or both, wherein the one or more secondary structures are not in the coding sequence for the N-terminal portion of the target protein nor in the coding sequence for the C-terminal portion of the target protein; or one or more elements to recruit nuclear RNA binding proteins in the first RNA molecule, one or more elements to recruit nuclear RNA binding proteins in the second RNA molecule, or both, wherein the one or more elements are not in the coding sequence for the N-terminal portion of the target protein nor in the coding sequence for the C-terminal portion of the target protein.
[0015] Also provided are compositions for expressing a nucleic acid editing protein, which include (a) a first RNA molecule comprising from 5’ to 3’: (i) a coding sequence for an N-terminal portion of the nucleic acid editing protein; (ii) a splice donor; and (iii) a first dimerization domain; and (b) a second RNA molecule comprising from 5’ to 3’: (i) a second dimerization domain, wherein the second dimerization domain binds to the first dimerization domain; (ii) a branch point sequence; (iii) a polypyrimidine tract; (iv) a splice acceptor; and (v) a coding sequence for a C-terminal portion of the nucleic acid editing protein; and (c) one or more secondary structures in the first RNA molecule, one or more secondary structures in the second RNA molecule, or both, wherein the one or more secondary structures are not in the coding sequence for the N-terminal portion of the nucleic acid editing protein nor in the coding sequence for the C-terminal portion of the nucleic acid editing protein; or one or more elements to recruit nuclear RNA binding proteins in the first RNA molecule, one or more elements to recruit nuclear RNA binding proteins in the second RNA molecule, or both, wherein the one or more elements are not in the coding sequence for the N-terminal portion of the nucleic acid editing protein nor in the coding sequence for the C-terminal portion of the nucleic acid editing protein.
[0016] In some examples, the target protein comprises two or more splice variants of the N-terminal portion of the target protein, and the first RNA molecule comprises two or more separate RNA molecules, wherein two or more separate RNA molecules each encode a distinct form of the N-terminal portion of the target protein; the target protein comprises two or more splice variants of the C-terminal portion of the target protein, and the second RNA molecule comprises two or more separate RNA molecules, wherein two or more separate RNA molecules each encode a distinct form of the C-terminal portion of the target protein; or combinations thereof.SLR 7158-102574-27 05 / 08 / 25 S2024-016A
[0017] Also provided are systems for expressing a nucleic acid editing protein comprising the described compositions.
[0018] Also provided are methods of using the disclosed systems or the RNAs encoded by the systems to express a nucleic acid editing protein in a cell, for example in combination with appropriate guide nucleic acid molecules that hybridize to a target nucleic acid molecule. Such a method can include introducing the system into a cell, and expressing the synthetic first and second RNA molecules in the same cell. In some examples, the cell is in a subject, and the method treats a disease in the subject such as a genetic disease caused by a mutation in a target DNA or RNA (e.g., gene). In some examples the genetic disease is Duchenne Muscular Dystrophy, Hemophilia A, Stargardt’s Disease, or Usher Syndrome (such as USH1F).
[0019] The foregoing and other objects and features of the disclosure will become more apparent from the following detailed description, which proceeds with reference to the accompanying figures.
[0020] BRIEF DESCRIPTION OF THE DRAWINGS
[0021] The patent or application file contains at least one drawing executed in color. Copies of this patent or patent application publication with color drawing(s) will be provided by the Office upon request and payment of the necessary fee.
[0022] FIG. 1 A depicts a schematic of vector designs (left) and RNA interactions and splicing (right). Left: 5’ trans-splice (trsp) DNA vector: Open arrows are two opposing promoters. RFP coding domain and 3’UTR with poly adenylation elements are expressed opposite from the N-terminal portion of YFP (n-yfp), followed by a splice donor sequence (SD), a downstream intronic splicing enhancer (DISE), and two intronic splicing enhancers (2xISE), a binding domain (BD, also referred to as dimerization domain), and a stable stem loop BoxB element (boxB), a self-cleaving hammerhead ribozyme (HHrz), ending with a 3’ UTR containing poly adenylation elements. The n-yfp segment has a small intron inserted (white segment within n-yfp). 3’ trsp DNA vector: Open arrows are two opposing promoters. BFP coding domain and 3’UTR with poly adenylation elements are expressed opposite from complementary binding domain (anti-BD, also referred to as dimerization domain), followed by three intronic splicing enhancer sequences (3xISE), a branch point (BP), a polypyrimidine tract (PPT), a splice acceptor sequence (SA), the c-terminal portion of the YFP coding sequence, ending with a 3’ UTR containing poly adenylation elements. Right: pre-mRNA interactions (5’ trsp-RNA + 3’ trsp-RNA) and trans-splicing to generate an mRNA encoding YFP protein are shown.
[0023] FIG. IB depicts transfection of only the N-terminal expression plasmid does not lead to YFP fluorescence.
[0024] FIG. 1C depicts transfection of only the C-terminal expression plasmid does not lead to YFP fluorescence.
[0025] FIG. ID depicts expression of N-terminal and C-terminal fragments without binding domains shows low levels of YFP induction.SLR 7158-102574-27 05 / 08 / 25 S2024-016A
[0026] FIG. IE depicts rationally designed dimerization / binding domain in a looped configuration (hypodiverse sequence consisting of either all pyrimidines or all purines that are interrupted by complementary sequences that form double stranded stem structures).
[0027] FIG. IF depicts 3D rendering of the “looped” dimerization domain configuration.
[0028] FIG. 1G depicts negative control with no binding domain on the C-terminal half.
[0029] FIG. 1H depicts negative control with no binding domain on the N-terminal half.
[0030] FIG. II depicts matching binding domains in a looped configuration on both N- and C-terminal half shows strong YFP induction in 90% of the cells.
[0031] FIGS. 1J-1N depict data equivalent to that in FIGS. 1E-1I for a configuration of a binding domain with a 150 nucleotide hypodiverse sequence comprised exclusively of pyrimidine (or alternatively exclusively purine) containing sequence resulting in a fully open configuration.
[0032] FIG 1J depicts a 150 nucleotide hypodiverse pyrimidine sequence resulting in a fully open configuration for complimentary base pairing.
[0033] FIG IK depicts a 3D rendering of the 150 nucleotide hypodiverse pyrimidine sequence from (1J). FIG 1L depicts a control HEK293T cell transfection with the C-terminal- YFP encoding construct lacking a complimentary hypodiverse binding domain. Few transfected cells express YFP.
[0034] FIG IM depicts a control HEK293T cell transfection with the N-terminal -YFP encoding construct lacking a complimentary hypodiverse binding domain. Few transfected cells express YFP.
[0035] FIG IN depicts a HEK293T cell transfection with N-terminal-YFP and C-terminal- YFP constructs that both have complimentary hypodiverse dimerization binding domains. Many cells express YFP at high levels.
[0036] FIG. 1O depicts representative fluorescence images for cells shown in FIG. 1G. The positive markers for transfection (RFP-i-BFP) are expressed, but YFP protein is not reconstituted efficiently.
[0037] FIG. 1P depicts representative fluorescence images for cells shown in FIG. 1L. The positive markers for transfection (RFP+BFP) are expressed, and YFP protein is reconstituted at high levels in cells that are both RFP and BFP double positive.
[0038] FIG. 1Q depicts a comparison of conditions shown in FIG. 1D, FIGS. 1G-1I, and FIGS. 1L-1N. N: no binding domain, Loop: looped hypodiverse binding domain configuration, Lin: linear hypodiverse configuration.
[0039] FIG. 2A depicts schematic of vector designs. The protein coding sequence of a yellow fluorescent protein (YFP) is split into an N-terminal, a middle fragment (m-yfp) and a C-terminal fragment. The junction of RNAs encoding the n and m fragments is joined by a looped design binding domain (BD1) and the junction between m and c fragments is joined by a looped binding domain (BD2). The pyrimidine (Y) and purine (R) sequences are arranged in such a way as to avoid self-circularization of the m-fragment and avoid direct recombination of the N- and C-fragment. The N-terminal fragment is co-expressed with red fluorescent protein as a transfection control, the C-terminal fragment is coexpressed with blue fluorescent protein as a transfection control. Promoter sequences are indicated with open arrows. Splice donor (SD) andSLR 7158-102574-27 05 / 08 / 25 S2024-016A
[0040] splice acceptor (SA) sites are indicated. Intronic splicing elements including splice enhancers, polypyrimidine tracts and branch points are included, analogous to the elements used upstream (5’) of the SA and downstream (3’) of the SD in FIG. 1A.
[0041] FIG. 2B depicts human cell line transfection of plasmids I+II+III (see FIG. 2A) efficiently reconstituting high level YFP expression in 80% of the transfected cells.
[0042] FIG. 2C depicts representative fluorescent image of expression of the n and m fragment (plasmid I+II, see FIG. 2A) shows no yfp fluorescence (negative control).
[0043] FIG. 2D depicts representative fluorescent image of expression of the m and c fragment (plasmid II+III, see FIG. 2A) shows no yfp fluorescence (negative control).
[0044] FIG. 2E depicts representative fluorescent image showing that strong YFP fluorescence is induced by co-transfection of all three fragments (plasmid I+II+III, see FIG. 2A).
[0045] FIGS. 3A-3D depict efficient reconstitution of yellow fluorescent protein (YFP) from two fragments (SEQ ID NOS: 1 and 2) expressed from two AAV2 / 8s after systemic administration in the newborn (P3) mouse pup. (A) depicts AAV 1 encoding the n-terminal half fragment of YFP, and AAV 2 encoding the c-terminal half fragment. AAV 1+AAV 2 were mixed at equal titer and injected intravenously into mice. Tissue sample were collected 3 weeks following injection. (B) depicts YFP fluorescence in the liver of the juvenile mouse at the time of sacrifice (green). Uninjected liver is shown for comparison (control: no YFP detected). DRAQ5 nuclear stain is shown in magenta for context. (C) depicts strong YFP fluorescence in the heart muscle at the time of sacrifice (green). Top panels show macroscopic view and red autofluorescence for context (in magenta). Bottom panel shows cross-section with DRAQ5 nuclear stain for context (in magenta). Uninjected mouse heart lacking YFP is shown for control. (D) depicts strong YFP fluorescence in the skeletal muscles of the leg at the time of sacrifice. Uninjected mouse legs are shown for comparison (negative control, no YFP detected). Top panels show macroscopic view with red autofluorescence in magenta. Bottom panel shows microscopic image of a cross-section through the leg. Bottom panel shows DRAQ5 nuclear stain in magenta for context.
[0046] FIGS. 4A-4B depict efficient reconstitution of yellow fluorescent protein (YFP) from three fragments (SEQ ID NOS: 145, 146 and 2, respectively) in the mouse tibialis anterior muscle after intramuscular injection of three AAV2 / 8 in the newborn (P3) mouse pup. (A) depicts a schematic of three AAV particles with separate N-, M-, and C-terminal fragments of YFP (analogous to Fig 2A). (B) Shows strong YFP fluorescence in a longitudinal section of the tibialis anterior muscle of a mouse injected with all three viral particles. DRAQ5 nuclear stain is shown in magenta for context.
[0047] FIGS. 5A-5F depict efficient reconstitution of yellow fluorescent protein (YFP) from two and from three fragments in adult mouse tibialis anterior muscle. (A) depicts N-terminal and C-terminal halves of YFP coding sequence are equipped with synthetic RNA-dimerization and recombination domains. (B) depicts two AAV transfer plasmids expressing these two fragments were electroporated transcutaneously into adult mouse tibialis anterior (TA) muscle and strong fluorescence was detected at 5 days post electroporation. (C) depicts no fluorescence was detectable in contralateral non-injected TA. (D) depicts n-SLR 7158-102574-27 05 / 08 / 25 S2024-016A
[0048] terminal, middle, and c-terminal YFP coding sequence are equipped with synthetic RNA-dimerization and recombination domains linking each fragment to its adjacent fragment(s). (E) depicts transcutaneous electroporation of three AAV transfer plasmids expressing these three fragments. Strong YFP fluorescence is detected indicating efficient reconstitution of YFP from three fragments. (F) depicts fluorescence in contralateral non-injected TA. Fluorescent channel is overlaid onto grey scale photographs for context.
[0049] FIG. 6A is a schematic drawing providing an exemplary system for the disclosed RNA recombination methods, using two nucleic acid molecules 110, 150, wherein the target protein is divided into two portions and each portion is encoded by a different nucleic acid molecule. In some examples, the nucleic acid molecules 110, 150, of the system are DNA, and include promoters 112, 152. In some examples, the nucleic acid molecules 110, 150, of the system are RNA, and thus lack the promoters 112, 152. Drawing not to scale.
[0050] FIG. 6B is a schematic drawing providing an exemplary dimerization domain (e.g., 122, 154 of FIG.
[0051] 6A) that includes hypodiverse sequences interspersed with sequences that can form a stem, which results in local RNA loops that are open and available for basepairing in the absence of pseudoknot formation.
[0052] Drawing not to scale.
[0053] FIG. 6C is a schematic drawing showing the interaction and hybridization (base pairing) between a pre-mRNA dimerization domain 122 of molecule 1 10 (FIG. 6A) and a pre-mRNA dimerization domain 154 of molecule 150 (FIG. 6A) allows the spliceosome components to recombine N-terminal coding sequence 114 and C-terminal coding sequence 164. This results in the 3’ end of the N-terminal protein coding sequence 114 fusing to the 5’ end of the C terminal protein sequence 164, and a seamless junction between the N- and C-terminal portions. Drawing not to scale.
[0054] FIG. 6D is a schematic drawing providing an exemplary system for the disclosed RNA recombination methods, using three nucleic acid molecules 110, 200, 150, wherein the target protein is divided into three portions (N-terminal, middle, C-terminal) and each portion is encoded by a different nucleic acid molecule. Prior to transcription, nucleic acid molecules 110, 150, 200 of the system are DNA, and include promoters 112, 152, 202. Following transcription, nucleic acid molecules 110, 150, 200 of the system are RNA, and thus lack the promoters 112, 152, 202. Drawing not to scale.
[0055] FIG. 6E is a schematic drawing showing the interaction and hybridization (base pairing) between dimerization domain 122 of molecule 110 (FIG. 6D) and dimerization domain 204 of molecule 200 (FIG 6D), and between dimerization domain 226 of molecule 200 (FIG. 6D) and dimerization domain 154 of molecule 150 (FIG 6D), allows the spliceosome components to recombine N-terminal coding sequence 114, middle coding sequence 216, and C-terminal coding sequence 164. This results in the 3’ end of the N terminal coding sequence 114 fusing to the 5’ end of the middle protein sequence 216, and the 3’ end of the middle coding sequence 216 fusing to the 5’ end of the C-terminal sequence 216, and a seamless junction between the N-, middle, and C-terminal portions. In some examples, for example following transcription, the elements shown are RNA. Drawing not to scale.SLR 7158-102574-27 05 / 08 / 25 S2024-016A
[0056] FIG. 6F is a schematic drawing providing an exemplary system for the disclosed RNA recombination methods, using two nucleic acid molecules 110, 150, wherein the target protein is divided into two portions and each portion is encoded by a different nucleic acid molecule. In this example, the DNA has been transcribed into RNA, such that nucleic acid molecules 110, 150, of the system are RNA, and thus lack the promoters 112, 152 present in the DNA (see FIG. 6A). Drawing not to scale.
[0057] FIG. 6G is a schematic drawing providing an exemplary composition or system for the disclosed RNA recombination methods, using two DNA molecules 110, 150, wherein the nucleic acid editing protein is divided into two portions 114, 164 and each portion is encoded by a different DNA molecule 110, 150 respectively. In some examples, promoters 112, 152 drive expression of each coding sequence 114, 164. The system in this example includes one or more guide nucleic acid molecules (e.g., gRNA or gRNA coding sequence) 140, 141, 171, 172 which are specific for one or more target nucleic acid molecules. However, such guide nucleic acid molecules 140, 141, 171, 172 are optional, and in some examples instead of being provided as part of molecules 110, 150, are provided separately, for example as part of a separate vector. This example shows optional guide nucleic acid molecules 140, 141, 171, 172 near the 5’ and 3’ -ends of nucleic acid molecules 110, 150, but the disclosure is not limited to such locations. If present, expression of one or more guide nucleic acid molecules 140, 141, 171, 172 can be driven by promoters 142, 143, 173, 174. The system can optionally include a parvovirus inverted terminal repeat (ITR) 176, 177, 178, 179 at each 5’ and 3’ -end of molecules 110, 150. Although this figure shows an embodiment where molecules 110, 150 are DNA, in some examples (for example following transcription), the nucleic acid molecules 110, 150 of the system are RNA, and thus lack the guide nucleic acid molecules 140, 141, 171, 172, promoters 112, 152, 142, 143, 173, 174 and parvovirus ITR 176, 177, 178, 179. Drawing not to scale.
[0058] FIG. 6H is a schematic drawing providing an exemplary system or composition for the disclosed RNA recombination methods, using three DNA molecules 110, 200, 150, wherein the nucleic acid editing protein is divided into three portions (N-terminal, middle, C-terminal, 114, 216, 154, respectively) and each portion is encoded by a different nucleic acid molecule 110, 200, 150, respectively. In some examples, promoters 112, 202, 152 drive expression of each coding sequence 114, 216, 164. The system in this example includes one or more guide nucleic acid molecules (e.g., gRNA or gRNA coding sequence) 140, 141, 231, 232, 171, 172 which are specific for one or more target nucleic acid molecules. However, such guide nucleic acid molecules 140, 141, 231, 232, 171, 172 are optional, and in some examples instead of being provided as part of molecules 110, 200, 150, are provided separately, for example as part of a separate vector. This example shows optional guide nucleic acid molecules 140, 141, 231, 232, 171, 172 near the 5’ and 3’-ends of nucleic acid molecules 110, 200, 150, but the disclosure is not limited to such locations. If present, expression of one or more guide nucleic acid molecules 140, 141, 231, 232, 171, 172 can be driven by promoters 142, 143, 233, 234, 173, 174. The system can optionally include a parvovirus inverted terminal repeat (ITR) 176, 177, 235, 236, 178, 179 at each 5’ and 3’-end of molecules 110, 200, 150.
[0059] Although this figure shows an embodiment where molecules 110, 150, 200 of the system are DNA, following transcription, nucleic acid molecules 110, 150, 200 of the system are RNA, and thus lack theSLR 7158-102574-27 05 / 08 / 25 S2024-016A
[0060] guide nucleic acid molecules 140, 141, 171, 172, 321, 232 promoters 112, 152, 202, 142, 143, 233, 234, 173, 174 and parvovirus ITR 176, 177, 235, 236, 178, 179. Drawing not to scale.
[0061] FIG. 7A is a schematic drawing providing an exemplary system for the disclosed RNA recombination methods, that like FIG. 6A uses two nucleic acid molecules 500, 600, but the dimerization domains are aptamers 512, 602, that recognize the same target molecule 700. In some examples, for example following transcription, the elements shown are RNA. Drawing not to scale.
[0062] FIG. 7B is a schematic drawing providing an exemplary system for the disclosed RNA recombination methods, that, related to FIG. 7A, uses dimerization domains that recognize the same target molecule. Here, the target recognized by the dimerization domain is a specific RNA molecule (instead of molecule 700 in FIG. 7A, e.g., protein or small molecule). Each domain recognizes a different portion of an mRNA molecule only expressed in target cells (i.e., cells where target protein expression is desired), such as a cancer-specific transcript. In some examples, for example following transcription, the elements shown are RNA. Drawing not to scale.
[0063] FIG. 7C is a schematic drawing providing an exemplary system for the disclosed RNA recombination methods, that like FIG. 6A and 7A, uses two nucleic acid molecules 800, 900, and shows the dimerization domains 812, 902 hybridizing to an oligonucleotide 1000 that prevents the dimerization domains from interacting with one another, and therefore prevents or reduces recombination of the N-terminal coding sequence 802 and C-terminal coding sequence 914. In some examples, for example following transcription, the elements shown are RNA. Drawing not to scale.
[0064] FIG. 8 is a bar graph comparing reconstitution of YFP protein expression in the presence (w / ) or absence (w / o) of a WPRE3 sequence in the 3’ untranslated region. N=3 replicates per sample are shown.
[0065] FIG. 9A is a schematic drawing providing an example for the use of dimerization domain (e.g., 122, 154 of FIG. 6A) that includes kissing loop interaction for high affinity dimerization. Using the teachings provided herein, one will appreciate that any of the disclosed coding portions (e.g., YFP) can be replaced with other target protein coding sequences. Drawing not to scale.
[0066] FIG. 9B depicts RFP, BFP, and YFP signal in HEK293T cells transfected with both halves of the split YFP. Equipped with either a linear dimerization domain adhering to the hypodiverse design principle or a structured dimerization domain designed for kissing loop-loop interactions. Strong yellow fluorescent signal indicates efficient reconstitution.
[0067] FIGS. 10A-10Z are exemplary synthetic nucleic acid molecules that can be used with the systems and methods. In some examples, a synthetic nucleic acid molecule as at least 80%, at least 85%, at least 90%, at least 95%, at least 98%, at last 99% or 100% sequence identity to the sequence of any one of SEQ ID NOS: 1 (FIGS. 10A-10B), 2 (FIGS. 10C-10E), 7 (FIG. 10E), 8 (FIG. 10F), 9 (FIG. 10G), 10 (FIG. 10H), 11 (FIG. 101), 12 (FIG. 10J), 13 (FIG. 10K), 14 (FIG. 10L), 15 (FIG. 10M), 16 (FIG. 10N), 17 (FIG. 10O), 18 (FIG. 10P), 19 (FIG. 10Q), 20 (FIGS. 10R-10U), and 21 (FIGS. 10V-10Z), but with a different target protein coding sequence. Thus an intronic region used with any of the systems or methods provided herein can have at least 80%, at least 85%, at least 90%, at least 95%, at least 98%, at last 99% or 100% sequenceSLR 7158-102574-27 05 / 08 / 25 S2024-016A
[0068] identity to any intronic sequence of SEQ ID NOS: 1, 2, 3, 4, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, or 21. For example, FIGS. 10A-D show exemplary (A, B) first (SEQ ID NO: 1) and (C, D) second (SEQ ID NO: 2) synthetic molecules that can be used to express full-length YFP, while SEQ ID NO: 3 and 4 provide the corresponding synthetic intron portion without the YFP coding portion. In some examples, a synthetic intron sequence has at least 80%, at least 85%, at least 90%, at least 95%, at least 98%, at last 99% or 100% sequence identity to SEQ ID NO: 3 or 4. Thus, the coding sequence portion of any synthetic molecule provided herein (e.g., nt 544 to 1032 of SEQ ID NO: 1 and nt 905 to 1141 of SEQ ID NO: 2), can be replaced with another coding sequence portion.
[0069] FIG. 11 is a bar graph showing the reconstitution efficiency of different length random complimentary base-pairing binding domains (50 bp, 100 bp, 150 bp, 200 bp, 300 bp, 400 bp, and 500 bp). YFP median fluorescence intensity is compared between cells with matching RFP and BFP transfection levels. n=3 samples per condition. n=3 samples per condition.
[0070] FIGS. 12A-12B show that inclusion of a splice enhancer into the synthetic intron increases the reconstitution efficiency. FIG. 12A is a schematic drawing of the 5’-N and 3’-C-terminal constructs used (SEQ ID NO: 1 and 2). (see FIG. 1A for abbreviations). FIG. 12B is a bar graph showing the resulting YFP fluorescence following transfection of SEQ ID NO: 1 and 2 into cells, or various truncations thereof. n=3 samples per condition.
[0071] FIGS. 13A-13D shows midline-crossing cortical neuron tracing by reconstitution of full-length flp recombinase (Flpo) from two fragments (SEQ ID NOS: 147 and 148). (A) Schematic representation of the 5’- and 3’ -sequences used to reconstitute flpo (analogous to constructs in Fig 12A) (B) Schematic representation of a flp-reporter mouse line injected with the N- and C-flpo encoding AAV virus injected into left and right regions of the cortex, respectively. (C and D) show neuronal cell body and axon labeling of cortical neurons that project to the contralateral hemisphere of the brain and therefore were infected by both the N-flpo and C-flpo viruses. Hoechst staining (nuclei) is shown for context.
[0072] FIGS. 14A-14D show expression of oversized cargo (i.e. proteins encoded by long RNAs) in cell culture and in vivo in the mouse primary motor cortex. (A) Schematic representation of the 5’- and 3’-sequences used to reconstitute YFP, which include long stuffer sequences (uninterrupted open reading frames; SEQ ID NOS: 22 and 23, respectively). (B) Quantitative real-time PCR analysis of reconstitution efficiency of the oversize YFP constructs in HEK 293t cells. N=3 per condition. (C) Reconstituted YFP protein expression from full-length oversized YFP expression and split-REJ expression assessed by flow cytometry of transiently transfected HEK 293t cells. Median yellow fluorescence intensity is compared between cell populations with equal transfection control (blue and red) fluorescence for the different conditions. Y-axis shows median yellow fluorescence intensity [a.u.]. N=3 per condition. (D) Schematic of injections into mouse primary motor cortex, and images of brain tissue 10 days following injection, showing successful reconstitution of a long (2401 aa) YFP protein in vivo.
[0073] FIGS. 15A-15C show efficient reconstitution of full-length human coagulation factor VIII (FVIII) with N-terminal HA tag (substituting the N-terminal signal peptide) (2317 aa). (A) Schematic representationSLR 7158-102574-27 05 / 08 / 25 S2024-016A
[0074] of the 5’- and 3’-sequences used to reconstitute FVIII (SEQ ID NOS: 24 and 25, respectively). (B) PCR amplification of the junction. (C) Western blot showing expression of FVIII. Lanes 1-3: expression of full-length FVIII (290kDa band shows full length, unprocessed FVIII). Lanes 4-6: expression of reconstituted FVIII (band at 290kDa shows successfully reconstituted FVIII). Lanes 7 and 8: expression of the N-terminus only shows absence of full-length FVIII band at 290 kDa. For all lanes: Expected proteolytic processing products are observed ranging from ~75kDa to ~210kDa. FVIII is probed for using a mouse anti-HA primary antibody. All lanes were loaded with 5micrograms of cleared cell protein extract. GAPDH (rabbit anti-GAPDH) is probed for as loading control.
[0075] FIGS. 16A-16F show efficient reconstitution of full-length human Abca4 with C-terminal FLAG-tag (2300 aa). (A) Schematic representation of the 5’- and 3’-sequences used to reconstitute Abca4 (SEQ ID NOS: 20 and 21, respectively), and a Sanger sequencing trace across the junction. (B) PCR amplification of the junction. (C) Schematic representation of the probes used to assay recombination of the 5’- and 3’-fragments. (D) PCR quantification of reconstitution efficiency after two days of expression in HEK 293t cells. N=2 per condition. (E) Western blot showing expression of Abca4. Lanes 1-3: expression of full-length Abca4 (~260kDa band shows full length Abca4). Lanes 4-6: expression of reconstituted Abca4 (band at 260kDa shows successfully reconstituted Abca4). Lanes 7 and 8: no transfection control (i.e., HEK 293t lysate only) shows absence of any signal. Abca4 is probed for using a mouse anti-FLAG primary antibody. All lanes were loaded with 5micrograms of cleared cell protein extract. GAPDH (rabbit anti-GAPDH) is probed for as loading control. (F) Quantification of the western blot in (E) normalized for differential BFP concentration. Data is shown as normalized to the average of full-length expression control.
[0076] FIGS. 17A and 17B provide (A) HIV-1 based kissing loop dimerization domain (N-fragment, SEQ ID NO: 139, C-fragment SEQ ID NO: 140); and (B) HIV-2 based kissing loop dimerization domain (N-fragment, SEQ ID NO: 141, C-fragment SEQ ID NO: 142).
[0077] FIGS. 18A-18C show efficient reconstitution of full-length murine Otof with C-terminal FLAG-tag (2019 aa). The DNA sequences of the 5’ and 3’ molecules used are shown in SEQ ID NOS: 155 and 156. (A) Western blot showing expression of Otof. Lanes 1-3: expression of full-length Otof (~250kDa band shows full length Otof). Lanes 4-6: expression of reconstituted Otof (band at 250k Da shows successfully reconstituted Otof). Lane 7: no transfection control (i.e., HEK 293t lysate only) shows absence of any signal. Otof is probed for using a mouse anti-FLAG primary antibody. All lanes were loaded with 5micrograms of cleared cell protein extract. GAPDH (rabbit anti-GAPDH) is probed for as loading control. (B) Raw quantification of the western blot and (C) normalized for differential BFP concentration. Data is shown as normalized to the average of full-length expression control.
[0078] FIGS. 19A-19C show efficient reconstitution of full-length human Myo7a with C-terminal FLAG-tag (2243 aa). The DNA sequences of the 5’ and 3’ molecules used are shown in SEQ ID NOS: 157 and 158. (A) Western blot showing expression of Myo7a. Lanes 1-3: expression of full-length Myo7a (~270kDa band shows full length Myo7a). Lanes 4-6: expression of reconstituted Myo7a (band at 270k Da shows successfully reconstituted Myo7a). Lane 7: no transfection control (i.e., HEK 293t lysate only) showsSLR 7158-102574-27 05 / 08 / 25 S2024-016A
[0079] absence of any signal. Myo7a is probed for using a mouse anti-FLAG primary antibody. All lanes were loaded with 5micrograms of cleared cell protein extract. GAPDH (rabbit anti-GAPDH) is probed for as loading control. (B) Raw quantification of the western blot and (C) normalized for differential BFP concentration. Data is shown as normalized to the average of full-length expression control.
[0080] FIGS. 20A-20D show efficient reconstitution of full-length DCas9-VPR (1951 aa). The DNA sequences of the 5' and 3’ molecules used are shown in SEQ ID NOS: 159 and 160. (A) Western blot showing expression of DCas9-VPR. Lanes 1-3: expression of full-length DCas9-VPR (~250kDa band shows full length DCas9-VPR). Lanes 4-6: expression of reconstituted DCas9-VPR (band at 250kDa shows successfully reconstituted DCas9-VPR). Lane 7: no transfection control (i.e.. HEK 293t lysate only) shows absence of any signal. DCas9-VPR is probed for using a mouse anti-Cas9 primary antibody. All lanes were loaded with 5micrograms of cleared cell protein extract. GAPDH (rabbit anti-GAPDH) is probed for as loading control. (B) Raw quantification of the western blot and (C) normalized for differential BFP concentration. Data is shown as normalized to the average of full-length expression control. (D) Example of transcriptional activation of a YFP expressing plasmid in HEK 293t cells. Full-length (upper panels) or two-way split REJ-dual dCas9-VPR (lower panels) is transiently transfected together with non-targeting guide RNA (left panels) or UAS-targeting guide RNA (right panels) expressing plasmids. All cells are also transfected with a UAS-YFP plasmid that is transcriptionally inactive until dCas9-VPR is targeted to the upstream region of a minimal promoter which results in expression of yellow fluorescent protein. Red fluorescent protein (RFP) is expressed with the N-terminal fragment of dCas9-VPR, Blue fluorescent protein (BFP) is expressed with the full-length dCas9-VPR or the C-terminal fragment of dCas9-VPR, respectively. RFP and BFP serve as transfection control. Upon expression of both full-length as well as two-way split dCas9-VPR paired with a UAS-targeting guide RNA, yellow fluorescent protein expression is observed, confirming functionality of the reconstituted full-length protein.
[0081] FIGS. 21A-21D show efficient reconstitution of full-length humanized Prime Editor (2118 aa). The DNA sequences of the 5’ and 3’ molecules used are shown in SEQ ID NOS: 161 and 162. (A) Western blot showing expression of Prime Editor. Lanes 1-3: expression of full-length Prime Editor (~260kDa band shows full length Prime Editor). Lanes 4-6: expression of reconstituted Prime Editor (band at 260k Da shows successfully reconstituted Prime Editor). Lane 7: no transfection control (i.e., HEK 293t lysate only) shows absence of any signal. Prime Editor is probed for using a mouse anti-Cas9 primary antibody. All lanes were loaded with 5micrograms of cleared cell protein extract. GAPDH (rabbit anti-GAPDH) is probed for as loading control. (B) Raw quantification of the western blot and (C) normalized for differential BFP concentration. Data is shown as normalized to the average of full-length expression control. (D) Shows Prime Editor induced G to T transversion mutations induced in the FANCF and the VEGFA3 loci of HEK293t cells. The top panel shows the sequence context for the FANCF and VEGFA3 loci respectively. The grey arrow indicates the sequence targeted by the prime editor guide RNA (pegRNA). The protospacer adjacent motif (PAM) is indicated with a grey box. The G that is targeted for transversion to T is highlighted in the sequence. Genomic loci are sequenced using Sanger sequence in three conditions. The topSLR 7158-102574-27 05 / 08 / 25 S2024-016A
[0082] panel shows a representative sanger trace for unedited wild type condition. The second from the top panel shows a representative sanger trace that represents the full-length expressed prime editor construct. The area highlighted with the black box shows the appearance of a T band in the sanger sequence, indicative of successful incorporation of the edit in a portion of the cells. The lowest panels show representative sanger traces for cells edited with a two-way split reconstituted prime editor. The appearance of a T trace (black box) demonstrates functionality of the prime editor when reconstituted from two fragments.
[0083] FIGS. 22A-22C show efficient reconstitution of full-length humanized Cytosine Base Editor (AncBE4) (1854 aa). The DNA sequences of the 5’ and 3’ molecules used are shown in SEQ ID NOS: 163 and 164. (A) Western blot showing expression of AncBE4. Lanes 1-3: expression of full-length AncBE4 (~230kDa band shows full length AncBE4). Lanes 4-6: expression of reconstituted AncBE4 (band at 230k Da shows successfully reconstituted AncBE4). Lane 7: no transfection control (i.e.. HEK 293t lysate only) shows absence of any signal. AncBE4 is probed for using a mouse anti-Cas9 primary antibody. All lanes were loaded with 5micrograms of cleared cell protein extract. GAPDH (rabbit anti-GAPDH) is probed for as loading control. (B) Raw quantification of the western blot. Data is shown as normalized to the average of full-length expression control. (C) Shows AncBE4 induced C to T transition mutations induced in the EMX1 and the HEK site 3 loci of HEK293t cells. The top panel shows the sequence context for the EMX1 and HEK site 3 loci respectively. The grey arrow indicates the sequence targeted by the AncBE4 gRNA. The protospacer adjacent motif (PAM) is indicated with a grey box. The Cs that are targeted for transition to T are highlighted in the sequence. Genomic loci are sequenced using Sanger sequence in three conditions. The top panel shows a representative sanger trace for unedited wild type condition. The second from the top panel shows a representative sanger trace that represents the full-length expressed AncBE4 construct. The area highlighted with the black box shows the appearance of a T band in the sanger sequence, indicative of successful incorporation of the edit in a portion of the cells. The lowest panels show representative sanger traces for cells edited with a two-way split reconstituted AncBE4. The appearance of a T trace (black box) demonstrates functionality of the AncBE4 when reconstituted from two fragments.
[0084] FIGS. 23A-23C show efficient reconstitution of full-length humanized Adenine Base Editor (ABE8e) (1606 aa) (e.g., SEQ ID NOS: 225, 226). The DNA sequences of the 5’ and 3’ molecules used are shown in SEQ ID NOS: 165 and 166. (A) Western blot showing expression of ABE8e. Lanes 1-3: expression of full-length ABE8e (~230kDa band shows full length ABE8e). Lanes 4-6: expression of reconstituted ABE8e (band at 230k Da shows successfully reconstituted ABE8e). Lane 7: no transfection control (i.e., HEK 293t lysate only) shows absence of any signal. ABE8e is probed for using a mouse anti-Cas9 primary antibody. All lanes were loaded with 5micrograms of cleared cell protein extract. GAPDH (rabbit anti-GAPDH) is probed for as loading control. (B) Raw quantification of the western blot. Data is shown as normalized to the average of full-length expression control. (C) Shows ABE8e induced A to G transition mutations induced in the BCL11 A and the HGB1 / 2 loci of HEK293t cells. The top panel shows the sequence context for the BCL11A and HGB1 / 2 loci respectively. The grey arrow indicates the sequence targeted by the ABE8e guide RNA (gRNA). The protospacer adjacent motif (PAM) is indicated with a greySLR 7158-102574-27 05 / 08 / 25 S2024-016A
[0085] box. The As that are targeted for transition to G are highlighted in the sequence. Genomic loci are sequenced using Sanger sequence in three conditions. The top panel shows a representative sanger trace for unedited wild type condition. The second from the top panel shows a representative sanger trace that represents the full-length expressed ABE8e construct. The area highlighted with the black box shows the appearance of a G band in the sanger sequence, indicative of successful incorporation of the edit in a portion of the cells. The lowest panels show representative sanger traces for cells edited with a two-way split reconstituted ABE8e. The appearance of a G trace (black box) demonstrates functionality of the ABE8e when reconstituted from two fragments.
[0086] FIGS. 23D-23G show efficient reconstitution of full-length humanized Adenine Base Editor (ABE8e) (1606 aa) to correct a premature stop codon in the mdx mouse model for Duchenne muscular dystrophy. (D) Schematic drawing of the experiment. The ABE8e base editor is split into two fragments, the coding DNA of which are individually packaged into two separate adeno-associated virus capsids. CRISPR gRNAs are designed to target the locus of the premature stop-codon such that the A to G conversion converts the TAA stop codon into a CAA codon. (E) A yellow fluorescent protein (YFP) is split into two coding fragments that are connected with a short stretch of the mdx coding sequence surrounding the premature stop codon. The first half of the YFP is non-fluorescent if the open reading frame is terminated by the premature TAA stop codon. Stop codon correction results in translation of the full YFP sequence which is rendered fluorescent (this construct is referred to as YFP-editing-reporter). The left panel shows absence of YFP fluorescence when HEK293T cells were transfected with the YFP-editing-reporter (which co-expresses a red fluorescent transfection control), and the N-terminal ABE8e vector, and the C-terminal ABE8e vector, and a non-targeting gRNA. The right panel shows expression of YFP fluorescence in a high percentage of cells that are cotransfected with the YFP-editing-reporter (which co-expresses a red fluorescent transfection control), and the N-terminal ABE8e vector, and the C-terminal ABE8e vector, and a mdx locus targeting gRNA. (F) Shows in vivo editing of the mdx premature stop codon resulting in expression of dystrophin in treated muscle. A male mdx mutant mouse was injected with a mixture of N-and C-terminal ABE8e packaged in adeno-associated virus 8 vectors (5E10 viral genomes per vector for a total of 1 x 10” viral genomes per muscle). The virus genomes contain two gRNA expression cassettes (composed of a RNA polymerase III promoter and the gRNA sequence) in each of the two genomes. The virus mix was injected intra muscularly into the tibialis anterior muscle. Top right panel shows dystrophin staining in a wild-type tibialis anterior muscle cross section for reference. The bottom left panel shows untreated tibialis anterior muscle tissue. The bottom right panel shows expression of dystrophin in a tibialis anterior muscle that was injected with the two adeno-associated viruses. (G) Shows dystrophin expression (top panel), ABE8e expression (middle panel), and a GAPDH loading control (bottom panel) of tibialis anterior muscle treated with the adeno-associated virus mix for expression of ABE8e. This illustrates the rescue of dystrophin expression using the ABE8e base editor.
[0087] FIGS. 24A-24C Influence of downstream intronic splicing enhancers (DISE) and intronic splicing enhancers (ISE) and acceptor sequences on the efficiency of RNA end joining. (A) Schematic depiction ofSLR 7158-102574-27 05 / 08 / 25 S2024-016A
[0088] screen setup. The 5’ fragment is an RNA molecule which is transcribed from a DNA construct using the human CMV promoter and enhancer. The RNA molecule produced contains a long stuffer open reading frame to simulate large cargo size. This stuffer sequence ends in a 2A self-cleaving peptide sequence and is followed by the coding region for the 5’ fragment of a Yellow Fluorescent Protein (n-yfp). The 5’ fragment of yfp ends in a splice donor site (SD). This splice donor site is followed by the 5’ intronic portion of the RNA end joining module. For the purpose of determining the impact of DISE and ISE sequences on RNA end joining reaction efficiency, the 5’ intronic portion is subdivided into three fragments: from 5’ to 3’: ds: downstream segment; m: mid intronic segment; dd: donor distal segment. The 5’ intronic portion is followed by a trimodal kissing loop RNA dimerization domain. The message is terminated with a short poly adenylation signal. The overall length of this 5’ RNA molecule is ~4kb to simulate a large cargo reconstitution scenario. The 3’ fragment is an RNA molecule which is transcribed from a DNA construct using the human CMV promoter and enhancer. The 3’ fragment starts with a trimodal kissing loop RNA dimerization domain that is complementary to the one on the 5’ fragment encoding RNA molecule. The dimerization domain is followed by the 3’ intronic portion of the RNA end joining module. This 3’ intronic portion is subdivided into three segments: ad: acceptor distal segment; m: mid-intronic segment; ap: acceptor proximal segment. The acceptor proximal segment contains variations of the branch point and polypyrimidine tracts which are both essential for the spliceosome mediated RNA joining reaction. The splice acceptor (SA) site is followed by the 3’ yfp coding sequence which is followed by a self-cleaving 2A sequence that is followed by a long stuffer open reading frame. The message is terminated by an SV40 poly adenylation signal. The overall length of the 3’ RNA molecule is ~4kb to simulate a large cargo reconstitution scenario. The association of the two RNA molecules (the 5’ fragment and the 3’ fragment) is mediated by the trimodal kissing loop RNA dimerization domain, the recruitment of the spliceosome and the RNA end joining reaction are mediated by the intronic segments. Successful RNA end joining results in reconstitution of the yfp open reading frame and subsequent translation of YFP. (B) Median YFP fluorescence intensity as determined by flow cytometry is shown for a number of intron configurations. In the first grouping (bars 1 to 9) a selection of potential downstream intronic splicing enhancer sequences were paired with a consensus splice donor site (GTAAGTATT in the DNA construct and GUAAGUAUU in the RNA sequence), shown in bars 1-8. These are compared to a consensus splice donor that is followed by scrambled sequence composed of equal parts of all four bases (ds9). In the second grouping, ml-ml6 a selection of potential intronic splicing enhancers was compared to a scrambled sequence (ml6). In the last grouping a selection of potential strong branch point, polypyrimidine tract, and splice acceptors were compared. The reference constructs were composed of scrambled sequence in all non-variable positions with a consensus donor followed by scrambled sequence in the ds position and a consensus splice acceptor sequence (where the whole polypyrimidine tract is composed of Ts in the DNA construct and Us in the RNA fragment respectively). (C) Listing of different DISE, ISE, and splice acceptor elements used.
[0089] FIGS. 25A-25B show the in vivo expression of a base editor in mouse muscle, and the correction of a dystrophin gene mutation using the disclosed methods. (A) Western blots showing expression ofSLR 7158-102574-27 05 / 08 / 25 S2024-016A
[0090] dystrophin, Cas9-ABE, and GAPDH in muscle extracts from wildtype (wt), untreated mdx-4cv mice, and treated mdx-4cv mice. (B) Immunohistochemistry analysis of dystrophin expression in muscle in wildtype (wt), untreated mdx-4cv mice, and treated mdx-4cv mice.
[0091] FIG. 26 provides a schematic of RNA end joining (REJ) methods provided herein. In some examples, splitting a coding sequence into two or more fragments, results in the production of both full-length protein and unspliced REJ RNAs (un-joined fragments). The un-joined fragments can leave the nucleus and enter the cytosol. The present application provides exemplary methods to reduce the expression of, or increase the degradation of, un-joined fragments.
[0092] FIG. 27 provides an exemplary method, ribosome assembly inhibition, that can be used to reduce the expression of unspliced REJ RNAs. In this method, secondary structure elements are introduced within the REJ intron. Such secondary structure conditionally inhibits ribosome assembly, scanning, and translation when the RNA is unspliced. Incorporation of hairpins, kissing stem loops, and G-Quadruplex structures are shown.
[0093] FIG. 28 provides an exemplary method, nuclear retention of RNA, that can be used to reduce the expression of unspliced REJ RNAs. In this method, addition of sequence elements (such as a SIRLOIN element, SEQ ID NO: 227) to recruit nuclear RNA binding proteins within the REJ intron conditionally prevents nuclear export of the unspliced REJ RNA. These nuclear retained RNAs can subsequently undergo splicing and will lose the nuclear retention element, allowing export to the cytoplasm for translation.
[0094] FIG. 29 is a digital image showing expression of C-terminal minidystrophin and unspliced protein fragment visualized from western blot across varied fragment suppression designs. Fragment suppression designs tested were Linear (Lin), 3x Kissing Stem Loops (KSL), or G-Quadruplex with Kissing Stem Loops (G: KSL), either delivered in 1:1 or 3: 1 N: C plasmid ratios. Designs and delivery ratios vary in their minidystrophin and fragment expression.
[0095] FIG. 30 is a bar graph showing expression of C-terminal minidystrophin and unspliced protein fragment quantified from western blot across varied fragment suppression designs. The stoichiometric ratio of N: C terminal (Delivery Ratio) can be tuned to increase target protein expression while reducing fragments. Kissing Stem Loops (KSL) provide the highest protein expression level while minimizing fragments when REJ RNAs are delivered in a 3:1 N: C ratio.
[0096] FIG. 31 is a schematic drawing showing two PCDH15 constructs (SEQ ID NOS: 230, 231) that can be used to express PCDH15. Relevant REJ intronic features, such as Intronic Splice Enhancers (ISE), Branch Point (BP), Poly Proline Tract (PPT), Splice Donor (SD) and Splice Acceptor (SA) are shown.
[0097] FIG. 32 is a schematic drawing showing relative positioning of different RNA dimerization domains.
[0098] FIGS. 33A-33E provide exemplary sequences for the structures shown in FIG. 32: (A) Proximall or Prxl (SEQ ID NO: 237); (B) Proximal2 or Prx2 (SEQ ID NO: 238); (C) Distall or Dstl (SEQ ID NO: 239); (D) Distal2 or Dst2 (SEQ ID NO: 240); and (E) Doublel or Doul (SEQ ID NO: 241) with predicted secondary structure depicted.SLR 7158-102574-27 05 / 08 / 25 S2024-016A
[0099] FIG. 34 is a western blot image showing production of PCDH15 using REJ with different RNA dimerization domains shown in FIG. 32.
[0100] FIG. 35 is a bar graph showing quantification of PCDH15 expression of the western blot in FIG. 34. FIG. 36 is a schematic showing an exemplary plasmid encoding the N-terminal REJ construct for expressing PCDH15 (using the Dst2 dimerization domain, e.g., see FIG. 33D and SEQ ID NO: 240, as an example).
[0101] FIG. 37 is a schematic showing an exemplary plasmid encoding the C-terminal REJ construct for expressing PCDH15 (using the Dst2 dimerization domain, e.g., see FIG. 33D and SEQ ID NO: 240, as an example).
[0102] FIG. 38 shows the sequencing results confirming that the N-terminal PCDH15 sequence is joined with the C-terminal PCDH15 sequence.
[0103] FIGS. 39A-39B are schematic drawings showing combinatorial splicing of multiple N- or C- REJ constructs that can be used to produce multiple isoforms of a protein. This combinatorial splicing can be used with PCDH15 for creating both the CD1 and CD2 isoforms (39A).
[0104] FIGS 40A-40B show combination of one N-terminal construct (encoding N-PCDH15) and two C-terminal constructs (encoding C-PCDH15 and YFP respectively) results in simultaneous production of two proteins: full length PCDH15, and N-PCDH15-YFP chimeric protein. FIG. 40A is a representative western blot image; FIG. 40B is a bar graph showing quantification of the western blot.
[0105] FIG. 41A is a schematic drawing showing an exemplary method for expressing PCDH15 in mouse retinas.
[0106] FIG. 41B is a schematic showing an exemplary plasmid encoding an N-terminal construct for expressing human PCDH15 CD1 in mouse retinas.
[0107] FIG. 41C is a schematic showing an exemplary plasmid encoding a C-terminal construct for expressing human PCDH15 CD1 in mouse retinas.
[0108] FIGS 42A-42B show that full-length PCDH15 CD1 was expressed in mouse retinas. FIG. 42 A is a representative western blot image; FIG. 42B is a bar graph showing quantification of the western blot.
[0109] FIG. 43A is a schematic showing an exemplary plasmid encoding an N-terminal construct for expressing mouse PCDH15 CD2 in the ear of mouse.
[0110] FIG. 43B is a schematic showing an exemplary plasmid encoding an C-terminal construct for expressing mouse PCDH15 CD2 in the ear of mouse.
[0111] FIG. 44 are line graphs showing auditory brainstem response (ABR) of PCDH15 knockout mice, demonstrating recovery of hearing in mice treated with disclosed REJ vectors.
[0112] FIG. 45 is a western blot image showing expression of mini -dystrophin at high levels using the disclosed a dual-vector REJ delivery system (SEQ ID NOS: 228, 229).
[0113] FIG. 46 is a graph showing treatment of a mouse model of Duchenne muscular dystrophy using the dual-vector AAV8 REJ mini-dystrophin expression system. An in situ muscle force measurement assay wasSLR 7158-102574-27 05 / 08 / 25 S2024-016A
[0114] performed on tibialis anterior muscles. Treated mice were significantly stronger than untreated mice and not statistically different than wildtype mice.
[0115] FIGS. 47A-47E: AAV cargo limitations necessitate multi-vector strategies for DMD. FIG. 47A: The 4.7 kb cargo capacity of AAV precludes the transfer of large therapeutic proteins such as some CRISPR effectors and the native dystrophin sequence. FIG. 47B: Engineered variants of DMD proteins contain varying domains of the native protein. The AH2-R17 mini-dystrophin used herein contains the full N and C terminal domains and additionally the R18-R19 domains absent in asymptomatic AH2-R19 BMD patients. AlphaFold2 demonstrates the predicted protein structures of the ABD, hinges, rods, and CT domains of a micro-dystrophin. FIG. 48C: The dose-expression relationship of micro-dystrophin was modeled as a stochastic Poisson process. The model predicts high multiplicities of infection within conventional AAV dose ranges and a ~2.6x increase in dose to achieve 90% dual-infection compared to 90% single-infection. FIGS. 48D-48E: Using REJ, the oversized Abe8e Adenine base editor or mini-dystrophin can be delivered via dual AAVs to restore or replace endogenous dystrophin expression. Appended REJ modules colocalize transcribed RNA through RNA: RNA dimerization domains, recruit the endogenous spliceosome through intronic splice enhancers, and catalyze an efficient and scar-free trans-splicing reaction, producing full-length mRNA for translation.
[0116] FIGS 48A-48J: REJ construct design and protein expression validation. FIG. 48 A: Abe8e coding sequence was split into two trans-splicing exons with kissing stem loop RNA binding domains and the REJ intron modules, containing splice acceptor and donor and splice enhancers. The AAVs additionally carry three copies of the guide RNA. FIG. 48B: Mini-dystrophin coding sequence was similarly split with additional stimulatory introns (int.) installed to enhance splicing. Linear dimerization domains were used. FIG. 48C: HEK293T cells were transfected with either full-length or REJ Abe8e plasmids and either a guide RNA targeting an early stop codon in a YFP reporter or a non-targeting guide RNA. The targeting guides enabled restored YFP expression, demonstrating functional efficacy of the REJ Abe8e. FIG. 48D: HEK293T cells transfected with full or split mini-dystrophin show robust expression of the full-length minidystrophin (253 kDa). FIG. 48E: Densitometry shows high levels of mini-dystrophin expression (148% of single plasmid controls). FIG. 48F: Intramuscular treatment with dual AAVs shows robust Abe8e expression and resulting restoration of the full-length dystrophin protein in B6 DMD mice. FIG. 48G: Intramuscular treatment with mini-dystrophin in D2 DMD mice shows high levels of protein expression. FIG. 48H: Abe8e treated mice exhibit 19.3% of wildtype dystrophin levels compared to 0.46% in untreated mice (p=0.0160). FIG. 481: Mini-dystrophin treated mice express 71.6% of wildtype dystrophin levels compared to 1.7% in untreated mice (p=0.0003). FIG. 48J: The severe D2 DMD mice developed macroscopic plaques of damaged muscle which were not observed in mini-dystrophin treated mice.
[0117] FIGS 49A-49G: An advanced computer vision and machine learning framework enables histopathological imaging informatics. FIG. 49A-49B: Abe8e (A) and mini-dystrophin (B) treated mice display restored proper localization of restored dystrophin expression to the myofiber membrane (laminin-a2 in green, Hoescht nuclear stain in blue, dystrophin in red). FIG. 49C: An advanced computer vision andSLR 7158-102574-27 05 / 08 / 25 S2024-016A
[0118] machine learning histopathology pipeline enables robust statistical phenotyping across hundreds of thousands of individual myofibers. Using deep-learning models, individual myofibers are segmented. Phenotypes are quantified by computer vision (gray) and machine learning implementations. FIGS. 49D and 49F: Abe8e treated mice exhibited improvements in centro-nucleation (wildtype 1.7%, untreated 15.7%, treated 8.8%, p=0.0008 treated vs. untreated). Hypertrophy was improved to wildtype levels (wildtype 2.9%, untreated 12.9%, treated 3.0%, p<0.0001 treated vs. untreated). Both phenotypes were negatively associated with dystrophin level (both p<0.0001, top right). Data from n=8-11 mice per group with >30,000 myofibers per group. FIGS. 49E and 49G: Mini-dystrophin treated mice showed improved centro-nucleation (wildtype 2.1%, untreated 7.2%, treated 3.25%, p=0.0094 treated vs. untreated) and nuclear infiltration (wildtype 0.06%, untreated 14.2%, treated 2.1%, p=0.0001 treated vs. untreated). Both phenotypes were negatively associated with dystrophin level (both p<0.0001, top right). Data from n=4-7 mice per group with >7,500 myofibers per group.
[0119] FIGS 50A-50F: Machine learning enables systematic dystrophin classification, myofiber typicality assessment, and reveals spatial interaction cooperativity in Abe8e treated mice. FIG. 50A: dystrophin+ / -machine learning model was trained and validated on a 2,300 myofiber human annotated dataset (ROC AUC 0.86±0.15, PR AUC 0.85+0.02, 95% CI) and applied to the entire Abe8e dataset. This showed that 45.5% of treated fibers expressed dystrophin compared to 1 % in untreated and 89.2% in wildtype (p<0.0001, treated vs. untreated). FIG. 50B: A machine-learned myofiber typicality classifier was created using wildtype and untreated myofibers as reference populations. Myofiber typicality was negatively associated with centro-nucleation and hypertrophy across all fibers and each condition. FIG. 50C: Significantly higher percentage of treated fibers were classified as typical compared to untreated fibers (81.9% vs 70.0%, p=0.0135). FIG. 50D: To study spatial interactions in the edited mice, which displayed expression moasiacism, for each myofiber of interest (red), physical neighbors (white) were identified within the reconstructed spatial network mesh (blue). For visualization, this network is overlaid on laminin-a2 immunofluorescence (green). FIG. 50E: Spatial network analysis revealed a positive correlation between the percentage of dystrophin+ neighbors and typicality probability in dystrophin- fibers (right, rho=0.0405, p<0.0001). This correlation was not present in edited, dystrophin+ fibers (left, rho=0.01, p=0.2266). Null hypotheses and confidence intervals are shown in red. FIG. 50F: Mosiacism.
[0120] FIGS 51A-51G: Transcriptomic and physiological efficacy. FIG. 51 A: Expression levels of known DMD-associated genes were assessed by RNA-seq. Abe8e treated mice expression was improved in 21 of 29 genes (p=0.0242). FIG. 51B: Mini-dystrophin treated mice exhibited expression improvement in 24 of 29 genes (p<0.0001). FIG. 51C: Whole-transcriptome principal component analysis revealed strong ontological differences between wildtype and DMD mice, separated primarily by muscle metabolism and immune activation ontologies. Abe8e treated mice transcriptomes clustered closer to those of wildtype transcriptomes, indicating overall improvements in global gene expression signatures (p=0.0119). FIG. 51D: Mini-dystrophin treated mice and controls revealed a different strain specific disease-associated transcriptomic signatures, encompassing muscle metabolism and specifically phagocytosis and interferonSLR 7158-102574-27 05 / 08 / 25 S2024-016A
[0121] regulation. These signatures were improved by treatment (p=0.0468). FIG. 51E: Bioinformatically quantified unspliced, on-target splicing, and endogenous off-target splicing events in mini-dystrophin treated mice. The 5’ and 3’ RNA had on-target frequencies of 78% and 9% respectively. Endogenous off-target splicing was not detected and is not statistically distinguishable from 0% (95%CI 0.00-0.19%). FIG. 51F: Abe8e treated mice produced higher specific force (sPo) than untreated mice (177.0 vs 128.2 mN / mm2, p=0.0164) but not to the level of wildtype mice (232.9 mN / mm2) at 250 Hz. n=8-15 per group. FIG. 51G: Mini-dystrophin treated mice produced higher specific force than untreated mice (198.3 vs 153.8 mN / mm2, p=0.0013) and were not significantly different than wildtype mice (224.5 mN / mm2, p=0.0762) at 250 Hz. n=12-15 per group.
[0122] FIGS. 52A-52F: Intravenous treatment with myotropic AAV restores systemic phenotypes: FIG.
[0123] 52A: Intravenous delivery of Abe8e via AAVmyo resulted in widespread editing and dystrophin restoration body wide, spanning triceps, tibialis anterior, heart, and diaphragm. FIG. 52B: Similar delivery of minidystrophin resulted in robust expression across the same muscle groups. FIG. 52C: Abe8e treated mice exhibited higher specific force (sPo) than untreated mice (173.8 vs 128.2 mN / mm2, p=0.024) but not to wildtype levels (232.9mN / mm2). n=8-11 per group. FIG. 52D: Aged mini-dystrophin treated mice body weight was improved from untreated weight (28.7g vs 25.5g, p=0.0027) and to wildtype levels (31.45g, p=0.1366). n=10-11 per group. FIG. 52E: These mice exhibited improved Po force generation compared to untreated mice (1058.7 mN vs 731.6 mN, p<0.0001), but not to wildtype levels (1227.5 mN, p<0.0001). n=9-ll per group. FIG. 52F: On the voluntary running wheel assay, treated mice ran significantly farther than untreated mice (17.4 km vs 7.4 km, p=0.0443), similar to wildtype mice (17.4 km, p=0.2305). Night and day cycles are shaded gray and white respectively. n=3-5 per group.
[0124] FIG. 53A: Schematic drawing showing the REJ constructs used for testing different dimerization domains.
[0125] FIGS. 53B-53C: Mean fluorescence intensity (MFI) for the random dimerization domains tested (see Table 6).
[0126] FIG. 54: MFI for the kissing dimerization domains tested (see Table 7).
[0127] FIGS. 55A-55B: MFI for the hybrid dimerization domains tested (see Table 9).
[0128] FIG. 56: Graph showing normalized protein expression for different transgenes utilizing different dimerization domains.
[0129] FIG. 57: A bar graph showing relative fluorescence intensity for different YFP REJ constructs including a stimulatory intron in the 5’ RNA fragment, the 3’ RNA fragment, or both, “n” refers to no stimulatory intron, “int” refers to inclusion of stimulatory intron. The first notation refers to the 5’ RNA, and the second refers to the 3’ RNA.
[0130] FIG. 58: A western blot image comparing mini-dystrophin production of variants of the 3’ minidystrophin RNA with and without the cTNT sequence, which contains a cryptic splice site. The first two lanes have the original 5’ and 3’ mini-dystrophin REJ constructs transfected, showing the unwantedSLR 7158-102574-27 05 / 08 / 25 S2024-016A
[0131] expression of C-terminal fragment expression (visualized using an anti-C-dystrophin antibody). When the cTNT is removed from the 3’ RNA, expression of the fragment is significantly reduced.
[0132] SEQUENCE LISTING
[0133] The nucleic and amino acid sequences listed in the accompanying sequence listing are shown using standard letter abbreviations for nucleotide bases, and three letter code for amino acids, as defined in 37 C. F. R. 1.822. Only one strand of each nucleic acid sequence is shown, but the complementary strand is understood as included by any reference to the displayed strand. Both DNA and RNA are included by the nucleic acid sequence shown (RNA sequence is obtained from DNA sequence by replacing “T” with “U”, and vice versa).
[0134] SEQ ID NOS: 1 and 2 are N- and C-terminal sequences, respectively, used to express full-length YFP. SEQ ID NO: 1, CMV promoter nt 1 to 543, YFP coding sequence nt 544 to 1032, synthetic intron nt 1033 to 1436, and untranslated poly A region nt 1437 tol491. SEQ ID NO: 2, CMV promoter nt 1 to 522, synthetic intron nt 523 to 904, YFP coding sequence nt 905 to 1141, and nt 1142 to 1302 is the untranslated poly A region.
[0135] SEQ ID NOS: 3 and 4 are 5’- and 3’-intronic sequences, respectively, that can be used to express a desired full-length protein, wherein a N-terminal portion of the full-length protein can be added at nt 1 of SEQ ID NO: 3, and C-terminal portion of the full-length protein can be added at nt 382 of SEQ ID NO: 4.
[0136] SEQ ID NOS: 5 and 6 are N- and C-terminal coding sequences, respectively, used to express full-length YFP.
[0137] SEQ ID NO: 7 is an exemplary synthetic intron dimerization domain (FIG. 10E).
[0138] SEQ ID NO: 8 is an exemplary synthetic intron without intronic splicing enhancers (FIG. 10F). SEQ ID NO: 9 is an exemplary synthetic intron without intronic splicing enhancers (FIG. 10G). SEQ ID NO: 10 is an exemplary synthetic intron without intronic splicing enhancers (FIG. 10H).
[0139] SEQ ID NO: 11 is an exemplary synthetic intron without binding domain (FIG. 101).
[0140] SEQ ID NO: 12 is an exemplary synthetic intron with dimerization domain (FIG. 10J).
[0141] SEQ ID NO: 13 is an exemplary synthetic intron with dimerization domain (FIG. 10K).
[0142] SEQ ID NO: 14 is an exemplary synthetic intron without intronic splicing enhancers (FIG. 10L).
[0143] SEQ ID NO: 15 is an exemplary synthetic intron with DISE only (FIG. 10M).
[0144] SEQ ID NO: 16 is an exemplary synthetic intron without HHrz (FIG. 10N).
[0145] SEQ ID NO: 17 is an exemplary synthetic intron without intronic splicing enhancers (FIG. 10O).
[0146] SEQ ID NO: 18 is an exemplary U12 dependent intron with binding domain (FIG. 10P).
[0147] SEQ ID NO: 19 is an exemplary U12 dependent intron with binding domain (FIG. 10Q).
[0148] SEQ ID NOS: 20 and 21 are the N- and C-terminal DNA sequences, respectively, used to express RNAs (pre-mRNAs) resulting in full-length Abca4. In SEQ ID NO: 20, the sequence corresponding to the N-terminal Abca4 coding region is at nt 22 to 3702, and nt 3703 to 3912 is the synthetic intron, and 3921 to 3969 is the untranslated poly A region. SEQ ID NO: 20 also comprises a splice donor at nt 3703-3711, a RatSLR 7158-102574-27 05 / 08 / 25 S2024-016A
[0149] FGFR2 DISE at nt 3714-3737, a cTNT intronic splicing enhancer at nt 3747-3770, an M2 intronic splicing enhancer at nt 3782-3794, and a kissing loop dimerization domain at nt 3801-3975. In SEQ ID NO: 21, nt 1 to 228 is the synthetic intron, nt 229 to 3366 is the C-terminal Abca4 coding region, 3367 to 3447 is the FLAG epitope tag, and nt 3476 to 3607 is the untranslated poly A region (signal). SEQ ID NO: 21 also comprises a kissing loop dimerization domain at nt 3-114, an M2 intronic splicing enhancer at nt 121-133, a cTNT intronic splicing enhancer at nt 140-163, an M2 intronic splicing enhancer at nt 175-187, a Branch Point Motif at nt 194-201, a poly pyrimidine tract at nt 207-226, and a splice acceptor at nt 228.
[0150] SEQ ID NOS: 22 and 23 are the N- and C-terminal DNA sequences, respectively, used to express RNAs (pre-mRNAs) resulting in a long full-length YFP, wherein each includes splice enhancers. In SEQ ID NO: 22, the N-terminal YFP coding region is nt 22 to 3702, nt 3703 to 3912 is the synthetic intron, and 3921 to 3969 is the untranslated poly A region. SEQ ID NO: 22 also comprises a splice donor at nt 3703-3711, a Rat FGFR2 DISE at nt 3714-3737, a cTNT intronic splicing enhancer at nt 3747-3770, an M2 intronic splicing enhancer at 3782-3794, and a kissing loop dimerization domain at 3801-3975. In SEQ ID NO: 23, nt 1 to 225 is the synthetic intron, nt 226 to 3747 C-terminal YFP coding region, nt 3748 to 3912 is the untranslated poly A region. SEQ ID NO: 23 comprises a kissing loop dimerization domain at nt 3-114, an M2 intronic splicing enhancer at nt 118-130, a cTNT intronic splicing enhancer at nt 137-160, a M2 intronic splicing enhancer at nt 172-184, a Branch Point Motif at nt 191-198, a poly pyrimidine tract at nt 204-223, and a splice acceptor at nt 225.
[0151] SEQ ID NOS: 24 and 25 are the N- and C-terminal sequences, respectively, used to express RNAs (pre-mRNAs) resulting in full-length human Factor VIII. In SEQ ID NO: 24, N-terminal FVIII coding region with N-terminal HA epitope tag nt are at nt 22 to 3561, nt 3562 to 3771 is the synthetic intron, and nt 3780 to 3828 is the untranslated poly A region. SEQ ID NO: 24 also comprises a splice donor at nt 3562-3570, a Rat FGFR2 DISE at nt 3573-3596, a cTNT intronic splicing enhancer at nt 3606-3629, an M2 intronic splicing enhancer at nt 3641-3653, and a kissing loop dimerization domain at nt 3660-3834. In SEQ ID NO: 25, nt 1 to 225 is the synthetic intron, nt 226 to 3636 is the C-terminal FVIII coding region, and nt 3665 to 3797 is the untranslated poly A region. SEQ ID NO: 25 also comprises a splice donor at nt 3703-3711, a Rat FGFR2 DISE at nt 3714-3737, a cTNT intronic splicing enhancer at nt 3747-3770, an M2 intronic splicing enhancer at 3782-3794, and a kissing loop dimerization domain at nt 3801-3975.
[0152] SEQ ID NOS: 26-136 are exemplary splicing enhancers that can be used with the systems provided herein (e.g., 118, 120, 156 of FIG. 6A).
[0153] SEQ ID NOS: 137 and 138 are exemplary splice donor sequences.
[0154] SEQ ID NOS: 139 and 140 are the N- and C-fragment respectively, of an HIV-1 based kissing loop dimerization domain.
[0155] SEQ ID NOS: 141 and 142 are the N- and C-fragment, respectively, of an HIV-2 based kissing loop dimerization domain.
[0156] SEQ ID NO: 143 is an exemplary cryptic splice acceptor sequence.
[0157] SEQ ID NO: 144 is an exemplary branch point consensus sequence.SLR 7158-102574-27 05 / 08 / 25 S2024-016A
[0158] SEQ ID NOS: 145 and 146 are the N- and middle sequences, respectively, used to express a full-length YFP, along with SEQ ID NO: 2 (C-terminal fragment). In SEQ ID NO: 145, nt 1 to 543 is the CMV promoter sequence, nt 544 to 849 N-terminal YFP coding region, and nt 850 to 1305 is the synthetic intron. In SEQ ID NO: 146, nt 1 to 522 is the CMV promoter sequence, nt 523 to 901 is the synthetic intron, nt 902 to 1084 is the middle YFP coding region, and nt 1085 to 1543 is the untranslated poly A region.
[0159] SEQ ID NOS: 147 and 148 are the 5’ and 3’- synthetic sequences, respectively, used to express a full-length Flpo. In SEQ ID NO: 147, nt 1 to 540 is the CMV promoter sequence, nt 541 to 1112 N-terminal Flpo coding region, and nt 1113 to 1571 is the synthetic intron. In SEQ ID NO: 148, nt 1 to 522 is the CMV promoter sequence, nt 523 to 904 is the synthetic intron, nt 905 to 1604 is the C-terminal Flpo coding region, nt 1605 to 1765 is the untranslated poly A region.
[0160] SEQ ID NOS: 149 and 150 are exemplary hypodiverse sequences.
[0161] SEQ ID NOS: 151 and 152 are exemplary splice donor consensus sequences.
[0162] SEQ ID NO: 153 is an exemplary kissing loop based on the HIV-2 kissing loop dimerization domain (SEQ ID NOS: 141 and 142, FIG. 17B). SEQ ID NO: 378 is an exemplary kissing loop based on the HIV-1 DIS loop (SEQ ID NOS: 139 and 140, FIG. 17A).
[0163] SEQ ID NO: 154 is an exemplary Kozak enhanced start codon.
[0164] SEQ ID NOS: 155 and 156 are exemplary constructs that can be used to express a murine Otof coding sequence in vivo. SEQ ID NO: 155 is used to produce the N-terminal Otof RNA. It comprises a human CMV enhancer and promoter at nt 1-522, a putative transcription start site at nt 523, and a poly adenylation signal at nt 4263-4311. It encodes the N-terminal Otof RNA elements as follows: 5’ untranslated region including Kozak sequence nt 523-546; 5’ Otoferlin coding sequence nt 547-4044; 5’ synthetic intron sequence nt 4045-4142; 5’ trimodal kissing loop dimerization domain nt 4143-4254; and linker at nt 4255-4262. SEQ ID NO: 155 is used to produce the C-terminal Otof RNA. It comprises a human CMV enhancer and promoter at nt 1-522, a putative transcription start site at nt 523, and a poly adenylation signal at nt 3335-3467. It encodes the C-terminal Otof RNA elements as follows: 3’ trimodal kissing loop dimerization domain nt 525-636; 3’ synthetic intron sequence nt 637-747; 3’ Otoferlin coding sequence nt 748-3225; C-terminal 3xFlag tag nt 3226-3306; and linker at nt 3307-3334.
[0165] SEQ ID NOS: 157 and 158 are exemplary constructs that can be used to express a human MYOSIN VIIA (Myo7a) coding sequence in vivo. SEQ ID NO: 157 is used to produce the N-terminal Myo7a RNA. It comprises a human CMV enhancer and promoter at nt 1-522, a putative transcription start site at nt 523, and poly adenylation signal at nt 4344-4392. It encodes the N-terminal Myo7A RNA elements as follows: 5’ untranslated region including Kozak sequence nt 523-543; 5’ Myo7a coding sequence nt 544-4125; 5’ synthetic intron sequence nt 4126-4223; 5’ trimodal kissing loop dimerization domain nt 4224-4335; and linker at nt 4336-4343. SEQ ID NO: 158 is used to produce the C-terminal Myo7a RNA. It comprises a human CMV enhancer and promoter at nt 1-522, a putative transcription start site at nt 523, and a poly adenylation signal at nt 3923-4055. It encodes the C-terminal Myo7a RNA elements as follows: 3’SLR 7158-102574-27 05 / 08 / 25 S2024-016A
[0166] trimodal kissing loop dimerization domain nt 525-636; 3’ synthetic intron sequence nt 637-747; 3’ Myo7a coding sequence nt 748-3813; C-terminal 3xFlag tag nt 3814-3894; and linker at nt 3895-3922.
[0167] SEQ ID NOS: 159 and 160 are exemplary constructs that can be used to express a full-length enzymatically dead Cas9 fused to a VPR transcriptional activator domain (dCas9-VPR) coding sequence in vivo. SEQ ID NO: 159 is used to produce the N-terminal DCas9-VPR RNA. It comprises a human CMV enhancer and promoter at nt 1-522, a putative transcription start site at nt 523, and poly adenylation signal at nt 4112-4161. It encodes the N-terminal DCas9-VPR RNA elements as follows: 5’ untranslated region including Kozak sequence nt 523-543; 5’ DCas9-VPR coding sequence nt 544-3894; 5’ synthetic intron sequence nt 3895-3992; 5’ trimodal kissing loop dimerization domain nt 3993-4104; and linker nt 4105-4112. SEQ ID NO: 160 is used to produce the C-terminal DCas9-VPR RNA. It comprises a human CMV enhancer and promoter at nt 1-522, a putative transcription start site at nt 523, and poly adenylation signal at nt 3278-3410. It encodes the C-terminal DCas9-VPR RNA elements as follows: 3’ trimodal kissing loop dimerization domain nt 525-636; 3’ synthetic intron sequence nt 637-747; 3’ DCas9-VPR coding sequence nt 748-3249; and linker at nt 3250-3277.
[0168] SEQ ID NOS: 161 and 162 are exemplary constructs that can be used to express a full-length humanized Cas9 Prime Editor (Prime Editor) coding sequence in vivo. SEQ ID NO: 161 encodes the N-terminal Prime Editor sequence as follows: Human CMV enhancer and promoter nt 1 -522; putative transcription start site nt 523; 5’ untranslated region including Kozak sequence nt 523-543; 5’ Prime Editor coding sequence nt 544-3894; 5' synthetic intron sequence nt 3895-3992; 5’ trimodal kissing loop dimerization domain nt 3993-4104; linker nt 4105-4112; poly adenylation signal nt 4112-4161. SEQ ID NO: 162 encodes the C-terminal Prime Editor sequence as follows: Human CMV enhancer and promoter nt 1-522; putative transcription start site nt 523; 3’ trimodal kissing loop dimerization domain nt 525-636; 3’ synthetic intron sequence nt 637-747; 3’ Prime Editor coding sequence nt 748-3750; linker nt 3751-3778; poly adenylation signal nt 3779-3911.
[0169] SEQ ID NOS: 163 and 164 are exemplary constructs that can be used to express a full-length humanized Cytosine Base Editor (AncBE4) coding sequence in vivo. SEQ ID NO: 163 encodes the N-terminal AncBE4 sequence as follows: Human CMV enhancer and promoter nt 1-522; putative transcription start site nt 523; 5’ untranslated region including Kozak sequence nt 523-540; 5’ AncBE4 coding sequence nt 541-2892; 5' synthetic intron sequence nt 2893-2990; 5’ trimodal kissing loop dimerization domain nt 2991-3102; linker nt 3103-3110; poly adenylation signal nt 3111-3159. SEQ ID NO: 164 encodes the C-terminal AncBE4 sequence as follows: Human CMV enhancer and promoter nt 1-522; putative transcription start site nt 523; 3’ trimodal kissing loop dimerization domain nt 525-636; 3’ synthetic intron sequence nt 637-747; 3’ AncBE4 coding sequence nt 748-3957; linker nt 3958-3982; poly adenylation signal nt 3983-4115.
[0170] SEQ ID NOS: 165 and 166 are exemplary constructs that can be used to express a full-length humanized Adenine Base Editor (ABE8e) coding sequence in vivo. SEQ ID NO: 165 encodes the N-terminal ABE8e sequence as follows: Human CMV enhancer and promoter nt 1-522; putative transcriptionSLR 7158-102574-27 05 / 08 / 25 S2024-016A
[0171] start site nt 523; 5’ untranslated region including Kozak sequence nt 523-540; 5’ ABE8e coding sequence nt 541-2706; 5’ synthetic intron sequence nt 2707-2804; 5’ trimodal kissing loop dimerization domain nt 2805-2916; linker nt 2917-2924; poly adenylation signal nt 2925-2973. SEQ ID NO: 166 encodes the C-terminal Abe8e sequence as follows: human CMV enhancer and promoter nt 1-522; putative transcription start site nt 523; 3’ trimodal kissing loop dimerization domain nt 525-636; 3’ synthetic intron sequence nt 637-747; 3’ ABE8e coding sequence nt 748-3399; linker nt 3400-3427; poly adenylation signal nt 3428-3560.
[0172] SEQ ID NO: 167 is an exemplary kissing loop domain (GATTTTTGACCTGCTCGATTGTCCACTGCGAGCAGGTCTTTTGGAGTCGGGCGAGGCGGAAGC CCGACTCCTTTTGGCATGCACGCTAGCCGCGTCGTGCATGCCTTTTATC).
[0173] SEQ ID NO: 168 is an exemplary ISE, M2 (GGGTTATGGGACC).
[0174] SEQ ID NO: 169 is an exemplary ISE, cTNT (GGCTGAGGGAAGGACTGTCCTGGG).
[0175] SEQ ID NO: 170 is an exemplary DISE, Rat FGFR2 (CTCTTTCTTTCCATGGGTTGGCCT). SEQ ID NOS: 171 and 172 are exemplary constructs that can be used to express a full-length YFP coding sequence. SEQ ID NO: 171 encodes the N-terminal YFP sequence as follows: Human CMV enhancer and promoter nt 1-522; putative transcription start site nt 523; 5’ untranslated region including Kozak sequence nt 523-543; 5’ Stuffer open reading frame nt 544-3654; self cleaving 2A sequence nt 3655-3729; 5’ yellow fluorescent protein segment nt 3730-4224; 5’ synthetic intron sequence (variable) nt 4225-4294; 5’ trimodal kissing loop dimerization domain (uppercase): 4295-4406; linker nt 4407-4414; poly adenylation signal nt 4415-4463. SEQ ID NO: 172 encodes the C-terminal YFP sequence as follows: Name: 3’ intron screening split YFP; Human CMV enhancer and promoter nt 1-522; Putative transcription start site nt 523; 3’ trimodal kissing loop dimerization domain nt 525-636; 3’ synthetic intron sequence (variable) nt 637-706; 3’ yfp coding sequence nt 707-940; self-cleaving 2A sequence nt 941-1006; 3’ stuffer open reading frame nt 1007-4228; linker nt 4229-4265; poly adenylation signal nt 4257-4388.
[0176] SEQ ID NOS: 173-180 are exemplary intronic splicing enhancer sequences.
[0177] SEQ ID NO: 181 is a scrambled sequence.
[0178] SEQ ID NOS: 182-196 are exemplary intronic splicing enhancer sequences.
[0179] SEQ ID NO: 197-198 are scrambled sequences.
[0180] SEQ ID NOS: 199-203 are exemplary intronic splicing enhancer sequences.
[0181] SEQ ID NO: 204 is a scrambled sequence.
[0182] SEQ ID NO: 205 is an exemplary branch point sequence (TACTAACA).
[0183] SEQ ID NO: 206 is an exemplary polyadenylation signal AATAAAATATCTTTATTTTCATTACATCTGTGTGTTGGTTTTTTGTGTG.
[0184] SEQ ID NOS: 207 and 208 are an exemplary Cas9 coding sequence and protein sequence, respectively.
[0185] SEQ ID NOS: 209 and 210 are an exemplary dCas9 coding sequence and protein sequence, respectively.SLR 7158-102574-27 05 / 08 / 25 S2024-016A
[0186] SEQ ID NOS: 211 and 212 are exemplary Casl3d nucleic acid and amino acid sequences, respectively.
[0187] SEQ ID NOS: 213 and 214 are exemplary Casl3d nucleic acid and amino acid sequences, respectively.
[0188] SEQ ID NOS: 215 and 216 are exemplary dead Casl3d (e.g., catalytically inactive) amino acid sequences.
[0189] SEQ ID NO: 217 is an exemplary native HEPN domain RXXXXH.
[0190] SEQ ID NOS: 218 -221 are exemplary nuclear localization signal coding and protein sequences.
[0191] SEQ ID NO: 222 is an exemplary Casl3d protein sequence.
[0192] SEQ ID NO: 223 is an exemplary DR sequence that can be included in a gRNA sequence and used with the Casl3d protein of SEQ ID NO: 212 for RNA editing.
[0193] SEQ ID NO: 224 is an exemplary DR sequence that can be included in a gRNA sequence and used with the Casl3d protein of SEQ ID NO: for RNA editing.
[0194] SEQ ID NOS: 225 and 226 are exemplary constructs that can be used to express a full-length humanized Adenine Base Editor (ABE8e) coding sequence in vivo. SEQ ID NO: 225 encodes the N-terminal ABE8e sequence and comprises two gRNA expression cassettes as follows: nt 1-141 AAV2 Inverted Terminal Repeat; nt 159-255 CRISPR gRNA (reverse orientation), nt 256-504 human U6 RNA polymerase III promoter (reverse orientation), nt 512-1019 CMV promoter, nt 1034-1051 5' untranslated region, nt 1052-3217 N-terminal ABE8e editor, nt 3218-3315 synthetic intron sequence, nt 3316-3427 Dimerization Domain, nt 3436-3485 poly adenylation signal, nt 3492-3706 Hl polymerase III promoter, nt 3707-3803 CRISPR gRNA, and nt 3825-3965 AAV2 Inverted Terminal Repeat. SEQ ID NO: 225 comprises SEQ ID NO: 165 and added gRNA expression cassettes. SEQ ID NO: 226 encodes the C-terminal ABE8e sequence and comprises two gRNA expression cassettes as follows: nt 1-141 AAV2 Inverted Terminal Repeat, nt 159-255 CRISPR gRNA (reverse orientation), nt 256-504 human U6 RNA polymerase III promoter (reverse orientation), nt 512-1019 CMV promoter, nt 1036-1147 Dimerization Domain, nt 1259-39103’ ABE8e coding sequence, nt 3939-4069 poly adenylation signal, nt 4078-4292 Hl polymerase III promoter, nt 4293-4389 CRISPR gRNA, and nt 4411-4551 AAV2 Inverted Terminal Repeat. SEQ ID NO: 226 comprises SEQ ID NO: 166 and added gRNA expression cassettes.
[0195] SEQ ID NO: 227 is an exemplary SIRLOIN element sequence.
[0196] SEQ ID NO: 228 is an exemplary construct that encodes an N-terminal REJ construct sequence that can be used with the C-terminal REJ construct of SEQ ID NO: 229 to express mini-dystrophin coding sequence in vivo. SEQ ID NO: 228 includes an HA-tag sequence, nt 544-627; an N-terminal coding sequence of mini-dystrophin, nt 628-3744 and 3832-3975; a stimulatory intron sequence, nt 3745-3831; and a synthetic intron sequence, nt 3976-4212 (including SD, ratFgfr2 DISE, cTNT ISE, ISE, and an RNA dimerization domain).
[0197] SEQ ID NO: 229 is an exemplary construct that encodes a C-terminal REJ construct sequence that can be used with the N-terminal REJ construct of SEQ ID NO: 228 to express mini-dystrophin codingSLR 7158-102574-27 05 / 08 / 25 S2024-016A
[0198] sequence in vivo. SEQ ID NO: 229 includes a synthetic intron sequence, nt 523-716 (including an RNA dimerization domain, ISE, cTNT, ISE, BP, PPT, and SA), a C-terminal coding sequence of mini-dystrophin, nt 717-884 and 972-3974; a stimulatory intron sequence, nt 885-971; and a FLAG-tag sequence, nt 3975-4055.
[0199] SEQ ID NOS: 230 and 231 are the N-terminus, and C-terminus, respectively, of REJ constructs that can be used to express full-length Pcdhl5. In these examples, these sequences include the distal2 binding domains (see FIGS. 32 and 33). SEQ ID NO: 230 includes a distal2 binding domain, nt 4107-4167, which can be replaced with a corresponding sequence from SEQ ID NOS: 232-241 to generate other sequences for expressing Pcdhl5. SEQ ID NO: 230 further includes a CMV promoter, nt 1-508; a GRK1 5’ UTR, nt 509-683; and a stimulatory intron sequence, nt 3790-3876. SEQ ID NO: 231 includes a distal 2 binding domain, nt 710-770, which can be replaced with a corresponding sequence from SEQ ID NOS: 232-241 to generate other sequences for expressing Pcdhl5. SEQ ID NO: 231 further includes a CMV promoter, nt 1-508; a GRK1 5’UTR, nt 509-683; and a stimulatory intron sequence, nt 1010-1096.
[0200] SEQ ID NO: 232 is the Prxl binding domain sequence (e.g., for joining with an N-terminal protein coding sequence, such as n-PCDH15).
[0201] SEQ ID NO: 233 is the Prx2 binding domain sequence (e.g., for joining with an N-terminal protein coding sequence, such as n-PCDH15).
[0202] SEQ ID NO: 234 is the Dstl binding domain sequence (e.g., for joining with an N-terminal protein coding sequence, such as n-PCDH15).
[0203] SEQ ID NO: 235 is the Dst2 binding domain sequence (e.g., for joining with an N-terminal protein coding sequence, such as n-PCDH15).
[0204] SEQ ID NO: 236 is the Doul binding domain sequence (e.g., for joining with an N-terminal protein coding sequence, such as n-PCDH15).
[0205] SEQ ID NO: 237 is the Prxl anti-binding domain sequence (e.g., for joining with a C-terminal protein coding sequence, such as C-PCDH15).
[0206] SEQ ID NO: 238 is the Prx2 anti-binding domain sequence (e.g., for joining with a C-terminal protein coding sequence, such as C-PCDH15).
[0207] SEQ ID NO: 239 is the Dstl anti-binding domain sequence (e.g., for joining with a C-terminal protein coding sequence, such as C-PCDH15).
[0208] SEQ ID NO: 240 is the Dst2 anti-binding domain sequence (e.g., for joining with a C-terminal protein coding sequence, such as C-PCDH15).
[0209] SEQ ID NO: 241 is the Doul anti-binding domain sequence (e.g., for joining with a C-terminal protein coding sequence, such as C-PCDH15).
[0210] SEQ ID NO: 242 is an exemplary 3' REJ intron RNA which contains linear design (e.g., see FIG.
[0211] 27).
[0212] SEQ ID NO: 243 is an exemplary 3' REJ intron RNA which contains 3X kissing loop design (e.g., see FIG. 27).SLR 7158-102574-27 05 / 08 / 25 S2024-016A
[0213] SEQ ID NO: 244 is an exemplary 3' REJ intron RNA which contains hairpin design (e.g., see FIG.
[0214] 27).
[0215] SEQ ID NO: 245 is an exemplary 3' REJ intron RNA which contains G-Quadruplex + 3X kissing loop design e.g., see FIG. 27).
[0216] SEQ ID NO: 246 is an exemplary MALAT1 M Region Sequence.
[0217] SEQ ID NO: 247 is an exemplary Prxl N-terminal sequence containing the proximal 1 conformation shown in FIG. 32, which can be used to generate full-length Pcdhl5.
[0218] SEQ ID NO: 248 is an exemplary Prxl C-terminal sequence containing the proximal 1 conformation shown in FIG. 32, which can be used to generate full-length Pcdhl5.
[0219] SEQ ID NO: 249 is an exemplary Prx2 N-terminal sequence containing the proximal 2 conformation shown in FIG. 32, which can be used to generate full-length Pcdhl5.
[0220] SEQ ID NO: 250 is an exemplary Prx2 C-terminal sequence containing the proximal 2 conformation shown in FIG. 32, which can be used to generate full-length Pcdhl5.
[0221] SEQ ID NO: 251 is an exemplary Dstl N-terminal sequence containing the distal 1 conformation shown in FIG. 32, which can be used to generate full-length Pcdhl5.
[0222] SEQ ID NO: 252 is an exemplary Dstl C-terminal sequence containing the distal 1 conformation shown in FIG. 32, which can he used to generate full-length Pcdhl5.
[0223] SEQ ID NO: 253 is an exemplary Dst2 N-terminal sequence containing the distal 2 conformation shown in FIG. 32, which can be used to generate full-length Pcdhl5.
[0224] SEQ ID NO: 254 is an exemplary Dst2 C-terminal sequence containing the distal 2 conformation shown in FIG. 32, which can be used to generate full-length Pcdhl5.
[0225] SEQ ID NO: 255 is an exemplary Dou N-terminal sequence containing the double 1 conformation shown in FIG. 32, which can be used to generate full-length Pcdhl5.
[0226] SEQ ID NO: 256 is an exemplary Dou C-terminal sequence containing the double 1 conformation shown in FIG. 32, which can be used to generate full-length Pcdhl5.
[0227] SEQ ID NOS: 257-336 and 377 are exemplary dimerization domain sequences.
[0228] SEQ ID NO: 337 is an exemplary N-terminal REJ construct sequence for expressing human Pcdhl5 CD1. The N-terminal REJ construct sequence includes a GRK1 promoter (nt 1-151), a GRK1 5’ UTR (nt 152-311), and a stimulatory intron sequence (nt 3455-3541).
[0229] SEQ ID NO: 338 is an exemplary C-terminal REJ construct sequence for expressing human Pcdhl5 CD1. The C-terminal REJ construct sequence includes a GRK1 promoter (nt 1-151), a GRK1 5’ UTR (nt 152-311), and a stimulatory intron sequence (nt 638-724).
[0230] SEQ ID NO: 339 is an exemplary N-terminal REJ construct sequence for expressing mouse Pcdhl5 CD2. The N-terminal REJ construct sequence includes a CMV enhancer and promoter sequences (nt 1-584).
[0231] SEQ ID NO: 340 is an exemplary C-terminal REJ construct sequence for expressing mouse Pcdhl5 CD2. The C-terminal REJ construct sequence includes a CMV enhancer and promoter sequences (nt 1-584).SLR 7158-102574-27 05 / 08 / 25 S2024-016A
[0232] SEQ ID NO: 341 is an exemplary CMV promoter sequence.
[0233] SEQ ID NO: 342 is an exemplary CMV enhancer sequence.
[0234] SEQ ID NO: 343 is an exemplary GRK1 promoter sequence.
[0235] SEQ ID NO: 344 is an exemplary GRK1 5’ UTR sequence.
[0236] SEQ ID NO: 345 is an exemplary stimulatory intron sequence.
[0237] SEQ ID NOS: 346 and 347 are N-terminal and C-terminal YFP coding sequences with a stimulatory intron.
[0238] SEQ ID NOS: 348-376 are splice donor and splice acceptor motifs.
[0239] SEQ ID NOS: 379 and 380 are N-terminal and C-terminal YFP coding sequences without a stimulatory intron.
[0240] SEQ ID NO: 381 is an exemplary coding sequence of AH2-R17 mini-dystrophin.
[0241] DETAILED DESCRIPTION
[0242] I. Summary of Terms
[0243] Unless otherwise noted, technical terms are used according to conventional usage. Definitions of common terms in molecular biology can be found in Benjamin Lewin, Genes VII, published by Oxford University Press, 1999; Kendrew et al. (eds.), The Encyclopedia of Molecular Biology, published by Blackwell Science Ltd., 1994; and Robert A. Meyers (edj, Molecular Biology and Biotechnology: a Comprehensive Desk Reference, published by VCH Publishers, Inc., 1995; and other similar references.
[0244] As used herein, the singular forms “a,” “an,” and “the,” refer to both the singular as well as plural, unless the context clearly indicates otherwise. As used herein, the term “comprises” means “includes.” Thus, “comprising a nucleic acid molecule” means “including a nucleic acid molecule” without excluding other elements. It is further to be understood that any and all base sizes given for nucleic acids are approximate, and are provided for descriptive purposes, unless otherwise indicated. Although many methods and materials similar or equivalent to those described herein can be used, particular suitable methods and materials are described below. In case of conflict, the present specification, including explanations of terms, will control. In addition, the materials, methods, and examples are illustrative only and not intended to be limiting. All references, including patent applications and patents, and GenBank Accession Nos., are herein incorporated by reference in their entireties.
[0245] In order to facilitate review of the various embodiments of the disclosure, the following explanations of specific terms are provided:
[0246] About: Unless context indicated otherwise, “about” refers to plus or minus 5% of a reference value. For example, “about” 100 refers to 95 to 105.
[0247] Administration: To provide or give a subject an agent, such as a therapeutic nucleic acid molecule provided herein (such as one encoding one or more portions of a nucleic acid editor protein, gRNA, or both), or other therapeutic agent, by any effective route. Exemplary routes of administration include, but are not limited to, injection (such as subcutaneous, intramuscular, intradermal, intraperitoneal, intrathecal,SLR 7158-102574-27 05 / 08 / 25 S2024-016A
[0248] intratumoral, intraosseous, and intravenous), transdermal, intranasal, and inhalation routes. Administration can be systemic or local.
[0249] Aptamer: Nucleic acid molecules (such as DNA or RNA) that bind a specific target agent or molecule with high affinity and specificity. Aptamers can be used in the disclosed nucleic acid molecules as a dimerization domain. In one example, two aptamers can bind to each other, e.g., by standard basepairing, non-canonical base pair interactions, non-base pairing interactions, or a combination thereof, to mediate dimerization. In one example, aptamers allow RNA dimerization (and subsequent recombination) only in the presence of one or more targets recognized by the aptamer. Aptamers have been obtained through a combinatorial selection process called systematic evolution of ligands by exponential enrichment (SELEX) (see for example Ellington et al.. Nature 1990, 346, 818-822; Tuerk and Gold Science 1990, 249, 505-510; Liu et al., Chem. Rev. 2009, 109, 1948-1998; Shamah et al.. Acc. Chem. Res. 2008, 41, 130-138; Famulok, et al., Chem. Rev. 2007, 107, 3715-3743; Manimala et al., Recent Dev. Nucleic Acids Res. 2004, 1, 207-231; Famulok et al., Acc. Chem. Res. 2000, 33, 591-599; Hesselberth, et al., Rev. Mol. Biotech. 2000, 74, 15-25; Wilson et al.. Annu. Rev. Biochem. 1999, 68, 611-647; Morris et al.. Proc. Natl. Acad. Sci. U. S. A. 1998, 95, 2902-2907). In such a process, DNA or RNA molecules that are capable of binding a target molecule of interest are selected from a nucleic acid library consisting of 1014-1015different sequences through iterative steps of selection, amplification and mutation. The affinity of the aptamers towards their targets can rival that of antibodies, with dissociation constants in as low as the picomolar range (Morris et al., Proc. Natl. Acad. Sci. U. S. A. 1998, 95, 2902-2907; Green et al.. Biochemistry 1996, 35, 14413-14424).
[0250] Aptamers that are specific to a wide range of targets from small organic molecules such as adenosine, to proteins such as thrombin, and even viruses and cells have been identified (Liu et al., Chem. Rev. 2009, 109, 1948-1998; Lee et al., Nucleic Acids Res. 2004, 32, D95-D100; Navani and Li, Curr. Opin. Chem. Biol. 2006, 10, 272-281; Song et al., TrAC, Trends Anal. Chem. 2008, 27, 108-117). For example, aptamers are available that recognize metal ions such as Zn(II) (Ciesiolka et al., RNA 1: 538-550, 1995) and Ni(II) (Hofmann et al., RNA, 3:1289-1300, 1997); nucleotides such as adenosine triphosphate (ATP) (Huizcnga and Szostak, Biochemistry, 34:656-665, 1995); and guanine (Kiga et al., Nucleic Acids Res., 26: 1755-60, 1998); co-factors such as NAD (Kiga et al., Nucleic Acids Res., 26: 1755-60, 1998) and flavin (Lauhon and Szostak, J. Am. Chem. Soc., 117:1246-57, 1995); antibiotics such as viomycin (Wallis et al., Chem. Biol. 4: 357-366, 1997) and streptomycin (Wallace and Schroeder, RNA 4:112-123, 1998); proteins such as HIV reverse transcriptase (Chaloin et al., Nucleic Acids Res., 30:4001-8, 2002) and hepatitis C virus RNA-dependent RNA polymerase (Biroccio et al., J. Virol. 76:3688-96, 2002); toxins such as cholera whole toxin and staphylococcal enterotoxin B (Bruno and Kiel, BioTechniques, 32: pp. 178-180 and 182-183, 2002); and bacterial spores such as the anthrax (Bruno and Kiel, Biosensors & Bioelectronics, 14:457-464, 1999).
[0251] Binding: An association between two substances or molecules, such as the hybridization of one nucleic acid molecule to another (or itself), such as between two dimerization domains, or the binding of an aptamer to its target. An oligonucleotide molecule binds or stably binds to another nucleic acid molecule ifSLR 7158-102574-27 05 / 08 / 25 S2024-016A
[0252] there are a sufficient number of complementary base pairs between the oligonucleotide molecule and the target nucleic acid to permit detection of that binding. In some examples, binding between nucleic acid molecules may occur directly. In some examples, binding between nucleic acid molecules may occur indirectly, e.g., through an intermediate molecule. Either direct binding or indirect binding may occur by standard base pairing, by non-canonical base pair interactions, by non-base pair interactions, or a combination thereof. Non-canonical base pair interactions may occur by any means of stabilization known to those of skill in the art, including but not limited to Hoogsteen base pairs and wobble base pairs. Nonbase pair interactions can include binding through an intermediate molecule. In some examples, direct binding is between kissing loop dimerization domains. In some examples, direct binding is between hypodiverse dimerization domains. In some examples, direct binding is between aptamer regions. In some examples, direct binding between aptamer regions involves non-canonical base pair interactions. In some examples, direct binding between aptamer regions involves standard base pairing and non-canonical base pair interactions. In some examples, indirect binding occurs through a nucleic acid bridge. In some examples the nucleic acid bridge is an mRNA. A nonlimiting example of a nucleic acid bridge is depicted in Fig. 7B. In some examples, indirect binding occurs through an aptamer molecule. A nonlimiting example of indirect binding through an aptamer molecule is depicted in Fig. 7A. In some embodiments, indirect binding through an aptamer molecule involves non-base pair interactions between the aptamer molecule and the binding regions. In some embodiments, indirect binding through an aptamer molecule involves non-base pair interactions between the aptamer molecule and the binding regions, and base pairing interactions between the binding regions.
[0253] C-terminal portion: A region of a protein sequence that includes a contiguous stretch of amino acids that begins at or near the C-terminal residue of the protein. A C-terminal portion of the protein can be defined by a contiguous stretch of amino acids (e.g., a number of amino acid residues).
[0254] Cancer: A malignant tumor characterized by abnormal or uncontrolled cell growth. Other features often associated with cancer include metastasis, interference with the normal functioning of neighboring cells, release of cytokines or other secretory products at abnormal levels and suppression or aggravation of inflammatory or immunological response, invasion of surrounding or distant tissues or organs, such as lymph nodes, etc. “Metastatic disease” refers to cancer cells that have left the original tumor site and migrate to other parts of the body for example via the bloodstream or lymph system.
[0255] Cas9: An RNA-guided DNA endonuclease enzyme that that participates in the CRISPR-Cas immune defense against prokaryotic viruses. Cas9 has two active cutting sites (HNH and RuvC), one for each strand of the double helix. An exemplary native Cas9 sequence from S. pyogenes is shown in SEQ ID NO: 208.
[0256] Catalytically inactive (deactivated or dead) Cas9 (dCas9) proteins, which have reduced or abolished endonuclease activity but still binds to dsDNA, are also encompassed by this disclosure. In some examples, a dCas9 includes one or more mutations in the RuvC and HNH nuclease domains, such as one or more of the following point mutations: D10A, E762A, D839A, H840A, N854A, N863A, and D986A (e.g., based onSLR 7158-102574-27 05 / 08 / 25 S2024-016A
[0257] numbering in SEQ ID NO: 208). An exemplary dCas9 sequence with D10A and H840A substitutions is shown in SEQ ID NO: 210. In one example, the dCas9 protein has mutations D10A, H840A, D839A, and N863A (see, e.g., Esvelt etal., Nat. Meth. 10:1116-21, 2013).
[0258] Cas9 and dCas9 sequences are publicly available. For example, GenBank® Accession Nos. nucleotides 796693..800799 of CP012045.1 and nucleotides 1100046..1104152 of CP014139.1 disclose Cas9 nucleic acids, and GenBank® Accession Nos. NP_269215.1, AMA70685.1, and AKP81606.1 disclose Cas9 proteins. In some examples, a deactivated form of Cas9 (dCas9) is nuclease deficient (e.g., those shown in GenBank® Accession Nos. AKA60242.1 and KR011748.1). Activatable Cas9 proteins are provided in US Publication No. 2018-0073002-Al. In certain examples, a Cas9 or dCas9 used in the disclosed compositions or methods has at least 80% sequence identity, for example at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99% sequence identity, or 100% to such sequences (such as SEQ ID NOS: 207, 208, 209, and 210), and retains the ability to be used in the disclosed compositions and methods (e.g., can be encoded by two or more separate molecules of the present disclosure, and subsequently recombined using the REJ methods provided herein).
[0259] Casl3d: An RNA-guided RNA endonuclease enzyme that can cut or bind RNA. Cas 13d proteins specifically recognize direct repeat (DR) sequences present in gRNA having a particular secondary structure. Casl3d proteins include one or two HEPN domains. Native HEPN domains include the sequence RXXXXH (SEQ ID NO: 217), wherein X is any amino acid. A catalytically inactive, or “dead” Casl3d, which include mutated HEPN domain(s) and thus cannot cut RNA, but can process gRNA, are also encompassed by this disclosure (e.g., see SEQ ID NOS: 215 and 216). Such a dead Cas 13d (dCas!3d) can be targeted to cis-elements of pre-mRNA to manipulate alternative splicing.
[0260] Exemplary native and variant Cas 13d protein sequences are provided in WO 2019 / 040664, US 10,876,101 and US 10,392,616 (all herein incorporated by reference in their entireties), as well as herein as SEQ ID NOS: 212, 214, 215, 216, and 222.
[0261] In one example, a full length (non-truncated) Casl3d protein is between 870-1080 amino acids long. In one example, the Cas 13d protein is derived from a genome sequence of a bacterium from the Order Clostridiales or a metagenomic sequence. In one example, the corresponding DR sequence of a Casl3d protein is located at the 5’ end of the spacer sequence in the molecule that includes the Cas 13d gRNA. In one example, the DR sequence in the Casl3d gRNA is truncated at the 5’ end relative to the DR sequence in the unprocessed Cas 13d guide array transcript (such as truncated by at least 1 nt, at least 2 nt, at least 3 nt, at least 4 nt, at least 5 nt, such as 1-3 nt, 3-6 nt, 5-7 nt, or 5-10 nt). In one example, the DR sequence in the Casl3d gRNA is truncated by 5-7 nt at the 5’ end by the Casl3d protein. In one example, the Casl3d protein can cut a target RNA flanked at the 3’ end of the spacer-target duplex by any of a, U, G or C ribonucleotide and flanked at the 5’ end by any of a, U, G or C ribonucleotide.
[0262] In one example, a Casl3d protein has at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 95%, at least 96%, at least 97%, at least 98% or at least 99% sequence identity to SEQ ID NO: 212, 214, 215, 216, or 222. In one example, a Casl3d coding sequence encodes a Casl3d protein having atSLR 7158-102574-27 05 / 08 / 25 S2024-016A
[0263] least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 95%, at least 96%, at least 97%, at least 98% or at least 99% sequence identity to SEQ ID NO: 212, 214, 215, 216, or 222. In one example, a Casl3d coding sequence has at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 95%, at least 96%, at least 97%, at least 98% or at least 99% sequence identity to SEQ ID NO: 211 or 213.
[0264] Complementarity: The ability of a nucleic acid to form hydrogen bond(s) with another nucleic acid sequence by either traditional Watson-Crick base pairing or other non-traditional types. Traditional Watson-Crick base pairing requires reverse complementarity, i.e., the two complementary strands being antiparallel (one running from 5' to 3', the other from 3' to 5'). A percent complementarity indicates the percentage of residues in a nucleic acid molecule (e.g., target DNA or RNA) which can form hydrogen bonds (e.g., Watson-Crick base pairing) with a second nucleic acid sequence, such as a gRNA (e.g., 5, 6, 7, 8, 9, 10 out of 10 being 50%, 60%, 70%, 80%, 90%, and 100% complementary), or between two dimerization domains. " Perfectly complementary" means that all the contiguous residues of a nucleic acid sequence will hydrogen bond with the same number of contiguous residues in a second nucleic acid sequence. " Substantially complementary" as used herein refers to a degree of complementarity that is at least 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 97%, 98%, 99%, or 100% over a region of 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 30, 35, 40, 45, 50, or more nucleotides, or refers to two nucleic acids that hybridize under stringent conditions. Thus, in some examples, a first dimerization domain and a second dimerization domain have perfect complementarity to one another (e.g., 100%). In other examples, a first dimerization domain and a second dimerization domain are substantially complementary to one another (e.g., at least 80%).
[0265] Contact: Placement in direct physical association, including a solid or a liquid form. Contacting can occur in vitro or ex vivo, for example, by adding a reagent to a sample (such as one containing cells), or in vivo by administering to a subject.
[0266] CRISPR / Cas system: A prokaryotic immune system that confers resistance to foreign genetic elements, such as plasmids and phages, and provides a form of acquired immunity. The system includes a Cas nuclease (e.g., Cas9, Casl3d) and a guide RNA (gRNA) that specifically binds to the target RNA or DNA and directs the Cas nuclease to a target site. The disclosed compositions, systems, and methods can be used to express a Cas nuclease from two or more different DNA molecules, which in some examples further encode one or more gRNAs to regulate gene expression, for example to increase or decrease expression of a target nucleic acid molecule, and / or to edit a sequence of a target nucleic acid molecule (for example to repair one or more mutations associated with a disease, such as a substitution, insertion or deletion).
[0267] Dead guide RNA (dgRNA): A guide RNA (gRNA) that can guide wild-type Cas nuclease (e.g., Cas9) to a target nucleic acid, but does not induce double strand DNA breaks. The shortened gRNAs contain shortened targeting sequences of about 14 to 15 nucleotides, whereas non-dead gRNAs contain targeting sequences of about 20 nucleotides. dgRNAs are further described, for example, in Dahlman et al. (2015) Nat. Biotechnol. 33:1159-1161; Kiani etal. (2015) Nat. Methods, 12:1051-1054; and Hsin-Kai Liao et al. (2017) Cell, 171:1495-1507, all herein incorporated by reference in their entirety. In some examples,SLR 7158-102574-27 05 / 08 / 25 S2024-016A
[0268] the dgRNA is an RNA molecule (for example, when expressed in a cell). In some examples, the dgRNA is encoded by a DNA molecule (for example, when in a vector, such as a viral vector).
[0269] Dimerization domain: An RNA sequence that binds with another RNA sequence, e.g., through base pairing. The term also encompasses DNA that encodes the RNA dimerization domain. A pair of dimerization domains (also referred to as binding domain and anti-binding domain) bring two RNA molecules together (e.g., an N-terminal REJ construct and a C-terminal REJ construct). In some aspects, the dimerization domains interact through complementarity. In some examples, the complementarity is formed between linear RNAs. In some examples, the complementarity is formed between one or more pairs of kissing loops.
[0270] DNA Editing: A type of genetic engineering in which a DNA molecule (or nucleotides of the DNA) is inserted, deleted or replaced in a cell or organism using a nucleases (such as Cas9, Casl3d or dead versions thereof), which create site-specific strand breaks at desired locations in the DNA. The induced breaks are repaired resulting in targeted mutations or repairs. CRISPR / Cas methods, for example using the REJ systems provided herein to express a Cas nuclease, can be used to edit the sequence of one or more target DNAs, such as one associated with cancer (e.g., breast cancer, colon cancer, lung cancer, prostate cancer, melanoma), infectious disease (such as HIV, hepatitis, HPV, and West Nile virus), or neurodegenerative disorder (e.g., Huntington’s disease or ALS). For example, DNA editing can be used to treat a disease or viral infection.
[0271] DNA insertion site: A site of the DNA that is targeted for, or has undergone, insertion of an exogenous polynucleotide. The disclosed methods include use of a nucleic acid editor expressed from two or more nucleic acid molecules provided herein, which can be used to target a DNA for manipulation at a DNA insertion site.
[0272] Downregulated or knocked down: When used in reference to the expression of a molecule, such as a target nucleic acid or protein, refers to any process which results in a decrease in production of the target nucleic acid or protein, but in some examples not complete elimination of the target RNA product or target nucleic acid function. In one example, downregulation or knock down does not result in complete elimination of detectable target nucleic acid / protein expression or activity. In some examples, downregulation or knock down of a target nucleic acid includes processes that decrease translation of the target RNA and thus can decrease the presence of corresponding proteins. The disclosed system can be used to downregulate any target nucleic acid / protein of interest.
[0273] Downregulation or knock down includes any detectable decrease in the target nucleic acid / protein. In certain examples, detectable target nucleic acid / protein in a cell or cell free system decreases by at least 10%, at least 20%, at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, at least 75%, at least 80%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% (such as a decrease of 40% to 90%, 40% to 80% or 50% to 95%) as compared to a control (such an amount of target nucleic acid / protein detected in a corresponding untreated cell or sample). In one example, a control is a relativeSLR 7158-102574-27 05 / 08 / 25 S2024-016A
[0274] amount of expression in a normal cell (e.g., a non-recombinant cell that does not include a nucleic acid molecule for RNA recombination provided herein).
[0275] Effective amount: The amount of an agent (such as a system providing multiple vectors, each encoding a different portion of a nucleic acid editing protein, such as a Cas9 or Casl3d protein, for example in combination with an effective amount of one or more gRNAs that can hybridize to the nucleic acid target) that is sufficient to effect beneficial or desired results. An effective amount also can refer to an amount of correctly joined RNA or nucleic acid editing protein produced that is sufficient to effect beneficial or desired results, for example in combination with one or more gRNAs that can hybridize to the nucleic acid target.
[0276] An effective amount (also referred to as a therapeutically effective amount) may vary depending upon one or more of: the subject and disease condition being treated, the weight and age of the subject, the severity of the disease condition, the manner of administration and the like, which can be determined by one of ordinary skill in the art. The beneficial therapeutic effect can include enablement of diagnostic determinations; amelioration of a disease, symptom, disorder, or pathological condition; reducing or preventing the onset of a disease, symptom, disorder or condition; and generally counteracting a disease, symptom, disorder or pathological condition.
[0277] In one embodiment, an “effective amount” of two or more synthetic nucleic acid molecules provided herein, sufficient to treat a disease, such as a genetic disease or cancer. In one embodiment, an “effective amount” of two or more synthetic nucleic acid molecules provided herein is amount sufficient to increase the survival time of a treated patient, for example by at least 10%, at least 20%, at least 25%, at least 50%, at least 70%, at least 75%, at least 80%, at least 90%, at least 95%, at least 99%, at least 100%, at least 200%, at least 300%, at least 400%, at least 500%, or at least 600% (as compared to no administration of the two or more synthetic nucleic acid molecules provided herein). In one embodiment, an “effective amount” of two or more synthetic nucleic acid molecules provided herein is an amount sufficient to increase the survival time of a treated patient, for example by at least 6 months, at least 9 months, at least 1 year, at least 1.5 years, at least 2 years, at least 2.5 years, at least 3 years, at least 4 years, at least 5 years, at least 10 years, at least 12 years, at least 15 years, or at least 20 years (as compared to no administration of the two or more synthetic nucleic acid molecules provided herein). In one embodiment, an “effective amount” of two or more synthetic nucleic acid molecules provided herein is an amount sufficient to increase mobility of a treated patient (such as a DMD patient), for example by at least 10%, at least 20%, at least 25%, at least 50%, at least 70%, at least 75%, at least 80%, at least 90%, at least 95%, at least 99%, at least 100%, at least 200%, at least 300%, at least 400%, at least 500%, or at least 600% (as compared to no administration of the two or more synthetic nucleic acid molecules provided herein). In one embodiment, an “effective amount” of two or more synthetic nucleic acid molecules provided herein is an amount sufficient to increase mobility of a treated patient (such as a DMD patient), for example by at least 10%, at least 20%, at least 25%, at least 50%, at least 70%, at least 75%, at least 80%, at least 90%, at least 95%, at least 99%, at least 100%, at least 200%, at least 300%, at least 400%, at least 500%, or at least 600% (as compared to no administration of the two or more synthetic nucleic acid molecules provided herein). In one embodiment, an “effective amount”SLR 7158-102574-27 05 / 08 / 25 S2024-016A
[0278] of two or more synthetic nucleic acid molecules provided herein is an amount sufficient to increase cognitive ability of a treated patient (such as a DMD patient), for example by at least 10%, at least 20%, at least 25%, at least 50%, at least 70%, at least 75%, at least 80%, at least 90%, at least 95%, at least 99%, at least 100%, at least 200%, at least 300%, at least 400%, at least 500%, or at least 600% (as compared to no administration of the two or more synthetic nucleic acid molecules provided herein). In one embodiment, an “effective amount” of two or more synthetic nucleic acid molecules provided herein is an amount sufficient to increase respiratory function of a treated patient (such as a DMD patient), for example by at least 10%, at least 20%, at least 25%, at least 50%, at least 70%, at least 75%, at least 80%, at least 90%, at least 95%, at least 99%, at least 100%, at least 200%, at least 300%, at least 400%, at least 500%, or at least 600% (as compared to no administration of the two or more synthetic nucleic acid molecules provided herein). In one embodiment, an “effective amount” of two or more synthetic nucleic acid molecules provided herein is an amount sufficient to increase blood clotting of a treated patient (such as a hemophilia patient), for example by at least 10%, at least 20%, at least 25%, at least 50%, at least 70%, at least 75%, at least 80%, at least 90%, at least 95%, at least 99%, at least 100%, at least 200%, at least 300%, at least 400%, at least 500%, or at least 600% (as compared to no administration of the two or more synthetic nucleic acid molecules provided herein). In one embodiment, an “effective amount” of two or more synthetic nucleic acid molecules provided herein is an amount sufficient to increase vision of a treated patient (such as a Usher or Stargardt patient), for example by at least 10%, at least 20%, at least 25%, at least 50%, at least 70%, at least 75%, at least 80%, at least 90%, at least 95%, at least 99%, at least 100%, at least 200%, at least 300%, at least 400%, at least 500%, or at least 600% (as compared to no administration of the two or more synthetic nucleic acid molecules provided herein). In one embodiment, an “effective amount” of two or more synthetic nucleic acid molecules provided herein is an amount sufficient to increase hearing of a treated patient (such as a Usher patient), for example by at least 10%, at least 20%, at least 25%, at least 50%, at least 70%, at least 75%, at least 80%, at least 90%, at least 95%, at least 99%, at least 100%, at least 200%, at least 300%, at least 400%, at least 500%, or at least 600% (as compared to no administration of the two or more synthetic nucleic acid molecules provided herein).
[0279] In one embodiment, an “effective amount” of two or more synthetic nucleic acid molecules provided herein is an amount sufficient to reduce calf muscle size of a treated DMD patient, for example by at least 10%, at least 20%, at least 25%, at least 50%, at least 70%, at least 75%, at least 80%, at least 90%, or at least 95% (as compared to no administration of the two or more synthetic nucleic acid molecules provided herein). In one embodiment, an “effective amount” of two or more synthetic nucleic acid molecules provided herein is an amount sufficient to reduce cardiomyopathy muscle size of a treated DMD patient, for example by at least 10%, at least 20%, at least 25%, at least 50%, at least 70%, at least 75%, at least 80%, at least 90%, or at least 95% (as compared to no administration of the two or more synthetic nucleic acid molecules provided herein). In some examples, combinations of these effects are achieved.
[0280] Guide RNA (gRNA): A synthetic nucleic acid sequence used to direct a Cas nuclease (or dead Cas nuclease) protein to a target nucleic acid sequence, such as a target DNA (e.g., genomic sequence) or targetSLR 7158-102574-27 05 / 08 / 25 S2024-016A
[0281] RNA sequence. gRNA molecules include, whether as part of a single nucleic acid molecule or divided into two or more nucleic acid molecules, (1) a portion with sequence complementarity to the target nucleic acid (such as at least 80%, at least 90%, at least 95%, or 100% sequence complementarity, and (2) a portion with secondary structure that binds to the Cas nuclease. Thus, one can change the target nucleic acid of the Cas protein by simply changing the target sequence present in the gRNA (See CRISPR-Cas9 Structures and Mechanisms. Fuguo Jiang and Jennifer A. Doudna, Annual Review of Biophysics, 46:1, 505-529 (2017)). In some examples, the gRNA is an RNA molecule (for example, when expressed in a cell). In some examples, a gRNA is encoded by a DNA molecule (for example, when part of a vector, such as a viral vector). A gRNA can include modified bases or chemical modifications (e.g., see Latorre et al., Angewandte Chemie 55:3548-50, 2016).
[0282] In some examples, a gRNA includes two or more MS2-binding loop sequences, which can be modified from the native MS2-binding loop sequence to increase GC content and / or shorten repetitive content. In some examples, the gRNA is modified to increase GC content and / or shorten repetitive content.
[0283] In some examples, the gRNA is a dead guide RNA (dgRNA). Increasing GC content and / or shortening the repetitive content of the gRNA can be used to convert an gRNA into a dgRNA, that is, a guide nucleic acid molecule that can direct a Cas nuclease to a target nucleic acid sequence, but does not induce a DNA double strand break (or RNA single strand break).
[0284] In one example, a gRNA directs a Cas DNA nuclease (such as Cas9) to a target DNA. In one such example, the gRNA includes from 5’ to 3’ (1) a CRISPR RNA (crRNA) region that includes a sequence designed to hybridize to a target DNA sequence (and in some examples edit the target DNA sequence) and a region that hybridizes with trans-activating crRNA (tracrRNA), and (2) a scaffold sequence (tracrRNA) necessary for Cas-binding. Thus, a gRNA can combine a crRNA and a tracrRNA into a single RNA transcript (referred to in the art as a single guide RNA, sgRNA; for simplicity, encompassed by the term “gRNA” herein). In this example, a region of the crRNA hybridizes with tracrRNA to form a unique dual-RNA hybrid structure that binds Cas endonuclease proteins and guides the protein to a target DNA molecule. In some examples, the cRNA and tracrRNA are two separate RNA molecules (e.g., 2 piece gRNA; for simplicity, also encompassed by the term “gRNA” herein). In some examples, a protospacer adjacent motif (PAM) immediately follows the site of the target DNA to be edited (e.g., Cas9 cleavage site about 3 nt upstream of PAM).
[0285] In another example, a gRNA directs a Cas RNA nuclease (such as Cas 13d) to a target RNA. In one such example, the gRNA includes from 5’ to 3’ (1) a crRNA containing a direct repeat (DR) region and (2) a spacer, for example for Casl3a, Casl3c, and Cas 13d nucleases. In one example includes about 36nt of DR followed by about 28-32nt of spacer sequence. In another such example, the gRNA includes from 5’ to 3’ (1) a spacer and (2) a crRNA containing a DR region, for example for Cas 13b nuclease. In some examples, the gRNA is processed (truncated / modified) by a Cas RNA nuclease or other RNases into the shorter “mature” form. The DR is the constant portion of the gRNA, containing secondary structure which facilitates interaction between the Cas RNA nuclease protein and the gRNA. The spacer portion is theSLR 7158-102574-27 05 / 08 / 25 S2024-016A
[0286] variable portion of the gRNA, and includes a sequence designed to hybridize to a target RNA sequence (and in some examples edit the target RNA sequence). In some examples, the full length spacer is about 28-32nt (such as 30-32 nt) long while the mature (processed) spacer is about 14-30nt.
[0287] Hybridization: Hybridization of a nucleic acid occurs when two nucleic acid molecules undergo an amount of hydrogen bonding to each other. The stringency of hybridization can vary according to the environmental conditions surrounding the nucleic acids, the nature of the hybridization method, and the composition and length of the nucleic acids used. Calculations regarding hybridization conditions required for attaining particular degrees of stringency are discussed in Sambrook et al., Molecular Cloning: A Laboratory Manual (Cold Spring Harbor Laboratory Press, Cold Spring Harbor, NY, 2001); and Tijssen, Laboratory Techniques in Biochemistry and Molecular Biology — Hybridization with Nucleic Acid Probes Part I, Chapter 2 (Elsevier, New York, 1993). The Tmis the temperature at which 50% of a given strand of nucleic acid is hybridized to its complementary strand.
[0288] Increase or Decrease: A statistically significant positive or negative change, respectively, in quantity from a control value (such as a value representing no therapeutic agent, such as no administration of the two or more synthetic nucleic acid molecules provided herein). An increase is a positive change, such as an increase at least 50%, at least 100%, at least 200%, at least 300%, at least 400% or at least 500% as compared to the control value. A decrease is a negative change, such as a decrease of at least 20%, at least 25%, at least 50%, at least 75%, at least 80%, at least 90%, at least 95%, at least 98%, at least 99%, or at least 100% decrease as compared to a control value. In some examples the decrease is less than 100%, such as a decrease of no more than 90%, no more than 95%, or no more than 99%.
[0289] Isolated: An “isolated” biological component (such as a nucleic acid molecule or a protein) has been substantially separated, produced apart from, or purified away from other biological components in the cell or tissue of an organism in which the component occurs, such as other cells (e.g., RBCs), chromosomal and extrachromosomal DNA and RNA, and proteins. Nucleic acids and proteins that have been “isolated” include nucleic acids and proteins purified by standard purification methods. The term also embraces nucleic acids and proteins prepared by recombinant expression in a host cell as well as chemically synthesized nucleic acids and proteins.
[0290] Kissing loop / kissing stem loop: An RNA structure that forms when bases between two hairpin loops form pair interactions. These intermolecular “kissing interactions” occur when the unpaired nucleotides in one hairpin loop, base pair with the unpaired nucleotides in another hairpin loop to form a stable interaction complex. See FIGS. 9 A and 32 for examples. Kissing loop is also used to refer to either of the two hairpin loops that together can form the kissing interactions.
[0291] N-terminal portion: A region of a protein sequence that includes a contiguous stretch of amino acids that begins at the N-terminal residue of the protein. An N-terminal portion of the protein can be defined by a contiguous stretch of amino acids (e.g., a number of amino acid residues).
[0292] Non-naturally occurring, synthetic, or engineered: Terms used herein as interchangeably and indicate the involvement of the hand of man. The terms, when referring to nucleic acid molecules orSLR 7158-102574-27 05 / 08 / 25 S2024-016A
[0293] polypeptides indicate that the nucleic acid molecule or the polypeptide is at least substantially free from at least one other component with which they are naturally associated in nature and as found in nature. In addition, the terms can indicate that the nucleic acid molecules or polypeptides have a sequence not found in nature.
[0294] Nucleic acid molecule: A deoxyribonucleotide (DNA) or ribonucleotide (RNA) polymer, which can include natural nucleotides / ribonucleotides and / or analogues of natural nucleotides / ribonucleotides that hybridize to nucleic acid molecules in a manner similar to naturally occurring nucleotides. A nucleic acid molecule can be a single stranded (ss) DNA or RNA molecule or a double stranded (ds) nucleic acid molecule. RNA or mRNA as used herein may refer to a pre-mRNA molecule, or a mature RNA transcript. A pre-mRNA molecule comprises sequences to be removed by processing, e.g., intron sequences removed by splicing following binding of the dimerization domains described herein. Nucleic acid molecules described herein can be DNA molecules from which an RNA is transcribed from a promoter on the DNA, e.g., in the context of a DNA expression vector.
[0295] Operably linked: A first nucleic acid sequence is operably linked with a second nucleic acid sequence when the first nucleic acid sequence is placed in a functional relationship with the second nucleic acid sequence. For instance, a promoter sequence is operably linked to a nucleic acid sequence if the promoter affects the expression of the nucleic acid sequence, for example, the promoter effects transcription of a pre-mRNA, which when spliced may result in expression of a protein (such as a portion of a nucleic acid editing protein coding sequence).
[0296] Pharmaceutically acceptable carriers: The pharmaceutically acceptable carriers useful in this disclosure are conventional. Remington’s Pharmaceutical Sciences, by E. W. Martin, Mack Publishing Co., Easton, PA, 15th Edition (1975), describes compositions and formulations suitable for pharmaceutical delivery of a therapeutic agent, such as a nucleic acid molecule disclosed herein.
[0297] In general, the nature of the carrier will depend on the particular mode of administration being employed. For instance, parenteral formulations usually comprise injectable fluids that include pharmaceutically and physiologically acceptable fluids such as water, physiological saline, balanced salt solutions, aqueous dextrose, glycerol or the like as a vehicle. In addition to biologically-neutral carriers, pharmaceutical compositions to be administered can contain minor amounts of non-toxic auxiliary substances, such as wetting or emulsifying agents, preservatives, and pH buffering agents and the like, for example sodium acetate or sorbitan monolaurate.
[0298] Polypeptide, peptide and protein: Refer to polymers of amino acids of any length. The polymer may be linear or branched, it may include modified amino acids, and it may be interrupted by non-amino acids. The terms also encompass an amino acid polymer that has been modified; for example, disulfide bond formation, glycosylation, lipidation, acetylation, phosphorylation, or any other manipulation, such as conjugation with a labeling component. As used herein the term "amino acid" includes natural and / or unnatural or synthetic amino acids, including glycine and both the D or L optical isomers, and amino acid analogs and peptidomimetics. In one example, a protein is nucleic acid editing protein, such as Cas9,SLR 7158-102574-27 05 / 08 / 25 S2024-016A
[0299] Casl3d, or a zinc finger nuclease. In one example, a protein is one associated with disease, such as a genetic disease (e.g., see Table 1-4). In one example, a protein is a therapeutic protein, such as one used in the treatment of a disease, such as cancer. In one example a protein is at least 50 aa in length, at least 100 aa in length, at least 500 aa in length, at least 1000 aa in length, at least 1500 aa in length, such as at least 2000 aa, at least 2500 aa, at least 3000 aa, or at least 5000 aa.
[0300] Polypyrimidine tract: A region of pre-messenger RNA (mRNA) that promotes the assembly of the spliceosome, the protein complex specialized for carrying out RNA splicing during the process of post-transcriptional modification. This tract can be primarily pyrimidine nucleotides, such as uracil, and in some examples is 15-20 base pairs long, located about 5–40 base pairs before the 3' end of the intron to be spliced.
[0301] Promoter / Enhancer: An array of nucleic acid control sequences which direct transcription of a nucleic acid sequence. A promoter includes necessary nucleic acid sequences near the start site of transcription, such as, in the case of a polymerase II type promoter, a TATA element. A promoter also optionally includes distal enhancer or repressor elements which can be located as much as several thousand base pairs from the start site of transcription. In some examples a promoter sequence + its corresponding coding sequence is larger than the capacity for an AAV. In some examples a promoter sequence of a target protein is at least 3500 nt, at least 4000 nt, at least 5000 nt, or even at least 6000 nt.
[0302] A “constitutive promoter” is a promoter that is continuously active and is not subject to regulation by external signals or molecules. In contrast, the activity of an “inducible promoter” is regulated by an external signal or molecule (for example, a transcription factor). Both constitutive and inducible promoters can be used in the methods and systems provided herein (see e.g., Bitter et al., Methods in Enzymology 153:516-544, 1987). A tissue-specific promoter can be used in the methods and systems provided herein, for example to direct expression primarily in a desired tissue or cell of interest, such as muscle, neuron, bone, skin, blood, specific organs (e.g., liver, pancreas), or particular cell types (e.g., lymphocytes). In some examples, a promoter used herein is endogenous to the target protein expressed. In some examples, a promoter used herein is exogenous to the target protein expressed.
[0303] Also included are promoter elements which are sufficient to render promoter-dependent gene expression controllable for cell-type specific, tissue-specific, or inducible by external signals or agents; such elements may be located in the 5' or 3' regions of the gene. Promoters produced by recombinant DNA or synthetic techniques can also be used to provide for transcription of the nucleic acid sequences.
[0304] Exemplary promoters that can be used with the methods and systems provided herein include, but are not limited to an SV40 promoter, cytomegalovirus (CMV) promoter (optionally with the CMV enhancer), a pol III promoter (e.g., U6 and Hl promoters), a pol II promoter (e.g., the retroviral Rous sarcoma virus (RSV) LTR promoter (optionally with the RSV enhancer), the dihydrofolate reductase promoter, the β-actin promoter, the phosphoglycerol kinase (PGK) promoter, and the EFla promoter).
[0305] Recombinant: A recombinant nucleic acid molecule or protein sequence is one that has a sequence that is not naturally occurring or has a sequence that is made by an artificial combination of two otherwiseSLR 7158-102574-27 05 / 08 / 25 S2024-016A
[0306] separated segments of sequence (e.g., a viral vector that includes a portion of a nucleic acid editing protein coding sequence, such as about a third, half, or two-thirds of a coding sequence). This artificial combination can be accomplished by, for example, chemical synthesis or the artificial manipulation of isolated segments of nucleic acids, such as by genetic engineering techniques. Similarly, a recombinant or transgenic cell is one that contains a recombinant nucleic acid molecule.
[0307] RNA Editing: A type of genetic engineering in which a RNA molecule (or ribonucleotides of the RNA) is inserted, deleted or replaced in a cell or organism using engineered nucleases (such as the Casl3d and dCasl3d proteins), which create site-specific strand breaks at desired locations in the RNA. The induced breaks are repaired resulting in targeted mutations or repairs. CRISPR / Cas methods, for example using the REJ systems provided herein to express a Cas nuclease, can be used to edit the sequence of one or more target RNAs, such as one associated with cancer (e.g., breast cancer, colon cancer, lung cancer, prostate cancer, melanoma), infectious disease (such as HIV, hepatitis, HPV, and West Nile virus), or neurodegenerative disorder (e.g., Huntington’s disease or ALS). For example, RNA editing can be used to treat a disease or viral infection.
[0308] RNA insertion site: A site of the RNA that is targeted for, or has undergone, insertion of an exogenous polynucleotide or polyribonucleotide. The disclosed methods include use of a nucleic acid editor expressed from two or more nucleic acid molecules provided herein, which can be used to target a RNA for manipulation at an RNA insertion site.
[0309] Sequence identity: The similarity between amino acid (or nucleotide) sequences is expressed in terms of the similarity between the sequences, otherwise referred to as sequence identity. Sequence identity is frequently measured in terms of percentage identity (or similarity or homology); the higher the percentage, the more similar the two sequences are.
[0310] Methods of alignment of sequences for comparison are known. Various programs and alignment algorithms are described in: Smith and Waterman, Adv. Appl. Math. 2:482, 1981; Needleman and Wunsch, J. Mol. Biol. 48:443, 1970; Pearson and Lipman, Proc. Natl. Acad. Sci. U. S. A. 85:2444, 1988; Higgins and Sharp, Gene 73:237, 1988; Higgins and Sharp, CABIOS 5:151, 1989; Corpet et al., Nucleic Acids Research 16:10881, 1988; and Pearson and Lipman, Proc. Natl. Acad. Sci. U. S. A. 85:2444, 1988. Altschul et al., Nature Genet. 6:119, 1994, presents a detailed consideration of sequence alignment methods and homology calculations.
[0311] The NCBI Basic Local Alignment Search Tool (BLAST) (Altschul et al., J. Mol. Biol. 215:403, 1990) is available from several sources, including the National Center for Biotechnology Information (NCBI, Bethesda, MD) and on the internet, for use in connection with the sequence analysis programs blastp, blastn, blastx, tblastn and tblastx. A description of how to determine sequence identity using this program is available on the NCBI website on the internet.
[0312] Variants of a native protein or coding sequence (such as a DMD, factor 8, factor 9, or AABCA4 sequence) are typically characterized by possession of at least about 80%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98% or at least 99% sequence identity counted over the full length alignment withSLR 7158-102574-27 05 / 08 / 25 S2024-016A
[0313] the amino acid sequence using the NCBI Blast 2.0, gapped blastp set to default parameters. For comparisons of amino acid sequences of greater than about 30 amino acids, the Blast 2 sequences function is employed using the default BLOSUM62 matrix set to default parameters, (gap existence cost of 11, and a per residue gap cost of 1). When aligning short peptides (fewer than around 30 amino acids), the alignment should be performed using the Blast 2 sequences function, employing the PAM30 matrix set to default parameters (open gap 9, extension gap 1 penalties). Proteins with even greater similarity to the reference sequences will show increasing percentage identities when assessed by this method, such as at least 95%, at least 98%, or at least 99% sequence identity. When less than the entire sequence is being compared for sequence identity, homologs and variants will typically possess at least 80% sequence identity over short windows of 10-20 amino acids, and may possess sequence identities of at least 85% or at least 90% or at least 95% depending on their similarity to the reference sequence. Methods for determining sequence identity over such short windows are available at the NCBI website on the internet. These sequence identity ranges are provided for guidance only; it is possible that strongly significant homologs could be obtained that fall outside of the ranges provided.
[0314] Variants of the disclosed nucleic acid sequences (such as synthetic intron sequences and coding sequences) are typically characterized by possession of at least about 80%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98% or at least 99% sequence identity counted over the full length alignment with the nucleic acid sequence using the NCBI Blast 2.0, gapped blastn set to default parameters. One of skill in the art will appreciate that these sequence identity ranges are provided for guidance only; it is possible that functional sequences could be obtained that fall outside of the ranges provided.
[0315] Subject: A mammal, for example a human. Mammals include, but are not limited to, murines, simians, humans, farm animals, sport animals, and pets. In one embodiment, the subject is a non-human mammalian subject, such as a monkey or other non-human primate, mouse, rat, rabbit, pig, goat, sheep, dolphin, dog, cat, horse, or cow. In some examples, the subject is a laboratory animal / organism, such as a mouse, rabbit, or rat. In some examples, the subject treated using the methods disclosed herein is a human.
[0316] In some examples, the subject has genetic disease, such as one listed in Tables 1-4 or USH1F, that can be treated using the methods disclosed herein. In some examples, the subject treated using the methods disclosed herein is a human subject having a genetic disease. In some examples, the subject treated using the methods disclosed herein is a human subject having cancer. In some examples, the subject treated using the methods disclosed herein is a human subject having an infection, such as a bacterial or viral infection.
[0317] Target nucleic acid: A nucleic acid molecule, such as a DNA or RNA sequence, such as a gene, having a sequence that is to be altered and / or whose expression is to be modulated. In some examples, a target nucleic acid molecule is one nucleic acid molecule, two or more portions of the same nucleic acid molecule (e.g., same RNA or gene), two or more different nucleic acid molecules (e.g., two different genes or two different RNAs), or two or more portions of the two or more different nucleic acid molecules. The target nucleic acid molecule can include one or more target editing sites, that is a regions of the target nucleic acid molecule (such as one or more nucleotides or ribonucleotides, such as at least 10, at least 15, atSLR 7158-102574-27 05 / 08 / 25 S2024-016A
[0318] least 20, or at least 30 consecutive nucleotides or ribonucleotides of the target) to be altered, such as substituted, deleted, or where an insertion is to be made. In some examples, a target nucleic acid molecule is one whose expression is to be modulated, such as an increase or decrease in expression of the gene product (e.g., protein). In one example, a target nucleic acid molecule is a DNA, RNA, or gene whose activated expression is desired. In one example, a target nucleic acid molecule is a DNA, RNA, or gene whose reduced or abolished expression is desired. In one example, a target nucleic acid molecule is a DNA, RNA, or gene having one or more point mutations that results in disease (such as one listed in Tables 1-4 or USH1F). In some examples, a targeting sequence (for example of a gRNA) has complementarity to the target gene / nucleic acid. In other examples, a targeting sequence (for example of a gRNA) has complementarity to a promoter and / or regulatory element of a target nucleic acid molecule. In some examples, the target nucleic acid sequence is DNA, and is present immediately adjacent to a Protospacer Adjacent Motif (PAM). In some examples, the target nucleic acid sequence is unique as compared to other nucleic acid sequences in the cell.
[0319] Targeting sequence: The portion of a gRNA having complementarity with a target nucleic acid sequence. In some examples, the targeting sequence has complementarity to a promoter or regulatory element of a target nucleic acid whose activated or repressed expression is desired. In some examples, the targeting sequence of a gRNA is about 14-30 nt and has sufficient complementarity with a target nucleic acid sequence to hybridize with the target sequence and direct sequence-specific binding of a Cas nuclease to the target nucleic acid sequence. In some embodiments, the degree of complementarity between a targeting sequence of an gRNA and its corresponding target nucleic acid, when optimally aligned using a suitable alignment algorithm, is about or more than about 50%, 60%, 75%, 80%, 85%, 90%, 95%, 97.5%, 98%, 99%, or more. In some embodiments, the degree of complementarity is 100%. Optimal alignment may be determined with the use of any suitable algorithm for aligning sequences, non-limiting examples of which include the Smith-Waterman algorithm, the Needleman-Wunsch algorithm, algorithms based on the Burrows-Wheeler Transform (e.g., the Burrows Wheeler Aligner), ClustalW, Clustal X, BLAT, Novoalign (Novocraft Technologies, ELAND (Illumina, San Diego, Calif.), SOAP (available at soap.genomics.org.cn), and Maq (available at maq.sourceforge.net).
[0320] Therapeutic agent: Refers to one or more molecules or compounds that confer some beneficial effect upon administration to a subject. The disclosed synthetic nucleic acid molecules and systems provided herein are therapeutic agents. The beneficial therapeutic effect can include enablement of diagnostic determinations; amelioration of a disease, symptom, disorder, or pathological condition; reducing or preventing the onset of a disease, symptom, disorder or condition; and generally counteracting a disease, symptom, disorder or pathological condition.
[0321] Transcriptional activator: A protein or protein domain that increases transcription of a nucleic acid molecule, such as a gene. Such proteins can be used in the compositions, systems and methods provided herein. Such proteins and proteins domains can have a DNA binding domain and a domain for activation of transcription. These activators can be introduced into the system through attachment to a CasSLR 7158-102574-27 05 / 08 / 25 S2024-016A
[0322] nuclease or gRNA. Examples of such activators include VP64, p65, myogenic differentiation 1 (MyoDl), heat shock transcription factor (HSF) 1, RTA, CBP, SET7 / 9, or any combination thereof (such as p65 and HSF1).
[0323] Transduced, Transformed and Transfected: A virus or vector “transduces” a cell when it transfers nucleic acid molecules into a cell. A cell is “transformed” or “transfected” by a nucleic acid transduced into the cell when the nucleic acid becomes stably replicated by the cell, either by incorporation of the nucleic acid into the cellular genome, or by episomal replication.
[0324] These terms encompasses all techniques by which a nucleic acid molecule can be introduced into such a cell, including transfection with viral vectors, transformation with plasmid vectors, and introduction of naked DNA by electroporation, lipofection, particle gun acceleration and other methods in the art. In some example the method is a chemical method (e.g., calcium-phosphate transfection), physical method (e.g., electroporation, microinjection, particle bombardment), fusion (e.g., liposomes), receptor-mediated endocytosis (e.g., DNA-protein complexes, viral envelope / capsid-DNA complexes) and biological infection by viruses such as recombinant viruses (Wolff, J. A., ed, Gene Therapeutics, Birkhauser, Boston, USA, 1994). Methods for the introduction of nucleic acid molecules into cells are known (e.g., see U. S. Patent No. 6,110,743). These methods can be used to transduce a cell with the disclosed nucleic acid molecules.
[0325] Transgene: An exogenous gene, for example supplied by a vector, such as AAV. In one example, a transgene encodes a portion of a nucleic acid editing protein, such as about a third, half, or two-thirds of a nucleic acid editing protein, for example operably linked to a promoter sequence. In one example, a transgene includes a portion of a Cas nuclease coding sequence, such as about a third, half, or two-thirds of a Cas nuclease coding sequence, for example operably linked to a promoter sequence.
[0326] Treating, Treatment, and Therapy: Any success or indicia of success in the attenuation or amelioration of an injury, pathology or condition, including any objective or subjective parameter such as abatement, remission, diminishing of symptoms or making the condition more tolerable to the patient, slowing in the rate of degeneration or decline, making the final point of degeneration less debilitating, improving a subject’s physical or mental well-being, or prolonging the length of survival. The treatment may be assessed by objective or subjective parameters; including the results of a physical examination, blood and other clinical tests, and the like. In some examples, treatment with the disclosed methods results in a decrease in the number or severity of symptoms associated with a genetic disease, such as increasing the survival time of a treated patient with the genetic disease.
[0327] In some examples, treatment with the disclosed methods results in a decrease in the number or severity of symptoms associated with DMD or other genetic disease, such as increasing survival, increasing the mobility (e.g., walking, climbing), improving cognitive ability, reducing calf muscle size, reduce cardiomyopathy, improving vision, improving hearing, improving blood clotting, or improve respiratory function. In some examples, combinations of these effects are achieved.
[0328] Tumor, neoplasia, malignancy or cancer: A neoplasm is an abnormal growth of tissue or cells which results from excessive cell division. Neoplastic growth can produce a tumor. The amount of a tumorSLR 7158-102574-27 05 / 08 / 25 S2024-016A
[0329] in an individual is the “tumor burden” which can be measured as the number, volume, or weight of the tumor. A tumor that does not metastasize is referred to as “benign.” A tumor that invades the surrounding tissue and / or can metastasize is referred to as “malignant.” A “non-cancerous tissue” is a tissue from the same organ wherein the malignant neoplasm formed, but does not have the characteristic pathology of the neoplasm. Generally, noncancerous tissue appears histologically normal. A “normal tissue” is tissue from an organ, wherein the organ is not affected by cancer or another disease or disorder of that organ. A “cancer-free” subject has not been diagnosed with a cancer of that organ and does not have detectable cancer.
[0330] Exemplary tumors, such as cancers, that can be treated with the disclosed methods and systems include solid tumors, such as breast carcinomas (e.g. lobular and duct carcinomas), sarcomas, carcinomas of the lung (e.g., non-small cell carcinoma, large cell carcinoma, squamous carcinoma, and adenocarcinoma), mesothelioma of the lung, colorectal adenocarcinoma, stomach carcinoma, prostatic adenocarcinoma, ovarian carcinoma (such as serous cystadenocarcinoma and mucinous cystadenocarcinoma), ovarian germ cell tumors, testicular carcinomas and germ cell tumors, pancreatic adenocarcinoma, biliary adenocarcinoma, hepatocellular carcinoma, bladder carcinoma (including, for instance, transitional cell carcinoma, adenocarcinoma, and squamous carcinoma), renal cell adenocarcinoma, endometrial carcinomas (including, e.g., adenocarcinomas and mixed Mullerian tumors (carcinosarcomas)), carcinomas of the endocervix, ectocervix, and vagina (such as adenocarcinoma and squamous carcinoma of each of same), tumors of the skin (e.g., squamous cell carcinoma, basal cell carcinoma, malignant melanoma, skin appendage tumors, Kaposi sarcoma, cutaneous lymphoma, skin adnexal tumors and various types of sarcomas and Merkel cell carcinoma), esophageal carcinoma, carcinomas of the nasopharynx and oropharynx (including squamous carcinoma and adenocarcinomas of same), salivary gland carcinomas, brain and central nervous system tumors (including, for example, tumors of glial, neuronal, and meningeal origin), tumors of peripheral nerve, soft tissue sarcomas and sarcomas of bone and cartilage, and lymphatic tumors (including B-cell and T- cell malignant lymphoma). In one example, the tumor is an adenocarcinoma.
[0331] The methods and systems can also be used to treat liquid tumors, such as a lymphatic, white blood cell, or other type of leukemia. In a specific example, the tumor treated is a tumor of the blood, such as a leukemia (for example acute lymphoblastic leukemia (ALL), chronic lymphocytic leukemia (CLL), acute myelogenous leukemia (AML), chronic myelogenous leukemia (CML), hairy cell leukemia (HCL), T-cell prolymphocytic leukemia (T-PLL), large granular lymphocytic leukemia, and adult T-cell leukemia), lymphomas (such as Hodgkin’s lymphoma and non-Hodgkin’s lymphoma), and myelomas).
[0332] Upregulated: When used in reference to the expression of a molecule, such as a target nucleic acid / protein, refers to any process which results in an increase in production of the target nucleic acid / protein. In some examples, upregulation or activation of a target RNA includes processes that increase translation of the target RNA and thus can increase the presence of corresponding proteins. The upregulated molecule may be a target nucleic acid or protein that is expressed from the nucleic acid molecules of theSLR 7158-102574-27 05 / 08 / 25 S2024-016A
[0333] composition and methods described herein, e.g., a nucleic acid editor protein produced from recombined transcripts in a REJ split system; an edited target nucleic acid and / or the resulting protein produced therefrom; or a representative marker, surrogate, or functional indicator of a target nucleic acid or protein or an edited target nucleic acid and / or the resulting protein.
[0334] Upregulation includes any detectable increase in target nucleic acid / protein. In certain examples, detectable target nucleic acid / protein expression in a cell or cell free system increases by at least 20%, at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, at least 75%, at least 80%, at least 90%, at least 95%, at least 100%, at least 200%, at least 400%, or at least 500% as compared to a control. A control may be an amount of target nucleic acid / protein detected in a corresponding sample not treated with a nucleic acid molecule provided herein. In one example, a control is a relative amount of expression in a normal cell (e.g., a non-recombinant cell that does not include a system provided herein). A control can be compared with: a target nucleic acid or protein that is expressed from the nucleic acid molecules of the composition and methods described herein, e.g., a nucleic acid editor protein produced from recombined transcripts in a REJ split system; an edited target nucleic acid and / or the resulting protein produced therefrom; or a representative marker, surrogate, or functional indicator of a target nucleic acid or protein or an edited target nucleic acid and / or the resulting protein. A control used for comparison may be any appropriate control as determined by one of skill in the art. A positive control may be an amount of a corresponding mRNA and / or protein produced from a full-length construct. A negative control may be an amount of a corresponding mRNA and / or protein produced from an empty or otherwise defective construct. A control can be compared with a target nucleic acid or protein produced in a cell in the presence of a nucleic acid editor protein expressed from the nucleic acid molecules using the compositions and methods provided herein. As described herein, when a nucleic acid editing protein is expressed in a cell from the nucleic acid molecules using the present compositions and methods, the recombined nucleic acid editing protein may edit its target nucleic acid. The edited nucleic acid and / or a protein produced therefrom may be compared to a positive control, which may be an amount of corresponding nucleic acid, e.g., mRNA, and / or protein produced from the target nucleic acid in a normal cell (e.g., a non-recombinant “normal” cell that does not have the target nucleic acid in need of editing and that does not include a system provided herein). The edited nucleic acid and / or a protein produced therefrom may be compared to a negative control, which may be an amount of corresponding nucleic acid, e.g., mRNA, and / or protein produced from the target nucleic acid in an untreated cell (e.g., a non-recombinant “mutant” cell having the target nucleic acid in need of editing and that does not include a system provided herein). The detectable target nucleic acid and / or protein expression relative to the detectable positive control target nucleic acid and / or protein expression (e.g., cell or cell free system), respectively, may be 20% to 500% of the expression in a positive control. The detectable target nucleic acid and / or protein expression in a cell or cell free system relative to expression in a positive control may be about 20% to about 500%. The detectable target nucleic acid and / or protein expression in a cell or cell free system relative to expression in a positive control may be about 20% to about 30%, about 20% to about 40%, about 20% to about 50%, about 20% to about 60%, about 20% toSLR 7158-102574-27 05 / 08 / 25 S2024-016A
[0335] about 70%, about 20% to about 90%, about 20% to about 95%, about 20% to about 100%, about 20% to about 150%, about 20% to about 200%, about 20% to about 500%, about 30% to about 40%, about 30% to about 50%, about 30% to about 60%, about 30% to about 70%, about 30% to about 90%, about 30% to about 95%, about 30% to about 100%, about 30% to about 150%, about 30% to about 200%, about 30% to about 500%, about 40% to about 50%, about 40% to about 60%, about 40% to about 70%, about 40% to about 90%, about 40% to about 95%, about 40% to about 100%, about 40% to about 150%, about 40% to about 200%, about 40% to about 500%, about 50% to about 60%, about 50% to about 70%, about 50% to about 90%, about 50% to about 95%, about 50% to about 100%, about 50% to about 150%, about 50% to about 200%, about 50% to about 500%, about 60% to about 70%, about 60% to about 90%, about 60% to about 95%, about 60% to about 100%, about 60% to about 150%, about 60% to about 200%, about 60% to about 500%, about 70% to about 90%, about 70% to about 95%, about 70% to about 100%, about 70% to about 150%, about 70% to about 200%, about 70% to about 500%, about 90% to about 95%, about 90% to about 100%, about 90% to about 150%, about 90% to about 200%, about 90% to about 500%, about 95% to about 100%, about 95% to about 150%, about 95% to about 200%, about 95% to about 500%, about 100% to about 150%, about 100% to about 200%, about 100% to about 500%, about 150% to about 200%, about 150% to about 500%, or about 200% to about 500%. The detectable target nucleic acid and / or protein expression in a cell or cell free system relative to expression in a positive control may be about 20%, about 30%, about 40%, about 50%, about 60%, about 70%, about 90%, about 95%, about 100%, about 150%, about 200%, or about 500%. The detectable target nucleic acid and / or protein expression in a cell or cell free system relative to expression of a positive control may be at least about 20%, about 30%, about 40%, about 50%, about 60%, about 70%, about 90%, about 95%, about 100%, about 150%, or about 200%. The detectable target nucleic acid and / or protein expression in a cell or cell free system relative to expression in a positive control may be at most about 30%, about 40%, about 50%, about 60%, about 70%, about 90%, about 95%, about 100%, about 150%, about 200%, or about 500%.
[0336] The edited nucleic acid and / or a protein produced therefrom may be compared to a negative control, which may be an amount of corresponding nucleic acid, e.g., mRNA, and / or protein produced from the target nucleic acid in an untreated cell (e.g., a non-recombinant “mutant” cell having the target nucleic acid in need of editing and that does not include a system provided herein). The detectable target nucleic acid and / or protein expression relative to the detectable negative control target nucleic acid and / or protein expression, respectively, may be increased by 20% to 500%. The detectable target nucleic acid / protein expression in a cell or cell free system may increase relative to a negative control by about 20% to about 30%, about 20% to about 40%, about 20% to about 50%, about 20% to about 60%, about 20% to about 70%, about 20% to about 90%, about 20% to about 95%, about 20% to about 100%, about 20% to about 150%, about 20% to about 200%, about 20% to about 500%, about 30% to about 40%, about 30% to about 50%, about 30% to about 60%, about 30% to about 70%, about 30% to about 90%, about 30% to about 95%, about 30% to about 100%, about 30% to about 150%, about 30% to about 200%, about 30% to about 500%, about 40% to about 50%, about 40% to about 60%, about 40% to about 70%, about 40% to about 90%,SLR 7158-102574-27 05 / 08 / 25 S2024-016A
[0337] about 40% to about 95%, about 40% to about 100%, about 40% to about 150%, about 40% to about 200%, about 40% to about 500%, about 50% to about 60%, about 50% to about 70%, about 50% to about 90%, about 50% to about 95%, about 50% to about 100%, about 50% to about 150%, about 50% to about 200%, about 50% to about 500%, about 60% to about 70%, about 60% to about 90%, about 60% to about 95%, about 60% to about 100%, about 60% to about 150%, about 60% to about 200%, about 60% to about 500%, about 70% to about 90%, about 70% to about 95%, about 70% to about 100%, about 70% to about 150%, about 70% to about 200%, about 70% to about 500%, about 90% to about 95%, about 90% to about 100%, about 90% to about 150%, about 90% to about 200%, about 90% to about 500%, about 95% to about 100%, about 95% to about 150%, about 95% to about 200%, about 95% to about 500%, about 100% to about 150%, about 100% to about 200%, about 100% to about 500%, about 150% to about 200%, about 150% to about 500%, or about 200% to about 500%. The detectable target nucleic acid / protein expression in a cell or cell free system may increase relative to a negative control by about 20%, about 30%, about 40%, about 50%, about 60%, about 70%, about 90%, about 95%, about 100%, about 150%, about 200%, or about 500%. The detectable target nucleic acid / protein expression in a cell or cell free system may increase relative to a negative control by at least about 20%, about 30%, about 40%, about 50%, about 60%, about 70%, about 90%, about 95%, about 100%, about 150%, or about 200%. The detectable target nucleic acid / protein expression in a cell or cell free system may increase relative to a negative control by at most about 30%, about 40%, about 50%, about 60%, about 70%, about 90%, about 95%, about 100%, about 150%, about 200%, or about 500%.
[0338] Under conditions sufficient for: A phrase that is used to describe any environment that permits a desired activity. In one example the desired activity is increased expression or activity of a protein needed to treat a disease. In one example the desired activity is decreased expression or activity of a protein needed to treat a disease. In one example the desired activity is expression of a corrected protein sequence needed to treat a disease. In one example the desired activity is treatment of or slowing the progression of a genetic disease such as DMD (or other genetic disease listed in Tables 1-4 or USH1F) in vivo, for example using the disclosed methods and systems.
[0339] Vector: A nucleic acid molecule into which a foreign nucleic acid molecule can be introduced without disrupting the ability of the vector to replicate and / or integrate in a host cell. Vectors include, but are not limited to, nucleic acid molecules that are single- stranded, double-stranded, or partially doublestranded; nucleic acid molecules that comprise one or more free ends, no free ends (e.g., circular); nucleic acid molecules that comprise DNA, RNA, or both; and other varieties of polynucleotides.
[0340] A vector can include nucleic acid sequences that permit it to replicate in a host cell, such as an origin of replication. A vector can also include one or more selectable marker genes and other genetic elements. An integrating vector is capable of integrating itself into a host nucleic acid. An expression vector is a vector that contains the necessary regulatory sequences to allow transcription and translation of inserted gene or genes.SLR 7158-102574-27 05 / 08 / 25 S2024-016A
[0341] One type of vector is a "plasmid," which refers to a circular double stranded DNA loop into which additional DNA segments can be inserted, such as by standard molecular cloning techniques. Another type of vector is a viral vector, wherein virally-derived DNA or RNA sequences are present in the vector for packaging into a virus (e.g., retroviruses, replication defective retroviruses, adenoviruses, replication defective adenoviruses, and adeno-associated viruses). Viral vectors also include polynucleotides carried by a virus for transfection into a host cell. In some embodiments, the vector is a lentivirus (such as an integration-deficient lentiviral vector) or adeno-associated viral (AAV) vector.
[0342] In some embodiments, the vector is an AAV, such as AAV serotypes AAV2, AAV8, AAV2 / 5, AAV9 or AAVrh.10. In some embodiments, the vector is one that can penetrate the blood-brain barrier, for example following intravenous administration. The adeno-associated virus serotype rh.10 (AAV.rhlO) vector partially penetrates the blood-brain barrier, providing high levels and spread of transgene expression.
[0343] II. Overview of Several Embodiments
[0344] One approach to curing patients who suffer from genetic diseases is gene editing therapy (generally referred to as gene therapy). In such an approach, the defective gene is replaced by an intact version of it, delivered through e.g., a viral vector, which achieves sustained expression from months to years. Although adeno associated viruses (AAVs) have been used for clinical gene replacement therapy, they have a limited packaging capacity (e.g., about less than 5 kb). Thus, strategies to overcome this packaging limitation are needed to achieve gene replacement of genes that exceed the about 5 kb size limit. For example some promoters alone, coding sequences alone, or the combined promoter + coding sequence, exceed the about 5 kb size limit of an AAV. Thus, such proteins encoded by such promoters and coding sequences can be expressed using the disclosed systems.
[0345] Several methods of nucleic acid editing, such as editing of target DNA or RNA molecules, such as gene editing, can be used to upregulate or downregulate expression of a target, as well as correct mutations in a target. Examples of such methods include CRISPR / Cas methods of editing DNA (e.g., using Cas9 or dCas9 DNA endonucleases), CRISPR / Cas methods of editing RNA (e.g., using Casl3d or dCas13d RNA endonucleases), zinc finger nuclease methods of genome editing (e.g., using zinc finger nucleases which include a zinc finger DNA-binding domain and a DNA-cleavage domain), and transcription activator-like effector nucleases (TALENs) based methods of genome editing, (e.g., using transcription activator-like effector nuclease (TALEN) proteins). All of such methods rely on the use of a nucleic acid editing protein, which is a nuclease that can insert, delete, and / or transverse a target nucleic acid sequence (such as a target DNA or RNA sequence) in a cell. Thus, a nucleic acid editing protein can affect the insertion, deletion, and / or substitution of one or more selected nucleotides or ribonucleotides in a target DNA or RNA sequence. However, as discussed above, the cargo limitations of vectors, such as AAV, can make it difficult to produce adequate levels of nucleic acid editing proteins (and in some examples also corresponding gRNAs) to treat disease.SLR 7158-102574-27 05 / 08 / 25 S2024-016A
[0346] Splicing mediated recombination of two RNA molecules using naturally occurring intron sequences for one or both of the RNA fragments is inefficient. First, these natural intron sequences are sequences from naturally occurring introns and are comprised of a mix of all four RNA nucleotides. Such sequences tend to fold up into structures that can obstruct trans-interaction by forming strong intramolecular base pairs rather than being available for intermolecular interactions. Second, these naturally occurring intron sequences have not evolved to strongly attract the spliceosome components, since exon rather than introns drive the exon definition in higher eukaryotes. These two limitations of previous strategies are addressed herein by designing synthetic intronic sequences that are not found in nature. These synthetic sequences contain elements that strongly attract and stimulate spliceosome recruitment on the one hand while minimizing the secondary structure (and in some examples other structure, such as tertiary structure) that obstructs bringing the two RNA fragments together.
[0347] The inventors developed a novel nucleic acid based element that can be used to efficiently reconstitute the coding sequence of large genes from multiple serial fragments. The compositions, systems and method provided herein allow reconstitution of a full-length RNA from the multiple serial fragments. Reconstitution of a full-length RNA, e.g., a messenger RNA transcript, in turn leads to production of the full-length “reconstituted” protein. The disclosed methods and systems differ from prior methods. The disclosed highly efficient synthetic introns utilize an optimal arrangement of RNA elements (or DNA encoding these elements) that efficiently drive the RNA splicing reaction between non-covalently linked RNAs (pre-mRNAs). The method / system is a significant advancement over previous attempts to harness trans- splicing because it generates high levels of functional nucleic acid editing protein that more closely approximate the therapeutic levels of a nucleic acid editing protein to edits a target nucleic acid molecule to treat genetic diseases. The innovation is based on selecting non-natural RNA domains that inherently are incapable of forming strong cis-binding interactions that interfere with trans-interactions with a second RNA having a complementary strand (also having inherently low cis-binding capacity). Intermolecular interactions between binding regions of partnered dimerization domains, e.g., a first dimerization domain and a second dimerization domain, a second dimerization domain and a third dimerization domain, a third dimerization domain and a fourth dimerization domain, a fourth dimerization domain and a fifth dimerization domain, or a fifth dimerization domain and a sixth dimerization domain, may be stronger than intramolecular interactions between a binding region on a single dimerization domain with other sequences in the same dimerization domain. For example, a single stranded kissing loop structure in a kissing loop first dimerization domain may more strongly bind to, hybridize, or associate with its complementary kissing loop structure on a second kissing loop dimerization domain than to other sequences on the same dimerization domain. It is understood that there can be regions within the same dimerization domain that bind more strongly intramolecularly than intermolecularly, for example the stem of a stem-loop structure. The resulting dimerization domains comprise stably available single-stranded binding regions that bind selectively and efficiently to their intended target, i.e., a complementary binding region on a partner dimerization domain. This strategy allows unprecedented reconstitution efficiencies of transcriptsSLR 7158-102574-27 05 / 08 / 25 S2024-016A
[0348] synthesized from separate templates, resulting in high levels of target protein or therapeutic nucleic acid production. Many examples of such optimized synthetic nucleic acid molecules and methods for their use are provided herein. These optimized dimerization domains and / or synthetic introns can include non-natural sequences (e.g., sequences not found in human cells and / or not found in another biological system) used in combination with optimized motifs that facilitate RNA splicing (including splice donor, splice acceptor, splice enhancer, and splice branch point sequences). A synthetic nucleic acid can be a non-natural nucleic acid sequence, e.g., a sequence not found in human cells and / or not found in another biological system). By optimizing the trans-dimerization of the RNA strands in the context of the appropriate RNA motifs that mediate efficient splicing, it is demonstrated herein for the first time that two or three different RNAs can be precisely and efficiently covalently linked in the same cell producing high levels of functional nucleic acid editing proteins in vivo and in vitro. Unlike the “hybrid” approach that provides an inefficient combination at the DNA level via DNA recombination that is ultimately followed by RNA splicing in cis to excise the DNA recombination site from the mature transcript, the disclosed method / system promotes a more efficient reaction in which two protein coding RNA fragments are joined together on the pre-mRNA level with less risk of producing recombination products that encode non-functional and / or deleterious products.
[0349] The data demonstrate that by using efficient synthetic RNA-dimerization and recombination domains (sRdR domains, also referred to as RNA end-joining (REJ) domains), a nucleic acid editing protein can be efficiently produced by reconstitution of its full-length mRNA from two separate gene fragments expressed from two separate nucleic acid constructs in the same cell. A desired guide RNA, e.g., a gRNA, can be expressed from one of the two constructs, or from a different construct. The disclosed methods and systems can be used to reconstitute transcripts encoding large genes like ABE8e, in order to edit any target nucleic acid to treat any genetic disease. Based on these observations, any genetic disease can be treated, such as ones benefiting from expression of a nucleic acid editing protein (e.g., see disorders listed in Tables 1-4 or USH1F). Other diseases that can be treated using nucleic acid editing proteins include cancer and infectious diseases (such as a bacterial or viral infection). Other applications include research and biotechnology applications.
[0350] In some embodiments, the disclosure provides methods for using the nucleic acid editing compositions and systems of the disclosure for treating a subject in need thereof by altering a target nucleic acid sequence in a cell of the subject. Uses for nucleic acid editing are described in the literature, e.g., in U. S. Pat. App. No. 2020 / 392473, “Novel CRISPR enzymes and systems,” incorporated herein by reference in its entirety.
[0351] In some embodiments, the compositions, systems, and methods described herein are used for treating a subject having a disease or disorder by repairing a nucleic acid mutation that causes defective or decreased production of a protein or another gene product, e.g., an RNA. In some embodiments, the defective protein or gene product has reduced activity, is nonfunctional, or is toxic. In some embodiments, the mutation is in a coding or noncoding region of the gene. In some embodiments, the disease or disorder is caused by a mutation in a coding region of a gene that results in a nonfunctional or otherwise defectiveSLR 7158-102574-27 05 / 08 / 25 S2024-016A
[0352] gene product, and repair of the mutation restores the expression of the functional gene product. In some embodiments, the disease or disorder is caused by a mutation in a regulatory region of a gene, for example, a promoter region or splicing element, and repair of the mutation restores the normal or a desirable level expression of the gene product by upregulation or downregulation. In some embodiments, mutations can be introduced to alter splicing to skip exons, thereby to restoring a normal or desirable level of a functional gene product. In some embodiments, the level of functional gene product is modulated in comparison to a control, e.g., an untreated control.
[0353] In some embodiments, a disease gene having a loss-of-function mutation is a disease gene listed in Table 1, reproduced from Table 2 of Chen and Altman, 2017, “Opportunities for developing therapies for rare genetic diseases: focus on gain of function and allostery,” Orphanet Journal of Rare Diseases 12:61, incorporated herein by reference. In some embodiments, the nucleic acid editing compositions, systems, and methods described herein are used to treat a disease listed in Table 1.
[0354] Table 1. Genetic Diseases Resulting from Loss-of-Function Mutations
[0355]
[0356] Another exemplary disease resulting from a loss-of function mutation is Usher Syndrome IF (USH1F). Usher Syndrome IF (USH1F) is a debilitating disease where patients experience congenital, bilateral, profound sensorineural hearing loss, vestibular areflexia, and adolescent-onset retinitis pigmentosa (RP). USH1F is caused by mutations in protocadherin-15 (PCDH15). Thus, the disclosed REJ systems and methods can be used to express a full-length PCDH15 protein to treat USH1F (e.g., see FIGS. 31 and 36). PCDH15 is a cell membrane binding protein that is primarily present in stereocilia and the retina and its coding sequence is approximately 5.8 kb, which exceeds the maximum packaging capacity of AAVs. To overcome the size limitations of AAVs for delivery of PCDH15 in a gene replacement strategy, REJ can be used (e.g., using SEQ ID NOS: 230 and 231, which can further include one or more of the binding domains SEQ ID NOS: 232-241).SLR 7158-102574-27 05 / 08 / 25 S2024-016A
[0357] In some embodiments, the compositions and methods described herein are useful for treating a subject having a disease or disorder by introducing a nucleic acid mutation to disrupt an undesirable genomic sequence and / or downregulate the production of an undesirable gene product in a cell of the subject. In some embodiments, the production of the undesirable genomic sequence and / or undesirable gene product is downregulated by altering a coding region or a noncoding region. In some embodiments, one or more regulatory sequence (e.g., a promoter or enhancer) is altered. In some embodiments, sequences that modulate splicing or other aspects of RNA processing (e.g., poly- adenylation, subcellular localization, or half-life) are altered to disrupt the undesirable genomic sequence and / or downregulate the level of an undesirable gene product. In some embodiments the coding region of the undesirable genomic sequence is altered to introduce a mutation, e.g., a deletion, missense mutation, frameshift mutation, or stop codon, resulting in downregulation of an activity of the undesirable genomic sequence and / or downregulation of production of an undesirable gene product. In some embodiments, the level of the activity and / or gene product is downregulated in comparison to a control, e.g., an untreated control. In some embodiments, an undesirable genomic sequence for targeting using the nucleic acid editing compositions, systems, and methods described herein can include any described in the literature, e.g., an oncogene or a disease gene, having a gain-of function mutation. In some embodiments, the oncogene is any listed in Table 2. In some embodiments, the nucleic acid editing compositions, systems, and methods described herein are used to treat a cancer listed in Table 2.
[0358] Table 2. Oncogenes
[0359] Oncogene Function / Activation Cancer*
[0360] Promotes cell growth through tyrosine
[0361] ABL1 Chronic myelogenous leukemia kinase activity
[0362] Fusion affects the MLLT11 transcription
[0363] AFF4 / MLLT11 factor / methyltransferase. MLLT11 is Acute leukemias
[0364] also called HRX, ALLI and HTRX1
[0365] Encodes a protein-serine / threonine
[0366] AKT2 Ovarian cancer
[0367] kinase
[0368] ALK Encodes a receptor tyrosine kinase Lymphomas
[0369] Translocation creates fusion protein with
[0370] ALK / NPM Large cell lymphomas nucleophosmin(npm)
[0371] RUNX1 (AML1) Encodes a transcription factor Acute myeloid leukemia
[0372] New fusion protein created by
[0373] RUNX1 / MTG8(ETO) Acute leukemias
[0374] translocation
[0375] AXL Encodes a receptor tyrosine kinase Hematopoietic cancers
[0376] BCL-2, 3, 6 Block apoptosis (programmed cell death) B-cell lymphomas and leukemias New protein created by fusion of bcr and Chronic myelogenous and acute BCR / ABL
[0377] abl triggers unregulated cell growth lymphocytic leukemia
[0378]
[0379] MYC (c-MYC) Transcription factor that promotes cell Leukemia; breast, stomach, lung,SLR 7158-102574-27 05 / 08 / 25 S2024-016A
[0380] proliferation and DNA synthesis cervical, and colon carcinomas;
[0381] neuroblastomas and glioblastomas MCF2 (DBL) Guanine nucleotide exchange factor Diffuse B-cell lymphoma DEK / NUP214 New protein created by fusion Acute myeloid leukemia
[0382] Acute pre B-ccll leukemia; TCF3 also TCF3 / PBX1 New protein created by fusion
[0383] called E2A
[0384] Cell surface receptor that triggers cell Squamous cell carcinoma, glioblastomas, EGFR
[0385] growth through tyrosine kinase activity lung cancer
[0386] Fusion protein created by a translocation
[0387] MLLT11 Acute leukemias
[0388] t(ll;19).
[0389] Fusion protein created by t( 16:21 )
[0390] ERG / FUS translocation. The erg protein is a Myeloid leukemia
[0391] transcription factor.
[0392] Cell surface receptor that triggers cell
[0393] Breast, salivary gland, and ovarian ERBB2 growth through tyrosine kinase activity;
[0394] carcinomas
[0395] also known as HER2 or neu
[0396] ETS1 Transcription factor Lymphoma
[0397] Fusion protein created by t( 11: 22)
[0398] EWSR1 / FLI1 Ewing Sarcoma
[0399] translocation.
[0400] CSF1R Tyrosine kinase Sarcoma
[0401] FOS Transcription factor for API Osteosarcoma
[0402] FES Tyrosine kinase Sarcoma
[0403] GLI1 Transcription factor Glioblastoma
[0404] GNAS (GSP) Membrane associated G protein Thyroid carcinoma
[0405] Overexpression of signaling kinase due
[0406] HER2 / neu Breast and cervical carcinomas
[0407] to gene amplification
[0408] TLX1 Transcription factor; aka HOX11 Acute T-cell leukemia
[0409] Encodes fibroblast growth factor; aka
[0410] FGF4 Breast and squamous cell carcinomas
[0411] HST1
[0412]
[0413] SLR 7158-102574-27 05 / 08 / 25 S2024-016A
[0414] IL3 Cell signaling molecule Acute pre B-cell leukemia
[0415] FGF3 (INT-2) Encodes a fibroblast growth factor Breast and squamous cell carcinomas
[0416] JUN Transcription factor for API Sarcoma
[0417] KIT Tyrosine kinase Sarcoma
[0418] FGF4 (KS3) Herpes virus encoded growth factor Kaposi's sarcoma
[0419] K-SAM Fibroblast growth factor receptor Stomach carcinomas
[0420] Guanine nucleotide exchange factor; aka
[0421] AKAP13 Myeloid leukemias
[0422] LBC
[0423] LCK Tyrosine kinase T-cell lymphoma
[0424] LMO1, LMO2 Transcription factors T-cell lymphoma
[0425] MYCL Transcription factor Lung carcinomas
[0426] LYL1 Transcription factor Acute T-cell leukemia NFKB2 Transcription factor. Also called LYT-10 B-cell lymphoma
[0427] Fusion protein formed by
[0428] the (10;14)(q24;q32) translocation of
[0429] NFKB2 / Cα1
[0430] NFKB2 next to the C alpha 1
[0431] immunoglobulin locus.
[0432] MASI Angiotensin receptor Mammary carcinoma Encodes a protein that inhibits and leads
[0433] MDM2 Sarcomas
[0434] to the degradation of p53
[0435] Transcription factor / methyltransferase
[0436] MLLT11 Acute myeloid leukemia
[0437] (also called HRX and ALE1)
[0438] MOS Serine / threonine kinase Lung cancer
[0439] Fusion of transcription repressor to
[0440] RUNX1T1 factor to a transcription factor. Also Acute leukemias
[0441] known as MTG8 and AML1-MTG8
[0442] MYB Transcription factor Colon carcinoma and leukemias
[0443]
[0444] MYH11 / CBFB New protein created by fusion of Acute myeloid leukemiaSLR 7158-102574-27 05 / 08 / 25 S2024-016A
[0445] transcription factors via an inversion in
[0446] chromosome 16.
[0447] Tyrosine kinase. Also called ERBB2 or Glioblastomas, and squamous cell NEU HER2 carcinomas
[0448] Neuroblastomas, retinoblastomas, and MYCN Cell proliferation and DNA synthesis
[0449] lung carcinomas
[0450] MCF2L (OST) Guanine nucleotide exchange factor Osteosarcomas
[0451] PAX-5 Transcription factor Lympho-plasmacytoid B-cell lymphoma Fusion protein formed via t( 1: 19)
[0452] PBX1 / E2A Acute pre B-cell leukemia translocation. Transcription factor
[0453] PIMl Serine / threonine kinase T-cell lymphoma
[0454] Encodes cyclin DI. Involved in cell
[0455] CCND1 Breast and squamous cell carcinomas cycle regulation. Also called PRAD1
[0456] RAFI Serine / threonine kinase Many cancer types
[0457] Fusion protein caused by t(15: 17)
[0458] RARA / PML Acute premyelocytic leukemia translocation. Retinoic acid receptor.
[0459] HRAS G-protein. Signal transduction. Bladder carcinoma
[0460] KRAS G-protein. Signal transduction Lung, ovarian, and bladder carcinoma NRAS G-protein. Signal transduction Breast carcinoma
[0461] Fusion protein formed by deletion in
[0462] REL / NRG B-cell lymphoma
[0463] chromosome 2. Transcription factor.
[0464] Thyroid carcinomas, multiple endocrine RET Cell surface receptor. Tyrosine kinase
[0465] neoplasia type 2
[0466] Transcription factors aka LMO1 and
[0467] RHOM1, RHOM2 Acute T-cell leukemia
[0468] LMO2
[0469] ROS1 Tyrosine kinase Sarcoma
[0470] SKI Transcription factor Carcinomas
[0471] SIS (aka PDGFB) Growth factor Glioma, fibrosarcoma
[0472] Fusion protein formed by rearrangement
[0473] SET / CAN Acute myeloid leukemia
[0474] of chromosome 9. Protein localization
[0475] SRC Tyrosine kinase Sarcomas
[0476] Transcription factor. TAL1 is also called
[0477] TALI, TAL2 Acute T-cell leukemia
[0478] SCL
[0479] Altered form of Notch (a cellular
[0480] NOTCH1 (TAN1) Acute T-cell leukemia
[0481] receptor) formed by t(7:9) translocation
[0482] TIAM1 Guanine nucleotide exchange factor T-lymphoma
[0483] TSC2 GTPase activator Renal and brain tumors
[0484]
[0485] NTRK1 Receptor tyrosine kinase Colon and thyroid carcinomas
[0486] In some aspects, a disease gene having a gain-of-function mutation is a neurodegenerative disease gene. In some embodiments, the neurodegenerative disease gene is any listed in Table 3, reproduced from Table 1 of Chen and Altman, 2017. In some embodiments, the nucleic acid editing compositions, systems, and methods described herein are used to treat a neurodegenerative disease listed in Table 3.SLR 7158-102574-27 05 / 08 / 25 S2024-016A
[0487] Table 3. Genetic Diseases Resulting from Gain-of-Function Mutations
[0488]
[0489] To address some of the limitations with existing strategies for reconstitution of fragmented genes from multiple AAVs, provided herein is a system that serially aligns and recombines two or more individual synthetic RNA molecules in the target cell. Each individual synthetic RNA molecule includes a synthetic intron sequence, containing a dimerization domain and elements needed for RNA splicing, which upon binding of dimerization domains to one another in the correct order, mediates efficient RNA recombination of individual fragments. In one example, reconstitution of a coding sequence from two fragments is achieved by appending a first synthetic intron (A) to the 3’ end of the N-terminal coding fragment and a complimentary second synthetic domain (A’) to the 5’ end of the C-terminal coding fragment. The two RNAs are recombined by a cell’s intrinsic RNA splicing machinery (i.e., the spliceosome machinery). The synthetic intron domains contain two functional elements: (1) a dimerization domain to mediate base pairing between the two halves that are to be recombined and (2) a domain optimized to efficiently recruit the splicing machinery to mediate efficient reconstitution of the two RNA molecules. The synthetic intron domain can include elements to prevent unspliced RNA from encoding protein. In some examples, a synthetic intron includes a sequence having at least 50% at least 60%, at least 70%, at least 75%, 80%, at least 85%, at least 90%, at least 95%, at least 98%, at least 99%, or 100% sequence identity to any synthetic intron provided in SEQ ID NOS: 159, 160, 161, 162, 163, 164, 165, 166 e.g., see FIGS. 10A-10Z), 242, 243, 244, and 245. In some examples, a synthetic intron is an RNA molecule encoded by a sequence having at least 50% at least 60%, at least 70%, at least 75%, 80%, at least 85%, at least 90%, at least 95%, at least 98%, at least 99%, or 100% sequence identity to any synthetic intron provided in SEQ ID NOS: 159, 160, 161, 162, 163, 164, 165, 166, 242, 243, 244, and 245, but without the provided promoter sequence). One skilled in the art will appreciate that any of the molecules provided in SEQ ID NOS: 159, 160, 161, 162, 163, 164, 165, 166, 242, 243, 244, and 245 can be modified to replace the protein coding portions (e.g., 114SLR 7158-102574-27 05 / 08 / 25 S2024-016A
[0490] and 164 of FIG. 6A) with another nucleic acid editing protein coding sequence of interest (e.g., YFP coding sequence of SEQ ID NO: 1, 2, 22 or 23 can be replaced with a Cas9 or Casl3d protein coding sequence). Thus, also provided herein are synthetic intron molecules having at least 80%, at least 85%, at least 90%, at least 95%, at least 98%, at least 99%, or 100% sequence identity to any synthetic intron portion provided in SEQ ID NO: 159, 160, 161, 162, 163, 164, 165, 166, 242, 243, 244, and 245. Also provided are synthetic intron RNA molecules encoded by a sequence having at least 50% at least 60%, at least 70%, at least 75%, 80%, at least 85%, at least 90%, at least 95%, at least 98%, at least 99%, or 100% sequence identity to any synthetic intron provided in SEQ ID NOS: 159, 160, 161, 162, 163, 164, 165, 166, 242, 243, 244, and 245, but without the provided promoter sequence).
[0491] Exemplary dimerization domains were bioinformatically selected to minimize / optimize their internal secondary / tertiary structure. The dimerization domains tested contained long stretches of low diversity nucleotide sequences to avoid intramolecular annealing. By avoiding intramolecular annealing, these dimerization domains are present in an open configuration and therefore are available for pairing with the corresponding complementary dimerization domain sequence. The synthetic intron domains contain intronic splice enhancing elements which lead to efficient recruitment of the splicing machinery.
[0492] The disclosed RNA molecules are designed to have at least an open and available single-stranded region that is available to bind to the complementary dimerization domain to allow efficient splicing and recombination of the RNAs. In some examples, this is achieved by utilizing only purines or only pyrimidines for the binding domains. Due to the inability of purines to pair with themselves (and pyrimidines likewise) these stretches of RNA have an open predicted structure.
[0493] RNA molecules are present as a single strand in the cells. Being single stranded they are inherently prone to hybridize to themselves and thereby form strong secondary and tertiary structures. The most stable base pairs will be G with C, A with U, and the G with U wobble pair. Thermodynamically, the pairing of two bases is favored over an open configuration. To design efficient synthetic nucleic acid molecules, two dimerization domains having complementarity to one another are present in an open configuration such that the dimerization domains are available for inter-molecular base pairing. To avoid intra-molecular base pairing in between other parts of the synthetic nucleic acid molecules, a long stretch of non-diverse sequences containing incompatible bases can be included. For example, a long stretch of pyrimidines (i.e., C and T) or purines (i.e., A and G) can be present in the synthetic nucleic acid molecules. Pyrimidines cannot form canonical base pairs with other pyrimidines, purines cannot form canonical base pairs with other purines. Such a stretch of purines or pyrimidines can range from a couple bases to a couple hundreds of bases. Since these stretches cannot intra-molecularly bind, they are available for inter-molecular base pairing with a complementary fragment. For example, the synthetic nucleic acid molecules A and A’ may be configured with A containing a pyrimidine stretch (e.g., 5’-CCUU(..,)CCUU-3’) and A’ containing the complementary purine sequence (e.g., 5’-AAGG(...)AAGG-3’).SLR 7158-102574-27 05 / 08 / 25 S2024-016A
[0494] The disclosed synthetic nucleic acid molecules (e.g., RNA or DNA encoding the RNA) are designed to minimize any off-target binding to incorrect sites in the genome. Off target binding can be reduced by altering the sequence of the nucleic acid molecule.
[0495] The same design principle, that is the use of hypodiverse stretches of RNA bases to achieve open synthetic nucleic acid configurations, can be extended to using stretches of single bases e.g. using a series of Gs that would base pair with a series of Cs and a series of As that would base pair with a series of Us, in the dimerization domains.
[0496] To increase recombination of two or more synthetic nucleic acid molecules, the following methods can be used. RNA splicing depends on the recruitment of spliceosome components to the 5’ end of the intron (the splice donor site) and the 3’ end of the intron (the splice acceptor site, with its associated branch point sequence and the polypyrimidine tract). Different ribonucleoproteins are recruited to the intron through base pairing of protein associated small nuclear RNA (snRNA) with intronic sequences. By placing perfect match consensus sequences into the RNA dimerization and recombination domains, the recruitment of spliceosome components can be facilitated which in turn enhances the efficiency of spliceosome mediated recombination. Previously characterized intronic splice enhancer sequences can recruit additional splicing promoting factors that are referred to as intronic splice enhancers.
[0497] In some examples, instead of using naturally occurring RNA sequences for the RNA splicing sequences, consensus sequences are used. For example, consensus sequences can be used for any of the sequences that are involved in splicing, including splice donor, splice acceptor, splice enhancer and splice branch point sequences. With these synthetic nucleic acid molecules, two (or more) RNA molecules can be serially joined together in a cell ex vivo, in vitro, or in vivo. Outside of the encoded synthetic intronic domains, synthetic nucleic acid molecules can include any promoter and coding sequence. For example, two synthetic nucleic acid molecules could carry two halves of a single nucleic acid editing gene. This was tested in vitro and in vivo by reconstituting two halves of a yellow fluorescent protein (YFP), and was shown to be efficient (see FIGS. 3A-3D).
[0498] The modular nature of the synthetic nucleic acid molecules allowed for testing the efficiency of achieving serial recombination (i.e., >2) of multiple RNA fragments using a combinatorial set of optimized complimentary dimerization domains (FIGS. 4A-4B). A three-way split yellow fluorescent protein was efficiently reconstituted and expressed at high levels in >80% of transfected cells.
[0499] These results demonstrate that a single RNA molecule can be reconstituted from at least three different synthetic nucleic acid molecules, such as when expression of a nucleic acid editing protein that has a promoter and / or a coding sequence that is too long to fit into a single gene therapy vector such as AAV.
[0500] In some examples, the synthetic nucleic acid molecules, e.g., synthetic DNA molecules, of the inventive compositions, systems, kits, and methods, are produced by transcription of an RNA virus genome by reverse transcriptase.
[0501] The disclosed system allows for the efficient RNA recombination between individual fragments. In some examples, reconstitution (i.e., splicing or recombination) efficiency achieved using the compositions,SLR 7158-102574-27 05 / 08 / 25 S2024-016A
[0502] systems or methods of the disclosure is determined using any suitable method known to one of skill in the art. In some examples, reconstitution efficiency is represented by a measure of correctly joined RNA relative to a control RNA, or a measure of full-length nucleic acid editing protein or protein activity relative to that of a control protein. In some examples the control RNA is the unjoined RNA, wherein reconstitution efficiency is represented by a measure of joined RNA relative to unjoined RNA. This measurement can be made by detecting and comparing junction RNA and the unjoined 3’ RNA species 3’ (e.g., junction RNA: 3’ RNA). In some examples wherein more than two RNAs are joined, joining at either or all junctions are evaluated. In some examples, reconstitution efficiency is represented by a measure of full-length or active nucleic acid editing protein relative to a protein fragment or inactive protein.
[0503] In some examples, the reconstitution, recombination or splicing efficiency (a measure of the correct joining of the two or more different coding sequences present on different RNA molecules, and / or the production of the desired full-length protein) is about 10% to about 100%. The synthetic intron domain can inlcude elements to prevent unspliced RNA from encoding protein. In some examples, the reconstitution efficiency is about 10% to about 15%, about 10% to about 20%, about 10% to about 25%, about 10% to about 30%, about 10% to about 40%, about 10% to about 50%, about 10% to about 60%, about 10% to about 70%, about 10% to about 80%, about 10% to about 90%, about 10% to about 100%, about 15% to about 20%, about 15% to about 25%, about 15% to about 30%, about 15% to about 40%, about 15% to about 50%, about 15% to about 60%, about 15% to about 70%, about 15% to about 80%, about 15% to about 90%, about 15% to about 100%, about 20% to about 25%, about 20% to about 30%, about 20% to about 40%, about 20% to about 50%, about 20% to about 60%, about 20% to about 70%, about 20% to about 80%, about 20% to about 90%, about 20% to about 100%, about 25% to about 30%, about 25% to about 40%, about 25% to about 50%, about 25% to about 60%, about 25% to about 70%, about 25% to about 80%, about 25% to about 90%, about 25% to about 100%, about 30% to about 40%, about 30% to about 50%, about 30% to about 60%, about 30% to about 70%, about 30% to about 80%, about 30% to about 90%, about 30% to about 100%, about 40% to about 50%, about 40% to about 60%, about 40% to about 70%, about 40% to about 80%, about 40% to about 90%, about 40% to about 100%, about 50% to about 60%, about 50% to about 70%, about 50% to about 80%, about 50% to about 90%, about 50% to about 100%, about 60% to about 70%, about 60% to about 80%, about 60% to about 90%, about 60% to about 100%, about 70% to about 80%, about 70% to about 90%, about 70% to about 100%, about 80% to about 90%, about 80% to about 100%, or about 90% to about 100%. In some examples, the reconstitution efficiency is about 10%, about 15%, about 20%, about 25%, about 30%, about 40%, about 50%, about 60%, about 70%, about 80%, about 90%, or about 100%. In some examples, the reconstitution efficiency is at least about 10%, about 15%, about 20%, about 25%, about 30%, about 40%, about 50%, about 60%, about 70%, about 80%, or about 90%. In some examples, the reconstitution efficiency is at most about 15%, about 20%, about 25%, about 30%, about 40%, about 50%, about 60%, about 70%, about 80%, about 90%, or about 100%.SLR 7158-102574-27 05 / 08 / 25 S2024-016A
[0504] In some examples, the reconstitution, recombination or splicing efficiency (in this example a measure of the correct joining of two different nucleic acid editing protein coding sequences present on different RNA molecules, and / or the production of the desired nucleic acid editing full-length protein, wherein the two different nucleic acid editing protein coding sequences encode a transcript of about 3200 nt to 9000 nt, such as about 4000 to 9000 nt, about 4400 to 9000 nt, about 3200 to 4000 nt, about 3200 to 3600 nt, for example about 4500 nt, about 4000 nt, about 3800 nt, about 3600 nt, or about 3200 nt), is about 10% to about 100%. In some examples, the reconstitution efficiency using a two-part system is about 10% to about 15%, about 10% to about 20%, about 10% to about 25%, about 10% to about 30%, about 10% to about 40%, about 10% to about 50%, about 10% to about 60%, about 10% to about 70%, about 10% to about 80%, about 10% to about 90%, about 10% to about 100%, about 15% to about 20%, about 15% to about 25%, about 15% to about 30%, about 15% to about 40%, about 15% to about 50%, about 15% to about 60%, about 15% to about 70%, about 15% to about 80%, about 15% to about 90%, about 15% to about 100%, about 20% to about 25%, about 20% to about 30%, about 20% to about 40%, about 20% to about 50%, about 20% to about 60%, about 20% to about 70%, about 20% to about 80%, about 20% to about 90%, about 20% to about 100%, about 25% to about 30%, about 25% to about 40%, about 25% to about 50%, about 25% to about 60%, about 25% to about 70%, about 25% to about 80%, about 25% to about 90%, about 25% to about 100%, about 30% to about 40%, about 30% to about 50%, about 30% to about 60%, about 30% to about 70%, about 30% to about 80%, about 30% to about 90%, about 30% to about 100%, about 40% to about 50%, about 40% to about 60%, about 40% to about 70%, about 40% to about 80%, about 40% to about 90%, about 40% to about 100%, about 50% to about 60%, about 50% to about 70%, about 50% to about 80%, about 50% to about 90%, about 50% to about 100%, about 60% to about 70%, about 60% to about 80%, about 60% to about 90%, about 60% to about 100%, about 70% to about 80%, about 70% to about 90%, about 70% to about 100%, about 80% to about 90%, about 80% to about 100%, or about 90% to about 100%. In some examples, the reconstitution efficiency is about 10%, about 15%, about 20%, about 25%, about 30%, about 40%, about 50%, about 60%, about 70%, about 80%, about 90%, or about 100%. In some examples, the reconstitution efficiency is at least about 10%, about 15%, about 20%, about 25%, about 30%, about 40%, about 50%, about 60%, about 70%, about 80%, or about 90%. In some examples, the reconstitution efficiency is at most about 15%, about 20%, about 25%, about 30%, about 40%, about 50%, about 60%, about 70%, about 80%, about 90%, or about 100%.
[0505] In some examples, the reconstitution, recombination or splicing efficiency (in this example a measure of the correct joining of the two different nucleic acid editing protein coding sequences present on different RNA molecules, and / or the production of the desired full-length nucleic acid editing protein, wherein the two different nucleic acid editing protein coding sequences encode a transcript of about 4000 nt), is about 40% to about 60%, such as about 40% to about 50%, about 42% to about 47%, for example about 45%.
[0506] In some examples, the reconstitution, recombination or splicing efficiency (in this example a measure of the correct joining of the two different nucleic acid editing protein coding sequences present onSLR 7158-102574-27 05 / 08 / 25 S2024-016A
[0507] different RNA molecules, and / or the production of the desired nucleic acid editing full-length protein, wherein the two different nucleic acid editing protein coding sequences encode a transcript of about 3800 nt), is about 40% to about 60%, such as about 40% to about 50%, about 42% to about 47%, for example about 45%.
[0508] In some examples, the reconstitution, recombination or splicing efficiency (in this example a measure of the correct joining of the two different nucleic acid editing protein coding sequences present on different RNA molecules, and / or the production of the desired nucleic acid editing full-length protein, wherein the two different nucleic acid editing protein coding sequences encode a transcript of about 3600 nt), is about 25% to about 50%, such as about 30% to about 40%, for example about 35%.
[0509] In some examples, the reconstitution, recombination or splicing efficiency (in this example a measure of the correct joining of the two different nucleic acid editing protein coding sequences present on different RNA molecules, and / or the production of the desired nucleic acid editing full-length protein, wherein the two different nucleic acid editing protein coding sequences encode a transcript of about 3200 nt), is about 25% to about 50%, such as about 30% to about 40%, for example about 35%.
[0510] In some examples, the reconstitution, recombination or splicing efficiency (in this example a measure of the correct joining of three different nucleic acid editing protein coding sequences present on different RNA molecules, and / or the production of the desired full-length nucleic acid editing protein, wherein the three different nucleic acid editing protein coding sequences encode a transcript of about 3200 nt to about 13,500 nt, such as about 4000 nt to about 5,000 nt, about 4000 nt to about 13,500 nt, about 6000 nt to about 12,000 nt, about 6000 nt to about 10,000 nt, or about 8000 nt to about 12,000 nt, for example up to about 13,500 nt), is about 10% to about 100%. In some examples, the reconstitution efficiency using a three-part system is about 10% to about 15%, about 10% to about 20%, about 10% to about 25%, about 10% to about 30%, about 10% to about 40%, about 10% to about 50%, about 10% to about 60%, about 10% to about 70%, about 10% to about 80%, about 10% to about 90%, about 10% to about 100%, about 15% to about 20%, about 15% to about 25%, about 15% to about 30%, about 15% to about 40%, about 15% to about 50%, about 15% to about 60%, about 15% to about 70%, about 15% to about 80%, about 15% to about 90%, about 15% to about 100%, about 20% to about 25%, about 20% to about 30%, about 20% to about 40%, about 20% to about 50%, about 20% to about 60%, about 20% to about 70%, about 20% to about 80%, about 20% to about 90%, about 20% to about 100%, about 25% to about 30%, about 25% to about 40%, about 25% to about 50%, about 25% to about 60%, about 25% to about 70%, about 25% to about 80%, about 25% to about 90%, about 25% to about 100%, about 30% to about 40%, about 30% to about 50%, about 30% to about 60%, about 30% to about 70%, about 30% to about 80%, about 30% to about 90%, about 30% to about 100%, about 40% to about 50%, about 40% to about 60%, about 40% to about 70%, about 40% to about 80%, about 40% to about 90%, about 40% to about 100%, about 50% to about 60%, about 50% to about 70%, about 50% to about 80%, about 50% to about 90%, about 50% to about 100%, about 60% to about 70%, about 60% to about 80%, about 60% to about 90%, about 60% to about 100%, about 70% to about 80%, about 70% to about 90%, about 70% to about 100%, about 80% toSLR 7158-102574-27 05 / 08 / 25 S2024-016A
[0511] about 90%, about 80% to about 100%, or about 90% to about 100%. In some examples, the reconstitution efficiency is about 10%, about 15%, about 20%, about 25%, about 30%, about 40%, about 50%, about 60%, about 70%, about 80%, about 90%, or about 100%. In some examples, the reconstitution efficiency is at least about 10%, about 15%, about 20%, about 25%, about 30%, about 40%, about 50%, about 60%, about 70%, about 80%, or about 90%. In some examples, the reconstitution efficiency is at most about 15%, about 20%, about 25%, about 30%, about 40%, about 50%, about 60%, about 70%, about 80%, about 90%, or about 100%.
[0512] In some examples, the reconstitution, recombination or splicing efficiency (in this example a measure of the correct joining of four different nucleic acid editing protein coding sequences present on different RNA molecules, and / or the production of the desired full-length nucleic acid editing protein, wherein the four different nucleic acid editing protein coding sequences encode a transcript of about 3200 nt to about 18,000 nt, such as about 4000 nt to about 18,000 nt, about 4000 nt to about 5,000 nt, about 10,000 nt to about 18,000 nt, about 15,000 nt to about 18,000nt, or about 12,000 nt to about 15,000 nt, for example up to about 18,000 nt), is about 10% to about 100%. In some examples, the reconstitution efficiency using a four-part system is about 10% to about 15%, about 10% to about 20%, about 10% to about 25%, about 10% to about 30%, about 10% to about 40%, about 10% to about 50%, about 10% to about 60%, about 10% to about 70%, about 10% to about 80%, about 10% to about 90%, about 10% to about 100%, about 15% to about 20%, about 15% to about 25%, about 15% to about 30%, about 15% to about 40%, about 15% to about 50%, about 15% to about 60%, about 15% to about 70%, about 15% to about 80%, about 15% to about 90%, about 15% to about 100%, about 20% to about 25%, about 20% to about 30%, about 20% to about 40%, about 20% to about 50%, about 20% to about 60%, about 20% to about 70%, about 20% to about 80%, about 20% to about 90%, about 20% to about 100%, about 25% to about 30%, about 25% to about 40%, about 25% to about 50%, about 25% to about 60%, about 25% to about 70%, about 25% to about 80%, about 25% to about 90%, about 25% to about 100%, about 30% to about 40%, about 30% to about 50%, about 30% to about 60%, about 30% to about 70%, about 30% to about 80%, about 30% to about 90%, about 30% to about 100%, about 40% to about 50%, about 40% to about 60%, about 40% to about 70%, about 40% to about 80%, about 40% to about 90%, about 40% to about 100%, about 50% to about 60%, about 50% to about 70%, about 50% to about 80%, about 50% to about 90%, about 50% to about 100%, about 60% to about 70%, about 60% to about 80%, about 60% to about 90%, about 60% to about 100%, about 70% to about 80%, about 70% to about 90%, about 70% to about 100%, about 80% to about 90%, about 80% to about 100%, or about 90% to about 100%. In some examples, the reconstitution efficiency is about 10%, about 15%, about 20%, about 25%, about 30%, about 40%, about 50%, about 60%, about 70%, about 80%, about 90%, or about 100%. In some examples, the reconstitution efficiency is at least about 10%, about 15%, about 20%, about 25%, about 30%, about 40%, about 50%, about 60%, about 70%, about 80%, or about 90%. In some examples, the reconstitution efficiency is at most about 15%, about 20%, about 25%, about 30%, about 40%, about 50%, about 60%, about 70%, about 80%, about 90%, or about 100%. In some examples, the compositions, systems or methods of the disclosure are evaluated bySLR 7158-102574-27 05 / 08 / 25 S2024-016A
[0513] determining an RNA or protein production level using any suitable method known to one of skill in the art. In some examples, the RNA production level is represented by a measure of correctly joined RNA relative to a control RNA, or a measure of full-length protein relative to a control. In some examples the control RNA is a corresponding mutant RNA or an endogenous RNA. For example, the ratio of the amount of joined RNA to the amount of mutant or endogenous RNA produced in the transfected cell is compared with same ratio in nontransfected cells, to determine the production level of the correctly joined RNA. In some examples, the ratio of the amount of the correctly joined RNA, full-length protein, or the protein activity, to the amount of the control RNA, or the amount or activity of the control protein, are compared.
[0514] In some examples, the RNA production level achieved is 5% to 100%. In some examples, the RNA production level achieved is about 5% to about 100%. In some examples, the RNA production level achieved is about 5% to about 10%, about 5% to about 20%, about 5% to about 25%, about 5% to about 30%, about 5% to about 40%, about 5% to about 50%, about 5% to about 60%, about 5% to about 70%, about 5% to about 80%, about 5% to about 90%, about 5% to about 100%, about 10% to about 20%, about 10% to about 25%, about 10% to about 30%, about 10% to about 40%, about 10% to about 50%, about 10% to about 60%, about 10% to about 70%, about 10% to about 80%, about 10% to about 90%, about 10% to about 100%, about 20% to about 25%, about 20% to about 30%, about 20% to about 40%, about 20% to about 50%, about 20% to about 60%, about 20% to about 70%, about 20% to about 80%, about 20% to about 90%, about 20% to about 100%, about 25% to about 30%, about 25% to about 40%, about 25% to about 50%, about 25% to about 60%, about 25% to about 70%, about 25% to about 80%, about 25% to about 90%, about 25% to about 100%, about 30% to about 40%, about 30% to about 50%, about 30% to about 60%, about 30% to about 70%, about 30% to about 80%, about 30% to about 90%, about 30% to about 100%, about 40% to about 50%, about 40% to about 60%, about 40% to about 70%, about 40% to about 80%, about 40% to about 90%, about 40% to about 100%, about 50% to about 60%, about 50% to about 70%, about 50% to about 80%, about 50% to about 90%, about 50% to about 100%, about 60% to about 70%, about 60% to about 80%, about 60% to about 90%, about 60% to about 100%, about 70% to about 80%, about 70% to about 90%, about 70% to about 100%, about 80% to about 90%, about 80% to about 100%, or about 90% to about 100%. In some examples, the RNA production level achieved is about 5%, about 10%, about 20%, about 25%, about 30%, about 40%, about 50%, about 60%, about 70%, about 80%, about 90%, or about 100%. In some examples, the RNA production level achieved is at least about 5%, about 10%, about 20%, about 25%, about 30%, about 40%, about 50%, about 60%, about 70%, about 80%, or about 90%. In some examples, the RNA production level achieved is at most about 10%, about 20%, about 25%, about 30%, about 40%, about 50%, about 60%, about 70%, about 80%, about 90%, or about 100%.
[0515] In some examples, the protein production level is represented by a measure of the amount of full-length nucleic acid editing protein or nucleic acid editing protein activity relative to that of a control protein. In some examples the control protein is a corresponding mutant protein or an endogenous protein. For example, the ratio of the amount of full-length nucleic acid editing protein or protein activity to the amountSLR 7158-102574-27 05 / 08 / 25 S2024-016A
[0516] of mutant or endogenous protein produced in the transfected cell is compared with same ratio in nontransfected cells. In some examples, the control protein is the full-length nucleic acid editing protein produced in, e.g., a cell that is engineered to express a control full-length protein (wherein the cell is not transfected with the inventive constructs) or a non-transfected cell from a normal subject that expresses a control full-length protein, and the protein production level is determined by measuring the amount or activity of the nucleic acid editing protein in the transfected cell and comparing it to that of the control protein. In some examples, the control protein is a mutant form of the protein, produced in a cell that is transfected or nontransfected with the construct, and the amount of full-length protein or protein activity is compared with that of the control protein to determine the protein production level. In some examples, the amount of full-length protein or protein activity is compared with that of an endogenous, or housekeeping, protein to determine the protein production level.
[0517] In some examples, the protein production level achieved is about 1% to about 100%. In some examples, the protein production level achieved is about 10% to about 100%. In some examples, the protein production level achieved is about 10% to about 20%, about 10% to about 30%, about 10% to about 40%, about 10% to about 50%, about 10% to about 60%, about 10% to about 70%, about 10% to about 75%, about 10% to about 80%, about 10% to about 85%, about 10% to about 90%, about 10% to about 100%, about 20% to about 30%, about 20% to about 40%, about 20% to about 50%, about 20% to about 60%, about 20% to about 70%, about 20% to about 75%, about 20% to about 80%, about 20% to about 85%, about 20% to about 90%, about 20% to about 100%, about 30% to about 40%, about 30% to about 50%, about 30% to about 60%, about 30% to about 70%, about 30% to about 75%, about 30% to about 80%, about 30% to about 85%, about 30% to about 90%, about 30% to about 100%, about 40% to about 50%, about 40% to about 60%, about 40% to about 70%, about 40% to about 75%, about 40% to about 80%, about 40% to about 85%, about 40% to about 90%, about 40% to about 100%, about 50% to about 60%, about 50% to about 70%, about 50% to about 75%, about 50% to about 80%, about 50% to about 85%, about 50% to about 90%, about 50% to about 100%, about 60% to about 70%, about 60% to about 75%, about 60% to about 80%, about 60% to about 85%, about 60% to about 90%, about 60% to about 100%, about 70% to about 75%, about 70% to about 80%, about 70% to about 85%, about 70% to about 90%, about 70% to about 100%, about 75% to about 80%, about 75% to about 85%, about 75% to about 90%, about 75% to about 100%, about 80% to about 85%, about 80% to about 90%, about 80% to about 100%, about 85% to about 90%, about 85% to about 100%, or about 90% to about 100%. In some examples, the protein production level achieved is about 10%, about 20%, about 30%, about 40%, about 50%, about 60%, about 70%, about 75%, about 80%, about 85%, about 90%, or about 100%. In some examples, the protein production level achieved is at least about 10%, about 20%, about 30%, about 40%, about 50%, about 60%, about 70%, about 75%, about 80%, about 85%, or about 90%. In some examples, the protein production level achieved is at most about 20%, about 30%, about 40%, about 50%, about 60%, about 70%, about 75%, about 80%, about 85%, about 90%, or about 100%.SLR 7158-102574-27 05 / 08 / 25 S2024-016A
[0518] In some examples, the protein activity level achieved is about 50% to about 100%. In some examples, the protein activity level achieved is about 50% to about 100%. In some examples, the protein activity level achieved is about 50% to about 55%, about 50% to about 60%, about 50% to about 65%, about 50% to about 70%, about 50% to about 75%, about 50% to about 80%, about 50% to about 85%, about 50% to about 90%, about 50% to about 95%, about 50% to about 100%, about 55% to about 60%, about 55% to about 65%, about 55% to about 70%, about 55% to about 75%, about 55% to about 80%, about 55% to about 85%, about 55% to about 90%, about 55% to about 95%, about 55% to about 100%, about 60% to about 65%, about 60% to about 70%, about 60% to about 75%, about 60% to about 80%, about 60% to about 85%, about 60% to about 90%, about 60% to about 95%, about 60% to about 100%, about 65% to about 70%, about 65% to about 75%, about 65% to about 80%, about 65% to about 85%, about 65% to about 90%, about 65% to about 95%, about 65% to about 100%, about 70% to about 75%, about 70% to about 80%, about 70% to about 85%, about 70% to about 90%, about 70% to about 95%, about 70% to about 100%, about 75% to about 80%, about 75% to about 85%, about 75% to about 90%, about 75% to about 95%, about 75% to about 100%, about 80% to about 85%, about 80% to about 90%, about 80% to about 95%, about 80% to about 100%, about 85% to about 90%, about 85% to about 95%, about 85% to about 100%, about 90% to about 95%, about 90% to about 100%, or about 95% to about 100%. In some examples, the protein activity level achieved is about 50%, about 55%, about 60%, about 65%, about 70%, about 75%, about 80%, about 85%, about 90%, about 95%, or about 100%. In some examples, the protein activity level achieved is at least about 50%, about 55%, about 60%, about 65%, about 70%, about 75%, about 80%, about 85%, about 90%, or about 95%. In some examples, the protein activity level achieved is at most about 55%, about 60%, about 65%, about 70%, about 75%, about 80%, about 85%, about 90%, about 95%, or about 100%.
[0519] In some examples, the amount of correctly joined RNA or full-length nucleic acid editing protein produced in a cell, for example in combination with expression of one or more gRNAs, is sufficient to ameliorate or cure a condition or disease in a subject, as understood by one of skill in the art for the particular condition or disease. In some examples, the amount of correctly joined RNA or full-length nucleic acid editing protein (for example in combination with expression of one or more gRNAs) produced in a cell is an effective amount. In some examples, this amount is equivalent to about 50% to 100% the amount of the RNA or protein produced in a normal cell. In some examples, this amount is equivalent to about 40% to about 100% the amount of the RNA or protein produced in a normal cell. In some examples, this amount is equivalent to about 40% to about 45%, about 40% to about 50%, about 40% to about 55%, about 40% to about 60%, about 40% to about 65%, about 40% to about 70%, about 40% to about 75%, about 40% to about 80%, about 40% to about 85%, about 40% to about 90%, about 40% to about 100%, about 45% to about 50%, about 45% to about 55%, about 45% to about 60%, about 45% to about 65%, about 45% to about 70%, about 45% to about 75%, about 45% to about 80%, about 45% to about 85%, about 45% to about 90%, about 45% to about 100%, about 50% to about 55%, about 50% to about 60%, about 50% to about 65%, about 50% to about 70%, about 50% to about 75%, about 50% to about 80%,SLR 7158-102574-27 05 / 08 / 25 S2024-016A
[0520] about 50% to about 85%, about 50% to about 90%, about 50% to about 100%, about 55% to about 60%, about 55% to about 65%, about 55% to about 70%, about 55% to about 75%, about 55% to about 80%, about 55% to about 85%, about 55% to about 90%, about 55% to about 100%, about 60% to about 65%, about 60% to about 70%, about 60% to about 75%, about 60% to about 80%, about 60% to about 85%, about 60% to about 90%, about 60% to about 100%, about 65% to about 70%, about 65% to about 75%, about 65% to about 80%, about 65% to about 85%, about 65% to about 90%, about 65% to about 100%, about 70% to about 75%, about 70% to about 80%, about 70% to about 85%, about 70% to about 90%, about 70% to about 100%, about 75% to about 80%, about 75% to about 85%, about 75% to about 90%, about 75% to about 100%, about 80% to about 85%, about 80% to about 90%, about 80% to about 100%, about 85% to about 90%, about 85% to about 100%, or about 90% to about 100% the amount of the RNA or protein produced in a normal cell. In some examples, this amount is equivalent to about 40%, about 45%, about 50%, about 55%, about 60%, about 65%, about 70%, about 75%, about 80%, about 85%, about 90%, or about 100% the amount of the RNA or protein produced in a normal cell. In some examples, this amount is equivalent to about at least about 40%, about 45%, about 50%, about 55%, about 60%, about 65%, about 70%, about 75%, about 80%, about 85%, or about 90% the amount of the RNA or protein produced in a normal cell. In some examples, this amount is equivalent to about at most about 45%, about 50%, about 55%, about 60%, about 65%, about 70%, about 75%, about 80%, about 85%, about 90%, or about 100% the amount of the RNA or protein produced in a normal cell.
[0521] The measurements of RNA or protein used to determine recombination efficiency or production level can be made by any suitable method known to those of skill in the art. In some examples, recombination efficiency or production level is determined by measuring an amount of functional protein expressed, for example by Western blotting. In some examples, recombination efficiency or production level is determined by measuring the RNA transcript, for example using two probe based quantitative realtime PCR. For example, the first assay spans a sequence fully contained in the 3’ exonic coding sequence (labelled 3’ probe). The second assay spans the junction between the 5’ and the 3’ exonic coding sequence (labelled junction probe). Reconstitution efficiency can be calculated as the ratio of (junction probe count) / (3’ probe count). “Reconstitution efficiency,” “recombination efficiency,” and “splicing efficiency” are used interchangeably herein.
[0522] The level of expression of a reconstituted protein, or the level of successful editing of a nucleic acid sequence (and, e.g., production of an encoded protein) achieved using the compositions and methods provided herein may be evaluated based on an indirect measurement, e.g., level of a representative marker, surrogate, or functional indicator. Any such marker known to those of skill in the art may be used for evaluating the level of a protein reconstituted or edited according to the present disclosure, and compared with a control accordingly. For example, the formation of Dystrophin-glycoprotein complex (DGC) or a subcomplex thereof can be representative of restoration of dystrophin function. See, e.g., Gao and McNally, 2015, “The Dystrophin Complex: structure, function and implications for therapy” Compr Physiol. 2015 5(3): 1223-39, and Omairi, et al., 2019, “Regulation of the dystrophinassociated glycoprotein complexSLR 7158-102574-27 05 / 08 / 25 S2024-016A
[0523] composition by the metabolic properties of muscle fibres,” Scientific Reports (2019) 9:2770, each incorporated herein by reference in its entirety.
[0524] In some examples, a dimerization domain is about 20 to about 1000 nt, or about 20 to 160 nt, or about 50 to about 160 nt, or about 30 to about 130 nt, or about 50 to about 500 nt, or about 50 to 1000 nt, or about 20 to 50 nt, or about 30 to 50 nt, wherein reconstitution efficiency results in production of an effective amount of correctly joined RNA or full-length nucleic acid editing protein. In some examples, a dimerization domain is about 50 to about 160 nt, wherein reconstitution efficiency results in production of an effective amount of correctly joined RNA or full-length nucleic acid editing protein.
[0525] Achieving efficient recombination between multiple RNA molecules allows for packaging and delivery of transgenes into AAVs, which exceed the packaging limit of a single AAV. AAV packaging limits represent a major hurdle for gene therapy approaches for diseases caused by the absence / defect of large genes. One application of this system is expression of a nucleic acid editing protein and one or more gRNAs specific for the target gene, using viral vectors with restricted packaging capacity. Disease and genes include but are not limited to (Disease (gene, OMIM gene identifier)): 1) Duchenne muscular dystrophy and Becker muscular dystrophy (dystrophin, OMIM:300377); 2) Dysferlinopathies (Dysferlin, OMIM:603009); 3) Cystic fibrosis (CFTR, OMIM:602421); 4) Usher’s Syndrome IB (Myosin VIIA, OMIM:276903); 5) Stargardt disease 1 (AABCA4, OMIM:601691); 6) Hemophilia A (Coagulation Factor VIII, OMIM:300841); 7) Von Willebrand disease (von Willebrand Factor, OMIM:613160): 8) Marfan Syndrome (Fibrillin 1, OMIM: 134797); 9) Von Recklinghausen disease (neurofibromatosis- 1, OMIM: 162200), 10) USH1F, OMIM: 602083, and 11) hearing loss (OTOF, OMIM: 603681). In some embodiments, the target nucleic acid is in a gene that (when wild-type) encodes a protein selected from Dystrophin; Dysferlin; Myosin VIIA; Fibrillin 1; Neurofibromatosis- 1; β-globin chain of hemoglobin; Clotting factor I; Clotting factor II; Clotting factor III; Clotting factor IV; Clotting factor V; Clotting factor VI; Clotting factor VII; Clotting factor VIII; Clotting factor IX; Clotting factor X; Clotting factor XI; Clotting factor XII; Clotting factor XIII; HBA1; HBA2; HBB; HBD; von Willebrand factor; MTHFR; FANCA; FANCC; FANCD2; FANCG; FANCJ; AD AMTS 13; Factor V Leiden Prothrombin; IL-2RG, JAK3, IL-2 receptor gamma chain; IL-4 receptor gamma chain; IL-7 receptor gamma chain; IL-9 receptor gamma chain; IL-15 receptor gamma chain; IL-21 receptor gamma chain; RAG1; RAG2; CXCR4; IL7 receptor; ADA; PNP; WAS; CYBA, CYBB, NCF1, NCF2, NCF4; Beta-2 integrin; C-C chemokine receptor type 5 (CCR5), MSRB1; CSCR4; P17; PSIP1; CCR5; DMD; G6Pase; CEP290; AABCA4; MAGT1; arylsulfatase A (ARSA); ABCD1; IDS; IDUA; IDUA; SGSH; NAGLU; HGSNAT; GNS; GALNS; GLB1; ARSB; GUSB; HYAL1; MAN2B1; SMPD1; NPC1; NPC2; CFTR; PKD-1; PDK-2; PDK-3; HEXA; GBA; HTT; NF-1; NF2; APOB; LDLR; LDLRAP1; PCSK9; BCR-ABL; ASXL1; RUNX2; EPHA1; PD-1; Androgen receptor; E6; E7; CD; NGF; ARSA; MBP; WASP; AADC; CLN2; ASPA; GAN; MT-ND4; SGSH; SUMF1; GAD; NTRN; TH; CHI; GDNF; GAA; SMN; PCDH15, and thymidine kinase. Others are provided in Tables 1-4. Expression of the nucleic acid editing protein, and for example or more gRNAs specific for the target nucleic acid, can be expressed using the disclosed systems provided herein, forSLR 7158-102574-27 05 / 08 / 25 S2024-016A
[0526] example to treat genomic point mutations or activate or overexpress genes. Delivery of a nucleic acid editing protein can be achieved by splitting it into multiple fragments using the approach provided herein.
[0527] Additional applications of the disclosed methods and systems include intersectional gene delivery for targeted gene expression. One can make use of differential infection / expression patterns of two viruses encoding a fragmented gene. The reconstituted protein will get expressed in an overlapping population of cells that represents the intersection of what either virus would express in on its own. Examples for such an application may include: (1) delivery of two halves (or three thirds, or other portions) of a protein using retrogradely transported viral vectors from two (or more) projection targets to label bifurcating dual projection neurons, (2) delivery of one fragment under the control of a promoter that is active in population A and the second fragment from a promoter active in population B to specifically tag / manipulate the AUB population, (3) delivery of the first half of a protein with a viral vector that has a tropism for population A and the second half with a viral vector that has a tropism for population B to specifically tag / manipulate the AUB population. Or, combinations of these approaches.
[0528] In one example the dimerization domains are aptamer sequences, for example to facilitate dimerization in the presence of a (a) small molecular trigger recognized by the aptamers, or a (b) protein that is present in the cell binding to the two halves and therefore stimulating dimerization.
[0529] In some embodiments the RNA-RNA interactions necessary for end-joining can be controlled positively or negatively by other nucleotides such as (a) an antisense oligonucleotide sequence with homology to the two halves (ssDNA triggered dimerization). In such an example, an antisense oligonucleotide having a complementary sequence to both halves bridges the two molecules together, thus facilitating spliceosome mediated recombination of the two molecules, (b) an antisense oligonucleotide sequence with homology to one of the two joining-RNAs could occlude RNA-dimerization of the two molecules and serve as an off-switch for gene expression, or (c) an endogenous cellular RNA with homology to the two halves (RNA triggered dimerization). In such an example, a cellular RNA (e.g., mRNA or retroelement) having a complementary sequence to both halves bridges the two molecules together, thus facilitating spliceosome mediated recombination of the two molecules.
[0530] These molecule, protein, or RNA mediated interactions allow for controllable / fine tuned gene expression levels: Through titrating in molecules that interact with the binding domains (e.g., antisense oligonucleotides, small molecules, endogenous cellular RNAs), dimerization efficiency between the two halves can be modulated to regulate expression levels independent of promoter activity. Such an installment can be used if a narrow range of protein expression levels are needed.
[0531] In some examples, splitting a coding sequence into two or more fragments, results in the production of both full-length protein and unspliced REJ RNAs (un-joined fragments) (FIG. 26). The un-joined fragments can leave the nucleus and enter the cytosol. The presence of such un-joined fragments can result in the expression of dominant negative proteins; those that bind or interact with a target but do not function properly. Thus, in some aspects, provided herein are systems and methods to reduce the expression of, or increase the degradation of, un-joined fragments.SLR 7158-102574-27 05 / 08 / 25 S2024-016A
[0532] In one aspect, the disclosed systems include one or more secondary structures in the REJ introns of the donor and acceptor (130 and 170 of FIG. 6A), to inhibit ribosome assembly (FIG. 27). This secondary structure is spliced out, and thus not present in the expressed protein. Exemplary secondary structures that can be used include hairpins, kissing stem loops, and quadruplexes. In one aspect, the secondary structure is present in intronic sequence 130, for example in DISE 118, ISE 120, dimerization domain 122, poly adenylation sequence 124. or splice donor sequence 116. In one aspect, the secondary structure is present in intronic sequence 170, for example in second dimerization domain 154, ISE 156, branching point 158, polypyrimidine tract 160, or splice acceptor sequence 162. Thus, in some examples, an REJ intron 130 and / or 170 includes one or more hairpins, kissing stem loops, quadruplexes, or combinations thereof. In one aspect, REJ intron 130 and / or 170 includes one or more hairpins, such as at least 2, at least 3, or at least 4 hairpins, such as 1, 2, 3, 4 or 5 hairpins. In one aspect, REJ intron 130 and / or 170 includes one or more kissing stem loops, such as at least 2, at least 3, or at least 4 kissing stem loops, such as 1, 2, 3, 4 or 5 kissing stem loops. In one aspect, REJ intron 130 and / or 170 includes one or more quadruplexes, such as at least 2, at least 3, or at least 4 quadruplexes, such as 1, 2, 3, 4 or 5 quadruplexes. In one aspect, REJ intron 130 and / or 170 includes one or more hairpins and one or more kissing stem loops. In one aspect, REJ intron 130 and / or 170 includes one or more hairpins and one or more quadruplexes. In one aspect, REJ intron 130 and / or 170 includes one or more quadruplexes and one or more kissing stem loops. In one aspect, REJ intron 130 and / or 170 includes one or more hairpins, one or more quadruplexes and one or more kissing stem loops. In one aspect, a hairpin is about 10-20 bp, such as 10, 11, 12, 13, 14, 15, 16, 17, 18, 19 or 20 bp. In one aspect, a kissing stem loop is about 8-20 bp, such as 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19 or 20 bp. In one aspect, a quadruplex is about 20-40 bp, such as 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39 or 40 bp.
[0533] In one aspect, the disclosed systems include one or more elements in the REJ introns of the donor and acceptor (130 and 170 of FIG. 6A), to recruit nuclear RNA binding proteins within the REJ intron to prevent nuclear export of unspliced REJ RNA (FIG. 28). Such aspects can retain RNA in the nucleus longer, until it is spliced, thereby reduce unjoined fragments in the cytoplasm. In one aspect, the one or more elements are present in intronic sequence 130, for example in DISE 118, ISE 120, dimerization domain 122, polyadenylation sequence 124. or splice donor sequence 116. In one aspect, the one or more elements are present in intronic sequence 170, for example in second dimerization domain 154, ISE 156, branching point 158, polypyrimidine tract 160, or splice acceptor sequence 162. The elements may also be placed adjacent to the above listed sequences. Examples of such elements include those that bind to nuclear RNA, such as the MALAT1 Region M Sequence (SEQ ID NO: 246) and a SIROLIN element (SEQ ID NO: 227). In some examples, the element is about 50-200 ribonucleotides / nucleotides in length, such as 50-150, 75-125, or about 100 ribonucleotides / nucleotides in length. In some aspects, two or more different elements are included, such as a MALAT 1 Region M sequence and a SIROLIN element. In some aspects, multiple of the same element is included, such as a least two SIROLIN elements, such as 2, 3, 4, 5, 6, 7, 8, 9 or 10SLR 7158-102574-27 05 / 08 / 25 S2024-016A
[0534] SIROLIN elements, such as a least two MALAT1 Region M sequences, such as 2, 3, 4, 5, 6, 7, 8, 9 or 10 MALAT1 Region M sequences.
[0535] In one aspect, a protein to be expressed by the disclosed REJ system is a protein that has more than one splice variant (e.g., more than one isoform), such as at least two splice variants, such as 2, 3, 4, 5, 6, 7, 8, 9 or 10 splice variants. In one aspect, a non varied end of the protein can be expressed from a single vector, and each varied end can be expressed from a different vector. For example, if the protein has two splice variants at the C-terminus, 3 vectors can be used: 1 to express the constant N-terminus, 1 to express variant 1 C terminus, and 1 to express variant 2 C terminus. In another example, if the protein has two splice variants at the N-terminus, 3 vectors can be used: 1 to express the constant C-terminus, 1 to express variant 1 N terminus, and 1 to express variant 2 N terminus. In one example, both portions of a target protein have more than one isoform, in which case multiple N-terminus and C-terminus vectors can be generated and expressed to generate combinations between each N and C terminus. Thus, multiple RNA molecules can be used to express distinct forms of the N-terminal portion of the target protein and / or multiple RNA molecules can be used to express distinct forms of the C-terminal portion of the target protein. In one example, the protein expressed is PCDH15 (see FIGS. 31 and 36). In one example, the protein with multiple isoforms expressed is ABCA4, dystrophin, or Myo7A.
[0536] III. Systems
[0537] Provided herein is a system that can be used to recombine two or more RNA molecules, such as at least two, at least three, at least four, or at least five different RNA molecules (such as 2, 3, 4, 5, 6, 7, 8, 9 or 10 different RNA molecules) using synthetic introns containing dimerization sequences. Unlike fragmentation and reconstitution of two fragments at the protein level, the disclosed approach does not require extensive protein engineering to find a suitable split point. Reconstitution on an RNA level allows for seamless joining of two fragments of a protein. The disclosed methods and systems allow for large genes (and corresponding proteins), such as those greater than about 4.5 kb, at least 5 kb, at least 5.5 kb, at least 6 kb, at least kb, at least 8 kb, at least 8 kb, at least 10 kb, at least 13.5 kb, or at least 18 kb, to be divided into two or more fragments or portions, which can each be introduced into a cell or subject via separate vectors, such as multiple AAV. In some examples, the disclosed methods and systems allow for genes (and corresponding proteins), such as those at least greater than about 2 kb, at least 2.5 kb, at least 3 kb, at least 3.5 kb, or at least 4 kb, to be divided into two or more fragments or portions, which can each be introduced into a cell or subject via separate vectors, such as multiple AAV. In some examples, the nucleic acid and protein are a nucleic acid editing nucleic acid / protein. In some examples, the nucleic acid and protein are a target nucleic acid / protein, such as one that causes a disease. In some examples, such vectors can further include one or more gRNA coding sequences specific for one or more target nucleic acid molecules (e.g., specific for one or more target sites to be edited, such as deletion, insertion, or substitution of one or more nucleotides or ribonucleotides). In some examples, multiple copies of the same one or more gRNA coding sequences are present, for example to increase the number of gRNA molecules expressedSLR 7158-102574-27 05 / 08 / 25 S2024-016A
[0538] from the vector. In one example, the system includes two portions for recombining two RNA molecules, for example wherein the target protein is encoded by at least about 4500 nt to about 9000 nt, such as 4000 nt to 5000 nt. In one example, the system includes three portions for recombining three RNA molecules, for example wherein the nucleic acid editing protein is encoded by up to about 13,500 nt, such as about 2000 nt to about 13,500 nt or 3000 nt to 5000 nt. In one example, the system includes four portions for recombining four RNA molecules, for example wherein the nucleic acid editing protein is encoded by up to about 18,000 nt, such as about 2000 nt to about 18,000 nt or 2000 nt to 5000 nt. This helps to overcome the limited space available in vectors. In some examples, an endogenous promoter length limits the capability of its corresponding gene to be expressed in an AAV. In some examples, a coding sequence length limits its capability to be expressed in an AAV. In some examples, an endogenous promoter length and its coding sequence length limits their capability to be expressed together in an AAV. The disclosed systems can be used to express such long sequences that have been previously difficult to express in AAV. The disclosed systems can also be used to express numerous copies of one or more gRNAs, for example in combination with a nucleic acid editing protein, as the amount of gRNAs can be rate limiting. In some examples, the disclosed DNAs and systems express at least 2 gRNAs, at least 3 gRNAs, at least 4 gRNAs, at least 5 gRNAs, at least 10 gRNAs, at least 20 gRNAs, at least 50 gRNAs, at least 75 gRNAs, at least 100 gRNAs, at least 200 gRNAs, at least 500 gRNAs, or at least 1000 gRNAs, wherein the gRNAs can target the same nucleic acid molecule and the same target site, target the same nucleic acid molecule and two or more target sites within the same target nucleic acid molecule, target two or more different nucleic acid molecules (such as 1 or more target sites within each different target nucleic acid molecule), or combinations thereof.
[0539] In some examples, the one or more gRNAs target a nucleic acid molecule, such as a gene, associated with disease, such as a monogenic disease, recessive genetic disease, a disease caused by a mutation in a gene. Examples of such diseases include, but are not limited to, hemophilia A (caused by mutations in the F8 gene, 7kb coding region, also referred to as Coagulation Factor VIII), hemophilia B (caused by mutations in the F9 gene), Duchenne muscular dystrophy (caused by mutations in the dystrophin gene, 11 kb coding region), sickle cell anima (caused by mutation in beta globin domain of hemoglobin, which has a promoter of about 3.5 kb), Stargardt disease (caused by mutations in the AABCA4 gene, 6.9 kb coding region), Usher syndrome (caused by a mutation in MYO7A, 7 kb coding region, resulting in hearing loss and visual impairment). In one example, the gene is one caused by a point mutation in a gene.
[0540] In one example, the one or more gRNAs target a nucleic acid molecule, such as a gene, to treat a disease, such as a cancer, such as a cancer of the breast, lung, prostate, liver, kidney, brain, bone, ovary, uterus, skin, or colon.
[0541] In some examples, an RNA sequence encoding the target nucleic acid editor and used in the disclosed methods and systems are codon optimized for expression in a target organism or cell, such as codon optimized for expression in a human, canine, pig, feline, mouse, or rat cell. Thus, in some examples, the RNA coding sequence includes preferred codons (e.g., does not include rare codons with low utilization). Codon optimization can be performed by identifying abundant tRNA levels in the targetSLR 7158-102574-27 05 / 08 / 25 S2024-016A
[0542] organism or cells. In some examples, an RNA sequence encoding the protein is de-enriched for cryptic splice donor and acceptor sites to maximize an RNA recombination reaction.
[0543] In some examples, the first RNA molecule of a disclosed REJ system (e.g., N-terminal encoding) includes one or more secondary structures, the second RNA molecule of a disclosed REJ system (e.g., C-terminal encoding) includes one or more secondary structures, or both. The one or more secondary structures are not in the coding sequence for the N-terminal portion of the target protein nor in the coding sequence for the C-terminal portion of the target protein. In some examples, the first RNA molecule of a disclosed REJ system (e.g., N-terminal encoding) includes one or more one or more elements to recruit nuclear RNA binding proteins, the second RNA molecule of a disclosed REJ system (e.g., C-terminal encoding) includes one or more one or more elements to recruit nuclear RNA binding proteins, or both. The one or more elements are not in the coding sequence for the N-terminal portion of the target protein nor in the coding sequence for the C-terminal portion of the target protein.
[0544] In some examples, a target protein includes two or more splice variants (e.g., isoforms) of the N-terminal portion of the target protein, and the first RNA molecule includes two or more separate RNA molecules, wherein two or more separate RNA molecules each encode a distinct form of the N-terminal portion of the target protein. In some examples, the target protein includes two or more splice variants of the C-terminal portion of the target protein, and the second RNA molecule comprises two or more separate RNA molecules, wherein two or more separate RNA molecules each encode a distinct form of the C-terminal portion of the target protein. In some examples, a target protein includes two or more splice variants (e.g., isoforms) of the N-terminal portion of the target protein, and two or more splice variants (e.g., isoforms) of the C-terminal portion of the target protein. In such examples, multiple vectors, each encoding one isoform, can be expressed using the REJ systems provided herein.
[0545] In some examples, a nucleic acid editing protein or a target protein is divided into two portions, such as about two equal halves (or other proportions, such as portion A expressing about 1 / 3 and portion B expressing about 2 / 3, or portion A expressing about 1 / 4 and portion B expressing about 3 / 4, etc.). However, it is not required that each portion be the same number of nucleotides (or encode the same number of amino acids). In such an example, the method can use two synthetic nucleic acid molecules (e.g., RNA or DNA encoding such RNA), one which includes a coding sequence for an N-terminal portion of the protein, and another which includes a coding sequence for a C-terminal portion of the protein. Based on this foundation, one skilled in the art will appreciate that in addition to dividing a protein into two fragments or portions, proteins can be divided or split into more than two fragments, such as three fragments. The design principle of the intronic sequences of three RNA molecules is similar to that of the two, but instead a different pair of dimerization domains for one of the two junctions is utilized. Thus, for example, an N-terminal protein coding sequence is followed by an intronic sequence with a specific binding domain (e.g., first dimerization sequence), the middle coding sequence includes an intronic sequence with a complementary sequence to the first dimerization sequence (second dimerization sequence). The middle coding fragment is followed by another intronic fragment with another dimerization sequence (third dimerization sequence, different fromSLR 7158-102574-27 05 / 08 / 25 S2024-016A
[0546] the second dimerization sequence). The third fragment includes the C-terminal coding sequence of the protein, and includes an intronic region with a dimerization sequence (fourth dimerization sequence) complementary to the third dimerization sequence. In the use of more than one middle portion, the two middle portions may be referred to as a middle portion and a first middle portion, or as a first middle portion and a second middle portion, or as a first middle portion, a second middle portion and a third middle portion, etc., in a way understood to distinguish the respective portions.
[0547] In one example, a nucleic acid editing protein or target protein is divided into an N-terminal portion and a C-terminal portion (e.g., divided in roughly half, or unequal apportionment, such as 1 / 3 and 2 / 3 or 1 / 4 and 3 / 4), which can be reconstituted using the disclosed systems and methods. Referring to FIGS. 6A and 6G, in such an example, the system includes at least two synthetic nucleic acid molecules 110, 150. Each nucleic acid molecule 110, 150 can be composed of DNA or RNA (if RNA, corresponding promoters 112, 152 are absent). In some examples, each of 110, 150 is about at least 100 nucleotides / ribonucleotides (nt) in length, such as at least 200, at least 300, at least 500, at least 1000, at least 2000, at least 3000, at least 4000, at least 5000, at least 6000, at least 7000, at least 8000 nt, at least 10,000 nt, such as 200 to 10,000 nt, 200 to 8000 nt, 500 to 5000 nt, 800 to 3000 nt, 1000 to 300 nt, or 200 to 1000 nt. The molecules 110, 150 can include natural and / or non-natural nucleotides or ribonucleotides.
[0548] Molecule 1 10 is the 5’-located molecule of the system, as it includes a splice donor 116. In embodiments where molecule 110 is DNA, it includes a promoter 112 operably linked to a sequence encoding an RNA molecule, the RNA molecule comprising from 5’ to 3’: a coding sequence for an N-terminal portion of the target protein 114, wherein the coding sequence for an N-terminal portion of the target protein 114 comprises a splice junction at a 3’-end of the target protein coding sequence, SD 116, optional DISE 118, optional ISE 120, dimerization domain 122, and optional polyadenylation sequence 124. Any promoter 112 (or enhancer) can be used, such as one that utilizes RNA polymerase II, such as a constitutive or inducible promoter. In some examples, promoter 112 is a tissue-specific promoter, such as one constitutively active in muscle tissue (such as skeletal or cardiac), optical tissue (such as retinal tissue), inner ear tissue, liver tissue, pancreatic tissue, lung tissue, skin tissue, bone, or kidney tissue. In some examples, promoter 112 is a cell-specific promoter, such as one constitutively active in a cancer cell, or a normal cell. In some examples, promoter 112 is an endogenous promoter of the protein expressed, and in some example is long (e.g., at least 2500 nt, at least 3000 nt, at least 4000 nt, at least 5000 nt, or at least 7500 nt). In some examples, promoter 112 is at least about 50 nucleotides (nt) in length, such as at least 100, at least 200, at least 300, at least 500, at least 1000, at least 2000, at least 3000, at least 4000, at least 5000, at least 6000, at least 7000, at least 8000 nt, at least 9000 nt, or at least 10,000 nt, such as 50 to 10,000 nt, 100 to 5000 nt, 500 to 5000 nt, or 50 to 1000 nt in length. In some examples, molecule 110 is DNA, and is at least 200, at least 300, at least 500, at least 800, at least 1000, at least 2000, at least 3000, at least 4000, at least 5000, at least 6000, at least 7000, or at least 8000 nt, such as 200 to 10,000 nt, 200 to 8000 nt, 500 to 5000 nt, 800 to 3000 nt, 1000 to 300 nt, or 200 to 1000 nt in length. As shown in FIG. 6F, in embodiments where molecule 110 is RNA, for example after transcription of the DNA into RNA, molecule 110 does notSLR 7158-102574-27 05 / 08 / 25 S2024-016A
[0549] include promoter 112, and 114 is the RNA encoded by the coding sequence for an N-terminal portion of the nucleic acid editing protein. In some examples, molecule 110 is RNA, does not include promoter 112, and is at least 200, at least 300, at least 500, at least 1000, at least 2000, at least 3000, at least 4000, at least 5000, at least 6000, at least 7000, or at least 8000 nt, such as 200 to 10,000 nt, 200 to 8000 nt, 500 to 5000 nt, 800 to 3000 nt, 1000 to 300 nt, or 200 to 1000 nt in length. The molecule 110 (with or without promoter 112) can include natural and / or non-natural nucleotides or ribonucleotides.
[0550] The splice junction around the 3’ end of the N-terminal coding sequence (or RNA sequence encoded thereby) 114 can match the consensus sequence found in the target cell or organism into which molecules 110, 150 are introduced. In humans the splice junction sequence is AG (adenine-guanine) or UG (uracil-guanine) at position -1 and -2 of the 5’ splice site for U2-dependent introns or AG, UG, CU (cytosineuracil), or UU for U12-dependent introns. Thus, in some examples, the splice junction is 2 nt in length, and the 3’ end of the N-terminal coding portion 114 is AG, UG, CU or UU. In some examples a DNA molecule encoding a portion of a nucleic acid editing protein comprises sequences that encode parts of multiple splice junctions, e.g., at the 3’ end of the DNA molecule encoding the N-terminal portion of the nucleic acid editing protein, and at the 5’ end of the DNA molecule encoding the C-terminal portion of the nucleic acid editing protein.
[0551] The remaining 3’ -terminal portion of molecule 110 is intronic, 130. In some examples, intronic sequence 130 is about at least 10 nt, such as at least 20 nt, at least 50 nt, at least 100 nt, at least 250 nt, at least 250 nt, at least 300 nt, at least 400 nt, or at least 500 nt in length, such as 20 to 500, 20 to 250, 20 to 100, 50 to 100, or 50 to 200 nt in length. Immediately following N-terminal coding sequence (or RNA encoded thereby) 114 is a splice donor (SD) 116 (such as a SD consensus sequence, such as a SD human consensus sequence). Thus SD 116 of intronic sequence 130 is 3’ to N-terminal coding sequence 114. SD 116 forms a recognition sequence for the spliceosome components to bind to the RNA molecule. The sequence of SD 116 can be a SD consensus sequence found in the target cell or organism into which molecules 110, 150 are introduced. In some examples, SD 116 is at least 2 nt, such as at least 5 nt, or at least 10 nt in length, such as 2 to 10, 2 to 8, 2 to 5 or 5 to 10 nt. The SD 116 can be used to recruit U2 or U12 dependent splicing machinery. In one example, U2 dependent splicing is used in human cells, and the SD 116 sequence includes or is GUAAGUAUU. In one example, U12 dependent splicing is used in human cells, and the SD 116 sequence includes or is AUAUCCUUUUUA (SEQ ID NO: 137) or GUAUCCUUUUUA (SEQ ID NO: 138). Throughout, it is understood that RNA sequences can be described using nucleotides A, G, U and C, and that DNA sequences can be described using nucleotides A, G, T and C. It is also understood that sequences described herein as comprised by a DNA molecule that is transcribed to form one or more RNA molecules, and sequences described as comprised by an RNA molecule that may be translated (i.e., a protein coding sequence) or not translated (e.g., a gRNA), are the sequences with the intended function as can be recognized by one of skill in the art. For example, a gRNA sequence present in a DNA molecule of the disclosure has a sequence that is functional following its transcription. As another example, a target protein coding sequence present in a DNA molecule of theSLR 7158-102574-27 05 / 08 / 25 S2024-016A
[0552] disclosure has a sequence that can be translated into the target protein from mRNA that is transcribed from the DNA. A sequence referred to as encoded by the DNA or RNA molecule, or comprised by the DNA or RNA molecule, is one that results in the intended, functional product, as understood by one of skill in the art reading the present disclosure.
[0553] Intronic sequence 130 optionally includes one or both of a set of splicing enhancer sequences referred to as downstream intronic splice enhancer (DISE) 118 and intronic splice enhancer (ISE) 120, which stimulate action (e.g., increase activity) of the spliceosome. In some examples, intronic sequence 130 includes at least two splicing enhancer sequences, such as at least 3, at least 4, or at least 5 splicing enhancer sequences. Exemplary splicing enhancer sequences include DISE 118 and ISE 120. In some examples, inclusion of one or more splicing enhancer sequences 118, 120 in intronic sequence 130 increases splicing efficiency by at least 20%, at least 30%, at least 40%, at least 50%, at least 75%, at least 80%, at least 90% or at least 95%. Exemplary splicing enhancer sequences that can be used are provided in SEQ ID NOS: 26-136, 151, and 152, as well as GGGTTT, GGTGGT, TTTGGG, GAGGGG, GGTATT, GTAACG, GGGGGTAGG, GGAGGGTTT, GGGTGGTGT TTCAT, CCATTT, TTTTAAA, TGCAT, TGCATG, TGTGTT, CTAAC, TCTCT, TCTGT, TCTTT, TGCATG, CTAAC, CTGCT, TAACC, AGCTT, TTCATTA, GTTAG, TTTTGC, ACTAAT, ATGTTT, CTCTG, GGG, GGG(N)2-4GGG, TGGG, YCAY, UGCAUG, or 3x(G3–6N1–7). In some examples, if DISE 118 is present, can be at least 3 nt, at least 4 nt, at least 5 nt, at least 10 nt, at least 25 nt, at least 50 nt, at least 75 nt, or at least 100 nt in length, such as 3 to 10, 3 to 11, 4 to 11, 5 to 11, 10 to 50, 5 to 100, 10 to 25, 10 to 20, or 20 to 75 nt, the sequence of DISE 118 is or comprises CUCUUUCUUUTCCAUGGGUUGGCU (SEQ ID NO: 134), TGCATG, CTAAC, CTGCT, TAACC, AGCTT, TTCATTA, GTTAG, TTTTGC, ACTAAT, ATGTTT or CTCTG. In some examples, if ISE 120 is present, it can be about at least 3 nt, at least 4 nt, at least 5 nt, at least 10 nt, such as at least 20 nt, at least 25 nt, at least 30 nt, at least 40 nt, or at least 50 nt in length, such as 3 to 10, 3 to 11, 4 to 11, 5 to 11, 10 to 50, 20 to 25, 10 to 25, 10 to 20, or 20 to 40 nt in length. In one example, the sequence of ISE 120 is or comprises GGCUGAGGGAAGGACUGUCCUGGG (SEQ ID NO: 135), GGGUUAUGGGACC (SEQ ID NO: 136), TTCAT, CCATTT, TTTTAAA, TGCAT, TGCATG, TGTGTT, CTAAC, TCTCT, TCTGT, or TCTTT. In some examples, intronic sequence 130 includes at least two, at least 3, or at least 4 ISEs 120. In some examples, ISE 120 is or comprises at least one sequence having at least 80%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99% or 100% sequence identity to SEQ NO: 173, 174, 175, 176, 177, 178, 179, 180, 182, 183, 184, 185, 186, 187, 188, 189, 190, 191, 192, 193, 194, 195, 196, 199, 200, 201, 202, or 203, such as at least 2, at least 3 of such sequences, such as 1, 2, 3, 4 or 5 of such sequences. In some examples, DISE 118 is or comprises at least one sequence having at least 80%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99% or 100% sequence identity to SEQ NO: 173, 174, 175, 176, 177, 178, 179, 180, 182, 183, 184, 185, 186, 187, 188, 189, 190, 191, 192, 193, 194, 195, 196, 199, 200, 201, 202, or 203, such as at least 2, at least 3 of such sequences, such as 1, 2, 3, 4 or 5 of such sequences.SLR 7158-102574-27 05 / 08 / 25 S2024-016A
[0554] The SD 116 (and if present also enhancer sequences 118, 120) is followed 3’ by a dimerization domain 122 used to bring the N-terminal coding sequence (or RNA encoded thereby) 114, and C-terminal coding sequence 164 to be combined, together. Intronic sequence 130 portion of molecule 110 can optionally include at the 3’-end a polyadenylation site 124, which terminates transcription of that fragment. In some examples, poly adenylation sequence 124 is a poly A sequence of at least 15 As, such as 15 to 30 or 15 to 20 As.
[0555] In some examples, first dimerization domain 122 of molecule 110 (and second dimerization domain 154 of molecule 150) includes a plurality of unpaired nucleotides (that is, unpaired within the structure of the molecule 110 itself, or 150 itself). Having unpaired nucleotides in the dimerization domain allows the 5’ (or first) dimerization domain 122 and the 3’ (or second) dimerization domain 154 to interact through base pairing. Through this interaction, molecules 110 and 150 are kept in proximity which prompts the spliceosome to recombine the two molecules by joining the N-terminal coding region (or RNA encoded thereby) 114 and the C terminal coding region (or RNA encoded thereby) 164.
[0556] In one example, dimerization domain 122 (and 154) includes “hypodiverse sequences,” which contain a limited diversity of nucleotides and are thus unlikely to form stem loops with themselves in the secondary structure of each molecule 110, 150. Such a hypodiverse dimerization domain 122 (and 154) can be a relatively open configuration, independent of the sequences of the DNA encoding the N- and C-terminus of the protein (or RNA encoded thereby) 114, 164. This allows the nucleotides of the first dimerization domain 122 of molecule 110 to be available to form base pairs with the corresponding second dimerization domain 154 of molecule 150, allowing subsequent joining of the N-terminal coding sequence (or RNA encoded thereby) 114 and C-terminal coding sequence (or RNA encoded thereby) 164. In some examples, first and second dimerization domain 122, 154 includes hypodiverse sequences interspersed with sequences that can form a stem, which results in local RNA loops that are open and available for basepairing in the absence of pseudoknot formation (FIG. 6B). Exemplary hypodiverse sequences include a repeated series of Us (such as 30 to 500 Us), a repeated series of As (such as 30 to 500 As), a repeated series of Gs (such as 30 to 500 Gs), a repeated series of Cs (such as 30 to 500 Cs), a mixture containing only As and Gs (such as 30 to 500 As and Gs, e.g., AAAGAAGGAA(...) (SEQ ID NO: 149) which can be repeated), a mixture containing only Cs and Us (such as 30 to 500 Cs and Us, e.g., CUUUCUUUUCUU(...) (SEQ ID NO: 150) which can be repeated). Other exemplary hypodiverse sequences include complementary sequences that form helices flanked by hypodiverse sequences.
[0557] In some examples, first and second dimerization domain 122, 154 only include purines or only include pyrimidines. In one example, the first dimerization domain 122 only includes purines, while the second dimerization domain 154 only includes pyrimidines. In another example, the first dimerization domain 122 only includes pyrimidines, while the second dimerization domain 154 only includes purines. Due to the inability of purines to pair with themselves (and pyrimidines likewise) these stretches of RNA have an open predicted structure.SLR 7158-102574-27 05 / 08 / 25 S2024-016A
[0558] In some examples, first and second dimerization domain 122, 154 do not include cryptic splice acceptors that could compete with RNA recombination, such as sequences similar to the splice donor consensus sequence NNNAGGUNNNN (SEQ ID NO: 151) or NNNUGGUNNNN (SEQ ID NO: 152) (wherein N refers to any nucleotide). In some examples, first dimerization domain 122 is no more than 1000 nt, such as no more than 750 nt, or more than 500 nt, such as 6 to 1000 nt, 10 to 1000 nt, 20 to 1000 nt, 30 to 1000 nt, 30 to 750 nt, 30 to 500 nt, 50 to 500 nt, 50 to 100 nt, or 100 to 250 nt. In some examples, first dimerization domain 122 is greater than 50 nt, such as at least 51 nt, at least 100 nt, at least 150 nt, at least 161 nt, or at least 170 nt, such as 51 to 159 nt, 51 to 150 nt, 51 to 120 nt, 51 to 100 nt, or 51 to 70 nt. In some examples, first dimerization domain 122 is greater than 160 nt, such as at least 161 nt, at least 170 nt, at least 180 nt, at least 200 nt, at least 300 nt, at least 400 nt, at least 500 nt, at least 600 nt, at least 700 nt, at least 800 nt, at least 900 nt, or at least 1000 nt, such as 161 to 1000 nt, 161 to 500 nt, 161 to 300 nt, 161 to 200 nt, or 161 to 170 nt. In some examples, first dimerization domain 122 is less than 50 nt, such 6 to 49 nt, 6 to 45 nt, 6 to 40 nt, 6 to 30 nt, 6 to 20 nt, or 6 to 10 nt.
[0559] In some examples, a dimerization domain is 20 to 160 nt, 50-500 nt, or 500-1000 nt. In some examples, a dimerization domain is about 20 nt to about 160 nt. In some examples, a dimerization domain is about 20 nt to about 40 nt, about 20 nt to about 50 nt, about 20 nt to about 70 nt, about 20 nt to about 90 nt, about 20 nt to about 100 nt, about 20 nt to about 110 nt, about 20 nt to about 120 nt, about 20 nt to about 130 nt, about 20 nt to about 140 nt, about 20 nt to about 150 nt, about 20 nt to about 160 nt, about 40 nt to about 50 nt, about 40 nt to about 70 nt, about 40 nt to about 90 nt, about 40 nt to about 100 nt, about 40 nt to about 110 nt, about 40 nt to about 120 nt, about 40 nt to about 130 nt, about 40 nt to about 140 nt, about 40 nt to about 150 nt, about 40 nt to about 160 nt, about 50 nt to about 70 nt, about 50 nt to about 90 nt, about 50 nt to about 100 nt, about 50 nt to about 110 nt, about 50 nt to about 120 nt, about 50 nt to about 130 nt, about 50 nt to about 140 nt, about 50 nt to about 150 nt, about 50 nt to about 160 nt, about 70 nt to about 90 nt, about 70 nt to about 100 nt, about 70 nt to about 110 nt, about 70 nt to about 120 nt, about 70 nt to about 130 nt, about 70 nt to about 140 nt, about 70 nt to about 150 nt, about 70 nt to about 160 nt, about 90 nt to about 100 nt, about 90 nt to about 110 nt, about 90 nt to about 120 nt, about 90 nt to about 130 nt, about 90 nt to about 140 nt, about 90 nt to about 150 nt, about 90 nt to about 160 nt, about 100 nt to about 110 nt, about 100 nt to about 120 nt, about 100 nt to about 130 nt, about 100 nt to about 140 nt, about 100 nt to about 150 nt, about 100 nt to about 160 nt, about 110 nt to about 120 nt, about 110 nt to about 130 nt, about 110 nt to about 140 nt, about 110 nt to about 150 nt, about 110 nt to about 160 nt, about 120 nt to about 130 nt, about 120 nt to about 140 nt, about 120 nt to about 150 nt, about 120 nt to about 160 nt, about 130 nt to about 140 nt, about 130 nt to about 150 nt, about 130 nt to about 160 nt, about 140 nt to about 150 nt, about 140 nt to about 160 nt, or about 150 nt to about 160 nt. In some examples, a dimerization domain is about 20 nt, about 40 nt, about 50 nt, about 70 nt, about 90 nt, about 100 nt, about 110 nt, about 120 nt, about 130 nt, about 140 nt, about 150 nt, or about 160 nt. In some examples, a dimerization domain is at least about 20 nt, about 40 nt, about 50 nt, about 70 nt, about 90 nt, about 100 nt, about 110 nt, about 120 nt, about 130 nt, about 140 nt, or about 150 nt. In some examples, a dimerization domain is at most about 40 nt, about 50 nt, about 70 nt,SLR 7158-102574-27 05 / 08 / 25 S2024-016A
[0560] about 90 nt, about 100 nt, about 110 nt, about 120 nt, about 130 nt, about 140 nt, about 150 nt, or about 160 nt.
[0561] In some examples, a dimerization domain is about 50 nt to about 500 nt. In some examples, a dimerization domain is about 50 nt to about 100 nt, about 50 nt to about 150 nt, about 50 nt to about 200 nt, about 50 nt to about 250 nt, about 50 nt to about 300 nt, about 50 nt to about 350 nt, about 50 nt to about 400 nt, about 50 nt to about 500 nt, about 100 nt to about 150 nt, about 100 nt to about 200 nt, about 100 nt to about 250 nt, about 100 nt to about 300 nt, about 100 nt to about 350 nt, about 100 nt to about 400 nt, about 100 nt to about 500 nt, about 150 nt to about 200 nt, about 150 nt to about 250 nt, about 150 nt to about 300 nt, about 150 nt to about 350 nt, about 150 nt to about 400 nt, about 150 nt to about 500 nt, about 200 nt to about 250 nt, about 200 nt to about 300 nt, about 200 nt to about 350 nt, about 200 nt to about 400 nt, about 200 nt to about 500 nt, about 250 nt to about 300 nt, about 250 nt to about 350 nt, about 250 nt to about 400 nt, about 250 nt to about 500 nt, about 300 nt to about 350 nt, about 300 nt to about 400 nt, about 300 nt to about 500 nt, about 350 nt to about 400 nt, about 350 nt to about 500 nt, or about 400 nt to about 500 nt. In some examples, a dimerization domain is about 50 nt, about 100 nt, about 150 nt, about 200 nt, about 250 nt, about 300 nt, about 350 nt, about 400 nt, or about 500 nt. In some examples, a dimerization domain is at least about 50 nt, about 100 nt, about 150 nt, about 200 nt, about 250 nt, about 300 nt, about 350 nt, or about 400 nt. In some examples, a dimerization domain is at most about 100 nt, about 150 nt, about 200 nt, about 250 nt, about 300 nt, about 350 nt, about 400 nt, or about 500 nt.
[0562] In some examples, the sequence of first and second dimerization domains 122 and 154 are determined by in silico structure prediction screening (e.g., RNA folding structure prediction is used to screen a library of possible dimerization domain sequences; sequences with a large proportion of unpaired nucleotides in both the dimerization domain and the corresponding anti-dimerization domain are selected), hypodiverse nucleotide design (e.g., dimerization domain designed to include a stretch of hypodiverse sequence, such as a repeat sequence of only U, only A, only C, only G, only R (G and A), or only Y (U and C), the sequence cannot fold onto itself), or empirical screening (e.g., a library of dimerization domains and corresponding anti-dimerization domains are synthesized and screened for maximal recombination efficiency).
[0563] In some examples, the sequences of a first and a second dimerization domains are designed based on complementarity (e.g., reverse complementarity), so that the two dimerization domains are brought together through base pairing (e.g., FIG. 6C). In some examples, the dimerization domain is linear (i.e., does not form any intramolecular secondary structure), when it is bound to the other dimerization domain. In some examples, the dimerization domain is about 30 nt to about 70 nt, about 35 nt to about 70 nt, about 40 nt to about 70 nt, about 45 nt to about 70 nt, about 50 nt to about 70 nt, about 30 nt to about 65 nt, about 35 nt to about 65 nt, about 40 nt to about 65 nt, about 45 nt to about 65 nt, about 50 nt to about 65 nt, about 30 nt to about 60 nt, about 35 nt to about 60 nt, about 40 nt to about 60 nt, about 45 nt to about 60 nt, about 50 nt to about 60 nt, about 30 nt to about 55 nt, about 35 nt to about 55 nt, about 40 nt to about 55 nt, about 45 nt to about 55 nt, or about 50 nt to about 55 nt in length. In some examples, the sequences of the twoSLR 7158-102574-27 05 / 08 / 25 S2024-016A
[0564] dimerization domains include at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% complementarity. In some examples, the sequences of the two dimerization domains are 100% complementary.
[0565] In some examples, a first dimerization domain (e.g., domain 122) includes at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity to any of SEQ ID NOs: 257-271 and 336, and a second dimerization domain (e.g., domain 154) includes at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% complementarity to the sequence of the first dimerization domain. In particular examples, a first dimerization domain includes or consists of any of SEQ ID NOs: 257-271 and 336, and a second dimerization domain includes or consists of 100% complementarity to the sequence of the first dimerization domain. Exemplary sequences of a second dimerization domain are reverse complements of any of SEQ ID NOs: 257-271 and 336. The reverse complement can be readily determined for any given sequence. For example, the reverse complement of SEQ ID NO: 257 is ACGGGGCGGTGTTACGTGAGGCTTGCTTTTGTTTCGCTGTACCCGAAACG (5’ to 3’, SEQ ID NO: 377).
[0566] In some examples, the sequence of first and second dimerization domains 122, 154 are designed to contain one or more complementary RNA hairpin structures (also called stem loops) that can form strong kissing loop interactions with their counter parts. In some examples, kissing loops are used when two or more dimerization domains are used to join two or more portions of a coding sequence, such as three or more, four or more, or five or more dimerization domains, such as 3, 4, 5, 6, 7, 8, 9 or 10 dimerization domains (e.g., FIG. 6E). Each hairpin loop (or stem loop) of a kissing loop is composed of at least two complementary sequences (e.g., form a stem) separated by a region of non-complementary sequence (e.g., form a loop). In some examples, a dimerization domain can be composed of 1 or more (such as at least 2, at least 3, at least 4, or at least 5, such as 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20) loops (e.g., stem loops). In some examples with multiple loops, all or some of the loops can be repeated. In some examples with multiple loops, all or some loops can be different In some examples, each complementary sequence is about 4 to 100 nt, which are separated by a loop of about 3 to 20 nt. Base-pairing between the two complementary sequences results in a helix (or stem), for example of at least 4 bp, at least 5 bp, at least 10 bp, at least 20 bp, at least 30 bp, at least 40 bp, at least 50 bp, at least 75 bp, at least 90 bp, or at least 100 bp, such as 4 to 100 bp, 5 to 75 bp, or 10 to 50 bp. In some examples, the loop portion is at least 3 nt, at least 5 nt, at least 10 nt, at least 15 nt, or at least 20 nt, such as 3 to 20 nt, 5 to 15 nt or 5 to 10 nt, wherein the loop is not base paired. Complementary sequences between two hairpin loops result in base pairing, and generation of a kissing loop / kissing stem loop interaction. In some examples, the complementary sequences between the two hairpin loops occurs between at least 3 nucleotides of one loop with at least 3 nucleotides of a second loop, such as at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 11, at least 12, at least 13, at least 14, at least 15, at least 16, at least 17, at least 19, or at least 20 nt (such as 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19 or 20) of the first loop, with at least 4, at least 5, at least 6, atSLR 7158-102574-27 05 / 08 / 25 S2024-016A
[0567] least 7, at least 8, at least 9, at least 10, at least 11, at least 12, at least 13, at least 14, at least 15, at least 16, at least 17, at least 19, or at least 20 nt (such as 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19 or 20) of the second loop. In some examples, the complementary sequences between the two hairpin loops occurs between at least 15%, at least 20%, at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, or 100% of the total loop sequence.
[0568] In some examples, one dimerization domain includes at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity to any of SEQ ID NOS: 139-142, 153, 232-241, 272-335, and 378, or a reverse complement thereof.
[0569] In some examples, a first dimerization domain and a second dimerization domain include at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity to SEQ ID NOS: 272 and 273, respectively; SEQ ID NOS: 274 and 275, respectively; SEQ ID NOS: 276 and 277, respectively; SEQ ID NOS: 278 and 279, respectively; SEQ ID NOS: 280 and 281, respectively; SEQ ID NOS: 282 and 283, respectively; SEQ ID NOS: 284 and 285, respectively; SEQ ID NOS: 286 and 287, respectively; SEQ ID NOS: 288 and 289, respectively; SEQ ID NOS: 290 and 291, respectively; SEQ ID NOS: 292 and 293, respectively; SEQ ID NOS: 294 and 295, respectively; SEQ ID NOS: 296 and 297, respectively; SEQ ID NOS: 298 and 299, respectively; SEQ ID NOS: 300 and 301, respectively; SEQ ID NOS: 302 and 303, respectively; SEQ ID NOS: 304 and 305, respectively; SEQ ID NOS: 306 and 307, respectively; SEQ ID NOS: 308 and 309, respectively; SEQ ID NOS: 310 and 311, respectively; SEQ ID NOS: 312 and 313, respectively; SEQ ID NOS: 314 and 315, respectively; SEQ ID NOS: 316 and 317, respectively; SEQ ID NOS: 318 and 319, respectively; SEQ ID NOS: 320 and 321, respectively; SEQ ID NOS: 322 and 323, respectively; SEQ ID NOS: 324 and 325, respectively; SEQ ID NOS: 326 and 327, respectively; SEQ ID NOS: 328 and 329, respectively; SEQ ID NOS: 330 and 331, respectively; SEQ ID NOS: 332 and 333, respectively; SEQ ID NOS: 334 and 335, respectively; SEQ ID NOS: 324 and 327, respectively; SEQ ID NOS: 326 and 325, respectively; SEQ ID NOS: 330 and 333, respectively; SEQ ID NOS: 332 and 331, respectively; SEQ ID NOS: 232 and 237, respectively; SEQ ID NOS: 233 and 238, respectively; SEQ ID NOS: 234 and 239, respectively; SEQ ID NOS: 235 and 240, respectively; or SEQ ID NOS: 236 and 241, respectively.
[0570] In some instances, the stems of the kissing loops are chosen to base pair in trans between the two RNA molecules. In such an example, after forming a kissing loop interaction of one hairpin loop on one molecule with another hairpin loop on a second molecule, the respective stem (or helix) regions of the initial hairpin loops can base pair in trans between the two RNA molecules through strand replacement / invasion and extended duplex formation. In some examples, within the initial loop sequence, up to about 85% of nucleotides can remain unpaired after extended duplex formation (e.g., about 15% of the nt are paired between the two loops). In some examples, the kissing loop is based on the HIV-1 DIS loop (SEQ ID NOS: 139 and 140, FIG. 17A), and includes two A nucleotides on the 5’ side of 6 nucleotides of complementary sequence, followed by one A nucleotide on the 3’ side (e.g., AANNNNNNA where N can be any of A, U, G, or C (SEQ ID NO: 378)). In some examples, the kissing loop is based on the HIV-2 kissing loopSLR 7158-102574-27 05 / 08 / 25 S2024-016A
[0571] dimerization domain (SEQ ID NOS: 141 and 142, FIG. 17B), and includes a G and an A nucleotide on the 5’ side of six nucleotides of complementary sequence followed by three A nucleotides on the 3’ side (e.g., GANNNNNNAAA (SEQ ID NO: 153) where N can be A, U, G, or C).
[0572] In one configuration, extended duplex formation is favored by inclusion of mismatches in the initial stems that result in higher percentage of matching in the extended duplex. Thus, in some examples, the helix or stem region of a hairpin loop can contain up to 30% of base pairs that are not paired initially (e.g., no more than 30%, no more than 20%, no more than 15%, no more than 10%, no more than 5%, or no more than 1%, such as 1 to 30%, 5 to 30%, 10 to 30%, or 25 to 30% of base pairs are not paired initially). These regions of non-pairing can form bulges, mismatches, or internal loops.
[0573] In addition to an interaction of two hairpin loops (kissing loop interaction), other forms of loop interactions can be utilized for the first and second dimerization domains 122, 154. In one example the loops are bulges, where one strand of a base paired helix contains one or more nucleotides that bulge out from the stem structure. Exemplary bulges are at least 1 nt, at least 2 nt, at least 3 nt, at least 4 nt, at least 5 nt, at least 10 nt or at least 20 nt, such as 1 to 20 nt, 1 to 15 nt, 1 to 10 nt, or 5 to 10 nt, or 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19 or 20 nt. In one example the loops are internal loops, for example, where 1 or more nucleotides in a helix are mismatched, resulting in a helix interrupted by an internal loop at the positions of mismatch. In some examples the helix is at least 4 nt on each of the strands (e.g., at least 5 nt, at least 10 nt, at least 20 nt, at least 30 nt, at least 40 nt, at least 50 nt, at least 75 nt, at least 90 nt, or at least 100 nt, such as 4 to 100 nt, 5 to 75 nt, or 10 to 50 nt. such as 4 to 100 nt), on either side of the internal loop that is at least 1 nt (e.g., at least 2 nt, at least 3 nt, at least 4 nt, at least 5 nt, at least 10 nt or at least 20 nt, such as 1 to 20 nt, 1 to 15 nt, 1 to 10 nt, or 5 to 10 nt, or 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19 or 20 nt on each of the strands). In one example the loops are multi-branched loops, wherein three helices or stems from a triangle with one or more unpaired nucleotides connecting the three helices. In some examples, each of the helices is at least 4 bp (e.g., at least 5 bp, at least 10 bp, at least 20 bp, at least 30 bp, at least 40 bp, at least 50 bp, at least 75 bp, at least 90bp, or at least 100 bp, such as 4 to 100 bp, 5 to 75 bp, or 10 to 50 bp), and the unpaired nucleotides that form the triangle are at least 3 nt (e.g., at least 4 nt, at least 5 nt, at least 10 nt, at least 20, at least 15, at least 30, at least 40, at least 50, or at least 60 nt, such as 3 to 60 nt, 3 to 30 nt, 3 to 25 nt, or 5 to 20 nt, such as 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 2, 25, 30, 35, 40, 45, 50, 55 or 60 nucleotides). A kissing interaction can occur between any two of these types of loops (e.g., between two or more binding domains that each include one or more helices). In some examples, helices within one dimerization domain (e.g., first dimerization domain 122) have a direct counterpart in the other binding domain (e.g., second dimerization domain 154) to allow for extended duplex formation after initial loop kissing interaction. In some examples, dimerization domains containing helices to generate loops, form a single kissing stem loop upon interaction between the two or more dimerization domains (e.g., 122, 154 of FIG. 6A). In some examples, dimerization domains containing helices form multiple loops for kissing loop interactions upon interaction between the two or more dimerization domains (e.g., 122, 154 of FIG. 6A). In some examples, one or more dimerization domains (e.g., 122 of FIG. 6A)SLR 7158-102574-27 05 / 08 / 25 S2024-016A
[0574] contain helices destabilized by the inclusion of bulges, single base bulges, mismatches or internal loops, or G-U wobble pairs, but match to the other binding domain (e.g., 154 of FIG. 6A), to favor extended duplex formation after initial kissing / pairing. In some examples, one or more dimerization domains (e.g., 122 of FIGS. 6A, 6G) contain destabilized helices, which when stabilized (e.g., theophylline switch kissing loop) expose a loop that can interact with a second dimerization domain (e.g., 122 of FIGS. 6A, 6G) via loop-loop interactions (e.g., kissing / pairing).
[0575] In some examples these stem loops contain at least 10 nt, such as at least 20 nt, at least 25 nt, at least 50 nt, at least 75 nt, or at least 100 nt in length, such as 10 to 50, 20 to 25, 10 to 100, 10 to 20, or 20 to 40 nt in length. Each dimerization domain can contain at least 1 individual stem loop, such as at least 2, at least 5, at least 10, at least 15, or at least 20, such as 1 to 20, 2 to 5 or 1 to 10 individual stem loops.
[0576] In some examples, 3 to 10 portions of a coding sequence are joined by 2 to 9 kissing loops, e.g., 3 portions are joined by 2 kissing loops, 4 portions are joined by 3 kissing loops, etc., wherein each of the 2 to 9 kissing loops are different. In some examples, a kissing loop comprises multiple stem loops, e.g., 2 to 20 stem loops. In some examples, each of the multiple stem loops in the kissing loop are the same. In some examples, each of the multiple stem loops in the kissing loop are different. In some examples, a dimerization domain comprises 1 to 20 stem loops. In some examples, a dimerization domain comprises 1 stem loop to 20 stem loops. In some examples, a dimerization domain comprises 1 stem loop to 2 stem loops, 1 stem loop to 3 stem loops, 1 stem loop to 4 stem loops, 1 stem loop to 5 stem loops, 1 stem loop to 6 stem loops, 1 stem loop to 7 stem loops, 1 stem loop to 8 stem loops, 1 stem loop to 9 stem loops, 1 stem loop to 10 stem loops, 1 stem loop to 15 stem loops, 1 stem loop to 20 stem loops, 2 stem loops to 3 stem loops, 2 stem loops to 4 stem loops, 2 stem loops to 5 stem loops, 2 stem loops to 6 stem loops, 2 stem loops to 7 stem loops, 2 stem loops to 8 stem loops, 2 stem loops to 9 stem loops, 2 stem loops to 10 stem loops, 2 stem loops to 15 stem loops, 2 stem loops to 20 stem loops, 3 stem loops to 4 stem loops, 3 stem loops to 5 stem loops, 3 stem loops to 6 stem loops, 3 stem loops to 7 stem loops, 3 stem loops to 8 stem loops, 3 stem loops to 9 stem loops, 3 stem loops to 10 stem loops, 3 stem loops to 15 stem loops, 3 stem loops to 20 stem loops, 4 stem loops to 5 stem loops, 4 stem loops to 6 stem loops, 4 stem loops to 7 stem loops, 4 stem loops to 8 stem loops, 4 stem loops to 9 stem loops, 4 stem loops to 10 stem loops, 4 stem loops to 15 stem loops, 4 stem loops to 20 stem loops, 5 stem loops to 6 stem loops, 5 stem loops to 7 stem loops, 5 stem loops to 8 stem loops, 5 stem loops to 9 stem loops, 5 stem loops to 10 stem loops, 5 stem loops to 15 stem loops, 5 stem loops to 20 stem loops, 6 stem loops to 7 stem loops, 6 stem loops to 8 stem loops, 6 stem loops to 9 stem loops, 6 stem loops to 10 stem loops, 6 stem loops to 15 stem loops, 6 stem loops to 20 stem loops, 7 stem loops to 8 stem loops, 7 stem loops to 9 stem loops, 7 stem loops to 10 stem loops, 7 stem loops to 15 stem loops, 7 stem loops to 20 stem loops, 8 stem loops to 9 stem loops, 8 stem loops to 10 stem loops, 8 stem loops to 15 stem loops, 8 stem loops to 20 stem loops, 9 stem loops to 10 stem loops, 9 stem loops to 15 stem loops, 9 stem loops to 20 stem loops, 10 stem loops to 15 stem loops, 10 stem loops to 20 stem loops, or 15 stem loops to 20 stem loops. In some examples, a dimerization domain comprises 1 stem loop, 2 stem loops, 3 stem loops, 4 stem loops, 5 stem loops, 6 stem loops, 7 stem loops, 8 stem loops, 9 stem loops,SLR 7158-102574-27 05 / 08 / 25 S2024-016A
[0577] 10 stem loops, 15 stem loops, or 20 stem loops. In some examples, a dimerization domain comprises at least 1 stem loop, 2 stem loops, 3 stem loops, 4 stem loops, 5 stem loops, 6 stem loops, 7 stem loops, 8 stem loops, 9 stem loops, 10 stem loops, or 15 stem loops. In some examples, a dimerization domain comprises at most 2 stem loops, 3 stem loops, 4 stem loops, 5 stem loops, 6 stem loops, 7 stem loops, 8 stem loops, 9 stem loops, 10 stem loops, 15 stem loops, or 20 stem loops.
[0578] Other mechanisms can be used to allow the two or more dimerization domains (e.g., 122, 154 of FIGS. 6 A, 6G) to bind or interact with one another sufficient for recombination of the coding sequences to occur. In some examples, the two or more dimerization domains (e.g., 122, 154 of FIGS. 6A, 6G) are nucleic acid aptamers (such as RNA aptamers) that can interact with one another, for example through a non-base pairing interaction, or can bind to a common molecule (e.g., protein, ATP, metal ion, co-factor, or synthetic ligand). In some examples, two or more dimerization domains (e.g. 122, 154 of FIGS. 6A, 6G) do not hybridize to one another, but can both (or all) hybridize to the same bridge nucleic acid molecule. In some examples, such a bridge nucleic acid molecule can be exogenously provided to the cells, tissues, or organism. In some examples, such a bridge nucleic acid molecule can be a DNA or RNA sequence inside the cell, such as a transcript or genomic locus. In some examples, the two or more dimerization domains (e.g., 122, 154 of FIGS. 6A, 6G) are sequences that can interact with one another, for example through a non-base pairing interaction.
[0579] Molecule 150 is the 3'-localed molecule, and includes a splice acceptor (SA) 162 and a second dimerization domain 154. In embodiments where molecule 150 is DNA, it includes a second promoter 152 followed by intronic sequence 170. Promoter 152 can be is operably linked to intronic sequence 170. Any promoter 152 can be used, such as a constitutive or inducible promoter. In some examples, promoter 152 is a tissue-specific promoter, such as one constitutively active in muscle tissue (such as skeletal or cardiac), optical tissue (such as retinal tissue), inner ear tissue, liver tissue, pancreatic tissue, lung tissue, skin tissue, bone, or kidney tissue. In some examples, promoter 112 is a cell-specific promoter, such as one constitutively active in a cancer cell, or a normal cell. In some examples, promoter 112 is an endogenous promoter of the target protein expressed, and in some examples is long (e.g., at least 2500 nt, at least 3000 nt, at least 4000 nt, at least 5000 nt, or at least 7500 nt). In some examples, promoter 112 is at least about 50 nucleotides (nt) in length, such as at least 100, at least 200, at least 300, at least 500, at least 1000, at least 2000, at least 3000, at least 4000, at least 5000, at least 6000, at least 7000, at least 8000 nt, at least 9000 nt, or at least 10,000 nt, such as 50 to 10,000 nt, 100 to 5000 nt, 500 to 5000 nt, or 50 to 1000 nt in length. In some examples promoter 112 and promoter 152 are the same promoter. In other examples, promoter 112 and promoter 152 are the different promoters. In some examples, molecule 150 is DNA, and is at least 200, at least 300, at least 500, at least 1000, at least 2000, at least 3000, at least 4000, at least 5000, at least 6000, at least 7000, or at least 8000 nt, such as 200 to 10,000 nt, 200 to 8000 nt, 500 to 5000 nt, or 200 to 1000 nt in length. As shown in FIG. 6F, in embodiments where molecule 150 is RNA, for example after expression of the DNA into RNA, molecule 150 no longer includes promoter 152, and 164 is the RNA encoded by the coding sequence for a C-terminal portion of the nucleic acid editing protein. In some examples, moleculeSLR 7158-102574-27 05 / 08 / 25 S2024-016A
[0580] 150 is RNA, does not include promoter 152, and is at least 200, at least 300, at least 500, at least 1000, at least 2000, at least 3000, at least 4000, at least 5000, at least 6000, at least 7000, or at least 8000 nt, such as 200 to 10,000 nt, 200 to 8000 nt, 500 to 5000 nt, or 200 to 1000 nt in length. Molecule 150 (with or without promoter 152) can include natural and / or non-natural nucleotides or ribonucleotides.
[0581] The intronic sequence 170 includes a second dimerization domain 154, optional ISE 156, branching point 158, polypyrimidine tract 160, followed by a splice acceptor sequence 162. In some examples, intronic sequence 130 is about at least 10 nt, such as at least 20 nt, at least 30 nt, at least 50 nt, at least 100 nt, at least 250 nt, at least 250 nt, at least 300 nt, at least 400 nt, or at least 500 nt in length, such as 20 to 500, 20 to 250, 20 to 100, 50 to 100, 30 to 500, or 50 to 200 nt in length.
[0582] Second dimerization domain 154 has a sequence that is the reverse complement of first dimerization domain 122 sequence of molecule 110. Thus, same design features and considerations of first dimerization domain 122 discussed above also apply to second dimerization domain 154. For example, in some examples the second dimerization domain 154 contains a stem loop that can form a kissing loop interaction the first dimerization domain 122. In some examples, second dimerization domain 154 does not include cryptic splice acceptors (e.g., NNNAGGUNNN; SEQ ID NO: 143) that could compete with RNA recombination. In some example, second dimerization domain 154 has a hypodiverse sequence. In some examples, second dimerization domain 154 is no more than 1000 nt, such as no more than 750 nt, or more than 500 nt, such as 30 to 1000 nt, 30 to 750 nt, 30 to 500 nt, 50 to 500 nt, 50 to 100 nt, or 100 to 250 nt. In some examples, second dimerization domain 154 is greater than 50 nt, such as at least 51 nt, at least 100 nt, at least 150 nt, at least 161 nt, or at least 170 nt, such as 51 to 159 nt, 51 to 150 nt, 51 to 120 nt, 51 to 100 nt, or 51 to 70 nt. In some examples, second dimerization domain 154 is greater than 160 nt, such as at least 161 nt, at least 170 nt, at least 180 nt, at least 200 nt, at least 300 nt, at least 400 nt, at least 500 nt, at least 600 nt, at least 700 nt, at least 800 nt, at least 900 nt, or at least 1000 nt, such as 161 to 1000 nt, 161 to 500 nt, 161 to 300 nt, 161 to 200 nt, or 161 to 170 nt. In some examples, second dimerization domain 154 is less than 50 nt, such 6 to 49 nt, 6 to 45 nt, 6 to 40 nt, 6 to 30 nt, 6 to 20 nt, or 6 to 10 nt.
[0583] 3’- to second dimerization domain 154 is an optional ISE 156, branch point sequence 158 (such as a branch point consensus sequence), polypyrimidine tract 160, followed by a splice acceptor sequence 162. ISE 156, like ISE 120 and DISE 118 of molecule 110, stimulates the spliceosome to catalyze the recombination reaction. In some examples, intronic sequence 150 includes at least two ISE 156, such as at least 3, at least 4, or at least 5 ISEs 156. Exemplary splicing enhancer sequences include ISE 156. In some examples, inclusion of one or more splicing enhancer sequences 156 in intronic sequence 150 increases recombination or splicing efficiency by at least 10%, at least 20%, at least 30%, at least 40%, or at least 50%. Exemplary splicing enhancer sequences that can be used are provided in SEQ ID NOS: 26-136, 151, and 152, as well as GGGTTT, GGTGGT, TTTGGG, GAGGGG, GGTATT, GTAACG, GGGGGTAGG, GGAGGGTTT, GGGTGGTGT TTCAT, CCATTT, TTTTAAA, TGCAT, TGCATG, TGTGTT, CTAAC, TCTCT, TCTGT, TCTTT, TGCATG, CTAAC, CTGCT, TAACC, AGCTT, TTCATTA, GTTAG, TTTTGC, ACTAAT, ATGTTT, CTCTG, GGG, GGG(N)2-4GGG, TGGG, YCAY, UGCAUG, or 3x(G3-SLR 7158-102574-27 05 / 08 / 25 S2024-016A
[0584] 6N1-7). In some examples, if ISE 156 is present, it can be about least 3 nt, at least 4 nt, at least 5 nt, at least 10 nt, such as at least 20 nt, at least 25 nt, at least 30 nt, at least 40 nt, or at least 50 nt in length, such as 3 to 10, 3 to 11, 4 to 11, 5 to 11, 10 to 50, 20 to 25, 10 to 25, 10 to 20, or 20 to 40 nt in length. In one example, the sequence of ISE 156 is or comprises GGCUGAGGGAAGGACUGUCCUGGG (SEQ ID NO: 135), GGGUUAUGGGACC (SEQ ID NO: 136), TTCAT, CCATTT, TTTTAAA, TGCAT, TGCATG, TGTGTT, CTAAC, TCTCT, TCTGT, or TCTTT. In some examples ISE 120 and ISE 156 are the same sequence. In other examples, ISE 120 and ISE 156 are the different sequences.
[0585] 3’- to second dimerization domain 154 (and ISE 156 if present) is branch point sequence 158 (such as a branch point consensus sequence), a polypyrimidine tract 160, followed by a splice acceptor sequence 162 (such as a splice acceptor consensus sequence). The sequence of branch point 158 is based on the consensus sequence of the species of the target cell or organism. For example, for human splicing, the consensus sequence can include or be YUNAY. Thus, a sequence that it uses can be CUAAC for Independent introns, or for U12-dependent introns UUUUCCUUAACU (SEQ ID NO: 144).
[0586] Polypyrimidine tract 160 includes C, U, or both C and U nucleotides, such as CnUy, wherein n+y is greater than or equal to 10 nucleotides, and can include nucleotides -3 to -22 relative to the 3’ -splice junction. In some examples, polypyrimidine tract 160 includes at least 80% Y nucleotides ( / .<?., U, C, or both U and C). In some examples, polypyrimidine tract 160 is a polyC or polyU sequence. In some examples, polypyrimidine tract 160 is a polyU sequence of at least 15 Us, such as 15 to 30 or 15 to 20 Us. Branch point 158 and polypyrimidine tract 160 are essential splicing components. The sequence of SA 162 can be based on the consensus sequence of the species of the target cell or organism. For example, in humans, the SA sequence can be AG in positions -1 and -2 relative to the 3’ -splice site for U2-dependnet introns and AC or AG for U12-dependnet introns. Thus, in some examples, SA 162 can be 2 nt in length, such as AG or AC.
[0587] Immediately following SA 162 is an exonic sequence which includes a DNA sequence encoding a C-terminal portion of a target protein 164 having a splice junction at its 5’end. The splice junction at the 5’end of DNA sequence encoding a C-terminal portion of a nucleic acid editing protein 164, that can match the consensus sequence found in the target cell or organism into which molecules 110, 150 are introduced. In some examples splice junction can be GA or GU at positon +1 and +2 of the 3’ splice site for U2-dependent introns or GU or AU for U12-dependent introns. Thus, in some examples, the splice junction is 2 nt in length, and the 5’ end of the C-terminal coding portion 164 is GA, GU, or AU.
[0588] The exonic sequence following intro nic portion 170 of molecule 150 includes a second coding portion (e.g., half) of the nucleic acid editing protein, e.g., the C terminal fragment 164, and optional polyadenylation sequence 166. Thus, molecule 150 includes sequence 164 encoding a C-terminal portion of a nucleic acid editing protein. The 3’ -end of molecule 150 optionally includes a polyadenylation sequence 166, which promotes the assembly of the spliceosome. In some examples, polyadenylation sequence 166 is a polyA sequence of at least 15 As, such as 15 to 30 or 15 to 20 As. In some examples polyadenylationSLR 7158-102574-27 05 / 08 / 25 S2024-016A
[0589] sequence 166 and polyadenylation sequence 124 are the same sequence. In other examples, polyadenylation sequence 166 and polyadenylation sequence 124 are the different sequences.
[0590] In some examples, the N-terminal coding region 114 and / or the C terminal coding region 164 is a native coding sequence. For example, the coding sequence is one that is found in the cell or organism into which the disclosed system is introduced, (e.g., a human coding sequence when introduced into a human cell or subject). In some examples, the N-terminal coding region 114 and / or the C terminal coding region 164 is codon optimized relative to a native coding sequence, for example to maximize tRNA availability, or to deenrich for cryptic splice sites e.g., to reduce or avoid incorrect splicing and promote the correct junction formation). In some examples, a portion of the N-terminal coding region 114 and / or the C terminal coding region 164 is codon optimized relative to a native coding sequence, for example the about 200 nt adjacent to each junction (e.g., the 3’-end of 114, and the 5’end of 164) can be codon optimized or altered to contain exonic splice enhancer sites (ESE) (which would bind SR proteins). For example, the coding sequence can be one not found in the cell or organism into which the disclosed system is introduced (e.g., a human coding sequence when introduced into a mouse cell or subject).
[0591] In some examples, the N-terminal coding region 114 and / or the C terminal coding region 164 include an intron that is either natural or synthetic in nature and contains both a splice donor and acceptor site. For example, an intron embedded inside the coding sequence to he expressed can be included upstream (e.g., about 130 nt to about 200 nt upstream, such as about 140 nt, 150 nt, 160 nt, 170 nt, 180 nt, or 190 nt upstream) of sequence 116, inside the N-terminal coding region 114, and / or an intron embedded inside the coding sequence to be expressed can be included downstream (e.g., about 130 nt to about 200 nt downstream, such as about 140 nt, 150 nt, 160 nt, 170 nt, 180 nt, or 190 nt downstream) of the sequence 162 and inside the C-terminal coding region 164. In some examples, the embedded intron can be located about 130 nt to about 200 nt (such as about 140 nt, 150 nt, 160 nt, 170 nt, 180 nt, or 190 nt) upstream the end of the coding sequence (e.g., N-terminal coding sequence), or about 130 nt to about 200 nt (such as about 140 nt, 150 nt, 160 nt, 170 nt, 180 nt, or 190 nt) downstream the start of the coding sequence (e.g., C-terminal coding sequence). Inclusion of such introns can be used to stimulate splicing machinery attachment to the trans- splicing intron donor and acceptor. In some examples, such (stimulatory-) introns could be derived from the host in which 110 and 150 are expressed. In some examples, such (stimulatory-) introns could be derived from other organisms, or viral in origin, or synthetic in origin. In some examples, the stimulatory intron includes a sequence that is at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% id...
Claims
SLR 7158-102574-27 05 / 08 / 25 S2024-016AWe claim:
1. A composition for expressing a target protein comprising(a) a first RNA molecule, the RNA molecule comprising from 5’ to 3’: (i) a coding sequence for an N-terminal portion of the target protein; (ii) a splice donor; and (iii) a first dimerization domain;(b) a second RNA molecule, the RNA molecule comprising from 5’ to 3’: (i) a second dimerization domain, wherein the second dimerization domain binds to the first dimerization domain; (ii) a branch point sequence; (iii) a polypyrimidine tract; (iv) a splice acceptor; and (v) a coding sequence for a C-terminal portion of the target protein; and(c) one or more secondary structures in the first RNA molecule, one or more secondary structures in the second RNA molecule, or both, wherein the one or more secondary structures are not in the coding sequence for the N-terminal portion of the target protein nor in the coding sequence for the C-terminal portion of the target protein; orone or more elements to recruit nuclear RNA binding proteins in the first RNA molecule, one or more elements to recruit nuclear RNA binding proteins in the second RNA molecule, or both, wherein the one or more elements are not in the coding sequence for the N-terminal portion of the target protein nor in the coding sequence for the C-terminal portion of the target protein.
2. A composition for expressing a nucleic acid editing protein, comprising:(a) a first RNA molecule comprising from 5’ to 3’:(i) a coding sequence for an N-terminal portion of the nucleic acid editing protein;(ii) a splice donor; and(iii) a first dimerization domain;(b) a second RNA molecule comprising from 5’ to 3’:(i) a second dimerization domain, wherein the second dimerization domain binds to the first dimerization domain;(ii) a branch point sequence;(iii) a polypyrimidine tract;(iv) a splice acceptor; and(v) a coding sequence for a C-terminal portion of the nucleic acid editing protein; and (c) one or more secondary structures in the first RNA molecule, one or more secondary structures in the second RNA molecule, or both, wherein the one or more secondary structures are not in the coding sequence for the N-terminal portion of the nucleic acid editing protein nor in the coding sequence for the C-terminal portion of the nucleic acid editing protein; orone or more elements to recruit nuclear RNA binding proteins in the first RNA molecule, one or more elements to recruit nuclear RNA binding proteins in the second RNA molecule, or both, wherein the one or more elements are not in the coding sequence for the N-terminal portion of the nucleic acid editing protein nor in the coding sequence for the C-terminal portion of the nucleic acid editing protein;SLR 7158-102574-27 05 / 08 / 25 S2024-016A(d) optionally, a third RNA molecule comprising at least one first guide RNA (gRNA) specific for a first target nucleic acid molecule, wherein the at least one first gRNA directs the nucleic acid editing protein to a target editing site on the first target nucleic acid molecule;(e) optionally, a fourth RNA molecule comprising at least one second gRNA specific for (i) the first target nucleic acid molecule, wherein the at least one second gRNA directs the nucleic acid editing protein to the same or different target editing site on the first nucleic acid molecule, or (ii) a second target nucleic acid molecule, wherein the at least one second gRNA directs the nucleic acid editing protein to a target editing site on the second target nucleic acid molecule;(f) optionally, a fifth RNA molecule comprising at least one third gRNA specific for (i) the first target nucleic acid molecule, wherein the at least one third gRNA directs the nucleic acid editing protein to the same or different target editing site on the first nucleic acid molecule as the first and second gRNA, (ii) the second target nucleic acid molecule, wherein the at least one third gRNA directs the nucleic acid editing protein to the same or different target editing site on the second nucleic acid molecule as the second gRNA, or (iii) a third target nucleic acid molecule, wherein the at least one third gRNA directs the nucleic acid editing protein to a target editing site on the third target nucleic acid molecule; and(g) optionally, a sixth RNA molecule comprising at least one fourth gRNA specific for (i) the first target nucleic acid molecule, wherein the at least one fourth gRNA directs the nucleic acid editing protein to the same or different target editing site on the first nucleic acid molecule as the first, second, and third gRNA, (ii) the second target nucleic acid molecule, wherein the at least one fourth gRNA directs the nucleic acid editing protein to the same or different target editing site on the second nucleic acid molecule as the second and third gRNA, (iii) the third target nucleic acid molecule, wherein the at least one fourth gRNA directs the nucleic acid editing protein to the same or different target editing site on the third target nucleic acid molecule as the third gRNA, or (iv) a fourth target nucleic acid molecule, wherein the at least one fourth gRNA directs the nucleic acid editing protein to a target editing site on the fourth target nucleic acid molecule.
3. The composition of claim 1 or 2, wherein the first and second dimerization domains bind by direct binding, indirect binding, or a combination thereof.
4. The composition of claim 3, wherein direct binding or indirect binding comprises base pairing interactions, non-canonical base pairing interactions, non-base pairing interactions, or a combination thereof.
5. The composition of claim 3 or 4, wherein direct binding comprises base pairing interactions between kissing loops or hypodiverse regions.
6. The composition of claim 3or 4, wherein direct binding comprises non-canonical base pairing interactions, non-canonical base pairing interactions, non-base pairing interactions, or a combination thereof, between aptamer regions.
7. The composition of claim 3 or 4, wherein indirect binding comprises base pairing interactions through a nucleic acid bridge.SLR 7158-102574-27 05 / 08 / 25 S2024-016A8. The composition of claim 3, wherein indirect binding comprises non-base pairing interactions between an aptamer and an aptamer target, or between two aptamers.
9. The composition of any one of claims 1 to 8, wherein the first or second dimerization domain does not comprise a cryptic splice acceptor or cryptic splice donor.
10. The composition of any one of claims 1 to 9, wherein the dimerization domains are directly binding or indirectly binding aptamer sequence dimerization domains.
11. The composition of any one of claims 1 to 10, wherein the dimerization domains are kissing loop interaction domains.
12. The composition of any one of claims 2 to 11, wherein the target editing site is part of the target nucleic acid molecule or a regulatory region of a target nucleic acid molecule associated with disease.
13. The composition of claim 12, wherein the disease is a monogenic disease.
14. The composition of claim 12 or 13, wherein the first, second, third, and / or fourth target nucleic acid molecule comprises one or more point mutations that results in the disease.
15. The composition of any one of claims 12 to 14, wherein the disease and the first, second, third, and / or fourth target nucleic acid molecule are one listed in Tables 1-4; or wherein the disease is Usher Syndrome IF and the first, second, third, and / or fourth target nucleic acid molecule is PCDH15 gene or gene product, or wherein the disease is Duchenne muscular dystrophy (DMD) and the first, second, third, and / or fourth target nucleic acid molecule is DMD gene or gene product.
16. The composition of any one of claims 1 to 15, wherein the first RNA molecule further comprises one or both of a downstream intronic splice enhancer (DISE) 3’ to the splice donor and 5’ to the first dimerization domain, an intronic splice enhancer (ISE) 3’ to the splice donor and 5’ to the first dimerization domain; and / orthe second RNA molecule further comprises one or both of an ISE 3’ to the second dimerization domain and 5’ to the branch point sequence, and a DISE 3’ to the splice donor and 5’ to the dimerization domain;or any combination thereof.
17. The composition of any one of claims 1 to 16, whereinthe first RNA molecule further comprises a self-cleaving RNA sequence or an RNA-cleaving enzyme target sequence positioned anywhere 3’ to the splice donor such that it cleaves off the 3’ located polyadenylated tail to decrease or suppress protein fragment expression from a non-recombined RNA molecule;the second RNA molecule further comprises a self-cleaving RNA sequence or an RNA-cleaving enzyme target sequence positioned anywhere 5’ to the branch point sequence such that it cleaves off the 5’ located RNA cap to decrease or suppress protein fragment expression from a non-recombined RNA molecule;SLR 7158-102574-27 05 / 08 / 25 S2024-016Athe second RNA molecule further comprises a start codon anywhere 5’ to the branch point sequence that is shifted relative to the open reading frame 3’ of the splice acceptor to decrease or suppress translation of a nucleic acid editing protein fragment from a non-recombined RNA molecule;the first RNA molecule further comprises a micro RNA target site anywhere 3’ to the splice donor such that an un-joined RNA fragment undergoes micro RNA dependent degradation once outside the nucleus;the second RNA molecule further comprises a micro RNA target site anywhere 3’ to the coding sequence such that an un-joined RNA fragment undergoes micro RNA dependent degradation once outside the nucleus;the first RNA molecule further comprises a sequence encoding a degron protein degradation tag anywhere 3’ to the splice donor such that it is in frame with the nucleic acid editing protein open reading frame 5’ of the splice donor site such that an un-joined protein fragment is tagged for degradation;the second RNA molecule further comprises a start codon and an in-frame degron protein degradation tag anywhere 5’ to the branch point sequence such that it is in frame with the nucleic acid editing protein open reading frame 3’ of the splice acceptor site such that an un-joined protein fragment is tagged for degradation;or any combination thereof.
18. The composition of any one of claims 2-17, wherein the nucleic acid editing protein comprises a Cas nuclease, zinc finger nuclease, or transcription activator-like effector nuclease.
19. The composition of any one of claims 1-18, wherein the first, second, third, and / or fourth target nucleic acid molecules are target DNA molecules, and the at least one first, second, third, and / or fourth gRNA comprises a crRNA and tracrRNA20. The composition of any one of claims 2-19, wherein the nucleic acid editing protein comprises Cas9 or dead Cas9 (dCas9).
21. The composition of claim 20, whereinthe Cas9 protein comprises at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity to SEQ ID NO: 208, and can function as an RNA-guided DNA endonuclease;the Cas9 protein is encoded by a sequence comprising at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity to SEQ ID NO: 207, and encodes an RNA-guided DNA endonuclease;the dCas9 protein comprises at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity to SEQ ID NO: 210, and is catalytically inactive;the dCas9 protein is encoded by a sequence comprising at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity to SEQ ID NO: 209, and encodes a protein that is catalytically inactive.SLR 7158-102574-27 05 / 08 / 25 S2024-016A22. The composition of claim 20 or 21, wherein the Cas9 or dCas9 is part of a fusion protein.
23. The composition of claim 22, wherein the fusion protein comprises Cas9 or dCas9 and one or more of:a transcriptional activation domain, such as VP64, P65, MyoDl, HSF1, RTA, CBP, SET7 / 9, or any combination thereof;a cytosine base editor (CBE), such as one from sea lamprey [AID], CDA1, or APOBEC3G; bacteriophage protein Gam; andan adenine base editor (ABE), such as ABE 6.3, 7.8, 7.9, 7.10, ABEmax, ABE8s, such as ABE8e(TadA-8e V106W).
24. The composition of any one of claims 2-23, wherein the first, second, third, and / or fourth target nucleic acid molecule are target RNA, and the at least one first, second, third, and / or fourth gRNA comprises one or more direct repeats and one or more spacers.
25. The composition of any one of claims 2-24, wherein the nucleic acid editing protein comprises Cas13a, Cas13b, Cas13c, Cas13d or dead Cas13d (dCas13d).
26. The composition of claim 25, whereinthe Cas13d protein comprises at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity to SEQ ID NO: 212, 215, or 222, and can function as an RNA-guided RNA endonuclease;the Cas13d protein is encoded by a sequence comprising at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity to SEQ ID NO: 211 or 213, and encodes an RNA-guided RNA endonuclease;the dCasl3d protein comprises at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity to SEQ ID NO: 215 or 216, and is catalytically inacti ve;the dCas13d protein is encoded by a sequence that encodes a protein comprising at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity to SEQ ID NO: 215 or 216, and encodes a protein that is catalytically inactive.
27. The composition of claim 25 or 26, wherein the Cas13a, Cas13b, Cas13c, Cas13d or dCas13d is part of a fusion protein.
28. The composition of claim 27, wherein the fusion protein comprises Cas13a, Cas13b, Cas13c, Cas13d or dCas13d and one or more of:a transcriptional activation domain, such as VP64, P65, MyoDl, HSF1, RTA, CBP, SET7 / 9, or any combination thereof;a cytosine base editor (CBE), such as one from sea lamprey [AID], CDA1, or APOBEC3G; bacteriophage protein Gam; andan adenine base editor (ABE), such as ABE 6.3, 7.8, 7.9, 7.10, ABEmax, ABE8s, such as ABE8e(TadA-8e V106W).SLR 7158-102574-27 05 / 08 / 25 S2024-016A29. The composition of any one of claims 2-28, wherein the at least one first, second, third, and / or fourth gRNA comprise multiple copies of each of the first, second, third, and / or fourth gRNA.
30. The composition of any one of claims 1-28, further comprising a parvovirus inverted terminal repeat (ITR) at the 5’ and the 3’ -end of each of the first RNA molecule and the second RNA molecule.
31. The composition of any one of claims 1-30, wherein the one or more secondary structures comprise one or more hairpins, kissing stem loops, quadruplexes, or combinations thereof.
32. The composition of any one of claims 1-31, wherein the one or more elements comprise one or more MALAT1 Region M sequences (SEQ ID NO: 246), one or more SIROLIN elements (SEQ ID NO: 227). or combinations thereof.
33. The composition of any one of claims 1-32, wherein:the target protein comprises two or more splice variants of the N-terminal portion of the target protein, and the first RNA molecule comprises two or more separate RNA molecules, wherein two or more separate RNA molecules each encode a distinct form of the N-terminal portion of the target protein;the target protein comprises two or more splice variants of the C-terminal portion of the target protein, and the second RNA molecule comprises two or more separate RNA molecules, wherein two or more separate RNA molecules each encode a distinct form of the C-terminal portion of the target protein; or combinations thereof.
34. The composition of any one of claims 1 and 3-33, wherein the target protein is PCDH15, ABCA4, dystrophin, truncated dystrophin, or Myo7A.
35. A DNA molecule composition for expressing the composition of any one of claims 1-34, comprising:(i) a first synthetic DNA molecule encoding the RNA molecule of (a), and optionally (d), (e), (f), and / or (g); and(ii) a second synthetic DNA molecule encoding the RNA molecule of (b), and optionally and (d), (e), (f), and / or (g).
36. The DNA molecule of claim 35, further comprising:a first promoter operably linked to a DNA sequence encoding the N-terminal portion of the nucleic acid editing protein;a second promoter operably linked to a DNA sequence encoding the C-terminal portion of the nucleic acid editing protein;if the first or second synthetic DNA molecule includes a DNA molecule encoding the RNA molecule of (d), a third promoter operably linked to a sequence encoding the at least one first gRNA;if the first or second synthetic DNA molecule includes a DNA molecule encoding the RNA molecule of (e), a fourth promoter operably linked to a sequence encoding the at least one second gRNA; if the first or second synthetic DNA molecule includes a DNA molecule encoding the RNA molecule of (f), a fifth promoter operably linked to a sequence encoding the at least one third gRNA; andSLR 7158-102574-27 05 / 08 / 25 S2024-016A(if the first or second synthetic DNA molecule includes a DNA molecule encoding the RNA molecule of (g), a sixth promoter operably linked to a sequence encoding the at least one fourth gRNA.
37. The composition of claim 36, wherein each promoter is independently selected.
38. The composition of claim 36 or 37, wherein:the first and second promoter are the same promoter;the first and second promoter are different promoters,the third, fourth, fifth and sixth promoters are the same promoter;the third, fourth, fifth and sixth promoters are different promoters, orcombinations thereof.
39. The composition of any one of claims 35 to 37, wherein each of the first and second promoters is independently selected from: a constitutive promoter; a tissue-specific promoter, such as an ocularspecific promoter, such as a GRK1 promoter; and a promoter endogenous to the nucleic acid editing protein.
40. The composition of any one of claims 35 to 39, wherein each of the third, fourth, fifth and sixth promoters is a polymerase III promoter, such as a U6 or Hl promoter.
41. A system for expressing a nucleic acid editing protein comprising a composition of any one of claims 35 to 40.
42. The system of claim 41, wherein when the system is introduced into a cell the RNA molecules are produced and recombine in the proper order, resulting in a full-length coding sequence of the nucleic acid editing protein.
43. The system of claim 41 or 42, wherein each of the synthetic first and second RNA molecules are transcribed from a separate viral vector.
44. The system of claim 43, wherein the viral vector is AAV.
45. The system of any one of claims 41 to 44, wherein each of the synthetic DNA molecules has a size independently selected from: about 2500 nt to about 5000 nt, 2,500 nt to about 2,750 nt, about 2,500 nt to about 3,000 nt, about 2,500 nt to about 3,250 nt, about 2,500 nt to about 3,500 nt, about 2,500 nt to about 3.750 nt, about 2,500 nt to about 4,000 nt, about 2,500 nt to about 4,250 nt, about 2,500 nt to about 4,500 nt, about 2,500 nt to about 4,750 nt, about 2,500 nt to about 5,000 nt, about 2,750 nt to about 3,000 nt, about 2.750 nt to about 3,250 nt, about 2,750 nt to about 3,500 nt, about 2,750 nt to about 3,750 nt, about 2,750 nt to about 4,000 nt, about 2,750 nt to about 4,250 nt, about 2,750 nt to about 4,500 nt, about 2,750 nt to about 4.750 nt, about 2,750 nt to about 5,000 nt, about 3,000 nt to about 3,250 nt, about 3,000 nt to about 3,500 nt, about 3,000 nt to about 3,750 nt, about 3,000 nt to about 4,000 nt, about 3,000 nt to about 4,250 nt, about 3,000 nt to about 4,500 nt, about 3,000 nt to about 4,750 nt, about 3,000 nt to about 5,000 nt, about 3,250 nt to about 3,500 nt, about 3,250 nt to about 3,750 nt, about 3,250 nt to about 4,000 nt, about 3,250 nt to about 4,250 nt, about 3,250 nt to about 4,500 nt, about 3,250 nt to about 4,750 nt, about 3,250 nt to about 5,000 nt, about 3,500 nt to about 3,750 nt, about 3,500 nt to about 4,000 nt, about 3,500 nt to about 4,250 nt, about 3,500 nt to about 4,500 nt, about 3,500 nt to about 4,750 nt, about 3,500 nt to about 5,000 nt, about 3,750 nt to about 4,000 nt, about 3,750 nt to about 4,250 nt, about 3,750 nt to about 4,500 nt, about 3,750 nt to aboutSLR 7158-102574-27 05 / 08 / 25 S2024-016A4.750 nt, about 3,750 nt to about 5,000 nt, about 4,000 nt to about 4,250 nt, about 4,000 nt to about 4,500 nt, about 4,000 nt to about 4,750 nt, about 4,000 nt to about 5,000 nt, about 4,250 nt to about 4,500 nt, about 4.250 nt to about 4,750 nt, about 4,250 nt to about 5,000 nt, about 4,500 nt to about 4,750 nt, about 4,500 nt to about 5,000 nt, about 4,750 nt to about 5,000 nt, about 2,500 nt, about 2,750 nt, about 3,000 nt, about 3.250 nt, about 3,500 nt, about 3,750 nt, about 4,000 nt, about 4,250 nt, about 4,500 nt, about 4,750 nt, and about 5,000 nt.
46. The system of any one of claims 41 to 45, wherein the coding sequence for an N-terminal portion of the nucleic acid editing protein or the target protein, or a C-terminal portion of the nucleic acid editing protein or the target protein encoded by a synthetic DNA molecule of the system each has a size independently selected from: about 2500 to 4500 nt, about 2,500 nt to about 2,750 nt, about 2,500 nt to about 3,000 nt, about 2,500 nt to about 3,250 nt, about 2,500 nt to about 3,500 nt, about 2,500 nt to about 3.750 nt, about 2,500 nt to about 4,000 nt, about 2,500 nt to about 4,250 nt, about 2,500 nt to about 4,500 nt, about 2,750 nt to about 3,000 nt, about 2,750 nt to about 3,250 nt, about 2,750 nt to about 3,500 nt, about 2.750 nt to about 3,750 nt, about 2,750 nt to about 4,000 nt, about 2,750 nt to about 4,250 nt, about 2,750 nt to about 4,500 nt, about 3,000 nt to about 3,250 nt, about 3,000 nt to about 3,500 nt, about 3,000 nt to about 3.750 nt, about 3,000 nt to about 4,000 nt, about 3,000 nt to about 4,250 nt, about 3,000 nt to about 4,500 nt, about 3,250 nt to about 3,500 nt, about 3,250 nt to about 3,750 nt, about 3,250 nt to about 4,000 nt, about 3,250 nt to about 4,250 nt, about 3,250 nt to about 4,500 nt, about 3,500 nt to about 3,750 nt, about 3,500 nt to about 4,000 nt, about 3,500 nt to about 4,250 nt, about 3,500 nt to about 4,500 nt, about 3,750 nt to about 4,000 nt, about 3,750 nt to about 4,250 nt, about 3,750 nt to about 4,500 nt, about 4,000 nt to about 4,250 nt, about 4,000 nt to about 4,500 nt, about 4,250 nt to about 4,500 nt, about 2,500 nt, about 2,750 nt, about 3,000 nt, about 3,250 nt, about 3,500 nt, about 3,750 nt, about 4,000 nt, about 4,250 nt, and about 4,500 nt.
47. The system of any one of claims 41 to 46, wherein any one or both of the RNA molecules encoded by the synthetic DNA molecules of the system, respectively, has a size independently selected from: about 2500 to 4500 nt, about 2,500 nt to about 2,750 nt, about 2,500 nt to about 3,000 nt, about 2,500 nt to about 3,250 nt, about 2,500 nt to about 3,500 nt, about 2,500 nt to about 3,750 nt, about 2,500 nt to about 4,000 nt, about 2,500 nt to about 4,250 nt, about 2,500 nt to about 4,500 nt, about 2,750 nt to about 3,000 nt, about 2,750 nt to about 3,250 nt, about 2,750 nt to about 3,500 nt, about 2,750 nt to about 3,750 nt, about 2,750 nt to about 4,000 nt, about 2,750 nt to about 4,250 nt, about 2,750 nt to about 4,500 nt, about 3,000 nt to about 3,250 nt, about 3,000 nt to about 3,500 nt, about 3,000 nt to about 3,750 nt, about 3,000 nt to about 4,000 nt, about 3,000 nt to about 4,250 nt, about 3,000 nt to about 4,500 nt, about 3,250 nt to about 3,500 nt, about 3,250 nt to about 3,750 nt, about 3,250 nt to about 4,000 nt, about 3,250 nt to about 4,250 nt, about 3,250 nt to about 4,500 nt, about 3,500 nt to about 3,750 nt, about 3,500 nt to about 4,000 nt, about 3.500 nt to about 4,250 nt, about 3,500 nt to about 4,500 nt, about 3,750 nt to about 4,000 nt, about 3,750 nt to about 4,250 nt, about 3,750 nt to about 4,500 nt, about 4,000 nt to about 4,250 nt, about 4,000 nt to about 4.500 nt, about 4,250 nt to about 4,500 nt, about 2,500 nt, about 2,750 nt, about 3,000 nt, about 3,250 nt, about 3,500 nt, about 3,750 nt, about 4,000 nt, about 4,250 nt, and about 4,500 nt.SLR 7158-102574-27 05 / 08 / 25 S2024-016A48. The system of any one of claims 41 to 47, wherein the system comprises a composition of any one of claims 35 to 40:the synthetic DNA molecules have a total size selected from about 5000 nt to about 10,000 nt, about 5,000 nt to about 5,500 nt, about 5,000 nt to about 6,000 nt, about 5,000 nt to about 6,500 nt, about 5,000 nt to about 7,000 nt, about 5,000 nt to about 7,500 nt, about 5,000 nt to about 8,000 nt, about 5,000 nt to about 8.500 nt, about 5,000 nt to about 9,000 nt, about 5,000 nt to about 9,500 nt, about 5,000 nt to about 10,000 nt, about 5,500 nt to about 6,000 nt, about 5,500 nt to about 6,500 nt, about 5,500 nt to about 7,000 nt, about 5.500 nt to about 7,500 nt, about 5,500 nt to about 8,000 nt, about 5,500 nt to about 8,500 nt, about 5,500 nt to about 9,000 nt, about 5,500 nt to about 9,500 nt, about 5,500 nt to about 10,000 nt, about 6,000 nt to about 6.500 nt, about 6,000 nt to about 7,000 nt, about 6,000 nt to about 7,500 nt, about 6,000 nt to about 8,000 nt, about 6,000 nt to about 8,500 nt, about 6,000 nt to about 9,000 nt, about 6,000 nt to about 9,500 nt, about 6,000 nt to about 10,000 nt, about 6,500 nt to about 7,000 nt, about 6,500 nt to about 7,500 nt, about 6,500 nt to about 8,000 nt, about 6,500 nt to about 8,500 nt, about 6,500 nt to about 9,000 nt, about 6,500 nt to about 9.500 nt, about 6,500 nt to about 10,000 nt, about 7,000 nt to about 7,500 nt, about 7,000 nt to about 8,000 nt, about 7,000 nt to about 8,500 nt, about 7,000 nt to about 9,000 nt, about 7,000 nt to about 9,500 nt, about 7,000 nt to about 10,000 nt, about 7,500 nt to about 8,000 nt, about 7,500 nt to about 8,500 nt, about 7,500 nt to about 9,000 nt, about 7,500 nt to about 9,500 nt, about 7,500 nt to about 10,000 nt, about 8,000 nt to about 8.500 nt, about 8,000 nt to about 9,000 nt, about 8,000 nt to about 9,500 nt, about 8,000 nt to about 10,000 nt, about 8,500 nt to about 9,000 nt, about 8,500 nt to about 9,500 nt, about 8,500 nt to about 10,000 nt, about 9,000 nt to about 9,500 nt, about 9,000 nt to about 10,000 nt, about 9,500 nt to about 10,000 nt, about 5,000 nt, about 5,500 nt, about 6,000 nt, about 6,500 nt, about 7,000 nt, about 7,500 nt, about 8,000 nt, about 8.500 nt, about 9,000 nt, about 9,500 nt, and about 10,000 nt;the total nucleic acid editing protein or the target protein coding sequence is selected from about 2000 nt to about 8000 nt, about 2,000 nt to about 3,000 nt, about 2,000 nt to about 3,500 nt, about 2,000 nt to about 4,000 nt, about 2,000 nt to about 4,500 nt, about 2,000 nt to about 5,000 nt, about 2,000 nt to about 5.500 nt, about 2,000 nt to about 6,000 nt, about 2,000 nt to about 6,500 nt, about 2,000 nt to about 7,000 nt, about 2,000 nt to about 7,500 nt, about 2,000 nt to about 8,000 nt, about 3,000 nt to about 3,500 nt, about 3,000 nt to about 4,000 nt, about 3,000 nt to about 4,500 nt, about 3,000 nt to about 5,000 nt, about 3,000 nt to about 5,500 nt, about 3,000 nt to about 6,000 nt, about 3,000 nt to about 6,500 nt, about 3,000 nt to about 7,000 nt, about 3,000 nt to about 7,500 nt, about 3,000 nt to about 8,000 nt, about 3,500 nt to about 4,000 nt, about 3,500 nt to about 4,500 nt, about 3,500 nt to about 5,000 nt, about 3,500 nt to about 5,500 nt, about 3.500 nt to about 6,000 nt, about 3,500 nt to about 6,500 nt, about 3,500 nt to about 7,000 nt, about 3,500 nt to about 7,500 nt, about 3,500 nt to about 8,000 nt, about 4,000 nt to about 4,500 nt, about 4,000 nt to about 5,000 nt, about 4,000 nt to about 5,500 nt, about 4,000 nt to about 6,000 nt, about 4,000 nt to about 6,500 nt, about 4,000 nt to about 7,000 nt, about 4,000 nt to about 7,500 nt, about 4,000 nt to about 8,000 nt, about 4.500 nt to about 5,000 nt, about 4,500 nt to about 5,500 nt, about 4,500 nt to about 6,000 nt, about 4,500 nt to about 6,500 nt, about 4,500 nt to about 7,000 nt, about 4,500 nt to about 7,500 nt, about 4,500 nt to aboutSLR 7158-102574-27 05 / 08 / 25 S2024-016A8,000 nt, about 5,000 nt to about 5,500 nt, about 5,000 nt to about 6,000 nt, about 5,000 nt to about 6,500 nt, about 5,000 nt to about 7,000 nt, about 5,000 nt to about 7,500 nt, about 5,000 nt to about 8,000 nt, about 5.500 nt to about 6,000 nt, about 5,500 nt to about 6,500 nt, about 5,500 nt to about 7,000 nt, about 5,500 nt to about 7,500 nt, about 5,500 nt to about 8,000 nt, about 6,000 nt to about 6,500 nt, about 6,000 nt to about 7,000 nt, about 6,000 nt to about 7,500 nt, about 6,000 nt to about 8,000 nt, about 6,500 nt to about 7,000 nt, about 6,500 nt to about 7,500 nt, about 6,500 nt to about 8,000 nt, about 7,000 nt to about 7,500 nt, about 7,000 nt to about 8,000 nt, or about 7,500 nt to about 8,000 nt. the total target protein coding sequence is about 2,000 nt, about 3,000 nt, about 3,500 nt, about 4,000 nt, about 4,500 nt, about 5,000 nt, about 5,500 nt, about 6,000 nt, about 6,500 nt, about 7,000 nt, about 7,500 nt, and about 8,000 nt; and / orthe summed size of the RNA molecules encoded by the two synthetic DNA molecules is selected from about 5,000 nt to about 9000 nt, about 5,000 nt to about 5,500 nt, about 5,000 nt to about 6,000 nt, about 5,000 nt to about 6,500 nt, about 5,000 nt to about 7,000 nt, about 5,000 nt to about 7,500 nt, about 5,000 nt to about 8,000 nt, about 5,000 nt to about 8,500 nt, about 5,000 nt to about 9,000 nt, about 5,500 nt to about 6,000 nt, about 5,500 nt to about 6,500 nt, about 5,500 nt to about 7,000 nt, about 5,500 nt to about 7.500 nt, about 5,500 nt to about 8,000 nt, about 5,500 nt to about 8,500 nt, about 5,500 nt to about 9,000 nt, about 6,000 nt to about 6,500 nt, about 6,000 nt to about 7,000 nt, about 6,000 nt to about 7,500 nt, about 6,000 nt to about 8,000 nt, about 6,000 nt to about 8,500 nt, about 6,000 nt to about 9,000 nt, about 6,500 nt to about 7,000 nt, about 6,500 nt to about 7,500 nt, about 6,500 nt to about 8,000 nt, about 6,500 nt to about 8.500 nt, about 6,500 nt to about 9,000 nt, about 7,000 nt to about 7,500 nt, about 7,000 nt to about 8,000 nt, about 7,000 nt to about 8,500 nt, about 7,000 nt to about 9,000 nt, about 7,500 nt to about 8,000 nt, about 7.500 nt to about 8,500 nt, about 7,500 nt to about 9,000 nt, about 8,000 nt to about 8,500 nt, about 8,000 nt to about 9,000 nt, about 8,500 nt to about 9,000 nt, about 5,000 nt, about 5,500 nt, about 6,000 nt, about 6.500 nt, about 7,000 nt, about 7,500 nt, about 8,000 nt, about 8,500 nt, and about 9,000 nt.
49. The system of any one of claims 41 to 48, wherein the first dimerization domain and the second dimerization domain are each no more than 1000 nt, such as at least 50 nt, at least 100 nt, at least 150 nt, at least 200 nt, at least 300 nt, at least 400 nt, at least 500 nt, 50 to 1000 nt, 50 to 500 nt, 50 to 150 nt, 50, 100, 150, 200, 250, 300, 400, or 500 nt; and the system has a recombination efficiency of at least about 20%, at least about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, or about 100%.
50. The system of any one of claims 41 to 49, wherein each dimerization domain is no more than 1000 nt, such as at least 50 nt, at least 100 nt, at least 150 nt, at least 200 nt, at least 300 nt, at least 400 nt, at least 500 nt, 50 to 1000 nt, 50 to 500 nt, 50 to 150 nt, 50, 100, 150, 200, 250, 300, 400, or 500 nt; and the system has a recombination efficiency of at least 20%, at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, at least 75%, at least 80%, at least 90%, or about 100%.
51. The system of any one of claims 41 to 50, wherein the RNA recombination efficiency is about 10% to about 100%, about 10% to about 20%, about 10% to about 30%, about 10% to about 35%, aboutSLR 7158-102574-27 05 / 08 / 25 S2024-016A10% to about 40%, about 10% to about 45%, about 10% to about 50%, about 10% to about 55%, about 10% to about 60%, about 10% to about 70%, about 10% to about 80%, about 10% to about 90%, about 20% to about 30%, about 20% to about 35%, about 20% to about 40%, about 20% to about 45%, about 20% to about 50%, about 20% to about 55%, about 20% to about 60%, about 20% to about 70%, about 20% to about 80%, about 20% to about 90%, about 30% to about 35%, about 30% to about 40%, about 30% to about 45%, about 30% to about 50%, about 30% to about 55%, about 30% to about 60%, about 30% to about 70%, about 30% to about 80%, about 30% to about 90%, about 35% to about 40%, about 35% to about 45%, about 35% to about 50%, about 35% to about 55%, about 35% to about 60%, about 35% to about 70%, about 35% to about 80%, about 35% to about 90%, about 40% to about 45%, about 40% to about 50%, about 40% to about 55%, about 40% to about 60%, about 40% to about 70%, about 40% to about 80%, about 40% to about 90%, about 45% to about 50%, about 45% to about 55%, about 45% to about 60%, about 45% to about 70%, about 45% to about 80%, about 45% to about 90%, about 50% to about 55%, about 50% to about 60%, about 50% to about 70%, about 50% to about 80%, about 50% to about 90%, about 55% to about 60%, about 55% to about 70%, about 55% to about 80%, about 55% to about 90%, about 60% to about 70%, about 60% to about 80%, about 60% to about 90%, about 70% to about 80%, about 70% to about 90%, about 80% to about 90%, about 10%, about 20%, about 30%, about 35%, about 40%, about 45%, about 50%, about 55%, about 60%, about 70%, about 80%, about 90%, about 95%, or about 100%.
52. A composition comprising a system of any one of claims 41 to 51.
53. The composition of claim 52, wherein the composition comprises first, second, third and optionally fourth RNA molecules, each encoding at least a portion of a nucleic acid editing protein or a target protein.
54. A kit comprising the system of any one of claims 41 to 51, or composition of any one of claims 52 and 53, wherein any of the synthetic first, second, third and fourth nucleic acid molecules can be in separate containers, and optionally further comprising a buffer such as a pharmaceutically acceptable carrier.
55. A method of expressing a nucleic acid editing protein in a cell, comprising:introducing the DNA composition of any one of claims 35-40, the system of any one of claims 41 to 51, or the composition of claim 52 or 53, into a cell (such as a cell in the eye or ear, or a skeletal or cardiac muscle cell), and expressing the first and second RNA molecules in the cell, wherein the nucleic acid editing protein or the target protein is produced in the cell.
56. The method of claim 55, wherein the cell is in a subject, and introducing comprises administering a therapeutically effective amount the system to the subject.
57. The method of claim 56, wherein the method treats a genetic disease caused by a mutation in a target nucleic acid molecule in the subject, wherein the method results in expression of the nucleic acid editing protein or the target protein and optionally the first, second, third, and / or fourth gRNAs and correction of the mutation in the subject.
58. The method of claim 57, whereinSLR 7158-102574-27 05 / 08 / 25 S2024-016Athe genetic disease is one resulting from loss of function mutation, and the at least one first, second, third and / or fourth gRNA comprise a sequence that targets a nucleic acid listed in Table 1, and the method treats the corresponding disease listed in Table 1;the genetic disease is Usher Syndrome IF, and the at least one first, second, third and / or fourth gRNA comprise a sequence that targets a mutated PCDH15 gene, or the target protein is a functional Pcdhl5 protein;the genetic disease is DMD, and the at least one first, second, third and / or fourth gRNA comprise a sequence that targets a mutated DMD gene, or the target protein is a functional dystrophin protein;the target nucleic acid is an oncogene, and least one first, second, third and / or fourth gRNA comprise a sequence that targets an oncogene in Table 2 and the method treat the corresponding cancer listed in Table 2;the genetic disease is one resulting from gain of function mutation, and the at least one first, second, third and / or fourth gRNA comprise a sequence that targets a nucleic acid listed in Table 3, and the method treats the corresponding disease listed in Table 3; orthe genetic disease is one listed in Table 4, wherein at least one first, second, third and / or fourth gRNA comprise a sequence targets a nucleic acid listed in Table 4, and the method treats the corresponding disease listed in Table 4.
59. A system of any one of claims 41 to 51, a composition of any one of claims 1 to 40, 52 and 53, or a method of any one of claims 55 to 58, wherein one or both of the first and second RNA molecules comprise at least 80%, at least 85%, at least 90%, at least 95%, at least 98%, at least 99%, or 100% sequence identity to a synthetic intron provided in any one of SEQ ID NOS: 159, 160, 161, 162, 163, 164, 165, 166, 225, and 226.
60. A system of any one of claims 41 to 51 and 59, a composition of any one of claims 1 to 40, 52 and 53, or a method of any one of claims 55 to 58, wherein one or both of the first and second RNA molecules comprise a synthetic intron selected from the RNA encoded by SEQ ID NO: 159, 160, 161, 162, 163, 164, 165, 166, 225, or 226.
61. A system of any of claims 41 to 51 and 59-60, a composition of any one of claims 1 to 40, 52 and 53, or a method of any one of claims 55 to 58, wherein one or both of the first and second RNA molecules further comprise a portion of a protein coding sequence.
62. A system of any of claims 41 to 51 and 59-61, a composition of any one of claims 1 to 40, 52 and 53, or a method of any one of claims 55 to 58, wherein the portion of the protein coding sequence comprises an N-terminal half, an N-terminal portion, a C-terminal half, or a C-terminal portion, of the protein coding sequence.
63. A system of any one of claims 41 to 51 and 59-61, a composition of any one of claims 1 to 40, 52 and 53, or a method of any one of claims 55 to 58, comprising: (a) a first RNA molecule, the RNA molecule comprising from 5’ to 3’: (i) at least one first gRNA, (ii) a coding sequence for an N-terminal portion of the nucleic acid editing protein; (iii) a splice donor; (ii-2) a DISE, an ISE, or both; (iv) a firstSLR 7158-102574-27 05 / 08 / 25 S2024-016Adimerization domain; and (v) at least one second gRNA, and (b) a second RNA molecule, the RNA molecule comprising from 5’ to 3’: (i) at least one third gRNA, (ii) a second dimerization domain, wherein the second dimerization domain binds to the first dimerization domain; (i-2) at least one ISE sequence; (iii) a branch point sequence; (ii) a polypyrimidine tract; (v) a splice acceptor; (vi) a coding sequence for a C-terminal portion of the nucleic acid editing protein, and (vii) at least one fourth gRNA.
64. A system of any one of claims 41 to 51 and 59-61, a composition of any one of claims 1 to 40, 52 and 53, or a method of any one of claims 55 to 58, comprising: (a) a first RNA molecule, the RNA molecule comprising from 5’ to 3’: (i) at least one first gRNA, (ii) a coding sequence for an N-terminal portion of the nucleic acid editing protein; (iii) a splice donor; (ii-2) a DISE, an ISE, and an ISE; and (iv) a first dimerization domain; and (v) at least one second gRNA, and (b) a second RNA molecule, the RNA molecule comprising from 5’ to 3’: (i) at least one third gRNA, (ii) a second dimerization domain, wherein the second dimerization domain binds to the first dimerization domain; (i-2) three ISE sequences; (iii) a branch point sequence; (iv) a polypyrimidine tract; (v) a splice acceptor; (vi) a coding sequence for a C-terminal portion of the nucleic acid editing protein and (vii) at least one fourth gRNA.
65. The composition of any one of claims 1 to 40, wherein any one or two of the first and second RNA molecules each has a size independently selected from: about 2500 to 4500 nt, about 2,500 nt to about 2,750 nt, about 2,500 nt to about 3,000 nt, about 2,500 nt to about 3,250 nt, about 2,500 nt to about 3,500 nt, about 2,500 nt to about 3,750 nt, about 2,500 nt to about 4,000 nt, about 2,500 nt to about 4,250 nt, about 2,500 nt to about 4,500 nt, about 2,750 nt to about 3,000 nt, about 2,750 nt to about 3,250 nt, about 2,750 nt to about 3,500 nt, about 2,750 nt to about 3,750 nt, about 2,750 nt to about 4,000 nt, about 2,750 nt to about 4,250 nt, about 2,750 nt to about 4,500 nt, about 3,000 nt to about 3,250 nt, about 3,000 nt to about 3,500 nt, about 3,000 nt to about 3,750 nt, about 3,000 nt to about 4,000 nt, about 3,000 nt to about 4,250 nt, about 3,000 nt to about 4,500 nt, about 3,250 nt to about 3,500 nt, about 3,250 nt to about 3,750 nt, about 3,250 nt to about 4,000 nt, about 3,250 nt to about 4,250 nt, about 3,250 nt to about 4,500 nt, about 3,500 nt to about 3,750 nt, about 3,500 nt to about 4,000 nt, about 3,500 nt to about 4,250 nt, about 3,500 nt to about 4,500 nt, about 3,750 nt to about 4,000 nt, about 3,750 nt to about 4,250 nt, about 3,750 nt to about 4,500 nt, about 4,000 nt to about 4,250 nt, about 4,000 nt to about 4,500 nt, about 4,250 nt to about 4,500 nt, about 2,500 nt, about 2,750 nt, about 3,000 nt, about 3,250 nt, about 3,500 nt, about 3,750 nt, about 4,000 nt, about 4,250 nt, and about 4,500 nt.500 nt to about 3,750 nt, about 3,500 nt to about 4,000 nt, about 3,500 nt to about 4,250 nt, about 3.500 nt to about 4,500 nt, about 3,750 nt to about 4,000 nt, about 3,750 nt to about 4,250 nt, about 3,750 nt to about 4,500 nt, about 4,000 nt to about 4,250 nt, about 4,000 nt to about 4,500 nt, about 4,250 nt to about 4.500 nt, about 2,500 nt, about 2,750 nt, about 3,000 nt, about 3,250 nt, about 3,500 nt, about 3,750 nt, about 4,000 nt, about 4,250 nt, and about 4,500 nt.00 nt to about 3,750 nt, about 3,500 nt to about 4,000 nt, about 3,500 nt to about 4,250 nt, about 3.500 nt to about 4,500 nt, about 3,750 nt to about 4,000 nt, about 3,750 nt to about 4,250 nt, about 3,750 nt to about 4,500 nt, about 4,000 nt to about 4,250 nt, about 4,000 nt to about 4,500 nt, about 4,250 nt to aboutSLR 7158-102574-27 05 / 08 / 25 S2024-016A4,500 nt, about 2,500 nt, about 2,750 nt, about 3,000 nt, about 3,250 nt, about 3,500 nt, about 3,750 nt, about 4,000 nt, about 4,250 nt, and about 4,500 nt.
66. The composition of any one of claims 1 to 40, wherein:the total nucleic acid editing protein or target protein coding sequence size is about 2000 nt to about 8000 nt, about 2,000 nt to about 3,000 nt, about 2,000 nt to about 3,500 nt, about 2,000 nt to about 4,000 nt, about 2,000 nt to about 4,500 nt, about 2,000 nt to about 5,000 nt, about 2,000 nt to about 5,500 nt, about 2,000 nt to about 6,000 nt, about 2,000 nt to about 6,500 nt, about 2,000 nt to about 7,000 nt, about 2,000 nt to about 7,500 nt, about 2,000 nt to about 8,000 nt, about 3,000 nt to about 3,500 nt, about 3,000 nt to about 4,000 nt, about 3,000 nt to about 4,500 nt, about 3,000 nt to about 5,000 nt, about 3,000 nt to about 5,500 nt, about 3,000 nt to about 6,000 nt, about 3,000 nt to about 6,500 nt, about 3,000 nt to about 7,000 nt, about 3,000 nt to about 7,500 nt, about 3,000 nt to about 8,000 nt, about 3,500 nt to about 4,000 nt, about 3,500 nt to about 4,500 nt, about 3,500 nt to about 5,000 nt, about 3,500 nt to about 5,500 nt, about 3,500 nt to about 6,000 nt, about 3,500 nt to about 6,500 nt, about 3,500 nt to about 7,000 nt, about 3,500 nt to about 7,500 nt, about 3,500 nt to about 8,000 nt, about 4,000 nt to about 4,500 nt, about 4,000 nt to about 5,000 nt, about 4,000 nt to about 5,500 nt, about 4,000 nt to about 6,000 nt, about 4,000 nt to about 6,500 nt, about 4,000 nt to about 7,000 nt, about 4,000 nt to about 7,500 nt, about 4,000 nt to about 8,000 nt, about 4,500 nt to about 5,000 nt, about 4,500 nt to about 5,500 nt, about 4,500 nt to about 6,000 nt, about 4,500 nt to about 6,500 nt, about 4,500 nt to about 7,000 nt, about 4,500 nt to about 7,500 nt, about 4,500 nt to about 8,000 nt, about 5,000 nt to about 5,500 nt, about 5,000 nt to about 6,000 nt, about 5,000 nt to about 6,500 nt, about 5,000 nt to about 7,000 nt, about 5,000 nt to about 7,500 nt, about 5,000 nt to about 8,000 nt, about 5,500 nt to about 6,000 nt, about 5,500 nt to about 6,500 nt, about 5,500 nt to about 7,000 nt, about 5,500 nt to about 7,500 nt, about 5,500 nt to about 8,000 nt, about 6,000 nt to about 6,500 nt, about 6,000 nt to about 7,000 nt, about 6,000 nt to about 7,500 nt, about 6,000 nt to about 8,000 nt, about 6,500 nt to about 7,000 nt, about 6,500 nt to about 7,500 nt, about 6,500 nt to about 8,000 nt, about 7,000 nt to about 7,500 nt, about 7,000 nt to about 8,000 nt, about 7,500 nt to about 8,000 nt, about 2,000 nt, about 3,000 nt, about 3,500 nt, about 4,000 nt, about 4,500 nt, about 5,000 nt, about 5,500 nt, about 6,000 nt, about 6,500 nt, about 7,000 nt, about 7,500 nt, or about 8,000 nt; and / orbout 7,500 nt, about 6,000 nt to about 8,000 nt, about 6,500 nt to about 7,000 nt, about 6,500 nt to about 7,500 nt, about 6,500 nt to about 8,000 nt, about 7,000 nt to about 7,500 nt, about 7,000 nt to about 8,000 nt, about 7,500 nt to about 8,000 nt, about 2,000 nt, about 3,000 nt, about 3,500 nt, about 4,000 nt, about 4,500 nt, about 5,000 nt, about 5,500 nt, about 6,000 nt, about 6,500 nt, about 7,000 nt, about 7,500 nt, or about 8,000 nt; and / orout 7,500 nt, about 6,000 nt to about 8,000 nt, about 6,500 nt to about 7,000 nt, about 6,500 nt to about 7,500 nt, about 6,500 nt to about 8,000 nt, about 7,000 nt to about 7,500 nt, about 7,000 nt to about 8,000 nt, about 7,500 nt to about 8,000 nt, about 2,000 nt, about 3,000 nt, about 3,500 nt, about 4,000 nt, about 4,500 nt, about 5,000 nt, about 5,500 nt, about 6,000 nt, about 6,500 nt, about 7,000 nt, about 7,500 nt, or about 8,000 nt; and / orSLR 7158-102574-27 05 / 08 / 25 S2024-016Athe summed size of the two RNA molecules is about 5,000 nt to about 9000 nt, about 5,000 nt to about 5,500 nt, about 5,000 nt to about 6,000 nt, about 5,000 nt to about 6,500 nt, about 5,000 nt to about 7,000 nt, about 5,000 nt to about 7,500 nt, about 5,000 nt to about 8,000 nt, about 5,000 nt to about 8,500 nt, about 5,000 nt to about 9,000 nt, about 5,500 nt to about 6,000 nt, about 5,500 nt to about 6,500 nt, about 5,500 nt to about 7,000 nt, about 5,500 nt to about 7,500 nt, about 5,500 nt to about 8,000 nt, about 5,500 nt to about 8,500 nt, about 5,500 nt to about 9,000 nt, about 6,000 nt to about 6,500 nt, about 6,000 nt to about 7,000 nt, about 6,000 nt to about 7,500 nt, about 6,000 nt to about 8,000 nt, about 6,000 nt to about 8,500 nt, about 6,000 nt to about 9,000 nt, about 6,500 nt to about 7,000 nt, about 6,500 nt to about 7,500 nt, about 6,500 nt to about 8,000 nt, about 6,500 nt to about 8,500 nt, about 6,500 nt to about 9,000 nt, about 7,000 nt to about 7,500 nt, about 7,000 nt to about 8,000 nt, about 7,000 nt to about 8,500 nt, about 7,000 nt to about 9,000 nt, about 7,500 nt to about 8,000 nt, about 7,500 nt to about 8,500 nt, about 7,500 nt to about 9,000 nt, about 8,000 nt to about 8,500 nt, about 8,000 nt to about 9,000 nt, about 8,500 nt to about 9,000 nt, about 5,000 nt, about 5,500 nt, about 6,000 nt, about 6,500 nt, about 7,000 nt, about 7,500 nt, about 8,000 nt, about 8,500 nt, or about 9,000 nt.about 9,000 nt, about 6,000 nt to about 6,500 nt, about 6,000 nt to about 7,000 nt, about 6,000 nt to about 7,500 nt, about 6,000 nt to about 8,000 nt, about 6,000 nt to about 8,500 nt, about 6,000 nt to about 9,000 nt, about 6,500 nt to about 7,000 nt, about 6,500 nt to about 7,500 nt, about 6,500 nt to about 8,000 nt, about 6,500 nt to about 8,500 nt, about 6,500 nt to about 9,000 nt, about 7,000 nt to about 7,500 nt, about 7,000 nt to about 8,000 nt, about 7,000 nt to about 8,500 nt, about 7,000 nt to about 9,000 nt, about 7,500 nt to about 8,000 nt, about 7,500 nt to about 8,500 nt, about 7,500 nt to about 9,000 nt, about 8,000 nt to about 8,500 nt, about 8,000 nt to about 9,000 nt, about 8,500 nt to about 9,000 nt, about 5,000 nt, about 5,500 nt, about 6,000 nt, about 6,500 nt, about 7,000 nt, about 7,500 nt, about 8,000 nt, about 8,500 nt, or about 9,000 nt.nt, about 7,000 nt to about 8,500 nt, about 7,000 nt to about 9,000 nt, about 7,500 nt to about 8,000 nt, about 7,500 nt to about 8,500 nt, about 7,500 nt to about 9,000 nt, about 8,000 nt to about 8,500 nt, about 8,000 nt to about 9,000 nt, about 8,500 nt to about 9,000 nt, about 5,000 nt, about 5,500 nt, about 6,000 nt, about 6,500 nt, about 7,000 nt, about 7,500 nt, about 8,000 nt, about 8,500 nt, or about 9,000 nt.bout 9,000 nt, about 7,500 nt to about 8,000 nt, about 7,500 nt to about 8,500 nt, about 7,500 nt to about 9,000 nt, about 8,000 nt to about 8,500 nt, about 8,000 nt to about 9,000 nt, about 8,500 nt to about 9,000 nt, about 5,000 nt, about 5,500 nt, about 6,000 nt, about 6,500 nt, about 7,000 nt, about 7,500 nt, about 8,000 nt, about 8,500 nt, or about 9,000 nt.
67. The composition of any one of claims 1 to 40, wherein the first dimerization domain and the second dimerization domain are each no more than 1000 nt, such as at least 50 nt, at least 100 nt, at least 150 nt, at least 200 nt, at least 300 nt, at least 400 nt, at least 500 nt, 50 to 1000 nt, 50 to 500 nt, 50 to 150 nt, 50, 100, 150, 200, 250, 300, 400, or 500 nt; and the system has a recombination efficiency of at least about 20%, at least about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 45%, at leastSLR 7158-102574-27 05 / 08 / 25 S2024-016Aabout 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, or at least about 95%.
68. The composition of any one of claims 1 to 40, wherein each dimerization domain is no more than 1000 nt, such as at least 50 nt, at least 100 nt, at least 150 nt, at least 200 nt, at least 300 nt, at least 400 nt, at least 500 nt, 50 to 1000 nt, 50 to 500 nt, 50 to 150 nt, 50, 100, 150, 200, 250, 300, 400, or 500 nt; and the system has a recombination efficiency of at least 20%, at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, at least 75%, at least 80%, or at least 90%.
69. The composition of any one of claims 1 to 40, wherein the RNA recombination efficiency is about 10% to about 100%, about 10% to about 20%, about 10% to about 30%, about 10% to about 35%, about 10% to about 40%, about 10% to about 45%, about 10% to about 50%, about 10% to about 55%, about 10% to about 60%, about 10% to about 70%, about 10% to about 80%, about 10% to about 90%, about 20% to about 30%, about 20% to about 35%, about 20% to about 40%, about 20% to about 45%, about 20% to about 50%, about 20% to about 55%, about 20% to about 60%, about 20% to about 70%, about 20% to about 80%, about 20% to about 90%, about 30% to about 35%, about 30% to about 40%, about 30% to about 45%, about 30% to about 50%, about 30% to about 55%, about 30% to about 60%, about 30% to about 70%, about 30% to about 80%, about 30% to about 90%, about 35% to about 40%, about 35% to about 45%, about 35% to about 50%, about 35% to about 55%, about 35% to about 60%, about 35% to about 70%, about 35% to about 80%, about 35% to about 90%, about 40% to about 45%, about 40% to about 50%, about 40% to about 55%, about 40% to about 60%, about 40% to about 70%, about 40% to about 80%, about 40% to about 90%, about 45% to about 50%, about 45% to about 55%, about 45% to about 60%, about 45% to about 70%, about 45% to about 80%, about 45% to about 90%, about 50% to about 55%, about 50% to about 60%, about 50% to about 70%, about 50% to about 80%, about 50% to about 90%, about 55% to about 60%, about 55% to about 70%, about 55% to about 80%, about 55% to about 90%, about 60% to about 70%, about 60% to about 80%, about 60% to about 90%, about 70% to about 80%, about 70% to about 90%, about 80% to about 90%, about 10%, about 20%, about 30%, about 35%, about 40%, about 45%, about 50%, about 55%, about 60%, about 70%, about 80%, about 90%, about 95%, or about 100%.
70. The composition of any one of claims 1 to 40, wherein:(a) the first and second RNA molecules are each about 2500 nt to 4500 nt;(b) the total nucleic acid editing protein coding sequence size is about 2000 nt to about 8000 nt; and / or(c) the summed size of the two RNA molecules is about 5,000 nt to about 9000 nt;and the RNA recombination efficiency is about 10% to about 100%.
71. The composition of any one of claims 1 to 34, further comprising one or more additional RNA molecules comprising one or more gRNAs specific for one or more target nucleic acid molecules.
72. The composition of any one of claims 35 to 40, further comprising one or more additional DNA molecules encoding one or more gRNAs specific for one or more target nucleic acid molecules.SLR 7158-102574-27 05 / 08 / 25 S2024-016A73. A dual vector composition for expressing a target protein or a nucleic acid editing protein, comprising:(a) a first adeno- associated viral (AAV) vector packaging a first transgene, the first transgene comprising from 5’ to 3’:(i) a first promotor;(ii) a first coding sequence comprising a coding sequence for an N-terminal portion of the target protein or the nucleic acid editing protein;(iii) a first intronic sequence comprising a splice donor and a first dimerization domain; and (b) a second AAV vector packaging a second transgene, the second transgene comprising from 5’ to 3’:(i) a second promoter;(ii) a second intronic sequence comprising a second dimerization domain, a branch point sequence, polypyrimidine tract, and a splice acceptor; and(iii) a second coding sequence comprising a coding sequence for a C-terminal portion of the target protein or the nucleic acid editing protein,wherein the first promoter is operably linked to the first coding sequence,wherein the second promoter is operably linked to the second coding sequence, and wherein, RNA transcribed from the first intronic sequence form one or more first secondary structures, and / or RNA transcribed from the second intronic sequence form one or more second secondary structures.
74. A dual vector composition for expressing a target protein or a nucleic acid editing protein, comprising:(a) a first adeno- associated viral (AAV) vector packaging a first transgene, the first transgene comprising from 5’ to 3’:(i) a first promotor;(ii) a first coding sequence comprising a coding sequence for an N-terminal portion of the target protein or the nucleic acid editing protein;(iii) a first intronic sequence comprising a splice donor and a first dimerization domain; and (b) a second AAV vector packaging a second transgene, the second transgene comprising from 5’ to 3’:(i) a second promoter;(ii) a second intronic sequence comprising a second dimerization domain, a branch point sequence, polypyrimidine tract, and a splice acceptor; and(iii) a second coding sequence comprising a coding sequence for a C-terminal portion of the target protein or the nucleic acid editing protein,wherein the first promoter is operably linked to the first coding sequence,SLR 7158-102574-27 05 / 08 / 25 S2024-016Awherein the second promoter is operably linked to the second coding sequence, and wherein, RNA transcribed from the first intronic sequence, and / or RNA transcribed from the second intronic sequence comprises one or more elements that recruit nuclear RNA binding proteins to transcripts of the first and second transgenes.
75. The composition of claim 73 or 74, wherein the composition is for expressing a nucleic acid editing protein, and the first transgene and / or the second transgene further encodes at least one first guide RNA (gRNA) specific for a first target nucleic acid molecule, wherein the at least one first gRNA directs the nucleic acid editing protein to a target editing site on the first target nucleic acid molecule;optionally, the first transgene and / or the second transgene further encodes at least one second gRNA specific for (i) the first target nucleic acid molecule, wherein the at least one second gRNA directs the nucleic acid editing protein to the same or different target editing site on the first nucleic acid molecule, or (ii) a second target nucleic acid molecule, wherein the at least one second gRNA directs the nucleic acid editing protein to a target editing site on the second target nucleic acid molecule;optionally, the first transgene and / or the second transgene further encodes at least one third gRNA specific for (i) the first target nucleic acid molecule, wherein the at least one third gRNA directs the nucleic acid editing protein to the same or different target editing site on the first nucleic acid molecule as the first and second gRNA, (ii) the second target nucleic acid molecule, wherein the at least one third gRNA directs the nucleic acid editing protein to the same or different target editing site on the second nucleic acid molecule as the second gRNA, or (iii) a third target nucleic acid molecule, wherein the at least one third gRNA directs the nucleic acid editing protein to a target editing site on the third target nucleic acid molecule; andoptionally, the first transgene and / or the second transgene further encodes at least one fourth gRNA specific for (i) the first target nucleic acid molecule, wherein the at least one fourth gRNA directs the nucleic acid editing protein to the same or different target editing site on the first nucleic acid molecule as the first, second, and third gRNA, (ii) the second target nucleic acid molecule, wherein the at least one fourth gRNA directs the nucleic acid editing protein to the same or different target editing site on the second nucleic acid molecule as the second and third gRNA, (iii) the third target nucleic acid molecule, wherein the at least one fourth gRNA directs the nucleic acid editing protein to the same or different target editing site on the third target nucleic acid molecule as the third gRNA, or (iv) a fourth target nucleic acid molecule, wherein the at least one fourth gRNA directs the nucleic acid editing protein to a target editing site on the fourth target nucleic acid molecule.
76. The composition of any one of claims 73-75, wherein the first and second dimerization domains, when transcribed, bind by direct binding, indirect binding, or a combination thereof.
77. The composition of claim 76, wherein the direct binding or the indirect binding comprises base pairing interactions, non-canonical base pairing interactions, non-base pairing interactions, or a combination thereof.SLR 7158-102574-27 05 / 08 / 25 S2024-016A78. The composition of claim 76 or 77, wherein the direct binding comprises base pairing interactions between kissing loops or hypodiverse regions.
79. The composition of any one of claims 76-78, wherein the direct binding comprises non-canonical base pairing interactions, non-canonical base pairing interactions, non-base pairing interactions, or a combination thereof, between aptamer regions.
80. The composition of claim 76 or 77, wherein the indirect binding comprises base pairing interactions through a nucleic acid bridge.
81. The composition of claim 76 or 77, wherein the indirect binding comprises non-base pairing interactions between an aptamer and an aptamer target, or between two aptamers.
82. The composition of any one of claims 73-81, wherein the first or second dimerization domain does not comprise a cryptic splice acceptor or cryptic splice donor.
83. The composition of any one of claims 73-82, wherein the first and second dimerization domains are directly binding or indirectly binding aptamer sequence dimerization domains.
84. The composition of any one of claims 73-83, wherein the first and second dimerization domains, when transcribed, are kissing loop interaction domains.
85. The composition of any one of claims 75-84, wherein the target editing site is part of the target nucleic acid molecule or a regulatory region of a target nucleic acid molecule associated with disease.
86. The composition of claim 85, wherein the disease is a monogenic disease.
87. The composition of claim 85 or 86, wherein the first, second, third, and / or fourth target nucleic acid molecule comprises one or more point mutations that results in the disease.
88. The composition of any one of claims 85-87, wherein the disease and the first, second, third, and / or fourth target nucleic acid molecule are one listed in Tables 1-4; or wherein the disease is Usher Syndrome IF and the first, second, third, and / or fourth target nucleic acid molecule is a PCDH15 gene or gene product, or wherein the disease is Duchenne muscular dystrophy (DMD) and the first, second, third, and / or fourth target nucleic acid molecule is a DMD gene or gene product.
89. The composition of any one of claims 73-88, wherein the first intronic sequence further comprises one or both of a downstream intronic splice enhancer (DISE) 3’ to the splice donor and 5’ to the first dimerization domain, and an intronic splice enhancer (ISE) 3’ to the splice donor and 5’ to the first dimerization domain; and / orwherein the second intronic sequence further comprises one or both of an ISE 3’ to the second dimerization domain and 5’ to the branch point sequence, and a DISE 3’ to the splice donor and 5’ to the dimerization domain.
90. The composition of any one of claims 73-89, whereinthe first transgene further encodes a self-cleaving RNA sequence or an RNA-cleaving enzyme target sequence positioned anywhere 3’ to the splice donor such that it cleaves off the 3’ located polyadenylated tail to decrease or suppress protein fragment expression from a non-recombined RNA molecule;SLR 7158-102574-27 05 / 08 / 25 S2024-016Athe second transgene further encodes a self-cleaving RNA sequence or an RNA-cleaving enzyme target sequence positioned anywhere 5’ to the branch point sequence such that it cleaves off the 5’ located RNA cap to decrease or suppress protein fragment expression from a non-recombined RNA molecule; the second transgene further comprises a start codon anywhere 5’ to the branch point sequence that is shifted relative to the open reading frame 3’ of the splice acceptor to decrease or suppress translation of a nucleic acid editing protein fragment from a non-recombined RNA molecule;the first transgene further comprises a micro RNA target site anywhere 3’ to the splice donor such that an un-joined RNA fragment undergoes micro RNA dependent degradation once outside the nucleus; the second transgene further comprises a micro RNA target site anywhere 3’ to the coding sequence such that an un-joined RNA fragment undergoes micro RNA dependent degradation once outside the nucleus;the first transgene further comprises a sequence encoding a degron protein degradation tag anywhere 3’ to the splice donor such that it is in frame with the nucleic acid editing protein open reading frame 5’ of the splice donor site such that an un-joined protein fragment is tagged for degradation;the second transgene further comprises a start codon and an in-frame degron protein degradation tag anywhere 5’ to the branch point sequence such that it is in frame with the nucleic acid editing protein open reading frame 3’ of the splice acceptor site such that an un-joined protein fragment is tagged for degradation; orany combination thereof.
91. The composition of any one of claims 73-90, wherein the nucleic acid editing protein comprises a Cas nuclease, zinc finger nuclease, or transcription activator-like effector nuclease.
92. The composition of any one of claims 75-91, wherein the first, second, third, and / or fourth target nucleic acid molecules are target DNA molecules, and the at least one first, second, third, and / or fourth gRNA comprises a crRNA and tracrRNA93. The composition of any one of claims 73-92, wherein the nucleic acid editing protein comprises Cas9 or dead Cas9 (dCas9).
94. The composition of claim 93, wherein the Cas9 protein comprises at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity to SEQ ID NO: 208, and can function as an RNA-guided DNA endonuclease.
95. The composition of claim 93, wherein the Cas9 protein is encoded by a sequence comprising at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity to SEQ ID NO: 207, and encodes an RNA-guided DNA endonuclease.
96. The composition of claim 93, wherein the dCas9 protein comprises at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity to SEQ ID NO: 210, and is catalytically inactive.SLR 7158-102574-27 05 / 08 / 25 S2024-016A97. The composition of claim 93, wherein the dCas9 protein is encoded by a sequence comprising at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity to SEQ ID NO: 209, and encodes a protein that is catalytically inactive.
98. The composition of any one of claims 93-97, wherein the Cas9 or dCas9 is part of a fusion protein.
99. The composition of claim 98, wherein the fusion protein comprises Cas9 or dCas9 and one or more of:a transcriptional activation domain, such as VP64, P65, MyoDl, HSF1, RTA, CBP, SET7 / 9, or any combination thereof;a cytosine base editor (CBE), such as one from sea lamprey [AID], CDAI, or APOBEC3G; bacteriophage protein Gam; andan adenine base editor (ABE), such as ABE 6.3, 7.8, 7.9, 7.10, ABEmax, ABE8s, such as ABE8e(TadA-8e V106W).
100. The composition of any one of claims 75-99, wherein the first, second, third, and / or fourth target nucleic acid molecule are target RNA, and the at least one first, second, third, and / or fourth gRNA comprises one or more direct repeats and one or more spacers.
101. The composition of any of claims 73-92, wherein the nucleic acid editing protein comprises Cas13a, Cas13b, Cas13c, Cas13d, or dead Cas13d (dCas13d).
102. The composition of claim 101, wherein the Cas13d protein comprises at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity to SEQ ID NO: 212, 215, or 222, and can function as an RNA-guided RNA endonuclease.
103. The composition of claim 101, wherein the Cas13d protein is encoded by a sequence comprising at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity to SEQ ID NO: 211 or 214, and encodes an RNA-guided RNA endonuclease.
104. The composition of claim 101, wherein the dCas13d protein comprises at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity to SEQ ID NO: 215 or 216, and is catalytically inactive.
105. The composition of claim 101, wherein the dCas13d protein is encoded by a sequence that encodes a protein comprising at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity to SEQ ID NO: 215 or 216, and encodes a protein that is catalytically inactive.
106. The composition of any one of claims 101-105, wherein the Cas13a, Cas13b, Cas13c, Cas13d or dCas13d is part of a fusion protein.
107. The composition of claim 106, wherein the fusion protein comprises Cas13a, Cas13b, Cas13c, Cas13d or dCas13d and one or more of:SLR 7158-102574-27 05 / 08 / 25 S2024-016Aa transcriptional activation domain, such as VP64, P65, MyoDl, HSF1, RTA, CBP, SET7 / 9, or any combination thereof;a cytosine base editor (CBE), such as one from sea lamprey [AID], CDA1, or APOBEC3G; bacteriophage protein Gam; andan adenine base editor (ABE), such as ABE 6.3, 7.8, 7.9, 7.10, ABEmax, ABE8s, such as ABE8e(TadA-8e V106W).
108. The composition of any one of claims 73-107, wherein the at least one first, second, third, and / or fourth gRNA comprise multiple copies of each of the first, second, third, and / or fourth gRNA.
109. The composition of any one of claims 73-108, wherein the first AAV and second AAV further comprise a parvovirus inverted terminal repeat (ITR) at the 5’ and the 3’ -end of each of the first transgene and second transgene.
110. The composition of any one of claims 73-109, wherein the one or more first and second secondary structures comprise one or more hairpins, kissing stem loops, quadruplexes, or combinations thereof.
111. The composition of any one of claims 73-110, wherein the one or more elements comprise one or more MALAT1 Region M sequences (SEQ ID NO: 246), one or more SIROLIN elements (SEQ ID NO: 227). or combinations thereof.
112. The composition of any one of claims 73-111, wherein:the target protein comprises two or more splice variants of the N-terminal portion of the target protein, and the first RNA molecule comprises two or more separate RNA molecules, wherein two or more separate RNA molecules each encode a distinct form of the N-terminal portion of the target protein;the target protein comprises two or more splice variants of the C-terminal portion of the target protein, and the second RNA molecule comprises two or more separate RNA molecules, wherein two or more separate RNA molecules each encode a distinct form of the C-terminal portion of the target protein; or combinations thereof.
113. The composition of any one of claims 73-112, wherein the target protein is PCDH15, ABCA4, dystrophin, truncated dystrophin, or Myo7A.
114. The composition of any one of claims 73-113, wherein:the first or second transgene comprises a third promoter operably linked to a sequence encoding the at least one first gRNA;optionally, the first or second transgene comprises a fourth promoter operably linked to a sequence encoding the at least one second gRNA;optionally, the first or second transgene comprises a fifth promoter operably linked to a sequence encoding the at least one third gRNA; andoptionally, the first or second transgene comprises a sixth promoter operably linked to a sequence encoding the at least one fourth gRNA.SLR 7158-102574-27 05 / 08 / 25 S2024-016A115. The composition of any one of claims 73-114, wherein each promoter is independently selected.
116. The composition of any one of claims 73-115, wherein:the first and second promoter are the same promoter;the first and second promoter are different promoters,the third, fourth, fifth and sixth promoters are the same promoter;the third, fourth, fifth and sixth promoters are different promoters, orcombinations thereof.
117. The composition of any one of claims 73-116, wherein each of the first and second promoters is independently selected from: a constitutive promoter; a tissue-specific promoter, such as an ocularspecific promoter, such as a GRK1 promoter; and a promoter endogenous to the nucleic acid editing protein.1 18. The composition of any one of claims 73-117, wherein each of the third, fourth, fifth and sixth promoters is a polymerase III promoter, such as a U6 or Hl promoter.
119. A method of treating a subject, comprising administering the composition of any one of claims 73-118 to the subject.
120. The composition, system, kit, or method of any one of the claims 1-72, wherein the first and second dimerization domains comprise one, two, three, four, or five kissing loops.
121. The composition or method of claims 73-119, wherein the first and second secondary structures comprise one, two, three, four, or five kissing loops.
122. The composition, system, kit, or method of any one of the claims 1-72, wherein the first dimerization domain comprises a first RNA hairpin, wherein the first RNA hairpin comprises RNA complementary sequences separated by a region of a non-complementary RNA sequence, wherein base pairing between the complementary sequences forms a first stem and the region of the non-complementary RNA sequence forms a first loop,wherein the second dimerization domain, when transcribed, comprises a second RNA hairpin, wherein the second RNA hairpin comprises RNA complementary sequences separated by a region of a non-complementary RNA sequence, wherein base pairing between the complementary sequences forms a second stem and the region of the non-complementary RNA sequence of the second RNA hairpin forms a second loop, andwherein the first loop hybridizes with the second loop.
123. The composition or method of any one of claims 73-119, wherein the one or more first secondary structures comprise a first RNA hairpin, wherein the first RNA hairpin comprises RNA complementary sequences separated by a region of a non-complementary RNA sequence, wherein base pairing between the complementary sequences forms a first stem and the region of the non-complementary RNA sequence forms a first loop,wherein the one or more second secondary structures comprise a second RNA hairpin, wherein the second RNA hairpin comprises RNA complementary sequences separated by a region of a non-SLR 7158-102574-27 05 / 08 / 25 S2024-016Acomplementary RNA sequence, wherein base pairing between the complementary sequences forms a second stem and the region of the non-complementary RNA sequence of the second RNA hairpin forms a second loop, andwherein the first loop hybridizes with the second loop.
124. The composition, system, kit, or method of any one of the preceding claims, wherein the first or second dimerization domain includes at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity to any of SEQ ID NOS: 139-142, 153, 232-241, 272-335, and 378, or a reverse complement thereof.
125. The composition, system, kit, or method of any one of the preceding claims, wherein the first dimerization domain and the second dimerization domain comprises at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity to SEQ ID NOS: 272 and 273, respectively; SEQ ID NOS: 274 and 275, respectively; SEQ ID NOS: 276 and 277, respectively; SEQ ID NOS: 278 and 279, respectively; SEQ ID NOS: 280 and 281, respectively; SEQ ID NOS: 282 and 283, respectively; SEQ ID NOS: 284 and 285, respectively; SEQ ID NOS: 286 and 287, respectively; SEQ ID NOS: 288 and 289, respectively; SEQ ID NOS: 290 and 291, respectively; SEQ ID NOS: 292 and 293, respectively; SEQ ID NOS: 294 and 295, respectively; SEQ ID NOS: 296 and 297, respectively; SEQ ID NOS: 298 and 299, respectively; SEQ ID NOS: 300 and 301, respectively; SEQ ID NOS: 302 and 303, respectively; SEQ ID NOS: 304 and 305, respectively; SEQ ID NOS: 306 and 307, respectively; SEQ ID NOS: 308 and 309, respectively; SEQ ID NOS: 310 and 11, respectively; SEQ ID NOS: 312 and 313, respectively; SEQ ID NOS: 314 and 315, respectively; SEQ ID NOS: 316 and 317, respectively; SEQ ID NOS: 318 and 319, respectively; SEQ ID NOS: 320 and 321, respectively; SEQ ID NOS: 322 and 323, respectively; SEQ ID NOS: 324 and 325, respectively; SEQ ID NOS: 326 and 327, respectively; SEQ ID NOS: 328 and 329, respectively; SEQ ID NOS: 330 and 331, respectively; SEQ ID NOS: 332 and 333, respectively; SEQ ID NOS: 334 and 335, respectively; SEQ ID NOS: 324 and 327, respectively; SEQ ID NOS: 326 and 325, respectively; SEQ ID NOS: 330 and 333, respectively; SEQ ID NOS: 332 and 331, respectively; SEQ ID NOS: 232 and 237, respectively; SEQ ID NOS: 233 and 238, respectively; SEQ ID NOS: 234 and 239, respectively; SEQ ID NOS: 235 and 240, respectively; SEQ ID NOS: 236 and 241, respectively; SEQ ID NOS: 139 and 140, respectively; or SEQ ID NOS: 141 and 142, respectively.
126. The composition, system, kit, or method of any one of the preceding claims, wherein the coding sequence for the N-terminal portion of the target protein and / or the C-terminal portion of the target protein comprises a stimulatory intron,wherein the coding sequence for the N-terminal portion of the nucleic acid editing protein and / or the C-terminal portion of the nucleic acid editing protein comprises a stimulatory intron, orwherein the first coding sequence and / or the second coding sequence comprises a stimulatory intron.
127. The composition, system, kit, or method of claim 126, wherein the stimulatory intron comprises a sequence that is at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at leastSLR 7158-102574-27 05 / 08 / 25 S2024-016A93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identical to SEQ ID NO: 345.
128. The composition, system, kit, or method of any one of the preceding claims, wherein the first RNA molecule, and / or second RNA molecule does not comprise a cryptic splice site, orwherein the first transgene and / or second transgene does not comprise a cryptic splice site.