Optimized acceptor splice site module for biological and biotechnological applications

JP2025060848A5Inactive Publication Date: 2025-10-09VIGENERON GMBH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2024229104
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2018-12-20
Filing Date
2024-12-25
Publication Date
2025-10-09
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Current methods for mRNA splicing, particularly in trans-splicing applications like gene therapy, face challenges due to unreliable prediction of acceptor splice sites and inefficient reconstitution of split coding sequences in adeno-associated virus (AAV) dual vector systems, limiting the effectiveness of gene delivery.

Method used

A novel acceptor splice region comprising a pyrimidine tract and a specific acceptor splice site, along with a nucleotide sequence of interest, is introduced into pre-mRNA molecules to enhance splicing efficiency, enabling effective trans-splicing and reconstitution of coding sequences in AAV vectors.

Benefits of technology

The novel acceptor splice region significantly improves splicing efficiency, allowing for more effective gene therapy by enhancing the reconstitution of coding sequences in AAV vectors, particularly in retinal cells, thereby addressing the limitations of current AAV dual vector systems.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000080_0000
    Figure 00000080_0000
  • Figure 00000080_0001
    Figure 00000080_0001
  • Figure 00000081_0000
    Figure 00000081_0000
Patent Text Reader

Abstract

To identify an optimized and experimentally validated strong acceptor splice site (ASS), for an optimal performance of mRNA splicing and particularly mRNA trans-splicing, so as to meet an existing need for strong ASS regions, which can inter alia be used in the development of further AAV vectors, such as next-generation rAAV dual vector systems or AAVs that deliver highly specific and efficient pre-mRNA trans-splicing molecules, which can be used in gene therapy.SOLUTION: The present invention relates to a novel acceptor splice region, as well as uses and applications thereof.SELECTED DRAWING: None
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] The present invention relates to novel acceptor splice regions, and their uses and applications. [Background technology]

[0002] Most eukaryotic genes contain non-coding introns. These must be removed from precursor messenger RNAs (pre-mRNAs) to generate translatable mature messenger RNAs (mRNAs) in a process called "splicing". Splicing is mediated through the large ribonucleoprotein complex, which consists of five conserved small nuclear ribonucleoproteins (snRNPs), namely U1, U2, U4, U5, and U6 snRNPs, the spliceosome, and more than 300 other proteins. The spliceosome assembles and degrades each intron in a highly dynamic process. For this purpose, specific splicing sequences must be recognized, so that exon and intron boundaries are clarified by distinct sequence motifs that serve as binding sites for ribonucleoproteins involved in the regulation of mRNA splicing.

[0003] The most relevant splice motifs are the canonical donor and acceptor splice sites that define exon-intron boundaries. The 5' donor splice site (DSS) is at the 3' end of the exon and the 5' end of the downstream intron, and the 3' acceptor splice site (ASS) is at the 3' end of the intron and the 5' end of the downstream exon. A functional acceptor splice site (ASS) requires three distinct elements, usually located in a range of about 50 bp: a branch point (BP), a polypyrimidine tract (PPT), and a canonical ASS sequence.

[0004] The so-called GT-AG splicing is the most common type of mammalian mRNA splicing and defines the first two and the last two nucleotides of the intron at the DNA level. Thus, part of the canonical DSS sequence represents the first two bases of the 5' intron sequence, which are a guanine followed by a thymine (GT) in the DNA sequence (corresponding to GU in the RNA sequence), and the last two bases, which are an adenine followed by a guanine (AG), representing the most conserved part of the consensus ASS. The GT-AG nucleotides are essential for an efficient splice reaction, and their breakage or replacement leads to the loss of a functional splice site.

[0005] Each splicing cycle consists of two transesterification reactions. In the first reaction, known as the branch point, a branch point nucleoside (usually adenosine) attacks the phosphate bond at the 5' exon-intron junction. This results in the formation of an intron lariat-3' exon intermediate (containing the exon and intron downstream of the splice site) and a free 5' exon end (containing the exon upstream of the splice site). In the second reaction, called exon ligation, the free end of the 5' exon attacks the phosphate at the intron-3' exon junction, causing the ligation of the two exons and the release of the intron lariat structure. The splicing reaction is controlled by auxiliary cis-acting splicing regulatory elements in the pre-mRNA, consisting of up to 10 nucleotides. Depending on their function and location, they can be classified as exonic splicing enhancing elements (ESEs), exonic splicing enhancing elements (ESSs), intronic splicing enhancing elements (ISEs), and intronic splicing suppressing elements (ISSs). These elements have been proposed to play important roles in alternative splicing because they can recruit trans-acting proteins to promote or prevent exon inclusion.

[0006] The splicing efficiency of a particular exon is expected to depend on the strength of the donor and acceptor splice sites. The strength of this DSS or ASS depends on the intronic or exonic sequence elements upstream or downstream of the GT or AG. The strength of a functional DSS can be reliably predicted using standard in silico prediction software (e.g., NNSplice: http: / / www.fruitfly.org / seq_tools / splice.html or Human Splice Finder: http: / / www.umd.be / HSF3 / ). In contrast, due to its complexity, in silico prediction of ASS strength leads to unreliable results (e.g., Koller et al., (2011) “A novel screening system improves genetic correction by internal exon replacement.” Nucleic acids research. 39:e108; Lorain et al., (2013) “Dystrophin rescue by trans-splicing: a strategy for DMD genotypes not eligible for exon skipping approaches.” Nucleic acids research. 41:8391-8402). Therefore, the actual ASS strength needs to be experimentally validated. Many biological, biotechnological and therapeutic applications rely on the efficiency of classical GT-AG mRNA splicing and thus on the use of strong splice sites.

[0007] Apart from the usual cis-splicing events that generate translatable mature mRNA from pre-mRNA by removing non-coding introns, splicing can also occur in trans, thereby combining two separate pre-mRNA molecules to create non-co-linear chimeric RNAs (Lei et al., (2016) “Evolutionary insights into RNA trans-splicing in vertebrates”, Genome Biol. Evol. 8(3):562-577). This process was first found in trypanosomes. Since then, trans-splicing events have also been identified in many species, including mice (Hirano M and Noda T., (2004) “Genomic organization of the mouse Msh4 gene producing bicistronic chimeric and antisense mRNA”, Gene 342: 165-177) and human cells (Chuang et al., (2018) “Integrative transcriptome sequencing reveals extensive alternative trans-splicing and cis-back splicing in human cells”. Nucleic Acids Research, 46(7): 3671-3691), although trans-splicing appears to occur only infrequently in higher vertebrates.

[0008] Attempts have been made to utilize trans-splicing for gene therapy. Viral vectors are attractive vehicles for gene therapy, but they often have limited loading capacity, which limits the genes that can be exchanged. For example, adeno-associated viral vectors have a packaging capacity of up to about 5.0 kb, which allows for transgenes of about 4 kb or less. Trans-splicing, i.e., joining two physically separated pre-mRNAs to form a mature mRNA, may be one way to overcome these limitations. For example, spliceosome-mediated RNA trans-splicing (SMaRT), which involves an exogenous pre-mRNA trans-splicing molecule introduced into a target cell to replace only a portion of a mutated endogenous pre-mRNA, could be used as a tool for gene therapy. This may allow for the delivery of shorter coding sequences in viral vectors.

[0009] Another way to address the size limitations of viral vectors, especially adeno-associated virus (AAV)-based vectors, is the use of recombinant AAV (rAAV) dual vector technology. AAV is a single-stranded DNA virus. For rAAV dual vector technology, the coding sequence of a transgene (gene of interest) is split into at least two parts and packaged into two or more separate rAAV vectors. After co-transduction of target cells with the split genome vectors, the full-length coding sequence is reconstituted. Efficient delivery of both rAAVs to the same target cell does not appear to be limiting, since many cells, such as photoreceptor cells in the retina, show high co-transduction efficiencies of over 90%. However, efficient reconstitution of the two transgene halves remains challenging. In the years following the development of the rAAV dual vector system, reconstitution has only been addressed at the DNA level, and several strategies have been explored to improve this approach (McClements and MacLaren, (2017) “Adeno-associated virus (AAV) dual vector strategies for gene therapy encoding large transgenes, Yale Journal of Biology and Medicine 90: 611-623).

[0010] Although these rAAV dual vector systems are often misleadingly referred to as “trans-splicing dual vectors,” splicing of mRNA in these approaches does not actually occur in trans. Rather, reconstitution in the aforementioned rAAV dual vector systems relies on concatemerization of ITR structures and / or homologous recombination of overlapping sequences to generate a single pre-mRNA that is spliced ​​in cis to remove concatemerized ITR elements and / or artificial recombination elements. Such rAAV dual vector systems are disclosed, for example, by Trapani et al. (“Effective delivery of large genes to the retina by dual AAV vectors”, (2014) EMBO Molecular Medicine, 6(2): 194-211). Reconstitution of the rAAV dual vector system at the DNA level can be recognized by the presence of a promoter driving expression of the 5′ portion of the coding sequence of the first AAV vector and the absence of a promoter driving expression of the 3′ portion of the coding sequence of the second AAV vector. The reported in vivo efficiency of such rAAV dual vector reconstitution is relatively low, less than 10% (Carvalho et al., “Evaluating efficiencies of dual AAV approaches for retinal targeting”, (2017) Frontiers in Neuroscience, 11(503): 1-8). In contrast, for pre-mRNA splicing to occur in trans, both vectors in the rAAV dual vector system require promoters to generate two separate pre-mRNA molecules.

[0011] For optimal performance of mRNA splicing, especially mRNA trans-splicing, there is an unmet need for the identification of optimized and experimentally validated strong ASSs. In particular, there is a need for strong ASS regions that can be used for the development of further AAV vectors, such as next-generation rAAV dual vector systems or AAVs that deliver highly specific and efficient pre-mRNA trans-splicing molecules, which can be used for gene therapy. Summary of the Invention

[0012] The present invention, as described herein and illustrated in the examples, figures and claims, meets this need.

[0013] Provided herein is a pre-mRNA trans-splicing molecule that includes: (i) an acceptor splice region that includes: (ia) a pyrimidine tract that includes: (iaa) 5-25 nucleotides; (iab) wherein at least 60% of the nucleotides within the 5-25 nucleotides are pyrimidine bases, e.g., cytosine (C), thymine (T), and / or uracil (U); (ib) an acceptor splice site that includes: (iba) where the acceptable splice site is located 3' to the pyrimidine tract; and (ibb) where the acceptable splice site comprises a 5' to 3' sequence of NAGG, where N is A, C, T / U, or G; (ii) a nucleotide sequence of interest or a portion thereof; where the acceptable splice region is located 3' or 5' to the nucleotide sequence of interest or a portion thereof; (iii) a binding domain that targets a pre-mRNA located 3' or 5' to the nucleic acid sequence of interest or a portion thereof; and (iv) optionally a spacer sequence, where the spacer sequence is located between the binding domain and the acceptable splice region. In one embodiment, the pre-mRNA trans-splicing molecule of claim 1 comprises: (ii) a nucleotide sequence of interest or a portion thereof, wherein the acceptor splice region is located 3' or 5' of the nucleotide sequence of interest or a portion thereof; (iii) a donor splice site, wherein the donor splice site is located 3' of the nucleotide sequence of interest or a portion thereof; (iv) a first binding domain that targets a pre-mRNA located 5' of the nucleotide sequence of interest or a portion thereof; (v) a second binding domain that targets a pre-mRNA located 3' of the nucleotide sequence of interest or a portion thereof; (vi) optionally, a first spacer sequence, wherein the first spacer is located between the first binding domain and the acceptor splice region; (vii) optionally, a second spacer sequence, wherein the second spacer is located between the second binding domain and the donor splice site.

[0014] Preferably, the acceptor splice region is located 5' to the nucleotide sequence or part thereof of interest and the binding domain is located 5' to the acceptor splice region. The pre-mRNA trans-splicing molecule may further comprise a termination sequence, preferably a polyA sequence. Preferably, the pre-mRNA trans-splicing molecule according to the invention comprises a pyrimidine tract, wherein 5 to 25 nucleotides of the pyrimidine tract comprise a sequence encoded by the sequence TTTTTT or TCTTTT. Additionally or alternatively, the sequence between the last pyrimidine of the pyrimidine tract and the acceptor splice site is less than 10 bases, preferably less than 5 bases, more preferably less than 3 bases. Additionally or alternatively, the acceptor splice site has the sequence CAGG.

[0015] The acceptable splice region may comprise: (a) the 5' 7 nucleotides of a pyrimidine tract having at least 4 nucleotides of the sequence CAACGAG, where the first 5' nucleotide is C; (b) the 5' 7 nucleotides of a pyrimidine tract having at least 5 nucleotides of the sequence CAACGAG, where the first 5' nucleotide is C; (c) the 5' 7 nucleotides of a pyrimidine tract having at least 4 nucleotides of the sequence CAACGAG, where the first 5' nucleotide is CAA; or (d) the 5' 7 nucleotides of a pyrimidine tract having at least 5 nucleotides of the sequence CAACGAG, where the first 5' nucleotide is CAAC. In one embodiment, the acceptable splice region is encoded by a nucleotide sequence having the sequence of SEQ ID NO:3 or 4 or a nucleotide sequence at least 85%, 90%, 95%, 97%, 98%, 99%, or 100% identical to SEQ ID NO:3 or 4. Also provided is a DNA molecule comprising a promoter and a sequence encoding a pre-mRNA trans-splicing molecule according to the invention, wherein the DNA molecule is preferably a vector or a plasmid.

[0016] Another aspect of the invention relates to a method of making a nucleic acid sequence, the method comprising: (A) providing a first nucleic acid sequence comprising one or more donor splice site sequences; (B) providing a second nucleic acid sequence comprising: (i) an acceptor splice region comprising: (ia) a pyrimidine tract comprising: (iaa) 5-25 nucleotides; (iab) wherein at least 60% of the nucleotides within these 5-25 nucleotides are pyrimidine bases, e.g., cytosine (C), thymine (T), and / or uracil (U); (ib) an acceptor splice site, (iba) wherein the acceptor splice site is located 3' to the pyrimidine tract; and (ibb) wherein the acceptor splice site comprises a sequence 5' to 3' of NAGG, where N is A, C, T / U, or G; (C) cleaving a first nucleic acid sequence at one or more donor splice site sequences and cleaving a second nucleic acid sequence at the acceptor splice site; (D) ligating the first cleaved nucleic acid sequence to the second cleaved nucleic acid sequence, thereby obtaining a nucleic acid sequence. The first nucleic acid sequence further comprises a nucleotide sequence of interest or a portion thereof, wherein at least a portion of the nucleotide sequence of interest is located 5' to the donor splice site, and the second nucleic acid sequence further comprises a nucleotide sequence of interest or a portion thereof, wherein at least a portion of the nucleotide sequence of interest is located 3' to the acceptor splice region. Preferably, the first and second nucleic acid sequences are introduced into a host cell, and preferably the first and second nucleic acid sequences are recombinant nucleic acid sequences.

[0017] In another aspect of one embodiment of the method of the present invention, the method comprises the steps of: step (A) introducing into a host cell a first nucleic acid sequence comprising a pre-mRNA trans-splicing molecule sequence or a nucleic acid sequence encoding said pre-mRNA trans-splicing molecule, wherein the first pre-mRNA trans-splicing molecule comprises, from 5' to 3', (a) a 5' portion of a nucleotide acid sequence of interest; (b) a donor splice site; (c) optionally a spacer sequence; (d) a first binding domain; and (e) optionally a termination sequence, preferably a polyA sequence; step (B) introducing into a host cell a second nucleic acid sequence comprising a pre-mRNA trans-splicing molecule sequence or a nucleic acid sequence encoding said pre-mRNA trans-splicing molecule, wherein the second pre-mRNA trans-splicing molecule comprises, from 5' to 3', (i) a second binding domain complementary to a first target domain of the first nucleic acid sequence; (ii) an acceptor splice region sequence comprising; (iia) a pyrimidine dinucleotide nucleotide sequence comprising; (iiab) wherein at least 60% of the nucleotides within these 5-25 nucleotides are pyrimidine bases, such as cytosine (C), thymine (T), and / or uracil (U); (iib) an acceptor splice site, (iiba) wherein the acceptor splice site is located 3' to the pyrimidine tract; and (iibb) wherein the acceptor splice site comprises a sequence 5' to 3' of NAGG, where N is A, C, T / U, or G; (iii) a 3' portion of a nucleotide sequence of interest, and (iv) a termination sequence, preferably a polyA sequence; step (C) cleaving the first nucleic acid sequence at the donor splice site sequence and cleaving the second nucleic acid sequence at the acceptor splice site; and step (D) ligating the first cleaved nucleic acid sequence comprising the 5' portion of the nucleotide sequence of interest to a second cleaved nucleic acid sequence comprising the 3' portion of the nucleotide sequence of interest, thereby obtaining a nucleic acid sequence of interest.

[0018] In yet another aspect, the present invention relates to an adeno-associated virus (AAV) vector comprising at least two inverted terminal repeats and comprising a nucleic acid sequence between the two inverted terminal repeats, wherein the nucleic acid sequence comprises from 5' to 3': (i) a promoter; (ii) a binding domain; (iii) optionally a spacer sequence; (iv) an acceptor splice region sequence comprising: (a) a pyrimidine tract comprising: (aa) 5 to 25 nucleotides; (ab) wherein at least 60% of the nucleotides within these 5 to 25 nucleotides are pyrimidine bases, e.g., cytosine (C), thymine (T), and / or uracil (U); (b) an acceptable splice site, wherein the acceptable splice site is located 3' to the pyrimidine tract; and wherein the acceptable splice site comprises a sequence 5' to 3' of NAGG, where N is A, C, T / U, or G; (v) a nucleotide sequence of interest or a portion thereof; and (vi) optionally a polyA sequence.The AAV vector may also be part of an AAV vector system comprising: (I) a first AAV vector comprising at least two inverted terminal repeats, the nucleic acid sequence between the two inverted terminal repeats comprising from 5' to 3': (a) a promoter; (b) a nucleotide sequence encoding an N-terminal portion of a polypeptide of interest; (c) a donor splice site; (d) optionally a spacer sequence; (e) a first binding domain; and (f) optionally a termination sequence, preferably a polyA sequence; (II) a second AAV vector comprising at least two inverted terminal repeats, the nucleic acid sequence between the two inverted terminal repeats comprising from 5' to 3': (i) a promoter; (ii) a second binding domain, complementary to the first binding domain of the first AAV vector; (ii) optionally a spacer sequence; (iii) an acceptor splice region sequence comprising: (a) a pyrimidine tract comprising: (aa) 5-25 nucleotides; (ab) wherein at least 60% of the nucleotides within these 5-25 nucleotides are pyrimidine bases, e.g., cytosine (C), thymine (T), and / or uracil (U); (b) an acceptable splice site, wherein said acceptable splice site is located 3' to the pyrimidine tract; and wherein said acceptable splice site comprises, 5' to 3', the sequence NAGG, where N is A, C, T / U, or G; (iv) a nucleotide sequence encoding a C-terminal portion of a polypeptide of interest; (iva) wherein the C-terminal portion of the polypeptide of interest and the N-terminal portion of the polypeptide of interest reconstitute a polypeptide of interest; and (v) a termination sequence, preferably a polyA sequence. In one embodiment, the polypeptide is a full-length polypeptide, and the first AAV vector comprises an N-terminal portion of the full-length polypeptide of interest and the second AAV vector comprises a C-terminal portion of the polypeptide of interest.

[0019] Also provided are nucleic acid sequences comprising: an acceptable splice region sequence comprising: (ia) a pyrimidine tract comprising: (aa) 5-25 nucleotides; (ab) where at least 60% of the nucleotides within these 5-25 nucleotides are pyrimidine bases, e.g., cytosine (C), thymine (T), and / or uracil (U); (ib) an acceptable splice site, ba) where the acceptable splice site is located 3' to the pyrimidine tract; and bb) where the acceptable splice site comprises a sequence 5' to 3' of NAGG, where N is A, C, T / U, or G; and (ii) a nucleotide sequence of interest, where the nucleotide sequence of interest (iia) is located 3' or 5' to the acceptable splice region.

[0020] In a particular embodiment according to the invention, the acceptable splice region described herein comprises a pyrimidine tract, wherein 5-25 nucleotides of the pyrimidine tract comprise a sequence encoded by the sequence TTTTTT or TCTTTT. The sequence between the last pyrimidine of the pyrimidine tract and the acceptable splice site is less than 10 bases, preferably less than 5 bases, more preferably less than 3 bases. Furthermore, the acceptable splice site preferably has a sequence of CAGG. The acceptable splice region further comprises 7 nucleotides 5' of the pyrimidine tract having at least 4 nucleotides of the sequence CAACGAG, wherein the first 5' nucleotide is C. In a preferred embodiment, the splice acceptor region comprises or consists of the sequence of SEQ ID NO:3, or 4, or is encoded by a nucleic acid sequence comprising or consisting of the sequence of SEQ ID NO:3, or 4. [Brief description of the drawings]

[0021] [Figure 1-1]mRNA splicing of the RHO minigene in HEK293 cells and transduced photoreceptors. A, Schematic diagram of the RHO gene to scale. Boxes represent exons, and the thick lines between them represent introns, showing the respective start and stop codons. The asterisk represents the c.620T>G mutation in exon 3. B and C, The rhodopsin minigene, including the coding portion of the exon and the adjacent intron, was driven by either the CMV promoter for expression in HEK293 cells (B) or the human rhodopsin (hRHO) promoter for expression in photoreceptors (C). To allow packaging of the RHO minigene into AAV vectors, intron 1 was truncated as shown. For protein visualization and detection, RHO was fused at the C-terminus to citrine (in B) or a myc tag (in C). Primers indicated as arrows in B and C were used for specific detection of wild-type (WT) or mutant splicing products derived from HEK293 cells (D) or transduced photoreceptors (E). [Figure 1-2] D, RT-PCR analysis from HEK293 cells transiently transfected with WT and mutant RHO minigenes. E, Left, Schematic of subretinal RHO minigene delivery into postnatal day 14 (P14) mouse retina. Right, RT-PCR from injected mouse retina containing the respective WT or mutant RHO minigene. RT-PCR was performed 4 weeks after injection. All experiments were repeated once. [Figure 2-1]Identification of the most efficient ASS_620 sequence. A, Organization of a functional single element of ASS. Branch points (BPs) are underlined to represent the consensus sequence of human branch points, and the branch point nucleoside adenosine is highlighted in bold. Polypyrimidine tracts (Poly-C / T) are represented as empty boxes, and ASSs spanning intron-exon boundaries are represented as grey filled boxes. B, Exon-intron organization of wild-type (WT, top panel) and c.620T>G mutant (bottom panel) human RHO genes and a close-up of the DNA sequence of exon 3 of the human RHO gene. Intron sequences are shown in lower case and exon sequences in upper case (bold). Single ASS elements are highlighted according to the scheme shown in A, BP sequences are underlined (bold), Poly-C / T sequences are marked with empty boxes, and ASS sequences are marked with filled grey boxes. The WT ATG sequence converted to an AGG sequence in the c.620T>G variant is marked with a thin underline. Note that the c.620T>G variant generates a new canonical ASS sequence. The other two elements required for a functional ASS (BP and Poly-C / T) are already present in WT RHO exon 3. [Figure 2-2] C, Structure of the human ribosomal protein 27 (RPS27) minigene used for splicing experiments. The primers used for RT-PCR shown in (D) are indicated by arrows. Individual sequences containing potential elements of ASS_620, designated RHO_E3a-g below, were inserted into exon 3 of the RPS27 gene as indicated. RHO_E3a is the WT RHO sequence lacking a functional ASS. Potential BP, Poly-C / T or ASS sequences are indicated as above. [Figure 2-3] D, RT-PCR from HEK293 cells transfected with a single chimeric RPS27 minigene as indicated. CS, correctly spliced ​​RPS27 minigene (i.e., using native exon 3 ASS). AS, above spliced ​​RPS27 minigene (i.e., using ASS_620). [Figure 3-1]ASS_620 is a strong acceptor splice site. A, In silico acceptor splice site (ASS) strength prediction using two commonly available prediction tools, NNSplice (0.9 version; January 1997) (http: / / www.fruitfly.org / seq_tools / splice.html) or human splicing finder (version 3.1) (HSF, http: / / www.umd.be / HSF3 / ). B, Organization and sequence of a single element of a 26 bp sequence called vgASS_620 that contains ASS, PolyC / T, and seven additional nucleotides 5' of PolyC / T. C, Schematic diagram of the various minigenes used to determine the strength of vgASS_620. The position of the vgASS_620 insertion within a single minigene is indicated by an asterisk and a dashed line. The binding positions of the primers used for RT-PCR shown in D are indicated by arrows. [Figure 3-2] D, RT-PCR from HEK293 cells transfected with each minigene containing only the native (nat) acceptor splice site or both the native acceptor splice site and vgASS_620 (620). All bands were confirmed by sequencing. [Figure 4-1]Cerulean reconstitution assay testing different binding domains. A, Schematic of RHO intron 2 sequences used as binding domains for reconstitution of cerulean by trans-splicing of mRNA. The entire RHO intron 2 sequence (a), and different 5' (b, d, f, h) and 3' (c, e, g, i) parts, or a small 5' part fused to a small 3' part of intron 2 (h+i) were tested, with the results shown in B–G. B and C, Control qRT-PCR (n=3) from transfected HEK293 cells to compare the mRNA levels of single constructs containing the different binding domains shown in A. Delta CT (ΔCT) values ​​related to the housekeeper aminolevulinic acid synthase (ALAS) are shown. Primer binding positions (p1+p2 in B and p3+p4 in C) are represented in D. Statistical analysis (n=3 for each transfection) was performed by one-way ANOVA followed by Tukey's test for multiple comparisons. [Figure 4-2]D, Principle of the cerulean reconstitution assay with the example of the h+i binding domain, donor splice site (DSS) and acceptor splice site ASS_620. Cerulean was split at nucleotide position 154 downstream of the start codon of the full-length cerulean sequence (c1), as indicated by the dashed line. In the control construct (c2), an artificial intron was introduced at the indicated position 154 of the cerulean coding sequence to create two artificial cerulean exons. The antibody used for Western blotting (α-cerNT) binds to the N-terminal half of cerulean. The control construct (c2) served as a reference for quantification in confocal imaging and Western blotting experiments. The various sequences shown in (A) were tested as binding domains using the dual vector approach. A scheme of the dual vector approach for cerulean reconstitution is exemplarily shown for the h+i binding domain combined with vgASS_620 (c3) or the native RHO exon 3 (RHO_E3) acceptor splice site (c4). E, Cerulean reconstitution efficiency (CRE) of the different binding domains was calculated from Western blot band intensities obtained from three independent transfections. For quantification, band intensities were first normalized to an internal tubulin control. Normalized values ​​are given as a percentage of the reference construct intensity (c2). F, Representative confocal imaging of live HEK293 cells transfected with c1, c2, c3, or c4 constructs driven by the CMV promoter. Cerulean-specific laser and filter settings were used for imaging. Scale bar, 50 μm. G, Representative Western blot from protein lysates of transfected HEK293 cells (φ, non-transfected cells; IB, immunoblotting; α-Tub, beta-tubulin-specific antibody). [Figure 5-1]Sequence of the first dual AAV containing the 5' coding sequence of ABCA4 (SEQ ID NO:19): The 5' and 3' ITR sequences are highlighted in grey without underlining (NNN). The spacer sequence is written in lower case letters. The CMV promoter is highlighted in grey and underlined using a solid line (NNN). The 5' coding sequence of the ABCA4 protein is underlined with a dotted line (NNN). The DSS is italicized and underlined (Nnn, capital letters represent coding sequences, lower case letters represent non-coding sequences). The binding domain is highlighted in grey and underlined with a broken line (NNN). [Figure 5-2] Sequence of the first dual AAV containing the 5' coding sequence of ABCA4 (SEQ ID NO:19): The 5' and 3' ITR sequences are highlighted in grey without underlining (NNN). The spacer sequence is written in lower case letters. The CMV promoter is highlighted in grey and underlined using a solid line (NNN). The 5' coding sequence of the ABCA4 protein is underlined with a dotted line (NNN). The DSS is italicized and underlined (Nnn, capital letters represent coding sequences, lower case letters represent non-coding sequences). The binding domain is highlighted in grey and underlined with a broken line (NNN). [Figure 6-1] Sequence of the second dual AAV containing the 3' coding sequence of ABCA4 (SEQ ID NO:20): The 5' and 3' ITR sequences are highlighted in grey without underlining (NNN). The spacer sequences are written in lower case. The CMV promoter is highlighted in grey and underlined using a solid line (NNN). The binding domain is highlighted in grey and underlined with a broken line (NNN). The acceptor splice site is italicized and underlined (NNN). The 3' coding sequence of the ABCA4 protein is underlined using a dotted line (NNN). The polyA sequence is highlighted in dark grey and underlined with a wavy line (NNN). [Figure 6-2]Sequence of the second dual AAV containing the 3' coding sequence of ABCA4 (SEQ ID NO:20): The 5' and 3' ITR sequences are highlighted in grey without underlining (NNN). The spacer sequences are written in lower case. The CMV promoter is highlighted in grey and underlined using a solid line (NNN). The binding domain is highlighted in grey and underlined with a broken line (NNN). The acceptor splice site is italicized and underlined (NNN). The 3' coding sequence of the ABCA4 protein is underlined using a dotted line (NNN). The polyA sequence is highlighted in dark grey and underlined with a wavy line (NNN). [Figure 7-1] Sequence of the first dual AAV containing the 5' coding sequence of ABCA4 and the ABCA4 promoter (SEQ ID NO:21): The 5' and 3' ITR sequences are highlighted in grey without underlining (NNN). The spacer sequences are written in lower case letters. The ABCA4 promoter is highlighted in grey and underlined using a solid line (NNN). The 5' coding sequence of the ABCA4 protein is underlined using a dotted line (NNN). The DSS is italicized and underlined (Nnn, capital letters represent coding sequences and lower case letters represent non-coding sequences). The binding domain is underlined using a broken line (NNN). [Figure 7-2] Sequence of the first dual AAV containing the 5' coding sequence of ABCA4 and the ABCA4 promoter (SEQ ID NO:21): The 5' and 3' ITR sequences are highlighted in grey without underlining (NNN). The spacer sequences are written in lower case letters. The ABCA4 promoter is highlighted in grey and underlined using a solid line (NNN). The 5' coding sequence of the ABCA4 protein is underlined using a dotted line (NNN). The DSS is italicized and underlined (Nnn, capital letters represent coding sequences and lower case letters represent non-coding sequences). The binding domain is underlined using a broken line (NNN). [Figure 8-1]Sequence of the second dual AAV containing the 3' coding sequence of ABCA4 and the ABCA4 promoter (SEQ ID NO:22): The 5' and 3' ITR sequences are highlighted in grey without underlining (NNN). The spacer sequences are written in lower case. The ABCA4 promoter is highlighted in grey and underlined using a solid line (NNN). The binding domain is underlined using a broken line (NNN). The ASS is italicized and underlined (NNN). The 3' coding sequence of the ABCA4 protein is highlighted in grey and underlined using a dotted line (NNN). The polyA sequence is highlighted in dark grey and underlined with a wavy line (NNN). [Figure 8-2] Sequence of the second dual AAV containing the 3' coding sequence of ABCA4 and the ABCA4 promoter (SEQ ID NO:22): The 5' and 3' ITR sequences are highlighted in grey without underlining (NNN). The spacer sequences are written in lower case. The ABCA4 promoter is highlighted in grey and underlined using a solid line (NNN). The binding domain is underlined using a broken line (NNN). The ASS is italicized and underlined (NNN). The 3' coding sequence of the ABCA4 protein is highlighted in grey and underlined using a dotted line (NNN). The polyA sequence is highlighted in dark grey and underlined with a wavy line (NNN). [Figure 9]Influence of acceptor splice sites and binding domains on Cerulean reconstitution efficiency. A, In this experiment, binding domains (BDs) and acceptor splice sites (ASSs) were tested. All BD sequences were derived from the human RHO gene. BDs and ASSs expected to have high efficiency are shown in bold and italics. B, Confocal live images of HEK293 cells transfected with constructs containing different combinations of the three BDs and the three ASSs shown in A. The intensities of the respective BDs and ASSs are shown. Scale bar, 50 μm. C, Upper panel, RT-PCR of different BD and ASS combinations. Lower panel, Western blot of different BD and ASS combinations. GAPDH and beta-tubulin served as loading references. D, Quantification of reconstitution efficiency by radiometric analysis of Cerulean protein bands associated with cis-ctrl (n = 3–8). All protein bands were normalized to beta-tubulin before quantification. [Figure 10] mRNA trans-splicing rAAV dual vector approach in vivo. A, 5' and 3' vector constructs used for in vivo expression. As expression references, Citrine and mCherry sequences were fused 5' of the Cerulean 5' coding sequence (CDS) and 3' of the Cerulean 3' CDS, respectively. B, Representative confocal images of retinal sections 2 weeks post-injection expressing constructs containing BD_h+i. Expression of the fluorophores is present in the retinal pigment epithelium (RPE). ONL, outer nuclear layer. Scale bar, 20 μm. C, Confocal images of retinal pigment epithelial cells before (top panel) and after (bottom panel) selective photobleaching of the Citrine and mCherry fluorophores using a 514 nm laser. Scale bar, 2 μm. [Figure 11]Identification of a suitable BD derived from the lacZ gene. A, Binding domains (BDs) obtained from the bacterial lacZ gene and modified to have no detectable homology to the human genome. B, Confocal live images of HEK293 cells transiently co-transfected with constructs containing the BDs shown in A. Scale bar, 50 µm. C, Western blots obtained from transfected HEK293 cell lysates. D, Quantification of Cerulean reconstitution efficiency based on ratiometric analysis of Western blot band intensities (n = 3–8). BD_g efficiency served as a measure of the best reconstitution obtained so far (see Figure 9). [Figure 12] Removal of regulatory elements and the impact on mRNA splicing efficiency. A, Confocal live images of HEK293 cells transiently co-transfected with Cerulean constructs containing BD_k.5'polyAdel, a 5' vector without a polyA signal.3'promdel, a 3' vector without a promoter sequence, as indicated. Scale bar, 50 μm. B, Western blots from transfected HEK293 cells shown in A. Beta-tubulin served as a loading reference. [Figure 13] Reconstitution of SpCas9-VPR. A, Plasmids (5' and 3' vectors) used for mRNA trans-splicing of SpCas9-VPR in HEK293 cells. As a positive control, a plasmid containing the full-length (FL) SpCas9-VPR CDS (FL vector) was used. A primer pair spanning the junction was used for RT-PCR (black arrows). B, RT-PCR from HEK293 cells co-transfected with 5' and 3' vectors containing BD_k (n = 3) or FL vector (n = 3). C, Representative sequencing results of reconstituted SpCas9-VPR products. D, Western blots of protein lysates from the respective transfections. φ, untransfected cells. Beta-tubulin served as a loading control. [Figure 14]Reconstitution of ABCA4. A, Plasmids (5' and 3' vectors) used for mRNA trans-splicing of ABCA4 in HEK293 cells. In addition, six short introns from six different genes were integrated within the CDS of ABCA4; three in the 5' CDS (5' vector with intron, 5'wi) and three in the 3' CDS (3' vector with intron, 3'wi). Primer pairs spanning the junctions were used for RT-PCR (black arrows). myc, myc tag. B, RT-PCR at various cycle numbers from HEK293 cells co-transfected with the respective constructs as indicated. GAPDH served as a loading control. φ, untransfected cells. C, Representative sequencing results of reconstituted ABCA4. [Figure 15-1] rAAV dual vector mRNA trans-splicing of ABCA4 in vivo. A, 5'-vector and 3'-vector constructs containing BD_k used for in vivo expression. hRho, human rhodopsin promoter. B, RT-PCR performed at different cycle numbers from retinal lysates of C54Bl / 6J wild-type mice co-transduced with the respective constructs as indicated. Non-injected wild-type (WT) mice were used as negative controls. NN / GL, capsid variants of AAV2. -RT, negative control without reverse transcriptase. GAPDH was used as a loading control. [Figure 15-2] C, Representative sequencing results of reconstituted ABCA4. D, Preliminary results of qRT-PCR performed from the same retina shown in B. ABCA4 expression obtained upon co-transduction was normalized to non-injected WT retina. E, Western blot from retinal lysates of the same mice shown in B-D. Human ABCA4 protein was detected with an anti-myc antibody and is indicated by an arrow. Beta-tubulin was used as a loading control. [Figure 16]Sequence of the minigene encoding the 5' coding sequence of Cerulean and the binding domain BD_g (SEQ ID NO:35). The CMV promoter is highlighted in grey. The 5' coding sequence of the Cerulean protein (bold) is highlighted in grey and underlined using a thick solid line (NNN). The DSS is underlined using a wavy line (NNN). The binding domain is marked in italics and underlined using a dotted line (NNN) and the polyadenylation signal is underlined using a double line (NNN). [Figure 17] Sequence of the minigene encoding the 3' coding sequence and binding domain BD_g of Cerulean (SEQ ID NO:36). The CMV promoter is highlighted in grey. The binding domain is marked in italics and underlined using a dotted line (NNN). The ASS is underlined using a wavy line and bold (NNN). The 3' coding sequence of the Cerulean protein is highlighted in grey and underlined using a thick solid line (NNN) and the polyadenylation signal is underlined using a double line (NNN). [Figure 18-1] Sequence of the minigene encoding the 5' coding sequence of SpCas9-VPR and the binding domain BD_k (SEQ ID NO: 37). The CMV promoter is highlighted in grey. The 5' coding sequence of the SpCas9-VPR protein is highlighted in bold and underlined using a thick solid line (NNN). The DSS is underlined using a wavy line (NNN). The binding domain is marked in italics and underlined using a dotted line (NNN), the polyadenylation signal is underlined using a double line (NNN). [Figure 18-2] Sequence of the minigene encoding the 5' coding sequence of SpCas9-VPR and the binding domain BD_k (SEQ ID NO: 37). The CMV promoter is highlighted in grey. The 5' coding sequence of the SpCas9-VPR protein is highlighted in bold and underlined using a thick solid line (NNN). The DSS is underlined using a wavy line (NNN). The binding domain is marked in italics and underlined using a dotted line (NNN), the polyadenylation signal is underlined using a double line (NNN). [Figure 19-1] Sequence of the minigene encoding the 3' coding sequence of SpCas9-VPR and the binding domain BD_k (SEQ ID NO: 38). The CMV promoter is highlighted in grey. The binding domain is marked in italics and underlined using a dotted line (NNN). The ASS is underlined using a wavy line and bold (NNN). The 3' coding sequence of the SpCas9-VPR protein is highlighted in grey and underlined using a thick solid line (NNN), and the polyadenylation signal is underlined using a double line (NNN). [Figure 19-2] Sequence of the minigene encoding the 3' coding sequence of SpCas9-VPR and the binding domain BD_k (SEQ ID NO: 38). The CMV promoter is highlighted in grey. The binding domain is marked in italics and underlined using a dotted line (NNN). The ASS is underlined using a wavy line and bold (NNN). The 3' coding sequence of the SpCas9-VPR protein is highlighted in grey and underlined using a thick solid line (NNN), and the polyadenylation signal is underlined using a double line (NNN). [Figure 20-1] Sequence of the first dual AAV containing the 5' coding sequence of ABCA4 including the intron (SEQ ID NO: 39): The 5' and 3' ITR sequences are highlighted in grey without underlining (NNN). The spacer sequence is written in lower case letters. The CMV promoter is highlighted in grey and underlined using a solid line (NNN). The 5' coding sequence of the ABCA4 protein is underlined with a dotted line interrupted by an intron shown with a wavy underline (NNNNNNNNN). The DSS is italicized and underlined (Nnn, capital letters represent coding sequences, lower case letters represent non-coding sequences). The binding domain is highlighted in grey and underlined with a broken line (NNN). The polyA sequence is highlighted in dark grey and underlined with a wavy underline (NNN). [Figure 20-2]Sequence of the first dual AAV containing the 5' coding sequence of ABCA4 including the intron (SEQ ID NO: 39): The 5' and 3' ITR sequences are highlighted in grey without underlining (NNN). The spacer sequence is written in lower case letters. The CMV promoter is highlighted in grey and underlined using a solid line (NNN). The 5' coding sequence of the ABCA4 protein is underlined with a dotted line interrupted by an intron shown with a wavy underline (NNNNNNNNN). The DSS is italicized and underlined (Nnn, capital letters represent coding sequences, lower case letters represent non-coding sequences). The binding domain is highlighted in grey and underlined with a broken line (NNN). The polyA sequence is highlighted in dark grey and underlined with a wavy underline (NNN). [Figure 20-3] Sequence of the first dual AAV containing the 5' coding sequence of ABCA4 including the intron (SEQ ID NO: 39): The 5' and 3' ITR sequences are highlighted in grey without underlining (NNN). The spacer sequence is written in lower case letters. The CMV promoter is highlighted in grey and underlined using a solid line (NNN). The 5' coding sequence of the ABCA4 protein is underlined with a dotted line interrupted by an intron shown with a wavy underline (NNNNNNNNN). The DSS is italicized and underlined (Nnn, capital letters represent coding sequences, lower case letters represent non-coding sequences). The binding domain is highlighted in grey and underlined with a broken line (NNN). The polyA sequence is highlighted in dark grey and underlined with a wavy underline (NNN). [Figure 21-1]Sequence of the second dual AAV containing the 3' coding sequence of ABCA4 (SEQ ID NO: 40): The 5' and 3' ITR sequences are highlighted in grey without underlining (NNN). The spacer sequences are written in lower case. The CMV promoter is highlighted in grey and underlined using a solid line (NNN). The binding domain is highlighted in grey and underlined with a broken line (NNN). The acceptor splice site is italicized and underlined (NNN). The 3' coding sequence of the ABCA4 protein is underlined using a dotted line interrupted by an intron shown with a wavy underline (NNNNNNNNN). The polyA sequence is highlighted in dark grey and underlined with a wavy underline (NNN). [Figure 21-2] Sequence of the second dual AAV containing the 3' coding sequence of ABCA4 (SEQ ID NO: 40): The 5' and 3' ITR sequences are highlighted in grey without underlining (NNN). The spacer sequences are written in lower case. The CMV promoter is highlighted in grey and underlined using a solid line (NNN). The binding domain is highlighted in grey and underlined with a broken line (NNN). The acceptor splice site is italicized and underlined (NNN). The 3' coding sequence of the ABCA4 protein is underlined using a dotted line interrupted by an intron shown with a wavy underline (NNNNNNNNN). The polyA sequence is highlighted in dark grey and underlined with a wavy underline (NNN). [Figure 21-3]Sequence of the second dual AAV containing the 3' coding sequence of ABCA4 (SEQ ID NO: 40): The 5' and 3' ITR sequences are highlighted in grey without underlining (NNN). The spacer sequences are written in lower case. The CMV promoter is highlighted in grey and underlined using a solid line (NNN). The binding domain is highlighted in grey and underlined with a broken line (NNN). The acceptor splice site is italicized and underlined (NNN). The 3' coding sequence of the ABCA4 protein is underlined using a dotted line interrupted by an intron shown with a wavy underline (NNNNNNNNN). The polyA sequence is highlighted in dark grey and underlined with a wavy underline (NNN). DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0022] Detailed Description of the Invention The present inventors have found a novel acceptor splice region, as described in the Examples. This acceptor splice region is highly efficient and can be used in various applications where splicing is utilized. Thus, the acceptor splice region of the present invention can be used for trans-splicing using pre-mRNA trans-splicing molecules, cis-splicing, SMaRT technology (spliceosome-mediated RNA trans-splicing), inter alia, for the reconstitution of split AAV vectors, AAV vector systems including two AAV vectors and / or trap vectors. However, the acceptor splice region described herein can also be used for further applications where splicing is the objective.

[0023] The present invention relates to (i) an acceptor splice region comprising: (ia) a pyrimidine tract comprising: (aa) 5–25 nucleotides; (ab) wherein at least 60% of the nucleotides within these 5 to 25 nucleotides are pyrimidine bases, e.g., cytosine (C), thymine (T), and / or uracil (U); (ib) acceptor splice site; ba) wherein the acceptor splice site is located 3' to the pyrimidine tract; and bb) wherein the acceptor splice site comprises a sequence 5' to 3' of NAGG, where N is A, C, T / U, or G; (ii) a nucleotide sequence of interest; wherein the nucleotide sequence of interest is (iia) is located 3' or 5' to the acceptor splice region, preferably 3' to the acceptor splice region; wherein the nucleic acid sequence according to the invention is preferably cleaved at the acceptor splice region, thereby separating the nucleotide sequence of interest from the acceptor splice region sequence. The invention further relates to the use of the above-mentioned nucleic acid sequences for cleaving an acceptor splice region, thereby separating a nucleotide sequence of interest from the acceptor splice region sequence.

[0024] The terms "nucleic acid molecule", "nucleic acid sequence" or "nucleotide sequence" are used synonymously herein and encompass any nucleic acid molecule having a nucleotide sequence including purine and pyrimidine bases contained in said nucleic acid molecule / sequence, whereby said bases represent the primary structure of the nucleic acid molecule. The nucleic acid sequence may include DNA, cDNA, genomic DNA, RNA, sense and antisense strands. The RNA may be, for example, pre-mRNA, mRNA, tRNA or rRNA. The polynucleotides of the invention may be composed of any polyribonucleotide or polydeoxyribonucleotide and may be unmodified RNA or DNA or modified RNA or DNA. The skilled artisan will understand that thymine (T) in polydeoxynucleotides is transcribed to uracil (U) in polyribonucleotides. The sequences referred to herein are typically provided as DNA sequences that can be transcribed (before a splicing event) into the corresponding RNA sequence. Thus, a nucleic acid sequence that includes a T also discloses the corresponding (transcribed) RNA sequence, preferably a pre-mRNA sequence that includes a U.

[0025] Various modifications can be made to DNA and RNA. Thus, the term "nucleic acid molecule" or "nucleotide" can include chemically, enzymatically, or metabolically modified forms. "Modified" bases / nucleotides include, for example, tritylated bases and unusual bases such as inosine.

[0026] The acceptable splice region sequence of the present invention comprises two features: an acceptable splice site and a pyrimidine tract. The pyrimidine tract comprises 5-25 nucleotides, with at least 60% of the nucleotides within these (total) 5-25 nucleotides being pyrimidine bases such as cytosine (C), thymine (T) and / or uracil (U). It is also envisaged that the pyrimidine tract comprises 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24 or 25 nucleotides, preferably 10-18 nucleotides, more preferably 12-16 nucleotides. Additionally or alternatively, at least 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 97%, 98%, 99%, or 100% of the nucleotides within these (total) 5-25 nucleotides are pyrimidine bases, such as cytosine (C), thymine (T), and / or uracil (U). For example, the pyrimidine tract may include 6 nucleotides. It is further contemplated that the pyrimidine tract comprises or consists of the sequence TTTTTT. It is further contemplated that the pyrimidine tract comprises or consists of the sequence TCTTTT. It is further contemplated that the pyrimidine tract comprises or consists of the sequence TTTTTTGTCATTT (SEQ ID NO:11). It is further contemplated that the pyrimidine tract comprises or consists of the sequence TCTTTTGTCATCTA (SEQ ID NO:12). It is further contemplated that the pyrimidine tract is preceded by 7 nucleotides comprising or consisting of the sequence CAACGAGTCTTTTGTCATCTA (SEQ ID NO: 13). The pyrimidine tract may also comprise or consist of the sequence TCTTTTGTCATCT (SEQ ID NO: 1). It is also contemplated that the pyrimidine tract described herein is at least 70%, 75%, 80%, 85%, 90%, 95%, 97%, 98%, 99%, or 100% identical to the sequence of SEQ ID NO: 1, 11, 12, or 13, or the sequence TTTTTT, or TCTTTT. In one embodiment, 5 to 25 nucleotides of the pyrimidine tract comprise the sequence TTTTTT, or TCTTTT.Preferably, the sequence between the last pyrimidine of the pyrimidine tract and the acceptor splice site is less than 10 bases, preferably less than 5 bases, more preferably less than 3 bases. The term "pyrimidine tract" is used interchangeably herein with "polypyrimidine tract" and is abbreviated as PPT or PolyC / T. The "last pyrimidine of the pyrimidine tract" refers to the 3'-most pyrimidine of the 5-25 nucleotides of the pyrimidine tract.

[0027] According to the present invention, the term "identical" or "percent identity" in the context of two or more nucleic acid molecules refers to two or more sequences or subsequences that are identical or have a certain percentage of identical nucleotides (e.g., at least 95%, 96%, 97%, 98%, or 99% identity) over a comparison window or over a designated region measured using sequence comparison algorithms known in the art, or over a designated region measured by manual alignment and visual inspection, when compared and aligned for maximum correspondence. For example, sequences having 80% to 95% or more sequence identity are considered to be substantially identical. Such definition also applies to the complement of a test sequence. A person skilled in the art will know how to determine percent identity between / among sequences, as known in the art, using, for example, the CLUSTALW computer program (Thompson Nucl. Acids Res. 2 (1994), 4673-4680), or algorithms based on FASTDB (Brutlag Comp. App. Biosci. 6 (1990), 237-245).

[0028] BLAST and BLAST2.0 algorithms (Altschul Nucl. Acids Res. 25 (1977), 3389-3402) are also available to those skilled in the art. For nucleic acid sequences, the BLASTN program uses as default a word size (W) of 28, an expectation threshold (E) of 10, a match / mismatch score of 1, -2, a gap cost linear, and a comparison of both strands. For amino acid sequences, the BLASTP program uses as default a word size (W) of 6, an expectation threshold (E) of 10, and gap costs of Existence: 11 and Extension: 1. In addition, the BLOSUM62 scoring matrix can be used (Henikoff Proc. Natl. Acad. Sci., USA, 89, (1989), 10915).

[0029] For example, BLAST2.0 stands for Basic Local Alignment Search Tool (Altschul, Nucl. Acids Res. 25 (1997), 3389-3402; Altschul, J. Mol. Evol. 36 (1993), 290-300; Altschul, J. Mol. Biol. 215 (1990), 403-410) and can be used to search for local sequence alignments.

[0030] In addition to the pyrimidine tract, the acceptor splice region sequence of the present invention comprises an acceptor splice site, as outlined herein. The term "acceptor splice site" as used herein has the meaning known to those skilled in the art and described in Alberts B, Johnson A, Lewis J, et al. (2002) "Molecular Biology of the Cell. 4th edition." New York: Garland Science, inter alia, under the heading "DNA to RNA". The acceptor splice site of the present invention comprises or consists of the nucleotides NAGG, preferably CAGG or CAGGT (or CAGG or CAGGU in RNA). The acceptor splice site is usually located within a sequence defined as an "acceptor splice region". The term "acceptor splice site" (abbreviated ASS) as used herein refers to a consensus acceptor splice sequence, also called "splice acceptor site" (abbreviated SAS). In the context of the present invention, the terms "acceptable splice region" and "acceptable splice site" are used as separate terms to distinguish this region from the consensus acceptable splice site, which the acceptable splice region contains. In the literature, the acceptable splice region and the acceptable splice site are sometimes referred to as ASS or SAS.

[0031] The acceptor splice site is called the acceptor splice site because at this site the nucleic acid is cleaved by the so-called spliceosome or an artificial variant thereof. The spliceosome and how it functions are also known to those skilled in the art. The structure of the human spliceosome is described, for example, in Zhang et al. (2017) "An atomic structure of the human spliceosome" Cell 169, 918-926. The artificial spliceosome, for example, contains an enzyme that mediates the same cleavage at the splice site as the spliceosome. The artificial spliceosome may comprise a DNA enzyme as described in Coppins and Silvermann (2005) “Mimicking the First Step of RNA Splicing: An Artificial DNA Enzyme Can Synthesize Branched RNA Using an Oligonucleotide Leaving Group as a 5'-Exon Analogue” Biochemistry, 44 (41), pp 13439-13446 and Muller (2017) “Design and Experimental Evolution of trans-Splicing Group I Intron Ribozymes.” Molecules. 22(1). pii: E75. doi: 10.3390 / molecules22010075. The acceptor splice site described herein may have a sequence of NAGG, where N may be any nucleotide. For example, N is a nucleotide selected from A, C, T, G, or U. The acceptor splice site described herein may have a sequence of CAGG. The acceptor splice site described herein may also have the sequence CAGGT (or CAGGU in the case of RNA).

[0032] Cleavage at the acceptor splice site of the present invention results in two fragments containing NAG at the NAGG and / or NAGGT splice site, or CAG at the CAGG splice site and / or CAGGT splice site. Other fragments containing the final G at the NAGG or CAGG splice site, or the final GT at the NAGGT or CAGGT splice site. Thus, in the uses, nucleic acid sequences, AAV vectors, AAV vector systems and methods described herein, it is envisaged that the final (desired) nucleotide sequence (after cleavage / splicing) contains the final G at the NAGG or CAGG sequence, or the GT at the NAGGT or CAGGT sequence. The skilled artisan will understand that the splicing referred to herein occurs at the RNA level, and thus the final (desired) RNA nucleotide sequence (after cleavage / splicing) contains the final G at the NAGG or CAGG sequence, or the GU at the CAGGU sequence of the pre-mRNA nucleotide sequence. However, in the present methods and uses, it is also contemplated that the final nucleic acid sequence (of interest) comprises the NAGG splice site sequence NAG, and / or the CAGG or CAGGT splice site sequence CAG. It is therefore envisaged that cleavage at the acceptor splice site comprises cleavage between the NAGG, CAGG and / or CAGGT (or CAGGU) acceptor splice site sequences G and G, thereby separating the (interest) nucleotide sequence from the intron acceptor splice region sequence.

[0033] The acceptor splice region of the present invention can be used for applications involving, among others, trans-splicing as well as cis-splicing or other splicing. For example, splicing occurrence can be measured as described in Berger et al. (2016) "mRNA trans-splicing in gene therapy for genetic diseases" WIREs RNA 7: 487-498. For example, spliced ​​nucleotide sequences can be quantified by end-point quantitative RT-PCR using specific primers and probes to identify spliced ​​products of interest. Further measurement of splicing can be performed as described in the examples herein. For example, splicing can be measured using marker genes described in Orengo et al. (2006) "A bichromatic fluorescent reporter for cell-based screens of alternative splicing." Nucleic Acids Research 34(22):e148. Restriction enzyme analysis can be used to detect additional splicing, as described in Berger et al., (2016) “Repair of rhodopsin mRNA by spliceosome-mediated RNA trans-splicing: a new approach for autosomal dominant retinitis pigmentosa.”, Mol Ther.; 23(5):918-930.

[0034] The acceptable splice site, when present in the pre-mRNA, usually corresponds to the 3' end of the intron and the 5' end of the next exon. When an acceptable splice site, preferably an acceptable splice region, is artificially introduced, it is not necessary that the acceptable splice site is located at an intron-exon boundary, but may be located within an open reading frame, within an intron, or at the 5' end of a complete or partial open reading frame. It is also envisaged that the acceptable splice site is located at the 5' end of an intron. It is further envisaged that the acceptable splice site is located not in a pre-mRNA molecule, but in an artificial molecule, such as any nucleic acid molecule.

[0035] The acceptable splice site of the present invention is contained in an acceptable splice region. In particular, the acceptable splice site of the present invention is located 3' to the pyrimidine tract.

[0036] The acceptor splice region (or acceptor splice site module) may further comprise a branch point and / or a branch point sequence. In principle, the present invention contemplates any suitable branch point or branch point sequence. Exemplary branch points or branch point sequences are described, inter alia, in Gao et al. (2008) "Human branch point consensus sequence is yUnAy" Nucleic Acid Research, vol. 36, no.7, pp.2257-2267; Mercer et al. (2016) "Genome-wide discovery of human splicing branch points" Genome Research 25: 290-303.

[0037] For example, the branch point array is the array UACUA A Additionally or alternatively, the branch point sequence may comprise or consist of the sequence YNYUR, where Y is U or C and R is A or G. AAdditionally or alternatively, the branch point sequence may comprise or consist of the sequence YNCUR, where Y is U or C and R is A or G. A Additionally or alternatively, the branch point sequence may comprise or consist of the sequence CUR, where Y is U or C and R is A or G. A Additionally or alternatively, the branch point sequence may comprise or consist of the sequence YUV, where Y is U or C, R is A or G, and V is A, C or G. A Additionally or alternatively, the branch point sequence may comprise or consist of the sequence CUS, where Y is U or C and S is G or C. A Y. Additionally or alternatively, the branch point sequence may comprise or consist of the sequence CUG A Additionally or alternatively, the branch point sequence may comprise or consist of the sequence CUA A The branch point may be C / UUN- A -C / U, CAACGA or CUC A A or GUC A A may be present within the branch point sequence of A (wherein the branch point sequence is C / TTN- A -C / T, CAACGA or CTC A A or GTC A A is encoded by each of the DNA sequences. The underlined adenosine in all of the sequences represents the branch point nucleotide.

[0038] The acceptor splice region of the present invention may further comprise a branch point nucleotide sequence (c), the branch point nucleotide sequence being (ca) contains 1 to 15 nucleotides; (cb) includes a branch point nucleotide, preferably an adenosine (A); and (cc) Located 5' to the pyrimidine tract and the acceptor splice site.

[0039] It is also contemplated that the branchpoint sequence comprises about 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15 or more nucleotides. It is further contemplated that the branchpoint sequence comprises 6-8 nucleotides. It is also contemplated that the branchpoint sequence comprises 8 nucleotides. However, in particular, the nucleotides of the branchpoint sequence can be contiguous, and thus may comprise a total of more than 15 nucleotides, such as, for example, a total of 30 (15+15), 40, 50, 100, 200 or more nucleotides.

[0040] The branch point nucleotide may be any nucleotide. Thus, the branch point nucleotide may be any of A (adenosine), T (thymine), G (guanine), C (cytosine) and U (uracil). Preferably, the branch point nucleotide is A (adenosine). The branch point nucleotide is located within the branch point sequence.

[0041] The acceptor splice region may additionally or alternatively comprise an intronic splice enhancer. Such splice enhancers are known to those skilled in the art and are described, inter alia, in Wang et al. (2012) "Intronic splicing enhancers, cognate splicing factors and context-dependent regulation rules" Nature Structural & Molecular Biology, vol. 19, no. 10, pp. 1044-1053. Two examples are the splice enhancer sequences "AACG" (group F) and "CGAG" (group D). However, the term includes any suitable intronic splice region.

[0042] An exemplary intron splice enhancer has a sequence of TGGGGGGAGG (SEQ ID NO:2). Another exemplary intron splice enhancer has a sequence of GTAACGGC. It is further contemplated that the intron splice enhancer has a sequence of AACG. It is also contemplated that the acceptor splice region described herein has a nucleotide sequence that is at least 70%, 75%, 80%, 85%, 90%, 95%, 97%, 98%, 99%, or 100% identical to the sequence of SEQ ID NO:2, GTAACGGC, or AACG.

[0043] The acceptable splice region of the present invention preferably further comprises about 7 nucleotides (e.g., 5-12 nucleotides, preferably 6-10 nucleotides, more preferably 6-8 nucleotides) 5' to the polypyrimidine tract. In a preferred embodiment, the acceptable splice region comprises: (a) 7 nucleotides 5' to a pyrimidine tract having at least 4 nucleotides of the sequence CAACGAG, where the first 5' nucleotide is C; (b) 7 nucleotides 5' to a pyrimidine tract having at least 5 nucleotides of the sequence CAACGAG, where the first 5' nucleotide is C; (c) 7 nucleotides 5' to a pyrimidine tract having at least 6 nucleotides of the sequence CAACGAG, where the first 5' nucleotide is C; (d) 7 nucleotides 5' to a pyrimidine tract having at least 4 nucleotides of the sequence CAACGAG, where the first 5' nucleotide is CAA; (e) 7 nucleotides 5' to a pyrimidine tract having at least 4 nucleotides of the sequence CAACGAG, where the first 5' nucleotide is C; (f) the 7 nucleotides 5' to the pyrimidine tract having at least 5 nucleotides of the sequence CAACGAG, where the first 5' nucleotide is CAA; (g) the 7 nucleotides 5' to the pyrimidine tract having at least 5 nucleotides of the sequence CAACGAG, where the first 5' nucleotide is CAAC; (h) the 7 nucleotides 5' to the pyrimidine tract having at least 6 nucleotides of the sequence CAACGAG, where the first 5' nucleotide is CAA; (i) the 7 nucleotides 5' to the pyrimidine tract having at least 6 nucleotides of the sequence CAACGAG, where the first 5' nucleotide is CAAC; or (j) the 7 nucleotides 5' to the pyrimidine tract having the sequence CAACGAG. Without being bound by theory, this sequence may function as a splice enhancer.

[0044] It is further contemplated that the acceptable splice region may contain one, two, three, four, five, or more different / identical branch point sequences.It is further contemplated that the acceptable splice region may contain one, two, three, four, five, or more different or identical intronic splice enhancer (sequences).

[0045] Pyrimidine tracts, acceptable splice sites, and optionally further branch point sequences and / or intronic splice enhancers are contemplated, such that the entire acceptable splice region comprises a total of about 1000, 500, 250, 100, 50, 45, 40, 35, 30, 25, 20, 15 or less nucleotides, preferably 26 nucleotides.

[0046] In some embodiments, the branch point sequence or branch point nucleotides used by a particular acceptor splice site may be flexible. This means that a variety of branch points can be used, such as 3 or more, 5 or more, 7 or more, 9 or more, or 11 or more. In one embodiment, the frequency of use of each branch point is less than 30% or less than 20%. When the selection of the branch point sequence or branch point nucleotides used by a particular acceptor splice site is flexible, the strength of the acceptor splice region is not dependent from the presence of the particular branch point sequence or branch point nucleotides contained therein. The flexibility of using any branch point sequence or branch point nucleotide makes the acceptor splice region very efficient and versatile, and may be particularly suitable for use with a variety of nucleic acid sequences to induce efficient splicing in a sequence-independent manner.

[0047] The acceptable splice region described herein and present in the nucleic acids, pre-mRNA trans-splicing molecules, AVV vectors, AVV vector systems and those used in the methods according to the invention comprises, in addition to the acceptable splice site and the polypyrimidine tract, further nucleotides 5' to the polypyrimidine tract. In a preferred embodiment, the acceptable splice region comprises about 7 nucleotides (e.g., 5-12 nucleotides, preferably 6-10 nucleotides, more preferably 6-8 nucleotides) 5' to the polypyrimidine tract. In a particularly preferred embodiment, the acceptable splice region comprises the sequence CAACGAG 5' to the polypyrimidine tract. The acceptable splice region comprising the acceptable splice site, the polypyrimidine tract and further nucleotides 5' to the polypyrimidine tract, preferably about 7 further nucleotides 5' to the polypyrimidine tract, serves as a minimal acceptable splice region and can be inserted into any nucleic acid sequence to introduce a very strong functional acceptable splice region.

[0048] It is further contemplated that the acceptable splice region of the present invention comprises or consists of the sequence CAACGAGTCTTTTGTCATCTACAGGT (SEQ ID NO:3). It is further contemplated that the acceptable splice region of the present invention comprises or consists of the sequence CAACGAGTTTTTTGTCATCTACAGGT (SEQ ID NO:4). It is also contemplated that the acceptable splice region of the present invention comprises or consists of any of the sequences CTACACGCTCAAGCCGGAGGTCAACAACGAGTCTTTTGTCATCTACAGGT (SEQ ID NO:5), GCCGGAGGTCAACAACGTCTTTTGTCATCTACAGGT (SEQ ID NO:6), GTCTTTTGTCATCTACAGGT (SEQ ID NO:7), GTCTTTTGTCATCTACAGGTGTTCGTG (SEQ ID NO:26), or GTCTTTTGTCATCTACAGGTGTTCGTGGTTCGTGGTCCA (SEQ ID NO:8). Preferably, the acceptor splice region of the present invention comprises or consists of the sequence CAACGAGTCTTTTGTCATCTACAGGT (SEQ ID NO:3) or CAACGAGTTTTTTGTCATCTACAGGT (SEQ ID NO:4) (or is encoded by a nucleic acid sequence comprising or consisting of the sequence SEQ ID NO:3 or SEQ ID NO:4).

[0049] It is also contemplated that the acceptable splice regions described herein have a nucleotide sequence that is at least 70%, 75%, 80%, 85%, 90%, 95%, 97%, 98%, 99%, or 100% identical to any one of SEQ ID NOs:3, 4, 5, 6, 7, 26, or 8. Preferably, the acceptable splice region comprises a nucleotide sequence that is at least 85%, 90%, 95%, 97%, 98%, 99%, or 100% identical to any one of SEQ ID NOs:3, 4, 5, 6, 7, 26, or 8. More preferably, the acceptable splice region comprises a nucleotide sequence that is at least 85%, 90%, 95%, 97%, 98%, 99%, or 100% identical to any one of SEQ ID NOs:3 or 4. Even more preferably, the acceptor splice region comprises a nucleotide sequence that is at least 90%, 95%, 97%, 98%, 99%, or 100% identical to any of the sequences of SEQ ID NO:3 or 4.

[0050] The nucleic acid sequence according to the present invention may further comprise a nucleotide sequence of interest in addition to the acceptor splice site. The nucleotide sequence of interest may be any suitable nucleotide sequence. The terms "nucleotide sequence of interest" and "nucleic acid sequence of interest" are used interchangeably herein. Typically, this term is referred to as a "nucleotide sequence of interest" to better distinguish it from the nucleic acid sequence or nucleic acid molecule according to the present invention. For example, the nucleotide sequence of interest may be an intronic or exonic nucleotide sequence, or may include both intronic and exonic (i.e., non-coding and coding) sequences. It is further contemplated that the nucleotide sequence of interest is a cDNA, an mRNA, an rRNA, or a tRNA. The nucleotide sequence of interest or a portion thereof may be located 3' or 5', preferably 3', of the acceptor splice region. In certain embodiments, the nucleotide sequence of interest or a portion thereof included in any of the nucleic acid sequences, pre-mRNA trans-splicing molecules, AAV vectors, AAV vector systems, or used in any of the methods according to the present invention may be a transgene or a portion thereof. It has been found that splicing efficiency, especially trans-splicing efficiency, can be improved by the presence of an intron in the nucleotide sequence of interest or a portion thereof. In particular, when the nucleotide sequence of interest or a portion thereof is greater than 1000 bp, it may be advantageous to introduce or retain an intron in the sequence. Preferably, an intron is present about every 200-1000 bp of the nucleotide sequence of interest or a portion thereof, but may occur at lower or higher sequence intervals. The intron may be derived from the nucleotide sequence of interest and / or from a nucleotide sequence of a different gene and / or from an artificial nucleotide sequence.

[0051] As used herein, a nucleotide sequence of interest (also referred to as a nucleic acid sequence of interest) may be a coding sequence, e.g., a sequence that encodes a polypeptide of interest, or a portion thereof. In a preferred embodiment, the nucleotide sequence of interest is a coding sequence, more preferably a transgene, a sequence that encodes a polypeptide of interest, or a portion thereof.

[0052] It is further contemplated by the present invention that the nucleotide sequence of interest encodes a therapeutic polypeptide, a therapeutic nucleic acid (eg, a therapeutic RNA), or a portion thereof.

[0053] The therapeutic nucleic acid or therapeutic polypeptide may be used to treat ocular diseases, such as autosomal recessive severe early-onset retinal degeneration (Leber's congenital amaurosis), congenital color blindness, Stargardt's disease, Best's disease (vitelliform macular degeneration), Doyne's disease, and the like. disease), retinitis pigmentosa (especially autosomal dominant, autosomal recessive, X-linked, digenic or polygenic retinitis pigmentosa), (X-linked) retinoschisis, macular degeneration (AMD), age-related macular degeneration, atrophic age-related macular degeneration, neovascular AMD, diabetic maculopathy, proliferative diabetic retinopathy (PDR), cystoid macular edema, central serous retinopathy, retinal detachment, endophthalmitis, glaucoma, posterior uveitis, congenital stationary night blindness, total choroidal atrophy, early-onset retinal dystrophies, cone dystrophies, rod-cone or cone-rod dystrophy, pattern dystrophies, Usher syndrome and other syndromic ciliary disorders It can be used to treat certain ciliopathies, such as Bardet-Biedl syndrome, Joubert syndrome, Senior-Loken syndrome or Alstrom syndrome.

[0054] The nucleic acids, therapeutic nucleic acids, therapeutic proteins / polypeptides or therapeutic molecules can be used to treat disorders affecting photoreceptor cells such as rods and / or cones (photoreceptor cell diseases). Non-limiting examples of photoreceptor cell diseases include color vision deficiencies, age-related macular degeneration, retinal degeneration, retinal dystrophies, retinitis pigmentosa, cone dystrophies, rod-cone dystrophies, color blindness, macular degeneration, night blindness, retinoschisis, total choroidal atrophy, diabetic retinopathy, hereditary optic neuropathy, Smallmouth disease type I, retinitis punctata albescens (RPA), progressive retinal atrophy (PRA), fundus albescens (FA), or congenital stationary night blindness (CSNB).

[0055] Therapeutic nucleic acids or therapeutic polypeptides can be used to treat inherited retinal diseases (IRDs). IRDs are a group of genetically and phenotypically heterogeneous disorders. IRDs are the leading cause of blindness in people aged 15 to 45 years, with an estimated prevalence ranging from 1 in 1,500 to 1 in 3,000. IRDs can be classified according to the type of retinal cells primarily affected, i.e., rod or cone photoreceptors, and according to the disease state. Non-limiting examples of IRD are color vision deficiencies, age-related macular degeneration, retinal degeneration, retinal dystrophy, retinitis pigmentosa, cone dystrophy, rod-cone dystrophy, color blindness, macular degeneration, retinoschisis, total choroidal atrophy, diabetic retinopathy, hereditary optic neuropathy, Type I Smallmouth disease, Retinitis Punctata Albescens (RPA), Progressive Retinal Atrophy (PRA), Fundus Punctata Albescens (FA), or Congenital Stationary Night Blindness (CSNB).

[0056] The most common progressive IRD, which primarily affects the cones, is Stargardt's macular dystrophy, which is inherited in an autosomal recessive manner and is primarily caused by mutations in the ABCA4 gene, which encodes the retinal transporter ABCR. In one embodiment, the therapeutic nucleic acid encodes ABCA4 or a portion thereof, and the therapeutic nucleic acid is used to treat Stargardt's macular dystrophy.

[0057] The most common progressive IRD, which mainly affects rods, is retinitis pigmentosa (RP). RP is a progressive disease that leads to night blindness and visual field constriction. In later stages, secondary cone photoreceptor cell death is induced, eventually causing complete vision loss. Some forms of retinitis pigmentosa are associated with mutations in the gene encoding rhodopsin. In one embodiment, the therapeutic nucleic acid encodes rhodopsin or a portion thereof, and the therapeutic nucleic acid is used to treat retinitis pigmentosa.

[0058] The nucleic acid sequence according to the present invention may also include a donor splice site, where the donor splice site is located 3' of the nucleotide sequence of interest or a portion thereof. The term "donor splice region" is used interchangeably herein with "donor splice site" (abbreviated as DSS) or "splice donor site" (abbreviated as SDS). Thus, the nucleic acid sequence of the present invention may also include a donor splice site. Donor splice sites are known to those skilled in the art and are described, inter alia, in Buckley et al. (2009) "A method for identifying alternative or cryptic donor splice sites within gene and mRNA sequences. Comparisons among sequences from vertebrates, echinoderms and other groups" BMC Genomics 10: 318 and Qu et al. (2017) "A Bioinformatics-Based Alternative mRNA Splicing Code that May Explain Some Disease Mutations Is Conserved in Animals" Front. Genet., vol. 8 Art. 38. Any suitable donor splice site is contemplated by this term. Donor splice sites are usually located at the 3' end of an exon and the 5' end of the adjacent intron. However, in the context of the present invention, the donor splice site may also be located anywhere within the nucleotide sequence of interest. Similar to the acceptor splice site, the donor splice site allows the spliceosome to cleave at that site.

[0059] In principle, the donor splice site described herein may be located at any position of the nucleic acid molecule. It is also envisaged that the donor splice site is located at the 3' end of the (first) nucleic acid sequence. For example, the donor splice site sequence may comprise or consist of the sequence AAGGTAAGT, AAGGTGAGT, CAGGTAAGT, CAGGTGAGT, AAGGTAAG, AAGGTGAG, CAGGTAAG, or CAGGTGAG. It is also envisaged that the donor splice site sequence has a sequence with at least 70%, 80%, 85%, 90%, 95%, 98%, 99%, or 100% sequence identity to the sequence AAGGTAAGT, AAGGTGAGT, CAGGTAAGT, CAGGTGAGT, AAGGTAAG, AAGGTAG, CAGGTAAG, or CAGGTGAG.

[0060] Similarly, the acceptable splice regions described herein may be located anywhere in the nucleic acid molecule. It is also envisaged that the acceptable splice site is located at the 5' end of the (second) nucleic acid sequence.

[0061] The nucleotide sequence of interest may additionally or alternatively contain a polyadenylation signal and / or a promoter as described herein.

[0062] The present invention also provides (A) providing a first nucleic acid sequence comprising one or more donor splice site sequences; (B) providing a second nucleic acid sequence comprising: (i) an acceptor splice region comprising: (ia) a pyrimidine tract comprising: (iaa) 5–25 nucleotides; (iab) where at least 60% of the nucleotides within these 5 to 25 nucleotides are pyrimidine bases, e.g., cytosine (C), thymine (T), and / or uracil (U); (ib) acceptor splice site; (iba) wherein the acceptor splice site is located 3' to the pyrimidine tract; and (ibb) wherein the acceptor splice site comprises a sequence 5' to 3' of NAGG, where N is A, C, T / U, or G; thereby obtaining a nucleic acid sequence. The method may also include the steps of (C) cleaving the first nucleic acid sequence at one or more donor splice site sequences and cleaving the second nucleic acid sequence at the acceptor splice site, and (D) ligating the first cleaved nucleic acid sequence to the second cleaved nucleic acid sequence.

[0063] This method reflects the so-called trans-splicing. In particular, trans-splicing is a type of splicing in which exons from two separate nucleic acid molecules (preferably pre-mRNA molecules) are joined together to form one nucleic acid molecule. For example, when pre-mRNA is spliced ​​into the final mRNA, tRNA, or rRNA. This process is also known to those skilled in the art and is described, inter alia, under the heading "DNA to RNA" in Alberts B, Johnson A, Lewis J, et al. (2002) "Molecular Biology of the Cell. 4th edition." New York: Garland Science. Thus, the spliced ​​nucleic acid sequence is preferably an mRNA, rRNA, or tRNA sequence. Thus, the nucleotide sequence of interest may be a pre-mRNA.

[0064] The first and second nucleic acid sequences may be DNA or RNA sequences.When the first and second nucleic acid sequences are DNA sequences, both the first and second nucleic acid sequences further comprise a promoter, and at least the second, preferably both nucleic acid sequences also comprise a transcription termination sequence, such as a polyA sequence.When the first and second nucleic acid sequences are RNA sequences, at least the second, preferably both nucleic acid sequences comprise a termination sequence, such as a polyA sequence.Preferably, the nucleic acid sequence is a DNA sequence, and more preferably, it is a vector or plasmid that comprises a DNA sequence.

[0065] In one embodiment, the first and second nucleic acid sequences further comprise a binding domain. In yet another embodiment, the first and second nucleic acid sequences are pre-mRNA trans-splicing molecules or DNA sequences encoding pre-mRNA trans-splicing molecules described herein.

[0066] Without being bound by theory, cleavage of a first nucleic acid sequence at one or more donor splice site sequences and cleavage of a second nucleic acid sequence at an acceptor splice region sequence, and ligation of the first cleaved nucleic acid sequence with the second cleaved nucleic acid sequence are performed using two transesterification reactions in cis-splicing and trans-splicing, and are referred to herein as "splicing." As used herein, the term "trans-splicing" refers to RNA splicing, particularly pre-RNA splicing of two separate RNA or pre-RNA molecules to form a mature RNA, such as the final mRNA, tRNA, or rRNA.

[0067] The term "in vitro methods" refers to methods outside the body, ie, the human or animal body, but includes cell lines or primary cells outside the body.

[0068] The methods and uses described herein may be carried out in any suitable cell-free system, such as a cell lysate, or in any suitable host cell.

[0069] Such suitable host cells are known to those skilled in the art. For example, suitable host cells contain a splicing mechanism, i.e., a spliceosome. Those skilled in the art can select regulatory elements for use in a suitable host, e.g., a mammalian or human host cell. Regulatory elements include, for example, promoters, termination sequences (e.g., transcription termination sequences, translation termination sequences), enhancers, and polyadenylation elements. Furthermore, the expression level of the transgene (nucleic acid sequence of interest) can be enhanced by using regulatory elements such as enhancer sequences or the Woodchuck Hepatitis Virus Posttranscriptional Response Element (WPRE). However, these elements further limit the packaging capacity of the rAAV. Therefore, in some cases, especially when size limitations are important, it is recommended to omit such optional elements and preferably use small essential regulatory elements such as short polyadenylation signals as termination sequences for the AAV vectors according to the present invention, and especially for dual rAAV vector systems. The host cell may be a eukaryotic cell, or a mammalian cell, such as a human or rodent cell line. The host cell may be derived from a HEK cell, e.g., a HEK293(T) cell, a Cos7 cell, a CHO cell, a fibroblast cell, e.g., a fibroblast cell of human or mouse origin, a retinoblastoma cell, a 661W cell, an induced pluripotent stem cell (iPSC), e.g., a human iPSC, a photoreceptor cell, e.g., a photoreceptor cell of vertebrate origin, a neuronal cell, or a glial cell.

[0070] In one embodiment, the method is performed in vitro (i.e., outside the body) in a host cell and includes introducing the first nucleic acid sequence and introducing the second nucleic acid sequence, which may be introduced by transfection or transduction, more particularly by co-transfection or simultaneous transduction.

[0071] Thus, in one embodiment, the method comprises: (A) providing a first nucleic acid sequence comprising one or more donor splice site sequences; (B) providing a second nucleic acid sequence comprising: (i) an acceptor splice region comprising: (ia) a pyrimidine tract comprising: (iaa) 5–25 nucleotides; (iab) where at least 60% of the nucleotides within these 5 to 25 nucleotides are pyrimidine bases, e.g., cytosine (C), thymine (T), and / or uracil (U); (ib) acceptor splice site; (iba) wherein the acceptor splice site is located 3' to the pyrimidine tract; and (ibb) wherein the acceptor splice site comprises a sequence 5' to 3' of NAGG, where N is A, C, T / U, or G; (ii) a nucleotide sequence of interest, wherein the nucleotide sequence of interest (iia) is located 3' to the acceptor splice region; The method may also include (C) cleaving the first nucleic acid sequence at one or more donor splice site sequences and cleaving the second nucleic acid sequence at the acceptor splice site, and (D) ligating the first cleaved nucleic acid sequence to the second cleaved nucleic acid sequence, thereby obtaining the nucleic acid sequence.

[0072] When both nucleic acid sequences comprise a nucleotide sequence of interest, the method further comprises the feature that the first nucleotide sequence (ii) comprises the nucleotide sequence of interest, wherein the nucleotide sequence of interest is located 5' to the donor splice site.

[0073] Both the second and first nucleic acid sequences may comprise a nucleotide sequence of interest, as described herein. Such a nucleotide sequence of interest may be located 5' of the donor splice site. The nucleotide sequence of interest may additionally or alternatively be located 3' of the acceptor splice region. In another embodiment of the invention, the polynucleotide comprises the following order: preferably 5'-(sequence of interest)-(donor splice site)-3' and 5' junction-(acceptor splice region)-(sequence of interest)-3'. Alternatively, 5'-(sequence of interest)-(acceptor splice region)-3' and 5' junction-(donor splice site)-(sequence of interest)-3'.

[0074] In a preferred embodiment, the nucleotide sequence of interest is a sequence encoding a polypeptide, the first nucleic acid sequence comprises a 5' portion of the sequence encoding the polypeptide and the second nucleic acid sequence comprises a 3' portion of the sequence encoding the polypeptide, and trans-splicing reconstitutes the 5' and 3' portions of the sequence encoding the polypeptide.

[0075] Ligation of the cleaved nucleic acid sequences is a process known to those skilled in the art. For example, this ligation process can be carried out by a ligating molecule. For example, the ligating molecule may be an RNA ligase or a protein with RNA ligase function. Such RNA ligases are known to those skilled in the art and are described, inter alia, in, for example, Chambers and Patrick (2015) "Archaeal Nucleic Acid Ligases and Their Potential in Biotechnology" Hindawi Publishing Corporation Archaea Volume 2015, Article ID 170571, page 10. Some (t)RNA ligases are described, for example, in Popow et al. (2012) "Diversity and roles of (t)RNA ligases" Cell Mol Life Sci. 69(16):2657-70. Thus, a method comprising a ligation step may further comprise a step of adding an RNA ligase. The RNA ligase may be, for example, an mRNA ligase, a tRNA ligase, or an rRNA ligase. Methods illustrating how, for example, ligated mRNA molecules can be analyzed are provided in the Examples.

[0076] In a preferred embodiment, the method is carried out in a host cell, more preferably in vitro (i.e., outside the body). Thus, the first and second nucleic acid sequences are introduced into the host cell, where the first and second nucleic acid sequences are recombinant nucleic acid sequences and / or heterologous to the host cell. Introducing the first and second nucleic acid sequences into the host cell includes transfecting or transducing. Transfecting may be DNA or RNA transfecting, and methods of transfecting DNA or RNA are known to those skilled in the art. Preferably, the nucleic acid sequence is a DNA sequence, and the DNA sequence is transfected in the form of a plasmid or vector comprising said DNA sequence. Transduction as referred to herein means introducing the nucleic acid using a viral vector, where the viral vector may comprise DNA or RNA. In a preferred embodiment, the viral vector is a DNA viral vector, and may comprise single-stranded DNA (e.g., AAV) or double-stranded DNA.

[0077] In one embodiment, the method according to the invention comprises the steps of: (A) introducing into a host cell a first nucleic acid sequence comprising a pre-mRNA trans-splicing molecule sequence or a nucleic acid sequence encoding said pre-mRNA trans-splicing molecule, wherein the first pre-mRNA trans-splicing molecule comprises from 5' to 3': (a) a 5' portion of a nucleotide acid sequence of interest; (b) a donor splice site; (c) optionally a spacer sequence; (d) a first binding domain; and (e) optionally a termination sequence, preferably a polyA sequence; and (B) introducing into a host cell a second nucleic acid sequence comprising a pre-mRNA trans-splicing molecule sequence or a nucleic acid sequence encoding said pre-mRNA trans-splicing molecule, wherein the second pre-mRNA trans-splicing molecule comprises from 5' to 3': (i) a second binding domain complementary to the first binding domain of the first nucleic acid sequence; (ii) an acceptor splice site comprising: a rice domain sequence; (iia) a pyrimidine tract comprising: (iiaa) 5 to 25 nucleotides; (iiab) wherein at least 60% of the nucleotides within these 5 to 25 nucleotides are pyrimidine bases, e.g., cytosine (C), thymine (T), and / or uracil (U); (iib) an acceptable splice site, (iiba) wherein the acceptable splice site is located 3' to the pyrimidine tract; and (iibb) wherein the acceptable splice site comprises the sequence 5' to 3' of NAGG, where N is A, C, T / U, or G; and (iii) a 3' portion of a nucleotide sequence of interest. The method may further comprise the steps of (C) cleaving the first nucleic acid sequence at the donor splice site sequence and cleaving the second nucleic acid sequence at the acceptor splice site; (D) ligating the first cleaved nucleic acid sequence comprising the 5' portion of the nucleotide sequence of interest to the second cleaved nucleic acid sequence comprising the 3' portion of the nucleotide sequence of interest, thereby obtaining the nucleic acid sequence of interest. In a preferred embodiment, the first and second nucleic acid sequences are DNA sequences encoding said pre-mRNA trans-splicing molecule, the first nucleic acid sequence further comprising a promoter 5' of the 5' portion of the nucleic acid sequence of interest and the second nucleic acid sequence further comprising a promoter 5' of the second binding domain.

[0078] The present invention also relates to a nucleic acid sequence comprising the following, wherein the nucleotide sequence of interest is not exon 3 of the rhodopsin gene of SEQ ID NO:9, or the nucleotide sequence of interest does not comprise the sequence defined in SEQ ID NO:9 and / or 10, or the nucleic acid sequence of interest does not comprise exon 3 of the rhodopsin gene of SEQ ID NO:9: (i) an acceptor splice region sequence comprising: (ia) a pyrimidine tract comprising: (iaa) 5–25 nucleotides; (iab) where at least 60% of the nucleotides within these 5 to 25 nucleotides are pyrimidine bases, e.g., cytosine (C), thymine (T), and / or uracil (U); (ib) acceptor splice site; (iba) wherein the acceptor splice site is located 3' to the pyrimidine tract; and (ibb) wherein the acceptor splice site comprises a sequence 5' to 3' of NAGG, where N is A, C, T / U, or G; (ii) a nucleotide sequence of interest, wherein the nucleotide sequence of interest (iia) is located 3' or 5' to the acceptor splice region;

[0079] The present invention also relates to a nucleic acid sequence comprising the following, wherein the nucleotide sequence of interest is not the rhodopsin mRNA of SEQ ID NO:10: (i) an acceptor splice region sequence comprising: (ia) a pyrimidine tract comprising: (iaa) 5–25 nucleotides; (iab) where at least 60% of the nucleotides within these 5 to 25 nucleotides are pyrimidine bases, e.g., cytosine (C), thymine (T), and / or uracil (U); (ib) acceptor splice site; (iba) wherein the acceptor splice site is located 3' to the pyrimidine tract; and (ibb) wherein the acceptor splice site comprises a sequence 5' to 3' of NAGG, where N is A, C, T / U, or G; (ii) a nucleotide sequence of interest, wherein the nucleotide sequence of interest (iia) is located 3' or 5' to the acceptor splice region.

[0080] It is contemplated that the nucleic acid sequences of the present invention have a length of up to 5800 nucleotides. It is further contemplated that the nucleic acid sequences of the present invention have a length of up to 5500, 5000, 4500, 4000, 3500, 3000, 2500, 2000, 1500, 1000, 500, 400, 300, 200, or up to 150 nucleotides.

[0081] Another application that the present invention's acceptor splice region can be used in nucleic acid molecules is the so-called exogenous pre-mRNA trans-splicing molecule for Smart technology.Such pre-mRNA trans-splicing molecules that can introduce the present invention's acceptor splice region are described, inter alia, in WO 2011 / 042556, WO 2013 / 025461, and Berger et al. (2016) "mRNA trans splicing in gene therapy for genetic diseases" WIREs RNA, 7:487-498, Puttaraju et al. (1999) "Spliceosome-mediated RNA trans-splicing as a tool for gene therapy" Nature Biotechnology, vol. 17, pp. 246-252; and Mansfield et al. (2003) "5' Exon replacement and repair by spliceosome-mediated RNA trans-splicing" RNA, vol. 9: 1290-1297. From these references, the skilled artisan also knows how to construct such pre-mRNA trans-splicing molecules. Some exemplary pre-mRNA trans-splicing molecule nucleic acid molecules are also described herein.

[0082] The present invention also relates to a pre-mRNA trans-splicing molecule comprising: (i) an acceptor splice region sequence comprising: (ia) a pyrimidine tract comprising: (iaa) 5–25 nucleotides; (iab) where at least 60% of the nucleotides within these 5 to 25 nucleotides are pyrimidine bases, e.g., cytosine (C), thymine (T), and / or uracil (U); (ib) acceptor splice site; (iba) wherein the acceptor splice site is located 3' to the pyrimidine tract; and (ibb) wherein the acceptor splice site comprises a sequence 5' to 3' of NAGG, where N is A, C, T / U, or G; (ii) a nucleotide sequence of interest; wherein the acceptor splice region is located 5' to the nucleotide sequence of interest; (iii) a pre-mRNA-targeting binding domain located 5' to the nucleic acid sequence of interest; and (vi) optionally a spacer sequence, wherein the spacer sequence is located between the binding domain and the acceptor splice region.

[0083] In a preferred embodiment, the acceptor splice region is located 5' to the nucleotide sequence of interest and the binding domain is located 5' to the acceptor splice region. Preferably, the nucleic acid sequence or the pre-mRNA trans-splicing molecule further comprises a termination sequence, preferably a polyA sequence.

[0084] The present invention further relates to a pre-mRNA trans-splicing molecule comprising: (i) an acceptor splice region sequence comprising: (ia) a pyrimidine tract comprising: (iaa) 5–25 nucleotides; (iab) where at least 60% of the nucleotides within these 5 to 25 nucleotides are pyrimidine bases, e.g., cytosine (C), thymine (T), and / or uracil (U); (ib) acceptor splice site; (iba) wherein the acceptor splice site is located 3' to the pyrimidine tract; and (ibb) wherein the acceptor splice site comprises a sequence 5' to 3' of NAGG, where N is A, C, T / U, or G; (ii) a nucleotide sequence of interest, where the acceptor splice region is located 5′ to the nucleotide sequence of interest; (iii) a donor splice site, wherein the donor splice site is located 3' to the nucleotide sequence of interest; (iv) a first binding domain that targets the pre-mRNA located 5′ to the nucleotide sequence of interest; (v) a second binding domain that targets the pre-mRNA located 3′ to the nucleotide sequence of interest; (vi) optionally a first spacer sequence, wherein the first spacer is located between the first binding domain and the acceptor splice region; (vii) optionally a second spacer sequence, where the second spacer is located between the second binding domain and the donor splice site.

[0085] In one embodiment, the first binding domain is located 5' to the acceptor splice region and the second binding domain is located 3' to the donor splice site. Those skilled in the art will understand that the DNA sequence encoding the pre-mRNA trans-splicing molecule includes a promoter for transcribing the RNA molecule, the pre-mRNA trans-splicing molecule. The term "pre-mRNA targeting binding domain" may also be referred to as a "pre-mRNA targeting binding domain", which may be located 5' or 3' to the nucleotide sequence of interest.

[0086] The nucleic acid sequences described herein can reflect pre-mRNA trans-splicing molecules. Those skilled in the art will understand that the nucleic acid sequences and pre-mRNA trans-splicing molecules described herein are recombinant sequences or molecules. These pre-mRNA trans-splicing molecules are known in the art and are described, inter alia, in WO 2011 / 042556, WO 2013 / 025461, and Berger et al. (2016) “mRNA trans splicing in gene therapy for genetic diseases” WIREs RNA, 7:487-498, Puttaraju et al. (1999) “Spliceosome-mediated RNA trans-splicing as a tool for gene therapy” Nature Biotechnology, vol. 17, pp. 246-252;Mansfield et al. (2003) “5' Exon replacement and repair by spliceosome-mediated RNA trans-splicing” RNA, vol. 9: 1290-1297, and Berger et al. (2016) “mRNA trans-splicing in gene therapy for genetic diseases” WIREs RNA, 7:487-498. Thus, those skilled in the art know how to construct these pre-mRNA trans-splicing molecules. It is envisaged that the pre-mRNA trans-splicing molecule binds to a target pre-mRNA, where the pre-mRNA can be a natural pre-mRNA, particularly a pre-mRNA endogenous to a host cell, or a recombinant pre-mRNA, particularly another pre-mRNA trans-splicing molecule. It is also envisaged that the pre-mRNA trans-splicing molecule preferentially induces the trans-splicing reaction more efficiently than the cis-splicing reaction.

[0087] The term "pre-mRNA targeting binding domain" as used herein may be any suitable "pre-mRNA targeting binding domain", meaning complementary to a sequence of a target pre-mRNA located 5' or 3' to a nucleotide sequence of interest. As used herein, a "pre-mRNA targeting binding domain" may also be referred to as a "target binding domain", "binding domain" (abbreviated as BD), or "binding sequence". For example, the binding domain may recognize a target pre-mRNA or mRNA by base pairing. For example, the target of the mRNA may be an intron. The target binding domain of the pre-mRNA trans-splicing molecules described herein may include one or two binding domains of at least 15-30 nucleotides, preferably 80-120 nucleotides, more preferably about 100 nucleotides; or may have a longer target binding domain of up to several hundred nucleotides that is complementary to a target region of a selected (e.g., endogenous) pre-mRNA and in an antisense orientation, as described in U.S. Patent Publication No. US 2006-0194317 A1. This confers specificity of binding and tightly fixes the endogenous pre-mRNA in space such that, for example, the spliceosome in the nucleus of the host cell can trans-splice a portion of the pre-mRNA trans-splicing molecule to a portion of the (e.g., endogenous) pre-mRNA. Alternatively, the binding domain may be reverse-complementary to the binding domain of another recombinant pre-mRNA, such as another pre-mRNA trans-splicing molecule. This confers specificity of binding and tightly fixes the two pre-mRNA trans-splicing molecules in space such that, for example, the spliceosome in the nucleus of the host cell can trans-splice a portion of one pre-mRNA trans-splicing molecule to a portion of the other pre-mRNA trans-splicing molecule. Thus, it is envisioned that the binding domain comprises between about 15-250 nucleotides, between about 15-200 nucleotides, between about 100-200 nucleotides, or less than 500, 400, 300, or 200 nucleotides.In one embodiment, the binding domain comprises 50-150 nucleotides, preferably 80-120 nucleotides, even more preferably 90-110 nucleotides or about 100 nucleotides. Alternatively, or even more preferably, an efficient binding domain has a GC content of 45-65% and / or is derived from an intronic eukaryotic or bacterial sequence and / or comprises at least one branch point in the last 30 bp (e.g. as assessed using human splicing finder 3.1 (http: / / www.umd.be / HSF / index.html)).

[0088] In the host cell, the binding domain can bind to an endogenous pre-mRNA or a second heterologous or recombinant pre-mRNA (provided or introduced into the host cell) as long as the endogenous pre-mRNA or the second heterologous or recombinant pre-mRNA contains a sequence that is reverse-complementary to the target binding domain. Thus, with respect to the second heterologous or recombinant pre-mRNA, it contains a binding domain that is reverse-complementary to the binding domain of the first heterologous pre-mRNA. The terms "endogenous" and "heterologous" are used herein with respect to the host cell. Targeting an endogenous pre-mRNA is also called spliceosome-mediated RNA trans-splicing and is usually used to replace a portion of an endogenous pre-mRNA. Targeting a second heterologous or recombinant pre-mRNA is also called trans-splicing and generates a new recombinant mRNA that links a first nucleotide sequence or a portion thereof (5') from one pre-mRNA trans-splicing molecule with a second nucleotide sequence or a portion thereof (3') of another pre-mRNA trans-splicing molecule. As used herein, the term "recombinant" refers to a DNA molecule or sequence formed by experimental methods of genetic recombination, such as molecular cloning, as well as to an RNA or polypeptide molecule / sequence encoded by said recombinant DNA molecule or sequence.

[0089] The terms "pre-mRNA" or "pre-RNA" refer to RNA before trans-splicing, and typically refer to an RNA molecule before cis-splicing (e.g., containing introns and exons), but can also refer to an RNA molecule after cis-splicing (e.g., containing only exons), such as when a binding domain targets a sequence spanning an exon-exon boundary.

[0090] The term "heterologous" as used herein refers to a protein or nucleic acid molecule or sequence that is experimentally transferred or introduced into a cell and is therefore not endogenous to this cell. If a heterologous protein or nucleic acid molecule or sequence can be derived from a different cell type or different species than the recipient, it can be derived from the same cell type or species as the recipient as long as it is introduced into the recipient host cell. As referred to herein, the terms "heterologous" and "recombinant" are used interchangeably.

[0091] Exemplary binding domains suitable for use in the present invention include, but are not limited to, binding domains having an RNA sequence encoded by the sequence of SEQ ID NO: 27, 28, 32, 33 or 18. Additional target binding domains can be identified using reporter reconstitution assays such as the Cerulean reconstitution assay to determine reconstitution efficiency as described and used in Examples 2 and 5, or as described in Riedmayr LM. (2020, SMaRT for Therapeutic Purposes. Methods Mol Biol 2079:219-32. doi: 10.1007 / 978-1-4939-9904-0_17) or Dallinger G. et al. (2003, Development of spliceosome-mediated RNA trans-splicing (SMaRT) for the correction of inherited skin diseases. Exp Dermatol 12:37-46).

[0092] The second binding region may be located at the 3" end of the molecule and may be incorporated into the pre-mRNA trans-splicing molecule of the present invention. Absolute complementarity is desirable but not required. As referred to herein, a sequence "complementary" to a portion of an endogenous pre-mRNA refers to a sequence that has sufficient complementarity to be able to hybridize with the endogenous pre-mRNA to form a stable duplex. The ability to hybridize depends on both the degree of complementarity and the length of the nucleic acid (see, for example, Sambrook et al., 1989, Molecular Cloning, A Laboratory Manual, 2d Ed., Cold Spring Harbor Laboratory Press, Cold Spring Harbor, NY).

[0093] Thus, complementarity as used with respect to a binding domain that targets a pre-mRNA means that the binding domain is at least 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 98%, 99%, or 100% complementary to the target sequence on the pre-mRNA, and preferably at least 90%, 95%, 98%, 99%, or 100% complementary to the target sequence on the pre-mRNA.

[0094] A spacer region for separating the splice site from the binding domain is also preferably included in the pre-mRNA trans-splicing molecule. The spacer sequence is a region of the pre-mRNA trans-splicing molecule that can cover the 3' and / or 5' splice site elements of the pre-mRNA trans-splicing molecule by relatively weak complementarity, thereby preventing non-specific trans-splicing.

[0095] The spacer separating the 5' donor splice site and the 3' terminal binding domain may comprise between about 10-100 nucleotides, preferably between about 10-70 nucleotides, between about 20-70 nucleotides, between about 20-50 nucleotides, more preferably between about 30-50 nucleotides. The spacer may comprise a downstream intron splice enhancer (DISE). The spacer sequence separating the 3' acceptor splice region from the 5' terminal binding domain may comprise between about 2-100 nucleotides, or between about 10-100 nucleotides, preferably between about 2-50 nucleotides, more preferably between about 5-20 nucleotides.

[0096] The spacer may be a non-coding sequence, but may be designed to include features such as a stop codon that blocks translation of the spliced ​​pre-mRNA trans-splicing molecule. Additional features may be added to the pre-mRNA trans-splicing molecule and are known to those skilled in the art, as described above. It is further envisaged that the nucleic acid sequences or pre-mRNA trans-splicing molecules described herein are included in a (recombinant) vector. Such vectors are likewise known to those skilled in the art. The vector containing the nucleotide sequence of the invention, such as the pre-mRNA trans-splicing molecule of interest, may be a plasmid, a virus, or other known in the art, used for replication and expression in mammalian cells.

[0097] Expression of the nucleotide sequence of interest or described herein, for example, the pre-mRNA trans-splicing molecule, can be regulated by any promoter / enhancer sequence known in the art to act in mammalian cells, preferably human cells. Such promoter / enhancer sequences may be inducible or constitutive. Such promoters are described elsewhere herein. One exemplary promoter is the human rhodopsin promoter or mouse rhodopsin promoter or short wavelength sensitive opsin promoter or mid wavelength sensitive opsin promoter or long wavelength sensitive opsin promoter, etc. Any type of plasmid, cosmid, YAC, or viral vector can be used to prepare a recombinant DNA construct that can be directly introduced into a tissue site. Alternatively, a viral vector can be used that selectively infects the desired target cells. Vectors for use in the practice of the present invention include any eukaryotic expression vector, including, but not limited to, viral expression vectors, such as those derived from the retrovirus, adenovirus, or adeno-associated virus classes. In a preferred embodiment, the recombinant vector of the present invention is a eukaryotic expression vector.

[0098] In another specific embodiment, the invention includes delivery of a nucleic acid sequence, such as a pre-mRNA trans-splicing molecule of the invention or a nucleic acid sequence encoding a pre-mRNA trans-splicing molecule of the invention, to a target cell. Various delivery systems are known and can be used to transfer the composition of the invention into a cell, such as encapsulation in liposomes, microparticles, microcapsules, recombinant cells capable of expressing the composition, receptor-mediated endocytosis, construction of a nucleic acid as part of a retrovirus, adenovirus, adeno-associated virus or other vector, injection of DNA, electroporation, transfection via calcium phosphate, etc. The invention also relates to cells comprising a nucleic acid sequence, such as a pre-mRNA trans-splicing molecule of the invention, a nucleic acid sequence encoding a pre-mRNA trans-splicing molecule, or a recombinant vector comprising a nucleic acid sequence encoding a pre-mRNA trans-splicing molecule of the invention.

[0099] Using the nucleic acid sequence of interest or disclosed herein, genes encoding functional biologically active molecules can be provided to cells of individuals with an inherited genetic disorder, where expression of a defective or mutated gene product produces a normal phenotype. This can be accomplished, inter alia, by adeno-associated viral vectors, as also disclosed herein.

[0100] The present invention also relates to a deoxyribonucleic acid (DNA) molecule comprising a promoter and a sequence encoding the pre-mRNA trans-splicing molecule described herein.Preferably, the DNA molecule is a vector or a plasmid, wherein the vector may be a viral vector such as an AAV, adenovirus, or lentivirus vector or plasmid.The DNA molecule such as the pre-mRNA trans-splicing molecule or viral vector can be used in therapy, particularly gene therapy, more particularly gene therapy for treating eye diseases.

[0101] Further applications that can use the acceptor splice site of the present invention are in any AAV vector or AAV vector system. Such vector systems and how they can be constructed are known to those skilled in the art and are described, inter alia, in Carvalho et al. (2017) "Evaluating efficiencies of dual AAV approaches for retinal targeting" Frontiers in Neuroscience, vol. 11, Article 503, US Patent Application No. 2014 / 0256802, or Trapani et al. (2013) "Effective delivery of large genes to the retina by dual AAV vectors" EMBO Molecular Medicine, vol. 6, no. 2, pp. 194-211. Some exemplary AAV vectors and AAV vector systems are also described herein.

[0102] The present invention also relates to an adeno-associated virus (AAV) vector comprising at least two inverted terminal repeats comprising a nucleic acid sequence between these two inverted terminal repeats, wherein the nucleic acid sequence is, from 5' to 3', (i) the promoter; (ii) a binding domain; (iii) optionally a spacer sequence; (iv) an acceptor splice region comprising: (a) a pyrimidine tract containing: (aa) 5–25 nucleotides; (ab) wherein at least 60% of the nucleotides within these 5 to 25 nucleotides are pyrimidine bases, e.g., cytosine (C), thymine (T), and / or uracil (U); (b) the acceptor splice site; ba) wherein the acceptor splice site is located 3' to the pyrimidine tract; and bb) wherein the acceptor splice site comprises a sequence 5' to 3' of NAGG, where N is A, C, T / U, or G; (v) a nucleotide sequence of interest (vi) optionally a termination sequence, e.g., a polyA sequence; Includes.

[0103] Structurally, AAV is a small (25 nm), single-stranded DNA non-enveloped virus with an icosahedral capsid. Natural or modified AAV variants (also AAV serotypes) that differ in the composition and structure of the capsid (cap) protein have various tropisms, i.e. the ability to transduce different (retinal) cell types (Boye et al. (2013) “A comprehensive review of retinal gene therapy” Molecular therapy 21 509-519). When combined with a ubiquitously active promoter, this tropism defines the site of gene expression. On the other hand, in combination with cell type-specific promoters, the degree of site specificity (i.e., expression of the transgene only in rod or cone photoreceptors) is defined by the combination of both AAV serotype tropism and promoter specificity (Schon et al. (2015) “Retinal gene delivery by adeno-associated virus (AAV) vectors: Strategies and applications” Eur J Pharm Biopharm. 95(Pt B):343-52).

[0104] The terms "adeno-associated virus vector" or "recombinant AAV" or "rAAV", all used interchangeably, are meant to include any AAV that contains a heterologous polynucleotide sequence (also referred to herein as a nucleotide sequence of interest) in its viral genome. Generally, the heterologous polynucleotide / nucleotide sequence of interest is flanked by at least one, and generally two, natural or variant AAV inverted terminal repeats (ITRs). The term rAAV vector encompasses both rAAV vector particles and rAAV vector plasmids. Thus, for example, a rAAV that contains a heterologous polynucleotide may be a rAAV that contains a nucleic acid sequence that is not normally included in naturally occurring wild-type AAV, such as a transgene (e.g., a non-AAV RNA-encoding polynucleotide sequence, a non-AAV protein-encoding polynucleotide sequence), a non-AAV promoter sequence, a non-AAV polyadenylation sequence, etc.

[0105] Such recombinant AAV vectors are common knowledge in the art, and those skilled in the art also know how to construct such recombinant AAV.

[0106] A "rAAV vector genome" or "rAAV genome" is an AAV genome (i.e., vDNA) that contains one or more heterologous nucleic acid sequences. rAAV vectors typically only require terminal repeats (TRs) to generate virus. All other viral sequences are considered unnecessary and may be supplied in trans (Muzyczka, (1992) Curr Topics Microbiol. Immunol. 158:97). Typically, rAAV vector genomes retain only one or more TR sequences to maximize the size of the transgene / nucleotide sequence of interest / heterologous nucleic acid sequence that can be efficiently packaged by the vector / capsid. Structural and nonstructural protein coding sequences may be provided in trans (e.g., from a vector such as a plasmid or by stably integrating the sequences into the packaging cell). In embodiments of the invention, the rAVV vector genome comprises at least one TR sequence (e.g., an AAV TR sequence), optionally two TRs (e.g., two AAV TRs), which are typically at the 5' and 3' ends of the vector genome and are adjacent to, but not necessarily adjacent to, the heterologous nucleic acid. The TRs may be identical to each other or different.

[0107] The terms "inverted terminal repeat", "terminal repeat" or "TR", all used interchangeably, include any viral terminal repeat or synthetic sequence that forms a hairpin structure and functions as an inverted terminal repeat (i.e., via a desired function such as replication, viral packaging, integration and / or proviral rescue). The TR may be an AAV TR or a non-AAV TR. For example, non-AAV TR sequences such as other parvoviruses (e.g., canine parvovirus (CPV), mouse parvovirus (MVM), human parvovirus B-19), or any other suitable viral sequence (e.g., SV40 hairpin that functions as an origin of SV40 replication) can be used as a TR and can be further modified by truncation, substitution, deletion, insertion, and / or addition. Additionally, the TR may be partially or completely synthetic, such as the "double D sequence" described in U.S. Pat. No. 5,478,745. The terminal repeat may have a length of about 50, 100, 150, 200, 250, 300 or more nucleotides. For example, the terminal repeat has about 145 nucleotides. For example, the inverted terminal repeat may have the sequence of SEQ ID NO: 14 and / or 15. It is also envisioned that the inverted terminal repeat may have at least 60%, 70%, 75%, 80%, 85%, 90%, 95%, 98%, or 100% sequence identity to the sequence of SEQ ID NO: 14 and / or 15. Preferably, the AAV vector comprises both SEQ ID NO: 14 and 15, or a sequence having at least 60% sequence identity to SEQ ID NO: 14 and 15.

[0108] An "AAV terminal repeat" or "AAV TR" may be from any AAV currently known or hereafter discovered, including, but not limited to, serotypes 1, 2, 3, 3B, 4, 5, 6, 7, 8, 9, 10, 11, 12, or 13, or any other AAV. The AAV terminal repeat need not have a native terminal repeat sequence (e.g., the native AAV TR sequence may be altered by insertions, deletions, truncations, and / or missense mutations), so long as the terminal repeat mediates the desired function, such as replication, viral packaging, integration, and / or proviral rescue. The viral vector of the present invention may be a "targeted" viral vector (e.g., with a directed tropism) and / or a "hybrid" parvovirus (i.e., the viral TR and viral capsid are derived from different parvoviruses), as further described in International Patent Publication No. WO 00 / 28004.

[0109] It is also envisioned that a recombinant AAV or AAV vector comprises or has a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 97%, 99%, or 100% sequence identity to naturally occurring AAV type 1 (AAV-1), AAV type 2 (AAV-2), AAV type 3 (AAV-3), AAV3B, AAV type 4 (AAV-4), AAV type 5 (AAV-5), AAV type 6 (AAV-6), AAV type 7 (AAV-7), AAV type 8 (AAV-8), AAV9, AAV10, AAV11, AAV12, rh10, avian AAV, bovine AAV, canine AAV, equine AAV, primate AAV, non-primate AAV, and ovine AAV.

[0110] According to the present invention, the nucleic acid sequence between the two inverted terminal repeats further comprises a promoter as described herein.The promoter used in the AAV can be selected from among a number of constitutive or inducible promoters that can express the selected transgene / nucleotide sequence of interest in the desired target cell.The promoter can be from any species, including human.

[0111] The promoters described herein may be "cell-specific." The term "cell-specific" means that the particular promoter selected for the recombinant vector is capable of directing expression of a selected transgene / nucleotide sequence of interest in a particular cell or ocular cell type. Useful promoters include rod opsin promoter, red-green opsin promoter, blue opsin promoter, cGMP-P-phosphodiesterase promoter, mouse opsin promoter (Beltran et al. (2010) “rAAV2 / 5 gene-targeting to rods: dose-dependent efficiency and complications associated with different promoters.” Gene Ther. 17(9):1162-74), rhodopsin promoter (Mussolino et al, Gene Ther, July 2011, 18(7):637-45); alpha subunit of cone transducin (Morrissey et al, BMC Dev, Biol, Jan 2011,11:3); beta phosphodiesterase (PDE) promoter; retinitis pigmentosa (RP1) promoter (Nicord et al, J. Gene Med, Dec 2007, 9(12): 1015-23); NXNL2 / NXNL1 promoter (Lambard et al, PLoS One, Oct. 2010, 5(10):el3025), PE65 promoter; retinal degeneration slow / peripherin 2 (Rds / perph2) promoter (Cai et al, Exp Eye Res. 2010 Aug;91(2): 186-94); and VMD2 promoter (Achi et al, Human Gene Therapy, (2009) 20:3-9).

[0112] Useful promoters for use in the present invention also include rod opsin promoters (RHO), red-green opsin promoters, blue opsin promoters, cGMP-phosphodiesterase promoters, SWS promoters (blue short wavelength sensitive (SWS) opsin promoters), mouse opsin promoters (Beltran et al. 2010, supra), rhodopsin promoters (Mussolino et al, Gene Ther, July 2011, 18(7):637-45); alpha subunit of cone transducin (Morrissey et al, BMC Dev, Biol, Jan 2011,11:3); cone arrestin (ARR3) promoters (Kahle NA et al., Hum Gene Ther Clin Dev, September 2018, 29(3):121-131), beta phosphodiesterase (PDE) promoters; retinitis pigmentosa (RP1) promoters (Nicord et al, J. Gene Med, Dec 2007, 9(12): 1015-23); NXNL2 / NXNL1 promoter (Lambard et al, PLoS One, Oct. 2010, 5(10):el3025), RPE promoter; retinal degeneration slow / peripherin 2 (Rds / perph2) promoter (Cai et al, Exp Eye Res. 2010 Aug;91(2): 186-94); VMD2 promoter (Achi et al, Human Gene Therapy, 2009 (20:3-9)), and ABCA4 promoter or any hybrid promoter consisting of at least two different promoters.

[0113] It is also contemplated that the nucleic acid sequence between the two inverted terminal repeats further comprises a termination signal, such as a polyadenylation signal, located at the 3' end of the nucleic acid sequence of interest.Because pre-RNA can be trans-spliced ​​to endogenous pre-RNA or further recombinant pre-RNA to form the 5' end of mature RNA, such as mature mRNA, a termination signal, such as a polyadenylation signal, located at the 3' end of the nucleic acid sequence of interest is optional.However, it has been shown that the efficiency of trans-splicing is higher in the presence of a polyadenylation signal.

[0114] Although a promoter is required to transcribe the pre-RNA, the nucleic acid sequence does not necessarily require an ATG (translation initiation) and / or a Kozak consensus sequence. If the pre-RNA can be trans-spliced ​​to an endogenous pre-RNA or to an additional recombinant pre-RNA to form the 3' end of a mature RNA, such as a mature mRNA, the lack of an ATG and / or a Kozak consensus sequence may be advantageous to avoid undesired translation from the pre-RNA prior to trans-splicing.

[0115] The selection of these and other common vectors and regulatory elements is conventional, and many such sequences are available. For example, see Ausubel et al., Current Protocols in Molecular Biology, John Wiley & Sons, New York, 1989. Of course, not all vectors and expression regulatory sequences work equally well to express all transgenes as described herein. However, one skilled in the art can make a selection among these and other expression regulatory sequences without departing from the scope of the present invention.

[0116] The present invention also relates to an adeno-associated virus (AAV) vector system comprising: (I) a nucleic acid sequence comprising a nucleic acid sequence between the two inverted terminal repeats; (a) a promoter; (b) a nucleotide sequence encoding the N-terminal portion of the polypeptide of interest; (c) donor splice site; (d) optionally a spacer sequence; (e) a first binding domain; and (f) optionally a termination sequence, preferably a polyA sequence; A first AAV vector comprising: (II) a second AAV vector comprising a nucleic acid sequence comprising at least two inverted terminal repeats, the nucleic acid sequence comprising the nucleic acid sequence between the two inverted terminal repeats, wherein the nucleic acid sequence between the two inverted terminal repeats is 5' to 3' (i) the promoter; (ii) a second binding domain that is complementary to the first binding domain of the first AAV vector; (iii) optionally a spacer sequence; (iv) an acceptor splice region sequence comprising: (a) a pyrimidine tract comprising: (aa) 5–25 nucleotides; (ab) wherein at least 60% of the nucleotides within these 5 to 25 nucleotides are pyrimidine bases, e.g., cytosine (C), thymine (T), and / or uracil (U); (b) the acceptor splice site; (ba) wherein the acceptor splice site is located 3' to the pyrimidine tract; and (bb) wherein the acceptor splice site comprises a sequence 5' to 3' of NAGG, where N is A, C, T / U, or G; (v) a nucleotide sequence encoding the C-terminal portion of the polypeptide of interest; (vb) wherein the C-terminal portion of the polypeptide of interest and the N-terminal portion of the polypeptide of interest reconstitute the polypeptide of interest; and (vi) a termination sequence, preferably a polyA sequence; An adeno-associated virus (AAV) vector system comprising:

[0117] Thus, the C-terminal portion of the polypeptide of interest corresponds to the portion of the polypeptide of interest that is missing from the N-terminal portion of the polypeptide contained in the first AAV vector. In one embodiment, the polypeptide is a full-length polypeptide, the first AAV vector comprises the N-terminal portion of the full-length polypeptide of interest, and the second AAV vector comprises the C-terminal portion of the full-length polypeptide of interest, and the C-terminal portion of the full-length polypeptide of interest and the N-terminal portion of the full-length polypeptide of interest reconstitute the polypeptide of interest. Thus, the C-terminal portion of the full-length polypeptide of interest corresponds to the portion of the full-length polypeptide of interest that is missing from the N-terminal portion of the full-length polypeptide of interest contained in the first AAV vector. More specifically, following trans-splicing, the mRNA comprises a sequence that encodes (in frame) the N-terminal portion of the polypeptide of interest and the C-terminal portion of the polypeptide of interest, where the mRNA encodes the (full-length) polypeptide of interest, and thus the polypeptide of interest is reconstituted.

[0118] In particular, the packaging capacity of rAAV or AAV vector genomes has a size limit of approximately 5 kb. Since the cDNAs of many therapeutic proteins are large, devising a strategy to deliver large transgenes using rAAV vectors can greatly expand the clinical application of rAAV-mediated gene therapy. Thus, in one embodiment, the polypeptide of interest is a polypeptide encoded by a transgene. To provide transgenes that already exceed the AAV packaging capacity, a trans-splicing approach has been developed. Briefly, two separate rAAV vectors deliver two parts of a transgene to a target cell. One part contains a splicing donor signal at the 3' end of the 5' part of the transgene, and the other part contains a splicing acceptor signal at the 5' end of the 3' part of the transgene. In conventional systems, intermolecular recombination between the two vector genomes generates an intervening ITR junction or intervening sequence that contains the recombinant gene sequence. This is then excised by the cellular cis-splicing machinery from a single pre-mRNA molecule to form a complete full-length transgene cassette. In the AAV vector system according to the invention, the two vector genomes are transcribed separately, the pre-mRNAs interact via their complementary binding domains, and the two pre-mRNAs are spliced ​​in trans to form the complete (full-length) transgene transcript.

[0119] The efficiency of the vector system can be assessed by measuring the expression of the mRNA encoding the (full-length) protein of interest or the protein of interest provided by the dual AAV vector approach, for example by means and techniques described herein or known to those of skill in the art.

[0120] The second binding domain is complementary to the first binding domain of the first AAV vector. Complete complementarity is preferred but not required. As referred to herein, a sequence that is "complementary" to a portion of the first binding domain means a sequence that has sufficient complementarity to be able to hybridize with the first binding domain to form a stable duplex. The ability to hybridize depends on both the degree of complementarity and the length of the nucleic acid (see, for example, Sambrook et al., 1989, Molecular Cloning, A Laboratory Manual, 2d Ed., Cold Spring Harbor Laboratory Press, Cold Spring Harbor, NY). Thus, complementarity as used with respect to the second binding domain means that the second binding domain comprises a sequence that is at least 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 98%, 99%, or 100% complementary to the first binding domain contained in the first AAV, and preferably at least 90%, 95%, 98%, 99% or 100% complementary to the first binding domain contained in the first AAV.

[0121] The first binding domain may have or comprise a sequence as set forth in SEQ ID NO: 17 or a sequence having 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 98%, 99%, or 100% sequence identity to the sequence of SEQ ID NO: 17. The second binding domain may have or comprise a sequence as set forth in SEQ ID NO: 18 or a sequence having 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 98%, 99%, or 100% sequence identity to the sequence of SEQ ID NO: 18. One skilled in the art will understand that the first binding domain and the second binding domain are interchangeable, so long as the second binding domain is complementary to the first binding domain. Thus, the first binding domain may have or comprise a sequence as set forth in SEQ ID NO: 18, or a sequence with 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 98%, 99%, or 100% sequence identity to SEQ ID NO: 17. The first binding domain may have or comprise a sequence as set forth in SEQ ID NO: 17, or a sequence with 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 98%, 99%, or 100% sequence identity to SEQ ID NO: 18.

[0122] Alternatively, the second binding domain has or comprises the sequence of any one of SEQ ID NOs: 18, 27, 28, 32, and 33, or a sequence having 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 98%, 99% or 100% sequence identity to any one of SEQ ID NOs: 18, 27, 28, 32, and 33, and the first binding domain has or comprises a sequence complementary to the second binding domain, preferably a sequence complementary to the second binding domain, preferably a sequence that is at least 80%, 85%, 90%, 95%, 98%, 99% or 100% complementary to the second binding domain, more preferably a sequence that is at least 90%, 95%, 98%, 99% or 100% complementary to the second binding domain, even more preferably a sequence that is 98%, 99% or 100% complementary to the second binding domain. One skilled in the art will understand that the first and second binding domains may be swapped, so long as the binding domains are complementary to each other. Thus, the first binding domain may have or comprise a sequence of any one of SEQ ID NOs: 18, 27, 28, 32, and 33, or a sequence having 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 98%, 99% or 100% sequence identity to any one of SEQ ID NOs: 18, 27, 28, 32, and 33, and the second binding domain has or comprises a sequence complementary to the first binding domain, preferably a sequence that is at least 80%, 85%, 90%, 95%, 98%, 99% or 100% complementary to the first binding domain, more preferably a sequence that is at least 90%, 95%, 98%, 99% or 100% complementary to the first binding domain, and even more preferably a sequence that is 98%, 99% or 100% complementary to the first binding domain.

[0123] Furthermore, it has been shown that nucleotides 50-100 of SEQ ID NO: 27 or 28 are effective as a binding domain. Thus, in one embodiment, the second or first binding domain comprises a sequence of nucleotides 50-100 of SEQ ID NO: 27 or 28, or a sequence having 80%, 85%, 90%, 95%, 98%, 99% or 100% sequence identity to the sequence of nucleotides 50-100 of SEQ ID NO: 27 or 28. Preferably, the binding domain has at least 80 nucleotides comprising a sequence of nucleotides 50-100 of SEQ ID NO: 27 or 28, or a sequence having 80%, 85%, 90%, 95%, 98%, 99% or 100% sequence identity to the sequence of nucleotides 50-100 of SEQ ID NO: 27 or 28. One skilled in the art will understand that the other binding domain has or comprises a sequence which is complementary to the first or second binding domain, respectively, preferably at least 80%, 85%, 90%, 95%, 98%, 99% or 100% complementary to the first or second binding domain, respectively, more preferably at least 90%, 95%, 98%, 99% or 100% complementary to the first or second binding domain, respectively, even more preferably 98%, 99% or 100% complementary to the first or second binding domain, respectively.

[0124] In another embodiment, the second or first binding domain comprises nucleotides 1-50 of SEQ ID NO: 18, 32, or 33, or a sequence having 80%, 85%, 90%, 95%, 98%, 99% or 100% sequence identity to nucleotides 1-50 of SEQ ID NO: 18, 32, or 33. Preferably, the binding domain has at least 80 nucleotides comprising nucleotides 1-50 of SEQ ID NO: 18, 32, or 33, or a sequence having 80%, 85%, 90%, 95%, 98%, 99% or 100% sequence identity to nucleotides 1-50 of SEQ ID NO: 18, 32, or 33. One skilled in the art will understand that the other binding domain has or comprises a sequence which is complementary to the first or second binding domain, respectively, preferably at least 80%, 85%, 90%, 95%, 98%, 99% or 100% complementary to the first or second binding domain, respectively, more preferably at least 90%, 95%, 98%, 99% or 100% complementary to the first or second binding domain, respectively, even more preferably 98%, 99% or 100% complementary to the first or second binding domain, respectively.

[0125] The binding domain may have between about 15-250 nucleotides, between about 15-200 nucleotides, between about 100-200 nucleotides, or 500, 400, 300, or less. In one embodiment, the binding domain comprises 50-150 nucleotides, preferably 80-120 nucleotides, even more preferably 90-110 nucleotides, or about 100 nucleotides. The binding domains described herein can also be used in the pre-mRNA trans-splicing molecules or methods according to the invention.

[0126] The first and second binding domains used according to the invention may be derived from human or non-human sequences, such as bacterial sequences. For human use, such as in therapy (particularly gene therapy), non-human sequences are preferred to avoid off-target effects in human cells.

[0127] The second AAV vector and optionally the first AAV vector described herein may include a termination signal, such as a polyadenylation (polyA) signal located 3' to the nucleic acid sequence. PolyA signals / sequences are known to those of skill in the art and may be derived from many suitable species, including but not limited to SV-40, human and bovine. "PolyA" (A=adenylic acid) refers to a nucleic acid sequence that includes multiple adenosine monophosphates, such as a nucleic acid sequence that includes the AAUAAA consensus sequence that allows for polyadenylation of the processed transcript. In a gene disruption or selection cassette (GDSC), the polyA sequence is located downstream of the reporter and / or selectable marker gene and signals the end of the transcript to the RNA polymerase. It is envisioned that the AAV vector may include the polyA sequence of SEQ ID NO:16, or a sequence having at least 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 98% or 100% sequence identity to the sequence shown in SEQ ID NO:16.

[0128] The AAV vector or other vector of the present invention may optionally further comprise one or more transcription termination sequences, one or more translation termination sequences, one or more signal peptide sequences, one or more internal ribosome entry sites (IRES), and / or one or more enhancer elements, or any combination thereof. Transcription termination regions can usually be obtained from the 3'' untranslated region of eukaryotic or viral gene sequences. Transcription termination sequences can be placed downstream of the coding sequence to provide efficient termination. Signal peptide sequences are amino-terminal peptide sequences that code for information involved in one or more post-translational cellular destinations, including, for example, specific organelle compartments, or sites of protein synthesis and / or activity, as well as location of an operably linked polypeptide to the extracellular environment.

[0129] Enhancers - cis-acting regulatory elements that increase gene transcription - may be included in one of the disclosed AAV vectors or vectors. A variety of enhancer elements are known to those skilled in the art, including, but not limited to, CaMV 35S enhancer elements, cytomegalovirus (CMV) early promoter enhancer elements, SV40 enhancer elements, and combinations and / or derivatives thereof. One or more nucleic acid sequences that direct or regulate the polyadenylation of mRNA encoded by a structural gene of interest may also be optional in one or more vectors of the present invention.

[0130] The disclosed dual vector system may be introduced into one or more selected mammalian cells using any one or more of the methods known to those skilled in the art of gene therapy and / or viral technology. Such methods include, but are not limited to, transfection, microinjection, electroporation, lipofection, cell fusion, and calcium phosphate precipitation, as well as biolistics. In one embodiment, the vector of the present invention may be introduced in vivo, including, for example, lipofection (i.e., DNA transfection via liposomes prepared from one or more cationic lipids). Synthetic cationic lipids (LIPOFECTIN, Invitrogen Corp., La Jolla, Calif., USA) can be used to prepare liposomes that encapsulate the vector and facilitate its introduction into one or more selected cells. The vector system of the present invention may also be introduced in vivo as "naked" DNA, using methods known to those skilled in the art.

[0131] The present invention also relates to kits comprising the nucleic acid sequences and / or AAV vectors and / or AAV vector systems of the invention.

[0132] The present invention also relates to a method for producing a nucleic acid sequence of interest, said method comprising the steps of: (A) contacting a nucleic acid sequence with a host cell, the nucleic acid sequence comprising: (i) one or more donor splice site sequences; (ii) an acceptor splice region sequence comprising: (a) a pyrimidine tract containing: (aa) 5–25 nucleotides; (ab) wherein at least 60% of the nucleotides within these 5 to 25 nucleotides are pyrimidine bases, e.g., cytosine (C), thymine (T), and / or uracil (U); (b) the acceptor splice site; ba) wherein the acceptor splice site is located 3' to the pyrimidine tract; and bb) wherein the acceptor splice site comprises from 5' to 3' the sequence NAGG, where N is A, C, T / U, or G; (III) a nucleic acid sequence of interest or a portion thereof, said nucleotide sequence of interest or a portion thereof (a) is located 3' to the donor splice site and 5' to the acceptor splice region; and (B) cleaving the nucleic acid sequence at (i) one or more donor splice site sequences, and (ii) the acceptor splice region, thereby separating the nucleotide sequence of interest or a portion thereof from the donor splice site and the acceptor splice region.

[0133] As known to those skilled in the art, splicing can be performed in various ways. Known to those skilled in the art are, for example, cis-splicing and trans-splicing. Cis-splicing is the process of excising intron sequences from cis RNA transcripts (mRNA or other RNA). This process in host cells occurs in the nucleus. Cis-splicing is known to those skilled in the art and is described, inter alia, under the heading "DNA to RNA" in Alberts B, Johnson A, Lewis J, et al. (2002) "Molecular Biology of the Cell. 4th edition." New York: Garland Science.

[0134] In general, therapy or treatment or prevention of a disease involves administering to a mammalian subject in need thereof an effective amount of an AAV vector, AAV vector system, vector, or composition comprising a nucleic acid molecule as described herein, such as a pre-mRNA trans-splicing molecule carrying a nucleic acid sequence encoding a transgene / nucleotide sequence of interest or a fragment thereof under the control of a regulatory sequence that expresses a gene product in a target cell of the subject, such as an ocular cell, and optionally further a pharmaceutically acceptable carrier.

[0135] The present invention therefore also relates to a pharmaceutical composition comprising an AAV vector, an AAV vector system, a vector or a nucleic acid sequence (e.g., a pre-mRNA trans-splicing molecule) according to the invention. Such a pharmaceutical composition may further comprise a carrier, preferably a pharma- ceutically acceptable carrier.

[0136] The pharmaceutical composition may be in the form of an injectable solution. Injectable solutions or suspensions can be formulated according to known techniques using suitable non-toxic pharma- ceutically acceptable diluents or solvents, such as mannitol, 1,3-butanediol, water, Ringer's solution or isotonic sodium chloride solution, or suitable dispersing or wetting agents and suspending agents, such as sterile, aseptic, synthetic monoglycerides, synthetic diglycerides, fixed oils, including fatty acids, including oleic acid.

[0137] For injectable formulations, the pharmaceutical composition can be a lyophilized powder mixed with suitable excipients in a suitable vial or tube. Prior to use in the clinic, the drug can be reconstituted by dissolving the lyophilized powder in a suitable solvent system to form a composition suitable for intravenous or intramuscular injection, or subretinal, intravitreal or subconjunctival injection.

[0138] It is also contemplated that the pharmaceutical compositions of the present invention may be formulated / administered as eye drops.

[0139] The AAV vector, AAV vector system, vector, or nucleic acid sequence (e.g., pre-mRNA trans-splicing molecule) according to the present invention, or pharmaceutical composition of the present invention can be administered in a therapeutically effective amount. The "therapeutically effective amount" of the AAV vector, AAV vector system, nucleic acid sequence or vector, such as pre-mRNA trans-splicing molecule, can vary depending on factors including, but not limited to, the stability of the active compound in the patient's body, the severity of the condition to be alleviated, the total weight of the patient to be treated, the route of administration, the ease of absorption, distribution, and excretion of the active compound by the body, the age and sensitivity of the patient to be treated, adverse events, etc., as will be apparent to those skilled in the art.

[0140] In principle, the AAV vector, AAV vector system, vector, or nucleic acid sequence, such as a pre-mRNA trans-splicing molecule, of the present invention, or the pharmaceutical composition of the present invention can be administered in any suitable manner. For example, the AAV vector, AAV vector system, vector, nucleic acid sequence, or pharmaceutical composition of the present invention, which contains a desired transgene (i.e., a nucleotide sequence of interest) for targeting photoreceptor cells, can be formulated into a pharmaceutical composition for subretinal or intravitreal injection. Other forms of administration that may be useful in the methods described herein include, but are not limited to, direct delivery to the desired organ (e.g., the eye, e.g., eye drops), oral, inhalation, intranasal, intratracheal, intravenous, intramuscular, subcutaneous, intradermal, and other parenteral routes. Routes of administration can be combined as necessary.

[0141] Additionally, it may be desirable to perform non-invasive retinal imaging and functional studies to identify specific ocular cell regions to target for therapy.

[0142] The AAV vector, AAV vector system, vector, or nucleic acid sequence, such as a pre-mRNA trans-splicing molecule, described herein, or pharmaceutical composition of the invention can be administered to a subject in a physiologically acceptable carrier, as described herein. The concentration of AAV in the pharmaceutical composition or upon administration can be between 10E8 and 10E12 total vector genomes per μl, preferably between 6×10E8 and 6×0E10 per μl. The AAV vector can also be administered at a concentration of about 10E9 vector genomes per μl.

[0143] The AAV vector, AAV vector system, vector, or nucleic acid sequence (e.g., pre-mRNA trans-splicing molecule) or pharmaceutical composition of the invention may be administered with a carrier or in combination with other treatments. Thus, the pharmaceutical composition may also additionally or alternatively contain one or more further active ingredients.

[0144] The pharmaceutical composition, AAV vector, AAV vector system, vector, pre-mRNA trans-splicing molecule, or nucleic acid sequence for use in the present invention can be administered to a subject. The AAV vector, AAV vector system, vector, pre-mRNA trans-splicing molecule, nucleic acid sequence, or pharmaceutical composition described herein is applicable for both human therapy and veterinary use, preferably for the treatment of ocular diseases, in particular for gene therapy of ocular diseases. Examples of suitable ocular diseases are autosomal recessive severe early-onset retinal degeneration (Leber's congenital amaurosis), congenital color blindness, Stargardt's disease, Best's disease (vitelliform macular degeneration), Doyne's disease, and others. disease), retinitis pigmentosa (especially autosomal dominant, autosomal recessive, X-linked, digenic or polygenic retinitis pigmentosa), (X-linked) retinoschisis, macular degeneration (AMD), age-related macular degeneration, atrophic age-related macular degeneration, neovascular AMD, diabetic maculopathy, proliferative diabetic retinopathy (PDR), cystoid macular edema, central serous retinopathy, retinal detachment, endophthalmitis, glaucoma, posterior uveitis, congenital stationary night blindness, total choroidal atrophy, early-onset retinal dystrophies, cone dystrophies, rod-cone or cone-rod dystrophy, pattern dystrophies, Usher syndrome and other syndromic ciliary disorders ciliopathies, such as, but not limited to, Bardet-Biedl syndrome, Joubert syndrome, Senior-Loken syndrome or Alstrom syndrome.

[0145] The subject may be a mammal, or any other vertebrate. Examples of suitable mammals include, but are not limited to, mice, rats, cows, goats, sheep, pigs, dogs, cats, horses, guinea pigs, canines, hamsters, minks, seals, whales, camels, chimpanzees, rhesus monkeys, and humans, with humans being preferred. Examples of other vertebrates include, but are not limited to, zebrafish, salamanders, turkeys, chickens, geese, ducks, teals, mallards, starlings, pintails, seagulls, swans, guinea fowl, or waterfowl, to name a few.

[0146] The present invention also relates to an AAV vector, an AAV vector system, a vector, a nucleic acid sequence (e.g., a pre-mRNA trans-splicing molecule), or a pharmaceutical composition of the invention for use in treating a photoreceptor cell disease. In these embodiments, the AAV vector (e.g., rAAV) AAV vector, an AAV vector system, a vector, a nucleic acid sequence (e.g., a pre-mRNA trans-splicing molecule), or a pharmaceutical composition may comprise a nucleotide sequence of interest that is a heterologous nucleic acid encoding a therapeutic polypeptide or portion thereof, a therapeutic nucleic acid or portion thereof, a therapeutic protein / polypeptide or portion thereof, or a therapeutic molecule or portion thereof.

[0147] Similar to mRNA splicing, which evolved to remove non-coding RNA sequences, unwanted protein sequences can be excised during a process known as protein splicing. In this context, protein fragments called exteins (analogs of exons) can be spliced ​​together upon removal of so-called inteins (analogs of introns). In contrast to mRNA splicing, which requires a complex splicing machinery, protein splicing is an autocatalytic chemical process. Protein splicing can also occur between two different proteins; this is a process known as protein trans-splicing. For this purpose, an intein is split into two parts and each of these parts is tagged to the protein to be fused. This approach results in a flawless fusion of two proteins or polypeptides or parts of them without the need for additional factors. The split intein technology can be further combined with mRNA trans-splicing to further increase the efficiency of protein reconstitution. Thus, in one embodiment, the first nucleic acid sequence or first AAV vector further comprises a nucleic acid sequence encoding an N-terminal portion of an intein (or a 5' portion of the nucleic acid of interest) between nucleotide sequences encoding N-terminal portions of a polypeptide of interest and a donor splice site, and the second nucleic acid sequence or second AAV vector further comprises a nucleic acid sequence encoding a C-terminal portion of an intein (or a 3' portion of the nucleic acid of interest) between nucleotide sequences encoding C-terminal portions of a polypeptide of interest and an acceptor splice region.

[0148] It is clear that all possible embodiments described for the ASS of the present invention can be used mutatis mutandis in the methods, nucleic acid sequences such as pre-mRNA trans-splicing molecules, kits, AAV vectors, AAV vector systems and uses described herein.

[0149] The present invention is further characterized by the following:

[0150] 1. The use of a nucleic acid sequence for separating a nucleotide sequence of interest from an acceptor splice region sequence by cleavage of the acceptor splice region, comprising: (i) an acceptor splice region sequence comprising: (ia) a pyrimidine tract comprising: (iaa) 5–25 nucleotides; (iab) where at least 60% of the nucleotides within these 5 to 25 nucleotides are pyrimidine bases, e.g., cytosine (C), thymine (T), and / or uracil (U); (ib) acceptor splice site; ba) wherein the acceptor splice site is located 3' to the pyrimidine tract; and bb) wherein the acceptor splice site comprises from 5' to 3' the sequence NAGG, where N is A, C, T / U, or G; (ii) a nucleotide sequence of interest; wherein the nucleotide sequence of interest is (iia) It is located on the 3' or 5' side of the splice region.

[0151] 2. A nucleic acid sequence for separating a nucleotide sequence of interest from an acceptor splice region sequence, optionally by cleavage of the acceptor splice region, comprising: (i) an acceptor splice region sequence comprising: (ia) a pyrimidine tract comprising: (iaa) 5–25 nucleotides; (iab) where at least 60% of the nucleotides within these 5 to 25 nucleotides are pyrimidine bases, e.g., cytosine (C), thymine (T), and / or uracil (U); (ib) acceptor splice site; ba) wherein the acceptor splice site is located 3' to the pyrimidine tract; and bb) wherein the acceptor splice site comprises the sequence 5' to 3' of NAGG, where N is A, C, T / U, or G.

[0152] 3. Use of the nucleic acid sequence of item 1 or the nucleic acid sequence of item 2, wherein the acceptor splice region further comprises a branch point nucleotide, preferably an adenosine, and / or an intronic splice enhancer.

[0153] 4. The acceptor splice region further comprises a branch point nucleotide sequence (c), the branch point nucleotide sequence being (ca) contains 1 to 15 nucleotides; (cb) includes a branch point nucleotide, preferably an adenosine (A); and (cc) a pyrimidine tract and a 5' side of the acceptor splice site; 4. The use or nucleic acid sequence according to any one of items 1 to 3.

[0154] 5. The use of a nucleic acid sequence or a nucleic acid sequence according to any one of items 1 to 4, wherein the pyrimidine tract, the acceptor splice region and optionally further the branch point sequence and / or the intron splice enhancer comprise in total no more than about 200, 150, 1000, 500, 250, 100, 50, 45, 40, 35, 30, 25, 20, 15 nucleotides, preferably 26 nucleotides.

[0155] 6. (a) 5 to 25 nucleotides of the pyrimidine tract contain the sequence TTTTTT or TCTTTT; (b) the sequence between the last pyrimidine of the pyrimidine tract and the acceptor splice site is less than 10 bases, preferably less than 5 bases, more preferably less than 3 bases; and / or (c) the acceptor splice site has the sequence CAGG; 6. Use of a nucleic acid sequence or a nucleic acid sequence according to any one of items 1 to 5.

[0156] 7. The use of a nucleic acid sequence according to any one of items 1 to 6, wherein the acceptor splice region further comprises the following: (a) the 7 nucleotides 5' to the pyrimidine tract having at least 4 nucleotides of the sequence CAACGAG, where the first 5' nucleotide is a C; (b) the 7 nucleotides 5' to the pyrimidine tract having at least 5 nucleotides of the sequence CAACGAG, where the first 5' nucleotide is a C; (c) the 7 nucleotides 5' to the pyrimidine tract having at least 6 nucleotides of the sequence CAACGAG, where the first 5' nucleotide is a C; (d) the 7 nucleotides 5' to the pyrimidine tract having at least 4 nucleotides of the sequence CAACGAG, where the first 5' nucleotide is CAA; (e) the 7 nucleotides 5' to the pyrimidine tract having at least 4 nucleotides of the sequence CAACGAG, where the first 5' nucleotide is CAAC; (f) the 7 nucleotides 5' to a pyrimidine tract having at least 5 nucleotides of the sequence CAACGAG, where the first 5' nucleotide is CAA; (g) the 7 nucleotides 5' to a pyrimidine tract having at least 5 nucleotides of the sequence CAACGAG, where the first 5' nucleotide is CAAC; (h) the 7 nucleotides 5' to a pyrimidine tract having at least 6 nucleotides of the sequence CAACGAG, where the first 5' nucleotide is CAA; or (i) The 7 nucleotides 5' to the pyrimidine tract having at least 6 nucleotides of the sequence CAACGAG, where the first 5' nucleotide is CAAC.

[0157] 8. The use of a nucleic acid sequence or a nucleic acid sequence according to any one of items 1 to 7, wherein the acceptor splice region has the sequence of SEQ ID NO: 3 or 4.

[0158] 9. The use of a nucleic acid sequence or a nucleic acid sequence according to any one of items 1 to 8, wherein the nucleic acid sequence further comprises a donor splice site.

[0159] 10. The use of a nucleic acid sequence or a nucleic acid sequence according to any one of items 1 to 9, wherein the nucleic acid sequence further comprises a promoter, preferably wherein the nucleic acid sequence is a DNA sequence and further comprises a promoter.

[0160] 11. The use of a nucleic acid sequence according to any one of items 1 to 10, or a nucleic acid sequence comprising: (i) an acceptor splice region sequence comprising: (iia) a pyrimidine tract comprising: (iiaa) 5 to 25 nucleotides; (iiab) wherein at least 60% of the nucleotides within these 5 to 25 nucleotides are pyrimidine bases, e.g., cytosine (C), thymine (T), and / or uracil (U); (iib) acceptor splice site, (iiba) wherein the acceptor splice site is located 3' to the pyrimidine tract; and (iibb) wherein the acceptor splice site comprises a sequence 5' to 3' of NAGG, where N is A, C, T / U, or G; (ii) a nucleotide sequence of interest, wherein the acceptor splice region is located 3' or 5' to the nucleotide sequence of interest; (iii) a pre-mRNA-targeting binding domain located 3' or 5' to the nucleotide sequence of interest; and (iv) optionally a spacer sequence, where the spacer sequence is located between the binding domain and the acceptor splice region.

[0161] 12. The use of a nucleic acid sequence according to any one of items 1 to 11, or a nucleic acid sequence comprising: (i) an acceptor splice region sequence comprising: (ia) a pyrimidine tract comprising: (iiiaa) 5 to 25 nucleotides; (iiiab) wherein at least 60% of the nucleotides within these 5 to 25 nucleotides are pyrimidine bases, e.g., cytosine (C), thymine (T), and / or uracil (U); (ib) acceptor splice site; (iba) wherein the acceptor splice site is located 3' to the pyrimidine tract; and (ibb) wherein the acceptor splice site comprises a sequence 5' to 3' of NAGG, where N is A, C, T / U, or G; (ii) a nucleotide sequence of interest; wherein the acceptor splice region is located 5' to the nucleotide sequence of interest; (iii) a donor splice site located 3' to the nucleotide sequence of interest; (iv) a first binding domain that targets the pre-mRNA, located 5′ to the nucleic acid sequence of interest; (v) a second binding domain that targets the pre-mRNA, located 3′ to the nucleic acid sequence of interest; (vi) optionally a first spacer sequence, where the first spacer sequence is located between the first binding domain and the acceptor splice region; and (vii) optionally a second spacer sequence, where the second spacer sequence is located between the second binding domain and the acceptor splice region.

[0162] 13. An adeno-associated virus (AAV) vector comprising a nucleic acid sequence comprising at least two inverted terminal repeats and a nucleic acid sequence between the two inverted terminal repeats, wherein the nucleic acid sequence is 5' to 3' (i) the promoter; (ii) optionally a binding domain; (iii) an acceptor splice region sequence comprising: (a) a pyrimidine tract comprising: (aa) 5–25 nucleotides; (ab) wherein at least 60% of the nucleotides within these 5 to 25 nucleotides are pyrimidine bases, e.g., cytosine (C), thymine (T), and / or uracil (U); (b) the acceptor splice site; ba) wherein the acceptor splice site is located 3' to the pyrimidine tract; and bb) wherein the acceptor splice site comprises a sequence 5' to 3' of NAGG, where N is A, C, T / U, or G; (v) a nucleotide sequence of interest; (vi) optionally a termination sequence, preferably a polyA sequence; 13. The use of a nucleic acid sequence or a nucleic acid sequence according to any one of items 1 to 12, comprising:

[0163] 14. The use of a nucleic acid sequence or a nucleic acid sequence according to any one of items 1 to 13, wherein the acceptor splice region is located in a vector, in an AAV vector, or in a pre-mRNA trans-splicing molecule.

[0164] 15. The use of a nucleic acid sequence or a nucleic acid sequence according to any one of items 1 to 14, wherein the acceptor splice region is used for trans-splicing.

[0165] 16. A method for producing a nucleic acid sequence, comprising: (A) providing a first nucleic acid sequence comprising one or more donor splice site sequences; (B) providing a second nucleic acid sequence comprising: (i) an acceptor splice region sequence comprising: (ia) a pyrimidine tract comprising: (iaa) 5–25 nucleotides; (iab) where at least 60% of the nucleotides within these 5 to 25 nucleotides are pyrimidine bases, e.g., cytosine (C), thymine (T), and / or uracil (U); (ib) acceptor splice site; (iba) wherein the acceptor splice site is located 3' to the pyrimidine tract; and (ibb) wherein the acceptor splice site comprises a sequence 5' to 3' of NAGG, where N is A, C, T / U, or G; (ii) a nucleotide sequence of interest, where the nucleotide sequence of interest is (iia) located 3' to the acceptor splice region; and (C) Obtaining a nucleic acid sequence.

[0166] 17. The method according to item 16, further comprising the steps of: (D) cleaving the first nucleic acid sequence at the one or more donor splice site sequences and cleaving the second nucleic acid sequence at the acceptor splice site; (E) ligating the first cleaved nucleic acid sequence to the second cleaved nucleic acid sequence, thereby obtaining a nucleic acid sequence.

[0167] 18. The method of item 16 or 17, wherein the first nucleic acid sequence further comprises a nucleotide sequence of interest or a portion thereof, wherein at least a portion of the nucleotide sequence of interest is located 5' to the donor splice site, and wherein the second nucleic acid sequence further comprises a nucleotide sequence of interest or a portion thereof, wherein at least a portion of the nucleotide sequence of interest is located 3' to the splice acceptor splice region.

[0168] 19. The method according to any one of items 16 to 18, wherein the first and second nucleic acid sequences are introduced into a host cell, preferably the first and second nucleic acid sequences are recombinant nucleic acid sequences.

[0169] 20. A method for producing a nucleic acid sequence, comprising: (A) providing a first nucleic acid sequence comprising one or more donor splice site sequences; (B) providing a second nucleic acid sequence comprising: (i) an acceptor splice region sequence comprising: (ia) a pyrimidine tract comprising: (iaa) 5–25 nucleotides; (iab) where at least 60% of the nucleotides within these 5 to 25 nucleotides are pyrimidine bases, e.g., cytosine (C), thymine (T), and / or uracil (U); (ib) acceptor splice site; (iba) wherein the acceptor splice site is located 3' to the pyrimidine tract; and (ibb) wherein the acceptor splice site comprises a sequence 5' to 3' of NAGG, where N is A, C, T / U, or G; (C) cleaving the first nucleic acid sequence at one or more donor splice site sequences and cleaving the second nucleic acid sequence at an acceptor splice site; (D) ligating the first cleaved nucleic acid sequence to the second cleaved nucleic acid sequence, thereby obtaining a nucleic acid sequence.

[0170] 21. The method according to any one of items 16 to 20, wherein the first nucleic acid sequence further comprises a nucleotide sequence of interest or a portion thereof, wherein at least a portion of the nucleotide sequence of interest is located 5' to the donor splice site, and wherein the second nucleic acid sequence further comprises a nucleotide sequence of interest or a portion thereof, wherein at least a portion of the nucleotide sequence of interest is located 3' to the splice acceptor splice region.

[0171] 22. A method for producing a nucleic acid sequence or a method according to any one of items 16 to 21, comprising: (A) introducing into a host cell a first nucleic acid sequence comprising a pre-mRNA trans-splicing molecule sequence or a nucleic acid sequence encoding said pre-mRNA trans-splicing molecule, wherein the first pre-mRNA trans-splicing molecule comprises 5' to 3': (a) the 5' portion of the nucleotide acid sequence of interest; (b) donor splice site; (c) optionally a spacer sequence; (d) a first binding domain; and (e) optionally a termination sequence, preferably a polyA sequence, and (B) introducing into a host cell a second nucleic acid sequence comprising a pre-mRNA trans-splicing molecule sequence or a nucleic acid sequence encoding said pre-mRNA trans-splicing molecule, wherein the second pre-mRNA trans-splicing molecule comprises 5' to 3': (i) a second binding domain complementary to the first target domain of the first nucleic acid sequence; (ii) an acceptor splice region sequence comprising: (iia) a pyrimidine tract comprising: (iiaa) 5 to 25 nucleotides; (iiab) wherein at least 60% of the nucleotides within these 5 to 25 nucleotides are pyrimidine bases, e.g., cytosine (C), thymine (T), and / or uracil (U); (iib) acceptor splice site, (iiba) wherein the acceptor splice site is located 3' to the pyrimidine tract; and (iibb) the acceptor splice site comprises, 5' to 3', the sequence NAGG, where N is A, C, T / U, or G; (iii) a 3' portion of the nucleotide sequence of interest, and thereby obtaining a nucleic acid sequence of interest, optionally (C) cleaving the first nucleic acid sequence at the donor splice site sequence and cleaving the second nucleic acid sequence at the acceptor splice site; and (D) ligating the first cleaved nucleic acid sequence comprising the 5' portion of the nucleotide sequence of interest to a second cleaved nucleic acid sequence comprising the 3' portion of the nucleotide sequence of interest, thereby obtaining the nucleic acid sequence of interest; Further includes:

[0172] 23. (a) Nucleotides 5 to 25 of the pyrimidine tract contain the sequence TTTTTT or TCTTTT; (b) the sequence between the last pyrimidine of the pyrimidine tract and the acceptor splice site is less than 10 bases, preferably less than 5 bases, more preferably less than 3 bases; and / or (c) the acceptor splice site has the sequence CAGG; 23. The method according to any one of items 16 to 22.

[0173] 24. The method according to any one of items 16 to 23, wherein the acceptor splice region further comprises: (a) the 7 nucleotides 5' to the pyrimidine tract having at least 4 nucleotides of the sequence CAACGAG, where the first 5' nucleotide is a C; (b) the 7 nucleotides 5' to the pyrimidine tract having at least 5 nucleotides of the sequence CAACGAG, where the first 5' nucleotide is a C; (c) the 7 nucleotides 5' to the pyrimidine tract having at least 6 nucleotides of the sequence CAACGAG, where the first 5' nucleotide is a C; (d) the 7 nucleotides 5' to the pyrimidine tract having at least 4 nucleotides of the sequence CAACGAG, where the first 5' nucleotide is CAA; (e) the 7 nucleotides 5' to the pyrimidine tract having at least 4 nucleotides of the sequence CAACGAG, where the first 5' nucleotide is CAAC; (f) the 7 nucleotides 5' to a pyrimidine tract having at least 5 nucleotides of the sequence CAACGAG, where the first 5' nucleotide is CAA; (g) the 7 nucleotides 5' to a pyrimidine tract having at least 5 nucleotides of the sequence CAACGAG, where the first 5' nucleotide is CAAC; (h) the 7 nucleotides 5' to a pyrimidine tract having at least 6 nucleotides of the sequence CAACGAG, where the first 5' nucleotide is CAA; or (i) The 7 nucleotides 5' to the pyrimidine tract having at least 6 nucleotides of the sequence CAACGAG, where the first 5' nucleotide is CAAC.

[0174] 25. A nucleic acid sequence comprising: (i) an acceptor splice region sequence comprising: (ia) a pyrimidine tract comprising: (iaa) 5–25 nucleotides; (iab) where at least 60% of the nucleotides within these 5 to 25 nucleotides are pyrimidine bases, e.g., cytosine (C), thymine (T), and / or uracil (U); (ib) acceptor splice site; (iba) wherein the acceptor splice site is located 3' to the pyrimidine tract; and (ibb) wherein the acceptor splice site comprises a sequence 5' to 3' of NAGG, where N is A, C, T / U, or G; (ii) a nucleotide sequence of interest; wherein the nucleotide sequence of interest is (iia) It is located 3' or 5' to the acceptor splice region.

[0175] 26. Includes: (i) an acceptor splice region sequence comprising: (ia) a pyrimidine tract comprising: (iaa) 5–25 nucleotides; (iab) where at least 60% of the nucleotides within these 5 to 25 nucleotides are pyrimidine bases, e.g., cytosine (C), thymine (T), and / or uracil (U); (ib) acceptor splice site; (iba) wherein the acceptor splice site is located 3' to the pyrimidine tract; and (ibb) wherein the acceptor splice site comprises a sequence 5' to 3' of NAGG, where N is A, C, T / U, or G; (ii) a nucleotide sequence of interest; wherein the nucleotide sequence of interest is (iia) located 3′ or 5′ to the acceptor splice region; And optionally, wherein the nucleotide sequence of interest is not the rhodopsin gene of SEQ ID NO:10, and / or is not exon 3 of the rhodopsin gene of SEQ ID NO:9, and / or is not exon 3 of the rhodopsin gene of SEQ ID NO:9, and / or wherein the nucleotide sequence of interest does not include the sequence shown in SEQ ID NO:9 and / or SEQ ID NO:10.

[0176] 27. The nucleic acid sequence according to item 25 or 26, wherein the nucleic acid sequence has a length of up to 150 nucleotides.

[0177] 28. The nucleic acid sequence according to any one of items 25 to 27, wherein the nucleic acid sequence has a length of up to 5500 nucleotides.

[0178] 29. The nucleic acid sequence according to any one of items 25 to 28, wherein (a) 5 to 25 nucleotides of the pyrimidine tract comprise the sequence TTTTTT or TCTTTT, and / or (b) the sequence between the last pyrimidine of the pyrimidine tract and the acceptor splice site is less than 10 bases, preferably less than 5 bases, more preferably less than 3 bases.

[0179] 30. An acceptor splice region sequence comprising: (ia) a pyrimidine tract comprising: (iaa) 5–25 nucleotides; (iab) where at least 60% of the nucleotides within these 5 to 25 nucleotides are pyrimidine bases, e.g., cytosine (C), thymine (T), and / or uracil (U); (ib) acceptor splice site; (iba) wherein the acceptor splice site is located 3' to the pyrimidine tract; and (ibb) wherein the acceptor splice site comprises the sequence 5' to 3' of NAGG, where N is A, C, T / U, or G.

[0180] 31. (a) Nucleotides 5 to 25 of the pyrimidine tract contain the sequence TTTTTT or TCTTTT; (b) the sequence between the last pyrimidine of the pyrimidine tract and the acceptor splice site is less than 10 bases, preferably less than 5 bases, more preferably less than 3 bases; and / or (c) the acceptor splice site has the sequence CAGG; 31. The nucleic acid or acceptor splice region sequence according to any one of items 25 to 30.

[0181] 32. The nucleic acid or the acceptor splice region sequence according to any one of items 25 to 31, wherein the acceptor splice region further comprises: (a) the 7 nucleotides 5' to the pyrimidine tract having at least 4 nucleotides of the sequence CAACGAG, where the first 5' nucleotide is a C; (b) the 7 nucleotides 5' to the pyrimidine tract having at least 5 nucleotides of the sequence CAACGAG, where the first 5' nucleotide is a C; (c) the 7 nucleotides 5' to the pyrimidine tract having at least 6 nucleotides of the sequence CAACGAG, where the first 5' nucleotide is a C; (d) the 7 nucleotides 5' to the pyrimidine tract having at least 4 nucleotides of the sequence CAACGAG, where the first 5' nucleotide is CAA; (e) the 7 nucleotides 5' to the pyrimidine tract having at least 4 nucleotides of the sequence CAACGAG, where the first 5' nucleotide is CAAC; (f) the 7 nucleotides 5' to a pyrimidine tract having at least 5 nucleotides of the sequence CAACGAG, where the first 5' nucleotide is CAA; (g) the 7 nucleotides 5' to a pyrimidine tract having at least 5 nucleotides of the sequence CAACGAG, where the first 5' nucleotide is CAAC; (h) the 7 nucleotides 5' to a pyrimidine tract having at least 6 nucleotides of the sequence CAACGAG, where the first 5' nucleotide is CAA; or (i) The 7 nucleotides 5' to the pyrimidine tract having at least 6 nucleotides of the sequence CAACGAG, where the first 5' nucleotide is CAAC.

[0182] 33. A pre-mRNA trans-splicing molecule comprising: (i) an acceptor splice region comprising: (iia) a pyrimidine tract comprising: (iiaa) 5 to 25 nucleotides; (iiab) wherein at least 60% of the nucleotides within these 5 to 25 nucleotides are pyrimidine bases, e.g., cytosine (C), thymine (T), and / or uracil (U); (iib) acceptor splice site, (iiba) wherein the acceptor splice site is located 3' to the pyrimidine tract; and (iibb) wherein the acceptor splice site comprises a sequence 5' to 3' of NAGG, where N is A, C, T / U, or G; (ii) a nucleotide sequence of interest or a portion thereof; wherein the acceptor splice region is located 3' or 5' to the nucleotide sequence of interest or portion thereof; (iii) a binding domain that targets a pre-mRNA located 5' to a nucleic acid sequence of interest or a portion thereof; and (iv) optionally a spacer sequence, wherein the spacer sequence is located between the binding domain and the acceptor splice region.

[0183] 34. The pre-mRNA trans-splicing molecule of item 33, further comprising a donor splice site 3' to the nucleic acid molecule and the acceptor splice region.

[0184] 35. A pre-mRNA trans-splicing molecule according to item 33 or 34, comprising: (i) an acceptor splice region comprising: (ia) a pyrimidine tract comprising: (iiiaa) 5 to 25 nucleotides; (iiiab) wherein at least 60% of the nucleotides within these 5 to 25 nucleotides are pyrimidine bases, e.g., cytosine (C), thymine (T), and / or uracil (U); (ib) acceptor splice site; (iba) wherein the acceptor splice site is located 3' to the pyrimidine tract; and (ibb) wherein the acceptor splice site comprises a sequence 5' to 3' of NAGG, where N is A, C, T / U, or G; (ii) a nucleotide sequence of interest, where the acceptor splice region is located 5′ to the nucleotide sequence of interest; (iii) a donor splice site located 3' to the nucleotide sequence of interest; (iv) a first binding domain that targets the pre-mRNA located 5′ to the nucleotide sequence of interest; (v) a second binding domain that targets the pre-mRNA located 3′ to the nucleotide sequence of interest; (vi) optionally a first spacer sequence, where the first spacer is located between the first binding domain and the acceptor splice region; (vii) optionally a second spacer sequence, wherein the second spacer is located between the second binding domain and the acceptor splice region.

[0185] 36. The pre-mRNA trans-splicing molecule according to any one of items 33 to 35, wherein the acceptor splice region is located 5' to the nucleotide sequence of interest or a portion thereof, and the binding domain is located 5' to the acceptor splice region.

[0186] 37. The pre-mRNA trans-splicing molecule according to any one of items 33 to 36, further comprising a termination sequence, preferably a polyA sequence.

[0187] 38. (a) Nucleotides 5 to 25 of the pyrimidine tract contain the sequence TTTTTT or TCTTTT; (b) the sequence between the last pyrimidine of the pyrimidine tract and the acceptor splice site is less than 10 bases, preferably less than 5 bases, more preferably less than 3 bases; and / or (c) the acceptor splice site has the sequence CAGG; 38. The pre-mRNA trans-splicing molecule according to any one of items 33 to 37.

[0188] 39. The pre-mRNA trans-splicing molecule according to any one of items 33 to 38, wherein the acceptor splice region further comprises: (a) the 7 nucleotides 5' to the pyrimidine tract having at least 4 nucleotides of the sequence CAACGAG, where the first 5' nucleotide is a C; (b) the 7 nucleotides 5' to the pyrimidine tract having at least 5 nucleotides of the sequence CAACGAG, where the first 5' nucleotide is a C; (c) the 7 nucleotides 5' to the pyrimidine tract having at least 6 nucleotides of the sequence CAACGAG, where the first 5' nucleotide is a C; (d) the 7 nucleotides 5' to the pyrimidine tract having at least 4 nucleotides of the sequence CAACGAG, where the first 5' nucleotide is CAA; (e) the 7 nucleotides 5' to the pyrimidine tract having at least 4 nucleotides of the sequence CAACGAG, where the first 5' nucleotide is CAAC; (f) the 7 nucleotides 5' to a pyrimidine tract having at least 5 nucleotides of the sequence CAACGAG, where the first 5' nucleotide is CAA; (g) the 7 nucleotides 5' to a pyrimidine tract having at least 5 nucleotides of the sequence CAACGAG, where the first 5' nucleotide is CAAC; (h) the 7 nucleotides 5' to a pyrimidine tract having at least 6 nucleotides of the sequence CAACGAG, where the first 5' nucleotide is CAA; or (i) The 7 nucleotides 5' to the pyrimidine tract having at least 6 nucleotides of the sequence CAACGAG, where the first 5' nucleotide is CAAC.

[0189] 40. A DNA molecule comprising a promoter and a sequence encoding the pre-mRNA trans-splicing molecule according to any one of items 33 to 39.

[0190] 41. Adeno-associated virus (AAV) vector systems, including: (I) a first AAV vector comprising at least two inverted terminal repeats, the first AAV vector comprising a nucleic acid sequence between the two inverted terminal repeats, the nucleic acid sequence between the two inverted terminal repeats being arranged 5' to 3' from the first AAV vector to the second AAV vector; (a) a promoter; (b) a nucleotide sequence encoding the N-terminal portion of the polypeptide of interest; (c) donor splice site; (d) optionally a spacer sequence; (e) a first binding domain; and (f) optionally a termination sequence, preferably a polyA sequence; A first AAV vector comprising: (II) a second AAV vector comprising a nucleic acid sequence comprising at least two inverted terminal repeats, the nucleic acid sequence between the two inverted terminal repeats being arranged 5' to 3' from the first AAV vector to the second AAV vector, (i) the promoter; (ii) a second binding domain that is complementary to the first binding domain of the first AAV vector; (iii) an acceptor splice region sequence comprising: (a) a pyrimidine tract comprising: (aa) 5–25 nucleotides; (ab) wherein at least 60% of the nucleotides within these 5 to 25 nucleotides are pyrimidine bases, e.g., cytosine (C), thymine (T), and / or uracil (U); (b) the acceptor splice site; (ba) wherein the acceptor splice site is located 3' to the pyrimidine tract; and (bb) wherein the acceptor splice site comprises a sequence 5' to 3' of NAGG, where N is A, C, T / U, or G; (iv) a nucleotide sequence of interest encoding the C-terminal portion of a polypeptide of interest; (iva) wherein the C-terminal portion of the polypeptide of interest and the N-terminal portion of the polypeptide of interest reconstitute the polypeptide of interest; and (v) a termination sequence, preferably a polyA sequence.

[0191] 42. The AAV vector system of item 41, wherein the polypeptide is a full-length polypeptide, and the first AAV vector comprises an N-terminal portion of the full-length polypeptide of interest, and the second AAV vector comprises a C-terminal portion of the full-length polypeptide of interest.

[0192] 43. The AAV vector system of item 41 or 42, wherein the C-terminal portion of the polypeptide of interest corresponds to a portion of the polypeptide of interest, optionally a full-length protein of interest, that is deleted from the N-terminal portion contained in the first AAV vector.

[0193] 44. The AAV vector system according to any one of items 41 to 43, wherein the first and second AAV vectors comprise a termination sequence, preferably a polyA sequence.

[0194] 45. An adeno-associated virus (AAV) vector comprising at least two inverted terminal repeats, the nucleic acid sequence being between the two inverted terminal repeats, the nucleic acid sequence being 5' to 3' (i) the promoter; (ii) optionally a spacer sequence; (III) a binding domain; (iv) an acceptor splice region sequence, comprising: (a) a pyrimidine tract comprising: (aa) 5–25 nucleotides; (ab) wherein at least 60% of the nucleotides within these 5 to 25 nucleotides are pyrimidine bases, e.g., cytosine (C), thymine (T), and / or uracil (U); (b) the acceptor splice site; ba) wherein the acceptor splice site is located 3' to the pyrimidine tract; and bb) wherein the acceptor splice site comprises from 5' to 3' the sequence NAGG, where N is A, C, T / U, or G; (iv) a nucleotide sequence of interest; and (v) optionally a termination sequence, preferably a polyA sequence; An adeno-associated virus (AAV) vector comprising:

[0195] 46. ​​(a) Nucleotides 5 to 25 of the pyrimidine tract contain the sequence TTTTTT or TCTTTT; (b) the sequence between the last pyrimidine of the pyrimidine tract and the acceptor splice site is less than 10 bases, preferably less than 5 bases, more preferably less than 3 bases; and / or (c) the acceptor splice site has the sequence CAGG; 46. ​​The AAV vector or AAV vector system according to any one of items 41 to 45.

[0196] 47. The AAV vector or AAV vector system according to any one of items 41 to 46, wherein the acceptor splice region further comprises: (a) the 7 nucleotides 5' to the pyrimidine tract having at least 4 nucleotides of the sequence CAACGAG, where the first 5' nucleotide is a C; (b) the 7 nucleotides 5' to the pyrimidine tract having at least 5 nucleotides of the sequence CAACGAG, where the first 5' nucleotide is a C; (c) the 7 nucleotides 5' to the pyrimidine tract having at least 6 nucleotides of the sequence CAACGAG, where the first 5' nucleotide is a C; (d) the 7 nucleotides 5' to the pyrimidine tract having at least 4 nucleotides of the sequence CAACGAG, where the first 5' nucleotide is CAA; (e) the 7 nucleotides 5' to the pyrimidine tract having at least 4 nucleotides of the sequence CAACGAG, where the first 5' nucleotide is CAAC; (f) the 7 nucleotides 5' to a pyrimidine tract having at least 5 nucleotides of the sequence CAACGAG, where the first 5' nucleotide is CAA; (g) the 7 nucleotides 5' to a pyrimidine tract having at least 5 nucleotides of the sequence CAACGAG, where the first 5' nucleotide is CAAC; (h) the 7 nucleotides 5' to a pyrimidine tract having at least 6 nucleotides of the sequence CAACGAG, where the first 5' nucleotide is CAA; or (i) The 7 nucleotides 5' to the pyrimidine tract having at least 6 nucleotides of the sequence CAACGAG, where the first 5' nucleotide is CAAC.

[0197] 48. An acceptor splice region, comprising: (ia) a pyrimidine tract comprising: (iaa) 5–25 nucleotides; (iab) where at least 60% of the nucleotides within these 5 to 25 nucleotides are pyrimidine bases, e.g., cytosine (C), thymine (T), and / or uracil (U); (ib) acceptor splice site; (iba) wherein the acceptor splice site is located 3' to the pyrimidine tract; and (ibb) wherein the acceptor splice site comprises the sequence 5' to 3' of NAGG, where N is A, C, T / U, or G.

[0198] 49. An acceptor splice region according to any one of items 1 to 48 for use in a vector, an AAV vector, or a pre-mRNA trans-splicing molecule according to the invention.

[0199] 50. (a) 5 to 25 nucleotides of the pyrimidine tract contain the sequence TTTTTT or TCTTTT; (b) the sequence between the last pyrimidine of the pyrimidine tract and the acceptor splice site is less than 10 bases, preferably less than 5 bases, more preferably less than 3 bases; and / or (c) the acceptor splice site has the sequence CAGG; 50. The acceptor splice region sequence according to any one of items 48 or 49.

[0200] 51. The acceptable splice region sequence according to any one of items 48 to 50, wherein the acceptable splice region further comprises: (a) the 7 nucleotides 5' to the pyrimidine tract having at least 4 nucleotides of the sequence CAACGAG, where the first 5' nucleotide is a C; (b) the 7 nucleotides 5' to the pyrimidine tract having at least 5 nucleotides of the sequence CAACGAG, where the first 5' nucleotide is a C; (c) the 7 nucleotides 5' to the pyrimidine tract having at least 6 nucleotides of the sequence CAACGAG, where the first 5' nucleotide is a C; (d) the 7 nucleotides 5' to the pyrimidine tract having at least 4 nucleotides of the sequence CAACGAG, where the first 5' nucleotide is CAA; (e) the 7 nucleotides 5' to the pyrimidine tract having at least 4 nucleotides of the sequence CAACGAG, where the first 5' nucleotide is CAAC; (f) the 7 nucleotides 5' to a pyrimidine tract having at least 5 nucleotides of the sequence CAACGAG, where the first 5' nucleotide is CAA; (g) the 7 nucleotides 5' to a pyrimidine tract having at least 5 nucleotides of the sequence CAACGAG, where the first 5' nucleotide is CAAC; (h) the 7 nucleotides 5' to a pyrimidine tract having at least 6 nucleotides of the sequence CAACGAG, where the first 5' nucleotide is CAA; or (i) The 7 nucleotides 5' to the pyrimidine tract having at least 6 nucleotides of the sequence CAACGAG, where the first 5' nucleotide is CAAC.

[0201] 52. A kit comprising an acceptor splice region, an AAV vector, an AAV vector system, a vector, DNA, a nucleic acid sequence, or a pre-mRNA trans-splicing molecule according to the present invention.

[0202] 53. A method for producing a nucleic acid sequence of interest, comprising: (A) providing a nucleic acid sequence to a host cell, the nucleic acid sequence comprising: (i) one or more donor splice site sequences; (ii) an acceptor splice region sequence comprising: (a) a pyrimidine tract comprising: (aa) 5–25 nucleotides; (ab) wherein at least 60% of the nucleotides within these 5 to 25 nucleotides are pyrimidine bases, e.g., cytosine (C), thymine (T), and / or uracil (U); (b) the acceptor splice site; ba) wherein the acceptor splice site is located 3' to the pyrimidine tract; and bb) wherein the acceptor splice site comprises a sequence 5' to 3' of NAGG, where N is A, C, T / U, or G. (iii) a nucleotide sequence of interest, wherein the nucleotide sequence of interest; (a) is located 3' to the donor splice site and 5' to the acceptor splice region; and (B)(i) cleaving the nucleic acid sequence at one or more donor splice site sequences and between the G and G of the NAGG splice acceptor sequence (site), thereby separating the nucleotide sequence of interest from the donor splice site and the acceptor splice region.

[0203] 54. A method for producing / cleaving a nucleic acid sequence of interest, comprising: (a) providing a nucleic acid sequence comprising: (i) an acceptor splice region sequence comprising: (ia) a pyrimidine tract comprising: (aa) 5–25 nucleotides; (ab) wherein at least 60% of the nucleotides within these 5 to 25 nucleotides are pyrimidine bases, e.g., cytosine (C), thymine (T), and / or uracil (U); (ib) acceptor splice site; ba) wherein the acceptor splice site is located 3' to the pyrimidine tract; and bb) wherein the acceptor splice site comprises a sequence 5' to 3' of NAGG, where N is A, C, T / U, or G; and (ii) a nucleotide sequence of interest, where the nucleotide sequence of interest is (iia) located 3′ or 5′ to the splice region; (b) cleaving the acceptable splice region, thereby separating the nucleotide sequence of interest from the acceptable splice region sequence.

[0204] 55. A pharmaceutical composition comprising the nucleic acid molecule, pre-mRNA trans-splicing molecule, AAV vector, AAV vector system, or vector according to any of items 1 to 54.

[0205] 56. The nucleic acid molecule, pre-mRNA trans-splicing molecule, AAV vector, AAV vector system, vector, DNA molecule, or pharmaceutical composition according to any of items 1 to 54 for use in treatment, preferably in the treatment of an eye disease.

[0206] 57. A method of treating a subject for a nucleic acid sequence of interest, comprising: (a) administering to a subject (in need thereof) a (therapeutically effective amount) of the nucleic acid molecule, pre-mRNA trans-splicing molecule, AAV vector, AAV vector system, vector, DNA molecule or pharmaceutical composition according to any of items 1 to 56.

[0207] 58. A nucleic acid molecule, a pre-mRNA trans-splicing molecule, an AAV vector, an AAV vector system, a vector, a DNA molecule, or a pharmaceutical composition according to any of items 1 to 57 for the manufacture of a medicament.

[0208] 59. Eye diseases include autosomal recessive severe early-onset retinal degeneration (Leber's congenital amaurosis), congenital color blindness, Stargardt's disease, Best's disease (vitelliform macular degeneration), Doyne's disease, and disease), retinitis pigmentosa (especially autosomal dominant, autosomal recessive, X-linked, digenic or polygenic retinitis pigmentosa), (X-linked) retinoschisis, macular degeneration (AMD), age-related macular degeneration, atrophic age-related macular degeneration, neovascular AMD, diabetic maculopathy, proliferative diabetic retinopathy (PDR), cystoid macular edema, central serous retinopathy, retinal detachment, endophthalmitis, glaucoma, posterior uveitis, congenital stationary night blindness, total choroidal atrophy, early-onset retinal dystrophies, cone dystrophies, rod-cone or cone-rod dystrophy, pattern dystrophies, Usher syndrome and other syndromic ciliary disorders 59. The nucleic acid molecule, pre-mRNA trans-splicing molecule, AAV vector, AAV vector system, vector, DNA molecule, pharmaceutical composition, or method for use according to any of items 1 to 58, wherein the nucleic acid molecule, pre-mRNA trans-splicing molecule, AAV vector, AAV vector system, vector, DNA molecule, pharmaceutical composition, or method is selected from the group consisting of a pulmonary ...

[0209] It should be noted that, as used herein, the singular forms "a," "an," and "the" include plural references unless the context clearly dictates otherwise. Thus, for example, reference to a "reagent" includes one or more of such different reagents, and reference to a "method" includes reference to equivalent steps or methods known to those of skill in the art that can be modified or substituted for the method described herein.

[0210] Additionally, the term "about" as used herein when referring to a measurable value, such as an amount or length of a polynucleotide or polypeptide sequence, dosage, time, temperature, and the like, is meant to encompass variations of ±20%, ±10%, ±5%, ±1%, ±0.5%, or ±0.1% of the specified amount. Also, as used herein, "and / or" refers to and encompasses the associated listed items, as well as the lack of combination when interpreted in the alternative ("or").

[0211] It is specifically contemplated that the various features of the invention described herein can be used in any combination, unless the context clearly indicates otherwise.

[0212] Moreover, the present invention also contemplates that in some embodiments of the invention, any feature or combination of features described herein can be excluded or omitted.

[0213] All publications and patents cited in this disclosure are incorporated by reference in their entirety. In the event that the material incorporated by reference conflicts or is inconsistent with the present specification, the present specification controls over such material.

[0214] Unless otherwise specified, the term "at least" preceding a series of elements should be understood to refer to every element of the series. Those skilled in the art will recognize, or be able to ascertain using no more than routine experimentation, many equivalents to the specific embodiments of the invention described herein. Such equivalents are intended to be encompassed by the present invention.

[0215] Throughout this specification and the claims that follow, unless the context clearly dictates otherwise, the word "comprise", and variations such as "comprises" and "comprising", are understood to mean the inclusion of a recited integer or step or group of integers or steps, but not the exclusion of any other integer or step or group of integers or steps. As used herein, the term "comprising" can be replaced with the term "containing" or, as used herein, with the term "having". However, as used herein, "comprising of" or equivalent terms encompass "consisting of" or "consisting essentially of", as defined below, and thus can be replaced with the terms "consisting of" or "consisting essentially of".

[0216] As used herein, "comprising of" excludes any element, step, or ingredient not specified in the claim element. As used herein, "consisting essentially of" does not exclude materials or steps that do not materially affect the basic and novel characteristics of the claim.

[0217] Sequences used in the present invention:

[0218] The following sequences of more than 10 nucleotides are referred to herein:

[0219] [Table 1-1] [Table 1-2] [Table 1-3] [Table 1-4] Table 1: Putative branch point sequences are underlined, branch points are underlined twice and in bold. Pyrimidine tract sequences are in italics and the acceptor splice site is in bold. Underlined in SEQ ID NOs:9 and 10 are the mutations from 620T>G and the start and stop codons. EXAMPLES

[0220] The following examples illustrate the invention. These examples should not be construed as limiting the scope of the invention. The examples are included for illustrative purposes, and the invention is limited only by the claims.

[0221] material and method Cell culture and transfection Human Embryonic Kidney 293 (HEK293) cells (DMSZ) were maintained in DMEM medium + GlutaMAX + 1 g / l glucose + pyruvate + 10% FBS (Biochrom) + 1% penicillin / streptomycin (Biochrom) in a CO2 incubator (Heraeus, Thermo Fisher Scientific) at 37 °C and 10% CO2. HEK293-derived Lenti-X 293T (HEK293T) cells (Clontech, Takara) were cultured in DMEM medium + GlutaMAX + 4.5 g / l glucose + 10% FBS + 1% penicillin / streptomycin under identical conditions. Both cell lines were passaged twice a week at approximately 90% confluence.

[0222] Transient transfections were performed using the calcium phosphate method. For this purpose, cells were seeded in 6 cm cell culture plates. When transfecting plasmids containing SpCas9 for Western blotting, 10 cm cell culture plates were used. Cells were incubated overnight until they reached the desired confluence of approximately 70%. Transfection mix components were added to a 15 ml Falcon tube in the order indicated. 2x BBS was added dropwise while vortexing.

[0223] The transfection mix was incubated at room temperature for 3-4 min and added dropwise to the medium. In the initial experiment (Example 2, Figure 4), cells were incubated at 5% CO2 for 24 h and cells were harvested without changing the medium. In the optimized protocol (Example 3, 5-8; Figures 9 and 11-14), cells were incubated at 5% CO2 for 3-4 h, medium was changed and cells were maintained at 10% CO2 for approximately 48 h. When plasmids containing fluorophores were transfected, the success of transfection and expression was assessed via an EVOS® FL cell imaging system (Life Technologies, Thermo Fisher Scientific).

[0224] 661W cells were provided by Prof. Muayyad Al-Ubaidi (University of Houston). This cell line was cloned from a mouse retinal tumor and was found to exhibit molecular characteristics of cone photoreceptors (al-Ubaidi et al., 1992, Tan et al., 2004). 661W cells were maintained in DMEM medium + GlutaMAX + 1g / l glucose + pyruvate + 10% FBS (Biochrom) + 1% antibiotic-antimycotic in a CO2 incubator (Heraeus, Thermo Fisher Scientific) at 37°C and 10% CO2. They were passaged twice a week at approximately 90% confluence. Transient transfections were performed using the calcium phosphate method described above.

[0225] Mouse embryonic fibroblasts (MEFs) were generated as described (Jat et al., 1986, Xu, 2005). Cells were maintained in DMEM medium + GlutaMAX + 1g / l glucose + pyruvate + 10% FBS (Biochrom) + 1% penicillin / streptomycin (Biochrom) at 37°C and 5% CO2 in a CO2 incubator (Heraeus, Thermo Fisher Scientific). They were passaged once a week at approximately 90% confluence. MEFs were transiently transfected using TurboFectTM transfection reagent (Thermo Fisher Scientific). Cells were seeded in 6 cm cell culture plates and incubated until they reached 70–90% confluence. Reaction mixtures were prepared in the following order:

[0226] After the addition of each component, the solution was mixed vigorously by vortexing. The transfection mix was incubated at room temperature for 15-20 min and then added dropwise to the culture plate. The medium was changed after 3 h, and the cells were harvested 48 h after transfection. When plasmids containing fluorophores were transfected, the success of transfection and expression was assessed via an EVOS® FL cell imaging system (Life Technologies, Thermo Fisher Scientific).

[0227] Production of recombinant adeno-associated viruses Recombinant adeno-associated viruses (rAAV) were produced by triple calcium phosphate transfection of pAAV2.1 plasmids containing the gene of interest, pAD; a helper plasmid and a plasmid encoding the desired capsid. For subretinal injection into mouse retina, the 2 / 8Y733F capsid variant was chosen due to its high transduction efficiency of photoreceptors and RPE (Petrs-Silva et al., 2009, Mol. Ther, 17, 463-71). HEK293T cells were seeded in 15 x 15 cm cell culture plates and incubated overnight until they reached 60-80% confluence. Prior to transfection, the medium containing FBS was replaced with serum-free medium. The transfection reagents were added to a 50 ml Falcon tube in the order shown: 270 μg pAAV2.1 plasmid, X μg pAD helper plasmid, Y μg capsid plasmid, 11.85 ml H2O ad, 15 μl polybrene (8 mg / ml), 1.5 ml dextran (10 mg / ml), 1.5 ml CaCl2 (2.5 M), 15 ml 2×BBS. The required amount of pAD helper and capsid plasmid was calculated using the following formula: X μg = 270 μg × MM of pAD helper MM of pAAV2.1 Y μg = 270 μg × MM of capsid plasmid MM of pAAV2.1. CaCl2 and 2×BBS were added dropwise while vortexing. 2 ml of the transfection mix was added dropwise to each of the 15 culture plates. The plates were gently rocked and then placed in a 5% CO2 setting for 24 hours, after which the medium was changed and the plates were placed in a 10% CO2 setting for an additional 48 hours.

[0228] The virus-containing medium was harvested twice. The first harvest was performed 72 h after transfection by collecting the entire medium of all plates and adding fresh medium. The second harvest was performed after another 72 h incubation period. The medium was collected in a 500 ml centrifuge tube. Residual cells were removed from the medium by centrifugation at 4,000 rpm at 4 °C for 15 min (JA-10 rotor, J2-MC high-speed centrifuge, Beckman Coulter) and filtering the supernatant through a 0.45 μm PES filter unit (Nalgene, Thermo Fisher Scientific). A 40% polyethylene glycol (PEG) solution was added to the flow-through to a final concentration of 8% and kept overnight at 4 °C to precipitate the virus particles. The solution was subsequently centrifuged at 4,000 rpm at 4 °C for 15 min (JA-10 rotor, J2-MC high-speed centrifuge, Beckman Coulter). The supernatant was discarded and the virus-containing pellet was stored at -20 °C until further processing.

[0229] For iodixanol density gradient centrifugation, the pellet was resuspended in 7.5 ml of sterile filtered PBS and incubated for 30 min with Benzonase® (VWR) at a final concentration of 50 U / ml and a 37°C water bath (Haake) to remove residual unpackaged DNA. The virus suspension was then pipetted into quick seal polypropylene tubes (39 ml, Beckman Coulter) and a density gradient was established by adding solutions with lower iodixanol concentrations than the virus suspension in the following order: 7 ml of 15%, 6 ml of 25%, 5 ml of 40%, and 6 ml of 60% iodixanol solution. For this purpose, a MINIPULS 3 peristaltic pump (Gilson) and a long glass pipette were used. The tube was then sealed with a Beckman tube topper and centrifuged for 1 h 45 min at 70,000 rpm and 18°C ​​in an Optima L-80K ultracentrifuge (70 Ti rotor, Beckman Coulter). A cannula was then drilled into the top of the tube to allow airflow. The 40% iodixanol phase, rich in virus particles, was collected from the gradient by puncturing the tube laterally at the boundary between the 40% and 60% phases using a 20G cannula and a 20ml syringe. The virus-containing solution was stored at -80°C until further processing.

[0230] To further purify the virus, anion exchange chromatography was performed using an AKTAprimeplus chromatography system (GE Healthcare), a 5 ml HiTrapTM Q FF anion exchange chromatography column (GE Healthcare), and PrimeView 5.31 software (GE Healthcare). Before starting, the column was equilibrated with buffer A (20 mM Tris, 15 mM NaCl, pH 8.5) and the virus-containing solution was diluted in a 1:1 ratio with this buffer. The solution was loaded onto the column via a loop injector (50 ml Superloop, GE Healthcare). The UV light absorption and conductance properties of the collected fractions were monitored to provide information on the amount of virus contained. The remaining bound molecules were removed from the column using a 2.5 M NaCl solution. All virus-containing fractions were pooled and used for subsequent processing. 3.5.4 Increasing rAAV concentration

[0231] To increase the virus concentration, Amicon® Ultra-4 centrifugal filter units (Merck) with a molecular weight cut-off of 100 kDa were used. The virus-containing solution was loaded onto the top of the filter unit and centrifuged at 4,000 rpm (JA-10 rotor, J2-MC high-speed centrifuge, Beckman Coulter) and 4°C for 20 min intervals until the volume was reduced to 500 μl. The filter unit was subsequently washed with 1 ml of 0.014% Tween / PBS-MK (add 50 ml of 10×PBS, 500 μl of 1 M MgCl, 500 μl of 2.5 M KCl, and 500 ml of water). The solution was further centrifuged under the same conditions until the volume was reduced to 100 μl of concentrated virus solution. Aliquots of 10 μl were prepared and stored at −80°C until use.

[0232] To determine the titer of the generated rAAV, qPCR was performed using a StepOnePlus real-time PCR system (Applied Biosystems, Thermo Fisher Scientific). A standard curve was generated to serve as a reference. For this purpose, a fragment containing part of the ITR was amplified by PCR using the following primers: ITR2 forward: 5' GGAACCCCTAGTGATGGAGTT 3' (SEQ ID NO: 30) ITR2 reverse: 5' CGGCCTCAGTGAGCGA 3' (SEQ ID NO: 31)

[0233] The amplified products were then purified and the concentrations were determined using a Nanodrop™ 2000c spectrophotometer (Thermo Fisher Scientific). 10 ~10 1 A dilution series covering a range of copies was made. To achieve a standard curve, qPCR was performed using three technical replicates of the standard dilution series. For this purpose, MicroAmp™ fast optical 96-well reaction plates (Applied Biosystems, Thermo Fisher Scientific) and PowerUp™ SYBR™ Green Master Mix (Thermo Fisher Scientific) were used. The virus solution was diluted 100-fold with H2O and run on the same reaction plate in three technical replicates. The reaction mixtures were prepared as follows: The data obtained were analyzed using the StepOnePlus real-time PCR system software (Applied Biosystems, Thermo Fisher Scientific). The baseline settings and cycling threshold positions were adjusted manually, if necessary. The standard curve was obtained by plotting the obtained cycle threshold (Ct) values ​​against the logarithm of the dilution. The number of viral genomes per μl of the generated rAAV (vg / μl) can be inferred from the standard curve.

[0234] subretinal injection For subretinal injections, postnatal day 21 (P21) C57Bl6 / J mice were anesthetized by intraperitoneal injection of ketamine (40 mg / kg body weight) and xylazine (20 mg / kg body weight). After complete loss of paw withdrawal reflex, the pupils were dilated by administration of eye drops containing atropine (1%) and tropicamide (0.5%) (Mydriaticum Stulln, Pharma Stulln GmbH). The fundus was focused using an operating microscope (OPMI 1 FR pro, Zeiss). 1 μl containing 10 10 rAAV particles was injected subretinal by a single injection using a NANOFIL 10 μl syringe (World Precision Instruments) and a 34 G beveled needle (World Precision Instruments). The injected eye was treated with eye ointment containing 5 mg / g gentamicin and 0.3 mg / g dexamethasone. Mice were placed on a 37 °C heating plate (Leica HI1120, Leica Biosystems) until they had fully recovered from anesthesia. Two to four weeks after injection, all injected retinas were harvested and processed for RT-PCR analysis or immunohistochemistry.

[0235] immunohistochemistry For immunohistochemistry, subretinal injected mice were euthanized via cervical dislocation. Eyes were removed and placed in 0.1 M phosphate buffer (PB). Subsequently, the eyeball was punctured at the ora serrata using a 21 G cannula and fixed in 4% paraformaldehyde (PFA, Sigma Aldrich, adjusted to pH 7.4) for 5 min. The eye was then placed under a stereomicroscope (Stemi 2000, Zeiss) on filter paper soaked with 0.1 M PB. The cornea, lens, and vitreous were removed by cutting together with the ora serrata using surgical scissors (SuperFine Vannas, World Precision Instruments). The remaining part of the eyeball, including the retina, was fixed in 4% PFA for 45 min at room temperature, followed by three washes in 0.1 M PB for 5 min. For cryopreservation, the eyeball was placed in 30% sucrose solution (w / v) overnight at 4 °C.

[0236] The next day, the eyeballs were embedded in tissue freezing medium (Sakura) and cooled on dry ice until the medium solidified. The retinas were cut into 10 μm thick slices using a cryostat (Leica CM3050 S, Leica Biosystems), collected on coated glass object slides (Superfrost Plus microscope slides, Thermo Fisher Scientific), and stored at -20 °C.

[0237] For immunohistochemical staining, retinal sections were thawed at room temperature and framed using Super PAP Pen Liquid Blocker (Science Services). Sections were then rehydrated in 0.1 M PB for 5 min and fixed in 4% PFA for 10 min. Sections were washed three times for 5 min each in 0.1 M phosphate buffer, pH 7.4 (PB), after which a solution containing primary antibody in 0.1 M PB, 5% ChemiBLOCKER (Merck), and 0.3% TritonX-100 was applied. Frozen sections were incubated with the primary antibody solution overnight at 4°C. The next day, retinas were washed three times for 5 min in 0.1 M PB and incubated with a solution containing secondary antibody in 0.1 M PB and 2% ChemiBLOCKER for 1.5 h at room temperature. Following subsequent washes in 0.1 M PB for 5 min, cell nuclei were stained with 5 μg / ml Hoechst 33342 solution (Invitrogen). Finally, sections were washed in 0.1 M PB, embedded in Fluoromount-G Mounting Medium (Thermo Fisher Scientific), coverslipped, and stored at 4 °C.

[0238] Confocal microscopy Images of stained retinas were obtained using a Leica TCS SP8 inverted confocal laser scanning microscope (Leica Microsystems) equipped with a 405 nm diode and 552 nm and 633 nm optically pumped semiconductor lasers suitable for excitation of Hoechst 33342, Cy3, and Cy5, respectively. Filter settings were selected according to the emission spectra of the respective dyes. Images were acquired as z-stacks (1 μm steps) with a HC PL APO 40× / 1.30 oil CS2 objective (Leica Microsystems) and type F immersion fluid (Leica Microsystems) using LAS X software (Leica Microsystems). Using the same software, the z-stacks were condensed into 2D images by applying maximum intensity projections. Images were further processed with ImageJ 1.48v software (National Institutes of Health, USA).

[0239] Images of transiently transfected live cells were acquired using a Leica TCS SP8 spectral confocal laser scanning microscope (Leica Microsystems) equipped with 448 nm, 514 nm, and 552 nm optically pumped semiconductor lasers suitable for excitation of cerulean, citrine, and mCherry, respectively. Filter settings were selected according to the emission spectra of the respective fluorophores. Images were acquired using an HCX APO 20× / 1.00W objective (Leica Microsystems). All images were processed with ImageJ 1.48v software.

[0240] RNA extraction For RNA extraction from injected retinas, mice were euthanized via cervical dislocation. Blunt forceps were placed under the eye, the eyeball was dissected using a sterile scalpel (Swann-Morton), and the retinas were collected by gradually moving the forceps upwards. Three retinas per construct were pooled and RNA was extracted using the RNeasy Mini Kit (Qiagen) according to the manufacturer's instructions. For disruption, 350 μl of RLT buffer (Qiagen, provided in the kit) + 3.5 μl of β-mercaptoethanol (β-ME, Sigma Aldrich) was added and the tissue was homogenized by passing it through a 20 G needle attached to a sterile syringe at least five times. The remaining steps were performed according to the protocol. RNA was eluted in 30 μl of RNAse-free H2O.

[0241] For RNA extraction from transiently transfected cells, the RNeasy Mini Kit Plus (Qiagen) was used. For this purpose, the medium was removed from 6 cm culture plates and the cells were scraped off using a 16 cm cell scraper (Sarstedt). The cells were collected in 500 μl medium in 2 ml safelock tubes (Eppendorf) and centrifuged at 3,000 × g and 4 °C for 10 min. The supernatant was discarded and the pellet was resuspended in 600 μl RLT Plus buffer (Qiagen, provided in the kit) + 6 μl β-ME (Sigma Aldrich). A steel ball was placed in each tube and the cells were disrupted for 1 min at 30 Hz using a mixer mill MM400 (Retsch). The ball was then removed and the suspension was centrifuged for 5 min at 21,000 × g and room temperature. The remaining steps were performed according to the protocol, including the optional step of removing genomic DNA via a gDNA Eliminator spin column. RNA was eluted in 30 μl of RNAse-free H2O.

[0242] RNA concentration was measured using a NanodropTM 2000c spectrophotometer (Thermo Fisher Scientific). RNA was kept on ice until further use or stored at −20°C for short-term storage or −80°C for long-term storage.

[0243] cDNA synthesis For cDNA synthesis, the Revert Aid First Strand cDNA Synthesis Kit (Thermo Fisher Scientific) was used according to the manufacturer's instructions. Equal amounts of RNA were used per experiment. The cDNA reaction mix was incubated in a Mastercycler® Nexus gradient. The cDNA was kept on ice until further use or stored at -20°C for short-term storage or -80°C for long-term storage.

[0244] Reverse transcription PCR Reverse transcription-PCR (RT-PCR) was performed using Herculase II fusion DNA polymerase (Agilent Technologies) or VWR Taq DNA polymerase (VWR).

[0245] Branch point analysis For nested lariat RT-PCR, RNA was extracted as described above. 10 μg of RNA was incubated with RNase R (Lucigen) to remove all non-circular RNA, according to the manufacturer's instructions. Subsequent cDNA synthesis (Revert Aid First Strand cDNA Synthesis Kit, Life Technologies) was performed as described above. Only random hexamer primers were used in this reaction. The lariats were then amplified by nested RT-PCR using Herculase II fusion DNA polymerase (Agilent Technologies). For the first amplification, the reaction mix was prepared as described in 3.12. For the second amplification, 5 μl of the first PCR was added to the reaction mix instead of cDNA. In addition, a second primer pair binding 25–30 bp downstream of the first primer pair was used. The applied cycling conditions are shown in Table 11.

[0246] For TOPO cloning and lariat analysis, the products of nested RT-PCR were subcloned into plasmids. For this purpose, 3'-adenine overhangs were added to the DNA fragments after amplification by incubating 1 unit of Taq polymerase (VWR) with the PCR reaction at 72 °C for 10 min. Lariats were subsequently subcloned into TOPO vectors (TOPO TA cloning kit, Thermo Fisher Scientific) according to the manufacturer's instructions. Plasmids were transformed into bacteria and small-scale plasmid preparations were performed. The resulting plasmids were sequenced (Eurofins Genomics) and the resulting lariat sequences were analyzed by aligning them with the examined introns using DNAMAN software (Lynnon Biosoft).

[0247] Protein extraction For protein extraction, medium was removed from 6 cm or 10 cm culture plates. Cells were scraped off using a 16 cm cell scraper (Sarstedt) and collected in 500 μl medium in 2 ml safe-lock tubes (Eppendorf). Cell suspensions were centrifuged at 3,000 × g for 10 min at 4 °C, the supernatant was discarded, and the pellet was dissolved in 150 μl and 250 μl of TritonX-100 (TX) lysis buffer (2.5 ml TritonX-100, 15 ml 5 mM NaCl ml, 0.4 ml 2.5 M CaCl2, 500 ml water; c0mplete™ ULTRA protease inhibitor cocktail tablets (Roche) were added to 6 cm and 10 cm plates, respectively (1 tablet / 10 ml)) just before use. Steel balls were added to each safe-lock tube, and cells were disrupted using a mixer mill MM400 (Retsch) at 30 Hz for 1 min. The tube containing the ball was then rotated upside down (VWR™ tube rotator) for 20 min at 4° C. The ball was then removed and the lysate was centrifuged at 5,000 × g for 10 min at 4° C. The supernatant containing the protein was transferred to a new reaction tube and stored at -20° C.

[0248] Total protein concentration was quantified using the Bradford assay. 5 μl of protein lysate was mixed with 95 μl of 0.15 M NaCl solution and transferred to a PMMA standard disposable cuvette (BRAND). Subsequently, 1 ml of Coomassie blue solution was added, mixed thoroughly by pipetting, and incubated for 2 min at room temperature. The absorbance of the solution was measured using a BioPhotometer (Eppendorf) against a blank control containing 5 μl of TX lysis buffer. The obtained value represents the total amount of protein contained in 5 μl of lysate.

[0249] Example 1: Identification of an optimized ASS module The effects of disease-associated rhodopsin mutations on mRNA splicing have been analyzed using a human rhodopsin (RHO) minigene in HEK293 cells and transduced mouse photoreceptors. Among them, one mutation (c.620T>G) in exon 3 of the RHO gene creates a novel ASS (Figure 1). The sequence of the ASS and the predicted ASS elements are shown in Figure 2A and B.

[0250] In silico predictions showed that the acceptor splice site formed by the c.620T>G mutation (hereafter referred to as ASS_620) has a similar splice score as the native rhodopsin acceptor splice site in exon 3 (see Figure 3A). This suggests that both splice sites can be alternatively used by the splicing machinery. However, experimental data showed that ASS_620 was exclusively used in both HEK293 cells (Figure 1D) and mouse photoreceptors (Figure 1E). This indicates that ASS_620 is a strong acceptor splice site.

[0251] For promising use in biotechnological applications, the strength and function of ASSs should be independent of the gene environment. To test whether ASS_620 functions in alternative non-native environments and to determine the ASS_620 elements required for the most efficient splicing, sequences of variable length flanking the c.620T>G mutation were introduced into exon 3 of the RPS27 gene (Figure 2C). A single sequence resulted in variable splicing efficiency at the position of ASS_620 (Figure 2D). The highest splicing efficiency (close to 100%) was obtained when using RHO-E3d, a minigene containing a 26-bp sequence (CAACGAGTCTTTTGTCATCTACAGGT; SEQ ID NO: 3) consisting of a 7-bp sequence at the 5' end, a polypeptide, and a canonical ASS. Surprisingly, this sequence did not contain the predicted branch point present in RHO exon 3. The 5'-terminal 7 bp sequence (CAACGAG) may contain a currently uncharacterized effective branch point sequence or an intron splice enhancer recognition site required for efficient use of ASS. The 26 bp sequence (SEQ ID NO:3) is referred to herein as "vgASS_620."

[0252] To further confirm the splicing efficiency of vgASS_620, it was introduced into four additional genes (HBQ1, S100A12, CLRN1, and CNGB1) that were randomly selected and had variable acceptor splice site strengths (Figure 3C). Compared with vgASS_620, the predicted strength of the native exon acceptor splice site was higher in the case of RPS27, similar in the cases of S100A12, CLRN1, and HBQ1, and lower in the case of CNGB1 (Figure 3A). Subsequent RT-PCR experiments showed that vgASS_620 is not only used in RPS27, but also in all additional genes tested here (Figure 3D-E). These results indicate that vgASS_620 is a very strong acceptor splice site that is independent of the gene environment. It was therefore hypothesized that vgASS_620 could be utilized to improve the currently rather low trans-splicing efficiency in SMaRT or dual adeno-associated virus (AAV) vector hybrid technologies.

[0253] Example 2: Cerulean reporter assay for in vitro applications To test this hypothesis, a splice reporter assay was generated by splitting the coding sequence of Cerulean into two artificial exons interrupted by an artificial intron (Figure 4D). This intron contained a strong donor splice site, vgASS_620, separated from each other by an artificial 145 bp. Confocal imaging of transfected HEK293 cells revealed robust Cerulean fluorescence when both splice sites were provided in cis (Figure 4F(c1)). RT-PCR experiments and sequencing of the corresponding bands confirmed that both Cerulean exons were efficiently spliced ​​in this configuration (data not shown). As shown in Figure 4G, Western blotting experiments using a specific antibody against the N-terminal portion of Cerulean detected a specific immunosignal at the expected size (27 kDa). Taken together, these findings suggest that the reporter assay is functional and highly efficient in the cis configuration. Therefore, in subsequent experiments, we used this construct as a reference to determine the efficiency of Cerulean reconstitution when the two Cerulean splice fragments were provided in trans (i.e., on separate plasmids). The DSS used in this experiment has the sequence AAGGTAAG.

[0254] Next, the effects of binding domain and acceptor splice site strength on Cerulean reconstitution efficiency were analyzed using a reporter assay, where reconstitution of the reporter requires mRNA trans-splicing.

[0255] For this purpose, a fluorescent reporter-based assay was developed in which the coding sequence of the cyan fluorescent protein "Cerulean" was split again in two at position 154 (Figure 4D). As shown in Figure 4D(c3), the first DNA construct containing the 5' part under the control of the CMV promoter was further equipped with a strong DSS (AAGGTAAG) followed by a binding domain, and the second DNA construct containing the 3' part under the control of the CMV promoter was equipped with vgASS_620 as an ASS followed by a binding domain complementary to the binding domain of the first DNA construct. After transcription, base pairing of the BDs brings the splicing elements into close proximity, facilitating trans-splicing, resulting in the full-length mature Cerulean mRNA. Reconstitution of Cerulean can be detected optically, for example, by microscopy (as in Figure 4F) or flow cytometry, at the mRNA level (for example, using RT-PCR) or at the protein level (for example, using Western blot as in Figure 4G).

[0256] As an exemplary template for optimization of the binding domain, intron 2 of the human rhodopsin gene was selected for two reasons: (1) the c.620T>G splice mutation is localized in exon 3, so the optimized version of the RHO intron 2 binding domain can be used for SMaRT-based replacement of the c.620T>G splice mutation (Figure 1). (2) The RHO intron 2 sequence is not homologous to any sequence in the mouse genome (data not shown) and is therefore not expected to cause off-target effects. As a result, the optimized RHO intron 2 binding domain can also be utilized for the reconstruction of other large human genes in the mouse retina using a dual AAV vector approach.

[0257] HEK293 cells were transiently co-transfected with the first and second DNA constructs as described above, shown as an example in BD_h+i in Figure 4D(c3), and the presence of Cerulean fluorescence was assessed by confocal live-cell imaging. A construct containing both Cerulean halves in cis, mediated by an artificial intron containing the same splicing element, was used as a cis-splicing reference control (cis-ctrl). No fluorescence was detectable when the two halves were transfected separately. When both constructs were co-transfected, fluorescent cells were observed, indicating successful trans-splicing and reconstitution of the Cerulean coding sequence.

[0258] The RHO intron 2 binding domain was optimized by varying its size and location (Figure 4A). Furthermore, artificial BDs were created and tested by fusing sequences derived from the 5' and 3' ends of the intron, similar to BD_h+i.

[0259] Following co-transfection of HEK293 cells, mRNA expression of the two constructs was analyzed relative to the housekeeping gene ALAS, using primers p1 and p2 for the 5' construct and primers p3 and p4 for the 3' construct, showing only small differences at the pre-mRNA level (Figure 4B and C). Reconstitution efficiency, determined by radiometric analysis of the Cerulean protein band compared to cis-ctrl, varied greatly using the various BDs (Figure 4E). All protein bands were normalized to beta-tubulin before quantification. The two binding domains BD_g (SEQ ID NO:27) and BD_h+i (SEQ ID NO:28), both approximately 100 bp in length, resulted in high Cerulean reconstitution efficiency reaching >30%. Importantly, this high efficiency far exceeded what was known from previous studies using alternative approaches based on the reconstitution of split genes at the genome level (hybrid, overlapping, and genomic "trans" splicing approaches) (Carvalho et al., 2017, Frontiers in Neurosciences, 11, Article 503). All three strategies were tested in vitro using the reporter gene lacZ. The highest reconstitution efficiency reported in this setting was 17.7%.

[0260] Recent studies have hypothesized that the efficiency of trans-splicing is not affected by the strength of the splice site (Lorain et al., 2013). To test this assumption, we compared the cerulean reconstitution efficiency of vgASS_620 with the native acceptor splice site of RHO exon 3 combined with the most efficient binding domain BD_h+i (SEQ ID NO:28). As can be seen in Figure 4E, compared to vgASS_620, the reconstitution efficiency derived from the native ASS was significantly lower (35.7 ± 4.6% vs. 0.7 ± 0.1%). This finding provides clear evidence that the trans-splicing efficiency in the dual vector approach strongly depends on the strength of the acceptor splice site (Figure 4D-G).

[0261] The reconstitution efficiency may increase in later experiments (see, for example, Example 3, FIG. 9D). First of all, the number (n) of independent samples is still very low in the first experiment. By increasing the number (n) of independent samples and optimizing the transfection protocol, the reconstitution efficiency can be more reliably quantified, and a higher value such as 60% or even higher reconstitution efficiency (FIG. 9) can be reached by improving the co-transfection efficiency.

[0262] In addition to HEK293 cells, the reconstitution efficiency of trans-splicing has also been tested in 661W and MEF cells. Using the binding domain, a BD_g reconstitution efficiency of >40% in 661W cells and >50% in MEF cells was observed (data not shown). No significant differences in the trans-splicing efficiency could be detected when comparing the two cell lines, suggesting that the mRNA trans-splicing approach is cell type independent.

[0263] Taken together, our in vitro results indicate that trans-splicing efficiency can be significantly increased not only by altering the sequence and length of the binding domain but also by optimizing the strength of the acceptor splice site. These promising findings open new avenues for further optimization of trans-splicing-based technologies.

[0264] In combination with vgASS_620, the binding domain "g" shown in FIG. 4 can be used to establish an SMaRT-based gene therapy approach for rhodopsin mutations in or downstream of exon 3. Additionally, vgASS_620 can be combined with various other binding domains to treat other inherited retinal disease (IRD) genes via SMaRT or dual AAV vector approaches. The disclosure provided herein applies to the use of any DNA or RNA sequence that contains the vgASS_620 splice module apart from its naturally occurring context (i.e., apart from the rhodopsin gene in patients with the c.620T>G mutation). vgASS_620 of SEQ ID NO:3 or a nucleotide acid sequence having at least 70% sequence identity to SEQ ID NO:3 can be used for the following applications: 1) Minigene design (e.g., for analysis of mRNA splicing mutations) 2) targeting endogenous mRNAs in a biological or therapeutic context (e.g., using SMaRT technology) 3) Reconstitution of two mRNA fragments in trans. This is particularly important in overcoming the limited genome packaging capacity of adeno-associated virus (AAV) vectors (~5.0 kb, preferably <4.7 kb). Figures 5-8 show the sequences of AAV vectors that can be used in the dual AAV vector system described herein. 4) Design of gene expression or targeting cassettes to selectively increase the presence of specific splice products in alternatively spliced ​​genes that contain weak ASS sites.

[0265] Example 3: Influence of the acceptor splice site and binding domain (BD) on reconstitution efficiency. The length and sequence of the BD represent important determinants of reconstitution efficiency. The BD most likely influences the tight binding and possibly folding of the mRNA, but is not expected to directly promote the efficiency or accuracy of the subsequent splicing process. As mentioned above, DSS has been well characterized and predictions of its strength are in good agreement with experimental performance. Therefore, there is no obvious need to optimize this splice site in the framework of a split fluorophore reconstitution assay. In contrast, due to their complexity, the strength of ASSs cannot be predicted with certainty. The results suggest that vgASS_620 is a very strong acceptor splice site. Given that splice site strength can affect the reconstitution efficiency of the split fluorophore assay, vgASS_620 should yield high values ​​when compared to other acceptor splice sites.

[0266] To analyze this, we compared the reconstitution efficiency in the presence of vgASS_620 or two other ASSs, namely the native ASS of RHO exon 3 (S3, FIG. 9) and a hybrid ASS (S2) created by replacing the polypyrimidine tract (PPT) of vgASS_620 with the PPT of the native RHO exon 3 ASS. Furthermore, each ASS was combined with three different binding domains: a strong (BD_g, B1; SEQ ID NO:27) and a weak (BD_f, B3) binding domain derived from this study, and a published BD sequence (PTM1, B2) taken from RHO intron 1 (SEQ ID NO:29). This sequence was shown to provide high efficiency in repairing mutant RHO transcripts via spliceosome-mediated mRNA trans-splicing (Berger et al., “Repair of rhodopsin mRNA by spliceosome-mediated RNA trans-splicing: A new approach for autosomal dominant retinitis pigmentosa”, (2015) Mol. Ther. 23(5):918-930). All combinations were analyzed by confocal live-cell imaging, RT-PCR, and Western blotting (Figure 9). This experiment resulted in several important findings. First, we revealed that both BD and ASS are important factors that determine the reconstitution efficiency. Second, the most potent BD identified in this study (BD_g, B1) is superior to the published RHO-binding domain (B2). Third, the combination of the strongest BD with the strongest ASS (B3+S1) results in detectable trans-splicing, whereas the strongest BD with the weakest ASS (B1+S3) does not result in detectable rearrangement of the coding sequence.

[0267] vgASS_620 consists of the ASS, PPT, and an additional 7-bp sequence upstream, the deletion of which impairs ASS_620 recognition. It has therefore been speculated that this 7-bp sequence may contain a very strong branch point that may explain the universal and efficient recognition of vgASS_620 independent of the gene environment. However, there were no strong branch points predicted within this sequence. Instead, several other sequences were predicted to function as branch points located up to 40 bp upstream of c.620T>GASS. Nevertheless, simultaneous mutation of all predicted branch point nucleotides and all possible branch point adenines contained in the 7-bp sequence upstream of the PPT failed to alter the splicing of the c.620T>G mutant (data not shown). This finding indicates that the branch point is located elsewhere or that c.620T>GASS has high flexibility in branch point selection. To more directly identify the branch points utilized for splicing at c.620T>GASS, we performed nested lariat RT-PCR using HEK293 cells transiently expressing the RHO c.620T>G minigene. HEK293 cells transfected with the RHO WT minigene served as a reference. When performing lariat RT-PCR, one band each was obtained for the WT and mutant minigenes, both of which differed in size. Both bands appeared somewhat diffuse, suggesting that the corresponding lariats differed in size. To identify the single sequences contained within these diffuse bands, the lariat RT-PCR products were subcloned into a TOPO vector, and the resulting clones were analyzed individually by sequencing. When studying the RHO WT lariats, two major branch points (used in 42% and 33% of cases) and three minor branch points (each in 8% of cases) were identified (Table 2). The two major branch points were highly similar to the consensus sequence, resulting in high prediction scores. All RHO WT branch points were found 57–184 bp upstream of intron-exon junctions.RHO exon 3 mRNA splicing appears to be atypical, as more than 90% of human branch points are located 50 bp upstream of the ASS sequence (Corvelo et al., 2010, PLoS Comput. Biol., 6(11): e1001016).

[0268] [Table 2]

[0269] Nevertheless, the branchpoint profile obtained with RHO c.620T>G was remarkably different when compared to the WT minigene. First, various branchpoints for the c.620T>G mutant have been identified. However, no major branchpoints could be detected, none identical to those obtained with the WT minigene. Second, the branchpoints were located further upstream of the used ASS, that is, between 107 and 245 bp, corresponding to 21–159 bp upstream of the native exon 3 ASS. Third, almost half of the detected branchpoints bore little similarity to the consensus sequence and were therefore not predicted using the Human Splice Finder (HSF) splice prediction tool. Overall, these data suggest that the strength of vgASS_620 may be partially caused by a high flexibility in the choice of branchpoints rather than being due to the presence of very strong branchpoints included. This could explain its unusually efficient performance, making it a highly attractive tool for biotechnological applications requiring efficient splicing.

[0270] Example 4: Study of mRNA trans-splicing-based rAAV dual vectors in vivo The most powerful application of the mRNA trans-splicing-based assay evaluated in the previous section is the reconstitution of large genes in the framework of dual rAAV vectors. Consequently, the mRNA trans-splicing approach was tested in mouse retina using rAAV. For this purpose, a slightly modified version of the split fluorophore assay was used. To control for rAAV vector-derived expression in cells transduced with a single virus, both dual rAAV vector cassettes were equipped with fluorophore sequences, namely Citrine at the 5' end of the coding sequence of the 5' vector and mCherry at the 3' end of the coding sequence of the 3'' vector. One of the BDs yielding the highest reconstitution efficiency in vitro, namely BD_h+i, was used in vivo (Figure 10A). In this experimental setting, Cerulean fluorescence should be present in cells expressing Citrine as well as mCherry.

[0271] Titer-matched viruses were injected subretinally into WT C57Bl6 / J mice at postnatal day 21 (P21). After harvesting retinas 2 weeks after injection, solid fluorophore expression could be detected in the RPE (Figure 10B). Moreover, cerulean fluorescence could be observed in all areas where citrine and mCherry were expressed, indicating successful mRNA trans-splicing in cells co-transduced with both AAVs. However, citrine and cerulean have partially overlapping excitation and emission spectra. To exclude the possibility that cerulean fluorescence was an artifact caused, for example, by bleed-through of citrine or mCherry, both fluorophores were selectively bleached in a small area of ​​the RPE by exciting the fluorophores with a high-intensity 514 nm laser. This procedure allowed the complete elimination of citrine and mCherry fluorescence (Figure 10C). Nevertheless, cerulean fluorescence was unchanged, indicating that it originates exclusively from trans-spliced ​​cerulean mRNA. This experiment provides proof of principle of the utility of mRNA trans-splicing for the reconstitution of genes expressed from AAV in vivo.

[0272] Example 5: Identification of potent binding domains suitable for human gene therapy So far, all binding domain (BD) sequences were obtained from human intron regions. Therefore, when used for human gene therapy, they may also bind to endogenous mRNAs and induce trans-splicing with these transcripts. Therefore, to apply mRNA trans-splicing to human gene therapy, it is necessary to identify BDs that do not contain sequences homologous to the human genome. For this purpose, we obtained random 100 bp sequences from the bacterial lacZ gene and modified them via random insertions, deletions, and substitutions to obtain four sequences with no homology to the human genome (Figure 11A).

[0273] When HEK293 cells were co-transfected with 5' vector and 3' vector constructs containing the respective binding domains, a very high Cerulean reconstitution efficiency of 78.3% ± 2.1% was observed for one of the BDs (BD_k, Figure 11B-D). Thus, BD_k was used in preliminary experiments to evaluate the reconstitution of large genes by mRNA trans-splicing, as it was more efficient than the best-performing human BD shown in Figure 4.

[0274] Example 6: Effect of polyadenylation signals on reconstitution efficiency and involvement of DNA-based reconstitution in in vitro assays Since the reconstitution of the two pre-mRNA molecules derived from the 5' and 3' vectors should take place in the nucleus, the polyadenylation signal (pA), necessary for the stabilization and translation of the mature mRNA, could theoretically be removed in the 5' vector. The advantage of this deletion is that the remaining unspliced ​​5' pre-mRNA does not cause the translation of a truncated protein. Therefore, we studied the effect of the deletion of pA in the 5' vector on the Cerulean reconstitution efficiency. For this purpose, a 5' vector deleting pA was cotransfected with a normal 3' vector. Compared to the cotransfected vectors, both of which contain pA, the reconstitution seems to be slightly reduced (Figure 12A and B). However, this experiment shows that the pA signal can be removed if necessary. Moreover, the successful reconstitutions observed so far could theoretically be mediated by homologous recombination at the DNA level, as in the prior art hybrid dual vector approach, since all necessary components, i.e., recombination sequences and splicing elements to remove this sequence via cis-splicing, are present in the mRNA trans-splicing vector (Carvalho et al., 2017). To exclude this possibility, the promoter of the 3' vector was deleted to prevent transcription into pre-mRNA, thus resembling the 3' vector used in the previously known hybrid dual vector approach. After co-transfecting this construct with the normal 5' vector, no reconstitution could be observed. This result confirms that all observed reconstitutions of Cerulean are mediated exclusively via mRNA trans-splicing.

[0275] Example 7: Proof of principle for the reconstitution of large genes via mRNA trans-splicing In addition to the AAV Cerulean split reporter assay, assays have been tested to evaluate mRNA trans-splicing of therapeutically promising proteins. The transcription activator SpCas9-VPR, a catalytically inactive nuclease fused to the transcription activator domain VP64-p65-Rta (VPR), is a new tool recently developed for gene therapy. Due to its large size (5.8 kb), it needs to be delivered via a dual vector for in vivo applications. Thus, SpCas9-VPR is a suitable candidate for reconstitution via mRNA trans-splicing. The coding sequence of SpCas9-VPR was split into two at c.2185, and the two halves of the coding sequence were equipped with BD_k or its complementary sequence, and the appropriate splicing elements, i.e., DSS or vgASS_620. A full-length (FL) SpCas9-VPR construct was used as a positive control. RT-PCR from HEK293 cells co-transfected with the split constructs revealed that SpCas9-VPR mRNA was successfully reconstituted and no unwanted by-products were generated (Figure 13B). Sequencing of the PCR products confirmed accurate restoration of the reading frame (Figure 13C). Furthermore, FL SpCas9-VPR protein was detectable by Western blotting (Figure 13D). The reconstitution efficiency of SpCas9-VPR was 13.2% ± 0.9%. This provides a proof of principle of the applicability of the mRNA trans-splicing approach for the reconstitution of large genes.

[0276] Example 8: Reconstitution of ABCA4 (Figure 14) Finally, to also study the rearrangement of a large gene of human origin, we selected the ABCA4 gene (6.8 kb), which encodes a retinal ATP-binding cassette transporter. This gene is a suitable candidate for gene therapy because its mutations cause the inherited retinal dystrophy Stargardt's macular dystrophy. ABCA4 was split into two at c.3243, equipped with BD_k, DSS, and vgASS_620 (Figure 14A). Both halves contain a long intronless ABCA4 coding sequence (CDS) that is not similar to the native pre-mRNA composed of exons and introns. This may therefore hinder the recruitment of splice factors to the pre-mRNA, and consequently also reduce the mRNA trans-splicing efficiency. To make the split gene more similar to the endogenous human pre-mRNA, we designed an additional construct that contains three short (80 bp) intervening introns in the CDS of both halves. HEK293 cells were transiently cotransfected with intronless or intron-containing constructs to examine their reconstitution ability in vitro. RT-PCR revealed that ABCA4 was successfully reconstituted at the mRNA level in the split constructs with and without the intervening intron (Fig. S4B). Interestingly, the intron-containing split construct appeared to be trans-spliced ​​more efficiently: again, no unspecified splice products were detectable and the reading frame was correctly restored (Fig. S4C).

[0277] Example 9: rAAV dual vector mRNA trans-splicing of ABCA4 in vivo (Figure 15) In a final experiment, reconstitution of ABCA4 was further tested in vivo. To this end, the CMV promoter was replaced by the human rhodopsin (hRHO) promoter to ensure photoreceptor-specific expression (Figure 15A). AAVs were generated using in-house optimized NN and GL capsid variants derived from wild-type AAV2 capsids as described in WO 2019 / 076856. Titer-matched viruses containing the 5' or 3' CDS of ABCA4 were co-injected subretinally into 1-month-old C57Bl6 / J wild-type mice. After 4 weeks, retinas were harvested and the success of reconstitution was assessed at the mRNA and protein levels. RT-PCR performed using junction-spanning primer pairs revealed that both capsid variants, i.e., NN and GL, were successfully reconstituted with the NN capsid, resulting in higher levels of reconstituted mRNA, presumably due to higher (co)transduction efficiency (Figure 15B). Seamless ligation of two separate pre-mRNA molecules was confirmed by sequencing (Figure 15C). For more precise quantification, qRT-PCR was performed. These preliminary results (n=1) show that we reached relative ABCA4 expression ranging from 10-fold to 42-fold compared to uninjected C57Bl6 / J wild-type retinas, again confirming that the reconstitution at the mRNA level was successful and efficient (Figure 15D). Finally, to study the reconstitution of ABCA4 at the protein level, protein lysates of injected retinas were used for Western blotting. To ensure specific detection of transgene-derived protein, an anti-myc antibody was used. The results show that in both cases, the reconstituted mRNA led to successful in vivo protein expression (Figure 15E).

Claims

1. A pre-mRNA trans-splicing molecule, comprising: (i) an acceptor splice region comprising a sequence that is at least 75% identical to the sequence of SEQ ID NO: 3 or SEQ ID NO: 4, and further comprising: (ia) a pyrimidine tract comprising: (iaa) 5 to 25 nucleotides; (iab) wherein at least 60% of the nucleotides within these 5-25 nucleotides are pyrimidine bases, e.g., cytosine (C), thymine (T), and / or uracil (U); and (iac) wherein said 5 to 25 nucleotides of said pyrimidine tract comprise the sequence TTTTTT or TCTTTT; and (ib) an acceptor splice site; (iba) wherein the acceptor splice site is located 3' to the pyrimidine tract; and (ibb) wherein the acceptor splice site comprises the sequence CAGG 5' to 3'; (ii) a nucleotide sequence of interest; wherein the acceptor splice region is located 3' or 5' to the nucleotide sequence of interest; and (iii) a binding domain that targets a pre-mRNA located 3' or 5' to the nucleotide sequence of interest; A pre-mRNA trans-splicing molecule comprising:

2. A spacer sequence, wherein the spacer is located between the binding domain and the acceptor splice region. The pre-mRNA trans-splicing molecule of claim 1, further comprising:

3. 3. The pre-mRNA trans-splicing molecule of claim 1, wherein the acceptor splice region is located 5' to the target nucleotide sequence, and the binding domain is located 5' to the acceptor splice region.

4. The pre-mRNA trans-splicing molecule according to any one of claims 1 to 3, further comprising a termination sequence, preferably a polyA sequence.

5. (a) the acceptor splice region comprises a sequence that is at least 80% identical to the sequence of SEQ ID NO:3 or SEQ ID NO:4; (b) the sequence between the last pyrimidine of the pyrimidine tract and the acceptor splice site is less than 10 bases; (c) the sequence between the last pyrimidine of the pyrimidine tract and the acceptor splice site is less than 5 bases; and / or (d) the sequence between the last pyrimidine of the pyrimidine tract and the acceptor splice site is less than 3 bases.

6. The acceptor splice region is (a) the 7 nucleotides 5' to said pyrimidine tract having at least 4 nucleotides of the sequence CAACGAG, where the first 5' nucleotide is a C; (b) the 7 nucleotides 5' to said pyrimidine tract having at least 5 nucleotides of the sequence CAACGAG, where the first 5' nucleotide is a C; (c) the 5' seven nucleotides of said pyrimidine tract having at least four nucleotides of the sequence CAACGAG, where the first 5' nucleotide is CAA; or (d) the 7 nucleotides 5' to said pyrimidine tract having at least 5 nucleotides of the sequence CAACGAG, where the first 5' nucleotide is CAAC; The pre-mRNA trans-splicing molecule of any one of claims 1 to 5, further comprising:

7. A DNA molecule comprising a promoter and a sequence encoding a pre-mRNA trans-splicing molecule described in any one of claims 1 to 6.

8. A method for producing a nucleic acid sequence, the method comprising: (A) providing a first nucleic acid sequence comprising one or more donor splice site sequences; (B) providing a second nucleic acid sequence comprising: (i) an acceptor splice region sequence comprising a sequence that is at least 75% identical to the sequence of SEQ ID NO: 3 or SEQ ID NO: 4, wherein the acceptor splice region sequence further comprises: (ia) a pyrimidine tract comprising: (iaa) 5 to 25 nucleotides; (iab) wherein at least 60% of the nucleotides within these 5-25 nucleotides are pyrimidine bases, cytosine (C), thymine (T), and / or uracil (U); and (iac) wherein said 5 to 25 nucleotides of said pyrimidine tract comprise the sequence TTTTTT or TCTTTT; (ib) an acceptor splice site; (iba) wherein the acceptor splice site is located 3' to the pyrimidine tract; and (ibb) wherein the acceptor splice site comprises the sequence CAGG 5' to 3'; (C) cleaving the first nucleic acid sequence at the one or more donor splice site sequences and cleaving the second nucleic acid sequence at the acceptor splice site; (D) ligating the cleaved first nucleic acid sequence to the cleaved second nucleic acid sequence, thereby obtaining a nucleic acid sequence; A method comprising:

9. 9. The method of claim 8, wherein the first nucleic acid sequence further comprises a nucleotide sequence of interest, wherein at least a portion of the nucleotide sequence of interest is located 5' to the donor splice site, and wherein the second nucleic acid sequence further comprises a nucleotide sequence of interest, wherein at least a portion of the nucleotide sequence of interest is located 3' to the acceptor splice region.

10. 10. The method of claim 8 or 9, wherein the first and second nucleic acid sequences are introduced into a host cell, preferably the first and second nucleic acid sequences are recombinant nucleic acid sequences.

11. A method for producing a nucleic acid sequence of interest, the method comprising: (A) introducing into a host cell a first nucleic acid sequence comprising a pre-mRNA trans-splicing molecule sequence or a nucleic acid sequence encoding said pre-mRNA trans-splicing molecule, wherein said first pre-mRNA trans-splicing molecule comprises, 5' to 3': (a) the 5' portion of the nucleotide acid sequence of interest; (b) donor splice site; (c) a first binding domain; and (d) a termination sequence, and (B) introducing into a host cell a second nucleic acid sequence comprising a pre-mRNA trans-splicing molecule sequence or a nucleic acid sequence encoding said pre-mRNA trans-splicing molecule, wherein said second pre-mRNA trans-splicing molecule comprises, 5' to 3': (i) a second binding domain complementary to the first target domain of the first nucleic acid sequence; (ii) an acceptor splice region sequence comprising a sequence that is at least 75% identical to the sequence of SEQ ID NO: 3 or SEQ ID NO: 4, wherein the acceptor splice region sequence further comprises: (iia) a pyrimidine tract comprising: (ii aa) 5 to 25 nucleotides; (iiab) wherein at least 60% of the nucleotides within these 5-25 nucleotides are pyrimidine bases, cytosine (C), thymine (T), and / or uracil (U); and (iac) wherein said 5 to 25 nucleotides of said pyrimidine tract comprise the sequence TTTTTT or TCTTTT; (iib) an acceptor splice site; (iiba) wherein the acceptor splice site is located 3' to the pyrimidine tract; and (iibb) wherein the acceptor splice site comprises the sequence CAGG 5' to 3'; (iii) the 3' portion of the nucleotide sequence of interest, and (iv) a termination sequence, preferably a polyA sequence; (C) cleaving the first nucleic acid sequence at the donor splice site sequence and cleaving the second nucleic acid sequence at the acceptor splice site; (D) ligating the cleaved first nucleic acid sequence comprising the 5' portion of the nucleotide sequence of interest to the cleaved second nucleic acid sequence comprising the 3' portion of the nucleotide sequence of interest, thereby obtaining the nucleotide sequence of interest; A method comprising:

12. An adeno-associated virus (AAV) vector comprising at least two inverted terminal repeats and a nucleic acid sequence between the two inverted terminal repeats, wherein the nucleic acid sequence is 5' to 3' (i) a promoter; (ii) a binding domain; (iii) an acceptor splice region sequence comprising a sequence that is at least 75% identical to the sequence of SEQ ID NO: 3 or SEQ ID NO: 4, further comprising: (a) a pyrimidine tract comprising: (aa) 5 to 25 nucleotides; (ab) wherein at least 60% of the nucleotides within these 5-25 nucleotides are pyrimidine bases, e.g., cytosine (C), thymine (T), and / or uracil (U); (ac) wherein said 5 to 25 nucleotides of said pyrimidine tract comprise the sequence TTTTTT or TCTTTT; (b) an acceptor splice site; (ba) wherein the acceptor splice site is located 3' to the pyrimidine tract; and (bb) wherein the acceptor splice site comprises the sequence CAGG 5' to 3'; (iv) a nucleotide sequence of interest; and (v) poly(A) sequence An adeno-associated virus (AAV) vector comprising:

13. An adeno-associated virus (AAV) vector system, comprising: (I) a first AAV vector comprising at least two inverted terminal repeats and a nucleic acid sequence between the two inverted terminal repeats, wherein the nucleic acid sequence between the two inverted terminal repeats is (a) a promoter; (b) a nucleotide sequence encoding the N-terminal portion of the polypeptide of interest; (c) donor splice site; (d) a first binding domain; and (e) a termination sequence; a first AAV vector comprising: (II) a second AAV vector comprising at least two inverted terminal repeats and a nucleic acid sequence between the two inverted terminal repeats, wherein the nucleic acid sequence between the two inverted terminal repeats is (i) a promoter; (ii) a second binding domain complementary to the first binding domain of the first AAV vector; (iii) an acceptor splice region sequence comprising a sequence that is at least 75% identical to the sequence of SEQ ID NO: 3 or SEQ ID NO: 4, further comprising: (a) a pyrimidine tract comprising: (aa) 5 to 25 nucleotides; (ab) wherein at least 60% of the nucleotides within these 5-25 nucleotides are pyrimidine bases, cytosine (C), thymine (T), and / or uracil (U); (ac) wherein said 5 to 25 nucleotides of said pyrimidine tract comprise the sequence TTTTTT or TCTTTT; (b) an acceptor splice site; (ba) wherein the acceptor splice site is located 3' to the pyrimidine tract; and (bb) wherein the acceptor splice site comprises the sequence CAGG 5' to 3'; (iv) a nucleotide sequence encoding the C-terminal portion of the polypeptide of interest; (iv) wherein the C-terminal portion of the polypeptide of interest and the N-terminal portion of the polypeptide of interest reconstitute the polypeptide of interest; and (v) a termination sequence, preferably a polyA sequence a second AAV vector comprising: An adeno-associated virus (AAV) vector system comprising:

14. 14. The AAV vector system of claim 13, wherein the polypeptide is a full-length polypeptide, the first AAV vector comprises an N-terminal portion of the full-length polypeptide of interest, and the second AAV vector comprises a C-terminal portion of the polypeptide of interest.

15. (a) the acceptor splice region comprises a sequence that is at least 80% identical to the sequence of SEQ ID NO:3 or SEQ ID NO:4; (b) the 5'-most 7 nucleotides of the pyrimidine tract have at least 4 nucleotides of the sequence CAACGAG, where the first 5'-nucleotide is a C; (c) the sequence between the last pyrimidine of the pyrimidine tract and the acceptor splice site is less than 10 bases; (d) the sequence between the last pyrimidine of the pyrimidine tract and the acceptor splice site is less than 5 bases; (e) the sequence between the last pyrimidine of the pyrimidine tract and the acceptor splice site is less than 3 bases; and / or (f) the first AAV vector further comprises a spacer sequence between the donor splice site and the first binding domain, and / or the second AAV vector further comprises a spacer sequence between the second binding domain and the acceptor splice site.

15. An AAV vector system according to claim 13 or 14.

16. (a) the acceptor splice region comprises a sequence that is at least 80% identical to the sequence of SEQ ID NO:3 or SEQ ID NO:4; (b) the 5'-most 7 nucleotides of the pyrimidine tract have at least 4 nucleotides of the sequence CAACGAG, where the first 5'-nucleotide is a C; (c) the sequence between the last pyrimidine of the pyrimidine tract and the acceptor splice site is less than 10 bases; (d) the sequence between the last pyrimidine of the pyrimidine tract and the acceptor splice site is less than 5 bases; (e) the sequence between the last pyrimidine of the pyrimidine tract and the acceptor splice site is less than 3 bases; and / or (f) the AAV vector further comprises a spacer sequence between the donor splice site and the first binding domain; The AAV vector of claim 12.

17. A pharmaceutical composition comprising a pre-mRNA trans-splicing molecule described in any one of claims 1 to 6, an adeno-associated virus vector described in claim 12 or 16, or an adeno-associated virus vector system described in any one of claims 13 to 15.