MINI-promoter compositions
The CP040 mini-promoter addresses the challenges of gene therapy for DNA repair disorders by providing a compact and efficient promoter sequence that specifically targets neurological tissues and regulates gene expression in response to disease progression.
Patent Information
- Application Number
- PCT/US2024/037893
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-07-14
- Filing Date
- 2024-07-12
- Publication Date
- 2025-05-08
AI Technical Summary
Current gene therapy approaches for DNA repair disorders like Cockayne Syndrome face challenges such as limited space in gene delivery vectors, difficulty in targeting neurological tissues, and off-target effects.
Development of a novel mini-promoter sequence, CP040, derived from upstream untranslated regions of ERCC6/CSB and ERCC8/CSA genes, which is compact enough to fit within AAV vector packaging constraints and specifically targets neurological tissues.
The CP040 mini-promoter effectively drives transgene expression in neurological tissues, demonstrating high efficacy in vitro and in vivo, and is capable of regulating gene expression in response to disease progression.
Smart Images

Figure IMGF000034_0001 
Figure 00000045_0000 
Figure 00000046_0000
Abstract
Description
[0001] MINI-PROMOTER COMPOSITIONS
[0002] CROSS-REFERENCE TO RELATED APPLICATION
[0003] This application claims priority to United States Provisional Application Number 63 / 526,711 that was filed on July 14, 2023. The entire content of the application referenced above is hereby incorporated by reference herein.
[0004] BACKGROUND
[0005] Cockayne Syndrome (CS) is a rare genetic disorder that affects the body’s ability to repair damaged DNA. It is mainly caused by mutations in the ERCC6 and ERCC8 genes, which are responsible for encoding proteins involved in DNA repair. Symptoms of CS include growth failure, neurological problems, and premature aging. Currently, there is no cure for CS, but gene therapy may offer a potential treatment option which involves the introduction of a corrective gene into the patient’s cells, which can then repair the defective gene and restore normal function. Adeno-associated virus (AAV) mediated gene therapy may prove beneficial to CS patients. This method has been approved for other genetic disorders, including spinal muscular atrophy, Leber congenital amaurosis, and hemophilia. However, there are challenges associated with AAV gene therapy, such as the difficulty of delivering the corrective gene to the target tissue / cells or off-target effects, routes of delivery, packaging size, and ensuring correct expression levels. As CS is a neurological disorder, there is an ongoing need to develop a safe and effective AAV gene therapy that targets neurological tissues.
[0006] SUMMARY
[0007] In certain aspects, provided herein is a promoter sequence consisting of a nucleic acid of between 12 and 300 nucleotides in length, wherein the nucleic acid comprises a polynucleotide having at least 90% identity to SEQ ID NO: 1, SEQ ID NO:2, SEQ ID NO:3, SEQ ID NO:4, SEQ ID NO:5, and / or SEQ ID NO:6.
[0008] In certain aspects, provided herein is a promoter sequence consisting of a nucleic acid of between 12 and 300 nucleotides in length, wherein the nucleic acid consists essentially of a polynucleotide having at least 90% identity to SEQ ID NO: 1, SEQ ID NO:2, SEQ ID NO:3, SEQ ID NO:4, SEQ ID NO:5, and / or SEQ ID NO:6.
[0009] In certain aspects, provided herein is a promoter sequence consisting of a nucleic acid of between 12 and 300 nucleotides in length, wherein the nucleic acid comprises a polynucleotide having 100% identity to SEQ ID NO: 1, SEQ ID NO:2, SEQ ID NO:3, SEQ ID NO:4, SEQ ID NO:5, and / or SEQ ID NO:6.
[0010] In certain aspects, provided herein is an expression cassette comprising the novel promoters described herein, wherein the promoter is operably linked to a preselected DNA segment encoding a protein or heterologous RNA transcript.
[0011] In certain aspects, provided herein is an isolated transformed cell comprising the novel expression cassette or vector described herein.
[0012] In certain aspects, provided herein is a method for producing transformed cells comprising the steps of (i) introducing into cells a recombinant DNA which comprises the novel promoter described herein operably linked to a DNA segment or the novel expression cassette described herein so as to yield transformed cells, and (ii) identifying or selecting a transformed cell line.
[0013] In certain aspects, provided herein is a method of treating a genetic disorder in a mammal comprising administering (a) the novel vector described herein, or (b) the novel transformed cell described herein.
[0014] In certain aspects, provided herein is a use of the novel promoter described herein, wherein the promoter activates, enhances, or represses the expression of a genetic sequence at onset of or during pathological progression of a genetic disorder. In certain aspects, provided herein is a transformed cell made by this method.
[0015] BRIEF DESCRIPTION OF THE DRAWINGS
[0016] Figure 1A. Complete CP040 Mini-Promoter Sequence: ERCC6 UTR (indicated by enlarged bold and italic script) and ERCC8 UTR (indicated by enlarged bold script) (126 base pair length) (SEQ ID NO: 1).
[0017] Spacer 1 = CTCGCCCTGCTC (SEQ ID NO:2)
[0018] CSA UTR = AAAGGCGTGAGCCCGTCCCTCCAAACCACGTCTTGCCCCTTCTT (SEQ ID NO:3)
[0019] Spacer 2 = GACCGACAATATCTCAATTAGTCAGCAA (SEQ ID NO:4)
[0020] CSB UTR = CACCATTGGCCGGATATATCACC (SEQ ID NO: 5)
[0021] Spacer 3 = GACTGACCTCTGCTGTTCC (SEQ ID NO: 6)
[0022] Figure IB. ERCC6 / CSB Upstream UTR (indicated by enlarged bold and italic script). Genbank consensus match (query is ERCC6 UTR sequence). Query 13 = SEQ ID NO:7; Sbjct = SEQ ID NO:8. Figure 1C. ERCC8 / CSA Upstream UTR (indicated by enlarged bold script in Figure 1A). Genbank consensus match (query is ERCC8 UTR sequence). Query 85 = SEQ ID NOV; Sbjct = SEQ ID NO:9
[0023] Figure 2. Depiction of the different promoter regions. Variant 1 is the existing minipromoter and Variants 2-4 are alternatives. Each variant is missing a portion of the original design to facilitate determination of the critical region / s for promoter activity. The sequences for the various segments in each of the Promoter Variants are provided in the figure legend for Figure 1 A above. Schematic scaled representation of the CP040 mini promoter sequence (Fig. 2A). Different fragments (Fig. 2B) and variants (Fig. 2B) of the CP040 sequence were cloned between Kpnl and Xhol restriction sites upstream of firefly luciferase gene in pGL4.20 promoter-less vector.
[0024] Figure 3. Transcriptional efficacy of the CP040 mini-promoter: Representative data from reverse transcription qPCR experiments using probes from three different regions of the codon optimized ERCC6 / CSB gene. N=3 and data are presented as mean ± SEM, statistical analysis determined by one-way ANOVA *p < 0.05, **p < 0.01; ns = not significant.
[0025] Figure 4. Dual luciferase assay design to determine relative promoter activities between the mini-promoter variants. An efficient promoter results in higher luciferase expression levels detected using a plate reader.
[0026] Figure 5. Examples of further experimental minimalist constructs.
[0027] Figure 6. Analysis of Mini-promoter sequence using AliBaba tools with TRANSFEC database to determine potential binding sites. (SEQ ID NO: 10)
[0028] Figure 7. Analysis of the mini-promoter sequence SEQ ID NO: 1 using Firefly tools to determine promoter predictions for seqO. SeqO is the query sequence added to the firefly tools, which is full length promoter SEQ ID NO: 1.
[0029] Figure 8. Vector map.
[0030] Figure 9. Transcription factors present in CP040 sequence that are expected to enhance function (analysis done with PROMO online tool (Messeguer, Escudero et al 2022, Farre, Roset et al. 2023).
[0031] Figures 10A-10B. Diagrammatic representation to determine the efficiency of a promoter with luciferase assay. Figure 10A. Promoter-less test vector, pGL4.20 map. The promoter sequence of interest is cloned at multiple cloning sites upstream of firefly luciferase gene. Figure 10B. Vector map for pRL-CMV. It is used as a positive control in dual luciferase assay expressing Renilla luciferase under CMV promoter.
[0032] Figures 11A-11B. In vitro characterization of CP040 promoter fragments and variants by dual luciferase assay. The fragments or variant constructs transfected in HEK293 (Fig. 11 A) or neural precursors (NPCs) (Fig. 1 IB). Lysates were prepared and renilla (control) and firefly (test) luminescence were measured and fold change in relative light units observed with respect to empty vector was calculated and presented for each fragment and variant in the graphs. CBA promoter was used as a positive control and comparison for full length CP040 promoter (variantl). Data represent mean±SEM of three independent replicates (N=3, **** p<0.0001 by ANOVA). CBA = Chicken P-actin promoter, Variant 1 = Full length CP040.
[0033] Figures 12A-12B. In vivo characterization of the full length CP040 promoter packaged in AAV-DJ adeno-associated viral capsid. AAV-DJ-CP040-eGFP or vehicle (IxPBS) was injected in the temporal vein of wild-type FVB / NJ neonates (postnatal day 1-2) with IxlO13vg / kg of mouse bodyweight. Necropsy was performed on the mice after six weeks and tissue collected were analyzed for viral vector genome and eGFP transcript analysis. Fig. 12A. Graph presenting viral vector genome copy number in different organs- cortex, cerebellum, spinal cord, kidneys, spleen, quadriceps, liver and heart. Fig. 12B. Graph showing the fold change in eGFP transcripts which can be extrapolated as promoter strength in the same organs as in Fig. 12A panel. Data represent mean±SEM, N=10 mice, **p<0.01, ***p<0.001, ****p<0.0001, ns-not significant by one-way ANOVA.
[0034] Figure 13. Schematic representation of AAV vector construct. Four vectors expressing the reporter gene enhanced green fluorescent (eGFP) under different promoters (solid segment in schemes). The CBA promoter is 1600bp long, EFl promoter is 2500bp and human synapsin promoter is 450bp. The shortest CP040 mini promoter is 126bp. The packaging cassettes are flanked by inverted terminal repeats (ITR) and poly-A signals are not shown.
[0035] Figures 14A-14G. AAV vector genome (VG) estimations in tissues. gDNA was isolated and eGFP gene vector genome numbers were calculated per nanograms of gDNA for (Figs. 14A-E) neurological tissues (hippocampus, hypothalamus, cortex, cerebellum, spinal cord) and other tissues (Fig. 14F) (liver) and (Fig. 14G) kidney. Data are presented as mean±SEM of the values from six mice; *p<0.05, **p<0.01, ***p<0.001, ns-not significant (based on one-way ANOVA).
[0036] Figures 15A-15G. AAV vector genome (VG) estimations in tissues. gDNA was isolated and eGFP gene vector genome numbers were calculated per nanograms of gDNA for (Fig. 15A-15E) neurological tissues (hippocampus, hypothalamus, cortex, cerebellum, spinal cord) and other tissues (Fig. 15F) (liver) and (Fig. 15G) kidney. Data are presented as mean±SEM of the values from six mice; *p<0.05, **p<0.01, ***p<0.001, ns-not significant (based on one-way ANOVA).
[0037] DETAILED DESCRIPTION
[0038] A new gene therapy program is needed for DNA repair disorders (such as Cockayne syndrome (CS) and Xeroderma Pigmentosum (XP) with combined CS Xeroderma Pigmentosum-Cockayne Syndrome (XP-CS) and / or Xeroderma Pigmentosum-neurologic disease (XP-ND). In addition to the two classic versions of CS ERCC6 / CSB and ERCC8 / CSA, deficiencies are also observed in four Xeroderma Pigmentosum (XP) proteins (XPB, XPD, XPF, and XPG) that are associated with Xeroderma Pigmentosum-Cockayne Syndrome XP-CS. Gene therapies are developed for all of these DNA repair disorders.
[0039] In building a gene therapy program for DNA repair disorders (specifically CS and Xeroderma XP, including with combined CS (XP-CS) and / or neurologic disease (XP-ND)), a novel small promoter was needed that the promoter + transgene cDNA can fit within the packaging constraints of approximately 4.9 kilobases of adeno-associated virus (AAV) vectors for Cockayne syndrome type B (CSB) gene therapy (delivery of the ERCC6 cDNA sequence. Further, conserving space within gene therapy vectors is of widespread interest to the field both for gene replacement type gene therapies as well as for genome editing methodologies. The present novel CP040 mini-promoter was created from a combination of upstream 5' untranslated regions (UTRs) from two of the genes of interest, namely, ERCC6 / CSB and ERCC8 / CSA. In certain embodiments, the remainder of the promoter contains a variety of transcription factor binding sites. This technology solves a persistent problem in the gene therapy field related to the limited space available in gene delivery vectors.
[0040] Others have used the Synapsin promoter (448 bp) that typically restricts expression to neuronal cells or a CMV promoter (595 bp) that confers more ubiquitous expression across tissues and cell types. Both of these, however, are too large to fit within the packaging constraints for AAV when the cDNA to be delivered is larger than ~ 4600 bp. Thus, although these existing options are smaller than some promoters, they still remain too large for many applications - including the present use.
[0041] The present inventors engineered a novel mini -promoter. This mini -promoter was created from a combination of upstream 5' untranslated regions (UTRs) from two of the genes of interest to the ERCC6 / CSB and ERCC8 / CSA program (Figures 1A-1C). The remainder of the mini-promoter contains three spacer regions that each contain a high concentration of transcription factor binding site consensus sequences (Figure 2, Variant 1). The initial mini-promoter (Variant 1) was designed and synthesized into a plasmid along with a codon optimized (co) version of ERCC6 (the gene that encodes CSB) coERCC6. Variations of this initial mini-promoter are also engineered (Variants 2-4).
[0042] The present invention provides the use of gene promoter sequences that can activate, enhance, or repress the expression of a genetic sequence at the onset of or during the pathological progression of a genetic disorder. This invention is directly applicable to the field of gene therapy. It has potential commercial value as a tool in the development of regulated gene therapy approaches to genetic disorders. This invention provides a means on which to dose or regulate the expression of therapeutic genes during the course of disease based on the state or rate of progression of the genetic disorder. For example, during disease progression, the gene promoter sequences can activate the expression of a therapeutic gene. If the therapy halts the progression of disease, activity of the gene promoter sequence would wane, resulting in reduced / limited therapeutic gene expression. This disease-regulated dynamic approach to therapeutic gene expression is unique and needed in the gene therapy field.
[0043] The invention is applicable to the field of gene therapy. It satisfies a need for the regulation and / or dosing of therapeutic genes after delivery into the brain. Gene expression analyses have shown that at the onset and / or during the course of genetic disorder progression, a number of gene promoter sequences act to "turn-on," enhance, repress or "shut-off the expression of genes. The present inventors have identified and cloned gene promoter sequences that activate and / or enhance gene expression during genetic disorder progression in cell and animal models of genetic disorder. These gene promoter sequences can also activate and / or enhance the expression of a reporter gene during genetic disorder progression in models of genetic disorder. This approach is unique and different from the current systems used to artificially regulate therapeutic gene expression in that it relies on genetic disorder related molecular events to modulate the expression of the therapeutic genes. These events, and thus the activity of the specific gene promoter sequences, are controlled / induced by intracellular and / or extracellular signals that are dependent on the onset and / or progression of disease.
[0044] Promoter Sequences In certain embodiments, the present provides the following promoter sequences that activate and / or enhance gene expression: SEQ ID NO: 1, SEQ ID NO:2, SEQ ID NO:3, SEQ ID NON, SEQ ID NO:5, or SEQ ID NO:6.
[0045] Full-length mini-promoter =
[0046] CTCGCCCTGCT C AAAGGCGTGAGCCCGTCCCTCCAAACCACGTCTTGCCCCTTCT
[0047] TGAC C GACAAT AT C T CAAT TAG T CAGCAACACCAT TGGCCGGATATATCACCGAC T GA
[0048] CCTCTGCTGTTCC (SEQ ID NO: 1)
[0049] Spacer 1 = CTCGCCCTGCTC (SEQ ID NO:2)
[0050] CSA UTR =
[0051] AAAGGCGTGAGCCCGTCCCTCCAAACCACGTCTTGCCCCTTCTT (SEQ
[0052] ID NON)
[0053] Spacer 2 = GACCGACAATATCTCAATTAGTCAGCAA (SEQ ID NON)
[0054] CSB UTR = CACCATTGGCCGGATATATCACC (SEQ ID NO:5)
[0055] Spacer 3 = GACTGACCTCTGCTGTTCC (SEQ ID NO:6)
[0056] In certain aspects, provided herein is a promoter sequence consisting of a nucleic acid of between 12 and 300 nucleotides in length, wherein the nucleic acid comprises (or consists of) a polynucleotide having at least 90% identity to SEQ ID NON, SEQ ID NON, SEQ ID NON, SEQ ID NON, SEQ ID NON, or SEQ ID NON. In certain aspects, the nucleic acid is between 60 and 120 nucleotides in length.
[0057] In certain aspects, the polynucleotide sequence has at least 90% identity to SEQ ID NON.
[0058] In certain aspects, polynucleotide sequence has at least 90% identity to SEQ ID NON.
[0059] In certain aspects, polynucleotide sequence has at least 90% identity to SEQ ID NON.
[0060] In certain aspects, polynucleotide sequence has at least 90% identity to SEQ ID NON.
[0061] In certain aspects, polynucleotide sequence has at least 90% identity to SEQ ID NON.
[0062] In certain aspects, polynucleotide sequence has at least 90% identity to SEQ ID NON operably linked to SEQ ID NON. In certain aspects, polynucleotide sequence has at least 90% identity to SEQ ID NO:3 operably linked to SEQ ID NO:4 operably linked to CSB UTR SEQ ID NO:5.
[0063] In certain aspects, polynucleotide sequence has at least 90% identity to SEQ ID NO:2 operably linked to CSA UTR SEQ ID NO:3 operably linked to SEQ ID NO:4 operably linked to CSB UTR SEQ ID NO:5.
[0064] In certain aspects, polynucleotide sequence has at least 90% identity to CSA UTR SEQ ID NO:3 operably linked to SEQ ID NO:4 operably linked to CSB UTR SEQ ID NO:5 operably linked to SEQ ID NO:6.
[0065] In certain aspects, polynucleotide sequence has at least 90% identity to SEQ ID NO: 1.
[0066] In certain aspects, polynucleotide has at least 100% identity to SEQ ID NO: 2.
[0067] In certain aspects, polynucleotide has 100% identity to SEQ ID NO: 3.
[0068] In certain aspects, polynucleotide has 100% identity to SEQ ID NO: 4.
[0069] In certain aspects, polynucleotide has 100% identity to SEQ ID NO: 5.
[0070] In certain aspects, polynucleotide has 100% identity to SEQ ID NO: 6.
[0071] In certain aspects, polynucleotide sequence has 100% identity to SEQ ID NO: 3 operably linked to SEQ ID NO:4 operably linked to CSB UTR SEQ ID NO:5.
[0072] In certain aspects, polynucleotide sequence has 100% identity to SEQ ID NO:2 operably linked to CSA UTR SEQ ID NO:3 operably linked to SEQ ID NO:4 operably linked to CSB UTR SEQ ID NO:5.
[0073] In certain aspects, polynucleotide sequence has 100% identity to CSA UTR SEQ ID NO:3 operably linked to SEQ ID NO:4 operably linked to CSB UTR SEQ ID NO:5 operably linked to SEQ ID NO:6.
[0074] In certain aspects, polynucleotide sequence has 100% identity to SEQ ID NO: 1.
[0075] In certain aspects, polynucleotide sequence has at least 90% identity to SEQ ID NO:22, SEQ ID NO:23 or SEQ ID NO:24. In certain aspects, polynucleotide sequence has at 100% identity to SEQ ID NO:22, SEQ ID NO:23 or SEQ ID NO:24.
[0076] Codon-optimized ERCC6 / CSB sequence
[0077] ATGCCGAACGAGGGAATCCCTCACAGCAGCCAGACACAAGAGCAGGACT GCCTGCAGAGCCAGCCTGTGTCCAACAACGAGGAAATGGCCATCAAGCAAGAGT CTGGCGGCGACGGCGAGGTGGAAGAGTACCTGTCTTTTAGAAGCGTTGGCGACG GCCTGAGCACATCTGCTGTGGGATGTGCTTCTGCCGCTCCTAGAAGAGGACCTGC TCTGCTGCACATCGACCGGCATCAGATTCAGGCCGTGGAACCTTCTGCACAGGCC CTGGAACTGCAAGGCCTGGGAGTCGATGTGTACGACCAGGACGTTCTGGAACAG
[0078] GGCGTGCTGCAACAGGTGGACAATGCCATTCACGAGGCCAGCAGAGCCTCTCAG
[0079] CTGGTGGACGTGGAAAAAGAATATCGGAGCGTGCTGGACGACCTGACCAGCTGT
[0080] ACAACCAGCCTGCGGCAGATCAACAAGATCATCGAGCAGCTGTCTCCCCAGGCC
[0081] GCCACCAGCAGAGACATCAACAGAAAGCTGGACAGCGTGAAGCGCCAGAAGTA
[0082] CAACAAAGAGCAGCAGCTGAAGAAGATCACCGCCAAGCAGAAACATCTGCAGG
[0083] CCATTCTCGGCGGAGCCGAAGTGAAGATCGAACTGGATCACGCCAGCCTGGAAG
[0084] AGGATGCCGAACCTGGACCAAGCAGCCTGGGCTCTATGCTGATGCCCGTGCAAG
[0085] AGACAGCCTGGGAAGAACTGATCCGGACCGGCCAGATGACCCCTTTCGGCACAC
[0086] AGATCCCTCAGAAGCAAGAGAAGAAACCCCGGAAGATCATGCTGAATGAGGCCA
[0087] GCGGCTTCGAGAAGTACCTGGCCGATCAGGCCAAGCTGAGCTTCGAGAGAAAGA
[0088] AGCAGGGCTGCAACAAGAGAGCCGCCAGAAAAGCTCCCGCTCCTGTGACACCTC
[0089] CTGCTCCAGTGCAGAACAAGAACAAGCCCAACAAGAAAGCCCGGGTGCTGAGCA
[0090] AGAAAGAGGAACGCCTGAAGAAACACATCAAGAAGCTGCAGAAGCGGGCCCTG
[0091] CAGTTTCAAGGCAAAGTGGGCCTGCCTAAGGCCAGAAGGCCTTGGGAAAGCGAC
[0092] ATGAGGCCTGAGGCCGAGGGCGATTCTGAGGGCGAAGAGAGCGAGTACTTCCCC
[0093] ACCGAGGAAGAAGAAGAGGAAGAGGACGACGAGGTTGAGGGCGCCGAAGCTGA
[0094] TCTTAGCGGAGATGGCACCGACTACGAGCTGAAGCCTCTGCCTAAAGGCGGCAA
[0095] GAGGCAGAAAAAGGTGCCAGTGCAAGAAATCGACGACGACTTCTTCCCATCCTC
[0096] CGGCGAAGAAGCTGAGGCCGCCTCTGTTGGAGAAGGCGGCGGAGGCGGAAGAA
[0097] AAGTGGGCAGATACAGAGATGACGGCGACGAGGACTACTACAAGCAGCGGCTG
[0098] AGAAGGTGGAACAAGCTGAGACTGCAGGACAAAGAGAAGCGGCTCAAGCTGGA
[0099] AGATGACAGCGAGGAATCCGACGCCGAGTTCGACGAGGGCTTTAAGGTGCCCGG
[0100] CTTCCTGTTTAAGAAGCTGTTCAAGTACCAGCAGACCGGCGTGCGGTGGCTGTGG
[0101] GAACTGCATTGTCAACAGGCAGGCGGAATCCTGGGCGACGAAATGGGACTGGGC
[0102] AAGACAATCCAGATTATCGCCTTCCTGGCCGGCCTGTCCTACAGCAAGATCAGAA
[0103] CCCGGGGCAGCAACTACAGATTCGAAGGACTGGGCCCCACCGTGATCGTGTGTC
[0104] CTACAACAGTGATGCACCAGTGGGTCAAAGAATTTCACACCTGGTGGCCTCCATT
[0105] CCGCGTGGCCATTCTGCACGAGACAGGCAGCTACACCCACAAGAAAGAAAAGCT
[0106] GATCCGCGACGTGGCCCACTGCCACGGCATCCTGATCACAAGCTACAGCTACATC
[0107] CGGCTGATGCAGGACGACATCAGCAGATACGACTGGCACTACGTGATCCTGGAC
[0108] GAGGGCCACAAGATTCGGAACCCCAATGCCGCTGTGACCCTGGCCTGCAAGCAG
[0109] TTCAGAACCCCTCACCGGATCATCCTGAGCGGCAGCCCCATGCAAAACAACCTG
[0110] AGAGAGCTGTGGTCCCTGTTCGACTTCATCTTCCCTGGCAAGCTGGGCACCCTGC CTGTGTTCATGGAACAGTTCAGCGTGCCCATCACCATGGGCGGCTACAGCAATGC
[0111] CTCTCCTGTGCAAGTGAAAACCGCCTACAAGTGCGCCTGCGTGCTGCGGGATACC
[0112] ATCAATCCTTACCTGCTGCGGCGGATGAAGTCCGACGTGAAGATGAGCCTGAGC
[0113] CTGCCTGACAAGAACGAACAGGTGCTGTTCTGCCGGCTGACCGACGAGCAGCAC
[0114] AAGGTGTACCAGAACTTCGTGGACAGCAAAGAGGTGTACAGAATCCTGAACGGC
[0115] GAGATGCAGATTTTCAGCGGACTGATCGCCCTGCGGAAAATCTGTAATCACCCCG
[0116] ACCTGTTCAGCGGCGGACCCAAGAATCTGAAGGGCCTGCCAGACGACGAACTCG
[0117] AAGAGGATCAGTTTGGCTACTGGAAGCGGAGCGGCAAGATGATCGTGGTGGAAA
[0118] GCCTGCTGAAGATTTGGCACAAGCAGGGCCAGAGAGTGCTGCTGTTCAGCCAGA
[0119] GCAGACAGATGCTGGATATCCTGGAAGTGTTCCTGCGGGCCCAGAAGTATACCT
[0120] ACCTGAAGATGGACGGCACCACCACAATCGCCAGCAGACAGCCTCTGATCACCC
[0121] GGTACAATGAGGATACCTCCATCTTCGTCTTTCTGCTGACCACCAGAGTCGGCGG
[0122] ACTGGGCGTTAACCTGACTGGCGCCAATCGGGTCGTGATCTATGACCCCGACTGG
[0123] AACCCCAGCACCGACACACAGGCTAGAGAGAGAGCTTGGAGAATCGGCCAGAA
[0124] AAAGCAAGTGACCGTGTACCGGCTGCTGACCGCCGGAACCATCGAAGAGAAAAT
[0125] CTACCACCGGCAGATTTTTAAGCAGTTCCTGACCAACAGGGTGCTGAAGGACCCC
[0126] AAGCAGAGAAGATTCTTCAAGAGCAACGACCTGTATGAGCTGTTCACCCTGACA
[0127] AGCCCCGATGCCAGCCAGTCTACAGAGACAAGCGCCATCTTTGCCGGCACCGGC
[0128] TCTGATGTGCAGACCCCTAAGTGTCACCTGAAGCGGAGAATCCAGCCTGCCTTTG
[0129] GCGCCGATCACGACGTGCCCAAGCGGAAGAAGTTCCCCGCCAGCAACATCAGCG
[0130] TGAACGATGCCACCAGCTCCGAGGAAAAGAGCGAAGCCAAAGGCGCTGAAGTG
[0131] AACGCCGTGACCAGCAACAGAAGCGACCCTCTGAAGGACGACCCTCACATGAGC
[0132] AGCAACGTGACCTCCAACGACCGGCTGGGAGAAGAGACAAATGCCGTGTCTGGC
[0133] CCCGAGGAACTGAGCGTGATATCTGGCAATGGCGAGTGCAGCAACAGCTCTGGC
[0134] ACCGGCAAGACCTCTATGCCATCTGGCGACGAGAGCATCGACGAGAAGCTGGGC
[0135] CTGTCTTACAAGAGAGAGCGGCCTTCTCAGGCCCAGACCGAGGCCTTTTGGGAG
[0136] AACAAGCAGATGGAAAACAACTTCTACAAGCACAAGTCCAAGACCAAGCACCAC
[0137] AGCGTGGCCGAAGAGGAAACCCTGGAAAAGCACCTGAGGCCTAAGCAGAAGCC
[0138] CAAGAACTCCAAGCACTGCCGGGATGCCAAGTTCGAAGGCACAAGAATCCCACA
[0139] CCTCGTGAAGAAGCGCAGATACCAGAAGCAGGACAGCGAGAACAAGTCCGAGG
[0140] CCAAAGAACAGTCCAACGACGACTACGTGCTGGAAAAACTGTTCAAGAAAAGCG
[0141] TGGGCGTGCACAGCGTGATGAAGCACGACGCCATTATGGACGGCGCTAGCCCCG
[0142] ATTATGTGCTGGTGGAAGCCGAAGCCAACAGAGTGGCACAGGATGCCCTGAAGG
[0143] CCCTGAGACTGTCCAGACAGAGATGTCTGGGAGCTGTGTCCGGCGTGCCAACAT GGACAGGACACAGAGGAATTTCTGGCGCCCCTGCCGGCAAGAAGTCTCGCTTTG
[0144] GCAAGAAGAGAAACAGCAACTTCTCCGTGCAGCACCCCTCCAGCACAAGCCCTA
[0145] CAGAGAAGTGCCAGGACGGCATCATGAAGAAAGAAGGCAAGGACAACGTCCCC
[0146] GAGCACTTTAGCGGCAGAGCCGAGGATGCTGATAGCTCTAGTGGACCTCTGGCC
[0147] AGCAGTTCCCTGCTGGCTAAGATGAGAGCCCGGAACCACCTGATCCTGCCAGAG
[0148] AGACTGGAAAGCGAGTCCGGCCATCTGCAAGAAGCTAGCGCACTGCTGCCTACC
[0149] ACCGAGCACGATGATCTGCTGGTCGAGATGCGGAACTTCATTGCCTTTCAAGCCC
[0150] ACACAGACGGCCAGGCCTCTACCAGAGAAATCCTGCAAGAGTTTGAGAGCAAGC
[0151] TGTCCGCCTCTCAGAGCTGCGTGTTCAGAGAGCTGCTGAGAAACCTGTGCACCTT
[0152] CCACAGAACCAGCGGAGGCGAAGGCATCTGGAAGCTGAAACCCGAGTACTGCTG
[0153] ATGATAG (SEQ ID NO: 22)
[0154] Codon-optimized ERCC8 / CSA sequence
[0155] ATGCCACCATGCTGGGCTTTCTGAGCGCCAGACAGACCGGCCTGGAAGA
[0156] TCCCCTGAGACTGAGAAGGGCCGAGAGCACCAGAAGAGTGCTGGGCCTGGAACT
[0157] GAACAAGGACCGGGACGTGGAACGGATCCACGGCGGAGGCATCAACACCCTGG
[0158] ACATCGAGCCTGTGGAAGGCCGGTACATGCTGAGCGGCGGATCCGATGGCGTGA
[0159] TCGTGCTGTACGACCTGGAAAACAGCAGCCGGCAGAGCTACTACACATGCAAGG
[0160] CCGTGTGCAGCATCGGCAGGGACCACCCCGATGTGCACCGGTACAGCGTGGAAA
[0161] CCGTGCAGTGGTATCCCCACGACACCGGCATGTTCACCAGCAGCAGCTTCGACA
[0162] AGACCCTGAAAGTGTGGGACACCAATACCCTGCAGACCGCCGACGTGTTCAACT
[0163] TCGAGGAAACAGTGTACAGCCACCACATGAGCCCCGTGTCCACCAAGCACTGTC
[0164] TGGTGGCCGTGGGCACCAGAGGCCCTAAGGTGCAGCTGTGCGATCTGAAGTCCG
[0165] GCAGCTGCAGCCACATCCTGCAGGGCCACAGACAGGAAATCCTGGCCGTGTCCT
[0166] GGTCCCCCAGATACGACTACATCCTGGCCACCGCCAGCGCCGACAGCAGAGTGA
[0167] AACTGTGGGATGTGCGGAGAGCCAGCGGCTGCCTGATCACCCTGGATCAGCACA
[0168] ACGGCAAGAAAAGCCAGGCCGTGGAAAGCGCCAACACCGCCCACAATGGCAAA
[0169] GTGAACGGCCTGTGCTTCACCAGCGACGGCCTGCATCTGCTGACAGTGGGCACC
[0170] GACAACCGGATGCGGCTGTGGAACAGCTCCAACGGCGAGAATACCCTCGTGAAC
[0171] TACGGCAAAGTGTGCAACAACAGCAAGAAGGGCCTGAAGTTCACCGTGTCTTGC
[0172] GGCTGCAGCAGCGAGTTCGTGTTCGTGCCCTACGGCAGCACAATCGCCGTGTACA
[0173] CCGTGTACTCCGGCGAGCAGATCACAATGCTGAAGGGCCACTACAAGACCGTGG
[0174] ACTGCTGCGTGTTCCAGAGCAACTTCCAGGAACTGTACAGCGGCAGCCGGGACT
[0175] GCAATATCCTGGCCTGGGTGCCCAGCCTGTACGAGCCCGTGCCTGACGACGATG AGACAACCACCAAGAGCCAGCTGAACCCCGCCTTCGAGGATGCCTGGTCCAGCT
[0176] CCGATGA (SEQ ID NO: 23)
[0177] Codon-optimized ERCC5 / XPG sequence
[0178] ATGGGAGTACAGGGACTCTGGAAACTGCTGGAATGCAGTGGGAGACAAG
[0179] TGAGTCCGGAAGCACTGGAAGGTAAGATTCTGGCTGTCGATATCTCTATCTGGCT
[0180] CAATCAGGCCCTGAAAGGGGTCAGGGATAGACATGGTAACTCCATCGAAAATCC
[0181] CCACCCGCTTACCCTTTTTCACCGCCTGTGCAAGCTTCTGTTCTTCAGAATCAGAC
[0182] CAATCTTCGTGTTTGACGGTGACGCACCCCTGCTGAAGAAGCAAACACTCGTAAA
[0183] AAGGCGACAACGGAAGGATCTGGCGAGCTCTGACTCCCGGAAGACCACAGAGA
[0184] AGTTGCTCAAGACATTCCTTAAGAGGCAGGCCATTCAAACTAGTTTCCGCTCCCA
[0185] ACGCGATGAAGCCCTGCCTTCTTTGACACAAGTGAGAAGAGAGAATGACCTGTA
[0186] CGTCCTGCCACCGCTCCAGGAGGAAGAGAAACATTCTTCAGAGGAGGAAGATGA
[0187] AAAGGAGTGGCAGGAGCGAATGAATCAGAAACAAGCCCTGCAGGAAGAATTTTT
[0188] TCACAACCCCCAGGCCATCGACATCGAATCCGAGGACTTTAGCTCCCTCCCGCCC
[0189] GAAGTCAAACACGAAATCCTCACAGACATGAAAGAATTTACCAAGCGCCGAAGA
[0190] ACCTTGTTTGAAGCCATGCCAGAAGAGTCCGATGATTTCAGCCAATATCAACTTA
[0191] AGGGACTCTTGAAAAAGAATTATTTGAACCAGCATATAGAGCATGTTCAGAAGG
[0192] AGGTCAATCAGCAGCACAGCGGACATATCCGATCCAGCCATGAAGATGAAGGCG
[0193] GGTTCCTTAAGGAAGTGGAATCCCGAAGGGTCGTATCAGAAGACACTTCTCATTA
[0194] CATACTGATCAAAGGTATTCAGGCAAAAACGGTCGCAGAAGTGGATTCCGAAAG
[0195] TCTCCCAAGCAGCTCAAAAATGCATGGAATGTCATTCGATGTGAAGTCAAGTCCA
[0196] TGTGAGAAACTGAAAACAGAGAAAGAACCAGATGCTACACCCCCCAGCCCTCGA
[0197] ACTCTCTTGGCGATGCAGGCGGCCCTCCTTGGTTCCTCAAGCGAAGAGGAATTGG
[0198] AGTCAGAGAACCGGAGACAGGCACGCGGTCGGAACGCACCCGCAGCCGTTGAC
[0199] GAAGGATCCATTAGCCCCCGCACCCTTAGCGCCATCAAACGGGCCCTCGATGAC
[0200] GACGAGGACGTGAAAGTATGTGCTGGAGACGACGTTCAGACTGGAGGGCCGGGG
[0201] GCCGAAGAGATGAGAATTAACAGCAGCACAGAGAATAGCGATGAAGGCCTCAA
[0202] AGTGAGGGATGGTAAAGGGATTCCATTTACAGCGACTCTGGCGTCCAGCAGCGT
[0203] GAATTCCGCAGAAGAGCATGTCGCATCAACAAATGAAGGCCGCGAACCCACAGA
[0204] CAGTGTGCCGAAGGAGCAAATGTCACTTGTCCACGTCGGAACTGAAGCCTTTCCA
[0205] ATCAGCGACGAATCAATGATAAAAGATCGCAAAGACAGGCTGCCCCTCGAATCC
[0206] GCCGTAGTGAGGCACTCCGACGCTCCGGGGCTTCCCAATGGTAGAGAGCTTACTC
[0207] CAGCAAGCCCAACCTGTACTAATTCAGTTTCTAAAAACGAGACTCATGCAGAAGT GCTCGAGCAGCAGAACGAGCTTTGCCCATATGAAAGTAAATTTGATAGTTCACTC
[0208] CTTTCTTCTGACGACGAGACAAAGTGTAAGCCAAATTCCGCATCTGAAGTGATAG
[0209] GCCCTGTATCCCTCCAGGAAACCTCAAGCATAGTGTCCGTTCCATCTGAAGCGGT
[0210] TGACAACGTAGAGAATGTGGTTTCATTCAACGCAAAAGAACACGAGAACTTCCT
[0211] CGAGACCATTCAGGAGCAGCAAACAACTGAGTCTGCCGGCCAGGATCTGATTAG
[0212] CATACCAAAAGCCGTCGAACCAATGGAGATCGACAGCGAGGAGTCTGAATCCGA
[0213] TGGGAGCTTCATCGAAGTGCAGAGCGTGATTAGCGACGAGGAACTGCAGGCAGA
[0214] ATTCCCGGAAACCAGCAAACCACCGAGCGAACAGGGTGAGGAGGAACTTGTCGG
[0215] TACGAGGGAAGGTGAGGCTCCTGCCGAGAGTGAGAGCCTGCTTAGGGATAATAG
[0216] CGAGAGGGACGACGTCGATGGAGAGCCCCAGGAAGCCGAGAAGGACGCTGAAG
[0217] ATTCTCTCCACGAATGGCAGGACATTAATCTGGAGGAGCTGGAAACCTTGGAAA
[0218] GCAATCTGCTTGCCCAACAAAACTCACTCAAGGCGCAGAAACAACAGCAGGAAC
[0219] GCTTTGCAGCCACCGTTACCGGACAGATGTTTCTCGAATCTCAAGAGTTGCTTCG
[0220] CCTGTTCGGCATCCCTTACATCCAAGCACCTATGGAGGCTGAAGCCCAATGTGCA
[0221] GTGCTGGACCTCACAGATCAGACTAGCGGGACAATCACGGATGATTCAGATATC
[0222] TGGCTCTTTGGCGCTAGACACGTCTACCGCAATTTTTTCAACAAGAATAAGTTCG
[0223] TCGAGTATTATCAGTATGTGGACTTCCATAACCAGCTGGGACTCGACCGCAATAA
[0224] ACTGATCAATTTGGCGTACCTTCTGGGGAGCGACTACACAGGGAACACCAATTGC
[0225] GGCCTGTGCACAGCAATGGAAATTTTGAACGAATTCCCCGGCCACGGTCTGGAG
[0226] CCGCTCCTGAAATTTAGCGAGTGGTGGCACGAAGCTCAGAAGAATCCTAAGATT
[0227] AGACCAAACCCGCATGACACCAAGGTCAAAAAAAAGCTTAGGACACTGCAGCTC
[0228] ACACCCGGGTTCCCGAATCCCGCGGTTGCAGAAGCTTACCTCAAGCCAGTCGTTG
[0229] ACGATTCCAAGGGTAGCTTCTTGTGGGGAAAACCGGATCTTGATAAAATCTCCGA
[0230] GTTCTGTCAGAGATATTTTGGTTGGAATAGAACTAAGACCGACGAGTCTCTTTTC
[0231] CCTGTTCTGAAACAACTGGATGCACAGCAGACACAACTGAGAATCGATTCCTTTT
[0232] TCCGCCTGGCTCAACAGGAAAAAGAAGACGCCAAGCGCATAAAGAGCCAACGC
[0233] CTGAACCGGGCAGTGACCTGCATGCTGAGAAAAGAAAAGGAAGCAGCAGCAAG
[0234] TGAGATTGAAGCCGTGAGTGTCGCTATGGAGAAAGAATTTGAGCTCCTGGATAA
[0235] GGCTAAAAGAAAAACCCAAAAAAGAGGGATTACTAATACCCTTGAAGAGTCATC
[0236] TAGCTTGAAGCGAAAAAGATTGTCCGATTCCAAGCGCAAGAACACGTGCGGAGG
[0237] ATTTCTCGGCGAAACGTGCCTTAGCGAATCCTCCGACGGGTCAAGTTCAGAGGAC
[0238] GCCGAATCTTCTTCTTTGATGAATGTGCAAAGAAGAACTGCCGCCAAAGAGCCCA
[0239] AGACATCTGCAAGTGACAGCCAGAACAGCGTTAAGGAAGCTCCAGTTAAAAATG
[0240] GCGGCGCCACAACATCAAGCTCTTCCGACTCCGACGACGATGGTGGAAAAGAAA AGATGGTACTCGTGACAGCTCGGTCTGTTTTTGGTAAGAAAAGAAGAAAGCTTCG GAGGGCTAGAGGTCGGAAAAGAAAAACCGGGGATCCAGACATGATAAGATACA TTGATGAGTTTGGACAAACCACAACTAGAATGCAGTGA (SEQ ID NO: 24)
[0241] "Promoter" refers to a nucleotide sequence, usually upstream (5') to its coding sequence, that controls the expression of the coding sequence by providing the recognition for RNA polymerase and other factors required for proper transcription. "Promoter" includes a minimal promoter that is a short DNA sequence comprised of a TATA-box and other sequences that serve to specify the site of transcription initiation, to which regulatory elements are added for control of expression. "Promoter" also refers to a nucleotide sequence that includes a minimal promoter plus regulatory elements that is capable of controlling the expression of a coding sequence or functional RNA. This type of promoter sequence consists of proximal and more distal upstream elements, the latter elements often referred to as enhancers. Accordingly, an "enhancer" is a DNA sequence that can stimulate promoter activity and may be an innate element of the promoter or a heterologous element inserted to enhance the level or tissue specificity of a promoter. It is capable of operating in both orientations (normal or flipped) and is capable of functioning even when moved either upstream or downstream from the promoter. Both enhancers and other upstream promoter elements bind sequence-specific DNA-binding proteins that mediate their effects. Promoters may be derived in their entirety from a native gene or be composed of different elements derived from different promoters found in nature, or even be comprised of synthetic DNA segments. A promoter may also contain DNA sequences that are involved in the binding of protein factors that control the effectiveness of transcription initiation in response to physiological or developmental conditions.
[0242] As used herein, "biologically active" means that the promoter has at least about 0.1%, 10%, 25%, 50%, 75%, 80%, 85%, even 90% or more, e.g. 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100% of the activity of the CCT promoter comprising SEQ ID NO: 1 or SEQ ID NO:2. The activity of a promoter can be determined by methods well known in the art. For example, see Sambrook et al., Molecular Cloning: A Laboratory Manual (1989). Promoters of the present invention that are not identical to SEQ ID NO: 1 or SEQ ID NO:2, but retain comparable biological activity, are called variant promoters. The nucleotide sequences of the invention include both naturally occurring sequences as well as recombinant forms. The invention encompasses isolated or substantially purified nucleic acid compositions. In the context of the present invention, an “isolated” or “purified” DNA molecule or RNA molecule is a DNA molecule or RNA molecule that exists apart from its native environment and is therefore not a product of nature. An isolated DNA molecule or RNA molecule may exist in a purified form or may exist in a non-native environment such as, for example, a transgenic host cell. For example, an “isolated” or “purified” nucleic acid molecule or biologically active portion thereof, is substantially free of other cellular material, or culture medium when produced by recombinant techniques, or substantially free of chemical precursors or other chemicals when chemically synthesized. In one embodiment, an “isolated” nucleic acid is free of sequences that naturally flank the nucleic acid (i.e., sequences located at the 5' and 3' ends of the nucleic acid) in the genomic DNA of the organism from which the nucleic acid is derived. For example, in various embodiments, the isolated nucleic acid molecule can contain less than about 5 kb, 4 kb, 3 kb, 2 kb, 1 kb, 0.5 kb, or 0.1 kb of nucleotide sequences that naturally flank the nucleic acid molecule in genomic DNA of the cell from which the nucleic acid is derived. Fragments and variants of the disclosed nucleotide sequences are also encompassed by the present invention. As used herein, the terms “fragment” or “portion” mean a full length or less than full length of the nucleotide sequence.
[0243] The term "nucleic acid" refers to deoxyribonucleic acid (DNA) or ribonucleic acid (RNA) and polymers thereof in either single- or double-stranded form, composed of monomers (nucleotides) containing a sugar, phosphate and a base that is either a purine or pyrimidine. Unless specifically limited, the term encompasses nucleic acids containing known analogs of natural nucleotides that have similar binding properties as the reference nucleic acid and are metabolized in a manner similar to naturally occurring nucleotides. Unless otherwise indicated, a particular nucleic acid sequence also encompasses conservatively modified variants thereof (e.g., degenerate codon substitutions) and complementary sequences, as well as the sequence explicitly indicated. Specifically, degenerate codon substitutions may be achieved by generating sequences in which the third position of one or more selected (or all) codons is substituted with mixed-base and / or deoxyinosine residues. A “nucleic acid fragment” is a portion of a given nucleic acid molecule.
[0244] “Naturally occurring,” “native,” or “wild-type” is used to describe an object that can be found in nature as distinct from being artificially produced. For example, a protein or nucleotide sequence present in an organism (including a virus), which can be isolated from a source in nature and that has not been intentionally modified by a person in the laboratory, is naturally occurring.
[0245] Expression Cassettes and Vectors
[0246] In certain embodiments, the present invention provides vectors and expression cassettes containing the promoters described above.
[0247] Expression Cassettes
[0248] "Expression cassette" as used herein means a DNA sequence capable of directing expression of a particular nucleotide sequence in an appropriate host cell, comprising a promoter operably linked to the nucleotide sequence of interest that can be operably linked to termination signals. It also typically comprises sequences required for proper translation of the nucleotide sequence. The coding region usually codes for a protein of interest but may also code for a functional RNA of interest, for example antisense RNA or a non-translated RNA, in the sense or antisense direction. The expression cassette comprising the nucleotide sequence of interest may be chimeric, meaning that at least one of its components is heterologous with respect to at least one of its other components. The expression cassette may also be one that is naturally occurring but has been obtained in a recombinant form useful for heterologous expression. Such expression cassettes will comprise the transcriptional initiation region linked to a nucleotide sequence of interest. Such an expression cassette may be provided with a plurality of restriction sites for insertion of the gene of interest to be under the transcriptional regulation of the regulatory regions. The expression cassette may additionally contain selectable marker genes.
[0249] In certain aspects, provided herein is an expression cassette comprising the promoter of as described herein, wherein the promoter is operably linked to a preselected DNA segment encoding a protein or heterologous RNA transcript. The promoter in certain aspects, are the following: a nucleic acid of between 12 and 300 nucleotides in length, wherein the nucleic acid comprises a polynucleotide having at least 90% identity to SEQ ID NO: 1, SEQ ID NO:2, SEQ ID NO:3, SEQ ID NO:4, SEQ ID NO:5, or SEQ ID NO:6, a polynucleotide sequence having at least 90% identity to SEQ ID NO:2, a polynucleotide sequence having at least 90% identity to SEQ ID NO:3, a polynucleotide sequence having at least 90% identity to SEQ ID NO:4, a polynucleotide sequence having at least 90% identity to SEQ ID NO:5, a polynucleotide sequence having at least 90% identity to SEQ ID NO: 6, a polynucleotide sequence having at least 90% identity to SEQ ID NO: 3 operably linked to SEQ ID NO:5, a polynucleotide sequence having at least 90% identity to SEQ ID NO: 3 operably linked to SEQ ID NON operably linked to CSB UTR SEQ ID NO:5, a polynucleotide sequence having at least 90% identity to SEQ ID NO:2 operably linked to CSA UTR SEQ ID NON operably linked to SEQ ID NON operably linked to CSB UTR SEQ ID NON, a polynucleotide sequence having at least 90% identity to CSA UTR SEQ ID NON operably linked to SEQ ID NON operably linked to CSB UTR SEQ ID NON operably linked to SEQ ID NON, a polynucleotide sequence having at least 90% identity to SEQ ID NON, a polynucleotide having at least 100% identity to SEQ ID NO: 2, a polynucleotide having 100% identity to SEQ ID NO: 3, a polynucleotide having 100% identity to SEQ ID NO: 4, a polynucleotide having 100% identity to SEQ ID NO: 5, a polynucleotide having 100% identity to SEQ ID NO: 6, a polynucleotide sequence having 100% identity to SEQ ID NON operably linked to SEQ ID NON operably linked to CSB UTR SEQ ID NON, a polynucleotide sequence having 100% identity to SEQ ID NON operably linked to CSA UTR SEQ ID NON operably linked to SEQ ID NON operably linked to CSB UTR SEQ ID NON, a polynucleotide sequence having 100% identity to CSA UTR SEQ ID NON operably linked to SEQ ID NON operably linked to CSB UTR SEQ ID NON operably linked to SEQ ID NON, or a polynucleotide sequence having 100% identity to SEQ ID NON.
[0250] In certain aspects, the preselected DNA segment comprises a selectable marker gene or a reporter gene.
[0251] In certain aspects, the preselected DNA segment encodes a therapeutic composition.
[0252] In certain aspects, therapeutic composition is an RNAi molecule.
[0253] Vectors
[0254] In certain aspects, provided herein is a vector comprising an expression cassette described above. In certain aspects, the vector is an adeno-associated virus (AAV) vector. A “vector" is defined to include, inter alia, any viral vector, as well as any plasmid, cosmid, phage or binary vector in double or single stranded linear or circular form that may or may not be self-transmissible or mobilizable, and that can transform prokaryotic or eukaryotic host either by integration into the cellular genome or exist extrachromosomally (e.g., autonomous replicating plasmid with an origin of replication).
[0255] The selection and optimization of a particular expression vector for expressing a specific therapeutic composition (e.g., a protein) in a cell can be accomplished by obtaining the nucleic acid sequence encoding the protein, possibly with one or more appropriate control regions (e.g., promoter, insertion sequence); preparing a vector construct comprising the vector into which is inserted the nucleic acid sequence encoding the protein; transfecting or transducing cultured cells in vitro with the vector construct; and determining whether the protein is present in the cultured cells.
[0256] Vectors for cell gene therapy include viruses, such as replication-deficient viruses. Replication-deficient retroviruses are capable of directing synthesis of all virion proteins but are incapable of making infectious particles. Accordingly, these genetically altered retroviral expression vectors have general utility for high-efficiency transduction of nucleic acid sequences in cultured cells, and specific utility for use in the method of the present invention. Such retroviruses further have utility for the efficient transduction of nucleic acid sequences into cells in vivo. Retroviruses have been used extensively for transferring nucleic acid material into cells. Protocols for producing replication-deficient retroviruses (including the steps of incorporation of exogenous nucleic acid material into a plasmid, transfection of a packaging cell line with plasmid, production of recombinant retroviruses by the packaging cell line, collection of viral particles from tissue culture media, and infection of the target cells with the viral particles) are well known in the art.
[0257] An advantage of using retroviruses for gene therapy is that the viruses insert the nucleic acid sequence encoding the target protein into the host cell genome, thereby permitting the nucleic acid sequence encoding the target protein to be passed on to the progeny of the cell when it divides. Promoter sequences in the LTR region can enhance expression of an inserted coding sequence in a variety of cell types.
[0258] Another viral candidate useful as an expression vector for transformation of cells is an adenovirus (Ad), which is a double-stranded DNA virus. The adenovirus is infective in a wide range of cell types, including, for example, muscle and endothelial cells. Adenoviruses are double-stranded linear DNA viruses with a 36 kb genome. Several features of adenovirus have made them useful as transgene delivery vehicles for therapeutic applications, such as facilitating in vivo gene delivery. Recombinant adenovirus vectors have been shown to be capable of efficient in situ gene transfer to parenchymal cells of various organs, including the lung, brain, pancreas, gallbladder, and liver. This has allowed the use of these vectors in methods for treating inherited genetic diseases, such as cystic fibrosis, where vectors may be delivered to a target organ.
[0259] Like the retrovirus, the adenovirus genome is adaptable for use as an expression vector for gene therapy, z.e., by removing the genetic information that controls production of the virus itself. Because the adenovirus functions in an extrachromosomal fashion, the recombinant adenovirus does not have the theoretical problem of insertional mutagenesis.
[0260] Several approaches traditionally have been used to generate recombinant adenoviruses. One approach involves direct ligation of restriction endonuclease fragments containing a nucleic acid sequence of interest to portions of the adenoviral genome. Alternatively, the nucleic acid sequence of interest may be inserted into a defective adenovirus by homologous recombination results. The desired recombinants are identified by screening individual plaques generated in a lawn of complementation cells.
[0261] Examples of appropriate vectors include DNA viruses (e.g., adenoviruses), lentiviral, adeno-associated viral (AAV), poliovirus, HSV, or murine Moloney -based viral vectors, viral vectors derived from Harvey Sarcoma virus, ROUS Sarcoma virus, MPSV or hybrid transposon-based vectors. In one embodiment, the vector is AAV. AAV is a small nonpathogenic virus of the parvoviridae family. AAV is distinct from the other members of this family by its dependence upon a helper virus for replication. The approximately 5 kb genome of AAV consists of one segment of single stranded DNA of either plus or minus polarity. The ends of the genome are short, inverted terminal repeats which can fold into hairpin structures and serve as the origin of viral DNA replication. Physically, the parvovirus virion is non-enveloped and its icosohedral capsid is approximately 20 nm in diameter.
[0262] To-date many serologically distinct AAVs have been identified and have been isolated from humans or primates. For example, the genome of AAV2 is 4680 nucleotides in length and contains two open reading frames (ORFs). The left ORF encodes the non- structural Rep proteins, Rep 40, Rep 52, Rep 68 and Rep 78, which are involved in regulation of replication and transcription in addition to the production of single-stranded progeny genomes. Rep68 / 78 has also been shown to possess NTP binding activity as well as DNA and RNA helicase activities. The Rep proteins possess a nuclear localization signal as well as several potential phosphorylation sites. Mutation of one of these kinase sites resulted in a loss of replication activity. The ends of the genome are short, inverted terminal repeats (ITR) which have the potential to fold into T-shaped hairpin structures that serve as the origin of viral DNA replication. Within the ITR region two elements have been described which are central to the function of the ITR, a GAGC repeat motif and the terminal resolution site (trs). The repeat motif has been shown to bind Rep when the ITR is in either a linear or hairpin conformation. This binding serves to position Rep68 / 78 for cleavage at the trs which occurs in a site- and strand-specific manner. AAV vectors have several features that make it an attractive vector for gene transfer, such as possessing a broad host range, are capable of transduce both dividing and non-dividing cells in vitro and in vivo, and are capable of maintaining high levels of expression of transduced genes.
[0263] In certain embodiments, the viral vector is an AAV vector. An "AAV" vector refers to an adeno-associated virus and may be used to refer to the naturally occurring wild-type virus itself or derivatives thereof. The term covers all subtypes, serotypes and pseudotypes, and both naturally occurring and recombinant forms, except where required otherwise. As used herein, the term "serotype" refers to an AAV which is identified by and distinguished from other AAVs based on capsid protein reactivity with defined antisera, e.g., to date there are at least 12 natural AAV serotypes and many other variants have been generated. For example, serotype AAV9 is used to refer to an AAV which contains capsid proteins encoded from the cap gene of AAV9 and a genome containing 5' and 3' ITR sequences from the same AAV9 serotype. In certain embodiments, the AAV vector is AAV9.
[0264] The abbreviation "rAAV" refers to recombinant adeno-associated virus, also referred to as a recombinant AAV vector (or "rAAV vector"). In one embodiment, the AAV expression vectors are constructed using known techniques to at least provide as operatively linked components in the direction of transcription, control elements including a transcriptional initiation region, the DNA of interest and a transcriptional termination region. The control elements are selected to be functional in a mammalian cell. The resulting construct which contains the operatively linked components is flanked (5' and 3') with functional AAV ITR sequences.
[0265] By "adeno-associated virus inverted terminal repeats" or "AAV ITRs" is meant the art-recognized regions found at each end of the AAV genome which function together in cis as origins of DNA replication and as packaging signals for the virus.
[0266] The nucleotide sequences of AAV ITR regions are known. As used herein, an "AAV ITR" need not have the wild-type nucleotide sequence depicted, but may be altered, e.g., by the insertion, deletion or substitution of nucleotides. Additionally, the AAV ITR may be derived from any of several AAV serotypes, including without limitation, AAV1, AAV2, AAV3, AAV4, AAV5, AAV7, etc. Furthermore, 5' and 3' ITRs which flank a selected nucleotide sequence in an AAV vector need not necessarily be identical or derived from the same AAV serotype or isolate, so long as they function as intended, i.e., to allow for excision and rescue of the sequence of interest from a host cell genome or vector.
[0267] Nucleic acids encoding therapeutic compositions can be engineered into an AAV vector using standard ligation techniques, such as those described in Sambrook and Russell, Molecular Cloning: A Laboratory Manual, Cold Spring Harbor Laboratory Press Cold Spring Harbor, NY (2001). For example, ligations can be accomplished in 20 mM Tris-Cl pH 7.5, 10 mM MgC12, 10 mM DTT, 33 pg / ml BSA, 10 mM-50 mM NaCl, and either 40 pM ATP, 0.01-0.02 (Weiss) units T4 DNA ligase at 0°C (for "sticky end" ligation) or 1 mM ATP, 0.3- 0.6 (Weiss) units T4 DNA ligase at 14°C (for "blunt end" ligation). Intermolecular "sticky end" ligations are usually performed at 30-100 pg / ml total DNA concentrations (5-100 nM total end concentration). AAV vectors which contain ITRs have been described in, e.g., U.S. Pat. No. 5,139,941. In particular, several AAV vectors are described therein which are available from the American Type Culture Collection (" ATCC") under Accession Numbers 53222, 53223, 53224, 53225 and 53226.
[0268] In certain embodiments, the adeno- associated virus packages a full-length genome, i.e., one that is approximately the same size as the native genome and is not too big or too small. In certain embodiments the AAV is not a self-complementary AAV vector.
[0269] The viral vector further includes a promoter for controlling transcription of the heterologous gene. The promoter may be an inducible promoter for controlling transcription of the therapeutic composition. The expression system is suitable for administration to the mammalian recipient.
[0270] In certain embodiments, viral particles are administered. Viral particles are heat stable, resistant to solvents, detergents, changes in pH, temperature, and can be concentrated on CsCl gradients. AAV is not associated with any pathogenic event, and transduction with AAV vectors has not been found to induce any lasting negative effects on cell growth or differentiation. The ITRs have been shown to be the only cis elements required for packaging allowing for complete gutting of viral genes to create vector systems.
[0271] Certain embodiments of the present invention provide a vector that encodes a target molecule, such as an isolated RNAi molecule. As used herein the term “encoded by” is used in a broad sense, similar to the term “comprising” in patent terminology. RNAi molecules include siRNAs, shRNAs and other small RNAs that can or are capable of modulating the expression of a target gene, for example via RNA interference. Such small RNAs include without limitation, shRNAs and miroRNAs (miRNAs).
[0272] “Operably-linked” refers to the association of nucleic acid sequences on single nucleic acid fragment so that the function of one of the sequences is affected by another. For example, a regulatory DNA sequence is said to be “operably linked to” or “associated with” a DNA sequence that codes for an RNA or a polypeptide if the two sequences are situated such that the regulatory DNA sequence affects expression of the coding DNA sequence (i.e., that the coding sequence or functional RNA is under the transcriptional control of the promoter). Coding sequences can be operably-linked to regulatory sequences in sense or antisense orientation. Nucleic acid is "operably linked" when it is placed into a functional relationship with another nucleic acid sequence. Generally, "operably linked" means that the DNA sequences being linked are contiguous. However, enhancers do not have to be contiguous. Linking is accomplished by ligation at convenient restriction sites. If such sites do not exist, the synthetic oligonucleotide adaptors or linkers are used in accordance with conventional practice. Additionally, multiple copies of the nucleic acid encoding enzymes may be linked together in the expression vector. Such multiple nucleic acids may be separated by linkers.
[0273] "Expression" refers to the transcription and / or translation of an endogenous gene or a transgene in cells. For example, in the case of antisense constructs, expression may refer to the transcription of the antisense DNA only. In addition, expression refers to the transcription and stable accumulation of sense (mRNA) or functional RNA. Expression may also refer to the production of protein.
[0274] Transformed Cells
[0275] The present disclosure also provides a host cell containing an expression cassette or vector described herein. In certain aspects, the host cell is a eukaryotic cell.
[0276] In certain aspects, the eukaryotic cell is an animal cell.
[0277] In certain aspects, the animal cell is a mammalian cell.
[0278] In certain aspects, the mammalian cell is a human cell.
[0279] In certain aspects, the transformed cells exhibit significantly increased expression of a reporter gene when introduced into the cells derived from Cockayne Syndrome patients as compared to cells derived from control individuals.
[0280] The cell may be human, and may be from brain, spleen, kidney, lung, heart, or liver, or be a blood cell (e.g., peripheral blood mononuclear cell). The cell type may be a stem or progenitor cell population. In certain aspects, provided herein is a method for producing transformed cells comprising the steps of (i) introducing into cells a recombinant DNA which comprises a novel promoter as described herein that is operably linked to a DNA segment or the expression cassette described above so as to yield transformed cells, and (ii) identifying or selecting a transformed cell line. In certain aspects, the recombinant DNA is expressed so as to impart a phenotypic characteristic to the transformed cells. In certain aspects, provided herein is a transformed cell made by this method.
[0281] Methods of Administering
[0282] In certain aspects, provided herein is a method of treating a genetic disorder in a mammal comprising administering (a) a novel vector as described herein, or (b) a novel transformed cell as described herein.
[0283] In certain aspects, the genetic disorder is spinal muscular atrophy, Leber congenital amaurosis, hemophilia a neurological disorder.
[0284] In certain aspects, the genetic disorder is a neurological disorder.
[0285] In certain aspects, the neurological disorder is Cockayne Syndrome (CS).
[0286] The present invention provides a method of administering a therapeutic composition. In certain embodiments, the therapeutic composition is a transgene, which is a gene encoding a polypeptide that is foreign to the retrovirus from which the vector is primarily derived and has a useful biological activity in the organism into which it is administered (e.g., a therapeutic gene). As used herein, the term "therapeutic gene" refers to a gene whose expression is desired in a cell to provide a therapeutic effect, e.g., to treat a disease.
[0287] Methods of Use
[0288] Certain aspects of the disclosure relate to polynucleotides, polypeptides, vectors, and genetically engineered cells (modified in vivo), and the use of them. In particular, the disclosure relates to a method for gene or protein therapy that is capable of both systemic delivery of a therapeutically effective dose of the therapeutic agent.
[0289] The invention has been described as “comprising” certain steps and / or elements, which those of skill in the art also “consist of’ or “consist essentially of’ those steps and / or elements. As used herein, the transitional term “comprising” is synonymous with “including,” “containing,” or “characterized by,” is inclusive or open-ended and does not exclude additional, unrecited elements or method steps. Where the invention is intended to be more narrowly defined, the terms “consisting of’ or “consisting essentially of’ also are used to describe the invention. As used herein, the transitional phrase “consisting of’ excludes any element, step, or ingredient not specified in the claim. The transitional phrase “consisting essentially of’ limits the scope of a claim to the specified elements or steps and those that do not materially affect the basic and novel characteristics of the claimed invention. As used herein, a claim reciting “consisting essentially of’ occupies a middle ground between closed claims reciting a “consisting of’ format and fully open claims that recite “comprising.”
[0290] The invention may also be further described by means of the following non-limiting examples.
[0291] Example 1 Transfection Methodology
[0292] HEK293 (Human Embryonic Kidney) cells were maintained in DMEM, 10% fetal bovine serum and lx antibiotic / antimycotic at 37 °C in 5% CO2. A day before transfection, cells were seeded in 6-well plate at the density of 0.8xl06per well. Lipofectamine3000 (Thermo) standard transfection protocol was used to transfect two concentrations of CP040 mini-promoter plasmid DNA (250ng and 500ng) in HEK293 cells. After 24 hours of transfection, cells were harvested in Zymo RNA lysis buffer and total RNA was isolated, quantified and 500ng of total RNA was converted to cDNA using Thermo reverse transcription kit. To determine the efficiency of promoter, quantification of downstream gene (codon optimized CSB) levels was performed on the obtained cDNA using three specific Taqman probes (Thermo). 18S ribosomal RNA was used as internal control for the samples. Transfections and analysis were performed in triplicates on different batch / passage of HEK293 cells.
[0293] Example 2 CP040 Mini Promoter: in vitro characterization results
[0294] An exemplary plasmid was engineered that contained the CP040 mini promoter, as described below:
[0295] CP040 Plasmid Sequence
[0296] GAGCGCGACG TAA TACGACT CACTA TAGGGCGAATTGAAGGAAGGCCGTC AAGGCCGCATGTTTAAACTACGTACCTAGGCTTAAGCTCGCCCTGCTC
[0297] AAAGGCGTGAGCCCGTCCCTCCAAACCACGTCTTGCC CCTTCTTGACCGACAATATCTCAATTAGTCAGCAACA CCATTGGCCGGATATATCACCGACTGACCTCTGCTGT
[0298] TCCGAATTCTTAATTAAAAGCTTCTCGAGCTGCCACCATGCCGAACGAGGGAA
[0299] TCCCTCACAGCAGCCAGACACAAGAGCAGGACTGCCTGCAGAGCCAGCCTGTGT CCAACAACGAGGAAATGGCCATCAAGCAAGAGTCTGGCGGCGACGGCGAGGTG GAAGAGTACCTGTCTTTTAGAAGCGTTGGCGACGGCCTGAGCACATCTGCTGTGG GATGTGCTTCTGCCGCTCCTAGAAGAGGACCTGCTCTGCTGCACATCGACCGGCA TCAGATTCAGGCCGTGGAACCTTCTGCACAGGCCCTGGAACTGCAAGGCCTGGG AGTCGATGTGTACGACCAGGACGTTCTGGAACAGGGCGTGCTGCAACAGGTGGA CAATGCCATTCACGAGGCCAGCAGAGCCTCTCAGCTGGTGGACGTGGAAAAAGA ATATCGGAGCGTGCTGGACGACCTGACCAGCTGTACAACCAGCCTGCGGCAGAT CAACAAGATCATCGAGCAGCTGTCTCCCCAGGCCGCCACCAGCAGAGACATCAA CAGAAAGCTGGACAGCGTGAAGCGCCAGAAGTACAACAAAGAGCAGCAGCTGA AGAAGATCACCGCCAAGCAGAAACATCTGCAGGCCATTCTCGGCGGAGCCGAAG TGAAGATCGAACTGGATCACGCCAGCCTGGAAGAGGATGCCGAACCTGGACCAA GCAGCCTGGGCTCTATGCTGATGCCCGTGCAAGAGACAGCCTGGGAAGAACTGA TCCGGACCGGCCAGATGACCCCTTTCGGCACACAGATCCCTCAGAAGCAAGAGA AGAAACCCCGGAAGATCATGCTGAATGAGGCCAGCGGCTTCGAGAAGTACCTGG CCGATCAGGCCAAGCTGAGCTTCGAGAGAAAGAAGCAGGGCTGCAACAAGAGA GCCGCCAGAAAAGCTCCCGCTCCTGTGACACCTCCTGCTCCAGTGCAGAACAAG AACAAGCCCAACAAGAAAGCCCGGGTGCTGAGCAAGAAAGAGGAACGCCTGAA GAAACACATCAAGAAGCTGCAGAAGCGGGCCCTGCAGTTTCAAGGCAAAGTGGG CCTGCCTAAGGCCAGAAGGCCTTGGGAAAGCGACATGAGGCCTGAGGCCGAGGG CGATTCTGAGGGCGAAGAGAGCGAGTACTTCCCCACCGAGGAAGAAGAAGAGG AAGAGGACGACGAGGTTGAGGGCGCCGAAGCTGATCTTAGCGGAGATGGCACCG ACTACGAGCTGAAGCCTCTGCCTAAAGGCGGCAAGAGGCAGAAAAAGGTGCCAG TGCAAGAAATCGACGACGACTTCTTCCCATCCTCCGGCGAAGAAGCTGAGGCCG CCTCTGTTGGAGAAGGCGGCGGAGGCGGAAGAAAAGTGGGCAGATACAGAGAT GACGGCGACGAGGACTACTACAAGCAGCGGCTGAGAAGGTGGAACAAGCTGAG ACTGCAGGACAAAGAGAAGCGGCTCAAGCTGGAAGATGACAGCGAGGAATCCG ACGCCGAGTTCGACGAGGGCTTTAAGGTGCCCGGCTTCCTGTTTAAGAAGCTGTT CAAGTACCAGCAGACCGGCGTGCGGTGGCTGTGGGAACTGCATTGTCAACAGGC AGGCGGAATCCTGGGCGACGAAATGGGACTGGGCAAGACAATCCAGATTATCGC CTTCCTGGCCGGCCTGTCCTACAGCAAGATCAGAACCCGGGGCAGCAACTACAG ATTCGAAGGACTGGGCCCCACCGTGATCGTGTGTCCTACAACAGTGATGCACCAG TGGGTCAAAGAATTTCACACCTGGTGGCCTCCATTCCGCGTGGCCATTCTGCACG AGACAGGCAGCTACACCCACAAGAAAGAAAAGCTGATCCGCGACGTGGCCCACT GCCACGGCATCCTGATCACAAGCTACAGCTACATCCGGCTGATGCAGGACGACA TCAGCAGATACGACTGGCACTACGTGATCCTGGACGAGGGCCACAAGATTCGGA ACCCCAATGCCGCTGTGACCCTGGCCTGCAAGCAGTTCAGAACCCCTCACCGGAT CATCCTGAGCGGCAGCCCCATGCAAAACAACCTGAGAGAGCTGTGGTCCCTGTT CGACTTCATCTTCCCTGGCAAGCTGGGCACCCTGCCTGTGTTCATGGAACAGTTC AGCGTGCCCATCACCATGGGCGGCTACAGCAATGCCTCTCCTGTGCAAGTGAAA ACCGCCTACAAGTGCGCCTGCGTGCTGCGGGATACCATCAATCCTTACCTGCTGC GGCGGATGAAGTCCGACGTGAAGATGAGCCTGAGCCTGCCTGACAAGAACGAAC AGGTGCTGTTCTGCCGGCTGACCGACGAGCAGCACAAGGTGTACCAGAACTTCG TGGACAGCAAAGAGGTGTACAGAATCCTGAACGGCGAGATGCAGATTTTCAGCG GACTGATCGCCCTGCGGAAAATCTGTAATCACCCCGACCTGTTCAGCGGCGGACC CAAGAATCTGAAGGGCCTGCCAGACGACGAACTCGAAGAGGATCAGTTTGGCTA CTGGAAGCGGAGCGGCAAGATGATCGTGGTGGAAAGCCTGCTGAAGATTTGGCA CAAGCAGGGCCAGAGAGTGCTGCTGTTCAGCCAGAGCAGACAGATGCTGGATAT CCTGGAAGTGTTCCTGCGGGCCCAGAAGTATACCTACCTGAAGATGGACGGCAC CACCACAATCGCCAGCAGACAGCCTCTGATCACCCGGTACAATGAGGATACCTC CATCTTCGTCTTTCTGCTGACCACCAGAGTCGGCGGACTGGGCGTTAACCTGACT GGCGCCAATCGGGTCGTGATCTATGACCCCGACTGGAACCCCAGCACCGACACA CAGGCTAGAGAGAGAGCTTGGAGAATCGGCCAGAAAAAGCAAGTGACCGTGTA CCGGCTGCTGACCGCCGGAACCATCGAAGAGAAAATCTACCACCGGCAGATTTT TAAGCAGTTCCTGACCAACAGGGTGCTGAAGGACCCCAAGCAGAGAAGATTCTT CAAGAGCAACGACCTGTATGAGCTGTTCACCCTGACAAGCCCCGATGCCAGCCA GTCTACAGAGACAAGCGCCATCTTTGCCGGCACCGGCTCTGATGTGCAGACCCCT AAGTGTCACCTGAAGCGGAGAATCCAGCCTGCCTTTGGCGCCGATCACGACGTG CCCAAGCGGAAGAAGTTCCCCGCCAGCAACATCAGCGTGAACGATGCCACCAGC TCCGAGGAAAAGAGCGAAGCCAAAGGCGCTGAAGTGAACGCCGTGACCAGCAA CAGAAGCGACCCTCTGAAGGACGACCCTCACATGAGCAGCAACGTGACCTCCAA CGACCGGCTGGGAGAAGAGACAAATGCCGTGTCTGGCCCCGAGGAACTGAGCGT GATATCTGGCAATGGCGAGTGCAGCAACAGCTCTGGCACCGGCAAGACCTCTAT GCCATCTGGCGACGAGAGCATCGACGAGAAGCTGGGCCTGTCTTACAAGAGAGA GCGGCCTTCTCAGGCCCAGACCGAGGCCTTTTGGGAGAACAAGCAGATGGAAAA CAACTTCTACAAGCACAAGTCCAAGACCAAGCACCACAGCGTGGCCGAAGAGGA AACCCTGGAAAAGCACCTGAGGCCTAAGCAGAAGCCCAAGAACTCCAAGCACTG CCGGGATGCCAAGTTCGAAGGCACAAGAATCCCACACCTCGTGAAGAAGCGCAG ATACCAGAAGCAGGACAGCGAGAACAAGTCCGAGGCCAAAGAACAGTCCAACG ACGACTACGTGCTGGAAAAACTGTTCAAGAAAAGCGTGGGCGTGCACAGCGTGA TGAAGCACGACGCCATTATGGACGGCGCTAGCCCCGATTATGTGCTGGTGGAAG CCGAAGCCAACAGAGTGGCACAGGATGCCCTGAAGGCCCTGAGACTGTCCAGAC AGAGATGTCTGGGAGCTGTGTCCGGCGTGCCAACATGGACAGGACACAGAGGAA TTTCTGGCGCCCCTGCCGGCAAGAAGTCTCGCTTTGGCAAGAAGAGAAACAGCA ACTTCTCCGTGCAGCACCCCTCCAGCACAAGCCCTACAGAGAAGTGCCAGGACG GCATCATGAAGAAAGAAGGCAAGGACAACGTCCCCGAGCACTTTAGCGGCAGAG CCGAGGATGCTGATAGCTCTAGTGGACCTCTGGCCAGCAGTTCCCTGCTGGCTAA GATGAGAGCCCGGAACCACCTGATCCTGCCAGAGAGACTGGAAAGCGAGTCCGG CCATCTGCAAGAAGCTAGCGCACTGCTGCCTACCACCGAGCACGATGATCTGCTG GTCGAGATGCGGAACTTCATTGCCTTTCAAGCCCACACAGACGGCCAGGCCTCTA CCAGAGAAATCCTGCAAGAGTTTGAGAGCAAGCTGTCCGCCTCTCAGAGCTGCG TGTTCAGAGAGCTGCTGAGAAACCTGTGCACCTTCCACAGAACCAGCGGAGGCG AAGGCATCTGGAAGCTGAAACCCGAGTACTGCTGATGATAGAATAAAGGATCCA GATCTGCGGCCGCCTGGGCCTCATGGGCCTTCCTTTCACTGCCCGCTTTCCAGTCGGGA AACCTGTCGTGCCA
[0300] Restriction sites - Times new Roman font, bold
[0301] GTTTAAACTACGTACCTAGGCTTAAG (SEQ ID NO: 11) GAATTCTTAATTAAAAGCTTCTCGAG (SEQ ID NO: 12) GGATCCAGATCTGCGGCCGC (SEQ ID NO: 13) Mini-Promoter - enlarged font , bold CTCGCCCTGCTCAAAGGCGTGAGCCCGTCCCTCCAAA CCACGTCTTGCCCCTTCTTGACCGACAATATCTCAAT TAGTCAGCAACACCATTGGCCGGATATATCACCGACT GACCTCTGCTGTTCC (SEQ ID NO : 1)
[0302] T7 promoter- enlarged font, i talic TAATACGACTCACTATAGG (SEQ ID NO: 14)
[0303] Codon optimized CSBZERCC6 — Times new Roman font, regular font ATGCCGAACGAGGGAATCCCTCACAGCAGCCAGACACAAGAGCAGGACTGCCTG CAGAGCCAGCCTGTGTCCAACAACGAGGAAATGGCCATCAAGCAAGAGTCTGGC GGCGACGGCGAGGTGGAAGAGTACCTGTCTTTTAGAAGCGTTGGCGACGGCCTG AGCACATCTGCTGTGGGATGTGCTTCTGCCGCTCCTAGAAGAGGACCTGCTCTGC
[0304] TGCACATCGACCGGCATCAGATTCAGGCCGTGGAACCTTCTGCACAGGCCCTGG AACTGCAAGGCCTGGGAGTCGATGTGTACGACCAGGACGTTCTGGAACAGGGCG TGCTGCAACAGGTGGACAATGCCATTCACGAGGCCAGCAGAGCCTCTCAGCTGG TGGACGTGGAAAAAGAATATCGGAGCGTGCTGGACGACCTGACCAGCTGTACAA CCAGCCTGCGGCAGATCAACAAGATCATCGAGCAGCTGTCTCCCCAGGCCGCCA
[0305] CCAGCAGAGACATCAACAGAAAGCTGGACAGCGTGAAGCGCCAGAAGTACAAC AAAGAGCAGCAGCTGAAGAAGATCACCGCCAAGCAGAAACATCTGCAGGCCATT CTCGGCGGAGCCGAAGTGAAGATCGAACTGGATCACGCCAGCCTGGAAGAGGAT GCCGAACCTGGACCAAGCAGCCTGGGCTCTATGCTGATGCCCGTGCAAGAGACA GCCTGGGAAGAACTGATCCGGACCGGCCAGATGACCCCTTTCGGCACACAGATC
[0306] CCTCAGAAGCAAGAGAAGAAACCCCGGAAGATCATGCTGAATGAGGCCAGCGG CTTCGAGAAGTACCTGGCCGATCAGGCCAAGCTGAGCTTCGAGAGAAAGAAGCA GGGCTGCAACAAGAGAGCCGCCAGAAAAGCTCCCGCTCCTGTGACACCTCCTGC TCCAGTGCAGAACAAGAACAAGCCCAACAAGAAAGCCCGGGTGCTGAGCAAGA AAGAGGAACGCCTGAAGAAACACATCAAGAAGCTGCAGAAGCGGGCCCTGCAG
[0307] TTTCAAGGCAAAGTGGGCCTGCCTAAGGCCAGAAGGCCTTGGGAAAGCGACATG AGGCCTGAGGCCGAGGGCGATTCTGAGGGCGAAGAGAGCGAGTACTTCCCCACC GAGGAAGAAGAAGAGGAAGAGGACGACGAGGTTGAGGGCGCCGAAGCTGATCT TAGCGGAGATGGCACCGACTACGAGCTGAAGCCTCTGCCTAAAGGCGGCAAGAG GCAGAAAAAGGTGCCAGTGCAAGAAATCGACGACGACTTCTTCCCATCCTCCGG
[0308] CGAAGAAGCTGAGGCCGCCTCTGTTGGAGAAGGCGGCGGAGGCGGAAGAAAAG TGGGCAGATACAGAGATGACGGCGACGAGGACTACTACAAGCAGCGGCTGAGA AGGTGGAACAAGCTGAGACTGCAGGACAAAGAGAAGCGGCTCAAGCTGGAAGA TGACAGCGAGGAATCCGACGCCGAGTTCGACGAGGGCTTTAAGGTGCCCGGCTT CCTGTTTAAGAAGCTGTTCAAGTACCAGCAGACCGGCGTGCGGTGGCTGTGGGA
[0309] ACTGCATTGTCAACAGGCAGGCGGAATCCTGGGCGACGAAATGGGACTGGGCAA GACAATCCAGATTATCGCCTTCCTGGCCGGCCTGTCCTACAGCAAGATCAGAACC CGGGGCAGCAACTACAGATTCGAAGGACTGGGCCCCACCGTGATCGTGTGTCCT ACAACAGTGATGCACCAGTGGGTCAAAGAATTTCACACCTGGTGGCCTCCATTCC GCGTGGCCATTCTGCACGAGACAGGCAGCTACACCCACAAGAAAGAAAAGCTGA TCCGCGACGTGGCCCACTGCCACGGCATCCTGATCACAAGCTACAGCTACATCCG GCTGATGCAGGACGACATCAGCAGATACGACTGGCACTACGTGATCCTGGACGA GGGCCACAAGATTCGGAACCCCAATGCCGCTGTGACCCTGGCCTGCAAGCAGTT CAGAACCCCTCACCGGATCATCCTGAGCGGCAGCCCCATGCAAAACAACCTGAG AGAGCTGTGGTCCCTGTTCGACTTCATCTTCCCTGGCAAGCTGGGCACCCTGCCT GTGTTCATGGAACAGTTCAGCGTGCCCATCACCATGGGCGGCTACAGCAATGCCT CTCCTGTGCAAGTGAAAACCGCCTACAAGTGCGCCTGCGTGCTGCGGGATACCAT CAATCCTTACCTGCTGCGGCGGATGAAGTCCGACGTGAAGATGAGCCTGAGCCT GCCTGACAAGAACGAACAGGTGCTGTTCTGCCGGCTGACCGACGAGCAGCACAA GGTGTACCAGAACTTCGTGGACAGCAAAGAGGTGTACAGAATCCTGAACGGCGA GATGCAGATTTTCAGCGGACTGATCGCCCTGCGGAAAATCTGTAATCACCCCGAC CTGTTCAGCGGCGGACCCAAGAATCTGAAGGGCCTGCCAGACGACGAACTCGAA GAGGATCAGTTTGGCTACTGGAAGCGGAGCGGCAAGATGATCGTGGTGGAAAGC CTGCTGAAGATTTGGCACAAGCAGGGCCAGAGAGTGCTGCTGTTCAGCCAGAGC AGACAGATGCTGGATATCCTGGAAGTGTTCCTGCGGGCCCAGAAGTATACCTACC TGAAGATGGACGGCACCACCACAATCGCCAGCAGACAGCCTCTGATCACCCGGT ACAATGAGGATACCTCCATCTTCGTCTTTCTGCTGACCACCAGAGTCGGCGGACT GGGCGTTAACCTGACTGGCGCCAATCGGGTCGTGATCTATGACCCCGACTGGAAC CCCAGCACCGACACACAGGCTAGAGAGAGAGCTTGGAGAATCGGCCAGAAAAA GCAAGTGACCGTGTACCGGCTGCTGACCGCCGGAACCATCGAAGAGAAAATCTA CCACCGGCAGATTTTTAAGCAGTTCCTGACCAACAGGGTGCTGAAGGACCCCAA GCAGAGAAGATTCTTCAAGAGCAACGACCTGTATGAGCTGTTCACCCTGACAAG CCCCGATGCCAGCCAGTCTACAGAGACAAGCGCCATCTTTGCCGGCACCGGCTCT GATGTGCAGACCCCTAAGTGTCACCTGAAGCGGAGAATCCAGCCTGCCTTTGGCG CCGATCACGACGTGCCCAAGCGGAAGAAGTTCCCCGCCAGCAACATCAGCGTGA ACGATGCCACCAGCTCCGAGGAAAAGAGCGAAGCCAAAGGCGCTGAAGTGAAC GCCGTGACCAGCAACAGAAGCGACCCTCTGAAGGACGACCCTCACATGAGCAGC AACGTGACCTCCAACGACCGGCTGGGAGAAGAGACAAATGCCGTGTCTGGCCCC GAGGAACTGAGCGTGATATCTGGCAATGGCGAGTGCAGCAACAGCTCTGGCACC GGCAAGACCTCTATGCCATCTGGCGACGAGAGCATCGACGAGAAGCTGGGCCTG TCTTACAAGAGAGAGCGGCCTTCTCAGGCCCAGACCGAGGCCTTTTGGGAGAAC AAGCAGATGGAAAACAACTTCTACAAGCACAAGTCCAAGACCAAGCACCACAGC GTGGCCGAAGAGGAAACCCTGGAAAAGCACCTGAGGCCTAAGCAGAAGCCCAA GAACTCCAAGCACTGCCGGGATGCCAAGTTCGAAGGCACAAGAATCCCACACCT CGTGAAGAAGCGCAGATACCAGAAGCAGGACAGCGAGAACAAGTCCGAGGCCA AAGAACAGTCCAACGACGACTACGTGCTGGAAAAACTGTTCAAGAAAAGCGTGG GCGTGCACAGCGTGATGAAGCACGACGCCATTATGGACGGCGCTAGCCCCGATT ATGTGCTGGTGGAAGCCGAAGCCAACAGAGTGGCACAGGATGCCCTGAAGGCCC TGAGACTGTCCAGACAGAGATGTCTGGGAGCTGTGTCCGGCGTGCCAACATGGA CAGGACACAGAGGAATTTCTGGCGCCCCTGCCGGCAAGAAGTCTCGCTTTGGCA AGAAGAGAAACAGCAACTTCTCCGTGCAGCACCCCTCCAGCACAAGCCCTACAG AGAAGTGCCAGGACGGCATCATGAAGAAAGAAGGCAAGGACAACGTCCCCGAG CACTTTAGCGGCAGAGCCGAGGATGCTGATAGCTCTAGTGGACCTCTGGCCAGC AGTTCCCTGCTGGCTAAGATGAGAGCCCGGAACCACCTGATCCTGCCAGAGAGA CTGGAAAGCGAGTCCGGCCATCTGCAAGAAGCTAGCGCACTGCTGCCTACCACC GAGCACGATGATCTGCTGGTCGAGATGCGGAACTTCATTGCCTTTCAAGCCCACA CAGACGGCCAGGCCTCTACCAGAGAAATCCTGCAAGAGTTTGAGAGCAAGCTGT CCGCCTCTCAGAGCTGCGTGTTCAGAGAGCTGCTGAGAAACCTGTGCACCTTCCA CAGAACCAGCGGAGGCGAAGGCATCTGGAAGCTGAAACCCGAGTACTGCTGATG AT AG (SEQ ID N0:15) Poly A tail — Courier new font, italics
[0310] AATAAA (SEQ ID NO: 16)
[0311] Non-coding sequences - Courier new, regular font GAGCGCGACG (SEQ ID NO : 17 ) GCGAATTGAAGGAAGGCCGTCAAGGCCGCAT (SEQ ID NO : 18) CTGCCACC (SEQ ID NO : 19) CTGGGCCTCATGGGCCTTCCTTTCACTGCCCGCTTTCCAGTCGGGAAACCTGTCGTGCCA (SEQ ID NO : 20)
[0312] The Variant 1 plasmid was transfected into cells in tissue culture, total RNA was isolated and reverse transcriptase PCR demonstrated expression levels of ERCC6 transcripts in the transfected cells that corresponded to the amount of plasmid transfected in each well (Figure 3). HEK293 cells were transfected with the CP040 plasmid and total RNA was isolated and reverse transcription quantitative PCR was performed on un-transfected (controls) and compared to cells that were transfected with either 250 ng or 500 ng of plasmid using specific probes designed to recognize the codon optimized ERCC6 / CSB gene (coERCC6). Co CSB probel is specific against the 5’ end of the co-CSB, probe2 is specific against the middle portion and probe3 detects the 3’ end of co-CSB mRNA sequence. All the three probes for co-CSB were able to detect more co-CSB transcripts in the CP040 transfected cells as compared to non-transfected cells. This preliminary experiment suggests that our CP040 mini-promoter is able to drive transgene expression in HEK293 cells and supports subsequent testing through in vivo studies.
[0313] Further in vitro characterizations of our mini-promoter sequence are performed using a dual luciferase assay. Various fragments of this promoter sequence are cloned into a promoter-free vector containing a firefly luciferase reporter gene (Figure 2, variants 2-4). This plasmid is co-transfected with a plasmid expressing the renilla luciferase transgene as a control (Figure 4).
[0314] The top promoter from the dual luciferase experiment is packaged into AAV capsids for in vivo characterization studies to determine tissue and cell specificity as well as duration of expression. Wild type (WT) mice are injected with this AAV vector and 8 weeks later, the mice are euthanized, and transgene expression determined (using RT-qPCR and immunohistochemistry) in different tissue and cell types following full necropsies. The results are compared to an identical study we are currently performing to evaluate other (larger) existing promoters.
[0315] In order to determine how well these promoters work in human disease relevant cell types, the expression levels they can provide in induced pluripotent stem cell (iPSC) derived neuronal cells and cerebral brain organoids using the same luciferase assay are evaluated.
[0316] A second approach for use each component individually into a plasmid (e.g., with a luciferase reporter gene). These further constructs are tested individually using the aforementioned assays (Figure 5).
[0317] The sequence of the mini-promoter sequence of SEQ ID NO: 1 was evaluated using AliBaba tools with TRANSFEC database to determine potential binding sites (Matys, V., Kel-Margoulis, O.V., Fricke, E., Liebich, I., Land, S., Barre-Dirrie, A., Reuter, I., Chekmenev, D., Krull, M., Homischer, K., et al. (2006). TRANSFAC and its module TRANSCompel: transcriptional gene regulation in eukaryotes. Nucleic Acids Res 34, D108- 110.) (Figure 6).
[0318] The sequence of the mini-promoter sequence of SEQ ID NO: 1 was evaluated using Firefly tools ( Reese, M.G. (2001). Application of a time-delay neural network to promoter annotation in the Drosophila melanogaster genome. Comput Chem 26, 51-56) to determine promoter predications for seqO (Figure 7).
[0319] A vector map is provided in Figure 8.
[0320] Example 4
[0321] A plasmid has been engineered and synthesized that includes the mini-promoter along with a codon optimized version of ERCC6 (the gene that encodes CSB) (SEQ ID NO:21). This plasmid was transfected into cells in tissue culture and reverse transcriptase PCR demonstrated expression levels of ERCC6 transcripts in the transfected cells that corresponded to the amount of plasmid transfected in each well.
[0322] Example 5
[0323] The small size of the CP040 enables the packaging of some larger sequences needed for gene therapy or genome engineering applications when the material needed is too large to fit within the adeno-associated virus (AAV) packaging capacity with conventional promoters. The fully defined sequence of the CP040 promoter is derived from human sequences, there are no viral elements present. A large number of transcription factor binding sites are included to further enhance promoter function (Figure 9).
[0324] It has been demonstrated both in vitro as well as in vivo CP040 function. The in vivo study was an evaluation of CP040 function as compared to existing, well-characterized promoters and showed CP040 performs well following AAV administration using two different delivery routes.
[0325] Defining the Functional Regions of CP040:
[0326] The goal was to determine if a fragment or variant of the present CP040 minipromoter sequence could be used further to reduce its size yet also retain promoter activity levels. To do this, a dual luciferase assay was performed in two different cell types (human embryonic kidney [HEK293s] and human induced pluripotent stem cell derived neural progenitor cells [NPCs]) to determine the relative efficiency or strength of these promoter fragments and variants. First, the fragments and variants were cloned into a promoter-less vector pGL4.20, upstream of a firefly luciferase (Flue) reporter (Figure 10A). The renilla luciferase (Rluc) gene in a separate plasmid pRL-CMV (Figure 10B) was used to normalize transfection efficiency by co-transfection. The spacerl region was cloned in pGL4.20 as fragmentl; CSA UTR region as fragment2; spacer2 region as fragments; CSB UTR region as fragment4; spacers region as fragments; full length CP040 as variantl; variant2 as CP040 missing spacerl; variants as CP040 missing spacerS; variant4 as CP040 missing spacer2 (Figure 2A-2C). HEK293 or NPCs were seeded in 12-well plate and co-transfected with a Flue-tagged pGL4.20 vector (fragment and variant plasmids) and a Rluc control vector (containing CMV promoter) in 50: 1 ratio. Transfected cells were lysed, 24 h post-transfection for luminescence analysis. Luminescence was measured in a 96-well white bottom plate. The background luminescence for both Flue and Rluc activities was initially measured using the following controls: a no transfection, single transfections of each fragment and variant pGL4.20 vector, and transfection of the empty pGL4.20 vector. Next, luminescence was detected from the CP040 fragments, variants and positive control chicken P-actin construct (CBA) for Flue and Rluc activities (Figure 11A-11B). The fold change was calculated with respect to the empty pGL4.20 vector values. In both cell types, fold change in variantl (full- length CP040) activity was significantly higher as compared to the empty vector. These levels were comparable to the well -characterized CBA promoter. No activity was observed in the fragments or other variants except for variant2, which was still significantly lower than variantl (Figure 11A-11B). These results suggested that the full length CP040 mini-promoter (variantl) was most efficient when all the fragments and spacer regions are included.
[0327] In vivo CP040 Evaluation:
[0328] Once it was established that full length CP040 is efficient in vitro, the promoter activity was evaluated in vivo in mice. The promoter was cloned upstream of the green fluorescent protein (GFP) reporter gene in an adeno-associated virus inverted terminal repeat (ITR) containing vector and packaged into the AAV-DJ capsid. A dose of IxlO13vector genome (vg) copies per kg body weight of AAVDJ-CP040-eGFP was administered to wildtype mouse neonates via systemic, intravenous (IV) temporal vein injections. Six weeks after the injection, the mice were euthanized and tissues were collected to analyze vgs, eGFP transcript levels, and eGFP protein expression. The highest transcript levels were observed in cortex and cerebellum followed by spinal cord, liver and heart which correlated well with the vg levels found in these tissues (Figure 12A and 12B). There were very few transcripts observed in spleen, quadriceps (Quad) muscles and kidneys. These data suggest that CP040 can drive transgene expression in vivo with good efficiency in central nervous system (CNS) tissues.
[0329] In vivo CP040 Evaluation as Compared to Existing Promoters Using a Marker Gene Delivered via AAV:
[0330] Once the in vivo promoter activity for CP040 was confirmed, the efficiency was further evaluated by comparing CP040 with strong and ubiquitous promoters like chicken P- actin (CBA) and elongation factor la (EFla), as well as a neuronal specific promoter human synapsin (hSYN) promoter (Figure 13). These four promoters were all packaged in AAV-DJ and a dose of IxlO12vg per kg body weight per mouse was injected into six-week-old adult wild-type FVB / NJ mice via intracerebroventricular (ICV) and intrathecal (IT) administration routes. At eight weeks post-administration, the mice were euthanized for tissue vg and promoter activity level assessments via transgene expression analyses. As our goal was to primarily target central nervous system (CNS) tissues, specific brain regions were dissected and processed. As expected, since all four promoters were packaged in the same AAV serotype, the vgs for each promoter were similar across tissues within ICV and IT administration route cohorts (Figures 14A-14G). To establish the efficiency of the CP040 mini-promoter, eGFP gene expression in the targeted CNS organs as well as those organs we hope to de-target was determined and compared across all cohorts. Total RNA was isolated from each tissue and eGFP transcription levels were assessed by quantitative PCR (Figures 15A-15G). Following IT AAV administration, fold change in eGFP transcripts using the CP040 promoter was comparable to all three established promoters in the hypothalamus, hippocampus, and spinal cord. In the cortex and cerebellum CP040 mediated eGFP expression levels that were significantly elevated as compared to the CBA, EFla, and hSYN promoters. However, when using the ICV route, eGFP transcript levels from the CP040 minipromoter were approximately half that obtained from either of the ubiquitous CBA and EFla promoters and equivalent to the hSYN promoter in the hippocampus and hypothalamus. In the cortex, CP040 eGFP expression levels were equivalent to both the EFla and hSYN promoters and significantly lower as compared to CBA. While in cerebellum, it was equal to CBA and significantly lower than EFla and hSYN promoters. The transcription profiles of eGFP were also evaluated in other de-targeted tissues such as liver and kidneys following both delivery routes where CP040 mini-promoter mediated expression was found to be similar to hSYN and significantly less than both the CBA and EFla promoters. These data suggest that with regards to transcription levels, CP040 is comparable to these well- established promoters despite its small size.
[0331] METHODS
[0332] Cell Culture
[0333] All cell lines were maintained in a 5% CO2 maintaining temperature regulated incubator (37°C). Human embryonic kidney (HEK293) cell line was maintained in 10% FBS (fetal bovine serum) containing high glucose DMEM in a 10 cm dish. Neural progenitor cells (NPCs) were maintained in NPC expansion media [47 ml DMEM / F12, 1ml B27 supplement (Gibco), 0.5ml N2 supplement (Gibco), 0.5 ml NEAA (Gibco), 0.5 ml Glutamax (Gibco), 0.5 ml antibiotic antimycotic (Invitrogen), 30 ng / ml EFG (Peprotech) and 30 ng / ml FGF (Peprotech)] on a basement membrane (Matrigel, #354277, Coming) coated 10 cm dish. NPCs were always seeded on Matrigel coated 6-well plates.
[0334] Promoter constructs used in the Dual Luciferase Assay
[0335] Dual Luciferase Assay
[0336] Dual luciferase assays were performed with the dual luciferase reporter assay system (Promega Corp., Madison, WI). All reagents were prepared as described by the manufacturer. The 5* passive lysis buffer was supplied by the manufacturer and used for cell lysis. From a confluent master plate Neural Precursor cell (NPCs) or HEK293 cells (2xl06) were seeded in 6-well plates. Next day, the media was changed to serum free media 3 hours prior to transfection. Using Lipofectamine 3000 (L3000001, Invitrogen) the cells were transfected with 500 ng of either pGL4.20 empty plasmid or pGL4.20 CP040 variants or fragments or pGL4.20 CBA plasmid with pRL-CMV control plasmid (pGL4.20: pRL should be 50: 1). The media was changed back to normal growth media after 6 hours of transfection. After 24 hours, the cells were harvested in 500 pL per well of lx passive lysis buffer allowing lysis for 20 min on orbital shaker. The lysate was collected in 2 mL tube and centrifuged briefly. In a white bottom 96-well plate, 20 pL supernatant was added with 50 luciferase assay substrate (LAS) and luminescence was measured with a 10s equilibration time and 10s integration time, followed by addition of 50 pL of Renilla luciferase reagent and firefly quenching (lx Stop & Gio), 10s equilibration time, and measurement of luminescence with a 10s integration time. Subtract the respective values from non-transfected values. For respective sample, the ratio of firefly (first luminescence reading) to Renilla (second luminescence reading) luciferase activity (Fluc / Rluc) was calculated and presented as the fold change with respect to empty vector (pGL4.20) transfection.
[0337] Animals
[0338] All husbandry and animal related procedures were approved by the Institutional Animal Care and Use Committee (IACUC) at the University of Minnesota (UMN). WT FVB / NJ (stock# 001800) mice at 5 weeks of age were ordered from Jackson Laboratory and acclimatized to the vivarium for one week prior to initiation of the study. The mice were maintained in the UMN Research Animal Facility at 25°C and 50% humidity with a 14-hour light and 10-hour dark cycle, fed normal chow, and water ad libitum.
[0339] In vivo AAV delivery
[0340] Neonatal Temporal Vein Injections - Intravenous (IV)
[0341] Neonatal injections were administered through the superficial temporal vein (at 1-2 days of age). The mice were anesthetized by keeping them on ice for less than five minutes to induce hypothermia. A 29.5-gauge short needle syringe was used to deliver vector in a total volume of 20 pL directly into the left temporal vein.
[0342] Intrathecal Injections (IT)
[0343] At 6 weeks of age five male and five female WT mice were intrathecally injected with either lx PBS or for each promoter group AAV-DJ-CBA-eGFP, AAV-DJ-EFla-eGFP, AAV-DJ-hSYN-eGFP, AAV-DJ-CP040-eGFP. Intrathecal injections were performed as previously described in Chauhan et al., 2024. Briefly, mice were anesthetized and a 26-gauge Hamilton syringe was used to deliver a volume of 10 pl of AAV (~lxl012vg per kg body weight per mouse) directly by lumbar puncture.
[0344] Intracerebroventricular Injections (ICV)
[0345] Bilateral ICV injections, were performed as described in Chauhan et al., 2024.
[0346] Briefly, mice were anesthetized, fixed on a stereotaxic apparatus and using 26-gauge Hamilton syringe lateral ventricles were located: -0.5 mm in the anteroposterior axis, ±1 mm in the mediolateral axis, and -2.3 mm in the dorsoventral axis (coordinated according to Allen Brain Atlas for a ~2 months old mouse). Once the needle was placed inside each ventricle, 5 pL of vehicle (lx PBS), AAVDJ-CBA-eGFP, AAVDJ-EFla-eGFP, AAVDJ-hSYN-eGFP, or AAV-DJ-CP040-eGFP was injected slowly into each ventricle with a flow rate of 1 pL per minute (Taylor, Khand et al. 2021). Similarly, the other ventricle was injected with the remaining 5 pL for a total of 10 pL delivered. Groups of 5 male and 5 female mice were injected at 6 weeks of age for each promoter group and the final AAV dose injected was ~lx 1012vg / kg body weight per mouse. Animals were monitored carefully post-surgery according to IACUC guidelines at the University of Minnesota.
[0347] Vector genome (vg) Quantification
[0348] AAV vg copy numbers from tissues collected 9 weeks post AAV administration were determined by qPCR after extraction of total gDNA using a DNeasy Blood and Tissue Kit (Qiagen). Tissues were mechanically homogenized using ceramic beads in a bead basher (Benchmark BeadBlaster24) with provided buffer and Proteinase K. gDNA isolation was performed according to manufacturer instructions. Vg content in each tissue, normalized to 100 ng of gDNA, was determined using a qPCR method and a Taqman probe specific for the eGFP reporter sequence (Assay ID: Mr00660654, Thermo). A standard curve was created using eGFP plasmid DNA dilutions with known copy numbers, to determine absolute copy numbers in the sample gDNA. Statistical analyses were performed using GraphPad Prism (GraphPad Software) and significance was determined by two tailed unpaired t-tests where a p < 0.05 was considered to be significant.
[0349] Transcript analysis
[0350] Total RNA was isolated from mouse tissues 9 weeks post AAV administration using the Quick-RNA miniPrep kit (R1055, Zymo Research). Tissues were mechanically homogenized using ceramic beads in a bead basher (Benchmark BeadBlaster24) with provided RNA lysis buffer. The RNA kit includes a gDNA removal step and a subsequent DNase I clean up step, the remainder of the isolation was performed according to manufacturer’s instructions. 300 ng of total gDNA-free-RNA was used for cDNA synthesis using the High-Capacity RNA-to-cDNA Kit (#4388950, Applied Biosystems). This cDNA product was used to set up qPCR with GFP probes (Assay ID: Mr0066065, Thermo) and Taqman Fast Advanced Master Mix (Thermo) in the QuantStudio 3 (Applied Biosystems) thermal cycler. Ribosomal RNA 18s amplification (Assay ID: Mm03928990_gl, Thermo) on the same cDNA was used as an internal control to calculate fold change in eGFP transcripts using the AACT method.
[0351] All publications, patents and patent applications are incorporated herein by reference. While in the foregoing specification this invention has been described in relation to certain preferred embodiments thereof, and many details have been set forth for purposes of illustration, it will be apparent to those skilled in the art that the invention is susceptible to additional embodiments and that certain of the details described herein may be varied considerably without departing from the basic principles of the invention.
[0352] The use of the terms “a” and “an” and “the” and similar referents in the context of describing the invention are to be construed to cover both the singular and the plural, unless otherwise indicated herein or clearly contradicted by context. The terms “comprising,” “having,” “including,” and “containing” are to be construed as open-ended terms (z.e., meaning “including, but not limited to”) unless otherwise noted. Recitation of ranges of values herein are merely intended to serve as a shorthand method of referring individually to each separate value falling within the range, unless otherwise indicated herein, and each separate value is incorporated into the specification as if it were individually recited herein. All methods described herein can be performed in any suitable order unless otherwise indicated herein or otherwise clearly contradicted by context. The use of any and all examples, or exemplary language (e.g., “such as”) provided herein, is intended merely to better illuminate the invention and does not pose a limitation on the scope of the invention unless otherwise claimed. No language in the specification should be construed as indicating any non-claimed element as essential to the practice of the invention.
[0353] Embodiments of this invention are described herein, including the best mode known to the inventors for carrying out the invention. Variations of those embodiments may become apparent to those of ordinary skill in the art upon reading the foregoing description. The inventors expect skilled artisans to employ such variations as appropriate, and the inventors intend for the invention to be practiced otherwise than as specifically described herein. Accordingly, this invention includes all modifications and equivalents of the subject matter recited in the claims appended hereto as permitted by applicable law. Moreover, any combination of the above-described elements in all possible variations thereof is encompassed by the invention unless otherwise indicated herein or otherwise clearly contradicted by context.
Claims
WHAT IS CLAIMED IS:
1. A promoter sequence consisting of a nucleic acid of between 12 and 300 nucleotides in length, wherein the nucleic acid comprises a polynucleotide having at least 90% identity to SEQ ID NO: 1, SEQ ID NO:2, SEQ ID NO:3, SEQ ID NO:4, SEQ ID NO:5, or SEQ ID NO:6.
2. The promoter sequence of claim 1, wherein the polynucleotide sequence has at least 90% identity to SEQ ID NO:2.
3. The promoter sequence of claim 1, wherein the polynucleotide sequence has at least 90% identity to SEQ ID NO:3.
4. The promoter sequence of claim 1, wherein the polynucleotide sequence has at least 90% identity to SEQ ID NO:4.
5. The promoter sequence of claim 1, wherein the polynucleotide sequence has at least 90% identity to SEQ ID NO: 5.
6. The promoter sequence of claim 1, wherein the polynucleotide sequence has at least 90% identity to SEQ ID NO:6.
7. The promoter sequence of claim 3, wherein the polynucleotide sequence has at least 90% identity to SEQ ID NO: 3 operably linked to SEQ ID NO: 5.
8. The promoter sequence of claim 7, wherein the polynucleotide sequence has at least 90% identity to SEQ ID NO:3 operably linked to SEQ ID NO:4 operably linked to CSB UTR SEQ ID NO:5.
9. The promoter sequence of claim 8, wherein the polynucleotide sequence has at least 90% identity to SEQ ID NO:2 operably linked to CSA UTR SEQ ID NO:3 operably linked to SEQ ID NO:4 operably linked to CSB UTR SEQ ID NO:5.
10. The promoter sequence of claim 8, wherein the polynucleotide sequence has at least 90% identity to CSA UTR SEQ ID NO:3 operably linked to SEQ ID NO:4 operably linked to CSB UTR SEQ ID NO:5 operably linked to SEQ ID NO:6.
11. The promoter sequence of claim 1, wherein the polynucleotide sequence has at least 90% identity to SEQ ID NO: 1.
12. The expression cassette of claim 1, wherein the polynucleotide has at least 100% identity to SEQ ID NO: 2.
13. The expression cassette of claim 1, wherein the polynucleotide has 100% identity to SEQ ID NO: 3.
14. The expression cassette of claim 1, wherein the polynucleotide has 100% identity to SEQ ID NO: 4.
15. The expression cassette of claim 1, wherein the polynucleotide has 100% identity to SEQ ID NO: 5.
16. The expression cassette of claim 1, wherein the polynucleotide has 100% identity to SEQ ID NO: 6.
17. The promoter sequence of claim 13, wherein the polynucleotide sequence has 100% identity to SEQ ID NO:3 operably linked to SEQ ID NO:4 operably linked to CSB UTR SEQ ID NO:5.
18. The promoter sequence of claim 17, wherein the polynucleotide sequence has 100% identity to SEQ ID NO:2 operably linked to CSA UTR SEQ ID NO:3 operably linked to SEQ ID NO:4 operably linked to CSB UTR SEQ ID NO:5.
19. The promoter sequence of claim 17, wherein the polynucleotide sequence has 100% identity to CSA UTR SEQ ID NO:3 operably linked to SEQ ID NO:4 operably linked to CSB UTR SEQ ID NO:5 operably linked to SEQ ID NO:6.
20. The promoter sequence of claim 1, wherein the polynucleotide sequence has 100% identity to SEQ ID NO: 1.
21. The promoter of any one of claims 1-20, wherein the nucleic acid is between 60 and 120 nucleotides in length.
22. An expression cassette comprising the promoter of any one of claims 1-21, wherein the promoter is operably linked to a preselected DNA segment encoding a protein or heterologous RNA transcript.
23. The expression cassette of claim 22, wherein the preselected DNA segment comprises a selectable marker gene or a reporter gene.
24. The expression cassette of claim 22, wherein the preselected DNA segment encodes a therapeutic composition.
25. The expression cassette of claim 24, wherein therapeutic composition is an RNAi molecule.
26. A vector comprising the expression cassette of any one of claims 21 to 25.
27. The vector of claim 26, wherein the vector is an adeno-associated virus (AAV) vector.
28. An isolated transformed cell comprising the expression cassette of any one of claims 22 to 35, or the vector of claim 26 or 27.
29. The isolated transformed cell of claim 28, wherein the host cell is a eukaryotic cell.
30. The isolated transformed cell of claim 29, wherein the eukaryotic cell is an animal cell.
31. The isolated transformed cell of claim 30, wherein the animal cell is a mammalian cell.
32. The isolated transformed cell of claim 31, wherein the mammalian cell is a human cell.
33. A method for producing transformed cells comprising the steps of (i) introducing into cells a recombinant DNA which comprises a promoter of any one of claims 1 to 21 operably linked to a DNA segment or the expression cassette of any one of claims 22 to 25 so as to yield transformed cells, and (ii) identifying or selecting a transformed cell line.
34. The method of claim 33, wherein the recombinant DNA is expressed so as to impart a phenotypic characteristic to the transformed cells.
35. The method of claim 33, wherein the transformed cells exhibit significantly increased expression of a reporter gene when introduced into the cells derived from Cockayne Syndrome patients as compared to cells derived from control individuals.
36. A transformed cell made by the method of any one of claims 33 to 35.
37. A transformed cell comprising the promoter of any one of claims 1 to 21.
38. A transformed cell comprising the expression cassette of any one of claims 22 to 25.
39. A method of treating a genetic disorder in a mammal comprising administering(a) the vector of claim 26 or 27, or(b) the transformed cell of any one of claims 36 to 38.
40. The method of claim 39, wherein the genetic disorder is spinal muscular atrophy, Leber congenital amaurosis, hemophilia a neurological disorder.
41. The method of claim 40, wherein the genetic disorder is a neurological disorder.
42. The method of claim 41, wherein the neurological disorder is Cockayne Syndrome (CS).
43. A use of the isolated promoter of any one of claims 1 to 21, wherein the promoter activates, enhances or represses the expression of a genetic sequence at onset of or during pathological progression of a genetic disorder.
44. An isolated transformed cell made by the method of claim 33.