Selective, tunable and differential autoregulation of CFTR gene expression

WO2026178389A1PCT designated stage Publication Date: 2026-08-27JIANG HONG
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/US2026/016076
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2026-02-05
Filing Date
2026-02-20
Publication Date
2026-08-27

Smart Images

  • Figure US2026016076_27082026_PF_FP_ABST
    Figure US2026016076_27082026_PF_FP_ABST
Patent Text Reader

Abstract

Provided are plasmids comprising transcriptional control elements, including transcription factor binding motifs, telomeric repeat motifs, transcription factor repressor binding motifs and promoters for tunable protein expression in target cells; wherein the plasmid for controlled transcription in a target cell type comprises a transcription factor binding motif for NFIA and the cell expresses NFIA. In one embodiment, the cell type is lung cells and the transcription factor is Nuclear Factor IA (NFIA).
Need to check novelty before this filing date? Find Prior Art

Description

Attorney Docket No. 206678-0002-00WOSELECTIVE, TUNABLE AND DIFFERENTIAL AUTOREGULATION OF CFTR GENE EXPRESSIONCROSS REFERENCE TO RELATED APPLICATIONS

[0001] This application claims priority to U. S. Provisional Application No. 63 / 760,840, filed February 20, 2025, U. S. Provisional Application No. 63 / 795,954, filed April 28, 2025, U. S. Provisional Application No. 63 / 796,731, filed April 29, 2025, U. S. Provisional Application No.63 / 798,202, filed May 1, 2025, U. S. Provisional Application No. 63 / 892,226, filed October 2, 2025, and U. S. Provisional Application No. 63 / 976,539, filed February 5, 2026, each of which is hereby incorporated by reference herein in its entirety.REFERENCE TO A “SEQUENCE LISTING” SUBMITTED AS AN XML FILE

[0002] The present application hereby incorporates by reference the entire contents of the sequence listing xml document named “206678-0002-00WO_SequenceListing.xml”. The xml file containing the Sequence Listing of the present application was created on February 20, 2026, and is 288,537 bytes in size.BACKGROUND OF THE INVENTION

[0003] Cystic fibrosis (CF) is caused by mutations in the CFTR gene, many of which result in misfolded or nonfunctional CFTR proteins. The most common mutation, AF508, leads to CFTR misfolding and retention in the endoplasmic reticulum (ER), triggering proteostasis stress, inflammation, and defective ion transport. Even in the presence of a functional CFTR transgene, the mutant CFTR allele remains active, continuing to produce misfolded proteins that can interfere with cellular homeostasis. Misfolded CFTR can exert dominant-negative effects, impairing the trafficking and function of co-expressed wild-type CFTR. Additionally, chronic ER stress caused by mutant CFTR accumulation can activate the unfolded protein response (UPR), leading to widespread transcriptional changes and reduced overall protein synthesis. The presence of mutant CFTR also sustains aberrant inflammatory signaling, particularly through NF-KB activation, contributing to epithelial dysfunction in the lungs. Indeed, studies show that strongly driving CFTR with a constitutive promoter raises mRNA levels but does not206678-0002-00WOproportionally increase mature protein, due to inefficient maturation and ER-associated degradation. This not only limits therapeutic benefit but also risks engaging the unfolded protein response (UPR) if misfolded protein accumulates. Therefore, controlling CFTR transgene expression within a safe range is crucial.

[0004] A major challenge in CFTR gene therapy is controlling the level of CFTR transgene expression, particularly when delivered using lipid nanoparticles (LNPs). Ideally gene therapies should not only introduce a functional CFTR transgene but also suppress the expression of endogenous mutant CFTR.

[0005] LNP-based delivery often results in heterogeneous uptake, with some cells receiving multiple copies of the transgene. In the absence of a regulatory mechanism, cells that receive excessive copies of the CFTR transgene could produce far more CFTR protein than necessary, leading to unintended cytotoxic or functional consequences. While CFTR restoration is the therapeutic goal, excessive CFTR expression could be problematic. Overexpressed membrane proteins can overload the ER’s folding and trafficking machinery, increasing the risk of protein aggregation and proteostasis imbalance. High levels of CFTR may also disrupt epithelial ion homeostasis, potentially leading to unintended physiological consequences, such as excessive chloride efflux or sodium imbalance. Moreover, the immune system may detect and respond to excessive transgene expression, raising concerns about immunogenicity and longterm safety. Controlling transgene expression levels is therefore critical to ensuring therapeutic efficacy while minimizing potential risks.

[0006] There remains a need in the art for improved methods for delivery of therapeutic nucleic acid molecules, including for delivery of gene therapy agents for controlled expression. The present invention addresses this need.SUMMARY OF THE INVENTION

[0007] In some embodiments, the invention relates to plasmids comprising transcriptional control elements, including transcription factor binding motifs, telomeric repeat motifs, transcription factor repressor binding motifs and promoters for tunable protein expression in target cells.206678-0002-00WO

[0008] In one embodiment, the invention relates to a plasmid for controlled transcription in a target cell type comprising at least one transcription factor binding motif for a transcription factor that is expressed in the target cell type or a subset of cells of the target cell type.

[0009] In one embodiment, the cell type is lung cells, or a subset thereof. In one embodiment, the plasmid comprises at least one transcription factor binding motif is specific for binding to Nuclear Factor I A (NFIA), Forkhead Box A2 (FOXA2), Multiciliate Differentiation and DNA Synthesis- Associated Cell Cycle Protein (MCIDAS), SAM-pointed Domain ETS Factor (SPDEF), Forkhead Box II (FOXIl), Transcription Factor CP2-like 1 (TFCP2L1), Achaete-Scute Family bHLH 3 (ASCL3), Tumor Protein p63 (TP63), Forkhead Box QI (FOXQ1), Forkhead Box JI (FOXJ1), Regulatory Factor X3 (RFX3), SRY-box 9 (SOX9), Hepatocyte Nuclear Factor la (HNFla), Hepatocyte Nuclear Factor 4a (HNF4a), Forkhead Box A3 (FOXA3), or any combination thereof.

[0010] In one embodiment, the plasmid comprises SEQ ID NO:1, SEQ ID NO:2, SEQ ID NO:3, SEQ ID NO:4, SEQ ID NO:5, SEQ ID NO:6, SEQ ID NO:7, GCCACTTAA, CCCACTTAA, GCCACTTAG, ACCACTTAG, GGCACTTAA, AGCACTTAA, CCCACTTAG, TCCACTTAA, SEQ ID NO: 95 or SEQ ID NO: 96, or a fragment or variant thereof which serves as a binding site for NFIA. In one embodiment, the plasmid comprises AATAAAG, ATAAACA, GTAAATA, GTAAACA, GTAAACAA, ATAAAT, GTAAAT, TGTTTAC, TGTTTAT, SEQ IDN0:8, SEQ ID N0:9, SEQ ID NO: 10, SEQ ID NO: 11, SEQ ID NO: 12, SEQ ID NO: 13, SEQ ID NO: 14, SEQ ID NO: 15, SEQ ID NO: 16, SEQ ID NO: 17, SEQ ID NO: 18 or a fragment or variant thereof which serves as a binding site for FOXA2. In one embodiment, the plasmid comprises TTTCGCGC, TTTGGCGC, TTTCCGCC, TTTCCCGC, TTTGCCGC, or TGTCCCGC, or a fragment or variant thereof which serves as a binding site for MCIDAS. In one embodiment, the plasmid comprises GGAT, AGGAT, TGGAT, CGGAT, GGAA, AGGATTC, ATGCGGGC, GTGCGGGT, GTGCGGGC, ATACGGGT, ATGGGGGT, ATGCGGGG, CTGCGGGT, ATACGGGC, ATGCGGGA or SEQ ID NO: 19 or a fragment or variant thereof which serves as a binding site for SPDEF. In one embodiment, the plasmid comprises TGTTTAC, GTAAACA, GTAAATA, TATTTAT, TGTTTAT, TGTTTGT, TATTTAC, GTCAACA, GTAATCA, ATAAACA, ATCAACA, GTAAAAA, GTAAATAA, GTCAATA or SEQ ID NO:20, or a fragment or variant thereof which serves as a binding site for FOXIl. In one embodiment, the plasmid comprises206678-0002-00WOAAACCGGTT, SEQ ID NO:21, SEQ ID NO:22, CCAGTTCAA, CAGTTCAAC, or SEQ ID NO:23, or a fragment or variant thereof which serves as a binding site for TFCP2L1. In one embodiment, the plasmid comprises CAGGTG, CACCTG, CACGTG, GCACCTGCC, CCACCTGCC, GCACCTGCT, ACACCTGCC, GCACCTGCA, GCACCTGCG, CCACCTGCT, TCACCTGCC, GCAGCTGCC, or GCACCTGGC, or a fragment or variant thereof which serves as a binding site for ASCL3. In one embodiment, the plasmid comprises SEQ ID NO:24, SEQ ID NO:25, SEQ ID NO:26, SEQ ID NO:27, SEQ ID NO:28, SEQ ID NO:29, SEQ ID NO:30 or SEQ ID NO:31, or a fragment or variant thereof which serves as a binding site for TP63. In one embodiment, the plasmid comprises ACAAAG, ATAAAG, ACAAAT, ATAAAC, ATAAAT, GTAAAC, TGTTTAC, TCAATA, GTAAATAA or SEQ ID NO:32, or a fragment or variant thereof which serves as a binding site for FOXQ1. In one embodiment, the plasmid comprises TGTTTAC, GTAAATA, GTTTACA, ATAAATA, GTAAAC AAA, ATAAAC AAA, ATAAAC AA, TAAACAAA, or SEQ ID NO: 33, or a fragment or variant thereof which serves as a binding site for FOXJ1. In one embodiment, the plasmid comprises SEQ ID NO:34, SEQ ID NO:35, SEQ ID NO:36, GTTACCATG, GTTGCTATG, GTTACTATG, or SEQ ID NO:37, or a fragment or variant thereof which serves as a binding site for RFX3. In one embodiment, the plasmid comprises SEQ ID NO:38, SEQ ID NO:39, SEQ ID NO:40, CATTGAA, CTTTGTT, CTTTGAA, ACAAAG, or TTCAAAG, or a fragment or variant thereof which serves as a binding site for SOX9. In one embodiment, the plasmid comprises GTTAAT, TTGTTA, SEQ ID NO:41, SEQ ID NO:47, SEQ ID NO:43, SEQ ID NO:44, SEQ ID NO:45, or SEQ ID NO:46 or a fragment or variant thereof which serves as a binding site for HNFla. In one embodiment, the plasmid comprises SEQ ID NO:47, SEQ ID NO:48, SEQ ID NO:49, SEQ ID NO:50, or SEQ ID NO:51, or a fragment or variant thereof which serves as a binding site for HNF4a. In one embodiment, the plasmid comprises GTAAACA, ATAAATA, ATCAATA, ATAAAT, ATAAAC, GTAAATAA, or GTAAAT, or a fragment or variant thereof which serves as a binding site for FOXA3.

[0011] In one embodiment, the plasmid comprises a cluster of transcription factor binding sites for recognition by two or more of NFIA, FOXA2, MCIDAS, SPDEF, FOXI1, TFCP2L1, ASCL3, TP63, FOXQ1, FOXJ1, RFX3, SOX9, HNFla, HNF4a, and FOXA3. In one embodiment, the plasmid comprises SEQ ID NO: 121, SEQ ID NO: 122, or SEQ ID NO: 125.

[0012] In one embodiment, the target cell type is muscle cells, or a subset thereof. In one embodiment, the plasmid comprises at least one transcription factor binding motif specific for binding to Myocyte enhancer factor 2A (MEF2A), Myogenin (MYOG), Muscle-Specific Regulatory Factor 4 (MRF4), TEA Domain Transcription Factor 1 (TEAD1), Paired Box 7 (PAX7), or SIX Homeobox 1 (SIX1), or any combination thereof.

[0013] In one embodiment, the plasmid comprises SEQ ID NO:52, SEQ ID NO:53, SEQ ID NO:54, SEQ ID NO:55, SEQ ID NO:56, SEQ ID NO:57, SEQ ID NO:58, SEQ ID NO:59, SEQ ID NO: 60, SEQ ID NO:61 or SEQ ID NO: 62, or a fragment or variant thereof which serves as a binding site for MEF2A. In one embodiment, the plasmid comprises SEQ ID NO:63, SEQ ID NO:64, SEQ ID NO:65, SEQ ID NO:66, SEQ ID NO:67, SEQ ID NO:68, SEQ ID NO:69, SEQ ID NO:70, SEQ ID NO:71, SEQ ID NO:72 or SEQ ID NO:73, or a fragment or variant thereof which serves as a binding site for MYOG. In one embodiment, the plasmid comprises SEQ ID NO:74, SEQ ID NO:75, SEQ ID NO:76, SEQ ID NO:77, SEQ ID NO:78, SEQ ID NO: 79, SEQ ID NO: 80, SEQ ID NO:81, SEQ ID NO: 82, or SEQ ID NO: 83, or a fragment or variant thereof which serves as a binding site for MRF4. In one embodiment, the plasmid comprises SEQ ID NO:84, SEQ ID NO:85, SEQ ID NO:86, SEQ ID NO:87, SEQ ID NO:88, SEQ ID NO:89, SEQ ID NO:90, SEQ ID NO:91, SEQ ID NO:92, or SEQ ID NO:93, or a fragment or variant thereof which serves as a binding site for TEAD1. In one embodiment, the plasmid comprises SEQ ID NO:94, or a fragment or variant thereof which serves as a binding site for PAX7. In one embodiment, the plasmid comprises CTAATTA, CTCATTA, TTAATTA, CAAATTA, CCAATTA, TTCATTA, CACATTA, CCCATTA, CTAATTG, or TAAATTA, or a fragment or variant thereof which serves as a binding site for SIX1.

[0014] In one embodiment, the plasmid comprises a cluster of transcription factor binding sites for recognition by two or more of MEF2A, MYOG, MRF4, TEAD1, PAX7, and SIX1. In one embodiment, the plasmid comprises SEQ ID NO: 128.

[0015] In one embodiment, the invention relates to a plasmid comprising a promoter comprising a combination of:a) a TFIIB recognition motif (referred to as a BREu element); b) a TATA box motif;c) an initiator; andd) a downstream promoter element (DPE).206678-0002-00WO

[0016] In one embodiment, the BREu element comprises GGGCGCC, GCACGCC, or GCGCGCC. In one embodiment, the TATA box motif comprises TATATAA, TATAAAAG, TATATAAG, TAT AAA, SEQ ID NO:264 or SEQ ID NO:265. In one embodiment, the initiator comprises TCAGTT, TC4TTC, TC4GTCT, TC4TATC, TCAGTTCC, GC4GTT, or CC4CTT. In one embodiment, the DPE comprises GGACCT, ACCT, AGTCGC, GGACTGG, GGTTTC, AGACGTG or AGACGT.

[0017] In one embodiment, the promoter further comprises a Motif Ten Element (MTE) located between the initiator and the DPE. In one embodiment, the MTE comprises SEQ ID NO: 145, SEQ ID NO: 146, SEQ ID NO: 147, SEQ ID NO: 148, SEQ ID NO: 149, SEQ ID NO:271 or SEQ ID NO:272. In one embodiment, the final two nucleotides of the MTE and the first two nucleotides of the DPE overlap.

[0018] In one embodiment, the plasmid comprises a promoter sequence of SEQ ID NO:158, SEQ IDNO:159, SEQ IDNO:160, SEQ IDNO:161, SEQ IDNO:162, SEQ ID NO: 163, SEQ ID NO: 164, SEQ ID NO: 165, SEQ ID NO: 166, SEQ ID NO: 167, SEQ ID NO: 168, SEQ ID NO: 169, SEQ ID NO: 170, SEQ ID NO: 171, SEQ ID NO: 172, SEQ ID NO:173, SEQ IDNO:174, SEQ IDNO:175, SEQ IDNO:176, SEQ IDNO:177, SEQ ID NO: 178, SEQ ID NO: 179, SEQ ID NO: 180, SEQ ID NO: 181, SEQ ID NO: 182, SEQ ID NO: 183, SEQ ID NO: 184, SEQ ID NO: 185, SEQ ID NO: 186, SEQ ID NO: 187, SEQ ID NO:188, SEQ IDNO:189, SEQ IDNO:190, SEQ IDNO:191, SEQ IDNO:192, SEQ ID NO:193, SEQ IDNO:194, SEQ IDNO:195, SEQ ID NO:196, SEQ IDNO:197, SEQ ID NO: 198, SEQ ID NO: 199, SEQ IDNO:200, SEQ IDNO:201, SEQ IDNO:202, SEQ ID NO:203, SEQ ID NO:204, SEQ ID NO:205, SEQ ID NO:206, SEQ ID NO:207, SEQ ID NO:208, SEQ ID NO:209, SEQ ID NO:210, SEQ ID NO:211, SEQ ID NO:212, SEQ ID NO:213, SEQ ID NO:214, SEQ ID NO:215, SEQ ID NO:279, SEQ ID NO:280, SEQ ID NO:281, SEQ ID NO:282, SEQ ID NO:283, or SEQ ID NO:284.

[0019] In one embodiment, the invention relates to an expression plasmid comprising at least one telomeric repeat motif between a sequence which serves as a promoter and a start codon of a coding sequence. In one embodiment, the telomeric repeat motif comprises SEQ ID NO:216, SEQ ID NO:217, SEQ ID NO:218, SEQ ID NO:219, SEQ ID NO:220, SEQ ID NO:221, SEQ ID NO:222, SEQ ID NO:223, SEQ ID NO:224, SEQ ID NO:225, or SEQ ID NO: 226.206678-0002-00WO

[0020] In one embodiment, the expression plasmid further comprises an intron downstream of the promoter and upstream of the start codon. In one embodiment, the intron comprises SEQ ID NO:229, SEQ ID NO:230 or SEQ ID NO:231

[0021] In one embodiment, the telomeric repeat motif is within the intron. In one embodiment, the intron comprising the telomeric repeat comprises SEQ ID NO:232, SEQ ID NO:233, SEQ ID NO:234, SEQ ID NO:235, SEQ ID NO:236, SEQ ID NO:237, SEQ ID NO:238, or SEQ IDNO:239.

[0022] In one embodiment, the invention relates to an expression plasmid comprising expression plasmid comprising at least one post-polyA telomeric repeat motif following the polyA tail of a coding sequence. In one embodiment, the post-polyA telomeric repeat motif comprises SEQ ID NO:227 or SEQ ID NO:228.

[0023] In one embodiment, the invention relates to an expression plasmid comprising at least one telomeric repeat motif between a sequence which serves as a promoter and a start codon of a coding sequence and further comprising at least one post-polyA telomeric repeat motif.

[0024] In one embodiment, the invention relates to an expression plasmid comprising a 5’ UTR of SEQ ID NO:240, SEQ ID NO:241, SEQ ID NO:242, SEQ ID NO:243, SEQ ID NO:244, or SEQ ID NO:245.

[0025] In one embodiment, the invention relates to an expression plasmid comprising a 3’ UTR of SEQ ID NO:252, SEQ ID NO:253, SEQ ID NO:254 or SEQ ID NO:255.

[0026] In one embodiment, the invention relates to an expression plasmid comprising a polyA tail comprising a sequence as set forth in SEQ ID NO:256.

[0027] In one embodiment, the invention relates to a codon optimized sequence encoding cystic fibrosis transmembrane conductance regulator (CFTR) comprising SEQ ID NO: 248 or SEQ ID NO:249, or a variant thereof encoding SEQ ID NO:247.

[0028] In one embodiment, the invention relates to a polycistronic coding sequence encoding a combination of a codon optimized sequence encoding green fluorescent protein (GFP) and a codon optimized sequence encoding CFTR comprising SEQ ID NO: 248 or SEQ ID NO:249, or a variant thereof encoding SEQ ID NO:247. In one embodiment, the polycistronic coding sequence comprises a combination of a codon optimized sequence encoding CFTR comprising SEQ ID NO: 248 or SEQ ID NO:249, or a variant thereof encoding SEQ ID NO:247,206678-0002-00WOand a codon optimized sequence encoding GFP comprising SEQ ID NO:251, or a variant thereof encoding SEQ ID NO:250.

[0029] In one embodiment, the invention relates to a plasmid for controlled transcription in a target cell type comprising at least one transcription factor binding motif for a transcription factor that is expressed in the target cell type or a subset of cells of the target cell type, and further comprising at least one of:a) a promoter comprising a combination of:i) a BREu element;ii) a TATA box motif;iii) an initiator; andiv) a downstream promoter element (DPE);b) at least one telomeric repeat motif between the promoter and the start codon of a coding sequence;c) the 5’ UTR of SEQ ID NO: 240, SEQ ID NO: 241, SEQ ID NO: 242, SEQ ID NO:243, SEQ ID NO:244, or SEQ ID NO:245;d) the 3’ UTR of SEQ ID NO 252, SEQ ID NO:253, SEQ ID NO:254 or SEQ ID NO:255; ande) the polyA tail of SEQ ID NO:256.

[0030] In one embodiment, the plasmid further comprises the codon optimized sequence encoding CFTR comprising SEQ ID NO: 248 or SEQ ID NO:249, or a variant thereof encoding SEQ ID NO:247, or the polycistronic coding sequence comprising a combination of a codon optimized sequence encoding CFTR comprising SEQ ID NO: 248 or SEQ ID NO:249, or a variant thereof encoding SEQ ID NO:247, and a codon optimized sequence encoding GFP comprising SEQ ID NO:251, or a variant thereof encoding SEQ ID NO:250.

[0031] In one embodiment, the plasmid further comprises at least one transcription factor repressor binding site for EHF or IRF2 downstream of the promoter but upstream of the 5’ UTR.

[0032] In one embodiment, the plasmid comprises SEQ ID NO:257, SEQ ID NO:258, SEQ ID NO:259, SEQ ID NO:260, SEQ ID NO:261, SEQ ID NO:262 or SEQ ID NO:263.

[0033] In one embodiment, the invention relates to a delivery vehicle comprising a plasmid as described above. In one embodiment, the delivery vehicle comprises a liposome or lipid nanoparticle (LNP).

[0034] In one embodiment, the invention relates to a method of treating cystic fibrosis in a subject in need thereof, the method comprising administering to a subject in need thereof a plasmid as described above or a delivery vehicle comprising the plasmid, wherein the plasmid comprises the codon optimized sequence encoding CFTR comprising SEQ ID NO: 248 or SEQ ID NO:249, or a variant thereof encoding SEQ ID NO:247, or the polycistronic coding sequence comprising a combination of a codon optimized sequence encoding CFTR comprising SEQ ID NO: 248 or SEQ ID NO:249, or a variant thereof encoding SEQ ID NO:247, and a codon optimized sequence encoding GFP comprising SEQ ID NO:251, or a variant thereof encoding SEQ ID NO:250.BRIEF DESCRIPTION OF THE DRAWINGS

[0035] The following detailed description of embodiments of the invention will be better understood when read in conjunction with the appended drawings. It should be understood that the invention is not limited to the precise arrangements and instrumentalities of the embodiments shown in the drawings.

[0036] Fig. 1 depicts a schematic of a plasma DNA gene cassette with a DTS, IRF2 / EHF binding motifs, TF promoter, IME / Intron, telomeric motifs, Kozak sequence, polycistronic CDS, and vertical poly-A tail with repressor protein regulation.

[0037] Fig. 2 depicts images showing GFP expression in HEK293T cells transfected with a plasmid encoding PositivoBio_19 at 20H and 48H following transfection. Images of 10X objective, left:20H, right 48H.

[0038] Fig. 3 depicts images showing GFP expression in HEK293T cells transfected with a plasmid encoding PositivoBio_19 + IT element at 20H and 48H following transfection. Images of 10X objective, left:20H, right 48H.

[0039] Fig. 4 depicts images showing GFP expression in HEK293T cells transfected with a plasmid encoding PositivoBio_2 at 20H and 48H following transfection. Images of 10X objective, left:20H, right 48H.

[0040] Fig. 5 depicts images showing GFP expression in HEK293T cells transfected with a plasmid encoding PositivoBio_2 + IT element at 20H and 48H following transfection. Images of 10X objective, left:20H, right 48H.206678-0002-00WO

[0041] Fig. 6 depicts images showing GFP expression in HEK293T cells transfected with a plasmid encoding PositivoBio_25 at 20H and 48H following transfection. Images of 10X objective, left:20H, right 48H.

[0042] Fig. 7 depicts images showing GFP expression in HEK293T cells transfected with a plasmid encoding PositivoBio_25 + IT element at 20H and 48H following transfection. Images of 10X objective, left:20H, right 48H.

[0043] Fig. 8 depicts images showing GFP expression in HEK293T cells transfected with an Aldevron Nanoplasmid EFla-eGFP at 20H and 48H following transfection. Images of 10X objective, left:20H, right 48H.

[0044] Fig. 9 depicts images showing GFP expression in untreated HEK293T cells at 20H and 48H. Images of 10X objective, left:20H, right 48H.

[0045] Fig. 10 depicts the quantification of relative fluorescent units (RFU) for the treated HEK293T cells at 20H.

[0046] Fig. 11 depicts the quantification of relative fluorescent units (RFU) for the treated HEK293T cells at 48H.DETAILED DESCRIPTION

[0047] The present invention relates to plasmids comprising transcriptional control elements, including transcription factor binding motifs, telomeric repeat motifs, transcription factor repressor binding motifs and promoters for tunable protein expression in target cells, and compositions comprising the plasmids.

[0048] In some embodiments, it also relates to methods of use of the compositions described herein for treating diseases or disorders in subjects including, but not limited to, cystic fibrosis.Definitions

[0049] Unless defined otherwise, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention belongs.

[0050] As used herein, each of the following terms has the meaning associated with it in this section.206678-0002-00WO

[0051] The articles “a” and “an” are used herein to refer to one or to more than one (i.e., to at least one) of the grammatical object of the article. By way of example, “an element” means one element or more than one element.

[0052] “About” as used herein when referring to a measurable value such as an amount, a temporal duration, and the like, is meant to encompass variations of ±20%, ±10%, ±5%, ±1%, or ±0.1% from the specified value, as such variations are appropriate to perform the disclosed methods.

[0053] A “disease” is a state of health of an animal wherein the animal cannot maintain homeostasis, and wherein if the disease is not ameliorated then the animal’s health continues to deteriorate. In contrast, a “disorder” in an animal is a state of health in which the animal is able to maintain homeostasis, but in which the animal’s state of health is less favorable than it would be in the absence of the disorder. Left untreated, a disorder does not necessarily cause a further decrease in the animal’s state of health.

[0054] “Encoding” refers to the inherent property of specific sequences of nucleotides in a polynucleotide, such as a gene, a cDNA, or an mRNA, to serve as templates for synthesis of other polymers and macromolecules in biological processes having either a defined sequence of nucleotides (i.e., rRNA, tRNA and mRNA) or a defined sequence of amino acids and the biological properties resulting therefrom. Thus, a gene encodes a protein if transcription and translation of mRNA corresponding to that gene produces the protein in a cell or other biological system. Both the coding strand, the nucleotide sequence of which is identical to the mRNA sequence and is usually provided in sequence listings, and the non-coding strand, used as the template for transcription of a gene or cDNA, can be referred to as encoding the protein or other product of that gene or cDNA.

[0055] “Homologous” refers to the sequence similarity or sequence identity between two polypeptides or between two nucleic acid molecules. When a position in both of the two compared sequences is occupied by the same base or amino acid monomer subunit, e.g., if a position in each of two DNA molecules is occupied by adenine, then the molecules are homologous at that position. The percent of homology between two sequences is a function of the number of matching or homologous positions shared by the two sequences divided by the number of positions compared X 100. For example, if 6 of 10 of the positions in two sequences are matched or homologous then the two sequences are 60% homologous. By way of example,206678-0002-00WOthe DNA sequences ATTGCC and TATGGC share 50% homology. Generally, a comparison is made when two sequences are aligned to give maximum homology.

[0056] “Isolated” means altered or removed from the natural state. For example, a nucleic acid or a peptide naturally present in a living animal is not “isolated,” but the same nucleic acid or peptide partially or completely separated from the coexisting materials of its natural state is “isolated.” An isolated nucleic acid or protein can exist in substantially purified form, or can exist in a non-native environment such as, for example, a host cell.

[0057] By describing two polynucleotides as “operably linked” is meant that a singlestranded or double-stranded nucleic acid moiety comprises the two polynucleotides arranged within the nucleic acid moiety in such a manner that at least one of the two polynucleotides is able to exert a physiological effect by which it is characterized, upon the other. By way of example, a promoter operably linked to the coding region of a gene is able to promote transcription of the coding region.

[0058] In one embodiment, when the nucleic acid encoding the desired protein further comprises a promoter / regulatory sequence, the promoter / regulatory sequence is positioned at the 5’ end of the desired protein coding sequence such that it drives expression of the desired protein in a cell. Together, the nucleic acid encoding the desired protein and its promoter / regulatory sequence comprise a “transgene.”

[0059] In the context of the present invention, the following abbreviations for the commonly occurring nucleosides (nucleobase bound to ribose or deoxyribose sugar via N-glycosidic linkage) are used. “A” refers to adenosine, “C” refers to cytidine, “G” refers to guanosine, “T” refers to thymidine, and “U” refers to uridine.

[0060] Unless otherwise specified, a “nucleotide sequence encoding an amino acid sequence” includes all nucleotide sequences that are degenerate versions of each other and that encode the same amino acid sequence. The phrase nucleotide sequence that encodes a protein or an RNA may also include introns to the extent that the nucleotide sequence encoding the protein may in some version contain an intron(s).

[0061] By the term “modulating,” as used herein, is meant mediating a detectable increase or decrease in the level of a response in a subject compared with the level of a response in the subject in the absence of a treatment or compound, and / or compared with the level of a response in an otherwise identical but untreated subject. The term encompasses perturbing and / oraffecting a native signal or response thereby mediating a beneficial therapeutic response in a subject, preferably, a human.

[0062] Unless otherwise specified, a “nucleotide sequence encoding an amino acid sequence” includes all nucleotide sequences that are degenerate versions of each other and that encode the same amino acid sequence. Nucleotide sequences that encode proteins and RNA may include introns. In addition, the nucleotide sequence may contain modified nucleosides that are capable of being translated by translational machinery in a cell. For example, an mRNA where all of the uridines have been replaced with pseudouridine, 1 -methyl pseudouridine, or another modified nucleoside.

[0063] The terms “patient,” “subject,” “individual,” and the like are used interchangeably herein, and refer to any animal, or cells thereof whether in vitro or in situ, amenable to the methods described herein. In certain non-limiting embodiments, the patient, subject or individual is a human.

[0064] The term “polynucleotide” as used herein is defined as a chain of nucleotides. Furthermore, nucleic acids are polymers of nucleotides. Thus, nucleic acids and polynucleotides as used herein are interchangeable. One skilled in the art has the general knowledge that nucleic acids are polynucleotides, which can be hydrolyzed into the monomeric “nucleotides.” The monomeric nucleotides can be hydrolyzed into nucleosides. As used herein polynucleotides include, but are not limited to, all nucleic acid sequences which are obtained by any means available in the art, including, without limitation, recombinant means, i.e., the cloning of nucleic acid sequences from a recombinant library or a cell genome, using ordinary cloning technology and PCR™, and the like, and by synthetic means.

[0065] In certain instances, the polynucleotide or nucleic acid of the invention is a “nucleoside-modified nucleic acid,” which refers to a nucleic acid comprising at least one modified nucleoside. A “modified nucleoside” refers to a nucleoside with a modification. For example, over one hundred different nucleoside modifications have been identified in RNA (Rozenski, et al., 1999, The RNA Modification Database: 1999 update. Nucl Acids Res 27: 196-197).

[0066] As used herein, the terms “peptide,” “polypeptide,” and “protein” are used interchangeably, and refer to a compound comprised of amino acid residues covalently linked by peptide bonds. A protein or peptide must contain at least two amino acids, and no limitation is206678-0002-00WOplaced on the maximum number of amino acids that can comprise a protein’s or peptide’s sequence. Polypeptides include any peptide or protein comprising two or more amino acids joined to each other by peptide bonds. As used herein, the term refers to both short chains, which also commonly are referred to in the art as peptides, oligopeptides and oligomers, for example, and to longer chains, which generally are referred to in the art as proteins, of which there are many types. “Polypeptides” include, for example, biologically active fragments, substantially homologous polypeptides, oligopeptides, homodimers, heterodimers, variants of polypeptides, modified polypeptides, derivatives, analogs, fusion proteins, among others. The polypeptides include natural peptides, recombinant peptides, synthetic peptides, or a combination thereof.

[0067] By the term “specifically binds,” as used herein with respect to an affinity ligand, in particular, an antibody, is meant an antibody which recognizes a specific antigen, but does not substantially recognize or bind other molecules in a sample. For example, an antibody that specifically binds to an antigen from one species may also bind to that antigen from one or more other species. But, such cross-species reactivity does not itself alter the classification of an antibody as specific. In another example, an antibody that specifically binds to an antigen may also bind to different allelic forms of the antigen. However, such cross reactivity does not itself alter the classification of an antibody as specific. In some instances, the terms “specific binding” or “specifically binding,” can be used in reference to the interaction of an antibody, a protein, or a peptide with a second chemical species, to mean that the interaction is dependent upon the presence of a particular structure (e.g., an antigenic determinant or epitope) on the chemical species; for example, an antibody recognizes and binds to a specific protein structure rather than to proteins generally. If an antibody is specific for epitope “A”, the presence of a molecule containing epitope A (or free, unlabeled A), in a reaction containing labeled “A” and the antibody, will reduce the amount of labeled A bound to the antibody.

[0068] The term “therapeutic” as used herein means a treatment and / or prophylaxis. A therapeutic effect is obtained by suppression, diminution, remission, or eradication of at least one sign or symptom of a disease or disorder.

[0069] To “treat” a disease as the term is used herein, means to reduce the frequency or severity of at least one sign or symptom of a disease or disorder experienced by a subject.

[0070] The term “transfected” as used herein refers to a process by which exogenous nucleic acid is transferred or introduced into the host cell. A “transfected” cell is one which has206678-0002-00WObeen transfected with exogenous nucleic acid. The cell includes the primary subject cell and its progeny.

[0071] The phrase “under transcriptional control” or “operatively linked” as used herein means that the promoter is in the correct location and orientation in relation to a polynucleotide to control the initiation of transcription by RNA polymerase and expression of the polynucleotide.

[0072] A “vector” is a composition of matter which comprises an isolated nucleic acid and which can be used to deliver the isolated nucleic acid to the interior of a cell. Numerous vectors are known in the art including, but not limited to, linear polynucleotides, polynucleotides associated with ionic or amphiphilic compounds, plasmids, and viruses. Thus, the term “vector” includes an autonomously replicating plasmid or a virus. The term should also be construed to include non-plasmid and non-viral compounds which facilitate transfer of nucleic acid into cells, such as, for example, polylysine compounds, liposomes, lipid nanoparticles, and the like.Examples of viral vectors include, but are not limited to, adenoviral vectors, adeno-associated virus vectors, retroviral vectors, and the like.

[0073] “Optional” or “optionally” means that the subsequently described event of circumstances may or may not occur, and that the description includes instances where said event or circumstance occurs and instances in which it does not.

[0074] Ranges: throughout this disclosure, various aspects of the invention can be presented in a range format. It should be understood that the description in range format is merely for convenience and brevity and should not be construed as an inflexible limitation on the scope of the invention. Accordingly, the description of a range should be considered to have specifically disclosed all the possible subranges as well as individual numerical values within that range. For example, description of a range such as from 1 to 6 should be considered to have specifically disclosed subranges such as from 1 to 3, from 1 to 4, from 1 to 5, from 2 to 4, from 2 to 6, from 3 to 6 etc., as well as individual numbers within that range, for example, 1, 2, 2.7, 3, 4, 5, 5.3, and 6. This applies regardless of the breadth of the range.

[0075] The invention relates, in part, to a DNA construct for gene therapy comprising at least one motif or element for controlled or optimized transcription of an encoded product for gene therapy. Gene therapeutic constructs that can incorporate one or more motif or element forcontrolled or optimized transcription of an encoded product for gene therapy include, but are not limited to, circular plasmid DNA, minicircle DNA, nanoplasmids, linear DNA constructs, closed linear (“Doggybone”) DNA constructs, minivectors / microDNA, unidirectional expression plasmids, bidirectional expression plasmids or another form of DNA construct known in the art for use for gene therapy. The term “plasmid” as used herein refers to any form of DNA construct for gene therapy. In some embodiments, the plasmid of the invention comprises at least one transcriptional control element which restricts transcription of an encoded transcription product to a specific target cell type. In one embodiment, the plasmid of the invention comprises at least one transcriptional control element which promotes transcription of the encoded transcription product when a specific transcription factor or combination of transcription factors is present.

[0076] The invention relates, in part, to the use of the plasmid of the invention for the treatment of a disease or disorder. Exemplary disease categories where DNA gene therapies could apply, include, but are not limited to, genetic diseases or disorders, protein replacement therapy, gene silencing, oncology, infectious diseases, and anti-aging or longevity.Transcription Factor DNA Targeting Sequences

[0077] The invention relates, in part, to plasmids that have been engineered to contain a DNA targeting sequence (DTS) which serves as a binding site for an endogenous transcription factor present in a target cell. Exemplary transcription factor (TF) binding sites that can be included on the plasmid, and the target transcription factor are provided in Table 2.

[0078] In one embodiment, the plasmid comprises a sequence of SEQ ID NO: 1, SEQ ID NO:2, SEQ ID NO:3, SEQ ID NO:4, SEQ ID NO:5, SEQ ID NO:6, SEQ ID NO:7, GCCACTTAA, CCCACTTAA, GCCACTTAG, ACCACTTAG, GGCACTTAA, AGCACTTAA, CCCACTTAG, TCCACTTAA, SEQ ID NO:95 or SEQ ID NO:96, or a fragment or variant thereof which serves as a binding site for Nuclear Factor I A (NFIA). In one embodiment, the fragment or variant of SEQ ID NO: 1, SEQ ID NO:2, SEQ ID NO:3, SEQ ID NO:4, SEQ ID NO:5, SEQ ID NO:6, SEQ ID NO:7, GCCACTTAA, CCCACTTAA, GCCACTTAG, ACCACTTAG, GGCACTTAA, AGCACTTAA, CCCACTTAG, TCCACTTAA, SEQ ID NO:95 or SEQ ID NO:96 retains the ability to bind to NFIA. NFIA is expressed in Club cells, therefore in one embodiment, the plasmid comprising SEQ ID NO:1, SEQ ID NO:2, SEQ ID NO:3, SEQ ID NO:4, SEQ ID NO:5, SEQ ID NO:6, SEQ ID NO:7,GCCACTTAA, CCCACTTAA, GCCACTTAG, ACCACTTAG, GGCACTTAA, AGCACTTAA, CCCACTTAG, TCCACTTAA, SEQ ID NO:95 or SEQ ID NO:96 is administered to a subject for expression in Club cells. In one embodiment, the plasmid further comprises a coding sequence encoding a protein for expression in Club cells, a nucleic acid molecule (e.g., mRNA, siRNA, miRNA, shRNA) for delivery to Club cells, or a combination thereof. In one embodiment, the protein or nucleic acid molecule is a therapeutic agent for the treatment of a disease or disorder.

[0079] In one embodiment, the plasmid comprises a sequence of AATAAAG, ATAAACA, GTAAATA, GTAAACA, GTAAACAA, ATAAAT, GTAAAT, TGTTTAC, TGTTTAT, SEQ IDNO:8, SEQ IDNO:9, SEQ ID NO: 10, SEQ ID NO: 11, SEQ ID NO: 12, SEQ ID NO: 13, SEQ ID NO: 14, SEQ ID NO: 15, SEQ ID NO: 16, SEQ ID NO: 17, SEQ ID NO: 18 or a fragment or variant thereof which serves as a binding site for Forkhead Box A2 (FOXA2). In one embodiment, the fragment or variant of AATAAAG, ATAAACA, GTAAATA, GTAAACA, GTAAACAA, ATAAAT, GTAAAT, TGTTTAC, TGTTTAT, SEQ IDNO:8, SEQ IDNO:9, SEQ ID NO: 10, SEQ ID NO: 11, SEQ ID NO: 12, SEQ ID NO: 13, SEQ ID NO: 14, SEQ ID NO: 15, SEQ ID NO: 16, SEQ ID NO: 17, SEQ ID NO: 18 retains the ability to bind to FOXA2. FOXA2 is expressed in secretory epithelial cells including, but not limited to, Club cells and submucosal glandular cells, therefore in one embodiment, the plasmid comprising AATAAAG, ATAAACA, GTAAATA, GTAAACA, GTAAACAA, ATAAAT, GTAAAT, TGTTTAC, TGTTTAT, SEQ ID NO:8, SEQ ID NO:9, SEQ ID NO: 10, SEQ ID NO: 11, SEQ ID NO: 12, SEQ ID NO: 13, SEQ ID NO: 14, SEQ ID NO: 15, SEQ ID NO: 16, SEQ ID NO: 17, SEQ ID NO: 18 is administered to a subject for expression in secretory epithelial cells. In one embodiment, the plasmid further comprises a coding sequence encoding a protein for expression in secretory epithelial cells, a nucleic acid molecule (e.g., mRNA, siRNA, miRNA, shRNA) for delivery to secretory epithelial cells, or a combination thereof. In one embodiment, the protein or nucleic acid molecule is a therapeutic agent for the treatment of a disease or disorder.

[0080] In one embodiment, the plasmid comprises a sequence of TTTCGCGC, TTTGGCGC, TTTCCGCC, TTTCCCGC, TTTGCCGC, or TGTCCCGC, or a fragment or variant thereof which serves as a binding site for Multiciliate Differentiation and DNA Synthesis- Associated Cell Cycle Protein (MCIDAS). In one embodiment, the fragment or variant206678-0002-00WOof TTTCGCGC, TTTGGCGC, TTTCCGCC, TTTCCCGC, TTTGCCGC, or TGTCCCGC retains the ability to bind to MCIDAS. MCIDAS is expressed in multiciliated cell precursors, therefore in one embodiment, the plasmid comprising TTTCGCGC, TTTGGCGC, TTTCCGCC, TTTCCCGC, TTTGCCGC, or TGTCCCGC is administered to a subject for expression in multiciliated cell precursors or ciliated epithelial cells. In one embodiment, the plasmid further comprises a coding sequence encoding a protein for expression in multiciliated cell precursors or ciliated epithelial cells, a nucleic acid molecule (e.g., mRNA, siRNA, miRNA, shRNA) for delivery to multiciliated cell precursors or ciliated epithelial cells, or a combination thereof. In one embodiment, the protein or nucleic acid molecule is a therapeutic agent for the treatment of a disease or disorder.

[0081] In one embodiment, the plasmid comprises a sequence of GGAT, AGGAT, TGGAT, CGGAT, GGAA, AGGATTC, ATGCGGGC, GTGCGGGT, GTGCGGGC, ATACGGGT, ATGGGGGT, ATGCGGGG, CTGCGGGT, ATACGGGC, ATGCGGGA or SEQ ID NO: 19 or a fragment or variant thereof which serves as a binding site for SAM-pointed Domain ETS Factor (SPDEF). In one embodiment, the fragment or variant of GGAT, AGGAT, TGGAT, CGGAT, GGAA, AGGATTC, ATGCGGGC, GTGCGGGT, GTGCGGGC, ATACGGGT, ATGGGGGT, ATGCGGGG, CTGCGGGT, ATACGGGC, ATGCGGGA or SEQ ID NO: 19 retains the ability to bind to SPDEF. SPDEF is expressed in goblet cells and submucosal gland cells of airway epithelium, therefore in one embodiment, the plasmid comprising GGAT, AGGAT, TGGAT, CGGAT, GGAA, AGGATTC, ATGCGGGC, GTGCGGGT, GTGCGGGC, ATACGGGT, ATGGGGGT, ATGCGGGG, CTGCGGGT, ATACGGGC, ATGCGGGA or SEQ ID NO: 19 is administered to a subject for expression in goblet cells and submucosal gland cells of airway epithelium. In one embodiment, the plasmid further comprises a coding sequence encoding a protein for expression in goblet cells or submucosal gland cells of airway epithelium, a nucleic acid molecule (e.g., mRNA, siRNA, miRNA, shRNA) for delivery to goblet cells or submucosal gland cells of airway epithelium, or a combination thereof. In one embodiment, the protein or nucleic acid molecule is a therapeutic agent for the treatment of a disease or disorder.

[0082] In one embodiment, the plasmid comprises a sequence of TGTTTAC, GTAAACA, GTAAATA, TATTTAT, TGTTTAT, TGTTTGT, TATTTAC, GTCAACA, GTAATCA, ATAAACA, ATCAACA, GTAAAAA, GTAAATAA, GTCAATA or SEQ ID206678-0002-00WONO:20, or a fragment or variant thereof which serves as a binding site for Forkhead Box II (FOXI1). In one embodiment, the fragment or variant of TGTTTAC, GTAAACA, GTAAATA, TATTTAT, TGTTTAT, TGTTTGT, TATTTAC, GTCAACA, GTAATCA, ATAAACA, ATCAACA, GTAAAAA, GTAAATAA, GTCAATA or SEQ ID NO: 20 retains the ability to bind to FOXI1. FOXI1 is expressed in pulmonary ionocytes, therefore in one embodiment, the plasmid comprising TGTTTAC, GTAAACA, GTAAATA, TATTTAT, TGTTTAT, TGTTTGT, TATTTAC, GTCAACA, GTAATCA, ATAAACA, ATCAACA, GTAAAAA, GTAAATAA, GTCAATA or SEQ ID NO:20 is administered to a subject for expression in pulmonary ionocytes. In one embodiment, the plasmid further comprises a coding sequence encoding a protein for expression in pulmonary ionocytes, a nucleic acid molecule (e.g., mRNA, siRNA, miRNA, shRNA) for delivery to pulmonary ionocytes, or a combination thereof. In one embodiment, the protein or nucleic acid molecule is a therapeutic agent for the treatment of a disease or disorder.

[0083] In one embodiment, the plasmid comprises a sequence of AAACCGGTT, SEQ ID NO:21, SEQ ID NO:22, CCAGTTCAA, CAGTTCAAC, or SEQ ID NO:23, or a fragment or variant thereof which serves as a binding site for Transcription Factor CP2-like 1 (TFCP2L1). In one embodiment, the fragment or variant of AAACCGGTT, SEQ ID NO:21, SEQ ID NO:22, CCAGTTCAA, CAGTTCAAC, or SEQ ID NO:23 retains the ability to bind to TFCP2L1. TFCP2L1 is expressed in pulmonary ionocytes, therefore in one embodiment, the plasmid comprising AAACCGGTT, SEQ ID NO:21, SEQ ID NO:22, CCAGTTCAA, CAGTTCAAC, or SEQ ID NO:23 is administered to a subject for expression in pulmonary ionocytes. In one embodiment, the plasmid further comprises a coding sequence encoding a protein for expression in pulmonary ionocytes, a nucleic acid molecule (e.g., mRNA, siRNA, miRNA, shRNA) for delivery to pulmonary ionocytes, or a combination thereof. In one embodiment, the protein or nucleic acid molecule is a therapeutic agent for the treatment of a disease or disorder.

[0084] In one embodiment, the plasmid comprises a sequence of CAGGTG, CACCTG, CACGTG, GCACCTGCC, CCACCTGCC, GCACCTGCT, ACACCTGCC, GCACCTGCA, GCACCTGCG, CCACCTGCT, TCACCTGCC, GCAGCTGCC, or GCACCTGGC, or a fragment or variant thereof which serves as a binding site for Achaete-Scute Family bHLH 3 (ASCL3). In one embodiment, the fragment or variant of CAGGTG, CACCTG, CACGTG, GCACCTGCC, CCACCTGCC, GCACCTGCT, ACACCTGCC, GCACCTGCA,206678-0002-00WOGCACCTGCG, CCACCTGCT, TCACCTGCC, GCAGCTGCC, or GCACCTGGC retains the ability to bind to ASCL3. ASCL3 is expressed in ionocytes and glandular pogenitors including, but not limited to, salivary glands, therefore in one embodiment, the plasmid comprising CAGGTG, CACCTG, CACGTG, GCACCTGCC, CCACCTGCC, GCACCTGCT, ACACCTGCC, GCACCTGCA, GCACCTGCG, CCACCTGCT, TCACCTGCC, GCAGCTGCC, or GCACCTGGC is administered to a subject for expression in ionocytes or glandular pogenitors. In one embodiment, the plasmid further comprises a coding sequence encoding a protein for expression in ionocytes or glandular pogenitors, a nucleic acid molecule (e.g., mRNA, siRNA, miRNA, shRNA) for delivery to ionocytes or glandular pogenitors, or a combination thereof. In one embodiment, the protein or nucleic acid molecule is a therapeutic agent for the treatment of a disease or disorder.

[0085] In one embodiment, the plasmid comprises a sequence of SEQ ID NO:24, SEQ ID NO:25, SEQ ID NO:26, SEQ ID NO 27, SEQ ID NO:28, SEQ ID NO:29, SEQ ID NO:30 or SEQ ID NO: 31, or a fragment or variant thereof which serves as a binding site for Tumor Protein p63 (TP63). In one embodiment, the fragment or variant of SEQ ID NO:24, SEQ ID NO:25, SEQ ID NO:26, SEQ ID NO:27, SEQ ID NO:28, SEQ ID NO:29, SEQ ID NO:30 or SEQ ID NO:31 retains the ability to bind to TP63. TP63 is expressed in basal cells of airway epithelium, therefore in one embodiment, the plasmid comprising SEQ ID NO:24, SEQ ID NO:25, SEQ ID NO:26, SEQ ID NO 27, SEQ ID NO:28, SEQ ID NO:29, SEQ ID NO:30 or SEQ ID NO:31 is administered to a subject for expression in basal cells of airway epithelium. In one embodiment, the plasmid further comprises a coding sequence encoding a protein for expression in basal cells of airway epithelium, a nucleic acid molecule (e.g., mRNA, siRNA, miRNA, shRNA) for delivery to basal cells of airway epithelium, or a combination thereof. In one embodiment, the protein or nucleic acid molecule is a therapeutic agent for the treatment of a disease or disorder.

[0086] In one embodiment, the plasmid comprises a sequence of ACAAAG, ATAAAG, ACAAAT, ATAAAC, ATAAAT, GTAAAC, TGTTTAC, TCAATA, GTAAATAA or SEQ ID NO:32, or a fragment or variant thereof which serves as a binding site for Forkhead Box QI (FOXQ1). In one embodiment, the fragment or variant of ACAAAG, ATAAAG, ACAAAT, ATAAAC, ATAAAT, GTAAAC, TGTTTAC, TCAATA, GTAAATAA or SEQ ID NO:32 retains the ability to bind to FOXQ1. FOXQ1 is expressed in goblet cells of airway epithelium, therefore in one embodiment, the plasmid comprising ACAAAG, ATAAAG, ACAAAT,206678-0002-00WOATAAAC, ATAAAT, GTAAAC, TGTTTAC, TCAATA, GTAAATAA or SEQ ID NO:32 is administered to a subject for expression in goblet cells of airway epithelium. In one embodiment, the plasmid further comprises a coding sequence encoding a protein for expression in goblet cells of airway epithelium, a nucleic acid molecule (e.g., mRNA, siRNA, miRNA, shRNA) for delivery to goblet cells of airway epithelium, or a combination thereof. In one embodiment, the protein or nucleic acid molecule is a therapeutic agent for the treatment of a disease or disorder.

[0087] In one embodiment, the plasmid comprises a sequence of TGTTTAC, GTAAATA, GTTTACA, ATAAAT A, GTAAACAAA, ATAAACAAA, ATAAACAA, TAAACAAA, or SEQ ID NO:33, or a fragment or variant thereof which serves as a binding site for Forkhead Box JI (FOXJ1). In one embodiment, the fragment or variant of TGTTTAC, GTAAATA, GTTTACA, ATAAAT A, GTAAACAAA, ATAAACAAA, ATAAACAA, TAAACAAA, or SEQ ID NO:33 retains the ability to bind to FOXJ1. FOXJ1 is expressed in multi-ciliated cells of the airway, therefore in one embodiment, the plasmid comprising TGTTTAC, GTAAATA, GTTTACA, ATAAATA, GTAAACAAA, ATAAACAAA, ATAAACAA, TAAACAAA, or SEQ ID NO:33 is administered to a subject for expression in multi-ciliated cells of the airway. In one embodiment, the plasmid further comprises a coding sequence encoding a protein for expression in multi-ciliated cells of the airway, a nucleic acid molecule (e.g., mRNA, siRNA, miRNA, shRNA) for delivery to multi-ciliated cells of the airway, or a combination thereof. In one embodiment, the protein or nucleic acid molecule is a therapeutic agent for the treatment of a disease or disorder.

[0088] In one embodiment, the plasmid comprises a sequence of SEQ ID NO:34, SEQ ID NO: 35, SEQ ID NO: 36, GTTACCATG, GTTGCTATG, GTTACTATG, or SEQ ID NO: 37, or a fragment or variant thereof which serves as a binding site for Regulatory Factor X3 (RFX3). In one embodiment, the fragment or variant of SEQ ID NO:34, SEQ ID NO:35, SEQ ID NO:36, GTTACCATG, GTTGCTATG, GTTACTATG, or SEQ ID NO:37 retains the ability to bind to RFX3. RFX3 is expressed in multi-ciliated cells of the airway, therefore in one embodiment, the plasmid comprising SEQ ID NO:34, SEQ ID NO 35, SEQ ID NO:36, GTTACCATG, GTTGCTATG, GTTACTATG, or SEQ ID NO:37 is administered to a subject for expression in multi-ciliated cells of the airway. In one embodiment, the plasmid further comprises a coding sequence encoding a protein for expression in multi-ciliated cells of the airway, a nucleic acid molecule (e.g., mRNA, siRNA, miRNA, shRNA) for delivery to multi-ciliated cells of the206678-0002-00WOairway, or a combination thereof. In one embodiment, the protein or nucleic acid molecule is a therapeutic agent for the treatment of a disease or disorder.

[0089] In one embodiment, the plasmid comprises a sequence of SEQ ID NO:38, SEQ ID NO: 39, SEQ ID NO:40, CATTGAA, CTTTGTT, CTTTGAA, ACAAAG, or TTCAAAG, or a fragment or variant thereof which serves as a binding site for SRY-box 9 (SOX9). In one embodiment, the fragment or variant of SEQ ID NO:38, SEQ ID NO:39, SEQ ID NO:40, CATTGAA, CTTTGTT, CTTTGAA, ACAAAG, or TTCAAAG retains the ability to bind to SOX9. SOX9 is expressed in a subset of basal cells in airways, therefore in one embodiment, the plasmid comprising SEQ ID NO:38, SEQ ID NO:39, SEQ ID NO:40, CATTGAA, CTTTGTT, CTTTGAA, ACAAAG, or TTCAAAG is administered to a subject for expression in basal cells in airways. In one embodiment, the plasmid further comprises a coding sequence encoding a protein for expression in basal cells in airways, a nucleic acid molecule (e g., mRNA, siRNA, miRNA, shRNA) for delivery to basal cells in airways, or a combination thereof. In one embodiment, the protein or nucleic acid molecule is a therapeutic agent for the treatment of a disease or disorder.

[0090] In one embodiment, the plasmid comprises a sequence of GTTAAT, TTGTTA, SEQ ID NO:41, SEQ ID NO:47, SEQ ID NO:43, SEQ ID NO:44, SEQ ID NO:45, or SEQ ID NO:46 or a fragment or variant thereof which serves as a binding site for Hepatocyte Nuclear Factor la (HNFla). In one embodiment, the fragment or variant of GTTAAT, TTGTTA, SEQ ID NO:41, SEQ ID NO:47, SEQ ID NO 43, SEQ ID NO:44, SEQ ID NO:45, or SEQ ID NO:46 retains the ability to bind to HNFla. HNFla is expressed in submucosal glandular cells, therefore in one embodiment, the plasmid comprising GTTAAT, TTGTTA, SEQ ID NO:41, SEQ ID NO:47, SEQ ID NO:43, SEQ ID NO:44, SEQ ID NO:45, or SEQ ID NO:46 is administered to a subject for expression in submucosal glandular cells. In one embodiment, the plasmid further comprises a coding sequence encoding a protein for expression in submucosal glandular cells, a nucleic acid molecule (e.g., mRNA, siRNA, miRNA, shRNA) for delivery to submucosal glandular cells, or a combination thereof. In one embodiment, the protein or nucleic acid molecule is a therapeutic agent for the treatment of a disease or disorder.

[0091] In one embodiment, the plasmid comprises a sequence of SEQ ID NO:47, SEQ ID NO:48, SEQ ID NO:49, SEQ ID NO:50, or SEQ ID NO:51, or a fragment or variant thereof which serves as a binding site for Hepatocyte Nuclear Factor 4a (HNF4a). In one embodiment,206678-0002-00WOthe fragment or variant of SEQ ID NO:47, SEQ ID NO:48, SEQ ID NO:49, SEQ ID NO:50, or SEQ ID NO:51 retains the ability to bind to HNF4a. HNF4a is expressed in airway secretory cells and submucosal gland cells, therefore in one embodiment, the plasmid comprising SEQ ID NO:47, SEQ ID NO:48, SEQ ID NO:49, SEQ ID NO:50, or SEQ ID NO:51 is administered to a subject for expression in airway secretory cells and submucosal gland cells. In one embodiment, the plasmid further comprises a coding sequence encoding a protein for expression in airway secretory cells or submucosal gland cells, a nucleic acid molecule (e.g., mRNA, siRNA, miRNA, shRNA) for delivery to airway secretory cells or submucosal gland cells, or a combination thereof. In one embodiment, the protein or nucleic acid molecule is a therapeutic agent for the treatment of a disease or disorder.

[0092] In one embodiment, the plasmid comprises a sequence of GTAAACA, ATAAATA, ATCAATA, ATAAAT, ATAAAC, GTAAATAA, or GTAAAT, or a fragment or variant thereof which serves as a binding site for Forkhead Box A3 (FOXA3). In one embodiment, the fragment or variant of GTAAACA, ATAAATA, ATCAATA, ATAAAT, ATAAAC, GTAAATAA, or GTAAAT retains the ability to bind to FOXA3. FOXA3 is expressed in goblet cells of airway epithelium during mucus hyperplasia, therefore in one embodiment, the plasmid comprising GTAAACA, ATAAATA, ATCAATA, ATAAAT, ATAAAC, GTAAATAA, or GTAAAT is administered to a subject for expression in goblet cells of airway epithelium. In one embodiment, the plasmid further comprises a coding sequence encoding a protein for expression in goblet cells of airway epithelium, a nucleic acid molecule (e g., mRNA, siRNA, miRNA, shRNA) for delivery to goblet cells of airway epithelium, or a combination thereof. In one embodiment, the protein or nucleic acid molecule is a therapeutic agent for the treatment of a disease or disorder.

[0093] In one embodiment, the plasmid comprises a sequence of SEQ ID NO:52, SEQ ID NO:53, SEQ ID NO:54, SEQ ID NO:55, SEQ ID NO:56, SEQ ID NO:57, SEQ ID NO:58, SEQ ID NO:59, SEQ ID NO:60, SEQ ID NO:61 or SEQ ID NO:62, or a fragment or variant thereof which serves as a binding site for Myocyte enhancer factor 2A (MEF2A). In one embodiment, the fragment or variant of SEQ ID NO:52, SEQ ID NO:53, SEQ ID NO:54, SEQ ID NO:55, SEQ ID NO:56, SEQ ID NO:57, SEQ ID NO:58, SEQ ID NO:59, SEQ ID NO:60, SEQ ID NO:61 or SEQ ID NO:62 retains the ability to bind to MEF2A. MEF2A is expressed in skeletal muscle cells, therefore in one embodiment, the plasmid comprising SEQ ID NO:52, SEQ206678-0002-00WOID NO:53, SEQ ID NO:54, SEQ ID NO:55, SEQ ID NO:56, SEQ ID NO:57, SEQ ID NO:58, SEQ ID NO:59, SEQ ID NO:60, SEQ ID NO:61 or SEQ ID NO:62 is administered to a subject for expression in skeletal muscle cells. In one embodiment, the plasmid further comprises a coding sequence encoding a protein for expression in skeletal muscle cells, a nucleic acid molecule (e.g., mRNA, siRNA, miRNA, shRNA) for delivery to skeletal muscle cells, or a combination thereof. In one embodiment, the protein or nucleic acid molecule is a therapeutic agent for the treatment of a disease or disorder.

[0094] In one embodiment, the plasmid comprises a sequence of SEQ ID NO:63, SEQ ID NO:64, SEQ ID NO:65, SEQ ID NO:66, SEQ ID NO:67, SEQ ID NO:68, SEQ ID NO:69, SEQ ID NO:70, SEQ ID NO:71, SEQ ID NO:72 or SEQ ID NO:73, or a fragment or variant thereof which serves as a binding site for Myogenin (MYOG). In one embodiment, the fragment or variant of SEQ ID NO:63, SEQ ID NO:64, SEQ ID NO:65, SEQ ID NO:66, SEQ ID NO:67, SEQ ID NO:68, SEQ ID NO:69, SEQ ID NO:70, SEQ ID NO:71, SEQ ID NO:72 or SEQ ID NO:73 retains the ability to bind to MYOG. MYOG is a driver of differentiation in myocytes, therefore in one embodiment, the plasmid comprising SEQ ID NO:63, SEQ ID NO:64, SEQ ID NO:65, SEQ ID NO:66, SEQ ID NO:67, SEQ ID NO:68, SEQ ID NO:69, SEQ ID NO:70, SEQ ID NO:71, SEQ ID NO:72 or SEQ ID NO:73 is administered to a subject for expression in differentiating myocytes. In one embodiment, the plasmid further comprises a coding sequence encoding a protein for expression in differentiating myocytes, a nucleic acid molecule (e.g., mRNA, siRNA, miRNA, shRNA) for delivery to differentiating myocytes, or a combination thereof. In one embodiment, the protein or nucleic acid molecule is a therapeutic agent for the treatment of a disease or disorder.

[0095] In one embodiment, the plasmid comprises a sequence of SEQ ID NO: 74, SEQ ID NO:75, SEQ ID NO:76, SEQ ID NO:77, SEQ ID NO:78, SEQ ID NO:79, SEQ ID NO:80, SEQ ID NO:81, SEQ ID NO:82, or SEQ ID NO:83, or a fragment or variant thereof which serves as a binding site for Muscle-Specific Regulatory Factor 4 (MRF4). In one embodiment, the fragment or variant of SEQ ID NO:74, SEQ ID NO:75, SEQ ID NO:76, SEQ ID NO:77, SEQ ID NO:78, SEQ ID NO:79, SEQ ID NO:80, SEQ ID NO:81, SEQ ID NO:82, or SEQ ID NO:83 retains the ability to bind to MRF4. MRF4 is expressed in mature myofibers, therefore in one embodiment, the plasmid comprising SEQ ID NO:74, SEQ ID NO:75, SEQ ID NO:76, SEQ ID NO:77, SEQ ID NO:78, SEQ ID NO:79, SEQ ID NO:80, SEQ ID NO:81, SEQ ID NO:82, or206678-0002-00WOSEQ ID NO:83 is administered to a subject for expression in mature myofibers. In one embodiment, the plasmid further comprises a coding sequence encoding a protein for expression in mature myofibers, a nucleic acid molecule (e.g., mRNA, siRNA, miRNA, shRNA) for delivery to mature myofibers, or a combination thereof. In one embodiment, the protein or nucleic acid molecule is a therapeutic agent for the treatment of a disease or disorder.

[0096] In one embodiment, the plasmid comprises a sequence of SEQ ID NO: 84, SEQ ID NO:85, SEQ ID NO:86, SEQ ID NO:87, SEQ ID NO:88, SEQ ID NO:89, SEQ ID NO:90, SEQ ID NO:91, SEQ ID NO:92, or SEQ ID NO:93, or a fragment or variant thereof which serves as a binding site for TEA Domain Transcription Factor 1 (TEAD1). In one embodiment, the fragment or variant of SEQ ID NO:84, SEQ ID NO:85, SEQ ID NO:86, SEQ ID NO:87, SEQ ID NO:88, SEQ ID NO:89, SEQ ID NO:90, SEQ ID NO:91, SEQ ID NO:92, or SEQ ID NO:93 retains the ability to bind to TEAD1. TEAD1 is expressed in skeletal muscle cells, therefore in one embodiment, the plasmid comprising SEQ ID NO:84, SEQ ID NO:85, SEQ ID NO: 86, SEQ ID NO: 87, SEQ ID NO: 88, SEQ ID NO: 89, SEQ ID NO: 90, SEQ ID NO:91, SEQ ID NO:92, or SEQ ID NO:93 is administered to a subject for expression in skeletal muscle cells. In one embodiment, the plasmid further comprises a coding sequence encoding a protein for expression in skeletal muscle cells, a nucleic acid molecule (e.g., mRNA, siRNA, miRNA, shRNA) for delivery to skeletal muscle cells, or a combination thereof. In one embodiment, the protein or nucleic acid molecule is a therapeutic agent for the treatment of a disease or disorder.

[0097] In one embodiment, the plasmid comprises a sequence of SEQ ID NO:94, or a fragment or variant thereof which serves as a binding site for Paired Box 7 (PAX7). In one embodiment, the fragment or variant of SEQ ID NO:94 retains the ability to bind to PAX7. PAX7 is expressed in satellite cells, therefore in one embodiment, the plasmid comprising SEQ ID NO:94 is administered to a subject for expression in satellite cells. In one embodiment, the plasmid further comprises a coding sequence encoding a protein for expression in satellite cells, a nucleic acid molecule (e.g., mRNA, siRNA, miRNA, shRNA) for delivery to satellite cells, or a combination thereof. In one embodiment, the protein or nucleic acid molecule is a therapeutic agent for the treatment of a disease or disorder.

[0098] In one embodiment, the plasmid comprises a sequence of CTAATTA, CTCATTA, TTAATTA, CAAATTA, CCAATTA, TTCATTA, CACATTA, CCCATTA, CTAATTG, or TAAATTA, or a fragment or variant thereof which serves as a binding site for206678-0002-00WOSIX Homeobox 1 (SIX1). In one embodiment, the fragment or variant of CTAATTA, CTCATTA, TTAATTA, CAAATTA, CCAATTA, TTCATTA, CACATTA, CCCATTA, CTAATTG, or TAAATTA retains the ability to bind to SIX. SIX is expressed in satellite cells, therefore in one embodiment, the plasmid comprising CTAATTA, CTCATTA, TTAATTA, CAAATTA, CCAATTA, TTCATTA, CACATTA, CCCATTA, CTAATTG, or TAAATTA is administered to a subject for expression in satellite cells. In one embodiment, the plasmid further comprises a coding sequence encoding a protein for expression in satellite cells, a nucleic acid molecule (e.g., mRNA, siRNA, miRNA, shRNA) for delivery to satellite cells, or a combination thereof. In one embodiment, the protein or nucleic acid molecule is a therapeutic agent for the treatment of a disease or disorder.PTS Clusters

[0099] In some embodiments, the plasmid is transported to the cell nucleus as a result of the native NLS on a transcription factor that recognizes and binds to the DTS. To increase nuclear localization of the plasmid, two or more DTS may be included on the plasmid in tandem or separated by a linker sequence, forming a DTS cluster. In one embodiment, two or more DTS that target the same TF are included on the plasmid. In one embodiment, two or more DTS that target different TFs are included on the plasmid.

[0100] In some embodiments, the plasmid comprises a single DTS cluster. For example, in one embodiment the plasmid includes a DTS cluster upstream of the promoter.

[0101] In some embodiments, the plasmid comprises two or more DTS clusters. In one embodiment, the two or more DTS are positioned flanking the coding region of the plasmid. For example, in one embodiment the plasmid includes a first DTS cluster upstream of the promoter and a second DTS cluster downstream of the coding sequence following the 3 ’UTS and polyA tail.

[0102] In one embodiment, the plasmid is a bidirectional expression plasmid and comprises first DTS cluster upstream of the promoter and a second DTS cluster downstream of the coding sequence following the 3 ’UTS and polyA tail in the sense direction and a third DTS cluster downstream of the coding sequence following the 3 ’UTS and polyA tail in the antisense direction.206678-0002-00WO

[0103] In one embodiment, the DTS cluster comprises two or more of a NFIA-DTS, HNFla-DTS, FOXI1-DTS, TP63-DTS, TFCP2L1-DTS, ASCL3-DTS, FOXJ1-DTS, RFX3-DTS, SPDEF-DTS, FOXQ1-DTS, FOXA2-DTS, and an HNFla-DTS. In one embodiment, the plasmid comprising a DTS cluster comprising two or more of a NFIA-DTS, HNFla-DTS, FOXI1-DTS, TP63-DTS, TFCP2L1-DTS, ASCL3-DTS, FOXJ1-DTS, RFX3-DTS, SPDEF-DTS, FOXQ1-DTS, FOXA2-DTS, and an HNFla-DTS is administered for lung-specific expression of an encoded agent. In one embodiment, the DTS cluster comprises each of a FOXI1-DTS, TFCP2L1-DTS, ASCL3-DTS, NFIA-DTS, SPDEF-DTS and FOXA2-DTS. In one embodiment, the DTS cluster comprises each of a FOXI1-DTS, ASCL3-DTS, NFIA-DTS, SPDEF-DTS and FOXA2-DTS. In one embodiment, the DTS cluster comprises each of an NFIA-DTS, HNF1α-DTS, FOXI1-DTS, TP63-DTS, FOXJ1-DTS, RFX3-DTS, SPDEF-DTS, FOXQ1-DTS, FOXA2-DTS, and HNFla-DTS.

[0104] In one embodiment, the DTS cluster comprises two or more of a MEF2A-DTS, MYOG-DTS, MRF4-DTS, TEAD1-DTS, PAX7-DTS, and SIX1-DTS. In one embodiment, the plasmid comprising a DTS cluster comprising two or more of a MEF2A-DTS, MYOG-DTS, MRF4-DTS, TEAD1-DTS, PAX7-DTS, and SIX1-DTS is administered for muscle-specific expression of an encoded agent.

[0105] In some embodiments, two or more DTS in a DTS cluster are separated by a spacer sequence. Exemplary spacers that can be included within the DTS cluster include, but are not limited to, SEQ ID NO: 99, SEQ ID NO: 100, SEQ ID NO: 101, SEQ ID NO: 102, SEQ ID NO: 103, SEQ ID NO: 104, SEQ ID NO: 105, SEQ ID NO: 106, SEQ ID NO: 107, SEQ ID NO: 108, SEQ ID NO: 109, SEQ ID NO: 110, SEQ ID NO: 111, SEQ ID NO: 112, SEQ ID NO:113, SEQ ID NO:114, SEQ ID NO:115 or SEQ ID NO:116.

[0106] Exemplary DTS cluster sequences for inclusion on a plasmid for tunable and regulated lung expression include, but are not limited to, SEQ ID NO: 121, SEQ ID NO: 122, and SEQ ID NO: 125. In one embodiment, the DTS cluster comprises a variant of SEQ ID NO: 121, SEQ ID NO: 122, or SEQ ID NO: 125, in which one or more DTS is varied for a different DTS of the same length, thus maintaining the spacing between the DTS within the DTS cluster.

[0107] An exemplary DTS cluster sequence for inclusion on a plasmid for tunable and regulated muscle expression includes SEQ ID NO: 128. In one embodiment, the DTS cluster206678-0002-00WOcomprises a variant of SEQ ID NO: 128, in which one or more DTS is varied for a different DTS of the same length, thus maintaining the spacing between the DTS within the DTS cluster.

[0108] In one embodiment, a DTS Cluster upstream of a promoter sequence is linked to the promoter by a spacer of about 50 nt. An exemplary spacer for positioning the DTS upstream of the promoter comprises SEQ ID NO: 131. In one embodiment, a DTS Cluster downstream of a coding sequence, following the 3’UTR and polyA tail, is linked to the 3’ end of the polyA tail by a spacer of about 50 nt. An exemplary spacer for positioning the DTS downstream of the coding sequence, following the 3’UTR and polyA tail, comprises SEQ ID NO: 132.Repressor Circuit

[0109] In one embodiment, the plasmid comprises at least one transcriptional repressor binding motif which recruits a transcriptional repressor specific for regulating a transcript generated from transcription of the coding sequence of the plasmid, which serves as a repressor circuit for increased transcriptional control and modulation of transgene output. In one embodiment, the plasmid comprises a combination of two or more transcriptional repressor binding motif. In one embodiment, at least one more transcriptional repressor binding motif is included between a DTS cluster and a promoter sequence.

[0110] For example, in one embodiment, the plasmid comprises a coding sequence for cystic fibrosis transmembrane conductance regulator (CFTR), and the plasmid further comprises a binding site for EHF, IRF2, or a combination of EHF and IRF2 to modulate transgene CFTR output. Exemplary transcriptional repressor binding motifs for binding to IRF2 include, but are not limited to, AANNGAAA, AAGTGAAA, AGGTGAAA, GANNGAAA, TANNGAAA, and AGTCGAAA. Exemplary transcriptional repressor binding motifs for binding to EHF include, but are not limited to, GGAA, GGAT, GGGA, ACCGGAAGT, TGGAAA, TGGAAAT, and CGGAAGT.Gene Cassette

[0111] In one embodiment, the plasmid comprises a gene cassette comprising (in 5’ to 3’ direction):a) Promoter-5’ UTR-Kozak Sequence- Open Reading Frame (ORF)-Polyadenylation Signal;b) Promoter-5’ UTR-Intron-Kozak Sequence- Open Reading Frame (ORF)-Polyadenylation Signal;c) Promoter-5’ UTR-Telomeric Repeat Motif-Intron-Kozak Sequence- Open Reading Frame (ORF)-Polyadenylation Signal;d) Promoter-5’ UTR-Intron with Telomeric Repeat Motif-Kozak Sequence-Open Reading Frame (ORF)-Polyadenylation Signal;e) Promoter-5’ UTR-Telomeric Repeat Motif-Intron-Kozak Sequence-Open Reading Frame (ORF)-Polyadenylation Signal-Post-Poly(A) Telomeric Motif; orf) Promoter-5’ UTR-Intron with Telomeric Repeat Motif-Kozak Sequence-Open Reading Frame (ORF)-Polyadenylation Signal-Post-Poly(A) Telomeric Motif.Promoter

[0112] In one embodiment, the plasmid comprises an optimized promoter that includes at least one of a BREu motif, a TATA box, an initiator (Inr), a downstream promoter element (DPE), and a Motif Ten Element (MTE), or any combination thereof. The BREu, TATA box, Inr, MTE, and DPE are designed to recruit general transcription factors (GTFs) rather than sequence-specific regulatory TFs. These short motifs are recognized by components of the basal transcription machinery (e.g. TFIIB binds the BRE, TBP / TFIID binds TATA and Inr, TFIID subunits recognize MTE / DPE). Any promoter comprising any subset or combination of the described elements and sequences, or any derivative thereof that retain the described functional properties, is intended to fall within the scope of this invention. This includes variants that incorporate these core element sequences (or functionally equivalent sequences) with the same spatial arrangement to achieve enhanced transcriptional performance. Exemplary elements that can be included in the promoter of the invention include, but are not limited to, those described below.BREu motif

[0113] In one embodiment, the promoter comprises an upstream TFIIB recognition element is included to enhance TFIIB interaction. Exemplary BREu motifs that can be included in the promoter include, but are not limited to, GGGCGCC, GCACGCC, or GCGCGCC. In one embodiment the BREu motif is located at a position of -38 to -32 nucleotides upstream of the transcription start site.BREu-TATA Spacer

[0114] In some embodiments, the promoter comprises a spacer between the BREu motif and one or more downstream motif. In one embodiment, the promoter comprises a one base spacer between the BREu sequence and the downstream TATA-like motif. In one embodiment, the promoter comprises a single adenine (A) spacer between the BREu sequence and the downstream TATA-like motif.TATA box

[0115] In one embodiment, the promoter comprises a TATA box element which helps to position the transcription start site properly. In one embodiment, the TATA box element conforms to the consensus TATAWAAR (where W = A / T and R = A / G). Exemplary TATA box elements that can be included in the promoter include, but are not limited to, TATATAA, TATAAAAG, TATATAAG, TATAAA, AGGTCTATATAAG (SEQ ID NO:264) and GTACTTATATAAG (SEQ ID NO:265). In one embodiment the TATA box element is located at a position of -30 to -25 / -24 nucleotides upstream of the transcription start site.TATA-Initiator Spacer

[0116] In some embodiments, the promoter comprises a spacer between the TATA box element and one or more downstream motif. In one embodiment, the promoter comprises a linker of 21-22 nucleotides between the TATA box element and the downstream initiator. Exemplary linker sequences that can be included between the TATA box element and the downstream initiator include, but are not limited to, AGTCGGCAGTCGGATCTCGCGA (SEQ ID NO: 133), AGTCGGCAGTCGGATCTCGC (SEQ ID NO: 134), AGTCGGCAGTCGGATCTCGCG (SEQ ID NO: 135), CAGACGTCGCATCGATCTACA (SEQ ID NO: 136), ACTCACGTCAGCTAGTCGCACG (SEQ ID NO: 137), CAGATTTCGCATCGATCTACA (SEQ ID NO: 138), AGTCGACGTCGTAGTCAGCTA (SEQ ID NO:139), CAGAGCTCGTTTAGTGAACC (SEQ ID NO:266) and GGGGTGGGGGCGCGTTCGTC (SEQ ID NO:267).Initiator

[0117] In one embodiment, the promoter comprises an initiator at the transcription start site which maximizes recruitment of TFIID and RNA Polymerase II. Exemplary initiator sequences that can be included in the promoter include, but are not limited to, TCAGTT, TCATTC, TCAGTCT, TCATATC, TCAGTTCC, GCAGTT, and CCACTT. In one embodimentthe initiator is located at a position of -2 to +4 / +5Z+6 relative to the transcription start site (TSS), with the italic A in indicating the +1 TSS.Initiator-Motif Ten Element (MTE) Spacer

[0118] In some embodiments, the promoter comprises a spacer between the initiator element and one or more downstream motif. In one embodiment, the promoter comprises a linker of 12-13 nucleotides between the initiator and the downstream motif ten element (MTE).Exemplary linker sequences that can be included between the initiator and the downstream MTE include, but are not limited to, TCACACGACATA (SEQ ID NO: 140), ATCACACGACATA (SEQ ID NO: 141), AGTCAGTCAGTCA (SEQ ID NO: 142), ATCACACGACATC (SEQ ID NO:143), GCCTGGAGACC (SEQ ID NO:268), CGATCGAACAC (SEQ ID NO:269), and TTTTTCAACAC (SEQ ID NO:270).Motif Ten Element (MTE)

[0119] In one embodiment, the promoter comprises an MTE downstream of the transcription start site which contributes to elevated basal transcription. In one embodiment, the MTE conforms to the consensus CSARCSSAACGS (SEQ ID NO: 144; S = G or C). In one embodiment, the MTE comprises two tandem “AACGG” repeat motifs. Exemplary MTE sequences that can be included in the promoter include, but are not limited to, CGAACGGAACGG (SEQ ID NO: 145), CGAACGGAACAG (SEQ ID NO: 146), CGAACGGAAC (SEQ ID NO: 147), CCAGCCGAAC (SEQ ID NO: 148), CGAGCCGAAC (SEQ ID NO: 149), TCGAGCCGAGT (SEQ ID NO:271) and TCGAGCCGAGC (SEQ ID NO:272). In one embodiment the initiator is located at a position of +18 to +29 / +27 nucleotides downstream of the transcription start site.Initiator - downstream promoter element (DPE) Spacer

[0120] In one embodiment, the promoter comprises a linker of 22-23 nucleotides between the initiator and the DPE. In one embodiment, the linker is selected to position the DPE exactly at +28 to +33 downstream of the transcriptional start site, aligning it with the consensus location for downstream promoter elements and ensuring it overlaps appropriately with the MTE such that the MTE and DPE elements are arranged in the promoter so that they overlap at positions +28 and +29 (sharing the dinucleotide “GG”). Exemplary linker sequences that can be included between the initiator and the downstream DPE to appropriately position the DPE, include but are not limited to, ACTGATCGGCTGATCGAATCGA (SEQ ID NO: 150),AACTGATCGGCTGATCGAATCGA (SEQ ID NO: 151), GTGTACTCAGCCTAGTCGATCGT (SEQ ID NO: 152), CGTACTCAGCCTAGTCGATCGT (SEQ ID NO: 153), ACTGATCGGCTGATCGAATCGT (SEQ ID NO: 154), GACTTGCGTCGTACGGTTACGC (SEQ ID NO:155), GACTTGCGTCGTATTGTTACAG (SEQ ID NO: 156) and ACTTATCGGCTTATCGAATCGA (SEQ ID NO: 157).Downstream Promoter Element (DPE)

[0121] In one embodiment, the promoter comprises a DPE downstream of the transcription start site to enhance promoter activity. In one embodiment, the DPE a cytosine at position (+31). Exemplary DPE sequences that can be included in the promoter include, but are not limited to, GGACCT, ACCT, AGTCGC, GGACTGG, GGTTTC, AGACGTG and AGACGT. In one embodiment the initiator is located at a position of +28 to +32 / +33 nucleotides downstream of the transcription start site.Core Promoter Sequences

[0122] In one embodiment, the promoter comprises a BREu comprising GCGCGCC at position -38 to -32, a BREu-TATA spacer comprising A at position -31, a TATA box element comprising TATAAA at position -30 to -25, a TATA-Initiator spacer comprising SEQ ID NO: 133 at position -24 to -3, an initiator comprising TCAGTT at -2 to +4 relative to the transcription start site indicated in bold italics, an initiator-DPE spacer comprising SEQ ID NO: 151 at position +5 to +27, and a DPE comprising AGACGT at position +28 to +33. In one embodiment, the promoter comprises SEQ ID NO: 158.

[0123] In one embodiment, the promoter comprises a BREu comprising GCGCGCC at position -38 to -32, a BREu-TATA spacer comprising A at position -31, a TATA box element comprising TATAAA at position -30 to -25, a TATA-Initiator spacer comprising SEQ ID NO: 133 at position -24 to -3, an initiator comprising TCAGTT at -2 to +4 relative to the transcription start site indicated in bold italics, an initiator-MTE spacer comprising SEQ ID NO:141 at position +5 to +17, an MTE comprising SEQ ID NO:145 at position +18 to +29 and a DPE comprising GGACCT at position +28 to +33. In one embodiment, the promoter comprises SEQ ID NO: 159.

[0124] In one embodiment, the promoter comprises a BREu comprising GCGCGCC at position -38 to -32, a BREu-TATA spacer comprising A at position -31, a TATA box element comprising TATAAA at position -30 to -25, a TATA-Initiator spacer comprising SEQ ID206678-0002-00WONO: 133 at position -24 to -3, an initiator comprising TC4GTCT at -2 to +5 relative to the transcription start site indicated in bold italics, an initiator-DPE spacer comprising SEQ ID NO: 150 at position +6 to +27, and a DPE comprising AGACGT at position +28 to +33. In one embodiment, the promoter comprises SEQ ID NO: 160.

[0125] In one embodiment, the promoter comprises a BREu comprising GCGCGCC at position -38 to -32, a BREu-TATA spacer comprising A at position -31, a TATA box element comprising TATAAA at position -30 to -25, a TATA-Initiator spacer comprising SEQ ID NO: 133 at position -24 to -3, an initiator comprising TCAGTCT at -2 to +5 relative to the transcription start site indicated in bold italics, an initiator-MTE spacer comprising SEQ ID NO:140 at position +6 to +17, an MTE comprising SEQ ID NO:145 at position +18 to +29 and a DPE comprising GGACCT at position +28 to +33. In one embodiment, the promoter comprises SEQ ID NO: 161.

[0126] In one embodiment, the promoter comprises a BREu comprising GCGCGCC at position -38 to -32, a BREu-TATA spacer comprising A at position -31, a TATA box element comprising TATATAA at position -30 to -24, a TATA-Initiator spacer comprising SEQ ID NO: 135 at position -23 to -3, an initiator comprising TCAGTT at -2 to +4 relative to the transcription start site indicated in bold italics, an initiator-DPE spacer comprising SEQ ID NO: 151 at position +5 to +27, and a DPE comprising AGACGT at position +28 to +33. In one embodiment, the promoter comprises SEQ ID NO: 162.

[0127] In one embodiment, the promoter comprises a BREu comprising GCGCGCC at position -38 to -32, a BREu-TATA spacer comprising A at position -31, a TATA box element comprising TATATAA at position -30 to -24, a TATA-Initiator spacer comprising SEQ ID NO: 135 at position -23 to -3, an initiator comprising TCAGTT at -2 to +4 relative to the transcription start site indicated in bold italics, an initiator-MTE spacer comprising SEQ ID NO: 141 at position +5 to +17, an MTE comprising SEQ ID NO: 145 at position +18 to +29 and a DPE comprising GGACCT at position +28 to +33. In one embodiment, the promoter comprises SEQ ID NO: 163.

[0128] In one embodiment, the promoter comprises a BREu comprising GCGCGCC at position -38 to -32, a BREu-TATA spacer comprising A at position -31, a TATA box element comprising TATATAA at position -30 to -24, a TATA-Initiator spacer comprising SEQ ID NO:135 at position -23 to -3, an initiator comprising TC4GTCT at -2 to +5 relative to the206678-0002-00WOtranscription start site indicated in bold italics, an initiator-DPE spacer comprising SEQ ID NO: 150 at position +6 to +27, and a DPE comprising AGACGT at position +28 to +33. In one embodiment, the promoter comprises SEQ ID NO: 164.

[0129] In one embodiment, the promoter comprises a BREu comprising GCGCGCC at position -38 to -32, a BREu-TATA spacer comprising A at position -31, a TATA box element comprising TATATAA at position -30 to -24, a TATA-Initiator spacer comprising SEQ ID NO: 135 at position -23 to -3, an initiator comprising TC4GTCT at -2 to +5 relative to the transcription start site indicated in bold italics, an initiator-MTE spacer comprising SEQ ID NO:140 at position +6 to +17, an MTE comprising SEQ ID NO:145 at position +18 to +29 and a DPE comprising GGACCT at position +28 to +33. In one embodiment, the promoter comprises SEQ ID NO: 165.

[0130] In one embodiment, the promoter comprises a BREu comprising GCGCGCC at position -38 to -32, a BREu-TATA spacer comprising A at position -31, a TATA box element comprising TATATAAG at position -30 to -23, a TATA-Initiator spacer comprising SEQ ID NO: 134 at position -22 to -3, an initiator comprising TCAGTT at -2 to +4 relative to the transcription start site indicated in bold italics, an initiator-DPE spacer comprising SEQ ID NO: 151 at position +5 to +27, and a DPE comprising AGACGT at position +28 to +33. In one embodiment, the promoter comprises SEQ ID NO: 166.

[0131] In one embodiment, the promoter comprises a BREu comprising GCGCGCC at position -38 to -32, a BREu-TATA spacer comprising A at position -31, a TATA box element comprising TATATAAG at position -30 to -23, a TATA-Initiator spacer comprising SEQ ID NO: 134 at position -22 to -3, an initiator comprising TG4GTT at -2 to +4 relative to the transcription start site indicated in bold italics, an initiator-MTE spacer comprising SEQ ID NO: 141 at position +5 to +17, an MTE comprising SEQ ID NO: 145 at position +18 to +29 and a DPE comprising GGACCT at position +28 to +33. In one embodiment, the promoter comprises SEQ ID NO: 167.

[0132] In one embodiment, the promoter comprises a BREu comprising GCGCGCC at position -38 to -32, a BREu-TATA spacer comprising A at position -31, a TATA box element comprising TATATAAG at position -30 to -23, a TATA-Initiator spacer comprising SEQ ID NO: 134 at position -22 to -3, an initiator comprising TG4GTCT at -2 to +5 relative to the transcription start site indicated in bold italics, an initiator-DPE spacer comprising SEQ ID206678-0002-00WONO:150 at position +6 to +27, and a DPE comprising AGACGT at position +28 to +33. In one embodiment, the promoter comprises SEQ ID NO: 168.

[0133] In one embodiment, the promoter comprises a BREu comprising GCGCGCC at position -38 to -32, a BREu-TATA spacer comprising A at position -31, a TATA box element comprising TATATAAG at position -30 to -23, a TATA-Initiator spacer comprising SEQ ID NO: 134 at position -22 to -3, an initiator comprising TC4GTCT at -2 to +5 relative to the transcription start site indicated in bold italics, an initiator-MTE spacer comprising SEQ ID NO:140 at position +6 to +17, an MTE comprising SEQ ID NO:145 at position +18 to +29 and a DPE comprising GGACCT at position +28 to +33. In one embodiment, the promoter comprises SEQ ID NO: 169.

[0134] In one embodiment, the promoter comprises a BREu comprising GCACGCC at position -38 to -32, a BREu-TATA spacer comprising A at position -31, a TATA box element comprising TATAAA at position -30 to -25, a TATA-Initiator spacer comprising SEQ ID NO: 133 at position -24 to -3, an initiator comprising TG4GTT at -2 to +4 relative to the transcription start site indicated in bold italics, an initiator-DPE spacer comprising SEQ ID NO: 151 at position +5 to +27, and a DPE comprising AGACGT at position +28 to +33. In one embodiment, the promoter comprises SEQ ID NO: 170.

[0135] In one embodiment, the promoter comprises a BREu comprising GCACGCC at position -38 to -32, a BREu-TATA spacer comprising A at position -31, a TATA box element comprising TATAAA at position -30 to -25, a TATA-Initiator spacer comprising SEQ ID NO: 133 at position -24 to -3, an initiator comprising TCAGTT at -2 to +4 relative to the transcription start site indicated in bold italics, an initiator-MTE spacer comprising SEQ ID NO:141 at position +5 to +17, an MTE comprising SEQ ID NO:145 at position +18 to +29 and a DPE comprising GGACCT at position +28 to +33. In one embodiment, the promoter comprises SEQ ID NO:171.

[0136] In one embodiment, the promoter comprises a BREu comprising GCACGCC at position -38 to -32, a BREu-TATA spacer comprising A at position -31, a TATA box element comprising TATAAA at position -30 to -25, a TATA-Initiator spacer comprising SEQ ID NO: 133 at position -24 to -3, an initiator comprising TC4GTCT at -2 to +5 relative to the transcription start site indicated in bold italics, an initiator-DPE spacer comprising SEQ ID206678-0002-00WONO:150 at position +6 to +27, and a DPE comprising AGACGT at position +28 to +33. In one embodiment, the promoter comprises SEQ ID NO: 172.

[0137] In one embodiment, the promoter comprises a BREu comprising GCACGCC at position -38 to -32, a BREu-TATA spacer comprising A at position -31, a TATA box element comprising TATAAA at position -30 to -25, a TATA-Initiator spacer comprising SEQ ID NO: 133 at position -24 to -3, an initiator comprising TC4GTCT at -2 to +5 relative to the transcription start site indicated in bold italics, an initiator-MTE spacer comprising SEQ ID NO:140 at position +6 to +17, an MTE comprising SEQ ID NO:145 at position +18 to +29 and a DPE comprising GGACCT at position +28 to +33. In one embodiment, the promoter comprises SEQ ID NO: 173.

[0138] In one embodiment, the promoter comprises a BREu comprising GCACGCC at position -38 to -32, a BREu-TATA spacer comprising A at position -31, a TATA box element comprising TATATAA at position -30 to -24, a TATA-Initiator spacer comprising SEQ ID NO: 135 at position -23 to -3, an initiator comprising TG4GTT at -2 to +4 relative to the transcription start site indicated in bold italics, an initiator-DPE spacer comprising SEQ ID NO: 151 at position +5 to +27, and a DPE comprising AGACGT at position +28 to +33. In one embodiment, the promoter comprises SEQ ID NO: 174.

[0139] In one embodiment, the promoter comprises a BREu comprising GCACGCC at position -38 to -32, a BREu-TATA spacer comprising A at position -31, a TATA box element comprising TATATAA at position -30 to -24, a TATA-Initiator spacer comprising SEQ ID NO: 135 at position -23 to -3, an initiator comprising TCAGTT at -2 to +4 relative to the transcription start site indicated in bold italics, an initiator-MTE spacer comprising SEQ ID NO:141 at position +5 to +17, an MTE comprising SEQ ID NO:145 at position +18 to +29 and a DPE comprising GGACCT at position +28 to +33. In one embodiment, the promoter comprises SEQ ID NO: 175.

[0140] In one embodiment, the promoter comprises a BREu comprising GCACGCC at position -38 to -32, a BREu-TATA spacer comprising A at position -31, a TATA box element comprising TATATAA at position -30 to -24, a TATA-Initiator spacer comprising SEQ ID NO: 135 at position -23 to -3, an initiator comprising TC4GTCT at -2 to +5 relative to the transcription start site indicated in bold italics, an initiator-DPE spacer comprising SEQ ID206678-0002-00WONO:150 at position +6 to +27, and a DPE comprising AGACGT at position +28 to +33. In one embodiment, the promoter comprises SEQ ID NO: 176.

[0141] In one embodiment, the promoter comprises a BREu comprising GCACGCC at position -38 to -32, a BREu-TATA spacer comprising A at position -31, a TATA box element comprising TATATAA at position -30 to -24, a TATA-Initiator spacer comprising SEQ ID NO: 135 at position -23 to -3, an initiator comprising TC4GTCT at -2 to +5 relative to the transcription start site indicated in bold italics, an initiator-MTE spacer comprising SEQ ID NO:140 at position +6 to +17, an MTE comprising SEQ ID NO:145 at position +18 to +29 and a DPE comprising GGACCT at position +28 to +33. In one embodiment, the promoter comprises SEQ ID NO: 177.

[0142] In one embodiment, the promoter comprises a BREu comprising GCACGCC at position -38 to -32, a BREu-TATA spacer comprising A at position -31, a TATA box element comprising TATATAAG at position -30 to -23, a TATA-Initiator spacer comprising SEQ ID NO: 134 at position -22 to -3, an initiator comprising TG4GTT at -2 to +4 relative to the transcription start site indicated in bold italics, an initiator-DPE spacer comprising SEQ ID NO: 151 at position +5 to +27, and a DPE comprising AGACGT at position +28 to +33. In one embodiment, the promoter comprises SEQ ID NO: 178.

[0143] In one embodiment, the promoter comprises a BREu comprising GCACGCC at position -38 to -32, a BREu-TATA spacer comprising A at position -31, a TATA box element comprising TATATAAG at position -30 to -23, a TATA-Initiator spacer comprising SEQ ID NO: 134 at position -22 to -3, an initiator comprising TCAGTT at -2 to +4 relative to the transcription start site indicated in bold italics, an initiator-MTE spacer comprising SEQ ID NO:141 at position +5 to +17, an MTE comprising SEQ ID NO:145 at position +18 to +29 and a DPE comprising GGACCT at position +28 to +33. In one embodiment, the promoter comprises SEQ ID NO: 179.

[0144] In one embodiment, the promoter comprises a BREu comprising GCACGCC at position -38 to -32, a BREu-TATA spacer comprising A at position -31, a TATA box element comprising TATATAAG at position -30 to -23, a TATA-Initiator spacer comprising SEQ ID NO: 134 at position -22 to -3, an initiator comprising TC4GTCT at -2 to +5 relative to the transcription start site indicated in bold italics, an initiator-DPE spacer comprising SEQ ID206678-0002-00WONO:150 at position +6 to +27, and a DPE comprising AGACGT at position +28 to +33. In one embodiment, the promoter comprises SEQ ID NO: 180.

[0145] In one embodiment, the promoter comprises a BREu comprising GCACGCC at position -38 to -32, a BREu-TATA spacer comprising A at position -31, a TATA box element comprising TATATAAG at position -30 to -23, a TATA-Initiator spacer comprising SEQ ID NO: 134 at position -22 to -3, an initiator comprising TC4GTCT at -2 to +5 relative to the transcription start site indicated in bold italics, an initiator-MTE spacer comprising SEQ ID NO:140 at position +6 to +17, an MTE comprising SEQ ID NO:145 at position +18 to +29 and a DPE comprising GGACCT at position +28 to +33. In one embodiment, the promoter comprises SEQ IDNO:181.

[0146] Exemplary promoter sequences that can be included on the plasmid include, but are not limited to SEQ ID NO:158, SEQ ID NO:159, SEQ ID NO:160, SEQ ID NO:161, SEQ ID NO: 162, SEQ ID NO: 163, SEQ ID NO: 164, SEQ ID NO: 165, SEQ ID NO: 166, SEQ ID NO: 167, SEQ ID NO: 168, SEQ ID NO: 169, SEQ ID NO: 170, SEQ ID NO: 171, SEQ ID NO: 172, SEQ ID NO: 173, SEQ ID NO: 174, SEQ ID NO: 175, SEQ ID NO: 176, SEQ ID NO:177, SEQ ID NO:178, SEQ ID NO:179, SEQ ID NO:180, SEQ ID NO:181, SEQ ID NO: 182, SEQ ID NO: 183, SEQ ID NO: 184, SEQ ID NO: 185, SEQ ID NO: 186, SEQ ID NO:187, SEQ ID NO:188, SEQ ID NO:189, SEQ ID NO:190, SEQ ID NO:191, SEQ ID NO: 192, SEQ ID NO: 193, SEQ ID NO: 194, SEQ ID NO: 195, SEQ ID NO: 196, SEQ ID NO:197, SEQ ID NO:198, SEQ ID NO:199, SEQ ID NO:200, SEQ ID NO:201, SEQ ID NO:202, SEQ ID NO:203, SEQ ID NO:204, SEQ ID NO:205, SEQ ID NO:206, SEQ ID NO:207, SEQ ID NO:208, SEQ ID NO:209, SEQ ID NO:210, SEQ ID NO:211, SEQ ID NO:212, SEQ ID NO:213, SEQ ID NO:214, SEQ ID NO:215, SEQ ID NO:279, SEQ ID NO:280, SEQ ID NO:281, SEQ ID NO:282, SEQ ID NO:283, or SEQ ID NO:284.DTS-Promoter Spacer

[0147] In one embodiment, the plasmid comprises at least one spacer upstream of the promoter. For example, in one embodiment, the plasmid comprises a spacer between a DTS or DTS cluster and the promoter. Exemplary upstream spacers that can be included on the plasmid include, but are not limited to, SEQ ID NO:273, SEQ ID NO:274, SEQ ID NO:275 and SEQ ID NO:276.DPE - 5’UTR Spacer

[0148] In one embodiment, the plasmid comprises at least one spacer following the DPE. For example, in one embodiment, the plasmid comprises a spacer between the DPE and the 5’UTR or telomeric repeat motif. Exemplary downstream spacers that can be included on the plasmid include, but are not limited to, SEQ ID NO:277 and SEQ ID NO:278.Telomeric Repeats

[0149] In one embodiment, the plasmid comprises at least one telomeric repeat which suppresses immune activation. In one embodiment, the plasmid comprises a telomeric repeat motif comprising at least 2, 3, 4, 5, 6 or more than 6 telomeric repeats. Exemplary telomeric motifs that can be included on the plasmid of the invention include, but are not limited to, SEQ ID NO:216, SEQ ID NO:217, SEQ ID NO:218, SEQ ID NO:219, SEQ ID NO 220, SEQ ID NO:221, SEQ ID NO:222, SEQ ID NO:223, SEQ ID NO:224, SEQ ID NO:225, SEQ ID NO:226, SEQ ID NO:227, and SEQ ID NO:228.

[0150] In one embodiment, at least one telomeric repeat motif is an upstream telomeric repeat included upstream of the promoter, within the 5’UTR, or immediately after the 5’UTR. In one embodiment, the upstream telomeric motif comprises SEQ ID NO:216, SEQ ID NO:217, SEQ ID NO:218, SEQ ID NO:219, SEQ ID NO:220, SEQ ID NO:221, SEQ ID NO:222, SEQ ID NO:223, SEQ ID NO:224, SEQ ID NO:225, or SEQ ID NO:226 or a fragment or variant thereof.

[0151] In one embodiment the plasmid comprises a post-polyA telomeric motif immediately downstream of the poly(A) signal (and its cleavage site). In one embodiment, the post-polyA telomeric motif is outside the transcription unit, meaning when RNA polymerase reaches the poly(A) signal, it will terminate and not transcribe beyond it. The telomeric motif inserted here is therefore not transcribed, but servers as an immune-suppressive module at the end of the expression cassette. In one embodiment, the post-polyA telomeric motif comprises SEQ ID NO:227, or SEQ ID NO:228 or a fragment or variant thereof.

[0152] In one embodiment, the plasmid comprises a combination of an upstream telomeric repeat motif and a post-polyA telomeric motif. Therefore in one embodiment, the plasmid comprises a combination of an upstream telomeric repeat motif comprising SEQ IDNO:216, SEQ IDNO:217, SEQ ID NO:218, SEQ ID NO:219, SEQ ID NO:220, SEQ ID NO:221, SEQ ID NO:222, SEQ ID NO:223, SEQ ID NO:224, SEQ ID NO:225, or SEQ ID NO:226, or a fragment or variant thereof, and a post-polyA telomeric motif comprising SEQ ID NO:227, or SEQ ID NO:228, or a fragment or variant thereofIntron

[0153] In one embodiment, the plasmid comprises an intron sequence to enhance transgene expression. Exemplary intron sequences that can be included on the plasmid of the invention include, but are not limited to, SEQ ID NO:229, SEQ ID NO:230 and SEQ ID NO:231. In one embodiment, the telomeric motif is upstream of the intron within the 5'UTR.

[0154] In one embodiment, the intron comprises at least one telomeric repeat motif within the intron sequence. Exemplary intron sequences including a telomeric repeat motif that can be included on the plasmid of the invention include, but are not limited to, SEQ ID NO:232, SEQ ID NO:233, SEQ ID NO:234, SEQ ID NO:235, SEQ ID NO:236, SEQ ID NO:237, SEQ ID NO:238, and SEQ ID NO:239.5’ UTR

[0155] In one embodiment, gene expression cassette comprises a 5’UTR sequence between the core promoter sequence and the downstream coding sequence. Exemplary 5’ UTR sequences that can be included in the gene expression cassette include, but are not limited to, SEQ ID NO:240, SEQ ID NO:241, SEQ ID NO:242, SEQ ID NO:243, SEQ ID NO:244, or SEQ ID NO:245.Kozak

[0156] In one embodiment, gene expression cassette comprises a kozak sequence immediately upstream of the start codon. Exemplary kozak sequences that can be included on the plasmid include, but are not limited to, SEQ ID NO:246, and GCCACC.Codon Optimized CTFR Sequence

[0157] In one embodiment, the invention provides synthetic polynucleotides encoding human cystic fibrosis transmembrane conductance regulator (CFTR), designated SEQ ID206678-0002-00WONO:248 and SEQ ID NO:249, which have been codon-optimized for expression in human pulmonary epithelial cells. SEQ ID NO:248 and SEQ ID NO:249 differ from the native CFTR coding sequence by a plurality of synonymous nucleotide substitutions that collectively enhance translational efficiency and transgene stability, without altering the CFTR amino acid sequence (thus the encoded protein retains all wild-type functional domains and post-translational modification sites). In one aspect, the invention encompasses the full-length codon-optimized CFTR coding sequence (SEQ ID NO:248 and SEQ ID NO:249) as well as any variant or fragment that includes a subset of the synonymous codon changes of SEQ ID NO:248 and SEQ ID NO:249 and yields improved expression or stability. In one embodiment, the codon optimized nucleic acid molecule encoding CFTR comprises SEQ ID NO:248, or a fragment or variant thereof, encoding human CFTR (SEQ ID NO:247). In one embodiment, the codon optimized nucleic acid molecule encoding CFTR comprises SEQ ID NO:249, or a fragment or variant thereof, encoding human CFTR (SEQ ID NO:247).Codon Optimized GFP Sequence

[0158] In some embodiments, the plasmid comprises a codon optimized coding sequence encoding a reporter protein. In one embodiment, the reporter peptide comprises green fluorescent protein (GFP). In one embodiment, the codon optimized nucleic acid molecule encoding GFP encodes SEQ ID NO:250. In one embodiment, the codon optimized nucleic acid molecule encoding GFP comprises SEQ ID NO:251 or a fragment or variant thereof.

[0159] In some embodiments, the plasmid comprises a polycistronic coding sequence comprising a codon optimized GFP coding sequence linked to a codon optimized CFTR coding sequence. In one embodiment, the polycistronic coding sequence includes a protease cleavage sequence between the GFP coding sequence and the CFTR coding sequence. In one embodiment, the polycistronic coding sequence comprises a combination of SEQ ID NO:251 linked to SEQ ID NO:248 and further comprising a sequence encoding a protease cleavage site between SEQ ID NO:251 and SEQ ID NO:248. In one embodiment, the polycistronic coding sequence comprises a combination of SEQ ID NO:251 linked to SEQ ID NO:249 and further comprising a sequence encoding a protease cleavage site between SEQ ID NO:251 and SEQ ID NO:249.206678-0002-00WO3’ UTR

[0160] In one embodiment, gene expression cassette comprises a 3’ UTR sequence between the coding sequence and the polyA tail. Exemplary 3’ UTR sequences that can be included in the gene expression cassette include, but are not limited to, SEQ ID NO:252, SEQ ID NO:253, SEQ ID NO:254 or SEQ ID NO:255.Vertical PolyA Tail

[0161] In one embodiment, gene expression cassette comprises a polyA tail following the 3’ UTR. An exemplary polyA tail sequence that can be included in the gene expression cassette includes, but is not limited to, SEQ ID NO:256.Plasmid Sequences

[0162] In some embodiments, the composition comprises a plasmid comprising a combination of at least one DTS or DTS cluster linked to a gene expression cassette comprising a promoter operably linked to a coding sequence for expression is a target cell type or a subset of cells of a target cell type. In one embodiment, the plasmid comprises (in the 5’ to 3’ direction):a. DTS cluster-Promoter-5’ UTR-Kozak Sequence-Open Reading Frame (ORF)- Polyadenylation Signal;b. DTS cluster-Promoter-5’ UTR- Kozak Sequence-Open Reading Frame (ORF)- Polyadenylation Signal-DTS cluster;c. DTS cluster-Promoter-5’ UTR-Intron-Kozak Sequence-Open Reading Frame (ORF)- Polyadenylation Signal;d. DTS cluster-Promoter-5’ UTR- Intron-Kozak Sequence-Open Reading Frame (ORF)- Polyadenylation Signal-DTS cluster;e. DTS cluster -Promoter-5’ UTR-Telomeric Repeat Motif-Intron-Kozak Sequence- Open Reading Frame (ORF)-Polyadenylation Signal;f. DTS cluster -Promoter-5’ UTR-Telomeric Repeat Motif-Intron-Kozak Sequence- Open Reading Frame (ORF)-Polyadenylation Signal- DTS cluster;g. DTS cluster -Promoter-5’ UTR-Intron with Telomeric Repeat Motif-Kozak Sequence-Open Reading Frame (ORF)-Polyadenylation Signal;h. DTS cluster -Promoter-5’ UTR-Intron with Telomeric Repeat Motif-Kozak Sequence-Open Reading Frame (ORF)-Polyadenylation Signal- DTS cluster;i. DTS cluster -Promoter-5’ UTR-Telomeric Repeat Motif-Intron-Kozak Sequence-Open Reading Frame (ORF)-Polyadenylation Signal-Post Poly(A) Telomeric Motif;j. DTS cluster -Promoter-5’ UTR-Telomeric Repeat Motif-Intron-Kozak Sequence-Open Reading Frame (ORF)-Polyadenylation Signal-Post Poly(A) Telomeric Motif-DTS cluster;k. DTS cluster -Promoter-5’ UTR-Intron with Telomeric Repeat Motif-Kozak Sequence-Open Reading Frame (ORF)-Polyadenylation Signal-Post Poly(A) Telomeric Motif;l. DTS cluster -Promoter-5’ UTR-Intron with Telomeric Repeat Motif-Kozak Sequence-Open Reading Frame (ORF)-Polyadenylation Signal-Post Poly(A) Telomeric Motif- DTS cluster.

[0163] In one embodiment, the plasmid further comprises a transcriptional repressor binding motif (TR motif) between the DTS cluster and the promoter, for additional transcriptional control over the transcription product. Therefore in one embodiment, the plasmid comprises (in the 5’ to 3’ direction):a. DTS cluster-TR motif-Promoter-5’ UTR-Kozak Sequence-Open Reading Frame (ORF)- Polyadenylation Signal;b. DTS cluster- TR motif -Promoter-5’ UTR-Kozak Sequence-Open Reading Frame (ORF)- Polyadenylation Signal-DTS cluster;c. DTS cluster - TR motif -Promoter-5’ UTR-Telomeric Repeat Motif-Intron-Kozak Sequence- Open Reading Frame (ORF)-Polyadenylation Signal;d. DTS cluster - TR motif -Promoter-5’ UTR-Telomeric Repeat Motif-Intron-Kozak Sequence- Open Reading Frame (ORF)-Polyadenylation Signal- DTS cluster;e. DTS cluster - TR motif -Promoter-5’ UTR-Intron with Telomeric Repeat Motif-Kozak Sequence-Open Reading Frame (ORF)-Polyadenylation Signal;f. DTS cluster - TR motif -Promoter-5’ UTR-Intron with Telomeric Repeat Motif-Kozak Sequence-Open Reading Frame (ORF)-Polyadenylation Signal- DTS cluster;206678-0002-00WOg. DTS cluster - TR motif -Promoter-5’ UTR-Telomeric Repeat Motif-Intron-Kozak Sequence-Open Reading Frame (ORF)-Polyadenylation Signal-Post Poly(A) Telomeric Motif;h. DTS cluster - TR motif -Promoter-5’ UTR-Telomeric Repeat Motif-Intron-Kozak Sequence-Open Reading Frame (ORF)-Polyadenylation Signal-Post Poly(A) Telomeric Motif- DTS cluster;i. DTS cluster - TR motif -Promoter-5’ UTR-Intron with Telomeric Repeat Motif-Kozak Sequence-Open Reading Frame (ORF)-Polyadenylation Signal-Post Poly(A) Telomeric Motif;j. DTS cluster - TR motif -Promoter-5’ UTR-Intron with Telomeric Repeat Motif-Kozak Sequence-Open Reading Frame (ORF)-Polyadenylation Signal-Post Poly(A) Telomeric Motif- DTS cluster.

[0164] In some embodiments, the transcriptional control elements, including but not limited to, the DTS cluster(s), a TR motif, the telomeric repeat motif(s), and the promoter comprise spacers which position them in a manner such that each element performs its function for controlled transcription of the transcription product.

[0165] Exemplary plasmid sequences for controlled expression of CFTR in lung cells include, but are not limited to plasmids comprising SEQ ID NO:257, SEQ ID NO:258, SEQ ID NO:259, SEQ ID NO:260, SEQ ID NO:261, SEQ ID NO:262 and SEQ ID NO:263.Delivery Vehicle

[0166] In some embodiments, the composition comprises a delivery vehicle comprising the plasmid of the invention. In one embodiment, the delivery vehicle is a colloidal dispersion system, such as macromolecule complexes, nanocapsules, microspheres, beads, and lipid-based systems including oil-in-water emulsions, micelles, mixed micelles, and liposomes. An exemplary colloidal system for use as a delivery vehicle in vitro and in vivo is a liposome (e.g., an artificial membrane vesicle).

[0167] The use of lipid formulations is contemplated for the introduction of the at least one agent into a host cell (in vitro, ex vivo or in vivo). In another aspect, the at least one agent may be associated with a lipid. The at least one agent associated with a lipid may be encapsulated in the aqueous interior of a liposome, interspersed within the lipid bilayer of a206678-0002-00WOliposome, attached to a liposome via a linking molecule that is associated with both the liposome and the oligonucleotide, entrapped in a liposome, complexed with a liposome, dispersed in a solution containing a lipid, mixed with a lipid, combined with a lipid, contained as a suspension in a lipid, contained or complexed with a micelle, or otherwise associated with a lipid. Lipid, lipid / nucleic acid or lipid / expression vector associated compositions are not limited to any particular structure in solution. For example, they may be present in a bilayer structure, as micelles, or with a “collapsed” structure. They may also simply be interspersed in a solution, possibly forming aggregates that are not uniform in size or shape. Lipids are fatty substances which may be naturally occurring or synthetic lipids. For example, lipids include the fatty droplets that naturally occur in the cytoplasm as well as the class of compounds which contain long-chain aliphatic hydrocarbons and their derivatives, such as fatty acids, alcohols, amines, amino alcohols, and aldehydes.

[0168] Lipids suitable for use can be obtained from commercial sources. For example, dimyristyl phosphatidylcholine (“DMPC”) can be obtained from Sigma, St. Louis, MO; dicetyl phosphate (“DCP”) can be obtained from K & K Laboratories (Plainview, NY); cholesterol (“Choi”) can be obtained from Calbiochem-Behring; dimyristyl phosphatidylglycerol (“DMPG”) and other lipids may be obtained from Avanti Polar Lipids, Inc. (Birmingham, AL). Stock solutions of lipids in chloroform or chloroform / methanol can be stored at about -20°C.Chloroform is used as the only solvent since it is more readily evaporated than methanol.“Liposome” is a generic term encompassing a variety of single and multilamellar lipid vehicles formed by the generation of enclosed lipid bilayers or aggregates. Liposomes can be characterized as having vesicular structures with a phospholipid bilayer membrane and an inner aqueous medium. Multilamellar liposomes have multiple lipid layers separated by aqueous medium. They form spontaneously when phospholipids are suspended in an excess of aqueous solution. The lipid components undergo self-rearrangement before the formation of closed structures and entrap water and dissolved solutes between the lipid bilayers (Ghosh et al., 1991 Glycobiology 5: 505-10). However, compositions that have different structures in solution than the normal vesicular structure are also encompassed. For example, the lipids may assume a micellar structure or merely exist as nonuniform aggregates of lipid molecules. Also contemplated are lipofectamine-agent complexes.

[0169] In one embodiment, delivery of the at least one agent comprises any suitable delivery method, including exemplary delivery methods described elsewhere herein. In certain embodiments, delivery of the at least one agent to a subject comprises mixing the at least one agent with a transfection reagent prior to the step of contacting. In another embodiment, a method of the present invention further comprises administering the at least one agent together with the transfection reagent. In another embodiment, the transfection reagent is a cationic lipid reagent.

[0170] In another embodiment, the transfection reagent is a lipid-based transfection reagent. In another embodiment, the transfection reagent is a protein-based transfection reagent. In another embodiment, the transfection reagent is a polyethyleneimine based transfection reagent. In another embodiment, the transfection reagent is calcium phosphate. In another embodiment, the transfection reagent is Lipofectin®, Lipofectamine®, or TransIT®. In another embodiment, the transfection reagent is any other transfection reagent known in the art.

[0171] In another embodiment, the transfection reagent forms a liposome. Liposomes, in another embodiment, increase intracellular stability, increase uptake efficiency and improve biological activity. In another embodiment, liposomes are hollow spherical vesicles composed of lipids arranged in a similar fashion as those lipids which make up the cell membrane. In some embodiments, the liposomes comprise an internal aqueous space for entrapping water-soluble compounds. In another embodiment, liposomes can deliver the at least one agent to cells in an active form.

[0172] In one embodiment, the composition comprises a lipid nanoparticle (LNP) and at least one agent.

[0173] The term “lipid nanoparticle” refers to a particle having at least one dimension on the order of nanometers (e.g., 1-1,000 nm) which includes one or more lipids. In some embodiments, lipid nanoparticles are included in a delivery vehicle comprising at least one agent as described herein. In some embodiments, such lipid nanoparticles comprise a cationic lipid and one or more excipient selected from neutral lipids, charged lipids, steroids and polymer conjugated lipids (e.g., a pegylated lipid). In some embodiments, the at least one agent is encapsulated in the lipid portion of the lipid nanoparticle or an aqueous space enveloped by some or all of the lipid portion of the lipid nanoparticle, thereby protecting it from enzymaticdegradation or other undesirable effects induced by the mechanisms of the host organism or cells e.g. an adverse immune response.

[0174] In various embodiments, the lipid nanoparticles have a mean diameter of from about 30 nm to about 150 nm, from about 40 nm to about 150 nm, from about 50 nm to about 150 nm, from about 60 nm to about 130 nm, from about 70 nm to about 110 nm, from about 70 nm to about 100 nm, from about 80 nm to about 100 nm, from about 90 nm to about 100 nm, from about 70 to about 90 nm, from about 80 nm to about 90 nm, from about 70 nm to about 80 nm, or about 30 nm, 35 nm, 40 nm, 45 nm, 50 nm, 55 nm, 60 nm, 65 nm, 70 nm, 75 nm, 80 nm, 85 nm, 90 nm, 95 nm, 100 nm, 105 nm, 110 nm, 115 nm, 120 nm, 125 nm, 130 nm, 135 nm, 140 nm, 145 nm, or 150 nm. In one embodiment, the lipid nanoparticles have a mean diameter of about 83 nm. In one embodiment, the lipid nanoparticles have a mean diameter of about 102 nm. In one embodiment, the lipid nanoparticles have a mean diameter of about 103 nm. In some embodiments, the lipid nanoparticles are substantially non-toxic. In certain embodiments, the at least one agent, when present in the lipid nanoparticles, is resistant in aqueous solution to degradation by intra- or intercellular enzymes.

[0175] The LNP may comprise any lipid capable of forming a particle to which the at least one agent is attached, or in which the at least one agent is encapsulated. The term “lipid” refers to a group of organic compounds that are derivatives of fatty acids (e.g., esters) and are generally characterized by being insoluble in water but soluble in many organic solvents. Lipids are usually divided in at least three classes: (1) “simple lipids” which include fats and oils as well as waxes; (2) “compound lipids” which include phospholipids and glycolipids; and (3) “derived lipids” such as steroids.

[0176] In one embodiment, the LNP comprises one or more cationic lipids, and one or more stabilizing lipids. Stabilizing lipids include neutral lipids and pegylated lipids.

[0177] In one embodiment, the LNP comprises a cationic lipid. As used herein, the term “ionizable cationic lipid” refers to a lipid that is cationic or becomes cationic (protonated) as the pH is lowered below the pK of the ionizable group of the lipid, but is progressively more neutral at higher pH values. At pH values below the pK, the lipid is then able to associate with negatively charged nucleic acids. In certain embodiments, the cationic lipid comprises a zwitterionic lipid that assumes a positive charge on pH decrease.206678-0002-00WO

[0178] In certain embodiments, the cationic lipid or ionizable cationic lipid comprises any of a number of lipid species which carry a net positive charge at a selective pH, such as physiological pH or becomes cationic (protonated) at a selective pH. Such lipids include, but are not limited to, N, N-dioleyl-N, N-dimethylammonium chloride (DODAC); N-(2,3-dioleyloxy)propyl)-N, N, N-trimethylammonium chloride (DOTMA); N, N-distearyl-N, N-dimethylammonium bromide (DDAB); N-(2,3-dioleoyloxy)propyl)-N, N, N-trimethylammonium chloride (DOTAP); 3-(N — (N', N'-dimethylaminoethane)-carbamoyl)cholesterol (DC-Chol), N-(l-(2,3-dioleoyloxy)propyl)-N-2-(sperminecarboxamido)ethyl)-N, N-dimethylammonium trifluoracetate (DOSPA), dioctadecylamidoglycyl carboxyspermine (DOGS), l,2-dioleoyl-3-dimethylammonium propane (DODAP), N, N-dimethyl-2,3-dioleoyloxy)propylamine (DODMA), and N-(l,2-dimyristyloxyprop-3-yl)-N, N-dimethyl-N-hydroxyethyl ammonium bromide (DMRIE), l,2-dilinoleyloxy-N, N-dimethylaminopropane (DLinDMA), N, N-dimethyl-2,3 -bis(((9Z, 12Z, 15Z)-octadeca-9, 12,15 -trien- 1 -yl)oxy )propan- 1 -amine (DLenDMA), (6Z,9Z,28Z,3 lZ)-Heptatriaconta-6,9,28,31-tetraen- 19-yl 4-(dimethylamino)butanoate (DLin-MC3-DMA), heptadecan-9-yl 8-((2-hydroxyethyl)(6-oxo-6-(undecyloxy)hexyl)amino)octanoate (SM-102), ((4-hydroxybutyl)azanediyl)bis(hexane-6,l-diyl) bis(2 -hexyldecanoate) (ALC-0315), l,l'-((2-(4-(2-((2-(bis(2-hydroxydodecyl)amino)ethyl)(2-hydroxydodecyl)amino)ethyl)piperazin-1 -yl)ethyl)azanediyl)bi s(dodecan-2-ol) (C 12-200), 3, 6-bi s [4- [bi s(2-hydroxydodecyl)amino]butyl]-2,5-piperazinedione (cKK-E12), and Ethylenediamine-based Cysteinyl Oleoyl Lipid (ECO). Additionally, a number of commercial preparations of cationic lipids are available which can be used in the present invention. These include, for example, LIPOFECTIN® (commercially available cationic liposomes comprising DOTMA and 1,2-dioleoyl-sn-3-phosphoethanolamine (DOPE), from GIBCO / BRL, Grand Island, N. Y.);LIPOFECTAMINE® (commercially available cationic liposomes comprising N-(l-(2,3-dioleyloxy)propyl)-N-(2-(sperminecarboxamido)ethyl)-N, N-dimethylammonium tri fluoroacetate (DOSPA) and (DOPE), from GIBCO / BRL); and TRANSFECTAM® (commercially available cationic lipids comprising di octadecyl amidoglycyl carboxyspermine (DOGS) in ethanol from Promega Corp., Madison, Wis.).

[0179] In one embodiment, the cationic lipid is an amino lipid. Suitable amino lipids useful in the invention include those described in WO 2012 / 016184, incorporated herein by reference in its entirety. Representative amino lipids include, but are not limited to, 1,2-206678-0002-00WOdilinoleyoxy-3-(dimethylamino)acetoxypropane (DLin-DAC), 1,2-dilinoleyoxy-3-morpholinopropane (DLin-MA), l,2-dilinoleoyl-3 -dimethylaminopropane (DLinDAP), 1,2-dilinoleylthio-3-dimethylaminopropane (DLin-S-DMA), l-linoleoyl-2-linoleyloxy-3-dimethylaminopropane (DLin-2-DMAP), l,2-dilinoleyloxy-3 -trimethylaminopropane chloride salt (DLin-TMA. Cl), l,2-dilinoleoyl-3 -trimethylaminopropane chloride salt (DLin-TAP. Cl), 1,2-dilinoleyloxy-3-(N-methylpiperazino)propane (DLin-MPZ), 3-(N, N-dilinoleylamino)-l,2-propanediol (DLinAP), 3-(N, N-dioleylamino)-l,2-propanediol (DOAP), l,2-dilinoleyloxo-3-(2-N, N-dimethylamino)ethoxypropane (DLin-EG-DMA), and 2,2-dilinoleyl-4-dimethylaminomethyl-[1,3]-dioxolane (DLin-K-DMA).

[0180] In certain embodiments, the cationic lipid is present in the LNP in an amount from about 20 to about 95 mole percent. In one embodiment, the cationic lipid is present in the LNP in an amount from about 20 to about 70 mole percent. In one embodiment, the cationic lipid is present in the LNP in an amount from about 30 to about 60 mole percent. In one embodiment, the cationic lipid is present in the LNP in an amount of about 30 to about 50 mole percent.

[0181] In certain embodiments, the LNP comprises one or more additional lipids which stabilize the formation of particles during their formation.

[0182] Suitable stabilizing lipids include neutral lipids and anionic lipids.

[0183] The term “neutral lipid” refers to any one of a number of lipid species that exist in either an uncharged or neutral zwitterionic form at physiological pH. Representative neutral lipids include diacylphosphatidylcholines, diacylphosphatidylethanolamines, ceramides, sphingomyelins, dihydro sphingomyelins, cephalins, and cerebrosides.

[0184] Exemplary neutral lipids include, for example, distearoylphosphatidylcholine (DSPC), dioleoylphosphatidylcholine (DOPC), dipalmitoylphosphatidylcholine (DPPC), dioleoylphosphatidylglycerol (DOPG), dipalmitoylphosphatidylglycerol (DPPG), dioleoylphosphatidylethanolamine (DOPE), palmitoyloleoylphosphatidylcholine (POPC), palmitoyloleoyl-phosphatidylethanolamine (POPE) and dioleoyl-phosphatidylethanolamine 4-(N-maleimidomethyl)-cyclohexane-l -carboxylate (DOPE-mal), dipalmitoyl phosphatidyl ethanolamine (DPPE), dimyristoylphosphoethanolamine (DMPE), distearoylphosphatidylethanolamine (DSPE), 16-O-monomethyl PE, 16-O-dimethyl PE, 18-1-trans PE, 1-stearioyl-2-oleoyl-phosphatidy ethanol amine (SOPE), and l,2-dielaidoyl-sn-glycero-3-phophoethanolamine (transDOPE). In one embodiment, the neutral lipid is 1,2-di stearoyl -sn-glycero-3 -phosphocholine (DSPC).

[0185] In some embodiments, the LNPs comprise a neutral lipid selected from DSPC, DPPC, DMPC, DOPC, POPC, DOPE and SM. In certain embodiments, the neutral lipid is present in the LNP in an amount from about 20 to about 60 mole percent. In one embodiment, the neutral lipid is present in the LNP in an amount from about 25 to about 40 mole percent. In one embodiment, the neutral lipid is present in the LNP in an amount from about 25 to about 30 mole percent. In various embodiments, the molar ratio of the cationic lipid to the neutral lipid ranges from about 2: 1 to about 8:1.

[0186] In various embodiments, the LNPs further comprise a steroid or steroid analogue. In certain embodiments, the steroid or steroid analogue is cholesterol. In certain embodiments, the cholesterol is present in the LNP in an amount from about 10 to about 40 mole percent. In one embodiment, the cholesterol is present in the LNP in an amount from about 14 to about 30 mole percent. In some of these embodiments, the molar ratio of the cationic lipid to cholesterol ranges from about 2: 1 to 1: 1.

[0187] The term “anionic lipid” refers to any lipid that is negatively charged at physiological pH. These lipids include phosphatidylglycerol, cardiolipin, diacylphosphatidylserine, diacylphosphatidic acid, N-dodecanoylphosphatidylethanolamines, N-succinylphosphatidylethanolamines, N-glutarylphosphatidylethanolamines, lysylphosphatidylglycerols, palmitoyloleyolphosphatidylglycerol (POPG), and other anionic modifying groups joined to neutral lipids.

[0188] In some embodiments, the LNPs comprise a polymer conjugated lipid. The term “polymer conjugated lipid” refers to a molecule comprising both a lipid portion and a polymer portion. An example of a polymer conjugated lipid is a pegylated lipid. The term “pegylated lipid” refers to a molecule comprising both a lipid portion and a polyethylene glycol portion. Pegylated lipids are known in the art and include 1 (monomethoxy polyethyleneglycol) 2,3 dimyristoylglycerol (PEG s- DMG) and the like.

[0189] In certain embodiments, the LNP comprises an additional, stabilizing -lipid which is a polyethylene glycol-lipid (pegylated lipid). Suitable polyethylene glycol-lipids include PEG-modified phosphatidylethanolamine, PEG-modified phosphatidic acid, PEG-modified ceramides (eg., PEG-CerC14 or PEG-CerC20), PEG-modified di alkyl amines, PEG-modified206678-0002-00WOdi acylglycerols, PEG-modified dialkylglycerols. Representative polyethylene glycol-lipids include PEG-c-DOMG, PEG-c-DMA, and PEG-s-DMG. In one embodiment, the polyethylene glycol-lipid is N-[(methoxy poly(ethylene glycol)2000)carbamyl]-l,2-dimyristyloxlpropyl-3-amine (PEG-c-DMA). In one embodiment, the polyethylene glycol-lipid is PEG-c-DOMG). In other embodiments, the LNPs comprise a pegylated diacylglycerol (PEG-DAG) such as 1 (monomethoxy polyethyleneglycol) 2,3 dimyristoylglycerol (PEG-DMG), a pegylated phosphatidylethanoloamine (PEG-PE), a PEG succinate di acylglycerol (PEG-S-DAG) such as 4-O-(2’,3’-di(tetradecanoyloxy)propyl-l-O-(w-methoxy(polyethoxy)ethyl)butanedioate (PEG-S-DMG), a pegylated ceramide (PEG-cer), or a PEG dialkoxypropylcarbamate such as co-methoxy(polyethoxy)ethyl-N-(2,3-di(tetradecanoxy)propyl)carbamate or 2,3-di(tetradecanoxy)propyl-N-(co-methoxy(polyethoxy)ethyl)carbamate.

[0190] In certain embodiments, the pegylated lipid is present in the LNP in an amount from about 0 to about 20 mole percent. In one embodiment, the pegylated lipid is present in the LNP in an amount from about 0.1 to about 10 mole percent. In one embodiment, the pegylated lipid is present in the LNP in an amount from about 0.1 to about 5 mole percent. In various embodiments, the molar ratio of the cationic lipid to the pegylated lipid ranges from about 100:1 to about 25:1.

[0191] In certain embodiments, the LNP comprises one or more targeting moieties that targets the LNP to a cell or cell population. For example, in one embodiment, the targeting domain is a ligand which directs the LNP to a receptor found on a cell surface.

[0192] In certain embodiments, the LNP comprises one or more internalization domains. For example, in one embodiment, the LNP comprises one or more domains which bind to a cell to induce the internalization of the LNP. For example, in one embodiment, the one or more internalization domains bind to a receptor found on a cell surface to induce receptor-mediated uptake of the LNP. In certain embodiments, the LNP is capable of binding a biomolecule in vivo, where the LNP-bound biomolecule can then be recognized by a cell-surface receptor to induce internalization. For example, in one embodiment, the LNP binds systemic ApoE, which leads to the uptake of the LNP and associated cargo.

[0193] Exemplary LNPs and their manufacture are described in the art, for example in U.S. Patent Application Publication No. US20120276209, Semple et al., 2010, Nat Biotechnol., 28(2): 172- 176; Akinc et al., 2010, Mol Then, 18(7): 1357-1364; Basha et al., 2011, Mol Ther,206678-0002-00WO19(12): 2186-2200; Leung et al., 2012, J Phys Chem C Nanomater Interfaces, 116(34): 18440-18450; Lee et al., 2012, Int J Cancer., 131(5): E781-90; Belliveau et al., 2012, Mol Ther nucleic Acids, 1: e37; Jayaraman et al., 2012, Angew Chem Int Ed Engl., 51(34): 8529-8533; Mui et al., 2013, Mol Ther Nucleic Acids. 2, el39; Maier et al., 2013, Mol Ther., 21(8): 1570-1578; and Tam et al., 2013, Nanomedicine, 9(5): 665-74, each of which are incorporated by reference in their entirety.Targeting Domain

[0194] In various embodiments of the invention, the delivery vehicle is conjugated to a targeting domain. In one embodiment, the conjugation is a reversible conjugation, such that the delivery vehicle can be disassociated from the targeting domain upon exposure to certain conditions or chemical agents. In another embodiment, the conjugation is an irreversible conjugation, such that under normal conditions the delivery vehicle does not dissociate from the targeting domain.

[0195] In some embodiments, the conjugation comprises a covalent bond between an activated polymer conjugated lipid and the targeting domain. The term “activated polymer conjugated lipid” refers to a molecule comprising a lipid portion and a polymer portion that has been activated via functionalization of a polymer conjugated lipid with a first coupling group. In one embodiment, the activated polymer conjugated lipid comprises a first coupling group capable of reacting with a second coupling group. In one embodiment, the activated polymer conjugated lipid is an activated pegylated lipid. In one embodiment, the first coupling group is bound to the lipid portion of the pegylated lipid. In another embodiment, the first coupling group is bound to the polyethylene glycol portion of the pegylated lipid. In one embodiment, the second functional group is covalently attached to the targeting domain.

[0196] The first coupling group and second coupling group can be any functional groups known to those of skill in the art to together form a covalent bond, for example under mild reaction conditions or physiological conditions. In some embodiments, the first coupling group or second coupling group are selected from the group consisting of mal eimides, N-hydroxysuccinimide (NHS) esters, carbodiimides, hydrazide, pentafluorophenyl (PFP) esters, phosphines, hydroxymethyl phosphines, psoralen, imidoesters, pyridyl disulfide, isocyanates, vinyl sulfones, alpha-haloacetyls, aryl azides, acyl azides, alkyl azides, diazirines,206678-0002-00WObenzophenone, epoxides, carbonates, anhydrides, sulfonyl chlorides, cyclooctyne, aldehydes, and sulfhydryl groups. In some embodiments, the first coupling group or second coupling group is selected from the group consisiting of free amines (–NH2), free sulfhydryl groups (–SH), free hydroxide groups (–OH), carboxylates, hydrazides, and alkoxyamines. In some embodiments, the first coupling group is a functional group that is reactive toward sulfhydryl groups, such as maleimide, pyridyl disulfide, or a haloacetyl. In one embodiment, the first coupling group is a maleimide.

[0197] In one embodiment, the second coupling group is a sulfhydryl group. The sulfhydryl group can be installed on the targeting domain using any method known to those of skill in the art. In one embodiment, the sulfhydryl group is present on a free cysteine residue. In one embodiment, the sulfhydryl group is revealed via reduction of a disulfide on the targeting domain, such as through reaction with 2-mercaptoethylamine. In one embodiment, the sulfhydryl group is installed via a chemical reaction, such as the reaction between a free amine and 2-iminothilane or N-succinimidyl S-acetylthioacetate (SATA).

[0198] In some embodiments, the polymer conjugated lipid and targeting domain are functionalized with groups used in “click” chemistry. Bioorthogonal “click” chemistry comprises the reaction between a functional group with a 1,3-dipole, such as an azide, a nitrile oxide, a nitrone, an isocyanide, and the link, with an alkene or an alkyne dipolarophiles. Exemplary dipolarophiles include any strained cycloalkenes and cycloalkynes known to those of skill in the art, including, but not limited to, cyclooctynes, dibenzocyclooctynes, monofluorinated cyclcooctynes, difluorinated cyclooctynes, and biarylazacyclooctynone.

[0199] In some embodiments, the polymer conjugated lipid and targeting domain are functionalized with groups used in EDC / NHS (N-ethyl-N'-(3-(dimethylamino)propyl)carbodiimide / N-hydroxysuccinimide) crosslinking chemistry in which the intermediate molecule succinimidyl ester (NHS-ester) is used to immobilize biomolecules containing free primary amino groups via amide linkage.

[0200] In one embodiment, the composition comprises a targeting domain that directs the delivery vehicle to a target cell. The targeting domain may comprise a nucleic acid, peptide, antibody, small molecule, organic molecule, inorganic molecule, glycan, sugar, hormone, and the like that targets the particle to a site in particular need of the therapeutic agent. In certain embodiments, the particle comprises multivalent targeting, wherein the particle comprises206678-0002-00WOmultiple targeting mechanisms described herein. In certain embodiments, the targeting domain of the delivery vehicle specifically binds to a target associated with a site in need of an agent comprised within the delivery vehicle. For example, the targeting domain may be chosen to recognize a ligand that acts as a cell surface marker on target cells associated with a particular disease state. Such a target can be a protein, protein fragment, antigen, or other biomolecule that is associated with the targeted site. In some embodiments, the targeting domain is an affinity ligand which specifically binds to a target. In certain embodiments, the target (e.g. antigen) associated with a site in need of a treatment with an agent. In some embodiments, the targeting domain may be co-polymerized with the composition comprising the delivery vehicle. In some embodiments, the targeting domain may be covalently attached to the composition comprising the delivery vehicle, such as through a chemical reaction between the targeting domain and the composition comprising the delivery vehicle. In some embodiments, the targeting domain is an additive in the delivery vehicle. Targeting domains of the instant invention include, but are not limited to, antibodies, antibody fragments, proteins, peptides, and nucleic acids.Therapeutic Methods

[0201] In some embodiments, the invention provides methods for treatment or prevention of a disease or disorder. In some embodiments, the disease or disorder is a genetic disease or disorder resulting from a mutation in a gene. In some embodiments, the disease or disorder is cystic fibrosis.

[0202] To practice the methods of the invention; the skilled artisan would understand, based on the disclosure provided herein, how to formulate and administer the appropriate composition to a subject. The present invention is not limited to any particular method of administration or treatment regimen.

[0203] The invention encompasses delivery of a delivery vehicle comprising the plasmid of the invention for tunable expression of an encoded therapeutic agent. In one embodiment, the therapeutic agent is a known or established therapeutic for the disease being treated. In one embodiment, the therapeutic agent boosts the level or expression of a protein or gene product that is under-expressed or absent in a subject having a disease or disorder. In one embodiment, the therapeutic agent decreases the level or expression of a protein or gene product that is overexpressed or present in a subject having a disease or disorder.206678-0002-00WO

[0204] It will be appreciated by one of skill in the art, when armed with the present disclosure including the methods detailed herein, that the invention is not limited to treatment of diseases or disorders that are already established. Particularly, the disease or disorder need not have manifested to the point of detriment to the subject; indeed, the disease or disorder need not be detected in a subject before treatment is administered. That is, significant signs or symptoms of diseases or disorders do not have to occur before the present invention may provide benefit. Therefore, the present invention includes a method for preventing diseases or disorders, in that a composition, as discussed previously elsewhere herein, can be administered to a subject prior to the onset of diseases or disorders, thereby preventing diseases or disorders.

[0205] One of skill in the art, when armed with the disclosure herein, would appreciate that the prevention of a disease or disorder, encompasses administering to a subject a composition as a preventative measure against the development of, or progression of, a disease or disorder.

[0206] One of skill in the art will appreciate that the compositions of the invention can be administered singly or in any combination. Further, the compositions of the invention can be administered singly or in any combination in a temporal sense, in that they may be administered concurrently, or before, and / or after each other. One of ordinary skill in the art will appreciate, based on the disclosure provided herein, that the compositions of the invention can be used to prevent or to treat a disease or disorder, and that a composition can be used alone or in any combination with another composition to affect a therapeutic result. In various embodiments, any of the compositions of the invention described herein can be administered alone or in combination with other modulators of other molecules associated with diseases or disorders.

[0207] In one embodiment, the invention includes a method comprising administering a combination of compositions described herein. In certain embodiments, the method has an additive effect, wherein the overall effect of administering a combination of compositions is approximately equal to the sum of the effects of administering each individual inhibitor. In other embodiments, the method has a synergistic effect, wherein the overall effect of administering a combination of compositions is greater than the sum of the effects of administering each individual composition.

[0208] The method comprises administering a combination of composition in any suitable ratio. For example, in one embodiment, the method comprises administering twoindividual compositions at a 1:1 ratio. However, the method is not limited to any particular ratio. Rather any ratio that is shown to be effective is encompassed.Pharmaceutical Compositions

[0209] The formulations of the pharmaceutical compositions described herein may be prepared by any method known or hereafter developed in the art of pharmacology. In general, such preparatory methods include the step of bringing the active ingredient into association with a carrier or one or more other accessory ingredients, and then, if necessary or desirable, shaping or packaging the product into a desired single- or multi-dose unit.

[0210] Although the description of pharmaceutical compositions provided herein are principally directed to pharmaceutical compositions which are suitable for ethical administration to humans, it will be understood by the skilled artisan that such compositions are generally suitable for administration to animals of all sorts. Modification of pharmaceutical compositions suitable for administration to humans in order to render the compositions suitable for administration to various animals is well understood, and the ordinarily skilled veterinary pharmacologist can design and perform such modification with merely ordinary, if any, experimentation. Subjects to which administration of the pharmaceutical compositions of the invention is contemplated include, but are not limited to, humans and other primates, mammals including commercially relevant mammals such as non-human primates, cattle, pigs, horses, sheep, cats, and dogs.

[0211] Pharmaceutical compositions that are useful in the methods of the invention may be prepared, packaged, or sold in formulations suitable for ophthalmic, oral, rectal, vaginal, parenteral, topical, pulmonary, intranasal, buccal, intravenous, intracerebroventricular, intradermal, intramuscular, or another route of administration. Other contemplated formulations include projected nanoparticles, liposomal preparations, resealed erythrocytes containing the active ingredient, and immunogenic-based formulations.

[0212] In certain embodiments, the composition of the invention is administered by inhalation. In certain embodiments, the invention is conveniently delivered from an insufflator, nebulizer or a pressurized pack or other convenient means of delivering an aerosol spray.Pressurized packs may comprise a suitable propellant such as dichlorodifluoromethane, trichlorofluoromethane, dichlorotetrafluoroethane, carbon dioxide or other suitable gas. In thecase of a pressurized aerosol, the dosage unit may be determined by providing a valve to deliver a metered amount.

[0213] A pharmaceutical composition of the invention may be prepared, packaged, or sold in bulk, as a single unit dose, or as a plurality of single unit doses. As used herein, a “unit dose” is discrete amount of the pharmaceutical composition comprising a predetermined amount of the active ingredient. The amount of the active ingredient is generally equal to the dosage of the active ingredient which would be administered to a subject or a convenient fraction of such a dosage such as, for example, one-half or one-third of such a dosage.

[0214] The relative amounts of the active ingredient, the pharmaceutically acceptable carrier, and any additional ingredients in a pharmaceutical composition of the invention will vary, depending upon the identity, size, and condition of the subject treated and further depending upon the route by which the composition is to be administered. By way of example, the composition may comprise between 0.1% and 100% (w / w) active ingredient.

[0215] In addition to the active ingredient, a pharmaceutical composition of the invention may further comprise one or more additional pharmaceutically active agents.

[0216] In addition to the active ingredient, a pharmaceutical composition of the invention may further comprise one or more additional adjuvants. Exemplary adjuvants include, but are not limited to, aluminum-based adjuvant and monophosphoryl lipid A.

[0217] Controlled- or sustained-release formulations of a pharmaceutical composition of the invention may be made using conventional technology.

[0218] As used herein, “parenteral administration” of a pharmaceutical composition includes any route of administration characterized by physical breaching of a tissue of a subject and administration of the pharmaceutical composition through the breach in the tissue. Parenteral administration thus includes, but is not limited to, administration of a pharmaceutical composition by injection of the composition, by application of the composition through a surgical incision, by application of the composition through a tissue-penetrating non-surgical wound, and the like. In particular, parenteral administration is contemplated to include, but is not limited to, inhalation, intraocular, intravitreal, subretinal, suprachoroidal, subcutaneous, intraperitoneal, intramuscular, intradermal, intrasternal, intratumoral, intravenous, intracerebroventricular injections and kidney dialytic infusion techniques.206678-0002-00WO

[0219] Formulations of a pharmaceutical composition suitable for parenteral administration comprise the active ingredient combined with a pharmaceutically acceptable carrier, such as sterile water or sterile isotonic saline. Such formulations may be prepared, packaged, or sold in a form suitable for bolus administration or for continuous administration. Injectable formulations may be prepared, packaged, or sold in unit dosage form, such as in ampules or in multi-dose containers containing a preservative. Formulations for parenteral administration include, but are not limited to, suspensions, solutions, emulsions in oily or aqueous vehicles, pastes, and implantable sustained-release or biodegradable formulations. Such formulations may further comprise one or more additional ingredients including, but not limited to, suspending, stabilizing, or dispersing agents. In one embodiment of a formulation for parenteral administration, the active ingredient is provided in dry (i.e., powder or granular) form for reconstitution with a suitable vehicle (e g., sterile pyrogen-free water) prior to parenteral administration of the reconstituted composition.

[0220] The pharmaceutical compositions may be prepared, packaged, or sold in the form of a sterile injectable aqueous or oily suspension or solution. This suspension or solution may be formulated according to the known art, and may comprise, in addition to the active ingredient, additional ingredients such as dispersing agents, wetting agents, or suspending agents described herein. Such sterile injectable formulations may be prepared using a non-toxicparenterally-acceptable diluent or solvent, such as water or 1,3-butane diol, for example. Other acceptable diluents and solvents include, but are not limited to, Ringer's solution, isotonic sodium chloride solution, and fixed oils such as synthetic mono- or di-glycerides. Other parentally-administrable formulations which are useful include those which comprise the active ingredient in microcrystalline form, in a liposomal preparation, or as a component of biodegradable polymer systems. Compositions for sustained release or implantation may comprise pharmaceutically acceptable polymeric or hydrophobic materials such as an emulsion, an ion exchange resin, a sparingly soluble polymer, or a sparingly soluble salt.

[0221] A pharmaceutical composition of the invention may be prepared, packaged, or sold in a formulation suitable for pulmonary administration via the buccal cavity. Such a formulation may comprise dry particles which comprise the active ingredient and which have a diameter in the range from about 0.5 to about 7 nanometers, and preferably from about 1 to about 6 nanometers. Such compositions are conveniently in the form of dry powders for administration206678-0002-00WOusing a device comprising a dry powder reservoir to which a stream of propellant may be directed to disperse the powder or using a self-propelling solvent / powder-dispensing container such as a device comprising the active ingredient dissolved or suspended in a low-boiling propellant in a sealed container. Preferably, such powders comprise particles wherein at least 98% of the particles by weight have a diameter greater than 0.5 nanometers and at least 95% of the particles by number have a diameter less than 7 nanometers. More preferably, at least 95% of the particles by weight have a diameter greater than 1 nanometer and at least 90% of the particles by number have a diameter less than 6 nanometers. Dry powder compositions preferably include a solid fine powder diluent such as sugar and are conveniently provided in a unit dose form.

[0222] Low boiling propellants generally include liquid propellants having a boiling point of below 65°F at atmospheric pressure. Generally the propellant may constitute 50 to 99.9% (w / w) of the composition, and the active ingredient may constitute 0.1 to 20% (w / w) of the composition. The propellant may further comprise additional ingredients such as a liquid non-ionic or solid anionic surfactant or a solid diluent (preferably having a particle size of the same order as particles comprising the active ingredient).

[0223] As used herein, “additional ingredients” include, but are not limited to, one or more of the following: excipients; surface active agents; dispersing agents; inert diluents; granulating and disintegrating agents; binding agents; lubricating agents; sweetening agents; flavoring agents; coloring agents; preservatives; physiologically degradable compositions such as gelatin; aqueous vehicles and solvents; oily vehicles and solvents; suspending agents; dispersing or wetting agents; emulsifying agents, demulcents; buffers; salts; thickening agents; fillers; emulsifying agents; antioxidants; antibiotics; antifungal agents; stabilizing agents; and pharmaceutically acceptable polymeric or hydrophobic materials. Other “additional ingredients” which may be included in the pharmaceutical compositions of the invention are known in the art and described, for example in Remington's Pharmaceutical Sciences (1985, Genaro, ed., Mack Publishing Co., Easton, PA), which is incorporated herein by reference.EXPERIMENTAL EXAMPLES

[0224] The invention is further described in detail by reference to the following experimental examples. These examples are provided for purposes of illustration only, and are not intended to be limiting unless otherwise specified. Thus, the invention should in no way beconstrued as being limited to the following examples, but rather should be construed to encompass any and all variations which become evident as a result of the teaching provided herein.

[0225] Without further description, it is believed that one of ordinary skill in the art can, using the preceding description and the following illustrative examples, make and utilize the present invention and practice the claimed methods. The following working examples therefore are not to be construed as limiting in any way for the remainder of the disclosure.Example 1; Selective, Tunable and Differential Autoregulation of Endogenous Mutant and Transgene CFTR Gene Expressions

[0226] Cystic Fibrosis is caused by mutations in the CFTR gene, leading to defects in chloride and bicarbonate ion transport. Traditional treatments have relied on small-molecule CFTR modulators, which work only for select mutations and often have limited efficacy. Gene therapy offers a potentially universal solution by introducing functional CFTR into lung epithelial cells.

[0227] A gene therapy approach for treating Cystic Fibrosis (CF) was developed by introducing DNA vectors that encode a functional full-length wild-type CFTR protein into lung epithelial cells that are impaired in CF due to mutations in the endogenous CFTR gene. This gene therapy method has the potential to be universally applicable to all CF patients, regardless of their specific CFTR mutation.Targeting CFTR-producing Cells using Transcription Factor responsive - DNA Targeting Sequences

[0228] Cystic Fibrosis Transmembrane Conductance Regulator (CFTR) is expressed at varying levels across different cell types in the lung epithelium. These cell types can be ranked based on their relative contribution to CFTR production and their functional role in maintaining airway hydration and mucus clearance.Table 1: CFTR Expression in Lung Epithelial CellsCell Type Location & % of Total Lung Epithelium % of Total CFTR Output206678-0002-00WOClub Cells Bronchioles (-10-15%) -30%lonocytes Rare (< 1 %) -25-30%Basal Cells Airways (-30-40%) -15-20%Goblet Cells Airways (-5-10%) -10%SubmucosalSubmucosal glands (-5-10%) -10%Gland CellsCiliated Cells Airways (-50-60%) -5-7%

[0229] Table 1 shows that Club Cells, Goblet Cells, Ionocytes, Submucosal Gland Cells, Ciliated Cells, and Basal Cells contribute to approximately 80% of the total CFTR expression in the airway epithelium, with secretory cells being the dominant contributors. Basal Cells contribute approximately 5-10% of the total CFTR expression. Ionocytes, despite their high percell CFTR expression, are relatively rare and thus contribute less to the overall CFTR expression. Broad overexpression of CFTR in all lung cells risks off-target toxicity and immunogenicity, especially in cells that do not naturally produce significant amounts of CFTR. Lung epithelial cell heterogeneity complicates gene delivery, since CFTR expression is mostly localized in these cell types: Club Cells, Goblet Cells, Ionocytes, Submucosal Gland Cells, Ciliated Cells, and Basal Cells. Ionocytes, once believed to contribute -50-60% of total CFTR production, are now understood to account for a smaller percentage of total CFTR output, with goblet cells (-30%) and club cells (-50%) playing a dominant role. Submucosal gland cells (-10-20%) contribute mucus and fluid secretion, further underscoring the need for cell-type-specific gene targeting.

[0230] In recent years, Transcription Factor responsive - DNA Targeting Sequences (TF-DTSs) have emerged as a powerful strategy to confine transgene expression to specific cell types. For example, TFs such as FOXI1, HNF1α, FOXA2, SPDEF, and TP63 are selectively active in CFTR-producing cells. By selecting DTSs only present in cell types that secrete CFTR, it is possible to selectively enable nuclear import of the TF-DTS-DNA complex into the nucleus of these cell types.

[0231] DNA constructs designed for selective and efficient transgene expression in targeted cell types are provided. Transcription Factor (TF) Responsive - DNA Targeting206678-0002-00WOSequences (TF-DTSs) is a key aspect of the invention and is applicable to a broad range of gene therapy applications for the selective and efficient expression in targeted cells.

[0232] TF-DTSs are utilized to deliver CFTR transgene to the targeted CFTR-producing cell types in lung epithelial cells. The design details for this indication are focused on herein even though the technology has broad applications.Transcription Factor Responsive - DNA Targeting Sequence (TF-DTS) Design

[0233] TF-DTSs are incorporated into the construct to ensure CFTR transgene expression is restricted to CFTR-producing cells, including goblet cells, club cells, ciliated cells, ionocytes, basal cells, and submucosal gland cells. TFs such as FOXI1, HNF1α, FOXA2, SPDEF, and TP63 are leveraged for their high selectivity and nuclear transport efficiency, ensuring that CFTR expression occurs only in cells naturally involved in CFTR regulation while minimizing off-target effects in non-CFTR-expressing tissues.

[0234] A key feature utilization of TF-DTSs to achieve selective and precise activation of the CFTR gene in CFTR-producing cells while preventing expression in non-CFTR-expressing cells. The objective of this gene therapy approach is to restore functional CFTR expression exclusively in CFTR-producing cells, while avoiding unintended gene activation in non-target cell types.

[0235] TF-DTSs are incorporated into CFTR gene therapy constructs to selectively drive CFTR expression in CFTR-producing cells while minimizing off-target expression in non-CFTR-producing cells or tissues. Certain TFs are uniquely expressed in CFTR-producing cell populations, creating a strategic opportunity for highly specific cell type targeting.

[0236] These DTS elements are incorporated into the CFTR gene therapy construct, ensuring that only cells expressing the appropriate TFs can enter the cell nucleus, thereby enabling highly specific and effective gene therapy for cystic fibrosis.

[0237] The following six steps outline how Transcription Factor Responsive - DTS (TF-DTS) drive the uptake of the TF-DTS-DNA complex by the nucleus:

[0238] 1. Cell Entry: the therapeutic DNA construct containing TF-DTSs and the transgene is delivered into lung epithelial cells—typically via lipid nanoparticles (LNPs) inhalation. The LNPs fuse with the cell membrane or are endocytosed, releasing the DNA into the cytoplasm.206678-0002-00WO

[0239] 2. Cell-Specific TF Expression: different lung epithelial subpopulations naturally express distinct transcription factors (TFs) as shown in Table 1.

[0240] 3. TF-DTS Binding: a plasmid is engineered to contain a DNA targeting sequence (DTS) - a binding site for an endogenous transcription factor present in the cell. Once the DNA construct is in the cytoplasm, the relevant TF recognizes and binds the unique TF-DTS motif on the DNA, forming the TF-DTS-DNA complex.

[0241] 4. Nuclear Import (TF Piggyback): each TF contains a nuclear localization signal (NLS) that facilitates nuclear import of the TF. Once bound to the plasmid, the transcription factor’s NLS is exposed on the DNA-protein complex. The cell’s importin-a / p recognizes the NLS on the TF (just as if the TF were free) and mediates transport of the entire TF-DTS-DNA complex into the nucleus through the nuclear pore.

[0242] The nuclear import efficiency of DNA mediated by transcription factors (TFs) is influenced by multiple factors, including the cytoplasmic concentration of the TF, the binding affinity between the TF and its DNA target sequence (DTS), and the interaction strength between the nuclear localization signal (NLS) of the TF and importin-a. Enhancing nuclear import efficiency can be achieved through strategic design approaches, such as incorporating multiple DTSs, utilizing tandem repeats of DTS motifs, optimizing extended DTS motifs, and fine-tuning DTS spacer arrangements to facilitate enhanced TF binding and nuclear translocation.

[0243] 5. CFTR Transgene Transcription: inside the nucleus, the TF can recruit the transcriptional machinery including co-activators and RNA polymerase, or even lung-specific promoter upstream of the CFTR gene. This helps the initiation of CFTR transcription, enabling the cell to produce functional CFTR protein.

[0244] One major limitation in gene therapy is epigenetic silencing of transgenes via DNA methylation and histone modifications. TF-bound DTSs actively recruit chromatinmodifying enzymes, such as histone acetyltransferases (HATs), which could prevent repressive chromatin remodeling and maintain CFTR expression long-term.

[0245] 6. De-Targeting Non-Expressing Cells: cells that do not express the TFs that can bind the DNA’s DTS motifs. Lacking the TF-DTS binding, the DNA great than 300 bp in size typically can’t defuse through nuclear complex core and enter the nucleus. Consequently, no or minimal CFTR transgene activation occurs in these non-target cells.206678-0002-00WO

[0246] The system ensures precise CFTR restoration in targeted cells such as ionocytes, submucosal gland cells, club cells, goblet cells or basal cells, while de-targeting non-relevant lung epithelium such as alveolar cells.

[0247] One of the primary barriers in gene therapy is ensuring that the DNA construct efficiently reaches the nucleus. TF-DTSs play a critical role in enhancing nuclear import of the TF-DTS-DNA complex. Many TFs involved in CFTR regulation (e g., FOXI1, FOXJ1, SPDEF) contain nuclear localization signals (NLSs), which facilitate nuclear transport when bound to their corresponding DTSs.

[0248] Once inside the nucleus, CFTR gene expression efficiency depends on transcriptional activation.

[0249] TFs were selected based on the following criteria. First, TFs must be exclusive to CFTR-producing cells and avoid expression in non-CFTR-producing cells (Alveolar, Fibroblast, Endothelial, Smooth Muscle Cells). This prevents unintended CFTR transgene expression in inappropriate cell types. Not all the criteria have to be met by the same TF. Some TFs may have efficient nuclear import with high cell specificity, while others may drive high transcription with low variability.

[0250] Factors considered included: 1) efficient nuclear import; 2) high cell specificity; 3) high TF expression; 4) high transgene transcription; 5) Durable transgene transcription; and 6) low transgene transcription variability.Table 2: DTS Candidates for TS-DTS-DNATranscription Example TF DNA BindingTF Expression in Lung Cell Types Factor (TF) Sequences (5'— >3')TTGGCATTTTGCCAA(SEQ IDNO:1); TTGGCAAAAAGCCAA(SEQ IDNO: 2);• Enriched in Club cells of airways NFIA TTGGCACGTAGCCAA (SEQ ID(secretory cells).(Nuclear NO:3); TTGGCACCTGCCAA• NFIA is the first identified Club cell-Factor I A) (SEQ IDNO 4);enriched TF, required for Notch TTGGCTTTTTGCCAA(SEQ IDsignaling in Club cell maintenance.NO: 5); TTGGCAATAAGCCAA(SEQ ID NO: 6);TTGGCCAAATGCCAA (SEQ ID206678-0002-00WONO: 7); GCCACTTAA;CCCACTTAA; GCCACTTAG;ACCACTTAG; GGCACTTAA;AGCACTTAA; CCCACTTAG;TCCACTTAA AATAAAG; ATAAACA;GTAAATA; GTAAACA;GTAAACAA; ATAAAT; GTAAAT;TGTTTAC; TGTTTAT;TATTGACTTAG (SEQ ID NO: 8);ATGTAAACATA (SEQ ID NO:9); • Expressed in secretory epithelial cells ATGTAAACATG (SEQ ID NO: 10);FOXA2 (e.g., Club cells and submucosal ATGTAAAC AAA (SEQ ID NO: 11 );(Forkhead glandular cells);ATGTAAACAAG (SEQ ID NO: 12);BoxA2) • FOXA2 maintains airway mucus GTGTAAACATA (SEQ ID NO: 13); homeostasis, acting to limit goblet cell AAGTAAACATA (SEQ ID NO: 14); hyperplasia.GTGTAAACATG (SEQ ID NO: 15);AAGTAAACATG (SEQ ID NO: 16);GTGTAAACAAA (SEQ IDNO: 17); AAGTAAACAAA (SEQID NO: 18)• Transiently expressed in multiciliated cell precursors; a master regulator of TTTCGCGC; TTTGGCGC;MCIDAS multiciliated cell differentiation.TTTCCGCC; TTTCCCGC;(“Multicilin”) • MCIDAS (with E2F4 / 5-DP1)TTTGCCGC; TGTCCCGC activates the multiciliogenesis program (e.g., inducing FOXJ1) in ciliated epithelial cells.GGAT; AGGAT; TGGAT; CGGAT;GGAA; AGGATTC; • Highly expressed in goblet cells and SPDEF (SAM- TTAGGAATAT (SEQ ID NO: 19); submucosal gland cells of airway pointed ATGCGGGC; GTGCGGGT; epithelium.Domain ETS GTGCGGGC; ATACGGGT; • SPDEF is required for goblet cell Factor) ATGGGGGT; ATGCGGGG; differentiation; overexpression induces CTGCGGGT; ATACGGGC; goblet cell hyperplasia and mucin ATGCGGGA production in trachea / bronchi.206678-0002-00WOTGTTTAC; GTAAACA; TATTTAT;TGTTTAT; TGTTTGT; TATTTAC;GTAAATAA; TATTGACTTTGFOXI1 (SEQ ID NO 20); GTAAACA; • Specifically expressed in pulmonary (Forkhead GTCAACA; GTAAATA; ionocytes, the CFTR-rich rare cells.Box II) ATAAACA; GTAAAAA; • FOXI1 is required for ionocyte GCAAACA; GTCAATA; identity.GAAAACA; GTAATCA;ATCAACA AAACCGGTT; CAAACCGGTTTFCP2L1 • Highly expressed in pulmonary (SEQ ID NO 21); TAAACCGGTT(Transcription ionocytes co-expressing CFTR.(SEQ ID NO:22); CCAGTTCAA;Factor CPI- • TFCP2L1 (a pluripotency and CAGTTCAAC; CCAGTTCAAClike 1) developmental TF) marks the ionocyte (SEQ IDNO:23) lineage in mouse and human airways.• Detected in CFTR-expressing CAGGTG; CACCTG; CACGTG; ionocytes, and is also found in certain ASCL3 GCACCTGCC; CCACCTGCC; glandular progenitors (e.g., salivary (Achaete- GCACCTGCT; ACACCTGCC; glands).Scute Family GCACCTGCA; GCACCTGCG; • ASCL3 is a bHLH transcription factor bHLH 3) CCACCTGCT; TCACCTGCC; co-expressed with FOXI1 in ionocytes GCAGCTGCC; GCACCTGGC, suggesting a role in secretory epithelial differentiation.AGGCATGTCT (SEQ ID NO:24);AAACATGTTT (SEQ ID NO:25); • Enriched in basal cells of airway GGGCATGTCC (SEQ ID NO:26); epithelium.TP63 (Tumor GGGCAAGTTT (SEQ ID NO:27); • p63 is the defining marker of basal Protein p63) GGGCTCGTTT (SEQ ID NO:28); progenitor cells; TP63' basal cells GGGCGTGTTT (SEQ ID NO:29); underlie the CFTR-expressing luminal GGGCATGTTT (SEQ ID NO:30); layers and can give rise to other cell TAACATGTTA (SEQ ID NO: 31) types during regeneration.ACAAAG; ATAAAG; ACAAAT; • Specifically expressed in goblet cells FOXQ1 ATAAAC; ATAAAT; GTAAAC; of airway epithelium.(Forkhead TGTTTAC; TCAATA; • FOXQ1 is required for mucin Box QI) GTAAATAA; TATTGATTTTG production; it controls goblet cell (SEQ IDNO:32) differentiation.206678-0002-00WO• Exclusively expressed in multi- TGTTTAC; GTAAATA; ciliated cells of the airway.FOXJ1 GTTTACA; ATAAATA; • FOXJ1 is the key transcriptional (Forkhead GTAAACAAA; ATAAACAAA; activator of motile cilia assembly; it Box JI) ATAAACAA; TAAACAAA; marks ciliated cells and is necessary TATTGATTTAG (SEQ ID NO:33) for the ciliated cell differentiation program in CFTR-expressing airway epithelia.• Highly expressed in multi-ciliated GTTGCCAGCAAC (SEQ ID cells of the airway.NO:34); GTARCCGGTAAC (SEQ • RFX3 (with RFX2) redundantly RFX3 ID NO:36); GTARCCAGCAAC regulates motile ciliogenesis, (Regulatory (SEQ ID NO 37); GTTACCATG; activating ciliary gene expression.Factor X3) GTTGCTATG; GTTACTATG; RFX-binding X-box motifs are present GTTACCTAGTAAC (SEQ ID in many cilia gene promoters; RFX3 NO:37) ensures proper cilia formation in ciliated CFTR-expressing cells.AGAACAATGG (SEQ ID NO:38); • Found in a rare subset of basal cells in CGAACAATGG (SEQ ID NO: 39);SOX9 (SRY- airways.AGAACAATAG (SEQ ID NO:40);box 9) • SOX9+basal cells reside in airway CATTGAA; CTTTGTT;“rugae” and have progenitor capacity CTTTGAA; ACAAAG; TTCAAAG to regenerate alveoli.GTTAAT; TTGTTA; • Found in submucosal glandular cells. GTTAATGATTAAC (SEQ ID • Expression correlates with CFTR- NO:41); GTTAATCATTAAC (SEQ expressing epithelial tissues.HNF1α ID NO: 42); GTTAATTATTAAC • HNFla forms a complex with other (Hepatocyte (SEQ ID NO:43); factors (CDX2, HNF4a, FOXA2) on Nuclear GTTAATTTATTAAC (SEQ ID CFTR enhancers. In CFTR-expressing Factor 1α) NO:44); GGTTAATCATTAACC cells, HNF1α acts as a master regulator (SEQ ID NO:45); – its presence stabilizes an active GGTTAATGATTAAC (SEQ ID chromatin state; loss of HNF1α NO: 46) binding drastically reduces CFTR expression.AGGTCAAAGGTCA (SEQ ID • Present in airway secretory cells and HNF4α NO:47); AGGTCACAGGTCA submucosal gland cells (tissues with (Hepatocyte (SEQ ID NO:48); high CFTR).Nuclear AGGTCATAGGTCA (SEQ ID • HNF4α (a nuclear receptor) Factor 4α) NO:49); AGGTCACTAGGTCA cooperates with HNF1α and FOXA in (SEQ ID NO:50);regulating CFTR intronic enhancers.206678-0002-00WOAGGTCATTAGGTCA (SEQ ID It helps drive a gene program in NO:51) intestinal, pancreatic, and airway submucosal gland epithelia - all prominent sites of CFTR expression. • Strongly expressed in goblet cells of airway epithelium during mucus FOXA3 GTAAACA; ATAAATA; hyperplasia.(Forkhead ATCAATA; ATAAAT; ATAAAC; • FOXA3 is induced by IL- 13 and viral Box A3) GTAAATAA; GTAAAT infection; it drives goblet cell metaplasia and mucin gene expressionin asthmatic airways.

[0251] It should be noted that the TFs and their DNA binding sequences listed in Table 2 herein are exemplary and non-limiting. The selection of TFs and their combinations may be further optimized based on the specific needs of a given CF patient population, the mode of gene delivery, or the precise lung epithelial cell type targeted. Additional TFs and regulatory sequences may be incorporated to further enhance the specificity, efficiency, and durability of CFTR expression in the airway epithelium.

[0252] Additionally, the disclosed techniques for cell-type-specific gene targeting using TF-DTSs may be applied beyond CF gene therapy to other indications. TF-based gene therapy may be used for chronic obstructive pulmonary disease (COPD) to restore epithelial integrity by selectively driving the expression of airway repair genes in basal cells. TF-based targeting may be employed for asthma by modulating inflammatory gene expression in airway epithelial and immune cells. TF-DTS gene therapy may be used to treat bronchiectasis or other lung disorders requiring selective enhancement of mucociliary function, wherein TF -responsive elements regulate genes involved in cilia motility and airway surface liquid balance.

[0253] Accordingly, the invention is not limited to the specific TF combinations disclosed but includes any modifications, additions, or alternative embodiments that utilize transcription factor-based mechanisms to selectively drive transgene nuclear import and transcription that requires precise cell-type-specific gene expression.Nuclear Import Efficiency of TFs

[0254] NLS and Importin Interactions: All TFs in Table 2 contain classical nuclear localization signals (NLS) that mediate import via importin-α / β. Forkhead (FOX) family members (FOXA2, FOXA3, FOXI1, FOXJ1, FOXQ1) have two NLS sequences within their DNA-binding (forkhead) domain - a conserved C-terminal RK-rich NLS and a second N-terminal NLS. These NLS motifs bind importin-α, which then recruits importin-β to ferry the TF into the nucleus. For example, FOXA2’s dual NLS ensures efficient nuclear import. Similarly, HMG-domain factor SOX9 carries two NLS (one at each end of its HMG domain) that function independently - one route likely involves importin-α, and the other can directly engage importin-β. HNF1α and HNF4α (hepatic nuclear factors) also have strong NLS regions recognized by importins. Indeed, importin-α acts as an “HNF-1 receptor” binding HNFl’s NLS and transporting it into the nucleus. NFIA (Nuclear Factor I / A) contains a bipartite NLS in its C-terminal region, ensuring its dimeric form efficiently localizes to nuclei. Other TFs like RFX3, SPDEF, TP63, MCIDAS, ASCL3, and TFCP2L1 all have basic-domain or motif NLS signals. In summary, all TFs in Table 2 possess NLS sequences and utilize the importin-α / β pathway for nuclear import.

[0255] Effectiveness in Aiding DNA Nuclear Import: these TFs were ranked by their expected ability to chaperone DNA into the nucleus, considering importin-binding affinity, DNA-binding affinity, and cytosolic availability.

[0256] FOXA2 is highly effective and functions as a pioneer factor that can bind DNA even in condensed chromatin. Its import is constitutive, meaning some fraction could engage cytosolic plasmid and move it inward.

[0257] HNF1α is another top import driver – it binds DNA as a dimer and is actively transported by importin. HNF1 is abundant in many epithelial cells (especially submucosal gland cells and intestinal cells), so plasmid bound to HNF1 has a good chance of import.

[0258] HNF4α likewise dimerizes on DNA and has a strong NLS; in cells where HNF4α is present (e.g. airway submucosal or certain lung cells with a gastrointestinal-like profile), it can robustly ferry DNA to the nucleus.

[0259] NFIA is present in a relatively large fraction of airway cells (club / secretory cells), has strong DNA-binding and a constitutive nuclear presence. Its NLS ensures efficient import. Thus, NFIA would most broadly and reliably enhance nuclear import and slightly support CFTR transcription.is also ranked in the top tier: it binds its palindromic site with high affinity andforms dimers, and its NLS is efficient. NFIA proteins have been used as nuclear targeting factors in viral origins (e.g. SV40 DTS element binds NFIA) to enhance plasmid import, indicating NFIA can effectively aid import.

[0260] FOXI1 can be very effective, as it is highly expressed in ionocyte cells. FoxI1 has strong NLS motifs (as a FOX protein) and high DNA affinity, so a FoxI1 site on the plasmid could yield import into ionocytes.

[0261] FOXJ1 and RFX3 are abundant in multiciliated cells, which make up a large fraction of airway epithelium. Both have classical NLS and bind DNA strongly as noted.Because these factors are largely nuclear (they actively regulate cilia genes), their availability in the cytosol may be lower - however, during ciliogenesis (and possibly during each ciliary gene transcription cycle), some newly synthesized FoxJ1 / RFX3 could bind plasmid DNA before entering the nucleus. They are ranked mid-tier: in ciliated cells, multiple FoxJ1 / RFX3 binding sites on the DNA can “piggyback” the plasmid into the nucleus, leveraging the constant import of these factors to sustain cilia gene expression.

[0262] FOXQ1 - Although goblet cells are fewer than club cells, FOXQ1 is highly expressed in both of those cells and strongly nuclear. Its binding could activate transcription of the CFTR transgene in goblet cells. In CF airways (where goblet cells are more numerous), FOXQ1 becomes more relevant. Its import efficiency is high in those cells; however, its narrower expression pattern drops it just below NFIA in rank.

[0263] TP63 (e.g. ΔNp63α) is a sequence-specific DNA-binding protein (recognizing p53 / p63 response elements in gene promoters), and it carries a strong NLS. In basal epithelial cells (which have abundant ΔNp63α), this approach could enhance nuclear uptake of the therapeutic DNA.Lower Tier- Limited Import Impact

[0264] SPDEF and FOXQ1 are lower-ranked for import. SPDEF (an ETS factor) and FOXQ1 (forkhead) are typically found in secretory / goblet cells, where they drive mucus production. They localize to the nucleus to activate goblet cell genes, but under baseline conditions, goblet cells are fewer in normal airways. Additionally, SPDEF / FOXQ1 may not be as strongly or constitutively imported as others - SPDEF’ s import relies on its ETS-domain NLS and it can be nuclear or partially cytosolic depending on signaling. If IL- 13 inflammationinduces goblet cells (increasing SPDEF / F0XQ1 expression), these could bind a plasmid and import it, but that scenario might coincide with an undesirable inflammatory state.

[0265] FOXA3 (HNF3y) is similar to FOXA2 in possessing dual NLS and pioneer activity, but it is typically expressed somewhat later or in different tissues (e.g. liver, gut). In lung epithelium its levels are lower; if present, it would aid import comparably to FOXA2 (hence mid-tier in lung context).

[0266] ASCL3 (achaete-scute like 3) is expressed in certain epithelial precursors (possibly tuft or rare secretory cells). Its bHLH domain confers an NLS, but ASCL3 expression in the differentiated airway is low. It might contribute import in rare cells or during cell fate changes.

[0267] MCIDAS (multiciliate differentiation factor) is transiently expressed in basal cells that are about to become multiciliated (it triggers massive centriole replication). MCIDAS has an NLS (as it acts in the nucleus to drive gene expression for ciliogenesis). However, MCIDAS expression is cell-cycle-regulated and short-lived - it might bind the plasmid in a basal cell and import it during the window of differentiation, but once the cell becomes fully ciliated, MCIDAS is downregulated. Thus, its import help is limited to that transient phase.

[0268] TFCP2L1 (CP2-like) is expressed in certain lung progenitors and alveolar type II cells during regeneration. It has an NLS and DNA-binding domain able to recruit importin-a. If the DNA reaches an alveolar region with TFCP2L1^+ cells, it could assist import, but in normal airway epithelium TFCP2L1 is not highly expressed except during repair. Thus, mid-tier and context-dependent.

[0269] SOX9 contains two nuclear localization signals (NLSs) flanking its HMG-box DNA-binding domain. It is highly expressed in submucosal gland progenitor / duct cells.Impact of These TFs on CFTR Transcription

[0270] Once the DNA reaches the nucleus, the binding of these TFs to their DNA targeting sequence (DTS) can influence CFTR transgene expression. It was assessed whether each TF is likely to enhance CFTR transcription or interfere with it, and whether their effect is direct (binding CFTR regulatory sequences to modulate transcription) or indirect (altering chromatin context or cell state). Several of the approved TFs are known components of CFTR’s native regulatory network.206678-0002-00WO

[0271] HNFla is a key direct activator of CFTR transcription. It binds specific ciselements in CFTR introns (notably intron 1 and intron 11 enhancers) and is required for full tissue-specific CFTR expression. In intestinal cells, HNFla occupies CFTR enhancers cooperatively with FOXA and CDX2. Loss of HNF1 reduces CFTR expression significantly. In the target lung epithelial cells, HNFla is expressed in submucosal gland cells and some club (secretory) cells, which have relatively high CFTR levels. If the DNA includes an HNF1 binding site, HNFla will bind directly and likely boost transcription of the CFTR transgene, acting as a potent enhancer-binding factor. This is a direct positive role. HNFla is ranked as one of the strongest enhancers for CFTR among the list.

[0272] HNF4a can positively regulate epithelial genes in concert with HNF1. While not highlighted in airway CFTR regulation literature, HNF4a is co-expressed with CFTR in gastrointestinal epithelia and may regulate overlapping targets. HNF4 binds DR1 elements in promoters / enhancers, and if such a site is placed on the plasmid, HNF4a binding could activate transcription by recruiting coactivators (HNF4 has ligand-independent activation functions). In lung context, HNF4a is not highly expressed in airway epithelial cells, so its direct impact may be limited to any transfected cells that do express it (perhaps submucosal gland cells or alveolar type II cells). In those cells, HNF4a would enhance CFTR transcription (directly binding and activating the promoter). Overall, HNF4a is a potential direct enhancer if present, but its contribution in most airway cells is modest compared to HNF1 / FOXA.

[0273] RFX3 is a transcriptional activator for many cilia-related genes (it binds X-box motifs in promoters). CFTR is not a classic motile cilia structural gene, but it is expressed in ciliated cells (which RFX3 helps differentiate). There is no evidence that RFX3 directly binds CFTR’s promoter / enhancers. However, RFX3 might indirectly enhance CFTR expression by promoting a ciliated cell state that is permissive for CFTR transcription. Ciliated cells have moderate CFTR levels, and RFX3, along with FoxJl, ensures those cells are fully differentiated (with an active gene expression program). If RFX3 binds to a DTS on the plasmid (e.g. an X-box sequence included as a nuclear targeting sequence), it could act as a local enhancer on the vector - RFX family members typically recruit co-activators (e.g. RFX forms enhanceosome complexes in MHC-II gene regulation). So, an RFX3 site near the CFTR transgene promoter might modestly activate transcription in any cell where RFX3 is active (multiciliated cells). This effect is likely positive but not as potent as FOXA or HNF1, since CFTR isn’t a primary target ofRFX3. RFX3’s role is classified as indirectly positive - maintaining a cell type that expresses CFTR and possibly providing slight enhancer activity if its site is present.

[0274] Combining the above analysis, the TFs were ranked by their ability to enhance CFTR transcription in lung epithelial target cells (from most positive to most negative):

[0275] HNFla - Strong Enhancer: Directly binds CFTR cis-elements and drives high expression. Critical for achieving robust transcription (particularly in submucosal / intestinal-type epithelial cells).

[0276] FOXA2 (and FOXA1 / 3) - Pioneer Activators: Essential for opening CFTR chromatin and enabling other activators. Greatly enhances transcription in any epithelial cell where they are present (intestinal, airway); acts broadly as a co-activator with HNF1.

[0277] FOXI1 - Cell-Type Specific Booster: Indirectly causes very high CFTR expression in ionocytes. When present, leads to orders-of-magnitude increases in CFTR (via ionocyte differentiation). Its impact is outsized in those few cells (which can contribute a significant fraction of total CFTR function).

[0278] HNF4a - Context-Dependent Activator: Can activate CFTR in cells that express HNF4 (gut, possibly submucosal glands). Not ubiquitous in airway, but in those cells, it would partner with HNF1 for strong transactivation. Moderate enhancer overall (strong locally where present).

[0279] FOXA3 - Redundant Pioneer: Likely contributes similarly to FOXA2 in tissues where it’s expressed. Helps maintain an open chromatin state. (Included with FOXA2 as a positive factor, though FOXA3 levels in lung may be lower).

[0280] NFIA - Neutral / Mild Enhancer: NFIA might assist basal transcription machinery or chromatin structure (NFIA can act as a transcriptional activator on some promoters). While not documented for CFTR, including an NFIA site (like the SV40 origin sequence) on plasmid might help sustain an open DNA conformation and could assist initiation. Any enhancement is probably minor compared to HNF / FOX. It is ranked roughly neutral to mildly positive in the CFTR context.

[0281] In practice, using multiple top-tier TF binding sites in tandem can synergize - e.g. a plasmid with both an HNF1 site and a FOXA2 site at an end could recruit a dimeric HNFla (2 NLS) and FOXA2 (2 NLS) to form a protein-DNA complex, greatly increasing the “cargo” of NLS for import. Such synergy can markedly enhance nuclear import efficiency. Known206678-0002-00WOsynergistic TF pairs are deliberately arranged next to each other to exploit co-activator effects. FOXA2 and HNFla sites sit adjacent to each other, leveraging FOXA2’s chromatin-opening to enable HNFla binding and transcriptional activation. Similarly, SPDEF and F0XQ1 sites are placed in tandem to cooperatively drive goblet cell-specific expression (SPDEF induces goblet cell differentiation, and F0XQ1 boosts mucin gene transcription). This arrangement ensures that secretory cells (club / goblet) have their key TFs directly engaging the promoter for maximum CFTR gene activation. In ciliated cells, the FOXJ1 and RFX3 binding sites are adjacent, because FOXJ1 (the ciliation master regulator) works with RFX3 as a co-activator to induce cilia-related genes. By placing FOXJ1-DTS immediately beside RFX3-DTS, this synergy is recreated on the CFTR enhancer - any ciliated cell containing the DNA will have both factors binding cooperatively to amplify CFTR promoter activity. These strategic pairings preserve critical TF-TF interactions that enhance transcription above what each factor could do alone. HNF4α and FOXA3 can be valuable additions to a CFTR plasmid vector for nuclear targeting and tissue-specific expression enhancement, but they are best used in combination with other complementary factors. An optimal design might include a multi-factorial enhancer (HNF4α+FOXA3±HNF1α sites together) to maximize CFTR gene activation.

[0282] Two nonlimiting examples of DNA construct layouts are as follows:Unidirectional Transcriptional Schematic Organization[Insulator] — (DTS Cluster 1) — [TATA Promoter] — [5' UTR] — [CFTR ORF] — [3' UTR & poly(A)] — (DTS Cluster 2) — [Insulator]Bidirectional Transcriptional Schematic Organization[Insulator] <- (DTS Cluster 3) <— [poly(A) & 3' UTR] <- [GENE2 ORF] <- [TATA Promoter] <- (Antisense Direction) (DTS Cluster 1) - [TATA Promoter] (Sense Direction) — > [5' UTR] — > [CFTR ORF] — > [3' UTR & poly(A)] — > (DTS Cluster 2) — > [Insulator]

[0283] Cluster 1 & 2 are designed to optimize for nuclear import first and foremost, then they are optimized for transcription enhancement, ideally across all CFTR-producing cell types. Cluster 3 design could contain TF-DTSs that enhance the transcription of GENE2.Example Construct #1v) Cluster 15′ ... → FOXA3 → HNF4α → NFIA... (CFTR promoter) ... CFTR ORF ... 3′

[0284] The construct is arranged to promote nuclear import first, followed by robust transcriptional activation of the CFTR gene. A Nuclear Factor I (NFIA) DNA-targeting sequence is positioned at the 5' end as a nuclear targeting sequence (DTS) - NFIA is a ubiquitous nuclear factor whose binding site (consensus TTGGCN5GCCAA) may help tether the plasmid for import into the nucleus. This order of TF-DTSs places pioneer factors like FOXA2 (Forkhead box A2) up front to maintain open chromatin, and key activators like HNFla (a master regulator of CFTR expression) early on. Additional sites for FOXJ1, FOXI1, FOXQ1, TP63, RFX3, SPDEF, FOXA3, and HNF4a are provided to maximize CFTR promoter activation in the target cells. Notably, FOXJ1 and RFX3 are critical for ciliated cell gene expression, while SPDEF and FOXQ1 drive goblet / secretory cell programs - inclusion of all ensures robust, tissuespecific CFTR transcription. Each TF-DTS is separated by a neutral spacer sequence engineered to avoid secondary structure, CpG dinucleotides, unintended motifs, or immunogenic patterns (no CpG islands or TLR9-activating motifs). These spacers also provide sufficient DNA backbone length so that each factor can bind its site without steric hindrance from neighboring sites.

[0285] Cluster 2 (Downstream of CFTR 3' UTR)CFTR 3′ UTR ... NFIA-DTS → FOXA2-DTS → FOXI1-DTS → TP63-DTS → SPDEF → FOXQ1 → FOXJ1-DTS → RFX3-DTS → FOXA2-DTS → HNF1α-DTS → FOXA3 → HNF4α-DTS... 3′ end of plasmidDesigning Optimized Spacer Sequences for TF-DTS Clusters

[0286] In designing spacer sequences, several general factors were accounted for to ensure the spacers are “invisible” facilitators that do not introduce new problems:

[0287] Neutral sequence composition: The spacer sequences are composed of DNA that does not bind known factors or form problematic structures. It is verified that each spacer lacks cryptic transcription factor binding sites that might inadvertently recruit repressors or activators. For example, repeats of known core motifs like “TATA” (which could act as a weak promoterelement) or “CpG islands” (which could attract DNA methylation) are avoided. The goal is for spacers to be genetic “insulators” that do not themselves influence gene expression.

[0288] Avoidance of cryptic splice sites: It is ensured that spacer sequences (especially those within transcribed regions, like the 3' UTR) do not contain sequences that resemble consensus splice donor or acceptor sites (e.g., “GT... AG” at intron boundaries). If a cryptic 5' or 3' splice site appears in the mRNA, the spliceosome might mistakenly use it, leading to aberrant splicing of the CFTR transcript. Cryptic splice sites are usually weak, but they can be recognized by the splicing machinery under certain conditions. To prevent this, the spacer design avoids the canonical dinucleotides and surrounding context that define splice sites. This helps ensure the transcript produced is the full-length CFTR mRNA with no unintended deletions.

[0289] Avoidance of secondary structures: Spacers are designed to minimize the formation of stable secondary structures (DNA hairpins or RNA hairpins after transcription). Extremely GC-rich spacer sequences could form strong stem-loop structures in the nascent RNA or cause the DNA to fold on itself, which might impede transcription or processing. Therefore, a balanced GC content is used, and long inverted repeats in the spacer sequences are avoided. By keeping the spacers relatively AT-rich or using alternating base patterns, flexibility in the DNA and a lack of stable RNA structure are promoted. This “bland” composition helps RNA polymerase read through the spacer smoothly during transcription and ensures that no structures sequester the poly(A) signal or other elements from the processing enzymes.

[0290] Spacer DNA sequences that accommodate both cooperative and independent binding of transcription factors (TFs) were designed:

[0291] Cooperative TF pairs - short spacers for synergy: Certain TFs in Cluster 1 are known to work together on DNA. For example, FOXA2 and HNFla often co-bind liver gene enhancers, with HNF1 sites found tightly clustered near FOXA2 sites. Similarly, FOXJ1 and RFX3 physically interact as co-activators on cilia gene promoters. For such cooperative TF pairs (including the given examples FOXA2-HNFla, FOXJ1-RFX3, SPDEF-FOXQ1), only a short spacer is inserted between their binding sites. A short spacer on the order of ~10 bp (roughly one DNA helical turn) keeps their motifs adjacent and on the same face of the DNA helix, facilitating simultaneous binding and protein-protein contacts. This close positioning is intended to enhance synergistic binding and transcriptional activation, as cooperative TF interactions are strongest when their binding sites are in close proximity. In gene regulation, the arrangement oftranscription factor binding sites around the DNA helix can impact how multiple transcription factors interact with each other. Depending on whether the sites are on the same side of the helix or on opposite sides, it can facilitate or hinder the formation of protein complexes. For Aldevron’s nanoplasmid DNA, which is a small, supercoiled plasmid, the DNA will assume B-form once inside the nucleus and likely become associated with histones similarly to genomic DNA. Thus, its helical periodicity will be in the same range (-10-10.5 bp per turn). Since halfbase spacing cannot be used, choosing 10 bp or 11 bp intervals achieves nearly one helical rotation. Empirical studies show 10.5 bp / turn is most stable for free DNA, so alternating 10 and 11 bp spacers could even be used to maintain phase over long distances.

[0292] Non-cooperative TFs - longer spacers to prevent hindrance: Not all TFs in the cluster will interact with each other, so for TF-DTS combinations that do not cooperate directly, a longer spacer is included to physically separate their binding sites. This prevents steric overlap or competition between factors that bind nearby. If two large proteins attempt to bind DNA sites that are immediately adjacent, one can sterically block the other’s access. In synthetic promoter studies, it’s observed that negative interference between TF binding sites is stronger at very short distances. Therefore, spacers of sufficient length (e.g. 20-30 bp) were used for non-cooperating TFs. Such a gap exceeds the footprint of a typical DNA-bound TF and ensures that one factor binding will not physically hinder another. By spacing these sites farther apart, each TF can bind independently without occluding the neighboring site, avoiding unwanted competitive effects. This balanced design allows each TF-DTS in Cluster 1 to contribute to import or transcription as intended, while maximizing cooperative pairs’ synergy and minimizing non-cooperative interference.

[0293] A spacer is needed between Cluster 1 and the CFTR core promoter (which contains the TATA box) to ensure proper transcription initiation:

[0294] Maintain core promoter accessibility: The CFTR minimal promoter (including the TATA box and transcription start site) must remain accessible to general transcription factors (like TATA-binding protein, TBP) and RNA polymerase II. Mammalian promoters are typically organized such that upstream regulatory motifs are positioned some distance away from the TATA box and initiation site. A neutral spacer sequence is therefore inserted between the end of the upstream DTS clusters and the beginning of the CFTR promoter region. This spacer acts as a buffer so that proteins bound in the clusters do not physically overlap or block the core promoter.206678-0002-00WOIn practice, a spacer on the order of a few tens of base pairs (for example, -30-50 bp) is used here. This length is roughly the size of the pre-initiation complex footprint around a TATA box (the TATA box is usually -30 bp upstream of the TSS in many promoters). By leaving this gap, TBP / TFIID is allowed to bind the TATAAA sequence unimpeded and assemble the transcription initiation machinery correctly.Spacer Design for Cluster 1

[0295] Spacers are designed to be AT-rich, avoiding secondary structures, splicing errors, CpG islands, unintended binding motifs, and immunogenic sequences.

[0296] Cooperative TF pair spacers: ATATATATAT (SEQ ID NO:97)

[0297] Neutral spacers: ATATATATATATATATATATATATA (SEQ ID NO:98)

[0298] Final spacer to promoter:ATATATATATATATATATATATATATATATATATATATATATATATATAT (SEQ ID NO:131).Spacer Sequence Arrangement Example

[0299] NFIA-DTS → [Neutral Spacer] → FOXA2 → [Cooperative Spacer] → HNF1α-DTS → [Neutral Spacer] → FOXJ1 → [Cooperative Spacer] → FOXI1-DTS → [Neutral Spacer] → FOXQ1 → [Neutral Spacer] → TP63-DTS → [Cooperative Spacer] → FOXJ1-DTS → [Cooperative Spacer] → RFX3-DTS → [Cooperative Spacer] → SPDEF-DTS → [Cooperative Spacer] → FOXQ1-DTS → [Neutral Spacer] → FOXA2-DTS → [Cooperative Spacer] → HNF1α-DTS → [Neutral Spacer] → FOXA3 → [Neutral Spacer] → HNF4α → [Neutral Spacer] → NFIA → [50 bp Spacer] → (CFTR Promoter)Total Base Pairs could approximately be estimated at 600bp for ClusterSpacer Design for Cluster 2

[0300] A spacer is included between the poly(A) signal and the first DTS (NFIA) of Cluster 25’ to 3’:

[0301] Allow proper 3' end processing: The poly(A) signal (often the hexanucleotide AATAAA in the DNA coding strand, which transcribes to AAUAAA in the mRNA) is crucial for cleavage and polyadenylation of the transcript. It works in concert with downstream sequence206678-0002-00WOelements to ensure the transcript is cut and a poly(A) tail added at the correct site. A spacer DNA segment upstream of the poly(A) signal is added to insulate it from the nearby DTS cluster. This spacer ensures that the poly(A) hexamer and the surrounding region are free of any bound proteins or secondary structures from the cluster that might interfere with recognition by the cleavage / polyadenylation machinery. In practice, a spacer of around ~50 bp before the poly(A) signal is a safe choice. That distance prevents any DNA-bound factors in the cluster from obstructing the assembly of cleavage factors (CPSF, CstF, etc.) at the poly(A) site. By keeping the poly(A) signal region clear, guarantee that the pre-mRNA’s 3' end will be correctly recognized and processed is assured, yielding a stable, polyadenylated mRNA.

[0302] Preserve mRNA stability: the spacer near the poly (A) site (as well as the DTS cluster sequences that lie in the 3’ UTR) is designed to avoid elements that could reduce mRNA stability. In particular, the spacer is ensured to not introduce any unintended polyadenylation signals or termination signals. The presence of an “AATAAA” sequence upstream of the real poly(A) site could act as a cryptic poly(A) signal, causing premature cleavage / termination of the transcript. The hexamer AATAAA (and close variants) is avoided in the spacer and cluster sequences (except at the actual CFTR poly(A) signal) to prevent this. Creating AU-rich sequences resembling AU-rich elements (AREs) within the 3' untranslated region is also avoided. AREs (e.g., ATTTA repeats in U-rich contexts) are known to target mRNAs for rapid degradation. By excluding such motifs, it is ensured to not inadvertently decrease the CFTR mRNA’s half-life. The spacer is designed to be neutral - it provides the necessary physical distance for processing, but does not contain any sequence that could disrupt 3' end formation or post-transcriptional stability of the mRNA.

[0303] Maintaining the order of DTS elements given above, spacer DNA sequences are inserted according to the design rules. Neutral spacers (~25 bp, AT-rich) are used between transcription factor sites that do not have known cooperative interactions, while short ~10 bp spacers are used for cooperative TF pairs. For example, the FOXJ1 and RFX3 sites are placed 10 bp apart since RFX3 acts as a co-activator with FOXJ1 in driving cilia gene expression.Similarly, the FOXA2 and HNFla sites, and the FOXA3 and HNF4a sites, are spaced 10 bp apart based on evidence that these factors co-occupy regulatory regions cooperatively. All spacer sequences were designed to avoid CpG dinucleotides (unmethylated CpGs can trigger immune responses) and to contain no cryptic polyadenylation signals (the canonical hexamer AATAAA206678-0002-00WOor its major variant ATTAAA). AU-rich instability motifs (e.g. “ATTTA” that mark mRNAs for degradation) are also avoided. The final ordered cluster sequence with each DTS and its adjacent spacer is detailed below.

[0304] Spacer (poly (A) to NFIA):AATTATTAATATTAATAATAATTAATTAATAATAATTAATAATAATTAATAA(AT-rich neutral spacer) (SEQ ID NO: 132)1. NFIA-DTS (15 bp): TTGGCAATAAGCCAA - Nuclear Factor I binding site (CFTR 3' UTR enhancer) (SEQ ID NO:6)Spacer (25 bp): TTAATATATTAATATTAATTATATA (neutral spacer) (SEQ ID NO: 99)2. FOXA2-DTS (8 bp): GTAAACAA – FOXA2 (HNF3β) binding motifSpacer (25 bp): ATATTAATTATTAATTAATTATTAA (neutral spacer) (SEQ ID NO: 100)3. FOXI1-DTS (8 bp): GTAAATAA - FOXI1 binding motif (forkhead family) Spacer (25 bp): AATATATTAATTAATATATTAATAT (neutral spacer) (SEQ ID NO:101)4. TP63-DTS (10 bp): TAACATGTTA - TP63 (p63) binding sequence (contains a CATGT core) (SEQ ID NO:31)Spacer (25 bp): TTATATTAATTATAATATTAATTAT (neutral spacer) (SEQ ID NO: 102)5. SPDEF-DTS (10 bp): TTAGGAATAT - SPDEF (ETS family) site containing a GGAA core) (SEQ ID NO: 19)Spacer (10 bp): TATATATTAT (neutral spacer) (SEQ ID NO: 103)6. FOXQ1-DTS (8 bp): GTAAATAA - FOXQ1 binding motif (forkhead family) Spacer (25 bp): TAATAATAATATATATTAATTATAT (neutral spacer) (SEQ ID NO: 104)7. FOXJ1-DTS (8 bp): GTAAACAA - FOXJ1 ciliary cell-specific forkhead motif Spacer (10 bp, cooperative: FOXJ1-RFX3 ): ATATATTATT (short spacer for cooperative interaction) (SEQ ID NO: 105)206678-0002-00WO8. RFX3-DTS (13 bp): GTTACCTAGTAAC - RFX3 “X-box” consensus motif for cilia genes (co-activator with FOXJ1) (SEQ ID NO: 37)Spacer (25 bp): AATTATATATTAATTAATATATATA (neutral spacer) (SEQ ID NO: 106)9. FOXA2-DTS (8 bp): GTAAACAA - FOXA2 binding site (repeated)Spacer (10 bp, cooperative: FOXA2-HNFla ): TATTATATTA (short spacer for cooperative interaction) (SEQ ID NO: 107)10. HNFla-DTS (14 bp): GGTTAATGATTAAC -HNFla pseudopalindromic site (binds as dimer) (SEQ ID NO:46)Spacer (25 bp): TAATATTAATATATTAATATAATTA (neutral spacer) (SEQ ID NO: 108)11. FOXA3-DTS (8 bp): GTAAATAA - FOXA3 (HNF3y) binding motif (forkhead family)Spacer (10 bp, cooperative: FOXA3-HNF4a ): AATATATTAA (short spacer for cooperative interaction) (SEQ ID NO: 109)12. HNF4a-DTS (13 bp): AGGTCATAGGTCA - HNF4a response element (DRl-type nuclear receptor site) (SEQ ID NO:49)Total sequence (Cluster 2 spacers + PTS elements)TTGGCAATAAGCCAATTAATATATTAATATTAATTATATAGTAAACAAATAT TAATTATTAATTAATTATTAAGTAAATAAAATATATTAATTAATATATTAATATTAA CATGTTATTATATTAATTATAATATTAATTATTTAGGAATATTATATATTATGTAAA TAATAATAATAATATATATTAATTATATGTAAACAAATATATTATTGTTACCTAGTA ACAATTATATATTAATTAATATATATAGTAAACAATATTATATTAGGTTAATGATT AACTAATATTAATATATTAATATAATTAGTAAATAAAATATATTAAAGGTCATAGG TCA (SEQ ID NO:121)

[0305] In some embodiments, the sequence comprises 2, 3 or 4 copies of the FOXA2 binding site (GTAAACAA).Optimization of 5' and 3' UTRs for CFTR Gene

[0306] Untranslated regions (UTRs) at the 5’ and 3’ ends of mRNAs play crucial roles in regulating mRNA stability, localization, and translation efficiency. These regions contain binding sites for RNA-binding proteins (RBPs) and micro RNAs (miRNAs), influencing mRNA half-life. Additionally, the secondary structures within UTRs, such as strong stem-loops, can cause ribosome pausing and disengagement, thereby negatively affecting protein translation.

[0307] The 5’ UTR is an important component of an mRNA as it affects mRNA stability and translation initiation. Determination of the most efficient 5’ UTR for translation efficiency remains a complex topic in mRNA studies. Most eukaryotic mRNAs contain a 5’ cap structure, where several initiation factors assemble and recruit the 40S ribosome subunit. The ribosome complex scans the 5’ UTR until it reaches the start codon, where the 60S ribosomal subunit is engaged and translation is initiated. Several studies have described that the presence, stability, and location of secondary structures in the 5’ UTR can affect translation efficiency in a gene-and cell-type-dependent manner. More specifically, the 5’ UTR of the human CFTR mRNA has been shown to contain sequences and secondary structures correlated with reduced protein expression when compared to the same mRNA lacking these features.

[0308] When selecting 5’ UTR sequences key considerations were adopted:

[0309] mRNA secondary structure: Stable secondary structures can cause RNA polymerase II to stall or slowdown, decreasing translation efficiency. They can also interfere with proper mRNA capping at the 5’ end of nascent transcripts, impacting mRNA stability, nuclear export, translation efficiency, and protection against exonucleolytic degradation. Thus, minimizing strong secondary structures in the 5 'UTR is important for efficient and sustained translation.

[0310] Upstream ORFs (uORFs): uORFs can divert scanning ribosomes and repress translation of the main ORF. The human wild-type CFTR 5' UTR naturally harbors a uORF that, even with a weak Kozak context, reduces CFTR translation initiation. Eliminating any uORFs in the 5 'UTR ensures ribosomes focus on the correct CFTR start codon, improving protein yield.

[0311] Kozak Sequence: The consensus around the start AUG (GCCRCCaugG (SEQ ID NO:246) in mammals, where R is a purine) has been shown to affect translation initiation efficiency. The most efficient Kozak sequence contains a purine at -3 and G at +4 positions, promoting correct start codon recognition. Accordingly, the CFTR 5' UTR should include an optimal Kozak motif immediately upstream of the AUG to maximize ribosome loading. Addition206678-0002-00WOof G at position +4 would lead to a point mutation in the CFTR protein. Therefore, this change is not adopted.

[0312] 5' UTR Length and Nucleotide Content: A compact 5'UTR (under ~ 50 nt) is preferred to limit ribosome scanning time and avoid extensive secondary structure. Many highly expressed genes naturally have short 5’ UTRs. For example, human a-globin mRNA has a 5' UTR of only ~ 37 nt, which is among the most efficient UTRs for translation. High GC-rich sequences (> 60 % GC) can form strong hairpins, whereas very AU-rich sequences can attract RNA-binding proteins that modulate stability (e.g. AU-rich elements). A moderate GC content (~ 40-50 %) balances these factors, avoiding stable structures while also preventing AU-rich instability motifs.

[0313] The 3’ UTR plays a crucial role in mRNA stability as it contains sequences important for mRNA cellular localization, binding sites for RBPs and miRNAs as well as polyadenylation (polyA) signal. The human CFTR mRNA contains a long 3’ UTR (> 1.5 kb). Several miRNAs and RBPs have been shown to fine-tune the expression of CFTR mRNA through mechanisms such as mRNA degradation, translation repression, and stabilization.

[0314] When selecting 3’ UTR sequences key considerations were adopted:

[0315] Polyadenylation Signal: The 3' UTR must contain a polyA signal (usually the hexamer AAUAAA or a variant) to ensure proper poly(A) tail addition. Efficient polyadenylation is critical fortranscription termination and for mRNA export, stability, and translation.

[0316] mRNA Stability Elements: Certain 3' UTR sequence elements bind proteins that stabilize the transcript. For instance, the human a-globin 3' UTR contains C-rich motifs that bind aCP / PCBP proteins, conferring unusually long mRNA half-life in erythroid cells. Incorporating such stabilizing elements in the 3' UTR can prolong CFTR mRNA lifespan. Conversely, destabilizing motifs (like AU-rich elements, AREs with the AUUUA sequence) should be avoided, as they recruit decay factors. In CFTR’s native 3' UTR, stretches of uridine (poly-U) or AUUUA motifs reduce mRNA abundance by recruiting negative RBPs and microRNAs. Thus, the design should exclude common AREs and similar destabilizing signals.

[0317] miRNA Binding Sites: miRNAs can strongly suppress gene expression by binding 3' UTRs and promoting mRNA degradation or translational arrest. CFTR’s 3' UTR is known to be targeted by several miRNAs that lower CFTR expression. A key design goal is to206678-0002-00WOeliminate seed sites for abundant miRNAs. Selecting sequences with minimal predicted miRNA binding should increase mRNA stability.

[0318] In summary, the UTRs were designed to maximize CFTR mRNA yield and longevity (transcriptional efficiency + stability) and maximize translation per mRNA (translational efficiency). This involves a synergy of features: a short, unstructured 5' sequence with an optimal Kozak for robust ribosome recruitment, and a 3' UTR with poly(A) signal, but lacking binding sites for miRNAs or decay factors. When both 5' and 3' elements are optimized and coordinated, they can greatly prolong protein expression. All these principles guide the specific UTR designs below.5’ UTRs

[0319] Human a-globin: Sourced from human a-globin (HBA1) mRNA. This 5' UTR is 40 nt long. The sequence below has been modified to include an optimal Kozak consensus (underlined). Human globin genes are known for conferring high, stable expression. Globin mRNAs are extremely abundant in erythroid cells, indicating their UTRs are optimized for efficient translation and stability. Researchers often borrow these UTRs for synthetic constructs.

[0320] Human p-globin: Incorporates the UTR from a highly expressed human globin gene (historically used to maximize mRNA stability / translation). This 5’ UTR is 50 nt long. The sequence below has been modified to include an optimal Kozak consensus (underlined).

[0321] Synthetic Chimeric UTRs: UTR sequences identified through high-throughput screening for maximum expression in human cells (Cao etal. 2021). Engineered high-performance UTR (NeoUTR-2 and CoNeoUTR2-3), derived from human sequences, selected after researchers screened thousands of variants and identified “NeoUTR-2” (106 nt long) and “CoNeoUTR2-3” (212 nt long) as a top 5' UTRs enhancing protein production. Both sequences contain an optimal Kozak consensus (underlined).Table 3: 5’ UTRs5’ UTRsSource SequenceNative (wildtype) GTAGTAGGTCTTTGGCATTAGGAGCTTGAGCCCAGACGGCCCTAGCAG GGACCCCAGCGCCCGAGAGACC (SEQ ID NO 240)(70 nts)Human a-globinACTCTTCTGGTCCCCACAGACTCAGAGAGAACCCGCCACC (SEQ ID NO:241)(40 nts)Human P-globin ACATTTGCTTCTGACACAACTGTGTTCACTAGCAACCTCAAACAGCCACC (SEQ ID NO: 242)(50 nts)NeoUTR-2 CACTCGCGCTGCCATCACTCTTCCGCCGTCTTCGCCGCCATCCTCGGCG CGACTCGCTTCTTTCGGTTCTACCAGGTAGAGTCCGCCGCCATCCTCCA(106 nts) CCGCCACC (SEQ ID NO 243)CACTCGCGCTGCCATCACTCTTCCGCCGTCTTCGCCGCCATCCTCGGCGCoNeoUTR2-3 CGACTCGCTTCTTTCGGTTCTACCAGGTAGAGTCCGCCGCCATCCTCCA CCCAACAACTTGTCTCGCTCCGGGGAACGCTCGGAAACTCCCGGCCGC(212 nts) CGCCACCCGCGTCTGTTCTGTTACACAAGGGAAGAAAAGCCGCTGCCGCACTCCGAGTGTGCCACC (SEQ ID NO 244)3' UTRs

[0322] Human p-globin: Sourced from human P-globin (HBB) gene. The P-globin 3'UTR (133 nts) contains the AAUAAA poly(A) signal (underlined). This 3' UTR has been widely described in the literature and historically known for conferring high, stable expression. Globin mRNAs are extremely abundant in erythroid cells, indicating their UTRs are optimized for efficient translation and stability.

[0323] Hybrid of AES and mtRNRl motifs: Extracted from von Niessen et al. (2019), it contains two human-derived stabilizers back-to-back. This consists of a fragment from the 3 'UTR of the human AES gene (Amino-terminal Enhancer of Split) and a fragment from mtRNRl (human mitochondrial 12S rRNA gene). Individually, these elements were top performers in a cellular library screen for 3 'UTRs that augment mRNA stability and expression; combined, they had a synergistic effect, yielding the highest protein levels in primary human dendritic cells. Interestingly, the authors noted AES and mtRNRl motifs had the fewest predicted miRNA sites. The poly (A) signal is embedded in the mtRNRl sequence (underlined).

[0324] Human a-globin: The a-globin (HBA2) contains several C-rich stability elements that can be bound by poly(C)-binding proteins (PCBPs). These elements are known to be crucial for the long half-life of a-globin mRNA (which remains stable even as reticulocytes differentiate). It naturally includes an AAUAAA signal (underlined).Table: 3’ UTRs3’ UTRsSource Sequence206678-0002-00WOAGAGCAGCATAAATGTTGACATGGGACATTTGCTCATGGAATTGGAGCTCG TGGGACAGTCACCTCATGGAATTGGAGCTCGTGGAACAGTTACCTCTGCCT CAGAAAACAAGGATGAATTAAGTTTTTTTTTAAAAAAGAAACATTTGGTAA GGGGAATTGAGGACACTGATATGGGTCTTGATAAATGGCTTCCTGGCAATA GTCAAATTGTGTGAAAGGTACTTCAAATCCTTGAAGATTTACCACTTGTGTT TTGCAAGCCAGATTTTCCTGAAAACCCTTGCCATGTGCTAGTAATTGGAAA GGCAGCTCTAAATGTCAATCAGCCTAGTTGATCAGCTTATTGTCTAGTGAAA CTCGTTAATTTGTAGTGTTGGAGAAGAACTGAAATCATACTTCTTAGGGTTA TGATTAAGTAATGATAACTGGAAACTTCAGCGGTTTATATAAGCTTGTATTCC TTTTTCTCTCCTCTCCCCATGATGTTTAGAAACACAACTATATTGTTTGCTAA GCATTCCAACTATCTCATTTCCAAGCAAGTATTAGAATACCACAGGAACCAC AAGACTGCACATCAAAATATGCCCCATTCAACATCTAGTGAGCAGTCAGGA AAGAGAACTTCCAGATCCTGGAAATCAGGGTTAGTATTGTCCAGGTCTACC AAAAATCTCAATATTTCAGATAATCACAATACATCCCTTACCTGGGAAAGGGNative CTGTTATAATCTTTCACAGGGGACAGGATGGTTCCCTTGATGAAGAAGTTG (wild-type) ATATGCCTTTTCCCAACTCCAGAAAGTGACAAGCTCACAGACCTTTGAACT (1,557 nts) AGAGTTTAGCTGGAAAAGTATGTTAGTGCAAATTGTCACAGGACAGCCCTT CTTTCCACAGAAGCTCCAGGTAGAGGGTGTGTAAGTAGATAGGCCATGGGC ACTGTGGGTAGACACACATGAAGTCCAAGCATTTAGATGTATAGGTTGATG GTGGTATGTTTTCAGGCTAGATGTATGTACTTCATGCTGTCTACACTAAGAG AGAATGAGAGACACACTGAAGAAGCACCAATCATGAATTAGTTTTATATGC TTCTGTTTTATAATTTTGTGAAGCAAAATTTTTTCTCTAGGAAATATTTATTTT AATAATGTTTCAAACATATATAACAATGCTGTATTTTAAAAGAATGATTATGA ATTACATTTGTATAAAATAATTTTTATATTTGAAATATTGACTTTTTATGGCAC TAGTATTTCTATGAAATATTATGTTAAAACTGGGACAGGGGAGAACCTAGGG TGATATTAACCAGGGGCCATGAATCACCTTTTGGTCTGGAGGGAAGCCTTG GGGCTGATGCAGTTGTTGCCCACAGCTGTATGATTCCCAGCCAGCACAGCC TCTTAGATGCAGTTCTGAAGAAGATGGTACCACCAGTCTGACTGTTTCCAT CAAGGGTACACTGCCTTCTCAACTCCAAACTGACTCTTAAGAAGACTGCAT TATATTTATTACTGTAAGAAAATATCACTTGTCAATAAAATCX’ATACATTTGT GTGAAA (SEQ ID NO: 251)Human GCTCGCTTTCTTGCTGTCCAATTTCTATTAAAGGTTCCTTTGTTCCCTAAGTC P-globin CAACTACTAAACTGGGGGATATTATGAAGGGCCTTGAGCATCTGGATTCTGC (133 nts) CTAATAAAAAACATTTATTTTCATTGCA (SEQ ID NO:252)CTGGTACTGCATGCACGCAATGCTAGCTGCCCCTTTCCCGTCCTGGGTACCC CGAGTCTCCCCCGACCTCGGGTCCCAGGTATGCTCCCACCTCCACCTGCCCHvbrid human CACTCACCACCTCTGCTAGTTCCAGACACCTCCCAAGCACGCAGCAATGCA AES / mtRNRl GCTCAAAACGCTTAGCCTAGCCACACCCCCACGGGAAACAGCAGTGATTA(278 nts)ACCTTTAGCAATAAACGAAAGTTTAACTAAGCTATACTAACCCCAGGGTTG GTCAATTTCGTGCCAGCCACACC (SEQ ID NO:253)Human GCTGGAGCCTCGGTGGCCATGCTTCTTGCCCCTTGGGCCTCCCCCCAGCCC a-globin CTCCTCCCCTTCCTGCACCCGTACCCCCGTGGTCTTTGAATAAAGTCTGAGT(111 nts) GGGCGGCA (SEQ ID NO:254)PTS and Enhancer DESIGNHelical Phasing Analysis

[0325] Transcription factor binding sites were positioned using 10.5 x N for integer N equal to or greater than 4. This spacing corresponds to exactly 8 helical turns (10.5 bp / turn), ensuring transcription factors bind on the same face of the DNA helix, facilitating cooperative protein-protein interactions and synergistic transcriptional activation.

[0326] Key Design Features include: 1) 42bp x M spacing between motif centers (4 helical turns per interval). M can be any integer. Multiple motifs can be on the same helix phase.2) Two DNA Targeting Sequences (DTS): FOXI1 and FOXA2. 3) Four enhancer elements targeting distinct lung cell populations. 4) Intentional ionocyte redundancy (FOXI1 + ASCL3) for robust targeting. 5) Pioneer factor (FOXA2) positioned proximally (-84) for chromatin opening. 6) A promoter (e.g., Super Core Promoter (SCP)) at TSS for basal transcription.Promoter Integration Requirements

[0327] Although a Super Core Promoter (SCP) is used here, other core promoters may be used as well.1. The Anchor Principle: Position -31 and Phasing

[0328] The Anchor (31.5 bp): The distance from the first T of the TATA box (the physical docking point for the TBP / TFIIB complex) to the Geometric Center of the proximal motif is set at exactly 31.5 bp.

[0329] Rotational Facings: This distance represents exactly 3.0 helical turns (10.5x3), ensuring the activator and the basal machinery are perfectly aligned on the DNA helix.2. The Periodicity Principle: 42-bp Harmony

[0330] The Repeat: Every subsequent upstream motif is spaced exactly 42 bp from the center of the previous motif.

[0331] Integer Perfection: This represents 4.0 helical turns, maintaining a consistent rotational phase across the entire cluster so every protein faces the same direction.

[0332] Steric Clearance: This spacing provides approximately 4 nm of physical space between motifs, which is a reasonable safety margin to prevent large protein complexes (like Mediator) from bumping into each other.3. Geometric Center vs. Motif Start

[0333] Why the Center? Transcription factors are 3D molecules that "straddle" the DNA. The geometric center is the most reliable coordinate for the protein's center of mass.

[0334] Consistency: Because the motifs vary in length (e.g., FOXI1 is 7 bp, while NFI is 15 bp), anchoring to the center ensures that the protein itself is always rotationally centered, regardless of the motifs footprint size.4. Addressing Variability and DNA Pitch

[0335] Natural Flexibility: it is acknowledged that natural promoters often function with non-integer phasing and variable spacing.

[0336] Synthetic Optimization: the choice of a rigid 10.5 bp / turn pitch and the 42-bp / 31.5-bp rules is an engineering choice intended to maximize recruitment velocity by reducing the energetic cost of DNA twisting.Lung-Specific ConstructDesign Rationale

[0337] Airway epithelium exhibits remarkable cellular heterogeneity, with distinct transcriptional programs governing ionocytes (CFTR regulation), club cells (surfactant production), and secretory cells (mucin synthesis). The design leverages master regulators from each lineage to achieve pan-airway specificity while maintaining high expression levels.Table 5: TFBS (Motif) ArchitectureRole Sequence TF Cell Target DTS GTAAACA FOXI1 Ionocytes ENHANCER CCAGTTCAAC (SEQ ID NO:23) TFCP2L1 Secretory ENHANCER GCACCTGCC ASCL3 Ionocytes ENHANCER TTGGCNNNNNGCCAA (SEQ ID NFIA Club Cells NO:95), where N can be any nucleotideExample:TTGGCATAATGCCAA (SEQ ID NO:96)ENHANCER ATGCGGGT SPDEF Goblet Cells DTS ATGTAAACATA (SEQ ID NO:9) FOXA2 Pan-airway206678-0002-00WORole: DTS = DNA Targeting Sequence (nuclear import); TF ENHANCER = Transcription Factor binding siteAssembly of Motifs and SCP, and Position Rationale

[0338] The following motifs were used.ccaaaggttttccaaaggttttc [spacer] (SEQ ID NO: 110)GTAAACA (FOXI1-DTS)aaaggtttccaaaggtttccaaaggtttccaaa [spacer] (SEQ ID NO:111) CCAGTTCAAC (TFCP2L1) (SEQ ID NO:23) aaaggtttccaaaggtttccaaaggtttccaaa [spacer] (SEQ ID NO:111) GCACCTGCC (ASCL3)aaaggtttccaaaggtttccaaaggtttc [spacer] (SEQ ID NO: 112) TTGGCATAATGCCAA (NFIA) (SEQ ID NO:96) aaaggtttccaaaggtttccaaaggtttcc [spacer] (SEQ ID NO: 113)ATGCGGGT (SPDEF)aaaggtttccaaaggtttccaaaggtttccaaa [spacer] (SEQ ID NO:111) ATGTAAACATA (FOXA2-DTS) (SEQ ID NO: 9)

[0339] The combined sequence from above is as follows:5’ccaaaggttttccaaaggttttcGTAAACAaaaggtttccaaaggtttccaaaggtttccaaaCCAGTTC AACaaaggtttccaaaggtttccaaaggtttccaaaGCACCTGCCaaaggtttccaaaggtttccaaaggtttcTTGGCATA ATGCCAAaaaggtttccaaaggtttccaaaggtttccATGCGGGTggaggtttccaaaggtttccaaaggtttccaaaATGT AAACATA-3’ (SEQ ID NO: 122)

[0340] The combined sequence from above with a short spacer between the DTS cluster and an example core promoter (underlined) is as follows:5’ccaaaggttttccaaaggttttcGTAAACAaaaggtttccaaaggtttccaaaggtttccaaaCCAGTTC AACaaaggtttccaaaggtttccaaaggtttccaaaGCACCTGCCaaaggtttccaaaggtttccaaaggtttcTTGGCATA ATGCCAAaaaggtttccaaaggtttccaaaggtttccATGCGGGTggaggtttccaaaggtttccaaaggtttccaaaATGT AAACATAaaaggttccaaaggtttccGCACGCCTATAAAAGcagacctcgcatcgatctacaTCAGTTatcacac gacatcCGAACGGAACAGTCGC-3' (SEQ ID NO:123)206678-0002-00WO

[0341] The combined sequence from above with a longer spacer between the DTS cluster and the core promoter (underlined) is as follows:5’ccaaaggttttccaaaggttttcGTAAACAaaaggtttccaaaggtttccaaaggtttccaaaCCAGTTC AACaaaggtttccaaaggtttccaaaggtttccaaaGCACCTGCCaaaggtttccaaaggtttccaaaggtttcTTGGCATA ATGCCAAaaaggtttccaaaggtttccaaaggtttccATGCGGGTggaggtttccaaaggtttccaaaggtttccaaaATGT AAAC AT AaaaggtttccaaaggtttccaaaggtttccaaaggtttccGC AC GC C T AT A A A AGcagacctc gcatcgatct acaTCAGTTatcacacgacatcCGAACGGAACAGTCGC-3’ (SEQ ID NO: 124)

[0342] In B-DNA, the helix completes a full 360° turn every 10.5 base pairs. To achieve 'Turbo' performance, every recruitment hub must land on an integer turn (n) relative to the TATA box by optimizing Spacer 1 or 2.Table 6TF Identity Center Dist. from - Turns (n)(Rel. 31TSS)TATA Box -31.0 0.0 bp 0.0FOXA2-DTS -62.5 31.5 bp 3.0SPDEF -104.5 73.5 bp 7.0NFI -146.5 115.5 bp 11.0ASCL3 -188.5 157.5 bp 15.0TFCP2L1 -230.5 199.5 bp 19.0FOXI1-DTS -272.5 241.5 bp 23.0

[0343] Another sequence where TFCP2L1 is removed is provided as below ccaaaggttttccaaaggttttc [spacer] (SEQ ID NO: 110)GTAAACA (FOXI1-DTS)aaaggtttccaaaggtttccaaaggtttccaaa [spacer] (SEQ ID NO:111) GCACCTGCC (ASCL3)aaaggtttccaaaggttttccaaaggtttc [spacer] (SEQ ID NO:116) TTGGCATAATGCCAA (NFIA) (SEQ ID NO:96) aaaggtttccaaaggtttccaaaggtttcc [spacer] (SEQ ID NO: 113)206678-0002-00WOATGCGGGT (SPDEF)aaaggtttccaaaggtttccaaaggtttccaaa [spacer] (SEQ ID NO:111) ATGTAAACATA (FOXA2-DTS) (SEQ ID NO: 9)

[0344] The combined sequence from above is as follows:5’ccaaaggttttccaaaggttttcGTAAACAaaaggtttccaaaggtttccaaaggtttccaaaGCACCT GCCaaaggtttccaaaggtttccaaaggtttcTTGGCATAATGCCAAaaaggtttccaaaggtttccaa aggtttccATGCGGGTggaggtttccaaaggtttccaaaggtttccaaaATGTAAAC ATA-3’ (SEQ ID NO: 125)

[0345] The combined sequence from above with a short spacer between the DTS cluster and an example core promoter (underlined) is as follows:5’ccaaaggttttccaaaggttttcGTAAACAaaaggtttccaaaggtttccaaaggtttccaaaGCACCT GCCaaaggtttccaaaggtttccaaaggtttcTTGGCATAATGCCAAaaaggtttccaaaggtttccaa aggtttccATGCGGGTggaggtttccaaaggtttccaaaggtttccaaaATGTAAACATAaaaggtt ccaaaggtttccGCACGCCTATAAAAGcagacgtcgcatcgatctacaTCAGTTatcacacgaca taCGAACGGAACAGACGT -3’ (SEQ ID NO: 126)

[0346] The combined sequence from above with a longer spacer between the DTS cluster and the core promoter (underlined) is as follows:5’ccaaaggttttccaaaggttttcGTAAACAaaaggtttccaaaggtttccaaaggtttccaaaGCACCT GCCaaaggtttccaaaggtttccaaaggtttcTTGGCATAATGCCAAaaaggtttccaaaggtttccaa aggtttccATGCGGGTggaggtttccaaaggtttccaaaggtttccaaaATGTAAACATAaaaggttt ccaaaggtttccaaaggtttccaaaggtttccGCACGCCTATAAAAGcagacgtcgcatcgatctacaT CAGTTatcacacgacataCGAACGGAACAGACGT -3’ (SEQ ID NO: 127)Tunable and Differential Autoregulation of Endogenous Mutant and Transgene CFTR Gene Expressions Using Endogenous Repressors

[0347] This DNA construct design addresses two key problems with regard to CFTR gene therapy with the following objectives: (1) Mutant CFTR repression (-80% knockdown) -206678-0002-00WOreducing the patient’s endogenous mutant CFTR expression to minimize deleterious effects, and (2) Transgene copy -number control (-50-150% of normal levels) - preventing overexpression when multiple CFTR transgene copies enter a cell, by automatically dialing back expression to safe levels.

[0348] Certain autoregulatory promoter designs directly control the CFTR gene expression via its own transcription feedback loop. In practice, this can be achieved by embedding operator sequences or response elements in the CFTR promoter that are recognized by a CFTR-inducible regulator. For example, the CFTR expression cassette could co-express a transcriptional repressor that binds the CFTR promoter when CFTR levels rise, dampening further transcription (a form of negative autoregulation). This creates a self-limiting circuit: as CFTR protein accumulates, transcription is automatically dialed down, preventing runaway expression. Negative feedback of this sort is a common motif in natural gene networks and has been shown to accelerate response times and stabilize gene output in synthetic circuits.

[0349] These are many human transcription factors known to function as transcriptional repressors in their native contexts. In fact, a network of transcription factors (TFs) keeps CFTR mRNA levels in check. For example, an siRNA screen in airway cells found 40 TFs whose depletion increased CFTR expression >2x.

[0350] Of those 40 TFs, two native TF repressors stand out in the CFTR regulatory system: Interferon Regulatory Factor 2 (IRF2) and ETS Homologous Factor (EHF).

[0351] 1. IRF2 (Interferon Regulatory Factor 2)Chromosomal Location: Chr4q35.1Genomic Coordinates: 4:186,314,567-186,317,986 (GRCh38)Size: -3.4 kb (including UTRs, exons, and introns)Expression in Lung Epithelium: IRF2 is moderately expressed in airway epithelial cells. It binds directly to the -35 kb CFTR enhancer, where it competes with IRF1 to repress CFTR transcription.Regulatory Elements: IRF2 itself is controlled by IFN-responsive elements (e.g., STAT 1 -binding sites in its promoter), but in non-inflammatory conditions, it is constitutively expressed in epithelial cells. No strong lung-specific enhancers aredirectly linked to IRF2, suggesting its expression is relatively stable across tissues.

[0352] 2. EHF (ETS Homologous Factor)Chromosomal Location: Chr11p13Genomic Coordinates: 11:33,430,218-33,443,267 (GRCh38)Size: -13 kb (including introns, exon-only -720 bp)Expression in Lung Epithelium: Highly expressed in lung epithelial cells, particularly in ciliated, goblet, and basal cells. It binds to the CFTR promoter and the -35 kb enhancer, recruiting HDACs to repress transcription.Regulatory Elements: EHF is activated by airway epithelial differentiation factors, meaning its expression is stable in lung cells but varies under inflammatory conditions.

[0353] IRF2 CFTR Repression Mechanism: IRF2 binds an interferon-responsive element in the CFTR locus (e.g., an airway-selective enhancer — 35 kb upstream) and functions as a transcriptional repressor. It competes with the activator IRF1 for this binding site; IRF1 occupancy boosts CFTR transcription, whereas IRF2 binding represses it. Constitutive IRF2 at the enhancer keeps CFTR expression low until signals (e.g., interferon) raise IRF1 levels, which displace IRF2 and activate CFTR expression. IRF2-mediated repression also influences chromatin structure: in complex with NF-Y, IRF2 helps recruit histone-modifying enzymes (e.g., the co-repressor SIN3A and methyltransferase SETD7) to the enhancer. This interaction promotes a less transcriptionally active chromatin state, antagonizing IRF1 -driven chromatin opening and gene activation.

[0354] Safety Profile of IRF2 Modulation: IRF2 is constitutively expressed in epithelial cells and immune cells, where it acts to temper excessive inflammation. By antagonizing IRF1-driven cytokine and interferon production, IRF2 helps maintain epithelial integrity and prevents inflammation-induced tissue damage or fibrosis. Its regulation of immune signaling also limits chronic oxidative stress in tissues, while still permitting necessary oxidative bursts for pathogen clearance. Generally, IRF2’s activity is associated with controlled immune responses and reduced inflammatory pathology. However, extreme modulation of IRF2 could pose risks:206678-0002-00WOoveractive IRF2 in non-target cells might repress needed immune defenses (risking opportunistic infection due to blunted interferon responses), whereas insufficient IRF2 function could allow prolonged IRF1 activity and chronic inflammation. Overall, IRF2’s influence is immunomodulatory rather than overtly toxic, supporting epithelial health by preventing unchecked inflammation.

[0355] If a TF-DTS nuclear import mechanism is utilized in the DNA design, then it is important that the IRF2 coding sequence itself does not contain any known DTS that binds with another transcription factors that could drive off-target nuclear import to non-CFTR producing cell types.

[0356] In summary, IRF2 fulfills the key criteria for a CFTR-repressing TF with a favorable safety profile.

[0357] EHF CFTR Repression Mechanism: Ets homologous factor (EHF) binds to cis-regulatory elements of the CFTR gene and modulates its transcription. In airway epithelial cells, EHF and KLF5 co-occupy an enhancer — 35 kb upstream of the CFTR promoter. At this site EHF functions as a transcriptional repressor: its binding (often alongside KLF5) recruits repressive cofactors and alters local chromatin to reduce CFTR gene activity. EHF’s occupancy at the -35 kb element and other CFTR regulatory regions is thought to impede enhancerpromoter looping or otherwise impose a less accessible chromatin configuration, thereby dampening CFTR expression.

[0358] Safety Profile of EHF Modulation: Epithelial Integrity & Fibrosis: EHF is essential for normal epithelial barrier function. In primary bronchial cells, EHF depletion impaired wound closure, indicating its role in epithelial repair. Proper EHF levels help maintain epithelial phenotype - loss of EHF can trigger epithelial -to-mesenchymal transition (EMT), a process linked to fibrotic remodeling in lung tissue. Thus, dysregulation of EHF may compromise epithelial integrity and promote fibrosis.

[0359] Inflammation & Stress Response: EHF helps regulate airway inflammation by restraining pro-inflammatory genes. It directly targets cytokine / chemokine loci; for example, EHF normally represses the neutrophil chemoattractant IL-8, and EHF knockdown leads to elevated IL-8 levels, which can drive chronic neutrophilic inflammation in the airway. EHF also modulates stress-response pathways, counteracting AP-1 -driven transcription and other oxidativestress-activated signals. In sum, adequate EHF activity is important for controlling inflammation and oxidative stress in epithelial cells, and its dysregulation could lead to chronic inflammatory damage or impaired immune homeostasis.

[0360] Similarly, if a TF-DTS nuclear import mechanism is utilized in the DNA design, then it is important that the EHF coding sequence itself does not contain any known DTS that binds with another transcription factors that could drive off-target nuclear import to non-CFTR producing cell types.

[0361] In summary, EHF fulfills the key criteria for a CFTR-repressing TF with a favorable safety profile.

[0362] One implementation could use a four-component DNA system: 1) therapeutic transgene gene for functional rescue; 2) repressor gene to control transgene expression; 3) Core promoter(s) (e.g., basic TATA box or Turbo TATA), engineered to drive transcription of both genes simultaneously; 4) Regulatory Region. The regulatory region may include enhancers, operators and decoys. The enhancers may contain one or more transcription binding sites to enhance transcription. The operators may contain one or more binding sites for the repressor proteins. Decoys are the unmutated high-affinity repressor binding sites. The repressor gene, the operators and decoys form an autoregulatory feedback loop. As CFTR (and thus the repressor) accumulates, the repressor binds the operator sites and the endogenous binding sites to slow further the transcription of CFTR native gene and transgene, and repressor gene. In effect, CFTR expression self-titrates: low CFTR output means little repression (promoter stays active), whereas high CFTR output increases repression (promoter activity is curbed). This dynamic adjustment could keep CFTR around an optimal setpoint.

[0363] Both IRF2 and EHF could fulfill the repressor role. They act in an allele-agnostic manner: they will bind to any DNA (endogenous chromosome or plasmid) containing their target sequences. This will silence native CFTR genes as well as transgenes:

[0364] It is desirable to differentially control the transcription of mutant and transgene CFTR. Specifically, there is preference to repress mutant CFTR expression while mildly modulating transgene expression. Different mechanisms can be deployed individually or in combination to differentially repress endogenous mutant CFTR expression and modulate CFTR Transgene expression. These mechanisms include:

[0365] 1. One mechanism is to introduce point mutations in IRF2 and EHF binding sites (the operators) within the CFTR transgene promoter to weaken their binding affinity.

[0366] 2. There are levers to fine tune the repression and modulation levels of mutant and transgene CFTR expressions on the DNA layout. For instance, the distance between the IRF2 / EHF binding site and the core promoter can vary to adjust for the targeted expression level.

[0367] 3. Decoys placed away from the promoter can be used to sequester nearby repressor proteins, thus reducing the binding likelihood of repressors to the operator site.

[0368] To achieve selective repression of endogenous mutant CFTR while modulating tunable transgene expression, repressor binding motifs are strategically modified. High-affinity EHF and IRF2 binding sites are preserved in the endogenous CFTR locus, ensuring strong repression, while the transgene promoter is engineered with subtle point mutations in these sites to reduce binding affinity. This differential binding strategy allows endogenous CFTR to be silenced efficiently (-80% repression) while ensuring transgene expression remains in the therapeutic range (-50-150%) even if multiple copies of transgene are delivered.

[0369] Through specific promoter binding site mutations, it is intended to fine-tune the binding affinities of IRF2 and EHF to achieve an optimal balance: roughly 80% repression of endogenous CFTR (mutant and any wild-type allele) to eliminate interference, while allowing the CFTR transgene to produce -50-80% of normal CFTR levels for a single copy gene delivery, or <150% of normal CFTR levels in the case of multiple copy gene delivery.

[0370] Special attention is required to ensure that the point mutations in the IRF2 and EHF binding sequences do not inadvertently introduce new DNA nuclear targeting sequences (DTSs), splicing sites, or other transcription factor binding sites. Given that both IRF2 and EHF are transcription factors with nuclear localization signals (NLSs) present in nearly all lung epithelial cell types, the mutated binding sequences are less likely to act as DTSs for nuclear import. Therefore, the additional benefit of these point mutations could potentially reduce the likelihood of unintended nuclear import of DNA into non-CFTR producing cells.

[0371] Sequestration of Repressors: Decoys can be used to further refine repression specificity, high-affinity decoy binding sites placed away from the promoter can be strategically introduced to competitively sequester excess EHF and IRF2, thereby reducing their effective concentration at the operator sites. By varying the spacing with the promoter and sometimes in repeats, these decoy binding sites act as molecular sponges, preferentially attracting EHF orIRF2 away from the operator sites that govern transcription. This provides an additional layer of regulatory control, ensuring that CFTR expression remains finely tuned within the desired therapeutic range.

[0372] The strategic point mutations to the consensus IRF2 (ISRE) and EHF (ETS) motifs in the operators weaken their binding strengths with repressor IRF2 and EHF, relative to how they preferentially bind the endogenous CFTR locus. This way, IRF2 / EHF still bind and exert some repression on the transgene (preventing overexpression), but not as strongly as they bind the unmodified endogenous sites.EHF and IRF2 Binding Motifs in the CFTR Locus

[0373] In airway epithelial cells, a critical enhancer ~35 kb upstream of the CFTR promoter (the “-35 kb DHS” site) recruits a complex of transcription factors including IRF1 / IRF2 and EHF. IRF1 is an activator at this site, whereas IRF2 is a repressor that antagonizes IRFl’s function.

[0374] EHF (Ets Homologous Factor) is an ETS-family transcription factor that also binds this enhancer and acts as a repressor of CFTR in the lung.

[0375] Both IRF2 and EHF binding motifs lie in the CFTR promoter / enhancer region and contribute to downregulating CFTR in the airway.

[0376] IRF2 Motif: IRF2 Binding Sequence: Interferon Regulatory Factor 2 (IRF2) recognizes the interferon-stimulated response element (ISRE) motif. The consensus DNA sequence for IRF1 / IRF2 binding is 5'-AANNGAAA-3' (with slight variability at the last two positions) in lung epithelial cells ( — 35 kb upstream of the promoter).

[0377] IRF2 (a constitutive repressor) and IRF1 (an inducer) both target this site - IRF2 competes with IRF1 for binding to the site, thereby repressing CFTR transcription.

[0378] Mild core mutation: Another approach is to change one of the adenines in the GAAA core to a different base that is less detrimental than G— > C. For instance, GAAA — GAGA or GAAA — > GATA. In these examples, the core still starts with “GA” but isn’t the perfect “GAAA” anymore. Such a mutation might still permit a very weak interaction with IRFs because part of the motif is recognizable (e.g. “GAGA” contains two “GA” repeats which an IRF might half-heartedly bind). However, it would no longer support strong binding. Experimental studies of IRF sites indicate that mutating even one of the three A’s markedly lowersbinding / activity - multiple GAAA sub-sites in promoters are often redundant, and loss of one makes that site nearly inactive. By keeping the G in place and altering an A, that a completely foreign sequence is not introduced, just a suboptimal one. This could allow, say, 10-20% of the original IRF binding instead of 0%, giving a tunable (much lower) repression capability.

[0379] Partial motif alterations: In light of the potential drastic reduction in binding affinity by mutating the core GAAA, one of the flanking bases or one letter of the core is changed to reduce IRF2’s affinity.

[0380] Alter flanking “AA The consensus starts with AA.

[0381] IRF2 site (original in CFTR enhancer: e.g. 5'-AAGTGAAA-3) - mutate to 5'- AGGTGAAA-3'. This example changes one of the leading A’s to G (A^G at position 2) and the second nucleotide of “N N” to G (making the flanking context less ideal), without altering the GAAA core. The result is a weakened IRF site (IRF2 binding reduced, IRF1 might still recognize GAAA). (Note: actual flanking sequence may vary; the principle is to disrupt the 5'-AA while keeping GAAA).

[0382] IRF2’s binding to DNA can involve contacts with these 5' adenines as well as the core. Changing one of these adenines to another nucleotide (G, C, or T) would produce a sequence like GANNGAAA or TANNGAAA (depending on which A is changed). For example, if the native site in the enhancer is AATCGAAA (hypothetically), one might mutate it to AGTCGAAA - here the second position A— G, leaving “GAAA” intact. This deviates from the IRF consensus at the flank but keeps the GAAA core unchanged. Without being bound by a particular theory, it is expected that this significantly weakens IRF2 binding: IRF2 prefers two A’s in those positions, so a G or T substitution reduces its binding efficiency. Importantly, IRF1 (the activator) primarily recognizes the GAAA part; it may still bind somewhat to the mutated site (since GAAA is present), possibly allowing a bit of enhancer activation to continue. In this way, the balance shifts - IRF2 can’t bind well to repress, but IRF1 might still provide some activation if needed. This creates a selective advantage for the transgene’s site: it’s no longer a good IRF2 target, whereas the endogenous site (still “AANNGAAA”) remains strongly repressive.

[0383] EHF Motif: EHF is a member of the ETS transcription factor family, which characteristically recognize DNA sequences centered on a 5'-GGAA-3' core in lung epithelial cells ( — 35 kb upstream of the promoter).

[0384] In many ETS sites, the core is 5'-GGAA-3' (sometimes 5'-GGAT-3'; collectively written as GGAA / T).

[0385] EHF Site Mutation:

[0386] Conservative core substitution: If there is preference to tweak the core itself, a more conservative change than GTAA / GGTA can be tried.

[0387] GGAA — GGAT: This swaps the last A for T. Many ETS factors (including EHF) tolerate a terminal T almost as well as A. So this change might only slightly reduce EHF binding affinity (depending on EHF’s exact preference, it might even bind GGAT nearly as well as GGAA). If EHF has a slight bias for A, then GGAT would be bound somewhat more weakly - providing a moderate reduction rather than a complete loss. (For context, some ETS factors actually prefer GGAT; but given EHF’s consensus is often noted as GGAA, it is suspected GGAT is somewhat less favored by EHF, making this a minor attenuation strategy.)

[0388] GGAA GGAG (i.e. change the fourth base to G): “GGAG” is not a typical ETS site - most ETS proteins show poor binding if the last position is G or C, preferring A / T. This mutation would likely reduce EHF binding substantially, but possibly not quite as completely as breaking the GG pair. The site still starts with “GGA”, which means the ETS domain’s major groove contacts at the first three positions could form, but the mismatch at the 4th position could destabilize the complex. EHF might bind “GGAG” weakly or transiently. This could be useful if just a trace of residual binding for tuning is desired. (Another similar option is GGAA — > GGAC; a C at the fourth position would likewise be highly unfavorable for EHF. Either G or C breaks the preferred A / T at that position.)

[0389] GGAA GGGA (change the third base to G): This gives “GGGA”. Here the double G is intact and the last A is intact, but there are three G’s in a row. ETS factors generally require a 5'-GG and then an A or T at the 3rd position. “GGG” in positions 1-3 is off-consensus (the core would be GGG*A*) Without being bound by a particular theory, it is expected that EHF’s affinity for GGGA is much lower than for GGAA, because the extra G at position 3 interferes with the specific hydrogen bond network an ETS domain usually forms with an A / T at that position. However, since the site still has GGAA, EHF is not blind to it - it may still bind to GGGA if present at high levels, albeit weakly. This mutation thus creates a low-affinity version of the ETS site.

[0390] The ETS motif bound by EHF (with core GGAA) will be similarly altered. For example, a native ETS-like sequence 5'-ACCGGAAGT-3' could be mutated to 5'-ACCGCTAGT-3', changing the core GGAA to GCTA. This disrupts the invariant GGAA / T element required for ETS factor binding.

[0391] Partial motif alterations: To weaken EHF binding without completely abolishing it, the ETS site can be made suboptimal but not unrecognizable:

[0392] Change flanking context: EHF and other ETS factors often have preferences for bases immediately flanking the core GGAA. In particular, AT-rich flanks enhance binding affinity for the epithelial ETS factors. If the native sequence around the EHF site is, for example, 5'-TGGAAA-3' (T at the 5' flank and A at the 3' flank, both favorable), those flanks can be altered to reduce affinity. Changing the 5' flank to C or G and / or the 3' flank to C or G would create a less hospitable context (e.g. CGGAAC). This keeps the GGAA core unchanged, so in principle ETS proteins can still bind, but the neighboring nucleotides no longer optimally support EHF binding. The result is a weaker EHF-DNA interaction. EHF might still bind at high concentration, but its occupancy will be much lower than at the endogenous site (which has the ideal flanks). This kind of mutation is useful because it doesn’t touch the core motif at all -meaning other ETS family members or cooperative factors might still bind there if needed, just with less frequency.

[0393] EHF site (original: e.g. 5'-TGGAAAT-3) - mutate to 5'-CGGAAGT-3'. This changes a T^C at the 5' flank and A^G at the 3' flank in this hypothetical sequence, plus changes the last A of the core to G (“GGAA” >“GGAG”). The core GG is intact and “GGA_” mostly intact, but the context is now GC-rich and the last base is non-consensus, yielding a low-affinity ETS site. EHF binding here will be much weaker than to the native 5'-TGGAAAT-3' sequence, but the site could still, in theory, bind an ETS protein weakly if high levels are present. This ensures the transgene doesn’t attract EHF under normal conditions, yet the general enhancer DNA sequence is there for other factors to use.

[0394] Three repressor circuit options are proposed — EHF, IRF2, and a combination of both — to potentially repress mutant CFTR expression while modulating transgene CFTR output. EHF offers a linear, proportional repression, ensuring steadier control over transgene activity, but may not be able to control overexpression effectively in the presence of high transgene copy numbers in the nucleus. In contrast, IRF2, exhibiting cooperative (Hill-type) repression, kicks into stabilize CFTR expression at high transgene copy numbers, providing a nonlinear, protective feedback against transgene overexpression. When combined, EHF and IRF2 provide robust modulation of transgene expression across a range of transgene copy numbers.Combined IRF2 + EHF

[0395] Combining IRF2 and EHF could offer synergistic control over CFTR transcription. IRF2 effectively shuts down transcription past a threshold, silencing mutant CFTR alleles and preventing runaway transgene expression in high-copy scenarios. EHF provides continuous basal repression, smoothing the overall response. This dual-repressor setup creates a robust feedback system where EHF establishes baseline control, and IRF2 introduces a nonlinear, high-gain response to curtail excessive rises in CFTR levels.

[0396] IRF2's moderation avoids the binary on / off behavior associated with a high Hill coefficient, opting instead for a coefficient of ~2-3. This provides a steep response to prevent overshoot without causing bistability or oscillations. EHF contributes a proportional control, enhancing system stability by tempering fluctuations.

[0397] With only EHF, CFTR expression would decrease smoothly with increasing plasmid copy or activating signals, but may not clamp down sufficiently under high-driving forces. Conversely, IRF2 alone maintains expression near a set point until surpassed, potentially leading to strong repression that could cause oscillatory behavior.

[0398] The combined EHF and IRF2 approach quickly dampens any upward deviations; EHF starts the moderation early, and IRF2 applies robust brakes when necessary, ensuring a balanced and stable expression state. Together, they prevent drastic overshoots that could trigger an oscillatory cycle, with IRF2 acting as a safety valve that strongly intervenes only when CFTR levels attempt to exceed the normal range significantly.

[0399] This cooperative dynamic between EHF and IRF2 is designed to maintain CFTR transgene expression within a therapeutic window while effectively silencing harmful mutant gene activity. The next section explores broader implications and compares this strategy to alternative gene regulation methods.Plasmid DNA (pDNA) Construct Design206678-0002-00WO

[0400] Two transcription directions can be envisioned: a bidirectional transcription and a unidirectional transcription. A version of the Turbo TATA can be used as the core promoter since Turbo TATA boxes don’t possess any IRF2 and EHF binding motifs, which allows mutated IRF2 and EHF binding motifs to be strategically inserted at the desired location.

[0401] The spatial configuration of the CFTR expression cassette requires optimization. For instance, in the bidirectional configuration, two promoters are placed in opposing orientations, and repressor binding sites in between them. This setup allows for spacing adjustment among the two promoters and the repressor binding sites. By adjusting the spacing, one might be able to favor the transcription of one gene vs the other. Similarly, in a unidirectional design, the distance between the repressor binding sites and the promoter can be varied to adjust repression level.

[0402] Bidirectional transcription: In this layout, two opposing core promoters (e g., Turbo TATA boxes) are set head-to-head to initiate transcription in both directions. One promoter drives CFTR transcription, while the opposite controls a repressor transcript (e.g., IRF2 or EHF, or both). Between them, a shared regulatory region, containing operators, enhancers, and decoys, manages transcriptional activity in both directions. Adjusting the spacing among the TATA boxes, enhancers, and operator sites allows precise tuning of CFTR and repressor transcription levels respectively. By altering the operator's proximity to the CFTR promoter, transcription can be selectively repressed more heavily for the CFTR transgene compared to the repressor, or vice versa.

[0403] Unidirectional transcription: This is a more conventional layout: first operators (mutated IRF2 / EHF binding sites, etc.), regulatory region, a single core promoter, the CFTR transgene, T2A / P2A and then EHF / IRF2 transgene in a single direction. The regulatory region may include enhancers, operators and decoys. The operators may include either EHF / IRF2 binding sites or both. The CFTR and EHF / IRF2 transgene placement can be reversed with pros and cons. The operator is placed within 100 bp upstream of the core promoter to allow local repressor binding to interfere with transcription initiation. The spacing can be tuned here as well - for instance, an IRF2 operator site 20 bp upstream of promoter versus 60 bp upstream can have different repression efficacy. Empirically spacing can be tested and chosen to achieve the desired -20-50% transcriptional dampening of CFTR transgene.206678-0002-00WO

[0404] Achieving Tight Gene Regulation in Therapy: The strategy outlined for CFTR serves as a blueprint for other genetic diseases where gene dosage must be carefully controlled. Many disorders require Goldilocks expression of a therapeutic gene - not too little, not too much. This approach of engineering feedback into the DNA level is broadly applicable. For example, consider metabolic enzyme deficiencies: delivering the gene for a missing enzyme (like in urea cycle disorders or phenylketonuria) could benefit from feedback control if the enzyme’s activity needs to stay within physiological norms (to avoid metabolite imbalances). Another case is hormone or growth factor delivery - gene therapy for conditions like diabetes (insulin) or anemia (EPO) should ideally respond to the body’s needs. While this CFTR system uses native transcriptional repressors, a similar design could use tissue-specific factors or synthetic regulators for other genes. The concept of using the cell’s own regulatory network to manage a therapeutic transgene is powerful: it leverages evolutionary tuned feedback loops instead of reinventing the wheel.

[0405] For example, in X-linked diseases where females have one normal allele, gene therapy must avoid overexpression in cells where the normal allele is active. A feedback promoter could automatically dial down the transgene in those cells, while ramping up in mutant cells. In protein misfolding diseases like Alpha-1 antitrypsin (A1AT) deficiency, delivering A1AT gene with feedback control might prevent accumulation of misfolded protein (Z mutant Al AT can aggregate). A similar IRF2 / EHF-like strategy (or using the unfolded protein response sensors as regulators) could adjust expression to what the ER can handle. For muscle or CNS gene therapies, where immunogenicity is a concern, a regulated promoter might keep expression low until needed.

[0406] The CFTR design specifically addresses the unique challenge in CF: the presence of a dominant negative or misfolded mutant protein (AF508 CFTR) that one might want to repress. By uniformly knocking down all CFTR transcripts in the cell by -80% and then supplying a functional CFTR via a regulated promoter, “replace and control” is essentially done in one step - this is akin to a combined gene-silencing and gene-replacement therapy.

[0407] Advantages of DNA-Based autoregulatory system include: 1) Dynamic Adaptability: DNA-level regulation (promoter feedback) means the system can respond to intracellular conditions. For instance, if a cell for some reason starts producing too much CFTR (maybe due to multiple vector insertions or a burst of transcription), the repressors automaticallycounteract this. If production drops, repression eases. This dynamic equilibrium is harder to achieve with one-time adjustments; 2) Single Construct Simplicity: an external drug or trigger for regulation is not needed (unlike inducible promoters that need a small molecule). The control is intrinsic to the construct and host cell environment; 3) Spatial and Temporal Specificity: By using transcription factors that are naturally present only in certain tissues or conditions, where and when the gene is active is inherently targeted. In this case, EHF is mainly in airway epithelial cells, so the promoter will function as intended predominantly in the lungs (the primary target for CFTR therapy). IRF2 is ubiquitous, but its interplay with IRF1 adds a temporal component (e.g., during infection / inflammation, IRF dynamics change); and 4) Preventing Overshoot Without Constant Oversight: Unlike an external gene switch that might need monitoring, feedback promoters self-regulate. This is particularly useful in gene therapy, where once the vector is delivered, you cannot easily adjust the dose at each cell.Comparison with RNA-Based Static Systems

[0408] Traditional gene therapy and mRNA therapy approaches often utilize static expression systems. In these systems, the expression level is predetermined by the design of the construct or the administered dose and remains fixed post-delivery. Although these systems can be finely tuned during the design phase, they lack the capability to adjust dynamically in vivo. In contrast, a DNA-based tunable and differential autoregulatory system like ours introduces several key advantages.

[0409] Self-Correction: Static systems, if they overshoot the intended number of transcripts, lack internal mechanisms to correct this excess; the only recourse is the natural degradation of the RNA or the hope that it doesn't lead to adverse effects. This system, however, is designed to respond to such overshoots actively. If transcription levels exceed the desired threshold, increased repressor binding automatically initiates, bringing the expression back into the desired range. This autoregulatory feature effectively mitigates the risk of overexpression-related toxicity.

[0410] Coping with Variability: Gene therapy applications often face challenges due to uneven distribution of vector copies among target cells. Static systems typically produce a wide range of expression levels, from very low to potentially harmful highs, which might trigger immune responses or other side effects. This feedback-regulated system reduces this variability,206678-0002-00WOleading to a more consistent expression level across all treated cells. This not only enhances the overall safety by reducing outliers but also improves therapeutic efficacy by ensuring that each cell achieves at least the minimum required expression level.

[0411] Precision versus One-Size-Fits- All: Static systems, whether RNA-based or DNA-based, often require dosing that aims to achieve an average therapeutic expression level across a population of cells. This approach can result in underdosing some cells while overdosing others. This autoregulatory design, however, enables each cell to fine-tune its expression around a precise setpoint. This method offers a more tailored control at the cellular level, significantly improving the precision of treatment. For conditions like cystic fibrosis, where achieving a specific fraction of functional CFTR channels per cell is crucial, enabling each cell to autonomously reach and maintain this critical threshold is highly beneficial.

[0412] Differential Regulation: The system not only maintains expression levels within a therapeutic window but also differentially regulates genetic elements. It can repress mutant gene expression while simultaneously modulating the output of a therapeutic transgene. This dual capability is particularly advantageous for conditions where suppressing a harmful native gene product while promoting a beneficial replacement is required.

[0413] This sophisticated dynamic control system offers a significant improvement over static gene therapies by providing adaptable, precise, differential, and safe gene expression regulation, which is particularly valuable for complex genetic disorders including but not limited to cystic fibrosis.Example 2: Turbo TATA Promoter

[0414] Turbo TATA promoter variants incorporate a specific combination of core promoter elements arranged to maximize transcription efficiency and stability. These core elements include the following:

[0415] BREu (TFIIB-Recognition Element): This 7-nt element (consensus SSRCGCC) lies just upstream of the TATA box where 'S' represents either G or C, and 'R' represents either A or G. It binds TFIIB, stabilizing TFIIB’s interaction with TBP (the TATA-binding protein) and the DNA. BREu sequence is chosen to avoid the GGGCGCC motif (thus eliminating potential Spl binding sites). Its primary function is to facilitate the binding of TFIIB, aiding the formation of the transcription pre-initiation complex.

[0416] The BREu is positioned adjacent to the TATA box and is primarily a TFIIB interaction site, not a typical enhancer or upstream activator sequence. This means that even a GC-rich BREu like GGGCGCC is not guaranteed to recruit Spl in vivo if the context is unfavorable (for example, if TFIIB and other basal factors occupy the site). However, introducing GCACGCC in place of GGGCGCC provides an extra safeguard by eliminating the Spl consensus motif altogether. This ensures that the promoter will be driven through the intended basal transcription machinery without any secondary influence from Spl or other GC-box binding factors. The result is an optimized core promoter fidelity.

[0417] TATA Box: The TATA box element conforms to the consensus TATAWAAR (where W = A / T and R = A / G). The TATA box helps position TFIID precisely at the promoter, directing RNA polymerase II to the correct start site. A correctly positioned TATA box is critical for strong transcription: mutating or removing TATA can decrease promoter activity by an order of magnitude.

[0418] Initiator (Inr): The initiator sequence is designed to conform to the human consensus YYA+1NWYY at the transcription start site, which maximizes recruitment of TFIID and RNA Polymerase II. A core sequence analogous to “TCAGTT” is preferred for the Inr element to optimize transcription initiation and expression levels. Incorporating an optimized initiator enhances transcriptional precision and efficiency.

[0419] MTE (Motif Ten Element): An optional MTE sequence can be included in the promoter design that confers strong functional activity. The MTE is positioned downstream of the initiator and contributes to elevated basal transcription.

[0420] DPE (Downstream Promoter Element): The DPE is a core motif located about +28 to +33 nucleotides downstream of the TSS. The consensus DPE sequence is 5'-RGWYVT-3', i.e. (A / G)G(A / T)(C / T)(A / C / G)T, as a broader consensus that applies across species including humans. Human CKS2 GGACTGG Identified as a high-activity natural human DPE. Optimized Human AGTCGC High-scoring motif in human-specific SVR models. Human IRF-1AG AC GT G Natural DPE from the human Interferon Regulatory Factor-1.

[0421] The chosen DPE sequence is positioned such that it spans nucleotides +28 to +33 relative to the transcription start site. Importantly, this sequence includes a cytosine at the fourth position (+31), a feature shown to enhance promoter activity.206678-0002-00WOElement Positioning and Overlap

[0422] The MTE and DPE elements are arranged in the promoter so that they overlap at positions +28 and +29 (sharing the dinucleotide “GG”). This overlap satisfies the consensus requirements of both the MTE and DPE simultaneously, allowing both elements to function in tandem without increasing the overall promoter length.Spacer Sequences and Arrangement

[0423] Proper spacing between these core motifs is critical for function, so the spacer sequences (the DNA between the core elements) were designed to achieve ideal distances and avoid any extraneous signals:

[0424] BREu-TATA Spacer (position -31): A single nucleotide adenine (A) is inserted between the BREu element and the TATA box. This one-base spacer prevents extension of the BREu sequence into the TATA region and avoids formation of a secondary TATA-like motif.

[0425] TATA-Inr Spacer (positions -24 / -23 to -3): A linker of 21-22 nucleotides that separates the TATA box from the initiator. This spacer sequence ensures the correct distance (~30 bp) between the TATA box and the Inr, preserving the proper helical spacing required for efficient transcription initiation. The spacer has a balanced, non-repetitive nucleotide composition to prevent the formation of unintended binding sites. Sequences like “CCAAT” (NF-Y / CAAT-box), GC-rich stretches (Spl sites), or T-rich TATA-like motifs were specifically avoided in this region. The spacer is AT / GC-mixed and lacks known core promoter motifs, ensuring that TFIID focuses on the intended TATA and Inr sites. By maintaining the correct distance and neutral composition, this spacer allows simultaneous recognition of TATA and Inr by TFIID with no interference.

[0426] Inr-MTE Spacer (positions +5 / +6 to +17): Twelve / Thirteen nucleotides are inserted between the end of the initiator and the start of the MTE. This spacer provides the necessary gap so that the MTE begins at the optimal downstream position, without introducing any extraneous regulatory signals in between.

[0427] Inr-DPE Spacer (positions +5 / +6 to +27): Twenty-two / Twenty-three carefully selected nucleotides lie between the initiator and the DPE. This fixes the DPE at +28-+33, exactly aligning with the required spacing from the +1 start site. The sequence was optimized to avoid “GT” or “AG” dinucleotides that could act as cryptic splice sites if transcribed. It was206678-0002-00WOconfirmed that nowhere in +6 to +27 does the sequence form a “GT... AG” pattern or a canonical “GU” donor site in the mRNA. For example, the sequence contains no GT at all in the sense strand, thereby eliminating any 5 '-splice donor consensus (GU) in the nascent RNA. It also avoids creating an inadvertent start codon (ATG) that could initiate a uORF - every occurrence of “ATG” was disrupted (e.g. the segment... AATCGA... contains no contiguous ATG).Additionally, no known transcription factor binding sites (AP-1, CRE, ETS, etc.) are present -the spacer primarily consists of unique 2-3 bp sub-sequences that do not match consensus binding motifs in databases. This neutrality ensures that the core promoter’s performance is consistent in different cellular environments, without cell-type-specific factors binding to the spacer. Crucially, the 22 bp length preserves the Inr-DPE phasing that TFIID requires for cooperative binding.

[0428] This spacing positions the DPE exactly at +28 to +33, aligning it with the consensus location for downstream promoter elements and ensuring it overlaps appropriately with the MTE if one is included.

[0429] Overall, the promoter sequence design maintains precise spacing between all core elements to optimize transcription efficiency while avoiding the creation of unintended binding sites or other extraneous regulatory motifs. Each spacer is engineered to be neutral (having no negative effect on promoter function) so that it simply maintains distance without interfering with core element activity.Table 7: Example Sequence #1Component Position (rel. TSS) SequenceTATA Box -30 to -25 TATAAAUpstream Spacer -24 to -3 cagacgtcgcatcgatctaca (SEQ ID NO: 136)Initiator (Inr) -2 to +5 TCAGTTDownstream Spacer +6 to +27 atcacacgacatc (SEQ ID NO:143)DPE Motif +28 to +33 AGACGTTable 8: Example Sequence #2BREu–TATA–Inr–DPE composite (5′→3′) with each element labeled and the +1 TSS in bold.Element Position (rel. TSS) SequenceGGGCGCC BREu -38 to -32 CCGCGCCGCACGCCSpacer (BREu-TATA) -31 ATATA box -31 to -24 TATATAASpacer (TATA-Inr) -23 to -3 cagacgtcgcatcgatctaca (SEQ ID NO: 136)Inr -2 to +5 (A—l) TCAGTCTSpacer (Inr-DPE) +6 to +27 gacttgcgtcgtacggttacgc (SEQ ID NO: 155)+28 to +33 AGACGTDPE+28 to +34 AGTCGTG

[0430] BRE-TATA Spacing accurately initiates at +1 via the Inr and enhances Pol II recruitment / stability via the DPE (especially in conjunction with TFIID).

[0431] TATA-Inr spacing is at the proper distance from the TATA box to the Initiator is critical. Classic studies found the TATA box (consensus around TATAA / TW; e.g. TATAAA) works best -25-30 bp upstream of the transcription start site (TSS). More recent analyses show 30-31 bp from the TATA “T” to the +1 base is optimal. Shorter or longer spacings can reduce efficiency, as they misalign TBP on DNA relative to the start site. In practice, placing the TATA box such that its first T is -30 bp from the +1 position yields highest activity. For instance, a TATA at -30 to -25 (if 6 nt long) or -30 to -24 (7 nt) with +1 at the Inr will provide -30 bp spacing. This positioning allows the DNA to bend appropriately around TBP and presents the Inr in the correct location to TFIID.

[0432] Inr-DPE spacing is a strict gap of -28-32 bp between the Inr and DPE is required for DPE function. The downstream promoter element (DPE) is located precisely +28 to +32 (up to +33) relative to the A+l of the Inr in Drosophila and human focused promoters. This spacing is “hard-wired”: all known DPE-dependent promoters maintain the same Inr-to-DPE distance. In fact, inserting or deleting even a single nucleotide between the Inr and DPE consensus motifs markedly reduces transcription and TFIID binding. Thus, in designing a core promoter that includes a DPE, one must ensure the DPE motif starts -28 nucleotides downstream of the +1 base. Any deviation can disrupt the cooperative recognition of Inr and DPE by TAF subunits of TFIID. In an optimized layout, if +1 is the A of the Inr, the DPE should span roughly +28 to +33.206678-0002-00WOExperiments confirm that altering this distance by even 2 bp drops promoter activity significantly.Avoiding Unintended TF Binding in Core Promoter Sequences

[0433] To ensure the designed promoter (with BREu TATA, Inr, DPE, and spacer DNA) has no unwanted transcription factor binding sites or cryptic enhancers, the following guidelines should be taken into consideration:

[0434] Use only known core promoter motifs: The BREu TATA box, Initiator (Inr), Motif Ten Element (MTE), and Downstream Promoter Element (DPE) are canonical core promoter elements that recruit general transcription factors (GTFs) rather than sequence-specific regulatory TFs. These short motifs are recognized by components of the basal transcription machinery (e g. TFIIB binds the BRE, TBP / TFIID binds TATA and Inr, TFIID subunits recognize MTE / DPE). Because they are designed to interact with the general Pol II initiation complex, they are less likely to coincidentally serve as binding sites for other TFs. This minimizes unintended regulatory interactions.

[0435] Scan for known TF motifs: It’s good practice to screen the composite sequence (including spacer regions) against transcription factor motif databases (e.g. JASPAR or TRANSFAC) to verify it doesn’t contain high-affinity sites for common transcriptional activators or repressors. Ensure that none of the spacers or overlaps between core elements form sequences like AP-1 (TGASTCA), Spl (GGGCGG), NF-KB (GGGRNNYYCC) (SEQ ID NO:285), etc. In a well-designed core promoter, spacer sequences are often chosen to be innocuous (lacking repeats or CpG islands) and just serve to maintain proper spacing between elements. For example, the spacing between Inr and DPE (∼+28 to +32) or Inr and MTE (+18 to +27) is critical, so the intervening bases should be only those needed for spacing and not encode any known enhancer element.

[0436] Avoid multiple initiator-like motifs: To prevent cryptic promoters, make sure there is only one strong Inr element (at the intended +1 TSS) and no similar YYANWYY sequence nearby that could act as an alternative start site. Likewise, use only one TATA motif in the expected position (around ∼−30) so that TFIID is focused there and does not initiate at a secondary location. A single, strong core promoter focus (as in a “focused” promoter) helps ensure transcription initiates only at the designed site, rather than dispersed initiation.206678-0002-00WO

[0437] Optimize for basal strength, not regulatory input: The chosen combination (BREu-TATA-Inr-MTE-DPE) is known to produce a high level of basal transcriptional activity without additional upstream activators. This super core promoter architecture (see SCP1 below) was rationally designed for maximal Pol II recruitment and productive transcription, outperforming even strong viral core promoters. By confining the sequence to core motifs, you avoid adding any enhancer modules. In other words, the promoter should drive efficient transcription initiation while remaining largely unresponsive to other regulatory proteins (aside from the basal GTFs). Ensuring the sequence contains only these core elements and neutral spacers means it won’t inadvertently behave as an enhancer or repressor element.

[0438] By following these guidelines, the promoter sequence will be optimized for efficient transcription initiation and have a low probability of recruiting unintended transcription factors or regulatory signals. The result is a strong, focused core promoter that reliably initiates at +1 and minimizes spurious interactions.Table 9: Example Sequence #3:BREu–TATA–Inr–MTE–DPE composite (5′→3′) with each element labeled and the +1 TSS in boldElement SequenceBREu GCACGCCSpacer (BREu-TATA) A (optional)TATAAAAG TATA boxTATATAAGcagacgtcgcatcgatctaca (SEQ ID NO: 136)Spacer (TATA-Inr)agtcgacgtcgtagtcagcta (SEQ ID NO: 139)TCAGTTInrTCATTCagtcagtcagtca (SEQ ID NO: 142)Spacer (Inr-MTE)atcacacgacatc (SEQ ID NO:143)CGAACGGAAC (SEQ ID NO: 147)MTE CCAGCCGAAC (SEQ ID NO: 148)AGTCGC DPE GGACTGGAGACGTG

[0439] Based on the above structure, following promoter sequences were created:SEQ ID NO: 207:GCACGCCTATAAAAGcagacctcgcatcgatctacaTCAGTTatcacacgacatcCGAACGGAACAGTC GC SEQ ID NO:208 (Spacer (BREu-TATA) a is added):GCACGCCaTATAAAAGcagacctcgcatcgatctacaTCAGTTatcacacgacatcCGAACGGAACAGTC GC SEQ ID NO: 209:GCACGCCTATATAAGagtcgacgtcgtagtcagctaTCATTCagtcagtcagtcaCCAGCCGAACGGACT GG SEQ ID NO:210 (Spacer (BREu-TATA) a is added):GCACGCCaTATATAAGagtcgacgtcgtagtcagctaTCATTCagtcagtcagtcaCCAGCCGAACGGAC TGG.Validated Super Core Promoter (SCP) Sequences

[0440] Additionally, the following sequences represent optimized sequences focused on basal initiation. These may be used in conjunction with other elements.SCP-01 (SEQ ID NO:211):AGGTCTATATAAGCAGAGCTCGTTTAGTGAACCGTCAGTCCGCCTGGAGACCTCGAG CCGAGTGGTCGTGCCTCCATAGAA SCP-1' (BREu-Enhanced) (SEQ ID NO:212):GCACGCCGTACTTATATAAGGGGGTGGGGGCGCGTTCGTCCTCAGTCGCGATCGAAC ACTCGAGCCGAGCAGACGTGCCTACGGACCG SCP-02 (SEQ ID NO:213) GTACTTATATAAGGGGGTGGGGGCGCGTTCGTCCTCAGTCGCGATCGAACACTCGAG CCGAGCAGACGTGCCTACGGACCG SCP-03 (SEQ ID NO:214):GTACTTATATAAGGGGGTGGGGGCGCGTTCGTCTTCAGTTTTTTTTCAACACTCGAG CCGAGCAGACGTGCCTACGGACCG.Example 3; Muscle Specific Construct

[0441] Design Rationale: Skeletal muscle transcriptional control requires coordinated activation of myogenic regulatory factors (MRFs) and MEF2 family members. The design incorporates both differentiated myofiber regulators (MYOG, MRF4) and satellite cell factors (PAX7) to ensure broad muscle targeting including regenerative capacity.

[0442] Key design features include: 1) 42bp x M spacing between motif centers (4 helical turns per interval), M is an integer. Multiple motifs can be on the same helix phase. 2) Two DNA Targeting Sequences (DTS): MEF2A (distal) and SIX1 (proximal). 3) Four enhancer elements targeting distinct muscle cell populations. 4) Satellite cell targeting (PAX7 + SIX1) for regenerative capacity. 5) Differentiation-stage coverage from myocytes to mature myofibers. 6) A promoter (e.g., Super Core Promoter (SCP)) at TSS for basal transcription.Table 10: TFBS Architecture (SCP-Compatible)Role Sequence TF Cell Target DTS TTCTAAAAATAGAAA (SEQ ID MEF2A Skeletal Muscle NO:62)ENHANCER CAGCACCTGTCCC (SEQ ID NO:73) MYOG Diff. Myocytes ENHANCER AACAGCTGTT (SEQ ID NO:74) MRF4 Mature Myofibers ENHANCER CCACATTCCAGGC (SEQ ID NO: 84) TEAD1 Skeletal Muscle ENHANCER TTTGCACACGGCAC (SEQ ID NO: 94) PAX7 Satellite Cells DTS CTAATTA SIX1 Satellite CellsRole: DTS = DNA Targeting Sequence (nuclear import); TF = Transcription Factor binding siteAssembly of Motifs and SCP, and Position Rationale

[0443] The 42-bp repeat (exactly 4.0 helical turns) ensures that the centers of all upstream sites stay in phase and don’t have steric hindrance. This eliminates 'phasing drift,' preventing the activators from rotating away from the TFIID recruitment face.

[0444] FOXA2-DTS (n=3.0): Precisely 3 turns from TATA. This is the optimal 'anchor' distance for bridging the enhancer cluster to the basal machinery.ccaaaggtttccaaaggtttc [spacer] (SEQ ID NO:117) aaaggtttccaaaggtttccaaaggttt [spacer] (SEQ ID NO:118) CAGCACCTGTCCC (MYOG) (SEQ ID NO:73) aaaggtttccaaaggtttccaaaggtttc [spacer] (SEQ ID NO: 112)206678-0002-00WOAACAGCTGTT (MRF4) (SEQ ID NO:74)aaaggtttccaaaggtttccaaaggtttc [spacer] (SEQ ID NO: 112) CCACATTCCAGGC (TEAD1) (SEQ ID NO: 84) aaaggtttccaaaggtttccaaagggtt [spacer] (SEQ ID NO: 119) TTTGCACACGGCAC (PAX7) (SEQ IDNO:94) aaaggtttccaaaggtttccaaaggtttccaa [spacer] (SEQ ID NO: 120)CTAATTA (SIX1-DTS)

[0445] Combined sequences from above:5 ’ ccaaaggtttccaaaggtttcTT C T A A A A AT AGA AAaaaggtttccaaaggtttccaaaggtttC AGC AC C T GT C CCaaaggtttccaaaggtttccaaaggtttcAACAGCTGTTggaggtttccaaaggtttccaaaggtttcCCACATTCCA GGCaaaggtttccaaaggtttccaaagggttTTTGCACACGGCACaaaggtttccaaaggtttccaaaggtttccaaCTAA TTA-3’ (SEQ ID NO: 128)

[0446] Combined sequences with short spacer between DTS cluster and promoter (underlined SCP):5’ccaaaggtttccaaaggtttcTTCTAAAAATAGAAAaaaggtttccaaaggtttccaaaggtttCAGCACCTGTC CCaaaggtttccaaaggtttccaaaggtttcAACAGCTGTTggaggtttccaaaggtttccaaaggtttcCCACATTCCA GGCaaaggtttccaaaggtttccaaagggttTTTGCACACGGCACaaaggtttccaaaggtttccaaaggtttccaaCTAA TTAaaaggtttccaaaggttccaGCACGCCTATAAAAGcagacgtcgcatcgatctacaTCAGTTatcacacgacat aCGAACGGAACAGACGT 3’ (SEQ ID NO: 129)

[0447] Combined sequences with longer spacer between DTS cluster and promoter (underlined SCP):5’ccaaaggtttccaaaggtttcTTCTAAAAATAGAAAaaaggtttccaaaggtttccaaaggtttCAGCACCTGTC CCaaaggtttccaaaggtttccaaaggtttcAACAGCTGTTggaggtttccaaaggtttccaaaggtttcCCACATTCCA GGCaaaggtttccaaaggtttccaaagggttTTTGCACACGGCACaaaggtttccaaaggtttccaaaggtttccaaCTAA TTAaaaggtttccaaaggtttccaaaaggtttccaaaggtttccaGCACGCCTATAAAAGcagacgtcgcatcgatctacaT C4GTTatcacacgacataCGAACGGAACAGACGT 3’ (SEQ ID NO: 130)206678-0002-00WO

[0448] Helical phasing is calculated as the distance (D) from the TATA start (-31) to the geometric center of each motif, divided by the B-DNA pitch of 10.5 bp / turn. An integer value (n) confirms the motif is aligned with the core recruitment face.Table 11TF Identity Center (Rel. Di st. from - Turns (n)TSS) 31TATA Box -31.0 0.0 bp 0.0SIX1-DTS -62.5 31.5 bp 3.0PAX7 -104.5 73.5 bp 7.0TEAD1 -146.5 115.5 bp 11.0MRF4 -188.5 157.5 bp 15.0MYOG -230.5 199.5 bp 19.0MEF2A- -272.5 241.5 bp 23.0DTSTable 12: Muscle-Specific Alternative MotifsMEF2A MYOG MRF4 TEAD1 SIX1 TGCTAAAAATAG CGGCACCTGTCCC AACAGCTGTT GCACATTCCAGGC CTAATTA AAC (SEQ ID NO:63) (SEQ ID NO:74) (SEQ ID NO:84)(SEQ ID NO:52)GGCTAAAAATAG CCGCACCTGTCCC AACAGCTGTC CCACATTCCAGGG CTCATTA AAC (SEQ ID (SEQ ID NO:64) (SEQ ID NO:75) (SEQ IDNO:85)NO:53)TTCTAAAAATAG GGGCACCTGTCCC AACAACTGTT GCACATTCCAGGC TTAATTA AAC (SEQ ID (SEQ ID NO:65) (SEQ ID NO:76) (SEQ ID NO:86)NO:54)GTCTAAAAATAG CGGCACCTGTCCG GACAGCTGTT GCACATTCCAGGC GAAATTA AAC (SEQ ID (SEQ ID NO:66) (SEQ ID NO:77) (SEQ ID NO:87)NO:55)TGCTAAAAATAG CGGCACCTGTCAC AACAACTGTC CTACATTCCAGGC CCAATTA ACC (SEQ ID (SEQ ID NO:67) (SEQ ID NO:78) (SEQ ID NO:88)NO:56)GGCTAAAAATAG GCGCACCTGTCCC AACACCTGTT CCACATTCCAGCG: TTCATTAACC (SEQ ID (SEQ ID NO:68) (SEQ ID NO:79) (SEQ ID NO:89)NO:57)TGCTAAAAATAG CCGCACCTGTCCG GACAGCTGTC GCACATTCCAGGC CACATTA CAC (SEQ ID (SEQ ID NO:69) (SEQ ID NO:80) (SEQ ID NO:90)NO:58)GGCTAAAAATAG CCGCACCTGTCAC AACAGTTGTT GCACATTCCAGGC CCCATTA CAC (SEQ ID (SEQID NQ:70) (SEQ ID NO:81) (SEQ ID NO:91)NO:59)206678-0002-00WOTTCTAAAAATAG GGGCACCTGTCCG AACAGGTGTT CTACATTCCAGGG CTAATTG ACC (SEQ ID (SEQ ID NO:71) (SEQ ID NO:82) (SEQ ID NO:92) NO:60)GTCTAAAAATAG GGGCACCTGTCAC GACAACTGTT CTACATTCCAGCC TAAATTA ACC (SEQ ID (SEQ ID NO:72) (SEQ ID NO:83) (SEQ ID NO:93)NO:61)Example 4: Intron and Telomeric Motif Design in DNA Gene Therapy

[0449] In gene therapy, a significant challenge is the host’s innate immune system recognizing the delivered DNA as a foreign threat, regardless of the gene delivery methods. Unmethylated CpG dinucleotides and other DNA features can activate receptors like Toll-like receptor 9 (TLR9) in endosomes and AIM2 or cGAS in the cytosol, leading to inflammation. These DNA-sensing pathways trigger production of interferons, cytokines, and activation of immune cells that can reduce the efficacy of the therapy (by silencing transgene expression or killing transfected cells) and cause side effects.

[0450] Traditionally Lipid NanoParticles (LNPs) deliver nucleic acid payload by entering cells primarily through endocytosis and become sequestered in endosomal vesicles. Within these compartments, host endosomal Toll-like receptor 9 (TLR9) can recognize foreign DNA (especially unmethylated CpG motifs) and initiate innate immune signaling. Efficient gene delivery requires the DNA to escape from endosomes into the cytosol so it can reach the nucleus; however, the process of endosomal escape can itself provoke additional innate immune activation. When endosomal membranes are disrupted and DNA leaks into the cytosol, the “naked” double-stranded DNA is exposed to cytosolic DNA sensors such as cyclic GMP-AMP synthase (cGAS) and AIM2. Cytosolic dsDNA binding to cGAS triggers the cGAS-STING pathway, leading to downstream induction of type I interferons and other proinfl ammatory cytokines. Similarly, cytosolic dsDNA can directly engage AIM2, causing assembly of the AIM2 inflammasome and activation of caspase- 1, which generates IL-1 and can induce pyroptotic cell death. Activation of these innate pathways by LNP-delivered DNA often results in robust inflammatory responses and cellular toxicity, posing a challenge for safe DNA therapy.

[0451] Described herein is a DNA gene therapy construct that includes two key features to improve therapeutic outcomes:

[0452] (1) Telomeric Motifs: at least one block of telomeric repeat DNA (one or more TTAGGG repeats, optionally linked by spacers) strategically in the cassette sequence to mitigate206678-0002-00WOinnate immune sensing. Mammalian telomeres consist of repeated TTAGGG motifs, and remarkably, DNA containing TTAGGG repeats can broadly suppress immune activation. By embedding these telomeric motifs within non-coding regions of a gene therapy cassette, the DNA can “cloak” itself from DNA-sensing immune receptors without interfering with the expression of the therapeutic gene.

[0453] (2) Intron-Mediated Enhancement: an intron sequence to enhance transgene expression. The key innovation is the use of multiple TTAGGG repeats in tandem as a built-in immune suppressor element within the cassette, combined with the inclusion of an intron that promotes higher transgene expression via intron-mediated enhancement. In parallel, the presence of a short intron in the cassette (for instance, immediately the 5’ UTR) leverages intron-mediated enhancement to boost the production of the therapeutic protein.

[0454] Overall, a gene therapy cassette engineered with an intron and telomeric motifs maintains high transgene expression in the target cells - both because the intron intrinsically boosts expression and because the cassette is less attacked or silenced by the immune system. This effect is especially beneficial during critical delivery stages (such as endosomal escape when the DNA cassette is delivered via lipid nanoparticles), and it improves the safety profile of the gene therapy by reducing acute inflammatory responses.Intron-Mediated Enhancement (IME):

[0455] The inclusion of a short intron immediately after the 5'UTR can significantly increase gene expression levels. The presence of an intron in the transcript promotes more efficient mRNA maturation - facilitating proper capping, splicing, nuclear export, and mRNA stability. Splicing of the intron recruits exon junction complexes and other factors that enhance nuclear export and translation of the mRNA. As a result, the intron-containing cassette yields higher steady-state mRNA and protein levels compared to an identical intron-less cassette.Modified Human β-Globin Intron I (130 bp, truncated)

[0456] This sequence shares <90% identity with the wild-type intron I:GTAAGAGAATGACAGTAGTATCTCCAAAGCATATTGAGATTCCTCTGTACATACCCC CTCCTTCCCAGTGTTTGTGGGCTGGTCCCCCAATCAAGGCTTTACTAACCCTCTTTTT GCAACAGCCAAACAG (SEQ ID NO:229)Modified Human β-Globin Intron II (105 bp, truncated)

[0457] Like intron I, this variant stays below 90% homology to the wild-type β-globin intron II sequence:GTATCATGTGTGGATCTCACAAAAAACCACCTTTCATCTAAAACCATGGTAGTGCAG GCTAGAGTCTACACCCCATACTAACTATTCTTGGGATCCCATTTAAAG (SEQ ID NO:230)Truncated Human Factor IX Intron I (90 bp)

[0458] This intron already omits non-critical midsections and retains the necessary 5' GT... AG splice signals and a branch point motif for proper splicing:GTGATTCCTCTAGCTGGTGCTCTAGTCAGTTGGCTGGAAGCAATCTTGGACCCCCCTT TCTGATATACTAACCCACCTCTATAGGGCAAG (SEQ ID NO:231)Telomeric Motif (TM)

[0459] Telomeric repeats, with the sequence 5'-TTAGGG-3' (and reverse complementary 5'-CCCTAA-3’), are a natural component of chromosome ends in mammals. These repeats have an intrinsic ability to suppress immune activation when present in DNA fragments. The mechanism, as understood in the field, is that certain DNA sensors require specific DNA patterns (such as CpG motifs or long dsDNA stretches) to trigger an immune response. Telomeric DNA sequence, being repetitive and GC-rich in a specific pattern, can bind to these receptors in a nonstimulatory way:

[0460] TLR9 Suppression: TLR9 normally detects unmethylated CpG motifs (e.g. 5'-TCGTT... sequences) in endosomal DNA and dimerizes to initiate signaling. Telomeric sequences like TTAGGG act as antagonists for TLR9. They can occupy the DNA binding site of TLR9 without activating it, preventing TLR9 from binding to true immunostimulatory DNA. In effect, telomeric repeats block TLR9 activation even if stimulatory DNA is present. This property has been demonstrated with synthetic oligodeoxynucleotides containing TTAGGG repeats, which inhibit TLR9-mediated responses.

[0461] Cytosolic DNA Sensor Inhibition: DNA in the cytoplasm is monitored by sensors such as AIM2 (which forms an inflammasome upon binding DNA) and cGAS (which produces a second messenger to trigger interferon production when it binds DNA). Telomeric repeat DNA has been found to bind these sensors without triggering them, thereby acting as a decoy. For example, a DNA sequence with multiple TTAGGG motifs can attach to AIM2 and prevent itfrom assembling the inflammasome, and likewise can occupy cGAS’s DNA-binding interface, reducing the enzyme’s activation. This means the presence of telomeric motifs in a plasmid can dampen the AIM2 and cGAS-STING pathways that would normally respond to foreign DNA in the cytosol.

[0462] Importantly, a single TTAGGG motif alone is not sufficient to cause a strong immunosuppressive effect. The inhibitory action requires multiple repeats in proximity. Research and experimental evidence indicate that three or more TTAGGG repeats in tandem are needed for potent immune pathway suppression. This is presumably because the receptors have multiple binding sites or require a certain DNA length / configuration to engage in an inhibitory conformation. Natural coding sequences occasionally contain TTAGGG segments; interestingly, those instances have been correlated with lowered innate immune recognition of those genes. For instance, if a gene’s coding sequence inherently has several TTAGGG segments, it tends to provoke less TLR9 activation. This underlies the strategy of intentionally adding telomeric segments to a therapeutic DNA.Design of the Telomeric Motif

[0463] A critical aspect of this invention is the design of the telomeric repeat cassette that is inserted into the DNA cassette. The number of repeats, their arrangement (tandem vs. with spacers), and sequence composition are chosen to maximize immune suppression while maintaining genomic stability and ease of manufacturing.

[0464] Number of Repeats: In various embodiments, the insert contains roughly 3 to 6 copies of the TTAGGG motif in a row. Using four repeats (which yields a 24 base pair sequence: TTAGGGTTAGGGTTAGGGTTAGGG (SEQ ID NO:216)) has proven to be a convenient and effective design, as it mirrors known immunosuppressive oligodeoxynucleotides used experimentally. Fewer than three repeats (e.g. a single TTAGGG or two repeats) has little effect on immune sensors, so at least three are preferred. Adding more than six repeats (creating sequences longer than ~36 bp of pure telomeric DNA) is generally not necessary for immune inhibition and could potentially introduce unwanted secondary structures or make the plasmid slightly larger than needed. However, the invention is not strictly limited to six repeats — longer tracts (even 8, 10, or more repeats) could be used if needed, but with diminishing returns.Therefore, an ideal range is 3-6 telomeric repeats.

[0465] Tandem vs. Spacer-Separated Repeats: Telomeric repeats can be inserted contiguously or with short spacer sequences between each repeat. A tandem contiguous sequence (e g. TTAGGGTTAGGGTTAGGG (SEQ ID NO:217)) closely mimics natural telomeres and ensures a high local density of the motif. This contiguity maximizes the chance that a DNA sensor protein binding along the DNA will encounter multiple TTAGGG motifs in one area, leading to strong inhibition. Alternatively, one can include short spacer sequences (for example, 2-5 nucleotides of adenine or thymine) between the TTAGGG units. These spacers can disrupt any potential formation of certain secondary structures (like G-quadruplexes, which are four-stranded DNA formations that G-rich sequences can form under some conditions). Spacers also avoid creating any unintended binding sites or palindromic sequences. For instance, an insert sequence could be designed as:i) Contiguous example:TTAGGGTTAGGGTTAGGGTTAGGGTTAGGG (SEQ IDNO:218) (5 repeats back-to-back, 30 bp total).ii) Spacer-separated example:TTAGGGAAAATTAGGGAAAATTAGGG (SEQ ID NO:219)(where AAAA represents a spacer of four adenines between each TTAGGG repeat; this example shows 3 repeats with spacers,totaling 26 bp).

[0466] The contiguous version offers simplicity and maximal motif density, whereas the spacer-separated version offers structural stability (the poly-A spacers are very unlikely to form any structure or interact with proteins). Poly(dA) or poly(dT) spacers are preferred if spacers are used, because homopolymeric A / T stretches are neutral and do not trigger immune responses (and they carry no CpG sequences). Spacers of other compositions (e.g. AT or TA dinucleotide repeats) can also be used as long as they do not introduce motifs that negate the immunosuppressive effect. The length of spacer can vary (2 to 6 bases is a practical range; the example above uses 5 or 6 adenines as a spacer).

[0467] Sequence Integrity: The telomeric insert is designed to lack any immunostimulatory motifs itself. Notably, the sequence TTAGGG contains no CpG dinucleotide, so it does not activate TLR9 (which specifically looks for CpG). By using only T, A, and G in that pattern (and possibly A / T in spacers), it is ensured the insert does notaccidentally add the very triggers that are being suppressed. Sequences like TTTT (long poly-T) that could potentially affect transcription termination are also avoided, unless they are very short (a 4-6 T stretch as a spacer is fine and commonly used in plasmid designs without issue). In summary, the insert is a compact, self-contained DNA element optimized to engage immune sensors in a suppressive manner and free of any pro-inflammatory signals.Optimal TM Insertion Sites in the Cassette

[0468] Selecting where to place the telomeric repeat cassette in the cassette is crucial. The insertion site should be “immune-neutral” (meaning the sequence will be present in the DNA delivered to the cell and sensed by the immune system) but biologically silent (meaning it does not disrupt the cassette’s primary function of gene delivery and expression). There are several optimal locations for embedding the TTAGGG repeat cluster in a typical gene therapy plasmid or DNA construct:

[0469] Immediately after 5’ UTR: Placement of the telomeric motif insert is preferably upstream of the intron within the 5 'UTR, so that these repeats are present in the plasmid and the initial transcript without disrupting any splice junction. For example, a series of TTAGGG repeats can be placed immediately after the promoter (as part of the first exon, just before the intron’s 5' splice site). In this configuration, the telomeric sequence will be transcribed as part of the 5 'UTR exon. It may remain in the final mRNA (if it is entirely within an exon region outside the intron) or be partially removed if included at an intron boundary, but in either case it serves its purpose of being present in the plasmid DNA to dampen immune recognition. By keeping the telomeric motifs outside the intron itself (i.e., not inside the intron sequence), it was ensured they do not interfere with splicing signals. However, in some embodiments, telomeric repeats could be placed within the intron as well, provided that the intron’s critical splice signal integrity is maintained and the overall intron length remains in the desired range. Whether in an exon or within the intron, these telomeric inserts do not encode protein and do not alter the coding sequence; their function is to act as a benign spacer that helps the plasmid mimic genomic DNA segments known to be non-inflammatory. The number of telomeric repeat units can be adjusted (for example, 4-10 repeats of TTAGGG in tandem) to achieve the desired effect without adding excessive length.206678-0002-00WO

[0470] Within an Intron: If the expression cassette includes a synthetic intron (as is often done to enhance gene expression, especially in eukaryotic expression cassettes), this intron provides an excellent hiding spot for the telomeric repeats. The cluster can be inserted into the intronic sequence, away from critical splicing signals. Introns have defined elements: a 5' splice donor site (typically starting with GT), a branch point sequence (somewhere in the middle), and a 3' splice acceptor site (ending with AG). The telomeric insert should be placed midway or in a region of the intron that does not disrupt these elements. For example, one can introduce the TTAGGG repeat cassette roughly in the middle of an intron, ensuring that the splice donor and acceptor sites remain untouched. By doing so, during mRNA processing the intron (with the embedded telomeric repeats) will be completely spliced out. Thus, the telomeric sequence will not appear in the mature mRNA at all, and the coding sequence of the gene remains exactly as intended. The presence of the insert in the intron only affects the plasmid DNA (and pre-mRNA temporarily), which is sufficient for it to exert the immune evasion effect.

[0471] Downstream of Poly(A) Signal: Another ideal location is immediately after the 3 ' untranslated region (UTR) of the gene, following the polyadenylation [poly(A)] signal sequence. A typical eukaryotic expression cassette ends with a poly(A) signal, which directs the cellular machinery to cleave and add a poly(A) tail to the mRNA. By inserting the TTAGGG repeat cluster right after this poly(A) sequence, one ensures that the insert lies outside the actual transcript (transcription will terminate at the poly(A) site, so the telomeric sequence will not be transcribed into mRNA). This region is part of the plasmid’s non-coding DNA - often there is a short spacer or multiple cloning site region after the poly(A) for flexibility. It is a “safe harbor” to add extra sequences. The telomeric insert here will be present in the plasmid DNA delivered to cells (so immune sensors can detect it), but it will have zero impact on the gene’s expression, since it’s downstream of the gene’s termination signal.

[0472] In the Plasmid Backbone (Non-Expression Region): The plasmid backbone includes elements needed for plasmid propagation in bacteria (like the origin of replication, selection marker, etc.), which are not expressed in the patient’s cells (assuming a non-viral plasmid delivery). Any neutral segment of the backbone can accommodate the telomeric repeat insert. For instance, if there is a stretch of DNA in the backbone that does not encode a protein and is not part of an essential regulatory sequence, the insert can be placed there. Some advanced gene therapy plasmids use a minimal backbone (for example, nanoplasmids or minicircles thatremove antibiotic resistance genes and minimize CpG content). In such cases, adding a telomeric sequence to the remaining backbone can counteract even the low residual immunogenicity. The backbone insertion is especially useful because it applies to virtually any cassette (even viral cassettes have “stuffer” or intergenic regions that could be analogous to a plasmid backbone segment where an insert won’t affect the cassette’s life cycle).

[0473] Multiple Insertion Sites: The above options are not mutually exclusive. In some embodiments of this invention, multiple telomeric repeat cassettes are inserted at different locations in the same cassette for an additive effect. For instance, one could insert a TM in an intron and another in the backbone, or one in the intron and one after the poly(A) signal, etc. Using more than one such insert can further ensure that wherever the DNA is encountered by sensors (nucleus, cytosol, endosome), there is a suppressive sequence nearby to engage the receptors. Combining them can produce a cassette with a very robust immune-evasive profile.Orientation and Sequence Considerations for Inserted Motifs

[0474] When adding the telomeric repeat sequence into a cassette, certain sequence orientation and context considerations help ensure that the insert does not create any unintended effects:

[0475] Strand Orientation (Sense vs. Antisense): The TTAGGG motif is specific in sequence; its reverse complement is CCCTAA. In double-stranded DNA, one strand will contain TTAGGG repeats and the opposite strand will contain CCCTAA repeats. One can insert the sequence in either or both orientations and still achieve immune inhibition. However, if the insert lies within a transcribed region (e.g., an intron or UTR), orientation can be important to avoid any impact on transcription. Specifically, a string of G-rich sequences on the non-template (coding) strand of a gene can, in some cases, form stable secondary structures (like G-quadruplexes) when that region is being transcribed, potentially pausing RNA polymerase. To prevent this, it is preferable to place the telomeric repeat in the antisense orientation relative to the gene’s coding strand when inserting into or near an expressed sequence. In practical terms, that means the strand of DNA that is not being used as the mRNA template contains the CCCTAA repeats, while the template strand (which RNA polymerase II reads) contains the TTAGGG repeats. In this configuration, as the polymerase transcribes the region, the strand that gets displaced (the non-template strand) is C-rich (CCCTAA repeats), which has a very low206678-0002-00WOtendency to form problematic structures. The G-rich strand (TTAGGG) is the one being actively paired with RNA, so it’s kept linear during transcription. This strategy was found to preserve normal transcription efficiency. In summary, for intronic inserts or any insert within the transcription unit, the TTAGGG motif is ideally on the template (antisense) strand of the gene. If the insert is in a completely non-transcribed region (like the backbone or after the polyA tail), then orientation does not matter for transcription, and one can arbitrarily choose an orientation (typically whatever is convenient for cloning) because it will not be transcribed at all. Either orientation will present identical TTAGGG motifs to the immune system in the double-stranded DNA.

[0476] Flanking and Junction Sequences: To ensure seamless integration, the telomeric repeat cassette should have a bit of “buffer” sequence at its edges, especially if near functional elements. Adding a few neutral bases (such as 2-6 nucleotides of AT-rich sequence) on each side of the insert can act as a genetic cushion. These flanking bases can prevent the creation of any cryptic signals at the junctions. For example, if one inserts a sequence bluntly, there’s a chance that the junction between the insert and the original sequence accidentally forms a new CG dinucleotide or a sequence resembling a splice site or a poly adenylation signal. By designing the insert with intentional flanks that disrupt such motifs, unintended consequences are avoided. A common practice is to use a short sequence like TA or TATA on both ends of the insert as a benign linker to the host sequence. These linkers do not code for anything and do not trigger immune sensors. In this design, it is ensured that no new CpG dinucleotide is created at the insertion site - this is easily achieved by avoiding C followed by G at the junction. Similarly, by analyzing the sequences around the insertion point, it is confirmed that no sequences like the splice donor GT or acceptor AG are inadvertently duplicated or destroyed (when inserting in introns), and that no AATAAA (polyA hexamer) or similar motif is created (when inserting near the 3' end). These checks are straightforward and part of the sequence design process.

[0477] Secondary Structure Avoidance: As mentioned, G-rich sequences can form four-stranded G-quadruplex structures under certain conditions. The telomeric insert, if made of four TTAGGG repeats, contains runs of 3 G’s separated by other bases - this can form a G-quadruplex if there are four or more TTAGGG in a row and if the DNA is single-stranded or supercoiled in certain ways. To minimize any risk of such structures affecting the plasmid, the design optionally uses spacers (as discussed) to break up long G tracts. In the contiguous 24-mer(TTAGGG×4), there are at most three consecutive G’s at any given point (the sequence is GGG, then that pattern repeats after some bases), which is on the threshold of potential quadruplex formation. Empirically, short telomeric sequences on plasmids do not appear to hinder replication or transcription significantly - and adding the spacers further reduces any chance. If one were particularly concerned, an alternative design is to split the insert into two parts that are inverse complements, which would make a perfect double-stranded structure with internal complementarity. For example, having a sequence like TTAGGG... and immediately its complement CCCTAA... in the insert. However, such design can inadvertently cause the insert to self-anneal or form a hairpin, which is undesirable. Thus, the simpler approach of moderate repeat count and optional A / T spacers for structural stability is favored. Notably, the length of the insert (-20-30 bp) is short enough that any secondary structure will be transient and unlikely to cause plasmid instability. Standard DNA oligonucleotides of this length with telomeric content have been used in laboratory and remain fairly stable.

[0478] Orientation is chosen to not interfere with transcription (when relevant), flanks are chosen to avoid creating signals, and the insert’s composition is chosen to avoid introducing any new immune triggers or structural problems.Alternative Embodiments and Variations

[0479] The invention has been described with specific preferred embodiments (such as the CFTR plasmid example with particular insert placements and sequences), but it is not limited to those examples. Many variations are possible without departing from the core concept of using telomeric DNA sequences to suppress immunity. Some alternative embodiments and design options include:

[0480] Different Repeat Lengths: While four TTAGGG repeats are exemplified, any number of repeats that provides a suppressive effect can be used. For instance, a construct might use three repeats (18 bp insert) or six repeats (36 bp insert). In some embodiments, even longer repeats (8-10 copies) could be inserted if space allows, or two shorter clusters could be placed in tandem separated by a short spacer to form a longer composite insert. The repeats could also be split into two parts of the cassette if desired (for example, two repeats in one region and two repeats in another).

[0481] Use of Linkers: The telomeric repeats may be separated by linker nucleotides as described. One embodiment might use a sequence like TTAGGG-NNN-TTAGGG-NNN-TTAGGG (SEQ ID NO:220) (where NNN is a short neutral sequence) to break up the motif. Another embodiment might use completely contiguous repeats. Both strategies fall within the invention’s scope; the choice may depend on the specific cassette’s tolerance for repeats or the ease of DNA synthesis for cloning.

[0482] Orientation Choices: In intronic or UTR contexts, the preferred orientation is antisense (as discussed) for technical reasons, but an embodiment could also place the sequence in sense orientation if testing shows no significant issue with that particular cassette (perhaps the intron is short or the repeats are few such that polymerase doesn’t pause). For inserts in nontranscribed regions, orientation is arbitrary - one could even include palindromic arrangements where one half of the insert is the complement of the other, ensuring both strands have a TTAGGG motif. As an example, an insert sequence like TTAGGGCCCTAA (SEQ ID NO:221) contains a TTAGGG on one strand and CCCTAA on the other within a short sequence. The invention encompasses any orientation that results in at least one strand of the DNA containing the TTAGGG motif.

[0483] Multiple Motif Types: While TTAGGG is the focus (being a known telomeric sequence in mammals), the invention could utilize variants or analogous repeat motifs that have similar immunosuppressive properties. For instance, certain other G-rich suppressive sequences identified in immunology (like other telomere-inspired sequences or inhibitory oligonucleotides) could be adapted and inserted. The common feature is a repetitive sequence rich in T / A and G (and lacking CpG) that can compete with immunostimulatory DNA for receptor binding. So, an embodiment might use a slightly different hexamer repeat such as TTAGGC or CTAGGG if found effective, although TTAGGG is a proven choice.

[0484] Independent of DNA Structure: The gene cassette can be cloned onto a plasmid DNA, linear DNA, viral vector, or a single strand DNA.

[0485] Plurality of Inserts: Some embodiments of the invention may include more than one telomeric insert in the same cassette. One could also imagine three or more inserts if needed - for instance, an insert in each intron of a multi-intron gene, plus one in the backbone. While likely not necessary to go to that extent, the invention allows for multiple occurrences of the immunosuppressive motif. Each instance increases the chance that no matter how the DNA is206678-0002-00WOprocessed or which part of it a sensor encounters, the sensor will find a TM nearby that tempers its activation.Example 5; Codon Optimization of Human CFTR Gene

[0486] CFTR is a large membrane glycoprotein (1480 amino acids) functioning as a cAMP -regulated chloride channel in epithelial tissues. Mutations in the CFTR gene cause cystic fibrosis, a lethal genetic disease characterized by defective ion transport in the airway epithelium. Therapeutic expression of CFTR in lung epithelial cells via gene therapy or mRNA delivery requires efficient and sustained protein production. However, the native human CFTR coding sequence is not inherently optimized for maximal expression in heterologous systems or therapeutic contexts. Factors such as suboptimal codon usage, mRNA structural elements, and spurious regulatory signals in the coding sequence can limit protein yield and stability.

[0487] Codon optimization is a proven strategy to improve recombinant protein expression by altering the DNA sequence of a gene without changing the encoded amino acids. By exploiting the redundancy of the genetic code, one can replace “rare” codons with synonymous “preferred” codons that better match the tRNA abundance and codon bias of the target host cells. Careful codon redesign should also remove inhibitory motifs in the mRNA (such as cryptic splice sites, premature polyadenylation signals, and AU-rich elements) that otherwise may trigger mRNA processing events or degradation. Similarly, adjustments to the nucleotide composition (e.g., reduced CpG dinucleotides) can enhance mRNA stability and avoid immune sensing. Importantly, synonymous codon changes must preserve the protein’s amino acid sequence and functional motifs, ensuring that the CFTR produced is identical to wild-type in structure and regulation.

[0488] Prior attempts to express CFTR at therapeutic levels have highlighted challenges that the present invention addresses. For example, certain silent polymorphisms in CFTR can markedly alter its folding and function by introducing codons ill-suited to the cellular tRNA pool. Other studies have noted that standard codon optimization (maximizing Codon Adaptation Index) can sometimes be counterproductive, potentially inducing mRNA instability or triggering cellular surveillance pathways. The disclosed invention builds on these insights by providing a comprehensive, rationally designed codon-optimized CFTR sequence that integrates multiple206678-0002-00WOdesign considerations to maximize expression, proper folding, and mRNA persistence in lung epithelial cells.

[0489] The present invention relates to genetic engineering and molecular therapeutics, specifically to a codon-optimized nucleic acid sequence encoding the human cystic fibrosis transmembrane conductance regulator (CFTR) protein. In particular, the invention provides synthetic CFTR coding sequences (SEQ ID NO:248 and SEQ ID NO:249) optimized for high-level, long-term expression in human lung epithelial cells, while preserving the wild-type CFTR amino acid sequence and proper protein folding.

[0490] Key features of the optimized sequence include: (a) codons selected according to the usage bias of human lung epithelial cells (reflecting abundant tRNAs in that cell type), (b) strategic rare codon distribution to modulate translation elongation rate and allow co-translational protein folding at domain boundaries, (c) minimized mRNA secondary structure, (e) elimination of cryptic splice donor and acceptor sequence motifs that could cause aberrant mRNA splicing, (f) removal of internal premature polyadenylation signals that could truncate the transcript, (g) removal or substantial reduction of CpG dinucleotides to mitigate CpG-mediated immunogenicity and methylation-induced gene silencing, (h) avoidance of tandem repeat sequences or homopolymeric runs that might induce polymerase slippage or genetic instability, (i) retention of all essential post-translational modification motifs (such as phosphorylation sites and trafficking signals) inherent to the CFTR protein sequence, and (j) avoidance of AU-rich instability motifs within the coding sequence. Collectively, these design elements synergistically confer robust and durable CFTR expression in target cells, thereby supporting sustained therapeutic effect with reduced need for re-dosing.Codon Usage Optimized Based on tRNA Bias

[0491] The codon-optimized sequences (SEQ ID NO:248 and SEQ ID NO:249) are tailored to the codon usage preferences of human cells, where CFTR primarily needs to be expressed for cystic fibrosis therapy. Optimal codons were chosen based on their prevalence in highly expressed human genes and correspondence to abundant tRNAs. By aligning the codon usage with the host cell’s tRNA pool, translation efficiency is maximized and ribosome stalling is minimized. In SEQ ID NO:248 and SEQ ID NO:249, many rare codons present in the native CFTR gene are replaced with synonymous codons more frequently used in (and efficientlytranslated by) human cells. This increases the Codon Adaptation Index (CAI) for human expression and ensures that the translational machinery in airway epithelia can synthesize CFTR protein rapidly and in ample quantity.

[0492] Notably, the codon changes do not alter the amino acid sequence of CFTR in any way - all 1480 amino acids encoded by SEQ ID NO:248 and SEQ ID NO:249 are identical to those of the wild-type CFTR protein. The numerous protein phosphorylation sites in the regulatory (R) domain, the N-linked glycosylation site in the extracellular loop, the PDZ-binding motif at the C-terminus, and other CFTR-specific motifs remain intact at the amino acid level. By maintaining the exact protein sequence, the optimization preserves CFTR’s normal biochemical functions and regulatory interactions, changing only the nucleic acid sequence to enhance expression.Controlled Translation Elongation and Folding (Codon Pausing Strategy)

[0493] The invention also introduces a deliberate tuning of translation elongation rates along the CFTR coding sequence. Rather than simply maximizing translation speed at every position, SEQ ID NO:248 and SEQ ID NO:249 employ strategic placements of moderately rare codons in regions corresponding to domain junctions and structurally important regions of the protein. The human CFTR protein comprises multiple domains - two membrane-spanning domains (MSD1 and MSD2), two nucleotide-binding domains (NBD1 and NBD2), and a regulatory R domain - that must fold and assemble co-translationally. Translational “pause” sites have been engineered near the boundaries of these domains to mimic the natural slowdowns observed in many multi-domain proteins.

[0494] Scientific evidence suggests that stretches of suboptimal codons at domain boundaries can facilitate correct folding by giving nascent polypeptide domains time to form structure before the next domain is synthesized. Naturally, evolutionary codon usage often encodes translationally slow regions at such junctions. In SEQ ID NO:248 and SEQ ID NO:249, synonymous codons with slightly lower translation rates (as determined by tRNA availability or known ribosome kinetics) are inserted just before or within the linkers connecting CFTR’s major domains. These controlled pauses are short and carefully placed so as not to trigger ribosomal drop-off or mRNA decay, but sufficient to allow proper folding of MSD1 before NBD1 is fully translated, NBD1 before the R domain, and so forth. This approach reduces the probability of206678-0002-00WOmisfolding and proteostatic stress, thereby increasing the yield of functional CFTR protein. The overall result is a smoother translation process that aligns with CFTR’s folding timeline, producing a properly conformed channel protein despite its complexity.mRNA Secondary Structure Minimization

[0495] The synonymous substitutions in SEQ ID NO:248 and SEQ ID NO:249 are additionally chosen to minimize stable mRNA secondary structures within the coding region, especially near the 5' end (starting from the AUG start codon). Stable hairpins or stem-loop structures in the mRNA can impede ribosome binding and scanning, or stall the ribosome during elongation. The optimized sequence avoids nucleotide combinations that w...

Claims

206678-0002-00WOCLAIMSWhat is claimed is:

1. A plasmid for controlled transcription in a target cell type, the plasmid comprising at least one transcription factor binding motif for a transcription factor that is expressed in the target cell type or a subset of cells of the target cell type.

2. The plasmid of claim 1, wherein the target cell type is lung cells, or a subset thereof, and the transcription factor binding motif is specific for binding to a transcription factor selected from the group consisting of Nuclear Factor I A (NFIA), Forkhead Box A2 (FOXA2), Multiciliate Differentiation and DNA Synthesis-Associated Cell Cycle Protein (MCIDAS), SAM-pointed Domain ETS Factor (SPDEF), Forkhead Box II (FOXI1), Transcription Factor CP2-like 1 (TFCP2L1), Achaete-Scute Family bHLH 3 (ASCL3), Tumor Protein p63 (TP63), Forkhead Box QI (FOXQ1), Forkhead Box JI (FOXJ1), Regulatory Factor X3 (RFX3), SRY-box 9 (SOX9), Hepatocyte Nuclear Factor la (HNFla), Hepatocyte Nuclear Factor 4a (HNF4a), Forkhead Box A3 (FOXA3), or any combination thereof.

3. The plasmid of claim 2, wherein the plasmid comprises at least one transcription factor binding motif selected from the group consisting of:a) SEQ ID NO: 1, SEQ IDNO:2, SEQ IDN0:3, SEQ IDN0:4, SEQ ID NO:5, SEQ IDN0:6, SEQ IDN0:7, GCCACTTAA, CCCACTTAA, GCCACTTAG, ACCACTTAG, GGCACTTAA, AGCACTTAA, CCCACTTAG, TCCACTTAA, SEQ ID NO:95 or SEQ ID NO:96, or a fragment or variant thereof which serves as a binding site for NFIA;b) AATAAAG, ATAAACA, GTAAATA, GTAAACA, GTAAACAA, ATAAAT, GTAAAT, TGTTTAC, TGTTTAT, SEQ ID NO:8, SEQ ID NO:9, SEQ ID NO: 10, SEQ ID NO: 11, SEQ ID NO: 12, SEQ ID NO: 13, SEQ ID NO: 14, SEQ ID NO: 15, SEQ ID NO: 16, SEQ ID NO: 17, SEQ ID NO: 18 or a fragment or variant thereof which serves as a binding site for FOXA2;c) TTTCGCGC, TTTGGCGC, TTTCCGCC, TTTCCCGC, TTTGCCGC, or TGTCCCGC, or a fragment or variant thereof which serves as a binding site for MCIDAS;206678-0002-00WOd) GGAT, AGGAT, TGGAT, CGGAT, GGAA, AGGATTC, ATGCGGGC, GTGCGGGT, GTGCGGGC, ATACGGGT, ATGGGGGT, ATGCGGGG, CTGCGGGT, ATACGGGC, ATGCGGGA or SEQ ID NO: 19 or a fragment or variant thereof which serves as a binding site for SPDEF;e) TGTTTAC, GTAAACA, GTAAATA, TATTTAT, TGTTTAT, TGTTTGT, TATTTAC, GTCAACA, GTAATCA, ATAAACA, ATCAACA, GTAAAAA, GTAAATAA, GTCAATA or SEQ ID NO:20, or a fragment or variant thereof which serves as a binding site for FOXI1;f) AAACCGGTT, SEQ ID NO:21, SEQ ID NO:22, CCAGTTCAA, CAGTTCAAC, or SEQ ID NO:23, or a fragment or variant thereof which serves as a binding site for TFCP2L1;g) CAGGTG, CACCTG, CACGTG, GCACCTGCC, CCACCTGCC, GCACCTGCT, ACACCTGCC, GCACCTGCA, GCACCTGCG, CCACCTGCT, TCACCTGCC, GCAGCTGCC, or GCACCTGGC, or a fragment or variant thereof which serves as a binding site for ASCL3;h) SEQ ID NO:24, SEQ ID NO:25, SEQ ID NO:26, SEQ ID NO:27, SEQ ID NO:28, SEQ ID NO:29, SEQ ID NO:30 or SEQ ID NO:31, or a fragment or variant thereof which serves as a binding site for TP63;i) ACAAAG, ATAAAG, ACAAAT, ATAAAC, ATAAAT, GTAAAC, TGTTTAC, TCAATA, GTAAATAA or SEQ ID NO:32, or a fragment or variant thereof which serves as a binding site for FOXQ1;j) TGTTTAC, GTAAATA, GTTTACA, ATAAAT A, GTAAAC AAA, ATAAACAAA, ATAAACAA, TAAACAAA, or SEQ ID NO:33, or a fragment or variant thereof which serves as a binding site for FOXJ1;k) SEQ ID NO:34, SEQ ID NO:35, SEQ ID NO:36, GTTACCATG, GTTGCTATG, GTTACTATG, or SEQ ID NO:37, or a fragment or variant thereof which serves as a binding site for RFX3;l) SEQ ID NO:38, SEQ ID NO:39, SEQ ID NO:40, CATTGAA, CTTTGTT, CTTTGAA, ACAAAG, or TTCAAAG, or a fragment or variant thereof which serves as a binding site for SOX9;206678-0002-00WOtn) GTTAAT, TTGTTA, SEQ ID NO:41, SEQ ID NO:47, SEQ ID NO:43, SEQ ID NO:44, SEQ ID NO:45, or SEQ ID NO:46 or a fragment or variant thereof which serves as a binding site for HNFl;n) SEQ ID NO:47, SEQ ID NO:48, SEQ ID NO:49, SEQ ID NO:50, or SEQ ID NO: 51, or a fragment or variant thereof which serves as a binding site for HNF4a; and o) GTAAACA, ATAAATA, ATCAATA, ATAAAT, ATAAAC, GTAAATAA, or GTAAAT, or a fragment or variant thereof which serves as a binding site for FOXA3.

4. The plasmid of claim 3, wherein the plasmid comprises a cluster of transcription factor binding sites for recognition by two or more of NFIA, FOXA2, MCIDAS, SPDEF, FOXI1, TFCP2L1, ASCL3, TP63, FOXQ1, FOXJ1, RFX3, SOX9, HNFla, HNF4a, and FOXA3.

5. The plasmid of claim 4, wherein the plasmid comprises a sequence selected from the group consisting of SEQ ID NO: 121, SEQ ID NO: 122, and SEQ ID NO: 125.

6. The plasmid of claim 1, wherein the target cell type is muscle cells, or a subset thereof, and the transcription factor binding motif is specific for binding to a transcription factor selected from the group consisting of Myocyte enhancer factor 2A (MEF2A), Myogenin (MYOG), Muscle-Specific Regulatory Factor 4 (MRF4), TEA Domain Transcription Factor 1 (TEAD1), Paired Box 7 (PAX7), and SIX Homeobox 1 (SIX1), or any combination thereof.

7. The plasmid of claim 6, wherein the plasmid comprises at least one transcription factor binding motif selected from the group consisting of:a) SEQ ID NO:52, SEQ ID NO:53, SEQ ID NO:54, SEQ ID NO:55, SEQ ID NO:56, SEQ ID NO:57, SEQ ID NO:58, SEQ ID NO:59, SEQ ID NO:60, SEQ ID NO:61 or SEQ ID NO: 62, or a fragment or variant thereof which serves as a binding site for MEF2A;b) SEQ ID NO:63, SEQ ID NO:64, SEQ ID NO:65, SEQ ID NO:66, SEQ ID NO:67, SEQ ID NO:68, SEQ ID NO:69, SEQ ID NO:70, SEQ ID NO:71, SEQ ID NO:72 or SEQ ID NO: 73, or a fragment or variant thereof which serves as a binding site for MYOG;c) SEQ ID NO: 74, SEQ ID NO: 75, SEQ ID NO: 76, SEQ ID NO: 77, SEQ ID NO:78, SEQ ID NO:79, SEQ ID NO:80, SEQ ID NO:81, SEQ ID NO:82, or SEQ ID NO:83, or a fragment or variant thereof which serves as a binding site for MRF4;d) SEQ ID NO:84, SEQ ID NO:85, SEQ ID NO:86, SEQ ID NO:87, SEQ ID NO:88, SEQ ID NO:89, SEQ ID NO:90, SEQ ID NO:91, SEQ ID NO:92, or SEQ ID NO:93, or a fragment or variant thereof which serves as a binding site for TEAD1;e) SEQ ID NO:94, or a fragment or variant thereof which serves as a binding site for PAX7; andf) CTAATTA, CTCATTA, TTAATTA, CAAATTA, CCAATTA, TTCATTA, CACATTA, CCCATTA, CTAATTG, or TAAATTA, or a fragment or variant thereof which serves as a binding site for SIX1.

8. The plasmid of claim 7, wherein the plasmid comprises a cluster of transcription factor binding sites for recognition by two or more of MEF2A, MYOG, MRF4, TEAD1, PAX7, and SIX1.

9. The plasmid of claim 8, wherein the plasmid comprises a sequence of SEQ ID NO: 128.

10. A promoter comprising a combination of:a) a TFIIB recognition motif (referred to as a BREu element); b) a TATA box motif;c) an initiator; andd) a downstream promoter element (DPE).

11. The expression plasmid of claim 10, wherein:a) the BREu element comprises a sequence selected from the group consisting of GGGCGCC, GCACGCC, and GCGCGCC;b) the TATA box motif comprises a sequence selected from the group consisting of TATATAA, TATAAAAG, TATATAAG, TATAAA, SEQ ID NO:264 and SEQ ID NO:265;c) the initiator comprises a sequence selected from the group consisting of TCAGTT, TG4TTC, TG4GTCT, TG4TATC, TCAGTTCC, GC4GTT, and CCACTT; and d) the DPE comprises a sequence selected from the group consisting of GGACCT, ACCT, AGTCGC, GGACTGG, GGTTTC, AGACGTG and AGACGT.

12. The expression plasmid of claim 10, wherein the plasmid further comprises a Motif Ten Element (MTE) located between the initiator and the DPE.

13. The expression plasmid of claim 12, wherein the MTE comprises a sequence selected from the group consisting of SEQ ID NO: 145, SEQ ID NO: 146, SEQ ID NO: 147, SEQ ID NO: 148, SEQ ID NO: 149, SEQ ID NO:271 and SEQ ID NO:272.

14. The expression plasmid of any one of claims 10-13, wherein the plasmid comprises a sequence selected from the group consisting of: SEQ ID NO: 158, SEQ ID NO: 159, SEQ ID NO: 160, SEQ ID NO: 161, SEQ ID NO: 162, SEQ ID NO: 163, SEQ ID NO: 164, SEQ ID NO: 165, SEQ ID NO: 166, SEQ ID NO: 167, SEQ ID NO: 168, SEQ ID NO: 169, SEQ ID NO: 170, SEQ ID NO: 171, SEQ ID NO: 172, SEQ ID NO: 173, SEQ ID NO: 174, SEQ ID NO:175, SEQ ID NO:176, SEQ ID NO:177, SEQ ID NO:178, SEQ ID NO:179, SEQ ID NO:180, SEQ ID NO:181, SEQ ID NO:182, SEQ ID NO:183, SEQ ID NO:184, SEQ ID NO: 185, SEQ ID NO: 186, SEQ ID NO: 187, SEQ ID NO: 188, SEQ ID NO: 189, SEQ ID NO: 190, SEQ ID NO: 191, SEQ ID NO: 192, SEQ ID NO: 193, SEQ ID NO: 194, SEQ ID NO: 195, SEQ ID NO: 196, SEQ ID NO: 197, SEQ ID NO: 198, SEQ ID NO: 199, SEQ ID NO: 200, SEQ ID NO: 201, SEQ ID NO: 202, SEQ ID NO: 203, SEQ ID NO: 204, SEQ ID NO:205, SEQ ID NO:206, SEQ ID NO:207, SEQ ID NO:208, SEQ ID NO:209, SEQ ID NO:210, SEQ ID NO:211, SEQ ID NO:212, SEQ ID NO:213, SEQ ID NO:214, SEQ ID NO:215, SEQ ID NO:279, SEQ ID NO:280, SEQ ID NO:281, SEQ ID NO:282, SEQ ID NO:283, or SEQ ID NO:284.

15. An expression plasmid comprising at least one telomeric repeat motif between a sequence which serves as a promoter and a start codon of a coding sequence.206678-0002-00WO16. The expression plasmid of claim 15, wherein the telomeric repeat motif comprises a sequence selected from the group consisting of SEQ ID NO:216, SEQ ID NO:217, SEQ ID NO:218, SEQ ID NO:219, SEQ ID NO:220, SEQ ID NO:221, SEQ ID NO:222, SEQ ID NO: 223, SEQ ID NO: 224, SEQ ID NO: 225, and SEQ ID NO: 226.

17. The expression plasmid of claim 15 further comprising an intron downstream of the promoter and upstream of the start codon.

18. The expression plasmid of claim 17, wherein the intron comprises a sequence selected from the group consisting of SEQ ID NO:229, SEQ ID NO:230 and SEQ ID NO:231.

19. The expression plasmid of claim 17, wherein the telomeric repeat motif is within the intron.

20. The expression plasmid of claim 19, wherein the intron comprising the telomeric repeat comprises a sequence selected from the group consisting of SEQ ID NO:232, SEQ ID NO:233, SEQ ID NO:234, SEQ ID NO:235, SEQ ID NO:236, SEQ ID NO:237, SEQ ID NO:238, and SEQ ID NO:239.

21. An expression plasmid comprising at least one post-polyA telomeric repeat motif following the polyA tail of a coding sequence.

22. The expression plasmid of claim 21, wherein the post-polyA telomeric repeat motif comprises SEQ ID NO:227 or SEQ ID NO:228.

23. The expression plasmid of any one of claim 15 to 20 further comprising at least one post-polyA telomeric repeat motif of any one of claim 21 and 22.

24. An expression plasmid comprising a 5’ UTR selected from the group consisting of SEQ ID NO: 240, SEQ ID NO: 241, SEQ ID NO: 242, SEQ ID NO: 243, SEQ ID NO:244, and SEQ ID NO:245.206678-0002-00WO25. An expression plasmid comprising a 3’ UTR selected from the group consisting of SEQ ID NO:252, SEQ ID NO:253, SEQ ID NO:254 or SEQ ID NO:255.

26. An expression plasmid comprising a polyA tail comprising a sequence as set forth in SEQ ID NO:256.

27. A codon optimized sequence encoding cystic fibrosis transmembrane conductance regulator (CFTR) comprising SEQ ID NO: 248 or SEQ ID NO:249, or a variant thereof encoding SEQ ID NO: 247.

28. A polycistronic coding sequence encoding a combination of a codon optimized sequence encoding green fluorescent protein (GFP) and a codon optimized sequence encoding CFTR of claim 27.

29. The polycistronic coding sequence of claim 28, wherein the coding sequence comprises SEQ ID NO:251, or a variant thereof encoding SEQ ID NO:250.

30. A plasmid for controlled transcription in a target cell type of any one of claims 1 to 9 further comprising at least one of:i) the promoter of the plasmid of any one of claims 10 to 14;ii) the telomeric repeat motif of the plasmid of any one of claims at least one telomeric repeat motif of any one of claims 15 to 23; iii) the 5’ UTR of the plasmid of claim 24;iv) the 3’ UTR of the plasmid of claim 25; and v) the polyA tail of the plasmid of claim 26.

31. The plasmid of claim 30 further comprising the codon optimized sequence encoding CFTR of claim 27 or the polycistronic coding sequence of any one of claims 28 or 29.

32. The plasmid of claim 31, further comprising at least one transcription factor repressor binding site for EHF or IRF2 downstream of the promoter but upstream of the 5’ UTR.206678-0002-00WO33. The plasmid of claim 30, comprising a sequence selected from the group consisting of SEQ ID NO:257, SEQ ID NO:258, SEQ ID NO:259, SEQ ID NO:260, SEQ ID NO:261, SEQ ID NO:262 and SEQ ID NO:263.

34. A delivery vehicle comprising a plasmid of any one of claims 1 to 33.

35. The delivery vehicle of claim 34 comprising a liposome or lipid nanoparticle (LNP).

36. A method of treating cystic fibrosis in a subject in need thereof, the method comprising administering to a subject in need thereof a plasmid of any one of claim 31 to 33 or a delivery vehicle comprising a plasmid of any one of claim 34 or 35.