Genome orthogonal zinc finger DNA binding domains and related enhancer elements and uses thereof to modulate gene expression

Synthetic transcription factors with orthogonal enhancer elements and zinc finger proteins allow precise modulation of gene expression, addressing gene therapy challenges by targeting specific genes without affecting others, ensuring effective and targeted therapeutic outcomes.

WO2026161539A2PCT designated stage Publication Date: 2026-07-30GENEFAB LLC
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
GENEFAB LLC
Filing Date
2026-01-22
Publication Date
2026-07-30

AI Technical Summary

Technical Problem

Existing gene therapy approaches face challenges such as unwanted immune responses, off-target effects, limitations on vector capacity, and lack of sustained therapeutic effect, necessitating a need for targeted and specific methods to modulate gene expression without impacting non-target genes.

Method used

Development of synthetic transcription factors comprising a transcriptional effector domain, synthetic enhancer elements, and DNA binding domains, specifically engineered to target and modulate expression of a gene of interest without affecting other genes, using zinc finger proteins and orthogonal enhancer elements.

Benefits of technology

Enables precise modulation of gene expression by initiating or terminating therapeutic gene expression in response to stimuli, minimizing off-target effects and maintaining specificity to the target gene, thus enhancing therapeutic efficacy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2026012126_30072026_PF_FP_ABST
    Figure US2026012126_30072026_PF_FP_ABST
Patent Text Reader

Abstract

The present disclosure provides DNA binding proteins engineered to bind selectively to engineered synthetic enhancer elements that are orthogonal to a host genome. Introduction of the DNA binding protein and the synthetic enhancer is used, in several embodiments, to modulate gene expression of a gene of interest without disruption or alteration of other genes within the host cell genome. In several embodiments, the DNA binding protein selected is a zinc finger protein.
Need to check novelty before this filing date? Find Prior Art

Description

Patent Application 97157.00116GENOME ORTHOGONAL ZINC FINGER DNA BINDING DOMAINS AND RELATED ENHANCER ELEMENTS AND USES THEREOF TO MODULATE GENE EXPRESSIONCross Reference to Related Applications

[0001] This application claims the benefit of United States Provisional Patent Application No. 63 / 749,120, filed on January 24, 2025, the entire contents of which is incorporated by reference herein.Incorporation by Reference of Material in Sequence Listing File

[0002] This application incorporates by reference the Sequence Listing contained in the following XML file being submitted concurrently herewith: File name: 97157.00116.xml; created on January 20, 2026, and is 1,421,169 bytes in size.Field

[0003] This application relates to a variety of zinc finger proteins (ZFPs), related synthetic enhancer elements, and related methods utilizing such proteins and / or enhancer elements for methods and / or use in regulating gene expression.Background

[0004] Zinc finger proteins (ZFPs) are proteins that can bind to DNA in a sequencespecific manner. Various methods and compositions for targeted binding of genomic DNA are known. There is a need for enhanced specificity for targeting genomic DNA for a variety of purposes, such as targeted cleavage of DNA(e.g., gene editing) or modulation of expression of genes and / or related proteins. Development of specifically targeted ZFPs to modulate gene and / or related protein expression are provided for herein.Summary

[0005] Numerous diseases are associated with abnormal expression of genes. Depending on the disease, a genetic mutation in a gene can lead to dysregulated expression, downregulated expression, or substantially complete or complete lack of expression. Likewise, while expression may be normal, a mutation in a gene may lead to compromised or abnormal (e.g., over or under) functionality of the encoded protein(s). In some cases, a genetic mutation in a gene causes it to be upregulated. resulting in overexpression of the gene and in some cases the encoded protein(s). A host of challenges face successful treatment of genetic disorders orPatent Application 97157.00116diseases. One approach is gene therapy, which involves therapeutic delivery of a nucleic acid into a patient's cell. However, certain gene therapy approaches remain hampered by unwanted immune response elicited by gene therapy, off-target effects, limitations on gene therapy vector capacity, and / or lack of sustained therapeutic effect. In some instances, a targeted approach can be used to introduce a regulatable expression system, also referred to as a gene circuit, in which expression of a desired therapeutic gene is initiated (or terminated) in response to a particular stimulus. There remains a need for targeted and specific methods of modulation of gene expression in a host-genome orthogonal fashion (i.e., without impacting the expression of other, endogenous, non-target genes of the host).

[0006] In several embodiments, there is provided for herein a synthetic transcription factor, comprising a transcriptional effector domain, a synthetic enhancer element, and a DNA binding domain, wherein the synthetic enhancer element comprises a structure of Formula 1 :(Target-(Spacer)(x-l))(x)Formula 1

[0007] wherein the Target of Formula 1 comprises a sequence of (GXiX2)nNy, wherein Xi and X2 each independently represent one of adenine, guanine, cytosine, and thymine, and wherein the synthetic enhancer element (i) does not have 100% sequence identity to any human or murine genomic DNA sequence, (ii) does not have fewer than 3 base pair mismatches with any human or murine genomic DNA sequence, and (iii) does not have any putative transcription factor binding sites (a) within the Target or Spacer sequences or (b) spanning a junction between a Target and its corresponding Spacer, wherein the DNA binding domain binds a target sequence and allows the transcriptional effector to modulate transcription of a target gene, and wherein the synthetic transcription factor does not modulate transcription of non-target genes with a cell. In several embodiments, n is an integer between 2 and 8, such as 2. 3, 4, 5, 6. 7, or 8. In several embodiments, N is selected from adenine, guanine, cytosine, and thymine. In several embodiments, y is an integer between 0 and 6, such as 0, 1, 2, 3, 4, 5, or 6. In several embodiments, the Spacer of Formula 1 comprises a sequence of betw een 6 and 12 nucleotides, optionally, 6, 7, 8, 9, 10, 11 or 12 nucleotides, depending on the embodiment. In several embodiments, x is an integer between 3 and 8, such as 3, 4, 5, 6, 7, or 8. According to several embodiments, the DNA binding domain comprises a zinc finger protein. In several embodiments, within the Target portion of Formula 1, n=6. In several embodiments, the synthetic enhancer element comprises at least one copy of SEQ ID NO: 24 (GXXGXXGXXGXXGXXGXX, wherein each instance of X is any nucleotide). In some embodiments, two, three, four, or more copies of the SEQ ID NO: 24 are used in the syntheticPatent Application 97157.00116enhancer element.

[0008] In some embodiments of the synthetic transcription factor, the first GX1X2 sequence does not comprise GGC, GGA, GAT, GCA, or GAA. In several embodiments, the GX1X2 sequence does not comprise GTT, GCG, GGC, GGA, GAT, GCA, or GAA. In several embodiments, the third GX1X2 sequence does not comprise GTT, GGA, GAT, GCA, or GAA. In several embodiments, the fourth GX1X2 sequence does not comprise GTT, GCA, or GAA. In several embodiments, the fifth GX1X2 sequence does not comprise GTT, GCA, or GAA. In several embodiments, the sixth GX1X2 sequence does not comprise GTC, GGC, or GCC. According to several embodiments, none of the GX1X2 sequences comprises GGG, GGT, GAG, GTG, or GCT

[0009] In several embodiments, the Target of Formula 1 is selected from the group consisting of SEQ ID NO: 1-23. In several embodiments, the Target of Formula 1 is selected from the group consisting of SEQ ID NO: 25-47.

[0010] In several embodiments of the synthetic transcription factors provided for herein, x=4, resulting in four Target repeats intercalated with three Spacer repeats. In several embodiments, y is 1 and the Spacer of Formula 1 comprises 10 nucleotides. In several embodiments, the synthetic enhancer element is selected from a sequence having at least 70%, at least 75%, or at least 80% identity to one or more of SEQ ID NO: 50-73. In several embodiments, the synthetic enhancer element is selected a from sequence having at least 90% or at least 95% identity to one or more of SEQ ID NO: 50-73. In several embodiments, the synthetic enhancer element is selected from the group consisting of SEQ ID NO: 50-73. In several embodiments, the synthetic enhancer element is selected from the group consisting of SEQ ID NO: 55, 65, and 68. In several embodiments, the synthetic enhancer element comprises SEQ ID NO: 55. In several embodiments, the synthetic enhancer element comprises SEQ ID NO: 65. In several embodiments, the synthetic enhancer element comprises SEQ ID NO: 68.

[0011] In several embodiments, the synthetic transcription factors provided for herein further comprise a promoter sequence. In several embodiments, the promoter sequence comprises a YBTATA minimal promoter sequence. In several embodiments, the promoter sequence has at least 85%, at least 90%, or at least 95% sequence identity to SEQ ID NO. 49.

[0012] In several embodiments, the synthetic transcription factors provided for herein comprise a synthetic enhancer element further comprising a promoter and a gene of interest. In several embodiments, the synthetic enhancer element comprises a nucleic acid sequence having at least 85% sequence identity to one or more of the nucleic acid sequences of SEQ IDPatent Application 97157.00116NOs 50-73. In several embodiments, the synthetic enhancer element comprises a nucleic acid sequence having at least 90% sequence identity to one or more of the nucleic acid sequences of SEQ ID NOs 50-73. In several embodiments, the synthetic enhancer element comprises a nucleic acid sequence having at least 95% sequence identity to one or more of the nucleic acid sequences of SEQ ID NOs 50-73. In several embodiments, the synthetic enhancer element comprises a nucleic acid sequence of one or more of the nucleic acid sequences of SEQ ID NOs 50-73.

[0013] In several embodiments, the synthetic transcription factor provided for herein comprise spacers having different sequences from one another.

[0014] In several embodiments, there is provided for herein a method of identifying at least one candidate genome orthogonal synthetic enhancer element, the method comprising (i) generating a plurality of oligonucleotides comprising the following sequence: (GXiX2)nNy, wherein Xi and X2 each independently represent one of adenine, guanine, cytosine, and thymine, wherein n is an integer between 2 and 8 (such as 2, 3, 4, 5, 6, 7, or 8), wherein N is selected from adenine, guanine, cytosine, and thymine, wherein y is an integer between 0 and 6 (such as 0, 1, 2, 3, 4, 5, or 6), (li) screening the plurality of oligonucleotides from (1) for (a) an exact matching sequence for the oligonucleotide from the human and / or murine genome and (b) a sequence within the human and / or murine genome having 1 or 2 base pair mismatches from an individual oligonucleotide selected from the plurality of oligonucleotides from (i), (iii) removing from consideration an oligonucleotide meeting either (a) and / or (b) from (ii), (iv) screening the plurality of oligonucleotides from (iii) for putative binding sites for one or more transcription factors, and (v) removing from consideration an oligonucleotide having at least one putative transcription factor binding site, thereby generating at least one candidate genome orthogonal synthetic enhancer element.

[0015] In several embodiments, when n>2 and the method is for designing a spacer oligonucleotide between 8 to 12 nucleotides long, the spacer does not comprise any putative transcription factor binding sites within the 8 to 12 nucleotides and when inserted between a first GX1X2 sequence and a second GX1X2 sequence does not introduce any putative transcription factor binding sites spanning a junction between the first GX1X2 sequence and the spacer or the second GX1X2 sequence and the spacer. In several embodiments, the synthetic enhancer element has a sequence of GX1X2GX1X2GX1X2GX1X2GX1X2GX1X2N (SEQ ID NO: 48).

[0016] In several embodiments, wherein when n=6. the method further comprising removing an oligonucleotide from consideration when:(i) the first GX1X2 sequence comprisesPatent Application 97157.00116GGC, GGA, GAT. GCA, or GAA; (ii) the second GX1X2 sequence comprises GTT, GCG, GGC, GGA, GAT, GCA. or GAA; (iii) the third GX1X2 sequence comprises GTT, GGA, GAT, GCA, or GAA; (iv) the fourth GX1X2 sequence comprises GTT, GCA, or GAA; (v) the fifth GX1X2 sequence comprises GTT, GCA, or GAA; or (vi) the sixth GX1X2 sequence comprises GTC, GGC, or GCC. In several embodiments, the method further comprises removing an oligonucleotide from consideration when any of the GX1X2 sequences comprises GGG, GGT, GAG, GTG, or GCT

[0017] In several embodiments, there is provided a method of generating at least one candidate genome orthogonal synthetic enhancer element, comprising (i) generating a plurality of in silico oligonucleotides comprising the following sequence: (GXiX2)nN, (ii) screening the plurality of in silico oligonucleotides from (i) for (a) an exact matching sequence for the oligonucleotide from the human and / or murine genome and (b) a sequence within the human and / or murine genome having 1 or 2 base pair mismatches from an individual oligonucleotide selected from the plurality of oligonucleotides from (i); (iii) removing from consideration an in silico oligonucleotide meeting either (a) and / or (b) from (ii); (iv) screening the plurality’ of in silico oligonucleotides from (iii) for putative binding sites for one or more transcription factor; (v) removing from consideration an in silico oligonucleotide having at least one putative transcription factor binding site; (vi) designing and inserting, in silico, a spacer oligonucleotide between 8 to 12 nucleotides long, wherein the spacer does not: (a) comprise any putative transcription factor binding sites within the 8 to 12 nucleotides or, (b) when inserted between a first GX1X2 sequence and a second GX1X2 sequence, introduce any putative transcription factor binding sites spanning a junction between the first GX1X2 sequence and the spacer or the second GX1X2 sequence and the spacer; (vii) removing from consideration any in silico oligonucleotide meeting either (a) or (b) from (vi); and (viii) synthesizing at least one oligonucleotide from those in silico oligonucleotides remaining after (vii), thereby synthesizing at least one candidate genome orthogonal synthetic enhancer element. In several embodiments, Xi and X2 each independently represent one of adenine, guanine, cytosine, and thymine. In several embodiments, n is an integer between 2 and 8 (such as 2, 3, 4, 5, 6, 7, or 8). In several embodiments, N is selected from adenine, guanine, cytosine, and thymine.

[0018] In several embodiments, there is provided a method of generating a genome orthogonal synthetic transcription factor, comprising (i) generating a plurality’ of genome orthogonal synthetic enhancer element according to the methods provided for herein; (ii) coupling at least one of the plurality of genome orthogonal synthetic enhancer elements of (i) to a promoter and a gene of interest; (iii) designing a DNA binding protein that specificallyPatent Application 97157.00116binds to only one of the plurality of genome orthogonal synthetic enhancer elements of (i); and (iv) coupling the DNA binding protein to a promoter and a transcriptional activator domain. In several embodiments, the binding of the DNA binding domain to the synthetic enhancer element is configured to initiate transcription of the gene of interest without modulation of transcription of genes with a cell that are not the gene of interest. In several embodiments, the DNA binding domain comprises a zinc finger protein; and

[0019] In several embodiments, there is provided a for modulating expression of a gene of interest, comprising: (i) introducing into a host cell a synthetic transcription factor, the synthetic transcription factor comprising a synthetic enhancer assembly and a zinc finger protein assembly, wherein the synthetic enhancer assembly comprises: a synthetic enhancer sequence comprising a sequence that is unique or has more than one or two base pair mismatches with respect the sequence of a host genome; a promoter element; and a gene of interest; wherein the zinc finger protein assembly comprises: a promoter element; azinc finger array comprising a plurality of zinc finger proteins engineered to bind specifically to the synthetic enhancer; and a transcriptional activator coupled to the zinc finger array; (ii) allowing integration of the synthetic enhancer assembly into the host genome; (iii) allowing expression of the zinc finger array and transcriptional activator, wherein the expressed zinc finger array allows the transcriptional activator to initiate transcription of the gene of interest, but does not alter transcription of non-gene of interest genes of the host genome.

[0020] In several embodiments, the synthetic transcription factor is introduced into the host cell via viral delivery. In several embodiments, the viral delivery comprises use of a lentiviral vector. In several embodiments, the gene of interest encodes a therapeutic protein. In several embodiments, the therapeutic protein comprises a chimeric antigen receptor, a cytokine, an antibody, or a protein for which the expression and / or function of a corresponding endogenous protein of the host is compromised.

[0021] In several embodiments, modulation of expression comprises introducing or overexpressing the gene of interest. In several embodiments, the expression is regulated, e.g., by an additional molecule, such as a drug or chemical compound. In several embodiments, modulation of expression comprises reducing or eliminating expression of another host protein due to transcription of the gene of interest.

[0022] In several embodiments, there is provided a zinc finger protein array comprising a plurality of zinc fingers separated by at least one linker element. In several embodiments, the zinc finger array is configured to bind with specificity to a nucleic acid sequence of SEQ ID NO: 24. In several embodiments, there is provided a zinc finger protein array comprisingPatent Application 97157.00116a plurality of zinc fingers separated by at least one linker element, the zinc finger array configured to bind with specificity to a nucleic acid sequence selected from one or more of SEQ ID NOs 1-23, 25-47, or 74-96. In several embodiments, the individual ZFPs within the array have at least 85% sequence identity to one or more of the amino acids of SEQ ID NOs 347-622. In several embodiments, the individual ZFPs within the array have at least 90% sequence identity to one or more of the amino acids of SEQ ID NOs 347-622. In several embodiments, the individual ZFPs within the array have at least 95% sequence identity to one or more of the amino acids of SEQ ID NOs 347-622. In several embodiments, the individual ZFPs within the array are selected from one or more of the amino acids of SEQ ID NOs 347-622. In several embodiments, the ZFP array comprises an amino acid sequence having at least 85% sequence identity to one or more of the amino acids of SEQ ID NOs 244-289. In several embodiments, the ZFP array comprises an amino acid sequence having at least 90% sequence identity7to one or more of the amino acids of SEQ ID NOs 244-289. In several embodiments, the ZFP array comprises an amino acid sequence having at least 95% sequence identity to one or more of the amino acids of SEQ ID NOs 244-289. In several embodiments, the ZFP array comprises an amino acid sequence of one or more of SEQ ID NOs 244-289. In several embodiments, the ZFP array is encoded by a nucleic acid having at least 85%, at least 90%, or at least 95% sequence identity’ to one or more of the nucleic acids of SEQ ID NOs 98-143. In several embodiments, the ZFP array is encoded by a nucleic acid selected from the group consisting of one or more of the nucleic acids of SEQ ID NOs 98-143.

[0023] In several embodiments, the zinc finger arrays provided for herein further comprise a mini-VPR transcriptional activator is coupled to the array and being encoded by a nucleic acid having at least 85%, at least 90%, or at least 95% sequence identity to one or more of the nucleic acids of SEQ ID NO 146 or 147. In several embodiments, the mini-VPR transcriptional activator is encoded by SEQ ID NO 146 or 147. In several embodiments, the ZFP array and mini-VPR transcriptional activator are encoded by a nucleic acid having at least 85%, at least 90%, or at least 95% sequence identity to one or more of the nucleic acids of SEQ ID NO 150-195. In several embodiments, the ZFP array and mini-VPR transcriptional activator is encoded one or more of the nucleic acids of SEQ ID NO 150-195.

[0024] In several embodiments, the zinc finger arrays provided for herein further comprise backbone structures. In several embodiments, ZFP array and mini-VPR transcriptional activator with backbone structure are encoded by a nucleic acid having at least 85%. at least 90%, or at least 95% sequence identity to one or more of the nucleic acids of SEQ ID NO 196-242. In several embodiments, the ZFP array comprising backbone structures andPatent Application 97157.00116mini-VPR transcriptional activator with backbone structure are encoded by one or more of the nucleic acids of SEQ ID NO 196-242. In several embodiments, the ZFP array has an amino acid sequence having at least 85%, at least 90%, or at least 95% sequence identity to one or more of the amino acids of SEQ ID NOs 623-668. In several embodiments, the zinc finger array comprises an amino acid sequence of one or more of the amino acids of SEQ ID NOs 623-668.

[0025] In several embodiments, there is provided a synthetic drug switch comprising at least one drug interaction domain, at least one transcriptional effector domain, and a DNA binding protein. In several embodiments, the DNA binding domain comprises a zinc finger protein, wherein the zinc finger protein binds a target sequence and allows the transcriptional effector to modulate transcription of a target gene when a drug that interacts with the at least one drug interaction domain is present.

[0026] In several embodiments, the drug interaction domain interacts with tamoxifen or a metabolite thereof. In several embodiments, the tamoxifen metabolite comprises 4-OHT.

[0027] In several embodiments, the drug interaction domain comprises a hormone receptor. In several embodiments, the hormone receptor is selected from an estrogen receptor, a glucocorticoid receptor, and a progesterone receptor.

[0028] In several embodiments, the drug interaction domain comprises an estrogen receptor, wherein the estrogen receptor comprises the ERT2 estrogen receptor. In several embodiments, the ERT2 estrogen receptor is encoded by a nucleic acid having at least 85% sequence identity to SEQ ID NO: 758, 775, or 792. In several embodiments, the ERT2 estrogen receptor is encoded by the nucleic acid of SEQ ID NO: 758, 775, or 792. In several embodiments, the ERT2 estrogen receptor comprises an amino acid sequence having at least 85% sequence identity to SEQ ID NO: 767, 784, or 801. In several embodiments, the ERT2 estrogen receptor comprises an amino acid sequence of SEQ ID NO: 767, 784, or 801.

[0029] In several embodiments, the drug interaction domain further comprises a p65 protein. In several embodiments, the p65 protein is encoded by a nucleic acid having at least 85% sequence identity to SEQ ID NO: 756, 773, or 790. In several embodiments, the p65 protein is encoded by the nucleic acid of SEQ ID NO: 756, 773, or 790. In several embodiments, the p65 protein comprises an amino acid sequence having at least 85% sequence identity7to SEQ ID NO: 765, 782, or 799. In several embodiments, the p65 protein comprises an amino acid sequence of SEQ ID NO: 765, 782, or 799.

[0030] In several embodiments, the synthetic drug switch comprises an amino acid having at least 85% sequence identity to one or more of SEQ ID NO: 770, 787, and 804.Patent Application 97157.00116

[0031] In several embodiments, the drug interaction domain comprises a cereblon protein. In several embodiments, the cereblon protein is engineered to (i) not express a DDB1 subdomain or (ii) express a DDB1 subdomain that does not allow the cereblon to interact with an E3 ubiquitin ligase complex. In several embodiments, the cereblon protein is encoded by a nucleic acid having at least 85% sequence identity to SEQ ID NO: 684, 708, or 732. In several embodiments, the cereblon protein is encoded by SEQ ID NO: 684, 708, or 732. In several embodiments, the cereblon protein comprises an amino acid sequence having at least 85% sequence identity to SEQ ID NO: 697, 721, or 745. In several embodiments, the cereblon protein comprises an amino acid sequence of SEQ ID NO: 697, 721, or 745. In several embodiments, the cereblon protein is fused to the transcriptional effector domain. In several embodiments, the transcriptional effector domain comprises a mini-VPR domain and the cereblon-transcriptional effector domain fusion is encoded by a nucleic acid sequence having at least 85% sequence identity to SEQ ID NO: 693, 717, or 741. In several embodiments, the cereblon-transcriptional effector domain fusion is encoded by SEQ ID NO: 693, 717, or 741. In several embodiments, the transcriptional effector domain comprises a mini-VPR domain and the cereblon-transcriptional effector domain fusion has an amino acid sequence having at least 85% sequence identity to SEQ ID NO: 700, 724, or 748.

[0032] In several embodiments, the drug interaction domain further comprises a CRIMP domain. In several embodiments, the CRIMP domain is encoded by a nucleic acid having at least 85% sequence identity to one or more of SEQ ID NO: 688. 712, or 736. In several embodiments, the CRIMP domain is encoded by a nucleic acid of SEQ ID NO: 688, 712, or 736. In several embodiments, the CRIMP domain has an amino acid sequence having at least 85% sequence identity to one or more of SEQ ID NO: 702, 726, or 750. In several embodiments, the the CRIMP domain has an amino acid sequence of SEQ ID NO: 702, 726, or 750. In several embodiments, the CRIMP domain is fused to the ZFP. In several embodiments, the CRIMP domain-ZFP fusion has a nucleic acid sequence having at least 85% sequence identity to one or more of SEQ ID NO: 694, 718, or 742. In several embodiments, the CRIMP domain-ZFP fusion has a nucleic acid sequence of SEQ ID NO: 694, 718, or 742. In several embodiments, the CRIMP domain-ZFP fusion has an amino acid sequence having at least 85% sequence identity’ to one or more of SEQ ID NO: 705, 729, or 753. In several embodiments, the CRIMP domain-ZFP fusion has an amino acid sequence of SEQ ID NO: 705, 729, or 753.

[0033] In several embodiments, the drug switch is configured to activate transcription in response to an immunomodulatory (IMiD) drug. In several embodiments, the IMiD drug isPatent Application 97157.00116selected from lenalidomide, pomalidomide, thalidomide, or combinations thereof.

[0034] In several embodiments, the nucleotide sequences provided for herein comprise one or more stop codons (e.g., nucleic acid sequence = TAA) selected from the group consisting of SEQ ID NO. 691, 715, 739, 755, 772, and 789. In several embodiments, the amino acid sequences provided for herein comprise one or more linker sequences. In several embodiments, the linker sequence (amino acid sequence = TCR) has the amino acid sequence of SEQ ID NO: 764, 781. or 798.Brief Description of the Drawings

[0035] Figure 1 schematically depicts a non-limiting example of a target DNA sequence (1) and a corresponding zinc finger protein (2). Such non-limiting embodiments provided for herein can be used to modulate transcription of genes of interest and production of the corresponding protein encoded by the gene(s) of interest.

[0036] Figures 2A-2D provide a summary of a non-limiting example of a process flow for generation of synthetic enhancer sequences that are orthogonal (e.g., not present within) to the genome of the target (e.g., host) species. Figure 2A represents the first step in the process in which a plurality7of candidate core synthetic enhancer sequences is generated with a plurality of GNN repeats (six repeats in this non-limiting example) an additional randomized nucleotide. Figure 2B represents a first screening step in which those candidate core synthetic enhancers that either have 100% identity to one or more sites in the target genome (e.g.. human and murine genomes in this non-limiting example) are removed from candidacy, as well as those having 1 or 2 mismatches to one or more sites in the target genome. Those candidates passing through this first screening step are then processed through the second screening step of Figure 2C. Figure 2C represents a second screening step in which any of the remaining candidate core synthetic enhancer elements that contain one or more vertebrate transcription factor binding sites are removed from the pool. Figure 2D represents the final step in which synthetic enhancers are generated that comprise a plurality7of surviving candidate core sequences are assembled. In this non-limiting example, four repeats of the candidate core synthetic enhancer elements are assembled in a repeat fashion, separated by710 base pair spacers. The spacers are computationally screened such that (i) the individual spacers do not contain any binding sites for any known vertebrate transcription factors and (ii) when put adjacent to the GNN repeats of the candidate core sequences, the resulting sequence does not create any transcription factor binding sites at the junction core sequence and the spacer.

[0037] Figures 3A-3B schematically show the results of the generation and screeningPatent Application 97157.00116process to identify synthetic genome orthogonal enhancer elements. Figure 3A shows a bar chart of data related to the number of synthetic enhancer elements and the number of mismatches to a target genome (here the human genome as a non-limiting example). The pool of candidate sequences that have 100% sequence identity to one or more location of the human genome was nearly 108sequences (total). After removing those sequences having 1 mismatch (~0.6xl08total) and 2 mismatches (just over 106sequences total) to the human genome, there remained twenty-five 19-mers with 3 or more mismatches. Of that pool, two of the sequences contained at least putatively weak transcription factor binding sites (screened against the HOCOMOCOvll full collection), which were removed. The final pool of candidate core sequences was 23 19-mers. Figure 3B shows an alignment table of the final pool of candidate core 19-mers organized by sequence identity, along with non-limiting examples of sequences with 90% sequence identity, 80% sequence identity, and 70% sequence identity to the consensus sequence.

[0038] Figures 4A-4C depict anon-limiting embodiment of a process flow of designing and testing ZFPs that target genome orthogonal synthetic enhancer elements.

[0039] Figure 4A shows a schematic of a flow diagram that involves (i) design ZFPs targeted to each GNN triplet of the candidate core synthetic enhancer elements (e.g., the target 19mer); (ii) linkage of the ZFPs of (i) into a 6 factor array with 3x2 arrangement (see Figures 4A / 4B); and (iii) creation of synthetic transcription factors (recognized by the ZFPs) and subsequent screening for ZFP / enhancer function and specificity.

[0040] Figure 4B shows a schematic of an experimental setup fortesting for zinc finger transcription factor strength (e.g., degree of binding to target enhancer). Two separate constructs are generated. The first comprises a promoter sequence coupled to a ZFP binding domain (a 6x array of ZPF repeats) and further coupled to a transcriptional activator. The second comprises a synthetic enhancer element (e.g., an enhancer for which a ZFP is designed to bind), a YBTATA minimal promoter and GFP (a reporting element). These two constructs are separately introduced into a cell and upon binding of the ZFP to its cognate synthetic enhancer, transcription of the GFP reporter gene is initiated, allowing for detection of GFP.

[0041] Figure 4C shows a schematic of an experimental setup for testing for zinc finger transcription factor specificity (e.g., lack of binding to a non-target enhancer). Two separate constructs are generated. The first comprises a promoter sequence coupled to a ZFP binding domain (a 6x array of ZFP repeats) and further coupled to a transcriptional activator (as in Figure 4B). The second comprises an off-target synthetic enhancer element (e.g., an enhancer for which a ZFP is not designed to bind), a YBTATA minimal promoter and GFP (a reportingPatent Application 97157.00116element). These two constructs are separately introduced into a cell and because there should be limited to no binding of the ZFP to the non-target synthetic enhancer, transcription of the GFP reporter gene would not be initiated, resulting in a failure of detection of GFP.

[0042] Figures 5A-5C depict data related to the strength and specificity of ZFPs as provided for herein. Figure 5 A shows a scatterplot of specificity' versus strength (of binding to the target synthetic enhancer). Figure 5B shows a second scatterplot of specificity’ vs. strength. Figure 5C shows a plot of a cytometric analysis of cells for GFP expression from the second analysis.

[0043] Figures 6A-6C shoyv an experimental setup and data resulting from a bulk RNA-seq analy sis to examine transcriptome impacts of three non-limiting examples of ZFP arraybased transcription factors described herein. Figure 6A shows a schematic of the various combinations of ZFP and enhancer combinations tested (in triplicate). Figure 6B shows a bar chart of data related to RNA-seq based strength of ZF transcription factors. Figure 6C shoyvs a bar chart of data related to RNA-seq based specificity of ZF transcription factors.

[0044] Figures 7A-7C show differential gene expression data from an additional bulk RNA-seq analysis to examine transcriptome impacts of non-limiting examples of ZFP arraybased transcription factors described herein against Jurkat cells. Figure 7A shoyvs a box plot depicting the number of RNA-seq reads associated with the ZF transcription factor (ZF linked to miniVPR) construct (denoted as x474, yvith X being representative of GF000 - in other words x474 is the GF000474 ZF construct), a known literature control ZFTF (x515). or a negative control (x717; containing mCherry linked to miniVPR (no ZF). Figure 7B shoyvs a bar chart depicting ZF strength in terms of the number of GFP RNA counts measured per ZFTF RNA count (e.g., for each ZFTF RNA counted, how many GFPs are counted. Figure 7C shows a line graph depicting (on the X-axis) the log 2-fold change in expression of each gene observed in the RNA seq analysis relative to the no ZF control (x717) and, on the Y axis, the number of genes displaying that degree of fold change.

[0045] Figures 8A-8C show corresponding data as that of 7A-7C but using HepB3 cells.

[0046] Figures 9A-9C show corresponding data as that of 7A-7C but using B16F10 cells.

[0047] Figures 10A-10B schematically illustrate drug switches designed in accordance with embodiments provided for herein. Figure 10A schematically illustrates a non-limiting embodiment of a drug switch employing a modified estrogen receptor (ERT2) to modulate gene expression triggered by metabolites of tamoxifen. Figure 10B schematically illustrates a non-Patent Application 97157.00116limiting embodiment of an drug switch employing IMiD co-binders to modulate gene expression triggered by the presence of an IMiD drug.

[0048] Figures 11A-11E show various data related to the non-limiting ERTA2 drug switch. Figure 11 A shows a bar chart of GFP expression as a function of 4-hydroxytamioxfen (4-OHT) concentration for the indicated drug sw itch (made up of a non-limiting example of a zinc finger linked to a modified estrogen receptor and a non-limiting example of a reporter). Figure 11B shows a short description of the components of the drug switches tested. Figures 11C-1 IE each show a plot of a cytometric analysis of cells for GFP expression as a function of 4-OHT concentration.

[0049] Figures 12A-12B depict schematics of constructs for drug switches and reporters. Figure 12A shows a schematic for the structure of an ERT2 drug switch (top) and an IMiD drug switch (bottom). Figure 12B shows a schematic for the structure of an enhanced Blue Fluorescent Protein (EBFP or BFP) reporter construct (top) and a mCherry reporter construct (bottom). By way of example only IMiD switches were paired with mCherry and ERT2 switches were paired with BFP.

[0050] Figure 13 provides a brief description of non-limiting examples of drug switch constructs tested.

[0051] Figures 14A-14C show data related to expression of the respective reporters for the indicated drug switch constructs tested. Figure 14A shows data regarding the ability to induce ERT2 drug switches. Figure 14B shows data regarding the ability to induce IMID drug switches, together Figures 14A and 14B also indicate that drug switches expressed within the same cell can be induced independently from one another and with minimal cross talk. Figure 14C show s data regarding the induction of ERT2 and IMiD drug switches in tandem (e.g., both switches in one cell at the same time).Detailed Description

[0052] The following detailed description discusses non-limiting embodiments of the zinc finger proteins and corresponding engineered synthetic enhancer elements that, in several embodiments, are used to modulate gene expression of a gene of interest without disruption or modulation of other genes w ithin a host genome (i.e., genes that are not the gene of interest).Definitions

[0053] The terms described below-, or elsewhere herein, shall be understood to have their ordinary’ meaning and shall also be understood to have the meanings specifically described herein, unless otherwise specifically indicated.Patent Application 97157.00116

[0054] Conventional and well established techniques used in molecular biology, biochemistry, cell culture, recombinant DNA, and other related fields are known to those of skill in the art and are discussed, for example, in the following literature references: Sambrook et al. MOLECULAR CLONING: A LABORATORY MANUAL, Second edition, Cold Spring Harbor Laboratory Press, 1989; Ausubel et al., CURRENT PROTOCOLS IN MOLECULAR BIOLOGY, John Wiley & Sons, New York, 1987 and periodic updates; the series METHODS IN ENZYMOLOGY, Academic Press, San Diego; and METHODS IN MOLECULAR BIOLOGY, Vol. 119, ‘'Chromatin Protocols” (P. B. Becker, ed.) Humana Press, Totowa, 1999, all of which are incorporated by reference in their entireties.

[0055] The term “zinc finger protein” or “ZFP” refers to a protein having DNA binding domains that are stabilized by zinc. The individual DNA binding domains are typically referred to as '‘fingers” A ZFP has least one finger, typically two, three, four, five, six or more fingers. Each finger binds from two to four base pairs of DNA, typically three or four base pairs of DNA. A ZFP binds to a nucleic acid sequence called a target site or target segment. Each finger typically comprises an approximately 30 amino acid, zinc-chelating, DNA-binding subdomain.

[0056] A “target site” is the nucleic acid sequence recognized by a ZFP. Engineered target sites are also referred to herein as synthetic enhancer elements.

[0057] “Kd” refers to the dissociation constant for a binding molecule - in other words, the concentration of a compound (e.g., a zinc finger protein) that gives half maximal binding of the compound to its target, meaning that one-half of the compound molecules are bound to the target under the particular assay or physiologic conditions. According to some embodiments, the Kd of a ZFP used to modulate transcription of a gene is less than about 100 nM, more preferably less than about 75 nM, more preferably less than about 50 nM. most preferably less than about 25 nM.

[0058] The phrase “substantially identical,” in the context of two nucleic acids or polypeptides, refers to tw o or more sequences or subsequences that have at least 75%, at least 85%, at least 90%, 95% or higher or any integral value there between nucleotide or amino acid residue identity, when compared and aligned for maximum correspondence, as measured using a sequence comparison algorithm such as those described below for example, or by visual inspection. Preferably, the substantial identity' exists over a region of the sequences that is at least about 10, about 20, about 40-60 residues in length or any integral value therebetween. In some embodiments, the identity is similar over a longer region of 60-80 residues, about 90-100 residues. In several embodiments, the sequences are substantially identical over the full lengthPatent Application 97157.00116of the sequences being compared.

[0059] For sequence comparison, typically one sequence acts as a reference sequence, to which test sequences are compared. When using a sequence comparison algorithm, test and reference sequences are input into a computer, subsequence coordinates are designated, if necessary, and sequence algorithm program parameters are designated. The sequence comparison algorithm then calculates the percent sequence identity for the test sequence(s) relative to the reference sequence, based on the designated program parameters. Optimal alignment of sequences for comparison can be conducted by well-established computerized algorithms such as GAP, BESTFIT, FASTA, and / or TFASTA, typically conducted using the default parameters specific for each program.

[0060] Another example of an algorithm that is suitable for determining percent sequence identity and sequence similarity is the BLAST algorithm, which is described in Altschul et al., J. Mol. Biol. 215:403-410 (1990). Software for performing BLAST analyses is publicly available through the National Center for Biotechnology Information.

[0061] Two nucleic acid sequences can be determined to be substantially identical when the two molecules hybndize to each other under stringent conditions.

[0062] A polypeptide can be determined to be substantially identical to a second polypeptide, for example, where the two peptides differ only by conservative substitutions as defined herein and referring to a change in the amino acid composition of the protein that does not substantially alter the protein's activity.

[0063] A “functional fragment” or “functional equivalent” of a protein, polypeptide or nucleic acid is a protein, polypeptide or nucleic acid whose sequence is not identical to the full-length protein, polypeptide or nucleic acid, yet retains the same function as the full-length protein, polypeptide or nucleic acid.

[0064] The terms “nucleic acid,” “polynucleotide,” and “oligonucleotide” are used interchangeably and refer to a deoxyribonucleotide or ribonucleotide polymer in either single-or double-stranded form. For the purposes of the present disclosure, these terms are not to be construed as limiting with respect to the length of a polymer.

[0065] An “exogenous molecule” is a molecule that is not normally present in a cell, but can be introduced into a cell by one or more genetic, biochemical or other methods. Normal presence of a molecule in the cell may vary based on the particular developmental stage and environmental conditions of the cell. Thus, for example, a molecule that is present only during embryonic development of a selected tissue is considered an exogenous molecule with respect to an adult version of that tissue. An exogenous molecule can comprise, for example, aPatent Application 97157.00116functioning version of a malfunctioning endogenous molecule or a malfunctioning version of a normally functioning endogenous molecule.

[0066] An exogenous molecule can be a small molecule, such as those generated by a combinatorial chemistry process. An exogenous molecule can also be a macromolecule such as a protein, nucleic acid, carbohydrate, lipid, glycoprotein, lipoprotein, polysaccharide, as wells as modified derivatives of the preceding, or any complex comprising one or more of the preceding. Nucleic acids include DNA and RNA, can be single- or double-stranded; can be linear, branched or circular; and can be of any length. Proteins include, but are not limited to, DNA-binding proteins, transcription factors, chromatin remodeling factors, methylated DNA binding proteins, polymerases, methylases, demethylases, acetylases, deacetylases, kinases, phosphatases, integrases, recombinases, ligases, topoisomerases, gyrases and helicases.

[0067] An exogenous molecule can be the same type of molecule as an endogenous molecule, e.g., protein or nucleic acid (i.e., an exogenous gene), providing it has a sequence that is different from an endogenous molecule. Methods for the introduction of exogenous molecules into cells are known to those of skill in the art and include, but are not limited to, lipid-mediated transfer (i.e., liposomes, including neutral and cationic lipids), electroporation, direct injection, cell fusion, particle bombardment, calcium phosphate co-precipitation, DEAE-dextran-mediated transfer and viral vector-mediated transfer.

[0068] An “endogenous molecule’7is one that is normally present in a particular cell at a particular developmental stage under particular environmental conditions.

[0069] The phrase “adjacent to a transcription initiation site” refers to a target site that is within about 50 bases either upstream or downstream of a transcription initiation site. “Upstream” of a transcription initiation site refers to a target site that is more than about 50 bases 5' of the transcription initiation site (i.e., in the non-transcribed region of the gene). “Downstream” of a transcription initiation site refers to a target site that is more than about 50 bases 3' of the transcription initiation site.

[0070] A “fusion molecule” is a molecule in which two or more subunit molecules are linked. In several embodiments, the fusion is through a covalent linkage. The subunit molecules can be the same chemical type of molecule or can be different. For example, a fusion molecule comprises, in several embodiments, a fusion between a ZFP DNA-binding domain and a transcriptional activation domain.

[0071] A “gene” includes a DNA region encoding a gene product as well as all DNA regions which regulate the production of the gene product, whether or not suchPatent Application 97157.00116regulatory' sequences are adjacent to coding and / or transcribed sequences. Accordingly, a gene includes, but is not necessarily limited to, promoter sequences, terminators, translational regulatory sequences such as ribosome binding sites and internal ribosome entry sites, enhancers, silencers, insulators, boundary elements, replication origins, matrix attachment sites and locus control regions.

[0072] “Gene expression” refers to the conversion of the information, contained in a gene, into a gene product, which includes not only the direct transcriptional product of a gene (e.g., mRNA, tRNA, rRNA, antisense RNA, ribozyme, structural RNA or any other ty pe of RNA) but also a protein produced by translation of a mRNA. Modified RNAs, such as those subject to capping, poly adenylation, methylation, and editing, and proteins modified by, for example, methylation, acetylation, phosphorylation, ubiquitination, ADP-ribosylation, myristilation, and / or glycosylation are also gene products.

[0073] “Gene activation” refers to any process that results in an increase in production of a gene product, whether through increase in transcription of a gene and / or translation of a mRNA or disinhibition of transcription and / or translation. In several embodiments, gene activation comprises an increase in the production of a gene product of about 2-fold, about 2-to about 5-fold, between about 5- and about 10-fold, between about 10- and about 20-fold, between about 20- and about 50-fold, between about 50- and about 100-fold or any integer between those listed. In several embodiments, gene activation results in 100-fold or more increase in a gene product.

[0074] “Gene repression” and “inhibition of gene expression” refer to any process which results in a decrease in production of a gene product which includes not only the direct decrease in transcription of a gene and / or translation of a mRNA but also those which result in inhibition of formation of a transcription initiation complex or otherwise inhibit transcription. In several embodiments, gene repression comprises a decrease in the production of a gene product of about 2-fold, about 2- to about 5-fold, between about 5- and about 10-fold, between about 10- and about 20-fold, between about 20- and about 50-fold, between about 50- and about 100-fold or any integer between those listed. In several embodiments, gene repression results in 100-fold or more decrease in a gene product.

[0075] “Modulation” refers to a change in the level or magnitude of an activity or process. The change can be either an increase or a decrease. For example, modulation of gene expression includes both gene activation and gene repression.

[0076] A “regulatory domain” or “functional domain” refers to a protein or a protein domain that has transcriptional modulation activity when tethered to a DNA binding domain,Patent Application 97157.00116for example, a ZFP. Regulatory domains can be activation domains or repression domains. Activation domains include, but are not limited to, VP 16, VP64 and the p65 subunit of nuclear factor Kappa-B. Repression domains include, but are not limited to, KRAB MBD2B and v-ErbA. Additional regulator}7domains include, e.g., transcription factors and co-factors (e.g., MAD, ERD, SID, early growth response factor 1, and nuclear hormone receptors), endonucleases, integrases, recombinases, methyltransferases, histone acetyltransferases, histone deacetylases etc. Activators and repressors include co-activators and co-repressors . In several embodiments, a ZFP can act alone, without a regulatory domain, to effect transcription modulation.

[0077] The term “operably linked'’ or “operatively linked” is used with reference to a juxtaposition of two or more components (such as sequence elements), in which the components are arranged such that both components function normally and allow the possibility that at least one of the components can mediate a function that is exerted upon at least one of the other components.

[0078] With respect to fusion polypeptides, the term “operably linked” or “operatively linked” refers the context in which each of the components performs their respective function when linked as they would when not linked. For example, when ZFP DNA-binding domain is fused to a transcriptional activation domain ZFP DNA-binding domain portion is able to bind its target site and the transcriptional activation domain is able to activate transcription.

[0079] The term “recombinant,” when used with reference to a cell, indicates that the cell replicates an exogenous nucleic acid, or expresses a peptide or protein encoded by an exogenous nucleic acid. Recombinant cells can contain genes that are not found within the native (non-recombinant) form of the cell. Recombinant cells can also contain genes found in the native form of the cell wherein the genes are modified and re-introduced into the cell by artificial means. The term also includes cells that contain a nucleic acid endogenous to the cell that has been modified without removing the nucleic acid from the cell; such modifications include those obtained by gene replacement, site-specific mutation, and related techniques.

[0080] A “recombinant expression cassette,” “expression cassette” or “expression construct” is a nucleic acid construct, generated recombinantly or synthetically, that has control elements that are capable of effecting expression of a structural gene that is operatively linked to the control elements in host cells. Expression cassettes include at least promoters and optionally, transcription termination signals. Typically, the recombinant expression cassette includes at least a nucleic acid to be transcribed (e.g.. a nucleic acid encoding a desired polypeptide) and a promoter. In several embodiments, an expression cassette includesPatent Application 97157.00116nucleotide sequences that encode a signal sequence that directs secretion of an expressed protein from the host cell. Transcription termination signals, enhancers, and other nucleic acid sequences that influence gene expression, are also included in an expression cassette, according to several embodiments.

[0081] A “promoter’' is defined as an array of nucleic acid control sequences that direct transcription. As used herein, a promoter typically includes necessary nucleic acid sequences near the start site of transcription, such as, a TATA element, CCAAT box, and / or an SP-1 site, etc. Promoters may also include distal enhancer or repressor elements, which can be located at a distinct site from the start site of transcription.

[0082] A “constitutive'’ promoter is a promoter that is active under most environmental and developmental conditions. An “inducible” promoter is a promoter that is active under certain environmental or developmental conditions.

[0083] An “expression vector” is a nucleic acid construct, generated recombinantly or synthetically, with a series of specified nucleic acid elements that permit transcription of a particular nucleic acid in a host cell. In several embodiments, the expression vector is configured for integration or replication of the expression vector in a host cell. The expression vector can be part of a plasmid, virus, or nucleic acid fragment, of viral or non-viral origin.

[0084] The term “host cell” refers to a cell that contains an expression vector or nucleic acid, either of which optionally encodes a ZFP or a ZFP fusion protein. Host cells can be prokaryotic cells or eukaryotic (e.g., mammalian) cells. Host cells can refer to in vitro or in vivo settings.

[0085] The term “naturally occurring,” as applied to an object, means that the object can be found in nature, as distinct from being artificially produced by humans.

[0086] The terms “polypeptide” and “protein” are used interchangeably to refer to a polymer of amino acid residues and are not limited to a minimum length. Polypeptides, including the provided receptors and other polypeptides, e.g., linkers or peptides, may include amino acid residues including natural and / or non-natural amino acid residues. The terms also include post-expression modifications of the polypeptide, for example, glycosylation, sialylation, acetylation, and phosphorylation. In some aspects, the polypeptides may contain modifications with respect to a native or natural sequence, as long as the protein maintains the desired activity. These modifications may be deliberate, as through site-directed mutagenesis, or may be accidental, such as through mutations of hosts which produce the proteins or errors due to PCR amplification.

[0087] As used herein, a “subject” is a mammal, such as a human or other animal, andPatent Application 97157.00116typically is human. In some embodiments, the subject, e.g., patient, to whom the agent or agents, cells, cell populations, or compositions are administered, is a mammal, typically a primate, such as a human. In some embodiments, the primate is a monkey or an ape. The subject can be male or female and can be any suitable age, including infant, juvenile, adolescent, adult, and geriatric subjects. In some embodiments, the subject is anon-primate mammal, such as a rodent.

[0088] As used herein, "treatment” (and grammatical variations thereof such as “treat” or “treating”) refers to complete or partial amelioration or reduction of a disease or condition or disorder, or a symptom, adverse effect or outcome, or phenotype associated therewith. Desirable effects of treatment include, but are not limited to, alleviation of symptoms, diminishment of any direct or indirect pathological consequences of the disease, preventing metastasis, decreasing the rate of disease progression, amelioration or palliation of the disease state, and remission or improved prognosis. The terms do not imply complete curing of a disease or complete elimination of any symptom or effect(s) on all symptoms or outcomes.

[0089] As used herein, “preventing” (and grammatical variations thereof such as “prevent” or “prevention”) as used herein, includes providing prophylaxis with respect to the occurrence or recurrence of a disease in a subject that may be predisposed to the disease but has not yet been diagnosed with the disease. In some embodiments, the provided cells and compositions are used to delay development of a disease or to slow the progression of a disease.

[0090] A “therapeutically effective amount” of an agent, e.g., a pharmaceutical formulation or cells, refers to an amount effective, at dosages and for periods of time necessary, to achieve a desired therapeutic result, such as for treatment of a disease, condition, or disorder, and / or pharmacokinetic or pharmacodynamic effect of the treatment. The therapeutically effective amount may vary according to factors such as the disease state, age, sex, and weight of the subject, and the populations of cells administered. In some embodiments, the provided methods involve administering the cells and / or compositions at effective amounts, e.g., therapeutically effective amounts.

[0091] As used herein, the singular forms “a,” “an,” and “the” include plural referents unless the context clearly dictates otherwise. For example, “a” or “an” means “at least one” or “one or more.

[0092] Throughout this disclosure, various aspects of the claimed subject matter are presented in a range format. It should be understood that the description in range format is merely for convenience and brevity and should not be construed as an inflexible limitation on the scope of the claimed subject matter. Accordingly, the description of a range should bePatent Application 97157.00116considered to have specifically disclosed all the possible sub-ranges as well as individual numerical values within that range. Language such as "up to,” "at least,” "greater than,” “less than,” “between,” and the like includes the number recited. Numbers preceded by a term such as “about” or “approximately” include the recited numbers. For example, where a range of values is provided, it is understood that each intervening value, between the upper and lower limit of that range and any other stated or intervening value in that stated range is encompassed within the claimed subject matter. The upper and lower limits of these smaller ranges may independently be included in the smaller ranges, and are also encompassed within the claimed subject matter, subject to any specifically excluded limit in the stated range. Where the stated range includes one or both of the limits, ranges excluding either or both of those included limits are also included in the claimed subject matter. This applies regardless of the breadth of the range.

[0093] The term “about” as used herein refers to the usual error range for the corresponding value that is readily known. Reference herein to “about” a value or parameter includes (and describes) embodiments that relate to that value or parameter per se. For example, a description referring to “about X” includes a description of “X”. In certain embodiments, “about X” refers to a value of ± 25%, ± 10%, ± 5%, ± 2%, ± 1%, ± 0.1% or ± 0.01% of X.

[0094] In addition, when a sequence is disclosed as “comprising” a nucleotide or amino acid sequence, such a reference shall also include, unless otherwise indicated, that the sequence “comprises”, “consists of’ or “consists essentially of’ the recited sequence.

[0095] As used herein, a statement that a cell or population of cells is “positive” for a particular marker or target gene / protein of interest refers to the detectable presence of the particular marker (typically a surface marker) on or in the cell. When referring to a surface marker, the term refers to the presence of surface expression as detected by flow cytometry, for example by staining with an antibody that specifically binds to the marker and detecting the antibody, wherein the staining is detectable by flow cytometry at a level that is substantially higher than the staining detected by the same procedure with an isotype matched control under otherwise identical conditions, and / or that is substantially similar to the level of cells known to be positive for the marker, and / or that is substantially higher than the level of cells known to be negative for the marker.

[0096] As used herein, a statement that a cell or cell population is “negative” for a particular marker or target gene / protein of interest means that the particular marker (typically a surface marker) is not present on or in the cell in a substantially detectable presence. When referring to a surface marker, the term refers to the absence of surface expression as detectedPatent Application 97157.00116by flow cytometry, for example by staining with an antibody that specifically binds to the marker and detecting the antibody, wherein the staining is not detected by flow cytometry at a level that is substantially higher than the staining detected by the same procedure with an isotype matched control under otherwise identical conditions, and / or at a level that is substantially lower than the level of cells known to be positive for the marker, and / or at a level that is substantially similar compared to the level of cells known to be negative for the marker.

[0097] As used herein, "percent” (%) amino acid sequence identity” and "percent identity,” when used in reference to an amino acid sequence (a reference polypeptide sequence), is defined as the percentage of amino acid residues in a candidate sequence (e.g., a subject antibody or fragment) that are identical to the amino acid residues in the reference polypeptide sequence, after aligning the sequences and introducing gaps, if necessary, to achieve the maximum percent sequence identity, and without considering any conservative substitutions as part of the sequence identity. Alignment for the purpose of determining percent amino acid sequence identity can be accomplished in a variety of known ways, for example, using publicly available computer software. Appropriate parameters for aligning the sequences can be determined, including any algorithms necessary to achieve maximum alignment over the full length of the sequences being compared.

[0098] Amino acid substitutions can include the substitution of one amino acid for another in a polypeptide. Substitutions may be conservative amino acid substitutions or nonconservative amino acid substitutions. Amino acid substitutions may be introduced into the binding molecule of interest (e g., an antibody) and the product screened for the desired activity (e.g., specific binding to a target sequence orthogonal to a host genome).

[0099] Amino acids can be generally grouped according to the following common side chain properties.

[0100] In some embodiments, conservative substitutions may include exchanging a member of one of these classes for another member of the same class. In some embodiments, anon-conservative amino acid substitution may involve exchanging a member of one of these classes for another class.

[0101] (1) hydrophobicity: Norleucine, Met, Ala, Vai, Leu, He;

[0102] (2) neutral hydrophilicity: Cys, Ser, Thr, Asn, Gin;

[0103] (3) acidity: Asp and Glu;

[0104] (4) alkalinity: His, Lys, Arg;

[0105] (5) residues that influence chain orientation: Gly, Pro; and

[0106] (6) aromatic: Trp, Tyr, Phe.Patent Application 97157.00116

[0107] As used herein, a composition refers to any mixture of two or more products, substances, or compounds, including cells. It may be a solution, a suspension, liquid, powder, a paste, aqueous, non-aqueous or any combination thereof.Synthetic Enhancer Elements and Assemblies

[0108] According to several embodiments, there are provided herewith engineered target DNA domains, also referred to as synthetic enhancer elements. These synthetic enhancer elements are designed, in several embodiments, to be orthogonal or substantially orthogonal (e.g., unique) in comparison to a host genome, such as the human genome, murine genome or others. The synthetic enhancer elements can be used as a target for transcription factors (via DNA binding proteins that recognize the synthetic enhancer elements). For example, according to several embodiments, ZFP-based transcription factors are used. However, other types of transcription factors are also used, as disclosed herein, in some embodiments. Transcription factors may any of the natural DNA-binding transcription factors, such as helix-tum-helix transcription factors, leucine zippers, helix-loop-helix transcription factors, homeodomain proteins, or any combination thereof. In several embodiments, engineered transcription factors, such as those employing enzymatically dead Cas proteins (dCas), such as Cas 9 (dCas9) or Cas 12 (dCasl2) are used. Transcription factors may be activating or repressing in their function, depending on the embodiment. TALE (Transcription Activator-Like Effect or)-based TFs, which recognize single bases and can be fused to functional domains, are used as the transcription factor in some embodiments. Thus, the synthetic enhancer elements provided for herein can function as a target for any machinery7that modulates transcription (at the gene or epigenetic level) in order to modulate transcription of a gene of interest.Synthetic Enhancer Characteristics and Core Sequences

[0109] According to several embodiments, synthetic enhancer elements are configured to serve as a target for a DNA binding protein, such as a zinc finger protein (ZFP). Based on the known structural characteristics and binding elements of DNA binding proteins, synthetic enhancer elements are engineered to have a length betw een about 10 and about 200 nucleotides. In several embodiments, synthetic enhancer elements are between: about 10 and about 20 nucleotides in length, about 20 and about 30 nucleotides in length, about 30 and about 40 nucleotides in length, about 40 and about 50 nucleotides in length, about 50 and about 60 nucleotides in length, about 60 and about 70 nucleotides in length, about 70 and about 80 nucleotides in length, about 80 and about 90 nucleotides in length, about 90 and about 100Patent Application 97157.00116nucleotides in length, about 100 and about 110 nucleotides in length (including 100, 101, 102, 103, 104. 105, 106, 107. 108, 109 and 110 nucleotides), about 110 and about 120 nucleotides in length, about 120 and about 130 nucleotides in length, about 130 and about 140 nucleotides in length, about 140 and about 150 nucleotides in length, about 150 and about 160 nucleotides in length, about 160 and about 170 nucleotides in length, about 170 and about 180 nucleotides in length, about 180 and about 190 nucleotides in length, about 190 and about 200 nucleotides in length, including any length between those ranges listed.

[0110] In several embodiments, the synthetic enhancer element comprises a plurality of repeated elements, also referred to as enhancer core sequences, each core sequence comprising one or more GX1X2 motifs. Depending on the embodiment, the core enhancer sequences range in length between about 10 and about 60 nucleotides. In several embodiments, core enhancer sequences are between: about 10 and about 20 nucleotides in length (including 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, and 20 nucleotides), about 20 and about 30 nucleotides in length, about 30 and about 40 nucleotides in length, about 40 and about 50 nucleotides in length, about 50 and about 60 nucleotides in length, including any length between those ranges listed.

[0111] According to several embodiments, in order to provide precision with respect to the engagement of a DNA binding protein, such as a ZFP, the core enhancer sequences (as well as the spacers, which are discussed below) are designed such that the sequences do not share 100% sequence identity with any known sequence in the host (e.g.. within the human genome). Thus, as discussed below, any candidate core enhancer sequences that are an exact match to a sequence within the host genome are eliminated from the candidate pool. Additionally, according to several embodiments, those core enhancer sequences that have 1 mismatch or 2 mismatches with a sequence of the host genome are also eliminated from the candidate pool. This screening increases the fidelity of the synthetic enhancer element-DNA binding protein interaction and reduces the chances of off-target effects and inadvertent or otherwise undesired modulation of host genes.

[0112] In addition to removing exact and 1 or 2 base pair mismatched sequences, in several embodiments, core enhancer sequences are screened for actual and / or putative binding sites for transcription factors. Known transcription factor binding sites, as well as putative or predicted sites within the core enhancer sequences could result in unintended transcription of a gene of interest (e.g., the gene driven by a promoter operatively linked to the synthetic enhancer element). The removal of core enhancer sequences that have one or more improves the control of modulated gene expression by reducing the probability that an endogenousPatent Application 97157.00116transcription factor initiates or silences transcription of the gene of interest (optionally an exogenous gene).

[0113] In several embodiments, the core enhancer sequences are assembled as a series of concatemerized repeat units. Depending on the embodiment, the number of core enhancer sequence repeats can vary. In several embodiments, a longer set of repeats provides a higher probability of a sequence that has no corresponding sequence in a host genome (such as the human genome). In several embodiments, the core enhancer sequences comprise 2, 3, 4, 5, 6, 7, 8, 9, or 10 repeats of the GX1X2 motif. In several embodiments, the core enhancer sequence comprises 6 GX1X2 motifs, resulting in an 18-mer. According to several embodiments, an additional nucleotide (any nucleotide selected from adenine, guanine, cytosine, and thymine) is added to the 3’ end of the core enhancer sequence, yielding a 19-mer. In several embodiments, the core enhancer sequences are configured in a repeated fashion (discussed in more detail below) to provide a plurality of targets for a DNA binding protein to bind to, thereby increasing the likelihood that a DNA binding protein will bind the synthetic enhancer element by providing multiple target sites.Synthetic Enhancer Spacers

[0114] According to several embodiments, the core enhancer sequences are intercalated with one or more spacer sequences. The core enhancer-spacer configuration can be represented as follows: CS1-SS1-CS2-SS2-CS3-SS3-CS4 [...], where CS= core sequence and SS = spacer sequence. The number of core sequences varies, depending on the embodiment, with the number of spacer sequences being one less than the number of core sequences. For example, if there are 6 core sequences then there are 5 spacer sequences, with a core sequence in each of the 5’ and 3’ regions, though not necessarily at the terminus of the oligonucleotide.

[0115] As with the core sequences, the length of the spacer sequence vanes, depending on the embodiment. In several embodiments, the spacer sequence ranges from about 5 to about 25 nucleotides, including about 5, about 10, about 15, about 20, about 25, or any length between those listed. In several embodiments, the spacer sequence ranges from 8-12 nucleotides. In several embodiments, the spacer sequence comprises 10 nucleotides. In several embodiments, the spacer sequence is 10 nucleotides.

[0116] Similar to the core sequences, candidate spacer sequences are also screened. Each core sequence is screened to filter out any candidate sequences that comprise an actual orPatent Application 97157.00116putative transcription factor binding site. A secondary screening is also performed to identify (and remove) any spacer sequences that, when assembled with an upstream or downstream core sequence, would create an actual or putative transcription factor binding site across the junction of the core sequence and the spacer sequence. Such filtering, in conjunction with that performed for the core sequences, yields a pool of component nucleotides that are assembled into a synthetic enhancer assembly that does not include any sequence matching (or having 1 or 2 base pair mismatches) the host genome and does not include any actual or putative transcription factor binding sites.Synthetic Enhancer Assemblies

[0117] According to several embodiments, component parts of a synthetic enhancer (i.e., core sequences and spacer sequences) are assembled to generate a synthetic enhancer assembly. In several embodiments, the multiple core sequences and multiple spacer sequences are joined to create a synthetic enhancer assembly. In several embodiments, a synthetic enhancer assembly comprises two core sequences, three core sequences, five core sequences, six core sequences, seven core sequences, eight core sequences, nine core sequences, ten core sequences, or more. In several embodiments, a synthetic enhancer assembly comprises one spacer sequence, two spacer sequences, three spacer sequences, five spacer sequences, six spacer sequences, seven spacer sequences, eight spacer sequences, nine spacer sequences, or more. In several embodiments, a synthetic enhancer assembly comprises four core sequences and three spacer sequences. In additional embodiments, an additional spacer sequence is included and separates the final core sequence from a transcriptional effector molecule, such as a YBTATA minimal promoter (SEQ ID NO. 49). These synthetic enhancer assemblies have a configuration can be represented as follows: CS1-SS1-CS2-SS2-CS3-SS3-CS4-SS4-Promoter [...], where CS= core sequence and SS = spacer sequence. Thus, according to several embodiments, a series of four core sequence (19-mer, 6 GNN repeats and 1 additional nucleotide) are separated by three spacer sequences (10 nucleotides), followed by an additional spacer sequence and a promoter element, such as the YBTATA minimal promoter. According to several embodiments, each of the core sequences differ from one another. In some embodiments, the core sequences are repeats (e.g., 1 core sequence repeated multiple times, or multiple core sequences used one or more times). According to several embodiments, each of the spacer sequences differ from one another. In some embodiments, the spacer sequences are repeats (e g., 1 core sequence repeated multiple times, or multiple core sequences used one or more times). Each of the GX1X2 motifs within a core sequence are targeted by a DNA bindingPatent Application 97157.00116protein, such as a zinc finger. Thus, in such instances, a synthetic enhancer assembly comprising four core sequences provides up to twenty-four (24) target binding sites for a DNA binding protein, such as a ZFP, thereby not only providing a target sequence long enough to be unique with respect to the host genome, but also providing multiple target regions to enhance the probability of successful binding of the DNA binding protein (and related modulation of transcription).DNA Binding Proteins

[0118] According to embodiments disclosed herein, DNA binding proteins are used to bind target sequences, e.g., synthetic enhancer assemblies, which are unique to a host genome and allow modulation of transcription of a gene of interest in an orthogonal manner - in other words, without concurrent modification of endogenous gene expression. DNA binding proteins are divided into classes based on their function: (1) transcription factors, which are involved in transcriptional regulation: (2) DNA replication factors:, which perform DNA synthesis (whole genome or DNA fragments); (3) repair factors, which have a role in removing single base pairs or specific oligonucleotides and filling the gaps with suitable nucleotides, and (4) histones, which are involved in transcription and chromosome packaging in the cell nucleus. A variety7of DNA binding proteins are available and suitable for use in embodiments provided for herein. For example, zinc finger proteins (ZFPs) are small protein motifs with finger-like protrusions that bind to specific DNA sequences, RNA sequences, lipids, and other proteins. Helix-tum-helix motifs are proteins with two alpha helices joined by a short amino acid spacer, with one helix functioning to bind the major groove of DNA and the other functioning to stabilize the DNA-protein complex. Leucine zippers are DNA binding motifs of -60-80 amino acids structured in a coiled-coil pattern comprising two alpha helices with a repeating patters of leucine residues every seventh amino acid. Single-stranded DNA binding proteins are proteins that envelope or otherwise wrap single-stranded DNA (ssDNA) to protect it from degradation. Histones are a family of small, positively charged proteins that bind tightly to DNA based on the negative charge of DNA. HU is a histone-like protein that binds specifically to DNA, particularly preferentially binding to certain damaged or distorted DNAs. Transcription activator-like effector nucleases (TALENs) are a type of restriction enzyme that can be engineered to bind to specific DNA sequences. TALENs comprise a DNA-binding domain from a transcription activator-like effector (TALE) protein which is coupled to a DNA cleavage domain, or nuclease. Thus, in some embodiments, a TALE can be used (absent the cleavage domain) to target DNA. Likewise, CRISPR guide RNA molecules (e.g., optionallyPatent Application 97157.00116coupled to, for example a dead Cas protein) can be used as DNA targeting / binding motifs.

[0119] In several embodiments, ZFPs are employed as the DNA binding domain. In several embodiments, a lurality of ZFPs is engineered to bind with a high degree of strength and specificity to synthetic enhancers (e.g., assemblies) as provided for herein. The ZFPs can be organized in an array of multiple ZFPs (separated by linkers) with the ZFPs either being unique within an array, or repeated (e.g., one or more ZFPs repeated throughout an array). In several embodiments, an induvial ZFP is engineered to mimic the structure of a corresponding synthetic enhancer element, in particular the core sequence. In such embodiments, the ZFP comprises a number of individual fingers that matches the number of GX1X2 motifs within a core sequence. Thus, in several embodiments, a ZFP array comprises 2, 3, 4, 5, 6, 7, 8, 9, or 10 ZFPs. In several embodiments, the ZFP array comprises 6 ZFPs, one for each of 6 GX1X2 motifs within a core sequence of a synthetic enhancer assembly. In several embodiments, the individual ZFPs are unique. In several embodiments, the ZFPs are separated by spacers. In one embodiment, the ZFP array is a 3 by 2 configuration, in which ZFP 1 and 2 are linked by a first linker, ZFP 3 and 4 are linked by a second linker, ZFP 5 and 6 are linked by a third linker, and two additional linkers are used to couple each of the pairs. Each of Linkers 1, 2, and 3 are the same length and each of Linkers 4 and 5 are the same length, leading to the following configuration: ZFP1-L1-ZFP2-L4-ZFP3-L2-ZFP4-L5-ZFP5-L3-ZFP6. In several embodiments, the linkers comprise a peptide of between about 3 and about 10 amino acids, including 3, 4, 5, 6, 7, 8, 9. 10. or more amino acids. In several embodiments, the linkers are either 5 or 6 amino acids. In several embodiments, the linkers have at least 85% sequence identity to one or more of the amino acids of SEQ ID NOs 342-346. In several embodiments, the linkers have at least 90% sequence identity to one or more of the amino acids of SEQ ID NOs 342-346. In several embodiments, the linkers have at least 95% sequence identity to one or more of the amino acids of SEQ ID NOs 342-346. In several embodiments, the linkers are selected from one or more of the amino acids of SEQ ID NOs 342-346.

[0120] In several embodiments the ZFPs range from about 20 to about 35 amino acids, including 20, 21, 22, 23, 24, 25, 26, 27, 28, 29. 30, 31, 32, 33, 34, or 35 amino acids. In several embodiments, the individual ZFPs within an array vary in size from one another (though some may have the same length). In several embodiments, the individual ZFPs within an array vary in sequence from one another. In several embodiments, the individual ZFPs within an array vary in size and sequence. In several embodiments, the individual ZFPs within an array have similar or matching sizes but vary in sequence. In several embodiments, the individual ZFPs within an array have at least 85% sequence identity to one or more of the amino acids of SEQPatent Application 97157.00116ID NOs 347-622. In several embodiments, the individual ZFPs within an array have at least 90% sequence identity to one or more of the amino acids of SEQ ID NOs 347-622. In several embodiments, the individual ZFPs within an array have at least 95% sequence identity to one or more of the amino acids of SEQ ID NOs 347-622. In several embodiments, the individual ZFPs within an array are selected from one or more of the amino acids of SEQ ID NOs 347-622. In several embodiments, the ZFP array comprises an amino acid sequence having at least 85% sequence identity to one or more of the amino acids of SEQ ID NOs 244-289. In several embodiments, the ZFP array comprises an amino acid sequence having at least 90% sequence identity to one or more of the amino acids of SEQ ID NOs 244-289. In several embodiments, the ZFP array comprises an amino acid sequence having at least 95% sequence identity to one or more of the amino acids of SEQ ID NOs 244-289. In several embodiments, the ZFP array comprises an amino acid sequence of one or more of SEQ ID NOs 244-289.Transcriptional Effectors

[0121] In several embodiments, one ore transcriptional effectors are employed in a complex utilized to modulate transcription of a gene of interest (e.g., an exogenously introduced gene (noting that the exogenous gene may also exist as an endogenous gene). In several embodiments, the transcriptional effector comprises a transcriptional activator. Depending on the embodiment, a variety of individual activator proteins are used and include, but are not limited to the catabolite activator protein (CAP), glucocorticoid receptor. CREB (cAMP response element binding protein), MyoD, PPARs (peroxisome proliferator-activated receptors), VDR (vitamin D receptor), AR (androgen receptor) a VPR (VP64-p65-Rta tripartite activator) transcriptional activator, and a mini -VPR transcriptional activator (a truncated VP64-p65-Rta tripartite activator). In several embodiments, a mini-VPR transcriptional activator is used. In several embodiments, the mini-VPR transcriptional activator is complexed with a ZFP array and a promoter (see e.g., Figure 5A).

[0122] In several embodiments, the transcriptional effector comprises a promoter. Eukaryotic promoters are used in several embodiments. In some embodiments, however, prokaryotic promoters may be used. In several embodiments, constitutive promoters are used. In several embodiments, inducible promoters are used. A variety of promoters are readily compatible with the modulation systems provided for herein, including but not limited to the human cytomegalovirus (CMV) promoter, the human elongation factor 1 (EFl) promoter, a truncated human EFl promoter (EFS). the simian virus 40 promoter (SV40), the spleen focusforming virus (SFFV) promoter, the mammalian phosphoglycerate kinase (PKG1) promoter,Patent Application 97157.00116the mammalian ubiquitin C (UBC) promoter, the human beta actin promoter, the mammalian CAG promoter, the mammalian tetracycline response element (TRE) promoter, the yeast UAS promoter, the drosophila Actin 5c (Ac5) promoter, the baculovirus polyhedrin promoter, the mammalian Ca2+ / calmodulin-dependent protein kinase II promoter, the yeast Gal 1, 10 promoter, the yeast transcription elongation factor promoter, the yeast glyceraldehyde 3-phosphage dehydrogenase promoter, the yeast alcohol dehydrogenase I promoter, the Cauliflower Mosaic Virus (CaMV35S) promoter, the human polymerase III RNA promoter (Hl), the human U6 small nuclear promoter, the YBTATA promoter, the bacteriophage T7 promoter, the bacteriophage T7 lac promoter, the bacteriophage SP6 promoter, the arabinose metabolic operon (araBAD) promoter, the E. coli tryptophan operon promoter, the lac operon promoter, the Ptac hybrid lac and trp operon promoter, the bacteriophage lambda promoter, the T3 bacteriophage promoter, among others. In several embodiments, the YBTATA minimal promoter is used.Methods of Transcriptional Modulation

[0123] Provided for herein are methods of modulation of transcription of a gene of interest, as well as related components and systems to accomplish such modulation. As discussed herein, in several embodiments, the gene of interest comprises an exogenous gene, though it may be a copy or highly similar to an endogenous gene. The gene of interest may encode a therapeutic protein, such as a chimeric antigen receptor, a cytokine, or any other protein that has a desired effect in the host cell. In several embodiments, modulation of transcription of the gene of interest allows, for example, overproduction of an endogenous protein that is not expressed at sufficient levels for the desired physiologic effect. In several embodiments, the systems provided for herein can be used to modulate expression of a gene in a controlled manner through use of a stimulus, for example drug-controlled on- or off-switches.

[0124] Akin to Figure 5A, a pair of complexes are generated to accomplish modulation of gene expression. As shown in Figure 5A, a promoter is coupled to a ZFP array (tailored to target a corresponding synthetic enhancer structure) and a transcriptional activator. An additional construct comprising a synthetic enhancer assembly as provided for herein (e.g., 4 repeats of a 6 repeat GX1X2N motil), a promoter (such as a YBTATA minimal promoter, and a gene of interest (Figure 5A shows GFP as an example that was used for reporter purposes) is generated.

[0125] These constructs can be introduced into a host cell (e.g., a human cell) independently or concurrently. In several embodiments, the constructs are introducedPatent Application 97157.00116independently, with the synthetic enhancer containing construct being introduced first (to allow time for integration into the host genome). The constructs can be introduced for example, by viral deliver}- through use of adenovirus, adeno-associated virus, lentivirus, among others. In several embodiments, the constructs are delivered using lentivirus. In several embodiments, random insertion sites within the genome are employed. In several embodiments, known insertions sites are used, such as the AAVS1 site, located on human chromosome 19. In several embodiments, upon identification of a unique site in the host genome, ZFP arrays can optionally be designed to target this site and a synthetic enhancer assembly is not required.

[0126] In several embodiments, the synthetic enhancer assembly (not including a gene of interest) has at least 85% sequence identity to one or more of the nucleic acid sequences of SEQ ID NOs 50-73. In several embodiments, the synthetic enhancer assembly (not including a gene of interest) has at least 90% sequence identity to one or more of the nucleic acid sequences of SEQ ID NOs 50-73. In several embodiments, the synthetic enhancer assembly (not including a gene of interest) has at least 95% sequence identity to one or more of the nucleic acid sequences of SEQ ID NOs 50-73. In several embodiments, the synthetic enhancer assembly (not including a gene of interest) is selected from the group consisting of one or more of the nucleic acid sequences of SEQ ID NOs 50-73.

[0127] In several embodiments, the ZFP array is encoded by a nucleic acid having at least 85% sequence identity to one or more of the nucleic acids of SEQ ID NOs 98-143. In several embodiments, the ZFP array is encoded by a nucleic acid having at least 90% sequence identity to one or more of the nucleic acids of SEQ ID NOs 98-143. In several embodiments, the ZFP array is encoded by a nucleic acid having at least 95% sequence identity to one or more of the nucleic acids of SEQ ID NOs 98-143. In several embodiments, the ZFP array is encoded by a nucleic acid selected from the group consisting of one or more of the nucleic acids of SEQ ID NOs 98-143.

[0128] In several embodiments, a mini-VPR transcriptional activator is coupled to the ZFP array. In several embodiments, the mini-VPR transcriptional activator is encoded by a nucleic acid having at least 85% sequence i den tity to one or more of the nucleic acids of SEQ ID NO 146 or 147. In several embodiments, the mini-VPR transcriptional activator is encoded by a nucleic acid having at least 90% sequence identity to one or more of the nucleic acids of SEQ ID NO 146 or 147. In several embodiments, the mini-VPR transcriptional activator is encoded by a nucleic acid having at least 95% sequence identity- to one or more of the nucleic acids of SEQ ID NO 146 or 147. In several embodiments, the mini-VPR transcriptional activator is encoded by SEQ ID NO 146 or 147. In several embodiments, the mini-VPRPatent Application 97157.00116transcriptional activator is encoded by SEQ ID NO 146.

[0129] In several embodiments, the ZFP array and mini-VPR transcriptional activator is encoded by a nucleic acid having at least 85% sequence identity to one or more of the nucleic acids of SEQ ID NO 150-195. In several embodiments, the ZFP array and mini-VPR transcriptional activator is encoded by a nucleic acid having at least 90% sequence identity to one or more of the nucleic acids of SEQ ID NO 150-195. In several embodiments, the ZFP array and mini-VPR transcriptional activator is encoded by a nucleic acid having at least 95% sequence identity to one or more of the nucleic acids of SEQ ID NO 150-195. In several embodiments, the ZFP array and mini-VPR transcriptional activator is encoded by one or more of the nucleic acids of SEQ ID NO 150-195. In several embodiments, the ZFP array and mini-VPR transcriptional activator (including backbone structure) is encoded by a nucleic acid having at least 85% sequence identity to one or more of the nucleic acids of SEQ ID NO 196-242. In several embodiments, the ZFP array and mini-VPR transcriptional activator (including backbone structure) is encoded by a nucleic acid having at least 90% sequence identity to one or more of the nucleic acids of SEQ ID NO 196-242. In several embodiments, the ZFP array and mini-VPR transcriptional activator (including backbone structure) is encoded by a nucleic acid having at least 95% sequence identity to one or more of the nucleic acids of SEQ ID NO 196-242. In several embodiments, the ZFP array and mini-VPR transcriptional activator (including backbone structure) is encoded by one or more of the nucleic acids of SEQ ID NO 196-242.

[0130] In several embodiments, the zinc finger array, coupled to the mini-VPR transcriptional activator comprises an amino acid sequence having at least 85% sequence identity to one or more of the amino acids of SEQ ID NOs 623-668. In several embodiments, the zinc finger array, coupled to the mini-VPR transcriptional activator comprises an amino acid sequence having at least 90% sequence identity to one or more of the amino acids of SEQ ID NOs 623-668. In several embodiments, the zinc finger array, coupled to the mini-VPR transcriptional activator comprises an amino acid sequence having at least 95% sequence identity to one or more of the amino acids of SEQ ID NOs 623-668. In several embodiments, the zinc finger array, coupled to the mini-VPR transcriptional activator comprises an amino acid sequence of one or more of the amino acids of SEQ ID NOs 623-668.

[0131] In several embodiments, the zinc finger array, coupled to the mini-VPR transcriptional activator (and bicistronically to a puromycin resistance gene) comprises an amino acid sequence having at least 85% sequence identity to one or more of the amino acids of SEQ ID NOs 295-340. In several embodiments, the zinc finger array, coupled to the mini-Patent Application 97157.00116VPR transcriptional activator (and bicistronically to a puromycin resistance gene) comprises an amino acid sequence having at least 90% sequence identity to one or more of the amino acids of SEQ ID NOs 295-340. In several embodiments, the zinc finger array, coupled to the mini-VPR transcriptional activator (and bicistronically to a puromycin resistance gene) comprises an amino acid sequence having at least 95% sequence identity7to one or more of the amino acids of SEQ ID NOs 295-340. In several embodiments, the zinc finger array, coupled to the mini-VPR transcriptional activator (and bicistronically to a puromycin resistance gene) comprises an amino acid sequence durgs of one or more of the amino acids of SEQ ID NOs 295-340.

[0132] After delivery7of both the synthetic enhancer construct (if used) and the ZFP construct, the binding of the ZFP to the target sequences within the synthetic enhancer construct enables the transcriptional activator protein to initiate translation of the gene of interest encoded by the synthetic enhancer construct. Advantageously, the use of the synthetic enhancer construct, and the screen of the component parts thereof (the GX1X2 motifs and the spacers) allows for precise modulation of transcription of the gene of interest, without otherwise impacting transcription of other genes within the host genome.Drug Switches

[0133] As disclosed above, in several embodiments, the ZFPs provided for herein can be integrated into a drug switch, which can be used to turn on (or off) transcription of a gene of interest at a particular time, for example, administration of a sufficient quantity of a triggering drug (or compound). Generally speaking, a drug switch comprises one or more transcriptional effectors, one more of DNA binding domain, and one or more drug interaction domain. These components can be encoded together and tethered by, for example linker sequences, or can be encoded separately and linked by the presence of a drug that couples the components through the binding of the drug by a plurality of drug interaction domains.Drug Switch Mechanism

[0134] A variety of different drug switch approaches can be used, depending on the embodiment and the goal of the functionality of the switch. For example, in some embodiments, Immunomodulatory Drug (IMiD)-based switches are used. In several embodiments, the IMiD drugs comprise one or more of lenalidomide, pomalidomide, thalidomide, or any combination thereof. In addition, Proteolysis Targeting Chimeras (PROTACs) may be used, such as bi-functional PROTACs. In some embodiments, chemically induced dimerization (CID) switches are used. Such approaches employ, for example, heterodimerization between FKBP and FRB upon exposure to rapamycin or homo dimerizationPatent Application 97157.00116of FKBP upon exposure to Rimiducid, for example. In several embodiments, drug-gated transcription factors (e.g., non-IMiD) switches are used. In several embodiments, the estrogen receptor or a modified estrogen receptor (ERT2) is used based on its responsiveness to tamoxifen or derivatives / metabolites thereof. Additional embodiments employ the glucocorticoid receptor and / or the progesterone receptor (responsive to dexamethasone and mifepristone, respectively). In several embodiments, degron / stability-based switches are used. In such embodiments, a drug like Shield- 1 can be used to stabilize FKBP protein or trimethoprim can be used to destabilize the dihydrofolate receptor. In several embodiments, protease-based switches are used. For example, in several embodiments, drugs like grazoprevir or asunaprevir are used to inhibit HCV NS3 protease to turn on or off the switch. In several embodiments, G-protein coupled receptor (GPCR) or other receptor switches are used. Designer Receptors Exclusively Activated by Designer Drugs (or DREADs) are used in some embodiments, wherein modified G-protein coupled receptors (e.g., muscarinic receptors) are inserted into target cells and are modified to respond only to specifically designed drugs, such as, for example, Clozapine-N-oxide or Salvinorin B (SALB), not natural body chemicals or unmodified drugs, to allow activation of specific pathways (which can be harnessed to drive transcription as provided for herein). In several embodiments, transcriptional drug switches are used. For example, while often used in a system-wide or in vitro or experimental in vivo setting, doxycycline can be used in a Tet-On / Tet-Off system to regulate transcriptional activity. In several embodiments. CRISPR based drug gated switches are used. In such embodiments, for example, IMiD drugs can be used to induce degrons or dCas (e.g., dCas9) can be dimerized with an effector (e.g., a transcription factor) upon exposure to a drug.Drug Interaction Domains

[0135] As shall be appreciated, each of the drug switch mechanisms disclosed herein each involve a receptor or other drug interaction domain that interacts, directly or indirectly, with a drug that is the trigger for switch activation (and thus activation or reduction of transcription). By w ay of example, two non-limiting approaches are disclosed in additional detail, however, it shall be appreciated that one of ordinary skill in the art can readily select a drug switch mechanism that is desired and deploy it within the embodiments provided for herein to regulate expression.

[0136] In several embodiments, non-IMiD drugs are used as the operator of the switch. In several embodiments, the modified estrogen receptor, ERT2, is used based on its responsiveness to tamoxifen, or its metabolite.4-Hydroxytamoxifen (4-OHT). In this scenario, a DNA binding domain, such as the ZFP binding domains provided for herein, is tethered (e.g.,Patent Application 97157.00116via a linker sequence) to a transcriptional activator, that is likewise tethered (e.g., via a linker sequence), to ERT2. In the absence of tamoxifen, or metabolites thereof, transcription is not active. However, upon addition of tamoxifen, or metabolites thereof, the complex translocates to the nucleus where transcription of a target gene occurs. As discussed herein, the target gene can be identified, for example, by a synthetic enhancer element as disclosed herein that is targeted and bound by a ZFP domain. See Figure 10A. In several embodiments, the ERT2 estrogen receptor is encoded by a nucleic acid having at least 85% sequence identity to SEQ ID NO: 758, 775, or 792. In several embodiments, the ERT2 estrogen receptor is encoded by the nucleic acid of SEQ ID NO: 758, 775, or 792. In several embodiments, the ERT2 estrogen receptor comprises an amino acid sequence having at least 85% sequence identity to SEQ ID NO: 767, 784. or 801. In several embodiments, the ERT2 estrogen receptor comprises an amino acid sequence of SEQ ID NO: 767, 784, or 801. In several embodiments, the drug interaction domain further comprises a p65 protein. In several embodiments, the p65 protein is encoded by a nucleic acid having at least 85% sequence identity to SEQ ID NO: 756, 773, or 790. In several embodiments, the p65 protein is encoded by the nucleic acid of SEQ ID NO: 756, 773. or 790. In several embodiments, the p65 protein comprises an amino acid sequence having at least 85% sequence identity to SEQ ID NO: 765, 782, or 799. In several embodiments, the p65 protein comprises an amino acid sequence of SEQ ID NO: 765, 782, or 799. In several embodiments, the synthetic drug switch comprises an amino acid having at least 85% sequence identity to one or more of SEQ ID NO: 770, 787, and 804. Similar to ERT2, the glucocorticoid receptor or progesterone receptor can be employed as a drug interaction domain.

[0137] Another approach provided for herein involved use of a cereblon (protein that normally tags other proteins, by ubiquitination, for degradation) that is modified by removal of the DDB1 domain to prevent its association with the E3 ubiquitin ligase complex (del.CRBN). In several embodiments, del.CRBN is fused to a DNA binding domain, such as the zinc fingers provided for herein. In several embodiments, a second fusion complex is generated in which a transcriptional activator is fused to CRIMP, a binder of IMiD drugs derived from an IKZF3 protein. In several embodiments, the fusions can be reversed, if desired, for example the DNA binding domain can be fused to CRIMP and the transcriptional activator to del .CRBN. In either case, the two drug interaction domains (e.g., del.CRBN and CRIMP) form a complex by binding to an IMiD drug, which can then be used to activate transcription of a gene of interest (e.g., a gene targeted by way of a synthetic enhancer element used to attract the DNA binding domain). See Figure 10B.Patent Application 97157.00116

[0138] In several embodiments, the drug interaction domain comprises a cereblon protein, optionally a cereblon protein that is engineered to (i) not express a DDB1 subdomain or (ii) express a DDB1 subdomain that does not allow the cereblon to interact with an E3 ubiquitin ligase complex. In several embodiments, the cereblon protein is encoded by a nucleic acid having at least 85% sequence identity' to SEQ ID NO: 684, 708, or 732. In several embodiments, the cereblon protein is encoded by SEQ ID NO: 684, 708, or 732. In several embodiments, the cereblon protein comprises an amino acid sequence having at least 85% sequence identity to SEQ ID NO: 697, 721, or 745. In several embodiments, the cereblon protein comprises an amino acid sequence of SEQ ID NO: 697, 721, or 745.

[0139] In several embodiments, the cereblon protein is fused to the transcriptional effector domain. In several embodiments, the transcriptional effector domain comprises a mini-VPR domain and the cereblon-transcriptional effector domain fusion is encoded by a nucleic acid sequence having at least 85% sequence identify to SEQ ID NO: 693, 717, or 741. In several embodiments, the cereblon-transcriptional effector domain fusion is encoded by SEQ ID NO: 693, 717, or 741. In several embodiments, the transcriptional effector domain comprises a mini-VPR domain and the cereblon-transcriptional effector domain fusion has an amino acid sequence having at least 85% sequence identify to SEQ ID NO: 700, 724, or 748.

[0140] In several embodiments, the drug interaction domain further comprises a CRIMP domain. In several embodiments, the CRIMP domain is encoded by a nucleic acid having at least 85% sequence identify to one or more of SEQ ID NO: 688. 712, or 736. In several embodiments, the CRIMP domain is encoded by a nucleic acid of SEQ ID NO: 688, 712, or 736. In several embodiments, the CRIMP domain has an amino acid sequence having at least 85% sequence identify to one or more of SEQ ID NO: 702, 726, or 750. In several embodiments, the CRIMP domain has an amino acid sequence of SEQ ID NO: 702, 726, or 750. In several embodiments, the CRIMP domain is fused to the ZFP. In several embodiments, the CRIMP domain-ZFP fusion has a nucleic acid sequence having at least 85% sequence identify to one or more of SEQ ID NO: 694, 718, or 742. In several embodiments, the CRIMP domain-ZFP fusion has a nucleic acid sequence of SEQ ID NO: 694, 718, or 742. In several embodiments, the CRIMP domain-ZFP fusion has an ammo acid sequence having at least 85% sequence identify to one or more of SEQ ID NO: 705, 729, or 753. In several embodiments, the CRIMP domain-ZFP fusion has an amino acid sequence of SEQ ID NO: 705, 729, or 753.Examples

[0141] The following examples represent non-limiting embodiments of compositionsPatent Application 97157.00116and related experimental methods related to the present disclosure.Overview

[0142] As discussed herein, cell therapies that can regulate protein production have the potential to address numerous challenging diseases. To develop these augmented cell therapies, potent and highly specific transcriptional regulators are needed that control expression of the gene of interest without interfering with native transcriptional programs. Provided for herein, and laid out in the non-limiting examples that follow, is the design of highly specific synthetic zinc finger (ZF) based transcription factors (zfTFs) with unprecedented orthogonality to human and mouse genomes, enabling the design of regulatory gene circuits with minimal disruption of host transcription.

[0143] To build the zfTFs, synthetic, mammalian genome orthogonal, enhancers that the zfTFs would bind to control transcription were generated. The synthetic enhancers were designed by generating all possible 19 nucleotide (19mer) sequences consisting of 6 GNN trinucleotides plus one final N (N= A, T, C, or G). The 19mers were then computationally filtered to remove: 1) sequences that have matches or near matches in the human genome, 2) sequences matching mammalian transcription factor binding sites. From the remaining 19mers, synthetic enhancers were designed comprising 4 repeats of the 19mer separated by lObp spacers devoid of transcription factor binding sites. Synthetic enhancers were then inserted upstream of a minimal promoter driving expression of GFP, as a non-limiting example of a gene of interest to be expressed (here for detection purposes).

[0144] For each synthetic enhancer, 2 unique zinc finger arrays were designed. The arrays were fused to a minVPR transcriptional activator. The enhancers and zfTFs were then transfected into HEK293T cells to evaluate their strength (GFP expression) and specificity (target enhancer binding). Top enhancer / zfTF pairs were transduced into multiple cell types and further evaluated by bulk RNA-seq to quantify the off target transcriptional impact of the zfTFs.

[0145] Post computational filtering, 23 synthetic enhancers and 46 synthetic zfTFs were evaluated. 3 zfTF robustly outperformed a literature control zfTF in strength by >1.5x and specificity by >20x in a transfection-based screen. The bulk RNAseq data showed that the top 3 zfTFs drive greater GFP expression per zfTF RNA molecule than the literature control while affecting the transcription of significantly fewer off target genes (1, 7, or 9 genes) than the literature control (80 genes). The strongest zfTF was further shown to have reduced off target effects compared to the literature control in 2 additional human cell types (T cells and liver cells) and mouse melanoma cells (each as a non-limiting example of a target cell).Patent Application 97157.00116

[0146] As described in more detail herein, the present disclosure describes unique, mammalian genome orthogonal enhancers and zfTFs. The zlTFs can serve as potent transcriptional activators and have demonstrated exquisite safety profiles by activating very few non-targeted genes. According to some embodiments, such zfTFs can be incorporated into small molecule responsive transcriptional switches to enable next-generation cell therapies that can regulate the production of therapeutic payloads to respond dynamically to disease.Example 1 - Synthetic Enhancer Design and Screening

[0147] As discussed above, in several embodiments, the design and synthesis of synthetic enhancer elements allows for precise and regulated control of gene expression of a gene, or genes of interest. Advantageously, in several embodiments, such expression is achieved in a genome orthogonal manner wherein endogenous gene expression of the host (e.g., a human) is not affected by the synthetic enhancer, or a corresponding DNA binding motif.

[0148] The enhancer sequences, schematically depicted in Figure 1 as (1) can be generated and a corresponding ZFP (schematically depicted as (2) in Figure 1) is generated to specifically bind the enhancer sequence.

[0149] In this non-limiting example, a process is disclosed to design synthetic enhancer elements that are orthogonal to a host genome, such as a human or murine genome, by way of example. Figures 2A-2D depict the process flow undertaken to design such synthetic enhancers.

[0150] At the outset, apool of candidate sequences is generated. According to several embodiments, the candidate sequences (also referred to as enhancer core sequences) comprise a repeated GX1X2 motif, with between 3 and 8 repeats of the motif. In this non-limiting example, a 6x repeat is used, resulting in an 18 nucleotide candidate enhancer core sequence that is supplemented with at least one additional nucleotide (which is any of adenine, thymine, guanine, or cytosine). Thus, random 19-mers containing 6 GX1X2 motifs were generated.

[0151] To ensure orthogonality to the human genome (as discussed herein, other host / target genomes can be used, depending on the embodiment), the pool of 19-mer candidate core sequences were screened against the human genome to identify potential similar sequences within the genome. In this example, the screening was used to eliminate two subpools of the candidate core sequences: (i) those 19-mers having an exact sequence identify to one or more sites in the human genome and (ii) those 19-mers having 1 or 2 base pair mismatches to one or more sites in the human genome. Following this initial screening, aPatent Application 97157.00116second screening was performed in which those 19-mers having binding sites for transcription factors (including weak binding sites) were removed from candidacy. In this example, the 19-mers were screened versus the HOCOMOCOvll full collection, though other databases and / or predictive binding algorithms can be used in other embodiments.

[0152] Figure 3A shows the results of the initial generation of the candidate pool and the resultant number of candidate core 19-mers as the screening process continued. The initial pool of candidate core 19-mers that have 100% sequence identity to one or more location of the human genome (i.e., 0 mismatches) was nearly IOXsequences (total). After removing those sequences having 1 mismatch (~0.6xl08total) and 2 mismatches (just over 106sequences total) to the human genome, there remained twenty-five 19-mers with 3 or more mismatches.

[0153] Following the removal of candidate 19-mers with exact or 1-2 base pair mismatches, the remaining candidate 19-mers were screened for putative transcription factor binding sites. Of the remaining pool, two of the sequences contained at least putatively weak transcription factor binding sites, which were removed. The final pool of candidate core sequences was 23 (twenty -three) 19-mers. Figure 3B shows an alignment table of the final pool of candidate core 19-mers organized by the sequence identity of the 6x 18-mer repeat sequence (SEQ ID NO. 24 represents the 100% consensus sequence), along with non-limiting examples of sequences with 90% sequence identity, 80% sequence identity, and 70% sequence identity to the consensus sequence. It shall be appreciated that other sequences (i.e., mismatched from consensus at locations other than those depicted) with similar percent identity are within the scope of the present disclosure.

[0154] Moving from these candidate 19-mers that passed the screening process, as described in Figure 2D, the candidate 19-mers are used to generate synthetic enhancers comprising 4x repeats of the 19-mer with lObp spacers between each repeat. As discussed elsewhere herein, longer sequences may be used for either the core synthetic enhancer portions (e.g., the 19-mers may be longer or shorter) as well as the spacers (which may be longer or shorter).

[0155] In designing and selecting the spacer sequences, the putative sequences underwent a preliminary screening and a similar screening process to that of the core sequences upon in silico assembly into the concatemeric format of the synthetic sequence enhancer. The preliminary screening was performed to identify any candidate spacer sequences that comprises one or more transcription factor binding sites. Those candidate spacer sequences that include even a weak binding site for a transcription factor were removed. The remaining pool of candidate spacer sequences (i.e., those without a transcription factor binding site within theirPatent Application 97157.00116sequences) were then screened for the creation of a transcription factor binding site (or sites) at the junction of the candidate spacer sequence with the upstream and / or downstream 19-mer. By way of example, in a synthetic transcription enhancer with a 4x 19-mer repeat, there are three spacers that are intercalated between the 19-mers (spacer 1 is between 19-mer 1 and 19-mer 2, spacer 2 is between 19-mer 2 and 19-mer 3, and spacer 3 is between 19-mer 3 and 19-mer 4). Thus, the sequence for spacer 1 is screened for creation of a binding site that spans the 19-mer 1 -spacer 1 junction and the spacer 1-19-mer 2 junction. If a transcription factor binding site is so created at the junction(s), the spacer sequence is removed from the candidate pool, at least with respect to the selected 19-mer repeat (that spacer sequence may readily be usable in a different synthetic enhancer element wherein the 19-mer sequence is different).

[0156] The resultant generated synthetic enhancer elements, with their orthogonal design with respect to the human genome (as an example) were then used as targets for the design and generation of DNA binding proteins.Example 2 - DNA Binding Proteins Targeting Designed Synthetic Enhancers

[0157] As discussed above, in several embodiments, the design and synthesis of synthetic enhancer elements allows for precise and regulated control of gene expression of a gene, or genes of interest. In this example, DNA binding proteins, here, as a non-limiting example, a zinc finger protein or ZFP, were designed that target the genome orthologous synthetic enhancer elements. After design, the ZFPs were tested in an in vitro test system to evaluate whether the genome orthogonality predicted during design was realized in a functional expression system with human cells.

[0158] As schematically depicted in Figure 4A, ZFPs were designed to target each GX1X2 motif within a synthetic enhancer element (recall that the non-limiting example of synthetic enhancers comprised 4 repeats of the core sequence, with the core sequence comprising 6 repeats of a given GX1X2 motif, with an additional nucleotide on the 3’ end (any nucleotide).

[0159] The designed ZFPs were linked into a 6-factor array, configured in a 3 by 2 arrangement. In this arrangement, ZFP 1 and 2 are linked by a first linker, ZFP 3 and 4 are linked by a second linker, ZFP 5 and 6 are linked by a third linker, and two additional linkers are used to couple each of the pairs. Each of Linkers 1, 2, and 3 are the same length and each of Linkers 4 and 5 are the same length, leading to the following configuration: ZFP1-L1-ZFP2-L4-ZFP3-L2-ZFP4-L5-ZFP5-L3-ZFP6. These were then tested for both strength of binding and specificity of binding to the genome orthogonal synthetic transcription factors that werePatent Application 97157.00116generated.

[0160] Figure 4B schematically depicts an experimental setup configured to assess the strength of the binding of a ZFP to its cognate (e.g., complementary) synthetic enhancer element (e.g., the target of the ZFP). Two separate constructs were generated. The first construct included a promoter sequence (here, a well-known short, intron-less form of the EF 1 alpha promoter, known as EFS; see Rao et al.. “Systematic Comparison of the EF-1 Alpha Short (EFS) and Viral Promoters for Gene Modification of Human Primary Cells for Clinical Applications". Blood 2014 124 (21):3497.) was used as anon-limiting example) coupled to a ZFP binding domain (a 6x array of ZFPs) and further coupled to a transcriptional activator. In this example, a mini-VPR transcriptional activator (a truncated VP64-p65-Rta tripartite activator) was used. The second construct included a synthetic enhancer element (e.g.. an enhancer for which the ZFP is designed to bind), a YBTATA minimal promoter and GFP (a reporting element). These two constructs were separately introduced into a cell and upon binding of the ZFP to its cognate synthetic enhancer, transcription of the GFP reporter gene is initiated, allowing for detection of GFP.

[0161] Strength of binding specificity was calculated using the following variables:

[0162] TFX= tested ZF array-based transcription factor

[0163] Px= on target promoter for TFX

[0164] TFOFF = negative control ZF array-based transcription factor, not expected to bind Px.

[0165] POFF = negative control enhancer, not expected to be bound by TFx

[0166] TF515 = positive control ZF10-1 based transcription factor

[0167] Psi4= target enhancer for TF514

[0168] Strength = (TFX+PX)-(TFOFF+ PX) / (TF515+P514)-(TF515+POFF)

[0169] Figure 4C shows a schematic of an experimental setup fortesting for zinc finger transcription factor specificity (e.g., lack of binding to a non-target enhancer). Two separate constructs were generated. The first construct included a promoter sequence (here, the EFS promoter was used as a non-limiting example) coupled to a ZFP binding domain (a 6x array of ZFPs) and further coupled to a transcriptional activator (as in Figure 4B). The second construct included an off-target synthetic enhancer element (e.g., an enhancer for which a ZFP is not designed to bind), a YBTATA minimal promoter and GFP (a reporting element). These two constructs were separately introduced into a cell and because there should be limited to no binding of the ZFP to the non-target synthetic enhancer, transcription of the GFP reporter gene would not be initiated, resulting in a failure of detection of GFP.Patent Application 97157.00116

[0170] Specificity of binding was calculated using the variables as described above and the following formula:

[0171] Specificity = (TFX+PX)-(TFOFF+ PX) / (TFX+POFF)

[0172] Figure 5 A shows a scatterplot of strength versus specificity'. As compared to the positive control, several of the candidate ZFPs bound to their cognate with increased strength and specificity (e.g., those in the upper right portion of the scatter plot). Those ZFPs labeled correspond to the following DNA sequences:Table 1 - ZFP Identifiers

[0173] An additional strength versus specificity screen was performed on selected ZFPs that in the first screen appeared to demonstrate enhanced binding to the target synthetic enhancer with high specificity. Figure 5B show these data relative to the positive control (identified in the scatterplot at the intersection of the gridlines, GF000514). These more granular data identified three particular ZFPs (GF000474, GF000461, and GF000471) that bound to their cognate synthetic enhancer with enhanced strength and specificity. Figure 5C shows a cytometric analysis of cells measuring GFP expression for these ZFPs (as well as controls). The top three rows depict the amount of GFP measured for ZFPs GF000474, GF000471, and GF000461 (each targeting their respective synthetic enhancer (GF000497, GF000494, and GF000484, respectively). The total GFP detected was greater for each of these candidates than the corresponding positive control (GF000514-GF000515). Negative controlsPatent Application 97157.00116show little to no signal for GFP (GF000436 and GF0000482 is a mismatched ZFP-synthetic enhancer combination and "NV" is a no vector negative control). While these were the top performers in this particular non-limiting example of an assay, it shall be appreciated that other ZFP-synthetic enhancer combinations provided for herein also can exhibit enhanced strength and specificity of binding (and coordinate gene expression) under different conditions (e.g., driven by a different promotor and / or driving expression of a gene other than GFP).

[0174] Having established that ZFP-synthetic enhancer elements can interact with specificity and strength to result in enhanced expression of a gene of interest, bulk RNA-sequencing was performed to measure the quantity and presence of RNA molecules to evaluate a snapshot of gene expression, or transcriptome, in cells expressing ZFP-synthetic enhancer element combinations. Lentiviral constructs (employing a VSV-G envelope protein for pseudotyping) were generated for the following constructs:

[0175] GF000461, GF000471, GF000474 - selected ZF transcription factors;

[0176] GF000484, GF000494, GF000497 - respective synthetic enhancers targeted by selected ZF transcription factors, driving a GFP reporter;

[0177] GF000515 - ZF10-1 based transcription factor;

[0178] GF000514 - synthetic enhancer targeted by GF000515, driving a GFP reporter;

[0179] GF000717 - negative control - expresses miniVPR attached to mCherry (in place of a ZF array).

[0180] Human K562 cells were transfected in triplicate according to the assay layout shown in Figure 6A (NV = no vector control). Subsequently, RNA was collected, and RNA sequencing performed to evaluate gene expression. Figure 6B shows a bar chart of data related to the strength of ZFP binding to the respective synthetic enhancer element. GFP count (per transcription factor RNA count) was measured and compared to the GFP RNA count level of the positive control (GF000515). These data show that each of the selected ZFPs bind to their respective synthetic enhancer element to induce GFP transcription at least as strongly as the positive control with two ZFPs showing robustly increased transcription induction. With respect to specificity of binding (i.e., orthogonality to the host cell genome), Figure 6C shows data comparing the number of off target genes regulated by the selected ZFPs as compared to control. These data show that the selected ZFPs exhibit at least 10-fold less off-target gene regulation than the positive control, with one ZFP showing 100-fold less off-target effects. These data collectively demonstrate that ZFP-synthetic enhancer pairs as provided for herein exhibit enhanced binding to engineered target sequences encoded by the synthetic enhancer pairs while also exhibiting enhanced specificity which leads to reduced off-target genePatent Application 97157.00116regulation.

[0181] Building on the experiments above, additional data was collected when additional cell lines were transfected with selected ZFP constructs. The experimental setup corresponded to that shown in Figure 6A, but with either Jurkat (human T cell line), Hep3b (human liver cell line), or Bl 6F 10 (murine melanoma cell line) and using GF000717 (negative control, miniVPR driving mCherry), GF000474 (non-limiting example of a ZF transcription factor as provided for herein, along with its respective synthetic enhancer GF000497, driving a GFP reporter), or GF00051 (a ZF10-1 based transcription factor (with its respective synthetic enhancer GF000514, driving a GFP reporter).

[0182] Figures 7A, 8A, and 9A show box plots depicting the number of RNA-seq reads associated with the GF000474 ZF transcription factor construct, the GF000515 construct, or the GF000717 negative control in Jurkat (7 A), Hep3B (8A), or B16F10 (9 A). In each of these non-limiting examples of cell lines, including both human and murine, the non-limiting GF000474 construct was more robustly expressed than either of the other two constructs. These data suggest that the ZFP constructs as provided for herein can be expressed in a variety of host cell t pes, indicating their advantageous versatility.

[0183] Figures 7B, 8B, and 9B show bar charts depicting bar chart depicting ZF strength in terms of the number of GFP RNA counts measured per ZFTF RNA count (e.g., for each ZFTF RNA counted, how many GFPs are counted) for the GF000474 ZF transcription factor construct, the GF000515 construct, or the GF000717 negative control in Jurkat (7B), Hep3B (8B), or Bl 6F10 (9B). In the Jurkat and Bl 6F10 cell lines, GF000474 is stronger than the GF000515 literature control. In Hep3B, while the GF00051 literature control showed stronger GFP expression, the non-limiting GF000474 construct w as notably stronger than the negative control. These data suggest that the ZFP constructs as provided for herein can drive robust expression of a gene of interest in multiple cell types.

[0184] Figures 7C, 8C, and 9C show line graphs correlating the change of expression of each gene that was quantified in the RNA seq analysis relative to the number of genes dysregulated that show that degree of fold change. In other words, a point on the line that is a high value on the Y axis and a low value on the X axis represents a ZFP that dysregulates a relatively larger number of genes but the overall change in expression of those genes is relatively small. In contrast, a data point that has a low' Y axis value and a high X axis value represents a ZFP that dysregulates a relatively smaller number of genes, but each the change in expression of the gene is relatively large. In each of the three non-limiting examples of cell types used, including both human and murine, the non-limiting example of a ZFP as providedPatent Application 97157.00116for herein, here GF000474, dysregulated expression of fewer genes than the literature control, GF000515. Furthermore, of the genes with dysregulated expression, the non-limiting example of a ZFP as provided for herein, here GF000474, changed the expression of those genes to a lower degree than that of the GF000515 literature control. These data demonstrate that ZFP constructs as provided for herein, along with synthetic enhancers as provided for herein can be well expressed in a variety of cells, can drive efficient expression of a gene of interest, and dysregulate fewer genes and change expression to a lesser degree than other ZFPs. This advantageously allows the ZFPs provided for herein to be used to specifically target genomic DNA regions of interest to modulate expression of genes and / or related proteins of interest with specificity, leading to reduced off target effects.Example 3 - Genome Orthogonal Modulation of Chimeric Antigen Receptor Expression

[0185] This is a prophetic example.

[0186] Two constructs will be generated, one synthetic enhancer construct and one ZFP array construct. The synthetic enhancer construct will comprise one of the synthetic enhancers disclosed herein and selected from 6x repeats of any combination of SEQ ID NOs 25-48 and a YBTATA promoter of SEQ ID NO: 49 and a nucleic acid sequence encoding a CD19-directed Chimeric Antigen Receptor (CAR) selected from the CARs comprising the sequence of one or more of SEQ ID NO. 669681. The ZFP array construct comprises an EFS promotor driving expression of one of the ZFP arrays linked to a mini-VPR transcriptional activator of selected from those provided in SEQ ID NOs. 197-242.

[0187] Each of the two constructs will be packaged in a lentiviral vector and used to infect host cells, which are human T cells, by way of example. The synthetic enhancer element sequence will be located within the T cell genome. Expression of the ZFP construct will result in translation of the encoded ZFP array. The ZFP array will bind to the synthetic enhancer element with the transcriptional activator acting to initiate transcription at the YBTATA locus and result in the transcription and, ultimately, translation of the CD19-directed CAR.

[0188] Immunohistochemical staining of the T cells using a fluorescent die coupled to CD 19 will show surface expression of the CD 19-directed CAR.

[0189] Raji, NALM6, and Daudi (all CD19+ cell lines that will be obtained from ATCC) and K562 (a CD19- cell line to be obtained from ATCC) will be cultured using standard cell culture techniques. Exposure of the CD 19+ cell lines to the T cells expressing the two constructs will result in increased levels of cytotoxicity (as compared to the CD 19-negative cell line).Patent Application 97157.00116Example 4 - Modulation of Gene Expression Through Drug Switches

[0190] As discussed above, drug switches are provided for in some embodiments. Two non-hmiting types of switches are schematically depicted in Figures 10A-10B and discussed above. Using ZFPs screened in the prior examples as non-hmiting examples (GD000461,000471, 000474), switches were designed as schematically depicted in Figures 10A-10B and discussed above. Figure 11B provides a brief description of the various switch constructs, which in this example comprise a ZFP linked to ERT2 and p65 (a component of the NF-kB complex that is involved in ERT2 interaction with tamoxifen and it’s metabolites). A GFP reporter gene responsive to one of the three example ZFPs w as used (in combination with the corresponding ZFP).

[0191] HEK293 cells were transfected in triplicate with a pair of constructs (e.g., the ZFP-ERT2-p65 fusion and the corresponding reporter fusion). Figure 11 A show s a bar graph of the concentration dependent expression of the respective reporter in response to various concentrations of 4-OHT. As can be seen based on the mean fluorescence intensity, in each case, addition of 4-OHT upregulated transcription of the reporter gene. Figures 11 C- 1 IE show more detailed plot of a cytometric analysis of cells for GFP expression as a function of 4-OHT concentration. Figure 11C show s data for the 900 / 484 pairing, Figure 11D show-s data for the 901 / 494 pairing, and Figure HE shows data for the 902 / 497 pairing. These data demonstrate that, in accordance with several embodiments, a drug sw itch can be used to specifically trigger transcription of the gene of interest in response to exposure to a triggering drug.

[0192] Additionally, IMiD-based drug switches were tested. Figure 12A shows schematics of constructs for the structure of an ERT2 drug swatch (top) and an IMiD drug switch (bottom). Figure 12B shows a schematic for the structure of an enhanced Blue Fluorescent Protein (EBFP or BFP) reporter construct (top) and a mCherry reporter construct (bottom). By way of example only IMiD switches were paired with mCherry and ERT2 switches were paired with BFP, though any combination of drug switch and reporter can be used (with the reporter being optional, e.g., optionally omitted for clinical use).

[0193] Figure 13 provides a brief description of the various constructs. The ERT2-ZFP constructs are those described earlier in this example. The IMiD constructs comprise a del.CRBN linked to a transcriptional activator and, at the nucleic acid level, a ZFP linked to CRIMP. The nucleic acid constructs are multi-cistronic and thus the two fusion proteins are expressed as separate proteins in test cells. Corresponding BFP or mCherry reporters are listed for each ZFP (such that expression of the reporter corresponding to one ZFP can be detected through that reporter while expression tied to another ZFP can be detected through the otherPatent Application 97157.00116reporter).

[0194] HEK293 cells were again used and were transfected in triplicate with two pairs of constructs, one ERT2 switch construct and its corresponding reporter construct (based on the ZFP used) and one IMiD switch construct and its corresponding reporter construct (based on the ZFP used). Cells were first exposed to either tamoxifen alone (at 1 nM) to trigger the ERT2 drug switch or pomalidomide (at 1 OnM) to trigger the IMiD drug switch.

[0195] Figure 14A shows data in a bar graph from the tamoxifen exposure. In each instance, the left of the pair of bars for each group is the mCherry signal (fold change relative to no drug) and the right of the pair of bars is the BFP signal. As shown in the data, despite being expressed in the same cells, the ERT2 gene switch was largely specifically induced in response to tamoxifen, with only one instance of cross-talk with the IMiD switch (899+327 / 901+1323 bars on far right). Otherwise, the tamoxifen induced over 10-fold induction of reporter expression upon activation of the switch.

[0196] Figure 14B shows corresponding data from pomalidomide exposure (bar identity is the same as in Figure 14A). Here, no significant cross-talk was observed with respect to pomalidomide inducing expression through the ERT2 switch. Rather, in this experiment, pomalidomide induced expression increases ranging from about 5-fold to over 100-fold, depending on the ZFP used.

[0197] Figure 14C shows data resulting from the activation of both switches together, through use of both tamoxifen and pomalidomide (at the same concentrations each were used alone; bar identity is the same as in Figure 14A). In this experiment, these data show that two different swatches can be induced in tandem, which provides the possibility of driving two different genes of interest together (or either one alone if only one switch is activated).

[0198] Taken together these data show that drug switches of different mechanistic types can be employed alone, or in combination, to drive gene expression of a specific gene of interest. The use of drug switches such as these allow for, in several embodiments, control of gene expression by allowing induction of expression to be turned on (here through addition of the drug) or off (by removing the drug).

[0199] It is contemplated that various combinations or subcombinations of the specific features and aspects of the embodiments disclosed above may be made and still fall within one or more of the inventions. Further, the disclosure herein of any particular feature, aspect, method, property, characteristic, quality, attribute, element, or the like in connection with an embodiment can be used in all other embodiments set forth herein. Accordingly, it should be understood that various features and aspects of the disclosed embodiments can be combinedPatent Application 97157.00116with or substituted for one another in order to form varying modes of the disclosed inventions. Thus, it is intended that the scope of the present inventions herein disclosed should not be limited by the particular disclosed embodiments described above. Moreover, while the invention is susceptible to various modifications, and alternative forms, specific examples thereof have been shown in the drawings and are herein described in detail. It should be understood, however, that the invention is not to be limited to the particular forms or methods disclosed, but to the contrary, the invention is to cover all modifications, equivalents, and alternatives falling within the spirit and scope of the various embodiments described and the appended claims. Any methods disclosed herein need not be performed in the order recited. The methods disclosed herein include certain actions taken by a practitioner; however, they can also include any third-party instruction of those actions, either expressly or by implication. In addition, where features or aspects of the disclosure are described in terms of Markush groups, those skilled in the art will recognize that the disclosure is also thereby described in terms of any individual member or subgroup of members of the Markush group.

[0200] All references cited herein, including but not limited to published and unpublished applications, patents, and literature references, are incorporated herein by reference in their entirety and are hereby made a part of this specification. To the extent publications and patents or patent applications incorporated by reference contradict the disclosure contained in the specification, the specification is intended to supersede and / or take precedence over any such contradictory material.Sequences

[0201] In several embodiments, there are provided amino acid sequences that correspond to any of the nucleic acids disclosed herein (and / or included in the accompanying sequence listing), while accounting for degeneracy of the nucleic acid code. Furthermore, those sequences (whether nucleic acid or amino acid) that vary from those expressly disclosed herein (and / or included in the accompanying sequence listing), but have functional similarity' or equivalency are also contemplated within the scope of the present disclosure. The foregoing includes mutants, truncations, substitutions, codon optimization, or other types of modifications.

[0202] In accordance with some embodiments described herein, any of the sequences may be used, or a truncated or mutated form of any of the sequences disclosed herein (and / or included in the accompanying sequence listing) may be used and in any combination. Sequences provided for herein that include an identifier, such as a tag or other detectable sequence (e.g., a Flag tag) are also provided for herein with the absence of such a tag or otherPatent Application 97157.00116detectable sequence (e.g., excluding the Flag tag from the listed sequence). A Sequence Listing in electronic format is submitted herewith. Some of the sequences provided in the Sequence Listing may be designated as Artificial Sequences by virtue of being non-naturally occurring fragments or portions of other sequences, including naturally occurring sequences. Some of the sequences provided in the Sequence Listing may be designated as Artificial Sequences by¬ virtue of being combinations of sequences from different origins, such as humanized antibody sequences.

Claims

1. Patent Application 97157.00116What is claimed is:

1. A synthetic transcription factor, comprising:a transcriptional effector domain;a synthetic enhancer element comprising a structure of Formula 1 :(Target-(Spacer)(X-i))(x)Formula 1wherein the Target of Formula 1 comprises a sequence of (GXiX2)nNy, wherein Xi and X2 each independently represent one of adenine, guanine, cytosine, and thymine,wherein n is an integer between 2 and 8,wherein N is selected from adenine, guanine, cytosine, and thymine, wherein y is an integer between 0 and 6;wherein the Spacer of Formula 1 comprises a sequence of between 6 and 12 nucleotides,wherein x is an integer between 3 and 8;wherein the synthetic enhancer element (i) does not have 100% sequence identity to any human or murine genomic DNA sequence, (ii) does not have fewer than 3 base pair mismatches with any human or murine genomic DNA sequence, and (iii) does not have any putative transcription factor binding sites (a) within the Target or Spacer sequences or (b) spanning a junction between a Target and its corresponding Spacer; anda DNA binding domain,wherein the DNA binding domain comprises a zinc finger protein, wherein the zinc finger protein binds a target sequence and allows the transcriptional effector to modulate transcription of a target gene; and wherein the synthetic transcription factor does not modulate transcription of non-target genes with a cell.

2. The synthetic transcription factor of Claim 1, wherein n=6.

3. The synthetic transcription factor of Claim 1 , wherein the synthetic enhancer element comprises at least one copy of SEQ ID NO. 24 (GXXGXXGXXGXXGXXGXX, wherein each instance of X is any nucleic acid).Patent Application 97157.001164. The synthetic transcription factor of Claim 3, wherein the first GX1X2 sequence does not comprise GGC, GGA, GAT, GCA, or GAA.

5. The synthetic transcription factor of Claim 3, wherein the second GX1X2 sequence does not comprise GTT, GCG. GGC, GGA, GAT, GCA, or GAA.

6. The synthetic transcription factor of Claim 3, wherein the third GX1X2 sequence does not comprise GTT, GGA, GAT, GCA, or GAA.

7. The synthetic transcription factor of Claim 3, wherein the fourth GX1X2 sequence does not comprise GTT, GCA, or GAA.

8. The synthetic transcription factor of Claim 3, wherein the fifth GX1X2 sequence does not comprise GTT. GCA, or GAA.

9. The synthetic transcription factor of Claim 3, wherein the sixth GX1X2 sequence does not comprise GTC, GGC, or GCC.

10. The synthetic transcription factor Claim 2, wherein none of the GX1X2 sequences comprises GGG, GGT, GAG, GTG, or GCT11. The synthetic transcription factor of Claim 1, wherein the Target of Formula 1 is selected from the group consisting of SEQ ID NO: 1-23.

12. The synthetic transcription factor of Claim 1, wherein the Target of Formula 1 is selected from the group consisting of SEQ ID NO: 25-47.

13. The synthetic transcription factor of Claim 1. wherein x = 4, resulting in four Target repeats intercalated with three Spacer repeats.

14. The synthetic transcription factor of Claim 1, wherein y is 1 and the Spacer of Formula 1 comprises 10 nucleotides.Patent Application 97157.0011615. The synthetic transcription factor of Claim 1, wherein the synthetic enhancer element is selected from a sequence having at least 80% identity to one or more of SEQ ID NO: 50-73.

16. The synthetic transcription factor of Claim 1, wherein the synthetic enhancer element is selected from a sequence having at least 90% identity to one or more of SEQ ID NO: 50-73.

17. The synthetic transcription factor of Claim 1, wherein the synthetic enhancer element is selected from the group consisting of SEQ ID NO: 50-73.

18. The synthetic transcription factor of Claim 1, wherein the synthetic enhancer element is selected from the group consisting of SEQ ID NO: 55, 65, and 68.

19. The synthetic transcription factor of any one of Claims 1 to 18, further comprising promoter sequence.

20. The synthetic transcription factor of Claim 19, wherein the promoter sequence comprises a YBTATA minimal promoter sequence.

21. The synthetic transcription factor of Claim 19 or 20, wherein the promoter sequence has at least 90% sequence identity to SEQ ID NO. 49.

22. A method of identifying at least one candidate genome orthogonal synthetic enhancer element, comprising:(i) generating a plurality of oligonucleotides comprising the following sequence: (GXiX2)nNy,wherein Xi and X2 each independently represent one of adenine, guanine, cytosine, and thymine,wherein n is an integer between 2 and 8,wherein N is selected from adenine, guanine, cytosine, and thymine, wherein y is an integer between 0 and 6;(ii) screening the plurality of oligonucleotides from (i) for (a) an exact matching sequence for the oligonucleotide from the human and / or murine genome and (b) aPatent Application 97157.00116sequence within the human and / or murine genome having 1 or 2 base pair mismatches from an individual oligonucleotide selected from the plurality of oligonucleotides from (i);(iii) removing from consideration an oligonucleotide meeting either (a) and / or (b) from (ii);(iv) screening the plurality of oligonucleotides from (iii) for putative binding sites for one or more transcription factors; and(v) removing from consideration an oligonucleotide having at least one putative transcription factor binding site,thereby generating at least one candidate genome orthogonal synthetic enhancer element.

23. The method of Claim 19, wherein, when n>2 and designing a spacer oligonucleotide between 8 to 12 nucleotides long, the spacer does not comprise any putative transcription factor binding sites within the 8 to 12 nucleotides and when inserted between a first GX1X2 sequence and a second GX1X2 sequence does not introduce any putative transcription factor binding sites spanning a junction between the first GX1X2 sequence and the spacer or the second GX1X2 sequence and the spacer.

24. The method of Claim 22 or 23, wherein the synthetic enhancer element has a sequence of GX1X2GX1X2GX1X2GX1X2GX1X2GX1X2N (SEQ ID NO: 48).

25. The method of Claim 23 or 24, wherein when n=6, the method further comprising removing an oligonucleotide from consideration when:(i) the first GX1X2 sequence comprises GGC, GGA, GAT, GCA, or GAA; (ii) the second GX1X2 sequence comprises GTT, GCG, GGC, GGA, GAT, GCA, or GAA;(iii) the third GX1X2 sequence comprises GTT, GGA, GAT, GCA. or GAA; (iv) the fourth GX1X2 sequence comprises GTT, GCA, or GAA;(v) the fifth GX1X2 sequence comprises GTT, GCA, or GAA; or (vi) the sixth GX1X2 sequence comprises GTC, GGC, or GCC.

26. The method of any one of Claims 22 to 24. further comprising removing an oligonucleotide from consideration when any of the GX1X2 sequences comprises GGG, GGT,Patent Application 97157.00116GAG, GTG, or GCT.

27. A method of generating at least one candidate genome orthogonal synthetic enhancer element, comprising:(i) generating a plurality7of in silico oligonucleotides comprising the following sequence: (GXiX2)nN,wherein Xi and X2 each independently represent one of adenine, guanine, cytosine, and thymine,wherein n is an integer between 2 and 8,wherein N is selected from adenine, guanine, cytosine, and thymine; (ii) screening the plurality of in silico oligonucleotides from (i) for (a) an exact matching sequence for the oligonucleotide from the human and / or murine genome and (b) a sequence within the human and / or murine genome having 1 or 2 base pair mismatches from an individual oligonucleotide selected from the plurality of oligonucleotides from (i);(iii) removing from consideration an in silico oligonucleotide meeting either (a) and / or (b) from (ii);(iv) screening the plurality7of in silico oligonucleotides from (iii) for putative binding sites for one or more transcription factor;(v) removing from consideration an in silico oligonucleotide having at least one putative transcription factor binding site;(vi) designing and inserting, in silico, a spacer oligonucleotide between 8 to 12 nucleotides long, wherein the spacer does not: (a) comprise any putative transcription factor binding sites within the 8 to 12 nucleotides or, (b) when inserted between a first GX1X2 sequence and a second GX1X2 sequence, introduce any putative transcription factor binding sites spanning a junction between the first GX1X2 sequence and the spacer or the second GX1X2 sequence and the spacer;(vii) removing from consideration any in silico oligonucleotide meeting either (a) or (b) from (vi); and(viii) synthesizing at least one oligonucleotide from those in silico oligonucleotides remaining after (vii), thereby synthesizing at least one candidate genome orthogonal synthetic enhancer element.

28. A method of generating a genome orthogonal synthetic transcription factor,Patent Application 97157.00116comprising:(i) generating a plurality’ of genome orthogonal synthetic enhancer element according to the method of Claim 27;(ii) coupling at least one of the plurality of genome orthogonal synthetic enhancer elements of (i) to a promoter and a gene of interest;(iii) designing a DNA binding protein that specifically binds to only one of the plurality of genome orthogonal synthetic enhancer elements of (i);wherein the DNA binding domain comprises a zinc finger protein; and (iv) coupling the DNA binding protein to a promoter and a transcriptional activator domain,wherein binding of the DNA binding domain to the synthetic enhancer element is configured to initiate transcription of the gene of interest without modulation of transcription of genes with a cell that are not the gene of interest.

29. A method for modulating expression of a gene of interest, comprising:(i) introducing into a host cell a synthetic transcription factor, the synthetic transcription factor comprising a synthetic enhancer assembly and a zinc finger protein assembly,wherein the synthetic enhancer assembly comprises:a synthetic enhancer sequence comprising a sequence that is unique or has more than one or two base pair mismatches with respect the sequence of a host genome;a promoter element; anda gene of interest;wherein the zinc finger protein assembly comprises:a promoter element;zinc finger array comprising a plurality7of zinc finger proteins engineered to bind specifically to the synthetic enhancer; anda transcriptional activator coupled to the zinc finger array; (ii) allowing integration of the synthetic enhancer assembly into the host genome;(iii) allowing expression of the zinc finger array and transcriptional activator, wherein the expressed zinc finger array allows the transcriptional activator to initiate transcription of the gene of interest, but does not alterPatent Application 97157.00116transcription of non-gene of interest genes of the host genome.

30. The method of Claim 29, wherein the synthetic transcription factor is introduced into the host cell via viral delivery.

31. The method of Claim 30, wherein the viral delivery comprises use of a lentiviral vector.

32. The method of Claim 29, 30, or 31, wherein the gene of interest encodes a therapeutic protein.

33. The method of Claim 32, wherein the therapeutic protein comprises a chimeric antigen receptor, a cytokine, an antibody, or a protein for which the expression and / or function of a corresponding endogenous protein of the host is compromised.

32. The method of any one of Claims 29 to 33, wherein modulation of expression comprises introducing or overexpressing the gene of interest.

32. The method of any one of Claims 29 to 33, wherein modulation of expression comprises reducing or eliminating expression of another host protein due to transcription of the gene of interest.

33. A zinc finger protein array comprising a plurality of zinc fingers separated by at least one linker element, the zinc finger array configured to bind with specificity’ to a nucleic acid sequence of SEQ ID NO: 24.

34. A zinc finger protein array comprising a plurality' of zinc fingers separated by at least one linker element, the zinc finger array configured to bind with specificity to a nucleic acid sequence selected from one or more of SEQ ID NOs 1-23, 25-47, or 74-96.

35. The zinc finger array of Claim 33 or 34, wherein the individual ZFPs within the array have at least 85% sequence identity to one or more of the amino acids of SEQ ID NOs 347-622.Patent Application 97157.0011636. The zinc finger array of Claim 33 or 34, wherein the individual ZFPs within the array have at least 90% sequence identity to one or more of the amino acids of SEQ ID NOs 347-622.

37. The zinc finger array of Claim 33 or 34, wherein the individual ZFPs within the array have at least 95% sequence identity to one or more of the amino acids of SEQ ID NOs 347-622.

38. The zinc finger array of Claim 33 or 34, wherein the individual ZFPs within the array are selected from one or more of the amino acids of SEQ ID NOs 347-622.

39. The zinc finger array of any one of Claims 33 to 38, wherein the ZFP array comprises an amino acid sequence having at least 85% sequence identity to one or more of the amino acids of SEQ ID NOs 244-289.

40. The zinc finger array of any one of Claims 33 to 39, wherein the ZFP array comprises an amino acid sequence having at least 90% sequence identity to one or more of the amino acids of SEQ ID NOs 244-289.

41. The zinc finger array of any one of Claims 33 to 40. wherein the ZFP array comprises an amino acid sequence having at least 95% sequence identity to one or more of the amino acids of SEQ ID NOs 244-289.

42. The zinc finger array of any one of Claims 33 to 41. wherein the ZFP array comprises an amino acid sequence of one or more of SEQ ID NOs 244-289.

43. The zinc finger array of any one of Claims 33 to 42, wherein the ZFP array is encoded by a nucleic acid having at least 85%, at least 90%, or at least 95% sequence identity to one or more of the nucleic acids of SEQ ID NOs 98-143.

44. The zinc finger array of any one of Claims 33 to 43, wherein the ZFP array is encoded by a nucleic acid selected from the group consisting of one or more of the nucleic acids of SEQ ID NOs 98-143.Patent Application 97157.0011645. The zinc finger array of any one of Claims 33 to 44, further comprising a mini-VPR transcriptional activator is coupled to the array and being encoded by a nucleic acid having at least 85%, at least 90%, or at least 95% sequence identity to one or more of the nucleic acids of SEQ ID NO 146 or 147.

46. The zinc finger array of Claim 45, wherein the mini-VPR transcriptional activator is encoded by SEQ ID NO 146 or 147.

47. The zinc finger array of Claim 45 or 46, wherein the ZFP array and mini-VPR transcriptional activator are encoded by a nucleic acid having at least 85%, at least 90%, or at least 95% sequence identity to one or more of the nucleic acids of SEQ ID NO 150-195.

48. The zinc finger array of Claim 45, 46, or 47, wherein the ZFP array and mini-VPR transcriptional activator is encoded one or more of the nucleic acids of SEQ ID NO 150-195.

49. The zinc finger array of Claim 45, 46, or 47, further comprising backbone structures and wherein the ZFP array and mini-VPR transcriptional activator with backbone structure are encoded by a nucleic acid having at least 85%, at least 90%, or at least 95% sequence identity7to one or more of the nucleic acids of SEQ ID NO 196-242.

50. The zinc finger array of any one of Claims 45 to 49, wherein further comprising backbone structures and wherein the ZFP array and mini-VPR transcriptional activator with backbone structure are encoded by one or more of the nucleic acids of SEQ ID NO 196-242.

51. The zinc finger array of any one of Claims 33 to 50, wherein the array has an amino acid sequence having at least 85%, at least 90%, or at least 95 sequence identity to one or more of the amino acids of SEQ ID NOs 623-668.

52. The zinc finger array of any one of Claims 33 to 51, wherein the array comprises an amino acid sequence of one or more of the amino acids of SEQ ID NOs 623-668.

53. The synthetic transcription factor of any one of Claims 1 to 11 , wherein the synthetic enhancer element further comprises a promoter and a gene of interest.Patent Application 97157.0011654. The synthetic transcription factor of any one of Claims 1 to 11 , wherein the synthetic enhancer element comprises a nucleic acid sequence having at least 85% sequence identity to one or more of the nucleic acid sequences of SEQ ID NOs 50-73.

55. The synthetic transcription factor of any one of Claims 1 to 11 , wherein the synthetic enhancer element comprises a nucleic acid sequence having at least 90% sequence identity to one or more of the nucleic acid sequences of SEQ ID NOs 50-73.

56. The synthetic transcription factor of any one of Claims 1 to 11 , wherein the synthetic enhancer element comprises a nucleic acid sequence having at least 95% sequence identity to one or more of the nucleic acid sequences of SEQ ID NOs 50-73.

57. The synthetic transcription factor of any one of Claims 1 to 11 , wherein the synthetic enhancer element comprises a nucleic acid sequence of one or more of the nucleic acid sequences of SEQ ID NOs 50-73.

57. The synthetic transcription factor of any one of Claims 1 to 11, wherein the spacers have different sequences from one another.

58. A synthetic drug switch comprising:at least one drug interaction domain;at least one transcriptional effector domain; anda DNA binding protein;wherein the DNA binding domain comprises a zinc finger protein, and wherein the zinc finger protein binds a target sequence and allows the transcriptional effector to modulate transcription of a target gene when a drug that interacts with the at least one drug interaction domain is present.

59. The synthetic drug switch of Claim 58, wherein the drug interaction domain interacts with tamoxifen or a metabolite thereof.

60. The synthetic drug switch Claim 62, wherein the tamoxifen metabolite comprises 4-OHT.Patent Application 97157.0011661. The synthetic drug switch of any one of Claims 58 to 60, wherein the drug interaction domain comprises a hormone receptor.

62. The synthetic drug switch of Claim 61, wherein the hormone receptor is selected from an estrogen receptor, a glucocorticoid receptor, and a progesterone receptor.

63. The synthetic drug switch of Claim 62, wherein the drug interaction domain comprises an estrogen receptor, wherein the estrogen receptor comprises the ERT2 estrogen receptor.

64. The synthetic drug switch of Claim 63, wherein the ERT2 estrogen receptor is encoded by a nucleic acid having at least 85% sequence identity to SEQ ID NO: 758, 775, or 792.

65. The synthetic drug switch of Claim 64, wherein the ERT2 estrogen receptor is encoded by the nucleic acid of SEQ ID NO: 758, 775, or 792.

66. The synthetic drug switch of Claim 64 or 65, wherein the ERT2 estrogen receptor comprises an amino acid sequence having at least 85% sequence identity to SEQ ID NO: 767, 784, or 801.

66. The synthetic drug switch of Claim 66, wherein the ERT2 estrogen receptor comprises an amino acid sequence of SEQ ID NO:

767. 784, or 801.

67. The synthetic drug switch of any one of Claims 58 to 66, wherein the drug interaction domain further comprises a p65 protein.

68. The synthetic drug switch of Claim 67, wherein the p65 protein is encoded by a nucleic acid having at least 85% sequence identity to SEQ ID NO: 756, 773, or 790.

69. The synthetic drug switch of Claim 67 or 68, wherein the p65 protein is encoded by the nucleic acid of SEQ ID NO: 756, 773. or 790.Patent Application 97157.0011670. The synthetic drug switch of Claim 67, wherein the p65 protein comprises an amino acid sequence having at least 85% sequence identity’ to SEQ ID NO:

765. 782, or 799.

71. The synthetic drug switch of Claim 70, wherein the p65 protein comprises an amino acid sequence of SEQ ID NO: 765, 782, or 799.

72. The synthetic drug switch of any one of Claims 58 to 71, wherein the synthetic drug switch comprises an amino acid having at least 85% sequence identity to one or more of SEQ ID NO: 770, 787, and 804.

73. The synthetic drug switch of Claim 58, wherein the drug interaction domain comprises a cereblon protein.

74. The synthetic drug switch of Claim 73, wherein the cereblon protein is engineered to (i) not express a DDB1 subdomain or (ii) express a DDB1 subdomain that does not allow the cereblon to interact with an E3 ubiquitin ligase complex.

75. The synthetic drug switch of Claim 73 or 74, wherein the cereblon protein is encoded by a nucleic acid having at least 85% sequence identity to SEQ ID NO: 684, 708, or 732.

76. The synthetic drug switch of Claim 75, wherein the cereblon protein is encoded by SEQ ID NO: 684, 708, or 732.

77. The synthetic drug switch of any one of Claims 73 to 76, wherein the cereblon protein comprises an amino acid sequence having at least 85% sequence identity to SEQ ID NO: 697, 721, or 745.

78. The synthetic drug switch of Claim 77, wherein the cereblon protein comprises an amino acid sequence of SEQ ID NO: 697, 721, or 745.

79. The synthetic drug switch of any one of Claims 73 to 78, wherein the cereblon protein is fused to the transcriptional effector domain.Patent Application 97157.0011680. The synthetic drug switch of Claim 79, wherein the transcriptional effector domain comprises a mini-VPR domain and the cereblon-transcriptional effector domain fusion is encoded by a nucleic acid sequence having at least 85% sequence identity to SEQ ID NO: 693, 717, or 741.

81. The synthetic drug switch of Claim 80, wherein the cereblon-transcriptional effector domain fusion is encoded by SEQ ID NO: 693, 717, or 741.

82. The synthetic drug switch of Claim 79 or 80, wherein the transcriptional effector domain comprises a mini-VPR domain and the cereblon-transcriptional effector domain fusion has an amino acid sequence having at least 85% sequence identity to SEQ ID NO: 700, 724, or 748.

83. The synthetic drug switch of Claim 73, wherein drug interaction domain further comprises a CRIMP domain.

84. The synthetic drug switch of Claim 83, wherein the CRIMP domain is encoded by a nucleic acid having at least 85% sequence identity to one or more of SEQ ID NO: 688, 712, or 736.

85. The synthetic drug switch of Claim 84, wherein the CRIMP domain is encoded by a nucleic acid of SEQ ID NO: 688, 712, or 736.

86. The synthetic drug switch of Claim 83, wherein the CRIMP domain has an amino acid sequence having at least 85% sequence identity to one or more of SEQ ID NO: 702, 726, or 750.

87. The synthetic drug switch of Claim 86, wherein the CRIMP domain has an amino acid sequence of SEQ ID NO: 702, 726, or 750.

88. The synthetic drug switch of any one of Claims 83 to 87, wherein the CRIMP domain is fused to the ZFP.

89. The synthetic drug switch of Claim 88, wherein the CRIMP domain-ZFP fusion hasPatent Application 97157.00116a nucleic acid sequence having at least 85% sequence identity to one or more of SEQ ID NO: 694, 718. or 742.

90. The synthetic drug switch of Claim 88, wherein the CRIMP domain-ZFP fusion has a nucleic acid sequence of SEQ ID NO: 694, 718, or 742.

91. The synthetic drug switch of any one of Claims 88 to 90, wherein the CRIMP domain-ZFP fusion has an amino acid sequence having at least 85% sequence identity to one or more of SEQ ID NO: 705, 729, or 753.

92. The synthetic drug switch of any one of Claims 88 to 91, wherein the CRIMP domain-ZFP fusion has an amino acid sequence of SEQ ID NO:

705. 729, or 753.

93. The synthetic drug switch of any one of Claims 73 to 92, wherein the drug switch is configured to activate transcription in response to an immunomodulatory (IMiD) drug.

94. The synthetic drug switch of Claim 94, wherein the (IMiD) drug is selected from lenalidomide, pomalidomide, thalidomide, or combinations thereof.