Transactivators and methods and uses thereof

JP2024524696A5Pending Publication Date: 2025-07-22THE GOVERNING COUNCIL OF THE UNIV OF TORONTO
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2024501992
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2021-07-14
Filing Date
2022-07-14
Publication Date
2025-07-22

AI Technical Summary

Technical Problem

Current methods for understanding the transcriptional effects of transcription factors and chromatin-associated factors are limited by the inability to causally associate genomic binding sites with transcriptional events and are not scalable for identifying novel factors, particularly in mammalian systems.

Method used

A platform is developed to systematically identify and characterize the transcriptional regulatory potential of human proteins by screening over 13,000 proteins in a pooled format, using a chemically induced dimerization system with CRISPR-Cas and chemical inhibitors to delineate cofactor specificity of transcriptional activators.

Benefits of technology

This approach identifies hundreds of potent transcriptional activators, including novel chromatin-associated proteins, and enables the generation of 'superactivators' that robustly upregulate genes even in highly condensed genomic regions, providing a comprehensive understanding of transcriptional activation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

A heterologous transcriptional activator comprising a DNA targeting domain, preferably a catalytically inactive DNA targeting protein, such as a CRISPR-Cas protein, and an effector domain comprising at least one transactivation domain or a functional variant thereof as described herein. Also provided herein are expression constructs, vectors, and cells encoding or expressing said transcriptional activator, as well as systems and methods for transcriptional activation of target genes, and compositions, kits, and reagents for use in making and using the same.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims the benefit of priority to U.S. Patent Application No. 63 / 221,611, filed July 14, 2021, the disclosure of which is incorporated herein by reference in its entirety.

[0002] The present disclosure relates to reagents and methods for transcriptional activation, and in particular to the use of heterologous transactivator domains in transcriptional activators for targeted transcriptional activation.

[0003] introduction Transcription of protein-coding genes is orchestrated by the coordinated interplay of transcription factors (TFs) that bind DNA in a sequence-specific manner, the RNA polymerase II machinery that initiates transcription from promoters, and diverse chromatin-associated factors and complexes that regulate chromatin structure and act as bridges between TFs and RNA pol II (Cramer, 2019). The human genome encodes thousands of proteins involved in various stages of transcriptional regulation, and the ready availability of methods such as ChIP-seq has revealed genomic binding sites for hundreds of factors in diverse conditions (The ENCODE Project Consortium et al., 2020). At the same time, systematic studies have characterized or inferred the DNA-binding specificity of about three-quarters of human TFs (Badis et al., 2009; Jolma et al., 2013; Lambert et al., 2018; Najafabadi et al., 2015; Weirauch et al., 2014). Similarly, interaction proteomics approaches have revealed many chromatin-associated proteins and characterized the composition of transcriptional regulatory complexes in human cells ( Gao et al., 2012 ; Huttlin et al., 2020 ; Lambert et al., 2019 ; Li et al., 2015 ; Marcon et al., 2014 ; Mashtalir et al., 2018 ).

[0004] However, whether and how TFs and chromatin-associated factors promote transcriptional activation or repression (or modulate chromatin states by other means) has remained largely unclear due to the limited causal insight afforded by these methods. For example, the vast majority of genomic binding sites observed in ChIP-seq experiments are not causally associated with transcriptional events; i.e., knockdown or knockout of a given transcriptional regulator does not affect transcription of many genes to which the regulator binds. Meanwhile, sequence-based annotation of transcriptional regulators has been challenging, as most transcriptional effector functions are encoded by degenerate linear motifs rather than folded, conserved protein domains (Arnold et al., 2018; Erijman et al., 2020; Sigler, 1988; Staller et al., 2021).

[0005] Artificial recruitment, also known as activator bypass, is a powerful method to characterize the transcriptional effects of diverse proteins in a defined context (Ptashne and Gann, 1997; Sadowski et al., 1988). In this approach, a protein or a fragment thereof is ectopically recruited to a reporter gene by fusing the protein to the DNA-binding domain of a well-characterized TF, such as Gal4 or TetR. A defined context mitigates the challenges posed by endogenous gene regulation, where multiple factors bind to regulatory elements in a coordinated manner, hindering causal inference. Artificial recruitment has traditionally been used to identify transcriptional activators or transactivation domains (TADs) in individual transcriptional regulators (Ptashne and Gann, 1997). However, recent studies have characterized the transcriptional effects of large collections of regulators in Drosophila or yeast by individually linking them to reporter genes (Keung et al., 2014; Stampfel et al., 2015). Due to the limited scalability of array formats, these studies focused on known regulators rather than potentially novel factors. Moreover, classical model organisms lack several regulatory mechanisms and layers important for gene expression in mammals, such as enhancers (yeast) or DNA methylation (yeast and Drosophila). More recently, Tycko et al. conducted an unbiased pooled screening strategy to characterize the transcriptional activation potential of annotated human protein domains (Tycko et al., 2020). This study highlighted the value of an unbiased approach in identifying novel transcriptional regulators, e.g., the unexpected role of variant KRAB domains in transcriptional activation instead of repression. However, because most transcriptional activation domains are encoded by disordered regions (Arnold et al., 2018; Dyson and Wright, 2005), a domain-focused screen is likely to miss a significant fraction of transactivators.

[0006] Thus, despite increasing knowledge of the composition of transcription factor complexes and their genomic binding patterns, a complete understanding of the downstream effects elicited by TFs and diverse chromatin-associated factors is lacking. Summary of the Invention

[0007] As described herein, we established a platform to systematically identify and characterize the transcriptional regulatory potential of human proteins in an unbiased manner. By screening over 13,000 proteins in a pooled format, hundreds of potent activators were identified, many of which were previously poorly annotated. Transactivation domains were systematically revealed among the hits, including some that do not attach to the standard "acidic blob" model of activation domains. Furthermore, we combined interaction proteomics with chemical inhibitors to delineate the cofactor specificity of both novel and known transcriptional activators, highlighting how highly related TFs with virtually identical DNA-binding specificities can activate transcription through distinct cofactor complexes.

[0008] We describe here the first systematic screen for transcriptional activators in human cells. As expected, hundreds of transcriptional activators were identified that were enriched for sequence-specific transcription factors and other chromatin-associated proteins.

[0009] The results presented herein suggest that only a very limited number of TFs are potent transcriptional activators. This makes sense in relation to the known DNA-binding specificity and chromatin occupancy of human TFs (Jolma et al., 2013; The ENCODE Project Consortium et al., 2020; Yan et al., 2013). Most TFs recognize short (~6-10 bp) motifs and bind to thousands of sites throughout the genome. However, most binding events are not associated with transcriptional outcomes. Potent transcriptional activation domains bound to the limited sequence specificity of DNA-binding domains would likely cause spurious activation of a wide range of genes, interfering with cellular compatibility. Many of the strongest activators identified herein were not DNA-binding transcription factors but other chromatin-associated proteins.

[0010] The discovery of novel and highly potent human transactivation domains also has therapeutic implications. Components derived from viral proteins, such as the VP16 transactivation domain or the tripartite VPR activator, can induce immune responses in vivo and cause side effects in the clinic. Designing synthetic transcription regulators from fully human components is expected to be advantageous in therapeutic applications (Israni et al., 2021). Moreover, as shown herein, by combining activation domains from multiple different human proteins, it is possible to generate "superactivators" that can robustly upregulate genes even in highly compacted regions of the genome.

[0011] One embodiment is a heterologous transcriptional activator comprising: a DNA targeting domain, optionally an enzymatically inactive CRISPR-CAS protein, a zinc finger DNA binding domain, a tet-repressor, or a transcription activator-like effector (TALE) DNA binding domain; an effector domain comprising at least one transactivation domain (TAD) selected from a TAD listed in any one of Tables 1-6, optionally in Table 2 or Table 6, or a functional variant thereof, or at least two TADs selected from a TAD listed in any one of Tables 1-6, optionally in Table 1 or Table 3, or any functional variant thereof, preferably at least one TAD selected from a TAD listed in Table 4 or Table 5 or Table 6, or a functional variant thereof; It comprises a heterologous transcriptional activator in which the DNA targeting domain and the effector domain are operably linked.

[0012] One embodiment includes an isolated nucleic acid encoding an effector domain described herein.

[0013] One aspect includes an isolated nucleic acid encoding a heterologous transcriptional activator described herein.

[0014] One embodiment includes an expression construct comprising a nucleic acid described herein operably linked to one or more promoters and one or more transcription termination sites.

[0015] One aspect includes a vector comprising a nucleic acid or expression construct described herein, optionally wherein the vector is an adenoviral or lentiviral vector.

[0016] One aspect includes a cell comprising a transcriptional activator, nucleic acid, expression construct, or vector described herein.

[0017] One embodiment includes a transcriptional activation system comprising a heterologous transcriptional activator as described herein, wherein the DNA targeting domain comprises a CRISPR-Cas protein and at least one gRNA.

[0018] One embodiment includes a method of activating transcription of a target gene in a cell, the method comprising: a) introducing into a cell a transcriptional activator, nucleic acid, expression construct, or vector described herein; and b) culturing the cell under suitable conditions such that the effector domain activates transcription of the target gene.

[0019] One embodiment includes a screening method comprising: a) introducing into a plurality of cells a transcriptional activator, one or more nucleic acids, one or more expression constructs, or one or more vectors described herein, wherein the DNA targeting domain comprises a CRISPR-Cas protein and a plurality of gRNAs; or introducing into a population of cells described herein a plurality of gRNAs, wherein the DNA targeting domain comprises a CRISPR-Cas protein; b) culturing the plurality of cells such that the one or more gRNAs bind to the CRISPR-Cas protein and guide the transcriptional activator to the CRISPR target site where the effector domain activates transcription of the target gene; c) optionally treating with an amount of a test drug or toxin; d) optionally culturing the plurality of cells for a period of time to allow for shedding or enrichment of the gRNAs; and e) recovering the plurality of cells or a subset thereof.

[0020] One aspect includes a composition comprising a transcriptional activator, a nucleic acid, an expression construct, a vector, or a cell described herein.

[0021] One embodiment includes a kit comprising a vial and a heterologous transcriptional activator, nucleic acid, expression construct, vector, cell, or composition described herein, and optionally one or more of an inducer, gRNA, or gRNA expression construct.

[0022] The preceding paragraphs are provided merely as examples and are not intended to limit the scope of the present disclosure and the appended claims. Additional objects and advantages associated with the compositions and methods of the present disclosure will be understood by those skilled in the art in light of the claims, descriptions, and examples. For example, the various aspects and embodiments of the present disclosure may be utilized in many combinations, all of which are expressly contemplated by this specification. These additional advantageous objects and embodiments are expressly included within the scope of the present disclosure. Publications and other materials used herein to clarify the background of the present disclosure, particularly to provide additional details regarding implementation in some cases, are incorporated by reference and are listed in the attached reference section for convenience.

[0023] Features and advantages of the present disclosure will become more apparent from the following detailed description taken in conjunction with the accompanying drawings illustrating exemplary embodiments of the disclosure. [Brief description of the drawings]

[0024] [Figure 1A] Figure 1 shows screening of the pooled ORFeome for transcriptional activators.Figure 2 shows a schematic of a chemically induced dimerization system for characterizing transcriptional activators in human cells. [Figure 1B] Screening of the pooled ORFeome for transcriptional activators. The indicated constructs were transfected into HEK293T reporter cells and the percentage of high GFP cells is shown upon treatment with abscisic acid or DMSO for 48 hours. [Figure 1C] 1 shows pooled ORFeome screening for transcriptional activators. Schematic of pooled ORFeome screening for transcriptional activators. [Figure 1D] Pooled ORFeome screening for transcriptional activators. Enrichment of high GFP cells in the pooled ORFeome screening after 48 hours of ABA treatment. [Figure 1E]Screening of the pooled ORFeome for transcriptional activators. Enrichment of ORFs in the high GFP pool compared to the unsorted ORFeome. [Figure 1F] Screening of the pooled ORFeome for transcriptional activators. Gene Ontology category enrichment in positive screening hits. [Figure 1G] Screening of the pooled ORFeome for transcriptional activators. Enrichment of InterPro domains in positive screening hits. [Figure 1H] Screening of the pooled ORFeome for transcriptional activators. Enrichment of CORUM complexes in screening hits. [Figure 2A] Transcriptional activity of transcription factor families is shown. Transcriptional regulators were individually tested for reporter activation in an arrayed format. DNA binding specificity (shown as sequence logo) is derived from CisB (Weirauch et al., 2014). Asterisks indicate statistically significant activators (FDR<0.05). Homeoboc family proteins. Statistical significance was calculated with an unpaired two-tailed t-test assuming equal variances and corrected for multiple hypotheses using the false discovery rate (FDR) approach of Benjamini, Krieger, and Yekutieli (Benjamini et al., 2006). [Figure 2B] Transcriptional activity of transcription factor families is shown. Transcriptional regulators were individually tested for reporter activation in an arrayed format. DNA binding specificity (shown as sequence logo) is derived from CisB (Weirauch et al., 2014). Asterisks indicate statistically significant activators (FDR<0.05). Forkhead Box proteins. Statistical significance was calculated with an independent two-tailed t-test assuming equal variances and corrected for multiple hypotheses using the false discovery rate (FDR) approach of Benjamini, Krieger, and Yekutieli (Benjamini et al., 2006). [Figure 2C]Transcriptional activity of transcription factor families is shown. Transcriptional regulators were individually tested for reporter activation in an arrayed format. DNA binding specificity (shown as sequence logos) is derived from CisB (Weirauch et al., 2014). Asterisks indicate statistically significant activators (FDR<0.05). Kruppel-like factors. Statistical significance was calculated with an independent two-tailed t-test assuming equal variances and corrected for multiple hypotheses using the false discovery rate (FDR) approach of Benjamini, Krieger, and Yekutieli (Benjamini et al., 2006). [Figure 2D] Transcriptional activity of transcription factor families is shown. Transcriptional regulators were individually tested for reporter activation in an arrayed format. DNA binding specificity (shown as sequence logo) is derived from CisB (Weirauch et al., 2014). Asterisks indicate statistically significant activators (FDR<0.05). SRY-related HMG-box (SOX) proteins. Statistical significance was calculated by independent two-tailed t-test assuming equal variances and corrected for multiple hypotheses using the false discovery rate (FDR) approach of Benjamini, Krieger, and Yekutieli (Benjamini et al., 2006). [Figure 2E] Transcriptional activity of transcription factor families is shown. Transcriptional regulators were individually tested for reporter activation in an arrayed format. DNA binding specificity (shown as sequence logo) is derived from CisB (Weirauch et al., 2014). Asterisks indicate statistically significant activators (FDR<0.05). Statistical significance was calculated by independent two-tailed t-test assuming equal variances and corrected for multiple hypotheses using the false discovery rate (FDR) approach of Benjamini, Krieger, and Yekutieli (Benjamini et al., 2006). Polycomb group RING Finger (PCGF) proteins. The composition of the canonical (cPRC1) and non-canonical (ncPRC1) complexes is shown on the right. [Figure 2F]Transcriptional activity of transcription factor families is shown. Transcriptional regulators were individually tested for reporter activation in an arrayed format. DNA binding specificity (shown as sequence logo) is derived from CisB (Weirauch et al., 2014). Asterisks indicate statistically significant activators (FDR<0.05). Spy1 / RINGO family proteins. Statistical significance was calculated with an independent two-tailed t-test assuming equal variances and corrected for multiple hypotheses using the false discovery rate (FDR) approach of Benjamini, Krieger, and Yekutieli (Benjamini et al., 2006). [Figure 3A] Systematic discovery of transactivation domains in human proteins by TAD-seq. Schematic of the TAD-seq pooled assay. [Figure 3B] Systematic discovery of transactivation domains in human proteins by TAD-seq. Examples of known transactivation domains. Domain organization is shown on top. TAD-seq plots show fold enrichment of RNAseq reads in high or mid-GFP populations. Each circle indicates the midpoint (30th amino acid) of a 60-aa tile. Black circles indicate statistically significant hits. Grey boxes indicate previously described transactivation domains. [Figure 3C] Systematic discovery of transactivation domains in human proteins by TAD-seq. Examples of novel transactivation domains. Labeling is as in panel B. [Figure 3D] Systematic discovery of transactivation domains in human proteins by TAD-seq. Location of the transactivation domain of HOXA2. [Figure 3E] Systematic discovery of transactivation domains in human proteins by TAD-seq. Sequences of activation fragments identified in HOXA2. Activation fragments are in bold. Regions common to all three fragments enriched in the medium GFP population are shown as overlapping activation fragments. The position of the antennapedia-like hexapeptide sequence is shown as a hexapeptide. [Figure 3F]Systematic discovery of transactivation domains in human proteins by TAD-seq. Location of the transactivation domain in YAF2. [Figure 3G] Systematic discovery of transactivation domains in human proteins by TAD-seq. Crystal structures of the RING1B RAWUL domain bound to the YAF2_RYBP domain of RYBP (PDB 3IXS) and the CBX_C domain of CBX7 (PDB 3GS2). [Figure 3H] Systematic discovery of transactivation domains in human proteins by TAD-seq. YAF2_RYBP and CBX_C domains from the indicated proteins were individually tested for transcriptional activity. Asterisks indicate statistically significant activators (FDR<0.05). Statistical significance was calculated using an unpaired two-tailed t-test assuming equal variances and corrected for multiple hypotheses using the false discovery rate (FDR) approach of Benjamini, Krieger, and Yekutieli (Benjamini et al., 2006). Right, in vitro affinity of the domains for RING1B (Wang et al., 2008, 2010). Statistical significance was calculated using an unpaired two-tailed t-test. [Figure 4A] Cofactor specificity of transcriptional activators is shown. Proximity partners of the indicated transcriptional regulators were identified with BioID2. Enrichment of selected cofactor complexes is shown as a heatmap. The average spectral counts (n+1) were normalized to the background spectral counts (n+1) of EGFP and Nanoluc baits. [Figure 4B] Cofactor specificity of transcriptional activators. Interaction patterns of activating forkhead transcription factors based on AP-MS studies from (Li et al., 2015). [Figure 4C]Cofactor specificity of transcriptional activators is shown. Left, effect of p300 / CBP inhibition by A-485 on the activity of 83 transcriptional regulators. Known p300 interactors are shown as solid circles and known NuA4 interactors as open circles. Right, p300 interactors are significantly more affected than other transcriptional regulators by A-485 treatment. Statistical significance was calculated by one-way ANOVA with Dunnett's multiple testing correction. [Figure 4D] Cofactor specificity of transcriptional activators is shown. Left, effect of BET bromodomain protein inhibition by JQ1 on the activity of 83 transcriptional regulators. Known p300 interactors are shown as solid circles and known NuA4 interactors are shown as open circles. Right, p300 interactors are significantly more affected than other transcriptional regulators by A-485 treatment. Statistical significance was calculated by one-way ANOVA with Dunnett's multiple testing correction. [Figure 4E] Cofactor specificity of transcriptional activators is shown. Clustering of translational activators is based on their sensitivity to various inhibitors. Clustering was performed using Euclidean distance with average linkage. Clusters enriched for the intersection of known CBP / p300 interactors (black circles) and NuA4 interactors are shown. [Figure 5A] Figure 1. SRF-C3orf62 fusion generates a potent p300-dependent transcriptional activator that drives expression of SRF / MRTF target genes. Schematic diagram of JAZF1-SUZ12 fusion found in low-grade endometrial stromal sarcoma, showing transactivation domains identified by TAD-seq. [Figure 5B] Figure 1. SRF-C3orf62 fusions generate potent p300-dependent transcriptional activators that drive expression of SRF / MRTF target genes. Schematic of SRF-C3orf62 fusions found in myofibromas / myopericytomas, showing transactivation domains identified by TAD-seq. [Figure 5C]We show that SRF-C3orf62 fusions generate potent p300-dependent transcriptional activators that promote expression of SRF / MRTF target genes. JAZF1, SUZ12, and JAZF1-SUZ12 proximal interactors were identified in BioID2. JAZF1-SUZ12 interacts with both PRC2 and NuA4 components to generate a supercomplex. [Figure 5D] We show that SRF-C3orf62 fusions generate potent p300-dependent transcriptional activators that promote expression of SRF / MRTF target genes. SRF, C3orf62, C3orf62-Cterm, and SRF-C3orf62 proximal interactors were identified by BioID. SRF-C3orf62 fusions tightly interact with CBP and p300. [Figure 5E] We show that SRF-C3orf62 fusions generate potent p300-dependent transcriptional activators that drive expression of SRF / MRTF target genes. SRF-C3orf62 is a potent transcriptional activator. The indicated PYL1 fusions were individually tested for activation of a genomically integrated reporter. [Figure 5F] We show that SRF-C3orf62 fusions generate potent p300-dependent transcriptional activators that drive expression of SRF / MRTF target genes. SRF-C3orf62 activates expression of a serum response element reporter in the absence of cofactors. The indicated constructs C-terminally tagged with 3xFLAG-V5 were co-transfected with an SRE-firefly luciferase reporter and a constitutive Nanoluc reporter into NIH3T3 cells. Relative luciferase activity was measured to assess the activity of each construct. [Figure 5G]We show that SRF-C3orf62 fusions generate potent p300-dependent transcriptional activators that drive expression of SRF / MRTF target genes. NIH3T3 cells stably expressing doxycycline-inducible SRF-C3orf62-GFP or Nanoluc-GFP were treated with doxycycline for 46 h, the last 22 h in low serum conditions (0.5% FCS). Gene expression patterns were analyzed by RNA-seq. Significantly upregulated (filled circles, top) and downregulated (open circles, bottom) genes (absolute log2 fold change >1, FDR <0.05). Well-characterized targets of SRF / MRTF (top, labeled circles) and SRF / TCF (bottom, labeled circles) are shown. [Figure 5H] We show that SRF-C3orf62 fusions generate potent p300-dependent transcriptional activators that drive expression of SRF / MRTF target genes. Gene set enrichment analysis of SRF-C3orf62-GFP expressing cells compared to Nanoluc-GFP expressing cells. [Figure 6A] Screening of the ORFeome for transcriptional activators. Distribution of sequencing reads across the ORFeome in pooled plasmid DNA and infected cells. [Figure 6B] Screening of the ORFeome for transcriptional activators. ORF size distribution in plasmid pools and infected cells. [Figure 6C] Screening of the ORFeome for transcriptional activators. Transcriptional activity is dependent on abscisic acid treatment. ORFeome-PYL1 infected cells were treated with 100 μM ABA and the fraction of high GFP cells was measured over time by flow cytometry. [Figure 6D] Screening of the ORFeome for transcriptional activators. Effect of ABA concentration on transcriptional activity. Reporter cells transfected with the indicated constructs were treated with increasing amounts of ABA for 48 h. [Figure 6E] Screening of the ORFeome for transcriptional activators. No high GFP population is observed in ABA-treated cells that do not express the ORFeome-PYL1 library. [Figure 6F] Screening of the ORFeome for transcriptional activators. Enrichment of interaction hubs in activation screen hits. [Figure 6G] Screening of the ORFeome for transcriptional activators. Enrichment of yeast two-hybrid autoactivators among activator screen hits. [Figure 6H] Screening of the ORFeome for transcriptional activators. Independent validation of transcriptional activators identified in the activation screen. The indicated constructs were transfected into reporter cell lines and the high GFP cell fraction was measured by flow cytometry after 48 h treatment with ABA. Asterisks indicate a statistically significant ABA-dependent increase in the high GFP population (FDR<5%). Statistical significance was calculated with an unpaired two-tailed t-test assuming equal variances and corrected for multiple hypotheses using the false discovery rate (FDR) approach of Benjamini, Krieger, and Yekutieli (Benjamini et al., 2006). [Figure 7A] Transcriptional activity of transcription factor and chromatin-associated protein families is shown. Transcriptional regulators were individually tested for reporter activation in an arrayed format. DNA binding specificity (shown as sequence logo) is derived from CisBP (Weirauch et al., 2014). Asterisks indicate statistically significant activators (FDR<5%). Statistical significance was calculated with an unpaired two-tailed t-test assuming equal variances and corrected for multiple hypotheses using the false discovery rate (FDR) approach of Benjamini, Krieger, and Yekutieli (Benjamini et al., 2006). ETS family TFs. [Figure 7B]Transcriptional activity of transcription factor and chromatin-associated protein families is shown. Transcriptional regulators were individually tested for reporter activation in an arrayed format. DNA binding specificity (shown as sequence logo) is derived from CisBP (Weirauch et al., 2014). Asterisks indicate statistically significant activators (FDR<5%). Statistical significance was calculated with an independent two-tailed t-test assuming equal variances and corrected for multiple hypotheses using the false discovery rate (FDR) approach of Benjamini, Krieger, and Yekutieli (Benjamini et al., 2006). Atonal-related bHLH factors. [Figure 7C] Transcriptional activity of transcription factor and chromatin-associated protein families is shown. Transcriptional regulators were individually tested for reporter activation in an arrayed format. DNA binding specificity (shown as sequence logo) is derived from CisBP (Weirauch et al., 2014). Asterisks indicate statistically significant activators (FDR<5%). Statistical significance was calculated with an independent two-tailed t-test assuming equal variances and corrected for multiple hypotheses using the false discovery rate (FDR) approach of Benjamini, Krieger, and Yekutieli (Benjamini et al., 2006). Myogenic factors. [Figure 7D] Transcriptional activity of transcription factor and chromatin-associated protein families is shown. Transcriptional regulators were individually tested for reporter activation in an arrayed format. DNA binding specificity (shown as sequence logo) is derived from CisBP (Weirauch et al., 2014). Asterisks indicate statistically significant activators (FDR<5%). Statistical significance was calculated with an independent two-tailed t-test assuming equal variances and corrected for multiple hypotheses using the false discovery rate (FDR) approach of Benjamini, Krieger, and Yekutieli (Benjamini et al., 2006). Twist / Hand bHLH factors. [Figure 7E]Transcriptional activity of transcription factor and chromatin-associated protein families is shown. Transcriptional regulators were individually tested for reporter activation in an arrayed format. DNA binding specificity (shown as sequence logo) is derived from CisBP (Weirauch et al., 2014). Asterisks indicate statistically significant activators (FDR<5%). Statistical significance was calculated with an unpaired two-tailed t-test assuming equal variances and corrected for multiple hypotheses using the false discovery rate (FDR) approach of Benjamini, Krieger, and Yekutieli (Benjamini et al., 2006). Casein kinase. [Figure 7F] Transcriptional activity of transcription factor and chromatin-associated protein families is shown. Transcriptional regulators were individually tested for reporter activation in an arrayed format. DNA binding specificity (shown as sequence logo) is derived from CisBP (Weirauch et al., 2014). Asterisks indicate statistically significant activators (FDR<5%). Statistical significance was calculated with an independent two-tailed t-test assuming equal variances and corrected for multiple hypotheses using the false discovery rate (FDR) approach of Benjamini, Krieger, and Yekutieli (Benjamini et al., 2006). Mediator components. [Figure 7G]Transcriptional activity of transcription factor and chromatin-associated protein families is shown. Transcriptional regulators were individually tested for reporter activation in an arrayed format. DNA binding specificity (shown as sequence logo) is derived from CisBP (Weirauch et al., 2014). Asterisks indicate statistically significant activators (FDR < 5%). Statistical significance was calculated with an independent two-tailed t-test assuming equal variances and corrected for multiple hypotheses using the false discovery rate (FDR) approach of Benjamini, Krieger, and Yekutieli (Benjamini et al., 2006). Transcription factors identified in the primary activation screen are enriched for factors that can reprogram human iPS cells (right) or mouse ES cells (left) when ectopically expressed (Ng et al., 2021; Theodorou et al., 2009). Statistical significance was calculated using Wilcoxon rank sum test (left) or Fisher's exact test (right). [Figure 8A] Figure 1 shows that differential activity of transcription factors is not explained by expression levels. Schematic of a transcription activation assay that measures both reporter gene expression and effector protein expression. [Figure 8B] Figure 1. Differential activity of transcription factors is not explained by expression levels. Transactivation of Kruppel-like factors (left) was compared to the expression levels of each factor measured by RFP fluorescence (right). Background RFP intensity is shown as a dashed line. [Figure 8C] Figure 2 shows that differential activity of transcription factors is not explained by expression levels. Forkhead TF activity. [Figure 8D] Figure 2 shows that differential activity of transcription factors is not explained by expression levels. Homeodomain TF activity. [Figure 8E] Figure 2 shows that differential activity of transcription factors is not explained by expression levels. SOX TF activity. [Figure 9A]We show that TAD-seq identifies transactivation domains. High and medium GFP populations were assessed by flow cytometry after mobilizing the 60aa fragment to the reporter with ABA for 48 hours. High and medium GFP cells were sorted by FACS and ORFs enriched in the pools were identified by next generation sequencing. [Figure 9B] Figure 1 shows that TAD-seq identifies transactivation domains. Amino acid enrichment and depletion in identified transactivator fragments was compared to inactive fragments in the library. Amino acids shown in bold were statistically significantly enriched or depleted. [Figure 9C] Figure 1 shows that TAD-seq identifies transactivation domains. Enrichment of predicted transactivation domains in active fragments. 9aaTADs were predicted with the 9aaTAD prediction tool (https: / / www.med.muni.cz / 9aaTAD / ) using the "moderately stringent pattern". Only 100% confident matches were considered in the analysis. The ADpred algorithm is described in (Erijman et al., 2020). Statistical significance was calculated using Fisher's exact test. [Figure 9D] Figure 1 shows that TAD-seq identifies transactivation domains. The predicted TADs in the active fragment are longer than those in the inactive fragment. Statistical significance was calculated using a two-tailed t-test assuming equal variances. [Figure 9E] Figure 2 shows that TAD-seq identifies transactivation domains. Enrichment of fragments in the high GFP pool relative to the medium GFP pool. Significant hits are shown as open circles. [Figure 9F]Figure 1 shows that TAD-seq identifies transactivation domains. Independent validation of transactivation fragments. The indicated TADs were fused to PYL1 and transfected into reporter cells in a sequenced format. Asterisks indicate statistically significant activators (FDR < 5%). Statistical significance was calculated by unpaired two-tailed t-tests assuming equal variances and corrected for multiple hypotheses using the false discovery rate (FDR) approach of Benjamini, Krieger, and Yekutieli (Benjamini et al., 2006). [Figure 10A] Systematic discovery of transactivation domains in human proteins by TAD-seq. Examples of known transactivation domains. Domain organization is shown at the top. TAD-seq plots show the fold enrichment of RNAseq reads in high GFP populations (filled circles) or medium GFP populations (open circles). Each circle indicates the midpoint (30th amino acid) of a 60-aa tile. Filled circles indicate statistically significant hits. Grey boxes indicate previously described transactivation domains. [Figure 10B] Systematic discovery of transactivation domains in human proteins by TAD-seq. Examples of novel transactivation domains. Labeling is as in panel A. [Figure 11A] The transactivation domains of SPDYE4 and YAF2 are shown. Alignment of five Spyl / RINGO family proteins: SPDYE4 (SEQ ID NO:142); SPDYE1 (SEQ ID NO:143); SPDYE7P (SEQ ID NO:144); SPDYE2 (SEQ ID NO:145); and SPDYC (SEQ ID NO:146). The four fragments screened by TAD-seq (SPDYE4-2 (SEQ ID NO:154); SPDYE4-3 (SEQ ID NO:90); SPDYE4-4 (SEQ ID NO:91); and SPDYE4-5 (SEQ ID NO:155)) are shown above the alignment, with the active fragments in bold. The dashed box indicates the inferred minimal activation domain. Note that the region of the minimal activation domain is not conserved in SPDYC, the only Spyl / RINGO family member that did not activate transcription. [Figure 11B]The transactivation domains of SPDYE4 and YAF2 are shown. Alignment of the YAF2_RYBP domains of YAF2 (SEQ ID NO: 147) and RYBP (SEQ ID NO: 148), and the CBX_C domains of CBX family proteins (SEQ ID NOs: 149-153). The positions of the two β-sheets are indicated by arrows. [Figure 12A] Figure 1 shows the interaction network of transcription activators. BioID2 network of transcription activators. Bait proteins (e.g., FAM90A1, SPDYE4, SS18L2, SOX7, etc.) are shown as light grey rectangles. BAF complex members, p300 / CBP, NuA4 complex members, Mediator components, and TFIID components are shown. Edge width indicates average spectral counts of two replicates. For clarity, two highly bound prey proteins (ZNF518A and ZNF518B) were removed from this visualization. [Figure 12B] Figure 1 shows the interaction network of transcriptional activators. AP-MS network of transcriptional activators. Labeled as in panel A. For clarity, nine highly connected prey proteins (APEH, ACTC1, ALDH1L1, PSDM4, LSS, FLII, PACSIN2, RPS3A, QPCTL) have been removed from this visualization. [Figure 13A] Figure 1 shows the effect of small molecule inhibitors on the transcriptional activity of 83 transcriptional regulators. Effect of CDK9 inhibition with flavopiridol. Known p300 interactors are shown as solid circles and known NuA4 interactors are shown as open circles. [Figure 13B] Figure 1 shows the effect of small molecule inhibitors on the transcriptional activity of 83 transcription regulators. Effect of casein kinase 2 inhibition by CX4945. Known p300 interactors are shown as filled circles and known NuA4 interactors are shown as open circles. [Figure 13C]Figure 1 shows the effect of small molecule inhibitors on the transcriptional activity of 83 transcription regulators. Effect of DYRK1A / DYRK1B inhibition by AZ191. Known p300 interactors are shown as solid circles and known NuA4 interactors are shown as open circles. DYRK1A / DYRK1B interactor DCAF7 and DCAF7 interactor NCKISPD are highlighted. [Figure 13D] Figure 1 shows the effect of small molecule inhibitors on the transcriptional activity of 83 transcriptional regulators. Effect of p300 inhibition by A-485 on transcriptional regulators characterized by AP-MS and BioID. Asterisks indicate statistically significant activators (FDR<5%). Statistical significance was calculated by independent two-tailed t-test assuming equal variances and corrected for multiple hypotheses using the false discovery rate (FDR) approach of Benjamini, Krieger, and Yekutieli (Benjamini et al., 2006). [Figure 13E] Figure 1 shows the effect of small molecule inhibitors on the transcriptional activity of 83 transcriptional regulators. Effect of BET bromodomain inhibition by JQ1 on transcriptional regulators characterized by AP-MS and BioID. Asterisks indicate statistically significant activators (FDR<5%). Statistical significance was calculated by unpaired two-tailed t-test assuming equal variances and corrected for multiple hypotheses using the false discovery rate (FDR) approach of Benjamini, Krieger, and Yekutieli (Benjamini et al., 2006). [Figure 14A] Figure 2 shows transcriptional activation by SRF-C3orf62. SRF, C3orf62, C3orf62-Cterm, and SRF-C3orf62 interactors were characterized by AP-MS. SRF-C3orf62 fusions tightly interact with CBP and p300. [Figure 14B] Figure 2 shows transcriptional activation by SRF-C3orf62. Analysis of differentially expressed genes in NIH3T3 cells expressing SRF-GFP, C3orf62-GFP, or SRF-C3orf62-GFP compared to cells expressing Nanoluc-GFP. [Figure 14C]Figure 2 shows transcriptional activation by SRF-C3orf62. Gene Ontology enrichment analysis of significantly up- and down-regulated genes in NIH3T3 cells expressing SRF-C3orf62-GFP. [Figure 14D] Figure 1 shows transcriptional activation by SRF-C3orf62. The overlap between significantly upregulated genes by SRF-C3orf62-GFP and target genes of SRF / MRTF or SRF / TCF, as previously published (Esnault et al., 2014; Gualdrini et al., 2016). Statistical significance was calculated using the hypergeometric distribution test. [Figure 15] Figure 1 shows the effect of the minimal activation sequence of each individual component of the SPDYE4-CITED1-P65-HSF1 fusion on the transcriptional activity of the EGFP reporter. All components were fused to the PYL1 dimerization domain and recruited to the reporter by addition of 1 μM abscisic acid (ABA) for 48 h. The sequences of the respective fusions are shown in SEQ ID NOs: 121, 123, 127, 129, 131, 133, and 135. The sequences of each of the individual domains tested are provided in SEQ ID NOs: 47, 90, 101-104, 116, and 118. [Figure 16] An example is shown when fusing two or more transactivation domains and targeting them to the same reporter. The addition of the SPDYE4 activation domain showed a marked improvement over the activity of the CITED1 domain alone. This contrasts with the weak activity when SPDYE4 itself was linked to a reporter, suggesting synergistic activity when fused to the CITED1 activation domain. All fusion constructs exhibited activation domains and no full-length proteins, except for "full-length CITED1 / 2" as shown, which were tested to compare activity to their corresponding transactivation domains. P300core is the catalytic domain EP300 (amino acids 1048-1664). Reporter cells were treated with 1 μM abscisic acid (ABA) or the same volume of DMSO for 48 h and then harvested for flow cytometry. Error bars represent SD from four replicates. [Figure 17]Activity of different combinations and orientations of activation domains from human SPDYE4, CITED1, p65, and HSF1 proteins linked to an EGFP reporter (left) or the promoter of the CD133 gene in HEK293T cells. Mobilization was induced by adding 1 μM abscisic acid (ABA) for 24 or 48 h before harvesting cells. [Figure 18] We show the effect of replacing each part of the multicomponent SPDYE4-CITED1-P65-HSF1 (SCPH) activator with different activation domains. The SPDYE4 activation domain appeared to be the most essential, as replacing it with other individually stronger activation domains destroyed the activity of the SPCH activator. On the other hand, from our screen we saw that replacing the CITED1 component with other potent human transactivation domains had no detrimental effect on activity and even improved its potency with the ZXDC and C3orf62 activation domains. All constructs were linked to an EGFP reporter by adding 1 μM abscisic acid for 48 h. [Figure 19A] Figure 1 shows the effect of 117 different effector domains, including different combinations of transactivation domains or fragments thereof, when used in combination with rTetR or dCas9-based recruitment systems. Transcription activation using 117 effector domains in combination with the rTetR DNA targeting domain. [Figure 19B] Figure 1 shows the effect of 117 different effector domains, including different combinations of transactivation domains or fragments thereof, when used in combination with rTetR or dCas9-based recruitment systems. Transcription activation using 117 effector domains in combination with the dCas9 DNA targeting domain. [Figure 19C] Figure 1 shows the effect of 117 different effector domains, including different combinations of transactivation domains or their fragments, when used in combination with rTetR or dCas9-based recruitment systems. Correlation of transcription activation between rTetR and dCas9-based recruitment systems. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0025] The following is a detailed description provided to assist those skilled in the art in carrying out the present disclosure.Unless otherwise defined, technical and scientific terms used herein have the same meaning as those commonly understood by those skilled in the art to which the present disclosure belongs.The technical terms used in the description of this specification are only for the purpose of describing specific embodiments and are not intended to limit the present disclosure.All publications, patent applications, patents, drawings and other references mentioned in this specification are expressly incorporated by reference in their entirety.

[0026] I. Definition As used herein, the following terms may have the meanings described below unless otherwise specified. However, it should be understood that other meanings known or understood by those skilled in the art are possible and are within the scope of this disclosure. All publications, patent applications, patents, and other references mentioned herein are incorporated by reference in their entirety. In case of conflict, the present specification, including definitions, will control. Additionally, the materials, methods, and examples are illustrative only and are not intended to be limiting.

[0027] As used herein, the terms "nucleic acid", "oligonucleotide", and "primer" refer to two or more covalently linked nucleotides. Unless the context clearly indicates otherwise, the term generally includes, but is not limited to, deoxyribonucleic acid (DNA) and ribonucleic acid (RNA), which may be single-stranded (ss) or double-stranded (ds). For example, the nucleic acid molecules or polynucleotides of the present disclosure may be composed of single-stranded and double-stranded DNA, DNA that is a mixture of single-stranded and double-stranded regions, single-stranded and double-stranded RNA, and RNA that is a mixture of single-stranded and double-stranded regions, hybrid molecules containing DNA and RNA that may be single-stranded, more typically double-stranded, or a mixture of single-stranded and double-stranded regions. Additionally, the nucleic acid molecule may be composed of triple-stranded regions containing RNA or DNA or both RNA and DNA. As used herein, the term "oligonucleotide" generally refers to a nucleic acid of up to 200 base pairs in length, which may be single-stranded or double-stranded. The sequences provided herein may be DNA or RNA sequences, but unless the context clearly dictates otherwise, the sequences provided are understood to encompass both DNA and RNA, as well as complementary RNA and DNA sequences. For example, the sequence 5'-GAATCC-3' is understood to include 5'-GAAUCC-3', 5'-GGATTC-3', and 5'GGAUUC-3'.

[0028] As used herein, the term "functional variant" includes modifications of the polypeptide sequences disclosed herein that perform substantially the same function in substantially the same manner as the polypeptide molecules disclosed herein. For example, functional variants can include active fragments of the polypeptides described herein, such as N- and / or C-terminal truncations that retain transcriptional activation activity and / or coactivator interactions. Functional variants can include variants with one or more substituted amino acids and / or variants that retain at least a minimum sequence identity to the unmodified or unvariant sequence. For example, functional variants can include one, two, three or more amino acid substitutions for every 10 amino acids. For example, functional variants can include sequences that have at least 80%, or at least 90%, or at least 95% sequence identity to the sequences disclosed herein. Functional variants can also include conservatively substituted amino acid sequences of the sequences disclosed herein. Substitutional amino acid variants are those in which at least one residue of a sequence has been removed and a different residue has been inserted in its place. An example of a substitutional amino acid variant is a conservative amino acid substitution. Functional variants, such as active fragments, including minimal fragments that retain transcriptional activation activity and / or coactivator interactions, can be identified, for example, using the methods described herein.

[0029] As used herein, a "conservative amino acid substitution" is one in which an amino acid is replaced with another amino acid residue without losing the desired properties of the protein. Suitable conservative amino acid substitutions can be made by substituting amino acids with similar hydrophobicity, polarity, and R chain length for one another. Examples of conservative substitutions include the substitution of one non-polar (hydrophobic) residue (e.g., alanine, isoleucine, valine, leucine, or methionine) for another, the substitution of one polar (hydrophilic) residue for another (e.g., between arginine and lysine, between glutamine and asparagine, between glycine and serine), the substitution of one basic residue (e.g., lysine, arginine, or histidine) for another, or the substitution of one acidic residue (e.g., aspartic acid or glutamic acid) for another. The phrase "conservative substitution" also includes the use of chemically derivatized residues or unnatural amino acids in place of derivatized residues, provided that such polypeptides exhibit the required activity.

[0030] As used herein, the term "heterologous transcriptional activator" or "transcriptional activator" refers to an engineered fusion protein or an engineered dimer comprising at least one transactivation domain (TAD) selected from the TADs listed in Table 2 and functional variants thereof, or an effector domain comprising at least two TADs selected from the TADs listed in Table 1, Table 2, Table 3, Table 4, Table 5, and / or Table 6 and any one functional variant thereof, operably linked to a DNA-targeting domain.

[0031] As used herein, the term "operably linked" refers to the relationship of two components that allows them to function in the intended manner. For example, a first polypeptide can be operably linked to a second polypeptide by covalent bonding (e.g., as a fusion protein) or through one or more interacting components. Similarly, when a reporter gene is operably linked to a promoter, the promoter drives the expression of the reporter gene.

[0032] The transcription activator may further comprise one or more interaction components of an interaction system that provides a functional interaction between the effector domain and / or the DNA targeting domain and / or the target DNA. The term "interaction components" is used herein to encompass one or more components of an interaction system, which together provide said functional interaction. As used herein, the term "interaction system" is intended to encompass interaction components that allow covalent or non-covalent interactions and / or constitutive or inducible interactions. Such an interaction system may comprise, for example, a peptide linker, optionally a protease-sensitive peptide linker; one or more dimers, trimers, or higher order multimerization components, such as an interaction domain, optionally an inducible dimer, trimer, or multimerization component, optionally an inducible interaction domain; and / or one or more components that can regulate the subcellular localization of the transcription activator. An interaction system may comprise two or more components.

[0033] The DNA targeting domain and the effector domain can be covalently linked, for example, as a domain of a single polypeptide (e.g., a fusion protein), or can be linked by an interaction component, such as, for example, an interaction domain (e.g., a dimer) that interacts under certain conditions. Thus, a heterologous transcriptional activator can include a single polypeptide, or can include a first polypeptide that includes a DNA targeting domain and a first interaction component (e.g., a dimer interaction domain), and a second polypeptide that includes an effector domain and a second interaction component (e.g., a dimer interaction domain), where the first dimer interaction domain and the second dimer interaction domain can interact under certain conditions. Higher order multimerization systems, such as the SunTag system (Tenenbaum et al., 2014), are also contemplated herein.

[0034] The interaction between the effector domain and / or the DNA targeting domain and / or the target DNA can be controlled using various inducible interaction systems.For example, the effector domain and the DNA targeting domain can be linked by a protease-sensitive linker, such as the self-cleaving NS3 protease domain, which is stabilized in the presence of an NS3 inhibitor, such as grazoprevir.In another example, the localization of the DNA targeting domain and / or the effector domain to the nucleus can be controlled by an interaction component, such as a localization domain, e.g., tamoxifen-regulated nuclear localization using an estrogen receptor ligand binding domain variant.In a further example, the DNA targeting domain can be linked to a first interaction component, such as a first interaction domain, and the effector domain can be linked to a second interaction component, such as a second interaction domain, such that the first interaction domain and the second interaction domain interact.

[0035] As used herein, the term "interaction domain" refers to a sequence motif in a first polypeptide (e.g., a first dimer interaction domain) that can interact with a binding partner that contains a sequence motif in a second polypeptide (e.g., a second dimer interaction domain) to operably link the first and second polypeptides. In particular, the term is intended to encompass a first or second interacting dimer domain that together form a heterodimer pair that dimerizes under, for example, appropriate induction conditions. Other interacting domains are specifically contemplated and can be identified by the skilled artisan depending on the desired properties. Suitable inducible interaction domain pairs include, but are not limited to, FKBP / FRB (FK506 binding protein / FKBP rapamycin binding), which can be induced, for example, with rapamycin or AP21967; PYL / ABI, which can be induced, for example, with abscisic acid; GID1 / GAI, which can be induced, for example, with gibberellin or gibberellic acid; and pMag / nMag, which can be induced, for example, with blue light and / or temperature.

[0036] As used herein, "DNA targeting domain" refers to a polypeptide domain that binds to DNA under DNA-binding conditions, thereby targeting the polypeptide to said DNA.The DNA targeting domain can be any suitable DNA binding domain, for example, enzymatically inactive sequence-specific DNA targeting protein, for example, CRISPR-Cas protein (for example, dCas9, dCas12, or other Cas family proteins), zinc finger DNA binding domain, transcription activator-like effector (TALE) DNA binding domain, bromo domain, chromo domain, Tudor domain, WD40 domain, PHD domain, PWWP domain, or other DNA binding domain (DBD) from eukaryote or prokaryote (for example, forkhead, basic helix-loop-helix, leucine zipper, homeodomain, nuclear hormone receptor, or tet repressor), or their variants. The DNA targeting domain may bind to DNA in a sequence-specific manner (e.g., Cas family proteins, zinc finger DNA binding domains, TALE DNA binding domains) or to specific chromatin modifications (e.g., bromodomains (for acetylated histones) or chromodomains, Tudor domains, WD40 domains, PHD domains, PWWP domains, etc. (for methylated histones)). The DNA targeting domain may be a natural (e.g., unengineered) DNA binding domain, such as a DNA binding domain found in a naturally occurring (e.g., endogenous) transcription factor, or the DNA targeting domain may be engineered to provide, for example, custom sequence specificity (e.g., different sequence specificity than the unengineered DNA binding domain) or altered DNA binding affinity. Methods for engineering, for example, zinc finger DNA binding domains and TALE DNA binding domains to provide custom DNA binding specificity are known in the art (e.g., Maeder et al. 2008 and Sanjana et al. 2012.). Enzymatically active Cas9 can also be used to effect inhibition, e.g., when the guide is a truncated guide (see, e.g.,

[24] ).A DNA targeting domain may have inherent target sequence specificity, for example, in the case of zinc finger DNA binding domains and TALE DNA binding domains, or the target sequence specificity may be mediated by additional sequence-specific factors, such as, for example, a guide RNA, in the case of CRISPR-Cas proteins. Suitable DNA binding conditions depend on the DNA targeting domain and may include the presence of additional factors, such as, for example, tetracycline, in the case of tet-repressors, or a guide RNA, in the case of Cas-family proteins.

[0037] As used herein, the term "effector domain" refers to a polypeptide domain that includes at least one transactivation domain (TAD) as described herein, e.g., the TADs listed in Tables 1-5 and functional variants thereof, e.g., active fragments thereof. Optionally, the effector domain may include two or more, e.g., two, three, four, or more, transactivation domains as described herein. In the activators described herein, the active fragment may be about 15 amino acids, about 20 amino acids, about 30 amino acids, about 40 amino acids, about 50 amino acids, about 60 amino acids, about 70 amino acids, or any number between 15 and 70 amino acids, or more than 70 amino acids. For example, with respect to HSF1, the active fragment may include GFSVDTSALLDLFSP (SEQ ID NO: 104), which corresponds to amino acids 406-420 of HSF1. Thus, by way of example, the active fragment of HSF1 may include amino acids 401-427 of HSF1. Active fragments of other TADs may be identified by any suitable method, e.g., using the methods described herein.

[0038] The heterologous transcriptional activator may be an effector N-terminal fusion or a C-terminal fusion, for example, the order of fusion may be effector domain-DNA targeting domain or DNA targeting domain-effector domain (see, for example,

[25] ,

[26] ,

[27] , and

[28] ). The effector domain may be fused to the DNA targeting domain via a linker. Similarly, two or more TADs may be fused together by one or more linkers. For example, glycine and glycine serine linkers may be used. The transcriptional activators described in the examples used various glycine serine linkers, for example, SGGSGGS (SEQ ID NO: 6), GGS, SGGS (SEQ ID NO: 7), and / or GSGSGS (SEQ ID NO: 8). Other linkers may also be used, for example, INSRSSGS (SEQ ID NO: 9).

[0039] As used herein, the term "CRISPR-Cas" or "Cas" refers to CRISPR clustered regularly interspaced short palindromic repeats-CRISPR associated (CRISPR-Cas) protein that binds to RNA and is targeted to a specific DNA sequence by the RNA it binds to. CRISPR-Cas is a class II monomeric Cas protein, for example, type II Cas, such as Cas9. The Cas9 protein can be Cas9 from Streptococcus pyogenes, Francisella novicida, A. Naesulndii, Staphylococcus aureus, or Neisseria meningitidis. Optionally, Cas9 is from S.pyogenes. The Cas protein can also be, for example, Cas12a (e.g., dCas12a) from Acidaminococcus sp., Lachnospiraceae bacterium, or Francisella tularensis (which have been shown to function as dCas variants), CasΦ (Cas12j) and CasX (Cas12e) can also be used.

[0040] As used herein, the term "dCas9" refers to an enzymatically inactive (or dead) Cas9 that lacks DNA endonuclease activity but retains target DNA binding activity. For example, dCAS9 comprises a sequence of CAS9 and D10A / H840A mutations in the RuvC1 and HNH nuclease domains. Optionally, dCas9 is a protein that comprises an amino acid sequence having at least 80%, at least 90%, at least 95%, at least 99%, or 100% sequence identity with the protein encoded by SEQ ID NO:1, comprises D10A / H840A mutations, and retains Cas9 target DNA binding activity (e.g., gRNA binding to a target site). Similarly, dCas12a refers to an enzymatically inactive Cas12a.

[0041] As used herein, the term "guide RNA", "guide", or "gRNA" refers to an RNA molecule that hybridizes with a specific DNA sequence and minimally includes a spacer sequence. The guide RNA may further include a protein-binding segment that binds to a CRISPR-Cas protein. The portion of the guide RNA that hybridizes with the specific DNA sequence is referred to herein as a nucleic acid targeting sequence or a spacer sequence. The protein-binding segment of the guide may include, for example, a tracrRNA and / or a direct repeat. The term "guide" or "guide RNA" may refer to an RNA molecule that includes only a spacer sequence or a spacer sequence and a protein-binding segment, depending on the context. The guide RNA may be represented by the corresponding DNA sequence. The guide may be, for example, a truncated guide that includes 15 or fewer nucleotides of complementarity to the target site, as described in

[24] , when the enzyme is Cas9. For example, when Cas9 interacts with a truncated guide, the DNA-binding ability of Cas9 remains intact and its nucleolytic activity is eliminated. Any length of guide that maintains the Cas-binding ability may be used.

[0042] As used herein, the term "spacer" or "spacer sequence" refers to a portion of a guide that forms or can form an RNA-DNA duplex with a target sequence or a portion thereof. A spacer sequence can be complementary to or correspond to a specific CRISPR target sequence. The nucleotide sequence of the spacer sequence can determine the CRISPR target sequence and can be designed or configured to target a desired CRISPR target site.

[0043] As used herein, the term "tracrRNA" refers to a "transcoded crRNA" that can interact with a CRISPR-Cas protein, such as Cas9, and can be connected to or form part of a guide RNA. The tracrRNA may be, for example, a tracrRNA from S. pyogenes. The tracrRNA may have, for example, the sequence 5'-gtttcagagctatgctggaaacagcatagcaagttgaaataaggctagtccgttatcaacttgaaaaagtggcaccgagtcggtgc-3' (SEQ ID NO: 2). Other tracrRNAs may also be used. Suitable tracrRNAs can be identified by one of skill in the art and include, for example, 5'-GTTTTAGAGCTAGAAATAGCAAGTTAAAATAAGGCTAGTCCGTTATCAACTTGAAAAAGTGGCACCGAGTCGGTGC-3' (SEQ ID NO: 3), or 5'-GTTTCAGAGCTACAGCAGAAATGCTGTAGCAAGTTGAAAT-3' (SEQ ID NO: 4).

[0044] As used herein, the term "CRISPR target site" or "CRISPR-Cas target site" refers to a nucleic acid to which an activated CRISPR-Cas protein (e.g., a CRISPR-Cas protein such as dCas9 bound to a guide RNA) binds under appropriate conditions. A CRISPR target site includes a protospacer adjacent motif (PAM) and a CRISPR target sequence (i.e., corresponding to the spacer sequence of the guide to which the activated CRISPR-Cas protein binds). The sequence and relative position of the PAM to the CRISPR target sequence depends on the type of CRISPR-Cas protein. For example, a CRISPR target site for Cas9 or dCas9 can include, from 5' to 3', a 3-nucleotide PAM having 15-25, 16-24, 17-23, 18-22, or 19-21 nucleotides, optionally a 20-nucleotide target sequence, followed by the sequence NGG. Thus, a Cas9 target site can have the sequence 5'-N1NGG-3', where N1 is 15-25, 16-24, 17-23, 18-22, or 19-21 nucleotides in length, optionally 20 nucleotides in length.

[0045] CRISPR target site can be in any suitable genomic locus.For example, CRISPR target site can be in promoter, enhancer, 3'UTR or other regulatory element, in gene, optionally intron or exon, in the locus corresponding to non-coding RNA, or in intergenic region.Optionally, CRISPR target site is in promoter or enhancer.

[0046] Target DNA located in the nucleus of a cell requires a transcriptional activator that can enter the nucleus. Thus, the transcriptional activator may be nuclear localized and / or may include, for example, one or more nuclear localization signals (NLS), optionally one or more SV40 NLSs. Optionally, the transcriptional activator includes two or more NLSs. Optionally, the transcriptional activator may include one or more N-terminal NLSs, one or more C-terminal NLSs, one or more internal NLSs, or one or more N-terminal, one or more C-terminal NLSs, and / or one or more internal NLSs. Other configurations are specifically contemplated. In one embodiment, the NLS is an SV40 NLS with the sequence PKKKRKV (SEQ ID NO: 22). In one embodiment, the NLS further includes an N-terminal and / or C-terminal linker, such as INSRSSGS (SEQ ID NO: 9), optionally with the sequence INSRSSGSPKKRKVGS (SEQ ID NO: 141).

[0047] Transcriptional activator can also be labeled with tag.For example, suitable tags include, but are not limited to, Myc, FLAG, HA, V5, ALFA, T7, 6xHis, VSV-G, S-tag, AviTag, StrepTag II, CBP, GFP, mCherry.Tag can be fused at N-terminus, C-terminus, or between two components of heterologous transcriptional activator, such as between DNA targeting domain and effector domain.

[0048] When a range of values ​​is provided, unless the context clearly dictates otherwise, each intermediate value between the upper and lower limits of the range, and at any other stated or intermediate value within the stated range, is included in the description to the tenth of the unit of the lower limit. Ranges from any lower limit to any upper limit are contemplated. The upper and lower limits of these smaller ranges that may be independently included in the smaller ranges are also included in the description, subject to any specifically excluded limits in the stated range. When a stated range includes one or both limits, ranges excluding either or both of those included limits are also included in the description.

[0049] Please note that as used in this specification and the appended claims, the singular forms "a," "an," and "the" include plural referents unless the context clearly dictates otherwise.

[0050] All numerical values ​​within the detailed description and claims herein are modified by "about" or "approximately" the indicated value to account for experimental error and variations that would be expected by one of ordinary skill in the art.

[0051] As used herein in the specification and claims, the term "and / or" should be understood to mean "either or both" of the elements so conjoined, i.e., elements that are conjunctively present in some cases and disjunctively present in other cases. Multiple elements listed with "and / or" should be construed in the same manner, i.e., "one or more" of the elements so conjoined. Other elements other than the elements specifically identified by the "and / or" clause may be optionally present, whether or not related to the elements specifically identified.

[0052] As used herein in the specification and claims, "or" should be understood to have the same meaning as "and / or" as defined above. For example, when separating items in a list, "or" or "and / or" shall be interpreted as being inclusive; that is, including at least one, but more than one, of a number or list of elements, and, optionally, including additional unlisted items. Only where terms clearly indicated to the contrary, such as "only one of" or "exactly one of," or when used in the claims, "consisting of" refers to including exactly one element of a number or list of elements. In general, as used herein, the term "or" should only be interpreted as indicating exclusive alternatives (i.e., "both, but not one or the other") when preceded by terms of exclusivity, such as "either," "one of," "only one of," or "exactly one of."

[0053] In the claims and herein above, all transitional phrases, such as "comprising," "including," "bearing," "having," "containing," "involving," "holding," "consisting of," and the like, are understood to be open-ended, i.e., including but not limited to. Only the transitional phrases "consisting of" and "consisting essentially of" shall be closed or semi-closed transitional phrases, respectively.

[0054] As used herein in the specification and claims, the phrase "at least one" in connection with a list of one or more elements should be understood to mean at least one element selected from any one or more of the elements of the list of elements, but not necessarily including at least one of each and every element specifically listed in the list of elements, and not excluding combinations of elements in the list of elements. This definition also allows for elements other than those specifically identified in the list of elements to which the phrase "at least one" refers, may optionally be present, whether or not related to the specifically identified element.

[0055] Similarly, it is specifically contemplated herein that the phrase "one or more" in reference to a group of elements includes at least one member of the recited group, but not necessarily one of each of the recited group members. For example, if an element includes one or more of the groups A, B, and / or C, the element may include A; B; C; A and B; A and C; B and C; or A, B, and C. For example, with reference to the above example, there may be additional members not specifically recited in the group, and the element may further include an unrecited member D, and thus include A and D; B and D; A, C, and D, etc.

[0056] As used herein, the term "about" means the referenced number plus or minus 10% to 15%, 5 to 10%, or, optionally, about 5%.

[0057] It should also be understood that in certain methods described herein that include more than one step or action, the order of the steps or actions of the method is not necessarily limited to the order in which the steps or actions of the method are recited, unless the context indicates otherwise. Furthermore, it is intended that the definitions and embodiments described in certain sections are applicable to other embodiments described herein where they are suitable, as would be understood by one of ordinary skill in the art. For example, in the following text, various aspects of the present disclosure are defined in more detail. Each aspect so defined may be combined with any other aspect or aspects, unless expressly indicated to the contrary. In particular, any feature described herein may be combined with any other feature or features described herein.

[0058] Although any materials and methods similar or equivalent to those described herein can also be used in the practice or testing of the present disclosure, the following materials and methods are now described.

[0059] II. Materials and Methods Described herein is a collection of heterologous transcriptional activators comprising one or more transactivation domains (TADs) and combinations thereof that can be operably linked to a DNA targeting domain to generate a heterologous transcriptional activator and used to activate gene expression of a desired gene, including an endogenous gene, for example, for therapeutic purposes. In the heterologous transcriptional activators described herein, an effector domain is operably linked to a DNA targeting domain that can direct binding to any locus in the genome of the fusion construct. As shown in the Examples, a heterologous transcriptional activator comprising dCas9 or rTetR functionally associated with an effector domain comprising any one of the TADs listed in Table 1, active fragments thereof, such as those of SEQ ID NOs: 100-104, 107-115, 118-119, 156-160, 162, 164, and 166-185, as well as combinations of two or more TADs or active fragments thereof, such as those listed in Table 3, can be used to activate transcription of a target gene. The TAD or active fragment thereof may further comprise a linker and / or additional native sequence, e.g., 1, 2, 3, 4, 5, 6, 7 or more, e.g., up to 5, up to 10, up to 20, up to 30, or up to 40 (or any number in between) N-terminal and / or C-terminal amino acids, on the ends of the TAD or active fragment. For example, the TAD labeled "ZXDC Short" in Table 3 (SEQ ID NO: 107) comprises a 40 amino acid fragment found in both ZXDC-12 and ZXDC-13, and includes an additional 6 N-terminal amino acids of the ZXDC-12 native sequence and an additional 5 C-terminal amino acids of the ZXDC-13 native sequence. In one embodiment, the active fragment may be shorter than the identified fragment. For example, the active fragment may be, e.g., 20 or 25 amino acids of the 40 amino acid fragment found in both ZXDC-12 and ZXDC-13, optionally an internal portion of, e.g., SEQ ID NO: 107. A shorter active fragment is exemplified, for example, by SEQ ID NO: 119, which contains the 20 amino acid fragment found in both HSF1-20 and HSF1-21. Active fragments can be identified as described elsewhere herein. For example, a minimal fragment can be identified by comparing the active fragments.For example, for ATF6, the overlapping fragment shown in Table 5 is HRLDEDWDSALFAELGYFTDTDETDELQLEANETYENDFDNL, and for KLF7, the overlapping fragment is YFSALPSLEPTTWQQTTLERAYQTEPRRISETFGEDLDC.

[0060] Thus, one embodiment of the present disclosure includes a heterologous transcriptional activator comprising a DNA-targeting domain and an effector domain comprising at least one TAD selected from the group consisting of any one of the TADs listed in Tables 1-6, optionally in Table 1 or optionally in any one of Tables 2 or 6, active fragments thereof, and combinations thereof, e.g., at least two TADs selected from the TADs listed in Table 1 or Table 3 or functional variants thereof, preferably at least one TAD selected from the TADs listed in Table 4 or Table 5 or Table 6, and / or functional variants thereof, wherein the DNA-targeting domain and the effector domain are operably linked.

[0061] In one embodiment, the at least one TAD is selected from Table 1.

[0062] In one embodiment, the at least one TAD is selected from Table 2.

[0063] In one embodiment, the at least one TAD is selected from Table 3.

[0064] In one embodiment, the at least one TAD is selected from Table 4.

[0065] In one embodiment, the at least one TAD is selected from Table 5.

[0066] In one embodiment, at least one TAD is selected from Table 6.

[0067] It is understood that the at least two TADs or functional variants thereof can be selected from any table, e.g., one, two or more from each of Tables 3 and 4, e.g., three or four, such as from Table 3. It is also contemplated that the grouping can include any subcombination of the TADs listed in any of Tables 1-6, e.g., excluding one or more TADs.

[0068] The DNA targeting domain and the effector domain can be operably linked, for example, as a domain of a single polypeptide, by covalent bonding, and / or can be operably linked through one or more interaction components, for example, interaction domains, and / or can interact under certain conditions.Thus, in one embodiment, the heterologous transcriptional activator is a single polypeptide.In another embodiment, the heterologous transcriptional activator further comprises a pair of (i.e., first and second) interaction domains, optionally a dimer interaction domain, and optionally a pair of inducible dimer interaction domains that dimerize under appropriate conditions. For example, a heterologous transcriptional activator can comprise a first polypeptide comprising a DNA targeting domain and a first dimer interaction domain, optionally an inducible dimerization domain, and a second polypeptide comprising an effector domain and a second dimer interaction domain, optionally an inducible dimerization domain, where the first dimer interaction domain and the second dimer interaction domain interact, and optionally where the first inducible dimerization domain and the second inducible dimerization domain interact in the presence of one or more inducing agents.

[0069] As shown in the examples, the dimerization of heterologous transcriptional activators including ABI1 and PYL1 can be induced by the addition of abscisic acid. Thus, in one embodiment, the transcriptional activator comprises a first and a second inducible dimerization domain that provide inducible transcriptional activation in the presence of an inducer. Those skilled in the art can easily identify and select suitable inducible dimerization domains that can be used together. Any suitable inducible dimerization domain can be used, for example, the dimerization of ABI1 and PYL1 can be induced by the addition of abscisic acid. Other inducible systems include those based on induction by rapamycin, gibberellic acid / gibberellin, and split dCas9-based systems. For example, the dimerization of GID1 and GAI can be induced by gibberellin, and the dimerization of FKBP and FRB can be induced by rapamycin or its analogs, such as rapalogs. Higher order multimerization systems, such as the SunTag system (Tenenbaum et al., 2014), are also contemplated herein.

[0070] The interaction between the DNA targeting domain and the effector domain can also be controlled using other inducible systems. Other systems (not dependent on dimerization) include grazoprevir-induced stabilization (Tague et al. 2018) or tamoxifen-regulated nuclear localization using estrogen receptor ligand-binding domain variants. In the case of grazoprevir-induced stabilization, the DNA targeting domain and the effector domain are linked by the self-cleaving NS3 protease domain. Only in the presence of grazoprevir (which inhibits NS3 activity) do the DNA targeting domain and the effector domain stay together and regulate gene expression.

[0071] The DNA targeting domain can be selected from various DNA binding domains, such as zinc finger DNA binding domains, transcription activator-like effector (TALE) DNA binding domains, dCas9, dCas12 or other Cas family proteins, or other DNA binding domains from eukaryotes or prokaryotes (e.g., forkhead, basic helix-loop-helix, leucine zipper, homeodomain, nuclear hormone receptor, or tet-repressor), or variants thereof. The DNA targeting domain can be a natural (e.g., non-engineered) DNA binding domain, such as a DNA binding domain found in a naturally occurring (e.g., endogenous) transcription factor, or the DNA targeting domain can be engineered to provide, for example, custom sequence specificity (e.g., sequence specificity different from that of the non-engineered DNA binding domain) or modified DNA binding affinity. Methods for engineering, for example, zinc finger DNA binding domains and TALE DNA binding domains to provide custom DNA binding specificity are known in the art (e.g., Maeder et al. 2008 and Sanjana et al. 2012.). When the heterologous transcriptional activator described herein comprises a DNA targeting domain that includes a native DNA binding domain, the effector domain is targeted to all loci to which the transcription factor endogenously binds, thereby enhancing / replacing the function of the endogenous transcription factor. For example, it is known that replacing the Oct4 transactivation domain with VP16 increases the efficiency of reprogramming fibroblasts into iPS cells. Similarly, a heterologous transcriptional activator that comprises a native DNA binding domain operably linked to an effector domain can promote, for example, wound healing, transdifferentiation, or tissue regeneration by activating the transcription of target genes regulated by endogenous transcription factors. In the case of engineered (e.g., individual sequence specific) zinc finger DNA binding domains, TALE DNA binding domains, or Cas family proteins, the effector domain can be delivered in a controlled manner to one or more specific loci, or optionally to a single locus in the genome.

[0072] In one embodiment, the DNA targeting domain comprises a CRISPR-Cas protein, such as dCas9. An enzymatically inactive CRISPR-Cas protein that retains gRNA and target DNA binding activity can be used. For example, the D10A / H840A mutation in Cas9 introduces mutations into the RuvC1 and HNH nuclease domains, resulting in inactivation. In one embodiment, the CRISPR-Cas protein is dCas9 having an amino acid sequence of SEQ ID NO:1 or an amino acid sequence having at least 80%, at least 90%, at least 95%, or at least 99% sequence identity with SEQ ID NO:1, comprising D10A / H840A, and retaining gRNA and target DNA binding activity. Other enzymatically inactive CRISPR-Cas proteins are also contemplated and can be identified by one skilled in the art.

[0073] In one embodiment, the DNA targeting domain comprises a zinc finger DNA binding domain, hi one embodiment, the zinc finger DNA binding domain is an engineered zinc finger DNA binding domain that is engineered to bind to a specific DNA sequence.

[0074] The effector domain comprises at least one transactivation domain (TAD) or active fragment thereof as described herein. As shown in the Examples, various full-length ORFs and TADs identified herein can be used alone or in combination to activate transcription of endogenous genes such as GFP reporter constructs or CD133. Also, as shown in the Examples, the effector domain can comprise at least one TAD domain as shown in Table 1, and / or an active fragment thereof, for example as shown in Table 3. Thus, in one embodiment, the effector domain comprises at least one TAD as shown in Table 2, Table 4, and / or Table 5, and / or Table 6, and / or a functional variant of any one of these, or two or more TADs as shown in Table 1, Table 2, Table 3, Table 4, Table 5, and / or Table 6, and a functional variant of any one of these. In one embodiment, the TAD comprises a polypeptide having a sequence having at least 80%, at least 90%, at least 95%, or at least 99% sequence identity to any one of the TAD domains of Table 1, Table 2, Table 3, Table 4, Table 5, and / or Table 6, as well as functional variants of any one of them, and which retains transcriptional activation activity and / or interaction with a particular transcriptional co-activator, e.g., CBP / p300, NuA4, and / or BRD4.

[0075] These variants and combinations may also be used. In one embodiment, the effector domain comprises two or more tandem TADs, optionally two TADs, three TADs, four TADs, or more than four TADs, such as five TADs, ten TADs, fifteen TADs, twenty TADs, twenty-five TADs, thirty TADs, or any number of TADs between five and thirty TADs, or more than thirty TADs. In one embodiment, the effector domain comprises two or more TADs or functional variants thereof selected from those listed in Table 1, Table 2, Table 3, Table 4, Table 5, and / or Table 6, and functional variants of any one thereof. In one embodiment, the effector domain comprises three or four TADs selected from those listed in Table 3. In one embodiment, the effector domain comprises one or more of SEQ ID NO: 185, SEQ ID NO: 103, SEQ ID NO: 167, SEQ ID NO: 105, SEQ ID NO: 106, and / or SEQ ID NO: 104. In one embodiment, the effector domain comprises SEQ ID NO: 185, optionally SEQ ID NO: 90, 91, 102, or 157. In one embodiment, the effector domain comprises SEQ ID NO: 103, optionally SEQ ID NO: 46, 47, or 162. In one embodiment, the effector domain comprises SEQ ID NO: 167, optionally SEQ ID NO: 101, 110, 166, or 172. In one embodiment, the effector domain comprises SEQ ID NO: 105, optionally SEQ ID NO: 116, 117, or 165. In one embodiment, the effector domain comprises SEQ ID NO: 106, optionally SEQ ID NO: 116 or 117. In one embodiment, the effector domain comprises SEQ ID NO: 104, optionally SEQ ID NO: 118, 119, or 159.In one embodiment, the effector domain comprises SPDYE4-CITED1-RELA-HSF1 (SEQ ID NO: 121); SPDYE4-CITED1-RELA (SEQ ID NO: 123); HSF1-RELA-SPDYE4-CITED1 (SEQ ID NO: 125); SPDYE4-CITED1-p65-miniHSF1 (SEQ ID NO: 127); miniSPDYE4-CITED1-p65-HSF1 (SEQ ID NO: 129); SPDYE4-miniCITED1-p65-HSF1 (SEQ ID NO: 131); SPDYE4-CITED1-minip65(C)-HSF1 (SEQ ID NO: 133); or SPDYE4-CITED1-minip65(N)-HSF1 (SEQ ID NO: 135). In one embodiment, the effector domain is SPDYE4-CITED1-SERTAD2-HSF1;SPDYE4-CITED1-KLF6-HSF1;SPDYE4-CITED1-ZXDC-HSF1;SPDYE4-CITED1-ATF6-HSF1;SPDYE4-CITED1-FOXO1-HSF1;SPDYE4-CITED1-ATMIN-HSF1;SPDYE4-CITED1-p65-SERTAD2;SPDYE4-CITED1-p65-KLF6;SPDYE4-CITED1-p65-ZXDC;SPDYE4-CITED1-p65-ATF6;SPDYE4 -CITED1-p65-FOXM1;SPDYE4-CITED1-p65-ATMIN;SPDYE4-C3orf62-p65-HSF1;SPDYE4-DDIT3-p65-HSF1;SPDYE4-FOXO1-p65-HSF1;SPDYE4-ATMIN-p65-HSF1;SPDYE4-ZXDC-p65-HSF1;C3orf62-CITED1-p65-HSF1;C11orf74-CITED1-p65-HSF1;KLF6-CITED1-p65-HSF1;ZXDC-CITED1-p65-HSF1;or SOX7-CITED1-p65-HSF1.In one embodiment, the effector domain comprises SPDYE4-C3orf62.2-P_AD-HSF1 (SEQ ID NO: 174); SPDYE4-C3orf62.3-P_AD-HSF1 (SEQ ID NO: 176); SPDYE4-C3orf62_MT-P_AD-HSF1 (SEQ ID NO: 178); SPDYE4-DDIT3_MT-P_AD-HSF1 (SEQ ID NO: 180); SPDYE4-CITED1-P_AD-HSF1_MT (SEQ ID NO: 182); or 3xZNF473_KRAB (SEQ ID NO: 184). Other combinations are specifically contemplated herein.

[0076] An effector domain may comprise two or more TADs with different transcriptional coactivator preferences. For example, an effector domain may comprise a TAD that interacts with a CBP / p300 component, such as a FOXO TAD, and a TAD that interacts with a BET component, such as a SPDYE4 TAD. An effector domain may comprise two or more TADs with similar transcriptional coactivator preferences. For example, an effector domain may comprise two TADs that interact with a CBP / p300 component, such as a FOXO1 TAD and a CITED1 TAD. Other combinations are specifically contemplated herein.

[0077] As used herein, with respect to a functional variant, "effective" means that the functional variant retains at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, 100%, or more than 100% of the transcriptional activation activity and / or coactivator interaction compared to an unmodified or non-variant TAD (e.g., wild-type or full-length TAD). The transcriptional activation activity and / or coactivator interaction of variants such as truncations can be determined, for example, using methods described herein. For example, transcriptional activation activity can be determined using a GFP reporter system as described in the Examples. The variants can be anchored in the same reporter or endogenous context, controlling for the expression level of the respective DNA targeting domain (e.g., dCas9). Any difference detected in the induced expression of a reporter gene or target gene compared to the parent TAD can be attributed to the effect of the variant. Coactivator interaction can be determined, for example, by AP-MS and / or BioID, for example, as shown in the Examples.

[0078] Exemplary TAD and effector domain nucleic acids and polypeptides are provided in Tables 1-6 and SEQ ID NOs: 120-135 and 173-184. In one embodiment, the effector domain can comprise an amino acid sequence having at least 80%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to the amino acid sequence encoded by the nucleic acid or the amino acid sequence encoded by the TAD of SEQ ID NOs: 120-135 and 173-184. The activity of the encoded polypeptide of such a polypeptide (when fused or expressed and activated) is as effective (e.g., provides at least 80% of the effective transcriptional activation) as, for example, SEQ ID NOs: 120-135 and 173-184.

[0079] In one embodiment, the effector domain is fused to the DNA targeting domain via a linker. In one embodiment, two or more TADs are fused together by one or more linkers. For example, glycine and glycine serine linkers can be used. The transcription activators described in the examples used various glycine serine linkers, such as SGGSGGS (SEQ ID NO: 6), GGS, SGGS (SEQ ID NO: 7), and / or GSGSGS (SEQ ID NO: 8). Other linkers can also be used, such as INSRSSGS (SEQ ID NO: 9).

[0080] In one embodiment, the transcriptional activator comprises one or more nuclear localization signals (NLS). Any suitable NLS can be used. Optionally, the NLS is an SV40 NLS. The one or more NLS can be one or more N-terminal NLSs, one or more C-terminal NLSs, one or more internal NLSs, and / or combinations thereof. Optionally, the NLS can comprise the NLS of SEQ ID NO: 22. In one embodiment, the NLS further comprises an N-terminal and / or C-terminal linker, such as INSRSSGS (SEQ ID NO: 9), optionally having the sequence INSRSSGSPKKRKVGS (SEQ ID NO: 141).

[0081] As described herein, a transcriptional activator or effector domain can be encoded by a nucleic acid and / or expressed from an expression construct. Thus, one aspect of the disclosure is a nucleic acid encoding a transcriptional activator described herein. Another aspect of the disclosure is a nucleic acid encoding an effector domain of a transcriptional activator described herein. For example, the nucleic acid can encode a TAD provided in any one of Tables 1-6, optionally Tables 2, 4, 5, and / or 6, optionally Table 2 or Table 4 or Table 6, or two or more TADs provided in Tables 1-6. In one embodiment, the nucleic acid can comprise any one of the nucleic acids of SEQ ID NOs: 120, 122, 124, 126, 128, 130, 132, 134, 173, 175, 177, 179, and 180, or a sequence having at least 80%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity to SEQ ID NOs: 120, 122, 124, 126, 128, 130, 132, 134, 173, 175, 177, 179, and 180, where the heterologous transcriptional activator can be, for example, a sequence Activates transcription as effectively as the effector domains encoded by Nos. 120, 122, 124, 126, 128, 130, 132, 134, 173, 175, 177, 179, and 180, e.g., at least 80% effectively, at least 85% effectively, at least 90% effectively, at least 95% effectively, at least 96% effectively, at least 97% effectively, at least 98% effectively, at least 99% effectively, at least 100% effectively, or more than 100% effectively, e.g., when assessed in an assay described herein. The sequence identity is, for example, to the complete effector domain sequence or to one or more TADs or TAD fragments encoded therein. Other parts, linkers, NLSs, etc., may be completely different. The nucleic acid encoding the effector domain may be suitable for generating a nucleic acid encoding a transcription activator as described herein.For example, the nucleic acid encoding the effector domain can be flanked by suitable cloning sites, or an expression construct or vector containing the nucleic acid can include cloning sites to facilitate insertion of the DNA-targeting domain to operably link the effector domain and the DNA-targeting domain.

[0082] As used herein, the term "cloning site" refers to a portion of a nucleic acid molecule into which a nucleic acid molecule of interest can be inserted or to which a nucleic acid molecule of interest can be linked using recombinant DNA technology (cloning). In the context of an expression cassette, the cloning site can be located between a promoter and a polyadenylation signal, such that a nucleic acid molecule of interest can be cloned into the expression cassette in operably linked to the promoter and polyadenylation site. Several cloning techniques are known to those skilled in the art, and the cloning site contains the necessary features to allow the insertion of a nucleic acid molecule of interest at the cloning site, such as restriction endonuclease site(s), recombinase recognition site(s), or blunt or sticky end(s). The cloning site can be, for example, a multiple cloning site (MCS) or polylinker region that contains multiple unique restriction enzyme recognition sites that allow the nucleic acid molecule of interest to be inserted. Alternatively or additionally, the cloning site may contain one or more recombinase recognition sites that allow DNA insertion by recombinational cloning, using a site-specific recombinase(s) such as integrase or Cre recombinase to catalyze DNA insertion. Examples of recombinational cloning systems include Gateway® (integrase), Creator™ (Cre recombinase), and Echo Cloning™ (Cre recombinase). In some cloning strategies, an expression cassette or vector may be provided as a linear molecule, and blunt or overhanging ends of a nucleic acid molecule of interest may be joined, for example by ligation or polymerase activity, to the blunt or overhanging ends of the expression cassette or vector, thus forming a circular molecule. In this case, the blunt or overhanging ends of the expression cassette or vector together may be considered as a cloning site. Such an approach is commonly used to clone PCR products.

[0083] A related aspect is an expression construct comprising a nucleic acid encoding a transcription activator and a transcription termination site operably linked to a promoter. Any suitable promoter can be used. Suitable promoters can be identified by those skilled in the art and can include, for example, CMV, EF1A, or PGK. For example, the promoter and / or enhancer sequences of SEQ ID NO: 25, 26, 27, and / or 28 can be used in the expression construct. Inducible promoters can also be used.

[0084] In one embodiment, the construct is a vector. Any suitable vector can be used. Suitable vectors can be identified by those skilled in the art and can include viral vectors, optionally lentiviral vectors or adenoviral vectors. Suitable vectors can include, for example, a promoter for expressing the effector construct, a polyA tail, a 3'UTR element such as WPRE for increasing the stability of expression, an insulator sequence, a lentiviral packaging signal, a fluorescent protein, and / or an antibiotic resistance marker. Additional suitable components can be identified by those skilled in the art.

[0085] In another embodiment, the transcription activator, nucleic acid, construct, or vector is in a cell. Any suitable cell may be used and may be determined by the skilled artisan based on the desired application. The cell may be derived from any organism. Optionally, the cell is a mammalian cell, such as a human cell or a mouse cell. Optionally, the cell is a cell line. The cell line may be any suitable cell line.

[0086] The transcription activator, nucleic acid, construct, or vector can be introduced into the cell in any suitable manner, for example, by transfection. Suitable transfection reagents and methods are routinely practiced in the art and can be identified by the skilled artisan. Optionally, the construct is a viral vector, optionally a lentiviral vector, and is introduced into the cell by transduction. Suitable transduction methods are routinely practiced in the art and can be identified by the skilled artisan.

[0087] In some embodiments, the cells stably express a heterologous transcriptional activator, and optionally the cells are stably transduced, e.g., prepared with a virus that contains a nucleic acid encoding the heterologous transcriptional activator.

[0088] Another embodiment is a transcription activation system comprising the transcription activator described herein, the nucleic acid encoding the transcription activator, or the construct or vector comprising said nucleic acid, or the cell expressing the transcription activator.In the case of the system based on CRISPR-Cas, the system comprises at least one gRNA.In the case of the system based on inducible dimerization domain, the system optionally comprises at least one inducer.

[0089] Also provided are compositions comprising a heterologous transcriptional activator as described herein, a nucleic acid as described herein, a construct as described herein, a vector as described herein, a cell as described herein, and / or a transcriptional activation system as described herein. The compositions can include a carrier, such as BSA, or a suitable diluent depending on the composition components, optionally water or buffered saline. The compositions can include multiple components, such as a transcriptional activator, a nucleic acid, a construct, a vector, or a cell containing the same or different components.

[0090] Also provided herein are kits, e.g., for activating transcription of a target gene or for carrying out a method described herein, comprising a transcription activator as described herein, a nucleic acid, expression construct, or vector encoding a transcription activator as described herein, or a cell expressing a transcription activator as described herein, and optionally a vial containing the transcription activator, nucleic acid, expression construct, vector, cell, or composition. The kit may include a plurality of one or more of the aforementioned components. Optionally, the kit includes a gRNA expression construct, an inducer, and / or instructions for carrying out a method described herein.

[0091] Also described herein is a method for activating the transcription of a target gene in a cell. As shown in the Examples, the transcription activator of the present disclosure can target a genomic locus, such as a promoter, to activate the transcription of a target gene in a cell.

[0092] The transcriptional effectors identified herein may be full-length proteins, fragments thereof (transactivation domains), functional variants thereof, or combinations of transactivation domains or functional variants. They cover multiple different transcriptional activation strengths, from very strong to moderate to weak. This can be used to achieve a desired expression level of an endogenous gene, especially when too high expression may cause a deleterious phenotype. The activation strength can be adjusted by selecting different TADs or functional variants (e.g., active fragments) for inclusion in the effector domain. The activation strength of an effector domain can be determined by one of skill in the art, for example, using the MFI or percent GFP-positive cells in a recruitment assay as shown in the examples described herein. For example, the relative strength of an effector domain can be determined by comparing the MFI or percent GFP-positive cells of a particular effector domain combined with a particular DNA targeting domain and a particular DNA target to a control, such as Renilla, combined with the same DNA targeting domain and a particular DNA target. High activation can be considered, for example, at least 50 times or more than the control, at least 75 times or more than the control, at least 100 times or more than the control, or at least 150 times or more than the control. Moderate activation can be considered, for example, at least 10 times or more than the control, at least 20 times or more, at least 30 times or more, or at least 40 times or more, and up to 50 times, 75 times, 100 times, or up to 150 times, of the control. Low activation can be considered, for example, at least 2 times, at least 2.5 times, at least 3 times, at least 4 times, or at least 5 times, and up to 10 times, 20 times, 30 times, or up to 40 times, of the control. For example, as shown in Figures 19A and 19B, high activation can be considered to be more than 50 times, moderate is 20 times to 50 times, and low activation is at least 3 times to 20 times, relative to the Renilla control. The appropriate activation level can be selected depending on the desired application.

[0093] As described herein, another level of control for transcriptional regulation can be added with, for example, chemically induced dimerization by a rapalog or abscisic acid. In this case, one half (e.g., DNA targeting domain) is fused to FKBP or PYL1, and the other half (e.g., effector domain) is fused to FRB or ABI1. Treatment with a rapalog or abscisic acid induces the interaction between FKBP and FRB or PYL1 and ABI1, respectively, resulting in temporally regulated gene expression. As shown in the examples, dimerization of heterologous transcriptional activators including ABI1 and PYL1 can be induced by the addition of abscisic acid. Those skilled in the art can easily identify and select suitable inducible dimerization domains and inducers that can be used together. Any suitable combination of inducible dimerization domains and inducers can be used, for example, dimerization of ABI1 and PYL1 can be induced by the addition of abscisic acid. Other inducible systems include those based on induction by rapamycin, gibberellic acid / gibberellin, and split dCas9-based systems. For example, the dimerization of GID1 and GAI can be induced by gibberellin, and the dimerization of FKBP and FRB can be induced by rapamycin or its analogs, such as rapalogs. Higher order multimerization systems, such as the SunTag system (Tenenbaum et al., 2014), are also contemplated herein.

[0094] The interaction between the DNA targeting domain and the effector domain can also be controlled using other inducible systems. Other systems (not dependent on dimerization) include grazoprevir-induced stabilization (Tague et al. 2018) or tamoxifen-regulated nuclear localization using estrogen receptor ligand-binding domain variants. In the case of grazoprevir-induced stabilization, the DNA targeting domain and the effector domain are linked by the self-cleaving NS3 protease domain. Only in the presence of grazoprevir (which inhibits NS3 activity) do the DNA-binding domain and the effector domain stay together and regulate gene expression.

[0095] Thus, one aspect of the present disclosure is a method for activating the expression of a target gene in a cell, the method comprising: introducing a transcription activator as described herein into a cell; and culturing the cell under suitable conditions, such that the DNA targeting domain guides the transcription activator to the target site, and the effector domain activates the transcription of the target gene.In one embodiment, the target gene is an endogenous gene.In one embodiment, when the transcription activator comprises CRISPR-Cas, the method further comprises introducing at least one gRNA that targets a desired genomic locus in the cell into the cell; and culturing the cell under suitable conditions, such that the at least one gRNA binds to CRISPR-Cas protein and induces CRISPR-Cas protein to guide the transcription activator to the CRISPR target site, so that the effector domain activates the transcription of the target gene. In embodiments in which the transcriptional activator comprises an inducible dimerization domain in each of the DNA targeting domain and the effector domain, the method further comprises introducing at least one inducing agent into the cell and culturing the cell under suitable conditions in which the first and second inducible dimerization domains bind such that the at least one effector domain activates transcription of the target gene.

[0096] The methods described herein can be used to regulate gene expression of target genes, for example, to induce expression of endogenous genes, or to regulate chromatin opening in defined regions of the genome. For example, some TADs can promote chromatin opening in intergenic regions (i.e., not promoters or enhancers), which can lead to chromatin opening and chromosome folding rearrangements.

[0097] The methods described herein can be used to identify or screen one or more genomic loci that are important for cell viability or a phenotype of interest. As an example, the methods described herein can be used to screen for genes or their regulatory elements that are important for resistance or sensitivity to a toxin of interest, such as diphtheria toxin. In another example, the methods described herein can be used to identify regulatory elements that are important for the expression of a protein of interest, such as CD81. In a further example, the methods described herein can be used in high-throughput screening methods to identify essential or non-essential genes in a cell type by screening for gRNAs that are over- or under-expressed in a cell population under certain conditions, such as drug treatment over time. Other applications can be determined by those skilled in the art.

[0098] The above disclosure generally describes the present application. A more complete understanding can be obtained by reference to the following specific examples. These examples are set forth for illustrative purposes only and are not intended to limit the scope of the present disclosure. Changes in form and substitution of equivalents are contemplated as circumstances may suggest or imply practice. Although specific terms are used herein, these terms are intended to be in a descriptive sense and not for the purpose of limitation. EXAMPLES

[0099] The following non-limiting examples are illustrative of the present disclosure:

[0100] III. Examples Example 1. A platform for identifying transcriptional activators in human cells. To systematically identify transcriptional regulators, a chemical-induced dimerization (CID) system was used, in which catalytically inactive Cas9 (dCas9) is tagged at its N-terminus with the protein phosphatase ABI1, and potential transcriptional activators are fused to the abscisic acid receptor PYL1 (Figure 1A) (Gao et al., 2016; Liang et al., 2011). Treatment of cells with abscisic acid (ABA) induces the interaction between ABI1 and PYL1, thereby recruiting potential transcriptional activators to specific genomic loci defined by guide RNAs (gRNAs). As a reporter, a HEK293T cell line containing a stably integrated construct with a 7xTetO array and a basal CMV promoter driving the expression of EGFP was used (Gao et al., 2016). Monoclonal cell lines expressing ABI1-dCas9 and a single gRNA targeting TetO were generated by lentiviral infection. Transfection of this cell line with known transcriptional activators VPR, VP64, and p300 fused to PYL1 resulted in robust induction of GFP upon ABA treatment ( Figure 1B ), consistent with a previous report ( Gao et al., 2016 ). Recruitment of luciferase or RFP to the promoter did not induce GFP expression, demonstrating that ABA alone does not activate the reporter ( Figure 1B ).

[0101] To scale up from individual clones, we generated pooled libraries of human ORFeome 8.1 and ORFeome collaborative clones (ORFeome Collaboration, 2016; Yang et al., 2011) in lentiviral vectors containing a C-terminal PYL1. Together, these open reading frame (ORF) libraries contain 14,821 clones corresponding to 13,571 unique genes. Reporter cell lines were infected with the pooled libraries at a low multiplicity of infection, thus ensuring that most cells were infected with only one lentivirus (Figure 1C). The pooled libraries covered 96% and 93% of the ORFeome before and after infection, respectively, demonstrating broad coverage of the libraries despite significant diversity in insert size (Figure 6A and B). After treating infected cells with ABA, a clear induction of GFP in a subset of cells was observed (Figure 1D). Reporter induction was dependent on treatment time, ABA concentration, and the presence of the ORFeome (Figure 6C-6E). Furthermore, removal of ABA caused a rapid loss of the high GFP population (Figure 6C), demonstrating that continuous promoter occupancy is required for activation. To identify transcriptional activators, we sorted the top 1% of GFP-positive cells in ABA-treated cells in duplicate and identified ORFs in the high GFP population by sequencing (Figure 1D).

[0102] Methods are as described in Example 7.

[0103] Example 2. ORFeome-wide screening identifies known and novel transcription activators. 248 putative transcriptional activators were identified using a 5% false discovery rate cutoff and at least a 4-fold change in read counts between the top 1% GFP-positive and unsorted cells (Figure 1E). Gene Ontology (GO) analysis revealed significant enrichment for multiple functional categories related to transcriptional activation (Figure 1F). Hits were also highly enriched in protein domains found in many transcriptional regulators (Figure 1G), subunits of chromatin-associated protein complexes (Figure 1H), as well as interactors of the central hub of transcription such as RNA polymerase II and the histone acetyltransferases CBP and p300 (Figure 6F). Furthermore, hits significantly overlapped with human proteins that function as autoactivators in yeast two-hybrid assays (29% in hits vs. 11% in the proteome; p<0.0001, Fisher's exact test; Figure 6G), proteins that activate reporter gene expression in yeast when ectopically recruited to promoters (Luck et al., 2020). The screening results were also validated by assaying a collection of 90 hits and non-hits by transfecting them individually into the same reporter cell line. The results were highly consistent with the screen: almost all hits were reproduced when tested individually, whereas the majority of non-hits did not activate the reporter (Figure 6H), suggesting low false positive and false negative rates in this environment. Taken together, these results indicate that the screen identified functionally relevant transcriptional activators in a reproducible manner.

[0104] Individual screening hits included well-characterized factors that regulate different stages of transcription (Fig. 1E ). For example, among the sequence-specific transcription factor hits, the activators RELB and MYCL (Barrett et al., 1992; Ryseck et al., 1992) were known, in addition to several master TFs regulating stress responses, such as HSF1, ATF6, and DDIT3 / CHOP (Vihervaara et al., 2018). Coactivators that do not bind DNA themselves but bind TFs included STAT2, CITED1, and SERTAD1 (Bousoik and Montazeri Aliabadi, 2018; Hsu et al., 2001; Yahata et al., 2001). Both subunits of TFIIE (GTF2E1 and GTF2E2), a general transcription factor that regulates the assembly of the preinitiation complex (Cramer, 2019), were prominent hits, as were proteins promoting RNA polymerase II release (P-TEFb complex subunit CDK9) and transcription elongation (e.g., ELL3 and MLLT1) (Chen et al., 2019). Other prominent activators are subunits of chromatin-modifying complexes, such as SAGA (ATXN7L3, TADA3, SGF29) and NuA4 (MGRBP), as well as the GADD45 family proteins (GADD45A, GADD45B, GADD45G) that mediate DNA demethylation (Barreto et al., 2007). Finally, we identified 13 of the 15 Mediator subunits that were present in the library. These highlighted examples show that the screen revealed factors that promote transcription at different stages of the transcription cycle by multiple distinct mechanisms.

[0105] In addition to known transcriptional regulators, a large collection of proteins not previously associated with transcription (e.g., C11orf74 / IFTAP, NCKIPSD, DCAF7, HFM1) or not fully characterized (e.g., C3orf62, FAM90A1, SPDYE4, SS18L2, FAM9A, C21orf58) were identified in the screen (Figure 1H). These results suggest that the human genome encodes many previously unknown transcriptional regulators. A detailed characterization of some of these factors is described below.

[0106] Methods are as described in Example 7.

[0107] Example 3. Differential transcriptional activity within TF families. Transcription factors comprise multiple families characterized by their distinct DNA-binding and auxiliary domains (Lambert et al., 2018). TFs belonging to the same family often have highly similar or identical sequence specificities (Jolma et al., 2013; Lambert et al., 2018; Weirauch et al., 2014). Nevertheless, even highly related TFs may have distinct effects on transcription and chromatin due to unique auxiliary domains. Consistent with this, only some members of transcription factor families were identified as hits in the pooled screen. To exclude the sensitivity of the pooled approach as a cause, members of the forkhead box (FOX), SRY-related HMG-box (SOX), E-twenty-six (ETS), atonal-related basic helix-loop-helix (bHLH), Twist / Hand, Krueppel-like factor (KLF), and homeobox protein (HOX) families were assayed individually (Figures 2A and 7A). Also included are two protein families associated with transcription and chromatin (Polycomb group RING finger (PCGF) and the casein kinase family), one family not previously associated with transcription (Spy1 / RINGO) (Gastwirt et al., 2007), and 24 Mediator subunits, which are evolutionarily unrelated but are essential components of the same conserved complex.

[0108] Consistent with the pooled screen, the activation profiles of many highly homologous transcription factors were strikingly different even when tested one at a time. For example, only 5 of 37 forkhead TFs, 1 of 14 SOX, 5 of 14 KLFs, and 2 of 36 HOX proteins activated reporter expression (Figures 2A-2D and 7). To exclude the obvious explanation that activity differences were due to expression, several TFs were assayed as RFP-PYL1 fusions to account for fusion protein expression levels. Results were highly concordant with PYL1-only fusions and did not correlate with RFP levels (Figure 8), indicating that differences in transcriptional activity reflected the intrinsic potential of the TFs.

[0109] Notably, the activating TFs identified in the screen were significantly enriched for factors that could induce differentiation of mouse embryonic stem cells and human induced pluripotent cells (iPSCs) when ectopically expressed (Figure 7G) (Ng et al., 2021; Theodorou et al., 2009), suggesting that activation differences in the assay are related to differences in biological function. For example, among related bHLH family transcription factors, NEUROG1, NEUROG2, NEUROD1, and NEUROD2, but not NEUROD6, can induce neuronal differentiation of iPSCs (Goparaju et al., 2017). This pattern is consistent with their ability to activate reporter genes (Figure 7B). Similarly, only four HOX factors (HOXA1, HOXA2, HOXB1, and HOXB2) can activate the b1-ARE autoregulatory element located at the Hoxb1 locus (Di Rocco et al., 1997), and three of these were characterized as activators in the assay (HOXA2 and HOXB2 in the pooled and arrayed assays, HOXA1 in the screen) (Figure 2A). Furthermore, a recent study revealed a striking collinearity between repressibility, expression patterns, and genomic location of HOX genes (Tycko et al., 2020), with HOX transcription factors at the 5' end of the homeobox cluster being repressive. Consistent with this, HOX family activators in the assay are encoded by the most abundant or last penultimate 3' genes in their HOX gene clusters.

[0110] A particularly interesting case was that of the PCGF family proteins, which are mutually exclusive components of the canonical and non-canonical polycomb repressive complex 1 (PRC1) (Gahan et al., 2020; Gao et al., 2012) (Figure 2E). Although generally thought to act in chromatin compaction and gene silencing in the context of PRC1, PCGF3 was identified in the original screen as an activator. Similar to PCGF5 (which was absent in the original pooled screen), PCGF3 strongly activated the reporter when tested individually (Figure 2E). In contrast, three other PCGF family members (PCGF1, PCGF2, and PCGF4 / BMI1) were neither screen hits nor activated the reporter when tested individually (Figure 2E). PCGF5 had previously been shown to regulate transcriptional activation (Gao et al., 2014), but no such role had been described for PCGF3. Interestingly, in the phylogeny of the PCGF family, PCGF3 and PCGF5 form a distinct group that arose early during animal evolution ( Gahan et al., 2020 ), suggesting that their transcriptional activation function has an ancient origin.

[0111] Some studies have shown that subunits of the Mediator complex may have different regulatory functions, with some subunits promoting transcriptional activation and others promoting repression (Conaway and Conaway, 2011; Stampfel et al., 2015). In our assays, 20 of the 24 Mediator subunits assayed strongly activated the reporter, with no differences between the Mediator submodules (Figure 7F). For example, MED29 has been suggested to have transcriptional repressor activity (Wang et al., 2004), but it was a hit in the pooled screen and was validated as an activator when tested individually (Figure 7F). Thus, at least in the context of the reporter system, almost all Mediator subunits promote transcriptional activation.

[0112] The primary activation screen identified two proteins (SPDYE4 and SPDYE7P) that belong to the Spy1 / RINGO (rapid inducer of G2 / M progression in oocytes) family of cell cycle regulators. Spy1 / RINGO proteins bind and activate Cdk1 and Cdk2 in a cyclin-independent manner, thereby promoting cell cycle progression (Gonzalez and Nebreda, 2020). However, they have not been previously implicated in transcriptional regulation. As a result of recent expansion (Chauhan et al., 2012), the human genome contains at least 19 Spy1 / RINGO family genes and multiple pseudogenes. Five Spy1 / RINGO proteins were tested individually for transcriptional activation, and four of them strongly activated the reporter (Figure 2F). These results suggest that Spy1 / RINGO proteins may promote cell cycle progression in part by functioning as transcriptional activators.

[0113] Methods are as described in Example 7.

[0114] Example 4. TAD-seq reveals novel human transactivation domains. Transcription factors generally activate transcription through transactivation domains (TADs) that interact with coactivators such as Mediator, CBP / p300 acetyltransferase, or TFIID. Most TADs are short unstructured sequences rich in acidic and hydrophobic residues (Sigler, 1988). Despite intensive efforts, no clear consensus motif has emerged, making computational prediction of TADs challenging (Erijman et al., 2020; Ravarani et al., 2018; Staller et al., 2021). Several groups have recently implemented pooled recruitment approaches to identify TADs from known transcription regulators or random sequences (Arnold et al., 2018; Erijman et al., 2020; Ravarani et al., 2018; Sanborn et al., 2020). We modified the above screening platform to identify the region(s) involved in transcription activation among the hits. This approach was similar to a previously published TAD-seq method (Arnold et al., 2018), except that we used synthetic fragments instead of randomly fragmented DNA.

[0115] A fragment library of the 75 activators identified in our screen was generated, with 60 amino acid tiles every 20 amino acids, such that every amino acid was represented by three distinct fragments (Figure 3A). The pooled fragment library was fused to PYL1 and mobilized to the same GFP reporter used in the original ORFeome screen (Figure 3A). Again, ABA-dependent GFP expression was observed in a subset of reporter cells infected with the fragment library (Figure 9A). GFP-positive populations were sorted into two independent replicates, and fragments in each population were quantified by sequencing. However, in contrast to the ORFeome screen, both high GFP cells (top 1%) and moderate GFP cells (top 2-5%) were sorted to potentially identify transactivators of different strengths.

[0116] The pooled approach revealed 70 active fragments in 39 different proteins. As expected, these fragments were enriched in acidic and hydrophobic amino acids and depleted in positively charged (basic) amino acids (Figure 9B), indicating that many represent "canonical" TADs consistent with the "acidic blob" concept (Sigler, 1988). Indeed, active fragments contained more predicted TADs based on two different algorithms (Erijman et al., 2020; Piskacek et al., 2007) (Figure 9C), and predicted TADs were longer in active fragments than in inactive fragments (Figure 9D). However, many active fragments were not predicted by any algorithm, highlighting the need for experimental approaches.

[0117] Of the 70 active fragments, the majority (44 = 63%) were identified only in the high or medium GFP populations, suggesting that the fragments have different activation abilities (Figure 9E). For example, VP64 was enriched in the high GFP population but depleted in the medium GFP population when compared to the unsorted fragment library (Figure 9E). This is consistent with the unusual potency of VP16 in transcriptional activation (VP64 consists of four VP16 peptides in tandem) (Sadowski et al., 1988). Ten fragments identified in the screen were individually tested, and nine of them robustly activated the reporter (Figure 9F), suggesting a low false positive rate.

[0118] The fragment screen retrieved several known TADs. For example, we identified known TADs in KLF6, KLF7, KLF15, ATF6, and CITED2 (Figures 3B and 10A). Furthermore, overlap between active fragments may specify the minimal sequence required for activity, so that the region required for activity of previously identified TADs (e.g., for KLF6, KLF7, and SERTAD2) was narrowed (Figures 3B and 10A). Many previously unknown TADs in both canonical transcription regulators and novel proteins were discovered in the ORFeome-wide screen (Figures 3C and 10B). TADs were revealed in the previously uncharacterized proteins SPDYE4, C3orf62, FAM22F, and FAM90A1 (Figure 3C). Interestingly, the novel TAD in SPDYE4 is adjacent to its Spy1 domain, which interacts with and activates Cdk2, indicating that the potential transcriptional and Cdk regulatory activities are mediated by distinct domains in Speedy family proteins (Figures 3C and S11A) (McGrath et al., 2017). Moreover, consistent with the transcriptional activation profile of Spy1 / RINGO family proteins, the TAD region is conserved in all family members except SPDYC, the only Spy1 / RINGO family protein that was inactive in recruitment assays (Figures 2F and S11A).

[0119] Interestingly, some of the unrevealed TADs did not have the characteristics of a typical transactivation domain. For example, three overlapping fragments in HOXA2 that activated transcription span a polyalanine stretch between the homeobox DNA binding domain and an antennapedia-like hexapeptide motif (Figures 3D and 3E), which interacts with the PBX1 coactivator (Piper et al., 1999). However, fragments lacking the hexapeptide motif still activated transcription, suggesting that activity is not regulated through PBX1 (Figure 3E). Polyalanine stretches are functionally important in Drosophila Ultrabithorastal (Ubx) proteins, which have recently been suggested to have a role in driving phase separation of many TFs, such as HOXD13 (Basu et al., 2020). The results herein suggest that polyalanine stretches may have additional roles in transcription activation.

[0120] Another non-canonical activation region is from YAF2, a component of polycomb repressive complex 1 (PRC1) (Gao et al., 2012) and a prominent hit in the original ORFeome screen (Figure 3F). This region contains the YAF2_RYBP domain, which folds into an antiparallel β-sheet that binds to the core PRC1 subunit RING1B (Wang et al., 2010) (Figure 3G). The YAF2_RYBP domain of YAF2 and its close homolog RYBP share sequence and structural homology with the CBX-C domain present in the CBX family polycomb proteins (Wang et al., 2010) (Figure 11B). CBX-C domain containing proteins and YAF2 / RYBP interact in a mutually exclusive manner with RING1A and RING1B to form canonical (cPRC1) and non-canonical (ncPRC1) Polycomb complexes (Figure 3G).

[0121] To test whether all CBX-C and YAF2_RYBP domains could promote transcriptional activation, the YAF2_RYBP domains of YAF2 (SEQ ID NO: 96) and RYBP (SEQ ID NO: 140) and the CBX-C domains of CBX2 (SEQ ID NO: 136), CBX4, CBX6 (SEQ ID NO: 138), CBX7 and CBX8 (SEQ ID NO: 139) were cloned and assayed for activity by the reporter system. The YAF2_RYBP motifs from both YAF2 and RYBP strongly activated the reporter (Figure 3H), consistent with the TADseq results of YAF2. Some CBX-C domains were also activators: CBX-C from CBX2 was the most potent activator, whereas those from CBX6 and CBX8 weakly activated the reporter (Figure 3H). In contrast, the CBX-C domains of CBX4 and CBX7 had no effect on the reporter. This pattern reflected the evolutionary ancestry of the CBX-C domain proteins (Figure 3H). Interestingly, the CBX-C domains of CBX4 and CBX7 bind RING1B with approximately 10-fold higher affinity than the same domains from transcriptionally activating CBX proteins or the YAF2_RYBP domain of RYBP (Figure 3H) (p=0.008, two-tailed t-test) (Wang et al., 2008, 2010). Thus, the differences in transcriptional activity of CBX-C and YAF2_RYBP domains could be explained by their binding affinity to RING1B or by differential binding to other factors. In either case, these results suggest that differences in CBX protein function could be explained, at least in part, by intrinsic differences in the CBX-C domains (Morey et al., 2012; Vincenz and Kerppola, 2008). More broadly, these results reveal another layer of complexity in the assembly and function of non-canonical and canonical PRC2 complexes.

[0122] Methods are as described in Example 7.

[0123] Example 5. A novel transcriptional activator interacts with a known cofactor. ORFeome-wide screening revealed several potent transactivators that were poorly or incompletely characterized. To understand how these factors regulate transcription, we established stable tetracycline-inducible HEK293 cell lines and expressed nine poorly characterized screening hits (C3orf62, C11orf74 / IFTAP, NCKIPSD, DCAF7, SS18L2, SPDYE4, FAM90A1, FAM22F / NUTM2F, JAZF1), five known transcription regulator hits (SOX7, KLF6, KLF15, CTBP1, HOXA2, and HOXB2), two synthetic transactivators (VP64 and VPR), and negative controls (EGFP and Nanoluc) fused to the biotin ligase BirA from Aquifex aeolicus and a FLAG epitope tag. Those protein interactions were characterized with proximity partners (BioID2) with affinity purification and proximity-dependent biotinylation coupled to mass spectrometry (AP-MS) (Kim et al., 2016). While AP-MS is an ideal method for characterizing stable protein complexes, BioID is superior for identifying interactions involving weaker or less soluble proteins, such as those tightly bound to chromatin (Lambert et al., 2015).

[0124] Transcriptional activator interactions revealed two patterns. First, they showed that the transactivation potential of novel hits likely reflects their intrinsic function rather than an artifact of the tethering assay. Second, the interactome revealed a striking preference of activators for specific coactivator complexes, converging on five distinct cofactors (CBP / p300, BAF, NuA4, Mediator, and TFIID).

[0125] Supporting the natural role of novel activators in transcriptional regulation, eight of the nine poorly characterized hits were associated with known transcriptional cofactors in AP-MS, BioID, or both (Figure 4A, Figure 12). C3orf62, DCAF7, and FAM22F / NUTM2F interacted with known transcriptional coactivators p300 and / or CBP. In contrast, JAZF1 bound to multiple subunits of the NuA4 histone acetyltransferase complex in BioID, suggesting that it is a novel subunit of this highly conserved complex. In turn, SS18L2 interacted with the BAF chromatin remodeling complex, which includes the core BAF members SMARCA2 and SMARCA4, as well as subunits specific for canonical BAF (cBAF) and noncanonical BAF (ncBAF) (Figures 4A and 12) (Centore et al., 2020). SPDYE4 and FAM90A1 interacted with the BET family bromodomain proteins BRD2, BRD3, and BRD4, which regulate transcription elongation (Fujisawa and Filippakopoulos, 2017). NCKIPSD, also known as SPIN90, interacted with the survival of motor neuron (SMN) complex, which regulates the assembly of ribonucleoprotein complexes but is also associated with transcriptional activation (Pellizzoni et al., 2001; Singh et al., 2017; Strasswimmer et al., 1999) (Figure 12). Interestingly, NCKIPSD also interacted with DCAF7 (Figure 12B). Both DCAF7 and NCKIPSD are involved in regulating actin dynamics in the cytoplasm (Cao et al., 2020; Morita et al., 2006), suggesting that these proteins may have different roles in the cytoplasm and nucleus.

[0126] Known transcriptional regulators also bind to the coactivator complex (Figures 4A and 12). The KLF6 bait identified the NuA4 subunit EPC1 as a nearby interactor, whereas HOXB2, KLF15, CTBP1, and SOX7 baits had CBP and p300 as nearby partners. In addition, SOX7 also binds to the BAF subunit. In contrast to all natural activators except SOX7, the potent synthetic activators VP64 and VPR identified multiple different coactivators as nearby partners. VP64, which consists of four tandem copies of the viral VP16 transactivation motif that binds to CBP / p300 and the Mediator subunits MED14 and MED15, is consistent with previous reports (Kundu et al., 2000; Yang et al., 2004). VPR is a fusion of VP64, human p65 / RELA transactivation domain, and Epstein-Barr virus R transactivator (Chavez et al., 2015), and binds a number of additional cofactors, including CBP / p300 and multiple subunits of Mediator and TFIID (Figures 4A and 12A). Such binding of multiple cofactors likely explains the exceptional transactivation potency of VPR when fused to dCas9 (Chavez et al., 2015, 2016).

[0127] These results suggest that activating transcription factors and other transcription regulators have strong inherent preferences for certain cofactors. To investigate this further, we analyzed a previously published AP-MS interaction dataset of forkhead family TFs (Li et al., 2015) and compared it to the transcription activation results in this assay. Notably, only forkhead TFs that activated transcription when recruited to the reporter interacted with coactivators in AP-MS (Figure 4B). In addition, similar to the mass spectrometry results, activation of forkhead TFs had distinct cofactor preferences: FOXO1 and FOXO3 interacted specifically with CBP and p300, FOXN1 interacted with both p300 and subunits of the BAF complex, while FOXR1 and FOXR2 preferred the NuA4 complex (Figure 4B). These data strongly suggest that related transcription factors that recognize highly similar sequences and activate transcription to roughly the same extent may promote transcription through different coactivator complexes.

[0128] To functionally examine the binding between transcriptional activators and cofactors, we arrayed a panel of 83 robust activators. Their activation potential was tested using reporter assays in the presence of small molecule inhibitors targeting multiple transcriptional coregulators. Three kinase inhibitors targeting transcription kinases (flavopiridol for CDK9, CX-4945 for casein kinase 2, and AZ191 for DYRK1A and DYRK1B) and two compounds inhibiting transcriptional cofactors (A-485 for CBP / p300 and JQ1 for BET family bromodomain protein) were used.

[0129] The three kinase inhibitors affected almost all activators, but to different extents. Inhibition of CDK9 with flavopiridol resulted in an almost complete loss of activity of all activators, consistent with a key role for p-TEFb in promoter clearance (Figure 13A). Inhibition of casein kinase 2, which regulates transcriptional elongation (Basnet et al., 2014), caused a moderate but general attenuation of transcriptional activation (Figure 13B). DYRK1A / DYRK1B inhibition had a more subtle effect than the other two compounds, slightly attenuating the activity of most activators (Figure 13C). Interestingly, the two activators that were not affected by the DYRK1A / DYRK1B inhibitor AZ191 were DCAF7, which forms a conserved complex with DYRK1A (Breitkreutz et al., 2010; Yu et al., 2019), and NCKIPSD, which interacts with DCAF7 (Figure 13C).

[0130] In contrast to kinase inhibitors, which have broad effects on transcription, inhibiting the acetyltransferase activity of CBP / p300 strongly affected the activity of some, but not all, transactivators (Figure 4C). Importantly, these effects were consistent with our AP-MS and BioID results: the activity of four of six CBP / p300 interactors was significantly reduced, whereas only one of seven non-interactors was affected by A-485 (Figure 13D). For example, A-485 inhibited the activity of KLF15 but had no effect on the related Kruppel-like factor KLF6 (Figure 13D). More broadly, the activity of proteins known to interact with CBP / p300 was significantly more inhibited by A-485 than that of non-interacting transactivators (Figure 4C). In comparison, interactors of the NuA4 complex were not significantly affected by CBP / p300 inhibition (Figures 4D and 13D).

[0131] Similar to CBP / p300 inhibition, BET family inhibition by JQ1 had different effects on some transactivators. Interestingly, in most cases, JQ1 treatment led to an increase in reporter gene activity (Figure 4D). JQ1 treatment often causes rapid downregulation of BRD4 target genes (Loven et al., 2013; Muhar et al., 2018) and is also associated with transcriptional activation of reporter genes (Sdelci et al., 2016), likely explaining the observed effects. Nevertheless, not all transactivators responded similarly to JQ1 treatment. Notably, factors interacting with subunits of the NuA4 complex were not affected by JQ1 treatment (e.g., JAZF1 and KLF6; Figures 4D and 13D). Indeed, NuA4 interactors were significantly less affected by JQ1 treatment than CBP / p300 interactors or other transactivators (Figure 4D). Although the mechanism by which NuA4 interactors respond to JQ1 treatment in a unique manner requires further study, these results demonstrate how transcriptional activators promote transcription through distinct coactivator complexes in the context of a single promoter. Indeed, hierarchical clustering of activators based on their sensitivity to the five compounds revealed multiple distinct groups (Figure 4E). Many paralogous factors (e.g., CITED1 and CITED2, or PCGF3 and PCGF5) clustered adjacent to each other, indicating that the clustering generated functionally related groups. Furthermore, known CBP / p300 interactors were primarily in two distinct clusters, as were NuA4 interactors (Figure 4E). Other members of these clusters likely similarly use CBP / p300 or NuA4 as coactivators, e.g., YAF2 or MYOG for p300, or NOM1 for NuA4.

[0132] Methods are as described in Example 7.

[0133] Example 6. SRF-C3orf62 fusion interacts with CBP / p300 and promotes the SRF / MRTF transcriptional program. Fusion proteins involving transcriptional regulators are a common feature of certain cancers, such as leukemias and sarcomas. Hits from our ORFeome-wide screen were significantly enriched for genes recorded in the COSMIC database (cancer.sanger.ac.uk) as fusion partners in various cancers (p = 0.019; hypergeometric distribution test). These included well-characterized fusion partners, such as ERG fused to EWSR1 in Ewing's sarcoma and TMPRSS2 in prostate cancer; DDIT3 / CHOP fused to EWSR1 or FUS in myxoid liposarcoma; CRTC1 fused to MAML2 in mucoepidermoid carcinoma; and ENL / MLLT1 fused to MLL in mixed lineage leukemia. In addition, some hits have been described in the literature as fusion partners but have not been functionally characterized. For example, BTBD18 and NCKIPSD were identified as KMT2A / MLL fusion partners in leukemia (Alonso et al., 2010; Sano et al., 2000). Furthermore, a fragment of BTBD18 fused to MLL contains a transactivation domain identified by TAD-seq (Figure 10B), linking TADs to the oncogenic potential of the fusion product. Most MLL fusions involve genes that regulate transcription elongation, such as super elongation factor complexes (Winters and Bernt, 2017). Interestingly, BTBD18 has also been shown to promote transcription elongation (Zhou et al., 2017).

[0134] To gain more insight into the mechanisms by which activation screening hits may promote tumorigenesis as fusion partners, two poorly characterized fusions, JAZF1-SUZ12 and SRF-C3orf62, were selected for further characterization. JAZF1-SUZ12 fusion is a hallmark of low-grade endometrial stromal sarcoma (LG-ESS) (Hrzenjak, 2016) and links the polycomb protein SUZ12 with JAZF1 (Figure 5A) (Piunti et al., 2019). SRF-C3orf62 was recently described in a pediatric case of myofibroma / myopericytoma (Antonescus et al., 2017). In this fusion, the DNA-binding domain of serum response factor (SRF) is fused to the C-terminus of C3orf62 (Figure 5B). Notably, in both cases, the transactivation domains identified by TAD-seq (Figures 3C and 10B) were retained in the fusion constructs (Figures 5A and 5B).

[0135] SUZ12 fused to SRF, SRF, JAZF1-SUZ12, SRF-C3orf62, and the C-terminal fragment of C3orf62 (C3orf62-Cterm) were tagged with BirA-FLAG and their interactions were analyzed by BioID and AP-MS to complement the data obtained for JAZF1 and C3orf62. As expected, SUZ12 proximity partners included other components of the PRC2 complex, such as EZH2, MTF2 / PCL2, and C10orf12 (Alekseyenko et al., 2014) (Figure 5C). Strikingly, the JAZF1-SUZ12 fusion protein had both PRC2 and NuA4 subunits as proximity partners (Figure 5C), indicating that this fusion assembles into a supercomplex of two chromatin-associated complexes that are normally associated with opposing transcriptional activities. Consistent with this, EPC1-PHF1 fusions, which also associate with LG-ESS, similarly assemble the PRC2-NuA4 supercomplex, leading to aberrant expression of polycomb targets (Sudarshan et al., 2021). Moreover, both JAZF1-SUZ12 and EPC1-PHF1 have recently been shown to be potent transcriptional activators (Sudarshan et al., 2021). Thus, incorporation of NuA4 into oncogenic supercomplexes may override the normal repressive function of the PRC2 complex.

[0136] SRF proximity partners included multiple transcription factors such as ELK1, which forms a ternary complex with SRF on the serum response element (Buchwalter et al., 2004) (Figure 5D). However, it did not bind any transcriptional coreactivators. In contrast, both the C3orf62-Cterm construct and the SRF-C3orf62 fusion robustly identified both CBP and p300 as proximity partners (Figure 5D). Similarly, AP-MS identified CBP and p300 as prominent SRF-C3orf62 interactors (Figure S4A). Consistent with these results, both C3orf62-Cterm and SRF-C3orf62 robustly activated the GFP reporter when linked to the reporter (Figure 5E). This activity was highly sensitive to CBP / p300 inhibition by A-485, consistent with the interaction pattern of SRF-C3orf62 (Figure 5F).

[0137] SRF functions with either ternary complex factors (TCFs; e.g., ELK1) or myocardin-related transcription factors (MRTFs; e.g., MAL) to regulate target gene expression (Buchwalter et al., 2004; Olson and Nordheim, 2010). The results suggested that SRF-C3orf62 could activate SRF target genes without such cofactors. To test this further, we assayed the activity of the different constructs in NIH3T3 fibroblasts using a luciferase-based serum response element reporter (Vartiainen et al., 2007). Although SRF or C3orf62 alone did not activate the reporter, SRF-C3orf62 did so potently (Figure 5G), further confirming that the cancer-associated fusions circumvent the requirement for SRF cofactors in transcriptional activation.

[0138] The SRF / TCF pathway, regulated by MAP kinase signaling, regulates the expression of immediate early genes (Gualdrini et al., 2016), whereas the actin-Rho signaling-dependent SRF / MRTF pathway targets genes involved in cell motility and adhesion (Miralles et al. To test whether SRF-C3orf62 could regulate target genes of both pathways, we generated stable doxycycline-inducible NIH3T3 cell lines stably expressing SRF, C3orf62, SRF-C3orf62, and Nanoluc fused to GFP. Transgene expression was induced with doxycycline for 24 h, and gene expression changes were analyzed by RNA-seq. Expression of GFP-tagged C3orf62 or SRF had very limited or no effect on the transcriptome compared to Nanoluc (Figure 14B), whereas expression of SRF-C3orf62-GFP led to the upregulation of 564 genes and the downregulation of 471 genes (log2 fold change >1, FDR <0.05; Figures 5H and 14B). Gene enrichment analysis (GSEA) and gene ontology analysis revealed that the upregulated genes were highly enriched in genes involved in DNA replication, cell cycle progression, and mitosis (Figs. 5I and 14C), whereas the downregulated genes were related to the extracellular matrix (Fig. 14C). The upregulated genes were also significantly enriched in targets of E2F transcription factors, which are master regulators of cell proliferation (Fig. 5I). These signatures are consistent with the oncogenicity of SRF-C3orf62 fusions. Furthermore, many of the most upregulated genes were involved in actin dynamics and myogenesis. For example, one of the most upregulated genes was smooth muscle actin (ACTA2), a known SRF / MTRF target and a hallmark of SRF fusion-positive myofibromas / myopericytomas (Antonescu et al., 2013). et al., 2017) (Figure 5H). Consistent with this, another significant signature of upregulated genes was myogenesis (Figure 5I). In contrast, many well-characterized immediate-early target genes of SRF / TCF, such as FOS, EGR1, or EGR2, were not affected by ectopic SRF-C3orf62 expression (Figure 5H).Interestingly, fusion of SRF to VP16 can potently upregulate FOS, suggesting that there are inherent differences in the ability of different TADs to activate target gene expression (Schratt et al., 2002). Indeed, there was significant overlap between genes upregulated by SRF-C3orf62 and previously reported target genes of SRF / MRTF, but not SRF / TCF (Figure 14D) (Esnault et al., 2014; Gualdrini et al., 2016). Taken together, these results indicate that SRF-C3orf62 expression results in a proliferative and myogenic gene expression signature and preferential expression of SRF / MRTF target genes over SRF / TCF targets.

[0139] Methods are as described in Example 7.

[0140] Example 7. Method: Cell culture. All HEK293T cells, including the pTRE3G-EGFP reporter cell line used for screening (a gift from Lei Stanley Qi lab, Stanford University), were maintained in DMEM containing 10% fetal bovine serum (FBS). NIH-3T3 cells were obtained from Dr. Sachdev Sidhu's lab (University of Toronto) and maintained in Dulbecco's modified Eagle's medium (DMEM) containing 10% bovine calf serum (BCS). All culture media were supplemented with 1% penicillin-streptomycin. Cells were maintained at 37°C with 5% CO2 in a humidified incubator and routinely tested for mycoplasma contamination.

[0141] Lentivirus production. Lentiviral particles containing the pooled ORFeome and transactivation libraries were generated by transfecting 293T cells with pLX301-ORFs / TADs-PYL1, psPAX2 (Addgene #12260), and pVSV-G (Addgene #8454) at a ratio of 8:6:1. Transfections were performed using XtremeGENE 9 (Roche) on 15 cm dishes according to the manufacturer's protocol. 6-8 hours after transfection, the medium was replaced with harvest medium (DMEM + 1.1 g / 100 mL BSA). 72 hours after transfection, the supernatant was filtered (0.45 μM), pooled and harvested. A similar protocol was followed for small-scale virus production where Lipofectamine 2000 (Thermo Fisher Scientific, 11668019) reagent was used to transfect on 6-well plates to establish individual stable cell lines.

[0142] Generation of cell lines. A clonal line of EGFP reporter line (co-expressing EBFP2) expressing ABI-dCas9 (blasticidin, 6 μg / mL) and gRNA (SEQ ID NO: 10) targeting the pTRE3G promoter was generated. Single cells were sorted (FACS Aria IIIu, BD), expanded, and clones showing induction by strong transcriptional activators were selected for subsequent experiments. To generate NIH3T3 cells expressing doxycycline-inducible EGFP-tagged proteins, entry clones were taken from the hORFeome assembly and subcloned into a Gateway-adapted pSTV6-TetO-ccdB-EGFP lentiviral plasmid (kind gift from Payman Samavarchi-Tehrani). NIH-3T3 cells were infected in the presence of 8 μg / mL polybrene and selected with 2 μg / mL puromycin 24 hours post-infection.

[0143] Generation of pooled ORFeome libraries. Entry clones from the human ORFeome assembly (v8.1) were collected into 40 normalized subpools, each containing approximately 384 ORFs, and cloned into the lentiviral Gateway-compatible destination vector pLX301-DEST-PYL1. LR reactions were set up in duplicate with 150 ng of each entry ORF subpool, combined with 1 μl of Gateway LR clonase II in a total reaction volume of 5 μl, and incubated overnight at room temperature in TE buffer. For the next 2 days, 1 μl of additional LR enzyme was added to each reaction in 4 μl TE and 150 ng of destination vector. Colonies were transformed into chemically competent DH5α E. coli and spread on LB agar plates containing carbenicillin (100 μg / μl) at 30° C. overnight. Colonies were counted to ensure greater than 200x coverage, collected in SOC on ice, pelleted, and maxiprepped onto multiple columns based on dry pellet weight.

[0144] Generation of Activation Domain Tiling Library. A tiling library was generated from 75 proteins identified as activators in the ORFeome screen. Oligonucleotides containing the 5' adapter GGAAGTCAGGGTAGCGGAAGTATG (SEQ ID NO:23) and the 3' adapter GGAGGTAGTGTTGAACGCGAAGGC (SEQ ID NO:24) to generate the 5' adapter (GSQGSGSM) (SEQ ID NO:11) and the 3' adapter (GGSVEREG) (SEQ ID NO:12) were synthesized as a pooled library (Twist Biosciences). 6x50μl PCR reactions were set up using NEBNext Ultra II Adapter Q5 Master Mix (New England Biolabs) with 5nM oligos as templates. PCR conditions were optimized to find the lowest cycle with a clean visible product of expected 300bp length. Thermocycling conditions were an initial 30 s at 98°C, then 2 cycles of 98°C 10 s, 63°C 20 s, and 72°C 15 s, followed by 10 more cycles of 98°C 10 s and 72°C 30 s, with a final extension of 72°C 5 min. Primers were designed to have Gateway-compatible flanking sequences. The resulting library was loaded on a 2% TAE gel at 60V for 2 h, then gel extracted by QIAgen gel extraction kit, and then cloned into pDONR221 in a total of 5 μl reactions using 20 separate BP reactions. The entry plasmid pool was transformed into DH5α competent E. coli after overnight reaction and incubated overnight on LB agar plates containing kanamycin (100 μg / μl). Colonies were collected and plasmid DNA was purified. Twenty LR reactions were then set up as described in the previous section. Each reaction was transformed into NEB 10-beta chemically competent E. coli and grown overnight on LB agar plates containing carbenicillin (100 μg / μl) at 30° C. Colonies were counted to ensure greater than 200-fold coverage at each cloning step, pooled, and maxiprepped.

[0145] Pooled activation screening. ORFeome and transactivation tiling libraries tagged C-terminally with PYL1 were packaged into lentiviral particles. A clonal EGFP reporter cell line stably co-expressing ABI-dCas9 and gRNA targeting the promoter were transduced at a low multiplicity of infection (MOI) with approximately 30% cell survival after puromycin (1 μg / mL) selection. Untransduced cells under the same conditions were completely eliminated. Sufficient cells were transduced to maintain >500-fold coverage of the library. Mobilization was induced by treating cells with 100 μM abscisic acid (ABA, Sigma) for 48 h. In parallel, a control batch of cells was treated with an equal volume of DMSO. The cells were then washed with PBS, treated with dissociation buffer (1 mM EDTA, 10 mM KCl, 150 mM NaCl, 5 mM sodium bicarbonate, 0.1% glucose) and resuspended in flow buffer (5 mM EDTA, 25 mM HEPES pH 7, 1% BSA, PBS). High GFP populations for each library, the top 1% for ORFeome, and two bins of the top 1% and next 4% for TAD screening were selected and their genomic DNA was directly extracted using QIAmp DNA Blood Mini Kit (QIAGEN).

[0146] ORFeome Sequencing. Nested PCR was performed using all of the purified genomic DNA from the sorted population or at least 5 μg of genomic DNA from the pre-population. Target ORFeome regions were amplified from genomic DNA using primers targeting the T7 promoter and PYL1. The products of this reaction were pooled for each sample and amplified with primers targeting outside the Gateway attB site for an additional 10 cycles. Amplicons were then separated on a 1% agarose gel and any visible PCR products, except primer dimers, were gel purified. DNA was quantified using the Quant-iT 1X dsDNA HS kit (Thermo Fisher Scientific, Q33232) and then 50 ng per sample was processed using the Illumina DNA Prep, (M) Tagmenination kit (Illumina, 20018705) for 6 cycles of amplification. 2 μl of each purified final library was run on an Agilent TapeStation HS D1000 ScreenTape (Agilent Technologies, 5067-5584). Libraries were quantified using the Quant-iT 1X dsDNA HS kit (Thermo Fisher Scientific, Q33232) and pooled in equimolar ratios after size adjustment. Final pools were quantified using the NEBNext Library Quant kit for Illumina (New England Biolabs, E7630L) and paired-end sequenced on an Illumina MiSeq.

[0147] TAD Sequencing. In the first step, a nested PCR was performed on purified genomic DNA using primers targeting the T7 promoter and PYL1 of the backbone vector to generate a product of approximately 470 bp. The products of the first reaction were then pooled and amplified in 10 further steps using primers targeting outside the Gateway site to generate a product of approximately 300 bp. The library was quantified on a Qubit dsDNA Broad Range kit and paired-end sequenced on an Illumina MIseq using custom PAGE-purified R1 sequencing primers.

[0148] Analysis of sequencing data from pooled activation screens. An index of the ORFeome reference sequence was created using STAR aligner v2.7.8a with the pre-indexing string length set to 11 to account for the smaller "genome" size. Reads from the ORFeome library were aligned with STAR aligner, allowing up to 3 mismatches. For TAD sequencing reads, cloning adaptor sequences were first removed from both ends using cutadapt with CCAGTGTGGAATCTGCAATACAAGTTTGTACAAAAAAGTTGGCGGAAGTCAGGGTAGCGGAAGT (SEQ ID NO: 20) for the 5' adaptor and CCGCCACTTGCTGGATACACTTTGTACTAGAATGAATGGTTGGTTGGGATTCTCGTTCAACTACCTCC (SEQ ID NO: 21) for the 3' adaptor. Bowtie references were generated and reads were mapped using Bowtie v1.2.3, allowing 0 mismatches. To identify activators, we calculated the log2 fold change, p-value, and false discovery rate (FDR) for each ORF by comparing the change in counts from sorted samples to unsorted cells using the edgeR package (Robinson et al., 2010).

[0149] Array-based recruitment assay. Reporter cells stably co-expressing ABI-dCas9 and TetO gRNA were seeded in 48-well or 96-well plates to reach 50-70% confluency on the day of transfection. 150 ng of each construct to be tested was transfected with polyethyleneimine (PEI) at a reagent ratio of 0.6 μl. Recruitment was induced by treatment with ABA (100 μM) the day after transfection. For the linked reporter assay in the presence of inhibitors, recruitment was similarly induced with 100 μM ABA the day after transfection but in the presence of either inhibitors or the same volume of optional additional DMSO. Inhibitors were dissolved in DMSO to a stock concentration of 10 mM. Final concentrations used were 100 nM for flavopiridol, 300 nM for JQ1, 1 μM for A-485, 2.5 μM for CX-4945, and 3 μM for AZ191. All inhibitors were kind gifts from the Structural Genomics Consortium (SGC). 48 hours after induction, cells were dissociated and resuspended in flow buffer using a liquid handling robot (TECAN) and analyzed by LSR Fortessa (BD). Cells were gated on high EBFP2 and RFP as a measure of gRNA and transfection control, respectively. Flow cytometry data were analyzed using FlowJo (v10).

[0150] SRF reporter assay. 30,000 NIH-3T3 cells on 24-well plates were transfected with 8 ng of SRF reporter (p3DA.luc), 20 ng of reference reporter (pcDNA3.1-Nanoluc-3xFLAG-V5) and 50 ng of 3xFLAG tagged construct (Addgene #87063). Transfection was performed using Lipofectamine 3000 reagent (Thermo Fisher Scientific, L3000001) according to the manufacturer's protocol. Luciferase construct was a kind gift from Dr. Maria Vartiainen (University of Helsinki). Cells were maintained in low serum medium (0.5% BCS) for 18 h and stimulated (15% BCS) for 7 h, after which luciferase activity was measured. Data from four independent transfections were used to normalize firefly luciferase to Renilla luciferase activity.

[0151] RNA sequencing and analysis. NIH-3T3 cells harboring stable integration of SRF, C3orf62, SRF-C3orf62, or Nanoluc tagged at the C-terminus with EGFP were induced with 1 μg / mL doxycycline for 24 h. RNA was extracted from cells maintained in low serum conditions (0.5% bovine serum) for 22 h using the RNeasy purification kit (Qiagen) and treated with DNase on-column. Samples were induced and collected in technical duplicates from 6-well plates. Libraries were prepared using NEBNext Ultra II Directional RNA-seq with polyA selection kit, pooled, and sequenced on a 100-cycle NovaSeq 6000 SP. Reads were aligned to the Gencode mouse primary assembly (GRCm39) using STAR aligner v2.7.8a. Counts for each gene were generated using Gencode vM26 transcript annotation. Changes in gene expression relative to cells expressing Nanoluc-EGFP were quantified using the edgeR package ( Robinson et al., 2010 ).

[0152] Mass spectrometry samples. Entry clones were obtained from the human ORFeome assembly (Yang et al., 2011). Clones were introduced into the pDEST-pcDNA5 vector (Kim et al., 2016) with a C-terminal BioID2-FLAG tag using Gateway recombinase. Stable HEK293 Flp-In T-REx cell lines were generated as previously reported (Piette et al., 2021).

[0153] For AP-MS, cells were grown to 70% confluence on 150 mm dishes and then bait expression was induced with 1 μg / mL tetracycline for 24 h. Cells were then washed once with 1x PBS, scraped, pelleted, flash frozen, and stored at -80 °C until processing. AP-MS was performed as previously described (Lambert et al., 2015). Briefly, cells were resuspended in cold lysis buffer (50 mM HEPES-NaOH pH 8.0, 100 mM KCl, 2 mM EDTA, 0.1% NP40, 10% glycerol, 1 mM PMSF, 1 mM DTT, 15 nM calyculin A, and protease inhibitor cocktail (Sigma-Aldrich P8340)) using a 1:4 pellet weight:volume ratio. Cells were lysed by one freeze-thaw cycle and the lysate was sonicated at 4°C using three 10-second bursts with a 2-second pause at 35% amplitude. The sonicated lysate was treated with 100 U benzonase for 30 minutes at 4°C and then cleared by centrifugation at 20,000g for 20 minutes at 4°C. Equal amounts of supernatant from all samples processed in a batch were transferred to a tube containing 25 μL of pre-washed anti-Flag magnetic beads 50% slurry (Sigma, M8823) and incubated for 2 hours at 4°C. The beads were retrieved by magnetization and the supernatant was discarded. As previously described (Taipale et al., 2014), beads were washed once in lysis buffer, once in 20 mM Tris-HCl pH 8.0 containing 2 mM CaCl2, and digested on-bead with trypsin in two steps (1 μg trypsin for 4 h, followed by 0.5 μg trypsin added to the supernatant and incubated overnight at 37 °C). Finally, samples were acidified with 5% formic acid (final concentration) and stored at -80 °C.

[0154] For BioID, cells were grown to 70% confluence in 150 mm dishes and then gene expression was induced with 1 μg / mL tetracycline for 18 hours. 50 μM biotin was then added to each plate for 6 hours. Cell pellets were harvested as for AP-MS and resuspended in lysis buffer (50 mM Tris-HCl pH 7.5, 150 mM NaCl, 0.1% SDS, 1% Igepal CA-630, 1 mM EDTA, 1 mM MgCl2, protease inhibitor cocktail (Sigma-Aldrich P8340, 1:500), and 0.5% sodium deoxycholate) using a 1:10 pellet weight:volume ratio. After sonication, each sample was treated with 250U Turbonuclease (BioVision 9207-50KU) and 1 μL RNase A solution (Sigma-Aldrich R6148) and incubated at 4°C for 30 min. SDS was then added to a final concentration of 0.25% and after mixing, the samples were incubated for an additional 10 min at 4°C, followed by centrifugation at 20,000g for 20 min. The supernatant was transferred to a tube containing 30 μl of pre-washed loaded streptavidin beads (GE Healthcare, 17-5113-01). Streptavidin pull-down was performed at 4°C for 3 h. The beads were washed once in 1 ml SDS buffer (2% SDS / 50 mM Tris-HCl pH 7.5), once in 1 ml lysis buffer, and once in TAP buffer (50 mM HEPES-KOH pH 8.0, 100 mM KCl, 10% glycerol, 2 mM EDTA, 0.1% Igepal CA-630), followed by three washes with 1 ml of 50 mM ammonium bicarbonate pH 8.0. After washing, the beads were resuspended in ABC buffer containing 1 μg trypsin and incubated overnight at 37° C. The next day, the supernatant was collected and the streptavidin beads were washed with 50 μl water, which was combined with the first supernatant fraction. 0.5 μg trypsin was added to the combined supernatant sample, which was then incubated at 37° C. for 4 hours. The beads were then spun down and the supernatant collected. The beads were rinsed twice with ABC buffer and these rinses were combined with the original supernatant. The combined supernatants were dried by centrifugal evaporation.

[0155] Mass spectrometry data acquisition and analysis. Samples were analyzed on a TripleTOF 5600 instrument (AB SCIEX, Concord, Ontario, Canada) using Data-Dependent Acquisition (DDA) as previously described (Piette et al., 2021). Data were processed and analyzed as previously described (Piette et al., 2021) using Proteowizard (Adusumilli and Mallick, 2017) implemented in ProHits v4.0 (Liu et al., 2016), Mascot, and Comet (Eng et al., 2013). Results were then analyzed by the Trans-Proteomic Pipeline using iProphet (Shteynberg et al., 2011), and proteins with iProphet probability ≥ 0.95 were further analyzed.

[0156] Significant interactors were identified with SAINTexpress (Teo et al., 2014). EGFP-BioID2-FLAG and EGFP-BioID2-FLAG were used as negative controls with 2-fold compression at stringency as previously described (Mellacheruvu et al., 2013). SAINTexpress analysis used default parameters and prey proteins were considered significant if they passed the Bayesian FDR cutoff ≤ 5% for the calculation. Dot plot diagrams were generated using the ProHit-viz webserver (Knight et al., 2017).

[0157] Example 8. Multiple TADs can be combined to activate transcription Several TADs were fused together with PYL1 and tested for transcriptional activity using the PYL1-ABI EGFP reporter system described above. The following TAD fragments were used: CITED1-8 (SEQ ID NO: 47), CITED2-12 (SEQ ID NO: 49), C3orf62-12 (SEQ ID NO: 45), BRD8-25 (SEQ ID NO: 40), ZXDC-12 (SEQ ID NO: 97), KLF7-1 (SEQ ID NO: 72), ATXN7L3-1 (SEQ ID NO: 35), FAM90A1-20 (SEQ ID NO: 60), SPDYE4-3 (SEQ ID NO: 90), YAF2-7 (SEQ ID NO: 96).

[0158] As shown in FIG. 16, the SPDYE4-CITED1 TAD fusion protein led to a significant increase in transcriptional activation compared to either TAD alone.

[0159] Up to two additional TADs selected from p65 and HSF1 were combined in different orders with the SPDYE4-CITED1 TAD and tested for their ability to activate transcription of the PYL1-ABI EGFP reporter system (Figure 17, left panel) or the endogenous CD133 gene (Figure 17, right panel). The SPDYE4-CITED1-P65-HSF1 (SCPH) effector produced the strongest activation in both systems tested.

[0160] To test additional TAD combinations, each portion of the multi-component SPDYE4-CITED1-P65-HSF1 (SCPH) effector was individually replaced with another TAD identified in Example 4 (see, e.g., Table 1). The results are shown in FIG.

[0161] To test whether shorter TAD sequences could be used in SCPH effectors, each TAD was independently replaced with a short or mini TAD component. The results are shown in Figure 15. The mini-TAD sequences were selected as follows:

[0162] miniSPDYE4 (SEQ ID NO: 102): The sequence was selected based on the overlap between two enriched fragments from our tiling screen of the SPDYE4 full-length protein. The minimal sequence was designed to include either the acidic-rich region or the following β-hairpin:

[0163] miniHSF1: The 150 amino acid C-terminal region of HSF1, including its activation domain, contains two known TADs. The longer TAD is between amino acids 431-529, and the shorter one ("mini-HSF1" or "H(mini)"; SEQ ID NO: 119) is between amino acids 401-420. We chose the short domain, rich in hydrophobic residues, to test because of its compact size and, as previously reported, has higher potency than the longer TADs (Newton et al., 1996; 10.1128 / MCB.16.3.839).

[0164] miniP65: RelA or p65 is known to have two distinct transactivation domains within its C-terminus: the first TAD ("P(mini N-term)"; SEQ ID NO: 105) encompasses amino acids 428-520, and the second ("P(mini C-term)"; SEQ ID NO: 106) precedes residues 521-551 (Schmitz and Baeuerle, 1991; 10.1002 / j.1460-2075.1991.tb04950.x).

[0165] miniCITED1 ("C(mini)"; SEQ ID NO: 103) was selected to contain a C-terminal acidic-rich region within the overlapping region between CITED1-7 and CITED1-8.

[0166] Methods: Cells were seeded in 48-well plates. The next day, after cells were between 70-90% confluent, 250ng of each construct was transfected into each well using (Thermo Fisher Scientific, 11668019) reagent. Each construct was either directly fused to TagRFP, directly fused to each component, or co-expressed from the same plasmid used as the means of transfection. In each experiment, the same gate of RFP+ cells was used to control for the effect of expression levels of each construct on activity. EGFP reporter cells were expanded from clonal lines to control for the levels of ABI-dCas9 and gRNA targeting the TetO7 site upstream of the promoter. For CD133 induction experiments, a 293T cell line stably expressing ABI-dCas9 and a pool of five gRNAs targeting the promoter were used for all activators. CD133 antibody conjugated to APC (Miltenyi Biotec, 130-113-668) was used. All activator constructs were fused to PYL1 at their C-terminus and recruited to their targets by treating cells with 1 μM abscisic acid for 24 or 48 h.

[0167] Example 9. An additional 117 different combinations of activation domains or fragments of individual activation domains were assayed by fusing them to either PYL1 (to test in the dCas9-ABI1 system) or rTetR, a transcription factor that binds to DNA in the presence of doxycycline. All constructs were tested with the same TetO reporter construct. Large differences in the activity of these constructs were observed, ranging from no activity to very high potency (Figure 19A-B). In general, there was a good correlation between the results with the rTetR and dCas9-based recruitment systems (Figure 19C), suggesting that most activation domains function similarly in both contexts. However, for example, VPR was much less active as an rTetR fusion than as a dCas9 fusion. Consistent with previous experiments, SCPH and its variants were the most active constructs in the system. For example, a construct containing a miniaturized activation domain of C3orf62 instead of the activation domain from CITED1 ("C" in the SCPH construct) and a shorter version of the p65 activation domain ("P") was more potent than the original SCPH construct, despite being significantly smaller (1059 bp vs. 1611 bp) (Figure 19A-B).

[0168] Methods: 200 ng of plasmid expressing the superactivator fused to rTetR at its C-terminus and 50 ng of DsRed were transfected into HEK293T cells stably expressing the 7xTetO-EGFP reporter. One day after transfection, cells were treated with 1 μg / ml doxycycline for 48 h. After treatment, cells were analyzed for EGFP expression by flow cytometry.

[0169] Although the present application has been described with reference to what are currently considered to be the preferred examples, it should be understood that the present application is not limited to the disclosed examples. To the contrary, the present application is intended to cover various modifications and equivalent variations that fall within the spirit and scope of the appended claims. The claims should not be limited by the preferred embodiments and examples, but should be accorded the broadest interpretation consistent with the description as a whole.

[0170] All publications, patents, and patent applications are herein incorporated by reference in their entirety to the same extent as if each individual publication, patent, or patent application was specifically and individually indicated to be incorporated by reference in its entirety. Sequence Listing Glycine serine linker (SEQ ID NO: 6) SGGSGGS Glycine serine linker (SEQ ID NO: 7) SGGS Glycine serine linker (SEQ ID NO: 8) GSGSGS Linker (SEQ ID NO: 9) INSRSSGS NLS (SEQ ID NO:22) PKKKRKV Linker + NLS (SEQ ID NO: 141) INSRSSGSPKKRKVGS >YAF2_RYBP domain of YAF2 (SEQ ID NO: 147) RPRLKNVDRSSAQHLEVTVGDLTVIITDFKEKT >YAF2_RYBP domain of RYBP (SEQ ID NO: 148) RPRLKNVDRSTAQQLAVTVGNVTVIITDFKEKT >CBX-C domain of CBX2 (SEQ ID NO: 149) QDWKPTRSLIEHVFVTDVTANLITTVTVKESPTSV >CBX-C domain of CBX6 (SEQ ID NO: 150) GDWRPEMSPCSNVVVTDVTSNLLTVTIKEFCNPE >CBX-C domain of CBX8 (SEQ ID NO: 151) ESWSPSLTNLEKVVVTDVTSNFLTVTIKESNTDQ >CBX-C domain of CBX4 (SEQ ID NO: 152) SEFKPFFGNIIITDVTANCLTTVTFKEYVTV >CBX-C domain of CBX7 (SEQ ID NO: 153) PPWTPALPSSEVTVTDITANSITTVTFREAQAAE >SPDYE4-2 (SEQ ID NO: 154) VRSPEVVVDDEVPGPSAPWIDPSPQPQSLGLKRKSEWSDESEEELEEELELERAPEPEDT >SPDYE4-5 (SEQ ID NO: 155) WVVETLCGLKMKLKRKRASSVLPEHHEAFNRLLGDPVVQKFLAWDKDLRVSDKYLLAMVI [Table 1-1] [Table 1-2] [Table 1-3] [Table 1-4] [Table 1-5] [Table 1-6] [Table 1-7] [Table 1-8] [Table 1-9] [Table 1-10]

Table 1-11

Table 1-12

Table 1-13

Table 1-14

Table 1-15

Table 1-16

Table 1-17

Table 1-18

Table 1-19

Table 1-20

Table 1-21

Table 1-22

Table 1-23

Table 1-24

Table 1-25

Table 1-26

Table 1-27

Table 1-28

Table 1-29

Table 1-30

Table 1-31

Table 1-32

Table 2-1

Table 2-2

Table 3-1

Table 3-2

Table 3-3

Table 3-4

Table 3-5

Table 4-1

Table 4-2

Table 4-3

Table 5

Claims

**Claim 1** A heterologous transcriptional activator comprising: a DNA targeting domain, an optionally enzymatically inactive CRISPR-CAS protein, a zinc finger DNA binding domain, a tet-repressor, or a transcription activator-like effector (TALE) DNA binding domain; and an effector domain comprising a transactivation domain (TAD) or a functional variant thereof listed in Table 2 or Table 6, optionally at least one TAD selected from Table 6, or a TAD listed in Table 1 or Table 3, or a functional variant thereof, preferably at least two TADs selected from the TADs listed in Table 4 or Table 5 or Table 6, and / or at least one TAD selected from a functional variant thereof; wherein the DNA targeting domain and the effector domain are operably linked; said heterologous transcriptional activator. **Claim 2** The transcriptional activator according to claim 1, wherein the effector domain comprises at least three or at least four transactivation domains selected from the TADs listed in Table 1 or Table 3 or a functional variant thereof. **Claim 3** The transcriptional activator according to claim 1, further comprising at least one interaction component. **Claim 4** The transcriptional activator according to claim 1, wherein the DNA targeting domain and the effector domain are domains of a single polypeptide. **Claim 5** a first polypeptide comprising the DNA targeting domain and a first interaction component; and a second polypeptide comprising the effector domain and a second interaction component; wherein the first and second interaction components interact under appropriate conditions; optionally, the first and second interaction components form an inducible heterodimer pair that interacts under inducible conditions, optionally forming ABI1 and PYL1; the transcriptional activator according to claim 3. **Claim 6** The transcriptional activator according to claim 1, wherein the DNA targeting domain comprises a zinc finger domain. **Claim 7** The effector domain is at least one TAD selected from any one of the TADs of SEQ ID NOs: 103, 104, 105, 106, 167, and 185, and optionally at least one TAD selected from any one of the TADs of SEQ ID NOs: 90, 91, 46, 47, 101-106, 110, 116-119, 156, 157, 159, 162, 165, 166, 167, and 172. The transcriptional activator according to claim 1.

8. The transcriptional activator according to claim 1, further comprising one or more nuclear localization signals (NLS), and optionally the SV40 NLS.

9. The effector domain is the amino acid sequence of SEQ ID NOs: 121, 123, 125, 127, 129, 131, 133, 135, 174, 176, 178, 180, or 182, or contains at least 80%, 85%, 90%, 95%, or 99% sequence identity to the TAD therein. The transcriptional activator according to claim 1.

10. An isolated nucleic acid encoding the transcriptional activator or effector domain according to any one of claims 1 to 9.

11. An expression construct comprising the nucleic acid according to claim 10 operably linked to one or more promoters and one or more transcription termination sites.

12. A vector comprising the nucleic acid according to claim 10 or an expression construct encoding the nucleic acid operably linked to one or more promoters and one or more transcription termination sites, optionally wherein the vector is an adenovirus or lentivirus vector. The vector.

13. A cell comprising the transcriptional activator according to any one of claims 1 to 9, the nucleic acid encoding the transcriptional activator, an expression construct comprising the nucleic acid operably linked to one or more promoters and one or more transcription termination sites, or a vector comprising the nucleic acid or the expression construct.

14. A transcriptional activation system, comprising a) a heterologous transcriptional activator according to any one of claims 1 to 9, wherein the DNA targeting domain comprises a CRISPR-Cas protein, and b) at least one gRNA The transcriptional activation system.

15. The transcriptional activation system according to claim 14, wherein the at least one gRNA targets a regulatory element of a gene, and optionally, the regulatory element is a promoter region, an enhancer region, or a distal regulatory site.

16. A method for activating the transcription of a target gene in a cell, the method comprising: a) introducing into the cell the transcriptional activator according to any one of claims 1 to 9, the nucleic acid encoding the transcriptional activator, an expression construct comprising the nucleic acid operably linked to one or more promoters and one or more transcription termination sites, or a vector comprising the nucleic acid or the expression construct; b) culturing the cell under suitable conditions such that the effector domain activates the transcription of the target gene.

17. The method according to claim 16, wherein the DNA targeting domain comprises a CRISPR-Cas protein, and the method further comprises introducing at least one gRNA into the cell and culturing the cell under suitable conditions such that the at least one gRNA binds to the CRISPR-Cas protein and induces the transcriptional activator to a CRISPR target site.

18. A screening assay, the assay comprising: a) introducing into a plurality of cells the transcriptional activator according to any one of claims 1 to 9, one or more nucleic acids encoding the transcriptional activator, one or more expression constructs comprising the one or more nucleic acids operably linked to one or more promoters and one or more transcription termination sites, or one or more vectors comprising the nucleic acid or the expression construct, wherein the DNA targeting domain comprises a CRISPR-Cas protein; and introducing a plurality of gRNAs, or introducing a plurality of gRNAs into a cell population comprising the transcriptional activator, the one or more nucleic acids, the one or more expression constructs or the one or more vectors, wherein the DNA targeting domain comprises a CRISPR-Cas protein; b) culturing the plurality of cells such that the one or more gRNAs bind to the CRISPR-Cas protein and induce the transcriptional activator to a CRISPR target site such that the effector domain activates the transcription of the target gene. c) optionally, treating with a quantity of test drug or toxin; d) optionally, culturing the plurality of cells over a period of time that allows for dropout or enrichment of the gRNA; e) collecting the plurality of cells or a subset thereof; f) optionally, identifying one or more gRNAs that are over- or under-expressed in the plurality of cells or a subset thereof, said screening assay.

19. A composition comprising a transcriptional activator according to any one of claims 1 to 9, a nucleic acid encoding said transcriptional activator, an expression construct comprising said nucleic acid operably linked to one or more promoters and one or more transcription termination sites, a vector comprising said nucleic acid or said expression construct, or a cell comprising said transcriptional activator, said nucleic acid, said expression construct or said vector, and a pharmaceutically acceptable carrier or diluent.

20. A kit comprising a vial and a heterologous transcriptional activator according to any one of claims 1 to 9, a nucleic acid encoding said transcriptional activator, an expression construct comprising said nucleic acid operably linked to one or more promoters and one or more transcription termination sites, a vector comprising said nucleic acid or said expression construct, a cell comprising said transcriptional activator, said nucleic acid, said expression construct or said vector, or said transcriptional activator, said nucleic acid, said expression construct, said vector or said cell, and optionally one or more of an inducer, a gRNA, or a gRNA expression construct.