Optimization of human transcription factors

By modifying the periodicity and number of aromatic amino acid residues in the IDRs of transcription factors, it is possible to customize their transcriptional activity, addressing the challenges of current TF models and enhancing their functional capabilities.

WO2025120143A1PCT designated stage expired Publication Date: 2025-06-12MAX PLANCK GESELLSCHAFT ZUR FOERDERUNG DER WISSENSCHAFTEN EV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
PCT/EP2024/085045
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-12-08
Filing Date
2024-12-06
Publication Date
2025-06-12

AI Technical Summary

Technical Problem

Current models suggest that the activity and specificity of transcription factors (TFs) are encoded in separate protein portions, making it challenging to understand and modify their functional trade-offs, particularly in their intrinsically disordered regions (IDRs).

Method used

By altering the periodicity and/or number of aromatic amino acid residues in the IDRs of transcription factors, it is possible to modify their transcriptional activity, either increasing or reducing it compared to the wild-type TFs.

Benefits of technology

This approach allows for the rational design of transcription factors with customized reprogramming and functional capabilities, enhancing or reducing their transcriptional activity as needed, which can be beneficial for various cellular processes and applications.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure IMGF000050_0001
    Figure IMGF000050_0001
  • Figure IMGF000052_0001
    Figure IMGF000052_0001
  • Figure IMGF000013_0001
    Figure IMGF000013_0001
Patent Text Reader

Abstract

The present invention relates to methods for modifying the activity of transcription factors by altering the periodicity and / or number of the aromatic amino acid residues in their intrinsically disordered regions (IDRs) thereby obtaining altered transcription factors having modified, e.g., increased and / or reduced transcriptional activity. Further, the present invention relates to novel altered transcription factors and nucleic acid molecules encoding those altered transcription factors. Furthermore, applications for the altered transcription factors, e.g., in cell programming methods, are disclosed.
Need to check novelty before this filing date? Find Prior Art

Description

Optimization of human transcription factorsSpecificationField of the InventionThe present invention relates to methods for modifying the activity of transcription factors by altering the periodicity and / or number of the aromatic amino acid residues in their intrinsically disordered regions (IDRs) thereby obtaining altered transcription factors having modified, e.g., increased and / or reduced transcriptional activity. Further, the present invention relates to novel altered transcription factors and nucleic acid molecules encoding those altered transcription factors. Furthermore, applications for the altered transcription factors, e.g., in cell programming methods, are disclosed.Background of the InventionCell-specific transcriptional programs in metazoans are established by transcription factors (TFs) binding specific DNA elements mostly within transcriptional enhancers1'3. Small sets of TFs are able to reprogram the identities of various cell types, but the principles how thousands of enhancers and hundreds of transcription factors active in any cell type interact to produce cell-specific transcriptional programs are largely unknown3-5. One major challenge is that virtually all genome-scale efforts to-date have focused on characterizing sequences in enhancers and transcriptional regulators that have strong transcriptional activity measured in gene reporter systems6'13. However, emerging evidence suggest that critical developmental information is encoded in enhancers that drive weak, tissue-specific expression patterns in developing embryos14-16. Such weak enhancers contain suboptimal DNA binding motifs and spacing, and mutant enhancers with optimized motifs drive elevated but less-specific patterns of transcription, leading to developmental defects14-17. These results suggest an important evolutionary trade-off between activity and specificity encoded within weak enhancers, also referred to as “suboptimization”14. Whether such a trade-off is encoded in TFs themselves is unclear. If yes, understanding the sequence features that encode such a trade-off could enable rational design of natural TF variants with customized reprogramming and other functionalities.Investigating functional trade-offs in TFs is impeded by current models that TF specificity and activity are encoded in separate protein portions. The activity of mammalian TFs is thought to be mediated by sequence motifs that comprise a ‘minimal’ activation domain, that is distinct from the DNA-binding domain (DBD) that determines binding specificity (Fig. 1a). Minimal activation domains are typically short (9-40 amino acids) and tend to assume secondary structure upon binding to co-activators10’11’13. However, the minimal activation domains are almost invariably embedded within much longer intrinsically disordered regions (IDRs) that do not have stable secondary structure (Fig. 1a)6-11. An emerging view suggests that TF IDRs may contribute to transcriptional activity by engaging in multivalent, weak interactions. Such interactions can drive phase separation of TFs in vitro and partitioning of TFs into condensates (also referred to as ‘hubs’) enriched in co-activators and RNA Polymerase II in cells18-21. Whether the ability of TFs to form condensates is important for their in vivo function is debated18-20’22. Nevertheless, deleting I DRs of yeast TFs was shown to reduce genomic binding23, suggesting that TF IDRs may contribute to transcriptional activity and also to binding specificity in vivo.US 2022 / 120736 refers to condensates and their components including transcription factors, methods for identifying agents that modulate condensate structure and function, and methods for modulating condensate function activity. Example 3 (

[0650] -

[0676] ) describes the formation of phase-separated droplets with wild-type and modified transcription factors OCT4 and GCN4.In

[0664] , mutants of the mouse transcription factor OCT4 were prepared in which all acidic amino acids or all aromatic amino acid residues in the I DR were replaced with alanine. The acidic OCT4 mutant lost its function to activate transcription in a GAL4 transactivation assay. The aromatic OCT4 mutant still incorporated into MED1-IDR droplets. Apparently, removal of the aromatic amino acids has no effect.In

[0675] , a mutant of the yeast transcription factor GCN4 was prepared in which the 11 aromatic residues contained in hydrophobic patches of the activation domain were changed to alanine. According to

[0676] the GCN4 aromatic mutant was tested in a GAL4 transactivation assay and found to have lost its capability of transcriptional activation in vivo.The examples of US 2022 / 120736 teach that removal of all aromatic amino acid residues in the I DR of a transcription factor either has no effect on the transcriptional activity (OCT4) or completely abolishes transcriptional activity (CGN4). There is no disclosure of transcription factors with increased transcriptional activity or transcription factors having reduced detectable activity.Takahashi et al. (Cell 126 (2006), 663-676) describe the use of transcription factors in the reprogramming of cells such as to obtain induced pluripotent stem cells. The use of modified transcription factors is not disclosed or suggested.Martin et al. (Science 367 (2020), 694-699) describe that altering the valence and periodicity of aromatic amino acid residues determines the phase-behavior of prion-like domains of RNA- binding proteins. There is no specific mentioning of transcription factors. Further, it is not shown that enhancing the periodicity of aromatic amino acid residues in the I DR of an RNA-binding protein modifies its capacity to form condensates (c.f. Fig 4D).WO 2021 / 219721 , the content of which is herein incorporated by reference, discloses determining the capacity of a Cluster 1 mammalian TF for phase separation and / or the capacity for forming a transcriptional condensate, including determining the presence, localization and / or morphology of a transcriptional condensate comprising said TF, and / or determining the composition of a transcriptional condensate comprising at least one Cluster 1 mammalian TF and / or determining the transcriptional activity of the TF or a condensate comprising the TF.In this work we set out to investigate whether human TF IDRs are suboptimized (i.e., their activity and specificity are submaximal because they are in a trade-off). To do so, we took inspiration from recent insights into prion-like IDRs of RNA-binding proteins, to identify a single sequence feature in human TFs IDRs that contributes to both transcriptional activity and binding specificity. Prion-like IDRs of RNA-binding proteins (e.g., FUS, HNRNPA1 , TDP-43), encode regularly spaced aromatic residues, whose number and uniform spacing promote the proteins’ phase separation capacity2425. We provide evidence that hundreds of TF IDRs encode elements of aromatic periodicity. Optimizing aromatic dispersion enhanced the activity and reduced specificity of several TFs, with consistent changes in their in vitro phase separation behavior. Since most human proteins contain IDRs, this principle may be generalizable for many biopolymers.Summary of the InventionA first aspect of the present invention relates to a method of modifying the activity of a transcription factor comprising the steps:(a) providing a transcription factor including an intrinsically disordered region (I DR)(b) altering the periodicity and / or number of the aromatic amino acid residues in the I DR, thereby obtaining an altered transcription factor,(c) optionally measuring the transcriptional activity of the altered transcription factor, and(d) obtaining an altered transcription factor with modified transcriptional activity.In certain embodiments, the IDR comprises at least 4 aromatic amino acid residues dispersed therein.In certain embodiments, the method comprises increasing the periodicity and / or number of the aromatic amino acid residues in the IDR, thereby obtaining an altered transcription factor with increased transcriptional activity.In certain embodiments, the method comprises reducing the periodicity and / or number of the aromatic amino acid residues in the IDR, thereby obtaining an altered transcription factor with reduced transcriptional activity.A further aspect of the present invention relates to an altered transcription factor with modified transcriptional activity.In certain embodiments, the altered transcription factor is obtainable by a method as described herein.In certain embodiments, the altered transcription factor has an increased transcriptional activity.In certain embodiments, the altered transcription factor has a reduced transcriptional activity.In certain embodiments, the transcriptional activity is increased or reduced compared to the respective wild-type transcription factor.A further aspect of the present invention relates to the use of an altered transcription factor with modified transcriptional activity in research, diagnostics, and medicine.In certain embodiments, the altered transcription factor with modified transcriptional activity is used in transcription factor-mediated cell reprogramming.Embodiments of the specificationIn the following, certain embodiments of the specification are disclosed.1. A method of modifying the activity of a transcription factor comprising the steps:(a) providing an initial transcription factor including at least one intrinsically disordered region (I DR),(b) altering the periodicity and / or number of aromatic amino acid residues in the I DR, thereby obtaining an altered transcription factor,(c) optionally measuring the transcriptional activity of the altered transcription factor, and(d) obtaining an altered transcription factor with modified transcriptional activity.2. The method of embodiment 1 wherein the initial transcription factor is a mammalian transcription factor, particularly a human transcription factor.3. The method of embodiment 1 or 2 wherein the initial transcription factor comprises- at least one DNA binding domain (DBD) and- at least one intrinsically disordered region (I DR).4. The method of any one of the preceding embodiments wherein the initial transcription factor comprises an IDR including at least 4 aromatic amino acid residues dispersed therein.5. The method of any one of the preceding embodiments, wherein the aromatic amino acid residues are selected from the group consisting of Y, W, F, and any combination thereof.6. The method of any one of the previous embodiments wherein the transcriptional activity is measured in a GAL4 DBD-transactivation reporter system.7. The method of any one of the preceding embodiments, comprising:(a) providing an initial transcription factor including an IDR, particularly including an IDR comprising at least 4 aromatic amino acid residues dispersed therein,(b) increasing the periodicity and / or number of aromatic amino acid residues in the IDR, thereby obtaining an altered transcription factor,(c) optionally measuring the transcriptional activity of the altered transcription factor, and(d) obtaining an altered transcription factor with increased transcriptional activity.8. The method of embodiment 7, wherein the altered transcription factor with increased transcriptional activity is selected from the group consisting of an alteredHOXD4, HOXC4, OCT4, C / EBPa, NGN2, MYOD-1 , PDX1 , FOXA3, BRACHYURY (TBXT), MSGN1 , NKX2-5 and CDX2 transcription factor.9. The method of any one of embodiments 7-8, wherein step (b) comprises(i) increasing the periodicity by rearranging the positions of aromatic amino acid residues originally present in the I DR; or(ii) increasing the number of aromatic amino acid residues originally present in the IDR; or(iii) a combination of (i) and (ii).10. The method of any one of embodiments 7-9, wherein increasing the number of the aromatic amino acid residues comprises inserting at least one aromatic amino acid into the IDR such that the altered transcription factor comprises a number of aromatic amino acid residues in the IDR which is at least 1 , at least 2, at least 3 or at least 4 greater than the number of aromatic amino acids in the IDR of the initial transcription factor.11 . The method of any one of embodiments 7-10, wherein increasing the periodicity of the aromatic amino acid residues comprises positioning aromatic amino acids within the IDR such that a substantially uniform spacer length between individual aromatic amino acids is obtained, e.g., wherein the spacer length between individual aromatic amino acids deviates by up to 2, up to 1 or 0 amino acids throughout the IDR.12. The method of embodiment 11 , wherein the substantially uniform spacer length between individual aromatic amino acids is between 4 and 30 amino acids, particularly between 6 and 20 amino acids.13. The method of any one of embodiments 1-6, comprising the steps:(a) providing an initial transcription factor including an IDR, particularly including an IDR comprising at least 4 aromatic amino acid residues dispersed therein, and(b) reducing the periodicity and / or number of aromatic amino acid residues in the IDR, thereby obtaining an altered transcription factor,(c) measuring the transcriptional activity of the altered transcription factor, and(d) obtaining an altered transcription factor with reduced transcriptional activity.The method of any one of embodiments 1-6 and 13, wherein the altered transcription factor with reduced detectable transcriptional activity is selected from the group consisting of an altered HOXD4, HOXC4, C / EBPa, MYOD-1 , HOXB1 , EGR1 , NFAT5, NANOG, and OCT4 transcription factor. The method of any one of embodiments 1-6 and 13-14, wherein step (b) comprises(i) decreasing the periodicity by rearranging the positions of aromatic amino acid residues originally present in the I DR; or(ii) decreasing the number of aromatic amino acid residues originally present in the IDR; or(iii) a combination of (i) and (ii). The method of any one of embodiments 1-6 and 13-15, wherein step (b) comprises replacing at least one aromatic amino acid with a non-aromatic amino acid. The method of embodiment 16, wherein the non-aromatic amino acid is a neutral non-aromatic amino acid, particularly selected from the group consisting of A, S, G, and any combination thereof. The method of any one of embodiments 1-6 and 13-17, wherein the altered transcription factor comprises a number of aromatic amino acid residues which at least 1 , at least 2, at least 3 or at least 4 lower than the number aromatic amino acids in the initial transcription factor. An altered transcription factor including an intrinsically disordered region (IDR), which comprises an altered number and / or periodicity of aromatic amino acid residues in the IDR compared to a corresponding wild-type transcription factor. The altered transcription factor of embodiment 19 wherein the corresponding wildtype transcription factor comprises an IDR including at least 4 aromatic amino acid residues. The altered transcription factor of embodiment 19 or 20 which has a modified transcriptional activity compared to the corresponding wild-type transcription factor. An altered transcription factor with modified transcriptional activity obtainable by the method of any one of embodiments 1-18.23. The altered transcription factor of any one of embodiments 19-22, which has an increased transcriptional activity compared to the corresponding wild-type transcription factor.24. An altered mammalian transcription factor including an intrinsically disordered region (I DR), which comprises an increased number and / or periodicity of aromatic amino acid residues in the I DR compared to a corresponding wild-type transcription factor, wherein the altered transcription factor has an increased transcriptional activity compared to the corresponding wild-type transcription factor and wherein the corresponding wild-type transcription factor particularly comprises an I DR including at least 4 aromatic amino acid residues.25. The altered transcription factor of any one of embodiments 19-24 having a transcriptional activity that is significantly increased versus the transcriptional activity of the corresponding wild-type transcription factor as measured in a GAL4 DBD- transactivation reporter system.26. The altered transcription factor of any one of embodiments 19-25 having a transcriptional activity that is, e.g., by at least 20%, by at least 50%, by at least 75% or by at least 100% increased versus the transcriptional activity of the corresponding wild-type transcription factor as measured in a GAL4 DBD-transactivation reporter system.27. The altered transcription factor of any one of embodiments 19-26 including an intrinsically disordered region (I DR) comprising at least 4 aromatic amino acid residues dispersed therein, wherein there is a substantially uniform spacer length between individual aromatic amino acids, e.g., wherein the spacer length between individual aromatic amino acids deviates by up to 2, up to 1 or 0 amino acids throughout the I DR.28. The altered transcription factor of embodiment 27 wherein the substantially uniform spacer length between individual aromatic amino acids is between 4 and 30 amino acids, particularly between 6 and 20 amino acids.29. The altered transcription factor of any one of embodiments 19-28 comprising at least one aromatic amino acid residue in the I DR at a position different from the position of said aromatic amino acid in the wild-type transcription factor I DR.30. The altered transcription factor of any one of embodiments 19-29 comprising a number of aromatic amino acid residues in the I DR which is at least 1 , at least 2, at least 3 or at least 4 greater than the number of aromatic amino acids in the I DR of the corresponding wild-type transcription factor.31 . The altered transcription factor of any one of embodiments 22-30, which has an increased transcriptional activity and is selected from the group consisting of an altered HOXD4, HOXC4, OCT4, C / EBPa, NGN2, MYOD-1 , PDX1 , FOXA3, BRACHYURY (TBXT), MSGN1 , NKX2-5 and CDX2 transcription factor.32. The altered transcription factor of any one of embodiments 19-22, which has a reduced detectable transcriptional activity compared to the corresponding wildtype transcription factor.33. An altered mammalian transcription factor including an intrinsically disordered region (I DR), which comprises a decreased number and / or periodicity of aromatic amino acid residues in the I DR compared to a corresponding wild-type transcription factor, wherein the altered transcription factor has a reduced detectable transcriptional activity compared to the corresponding wild-type transcription factor and wherein the corresponding wild-type transcription factor particularly comprises an I DR including at least 4 aromatic amino acid residues.34. The altered transcription factor of any one of embodiments 19-22 and 32-33 having a transcriptional activity that is significantly decreased versus the transcriptional activity of the corresponding wild-type transcription factor as measured in a GAL4 DBD-transactivation reporter system.35. The altered transcription factor of any one of embodiments 19-22 and 32-34 having a transcriptional activity that is, e.g., by at least 20%, by at least 50%, by at least 80% or by at least 90% decreased versus the transcriptional activity of the corresponding wild-type transcription factor as measured in a GAL4 DBD-transactiva- tion reporter system.36. The altered transcription factor of any one of embodiments 32-35, which has a reduced transcriptional activity and is selected from the group consisting of an altered HOXD4, HOXC4, C / EBPa, MYOD-1 , HOXB1 , EGR1 , NFAT5, NANOG, and OCT4 transcription factor.37. The altered transcription factor of any one of embodiments 32-36, wherein at least one aromatic amino acid in the I DR is replaced by a non-aromatic amino acid,particularly selected from the group consisting of A, S, G, and any combination thereof.38. The altered transcription factor of any one of embodiments 32-37 comprising a number of aromatic amino acid residues in the I DR which at least 1 , at least 2, at least 3 or at least 4 lower than the number aromatic amino acids in the I DR of the corresponding wild-type transcription factor.39. A recombinant nucleic acid molecule encoding an altered transcription factor of any one of embodiments 19-38.40. The recombinant nucleic acid molecule of embodiment 39 in operative linkage with an expression control sequence.41 . A vector comprising a nucleic acid molecule of embodiment 39 or 40.42. A recombinant cell comprising a nucleic acid molecule of embodiment 39 or 40 or a vector of embodiment 41.43. Use of an altered transcription factor of any one of embodiments 19-38, a nucleic acid molecule of embodiment 39 or 40, a vector of embodiment 41 or a cell of embodiment 42 in a transcription factor mediated process, e.g., transcription factor-mediated cell reprogramming.44. The use of embodiment 43 for the reprogramming of cells to obtain macrophages, neurons, muscle-like cells, pancreatic B cells, hepatocytes, or pluripotent stem cells, e.g., induced pluripotent stem cells.45. An in vitro method of transcription factor-mediated cell reprogramming comprising the culturing of a starting cell under suitable conditions in the presence of at least one altered transcription factor of any one of embodiments 19-38 and obtaining a desired product cell.Detailed DescriptionThe present inventors identified more than 500 human TFs that encode short periodic blocks of aromatic residues, resembling imperfect prion-like sequences, and show that periodic regions are distinct from previously annotated functional domains. Mutation of periodic aromatic residues inhibited and increasing aromatic dispersion in the disordered regions of multiple TFs enhanced transcriptional activity. In three cellular reprogramming systems, optimizing aromaticdispersion enhanced reprogramming efficiencies of human TFs. Mechanistically, optimizing aromatic dispersion enhanced liquid-like features of phase-separated droplets, and led to more promiscuous DNA binding in cells. Rational engineering of amino acid features that alter phase separation may be a strategy to facilitate TF-mediated cellular reprogramming, and to modulate other TF-mediated processes such as RNA production, protein production, cellular responsiveness to extrinsic and intracellular signaling, physical contacts between genomic elements, enhancer activity, tissue cell composition, regeneration, tissue repair and others.According to the present invention, a method of modifying the activity of a TF including an intrinsically disordered region (IDR) is provided. Typically, the method involves altering the periodicity and / or number of the aromatic amino acid residues in the IDR of an initial TF, e.g., a wild-type TF, including an IDR comprising at least one aromatic amino acid. In particular embodiments, the initial TF comprises at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10 and up 20 or up to 30 aromatic amino acid residues dispersed therein. The aromatic amino acid residues to be altered are typically selected from the group consisting ofY (Tyr), W (Trp), F (Phe), H (His) and any combination thereof. In certain embodiments, the aromatic amino acid residues to be altered are typically selected from the group consisting ofY (Tyr), W (Trp), F (Phe) and any combination thereof. By altering the periodicity and / or number of the aromatic amino acid residues in the IDR of an initial TF, an altered TF with modified transcriptional activity may be obtained.The transcriptional activity of the altered TF may be measured and optionally compared to the transcriptional activity of the corresponding initial, e.g., wild-type TF. In certain embodiments, the transcriptional activity can be measured in a GAL4 DBD-transactivation reporter system or in any other suitable system such as a reporter system that relies on the production of a protein with enzymatic activity, a reporter system that relies on the production of a protein with spectroscopic properties measurable in cellular extracts, or a reporter system that utilizes DNA- binding domains other than GAL4-DBD (e.g., Lacl etc.).In certain embodiments, the initial TF to be altered is a mammalian TF, particularly a human TF. In certain embodiments, an altered mammalian or human TF is capable of modulating the transcriptional activity of a mammalian gene, e.g., a human gene, different from the corresponding wild-type TF. Typically, the TF comprises an IDR and a DNA binding domain (DBD). In certain embodiments, the IDR comprises a minimal activation domain (AD). In certain embodiments, the DBD of the altered TF is identical to the DBD of a corresponding wild-type TF. In further embodiments, the DBD of the altered TF is altered compared to the DBD of a corresponding wild-type TF, e.g., the alteration may result in an increased binding to its DNA target sequence or to a reduced binding to its DNA target sequence. Binding of a TF to its targetsequence can be determined by standard methods, e.g., chromatin immunoprecipitation followed by sequencing (ChlP-Seq), or any variation of such technologies relying on antibody probes (e.g. ChlP-mentation, Cut-and Tag), as well as indirect transcription factor foot printing methods (e.g. ATAC-Seq, ChEC-Seq).Specific examples of suitable TFs are described by Lambert et al., The Human Transcription Factors, Cell 172 (2018), 650-665, 2017, particularly in Supplementary Table 1 thereof, the content of which is herein incorporated by reference. Further specific examples of suitable TFs are disclosed in the following Table:The Table shows the names of the respective TFs and their Genl D, TrnlD, PepID and UniProt identification numbers as of the priority date of the present application. It should be noted, however, that the invention also encompasses functional variants of any of these TFs, including an allelic variant, a wild-type variant or disease-associated variant, and a recombinant variant, e.g., a variant fused to a heterologous peptide or polypeptide. Further, the invention encompasses orthologs of the indicated TFs from non-human organisms, particularly non-human mammals. In certain embodiments, orthologs have an amino acid sequence identity of at least 50%, at least 70%, or at least 90% with the respective human TF over the whole length of the amino acid sequence.Still further examples of human TFs are described in Fig. 1 b and Extended Data Fig. 1 of the present application.In certain embodiments, the invention relates to the manufacture of a TF with an increased transcriptional activity comprising the steps:(a) providing an initial TF including an I DR, particularly an I DR comprising at least 4 aromatic amino acid residues dispersed therein,(b) increasing the periodicity and / or number of aromatic amino acid residues in the I DR, thereby obtaining an altered TF,(c) optionally measuring the transcriptional activity of the altered TF, and(d) obtaining an altered TF with increased transcriptional activity.In certain embodiments, an increased transcriptional activity is obtained by:(i) increasing the periodicity by rearranging the positions of aromatic amino acid residues originally present in the I DR; or(ii) increasing the number of aromatic amino acid residues in the I DR; or(iii) a combination of (i) and (ii).The number of the aromatic amino acid residues may be increased by inserting at least one aromatic amino acid into the I DR such that the altered TF comprises a number of aromatic amino acid residues in the IDR which is at least 1 , at least 2, at least 3 or at least 4 greater than the number of aromatic amino acids in the IDR of an initial, e.g., wild-type TF.The periodicity of the aromatic amino acid residues may be increased by positioning aromatic amino acids within the IDR such that a substantially uniform spacer length between individual aromatic amino acids is obtained, e.g., wherein the spacer length between individual aromatic amino acids deviates by up to 2, up to 1 or 0 amino acids throughout the IDR. The substantially uniform spacer length between individual aromatic amino acids may be between 4 and 30 amino acids, particularly between 6 and 20 amino acids.In certain embodiments, the invention relates to the manufacture of a TF with a reduced detectable transcriptional activity comprising the steps:(a) providing an initial TF including an IDR, particularly an IDR comprising at least 4 aromatic amino acid residues dispersed therein, and(b) reducing the periodicity and / or number of the aromatic amino acid residues in the IDR, thereby obtaining an altered TF,(c) measuring the transcriptional activity of the altered TF, and(d) obtaining an altered TF with reduced transcriptional activity.In certain embodiments, a reduced transcriptional activity is obtained by:(i) decreasing the periodicity by rearranging the positions of aromatic amino acid residues originally present in the IDR; or(ii) decreasing the number of aromatic amino acid residues in the IDR; or(iii) a combination of (i) and (ii).The number of aromatic amino acid residues in the IDR may be decreased by replacing at least one aromatic amino acid with a non-aromatic amino acid. The non-aromatic amino acid may be selected from neutral non-aromatic amino acids such as A (Ala), S (Ser), G (Gly), L (Leu), V (Vai), T(Thr), I (lie), C (Cys), Q (Gin) and N (Asn) and / or charged non-aromatic aminoacids such as K (Lys), R (Arg), D (Asp), and E (Glu). In certain embodiments, the non-aromatic amino acid is a non-aromatic neutral amino acid, particularly selected from the group consisting of A (Ala), S (Ser), G (Gly), and any combination thereof.The number of the aromatic amino acid residues may be reduced by replacing at least one aromatic amino acid in the I DR by a non-aromatic amino acid such that the altered TF comprises a number of aromatic amino acid residues in the I DR which is at least 1 , at least 2, at least 3 or at least 4 lower than the number of aromatic amino acids in the I DR of the initial, e.g., wild-type TF.A further aspect of the present invention relates to an altered TF including an I DR, particularly an I DR comprising at least 4 aromatic amino acid residues dispersed therein, which comprises an altered number and / or periodicity of aromatic amino acid residues in the I DR compared to a corresponding wild-type TF. In particular embodiments, the altered TF has a modified transcriptional activity, e.g., an increased transcriptional activity or a reduced transcriptional activity compared to a corresponding wild-type TF. The altered TF with modified transcriptional activity may be obtained by a method as described herein.In certain embodiments, the altered TF comprises an IDR and a DNA binding domain (DBD) wherein the IDR differs in its amino acid sequence from the IDR of a corresponding wild-type TF. In certain embodiments, the DBD of the altered TF is identical to the DBD of a corresponding wild-type TF. In further embodiments, the DBD of the altered TF is altered compared to the DBD of a corresponding wild-type TF, wherein the alteration may result in an increased binding to its DNA target sequence or to a reduced binding to its DNA target sequence.The altered TF of the present invention differs in its amino acid sequence from the corresponding wild-type TF. In certain embodiments, the amino acid sequence identity between the altered TF and the wild-type TF is at least 90%, at least 92%, at least 94%, at least 96%, at least 98% or at least 99% up to less than 100% over the whole length of the polypeptide. The sequence identity may be determined by standard algorithms such as BLAST.In certain embodiments, the altered TF has an increased transcriptional activity compared to the corresponding wild-type TF. Specific examples of altered TFs with increased transcriptional activity include altered HOXD4, HOXC4, OCT4, C / EBPa, NGN2, MYOD-1 , PDX1 , and FOXA3 TFs, particularly altered TFs disclosed herein. Further specific examples of altered TFs with increased transcriptional activity include altered BRACHYURY (TBXT); MSGN1 , NKX2-5 and CDX2.In certain embodiments, the transcriptional activity of the altered TF is significantly increased versus the transcriptional activity of the corresponding wild-type TF as measured in a GAL4 DBD-transactivation reporter system. In certain embodiments, the transcriptional activity is, e.g., by at least 20%, by at least 50%, by at least 75% or by at least 90% increased versus the transcriptional activity of the corresponding wild-type TF as measured in a GAL4 DBD-trans- activation reporter system.The altered TF may include an I DR comprising at least 4 aromatic amino acid residues dispersed therein, wherein there is a substantially uniform spacer length between individual aromatic amino acids, e.g., wherein the spacer length between individual aromatic amino acids deviates by up to 2, up to 1 or 0 amino acids throughout the I DR. The substantially uniform spacer length between individual aromatic amino acids may be between 4 and 30 amino acids, particularly between 6 and 20 amino acids.In certain embodiments, the altered TF comprises at least one aromatic amino acid residue, e.g., 1 , 2, 3 , 4, 5 or even more aromatic amino acid residues in the IDR at a position different from the position of said aromatic amino acid in the IDR of a corresponding wild-type transcription factor.In certain embodiments, the altered TF comprises a number of aromatic amino acid residues in the IDR which is at least 1 , at least 2, at least 3 or at least 4 greater than the number of aromatic amino acids in the IDR of the corresponding wild-type TF.In certain embodiments, the altered TF has a reduced detectable transcriptional activity compared to the corresponding wild-type TF. Specific examples of altered TFs with reduced transcriptional activity are selected from the group consisting of altered HOXD4, HOXC4, C / EBPa, MYOD-1 , HOXB1 , EGR1 , NFAT5, NANOG, and OCT4 TFs, particularly altered TFs disclosed herein.In certain embodiments, the altered transcription factor has a transcriptional activity that is significantly decreased versus the transcriptional activity of the corresponding wild-type transcription factor as measured in a GAL4 DBD-transactivation reporter system. In certain embodiments, the altered TF has a transcriptional activity that is, e.g., by at least 20%, by at least 50%, by at least 75% or by at least 90% decreased versus the transcriptional activity of the corresponding wild-type TF as measured in a GAL4 DBD-transactivation reporter system. In certain embodiments, the altered TF has a reduced detectable transcriptional activity that is at least 5%, at least 10%, at least 25% and up to 80%, or up to 50% of the transcriptional activityof the corresponding wild-type TF as measured in a GAL4 DBD-transactivation reporter system.In certain embodiments, an altered TF with reduced transcriptional activity comprises an I DR wherein at least one aromatic amino acid in the IDR of a wild-type TF is replaced by a nonaromatic amino acid, particularly selected from the group consisting of A, S, G, and any combination thereof.In certain embodiments, an altered TF with reduced transcriptional activity comprises a number of aromatic amino acid residues in the IDR which at least 1, at least 2, at least 3 or at least 4 lower than the number aromatic amino acids in the IDR of the corresponding wild-type TF.A further aspect of the present invention relates to a recombinant nucleic acid molecule encoding an altered TF as described above. The nucleic acid molecule may be a single- or double-stranded nucleic acid molecule, e.g., a DNA or RNA molecule. In certain embodiments, the recombinant nucleic acid molecule is in operative linkage with an expression control sequence, e.g., an expression control sequence allowing expression in prokaryotic or eukaryotic cell, particularly in a mammalian cell including a human cell.A further aspect of the invention is a vector including a viral or non-viral vector, e.g., a plasmid comprising a nucleic acid molecule as described above.A further aspect of the present invention relates to a recombinant cell comprising a nucleic acid molecule or a vector as described above. The nucleic acid molecule or the vector may be present extra chromosomally or integrated into the cell genome. The recombinant cell may be obtained by introducing a nucleic acid molecule or vector as described above into the cell by transformation, transfection, or infection according to known procedures. Alternatively, the recombinant cell may be obtained by gene editing, wherein one or more copies of an endogenous gene encoding a wild-type TF are modified by using gene editing procedures, e.g., by using CRISPR / Cas or a similar nuclease.The recombinant cell may be a prokaryotic host cell or a eukaryotic host cell, e.g., an insect or mammalian cell, particularly a human cell. In certain embodiments, the recombinant cell is a cell, which can be subjected to a TF-mediated process, e.g., TF-mediated cell reprogramming. In certain embodiments, the recombinant cell is a cell, which is obtained by TF-mediated cell reprogramming.In certain embodiments, the recombinant cell is selected from lymphocytes such as T cells, B cells, e.g., pancreatic B cells, macrophages, myeloid cells, epithelial cells, connective tissue cells, extraembryonic cells, endocrine cells, cells of the kidney, lung, heart, neurons, musclelike cells, hepatocytes, or pluripotent stem cells, e.g., induced pluripotent stem cells or any other mammalian cell type originating from the three germ layers.The altered TF, the nucleic acid molecule, the vector, or the cell may be used for any procedure involving a TF-mediated process, e.g., TF-mediated cell reprogramming, RNA production, protein production, cellular responsiveness to extrinsic and intracellular signaling, physical contacts between genomic elements, enhancer activity, tissue cell composition, rejuvenation, regeneration, and tissue repair.In particular embodiments, the altered TF, the nucleic acid molecule, the vector, or the cell may be used for liver regeneration, e.g., for the treatment of hepatocellular carcinoma (OMIM: 114550), heart regeneration, muscle regeneration, brain regeneration after brain injury, vision repair, repair of retinal damage, beta-cell reprogramming, e.g., for the treatment of type 1 diabetes (OMIM: 222100), or immune cell reprogramming, e.g., for the treatment of autoimmune disorders. Suitable medical applications are described in references99’101’102, the contents of which are herein incorporated by reference.The altered transcription factors can be delivered either in the form of nucleic acids that encode them, or as purified, packed proteins. The potential delivery methods include but are not limited to delivery by viral vectors, e.g., retroviral vectors, lentiviral vectors, or adeno-associated viral vectors, or by non-viral vectors, e.g., lipid particles such as nucleic acid-containing lipid nanoparticles (LNPs) or protein containing LNPs. Suitable delivery approaches are described in references "•10°, the contents of which are herein incorporated by reference.Still a further aspect of the present invention relates to an in vitro method of TF-mediated cell reprogramming comprising the culturing of a starting cell under suitable conditions in the presence of at least one altered TF as described herein and obtaining a desired product cell.In the following, the invention shall be explained in more detail by the following Figures and Examples.Figure LegendsFigure 1. Human TF IDRs encode short periodic blocks of aromatic amino acids, but the overall periodicity of TF IDRs is limited a. Schematic model of a transcription factor, and the method to identify aromatic periodic blocks. DNA binding domain (DBD) is in yellow, the intrinsically disordered region (I DR) in dark blue, and ‘minimal’ activation domain (AD) in light blue. a. a: amino acid b. The top 80 TFs ranked on I DR periodicity score. The rank is noted in parentheses after the TFs’ names. The TFs are colored on the type of DNA binding domain they contain. On the outer circle, the height of the bar is proportional to the periodicity score. The inner circles include annotation whether the I DR contains a minimal activation domain (AD) identified in the four indicated studies. c. Positioning of aromatic residues in NFAT5. The DBD is denoted with a yellow bar, the IDR is denoted with a dark blue bar. A predicted minimal activation domain (AD) is denoted with a light blue bar. The positions of aromatic residues are denoted as yellow dots. Aromatic residues in periodic blocks are in red. d. Omega plot of the NFAT5 IDR. The aromatic residues in the NFAT5 IDR are more uniformly distributed than 100% of randomly shuffled sequences of the same length and amino acid composition (QAro=0.124, empirical P-value=0). e. Disorder plots (Metapredict) of HOXC4 in black, AlphaFold2 pLDDT score plots in yellow. Aromatic residues in periodic blocks are highlighted in red. f. Omega plots of the HOXC4 IDR (top sub-panel) and the portion encoding the periodic aromatic block (bottom sub-panel). Shown are the co-ordinates, QATO scores and the percentage of randomly generated sequences that have a lower QATO score than the actual sequence. g. Representative images of droplet formation of purified, recombinant HOXC4 IDR-mEGFP fusion proteins at the indicated concentrations in the presence of 10% PEG 8000. In the Aro- LITE mutant, aromatic residues in the entire IDR were substituted with serines. Scale bar is 5 pm. h. The relative amount of condensed protein per concentration quantified in the droplet formation assays. Data are displayed as mean ± SD. N = 10 images from at least 2 replicates. The curve was generated as a non-linear regression to a sigmoidal curve function. i. Schematic and results of luciferase reporter assays in V6.5 mouse embryonic stem cells. Luciferase values were normalized against an internal Renilla control, and the values are displayed as percentages normalized to the activity measured using an empty vector. Data are displayed as mean ± SD. Data are from three biological replicates with three technical replicates each. P-values are from Student’s t-tests. ***: P<10-3.j. Schematic of the pipeline for the identification of regions with significant periodicity in the human proteome. k. Density plot of all proteins that contain a region of significant periodicity. Top left: Intrinsically disordered regions (IDRs) that contain a region of significant periodicity. Top right: Prion-like domains that contain a region of significant periodicity. Bottom left: Aromatic rich prion-like domains (>10% aromatic content) that contain a region of significant periodicity. Bottom right: transcription factors that contain a region with significant periodicity. For each region of significant periodicity, the length of the region is plotted against the lowest P-value within the region. The depth of the color of the cloud is proportional to the density of the dots in the area. The numbers in the plots represent the number of proteins that contain a region of significant periodicity over the total number of proteins in each category. l. Omega scores of intrinsically disordered regions (IDRs) in the proteome that overlap PLDs, Aromatic-rich PLDs or TFs. Note that PLDs have on average a lower Omega score than TFs. P-values are from one-way ANOVA with Tukey’s multiple comparisons post-test. ***: P<10’4m. Schematic models of examples of PLDs and TF IDRs and their Omega scores. Aromatic residues are highlighted as orange dots.Figure 2. Increasing aromatic dispersion in TF IDRs enhances transactivation a. (left) Schematic models of HOXD4 IDRs. Aromatic residues are highlighted as orange dots. Alanine mutations are highlighted as white dots. Additionally introduced tyrosines are highlighted as red dots, (middle) Omega plots of the HOXD4 IDRs and QATO scores, (right) Results of luciferase reporter assays. Data are from three biological replicates with three technical replicates each. b. Representative images of droplet formation of purified HOXD4 IDR-mEGFP fusion proteins at the indicated concentrations in droplet formation buffer. Scale bar is 5 pm. c. The relative amount of condensed protein per concentration quantified in the droplet formation assays. Data are displayed as mean ± SD. N = 15 images from 3 replicates. The curve was generated as a non-linear regression to a sigmoidal curve function. d. Fluorescence intensity of HOXD4 wild type and HOXD4 AroPERFECT in vitro droplets before, during and after photobleaching. Data displayed as mean ± SD. N = 20. e. Results of a HOXD4 I DR tiling experiment by using luciferase reporter assays. Sequences were tiled into fragments of 40 amino acids with 20 amino acid overlaps. The activities of the full-length IDRs are indicated with dashed horizontal lines. A predicted activation domain in the HOXD4 wild type I DR is highlighted with a light blue bar. f. Results of luciferase reporter assays of the indicated I DR constructs.g. (left) Schematic models of synthetic sequences. Tyrosines are highlighted as orange dots, (right) Results of luciferase reporter assays. Data displayed as mean ± SD from four biological replicates with two technical replicates each.In panels a, e, f, g, Luciferase values were normalized against an internal Renilla control, and the values are displayed as percentages normalized to the activity measured using an empty vector. Data are displayed as mean ± SD. P-values are from Student’s t-tests. ***: P<10-3.Figure 3. Evidence for gain-of-function of periodic HOXD4 mutants in vivo a. (top) Differential interference contrast (DIG) microscopy of the indicated cell lines. Scale bar is 0.4 mm. (bottom) Representative fluorescence microscopy images of cell nuclei. b. Representative images of HAP1 HOXD4 wild type-mEGFP, HOXD4 AroPERFECT-mEGFP and HOXD4 AroPLUS-mEGFP nuclei after 24h of HOXD4 expression. The fusion proteins were visualized using mEGFP fluorescence in fixed cells. The normalized signal intensity was calculated by dividing standard deviation of mEGFP signal of each nucleus by the corresponding mean mEGFP signal. Number of individual nuclei per condition is displayed. c. Granularity scores of nuclei, with corresponding mean nuclear mEGFP intensities. Data are displayed as mean ± SD from two biological replicates. P-values are from Student’s t-tests. d. Principal component (PC) analysis of the RNA-Seq expression profiles of HAP1 wild type, HOXD4 knockout and indicated knock-in cell lines. e. Differential expression analysis of HAP1 HOXD4 AroPERFECT-mEGFP and HOXD4 AroPLUS-mEGFP cells versus HOXD4 wild type-mEGFP cells. HOXD4 target genes are highlighted in blue, non-HOXD4 target genes are highlighted in red. f. Western blot analysis of HOXD4-mEGFP, IFI16 and ARHGAP4 in the indicated cell lines. HOXD4-mEGFP proteins were probed with an anti-GFP anybody. HSP90 is shown as loading control. g. (left) schematic model of the condensate tethering system, (right) Fluorescence images of ectopically expressed YFP-RNAPII CTD in live U2OS cells co-transfected with the indicated - CFP-Lacl-HOXD4 I DR fusion constructs. Dashed line is the nuclear contour. Scale bar is 10 pm. h. Quantification of the relative YFP signal intensity in the tether foci. P-values are from a t- test. **P<0.01 , ***P<10-3.Figure 4. Optimizing aromatic dispersion in C / EBPa enhances transactivation a. (left) Schematic models of wild type and mutant C / EBPa proteins. The position of the bZIP DNA binding domain is highlighted with a grey box. Aromatic residues are highlighted as orange dots, (middle) Omega plots and QATO scores. I DR: intrinsically disordered region, (right) Results of luciferase reporter assays. Data are displayed as mean ± SD from three biologicalreplicates with three technical replicates each. P-values are from Student’s t-tests. ***: P<10’3 b. Representative images of droplet formation of purified C / EBPa IDR-mEGFP fusion proteins at the indicated concentrations in droplet formation buffer. Scale bar is 5 pm. c. Fluorescence intensity of C / EBPa wild type, AroLITE and AroPERFECT IS15 IDR in in vitro droplets before, during and after photobleaching. Data are displayed as mean ± SD. N = 15. d. Fluorescence images of ectopically expressed YFP-RNAPII CTD in live LI2OS cells cotransfected with the indicated CFP-Lacl-C / EBPa IDR fusion constructs. Dashed line is the nuclear contour. Scale bar is 10 pm. e. Quantification of the relative YFP signal intensity in the tether foci. P-values are from a t- test.*** p< Q-3 f. Results of a C / EBPa IDR tiling experiment using luciferase reporter assays. C / EBPa wild type and AroPERFECT IS15 IDR sequences were tiled into fragments of 40 amino acids with 20 amino acid overlaps. Data are displayed as mean ± SD from four biological replicates with two technical replicates each. The activities of the full-length IDRs are indicated with dashed horizontal lines. g. Results of luciferase reporter assays of the indicated IDR constructs. P-values are from a t- test. *P<0.05, **P<0.01, ***p<io-3.In panels, a, f, g, luciferase values were normalized against an internal Renilla control, and the values are displayed as percentages normalized to the activity measured using an empty vector.Figure 5. Optimizing aromatic dispersion in C / EBPa enhances macrophage reprogramming, and leads to stronger and more promiscuous genomic binding a. Schematic models of wild type and mutant C / EBPa proteins. The transactivation data is identical to the data displayed in Fig. 4a. b. Schematic model of the C / EBPa-mediated B-cell to macrophage transdifferentiation experiment. c. FACS quantification of GFP+RCH-rtTA cells carrying a C / EBPa wild type or mutant overexpression cassette. Cells were quantified for the level of the macrophage differentiation marker Mac1 and B-cell marker CD19, 48h, 96h and 168h after transgene induction. N = 3-5 biological replicate experiments. d. Graph-based clustering of the scRNA-seq data from C / EBPa mediated B-cell to macrophage transdifferentiation experiment. LIMAP (uniform manifold approximation projection) was used to represent the data. The inset is a pseudotime plot illustrating the relative time relationship between the cells.e. Quantification of mEGFP-positive cells in macrophage clusters. AroPERFECT IS15 has a higher fraction of late macrophage cells but also an overall higher fraction of differentiating macrophages. f. Heatmap representation of ChlP-Seq read densities of C / EBPa wild type and AroPERFECT IS15 within a 1.5kb window around all shared C / EBPa peaks (top), and differentially enriched peaks in C / EBPa AroPERFECT IS15. “Peaks unique to IS15 and reported before” denote binding sites that are differentially enriched in IS15-binding and overlap C / EBPa peaks reported in previous literature. FE: enrichment. g. Enrichment scores of bZIP TF motifs, and adjusted P-values of enrichment at the three indicated peak sets. h. j. C / EBPa AroPERFECT IS15 shows enhanced binding at the FAM98 (h) and GBP5 (j) loci. Displayed are genome browser tracks of ChlP-Seq data of C / EBPa wild type and AroPERFECT IS15 in RCH-rtTA cells, 24 and 48 hours after C / EBPa expression. Co-ordinates are hg38 genome assembly co-ordinates. i. k. UMAPs colored on FAM98 (i) and GBP5 (k) expression. The numbers denote the mean expression + / -SD in the whole samples.I. Luciferase assays using the indicated reporter plasmids co-transfected with expression vectors encoding either WT (red bars) or AroPERFECT IS15 (purple bars) C / EBPa proteins. Luciferase values were normalized against an internal Renilla control, and the values are displayed as percentages normalized to the activity measured using the ‘basic’ vector. Data are displayed as mean ± SD from three biological replicates with three technical replicates each. P-values are from Student’s t-tests. *: P<0.05, ***: P<10-3.Figure 6. Optimizing aromatic dispersion in NGN2 enhances neural differentiation a. (left) Schematic models of wild type and mutant NGN2 proteins. The position of the bHLH DNA binding domain is highlighted with a grey box. Aromatic amino acids are highlighted as orange dots, (right) Omega plots and QATO scores. I DR: intrinsically disordered region. b. Fluorescence intensity of NGN2 wild type and AroPERFECT I DR in in vitro droplets before, during and after photobleaching. Data are displayed as mean ± SD. N = 15. c. Schematic model of the NGN2-mediated hIPSC to neuron differentiation experiment. hIPSC: human induced pluripotent stem cell. ROCKi: Rho-kinase inhibitor. d. Representative fluorescence microscopy images of differentiating human iPSCs expressing the indicated NGN2 proteins at the indicated timepoints. T ubulin staining is in magenta, nuclear counterstain (Hoechst) in blue, NGN2-T2A-mEGFP is green. Scale bar is 0.1 mm. Scale bar of insets is 0.05 mm.e. Quantification of the number of cells based on Hoechst nuclear staining in the NGN2-di- rected differentiation experiments. Data are displayed as mean ± SD. N = 6 images from 2 independent experiments. P-value is from a Student’s t-test. *:P<0.05. f. Quantification of neurite density based on tubulin staining in the NGN2-directed differentiation experiments. Data are displayed as mean ± SD. N = 6 images from 2 independent experiments. P-value is from a Student’s t-test. *:P<0.05. g. Principal component analysis of the RNA-Seq expression profiles of parental ZIP13K2 hlP- SCs, and hIPSCs expressing the indicated NGN2 transgenes (PC1 vs. PC3). h. Differential expression analysis of hIPSCs expressing the indicated transgenes. NGN2 target genes are highlighted in blue. i. Heatmap representation of ChlP-Seq read densities of NGN2 wild type, AroLITE and AroPERFECT-expressing cells within a 1.5kb window around all shared NGN2 peaks (top), differentially enriched peaks in NGN2 AroPERFECT (center) and differentially enriched peaks in NGN2 wild type (bottom). F.o.l: fold over input. j. NGN2 AroPERFECT shows differential binding at the TMEM97 locus. Displayed are genome browser tracks of ChlP-Seq data in cells expressing the indicated NGN2 transgenes, after 24 and 48 hours of expression. Arrowhead highlights a differentially bound peak at 24h. Co-ordinates are hg38 genome assembly co-ordinates. k. Nascent transcription (TT-SLAM-Seq) meta-gene profiles at -9000 NGN2-target genes.Figure 7. Optimizing aromatic dispersion in MYOD1 enhances myotube differentiation a. (left) Schematic models of wild type and mutant MYOD1 proteins. The position of the bHLH DNA binding domain is highlighted with a grey box. Aromatic amino acids are highlighted as orange dots, (middle) Omega plots and QATO scores of the N-terminal and C-terminal IDRs. I DR: intrinsically disordered region, (right) Results of luciferase reporter assays in C2C12 mouse myoblasts. Luciferase values were normalized against an internal Renilla control, and the values are displayed as percentages normalized to the activity measured using an empty vector. Data are displayed as mean ± SD from three biological replicates with three technical replicates each. P-values are from Student’s t-tests. ***: P<10-3. b. Schematic model of the MYOD1 -mediated myotube differentiation experiment. c. Representative fluorescence microscopy images of differentiating myoblasts expressing the indicated MYOD1 proteins at day 3 after Dox induction. The mEGFP signal of the MYOD1 - T2A-mEGFP construct was used as a cytoplasmic marker and is shown in cyan. Nuclear counterstain (Hoechst) is shown in magenta. Scale bar is 0.5 mm. d. Quantification of MYOD1 driven myotube differentiation efficiency. Fusion coefficient was calculated as the percentage of nuclei in cells containing at least 3 nuclei. Data are displayedas mean ± SD. N = 15 images from 3 independent experiments. P-value is from a Student’s t- test. *:<0.05. e. Principal component analysis of RNA-Seq expression profiles of parental C2C12 cells, and cells expressing the indicated MYOD1 transgenes. f. Differential expression analysis of C2C12 MYOD1 AroLITE and C2C12 MYOD1 AroPER- FECT-C -expressing cells versus C2C12 cells expressing wild type MYOD1. MYOD1 target genes are highlighted in blue. Highlighted genes are differentially expressed and are involved in cell adhesion.Figure 8. Deletion of aromatic residues in the IDRs of the transcription factors Brachy- ury (TBXT, also called T), MSGN2, NKX2-5 and CDX2 reduces transcriptional activity.Enhancing the periodicity of aromatic residues in the IDRs of the transcription factors Brachyury (TBXT), MSGN2, NKX2-5 and CDX2 increases transcriptional activity. a. Disorder plot (Metapredict) for the mouse Brachyury (TBXT) transcription factor. The position of the IDR is noted with a blue bar. Positioning of aromatic residues in NFAT5. Yellow dots indicate the position of aromatic residues in the sequence. b. (left) Schematic model of the TBXT IDR constructs tested for transcriptional activity. Aromatic residues (orange dots) acids and alanine mutations (white dots) are highlighted, (right) Results of luciferase reporter assays. The luciferase values were normalized to an internal Renilla control and the values are displayed as percentages of the activity measured using an empty vector. Data are the mean ± s.d. of n = 3 biological replicates. P values are from two-sided unpaired Student’s t-tests **P<0.01 , *P<0.05 c. Same as (a) for MGSN1 d. Same as (b) for MGSN1 e. Same as (a) for NKX2-5 f. Same as (b) for NKX2-5 g. Same as (a) for CDX2 h. Same as (b) for CDX2Extended Data Figure 1. Functional characterization of periodic blocks in human TFs. a. Distribution plot of the 531 human TFs that contain short periodic blocks overlapping their intrinsically disordered regions (IDRs). Most TF IDRs overlap one short periodic block. b. Distribution plot of the 748 periodic blocks of aromatic amino acids in human TF I DRs. Most periodic blocks consist of 4 aromatic residues. c. Domain annotation of the 80 human TFs with the highest IDR periodicity score. Zinc finger TFs are shown on the left, members of all other TF families on the right. The majority of periodic blocks do not overlap ‘minimal’ activation domains.d. Frequency of amino acids in non-periodic, and periodic TF IDRs, relative to their frequencies in the full proteome. Note that periodic TF IDRs are relatively enriched for aromatic residues, depleted for charged residues, and enriched for neutral residues. e. Amino acid PWM and cumulative bar frequency plot around aromatic residues in periodic blocks, Colors represent disorder promoting (yellow), order promoting (blue) and neutral residues (grey). f. Variable length gapped or un-gapped motif analysis of periodic blocks and charged blocks from Lyons et. al, represented as PWM plot. Note that no motif could be found.Extended Data Figure 2. Aromatic residues in periodic TF IDRs are necessary for in vitro phase separation and transactivation. a. Disorder plots (Metapredict) of HOXB1 and HOXD4 in black, AlphaFold2 pLDDT score plots in yellow. Aromatic residues in periodic blocks are highlighted in red. Predicted activation domains are highlighted in light blue. b. Omega plots of HOXB1 and HOXD4 for full IDR regions (top sub-panel) and portions encoding periodic blocks of aromatic blocks (bottom sub-panels). Shown are the co-ordinates of the regions, QATO scores and the percentage of randomly generated sequences that have a lower QATO score than the actual sequence. c. Representative images of droplet formation of purified, recombinant TF IDR-mEGFP fusion proteins at the indicated concentrations in the presence of 10% PEG 8000. For HOXB1 and HOXD4 aromatic residues in the entire IDR were substituted with alanines in the AroLITE mutants. Scale bar is 5 pm. d. The relative amount of condensed protein per concentration quantified in the droplet formation assays. Data are displayed as mean ± SD. N = 10 images from at least 2 replicates. The curve was generated as a non-linear regression to a sigmoidal curve function.Luciferase values were normalized against an internal Renilla control, and the values are displayed as percentages normalized to the activity measured using an empty vector. Data are displayed as mean ± SD. Data are from three biological replicates with three technical replicates each. P-values are from Student’s t-tests. *: P<0.05, ***: P<10-3. f. Schematic model of HOXD4 wild type and mutant IDRs. g. Representative images of droplet formation of purified HOXD4 IDR-mEGFP fusion proteins at the indicated concentration in droplet formation buffer. Scale bar is 5 pm. h. The relative amount of condensed protein per concentration quantified in the droplet formation assays. Data are displayed as mean ± SD. N = 10 images from 2 replicates. The curve was generated as a non-linear regression to a sigmoidal curve function.Luciferase values were normalized against an internal Renilla control, and the values are displayed as percentages normalized to the activity measured using an empty vector. Data are displayed as mean ± SD from three biological replicates with three technical replicates each. P-values are from Student’s t-tests. *: P<0.05, ***: P<10-3. j. (left) Disorder plot (Metapredict) in black and AlphaFold2 pLDDT score plots in yellow for EGR1. Aromatic residues in periodic blocks are highlighted in red. IDR: intrinsically disordered region, DBD: DNA binding domain, (right) Results of luciferase reporter assays of the EGR1 C-IDR in V6.5 mouse embryonic stem cells. Luciferase values were normalized against an internal Renilla control, and the values are displayed as percentages normalized to the activity measured using an empty vector. Data are displayed as mean ± SD from three biological replicates with three technical replicates each. P-values are from Student’s t-tests. ***: P<10-3. k. (left) Disorder plot (Metapredict) in black and AlphaFold2 pLDDT score plots in yellow for NFAT5. Aromatic residues in periodic blocks are highlighted in red. IDR: intrinsically disordered region, DBD: DNA binding domain, AD: activation domain highlighted in light blue, (right) Results of luciferase reporter assays. Luciferase values were normalized against an internal Renilla control, and the values are displayed as percentages normalized to the activity measured using an empty vector. Data are displayed as mean ± SD from three biological replicates with three technical replicates each. P-values are from Student’s t-tests. ***: P<10-3. l. (left) Disorder plot (Metapredict) in black and AlphaFold2 pLDDT score plots in yellow for NANOG. Aromatic residues in periodic blocks are highlighted in red. IDR: intrinsically disordered region, DBD: DNA binding domain, (right) Results of luciferase reporter assays of the NANOG C-IDR. Luciferase values were normalized against an internal Renilla control, and the values are displayed as percentages normalized to the activity measured using an empty vector. Data are displayed as mean ± SD from three biological replicates with three technical replicates each. P-values are from Student’s t-tests. ***: P<10-3. m. Luciferase values were normalized against an internal Renilla control, and the values are displayed as percentages normalized to the activity measured using an empty vector. Data are displayed as mean ± SD from three biological replicates with three technical replicates each. P-values are from Student’s t-tests. **: P<0.01 , ***: P<10-3.Extended Data Figure 3. Proteins that contain regions with significant periodicity. a. Region of significant periodicity in HNRNPA1. Plotted is the disorder score (Metapredict) on the top, and the P-values of the periodicity algorithm on the bottom against the position of amino acids. b. Density plot of all proteins that contain a region of significant periodicity. For each region of significant periodicity, the length of the region is plotted against the lowest P-value within the region. A P-value cutoff of 0.01 was used to identify 2,202 regions.c. AlphaFold models of four proteins. Aromatic residues are colored in red, and all other residues are colored in yellow. Note that in DAZ1 , the periodic aromatic residues are in a structure of beta-sheets. EGR1 is the transcription factor with the highest ranked region of significant periodicity. d-e. Gene set enrichment analysis (GSEA) of the 2,202 human proteins that contain a region with significant periodicityExtended Data Figure 4. Characterization of periodic TF IDR mutants. a. Representative fluorescence microscopy images of the fluorescence recovery after photobleaching (FRAP) experiments of HOXD4 wild type IDR-mEGFP and HOXD4 AroPERFECT IDR-mEGFP droplets. Shown is one droplet of each genotype, in droplet formation buffer at a concentration of 20 pM. Images were taken every 3 seconds over a time span of 60 seconds. b. Western blot of GAL4-DBD and GAL4-DBD-HOXD4-IDR-fusion proteins in HEK293T cells 24 hours after transfection using a GAL4-DBD specific antibody. HSP90 is shown as a loading control. Except for GAL4-HOXD4 AroLITE A, GAL4-DBD-HOXD4-IDR fusion proteins are expressed at comparable levels. c. Schematic models of HOXD4 wild type and mutant IDRs. Aromatic amino acids are highlighted as orange dots. Alanine mutations are highlighted as white dots. Omega plots of the HOXD4 IDRs and QATO scores are shown next to the schematic models. d. The YPWM motif, ubiquitously found in HOX transcription factors, does not contribute to the transactivation potential of the HOXD4 IDR. Shown are results of luciferase reporter assays in V6.5 mouse embryonic stem cells of HOXD4 IDRs. Data are displayed as mean ± SD from three biological replicates with three technical replicates each. P-values are from Student’s t- tests. *: P<0.05, ***: P<10-3. e. (left) The activity of HOXD4 IDRs scales with the number of small inert residues adjacent to aromatic residues in the IDR constructs, (right) The activity of C / EBPa IDRs scales with the number of small inert residues adjacent to aromatic residues in the IDR constructs. f. (left) Schematic models of wild type and AroPERFECT HOXC4 IDRs. Aromatic residues are highlighted as orange dots, (middle) Omega plots and QATO scores of the IDRsLuciferase values were normalized against an internal Renilla control, and the values are displayed as percentages normalized to the activity measured using an empty vector (dashed orange line). Data are displayed as mean ± SD. Data are from three biological replicates with three technical replicates each. P-values are from Student’s t-tests. ***: P<10-3. g. Western blot of GAL4-DBD and GAL4-DBD-HOXC4-IDR fusion proteins in HEK293T cells 24 hours after transfection using a GAL4-DBD specific antibody. HSP90 is shown as a loading control.h. Representative images of droplet formation of purified HOXC4 IDR-mEGFP fusion proteins at the indicated concentration in droplet formation buffer Scale bar is 5 pm. For the wild type I DR, the exact same images are displayed in Fig. 1g. i. The relative amount of condensed protein per concentration quantified in the droplet formation assays. Data are displayed as mean ± SD. N = 10 images from at least 2 replicates. j. Representative fluorescence microscopy images of the fluorescence recovery after photobleaching (FRAP) experiments of HOXC4 wild type IDR-mEGFP and HOXC4 AroPERFECT IDR-mEGFP droplets. Shown is one droplet of each genotype, in droplet formation buffer at a concentration of 20 pM. Images were taken every 3 seconds over a time span of 60 seconds. k. Fluorescence intensity of HOXC4 wild type I DR and HOXC4 AroPERFECT I DR in vitro droplets before, during and after photobleaching. Data are displayed as mean ± SD. N = 20.Extended Data Figure 5. Optimizing aromatic dispersion enhances the activity of multiple TF IDRs a. AlphaFold models of OCT4 (top), PDX1 (middle) and FOXA3 (bottom). Highlighted are DBDs (dark blue), predicted activation domains (light blue) and IDRs (orange). b. (left) Schematic models of OCT4 (top), PDX1 (middle) and FOXA3 (bottom) wild type and mutant sequences. Aromatic amino acids are highlighted as orange dots. Alanine mutations are highlighted as white dots. The annotated DNA binding domains are highlighted with a grey background, (right) Results of luciferase reporter assays in V6.5 mouse embryonic stem cells of indicated OCT4 (top), PDX1 (middle) and FOXA3 (bottom) IDRs. Luciferase values were normalized against an internal Renilla control, and the values are displayed as percentages normalized to the activity measured using an empty vector. Data are displayed as mean ± SD from three biological replicates with three technical replicates each. P-values are from Student’s t-tests. *: P<0.05, ***: P<10-3. Note that shown AroPERFECT IDRs have stronger transactivation capacity as their respective wild type sequences. c. Western blot of GAL4-DBD and GAL4-DBD-OCT4-IDR- (top), GAL4-DBD-PDX1-IDR- (middle) and GAL4-DBD-FOXA3-IDR- (bottom) fusion proteins in HEK293T cells 24 hours after transfection using a GAL4-DBD specific antibody. HSP90 is shown as a loading control. Wild type and AroPERFECT mutants are expressed at comparable levels. d. Results of an OCT4 C-IDR tiling experiment by using luciferase reporter assays. Sequences were tiled into fragments of 40 amino acids with 20 amino acid overlaps. The activities of the full-length IDRs are indicated with dashed horizontal lines. e. (left) Schematic model of EGR1 I DR wild type and mutant sequences. Aromatic amino acids are highlighted as orange dots, (right) Results of luciferase reporter assays in V6.5 mouse embryonic stem cells of the EGR1 wild type I DR and mutants. Luciferase values were normalized against an internal Renilla control, and the values are displayed as percentages normalized to the activity measured using an empty vector. Data are displayed as mean ± SD fromthree biological replicates with three technical replicates each. P-values are from Student’s t- tests. *: P<0.05, ***: P<10'3. f. Results of an EGR1 IDR tiling experiment by using luciferase reporter assays. Sequences were tiled into fragments of 40 amino acids with 20 amino acid overlaps. The activities of the full-length IDRs are indicated with dashed horizontal lines. g. (left) Schematic model of HOXB1 IDR wild type and AroPERFECT sequences. Aromatic amino acids are highlighted as orange dots, (middle) Omega plots and QATO scores of the IDRs. (right) Results of luciferase reporter assays in V6.5 mouse embryonic stem cells of the HOXB1 wild type and AroPERFECT IDRs. Luciferase values were normalized against an internal Re- nilla control, and the values are displayed as percentages normalized to the activity measured using an empty vector. Data are displayed as mean ± SD from three biological replicates with three technical replicates each. P-values are from Student’s t-tests. *: P<0.05, ***: P<10-3.Extended Data Figure 6. Characterization of HAP1 HOXD4 knock-in and HOXD4 overexpression cells. a. Scheme of mEGFP knock-in strategy at the H0XD4 locus. b. Flow cytometry analysis of mEGFP expression in HAP1 HOXD4-mEGFP knock-in cell lines. A representative quantification is shown. Data normalized to mode. c. (left) PCR genotyping of HAP1 wild type and HAP1 HOXD4 knock-out lines. d. H0XD4 gene expression levels quantified as RQ value in HAP1 wild type and HAP1 HOXD4 knock-out cells by quantitative real-time PCR. The reduction of normalized H0XD4 expression in HOXD4 knock-out cells was reduced significantly. P-value is from a Student’s t-test. e. Heatmap analysis of RNA-Seq data in the five cell lines. Genes were clustered using k- means clustering on expression values. f. (left) Western blot analysis of GATA6 and ESX1 in the indicated HAP1 knock-in and knockout lines. HSP90 is shown as loading control, (right) Western blot analysis of HOXD4-mEGFP and I Fl 16 in the indicated HAP1 PiggyBac overexpression system. HSP90 is shown as loading control. g. (top) Differential interference contrast (DIC) microscopy of the indicated cell lines. Scale bar is 0.4 mm. (bottom) Fluorescence microscopy images of cells expressing the HOXD4-GFP fusion proteins. Cells were imaged 14 days after constant doxycycline induction. Note that AroPERFECT and AroPLUS cells acquire a similar morphology as the HOXD4 AroPERFECT and AroPLUS knock-in lines. h. Flow cytometry analysis of mEGFP expression in HAP1 HOXD4-mEGFP PiggyBac cell lines after 14 days of Dox induction. A representative quantification is shown. Data normalized to mode.i. Gene expression levels quantified as fold change value in HAP1 PiggyBac clones for HOXD4 wild type and mutant cells by quantitative real-time PCR after 14 days of constant doxycycline induction.Extended Data Figure 7. C / EBPa supporting data. a. The relative amount of condensed protein per concentration quantified in the droplet formation assays. Data are displayed as mean ± SD. N = 10 images from 2 replicates. The curve was generated as a non-linear regression to a sigmoidal curve function. b. Western blot of GAL4-DBD and GAL4-DBD-C / EBPa-IDR fusion proteins in HEK293T cells 24 hours after transfection using a GAL4-DBD specific antibody. HSP90 is shown as a loading control. Wild type and AroPERFECT IS15 mutants are expressed at comparable levels. c. (left) Schematic models of wild type and mutant C / EBPa proteins. The position of the bZIP DNA binding domain is highlighted with a grey box. Aromatic amino acids are highlighted as orange dots, (middle) Omega plots and QATO scores in the IDR. IDR: intrinsically disordered region, (right) Results of luciferase reporter assays. Data are displayed as mean ± SD from three biological replicates with three technical replicates each. P-values are from Student’s t- tests. ***: P<10-3. d. Scheme of FACS analysis strategy for quantification of macrophage differentiation efficiency. e. Flow cytometry analysis of Mac1 and CD19 expression in differentiating RCH-rtTA cells after induction of C / EBPa constructs with doxycycline.Extended Data Figure 8. C / EBPa single-cell RNA-seq supporting data. a. Characterization of scRNA-seq clusters using the data for various stages of B-cell macrophage differentiation from a previous study86. b. Quantification of the cluster’s genes for each k-cluster of the heatmap. Based on the quantification and expression profile of the heatmap the single cell clusters were manually assigned. c. RNA velocity stream plot was embedded to pre-computed LIMAP plot. d. Quantification of mEGFP-positive cells in the initial clusters. Cluster 0 and 2 contain virtually no mEGFP-positive cells and were therefore removed from downstream analyses. e. Sample proportions for each cluster. Differentiating macrophage 1 is wild type-specific and Differentiating macrophage 2 is AroPERFECT IS15-specific. AroPERFECT IS10 cells are absent from the macrophage clusters. f. (left to right) Combined LIMAP colored CD14 and PTPRC, CD19 and ITGAM (MAC1) gene expression. These markers are associated with macrophage differentiation. g. Top 5 differentially expressed genes per cluster.h. Stacked violin plots for select DEG genes for Late macrophage cluster between AroPER- FECT IS15 and Wild type.Extended Data Figure 9. C / EBPa ChlP-seq supporting data. a. Principal component analysis of the Chi P-Seq peak profiles for wild type and AroPERFECT IS15 C / EBPa expressing cells 24h and 48h after induction of C / EBPa expression (PC1 vs. PC2). b. Heatmap representation of Chi P-Seq read densities of C / EBPa Wild type and AroPERFECT IS15 within a 1.5kb window around all shared C / EBPa peaks (top), differentially enriched peaks in C / EBPa AroPERFECT IS15 (center) and differentially enriched peaks in C / EBPa Wild type (bottom). F.o.l: fold over input. c. Heatmap representation of Chi P-Seq read densities of C / EBPa wild type and AroPERFECT IS15 at 24h and 48h after induction of C / EBPa overexpression. Regions between the start and end co-ordinates were length-normalized. F.o.l: fold over input. d. Fluorescence microscopy images of differentiating RCH-rtTA cells expressing GFP-tagged versions of C / EBPa. (left) localization of C / EBPa-GFP fusion protein, (right) Nuclear counterstain. e. Flow cytometry analysis of GFP expression in RCH-rtTA cell lines expressing GFP-tagged versions of C / EBPa. Data normalized to mode. f.h. C / EBPa AroPERFECT IS15 shows enhanced binding at the HSPA8 (f) and PPP1R17 (h) loci. Displayed are genome browser tracks of Chi P-Seq data of C / EBPa Wild type and AroPERFECT IS15 in RCH-rtTA cells, 24 and 48 hours after C / EBPa expression. Co-ordinates are hg38 genome assembly co-ordinates. F.o.l: fold over input. g.i. UMAPs colored on HSPA8 (g) and PPP1R17 (i) expression. j. C / EBPa AroPERFECT IS15 shows enhanced binding at the CEACAM gene cluster. Displayed are genome browser tracks of ChlP-Seq data of C / EBPa Wild type and AroPERFECT IS15 in RCH-rtTA cells, 24 and 48 hours after C / EBPa expression. Co-ordinates are hg38 genome assembly co-ordinates. k. Combined LIMAP colored on CEACAM8 and CEACAM1 expression, the same markers highlighted as differentially upregulated in C / EBPa AroPERFECT IS15 cells in the Late Macrophage cluster. l. Flow cytometry analysis of CD66 expression in differentiating GFP+ RCH-rtTA cells Oh and 48h after induction of C / EBPa overexpression. Data normalized to mode.m. Displayed are genome browser tracks of ChlP-Seq data of C / EBPa wild type and AroPER- FECT IS15 in RCH-rtTA cells, 24 and 48 hours after C / EBPa overexpression. Co-ordinates are hg38 genome assembly co-ordinates. n. Combined LIMAP colored on FCGR2B and FCGR2A expression, the same markers highlighted as differentially upregulated in C / EBPa AroPERFECT IS15 cells in the Late Macrophage cluster. o. Flow cytometry analysis of FCGR2A expression in differentiating GFP+ RCH-rtTA cells Oh and 48h after induction of C / EBPa overexpression. Data normalized to mode.Extended Data Figure 10. NGN2 supporting data. a. (left) Schematic models of wild type and mutant NGN2 proteins. The position of the bHLH DNA binding domain is highlighted with a grey box. Aromatic residues are highlighted as orange dots, (middle) Omega plots and QATO scores of the IDRs. I DR: intrinsically disordered region, (right) Results of luciferase reporter assays in V6.5 mouse embryonic stem cells. Luciferase values were normalized against an internal Renilla control, and the values are displayed as percentages normalized to the activity measured using an empty vector (dashed orange line). Data are displayed as mean ± SD from three biological replicates with three technical replicates each. P-values are from Student’s t-tests. ***: P<10-3. b. Representative images of droplet formation of purified NGN2 C-terminal IDR-mEGFP fusion proteins at the indicated concentrations in droplet formation buffer. Scale bar is 5 pm. c. The relative amount of condensed protein per concentration quantified in the droplet formation assays. Data are displayed as mean ± SD. N = 10 images from 2 replicates. The curve was generated as a non-linear regression to a sigmoidal curve function. d. Western blot analysis of FLAG-NGN2 in the indicated ZIP13K2 PiggyBac overexpression lines. HSP90 is shown as loading control. e. Fluorescence microscopy images of differentiating ZIP13K2 cells expressing indicated FLAG-tagged versions of NGN2. (top) Nuclear counterstain (DAPI). (middle) Visualization of cellular localization of NGN2 wild type and mutants using an a-FLAG antibody, (bottom) Endogenous mEGFP signal of NGN2-T2A-mEGFP over-expressing cells. f. Heatmap analysis of RNA-Seq data in the four cell lines. Genes were clustered using k- means clustering on expression values. Expression values are represented by scaling and centering VST transformed read count normalized values (z-score). K-means clustering was used to define the clusters. g. Marker gene analysis from selected genes from single cell cluster markers in NGN2 induced neural differentiation. The expression values of RNA-Seq data in the four cell lines are shown for the genes identified as marker genes in the indicated populations in previous studies.h. Principal component analysis of the ChlP-Seq peak profiles (ChlP-Seq) of ZIP13K2 Wild type, ZIP13K2 NGN2 Wild type, ZIP13K2 NGN2 AroLITE and ZIP13K2 NGN2 AroPERFECT cells (PC1 vs. PC2 and PC1 vs. PC3). i. NGN2 AroLITE loss of binding at the SERTM1 locus. Displayed are genome browser tracks of ChlP-Seq data of NGN2 Wild type, AroLITE and AroPERFECT in ZIP13K2 cells, 24 and 48 hours after NGN2 overexpression. Co-ordinates are hg38 genome assembly co-ordinates. j. Enrichment scores of bHLH TF motifs, and adjusted P-values of enrichment at the three indicated peak sets. k. Heatmap analysis of TT-SLAM-seq data in the four cell lines 12h and 24h after transgene induction. Genes were clustered using k-means clustering on expression values. Expression values are represented by scaling and centering VST transformed read count normalized values (z-score). K-means clustering was used to define the clusters.Extended Data Figure 11. MYOD1 supporting data. a. Western blot of GAL4-DBD and GAL4-DBD-MYOD1 C-IDR-fusion proteins in HEK293T cells 24 hours after transfection using a GAL4-DBD specific antibody. HSP90 is shown as a loading control. Wild type and AroPERFECT mutants are expressed at comparable levels. b. Results of a MYOD1 C-IDR tiling experiment by using luciferase reporter assays. Sequences were tiled into fragments of 40 amino acids with 20 amino acid overlaps. The activities of the full-length IDRs are indicated with dashed horizontal lines. c. Fluorescence images of C2C12 myoblasts at day 0 and 1 after induction of MYOD1 wild type, MYOD1 AroLITE, MYOD1 AroPERFECT C or MYOD1 AroLITE C transgene with doxycycline. DAPI was used as DNA counterstain (magenta). Co-expressed mEGFP of the MYOD1- T2A-mEGFP fusion protein was used as cytoplasmic marker (cyan). Scale bar 0.5 mm. d. Principal component analysis of the RNA-Seq expression profiles. e. Differential expression analysis. f. Heatmap analysis of RNA-Seq data in the six cell lines. K-means clustering was used to define the clusters. g. Gene set enrichment analysis (GSEA) of differentially expressed genes in the MYOD1 AroPERFECT C RNA-Seq sample.Examples1. Material and Methods1.1 Generation of Doxycycline-inducible NGN2 overexpression systems in human iPS cells To generate a doxycycline-inducible overexpression system of NGN2, we randomly integrated the coding sequences of NGN2 wild type, AroLITE and AroPERFECT into ZIP13K2 cells by using the PiggyBac transposon system.N-terminally FLAG-tagged coding sequences of human NGN2 wild type, AroLITE or AroPERFECT (Twist Bioscience) with a downstream T2A tag (Sigma) were cloned into a backbone of the inducible Caspex expression vector (Addgene #97421) linearized by restriction digest with Ncol (NEB) and Kpnl (NEB). Carrier plasmids and PiggyBac transposase expression vector (SBI, PB210PA-1) were co-transfected into ZIP13K2 wild type cells using Lipofectamine Stem Transfection Reagent (Thermo Scientific) following the manufacturer’s instructions in a molar ratio of 6:1. The transfected bulk population was screened for integration by addition of 2 pg / ml puromycin (Gibco) to the cell culture medium 24h after transfection for a total of 4 days. Surviving cells were seeded at low density under addition of 1x Y-27632 Rho-kinase inhibitor (biogems, 1293823) for the first 24 hours and expanded for several days until colonies derived from single cells were big enough to be picked and cultured separately. Clones of every condition were induced by addition of 2 pg / ml Doxycycline (Sigma) and screened for matching mEGFP expression levels across conditions with flow cytometry.1.2 Generation of Doxycycline-inducible MYOD1 overexpression lines in C2C12 cellsTo generate a Doxycycline-inducible overexpression system of MYOD1 , we randomly integrated the coding sequences of MYOD1 wild type, AroLITE, AroPERFECT C and AroLITE C into C2C12 cells by using the PiggyBac transposon system.N-terminally FLAG-tagged coding sequences of human MYOD1 wild type, AroLITE, AroPERFECT C or AroLITE C (Twist Bioscience) with a downstream T2A (Sigma) were cloned into a backbone of the inducible Caspex expression vector (Addgene #97421) linearized by restriction digest with Ncol (NEB) and Kpnl (NEB). Carrier plasmids and PiggyBac transposase expression vector (SBI, PB210PA-1) were co-transfected into C2C12 wild type cells using Lipofectamine3000 transfection reagent (Thermo Scientific) following the manufacturer’s instructions in a molar ratio of 6:1. The transfected bulk population was screened for integration by addition of 2 pg ■ mL-1puromycin (Gibco) to the cell culture medium 24h after transfection for a total of 4 days. Cells of every condition were induced by addition of 2 pg / ml Doxycycline(Sigma) and screened for matching mEGFP expression levels across conditions by flow cytometry.1.3 MYOD1-mediated myogenic differentiation of C2C12 myoblastsC2C12 myoblasts with integrated MYOD1 overexpression cassette were seeded on chambered p-Slide 8 Well ibiTreat coverslips (Ibidi). Upon reaching 85-90% confluence, 2 pg / ml Doxycycline was added to culture medium to induce expression of MYOD1 transgene. Differentiation medium was changed every day over 3 days. For imaging, cells were washed with PBS and fixed with 4% PFA for 15 minutes at room temperature. Cells were counterstained with DAPI (Fig. 7c, Extended Data Fig. 11c).1.4 Protein purificationOverexpression of recombinant protein in BL21 (DE3) (NEB M0491S) was performed as described20.1.5 In vitro droplet assayFor in vitro droplet formation experiments (Fig. 1g, 2b, 4b, Extended Data Fig. 2c, 2g, 4h, 10b), we measured the concentration of purified mEGFP IDR fusion proteins with a NanoDrop2000 (Thermo Scientific) and subsequently diluted protein preparations to the required concentration in Storage Buffer (50 mM Tris pH 7.5, 125 mM NaCI, 1 mM DTT, 10% Glycerol). The in vitro droplet formation assay was performed as previously described21.1.6 Fluorescence recovery after photobleaching (FRAP)Fluorescence recovery after photobleaching experiments on transcription factor IDR droplets were formed as described above without 30 minutes of pre-assembly at room temperature at a protein concentration of 25 iM . Formed droplets were bleached immediately after pipetting the protein mixture onto the slide by using 488 nm light at 70% laser power in 10 iterations. Bleaching was performed on a central region of a settled single droplet. Fluorescence recovery was measured over a time course of 60 seconds in 2-second intervals. Quantification of FRAP data was based on at least ten images acquired in at least two independent image series per condition. The resulting signal recovery was normalized to background and fitted to a power law model in Microsoft Excel. All figures were generated using GraphPad PRISM9 (Fig. 2d, 4c, 6b, Extended Data Fig. 4a, 4j-k).1.7 Generation of DNA constructs for transactivation assaysTo study transactivation strength of transcription factor IDRs we amplified sequences from codon optimized gene fragments (Twist Bioscience). Amplified gene fragments were clonedinto a pGAL4 (Addgene #145245) backbone, linearized with AsiSI (NEB) and BsiWI (NEB) via NEBuilder HiFi Assembly.1.8 Generation of DNA constructs for TF-IDR tiling assaysTo control for the spontaneous creation of short linear motifs engendering transcriptional activation in TF-IDR mutants, we tiled up the HOXD4 wild type, HOXD4 AroPERFECT, C / EBPa wild type, C / EBPa AroPERFECT IS15, OCT4 wild type C-, OCT4 AroPERFECT c-, MYOD1 wild type C-, MYOD1 AroPERFECT C-, EGR1 wild type and EGR1 AroSCRAMBLED IDRs into 40 amino acid segments with 20 amino acid overlaps. Amplified gene fragments were cloned into a pGAL4 (Addgene #145245) backbone, linearized with AsiSI (NEB) and BsiWI (NEB) via NEBuilder HiFi Assembly (Fig. 2e, 4f, Extended Data Fig. 5d, 5f, 11 b).1.9 Transactivation assayThe transactivation activity of TF IDRs was assayed using the Dual-Glo Luciferase Assay system (Promega). Mouse embryonic stem cells were seeded on gelatin pre-coated 24-well plates with a density of 1x105cells per cm2. For feeder-free culture conditions, mESC medium was supplemented with 2x LIF. HEK-293T, SH-SY5Y, Kelly cells and C2C12 mouse myoblasts were seeded on 24-well plates with a density of 1x105cells per cm2. After 24 hours, every well was transfected with 200 ng pGal4 empty vector control or the equimolar amount of the expression construct carrying an I DR of interest, 250 ng of the Firefly luciferase expression vector (Promega) and 15 ng of the Renilla luciferase expression vector (Promega) using FuGENE HD transfection reagent (Promega) following the manufacturer’s instructions. After 24 hours, cells were washed once with PBS and lysed in 100 pl of 1x Lysis Passive Buffer (Promega) for 15 minutes on a shaker at room temperature. Subsequently, 10 pl of cell lysate was pipetted onto a white bottom 96-microwell plate in triplicates followed by quantification of Firefly and Renilla using the Dual-Glo Luciferase Assay System Quick Protocol for 96-well plates (Promega). Triplicate data was normalized to Renilla luminescence of the respective well and finally normalized to the empty vector control. Data are shown as mean ± SD. All data shown were generated of three independent transfections from at least two cell passages (Fig. 1 i, 2a, 2e-g, 4a, 4f-g, 5a, 7a, Extended Data Fig. 2e, 2i-m, 4d, 4f, 5a, 5e, 5g, 7c, 10a). All data were plotted with GraphPad PRISM9. To assess statistical significance, two-tailed t-tests were performed.1.10 LacO-LacIFor LacO-LacI tethering experiments (Fig. 3g, 4d), we used a vector containing CFP-Lacl followed by a multiple cloning site (MCS) from20. MED1-IDR and POLR2-CTD plasmids werecloned via digestion with AsiSI (NEB) and BsiWI (NEB) via NEBuilder HiFi Assembly Master Mix. Tethering experiments were adapted from201.11 NGN2-mediated neural differentiation of human iPSCsWe adapted our protocol for the differentiation of human iPSCs into neurons by overexpression of NGN2 from41. ZIP13K2 cells with integrated NGN2 overexpression cassette were cultured on Matrigel (Corning) pre-coated 10 cm culture plates. When cultures reached a confluency of approximately 80%, 2 pg / mL Doxycycline (Sigma) was added to the culture medium to induce expression of the NGN2 transgene. After 24 hours, induced cultures were sorted for mEGFP expressing cells by flow cytometry. Positive cells were seeded on Matrigel pre-coated 96-well microclear plates (Greiner bio-one) in mTeSR+ and 1x Rho-kinase inhibitor at a density of 2x104cells I cm2. On day 2, mTeSR+ medium was replaced by N2B27 neural cell culture medium supplemented with 5 pg / ml human BDNF (biotechne). Differentiation medium was changed every day for a total of 4 days. Living cells were stained with 0.25 pg / ml Hoechst and Spy650-TUB (1 :2000) (Spirochrome) and incubated in the microscope prior to image acquisition to equilibrate and thermalize all materials (Fig. 6d-f).1.12 KAPA Stranded mRNA-seq of ZIP13K2 NGN2 PiqqyBac cellsAt day 5 of NGN2-mediated neural differentiation, ZIP13K2 induced neurons were harvested following the Direct-zol™ RNA MicroPrep Kit (Zymo Research) standard protocol. 1 pg of RNA of each sample was used as input for library preparation using the KAPA Stranded mRNA-Seq Kit (Roche) following the manufacturer’s instructions. Unique Dual-Indexed Set-B (UDI; KAPA biosystems) adapters were ligated and the library was amplified for 8 cycles. Libraries were sequenced on a Novaseq6000 as Paired-end 100 with 50 million fragments per library (Fig. 6g-h, Extended Data Fig. 10f-g).1.13 Live-cell imaging of human iPSC derived neuronsLiving cells were imaged using the Celldiscoverer 7 Imaging Platform (Zeiss), in wide field mode running under ZEN Blue v 3.1 and full environmental control (5% v / v CO2, 100% humidity, 37°C). Final experiments were performed using the Plan-Apochromat 20x / NA=0.7 objective, a 2x tube lens (Zeiss), and captured on an Axiocam 506 (Zeiss) with 3x3 binning resulting in a lateral pixel resolution of 0.347 pm I pixel.1.14 Image analysis of nuclei and neurite densities in differentiated neuronsWide field images were acquired using a 20x Air Objective (NA=0.7) a 2x optical post magnification on a Celldiscoverer 7 under ZEN 3.2 Blue (Zeiss Germany). Object quantification was performed in the image analysis module in ZEN 3.4 (Zeiss, Germany). Briefly, within MIPsnuclei were identified by nuclear counter staining using Otsu intensity thresholds after faint smoothing (Gauss: 2,0), close by objects were segmented downstream by standard water shedding. Neurites were segmented by fixed intensity threshold on the respective staining without any water shedding (Fig. 6e-f).1.15 FLAG-NGN2 and C / EBPa-GFP Chromatin iTo study chromatin association of NGN2 wild type, AroLITE and AroPERFECT, we performed ChlP-seq experiments in ZIP13K2 NGN2 wild type, AroLITE and AroPERFECT -expressing cells 24 and 48 hours after induction of NGN2-mediated neural differentiation.Cells were detached with Accutase solution (Sigma), washed twice in PBS and fixed in rotation with 1 % formaldehyde for 10 minutes at room temperature. Subsequently, the reaction was quenched by addition of glycine leading to a final concentration of 125 mM. Per replicate, three million cells were used as starting material. In brief, we followed the ChlPmentation protocol version 3 for histone marks and transcription factors60. Sequencing libraries were amplified using the Kapa HiFi HotStart Ready Mix (Roche) and Nextera custom primers (Illumina)61for a total of 12 cycles and paired-end sequenced on an NovaSeq6000 (Illumina) with a depth of ~50 million fragments per library (Fig. 6i-j, Extended Data Fig. 10h-i).1.16 Image analysis of differentiated C2C12 myotubesWide field Images were acquired using a 20x Air Objective (NA=0.7) a 2x optical post magnification on a Celldiscoverer 7 under ZEN 3.2 Blue (Zeiss Germany). For each well and replicate a mosaic of 49 tile regions was covered. We defined the definite hardware focus as the center for 3 slices of a consecutive z-stack with a slice distance of 0.34 pm. Image acquisition was performed using a Zeiss Axiocam 506, in a 3x3 bininng mode, resulting in a lateral resolution of 0.34 pm / pixel. The resulting images were projected using Maximum Intensity Projection (MIP) in ZEN 3.4 (Zeiss, Germany) on a dedicated Zeiss analysis workstation. Quantification of Fusion scores was conducted by implementation of a simple hierarchy order which was built within the image analysis module in ZEN 3.4 (Zeiss, Germany). We designed two segregating parent-classes by fixed intensity thresholds based on meGFP signal resulting in fused myotubes (MT) and non-myotubes (NMT). Within these primary regions, nuclei were identified. Secondary objects were identified exclusively within primary objects (MT, NMT) by applying gaussian smoothing and fixed intensity thresholds on the nuclear counter staining, followed by standard water shedding the respective fluorescence image. All nuclei objects were filtered according to an area in between 30 and 300 pm2(Fig. 7d).1.17 C / EBPa mediated B-cell to macrophage transdifferentiationTo induce C / EBPa-mediated B-cell to macrophage transdifferentiation, infected RCH-rtTA cells were seeded at 0.3x106cells / ml in RCH culture medium supplemented with 10 ng / ml of each, IL-2 (200-03, Preprotech) and CSF-1 (315-03B, Preprotech), as well as 2 pg / ml of Doxycycline. The macrophage transdifferentiation was monitored by flow cytometry. Briefly, blocking was carried out for 10 min at room temperature using human a 1 :20 dilution of FcR binding inhibitor (eBiosciences, 16-9161-73). Subsequently, cells were stained with antibodies against CD19 (APC-Cy7 Mouse anti-Human CD19, BD Pharmingen, 557791) and Mac-1 (APC Mouse Anti-Human CD11 b / Mac-1 , BD Pharmingen, 550019) at 4 °C for 20 min in the dark. After washing, DAPI counterstaining was performed just before analysis. All analyses were performed using an LSR Fortessa instrument (BD Biosciences). Data analysis was completed using FlowJo software (Fig. 5c, Extended Data Fig. 7e).1.18 Single cell RNA-seq (sc-RNA-Seq) data generationOne week after induction of C / EBPa-mediated B-cell to macrophage transdifferentiation, cells were collected and washed twice in PBS to remove dead cells and debris. Cells were resuspended in solution at a density of 700 cells / pl. We used Chromium Next GEM Single Cell 3’ technology for generating gene expression libraries from single cells. The final libraries contained the P5 and P7 primers used in Illumina bridge amplification. Final libraries were analyzed using Agilent Bioanalyzer assay to estimate the quantity and check size distribution and were then quantified by qPCR using the KAPA Library Quantification Kit (ref. KK4835, Kapa- Biosystems).1.19 Identification of periodic blocks in TF IDRsWe used 1 ,392 full length TF protein sequences from AnimalTFDB3.062and determined the positions of all aromatic residues F, Y, W (stickers) within them. Next, we identified spacers, stretches of non-aromatic residues between the stickers. A periodic block of aromatic residues was defined as a region that comprises at least 4 aromatic amino acids. We considered spacer lengths of 4-9 amino acids, 10-20 amino acids or 21-30 amino acids. The ranges of different spacer lengths used for the analysis were chosen based on previous modeling studies on biopolymers using the stickers-and-spacers formalism63-65. Next, we identified periodic blocks that overlap I DR regions using the Metapredict v2 I DR prediction network66. This resulted in the identification of periodic blocks of aromatic residues in 531 TF IDRs (Extended Data Fig. 1a-b). For an internal ranking of periodic TF IDRs, we calculated a periodicity score comprising the number of periodic blocks that overlapped with the protein IDRs. The three spacer subgroups were weighed by 1 , 1.1 and 1.2 for the lengths of 4-9, 10-20 and 21-30 residues in asingle spacer, respectively. The weighing values were arbitrarily chosen with the assumption that uniform aromatic dispersion with long spacers may be less likely to occur randomly (Extended Data Fig. 1a-b).1.20 Identification of intrinsically disordered protein regionsIntrinsically disordered protein regions (IDRs) were predicted using Metapredict at default settings using the Metapredict v2 network66.1.21 Identification of regions with significant periodicity in the human proteomeWe developed an in-house method to identify regions with significant, albeit not necessarily perfect, periodicity. In brief, the number of residues between adjacent aromatic residues (i.e. , "spacer length") was calculated for each protein, and the observed distribution of spacer lengths within a seguence was compared to the expected geometric distribution using a Kolmogorov-Smirnov (K-S) test. Then, the mean of the geometric distribution was extrapolated from the proportion of aromatic residues, implicitly modeling their occurrence by a Poisson process. Next, the method was applied to every 100 amino acid-long regions using a sliding window approach, and the P-value of the K-S test was plotted against the position of each window in every protein. After plotting the P-value of every 100 acid-long regions for each protein, the consecutive points below a P-value threshold (=0.5 * average(P-value)) were identified periodic regions. Those regions were compared to the Metapredict IDRs and InterPro domain regions (htps: / / www.ebi.ac.uk / interpro / ), and overlap was defined as the overlap between regions of at least one amino acid. Only regions that contained at least 5 aromatic residues in the 100 amino acid window with the lowest P-value were included. Regions with significant periodicity were defined by the min P-value cutoff of 0.01. (Fig. 1 k-l, Extended Data Fig. 3a-b).1.22 Omega score (QATO) calculationOmega score was calculated using a modified localCIDER version68. Since the Omega score function is not length normalized, we adapted the python code to allow for variable interspace size referred in the package as the so-called blob size. This parameter is now calculated by dividing the seguence length over the fraction of aromatic residues. For this analysis only IDRs with a minimum of 3 aromatic residues were included. The mean of random score was defined as the mean of 1000 kappa score calculations of randomly shuffled seguence from the original seguence. ggplot269was used for plotting violin plots and custom R to generate a distribution plot for the mean of random (Fig. 1f, 2a, 4a, 6a, 7a, Extended Data Fig. 2b, 4c, 4f, 5g, 7c, 10a). One-way ANOVA with a post-Tu key test was used to compare IDR sets (Fig. 11).1.23 Spacer analysisI DR composition was measured by calculating the frequency of each amino acid as a probability with the ‘alphabetFrequency’ function from Biostrings package v 2.40.2, divided by the frequency of the amino acid calculated over the full human proteome in R. Quantification was performed for IDRs with and without periodic blocks. The frequency bar chart was plotted with ggplot (Extended Data Fig. 1d). To calculate the amino acid composition around the aromatic residues we extracted the sequence in fasta format of every periodic block for positions -2, -1 , 0, +1 and +2 around the aromatic residue (0 represents the aromatic residue) using custom python script. The Fasta file was then submitted to GLAM2 analysis to calculate frequency of amino acids and to output a PWM. The cumulative bar plot was plotted with ggplot masking the PWM table into disorder promoting, order-promoting and neutral residues (Extended Data Fig. 1e). Periodic block motif analysis was performed by extracting sequences of the periodic blocks in TF IDRs described in this study and charged- blocks from a previous study70, in fasta format and submitting them to GLAM2 analysis. Top 3 PWM were plotted (Extended Data Fig. 1f).1.24 Bulk RNA-seq analysisRNA-seq data from HAP1 and ZIP13K2 cells were mapped to a custom human genome hg38 including the cloned mEGFP sequence using STAR aligner72. Count read tables were generated by the same program. C2C12 RNA-seq data was mapped to mm10 mouse genome using the above-mentioned programs. Differential expression analysis was performed by the DEseq2 package73in R version 4.2.74. Principal component analysis was carried out using the PCAPIot function from the DEseq2 package on the normalized read matrix that was transformed by using the variance stabilizing transformation (VST) function from the DEseq2 package and plotted by ggplot2 (Fig. 3d, 6g, 7e, Extended Data Fig. 11 d). Heatmaps were plotted with the aid of the ComplexHeatmap package75in R, and cluster analysis was done by k- means clustering using the cluster76package in R (Extended Data Fig. 6e, 10f, 11f).Gene set enrichment analysis of the MYOD1 RNA-seq was conducted using GSEAPreranked v6.0.1277with 1000 permutations on ranked list of gene set from the comparisons AroPER- FECT-C vs wild type and Wild type vs Parental sorted by Wald statistic (stat)73against Wikipathways cell adhesion gene set in Mus musculus78(Extended Data Fig. 11g). Highest ranking genes in the AroPERFECT-C vs wild type that are MYOD1 targets were highlighted in the volcano plots (Fig. 7f, Extended Data Fig. 11e).Marker genes shown in Extended Data Fig. 10g were identified single cell cluster markers in NGN2 induced neural differentiation in previous studies79’80.1.25 Gene set enrichment analysisGene Ontology enrichment analysis for qlDR and the periodic block containing TFs was done on ql DR proteins of our ranked list using gProfiler81. GO categories for biological process were filtered for term size above 1000 genes to remove general categories. Cuttoff of 0.001 of P- value adjusted was used. For periodic block containing TFs analysis REAC and WikiPathways enrichment was also done with gProfiler. Gene-set enrichment analysis was done using clus- terProfiler82 83(Extended Data Fig. 3d-e).1.26 Single-cell RNA-seq analysisData pre-processingThe single-cell RNA-seq datasets were processed with 10x Genomics' Cell Ranger pipeline version 3.1.084and mapped to a custom human genome hg38 including meGFP and codon- optimized C / EBPa wild type, C / EBPa AroPERFECT IS15, and C / EBPa AroPERFECT IS10 sequences. The Cell Ranger hdf5 files were processed by Seurat package version 4.0.685.Filtering and normalizationWe kept cells with more than 2000 expressed genes and genes with more than 5 reads across the samples were considered for analysis. Further filtering was done by removing cells with more than 20% or more mitochondrial genes and less than 5% ribosomal gene expression. The top 10 genes associated with PCA components were then checked for mitochondrial and ribosomal genes. Finally, the Harmony package was used to batch correct the 3 libraries.Cluster identificationCluster identification was then carried out by Seurat's built-in functions FindvariableGenes, RunPCA, RunllMAP, and FindClusters by first identifying the genes with the highest variation across all samples and cell types, building a shared nearest neighbor graph, and then running the Louvain algorithm on it. The number of clusters was determined by the optimum of the modularity function from the Louvain algorithm. The number of mEGFP positive cells was then calculated for each cluster, and this was used to filter untransformed cell clusters, mainly cluster 0 and cluster 2.Assignment of cell types to clustersCell type cluster assignment was based on comparison of market set from bulk RNA-seq experiment from86and augmented by both RNA velocity analysis and known markers for both B-cell and macrophage cell types. In brief, RNA-Seq data and marker set was retrieved from86and raw fastq files, aligned and reads were counted with STAR aligner against the humangenome v38. Raw count data was then processed in DESeq2 and normalized by VST transformation. Marker set VST data was then retrieved and clustered according to the methods described in86and each gene was assigned a gene cluster for Early, Early-inter, Interl , Inter2, Inter-late, Latel , and Late2 as descripted in the publication above. This assignment was designated “Choi et al. differentiation clusters” in Extended Data Fig. 8a-b. As mentioned before clusters 0 and 2 were excluded based on mEGFP quantification (Extended Data Fig. 8d) and were considered as untransduced B-cells.Differential expression analysisInter-cluster differential expression analysis was performed by Wilcox test using the FindMark- ers function with default setting and inter-sample cluster differential expression analysis between wild type and IS15 cells in cluster 7 using the FindMarkers function with DESeq2 function. The differentially expressed genes within the clusters are listed in Table S5. A q-value cutoff of 0.05 was used to define differentially expressed genes for the Wilcox test, and an adjusted p-value of 0.05 was used for the inter-sample test, listed in Table S5. Volcano plots and bar plots were plotted in ggplot2, violin, LIMAP, and feature plots by Seurat's VlnPlot, FeaturePlot, and DimPlot function. Dot plot was made using a custom function to modify the output of the complexHeatmap package (Extended Data Fig. 8i-j).RNA velocityWe generated loop files necessary for RNA velocity using velocyto87and exported barcodes, expression matrix, metadata and LIMAP coordinates from Seurat to csv files. scVelo88was used to build the manifold, calculate and visualize the RNA velocity using generalized dynamical model to solve the full transcriptional dynamics. PAGA graph89was calculated from this model to visualize cell trajectory. Pseudotime was calculated by Markov diffusion process and plotted by scVelo bult-in function (Extended Data Fig. 8c).1.26 ChlP-seq analysisChlP-seq data from C / EBPa and NGN2 were mapped to a custom human genome hg38 using BWA vO.7.1790. Samtools91was used for SAM to BAM conversion, sorting and indexing and Genome Analysis Toolkit v492was used to remove duplicated reads. Peak calling was then performed using MACS3 v3.0.0 b193using the input of the respective sample. Analysis and differential peak calling were done with DiffBind v3.6.594. Normalization was done with native method and background input. Differential calling was done using the DEseq2 method and FDR threshold was set to 0.01. Peak visualization was performed using the DiffBind “plotprofile” function with default settings for general profiles if otherwise stated. Set of overlapping sites was done with bedtools v2.6.0 and intersect function. Profiles in Extended Data Fig. 9cwere plotted using “percentOfRegion” with 27 windows and 300% extension. The regions plotted correspond to a merged set of promoters, a merged set of enhancers and separate sets for B-cell and Macrophage super enhancers from a previous study86. PCA analysis was done on normalized count samples and plotted with DiffBind (Extended Data Fig. 9a, 10h).For track visualization MACS3 backgroup subtracted bigwig files from each replicate were merged using LICSC Bigwigmerge tool and then converted from bigbedgraph format back into bigwig using LICSC bedGraphToBigWig tool. Visualization was done with pygenometracks tool set95.2. Results2.1 Human TFs encode short periodic blocks of aromatic amino acidsPrion-like domains of RNA-binding proteins contain periodically arranged aromatic residues that promote phase separation25, but whether transcription factors (TFs) contain periodically arranged aromatic residues is not known. To gain initial insights into the extent of periodicity of aromatic residues in human TFs, we developed a computational pipeline to identify short blocks of periodically arranged aromatic residues with varying spacer lengths in -1 ,500 previously curated human TFs (Fig. 1a)1. We filtered for periodic blocks of at least four aromatic residues that overlap intrinsically disordered regions (IDRs), which revealed 531 TF IDRs that contain at least one periodic block (Fig. 1 b, Extended Data Fig. 1a-b). Only 60 / 531 TF IDRs that contained a short periodic block also contained a ‘minimal’ activation domain annotated from four recent studies8-11, and they overlapped in only 31 TF IDRs (Fig. 1a-c, Extended Data Fig. 1c), suggesting that the periodic blocks are distinct from minimal activation domains. TF IDRs with periodic blocks were enriched for aromatic residues and serines and depleted of charged residues (Extended Data Fig. 1d-f), consistent with typical sequence features of aromatic “stickers” and serine / glycine-rich “spacers” in prion-like domains24-26.To quantify the extent of periodicity, we generated a “periodicity score” as a weighted sum of the periodic blocks of aromatic residues and used the score to rank TFs based on the periodicity score of their IDRs (Fig. 1 b, Table S1). The periodicity score was further validated by calculating a previously described patterning parameter (QATO)25. The QAro score measures the extent of mixing of aromatic residues, where high dispersion leads to a low QArovalue, which is then compared to the mean dispersion of 1000 randomly generated sequences25. For example, the 30 aromatic residues in the NFAT5 IDR are more uniformly dispersed than in 1000 / 1000 randomly generated sequences of identical composition (QAro=0.124, empirical P- value=0) (Fig. 1c-d). These results suggest that at least 30% of human TF IDRs contain shortblocks of periodically arranged aromatic residues, and some of the observed periodicity appears non-random.Three TF IDRs that encode periodic blocks of aromatic residues were selected for functional testing (HOXB1 , HOXD4, HOXC4). All three purified, recombinant, GFP-tagged IDRs formed droplets in a concentration-dependent manner in the presence of a crowding agent (10% PEG 8000), with saturation concentrations (Csat) -0.5-5 pM (Fig. 1e-h, Extended Data Fig. 2a-d). The droplets underwent fusion, and wetted the surface of the microscopy slide, which are hallmarks of liquid-liquid phase separation27. Substitution of aromatic residues (“AroLITE” mutants) reduced droplet formation (Fig. 1g-h, Extended Data Fig 2c-d). Substitution of aromatic residues with alanines, serines or glycines had similar effects (Extended Data Fig. 2f-h). As a test of transcriptional activity, the wild type IDRs fused to the GAL4 DBD activated transcription of a 5X-UAS-driven luciferase reporter when transfected into murine embryonic stem cells (-1.5-100 -fold, P<0.05, t-test), and substitution of aromatic residues virtually abolished activity (Fig. 1 i, Extended Data Fig. 2e, 2i). Substitution of aromatic residues also reduced transactivation of the NFAT5, EGR1 , and NANOG IDRs that contain periodic aromatic blocks (Extended Data Fig. 2j-l). Measuring the transcriptional activity in various cell types led to similar results (Extended Data Fig. 2m). These findings suggest that aromatic residues are necessary for in vitro phase separation and transactivation capacities of TF IDRs that contain periodic blocks of aromatic residues.2.2 Submaximal periodicity of aromatic residues in TF IDRsWe noted that many TF IDRs contain short periodic blocks of aromatic residues, but the overall periodicity of TF IDRs tends to be limited (e.g., Fig. 1e-f). We thus hypothesized that TF IDRs might differ from well-characterized prion-like domains in RNA-binding proteins in that their aromatic residues are dispersed in a pattern substantially less uniform than the theoretical maximum. To test whether aromatic residues in TF IDRs are arranged with submaximal periodicity we quantified periodicity using several approaches. We developed a method to identify protein regions with significant, albeit not necessarily perfect periodicity, independent of sequence length and composition. The spacer length between adjacent aromatic residues was calculated for each protein, and the observed distribution of spacer lengths within a sequence was compared to the expected geometric distribution using a Kolmogorov-Smirnov (K-S) test (Fig. 1j). The mean of the geometric distribution was extrapolated from the proportion of aromatic residues, implicitly modeling their occurrence by a Poisson process. The method was applied to 100 amino acid-long regions using a sliding window approach, and the P-value of the K-S test was plotted against the position of each window in every protein in the human proteome. The P-value and length of the regions encompassing 100 residue windows belowa P-value threshold were used to define regions with significant periodicity (Fig. 1j). Of note, our approach captured the previously described periodic region in HNRNPA1 (Extended Data Fig. 3a)25.Regions of significant periodicity were identified in 2,202 human proteins. 396 / 2,202 of the periodic regions overlapped intrinsically disordered regions (IDRs) annotated by Metapredict (Extended Data Fig. 3b-c). The proteins containing regions of significant periodicity were enriched for prion-like protein, and were not enriched for TFs (Fig. 1k, Extended Data Fig. 3d- e). Only 134 / 1542 transcription factors contain a region of significant periodicity, and only 63 of these regions were in the I DR (Fig. 1k). Furthermore, the average QAroscore of IDRs in TFs was significantly higher than of prion-like domains (P<10-4, One-way ANOVA) (Fig. 11-m). These results demonstrate that transcription factor IDRs have lower periodicity than prion-like domains and suggest that the periodicity of transcription factor IDRs may be submaximal.2.3 Increasing aromatic dispersion in TF IDRs enhances transactivationIf TF IDRs have submaximal aromatic dispersion, one would expect that increasing their aromatic dispersion enhances activity. We tested this idea using the HOXD4 I DR as a proof-of- concept. The HOXD4 I DR contains 11 periodically arranged aromatic residues in its N-termi- nus, but the overall periodicity of the entire I DR is limited (Fig. 2a). To test the effect of increasing aromatic periodicity of the HOXD4 I DR, we substituted seven non-aromatic residues with aromatic residues in regions of spacer lengths >15 amino acids within the I DR (“AroPLUS”) (Fig. 2a). Purified GFP-tagged AroPLUS I DR protein formed droplets at a lower Csat than the wild type HOXD4 I DR in vitro (Fig. 2b-c) and had a 3-fold higher activity in the GAL4-DBD transactivation assay (P<10-3, t-test) (Fig. 2a), which was specific to adding aromatic residues in positions that increase periodicity (Fig. 2a-c). We also generated a HOXD4 IDR mutant in which the aromatic residues in the native sequence were re-shuffled so that they were separated by a uniform spacer length (“AroPERFECT”) (Fig. 2a). Purified GFP-tagged AroPER- FECT IDR formed liquid-like droplets at a similar Csat to the wild type HOXD4 IDR in vitro (Fig. 2b-c). Fluorescence recovery after photobleaching (FRAP) however revealed an increase in the recovery of fluorescence in AroPERFECT droplets after photobleaching (Fig. 2d, Extended Data Fig. 4a), suggesting enhanced liquid-like features of IDR droplets. Moreover, the AroPERFECT IDR had a 6-fold higher activity in the GAL4-DBD transactivation assay compared to the wild type IDR (P<10-3, t-test) (Fig. 2a, Extended Data Fig. 4b). These results suggest that increasing aromatic dispersion in the HOXD4 IDR enhances its activity.Further mutagenesis of the HOXD4 IDR revealed that increasing aromatic dispersion enhances transactivation within the confines of additional sequence features, but independent ofpredicted structural elements. The HOXD4 I DR is predicted to contain a minimal activation domain (Fig. 2e). Nevertheless, assaying the activity of 40 amino acid tiles in the GAL4-DBD luciferase system revealed that the activity of this element was lower in the AroPERFECT I DR (Fig. 2e). Furthermore, the elevated activity of the AroPERFECT I DR could not be explained by the creation of additional minimal activation domains (Fig. 2e) and no correlation with known short linear motifs (SLiMs)13was apparent (Extended Data Fig. 4c-d, Table S3). Shifting the uniformly spaced aromatic residues two positions towards the N-terminus led to moderately elevated activity, shifting them one position did not (Extended Data Fig. 4c-d), and the degree of enhancement correlated with the number of small, inert residues adjacent to the aromatic residues (Extended Data Fig. 4e), consistent with previous studies on prion-like sequences 26,28-30 Fina||ywecomplemented the I DR portion downstream of the minimal activation domain with a short periodic portion of the FUS I DR, which also led to an enhancement of activity (Fig. 2f). Taken together, these results suggest that increased aromatic dispersion of the HOXD4 I DR enhances its transcriptional activity. Consistent results were observed when increasing aromatic dispersion of the orthologous HOXC4 I DR (Extended Data Fig. 4f-k). For both IDRs, dispersing the native aromatic residues in a uniform pattern enhanced liquid-like features of in vitro I DR droplets.Increasing aromatic dispersion enhanced the transcriptional activity of the multiple other TF IDRs tested. The TFs included OCT4, a master regulator of pluripotent cells; PDX1 , as master regulator of pancreatic beta cells; and FOXA3, a master regulator of hepatocytes (Extended Data Fig. 5a-c). As a corollary, reducing the aromatic dispersion of the periodic EGR1 I DR reduced its activity (Extended Data Fig. 5e-f). As with the HOXD4 I DR, the nature of spacer residues appeared to pose constraints for the effect of optimizing aromatic dispersion, as increasing aromatic dispersion of the HOXB1 I DR did not enhance its (already strong) activity (Extended Data Fig. 5g). Supporting this model, in a synthetic neutral I DR backbone aromatic dispersion correlated with activity, but in a negatively charged backbone it did not (Fig. 2g). These results suggest that optimizing aromatic dispersion can enhance the activity of TFs, but not without limitations that remain to be further investigated.2.4 Evidence for gain-of-function of periodic HOXD4 mutants in vivoTo investigate the impact of the periodic HOXD4 mutants in vivo, we generated cell lines in which mEGFP-tagged full-length HOXD4 wild type, AroPERFECT, and AroPLUS variants were knocked-in into the endogenous locus using CRISPR / Cas9 in HAP1 cells (Extended Data Fig. 6a-d). To our surprise, knocking-in the AroPERFECT and AroPLUS HOXD4 mutants altered the morphology of the cell colonies suggesting a gain-of-function effect (Fig. 3a). Thewild type HOXD4-mEGFP protein was modestly enriched in the nucleus, while the AroPER- FECT and AroPLUS HOXD4 proteins localized to the nucleus and formed intense clusters (Fig. 3a). To probe nuclear HOXD4 clusters in cells that express the three variants at comparable levels, we integrated Dox-inducible mEGFP-tagged alleles using a PiggyBac transposon. The average granularity (i.e., normalized standard deviation of the fluorescence signal) in cells expressing AroPERFECT and especially AroPLUS HOXD4 transgenes was higher compared to cells expressing wild type HOXD4 (Fig. 3b-c). These results suggest that increased periodicity of aromatic residues in the HOXD4 I DR results in a partial gain of function in vivo.To gain insights into the genes deregulated by the AroPERFECT and AroPLUS HOXD4 proteins, we performed RNA-Seq on the HAP1 cell lines that encode stably-integrated HOXD4 variants at the endogenous locus. On a genome-wide level, principal component analysis of the -16,000 quantified transcripts revealed that the expression profile of AroPERFECT and AroPLUS HOXD4-expressing cells were highly similar to each other, and distinct from that of the wild type and HOXD4 knockout cells (Fig. 3d). We then annotated 1 ,133 HOXD4-target genes based on differential expression between the parental HAP1 and HOXD4 knockout cells. 76% of HOXD4-target genes were deregulated in the AroPERFECT and AroPLUS cells in the same direction as in knockout cells, (consistent with loss of heterodimerization with PBX- factors31) (Fig. 3e, Extended Data Fig. 6e). However, we also identified 396 genes upregu- lated in the AroPERFECT- and AroPLUS-expressing cells that were downregulated in the HOXD4 knockout cells, consistent with enhanced activity. One of the genes was HOXD4 itself, consistent with previous studies that HOX TFs, including HOXD4, frequently autoregulate their own genes32-34. The elevated levels of HOXD4 and another protein, ARHGAP4, were validated with Western blot (Fig. 3f). Furthermore, we also identified 43 genes upregulated in the AroPERFECT-expressing cells, and 64 genes upregulated in the AroPLUS-expressing cells, that were not HOXD4-targets, e.g., IFI16 (Fig. 3f, Extended Data Fig. 6e-f). The morphology and expression phenotypes were confirmed in the PiggyBac cells expressing comparable levels of wild type and periodic HOXD4 transgenes (Extended Data Fig. 6f-i). These results indicate that increased aromatic dispersion in the HOXD4 I DR is associated with enhanced activity and altered gene specificity.To further probe the link between the effects of aromatic dispersion on transcriptional activity and condensates, we measured RNAPII CTD recruitment into HOXD4 I DR condensates using a cell-based condensate system35. In this system, wild type or AroPERFECT HOXD4 IDR was tethered to a LacO array in U2OS cells expressing an ectopic RNAPII CTD-YFP fusion protein (Fig. 3g). RNAPII CTD was mildly enriched in the tethered HOXD4 wild type IDR condensates, and its enrichment was significantly higher in the AroPERFECT IDR condensates (Fig. 3g-h).These results suggest that the enhanced activity and altered gene specificity of periodic HOXD4 I DR is associated with reduced heterodimerization and enhanced RNAPII interaction.2.5 Optimizing aromatic dispersion in C / EBPa enhances transactivationTranscription factors play essential roles in cell type specification, and small sets of transcription factors can reprogram cell identity45. Therefore, we tested the impact of increasing the periodicity of aromatic residues on the ability of transcription factors to reprogram cells.C / EBPa is a master regulator of myeloid cell differentiation and forced expression of C / EBPa reprograms B cells into macrophages36. C / EBPa contains a C-terminal bZIP DNA binding domain, and an N-terminal I DR devoid of periodicity of aromatic residues (Fig. 4a). Purified recombinant mEGFP-tagged C / EBPa IDRs formed in vitro droplets with liquid-like features (Fig. 4b), and had transactivation capacity in the GAL4-DBD luciferase system (Fig. 4a). IDR droplet formation and transactivation was dependent on the presence of aromatic residues (Fig. 4a-b, Extended Data Fig. 7a). To test the impact of increased aromatic dispersion, we generated an IDR in which the aromatic residues were dispersed with perfectly uniform spacing (“AroPERFECT IS15”). Increased dispersion did not affect the Csat for droplet formation (Fig. 4a-b, Extended Data Fig. 7a), but enhanced recovery after photobleaching in droplets (Fig. 4c), and enhanced transactivation 2-fold in the GAL4-DBD -luciferase system compared to the wild type IDR (P<10-4, t-test) (Fig. 4a, Extended Data Fig. 7b). Similar to results with HOXD4, RNAPII CTD was more enriched in AroPERFECT IS15 condensates compared to wild type IDR condensates tethered onto the LacO array (Fig. 4d-e). In vitro, increasing both the number of aromatic residues and their dispersion (“AroPERFECT IS10”) resulted in a decrease in FRAP (Fig. 4c), and decreased transactivation in the GAL4-DBD -luciferase system compared to the wild type IDR (P<10-3, t-test) (Fig. 4a). These results suggest that increased aromatic dispersion enhances transactivation of the C / EBPa IDR, but the increase of aromaticity inhibits it, consistent with models that proteins evolve at the edge of surface hydrophobicity37.Further mutagenesis of the C / EBPa IDR revealed that increasing aromatic dispersion enhances transactivation within the confines of additional sequence features. The C / EBPa IDR encodes a known minimal activation domain38. The activity of this element however was lower in the AroPERFECT IS15 IDR, and the elevated activity of the AroPERFECT IS15 IDR was not caused by the creation of additional minimal activation domains (Fig. 4f). Second, when we increased the aromatic dispersion only of the portion of the C / EBPa IDR downstream of the activation domain [“WT (N)-IS15”], the IDR had 5-fold elevated activity over the wild type, and 3-fold activity over the N-terminal portion including the minimal activation domain (Fig.4g). Third, replacing the downstream IDR portion with the periodic FUS N-terminal IDR enhanced transcriptional activity over the wild type IDR (Fig. 4g), and the enhancement was also evident when only a portion of FUS IDR was used that contains the exact number of aromatic residues (i.e. eight) as the C / EBPa IDR that was replaced [“WT(N)-FUSNxs”] (Fig. 4g). Fourth, shifting the aromatic pattern of AroPERFECT IS15 IDR one amino acid towards the C-terminus resulted in higher transactivation compared to the wild type, and shifting by two positions did not (Extended Data Fig. 7c), and the magnitude of change correlated with the proportion of small inert residues adjacent to the aromatic residues (Extended Data Fig. 4e). Aromatic dispersion therefore appears to enhance transactivation independent of the known C / EBPa activation domain, and within confines of the spacer residues in the IDR.2.6 Optimizing aromatic dispersion in C / EBPa enhances macrophage reprogrammingTo test whether sequence-optimization impacts the activity of C / EBPa in cells, we measured cellular reprogramming of stably transduced C / EBPa variants in a leukemic human B cell line (RCH rtTA cells). In this system, induction of C / EBPa by doxycycline reprograms the B cells into terminally differentiated macrophages in a stepwise manner while arresting the cell cycle39’40. The B cell to macrophage conversion was monitored with FACS analysis of the B cell marker CD19, and the macrophage marker Mac1 (a.k.a. CD11 b encoded by the gene ITGAM) (Fig. 5a-b, Extended Data Fig. 7d)39’40. As expected, C / EBPa expression led to a gradual increase of the proportion of Mac1+CD19' macrophages among the GFP+cell population over seven days (Fig. 5c, Extended data Fig. 7d-e). Expression of the AroPERFECT IS15 C / EBPa mutant increased both the speed of appearance and proportion of Mac1+cells among the GFP+population (Fig. 5c, Extended data Fig. 7d-e). These results suggest that increased aromatic dispersion in the C / EBPa IDR enhances macrophage reprogramming.To gain insights into the transcriptional programs and differentiation trajectories driven by the C / EBPa proteins, we performed single-cell RNA-Sequencing of the cultures expressing wild type and AroPERFECT IS15 C / EBPa variants after seven days of transgene induction. As negative control, the culture expressing the transcriptionally inert AroPERFECT IS10 C / EBPa variant was also included (Fig. 5a). Cross-referencing the clusters on the combined scRNA cell state map of the three cultures with marker genes of known cell populations identified terminally differentiated macrophages, macrophage precursors, and various B cell subpopulations in our data (Fig. 5d, Extended Data Fig. 8a-g). Consistent with the FACS analysis, the proportion of late macrophages was higher among the GFP+cells in the “AroPERFECT IS15”- transduced population (Fig. 5e), indicating enhanced reprogramming capacity. A comparative analysis of the transcriptomes of WT and “AroPERFECT IS15” C / EBPa -expressing late mac-rophages revealed largely similar expression profiles; however, the “AroPERFECT IS15” macrophages expressed a small set of 31 genes not detected in the WT C / EBPa-expressing macrophages (Extended Data Fig. 8h-i, Table S5), suggesting slightly altered gene specificity.2.7 Optimizing aromatic dispersion in C / EBPa leads to stronger and more promiscuous genomic bindingTo dissect the molecular basis of enhanced reprogramming and differences in the expression profiles of WT and AroPERFECT IS15 -transduced cells, we performed ChlP-Seq of the C / EBPa-GFP proteins with an anti-GFP antibody after 24 and 48 hours of transgene induction. The majority of sites bound by WT C / EBPa were also bound by AroPERFECT IS15 C / EBPa, but the read densities at the bound sites were consistently higher in the AroPERFECT IS15 samples (Fig. 5f, Extended Data Fig. 9a-e). Overall, -100 times more differentially bound peaks had higher read densities in AroPERFECT IS15 than the other way around (Extended Data Fig. 9b).Differential genomic binding of AroPERFECT IS15 C / EBPa was associated with differences in motif composition at the binding sites. For these analyses, we used -28,000 ChlP-Seq peaks identified as “shared” by both WT and AroPERFECT IS15 C / EBPa, and -60,000 ChlP-Seq peaks uniquely bound by AroPERFECT IS15 C / EBPa at least at one time point (Fig. 5f). Cross-referencing the peaks with published C / EBPa ChlP-Seq datasets revealed that -50,000 of the sites were reported before as binding sites of WT C / EBPa (“peaks unique to IS15, reported before” in Fig. 5f), and -10,000 were specific to our AroPERFECT IS15 C / EBPa data (“peaks specific to IS15” in Fig. 5f). The shared binding peaks and peaks unique to IS15 reported before were highly enriched for the same canonical C / EBPa motif (Fig. 5g). However, the peaks specific to IS15 were less enriched for the C / EBPa motif and were instead more enriched for other bZIP TF motifs including C / EBp and NFIL3 (Fig. 5g).The impact of differential binding on gene expression was confirmed with multiple approaches. For example, IS15-specific binding at several loci was associated with detectable IS15-specific expression of the gene in the scRNA-Seq data in B-cell or in macrophage clusters (Fig. 5h-k, Extended Data Fig. 9f-i). Furthermore, cloning of IS15-specific peaks in a luciferase reporter revealed elevated activity when co-transfected with an AroPERFECT IS15 C / EBPa vector compared to wild type (Fig. 5I). Finally, differential expression was confirmed with FACS analysis of the products of two macrophage-restricted genes. These included CD66, the product of the CEACAM genes (Extended Data Fig. 9j-l), and FCGR2A (Extended Data Fig. 9m-o). Taken together these results suggest that the enhanced reprogramming capacity of AroPERFECT IS15 C / EBPa is associated with stronger and more promiscuous genomic binding.2.8 Optimizing aromatic dispersion in NGN2 enhances neural differentiationAs a second proof-of-concept, we tested the impact of optimizing aromatic dispersion on the reprogramming capacity of Neurogenin-2 (NGN2). NGN2 is a neurogenic transcription factor, whose forced expression reprograms pluripotent cells into induced neurons41. NGN2 contains a central basic helix-loop-helix DNA-binding domain, an N-terminal I DR that contains one aromatic residue, and a C-terminal I DR that contains five aromatic residues devoid of appreciable periodicity (Fig. 6a).The wild type recombinant, GFP-tagged NGN2 C-IDR formed liquid-like droplets in a concentration-dependent manner (Extended Data Fig. 10a-c). Droplet formation was reduced when the aromatic residues were substituted with alanines (“AroLITE C-IDR”) (Extended Data Fig. 10a-c). Similar to results with the IDRs of C / EBPa, HOXD4 and HOXC4, a mutant NGN2 C- IDR in which the five aromatic residues were rearranged to have uniform spacing (“AroPER- FECT C-IDR”) formed droplets similar to the wild type I DR in vitro, and had a small, statistically non-significant difference in FRAP (Fig. 6b). None of the IDRs had measurable activity in the GAL4-DBD-luciferase system (Extended Data Fig. 10a), consistent with a report that a minimal activation domain is located within the NGN2 DNA-binding domain42.To assay the reprogramming capacity of NGN2 mutants, Dox-inducible FLAG-tagged NGN2 transgenes were stably integrated in ZIP13K2 human induced pluripotent stem cells (iPSCs) using a PiggyBac transposon (Fig. 6c, Extended Data Fig. 10d-e). The transposon also encoded GFP separated by a T2A sequence from NGN2. After 24 hours of Dox-induction, GFP- positive cells were FACS-sorted and replated. After 48 hours, media was exchanged with media supporting neural differentiation, and cells were eventually characterized by staining nuclei and tubulin (Fig. 6c). Twice as many sorted cells expressing the “AroPERFECT” NGN2 mutant survived, and half as many cells expressing the “AroLITE” NGN2 mutant survived compared to the wild type NGN2-expressing cells after 5 days of transgene induction (P<0.05, t-test) (Fig. 6d-e). Consistently, the density of cell projections was significantly higher in the “AroPERFECT” NGN2-expressing cultures compared to wild type NGN2-expressing cultures after 5 days of transgene induction (P<0.05, t-test) (Fig. 6d, 6f). These results indicate that the increased aromatic dispersion in the C-terminal I DR of NGN2 enhances its capacity to reprogram iPSCs into neuron-like cells.To investigate the molecular basis of enhanced reprogramming by the NGN2 AroPERFECT mutant, we performed RNA-Seq after five days and NGN2 ChlP-Seq 24 and 48 hours after transgene induction. The global RNA-Seq profiles of cultures expressing wild type, AroLITEand AroPERFECT NGN2 proteins were largely similar, and included NGN2-target genes annotated based on previous studies (Fig. 6g-h, Extended Data Fig. 10f-g), consistent with media conditions promoting survival of neurons but not iPSCs after the media switch on day 2 (Fig. 6b). The ChlP-Seq data revealed that most sites bound by wild type NGN2 were also bound by the AroLITE and AroPERFECT protein (Fig. 6i, Extended Data Fig. 10h). However, the ChlP-Seq read densities at the binding sites were consistently lower in the AroLITE-ex- pressing cells, and moderately higher in AroPERFECT-expressing cells at 24h (Fig. 6i-j, Extended Data Fig. 10i). The bHLH TF motif composition of the binding peaks was largely similar (Extended Data Fig. 10j). Consistent with these results, measuring nascent transcription genome-wide after short-term NGN2-induction revealed elevated transcription of NGN2-target genes in AroPERFECT-expressing cells (Fig. 6k, Extended Data Fig. 10k). These results suggest that optimizing aromatic dispersion in the NGN2 C-terminal I DR enhances neural reprogramming and alters genomic binding. Changes in in vitro droplet features and genomic binding for optimized NGN2 were not as pronounced as for optimized C / EBPa, which may be explained by NGN2 containing two IDRs and the optimized C-IDR containing only six aromatic residues as opposed to the C / EBPa I DR that contains 16.2.9 Optimizing aromatic dispersion in MYOD1 enhances myotube differentiationFinally, we tested the impact of increased periodicity of aromatic residues on the reprogramming capacity of MYOD1. MYOD1 is a myogenic transcription factor, whose forced expression reprograms cells into muscle-like cells43. MYOD1 contains a central basic helix-loop-helix DNA-binding domain, an N-terminal I DR that contains eight aromatic residues, and a C-terminal I DR that contains nine aromatic residues, devoid of appreciable periodicity (Fig. 7a). Both the N-terminal and C-terminal MYOD1 IDRs had transactivation capacity in the GAL4-DBD - luciferase system in myoblasts (Fig. 7a). Mutation of aromatic residues into alanines in both IDRs abolished transactivation (Fig. 7a). Increased aromatic dispersion of aromatic residues abolished transactivation of the N-terminal IDR that contains a minimal activation domain, but increased transactivation of the C-terminal IDR (Fig. 7a, Extended Data Fig. 11a). Consistent with results on HOXD4, C / EBPa, and OCT4, the enhanced activity of the AroPERFECT C-IDR did not appear to be caused by creating minimal activation domains (Extended Data Fig. 11 b).To assay the reprogramming capacity of MYOD1 mutants, Dox-inducible MYOD1 transgenes were stably integrated into C2C12 murine myoblasts using a PiggyBac transposon (Fig. 7b) The transposon also encoded GFP separated by a T2A sequence from MYOD1 (Fig. 7b). In this system, forced expression of MYOD1 differentiates myoblasts into multi-nucleated myotubes within a few days44. Cell fusion was quantified as the percentage of DAPI-stained nuclei in multi-nucleated cells visualized using the mEGFP fluorescence signal as cytoplasmicmarker45. -50% of nuclei expressing wild type MY0D1 were found in fused cells after 3 days of transgene induction (Fig. 7c-d, Extended Data Fig. 11c). Mutating the aromatic residues into alanines in both IDRs (“AroLITE”) prevented fusion, while mutating the aromatic residues in the C-terminal IDR (“AroLITE C”) had a negligible effect (Fig. 7c-d). Expression of the MYOD1 mutant with enhanced periodicity in its C-terminal IDR (“AroPERFECT C”) led to a significant increase in fusion after 3 days (P<0.05, t-test) (Fig. 7c-d). These results suggest that increased periodicity of aromatic residues in the C-terminal IDR of MYOD1 enhances myotube differentiation.RNA-Seq analysis of differentiating cells expressing various MYOD1 proteins revealed signatures consistent with the observed morphological differences. Principal component analysis of the RNA-data showed that the global expression profiles of AroLITE-expressing cells were similar to that of the parental myoblasts (Fig. 7e, Extended Data Fig. 11 d). The expression profile of AroPERFECT C-expressing cells was largely similar to cell expressing wild type MYOD1, but included 290 differentially expressed genes, 197 of which were MYOD1 targets, and were enriched for genes implicated in cell adhesion (Fig. 7e-f, Extended Data Fig. 11 d- g). These results suggest that morphologies are associated with differences in gene expression profiles in differentiating myotubes expressing various MYOD1 proteins.2.9 Optimizing aromatic dispersion in BRACHYURY (TBXT), MSGN1, NKX2-5 and CDX2 Deletion of aromatic residues in the IDRs of the murine transcription factors Brachyury (TBXT, also called T), MSGN2, NKX2-5 and CDX2 reduces transcriptional activity. Enhancing the periodicity of aromatic residues in the IDRs of the murine transcription factors Brachyury (TBXT), MSGN2, NKX2-5 and CDX2 increases transcriptional activity. It is expected that these results also apply to the human orthologues (Fig. 8).3. DiscussionThe results presented here support a model that human transcription factors are suboptimal for sequence features associated with transcriptional activity. We present evidence that suboptimality in several transcription factors is encoded as submaximal dispersion of aromatic residues in their intrinsically disordered regions (IDRs). In several cellular reprogramming systems, increasing aromatic dispersion enhanced the activity and compromised gene specificity of key TFs. T ogether with previous work that enhancer DNA sequences are suboptimal for transcription factor binding14’16, the results suggest an important evolutionary trade-off between activity and specificity at multiple levels in eukaryotic transcriptional control.The results provide insights into how human transcription factors work. Some TFs encode short linear motifs that can fold into secondary structures and mediate specific interactions with effector proteins46. Such sequences are typically identified as ‘minimal’ activation domains that are sufficient to activate transcription of a reporter gene7-10.13 4748. Qur results suggest that some TF IDRs encode periodically arranged aromatic residues that contribute to activity via multivalent, weak interactions with other disordered protein regions. This mode of activity may be distinct from and complementary to the transcriptional activity conferred by minimal activation domains. Consistent with this proposal, hydrogels of periodic low complexity domains can bind RNAPII CTD that itself is highly periodic49, and we found that periodic TF IDRs recruit RNAPII CTD more efficiently than wild type TF IDRs in the cell-based condensate tethering system. This model may help explain why ‘minimal’ activation domains are typically embedded in large disordered sequences6'10, and why some TF IDRs can be substituted with the periodic FUS prion-like domain50 51. A prediction of this model is thus that important regulatory information may be encoded in sequences with weak or no activity.Increasing aromatic dispersion enhanced the transcriptional activity of multiple, but not all TF IDRs, suggesting that the intervening spacer sequences have significant contributions to the functions of aromatic residues26’28'30. Previous work has noted that charged residues may act to prevent intramolecular interactions between hydrophobic residues11, and we found that increasing hydrophobicity by adding aromatic residues into TF IDRs can inhibit the dynamics of their self-association and transcriptional activity. Therefore, aromatic dispersion may enhance the propensity of intermolecular interactions with effector partners52. Consistent with this notion, the dynamics of self-association was significantly enhanced when aromatic dispersion was increased in multiple TF IDRs in vitro in this study.Transcription factor-mediated differentiation and reprogramming are generally stochastic and inefficient, and the inefficiency is thought to be explained by chromatin barriers or lack of TF effector partners4553-58. Our results suggest that an additional impediment to directed differentiation and reprogramming may be the suboptimal activity of native transcription factors, and that reprogramming efficiency may be improved by enhancing a prion-like phase separation “grammar” in native TFs. In summary, we propose that altering phase separation capacity may be a universal strategy to rationally control the activity of any biomolecule.Protein sequencesAromatic residues are in bold face, short linear motifs (L / F / Y / W XX L / F / Y / W) are underlined. The sequences are the translated protein sequences used in the in vitro droplet formation and transactivation assays.I DR sequences designated as AROPERFECT or AROPLUS have increased numbers and / or periodicity of aromatic amino acid residues in their IDRs and show increased transcriptional activity compared to the respective wild-type sequences. I DR sequences designated as ARGUTE have decreased numbers and / or periodicity of aromatic amino acid residues in their IDRs and show reduced transcriptional activity compared to the respective wild-type sequences.The AroSCRAMBLED and AroPATCHY also have altered activities. These sequences provide an additional experimental approach to support the importance of aromatic dispersion for transcription factor activity.The FLISN sequences are used in a series of experiments to further validate the importance of aromatic dispersion for transcription factor activity. The FUS sequence originated from an RNA binding protein, and in this sequence the aromatic residues are known to be uniformly distributed. These sequences have activities similar to transcription factor IDRs optimized for aromatic dispersion.HOXD4 IDR Wild type (SEQ ID NO: 1)MVMSSYMVNSKYVDPKFPPCEEYLQGGYLGEQGADYYGGGAQGAD-FQPPGLYPRPDFGEQPFGGSGPGPGSALPARGHGQEPGGPGGHYAAPGEPCPAPPAPPPAPLPGARAYSQSDPKQPPSGTALKQPAWYPWMKKVHOXD4 IDR AroLITE A (SEQ ID NO: 2)MVMSSAMVNSKAVDPKAPPCEEALQG-GALGEQGADAAGGGAQGADAQPPGLAPRPDAGEQPAGGSGPGPGSALPARGHGQEPG GPGGHAAAPGEPCPAPPAPPPAPLPGARAASQSDPKQPPSGTALKQPAWAPAMKKVHOXD4 IDR AroLITE S (SEQ ID NO: 3)MVMSSSMVNSKSVDPKSPPCEESLQGGSLGEQGADSSGGGAQGADSQPPGL-SPRPDSGEQPSGGSGPGPGSALPARGHGQEPGGPGGHSAAPGEPCPAPPAPPPAPLPG ARASSQSDPKQPPSGTALKQPAVVSPSM KKVHOXD4 I DR AroLITE G (SEQ ID NO: 4)MVMSSGMVNSKGVDPKGPPCEEGLQGGGLGEQGADGGGGGAQGADG-QPPGLGPRPDGGEQPGGGSGPGPGSALPARGHGQEPGGPGGHGAAPGEPCPAPPAPPPAPLPGARAGSQSDPKQPPSGTALKQPAWGPGMKKVHOXD4 I DR AroPLUS (SEQ ID NO: 5)MVMSSYMVNSKYVDPKFPPCEEYLQGGYLGEQGADYYGGGAQGAD-FQPPGLYPRPDFGEQPFGGSGPGYGSALPARYHGQEPYGPGGHYAAPGEPCPYPPAPPPYPLPGARAYSQSDPKYPPSGTAYKQPAVVYPWMKKVHOXD4 I DR AroPLUS LITE (SEQ ID NO: 6)MVMSSYMVNSKYVDPKFPPCEEYLQGGYLGEQGADYYGGGAQGAD-FQPPGLYPRPDFGEQPFGGSGPGAGSALPARAHGQEPAGPGGHYAAPGEPCPAPPAPPPAPLPGARAYSQSDPKAPPSGTAAKQPAWYPWMKKVHOXD4 I DR AroPLUS patched (SEQ ID NO: 7)MVMSSYMVNSKYVDPKFYPPCEEYLQGGYYLGEQGADYYGGGAQGAD-FQPPGLYYPRPDFGEQPFYGGSGPGGSALPARHGQEPGPGGHYYAAPGEPCPPPAPPPPLPGARAYYSQSDPKPPSGTAKQPAVVYYPWMKKVHOXD4 I DR AroPLUS LITE patched (SEQ ID NO: 8)MVMSSYMVNSKYVDPKFAPPCEEYLQGGYALGEQGADYYGGGAQGAD-FQPPGLYAPRPDFGEQPFAGGSGPGGSALPARHGQEPGPGGHAYAAPGEPCPPPAPPPPLPGARAYASQSDPKPPSGTAKQPAVVAYPWMKKVHOXD4 I DR AroPERFECT (SEQ ID NO: 9)MVMSSMVNYSKVDPKPPFCEELQGGLYGEQGADGGYG-AQGADQPFPGLPRPDGFEQPGGSGPFGPGSALPAYRGHGQEPGYGPGGHAAPYGEPCPAPPYAPPPAPLPYGARASQSDYPKQPPSGTYALKQPAVVWPMKKVHOXD4 IDR AroPERFECT -1 (SEQ ID NO: 10)MVMSSMVYNSKVDPKPFPCEELQGGYLGEQGADGYG-GAQGADQFPPGLPRPDFGEQPGGSGFPGPGSALPYARGHGQEPYGGPGGHAAYPGEPCPAPYPAPPPAPLYPGARASQSYDPKQPPSGYTALKQPAVWVPMKKVHOXD4 IDR AroPERFECT -2 (SEQ ID NO: 11)MVMSSMYVNSKVDPKFPPCEELQGYGLGEQGADYGGGAQGAD-FQPPGLPRPFDGEQPGGSFGPGPGSALYPARGHGQEYPGGPGGHAYAPGEPCPAYPPAPPPAPYLPGARASQYSDPKQPPSYGTALKQPAWWPMKKVH0XD4 I DR Wild type YPWM(-) (SEQ ID NO: 12)MVMSSYMVNSKYVDPKFPPCEEYLQGGYLGEQGADYYGGGAQGAD-FQPPGLYPRPDFGEQPFGGSGPGPGSALPARGHGQEPGGPGGHYAAPGEPCPAPPAPPPAPLPGARAYSQSDPKQPPSGTALKQPAWAPAMKKVHOXD4 I DR AroLITE YPWM(+) (SEQ ID NO: 13)MVMSSAMVNSKAVDPKAPPCEEALQG-GALGEQGADAAGGGAQGADAQPPGLAPRPDAGEQPAGGSGPGPGSALPARGHGQEPGGPGGHAAAPGEPCPAPPAPPPAPLPGARAASQSDPKQPPSGTALKQPAWYPWMKKVHOXD4 I DR AroPERFECT YPWM (+) (SEQ ID NO: 14)MVMSSMVNYSKVDPKPPFCEELQGGLYGEQGADGGYG-AQGADQPFPGLPRPDGFEQPGGSGPFGPGSALPAYRGHGQEPGYGPGGHAAPYGEPCPAPPYAPPPAPLPYGARASQSDYPKQPPSGTALKQPAVVYPWMKKVHOXD4 wild type (N) (SEQ ID NO: 15)MVMSSMVNSKVDPKFPPCEEYLQGGYLGEQGADYYGGGAQGAD-FQPPGLYPRPDFGEQPFGGSGPGPGSALHOXC4 I DR Wild type (SEQ ID NO: 16)MIMSSYLMDSNYIDPKFPPCEEYSQNSYIPEHSPEYYGRTRESGFQHHHQELYPPPP-PRPSYPERQYSCTSLQGPGNSRGHGPAQAGHHHPEKSQSLCEPAPLSGASASPSPAPPACSQPAPDHPSSAASKQPIVYPWMKHOXC4 I DR AroLITE S (SEQ ID NO: 17)MIMSSALMDSNAIDPKAPPCEEASQNSAIPEHSPEAAGRTRESGAQHHHQELAPPPP-PRPSA-PERQASCTSLQGPGNSRGHGPAQAGHHHPEKSQSLCEPAPLSGASASPSPAPPACSQPAPDHPSSAASKQPI VAPAM KHOXC4 I DR AroPERFECT (SEQ ID NO: 18)MIMSSLMYDSNIDPKPPCFEESQNSIPEHYSPEGRTRESGFQHHHQELPPPYPPRP-SPERQSYCTSLQGPGNSYRGHGPAQAGHYHHPEKSQSLCYEPAPLSGASAYSPSPAPPACSYQPAPDHPSSAYASKQPIVPMWKH0XB1 I DR Wild type (SEQ ID NO: 19)MDYNRMNSFLEYPLCNRGPSAYSAHSAPTSFPPSSAQAVDSYASEGRYGG-GLSSPAFQQNSGYPAQQPPSTLGVPFPSSAPSGYAPAACSPSYGPSQYYPLGQSEGDGGYFHPSSYGAQLGGLSDGYGAGGAGPGPYPPQHPPYGNEQTASFAPAYA-DLLSEDKETPCPSEPNTPTARTFDWMKVKRNPPKTAKVSEPGLHOXB1 I DR AroLITE A (SEQ ID NO: 20)MDANRMNSALEAPLCNRGPSAASAHSAPTSAPPSSAQAVDSAASEG-RAGGGLSSPAAQQNSGAPAQQPPSTLGVPAPSSAPSGAAPAACSPSAGPSQAAPLGQSEGDGGAAHPSSAGAQLGGLSDGAGAGGAGPGPAPPQHP-PAGNEQTASAAPAAADLLSEDKETPCPSEPNTPTARTADAMKVKRNPPKTAKVSEPGLHOXB1 IDR AroPERFECT (SEQ ID NO: 21)MDYNRMNSLEYPLCNRGPYSASAHSAFPTSPPSSFAQAVDSAYSEGRGG-GYLSSPAQQFNSGPAQQYPPSTLGVFPPSSAPSYGAPAACSYPSGPSQPYLGQSEGDYGG H PSSGFAQLGGLSYDGGAGGAYG PG PPPQYH PPG N EQYTAS-APAAFDLLSEDKYETPCPSEYPNTPTARFTDMKVKRWNPPKTAKVSEPGLNANOG IDR Wild type (SEQ ID NO: 22)KQVKTWFQNQRMKSKRWQKNNWPKNSNGVTQKASAPTYPSLYSSYHQGCLVNPTGN-LPMWSNQTWNNSTWSNQTQNIQSWSNHSWNTQTWCTQSWNNQAWNSPFYNCGEESLQSCMQFQPNSPASDLEAALNANOG IDR AroLITE A (SEQ ID NO: 23)KQVKTAAQNQRMKSKRAQKNNAPKNSNGVTQKASAPTAPSLASSAHQGCLVNPTGNLP-MASNQTANNSTASNQTQNIQSASNHSANTQTACTQSANNQAANSPAANCGEESLQSCMQAQPNSPASDLEAALEGR1 IDR Wild type (SEQ ID NO: 24)LRQKDKKADKSVVASSATSSLSSYPSPVATSYPSPVTTSYPSPATTSYP-SPVPTSFSSPGSSTYPSPVHSGFPSPSVATTYSSVPPAFPAQVSSFPSSAVTNSFSASTGLSDMTATFSPRTIEICEGR1 I DR AroLITE A (SEQ ID NO: 25)LRQKDKKADKSVVASSATSSLSSAPSPVATSAPSPVTTSAPSPATTSAPSPVPTSASS-PGSSTAPSPVHSGAPSPSVATTASSVPPAAPAQVSSAPSSAVTNSASASTGLSDMTATASP RTIEICEGR1 I DR AroSCRAMBLED (SEQ ID NO: 26)LRQKDKKADKSVVASSATSSLSSYPSPVAFTYSPSPVTTSPSPYATYTSPSPVPTSSS-FPGSSYFTPSPVHSGPYSPSVATTSSVPPAPAQVSSPSSAVFTNSFSASTGFLSDMTATSP RTIEICEGR1 I DR AroPATCHY3 (SEQ ID NO: 27)LRQKDKKADKSVVASSATSSLSSYYYYP-SPVATSPSPVTTSPSPATTSPSPVPTSSSPGSSTPSPVHSGPSPSVATTYFFYSSVPPAPAQVSSPSSAVTNSFFFFSASTGLSDMTATSPRTIEICEGR1 I DR AroPATCHYl (SEQ ID NO: 28)LRQKDKKADKSVVAS-SATSSLSSPSPVATSPSPVTTSPSPATTSPSPVPTSSSPGSSTYYYYFYYFFFFFPSPVHSGPSPSVATTSSVPPAPAQVSSPSSAVTNSSASTGLSDMTATSPRTIEICNFAT5 I DR Wild type (SEQ ID NO: 29)TM VKKEISSPARPCSFEEAM KAM KTTGCN LDKVN 11 PNALMTPLI PSSM I KSEDVTPM EV- TAEKRSSTIFKTTKSVGSTQQTLENISNIAGNGSFSSPSSSHLPSENEKQQQIQPKAYNPETL TTIQTQDISQPGTFPAVSASSQLPNSDALLQQATQFQTRETQSREILQSDGTVVNLSQkTEASQQQQQSPLQEQAQTLQQQISSNIFPSPNSVSQLQNTIQQLQAGSFTGSTASGSSGSV DLVQQVLEAQQQLSSVLFSAPDGNENVQEQLSADI-FQQVSQIQSGVSPGMFSSTEPTVHTRPDNLLPGRAESVHPQSENTLSNQQQQQQQQQQV MESSAAMVMEMQQSICQAAAQIQSELFPSTASANGNLQQSPVYQQTSHMMSALSTNED-MQMQCELFSSPPAVSGNETSTTTTQQVATPGTTMFQTSSSGDGEETGTQAKQIQNSVFQTMVQMQHSGDNQPQVNLFSSTKSMM-SVQNSGTQQQGNGLFQQGNEMMSLQSGNFLQQSSHS-QAQLFHPQNPIADAQNLSQETQGSLFHSPNPIVHSQTSTTSSEQMQPPMFHSQSTIAVLQG SSVPQDQQSTNIFLSQSPMNNLQTNTVAQEAFFAAPNSISPLQSTSNSEQQAAFQQQAP- ISHIQTPMLSQEQAQPPQQGLFQPQVALGSLPPNPMPQSQQGTMFQSQHSIVAMQSNSPSQEQQPPPPRRPLPPLPLQQSILFSNQNTMATMASPKQPPPNMIFNPNQNPMAN-QEQQNQSIFHQQSNMAPMNQEQQPMQFQSQSTVSSLQNPGPTQSESSQTPLFHSSPQIQLVQGSPSSQEQQVTLFLSPASMSALQTSINQQDMQQSPLYSPQNN-MPGIQGATSSPQPQATLFHNTAGGTMNQLQNSPGSSQQTSGMFLFGIQNNCSQLLTSGPATLPDQLMAISQPGQPQNEGQPPVTTLLSQQMPENSPLASSINTNQNIEKID-LLVSLQNQGNNLTGSFNFAT5 IDR AroLITE A (SEQ ID NO: 30)TMVKKEISSPARPCSAEEAMKAMKTTGCNLDKVNIIPNALMTPLIPSSMIKSEDVTPMEV-TAEKRSSTIAKTTKSVGSTQQTLENISNIAGNGSASSPSSSHLPSENEKQQQIQPKAANPETLTTIQTQDISQPGTAPAVSASSQLPNSDALLQQATQAQTRETQSREILQSDGTVVNLSQLzTEASQQQQQSPLQEQAQTLQQQISSNIAPSPNSVSQLQNTIQQLQAGSATGSTASGSSGSVDLVQQVLEAQQQLSSVLASAPDGNENVQEQLSADIA-QQVSQIQSGVSPGMASSTEPTVHTRPDNLLPGRAESVHPQSENTLSNQQQQQQQQQQVMESSAAMVMEMQQSICQAAAQIQSELAPSTASANGNLQQSPVAQQTSHMMSALSTNED-MQMQCELASSPPAVSGNETSTTTTQQVATPGTTMAQTSSSGDGEETGTQAKQIQNSVAQTMVQMQHSGDNQPQVNLASSTKSMM-SVQNSGTQQQGNGLAQQGNEMMSLQSGNALQQSSHS-QAQLAHPQNPIADAQNLSQETQGSLAHSPNPIVHSQTSTTSSEQMQPPMAHSQSTIAVLQGSSVPQDQQSTNIALSQSPMNNLQTNTVAQEAAAAAPNSISPLQSTSNSEQQAAAQQQAP-ISHIQTPMLSQEQAQPPQQGLAQPQVALGSLPPNPMPQSQQGTMAQSQHSIVAMQSNSPSQEQQQQQQQQQQQSILASNQNTMATMASPKQPPPNMIANPNQNPMANQEQQNQSI-AHQQSNMAPMNQEQQPMQAQSQSTVSSLQNPGPTQSESSQTPLAHSSPQIQLVQGSPSSQEQQVTLALSPASMSALQTSINQQDMQQSPLASPQNN-MPGIQGATSSPQPQATLAHNTAGGTMNQLQNSPGSSQQTSGMALAGIQNNCSQLLTSGPATLPDQLMAISQPGQPQNEGQPPVTTLLSQQMPENSPLASSINTNQNIEKID-LLVSLQNQGNNLTGSAC / EBPa IDR Wild type (SEQ ID NO: 31)MRGRGRAGSPGGRRRRPAQAGGRRGSPCRENSNSPMESADFYEAEPRPPMSSHLQSP-PHAPSSAAFGFPRGAGPAQPPAPPAAPEPLGGICEHETSIDISAYIDPAAFNDEFLADLFQHSRQQEKAKAAVGPTGGGGGGDFDYPGAPAGPGGAVMPGGAHGPPPGYGCAAA-GYLDGRLEPLY-ERVGAPALRPLVIKQEPREEDEAKQLALAGLFPYQPPPPPPPSHPHPHPPPAHLAAPHLQFQIAHCGQC / EBPa IDR AroLITE A (SEQ ID NO: 32)MRGRGRAGSPGGRRRRPAQAGGRRGSPCRENSNSPMESADAAEAEPRPPMSSHLQSP-PHAPSSAAAGAPRGAGPAQPPAPPAAPEPLGGICEHETSIDISAAIDPAAANDEALADLAQHSRQQEKAKAAVGPTGGGGGGDADAPGAPAGPGGAVMPGGAHGPPPGAGCAAAGALD-GRLEPLAERVGAPALRPLVIKQEPREEDEAKQLALAGLAPAQPPPPPPPSHPHPHPPPAHL AAPHLQAQIAHCGQC / EBPa IDR AroPERFECT IS15 (SEQ ID NO: 33)MRGFRGRAGSPGGRRRRPAYQAGGRRGSPCRENSNFSPMESADEAEPRPP-MFSSHLQSP-PHAPSSAAFGPRGAGPAQPPAPPAYAPEPLGGICEHETSIFDISAIDPAANDELADFLQHSRQQEKAKAAVGFPTGGGGGGDDPGAPAYGPGGAVMPGGAHGPPYPGGCAAAGLD-GRLEPYL-ERVGAPALRPLVIKYQEPREEDEAKQLALAFGLPQPPPPPPPSHPHYPHPPPAHLAAPHLQQFIAHCGQC / EBPa IDR AroPERFECT IS15 +1 (SEQ ID NO: 34)MRGRFGRAGSPGGRRRRPAQYAGGRRGSPCRENSNSFPMESADEAEPRPPMSF;SHLQSP-PHAPSSAAGFPRGAGPAQPPAPPAAYPEPLGGICEHETSIDFISAIDPAANDELADLFQHSRQQEKAKAAVGPFTGGGGGGDDPGAPAGYPGGAVMPGGAHGPPPYGGCAAAGLD-GRLEPLY-ERVGAPALRPLVIKQYEPREEDEAKQLALAGFLPQPPPPPPPSHPHPYHPPPAHLAAPHLQQIFAHCGQC / EBPa IDR AroPERFECT IS15 +2 (SEQ ID NO: 35)MRGRGFRAGSPGGRRRRPAQAYGGRRGSPCRENSNSPFMESADEAEPRPP-MSSFHLQSP-PHAPSSAAGPFRGAGPAQPPAPPAAPYEPLGGICEHETSIDIFSAIDPAANDELADLQFHSRQQEKAKAAVGPTFGGGGGGDDPGAPAGPYGGAVMPGGAHGPPPGYGCAAAGLD-GRLEPLEYRVGAPALRPLVIKQEYPREEDEAKQLALAGLFPQPPPPPPPSHPHPHYPPPAHL AAPHLQQIAFHCGQC / EBPa IDR AroPERFECT IS10 (SEQ ID NO: 36)MRGRGRAGSFPGGRRRRPAQFAGGRRGSPCRYENSNSPMESAYDEAEPRPPMSF;SHLQSP-PHAPFSSAAGPRGAGFPAQPPAPPAAFPEPLGGICEHYETSIDISAIDYPAANDELADLFQHSRQQEKAKFAAVGPTGGGGFGGDDPGAPAGFPGGAVMPGGAFHGPPPGGCAAFAGLD-GRLEPLFERVGAPALRPFLVIKQEPREEYDEAKQLALAGYLPQPPPPPPPYSHPHPHPPPAY HLAAPHLQQIYAHCGQC / EBPa I DR Wild type (N) (SEQ ID NO: 37)MRGRGRAGSPGGRRRRPAQAGGRRGSPCRENSNSPMESADFYEAEPRPPMSSHLQSP-PHAPSSAAFGFPRGAGPAQPPAPPAAPEPLGGICEHETSIDISAYIDPAAFNDEFLADLFQHSC / EBPa IDR WT(N)-IS15 (SEQ ID NO: 38)MRGRGRAGSPGGRRRRPAQAGGRRGSPCRENSNSPMESADFYEAEPRPPMSSHLQSP-PHAPSSAAFGFPRGAGPAQPPAPPAAPEPLGGICEHETSIDISAYIDPAAFNDEFLADLFQHSRQQEKAKAAVGFPTGGGGGGDDPGAPAYGPGGAVMPGGAHGPPYPGGCAAAGLD-GRLEPYL-ERVGAPALRPLVIKYQEPREEDEAKQLALAFGLPQPPPPPPPSHPHYPHPPPAHLAAPHLQQFIAHCGQC / EBPa IDR WT(N)-FUSN (SEQ ID NO: 39)MRGRGRAGSPGGRRRRPAQAGGRRGSPCRENSNSPMESADFYEAEPRPPMSSHLQSP-PHAPSSAAFGFPRGAGPAQPPAPPAAPEPLGGICEHETSIDISAYIDPAAFNDEFLADLFQHSASNDYTQQATQSYGAYPTQPGQGYSQQSSQPYGQQSYSGYSQSTDTSGYGQSSYSSY-GQSQNTGYGTQSTPQGYGSTGGYGSSQSSQSSYGQQSSYPGYGQQPAPSSTSGSYGSSSQSSSYGQPQSGSYSQQPSYGGQQQSYGQQQSYNPPQGYG-QQNQYNSSSGGGGGGGGGGNY-GQDQSSMSSGGGSGGGYGNQDQSGGGGSGGYGQQDRGGRGRGGSGGGGGGGGGGYNRSSGGYEPRGRGGGRGGRGGMGGSDRGGFNKFGGPRDQGSRHDSEQDNSDNNTIC / EBPa IDR WT(N)-FUSNxs (SEQ ID NO: 40)MRGRGRAGSPGGRRRRPAQAGGRRGSPCRENSNSPMESADFYEAEPRPPMSSHLQSP-PHAPSSAAFGFPRGAGPAQPPAPPAAPEPLGGICEHETSIDISAYIDPAAFNDEFLADLFQHSASNDYTQQATQSYGAYPTQPGQGYSQQSSQPYGQQSYSGYSQSTDTSGYGQSSNGN2 Wild type (SEQ ID NO: 41)MFVKSETLELKEEEDVLVLLGSASPALAALTPLSSSADEEEEEEP-GASGGARRQRGAEAGQGARGGVAAGAEGCRPARLLGLVHDCKRRPSRARAVSRGAKTAETVQRIKKTRRLKANNRERNRMHNLNAALDALREVLPTFPEDAKLTKIETLRFAHNYI-WALTETLRLADHCGGGGGGLPGALFSEAVLLSPGGASAALSSSGDSPSPASTWSCTNSPAPSSSVSSNSTSPYSCTLSPASPAGSDMDYWQPPPPDKHRYAPHLPIARDCINGN2 AroLITE A (SEQ ID NO: 42)MAVKSETLELKEEEDVLVLLGSASPALAALTPLSSSADEEEEEEP-GASGGARRQRGAEAGQGARGGVAAGAEGCRPARLLGLVHDCKRRPSRARAVSRGAKTAETVQRIKKTRRLKANNRERNRMHNLNAALDALREVLPTFPEDAKLTKIETLRFAHNYI-WALTETLRLADHCGGGGGGLPGALFSEAVLLSPGGASAALSSSGDSPSPASTASCTNSPAPSSSVSSNSTSPASCTLSPASPAGSDMDAAQPPPPDKHRAAPHLPIARDCINGN2 AroPERFECT (SEQ ID NO: 43)MFVKSETLELKEEEDVLVLLGSASPALAALTPLSSSADEEEEEEP-GASGGARRQRGAEAGQGARGGVAAGAEGCRPARLLGLVHDCKRRPSRARAVSRGAKTAETVQRIKKTRRLKANNRERNRMHNLNAALDALREVLPTFPEDAKLTKIETLRFAHNYI-WALTETLRLADHCGGGGGGLPGALFSEAVLLSPGGASAAWLSSSGDSPSPASTSYCTNSPAPSSSVSSNYSTSPSCTLSPASPAWGSDMDQPPPPDKHRYAPHLPIARDCINGN2 N-IDR Wild type (SEQ ID NO: 44)MFVKSETLELKEEEDVLVLLGSASPALAALTPLSSSADEEEEEEP-GASGGARRQRGAEAGQGARGGVAAGAEGCRPARLLGLVHDCKRRPSRARAVSRGAKTAETVQRIKKTRRLKANNRERNRMHNLNAANGN2 N-IDR AroPERFECT (SEQ ID NO: 45)MFVKSETLELKEEEDVWLVLLGSASPALAALYT-PLSSSADEEEEEEYPGASGGARRQRGAEW-AGQGARGGVAAGAEYGCRPARLLGLVHDCWKRRPSRARAVSRGAYKTAETVQRIKKTRRYLKANNRERNRMHNLWNAANGN2 C-IDR Wild type (SEQ ID NO: 46)AVLLSPGGASAALSSSGDSPSPASTWSCTNSPAPSSSVSSNSTSPYSCTLSPAS-PAGSDMDYWQPPPPDKHRYAPHLPIARDCINGN2 C-IDR AroLITE A (SEQ ID NO: 47)AVLLSPGGASAALSSSGDSPSPASTASCTNSPAPSSSVSSNSTSPASCTLSPAS-PAGSDMDAAQPPPPDKHRAAPHLPIARDCINGN2 C-IDR AroPERFECT (SEQ ID NO: 48)AVLLSPGGASAAWLSSSGDSPSPASTSYCTNSPAPSSSVSSNYSTSPSCTLSPASPAW-GSDMDQPPPPDKHRYAPHLPIARDCIMY0D1 Wild type (SEQ ID NO: 49)MELLSPPLRDVDLTAPDGSLCSFATTDDFYDDPCFDSPDLRFFEDLDPRLMHVGALLK-PEEHSHFPAAVHPAPGAREDEHVRAPSGHHQAGRCLLWACKACKRKTTNADRRKAATMRERRRLSKVNEAFETLKRCTSSNPNQRLPKVEILRNAIRYIEGLQALLRDQDAAPP-GAAAAFYAPGPLPPGRGGEHYSGDSDASSPRSNCSDGMMDYSGPPSGARRRNCYEGAYYNEAPSEPRPGKSAAVSSLDCLSSIVERISTESPAAPALLLADVPSESPPRRQE-AAAPSEGESSGDPTQSPDAAPQCPAGANPNPIYQVLMY0D1 AroLITE A (SEQ ID NO: 50)MELLSPPLRDVDLTAPDGSLCSAATTDDAADDPCADSPDLRAAEDLDPRLMHVGALLK-PEEHSHAPAAVHPAPGAREDEHVRAPSGHHQAGRCLLAACKACKRKTTNADRRKAATMRERRRLSKVNEAAETLKRCTSSNPNQRLPKVEILRNAIRAIEGLQALLRDQDAAPP-GAAAAAAAPGPLPPGRGGEHASGDSDASSPRSNCSDGMMDASGPPSGARRRNCAEGAAANEAPSEPRPGKSAAVSSLDCLSSIVERISTESPAAPALLLADVPSESPPRRQE-AAAPSEGESSGDPTQSPDAAPQCPAGANPNPIAQVLMYOD1 AroPERFECT (SEQ ID NO: 51)MFELLSPPLRDVDLTAFPDGSLCSATTDDDDPYCDSPDLREDLDPRLMFHVGALLK-PEEHSHPAFAVHPAPGAREDEHVRFAPSGHHQAGRCLLACFKACKRKTTNADRRKAATMRERRRLSKVNEAFETLKRCTSSNPNQRLPKVEILRNAIRYIEGLQALLRDQDAAPP-GAAAAAPGPLPPGRGGEWHSGDSDASSPRSNCSFDGMMDSGPPSGARRRYNCEGANEAPSEPRPGYKSAAVSSLDCLSSIVYERISTESPAAPALLLYADVPSESPPRRQEAAYA-PSEGESSGDPTQSPYDAAPQCPAGANPNPIYQVLMYOD1 N-IDR Wild type (SEQ ID NO: 52)MELLSPPLRDVDLTAPDGSLCSFATTDDFYDDPCFDSPDLRFFEDLDPRLMHVGALLK-PEEHSHFPAAVHPAPGAREDEHVRAPSGHHQAGRCLLWACKACMYOD1 N-IDR AroLITE A (SEQ ID NO: 53)MELLSPPLRDVDLTAPDGSLCSAATTDDAADDPCADSPDLRAAEDLDPRLMHVGALLK-PEEHSHAPAAVHPAPGAREDEHVRAPSGHHQAGRCLLAACKACMYOD1 N-IDR AroPERFECT (SEQ ID NO: 54)MFELLSPPLRDVDLTAFPDGSLCSATTDDDDPYCDSPDLREDLDPRLMFHVGALLK-PEEHSHPAFAVHPAPGAREDEHVRFAPSGHHQAGRCLLACFKACMYOD1 C-IDR Wild type (SEQ ID NO: 55)AAAAFYAPGPLPPGRGGEHYSGDSDASSPRSNCSDGMMDYSGPPSGARRRNCYE-GAYYNEAPSEPRPGKSAAVSSLDCLSSIVERISTESPAAPALLLADVPSESPPRRQEAAAPSEGESSGDPTQSPDAAPQCPAGANPNPIYQVLMY0D1 C-IDR AroLITE A (SEQ ID NO: 56)AAAAAAAPGPLPPGRGGEHASGDSDASSPRSNCSDGMMDASGPPSGARRRNCAE-GAAANEAPSEPRPGKSAAVSSLDCLSSIVERISTESPAAPALLLADVPSESPPRRQEAAAPSEGESSGDPTQSPDAAPQCPAGANPNPIAQVLMYOD1 C-IDR AroPERFECT (SEQ ID NO: 57)AAAAAPGPLPPGRGGEWHSGDSDASSPRSNCSFDGMMDSGPPSGARRRYNCE-GANEAPSEPRPGYKSAAVSSLDCLSSIVYERISTESPAAPALLLYADVPSESPPRRQEAAYAPSEGESSGDPTQSPYDAAPQCPAGANPNPIYQVLOCT4 N-IDR Wild type (SEQ ID NO: 58)MAGHLASDFAFSPPPGGGDGSAGLEPGWVDPRTWLSFQGPPGGPGIGPGSEVLGISPCP-PAYEFCGGMAYCGPQVGLGLVPQVGVETLQPEGQAGARVESNSEGTSSEPCADRPNAVK LEKVEPTPEESQDMKALQKELEQOCT4 C-IDR Wild type (SEQ ID NO: 59)KGKRSSIEWSQREEYEATGTPFPGGAVSFPLPPGPHFGTPGYGSPHFT^TLYSVPFPEGEAFPSVPVTALGSPMHSNOCT4 N-IDR AroLITE (SEQ ID NO: 60)MAGHLASDAAASPPPGGGDGSAGLEPGAVDPRTALSAQGPPGGPGIGPGSEVLGISPCP- PAAEACGGMAACGPQVGLGLVPQVGVETLQPEGQAGARVESNSEGTSSEPCADRPNAVK LEKVEPTPEESQDMKALQKELEQOCT4 C-IDR AroLITE (SEQ ID NO: 61)KGKRSSIEASQREEAEATGTPAPGGAVSAPLPPGPHAGTPGAGSPHATTLASVPAPEGE-AAPSVPVTALGSPMHSNOCT4 N-IDR AroPERFECT (SEQ ID NO: 62)MAGHLASDFASPPPGGGDGSAGLFEPGVDPRTLSQGPPWGGPGIGPGSEVLG1WSPCP- PAECGGMACGFPQVGLGLVPQVGVEYTLQPEGQAGARVESFNSEGTSSEPCADRPYNAV KLEKVEPTPEEFSQDMKALQKELEQOCT4 C-IDR AroPERFECT (SEQ ID NO: 63)KGKRSSIEYSQREEEYATGTPPFG-GAVSPFLPPGPHFGTPGGSYPHTTLSFVPPEGEYAPSVPVFTALGSPFMHSNPDX1 I DR Wild type (SEQ ID NO: 64)MNGEEQYYAATQLYKDPCAFQRGPAPEFSASPPACLYMGRQPPPPPPHPFPGADGALEQG-SPPDISPYEVPPLADDPAVAHLHHHLPAQLALPHPPAGPFPEGAEPGVLEEPNRVQLPPDX1 I DR AroLITE (SEQ ID NO: 65)MNGEEQAAAATQLAKDPCAAQRGPAPEASASPPACLAMGRQPPPPPPHPAPGADGALEQG-SPPDISPAEVPPLADDPAVAHLHHHLPAQLALPHPPAGPAPEGAEPGVLEEPNRVQLPPDX1 I DR AroPERFECT (SEQ ID NO: 66)MNYGEEQAATQLKDPCYAQRGPAPESASPPYACLMGRQPPPPPPFHPPGALGALEQGS-FPPDISPEVPPLADYDPAVAHLHHHLPAFQLALPHPPAGPPEYGAEPGVLEEPNRVFQLPFFOXA3 I DR Wild type C (SEQ ID NO: 67)RRQKRFKLEEKVKKGGSGAATTTRNGTGSAASTTTPAATVTSPPQPPPPAPE-PEAQGGEDVGALDCGSPASSTPYFTGLELPGELKLDAPYNFNHPFSINNLMSEQTPAPPKLDVGFGGYGAEGGEPGVYYQGLYSRSLLNASFOXA3 I DR AroLITE C (SEQ ID NO: 68)RRQKRAKLEEKVKKGGSGAATTTRNGTGSAASTTTPAATVTSPPQPPPPAPE-PEAQGGEDVGALDCGSPASSTPAATGLELPGELKLDAPANANHPASINNLMSEQTPAPPKLDVGAGGAGAEGGEPGVAAQGLASRSLLNASFOXA3 I DR AroPERFECT C (SEQ ID NO: 69)RRQKRFKLEEKVKKGGSGYAATTTRNGTGSAFASTTTPAATVTSYPPQPPPPAPEPEFA-QGGEDVGALDCFGSPASSTPTGLEFLPGELKLDAPNNYHPSINNLMSEQTYPAPPKLDVGGGGYAEGGEPGVQGLSYRSLLNASS6Y AroPATCHYl (SEQ ID NO: 70)MSGSSSGSSGGSSSSGSSGSGGSSSYYYYYYYYYYYYYYSSGSGSSGSSSGGSSSGSSGSSSGSSGGSSSSGSSGSGGSSSSSGSGSSGSSSGGSSSGSS6Y AroPATCHY3 (SEQ ID NO: 71)MSGSSYYYYSGSSGGSSYYYYSSGSSGSGGSSSSSGSGSSGSSSGYYYYGSSSGSSGSSSGSSGGSSSSGSSGSGGSSSSSGSGSSGSSSGGSSSGSSSS6Y AroPERFECT (SEQ ID NO: 72)MSGSSSGYSSGGSSYSSGSSGYSGGSSSYSSGSGSYSGSSSGYGSSSGSYSGSSSGY-SSGGSSYSSGSSGYSGGSSSYSSGSGSYSGSSSGYGSSSGSYD6Y AroPERFECT (SEQ ID NO: 73)MSDDDSGYSSDGDSYSDDSDGYSDDSDSYDSDSDSYDGSDDGYGDDSDSYD-GDDSGYSDGDDSYSSDDDGYSDGDSDYSDDDGSYDGDSDGYGSDDGDYTBXT I DR wild type (SEQ ID NO: 74)RNDHKDVMEEPGDCQQPGYSQWGWLVPGAGTLCPPASSH-PQFGGSLSLPSTHGCERYPALRN-HRSSPYPSPYAHRNSSPTYADNSSACLSMLQSHDNWSSLGVPGHTSMLPVSHNASPPTGSSQYPSLWSVSNGTITPGSQTAGVSNGLGAQFFRGSPAHYT-PLTHTVSAATSSSSGSPMYEGAATVTDISDSQYDTAQSLLIASWTPVSPPSMTBXT I DR AroLITE (SEQ ID NO: 75)RNDHKDVMEEPGDCQQPGASQAGALVPGAGTLCPPASSHPQAGGSLSLPSTHGCERA-PALRN-HRSSPAPSPAAHRNSSPTAADNSSACLSMLQSHDNASSLGVPGHTSMLPVSHNASPPTGSSQAPSLASVSNGTITPGSQTAGVSNGLGAQAARGSPAHATPLTHTVSAATSSSSGSPMAE-GAATVTDISDSQADTAQSLLIASATPVSPPSMTBXT I DR AroPERFECT (SEQ ID NO:76)RYNDHKDVMEEPGDWCQQPGSQGLVPGWAGTLCPPASSHPFQGGSLSLPSTHGYCER-PALRN-HRSSYPPSPAHRNSSPTYADNSSACLSMLQYSHDNSSLGVPGHWTSMLPVSHNASPYPTGSSQPSLSVSWNGTITPGSQTAGFVSNGLGAQRGSPFAHTPLTHTVSAAY-TSSSSGSPMEGAYATVTDISDSQDTYAQSLLIASTPVSWPPSMMSGN1 I DR wild type (SEQ ID NO: 77)MDNLGETFLSLEDGLDSSDTAGLLASWDWKSRARPLELVQESPTQSLSPAPSLESYSE-VALPCGHSGASTGGSDGYGSHEAAGLVELDYSMLAFQPPYLHTAGGLKGQKGSKVKMSVMSGN1 I DR AroLITE (SEQ ID NO: 78)MDNLGETALSLEDGLDSSDTAGLLASADAKSRARPLELVQESPTQSLSPAPSLESASE-VALPCGHSGASTGGSDGAGSHEAAGLVELDASMLAAQPPALHTAGGLKGQKGSKVKMSVMSGN1 IDR AroPERFECT (SEQ ID NO: 79)MDNFLGETLSLEDGLDSSDWTAGLLASDKSRARPLWELVQESPTQSLSPAPYS-LESSEVALPCGHSGYASTGGSDGGSHEAAGYLVELDSMLAQPPLHTFAGGLKGQKGSKVKMSYVNKX2-5 IDR wild type (SEQ ID NO: 80)LGPPPPPARRIAVPVLVRDGKPCLGDPAAYAPAYGVGLNAYGYNAYPYPSYGGAACSPGY-SCAAYPAAPPAAQPPAASANSNFVNFGVGDLNTVQSPGMPQGNSGVSTLHGIRAWNKX2-5 IDR AroPERFECT (SEQ ID NO: 81)LYGPPPPPARRYIAVPVLVRDYGKPCLGDPAYAAPAGVGLNYAGNAPPSGGYAACSPG-SCAYAPAAPPAAQYPPAASANSNYVNGVGDLNTFVQSPGMPQGFNSGVSTLHGWIRACDX2 IDR wild type (SEQ ID NO: 82)MYVSYLLDKDVSMYPSSVRHSGGLNLAPQNFVSPPQYPDYGGYHVAAAAAATANLD-SAQSPGPSWPTAYGAPLREDWNGYAPGGAAAANAVAHGLNGGSPAAAMGYSSPAEYHAHHHPHHHPHH PAASPSCASG LLQTLN LG PPG PAATAAAEQLSPSGQCDX2 IDR AroPERFECT (SEQ ID NO: 83)MVSLLDKDVSYMPSSVRHSGGLYNLAPQNVSPPQYPDGGHVAAAAAFATANLD-SAQSPYGP-SPTAGAPLRYEDNGAPGGAAAYANAVAHGLNGGWSPAAAMGSSPAYEHAHHHPHHHPWH H PAASPSCASYG LLQTLN LG PPYG PAATAAAEQLYSPSGQReferences1 Lambert, S. A. et al. The Human Transcription Factors. Cell 175, 598-599, doi:10.1016 / j.cell.2O18.09.045 (2018).2 Lee, T. I. & Young, R. A. Transcriptional regulation and its misregulation in disease. Cell 152, 1237-1251 , doi:10.1016 / j. cell.2013.02.014 (2013).3 Levine, M., Cattoglio, C. & Tjian, R. Looping back to leap forward: transcription enters a new era. Cell 157, 13-25, doi:10.10167j.cell.2014.02.009 (2014).4 Graf, T. & Enver, T. Forcing cells to change lineages. Nature 462, 587-594, doi:10.1038 / nature08533 (2009).5 Takahashi, K. & Yamanaka, S. A decade of transcription factor-mediated reprogramming to pluripotency. Nature reviews. Molecular cell biology 17, 183-193, doi:10.1038 / nrm.2016.8 (2016).6 Stampfel, G. et al. Transcriptional regulators form diverse groups with context-dependent regulatory functions. Nature 528, 147-151 , doi:10.1038 / nature15545 (2015).7 Arnold, C. D. et al. A high-throughput method to identify trans-activation domains within transcription factor sequences. The EMBO journal 37, doi:10.15252 / embj.201798896 (2018).8 Alerasool, N., Leng, H., Lin, Z. Y., Gingras, A. C. & Taipale, M. Identification and functional characterization of transcriptional activators in human cells. Molecular cell, doi : 10.1016 / j.molcel.2021 .12.008 (2022).9 Erijman, A. etal. A High-Throughput Screen for Transcription Activation Domains Reveals Their Sequence Features and Permits Prediction by Deep Learning. Molecular cell78, 890-902 e896, doi:10.1016 / j.molcel.2020.04.020 (2020).10 Sanborn, A. L. et al. Simple biochemical features underlie transcriptional activation domain diversity and dynamic, fuzzy binding to Mediator. eLife 10, doi:10.7554 / eLife.68068 (2021).11 Staller, M. V. et al. Directed mutational scanning reveals a balance between acidic and hydrophobic residues in strong human activation domains. Cell Syst, doi:10.1016 / j.cels.2022.01.002 (2022).12 Arnold, C. D. etal. Genome-wide quantitative enhancer activity maps identified by STARR-seq. Science 339, 1074-1077, doi:10.1126 / science.1232542 (2013).13 Piskacek, S. et al. Nine-amino-acid transactivation domain: establishment and prediction utilities. Genomics 89, 756-768, doi:10.1016 / j.ygeno.2007.02.003 (2007).14 Farley, E. K. et al. Suboptimization of developmental enhancers. Science 350, 325-328, doi: 10.1126 / science. aac6948 (2015).15 Jiang, J. & Levine, M. Binding affinities and cooperative interactions with bHLH activators delimit threshold responses to the dorsal gradient morphogen. Cell 72, 741-752, doi:10.1016 / 0092- 8674(93)90402-c (1993).16 Crocker, J. et al. Low affinity binding site clusters confer hox specificity and regulatory robustness. Ce / / 160, 191-203, doi:10.1016 / j.cell.2O14.11.041 (2015).17 Ramos, A. I. & Barolo, S. Low-affinity transcription factor binding sites shape morphogen responses and enhancer evolution. Philos Trans R Soc Lond B Biol Sci 368, 20130018, doi : 10.1098 / rstb.2013.0018 (2013).18 Chong, S. et al. Imaging dynamic and selective low-complexity domain interactions that control gene transcription. Science 361 , doi:10.1126 / science.aar2555 (2018).19 Boija, A. et al. Transcription Factors Activate Genes through the Phase-Separation Capacity of Their Activation Domains. Cell 175, 1842-1855 e1816, doi:10.1016Zj.cell.2018.10.042 (2018).20 Basu, S. et al. Unblending of Transcriptional Condensates in Human Repeat Expansion Disease. Ce / / 181 , 1062-1079 e1030, doi:10.1016Zj.cell.2020.04.018 (2020).21 Asimi, V. et al. Hijacking of transcriptional condensates by endogenous retroviruses. Nature genetics, doi:10.1038 / s41588-022-01132-w (2022).22 Chong, S. et al. Tuning levels of low-complexity domain interactions to modulate endogenous oncogenic transcription. Molecular cell 82, 2084-2097 e2085, doi:10.1016 / j.molcel.2022.04.007 (2022).23 Brodsky, S. et al. Intrinsically Disordered Regions Direct Transcription Factor In Vivo Binding Specificity. Molecular cell 79, 459-471 e454, doi:10.1016 / j.molcel.2020.05.032 (2020).24 Wang, J. et al. A Molecular Grammar Governing the Driving Forces for Phase Separation of Prion-like RNA Binding Proteins. Cell 174, 688-699 e616, doi:10.1016Zj.cell.2018.06.006 (2018).25 Martin, E. W. et al. Valence and patterning of aromatic residues determine the phase behavior of prion-like domains. Science 367, 694-699, doi:10.1126 / science.aaw8653 (2020).26 Murthy, A. C. et al. Molecular interactions underlying liquid-liquid phase separation of the FUS low-complexity domain. Nature structural & molecular biology 26, 637-648, doi :10.1038 / s41594- 019-0250-x (2019).27 Alberti, S., Gladfelter, A. & Mittag, T. Considerations and Challenges in Studying Liquid-Liquid Phase Separation and Biomolecular Condensates. Cell 176, 419-434, doi:10.1016 / j.cell.2O18.12.035 (2019).28 Zhou, X. et al. Mutations linked to neurological disease enhance self-association of low- complexity protein sequences. Science 377, eabn5582, doi:10.1126 / science.abn5582 (2022).29 Kim, H. J. et al. Mutations in prion-like domains in hnRNPA2B1 and hnRNPAI cause multisystem proteinopathy and ALS. Nature 495, 467-473, doi:10.1038 / nature11922 (2013).30 Choi, J. M., Holehouse, A. S. & Pappu, R. V. Physical Principles Underlying the Complex Biology of Intracellular Phase Transitions. Annu Rev Biophys 49, 107-133, doi:10.1146 / annurev-biophys-121219-081629 (2020).31 Morgan, R., In der Rieden, P., Hooiveld, M. H. & Durston, A. J. Identifying HOX paralog groups by the PBX-binding region. Trends in genetics : TIG 16, 66-67, doi:10.1016 / s0168- 9525(99)01881-8 (2000).32 Kmita, M., van Der Hoeven, F., Zakany, J., Krumlauf, R. & Duboule, D. Mechanisms of Hox gene colinearity: transposition of the anterior Hoxbl gene into the posterior HoxD complex. Genes & development 14, 198-211 (2000).33 Popperl, H. et al. Segmental expression of Hoxb-1 is controlled by a highly conserved autoregulatory loop dependent upon exd / pbx. Cell 81 , 1031-1042, doi:10.1016 / s0092- 8674(05)80008-x (1995).34 Popperl, H. & Featherstone, M. S. An autoregulatory element of the murine Hox-4.2 gene. The EM BO journal 11 , 3673-3680 (1992).35 Janicki, S. M. et al. From silencing to gene expression: real-time analysis in single cells. Cell116, 683-698, doi:10.1016 / s0092-8674(04)00171-0 (2004).36 Xie, H., Ye, M., Feng, R. & Graf, T. Stepwise reprogramming of B cells into macrophages. Cell117, 663-676, doi:10.1016 / s0092-8674(04)00419-2 (2004).37 Garcia-Seisdedos, H., Empereur-Mot, C., Elad, N. & Levy, E. D. Proteins evolve on the edge of supramolecular self-assembly. Nature 548, 244-247, doi:10.1038 / nature23320 (2017).38 Friedman, A. D. & McKnight, S. L. Identification of two polypeptide segments of CCAAT / enhancer-binding protein required for transcriptional activation of the serum albumin gene. Genes & development 4, 1416-1426, doi:10.1101 / gad.4.8.1416 (1990).39 Rapino, F. et al. C / EBPalpha induces highly efficient macrophage transdifferentiation of B lymphoma and leukemia cell lines and impairs their tumorigenicity. Cell reports 3, 1153-1163, doi : 10.1016 / j.celrep.2O13.03.003 (2013).40 Stik, G. et al. CTCF is dispensable for immune cell transdifferentiation but facilitates an acute inflammatory response. Nature genetics 52, 655-661 , doi:10.1038 / s41588-020-0643-0 (2020).41 Zhang, Y. et al. Rapid single-step induction of functional neurons from human pluripotent stem cells. Neuron 78, 785-798, doi:10.1016 / j.neuron.2013.05.029 (2013).42 DelRosso, N. et al. Large-scale mapping and systematic mutagenesis of human transcriptional effector domains. bioRxiv, 2022.2008.2026.505496, doi:10.1101 / 2022.08.26.505496 (2022).43 Davis, R. L., Weintraub, H. & Lassar, A. B. Expression of a single transfected cDNA converts fibroblasts to myoblasts. Cell 51 , 987-1000 (1987).44 Weintraub, H. et al. Activation of muscle-specific genes in pigment, nerve, fat, liver, and fibroblast cell lines by forced expression of MyoD. Proceedings of the National Academy of Sciences of the United States of America 86, 5434-5438, doi:10.1073 / pnas.86.14.5434 (1989).45 Bajaj, P. et al. Patterning the differentiation of C2C12 skeletal myoblasts. Integr Biol (Camb) 3, 897-909, doi:10.1039 / c1ib00058f (2011).46 Feme, J. J., Karr, J. P., Tjian, R. & Darzacq, X. "Structure"-function relationships in eukaryotic transcription factors: The role of intrinsically disordered regions in gene regulation. Molecular cell 82, 3970-3984, doi:10.1016 / j.molcel.2022.09.021 (2022).47 Soto, L. F. etal. Compendium of human transcription factor effector domains. Molecular cell 82, 514-526, doi:10.1016 / j.molcel.2021 .11.007 (2022).48 Tycko, J. et al. High-Throughput Discovery and Characterization of Human Transcriptional Effectors. Cell 183, 2020-2035 e2016, doi: 10.10167j.cell.2020.11 .024 (2020).49 Kwon, I. et al. Phosphorylation-regulated binding of RNA polymerase II to fibrous polymers of low-complexity domains. Cell 155, 1049-1060, doi:10.1016 / j. cell.2013.10.033 (2013).50 Wang, Y. et al. A Prion-like Domain in Transcription Factor EBF1 Promotes Phase Separation and Enables B Cell Programming of Progenitor Chromatin. Immunity, doi:10.1016 / j. immuni.2020.10.009 (2020).51 Wang, J. et al. Phase separation of OCT4 controls TAD reorganization to promote cell fate transitions. Cell stem cell 28, 1868-1883 e1811 , doi:10.1016 / j.stem.2021 .04.023 (2021).52 Holehouse, A. S., Ginell, G. M., Griffith, D. & Boke, E. Clustering of Aromatic Residues in Prionlike Domains Can Tune the Formation, State, and Organization of Biomolecular Condensates. Biochemistry 60, 3566-3581 , doi:10.1021 / acs.biochem.1 c00465 (2021).53 Smith, Z. D., Sindhu, C. & Meissner, A. Molecular features of cellular reprogramming and development. Nature reviews. Molecular cell biology 17, 139-154, doi:10.1038 / nrm.2016.6 (2016).54 Vierbuchen, T. & Wernig, M. Molecular roadblocks for cellular reprogramming. Molecular cell 47, 827-838, doi:10.1016 / j.molcel.2012.09.008 (2012).55 Xu, J., Du, Y. & Deng, H. Direct lineage reprogramming: strategies, mechanisms, and applications. Cell stem cell 'iQ, 1 19-134, doi:10.1016 / j.stem.2015.01.013 (2015).56 Wang, H., Yang, Y., Liu, J. & Qian, L. Direct cell reprogramming: approaches, mechanisms and progress. Nature reviews. Molecular cell biology 22, 410-424, doi:10.1038 / s41580-021-00335- z (2021).57 Morris, S. A. & Daley, G. Q. A blueprint for engineering cell fate: current technologies to reprogram cell identity. Cell research 23, 33-48, doi:10.1038 / cr.2013.1 (2013).58 Bocchi, R., Masserdotti, G. & Gotz, M. Direct neuronal reprogramming: Fast forward from new concepts toward therapeutic approaches. Neuron 110, 366-393, doi: 10.1016 / j.neuron.2021 .11 .023 (2022).59 Schindelin, J. et al. Fiji: an open-source platform for biological-image analysis. Nature methods 9, 676-682, doi:10.1038 / nmeth.2019 (2012).60 Schmidl, C., Rendeiro, A. F., Sheffield, N. C. & Bock, C. ChlPmentation: fast, robust, low-input ChlP-seq for histones and transcription factors. Nature methods 12, 963-965, doi:10.1038 / nmeth.3542 (2015).61 Buenrostro, J. D., Giresi, P. G., Zaba, L. C., Chang, H. Y. & Greenleaf, W. J. Transposition of native chromatin for fast and sensitive epigenomic profiling of open chromatin, DNA-binding proteins and nucleosome position. Nature methods 10, 1213-1218, doi:10.1038 / nmeth.2688 (2013).62 Hu, H. etal. AnimalTFDB 3.0: a comprehensive resource for annotation and prediction of animal transcription factors. Nucleic acids research 47, D33-D38, doi:10.1093 / nar / gky822 (2019).63 Choi, J. M., Dar, F. & Pappu, R. V. LASSI: A lattice model for simulating phase transitions of multivalent proteins. PLoS computational biology 15, e1007028, doi:10.1371 / journal.pcbi.1007028 (2019).64 Harmon, T. S., Holehouse, A. S., Rosen, M. K. & Pappu, R. V. Intrinsically disordered linkers determine the interplay between phase separation and gelation in multivalent proteins. eLife 6, doi:10.7554 / eLife.30294 (2017).65 Harmon, T. S., Holehouse, A. S. & Pappu, R. V. Differential solvation of intrinsically disordered linkers drives the formation of spatially organized droplets in ternary systems of linear multivalent proteins. New J. Phys 20, 045002, doi:https: / / doi.orq / 10.1088 / 1367-2630 / aab8d9 (2018).66 Emenecker, R. J., Griffith, D. & Holehouse, A. S. Metapredict: a fast, accurate, and easy-to-use predictor of consensus disorder and structure. Biophys J 120, 4312-4319, doi:10.1016 / j.bpj.2O21 .08.039 (2021).67 Lancaster, A. K., Nutter-Upham, A., Lindquist, S. & King, O. D. PLAAC: a web and commandline application to identify proteins with prion-like amino acid composition. Bioinformatics 30, 2501-2502, doi:10.1093 / bioinformatics / btu310 (2014).68 Holehouse, A. S., Das, R. K., Ahad, J. N., Richardson, M. O. & Pappu, R. V. CIDER: Resources to Analyze Sequence-Ensemble Relationships of Intrinsically Disordered Proteins. Biophys J 112, 16-21 , doi:10.1016 / j.bpj.2016.11.3200 (2017).69 Wickham, H. ggplot2: Elegant Graphics for Data Analysis. Springer-Verlag New York ISBN 978- 3-319-24277-4 (2016).70 Lyons, H. et al. Functional partitioning of transcriptional regulators by patterned charge blocks. Cell 186, 327-345 e328, doi:10.1016 / j.cell.2022.12.013 (2023).71 Jumper, J. et al. Highly accurate protein structure prediction with AlphaFold. Nature 596, 583- 589, doi:10.1038 / s41586-021-03819-2 (2021).72 Dobin, A. et al. STAR: ultrafast universal RNA-seq aligner. Bioinformatics 29, 15-21 , doi:10.1093 / bioinformatics / bts635 (2013).73 Love, M. L, Huber, W. & Anders, S. Moderated estimation of fold change and dispersion for RNA-seq data with DESeq2. Genome biology 15, 550, doi: 10.1186 / s13059-014-0550-8 (2014).74 R-Core-Team. A language and environment for statistical computing. R Foundation for Statistical Computing http: / / www.R-proiect.org / . (2021).75 Gu, Z., Eils, R. & Schlesner, M. Complex heatmaps reveal patterns and correlations in multidimensional genomic data. Bioinformatics 32, 2847-2849, doi:10.1093 / bioinformatics / btw313 (2016).76 Langfelder, P., Zhang, B. & Horvath, S. Defining clusters from a hierarchical cluster tree: the Dynamic Tree Cut package for R. Bioinformatics 24, 719-720, doi:10.1093 / bioinformatics / btm563 (2008).77 Subramanian, A. et al. Gene set enrichment analysis: a knowledge-based approach for interpreting genome-wide expression profiles. Proceedings of the National Academy of Sciences of the United States of America 102, 15545-15550, doi:10.1073 / pnas.0506580102 (2005).78 Sepulveda, J. L., Gkretsi, V. & Wu, C. Assembly and signaling of adhesion complexes. Curr Top Dev Biol 68, 183-225, doi:10.1016 / S0070-2153(05)68007-6 (2005).79 Lin, H. C. etal. NGN2 induces diverse neuron types from human pluripotency. Stem cell reports 16, 2118-2127, doi:10.1016 / j.stemcr.2021 .07.006 (2021).80 Nehme, R. et al. Combining NGN2 Programming with Developmental Patterning Generates Human Excitatory Neurons with NMDAR-Mediated Synaptic Transmission. Cell reports 23, 2509-2523, doi:10.1016 / j.celrep.2018.04.066 (2018).81 Raudvere, U. et al. g : Profiler: a web server for functional enrichment analysis and conversions of gene lists (2019 update). Nucleic acids research 47, W191-W198, doi:10.1093 / nar / gkz369 (2019).82 Yu, G., Wang, L. G., Han, Y. & He, Q. Y. clusterProfiler: an R package for comparing biological themes among gene clusters. OMICS 16, 284-287, doi:10.1089 / omi.2011.0118 (2012).83 Wu, T. et al. clusterProfiler 4.0: A universal enrichment tool for interpreting omics data. Innovation (N Y) 2, 100141 , doi:10.1016 / j.xinn.2O21.100141 (2021).84 Zheng, G. X. et al. Massively parallel digital transcriptional profiling of single cells. Nature communications 8, 14049, doi:10.1038 / ncomms14049 (2017).85 Hao, Y. et al. Integrated analysis of multimodal single-cell data. Cell 184, 3573-3587 e3529, doi:10.1016Zj.cell.2021.04.048 (2021).86 Choi, J. et al. Evidence for additive and synergistic action of mammalian enhancers during cell fate determination. eLife 10, doi:10.7554 / eLife.65381 (2021).87 La Manno, G. et al. RNA velocity of single cells. Nature 560, 494-498, doi:10.1038 / s41586-018- 0414-6 (2018).88 Bergen, V., Lange, M., Peidli, S., Wolf, F. A. & Theis, F. J. Generalizing RNA velocity to transient cell states through dynamical modeling. Nature biotechnology 38, 1408-1414, doi:10.1038 / s41587-020-0591-3 (2020).89 Wolf, F. A. et al. PAGA: graph abstraction reconciles clustering with trajectory inference through a topology preserving map of single cells. Genome biology 20, 59, doi:10.1186 / s13059-019- 1663-x (2019).90 Li, H. & Durbin, R. Fast and accurate short read alignment with Burrows-Wheeler transform. Bioinformatics 25, 1754-1760, doi:10.1093 / bioinformatics / btp324 (2009).91 Danecek, P. et al. Twelve years of SAMtools and BCFtools. Gigascience 10, doi:10.1093 / gigascience / giab008 (2021).92 McKenna, A. et al. The Genome Analysis Toolkit: a MapReduce framework for analyzing nextgeneration DNA sequencing data. Genome research 20, 1297-1303, doi:10.1 101 / gr.107524.110 (2010).93 Zhang, Y. et al. Model-based analysis of ChlP-Seq (MACS). Genome biology 9, R137, doi:10.1 186 / gb-2008-9-9-r137 (2008).94 Ross-Innes, C. S. et al. Differential oestrogen receptor binding is associated with clinical outcome in breast cancer. Nature 481 , 389-393, doi:10.1038 / nature10730 (2012).95 Lopez-Delisle, L. et al. pyGenomeTracks: reproducible plots for multivariate genomic datasets. Bioinformatics 37, 422-423, doi:10.1093 / bioinformatics / btaa692 (2021).96 Neumann, T. et al. Quantification of experimentally induced nucleotide conversions in high- throughput sequencing datasets. BMC Bioinformatics 20, 258, doi:10.1186 / s12859-019-2849- 7 (2019).97 Liao, Y., Smyth, G. K. & Shi, W. featurecounts: an efficient general purpose program for assigning sequence reads to genomic features. Bioinformatics 30, 923-930, doi:10.1093 / bioinformatics / btt656 (2014).98 Ramirez, F. et al. deepTools2: a next generation web server for deep-sequencing data analysis. Nucleic acids research 44, W160-165, doi:10.1093 / nar / gkw257 (2016).99 Bocchi, R., Masserdotti, G., and Gotz, M. (2022). Direct neuronal reprogramming: Fast forward from new concepts toward therapeutic approaches. Neuron 110, 366-393.100 Hou, X., Zaks, T., Langer, R., and Dong, Y. (2021). Lipid nanoparticles for mRNA delivery. Nat Rev Mater 6, 1078-1094.101 Ofenbauer, A., and Tursun, B. (2019). Strategies for in vivo reprogramming. Current opinion in cell biology 61 , 9-15.102 Stephan, M.T. (2021). Empowering patients from within: Emerging nanomedicines for in vivo immune cell reprogramming. Semin Immunol 56, 101537.

Claims

Claims1. A method of modifying the activity of a transcription factor comprising the steps:(a) providing an initial transcription factor including an intrinsically disordered region (I DR) , particularly an IDR comprising at least 4 aromatic amino acid residues dispersed therein,(b) altering the periodicity and / or number of aromatic amino acid residues in the IDR, thereby obtaining an altered transcription factor,(c) optionally measuring the transcriptional activity of the altered transcription factor, and(d) obtaining an altered transcription factor with modified transcriptional activity.

2. The method of claim 1 wherein the initial transcription factor is a mammalian transcription factor, particularly a human transcription factor.

3. The method of claim 1 or 2, comprising:(a) providing an initial transcription factor including an IDR, particularly an IDR comprising at least 4 aromatic amino acid residues dispersed therein,(b) increasing the periodicity and / or number of aromatic amino acid residues in the IDR, thereby obtaining an altered transcription factor,(c) optionally measuring the transcriptional activity of the altered transcription factor, and(d) obtaining an altered transcription factor with increased transcriptional activity.

4. The method of claim 3, wherein the altered transcription factor with increased transcriptional activity is selected from the group consisting of altered HOXD4, HOXC4, OCT4, C / EBPa, NGN2, MYOD-1 , PDX1 , FOXA3, BRACHYURY (TBXT), MSGN1 , NKX2-5 and CDX2 transcription factor.

5. The method of claim 3 or 4, wherein step (b) comprises(i) increasing the periodicity by rearranging the positions of aromatic amino acid residues originally present in the IDR; or(ii) increasing the number of aromatic amino acid residues originally present in the IDR; or(iii) a combination of (i) and (ii).

6. The method of claim 1 or 2, comprising the steps:(a) providing an initial transcription factor including an I DR, particularly an I DR comprising at least 4 aromatic amino acid residues dispersed therein, and(b) reducing the periodicity and / or number of aromatic amino acid residues in the I DR, thereby obtaining an altered transcription factor,(c) measuring the transcriptional activity of the altered transcription factor, and(d) obtaining an altered transcription factor with reduced detectable transcriptional activity.

7. The method of claim 6, wherein the altered transcription factor with reduced transcriptional activity is selected from the group consisting of altered HOXD4, HOXC4, C / EBPa, MYOD-1 , HOXB1 , EGR1 , NFAT5, NANOG, and OCT4 transcription factors.

8. The method of claim 6 or 7, wherein step (b) comprises(i) decreasing the periodicity by rearranging the positions of aromatic amino acid residues originally present in the I DR; or(ii) decreasing the number of aromatic amino acid residues originally present in the IDR; or(iii) a combination of (i) and (ii).

9. An altered transcription factor including an intrinsically disordered region (I DR), which comprises an increased number and / or periodicity of aromatic amino acid residues in the IDR compared to a corresponding wild-type transcription factor, wherein the altered transcription factor has an increased transcriptional activity compared to the corresponding wild-type transcription factor and wherein the corresponding wild-type transcription factor particularly comprises an IDR including at least 4 aromatic amino acid residues.

10. An altered transcription factor including an intrinsically disordered region (I DR), which comprises a decreased number and / or periodicity of aromatic amino acid residues in the IDR compared to a corresponding wild-type transcription factor, wherein the altered transcription factor has a reduced detectable transcriptional activity compared to the corresponding wild-type transcription factor and whereinthe corresponding wild-type transcription factor particularly comprises an I DR including at least 4 aromatic amino acid residues.

11. A recombinant nucleic acid molecule encoding an altered transcription factor of claim 9 or 10, particularly in operative linkage with an expression control sequence, or a vector comprising said nucleic acid molecule.

12. A recombinant cell comprising a nucleic acid molecule or a vector of claim 11 .

13. Use of an altered transcription factor of claim 9 or 10, a nucleic acid molecule, a vector of claim 11 or a cell of claim 12 in a transcription factor-mediated process, e.g., in transcription factor-mediated cell reprogramming.

14. The use of claim 13 in a method for the reprogramming of cells to obtain macrophages, neurons, muscle-like cells, pancreatic B cells, hepatocytes, or pluripotent stem cells, e.g., induced pluripotent stem cells.

15. An in vitro method of transcription factor-mediated cell reprogramming comprising the culturing of a starting cell under suitable conditions in the presence of at least one altered transcription factor of claim 9 or 10 and obtaining a desired product cell.

Citation Information

Patent Citations

  • Unblending of transcriptional condensates in human repeat expansion disease

    WO2021219721A1

  • Methods and assays for modulating gene transcription by modulating condensates

    US20220120736A1