Compositions, systems, and methods for creating, identifying, and characterizing effector domains for activating and silencing gene expression.
A high-throughput system for discovering and characterizing effector domains addresses the limitation of a small toolbox in synthetic transcription factors, enabling the development of enhanced transcription factors for gene therapy and synthetic biology.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- THE BOARD OF TRUSTEES OF THE LELAND STANFORD JUNIOR UNIV
- Filing Date
- 2021-05-04
- Publication Date
- 2026-06-03
AI Technical Summary
Existing methods for engineering synthetic transcription factors are limited by a small toolbox of effector domains, necessitating the development of new methods to expand this toolbox for applications in gene therapy, cell therapy, and synthetic biology.
A high-throughput system for discovering and characterizing effector domains is provided, involving the preparation of a domain library, transformation of reporter cells, treatment with agents to induce DNA-binding domains, separation of cells based on marker presence, sequencing, and calculating ratios to identify transcriptional repressors or activators.
This approach significantly expands the toolbox of effector domains, enabling the creation of enhanced synthetic transcription factors for modulating gene expression, thereby enhancing applications in gene therapy and synthetic biology.
Smart Images

Figure 0007869753000092 
Figure 0007869753000093 
Figure 0007869753000094
Abstract
Description
Technical Field
[0001] Cross - reference to Related Applications This application claims the benefit of U.S. Provisional Application No. 63 / 019,706, filed May 4, 2020, and U.S. Provisional Application No. 63 / 074,793, filed September 4, 2020, each of which is hereby incorporated by reference in its entirety.
[0002] Field Provided herein are compositions, systems, and methods for creating, identifying, and characterizing effector domains for activating and silencing gene expression. In particular, a high - throughput system for discovering and characterizing effector domains is provided.
[0003] Statement Regarding Federally Sponsored Research This invention was made with government support under Grant No. GM128947 awarded by the National Institutes of Health. The United States Government has certain rights in this invention.
Background Art
[0004] Existing efforts to engineer synthetic transcription factors have taken activation and repression domains from a small toolbox of already - discovered effector domains. There is a need for new methods to expand this toolbox.
Summary of the Invention
Problems to be Solved by the Invention
[0005] (Summary of the Invention) Compositions, systems, and methods for the creation, identification, and characterization of effector domains for activating and silencing gene expression are provided herein. In particular, a high-throughput system for discovering and characterizing effector domains is provided. In some embodiments herein, a high-throughput method for discovering and characterizing effector domains that greatly expands the toolbox is provided. These domains meet the urgent need to engineer enhanced synthetic transcription factors for applications in gene therapy and cell therapy, synthetic biology, and functional genomics.
Means for Solving the Problems
[0006] In some embodiments, a method for identifying an effector domain comprises: a) preparing a domain library comprising a plurality of nucleic acid sequences configured to express fusion proteins each comprising a protein domain linked to an inducible DNA-binding domain; b) transforming reporter cells with the domain library, wherein the reporter cells comprise a bipartite reporter gene comprising a surface marker and a fluorescent protein under the control of a strong promoter, and the bipartite reporter gene is capable of being silenced by a putative transcriptional repression domain upon treatment with an agent configured to induce the inducible DNA-binding domain; c) treating the reporter cells with the agent for a length of time sufficient for degradation of intracellular proteins and mRNAs; d) separating the reporter cells based on the presence or absence of the surface marker, fluorescent protein, or a combination thereof; e) sequencing the protein domain from the separated reporter cells; f) calculating, for each protein domain sequence, the ratio of the sequencing count from reporter cells without the surface marker, fluorescent protein, or a combination thereof to the sequencing count from reporter cells with the surface marker, fluorescent protein, or a combination thereof; and g) identifying the protein domain as a transcriptional repressor.
[0007] In some embodiments, a method for identifying effector domains includes: a) preparing a domain library comprising a plurality of nucleic acid sequences configured to express a fusion protein, each containing a protein domain linked to an inducible DNA-binding domain; b) transforming reporter cells with the domain library, wherein the reporter cells contain a bipartite reporter gene comprising a surface marker and a fluorescent protein under the control of a weak promoter, and the bipartite reporter gene is capable of being activated by a putative transcriptional activation domain after treatment with a drug configured to induce an inducible DNA-binding domain; and c) transforming the reporter cells with the drug. The method comprises the steps of: d) treating the cells with an agent for a period of time necessary for the production of intracellular proteins and mRNA; e) isolating reporter cells based on the presence or absence of a surface marker, fluorescent protein, or combination thereof; f) sequencing protein domains from the isolated reporter cells; g) calculating the ratio of sequencing counts from reporter cells without a surface marker, fluorescent protein, or combination thereof to sequencing counts from reporter cells with a surface marker, fluorescent protein, or combination thereof for each protein domain sequence; and g) identifying the protein domains as transcription activators.
[0008] In some embodiments, the method further includes the step of stopping the drug treatment of reporter cells and repeating steps d-g one or more times. In some embodiments, steps d-g are repeated for at least 48 hours after stopping the drug treatment of reporter cells.
[0009] In some embodiments, each protein domain has 80 amino acids or less. In some embodiments, the protein domain is derived from a nuclear-localized protein. In some embodiments, the protein domain includes the amino acid sequence of a wild-type protein domain derived from a nuclear-localized protein. In some embodiments, the protein domain includes a mutant amino acid sequence of a protein domain derived from a nuclear-localized protein.
[0010] In some embodiments, the inducible DNA-binding domain includes a tag.
[0011] In some embodiments, the method further includes the step of measuring the expression level of the protein domain. In some embodiments, the expression level is determined by measuring the relative presence or absence of a tag on the DNA-binding domain.
[0012] In some embodiments, reporter cells are treated with the drug for at least 3 days. In some embodiments, reporter cells are treated with the drug for at least 5 days. In some embodiments, reporter cells are treated with the drug for at least 24 hours. In some embodiments, reporter cells are treated with the drug for at least 48 hours.
[0013] In some embodiments, a protein domain is identified as a transcriptional repressor if the log2 ratio is at least two standard deviations (e.g., higher) from the mean of the poorly expressed negative control.
[0014] In some embodiments, a protein domain is identified as a transcription activator if the log2 ratio is at least two standard deviations (e.g., lower) from the mean of a low-expression negative control.
[0015] This specification also provides synthetic transcription factors comprising one or more transcriptional activating domains, one or more transcriptional repressing domains, or a combination thereof, fused to a heterologous DNA-binding domain. In some embodiments, at least one of the one or more transcriptional activating domains or at least one of the one or more transcriptional repressing domains comprises an amino acid sequence having at least 70% identity to any of SEQ ID NOs: 1 to 896.
[0016] In some embodiments, the synthetic transcription factor comprises two or more transcriptional activation domains or two or more transcriptional repression domains fused to a heterologous DNA-binding domain.
[0017] In some embodiments, at least one of the one or more transcriptional activation domains contains an amino acid sequence having at least 70% identity to any of SEQ ID NOs. 563-664. In some embodiments, at least one of the one or more transcriptional activation domains is selected from those found in Table 2.
[0018] In some embodiments, at least one of the one or more transcriptional repression domains contains an amino acid sequence having at least 70% identity to any of SEQ ID NOs: 1-562 and 665-896. In some embodiments, at least one of the one or more transcriptional repression domains is selected from those found in Tables 1, 3, or 4.
[0019] In some embodiments, one or more transcriptional activation domains or one or more transcriptional repression domains are identified by the methods disclosed herein.
[0020] In some embodiments, the heterologous DNA-binding domain includes a programmable DNA-binding domain. In some embodiments, the DNA-binding domain is derived from a clustered regularly spaced short-chain palindromic repeat-associated (Cas) protein. In some embodiments, the DNA-binding domain is derived from a transcription activator-like effector (TALE) domain.
[0021] This specification also provides nucleic acids encoding synthetic transcription factors or effector domains disclosed herein. In some embodiments, the nucleic acids are under the control of an inducible promoter. In some embodiments, the nucleic acids are under the control of a tissue-specific promoter. In some embodiments, the nucleic acids encode at least one further transcription factor or effector domain.
[0022] This specification further provides compositions or systems comprising synthetic transcription factors, nucleic acids, vectors, or cells disclosed herein. In some embodiments, the composition comprises two or more synthetic transcription factors, nucleic acids, vectors, or cells. In some embodiments, the composition further comprises a guide RNA or a nucleic acid encoding a guide RNA.
[0023] In addition, a method for modulating the expression of at least one target gene within a cell is also provided. The method comprises the step of introducing at least one synthetic transcription factor, nucleic acid, vector, or composition or system described herein into a cell. The gene expression of at least one target gene is modulated when the gene expression level of at least one target gene increases or decreases compared to the normal gene expression level for at least one target gene. In some embodiments, the synthetic transcription factor comprises a DNA-binding domain of a Cas protein, and the method further comprises the step of bringing the cell into contact with at least one guide RNA.
[0024] In some embodiments, the cells are in vitro (e.g., ex vivo) cells or cells in a subject.
[0025] In some embodiments, the gene expression of at least two genes is modulated. [Brief explanation of the drawing]
[0026] [Figure 1A]This figure shows high-throughput recruitment to measure the transcriptional repressive activity of thousands of Pfam annotation domains derived from nuclear-localized proteins. The length of the Pfam annotation domain within nuclear-localized human proteins is shown. Domains with ≤80 amino acids were selected for inclusion in the library. [Figure 1B] This figure shows high-throughput recruitment measuring the transcriptional repressive activity of thousands of Pfam annotation domains derived from nuclear localized proteins. A schematic diagram of the screen for identifying transcriptional repressors is also shown. The repression reporter uses a potent pEF promoter, which can be silenced by doxycycline-mediated recruitment of repressive domains. Cells were treated with doxycycline for 5 days, and on-cells and off-cells were magnetically separated and the domains sequenced. Doxycycline was removed, and further time points were set on days 9 and 13. [Figure 1C] This figure shows high-throughput recruitment measuring the transcriptional repressive activity of thousands of Pfam annotation domains derived from nuclear localized proteins. The log2 (off:on) ratio is shown, indicating reproducibility from independently transduced biological replicates, and selected domain families are colored. [Figure 1D] This figure shows that high-throughput recruitment measures the transcriptional repressive activity of thousands of Pfam annotation domains derived from nuclear localized proteins. The box plots show the top repressive domain families, ranked by the maximum repressive intensity of the domains within the family on day 5. [Figure 1E] This figure shows high-throughput recruitment measuring the transcriptional repressive activity of thousands of Pfam annotation domains derived from nuclear localized proteins. The time progression of individual validation of hit RYBP domains, measured by flow cytometry, is also shown. [Figure 1F]This figure shows high-throughput recruitment measuring the transcriptional repressive activity of thousands of Pfam annotation domains derived from nuclear localized proteins. Further validation timelines for the panel of repressive domains are shown. Domain lengths are listed in parentheses, as some domains were examined as the exact 80-amino acid sequence derived from the library, and some domains were examined as shorter sequences trimmed to the region annotated as a domain by Pfam. 1000 ng / ml of doxycycline was added on day 0 and removed on day 5. [Figure 1G] This figure shows high-throughput recruitment measuring the transcriptional repressive activity of thousands of Pfam annotation domains derived from nuclear localized proteins. Correlation of screen measurements for a collection of KRAB effector domains with individual validated flow cytometry measurements. [Figure 2A] This figure shows that the repressive KRAB domain colocalizes with and binds to the KAP1 repressive cofactor within a more recent KRAB zinc finger protein. The silencing function of KRAB was compared between the architecture of the KRAB domain and that of naturally occurring KRAB zinc finger proteins. [Figure 2B] This figure shows that the repressive KRAB domain colocalizes with and binds to the KAP1 repressive cofactor within a more recent KRAB zinc finger protein. The silencing function of KRAB was compared with the evolutionary age of the KRAB zinc finger gene, which was determined by finding the most recent ortholog of the gene using its whole DNA-binding zinc finger array sequence (the age was published in Trono, 2017). [Figure 2C]This figure shows that the repressive KRAB domain colocalizes with the KAP1 repressive cofactor and is located within a more recent KRAB zinc finger protein that binds to it. The KRAB domains were classified as silencers or non-silencers, and their genomic localization in the ChIP-seq dataset was compared with the localization of the repressive cofactor, KRAB-related protein 1 (KAP1). [Figure 2D] This figure shows that the repressive KRAB domain is located within more recent KRAB zinc finger proteins that colocalize with and bind to the KAP1 repressive cofactor. The figure also shows the repressive intensity distribution of KRAB domains, categorized based on whether these KRAB zinc finger genes significantly interact with the repressive cofactor KAP1, using a mass spectrometry dataset (Helleboid, 2019). The dot color represents the quintile for the KRAB domain expression level. [Figure 3A] This figure shows that deep-depth mutant scanning of the ZNF10 KRAB domain identifies substitutions that reduce or enhance repressive activity. The deep-depth mutant scanning library includes all single substitutions, as well as consecutive double and triple substitutions within the KRAB domain derived from ZNF10. The DNA oligo is designed to differ significantly from the protein sequence due to variations in codon usage. Residues in red are different from the WT sequence. [Figure 3B] This figure shows that a deep mutation scan of the ZNF10 KRAB domain identifies substitutions that reduce or enhance repressive activity. The measured values of repressors from all single-substitution and triple-substitution mutants, compared to the wild-type (WT), are shown below the schematic diagram for the KRAB domain. [Figure 3C] This figure shows that a deep mutation scan of the ZNF10 KRAB domain identifies substitutions that reduce or enhance repressive activity. The average effect of the mutations on repression at day 9, compared to sequence conservation (calculated by ConSurf) with multiple sequence alignments for all human KRAB domains. [Figure 3D]This figure shows that a deep mutation scan of the ZNF10 KRAB domain identifies substitutions that reduce or enhance repressive activity. It also shows the correlation of high-throughput measurements with previously published low-throughput data using CAT assays in different cell types. [Figure 3E] This figure shows that a deep mutation scan of the ZNF10 KRAB domain identifies substitutions that reduce or enhance repressive activity. Individual time courses for KRAB mutants examine the effects of substitutions at the A-box / B-box and N-terminus. [Figure 3F] This figure shows that a deep mutagenesis scan of the ZNF10 KRAB domain identifies substitutions that reduce or enhance suppressive activity. For each position at each time point in Figure 3B, the distribution of all single substitutions was compared to the distribution of the wild-type effect (Wilcoxon rank-sum test). Positions where the signed log10(p) at day 5 is <-5 are colored red (highly significant decrease in silencing), positions where the signed log10(p) at day 9 is <-5 but the signed log10(p) at day 5 is not <-5 are colored green, and position W8 where the log10(p) at day 13 is >5 is colored blue (highly significant increase). The horizontal dashed line indicates the hit threshold. The ConSurf score for sequence conservation is shown in orange. [Figure 3G] This figure shows that a deep mutagenesis scan of the ZNF10 KRAB domain identifies substitutions that reduce or enhance repressive activity. When mutations occur, residues that invalidate silencing on day 5 are mapped to the ordered region of the NMR structure of the mouse KRAB A box (PDB:1v65). [Figure 4A]This figure shows the collinearity between homeodomain repression strength and Hox gene organization. It ranks homeobox gene families or classes based on median repression strength on day 5. The HOXL and NKL subclasses, as well as the PRD and LIM classes, of the ANTP class homeodomains containing the strongest homeodomain repressors, are separated into individual gene families (Holland BMC, 2007), while the remaining classes are aggregated. The dot color represents the quintile of homeodomain expression level as measured in the high-throughput expression assay. [Figure 4B] This figure shows the collinearity between homeodomain repression strength and Hox gene organization. It represents the repression strength of homeodomains derived from the Hox gene family on day 5. Arrows indicate genes found within four human Hox loci and point to the direction of Hox gene transcription. Graebers separate gene families. Spearman's Rhoe and p-values were calculated for all Hox genes, relating gene count to repression strength. The data was filtered to remove any domains with a count less than 10 in any of the sequencing samples on day 5. [Figure 5A] This figure shows that high-throughput recruitment discovers the activating domain within ZNF473, including potent, acidic, and branched variants of the KRAB domain. Schematic diagrams of an activation reporter using a weak minCMV promoter, which can be activated by doxycycline-mediated recruitment of the activating domain, and a schematic diagram of an activation screen are also shown. A pool of cells was treated with doxycycline for 48 hours, and on-cells and off-cells were magnetically separated using a ProG Dynabead, followed by domain sequencing. [Figure 5B]This figure shows that high-throughput recruitment discovers activating domains within ZNF473, including potent, acidic, and branched variants of the KRAB domain. Known activating domain families (FOXO-TAD, Myb LMSTEN, TORC_C) are color-coded to show reproducibility from independently transduced biological replicates using log2 (off:on) ratios. [Figure 5C] This figure shows that high-throughput recruitment discovers activating domains within ZNF473, including potent, acidic, and branched variants of the KRAB domain. This is GO-term enrichment for genes containing domains with activation strength below a threshold. [Figure 5D] This figure shows that high-throughput recruitment discovers the activating domain within ZNF473, including potent, acidic, and branched variants of the KRAB domain. The activating domain (red) is more acidic than the non-hit domain (gray). [Figure 5E] This figure shows that high-throughput recruitment discovers activating domains within ZNF473, including potent, acidic, and branched variants of the KRAB domain. It also shows a list of domain families ranked by their average activation intensity. [Figure 5F] This figure shows that high-throughput recruitment discovers the activating domain within ZNF473, including potent, acidic, and branched variants of the KRAB domain. Sequence alignment and clustering of the KRAB domain yielded results similar to the classification in Helleboid, 2019. The most branched KRAB sequence cluster is the mutant KRAB shown in green. The results from the screen are shown below the heatmap. Standard KRAB functions as a repressor when expression is favorable. Mutant KRAB shows mixed effects as both a repressor and activator in the screen, and does not show any transcriptional effect. [Figure 6A]This figure shows how a tiling library reveals novel autonomy-inhibiting domains within large chromatin regulatory proteins. The graph illustrates the library, where 80-amino acid tiles cover protein sequences with 10-amino acid sliding windows. [Figure 6B] This figure shows that tiling libraries reveal novel autonomy-repressing domains within large chromatin regulatory proteins. It demonstrates the reproducibility of log2 (off:on) ratios from independently transduced biological replicates. [Figure 6C] This figure shows that tiling libraries reveal novel autonomous repressive domains within large chromatin regulatory proteins. Repression on day 5 is compared to the known domain architecture for the MGA protein. Two repressive domains are found outside the existing annotation region. [Figure 6D] This figure shows that the tiling library reveals a novel autonomy-inhibiting domain within large chromatin regulatory proteins. Time-course flow cytometry examines individual MGA effectors as 80-amino acid tiles. [Figure 6E] This figure shows that tiling libraries reveal novel autonomous repressive domains within large chromatin regulatory proteins. Effectors were minimized to 10-30 amino acid subtiles by selecting sequences shared between tiles exhibiting repressive activity on the screen. These minimized sequences were individually validated over time using flow cytometry. [Figure 6F]This figure shows that the tiling library reveals a novel autonomous repressive domain within a large chromatin regulatory protein. Individual validation of an additional 80 amino acid repressor hit by tiling screening was performed. rTetR-tile fusions were delivered to K562 reporter cells via lentivirus, and the cells were treated with 100 ng / ml doxycycline for 5 days, followed by doxycycline removal. The cells were analyzed by flow cytometry, and the percentage of off-cells was measured by gated cells based on their citrin expression levels. [Figure 7A] This figure shows that the recruitment assay measures gene silencing by a lentiviral rTetR domain fusion accompanied by a fluorescent reporter. A schematic diagram of the lentiviral vector is also shown. [Figure 7B] This figure shows that the recruitment assay measures gene silencing by a lentiviral rTetR domain fusion accompanied by a fluorescent reporter. This is a pilot study in K562 reporter cells, showing a time-course citrin-off:on FACS histogram for ZNF10KRAB cloned into pJT050. 1000 ng / ml of doxycycline was added on day 0 and removed on day 5. [Figure 7C] This figure shows that the recruitment assay measures gene silencing by a lentiviral rTetR domain fusion accompanied by a fluorescent reporter. The percentage of on-cells over time is shown. [Figure 7D] This figure shows that the recruitment assay measures gene silencing by a lentiviral rTetR domain fusion accompanied by a fluorescent reporter. The reporter system was also established in HEK293T cells. Cells were transfected with plasmids encoding rTetR-KRAB or pOri controls and treated with or without 1000 ng / ml of doxycycline for 2 days (top) and 4 days (bottom) before analysis by flow cytometry. [Figure 8A] This figure shows high-throughput measurements of domain expression using FLAG staining, preparative sampling, and sequencing. It is a schematic diagram of a high-throughput method for measuring the expression level of each domain fusion within the library. [Figure 8B] This figure shows high-throughput measurements of domain expression using FLAG staining, preparative sampling, and sequencing. It represents the reproducibility of the domain expression measurements. [Figure 8C] This figure shows high-throughput measurements of domain expression using FLAG staining, preparative sampling, and sequencing. Verification was performed using Western blotting. [Figure 8D] This figure shows high-throughput measurements of domain expression by FLAG staining, preparative sampling, and sequencing. Stability of sublibraries: Randomization is destabilized, but tiled expression is similar to that of the Pfam domain. [Figure 8E] This figure shows high-throughput measurements of domain expression by FLAG staining, preparative sampling, and sequencing. Stability is related to the net charge of residues and residues classified as damage-promoting. [Figure 9A] This figure shows a screen illustrating the inhibitory function of the Pfam domain. The images also show flow cytometry of the cell library before and after magnetic separation. [Figure 9B] This figure shows a screen regarding the inhibitory function of the Pfam domain. It represents PANTHER protein class enrichment for the top 10 stability inhibitors regulated by logP, compared to transient inhibitors. [Figure 9C] This diagram shows a screen regarding the inhibitory function of the Pfam domain. It is a complete list of the domain family, ranked by inhibitory strength on day 5. [Figure 9D]This diagram shows a screen regarding the inhibitory function of the Pfam domain. The rTetR-SUMO fusion silences the reporter. Mutations in the SUMO conjugation site (GG91AA) reduce the silencing rate, and mutations in the SUMO-interacting non-covalent binding site reduce silencing memory. [Figure 9E] This diagram shows a screen regarding the inhibitory function of the Pfam domain. It examines the functionally unknown domain (DUF) that exhibits inhibitory activity. [Figure 10A] This figure shows a deep mutation scan for KRAB. It presents off-:on scores for a deep mutation library of ZNF10-derived KRAB domains at days 5, 9, and 13, using two consecutive biological replicates. [Figure 10B] This figure shows a deep-depth mutation scan for KRAB. FLAG tag staining for KRAB mutant expression levels: non-silencing mutants are degraded. B-box mutants are stable. [Figure 10C] This figure shows a deep-depth mutation scan for KRAB. FLAG tag staining correlates with Western blotting for FLAG tags. [Figure 11A] This figure shows the screen data for the activator. This is a pilot study in which rTetR-VP64 is electroporated into K562minCMV reporter cells. After the addition of doxycycline, the reporter cells are turned on, as measured by flow cytometry for citrin expression. [Figure 11B] This figure shows the screen data for activators. It represents the magnetic separation of the pooled library during activator screening, analyzed by flow cytometry. [Figure 11C]This figure shows screen data for activators. It compares measurements of transcriptional regulation by high-throughput recruitment using a Pfam domain library with two different reporter promoters. Each domain is represented by a dot, and the size of the dot is the expression quartile measured in the FLAG screen. [Figure 12A] This figure shows hundreds of repressors found in a screen of thousands of Pfam domains. It is a box plot of the top repression domain families, ranked by the maximum repression intensity on day 5 for domains within any given family. The straight line represents the median, the whiskers represent the high and low quartiles within a range of 1.5 times the interquartile range, and outliers are indicated by diamonds. The dashed line represents the hit threshold. Boxes are colored for domain families identified in the text. [Figure 12B] This figure shows hundreds of inhibitory factors discovered in a screen of thousands of Pfam domains. Individual validations of the RYBP domain and two functionally unknown domains (DUFs) with inhibitory activity were measured by flow cytometry. The distribution of untreated cells is shown in light gray, and doxycycline-treated cells are shown in color, with two sets of independently transduced biological replicates for each condition. Vertical lines indicate citrine gates used to determine the percentage of off-cells. [Figure 12C] This figure shows hundreds of repressors discovered in a screening of thousands of Pfam domains. The validation time course fitted by a gene silencing model is exponential silencing followed by exponential reactivation at rate ks. Doxycycline (1000 ng / ml) was added on day 0 and removed on day 5 (N=2 for biological replicates). The percentage of mCherry-positive cells with the citrin reporter off was determined by flow cytometry as shown in Figure 12B and normalized for background silencing using untreated time-matched controls. [Figure 12D]This figure shows hundreds of suppressors found in a screen of thousands of Pfam domains. The correlation between high-throughput measurements on day 5 and silencing rates ks is shown (R²=0.86, n=15 domains, N=2-3 biological replicates). The horizontal error bars are the standard deviation for the fitted rates, the vertical error bars are the range of biological replicates in the screen, and the dashed lines are the 95% confidence intervals for the linear regression. [Figure 13A] This figure shows that the repression strength of Hox homeodomains is collinear with the organization of Hox genes and is associated with positive charge. The ranking of homeobox gene classes is based on the median repression strength of these homeodomains on day 5. The horizontal lines indicate the hit threshold. All five homeodomains derived from the CERS class were poorly expressed. [Figure 13B] This figure shows that the repression strength of Hox homeodomains is collinear with the organization of Hox genes and is associated with positive charge. These are homeodomains derived from the Hox gene family. (Top) Hox gene expression patterns along the anterior-posterior axis are colored according to the number of Hox paralogs on the fitted embryo image (Hueber et al., 2010). Both Hox11 and Hox12 are expressed along the proximal-distal axis at the posterior end of the limbs (Wellik and Capecchi, 2003). (Middle) Repression strength after doxycycline treatment for 5 days. Dots are colored according to Hox clusters, and the number of paralogs is colored as shown in the embryo schematic diagram. Spearman's Rhoe and p-values were calculated for all Hox genes regarding the relationship between the number of paralogs and repression strength. (Bottom) Colored arrows represent genes found within four human Hox clusters, indicating the direction of Hox gene transcription from 5' to 3'. Graeber categorizes gene sequence similarities according to existing classifications (Hueber et al., 2010). [Figure 13C]This figure shows that the repression strength of the Hox homeodomain is collinear with the Hox gene's structure and is associated with positive charge. It presents multiple sequence alignments of the Hox homeodomain, ranked (by off:on ratio on day 5) with the strongest repressors, highlighted in red, being the most potent. Other base residues within the N-terminal arm are colored lavender. [Figure 13D] This figure shows that the repression strength of Hox homeodomains is collinear with the organization of Hox genes and is associated with positive charge. The correlation between the number of positively charged residues in the N-terminal arm upstream of helix 1 of each Hox homeodomain and the mean repression on day 5 is shown. The color of the dots indicates the number of paralogs. [Figure 13E] This figure shows that the repression strength of the Hox homeodomain is related to the organization and collinearity of the Hox gene and to its positive charge. The NMR structure of the HOXA13 homeodomain, retrieved from PDB ID: 2L7Z, with the RKKR motif highlighted in red, is shown. The sequence from G15 to S81 is shown, using coordinates derived from multiple sequence alignments. [Figure 14A] This figure shows the discovery of the activation domain. It is a schematic diagram of an activation reporter that uses a weak minCMV promoter that can be activated by doxycycline-mediated recruitment, activating the effector domain fused to rTetR. [Figure 14B]This figure shows the discovery of activation domains. It demonstrates the reproducibility of high-throughput activator measurement from two independently transduced biological replicates. The nuclear domain library was transduced into a pool of cells containing the activation reporter shown in Figure 14A. The cell pool was treated with doxycycline for 48 hours, and on-cells and off-cells were magnetically separated, and the domains were sequenced. The ratio of sequencing reads from off-cells to sequencing reads from on-cells is shown for well-expressed domains. Annotated Pfam activation domain families (FOXO-TAD, Myb LMSTEN, TORC_C) are colored with red shading. A straight line is drawn in relation to the KRAB domain derived from ZNF473, which is the strongest hit. The hit threshold is a dashed line drawn two standard deviations below the mean of the poorly expressed domain distribution. [Figure 14C] This figure shows the discovery of activating domains. It is a ranked list of domain families with at least one activating hit. Within Pfam, families already annotated as activators are shown in red. The dashed line represents the hit threshold, as shown in Figure 14B. Only domains with good expression are shown. [Figure 14D] This figure shows the discovery of the activating domain. The acidity of the effector domains derived from the Pfam library is calculated as the net charge per amino acid. (Left) Comparison of well-expressed Pfam domains (excluding KRAB and annotation activators) with activating hit domains. The annotated Pfam activating domain family is shown as the positive control group (orange). (Right) Comparison of activating hit domains with non-hit domains derived from the KRAB domain family. The P-values from the Mann-Whitney test are shown by the bars between the comparison groups. ns = not significant (p>0.05). [Figure 14E]This figure shows the discovery of the activating domain. The phylogenetic tree for all well-expressed KRAB domains, including the sequence-branched mutant KRAB cluster, is shown in green (top). High-throughput recruitment measurements for repression on day 5 are shown in blue (middle), and measurements for activation are shown in red (bottom). The horizontal dashed line indicates the hit threshold. Examples for repressive KRAB derived from ZNF10, repressive KRAB_1 derived from ZFP28, and all activating KRAB domains are called out in larger font. The starting point of the KRAB domain is indicated in parentheses. [Figure 14F] This figure shows the discovery of the activation domain. This is a separate validation of the mutant KRAB activation domain. The rTetR(SE-G72P)-domain fusion was delivered to K562 reporter cells by lentivirus, selected by blastosidine, and treated with 1000 ng / ml doxycycline for 2 days. Citrin reporter levels were then measured by flow cytometry. The distribution of untreated cells is shown in light gray, and doxycycline-treated cells are shown in color, with two sets of independently transduced biological replicates for each condition. The vertical lines indicate the citrin gate used to determine the percentage of on cells, and show the average percentage of on cells for doxycycline-treated cells. [Figure 14G]This figure shows the discovery of the activation domain. The distance from the nearest peak of H3K27ac, the KRAB zinc finger protein activity chromatin mark, to the ChIP peak position. KRAB proteins are classified as hits (blue) or no hits (green) on the repressor screen on day 5, depending on their status (left). In addition, data are shown individually for ZNF10 (black) containing repressive hit KRAB, ZNF473 (red) containing activating hit KRAB, and ZFP28 (yellow) (right) containing both activating and repressive hit KRAB. Each dot indicates the percentage of peaks in a 40-base pair bin. ChIP-seq and ChIP-exo data were retrieved from ENCODE Project Consortium et al., 2020; Imbeault et al., 2017; Najafabadi et al., 2015; Schmitges et al., 2016. In the aggregated data, only the single peaks bound to a single KRAB zinc finger are included (blue and green dots in the left figure). However, because the number of single peaks for each individual protein is small, all peaks are included for each individual protein (red, black, and yellow dots in the right figure). [Figure 15A] This figure shows compact repression domains discovered within nuclear proteins. It is a schematic diagram of an 80-amino acid tiling library covering a curated set of 238 nuclear-localized proteins. These tiles were fused with rTetR and recruited to reporters using the same workflow as in Figure 1 to measure repression intensity. [Figure 15B] This figure shows compact repression domains discovered within nucleoproteins. Each tile represents a tiling gene ranked by its maximum repression function on day 5, indicated by a dot. Hits are tiles where log2 (off:on) is ≥2 standard deviations above the mean of the negative control. Genes that were hits are colored with grayscale, while genes that were not hits are colored gray. [Figure 15C]This figure shows compact repression domains (CTCFs) found within nucleoproteins. The tiling is of CTCFs. The schematic diagram shows protein annotations retrieved from UniProt. Horizontal bars indicate the region covered by each tile, and vertical error bars indicate the standard error for the screen based on two sets of biological replicates. The strongest hit tile is highlighted with vertical gradation and annotated as a repression domain (orange). [Figure 15D] This figure shows a compact repression domain discovered within a nucleoprotein. It is a tiling of BAZ2A (also known as TIP5). [Figure 15E] This figure shows compact repressive domains discovered within nucleoproteins. Individual validations were performed. Lentivirus rTetR(SE-G72P)-tile fusions were delivered to K562 reporter cells, which were treated with 100 ng / ml doxycycline for 5 days (between the vertical dashed lines), and then the doxycycline was removed. Cells were analyzed by flow cytometry to determine the percentage of cells with the citrin reporter off, and the data were fitted using a gene silencing model (N=2 for biological replicates). Two KRAB repressive domains are shown as positive controls. Tiling screen data corresponding to the validations shown below (blue curves) are shown in Figure 22. [Figure 15F] This figure shows compact repression domains discovered within nucleoproteins. It is a tiling of MGAs. The two repression domains are found outside the existing annotation region and are shown as repressors 1 and 2 (dark red and purple, respectively). Minimized repression regions in the overlapping hit tile areas are highlighted with narrow vertical red tones. [Figure 15G] This figure shows the compact repression domain discovered within the nucleoprotein. The strongest repression tile, derived from two peaks within MGA, was individually examined using the method described in Figure 15E (N=2 for biological replicates). [Figure 15H]This figure shows the compact repression domains discovered within the nucleoprotein. The sequence of MGA repressor 1 was minimized and shaded red by selecting the region shared between all hit tiles within the peaks shown between the vertical dashed lines. The ConSurf score for protein sequence conservation is shown below by an orange line, and the confidence interval (25th to 75th percentile of the estimated evolutionary rate distribution) is shown in gray. Asterisks mark residues predicted by ConSurf to be functional (highly conserved and exposed). The sequence of repressor 2 was minimized using the same method, and the region overlapping with the predicted functional residue was also minimized (data not shown). [Figure 15I] This figure shows the compact repression domain discovered within the nucleoprotein. The MGA effector was minimized to a 10-30 amino acid subtile, cloned as a lentiviral rTetR(SE-G72P)-tile fusion as shown in Figure 15H, and delivered to K562 reporter cells. After selection, the cells were treated with 100 or 1000 ng / ml doxycycline for 5 days, and the percentage of cells silenced by the citrin reporter was measured by flow cytometry (N=2 for biological replicates). [Figure 16A] This figure illustrates the validation of a dual reporter for lentiviral recruitment assays and gene silencing. It is a schematic diagram of a lentiviral recruitment vector with a Golden Gate cloning site for creating a fusion of the effector domain with rTetR, a doxycycline-inducible DNA-binding domain. The constitutive pEF promoter drives the expression of the rTetR-effector fusion, separated by a T2A self-cleaving peptide, and mCherry-BSD (blasticidin S deaminase resistance gene). [Figure 16B]This figure shows the validation of a dual reporter for lentiviral recruitment assays and gene silencing. (Top) A schematic diagram of the recruitment of the rTetR-KRAB fusion to the dual reporter gene. The reporter is incorporated into the AAVS1 locus via TALEN-mediated homology-directed repair, and the PuroR resistance gene is driven by the endogenous AAVS1 promoter. The dual reporter consists of a synthetic surface marker (Igκ-hIgG1-Fc-PDGFRβ) and a citrin fluorescent protein. (Bottom) A pilot study in K562 reporter cells. Reporter cells were created by TALEN-mediated homology-directed repair, incorporating the reporter into the AAVS1 locus, and then selected with puromycin. Next, cells were spin-impacted with lentivirus to deliver rTetR-KRAB, and then either left untreated or treated with 1000 ng / ml doxycycline to induce binding of rTetR to DNA at the TetO site. The distribution of untreated cells is shown in light gray, and doxycycline-treated cells are shown in black or orange, with independently transduced, two-chain biological repeats under each condition. Lentivirus-treated cells were gated for mCherry as a delivery marker. A KRAB domain derived from human ZNF10 was used. [Figure 16C] This figure illustrates the validation of a dual reporter for lentiviral recruitment assays and gene silencing. It demonstrates magnetic separation of off-cells from on-cells using ProG Dynabeads bound to a synthetic surface marker. Ten million cells were subjected to magnetic separation using 30 μl of beads, and citrin reporter expression was measured by flow cytometry before and after separation. An example of magnetic separation of mixed on-cells and off-cells is shown on the right. [Figure 17A]This figure shows high-throughput measurements of domain expression by FLAG staining, parsing, and sequencing. (Top) A schematic diagram of a high-throughput strategy for measuring the expression level of each domain in the library. Using their native protein sequences, domains shorter than 80 amino acids were extended on both sides to reach 80 amino acids, thereby making all synthetic library elements the same length. (Middle) The library was cloned into a FLAG-tagged construct and delivered to K562 cells by lentivirus at a low infection multiplicity so that the majority of cells expressed a single library member. The mCherry-BSD fusion protein allows for selection by blastosiding and a fluorescent marker for delivery and selection efficiency without using a second 2A component. (Bottom) Cells were stained with anti-FLAG, high-expression and low-expression populations were parsed, domains were sequenced, and expression was measured by calculating the log2(FLAGhigh:FLAGlow) ratio. [Figure 17B] This figure shows high-throughput measurements of domain expression by FLAG staining, fractionation, and sequencing. It shows the distribution of FLAG staining levels, measured by flow cytometry, before and after fractionation into two bins (N=2 of the biological replicates in the cell library, indicated by shading in overlapping regions). [Figure 17C] This figure shows high-throughput measurements of domain expression by FLAG staining, preparative sampling, and sequencing. It represents the reproducibility of biological replicates derived from the domain expression screen (r² = 0.82). Domains with good expression exceeding the threshold (dashed line, one standard deviation above the median of the random controls) were selected for further analysis in the transcriptional regulatory screen. [Figure 17D]This figure shows high-throughput measurements of domain expression by FLAG staining, preparative sampling, and sequencing. It validates the expression levels of a panel of KRAB domains. Individual rTetR-3×FLAG-KRAB constructs were delivered to K562 cells via lentivirus. Cells were selected by blastoscientin, and >80% were confirmed to be mCherry-positive by flow cytometry. Expression levels were measured by Western blotting with anti-FLAG antibody. Anti-histone H3 was used as a loading control for normalization. Levels were quantified using ImageJ. [Figure 17E] This figure shows high-throughput measurements of domain expression by FLAG staining, preparative sampling, and sequencing. The high-throughput expression measurements are compared to protein levels determined by Western blotting. These six KRAB domains were cloned individually using precise 80-amino acid sequences derived from the Pfam domain library. [Figure 17F] This figure shows high-throughput measurements of domain expression by FLAG staining, preparative sampling, and sequencing. It shows the distribution of expression levels for different categories of library members. Random controls show poor expression (p < 1 × 10⁻⁵ by Mann-Whitney U test) compared to tiles across DMD proteins or Pfam domains. The dashed lines indicate thresholds for expression levels, as shown in Figure 17C. [Figure 18A] This figure shows the identification of domains with inhibitory function. Flow cytometry shows the distribution of citrin reporter levels in a pool of cells expressing the Pfam domain library before and after magnetic separation using a ProG DynaBead bound to a synthetic surface marker. A duplicate histogram is shown for two sets of biological repeats. The mean percentage of off-cells is shown to the left of the vertical line indicating the citrin level gate. 1000 ng / ml of doxycycline was added on day 0 and removed on day 5. [Figure 18B]This figure shows the identification of domains with repressive function. The enrichment by the PANTHER protein class is for nuclear proteins containing repressive domains that have strong or weak memory compared to the background set of all nuclear proteins in which the domain was incorporated into the library. [Figure 18C] This figure shows the identification of domains with suppressive function. The validation time course was fitted using the rTetR-SUMO gene silencing model. An 80-amino acid sequence, centered around the vicinity of the Rad60-SLD domain and trimming domain of SUMO3, was individually cloned into a lentivirus and delivered to reporter cells. 1000 ng / ml of doxycycline was added on day 0 and removed on day 5 (N=2 for biological replicates). The percentage of mCherry-positive cells with the citrin reporter off was determined by flow cytometry and normalized for background silencing using an untreated time-matched control. [Figure 18D] This figure shows the identification of domains with inhibitory function. It is a verification of the MPP8 Chromo domain, a member of the HUSH complex, using the complete 80-amino acid sequence and the sequence trimmed to match the Pfam and UniProt annotations used in the screen. [Figure 18E] This figure shows the identification of a domain with inhibitory function. It is a validation of the CBX1 Chromoshadow domain using a 52-amino acid sequence trimmed to match the Pfam annotation. [Figure 18F] This figure shows the identification of a domain with inhibitory function. It is a verification of the SCMH1 SAM1 domain (also known as SPM), a component of Polycomb 1, using a 65-amino acid sequence trimmed to match Pfam annotations. [Figure 18G]This figure shows the identification of the domain with inhibitory function. It is a validation of the HERC2 Cyt-b5 domain using the complete 80-amino acid sequence used in the screen and a 72-amino acid sequence trimmed to match the Pfam annotation. [Figure 18H] This figure shows the identification of domains with inhibitory functions. This is a verification of the BIN1 SH3_9 domain. [Figure 18I] This figure shows the identification of a domain with inhibitory function. It is a verification of the PCGF2 zf-C3HC4_2 domain, a component of Polycomb 1, using a 39-amino acid sequence trimmed to match the Pfam annotation. [Figure 18J] This figure shows the identification of domains with inhibitory function. It is a validation of the TOX HMG box domain using the complete 80-amino acid sequence used in the screen and a 68-amino acid sequence trimmed to match the Pfam annotation. [Figure 18K] This figure shows the identification of domains with inhibitory function. This involves verification of a random sequence of 80 amino acids that functions as an inhibitor. [Figure 19A] This figure shows that rTetR(SE-G72P) reduces the leakage of KRAB silencing in human cells. Silencing by the rTetR-KRAB fusion shows leakage of silencing without doxycycline treatment for a subset of the KRAB domain (dark gray bars). On day 0, the construct was delivered to reporter cells by lentivirus. Between days 3 and 11, the cells were selected by blastosepticin. On day 11, the cells were divided into doxycycline-treated and untreated conditions. On day 16, reporter levels were measured by flow cytometry. The results after gateding for mCherry-positive cells are shown. The KRAB domains were selected from three categories based on their measurements on the screen and are shown on the right. The bars represent the mean, and the error bars represent the standard deviation (N=3 of independently transduced biological replicates). [Figure 19B]This figure shows that rTetR(SE-G72P) reduces leakage of KRAB silencing in human cells. Leakage can be reduced by using rTetR(SE-G72P) or by introducing a 3×FLAG between rTetR and the KRAB domain derived from ZNF823. On day 0, the construct was delivered to reporter cells by lentivirus, on day 4 the cells were divided into doxycycline-treated and untreated conditions, and on day 7 reporter levels were measured by flow cytometry. The results after gateding for mCherry-positive cells are shown. A non-leaking KRAB domain derived from ZNF140 was used as a control. Bars represent the mean, and error bars represent the standard deviation (N=2 of independently transduced biological replicates). [Figure 19C] This figure shows that rTetR(SE-G72P) reduces the leakage of KRAB silencing in human cells. K562 reporter cell lines, which stabilize lentiviral expression of a leaky KRAB domain derived from ZNF823 or a non-leakage-suppressing KRAB domain derived from ZNF140 (cloned as rTetR or a fusion of rTetR(SE-G72P)), were treated with variable doses of doxycycline. After 4 days, reporter levels were measured by flow cytometry and represent the percentage of mCherry-positive cells with the citrin reporter off (N=2 for independently transduced biological replicates). Dose-response was fitted using least squares with a nonlinear variable gradient S-shaped curve using PRISM statistical analysis software. [Figure 19D]This figure shows that rTetR(SE-G72P) reduces the leakage of KRAB silencing in human cells. The silencing / memory dynamics are for all individual validations of the KRAB domain, fitted by a gene silencing model. The rTetR(SE-G72P)-KRAB fusion was delivered to K562 reporter cells via lentivirus, selected by blastosidine, and then 10 ng / ml doxycycline was added on day 0 and removed on day 5 (N=2 for biological replicates). The percentage of mCherry-positive cells with the citrin reporter off was determined by flow cytometry and normalized for background silencing using an untreated time-matched control. 10 ng / ml doxycycline was used to operate within a dynamic range that facilitates measuring differences in silencing / memory capacity between fast KRAB silencing domains. When doxycycline was administered at 1000 ng / ml, all hit-suppressing KRAB domains (green and orange) completely silenced the reporter within 5 days, and their kinetics were indistinguishable (data not shown). In particular, the KRAB (orange) that was leaky on rTetR did not exhibit significantly different memory kinetics when fused with rTetR(SE-G72P) compared to the KRAB (green) that was not leaky on rTetR. Importantly, none of the rTetR(SE-G72P)-KRAB fusions showed significant leakage of silencing under untreated conditions. [Figure 20A] This figure shows a deep mutation scan for ZNF10 KRAB used in CRISPRi. Flow cytometry shows citrin reporter levels in pooled cells of the KRAB library before and after magnetic separation using ProG DynaBeads bound to synthetic surface markers. A duplicate histogram is shown for two sets of biological repeats. The mean percentage of off-cells is shown to the left of the vertical line indicating the citrin level gate. [Figure 20B]This figure shows the deep-depth mutation scans for ZNF10 KRAB used in CRISPRi. The off:on scores are from two consecutive biological replicates of the deep-depth mutation scan library for the ZNF10 KRAB domain at days 5, 9, and 13. Cells were treated with 1000 ng / ml doxycycline for the first 5 days. The gray diagonal line indicates that the mean log2 (off:on) is the median for the WT domain (black dots). The black diagonal line shows the fitted linear model. [Figure 20C] This figure shows a high-depth mutation scan of ZNF10 KRAB used in CRISPRi. It shows the alignment of human ZNF10 KRAB with mouse KRAB used in the NMR structure (PDB:1v65) and KRAB-O used in the recombinant protein binding assay (Peng et al., 2009). The ordered region is used in Figure 3, and the alignment region containing all 12 necessary residues is used in Figure 20D. Residues necessary for silencing on day 5 are colored red within the ZNF10 sequence and the PDB:1v65 sequence. Residues necessary for binding to recombinant KAP1 are colored red, and residues unnecessary for binding to recombinant KAP1 are colored gray within the KRAB-O sequence, summarizing previously published results (Peng et al., 2009). [Figure 20D] This figure shows a high-depth mutation scan of ZNF10 KRAB used in CRISPRi. It is an aggregate of 20 states of the NMR structure of KRAB (PDB:1v65). Residues necessary for silencing on day 5 are colored red. [Figure 20E]This figure shows a deep mutation scan of ZNF10 KRAB used in CRISPRi. It shows the silencing / memory dynamics for all individual validations of KRAB ZNF10 mutants fitted to a gene silencing model. (Top) rTetR-KRAB fusions were delivered to K562 reporter cells via lentivirus, selected by blastosidine, and then 1000 ng / ml doxycycline was added on day 0 and removed on day 5 (N=2 biological replicates). (Bottom) rTetR(SE-G72P)-KRAB fusions were delivered to K562 reporter cells via lentivirus, selected by blastosidine, and then 10 ng / ml doxycycline was added on day 0 and removed on day 5 (N=2 biological replicates). Columns describe the mutant location within the KRAB domain and its impact on effector function. The percentage of mCherry-positive cells with the citrin reporter off was determined by flow cytometry and normalized for background silencing using untreated time-matched controls. All rTetR(SE-G72P)-KRAB fusions were also measured over 5 days of treatment with 1000 ng / ml doxycycline, but the results were indistinguishable from those with rTetR, and all KRAB variants completely silenced the reporter, with the exception of the EEW25AAA variant which did not (data not shown). [Figure 20F] This figure shows a high-depth mutation scan of ZNF10 KRAB used in CRISPRi. It correlates the expression level of the rTetR-KRAB fusion, as measured by the Pfam domain library, with the silencing score on day 13. Only KRAB domains that have been shown to interact with the repressive cofactor KAP1 by IP / MS (Helleboid et al., 2019) were included. [Figure 20G]This figure shows a deep mutation scan of ZNF10 KRAB used in CRISPRi. It correlates amino acid frequencies across the library and control for the Pfam domain with domain expression levels (Pearson's r-value is shown). [Figure 20H] This figure shows a deep-depth mutation scan for ZNF10 KRAB used in CRISPRi. It is a Western blot of FLAG-tagged rTetR-KRAB fusions after lentiviral delivery to K562. Cells were selected for delivery by blastoscidin, and >80% were confirmed to be mCherry-positive by flow cytometry. Expression levels were quantified using ImageJ compared to H3-loaded controls. [Figure 21A] This figure shows that high-throughput recruitment to the minimal promoter leads to the discovery of the activation domain. Flow cytometry of the Pfam domain in the activation reporter cell against a pooled library before and after magnetic separation is shown. On-cell percentages are shown to the right of the citrin-level gate, depicted by vertical lines. One to two sets of biological repeats are shown with shading in the overlapping regions. [Figure 21B] This figure shows that high-throughput recruitment to the minimum promoter leads to the discovery of the activation domain. The GO term enrichment is for genes containing the hit activation domain, compared to a background set of all proteins containing the well-expressed domain within the library after filtering for counts. Raw p-values are shown, but all GO terms shown had false-find rates below 10%. [Figure 21C]This figure shows that high-throughput recruitment to the minimal promoter leads to the discovery of the activating domain. Individual validation of the activating domain is performed. The rTetR(SE-G72P)-domain fusion was delivered to K562 reporter cells via lentivirus and selected by blastosidine. Cells were treated with 1000 ng / ml doxycycline for two days, and citrin reporter levels were measured by flow cytometry. The distribution of untreated cells is shown in light gray, and doxycycline-treated cells are shown in color, with two sets of independently transduced biological replicates for each condition. The vertical lines indicate the citrin gate used to determine the percentage of on cells, and show the mean percentage of on cells for doxycycline-treated cells. VP64 is the positive control. Each domain was examined as an 80-amino acid sequence, which was either a library sequence or a sequence extended from a trimmed Pfam annotation domain sequence, with the exception of Med9 and DUF3446, which had the smallest extension due to the Pfam annotation region being 75-69 amino acids. The corresponding results for the 80-amino acid library sequences for the KRAB domain are shown in Figure 14. [Figure 22A] This figure shows the identification of compact repression domains within nuclear proteins by tiling screen. Flow cytometry shows the distribution of citrin reporter levels in a pool of cells expressing the tiling library before and after magnetic separation using a ProG DynaBead bound to a synthetic surface marker. A duplicate histogram is shown for two sets of biological repeats. The mean percentage of off-cells is shown to the left of the vertical line indicating the citrin level gate. 1000 ng / ml of doxycycline was added on day 0 and removed on day 5. [Figure 22B]This figure shows the identification of compact repression domains within nuclear proteins using tiling screens. The figures represent high-throughput recruitment measurements from two consecutive biological replicates of a nuclear protein tiling library on day 5 after doxycycline treatment and day 13, 8 days after doxycycline removal. The hit threshold is defined as a value two standard deviations above the mean of the random control and DMD tiling control. [Figure 22C] This figure shows the identification of compact repression domains within nuclear proteins using a tiling screen. The tiling results are for ZNF57 and ZNF461, zinc finger proteins of KRAB. Each bar represents an 80-amino acid tile, and the vertical error bars represent the range due to two sets of biological repeats. Protein annotations are sourced from UniProt. [Figure 22D] This figure shows the identification of compact repression domains within nuclear proteins using a tiling screen. The tiling is of RYBP. The schematic diagram shows the protein annotations, retrieved using the UniProt IDs listed above. The vertical error bars indicate the standard error for two sets of biological replicates. [Figure 22E] This figure shows the identification of compact repression domains within nuclear proteins using a tiling screen. It is a tiling of REST. [Figure 22F] This figure shows the identification of compact repression domains within nuclear proteins using a tiling screen. The image is a tiling of CBX7. [Figure 22G] This figure shows the identification of compact repression domains within nuclear proteins using a tiling screen. The image is a tiling of DNMT3B. [Figure 22H]This figure shows the identification of compact repression domains within nuclear proteins using a tiling screen. (Top) Tiling of DMD. (Bottom) Dynamics of silencing and memory after recruitment of DMD hit tiles. Cells were treated with 1000 ng / ml doxycycline for the first 5 days, and citrin reporter levels were measured by flow cytometry. Off-cell percentages were normalized to account for background silencing, and the data (dots) were fitted to a gene silencing model (curve) (N=2 for biological replicates). [Modes for carrying out the invention]
[0027] A system and method are provided for generating a catalog of compact transcription effector domains. Furthermore, in some embodiments, this catalog of domains is fused to DNA-binding domains to manipulate synthetic transcription factors. These are used to implement targeted, tunable regulation of gene expression within eukaryotic (or other) cells. This technology utilizes a high-throughput platform to screen and characterize tens of thousands of synthetic transcription factors within cells. These synthetic transcription factors are fusions of DNA-binding domains and transcription effector domains. The system is used to generate hundreds of short effector domains (e.g., 80 amino acids) and then perform high-throughput steps to further shorten them to minimally sufficient sequences (e.g., 10 amino acids), which is advantageous for delivery (e.g., packaging within viral vectors). Targeting these fusions results in localized negative or positive regulation of mRNA transcription, depending on the effector domain. Some of these synthetic transcription factors mediate long-term epigenetic regulation that persists even after the factor itself has been released from its target.
[0028] Previously, a limited number of transcription effector domains were available for manipulating synthetic transcription factors. To address this limitation, this specification provides a high-throughput method for screening and quantifying the function of transcription effector domains. This method has enabled the discovery of hundreds of effector domains that, when fused to DNA-binding domains, may upregulate or downregulate transcription in a targeted manner. This step can also be used to identify mutants of effector domains with enhanced activity. These effector domains can be used to manipulate synthetic transcription factors for applications in gene therapy and cell therapy, synthetic biology, and functional genomics.
[0029] Exemplary applications include, but are not limited to, the following: Targeting of endogenous gene repression / activation by fusion of programmable DNA-binding domains (e.g., dCas9, dCas12a, zinc finger, TALE) with transcription effector domains; Gene therapy and cell therapy (e.g., silencing pathogenic transcripts in patients) or research; Synthetic transcription factors are used to simultaneously perturb the expression of multiple genes (for example, to perform high-throughput gene interaction mapping using multiple guide RNAs via CRISPRi / a screening); The use of synthetic transcription factors within genetic circuits, such as inducible gene expression or more complex circuits. These circuits are used in gene therapy (e.g., antibody delivery via AAV) and cell therapy (e.g., manipulation of CAR-T cells ex vivo) to achieve therapeutic gene expression outputs in response to small molecule inputs from the environment.
[0030] The novel transcription effector domains provided herein offer several advantages for applications relying on synthetic transcription factors. Short-chain domains (e.g., ≤80 amino acids) have been identified, and high-throughput steps have been developed to further shorten them to minimally sufficient sequences, which is advantageous for delivery (e.g., packaging in viral vectors). In some cases, potent effector domains as short as 10 amino acids have been identified. In some embodiments, the domains are extracted from human proteins, which offers the advantage of reduced immunogenicity compared to viral effector domains. The majority of the generated domains have not yet been reported as transcription effectors. In addition, high-throughput processes are also provided for examining mutations within these domains to identify enhancing variants. High-throughput methods are more readily supported by the development of artificial cell surface markers, which result in more efficient, inexpensive, and rapid screening of these libraries using magnetic separation. This is an advantage over the more conventional method of fractionating libraries based on the expression of fluorescent reporter genes.
[0031] The identified domain collection is vast and diverse, and the platform makes it easy to investigate new combinations of domains as fusions in high throughput, so as to create synthetic transcription factors with novel properties (e.g., compositions of two repressive domains that achieve a combination of fast silencing and permanent silencing).
[0032] Hundreds of uncharacterized or unknown effector domains, capable of silencing or activating transcription, can be fused to DNA-binding domains. For example, in human cells, lentiviral screening provides a high-throughput method for screening single domains and domain pairs. High-throughput methods are made easier by the development of artificial cell surface markers, which use magnetic separation, resulting in more efficient, inexpensive, and rapid screening of these libraries.
[0033] 1.Definition As used herein, the terms “comprise,” “include,” “have,” “may have,” “may contain,” and their variations are intended to be open-ended transitional phrases, terms, or words that do not exclude the possibility of further actions or structures. Unless otherwise indicated by context, the singular “a,” “an,” and “it” refer to multiple objects. Whether expressly or otherwise, this disclosure also assumes other aspects that “include,” “consist of,” and “essentially consist of” the aspects or elements presented herein.
[0034] For the purposes of enumerating numerical ranges in this specification, each number intervening between them is assumed to be of a similar degree of precision. For example, in the range of 6 to 9, in addition to 6 and 9, the numbers 7 and 8 are also assumed, and in the range of 6.0 to 7.0, the numbers 6.0, 6.1, 6.2, 6.3, 6.4, 6.5, 6.6, 6.7, 6.8, 6.9 and 7.0 are also explicitly assumed.
[0035] Unless otherwise specified herein, scientific and technical terms used herein shall have the same meaning as those generally understood by those skilled in the art. For example, any terminology used herein in connection with the cell and tissue culture, molecular biology, immunology, genetics, and protein chemistry, nucleic acid chemistry, and hybridization techniques described herein is well known and commonly used in the art. The meaning and scope of terms shall be clear, but in the event of any potential ambiguity, the definitions provided herein shall prevail over any dictionary or external definitions. Furthermore, unless otherwise required by context, singular terms shall include plural forms, and plural terms shall include singular forms.
[0036] As used herein, the term “antibody” refers to a protein endogenously used by the immune system to identify and neutralize foreign substances such as bacteria and viruses. Typically, an antibody is a protein containing at least one complementarity-determining region (CDR). The CDR forms the “hypervariable region” of the antibody (discussed further below), which is responsible for binding to the antigen. A total antibody typically consists of four polypeptides: two identical copies of a heavy (H) chain polypeptide and two identical copies of a light (L) chain polypeptide. Each of the heavy chains has one N-terminal variable (V) H ) region and three C-terminal steady (C H1 , C H2 and C H3 ) contains a region, and each light chain has one N-terminal variable (V L ) region and one C-terminal steady (C LThe antibody light chains contain a region. Based on the amino acid sequence of their constant domains, the light chains of an antibody can be assigned to one of two distinctly different types, kappa (κ) or lambda (λ). In a typical antibody, each light chain is linked to a heavy chain by a disulfide bond, and two heavy chains are linked to each other by disulfide bonds. The variable region of the light chain is aligned with the variable region of the heavy chain, and the constant region of the light chain is aligned with the first constant region of the heavy chain. The remainder of the constant region of the heavy chain is aligned with each other. The variable regions of each pair of light and heavy chains form the antigen-binding site of the antibody. H Region and V L Each region has the same general structure, containing four framework (FW or FR) regions. As used herein, the term “framework region” refers to a relatively conserved amino acid sequence within the variable region, located between CDRs. Within each variable domain, there are four framework regions, designated FR1, FR2, FR3, and FR4. The framework regions form a β-sheet that provides the structural framework for the variable region (see, for example, CAJaneway et al., “Immunobiology,” 5th edition, Garland Publishing, New York, NY (2001)). The framework regions are connected by three CDRs. As discussed above, the three CDRs, known as CDR1, CDR2, and CDR3, form the “hypervariable region” of the antibody, which contributes to antigen binding. The CDRs connect the β-sheet structures formed by the framework regions, sometimes forming loops that include parts of these structures. While the constant regions of the light and heavy chains do not directly participate in antibody binding to antigens, they can influence the orientation of the variable regions. The constant regions also exhibit diverse effector functions, such as participation in antibody-dependent complement-mediated lysis or antibody-dependent cytotoxicity through interactions with effector molecules and cells.
[0037] As used herein, the terms “antibody fragment,” “antibody fragment,” and “antigen-binding fragment” of an antibody are used interchangeably to refer to fragments of one or more antibodies that retain the ability to specifically bind an antigen (see generally Holliger et al., Nat. Biotech., 23(9):1126-1129 (2005)). Any antigen-binding fragment of an antibody described herein is within the scope of the present invention. Antibody fragments are desired to include, for example, one or more CDRs, variable regions (or portions thereof), constant regions (or portions thereof), or combinations thereof. Examples of antibody fragments include (i) a Fab fragment, which is a monovalent fragment consisting of a V L domain, a V H domain, a C L domain, and a C H1 domain; (ii) an F(ab’)2 fragment, which is a bivalent fragment comprising two Fab fragments linked by a disulfide bridge in the hinge region; (iii) an Fv fragment consisting of the V L domain and the V H domain of a single arm of an antibody; (iv) a Fab’ fragment resulting from cleavage of the disulfide bridge of an F(ab’)2 fragment using mild reducing conditions; (v) a disulfide-stabilized Fv fragment (dsFv); and (vi) a domain antibody (dAb), which is a single-chain variable region domain (V H or V L ) polypeptide of an antibody that specifically binds an antigen, but are not limited thereto.
[0038] As used herein, “nucleic acid” or “nucleic acid sequence” refers to polymers or oligomers of pyrimidine bases and / or purine bases, preferably cytosine, thymine, and uracil, and adenine and guanine, respectively (see Albert L. Lehninger, “Principles of Biochemistry,” 793-800 (Worth Pub., 1982)). This technique assumes any deoxyribonucleotide, ribonucleotide, or peptide nucleic acid components and any chemical variants thereof, such as methylated, hydroxymethylated, or glycosylated forms of these bases. Polymers or oligomers may be heterogeneous or homogeneous in a composition, may be isolated from naturally occurring sources, or may be artificially or synthetically produced. In addition, nucleic acids can be DNA or RNA or mixtures thereof, and may be constitutively or transiently present in single-stranded or double-stranded forms, including homo-double-stranded, hetero-double-stranded, and hybrid states thereof. In some embodiments, the nucleic acid or nucleic acid sequence includes other types of nucleic acid structures, such as DNA / RNA helices, peptide nucleic acids (PNAs), morpholino nucleic acids (see, for example, Braasch and Corey, Biochemistry, 41(14):4503-4510 (2002) and U.S. Patent No. 5,034,506), located nucleic acids (LNAs; see Wahlestedt et al., Proc. Natl. Acad. Sci. USA, 97:5633-5638 (2000)), cyclohexynyl nucleic acids (see Wang, J. Am. Chem. Soc., 122:8595-8602 (2000)), and / or ribozymes.Therefore, the terms “nucleic acid” or “nucleic acid sequence” can also include chains containing non-natural nucleotides, modified nucleotides, and / or non-nucleotide components (e.g., “nucleotide analogs”) that may perform the same function as natural nucleotides; furthermore, as used herein, the term “nucleic acid sequence” refers to oligonucleotides, nucleotides, or polynucleotides and fragments or parts thereof, as well as DNA or RNA of genomic or synthetic origin, which may be single-stranded or double-stranded, and may represent a sense strand or an antisense strand. The terms “nucleic acid,” “polynucleotide,” “nucleotide sequence,” and “nucleotide” are used interchangeably. These terms refer to polymeric forms of nucleotides of any length, which are deoxynucleotides or nucleotides or analogs thereof.
[0039] A "peptide" or "polypeptide" is a linked sequence of two or more amino acids linked by peptide bonds. A peptide or polypeptide may be a natural peptide or polypeptide, a synthetic peptide or polypeptide, a modified peptide or polypeptide, or a combination of a natural peptide or polypeptide and a synthetic peptide or polypeptide. Polypeptides include proteins such as binding proteins, receptors, and antibodies. Proteins can be modified by the addition of sugars, lipids, or other parts not present in the amino acid chain. The terms "polypeptide" and "protein" are used interchangeably herein.
[0040] As used herein, the term “sequence identity percentage” refers to the percentage of nucleotides or nucleotide analogues in a nucleic acid sequence, or amino acids in an amino acid sequence, that are identical to the corresponding nucleotides or amino acids in a reference sequence, after aligning two sequences to achieve the maximum identity percentage and introducing gaps where necessary. Therefore, for nucleic acids longer than the reference sequence that conform to this technique, further nucleotides in the nucleic acid that are not aligned with the reference sequence are not considered for determining sequence identity. Numerous mathematical algorithms for obtaining optimal alignment and calculating identity between two or more sequences are known and incorporated into numerous available software programs. Examples of such programs include CLUSTAL-W, T-Coffee, and ALIGN (for nucleic acid and amino acid sequence alignment), the BLAST program (e.g., BLAST 2.1, BL2SEQ, and their latest versions), and the FASTA program (e.g., FASTA3x, FAS®, and SSEARCH) (for sequence alignment and sequence similarity searching). Sequence alignment algorithms are also disclosed, for example, in Altschul et al., J. Molecular Biol., 215(3):403~410 (1990); Beigert et al., Proc. Natl. Acad. Sci. USA, 106(10):3770~3775 (2009); Durbin et al., eds., "Biological Sequence Analysis: Probabilistic Models of Proteins and Nucleic Acids", Cambridge University Press, Cambridge, UK (2009); Soding, Bioinformatics, 21(7):951~960 (2005); Altschul et al., Nucleic Acids Res., 25(17):3389~3402 (1997); and Gusfield, "Algorithms on Strings, Trees and Sequences", Cambridge University Press, Cambridge, UK (1997).
[0041] A "vector" or "expression vector" is a replicon, such as a plasmid, phage, virus, or cosmid, into which another DNA segment, such as an "insert," may be joined or incorporated to result in the replication of the joined segment within a cell.
[0042] The term "wild-type" refers to a gene or gene product that possesses the characteristics of the gene or gene product as it was when isolated from a naturally occurring source. Since wild-type genes are the most frequently observed genes in a population, they are arbitrarily referred to as the "normal" or "wild-type" form of the gene. In contrast, the terms "modified," "mutant," or "polymorphic" refer to a gene or gene product that exhibits modifications to its sequence and / or functional characteristics (e.g., alterations in features) compared to the wild-type gene or gene product. It should be noted that naturally occurring mutants can also be isolated; these are identified by the fact that their characteristics are altered compared to the wild-type gene or gene product.
[0043] 2. Methods for identifying transcription modification domains This specification discloses a method for identifying transcription effector (e.g., activation and repression) domains. In some embodiments, the method includes the steps of: preparing a domain library comprising a plurality of nucleic acid sequences configured to express a fusion protein comprising a protein domain derived from a nuclear localized protein, each of which is linked to an inducible DNA-binding domain; transforming reporter cells with the domain library, wherein the reporter cells comprise a bipartite reporter gene comprising a surface marker and a fluorescent protein under the control of a promoter, and the bipartite reporter gene is moduloable by a putative transcription effector domain after treatment with a drug configured to induce an inducible DNA-binding domain; treating the reporter cells with the drug for a period of time required to alter intracellular levels of protein and mRNA (e.g., for an increase due to production or a decrease due to degradation); sequencing the protein domains from the isolated reporter cells; calculating the ratio of sequencing counts from reporter cells without a surface marker, fluorescent protein, or combination thereof to sequencing counts from reporter cells with a surface marker, fluorescent protein, or combination thereof for each protein domain sequence; and identifying the protein domains as transcription repressors or activators.
[0044] The method includes the step of preparing a domain library comprising a plurality of nucleic acid sequences configured to express a fusion protein, each containing a protein domain derived from a nuclear localized protein, each linked to an inducible DNA-binding domain. The protein domain may be 80 amino acids or less. In some embodiments, the protein domain may be about 75 amino acids, about 70 amino acids, about 65 amino acids, about 60 amino acids, about 55 amino acids, about 50 amino acids, about 45 amino acids, about 40 amino acids, about 35 amino acids, about 30 amino acids, about 25 amino acids, about 20 amino acids, about 15 amino acids, about 10 amino acids, or about 5 amino acids.
[0045] The protein domain may be derived from any known protein. In some embodiments, the protein domain is derived from a nuclear localization protein. Nuclear localization proteins include nuclear localization proteins that are completely or partially localized to the nucleus, or can be localized to the nucleus, during the lifespan of the protein. In some embodiments, the protein domain includes the amino acid sequence of a wild-type protein domain derived from a nuclear localization protein. In some embodiments, the protein domain includes a mutant amino acid sequence of a protein domain derived from a nuclear localization protein.
[0046] The inducible DNA-binding domain may be any system used for inducing binding to DNA, including but not limited to the Tet / DOX tetracycline-inducible system, the photo-inducible system, the abscisic acid (ABA)-inducible system, the k-mate system, the 40HT / estrogen-inducible system, the ecdysone-based-inducible system, and the FKBP12 / FRAP (FKBP12-rapamycin complex)-inducible system.
[0047] In some embodiments, the inducible DNA-binding domain includes a tag. The tag may include any tag known in the art, including tags that can be removed by chemical or enzymatic means. Tags suitable for use in this method include chitin-binding proteins (CBPs), maltose-binding proteins (MBPs), Strep tags, glutathione-S-transferase (GST), polyhistidine (PolyHis) tags, ALFA tags, V5 tags, Myc tags, hemagglutinin (HA) tags, spot tags, T7 tags, NE tags, calmodulin tags, polyglutamic acid tags, polyarginine tags, and FLAG tags.
[0048] The method comprises the step of transforming reporter cells with a domain library, wherein the reporter cells contain a bipartite reporter gene comprising a surface marker and a fluorescent protein under the control of a promoter, and the bipartite reporter gene is moduloable by a putative transcription effector domain after treatment with a drug configured to induce an inducible DNA-binding domain.
[0049] Promoters can confer high transcription rates (strong promoters) or low transcription rates (weak promoters). Many promoter libraries have been established experimentally, and the selection of promoters and promoter strengths is cell type-dependent. In some embodiments, weak promoters may be used to identify transcriptional activation domains. In some embodiments, strong promoters may be used to identify transcriptional repression domains.
[0050] Cell surface markers include proteins and carbohydrates conjugated to the cell membrane. In the art, cell surface markers are generally known for various cell types and can be expressed in selected reporter cells based on known molecular biology methods. Surface markers can be synthetic surface markers, comprising a marker polypeptide conjugated to a transmembrane domain. For example, the marker polypeptide may comprise an antibody or a fragment thereof (e.g., an Fc region) conjugated to a transmembrane domain. In some embodiments, the marker polypeptide is a human IgG1 Fc region, and the synthetic surface marker comprises a human IgG1 Fc region conjugated to a transmembrane domain.
[0051] In the art, fluorescent proteins are well known and include proteins adapted to fluoresce in various cellular compartments as a result of variations in the wavelength of incident light. Examples of fluorescent proteins include phycobiliproteins, cyan fluorescent protein (CFP), green fluorescent protein (GFP), yellow fluorescent protein (YFP), enhanced orange fluorescent protein (OFP), enhanced green fluorescent protein (eGFP), modified green fluorescent protein (emGFP), enhanced yellow fluorescent protein (eYFP), and / or monomeric red fluorescent protein (mRFP), as well as their derivatives and variants.
[0052] The method includes the step of separating reporter cells based on the presence or absence of a surface marker, a fluorescent protein, or a combination thereof. In the art, numerous cell separation methods are known that are suitable for use by the methods disclosed herein, including, for example, immunomagnetic cell separation, fluorescence-activated cell sorting (FACS), and microfluidic cell sorting. In some embodiments, the cell separation includes immunomagnetic cell separation.
[0053] In some embodiments, the method further includes the step of stopping the drug treatment of reporter cells, and repeating the steps of separation, sequencing, computation and identification one or more times. In some embodiments, the steps are repeated for at least 48 hours after stopping the drug treatment of reporter cells.
[0054] In some embodiments, the method further includes the step of measuring the expression level of the protein domain. The expression level of the protein domain may be determined using any method known in the art, including immunoblotting and immunoassays of the protein itself or any tag or label thereof. In some embodiments, the expression level is determined by measuring the relative presence or absence of the tag on the DNA-binding domain.
[0055] In some embodiments, the method identifies transcriptional repression domains. In some embodiments, the method includes the steps of: a) preparing a domain library comprising a plurality of nucleic acid sequences configured to express a fusion protein, each comprising a protein domain linked to an inducible DNA-binding domain; b) transforming reporter cells with the domain library, wherein the reporter cells comprise a bipartite reporter gene comprising a surface marker and a fluorescent protein under the control of a strong promoter, and the bipartite reporter gene is capable of being silenced by the putative transcriptional repression domain after treatment with a drug configured to induce an inducible DNA-binding domain; and c) transforming the reporter cells with the drug into cells The process includes the steps of: d) processing the cells for a period of time necessary for the degradation of proteins and mRNA within them; e) isolating reporter cells based on the presence or absence of a surface marker, fluorescent protein, or combination thereof; f) sequencing protein domains from the isolated reporter cells; g) calculating the ratio of sequencing counts from reporter cells without a surface marker, fluorescent protein, or combination thereof to sequencing counts from reporter cells with a surface marker, fluorescent protein, or combination thereof for each protein domain sequence; and g) identifying the protein domains as transcriptional repressors.
[0056] In some embodiments, reporter cells are treated with the drug for at least three days. For example, reporter cells may be treated with the drug for at least three days, at least four days, at least five days, at least six days, at least seven days, at least eight days, at least nine days, at least ten days, or at least fourteen days or more. In some embodiments, reporter cells are treated with the drug for 3 to 12 days, 3 to 10 days, 3 to 7 days, or 3 to 5 days.
[0057] A protein domain is identified as a transcriptional repressor if the log2 ratio of the sequencing count from reporter cells without a surface marker, fluorescent protein, or combination thereof to the sequencing count from reporter cells with a surface marker, fluorescent protein, or combination thereof is at least 2 standard deviations (e.g., large) from the mean of the negative control (see, for example, Figure 1C).
[0058] In some embodiments, the method identifies a transcriptional activation domain. In some embodiments, the method includes: a) preparing a domain library comprising a plurality of nucleic acid sequences configured to express a fusion protein, each comprising a protein domain linked to an inducible DNA-binding domain; b) transforming reporter cells with the domain library, wherein the reporter cells comprise a bipartite reporter gene comprising a surface marker and a fluorescent protein under the control of a weak promoter, and the bipartite reporter gene is capable of being activated by the putative transcriptional activation domain after treatment with a drug configured to induce an inducible DNA-binding domain; and c) transforming the reporter cells with the drug into intracellular The method includes the steps of: d) processing for a period of time necessary for the production of the protein and mRNA; e) isolating reporter cells based on the presence or absence of a surface marker, fluorescent protein, or combination thereof; f) sequencing protein domains from the isolated reporter cells; g) calculating the ratio of sequencing counts from reporter cells without a surface marker, fluorescent protein, or combination thereof to sequencing counts from reporter cells with a surface marker, fluorescent protein, or combination thereof for each protein domain sequence; and g) identifying the protein domain as a transcriptional repressor.
[0059] In some embodiments, reporter cells are treated with the drug for at least 24 hours. For example, reporter cells may be treated with the drug for at least 24 hours (1 day), at least 36 hours, at least 48 hours (2 days), at least 60 hours, at least 72 hours (3 days), at least 94 hours, or at least 106 hours (4 days) or longer. In some embodiments, reporter cells are treated for between 24 and 72 hours or between 36 and 60 hours.
[0060] A protein domain is identified as a transcription activator if the log2 ratio of the sequencing count from reporter cells without a surface marker, fluorescent protein, or combination thereof to the sequencing count from reporter cells with a surface marker, fluorescent protein, or combination thereof is at least 2 standard deviations (e.g., small) from the mean of the negative control (see, for example, Figure 5B).
[0061] 3. Transcription factors The Disclosure also provides synthetic transcription factors comprising one or more transcription effector domains fused to heterogeneous DNA-binding domains. As used herein, the term “transcription factor” refers to a protein or polypeptide that directly or indirectly interacts with a specific DNA sequence associated with a genomic locus or gene of interest to block RNA polymerase activity or recruit it to a promoter site for a gene or set of genes.
[0062] In some embodiments, the synthetic transcription factor comprises one or more transcriptional activation domains, one or more transcriptional repression domains, or a combination thereof, fused to a heterologous DNA-binding domain. In some embodiments, at least one of the one or more transcriptional activation domains or at least one of the one or more transcriptional repression domains comprises an amino acid sequence having at least 70% (e.g., at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 98%, 99%) identity to any of SEQ ID NOs: 1 to 896. In some embodiments, one or more transcriptional activation domains, one or more transcriptional repression domains, or a combination thereof are identified by the methods disclosed herein.
[0063] In some embodiments, the synthetic transcription factor comprises two or more transcription effector domains (e.g., a transcription activating domain, a transcription repressing domain, or a combination thereof) fused to a heterogeneous DNA-binding domain. The two or more effector domains can be fused to the DNA-binding domain in any orientation and can be separated from each other by amino acid linkers.
[0064] In some embodiments, if the synthetic transcription factor includes more than one transcription effector domain, the synthetic transcription factor may include at least one transcriptional activating domain or at least one transcriptional repressing domain as disclosed herein, along with at least one further effector domain known in the Art. See, for example, Tycko J. et al., Cell, December 23, 2020, 183(7):2020-2035, incorporated herein by reference in its entirety. In some embodiments, one or more transcriptional activating domains and one or more transcriptional repressing domains are identified by the methods described herein.
[0065] In some embodiments, if the synthetic transcription factor contains more than one transcription effector domain, at least one of the one or more transcription activation domains contains an amino acid sequence having at least 70% identity to any of SEQ ID NOs. 563-664. In some embodiments, at least one of the one or more transcription activation domains contains an amino acid sequence having at least 70% identity to any of SEQ ID NOs. 563-596. In some embodiments, at least one of the one or more transcription activation domains is selected from those found in Table 2.
[0066] In some embodiments, if the synthetic transcription factor contains more than one transcription effector domain, at least one of the one or more transcription repression domains contains an amino acid sequence having at least 70% identity to any of SEQ ID NOs: 1-562 and 665-896. In some embodiments, at least one of the one or more transcription repression domains contains an amino acid sequence having at least 70% identity to any of SEQ ID NOs: 666. In some embodiments, at least one of the one or more transcription repression domains is selected from those found in Tables 1, 3, or 4.
[0067] A DNA-binding domain is any polypeptide capable of binding to double-stranded or single-stranded DNA, either generally or in a sequence-specific manner. DNA-binding domains include polypeptides having helix-turn-helix motifs, zinc fingers, leucine zippers, HMG (high mobility group box) domains, winged helix regions, winged helix-turn-helix regions, helix-loop-helix regions, immunoglobulin folds, B3 domains, Wor3 domains, TAL effector DNA-binding domains, and the like. Heterogeneous DNA-binding domains can be innate binding domains. In some embodiments, heterogeneous DNA-binding domains include programmable DNA-binding domains, such as those manipulated by modifying one or more amino acids of an innate DNA-binding domain to bind to a given nucleotide sequence.
[0068] In some embodiments, the DNA-binding domain can directly bind to a target DNA sequence.
[0069] The DNA-binding domain may originate from domains found within spontaneously occurring transcription activator-like effectors (TALEs), such as AvrBs3, Hax2, Hax3, or Hax4 (Bonas et al., 1989, Mol Gen Genet, 218(1):127~36; Kay et al., 2005, Mol Plant Microbe Interact 18(8):838~48). TALEs possess a modular DNA-binding domain consisting of repeating residue sequences, with each repeat region comprising 34 amino acids. The residue pairs at positions 12 and 13 of each repeat region determine nucleotide specificity, and the combination of regions enables the synthesis of sequence-specific TALE DNA-binding domains. In some embodiments, TALE DNA-binding domains may be manipulated using known methods to impart selective specificity to any target sequence. The DNA-binding domain may contain multiple (e.g., two, three, four, five, six, ten, twenty or more) Tal effector DNA-binding motifs. In particular, any number of nucleotide-specific Tal effector motifs can be combined to form the sequence-specific DNA-binding domain utilized in this transcription factor.
[0070] In some embodiments, the DNA-binding domain associates with target DNA in response to an exogenous factor.
[0071] In some embodiments, the DNA-binding domain is derived from a clustered regularly spaced short-chain palindromic repeat-associated (Cas) protein (e.g., catalytically inactivated Cas9) and associates with the target DNA via a guide RNA. The gRNA itself is a sequence complementary to one strand and a scaffold sequence of the target DNA sequence, and includes a sequence that binds to the target DNA sequence and recruits it to Cas9. The transcription factors described herein may be useful for CRISPR interference (CRISPRi) or CRISPR activation (CRISPRa).
[0072] Guide RNA (gRNA) can be crRNA, crRNA / tracrRNA (or single-stranded guide RNA, sgRNA). gRNA can be non-spontaneous gRNA. The terms “gRNA,” “guide RNA,” and “guide sequence” may be used interchangeably throughout this specification and refer to nucleic acids containing the sequence that determines the specificity of binding to the Cas protein. gRNA hybridizes (partially or completely complementarily) with the DNA target sequence.
[0073] The gRNA or a portion thereof that hybridizes with the target nucleic acid (target site) can be of any length necessary for selective hybridization. The gRNA or sgRNA(s)(s)(s)(s)(s)(s)(s)(s)(s))(s))))))))))))))))))))(s)(s)(s)(s)(s)))))))))))))(s)) It can be 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, or 100 nucleotides or longer.
[0074] Numerous computational tools have been developed to facilitate gRNA design (see Prykhozhij et al. (PLoS ONE, 10(3):(2015)); Zhu et al. (PLoS ONE, 9(9)(2014)); Xiao et al. (Bioinformatics., January 21 (2014)); Heigwer et al. (Nat Methods, 11(2):122-123 (2014))). Methods and tools for designing guide RNAs are discussed by Zhu (Frontiers in Biology, 10(4), pp. 289-296 (2015)), which are incorporated herein by reference. In addition, many software tools have been published that can be used to facilitate sgRNA design(s). Genome-wide gRNA databases, including but not limited to IDT DNA Predesigned Alt-R CRISPR-Cas9 guide RNA, Addgene Validated gRNA Target Sequences, and GenScript, have also been published. These are pre-designed gRNA sequences that target many genes and locations within the genomes of many species (humans, mice, rats, zebrafish, and C. elegans).
[0075] This disclosure also provides nucleic acids that encode synthetic transcription factors or transcription effector domains (e.g., activating domains or repressive domains) as disclosed herein. For example, effector domains may be encoded by nucleic acids disclosed in Tables 1 to 3. In some embodiments, effector domains may be encoded by nucleic acids having at least 70% identity to any of sequence numbers 897 to 1329. In some embodiments, the nucleic acid encodes one or more synthetic transcription factors or one or more effector domains.
[0076] The nucleic acids of this disclosure may include any of the numerous promoters known in the art, in which case the promoter is a constitutive promoter, a regulatory promoter or an inductive promoter, a cell type-specific promoter, a tissue-specific promoter or a species-specific promoter. In addition to sequences sufficient to direct transcription, the promoter sequences of the present invention may also include sequences of other regulatory elements (e.g., enhancers, Kozak sequences and introns) that are involved in modulating transcription. In this field, many promoter / regulatory sequences useful for driving constitutive gene expression are available, including, but not limited to, CMV (cytomegalovirus promoter), EF1a (human elongation factor 1 alpha promoter), SV40 (monkey vacuolated virus 40 promoter), PGK (mammalian phosphoglycerate kinase promoter), Ubc (human ubiquitin C promoter), human beta-actin promoter, rodent beta-actin promoter, CBh (chicken beta-actin promoter), CAG (hybrid promoter containing CMV enhancer, chicken beta-actin promoter and rabbit beta-globin splice acceptor), TRE (tetracycline response element promoter), H1 (human polymerase III RNA promoter), U6 (human U6 micronucleus promoter), etc. Further promoters that may be used for the expression of components of this system include, but are not limited to, viral LTRs such as the cytomegalovirus (CMV) intermediate early promoter, Rous sarcoma virus LTR, HIV-LTR, high-throughput LV-1 LTR, Moloney's mouse leukemia virus (MMLV) LTR, myeloproliferative sarcoma virus (MPSV) LTR, and splenic fociforming virus (SFFV) LTR, as well as the simian virus 40 (SV40) early promoter, herpes simplex virus tk promoter, and elongation factor 1-alpha (EF1-α) promoter with or without the EF1-α intron. Further promoters include any constitutively active promoters. Alternatively, any regulatory promoter may be used so that its expression can be modulated intracellularly.
[0077] Furthermore, inducible expression can be achieved by placing nucleic acids encoding such molecules under the control of an inducible promoter / regulatory sequence. Promoters known in the art are capable of induction in response to inducers such as metals, glucocorticoids, tetracyclines, and hormones, and are also envisioned for use according to the present invention. Accordingly, it is perceived that this disclosure includes the use of any promoter / regulatory sequence known in the art that is operably linked to these and capable of driving the expression of a desired protein.
[0078] This disclosure also provides nucleic acid-containing vectors and cells containing nucleic acids or such vectors. Vectors may be used to propagate nucleic acids within suitable cells and / or to enable expression from nucleic acids (e.g., expression vectors). Those skilled in the art are familiar with a variety of vectors available for the propagation and expression of nucleic acid sequences.
[0079] To construct cells expressing this transcription factor, an expression vector for stable or transient expression of the system can be constructed and introduced into cells via conventional methods. For example, nucleic acids or other nucleic acids or proteins encoding components of the transcription factor of this disclosure can be cloned into a suitable expression vector, such as a plasmid or viral vector, under operable ligation to a suitable promoter. The choice of expression vector / plasmid / viral vector should be suitable for integration and replication within eukaryotic cells.
[0080] In certain embodiments, the vectors of this disclosure may be used to drive the expression of one or more sequences in mammalian cells using mammalian expression vectors. Examples of mammalian expression vectors include pCDM8 (Seed, Nature (1987), 329:840, incorporated herein by reference) and pMT2PC (Kaufman et al., EMBO J. (1987) 6:187, incorporated herein by reference). When used in mammalian cells, the regulatory function of the expression vector is typically provided by one or more regulatory elements. For example, commonly used promoters are derived from polyomaviruses, adenovirus type 2, cytomegalovirus, simian virus 40, and other viruses disclosed herein and known in the art. For other expression systems suitable for both prokaryotic and eukaryotic cells, see, for example, Chapters 16 and 17 of Sambrook et al., "MOLECULAR CLONING: A LABORATORY MANUAL," 2nd edition, Cold Spring Harbor Laboratory, Cold Spring Harbor Laboratory Press, Cold Spring Harbor, NY, 1989, which is incorporated herein by reference.
[0081] The vectors of this disclosure can direct the expression of nucleic acids within a specific cell type (for example, tissue-specific regulatory elements are used to express nucleic acids). Such regulatory elements may include promoters that are tissue-specific or cell-specific. When applied to a promoter, the term “tissue-specific” refers to a promoter capable of directing the selective expression of a target nucleotide sequence to a specific tissue type (e.g., seed) in the relative absence of expression of the same nucleotide sequence in a different tissue type. When applied to a promoter, the term “cell-type specific” refers to a promoter capable of directing the selective expression of a target nucleotide sequence within a specific cell type in the relative absence of expression of the same nucleotide sequence in a different cell type within the same tissue. When applied to a promoter, the term “cell-type specific” also means a promoter capable of promoting the selective expression of a target nucleotide sequence in a region within a single tissue. The cell-type specificity of a promoter can be evaluated using methods well known in the art, such as immunohistochemical staining.
[0082] In addition, the vector may contain, for example, some or all of the following: selection marker genes for selecting stable or transient transformants in host cells; transcription termination signals and RNA processing signals; 5' and 3' untranslated regions; an internal ribosome binding site (IRES), which is a multipurpose multiple cloning site; and a reporter gene for evaluating the expression of a chimeric receptor. Vectors and methods suitable for constructing vectors containing transgenes are well known and available in the art. Selection markers include chloramphenicol resistance, tetracycline resistance, spectinomycin resistance, neomycin, streptomycin resistance, erythromycin resistance, rifampicin resistance, bleomycin resistance, thermocompatible kanamycin resistance, gentamicin resistance, hygromycin resistance, trimethoprim resistance, dihydrofolate reductase (DHFR), GPT; and genes from S. cerevisiae, namely URA3, HIS4, LEU2, and TRP1.
[0083] When introduced into cells, the vector may be maintained as an autonomous replicating sequence or an extrachromosomal element, or it may be integrated into the host DNA. Accordingly, this disclosure further provides cells comprising synthetic transcription factors, nucleic acids, or vectors as disclosed herein.
[0084] Conventional virus-based and non-virus-based gene transfer methods can be used to introduce nucleic acids into cells, tissues, or subjects. Such methods can be used to administer nucleic acids to cells in culture or to cells in a host organism. Non-viral vector delivery systems include DNA plasmids, cosmids, RNA (e.g., transcripts of vectors described herein), nucleic acids, and nucleic acids complexed with a delivery medium.
[0085] Viral vector delivery systems include DNA viruses and RNA viruses that are episomal or integrated into the genome after delivery to cells. Various viral constructs can be used to deliver these nucleic acids to cells, tissues and / or targets. Viral vectors include, for example, retroviral vectors, lentiviral vectors, adenovirus vectors, adeno-associated virus vectors, and herpes simplex virus vectors. Non-limited examples of such recombinant viruses include recombinant adeno-associated virus (AAV), recombinant adenovirus, recombinant lentivirus, recombinant retrovirus, recombinant herpes simplex virus, recombinant poxvirus, and phages. This disclosure presents vectors, such as retroviruses or lentiviruses, that can be integrated into the host genome. See, for example, Ausubel et al., "Current Protocols in Molecular Biology," John Wiley & Sons, New York, 1989; Kay, MA et al., 2001, Nat. Medic., 7(1):33-40; and Walther W. and Stein U., 2000, Drugs, 60(2):249-71, which are incorporated herein by reference.
[0086] Nucleic acids or transcription factors may be delivered by any suitable means. In certain embodiments, nucleic acids or their proteins are delivered in vivo. In other embodiments, nucleic acids or their proteins are delivered in vitro or ex vivo to isolated / cultured cells to yield modified cells useful for in vivo delivery to patients suffering from a disease or condition.
[0087] A wide variety of host cells may be transformed or transfected with vectors relating to this disclosure, and vectors relating to this disclosure may be introduced into a wide variety of host cells in other ways. Transfection refers to the uptake of a vector into a cell, whether or not any coding sequence is actually expressed. To those skilled in the art, numerous transfection methods are known, e.g., lipofectamine, calcium phosphate coprecipitation, electroporation, DEAE dextran treatment, microinjection, viral infection, and other methods known in the art. Transduction refers to the entry of a virus into a cell and the expression (e.g., transcription and / or translation) of the sequence delivered by the viral vector genome. In the case of recombinant vectors, "transduction" generally refers to the entry of a recombinant viral vector into a cell and the expression of the target nucleic acid delivered by the vector genome.
[0088] In the art, methods for delivering vectors to cells are well known and may include: electroporation of DNA or RNA; transfection reagents such as liposomes or nanoparticles for delivering DNA or RNA; delivery of DNA, RNA, or proteins by mechanical deformation (see, for example, Sharei et al., Proc. Natl. Acad. Sci. USA (2013), 110(6):2082-2087, incorporated herein by reference); or viral transduction. In some embodiments, the vector is delivered to the host cell via viral transduction. Nucleic acids may also be delivered as part of a larger construct, such as a plasmid or viral vector, or directly by, for example, electroporation, lipid vesicles, viral transporters, microinjection, and bioristic methods (high-speed particle bombardment). Similarly, constructs containing one or more transgenes may be delivered by any method suitable for introducing the nucleic acid into a cell. In some embodiments, the construct or nucleic acid encoding the components of the system is a DNA molecule. In some embodiments, the nucleic acid encoding the components of the system is a DNA vector that can be electroporated into cells. In some embodiments, the nucleic acid encoding the components of the system is an RNA molecule that can be electroporated into cells.
[0089] In addition, delivery media such as nanoparticle-based and lipid-based delivery systems may also be used. Further examples of delivery media include lentiviral vectors, ribonucleoprotein (RNP) complexes, lipid-based delivery systems, gene guns, hydrodynamic methods, electroporation or nucleofection, microinjection, and bioristic methods. A variety of gene delivery methods are discussed in detail by Nayerossadat et al. (Adv Biomed Res., 2012, 1:27) and Ibraheem et al. (Int J Pharm., January 1, 2014, 459(1~2):70~83), which are incorporated herein by reference.
[0090] Thus, this disclosure provides isolated cells containing the vector(s) or nucleic acid(s) disclosed herein. Preferred cells are those that proliferate readily and reliably, have a reasonably fast growth rate, have a well-characterized expression system, and can be easily and efficiently transformed or transfected. Suitable prokaryotic cells include, but are not limited to, cells derived from the genera Bacillus (such as Bacillus subtilis and Bacillus brevis), Escherichia (such as E. coli), Pseudomonas, Streptomyces, Salmonella, and Erwinia. Suitable eukaryotic cells are known in the art and include, for example, yeast cells, insect cells, and mammalian cells. Examples of suitable yeast cells include those derived from the genera Kluyveromyces, Pichia, Rhinosporidium, Saccharomyces, and Schizosaccharomyces. Exemplary insect cells include Sf-9 cells and HIS cells (Invitrogen, Carlsbad, Calif.), which are described, for example, Kitts et al., Biotechniques, 14:810-817 (1993); Lucklow, Curr. Opin. Biotechnol., 4:564-572 (1993), and Lucklow et al., J. Virol., 67:4566-4579 (1993), which are incorporated herein by reference. The cells are preferably mammalian cells, and in some embodiments, the cells are human cells. In this field, numerous suitable mammalian and human host cells are known, many of which are available from the American Type Culture Collection (ATCC, Manassas, Va.).Examples of suitable mammalian cells include, but are not limited to, Chinese hamster ovary cells (CHO) (ATCC accession number: CCL61), CHO DHFR cells (Urlaub et al., Proc. Natl. Acad. Sci. USA, 97:4216~4220 (1980)), human embryonic kidney (HEK) 293 cells or HEK 293T cells (ATCC accession number: CRL1573), and 3T3 cells (ATCC accession number: CCL92). Other suitable mammalian cell lines include the monkey COS-1 cell line (ATCC accession number: CRL1650) and COS-7 cell line (ATCC accession number: CRL1651), as well as the CV-1 cell line (ATCC accession number: CCL70). Further exemplary mammalian host cells include primate cell lines, rodent cell lines, and human cell lines, including transformed cell lines. Suitable cell lines include normal diploid cells, cell lines derived from in vitro cultures of primary tissue, and primary explants. Other suitable mammalian cell lines include, but are not limited to, mouse neuroblastoma N2A cells, HeLa cells, HEK cells, A549 cells, HepG2 cells, mouse L-929 cells, and the BHK hamster cell line or HaK hamster cell line.
[0091] In this technical field, methods for selecting mammalian cells and methods suitable for cell transformation, culture, amplification, screening, and purification are known.
[0092] The present invention also covers compositions or systems comprising synthetic transcription factors, nucleic acids, vectors, or cells as described herein. In some embodiments, the composition or system comprises two or more synthetic transcription factors, nucleic acids, vectors, or cells.
[0093] In some embodiments, the composition or system further comprises gRNA. The gRNA may be encoded on the same nucleic acid as the synthetic transcription factor, or on a different nucleic acid. In some embodiments, the vector encoding the synthetic transcription factor may further encode the gRNA under the same promoter, or under a different promoter. In some embodiments, the gRNA is encoded on its own vector, separated from the transcription factor vector.
[0094] 4. Methods for modulating gene expression This disclosure also provides a method for modulating the expression of at least one target gene in a cell, comprising the step of introducing at least one synthetic transcription factor, nucleic acid, vector, or composition or system described herein into the cell. In some embodiments, the gene expression of at least two genes is modulated.
[0095] Expression modulation involves increasing or decreasing gene expression compared to normal gene expression for a target gene. When the expression of at least two genes is modulated, the expression of any of the genes may increase, the expression of any of the genes may decrease, or the expression of one gene may increase while the expression of the other decreases.
[0096] The cells may be prokaryotic or eukaryotic. In a preferred embodiment, the cells are eukaryotic. In some embodiments, the cells are in vitro. In some embodiments, the cells are ex vivo.
[0097] In some embodiments, the cells are cells in an organism or host, such that the introduction of the disclosed systems, compositions, or vectors into the cells includes administration to a subject. The method may include the step of providing or administering to a subject at least one synthetic transcription factor, nucleic acid, vector, or composition or system described herein by transferring cells treated in vivo or ex vivo.
[0098] The “subject” may be human or non-human, and may include animal strains or species used as “model systems” for research purposes, such as the mouse models described herein. Similarly, the subject may include adults or young people (e.g., children). Furthermore, the subject may mean any organism, preferably a mammal (e.g., human or non-human mammal), that may benefit from the administration of the compositions envisioned herein. Examples of mammals include, but are not limited to, any member of the following classes of mammals: non-human primates such as humans, chimpanzees and other ape and monkey species; farm animals such as cattle, horses, sheep, goats and pigs; pets such as rabbits, dogs and cats; and laboratory animals such as rodents such as rats, mice and guinea pigs. Examples of non-mammals include, but are not limited to, birds and fish. In one embodiment of the methods and compositions provided herein, the mammal is human.
[0099] As used herein, the terms “to administer,” “to introduce,” and “to introduce” are interchangeable herein and refer to the placement of the System of the Disclosure into a subject by a method or route that results in at least partial localization of the System to a desired site. The System may be administered by any suitable route that results in delivery to a desired site in the subject.
[0100] 5. Kit The scope of this disclosure also includes at least one or all of at least one nucleic acid encoding an effector domain or a DNA-binding domain or a combination thereof, at least one synthetic transcription factor or nucleic acid encoding such a factor, a vector encoding at least one effector domain or at least one synthetic transcription factor, compositions or systems described herein, cells comprising an effector domain, a DNA-binding domain, a synthetic transcription factor, or nucleic acid encoding any of these, reporter cells described herein, and kits comprising a bipartite reporter gene or nucleic acid encoding such a factor, as described herein.
[0101] The kit may also include instructions for using the components of the kit. Instructions are the relevant materials or methods relating to the kit. Materials may include any combination of: background information, a list of components, abbreviated or detailed protocols for using the compositions, troubleshooting, references, technical support, and any other relevant literature. Instructions may be provided with the kit, as a separate component, in paper form, on a computer-readable memory device, downloadable from a website on the internet, or presented as a record of display.
[0102] It is understood that the disclosed kit may be used in connection with the disclosed methods. The kit may include instructions for use in any of the methods described herein. The instructions may include descriptions of the use of components for methods of identifying repression domains or modulating gene expression.
[0103] The kits provided herein are appropriately packaged. Appropriate packaging includes, but is not limited to, vials, bottles, jars, flexible packaging, etc.
[0104] The kit may optionally provide further components, such as information for interpreting the buffer. Typically, the kit includes a container and labeling on the container or package insert accompanying the container. In some embodiments, the disclosure provides a product containing the contents of the kit described above.
[0105] The kit may further include devices for holding or administering the system or the composition. The devices may include an infusion device, an intravenous solution bag, a subcutaneous injection needle, a vial, and / or a syringe.
[0106] This disclosure also provides kits for carrying out the method or preparing components in vitro. The kits may include components of the system. Optional components of the kit include one or more of the following: (1) buffer components, (2) control plasmids, and (3) sequencing primers. [Examples]
[0107] 6. Examples Human gene expression is regulated by thousands of proteins that activate or repress transcription. We lack a complete and quantitative description of effector domains, which are sufficient domains to mediate changes in the gene expression of these proteins. To systematically measure transcriptional effector domains in human cells, a high-throughput assay is provided herein in which a library of protein domains is fused to a DNA-binding domain and recruited to a reporter gene. Cells are then separated according to reporter expression levels, and the library of protein domains is sequenced. The reporter is a synthetic surface marker that facilitates simple separation of tens of millions of cells into high-expression and low-expression populations using magnetic beads.
[0108] Gene silencing and epigenetic memory were quantified after the recruitment of all nucleoprotein domains of ≤80 amino acids. Measurements for the entire family of >300 KRAB domains and >200 homeodomains were used to discover the relationship between the repressive domain strength of transcription factors and their evolutionary history and developmental roles. Furthermore, a deep mutation scan was performed on ZNF10 KRAB effector function to identify substitutions that enhance stability and repression compared to KRAB domains used in CRISPRi. To explore effector domains beyond existing annotation regions, the sequences of 238 repressive complex proteins were tiled, and a novel, short-chain repressive domain of 10 amino acids was discovered within the non-annotation region of large chromatin regulators, including the non-canonical Polycomb 1.6 recruitment protein MGA. We individually characterized over 20 repressors and found that all of them, although their silencing and epigenetic memory dynamics differed significantly, silenced reporter genes completely or to zero at the single-cell level.
[0109] In addition, we discovered novel activation domains within nucleoproteins, including highly branched, acidic KRAB domain variants.
[0110] In summary, these results support strategies for the systematic measurement of transcription effector domain activity in human cells and expand the number of compact transcription effector domains that can be applied in synthetic transcription and epigenetic perturbation techniques.
[0111] The problems addressed by this technology are: i. The unknown question of which genes possess effector functions. ii. Problems within known TF / CR genes where it is often unknown which domain is responsible for this function. iii. Issues within a domain family, including known effector domains, where it is unknown which family members possess this functionality. iv. Known problems within effector domains where it is unknown which residues are necessary and how mutations reduce or enhance function. That is the case.
[0112] The systems and methods provided herein can measure regulatory domains that alter the output from a reporter promoter, affecting activating and repressive capabilities. Historically, this has required low-throughput work, and relatively few effector domains have been measured. The systems and methods provided herein offer an alternative high-throughput assay.
[0113] The systems and methods are used, for example, a. to understand gene regulation and predict the function of non-coding regulatory elements to which these proteins bind; and b. to identify effector domains for epigenetic perturbation tools.
[0114] Previously, a limited number of transcription effector domains were available for manipulating synthetic transcription factors. To address this limitation, a high-throughput method for screening and quantifying the function of transcription effector domains is provided herein. This method enables the discovery of hundreds of effector domains that, when fused to DNA-binding domains, may upregulate or downregulate transcription in a targeted manner. This step also identifies mutants of effector domains with enhanced activity. These effector domains can be used to manipulate synthetic transcription factors for applications in gene therapy and cell therapy, synthetic biology, and functional genomics.
[0115] The novel transcription effector domains provided herein offer several advantages for applications relying on synthetic transcription factors. The inventors identify high-throughput steps for further shortening to minimally sufficient sequences, which are advantages for short-chain domains (≤80 amino acids) and delivery (e.g., packaging within viral vectors). In some cases, the inventors identify potent effector domains as short as 10 amino acids. These domains are extracted from human proteins, which offers the advantage of reduced immunogenicity compared to viral effector domains. The majority of these domains have not yet been reported as transcription effectors.
[0116] High-throughput recruitment with Pfam domain libraries for both strong pEF promoters and weak minCMV promoters allowed for the measurement of both repressive and activating domains. One possible reason for the discovery of numerous further repressors is that TADs are more often unannotated, disordered, or low-complexity regions, while being autonomously stable folding sequences that satisfy Pfam's domain definition. Another possible reason is that, in the nucleus, activating cofactors are more limited than repressive cofactors (Gillespie, Mol Cell, 2020), which implies that reduced expression of activating domains may result in increased activation intensity, but this effect is not expected to completely shield signals within the screen. Novel library designs that tile transcription factors or focus on regions with TAD-like signatures (e.g., acidity) will reveal further activating domains.
[0117] In addition, this specification discloses high-throughput steps for examining mutations within these domains to identify enhancing mutants. High-throughput methods are made more readily possible by the development of artificial cell surface markers, which result in more efficient, inexpensive, and rapid screening of these libraries using magnetic separation. This is an advantage over the more conventional method of separating libraries based on the expression of fluorescent reporter genes.
[0118] [Example 1] High-throughput recruitment identifies hundreds of repressive domains within human proteins. To transform classical recruitment reporter assays for transcription domains into high-throughput assays, we addressed two problems: (1) modifying the reporter to make it suitable for rapid screening of libraries containing tens of thousands of domains; and (2) developing a strategy for creating a library of candidate effector domains. To improve upon already published fluorescent reporters (Bintu et al., 2016), we manipulated synthetic surface markers to enable easy magnetic separation of a large number of cells and incorporated the reporter into a suspension cell system suitable for cell culture in large-capacity spinner flasks. Specifically, we created K562 reporter cells with a 9×TetO binding site upstream of a strongly constitutive pEF1a promoter that drives the expression of a two-part reporter (Figure 1) consisting of a synthetic surface marker (human IgG1 Fc region linked to an Igκ reader and PDGFRβ transmembrane domain) and a fluorescent citrin protein. Flow cytometry confirmed that within 5 days, recruitment of the KRAB domain, a known repressive domain derived from the zinc finger transcription factor ZNF10, at the TetO site silenced this reporter in a doxycycline-dependent manner (Figures 7 and 16A and 16B). Magnetic separation using ProG Dynabead bound to the synthetic surface marker separated reporter-on cells from reporter-off cells (Figures 7 and 16C).
[0119] Sequences were extracted from the UniProt database for Pfam annotation domains within human proteins (including human proteins, not limited to nuclear-localized proteins) that can localize to the nucleus. A total of 14,657 domains were searched. Of these, 72% were 80 amino acids (AA) or less in length (Figure 1), which indicated compatibility of these domains with synthetic oligonucleotides pooled as 300-base oligonucleotides. For domains shorter than 80 amino acids, the domain sequence was extended to 80 amino acid length at both ends by adjacent residues derived from the native protein sequence to avoid PCR amplification bias. 861 negative controls were added, which were 80-amino acid sequences tiled along the DMD protein with a random sequence of 80 amino acids or a tiling window of 10 amino acids. Since the DMD protein did not localize in the nucleus (Chevron et al., 1994), its potential to characterize domains with transcriptional activity was small. The library was cloned for lentiviral expression as a fusion protein with either the rTetR doxycycline-inducible DNA-binding domain alone or with 3×FLAG-tagged rTetR (Figures 17A and 178), and delivered to K562 reporter cells (Figure 1).
[0120] Before assaying transcriptional activity, a high-throughput method was used to determine which protein domains were well-expressed in K562 cells (Figures 17A and 178). The cell library was stained with an anti-FLAG fluorescently labeled antibody, and the cells were separated into two vials (Figures 17B and 178). Genomic DNA was extracted, and the frequency of each domain was counted by unit replication sequence sequencing. Using the sequencing count, the expression level of each domain was measured using FLAG low FLAG, in contrast to the group. high The enrichment ratio of the population was calculated. These measurements were reproducible between individually transfected biological replicates (r 2=0.82, Figures 17C and 8), highly correlated with the expression levels of individual domain fusions as measured by Western blotting (r 2 =0.92, Figures 17D and 17E and 8). The natural Pfam domain was significantly better expressed than the random sequence control (p<1×10 by Mann-Whitney test). -5 In contrast, the Pfam domain and DMD tiling controls showed equally good expression (Figures 17F and 178). FLAG expressed one standard deviation above the median of the random controls. high :FLAG low A threshold was set to identify well-expressed domains based on the ratio. According to this definition, 66% of the Pfam domains were well-expressed, and these domains were targeted for further analysis.
[0121] The Pfam domain library was screened for transcriptional repressors. The pooled cell library was treated with doxycycline for 5 days, allowing sufficient time for transcriptional silencing and subsequent degradation and dilution of reporter mRNA and protein for cell division, resulting in a clear biphasic mixture of "on" and "off" cells (Figures 18A and 189). Magnetic cell separation (Figures 18A and 189) and domain sequencing were then performed, and the log2(off:on) ratio was calculated for each library member using read counts within the unbound and bead-bound populations (Figure 1). For clarity, the bead-bound population was referred to as the "on" population, and the unbound population as the "off" population. The measurements were highly reproducible between individually transduced biological replicates (r 2=0.96, Figure 1). A domain was considered a hit if it caused suppression exceeding two standard deviations of the mean of the poorly expressed negative control. This resulted in 446 repressor hits on day 5 for domains derived from 63 domain families (Figure 12A). These repressor domains were found in 451 human proteins, as the exact same domain sequence can occur in multiple genes in some cases. Known repressor domains derived from 10 domain families and described by Pfam as repressor or repressor cofactor-binding domains (e.g., KRAB from human ZNF10, Chromoshadow from CBX5) were among the hits. Further time points were set on days 9 and 13 to measure epigenetic memory. The set of proteins containing hits was significantly enriched for transcription factors and chromatin regulators compared to all nuclear proteins used in the library, but proteins from different categories were also differentially enriched when classified by their memory levels (Figures 18B and 9). Specifically, on day 13, highly memory-enriched repressors (which the cells maintained in the off state) were most highly enriched in C2H2-type zinc finger transcription factors, including the KRAB ZNF protein, while less memory-enriched repressors were most highly enriched in homeodomain transcription factors, including the Hox protein. Overall, the extremely high reproducibility between hits and the identification of predicted positive control repressor domains suggested that a screening method called high-throughput recruitment yields reliable results. Table 1 shows the amino acid and nucleic acid sequences of the repressors identified in the nuclear Pfam domain library, with high scores indicating high repression.
[0122] One of the strongest hits was YAF2_RYBP, a domain present in RING1 / YY1-binding protein (RYBP) and its paralog, YY1-related factor 2 (YAF2), both of which are components of Polycomb repression complex 1 (PRC1) (Chittock et al., 2017; Garcia et al., 1999). The RYBP protein domain annotated by Pfam (only 32 amino acids, therefore a short chain synthesized within an 80-amino acid domain library) was examined individually to confirm rapid silencing of the reporter gene (Figure 12B). RYBP-mediated silencing has also been supported by recent reports on the recruitment of full-length RYBP protein in mouse embryonic stem cells (Moussa et al., 2019; Zhao et al., 2020). The results established that the 32-amino acid RYBP domain (Wang et al., 2010), which was shown by surface plasmon resonance to be the minimum domain required to bind to the polycomb histone modifying enzyme RING1B, is sufficient to mediate intracellular silencing.
[0123] To quantify the inhibitory response rate, the distribution of citrin levels was gated, and the percentage of silencing cells, normalized by uniform low levels of background silencing within untreated cells, was calculated. The data were then fitted to a model that reached a plateau with a constant irreversible silenced percentage of cells, accompanied by an exponential silencing rate during doxycycline treatment and exponential decay (or reactivation) after doxycycline removal (Figure 12C). Using this method, the inhibitory functions of the Chromo domain derived from SUMO3 and MPP8, the Chromoshadow domain derived from CBX1, and the SAM_1 / SPM domain derived from SCMH1 (Figures 18C-18F and 9), all of which have been supported by recruitment assays or inhibitory cofactor binding assays, were also verified (Chang et al., 2011; Chupreta et al., 2005; Frey et al., 2016; Lechner et al., 2000). The silencing rate from all individual measurements (for the suppressor hits mentioned above and other hits discussed below; Figures 18C-18K and 9) correlated well with the high-throughput silencing measurements on day 5 (R 2 =0.86; Figure 12D). These individual validations were performed using a novel mutant (SE-G72P) of the DNA-binding domain rTetR, which has been engineered to reduce leakage in the absence of doxycycline in yeast (Roney et al., 2016), found not to leak in human cells (Figures 19A and 19B), and has become a useful tool in mammalian synthetic biology. This novel rTetR mutant exhibited the same silencing intensity as the original rTetR in maximum doxycycline recruitment (Figure 19C), which was also supported by a high degree of correlation between the individual validations and the screen scores (Figure 12D). In summary, these validation experiments confirmed that high-throughput recruitment succeeded in both identifying the true repressor and quantifying the repression intensity for each domain with the same precision as individual flow cytometry experiments.
[0124] [Example 2] Identification of domains with unknown function that suppress transcription. While over 22% of the Pfam domain family are classified as functionally unknown domains (DUFs), other domains are not classified using this classification but are nevertheless DUFs (El-Gebali et al., 2019). These domains possess conserved recognizable sequences but lack experimental characterization. Thus, the high-throughput domain screen described herein provided an opportunity to associate initial functions with DUFs. First, the DUF3669 domain was identified as a repressor hit and individually validated by flow cytometry (Figures 12A-12C). These DUFs are found natively within the KRAB zinc finger protein, a gene family containing many repressive transcription factors. Consistent with this, results supporting transcriptional repression after the recruitment of two DUF3669 family domains have been published recently (Al Chiblak et al., 2019), and the high-throughput results extend this finding to include the remaining four unexamined DUF3669 sequences. The C-terminal domain of HNF3, HNF_C, is a different DUF, though it has a more specific name, as it is found only in hepatocyte nuclear factor 3 alpha and hepatocyte nuclear factor 3 beta (also known as FOXA1 and FOXA2). The HNF_C domain derived from either FOXA1 or FOXA2 was also found as an inhibitory hit. Both of these contain the EH1 (engrailed homology 1) motif, characterized by an FxIxxIL sequence, which has been listed as a candidate inhibitory motif (Copley, 2005).
[0125] All three uncharacterized domains found in IRF2BP1, IRF2BP2, and IRF2BPL, which are repressor cofactors of interferon regulator 2 (IRF2) and are the N-terminal zinc finger domain of IRF-2BP1_2 (Childs and Goodbourn, 2003), were repressor hits. The Cyt-b5 domain in the DNA repair factor HERC2 E3 ligase (Mifsud and Bateman, 2002) was validated as a potent repressor hit, but it was a different domain that had not been functionally characterized (Figures 18G and 9). The SH3_9 domain in BIN1, which is a largely uncharacterized variant of the SH3 protein-binding domain, was also validated as a repressor (Figures 18H and 9). BIN1 is a Myc-interacting protein and tumor suppressor (Elliott et al., 1999) also associated with the risk of Alzheimer's disease (Nott et al., 2019). Both full-length BIN1 and Myc-binding domain deletion mutants have already been shown to repress transcription in Gal4 recruitment assays in HeLa cells (Elliott et al., 1999), and the finding that hob1, a yeast homolog of BIN1, was associated with transcriptional repression and histone methylation (Ramalingam and Prendergast, 2007) is consistent with these results. In addition, the repressive activity of the HMG box domain derived from the transcription factor TOX and the zf-C3HC4_2 RING finger domain derived from PCGF2, a component of Polycomb, was also examined (Figures 18I and 18J). Finally, DUF1087 was found within the chromatin remodeler CHD, and its high-throughput measurement was slightly below the screen significance threshold (Figure 12A), but CHD3 DUF1087 was validated as a weak repressor by individual flow cytometry (Figures 12B and 12C). In summary, these results support the idea that high-throughput protein domain screening can assign initial functions to DUFs and extend our understanding of the functions of domains with incomplete characterization.
[0126] [Example 3] A random sequence with strong inhibitory activity Random sequences have not yet been investigated for their inhibitory activity. Surprisingly, one of the 80-amino acid random sequences designed as a negative control was a potent inhibitory hit, with a mean log2 (off:on) of 4.0, despite having low expression levels below the threshold. Individual validation by flow cytometry confirmed that this sequence completely silenced a population of reporter cells after a 5-day recruitment period and resulted in moderate epigenetic memory lasting less than two weeks after doxycycline removal (Figures 18K and 9). One further random sequence showed an inhibitory score slightly above the hit threshold.
[0127] [Example 4] Inhibitory KRAB domains are found in more recent proteins. The data provided an opportunity to analyze the function of all effector domains within the largest transcription factor family: the KRAB domain. The KRAB gene family represented some of the most powerful known repressor domains (such as KRAB in ZNF10). Previous studies on a subset of repressive KRAB domains have shown that they can repress transcription by interacting with the repressor cofactor KAP1, which in turn interacts with chromatin regulators such as SETDB1 and HP1 (Cheng et al., 2014). However, it remains unclear how many KRAB domains are repressors and whether KAP1 recruitment is necessary or sufficient for repression across all KRAB domains.
[0128] The library contained 335 human KRAB domains, and after filtering for well-expressed domains, 92.1% were identified as repressor hits. Nine hit-repressive KRAB domains and two non-hit KRAB domains were individually validated by flow cytometry, and their classification was confirmed in all cases (Figure 19D). Subsequently, the results of domain recruitment were compared with previously published immunoprecipitation mass spectrometry data generated from full-length KRAB protein pulldown (Helleboid et al., 2019). All but one non-repressive KRAB were found to be in proteins that did not interact with KAP1 (one exceptional KRAB had low expression), and all hit-repressive KRAB domains were KAP1-interacting (p<1×10⁻⁶). -9 Fisher's exact test; Figure 2). Furthermore, analysis of available ChiP-seq and ChIP-exo datasets (ENCODE Project Consortium et al., 2020; Imbeault et al., 2017; Najafabadi et al., 2015; Schmitges et al., 2016) revealed that the repressive KRAB domain, in contrast to the non-repressive KRAB domain, originated from the KRAB zinc finger protein, which colocalizes with KAP1 (Figure 2).
[0129] Interestingly, repressive KRAB domains were mostly found in proteins with extremely simple domain architectures consisting only of a KRAB domain and a zinc finger array, while non-repressive KRAB domains were mostly found in genes that also contained a DUF3669 domain or a SCAN domain (Figure 2). In fact, within DUF3669-containing genes, only one KRAB, ZNF783, was a repressor. ZNF783 is an uncharacterized DUF3669-KRAB-containing gene that (despite its name) lacks a zinc finger array in a unique form, suggesting that it differs significantly from other transcription factors in this class in both its effector function and the manner in which it localizes to its target.
[0130] Complex domain architectures, including SCAN or DUF3669, are more common in older KRAB genes (Imbeault et al., 2017). In this embodiment, a clear relationship was observed between the evolutionary age of the KRAB gene and the repressive strength of KRAB. KRAB domains derived from genes predating the common marsupial-human ancestor lacked repressive activity, while KRAB domains derived from later evolved genes functioned as strong repressors (Figure 2). In summary, these results support a model in which older, non-repressive KRAB genes recruit KAP1, leading to a large-scale expansion of repressive KRAB genes in later generations that silence genomic targets.
[0131] [Example 5] High-depth mutation scanning of the CRISPRi ZNF10 KRAB effector to identify mutations that modulate gene silencing. The KRAB domain derived from ZNF10 is widely used in synthetic biology applications for gene repression and is fused to dCas9 in programmable epigenetic / transcriptional control tools known as CRISPR interference (Gilbert et al., 2014). To better understand its sequence-function relationship, we performed a deep mutation scan (DMS) of this KRAB domain using high-throughput recruitment. We designed libraries with all possible single substitutions and all consecutive double and triple substitutions (Figure 3). To improve the ability to clearly align sequencing reads, we implemented silent barcodes within the domain's coding sequence using variable codon usage so that the DNA sequence has greater specificity than the amino acid sequence (Figure 3). Using the reporter and workflow in Figure 1, we performed high-throughput recruitment: doxycycline induction over 5 days and magnetic separation of on-cells and off-cells on days 5, 9, and 13 (Figures 20A and 20A). These measurements were highly reproducible and, as predicted, showed a general trend of increasing toxicity as the mutation length increased from single to triple (Figures 20B and 20B). Furthermore, when these results were compared with KRAB amino acid conservation, a remarkable correlation was found between conservation and the toxicity of the mutations (Figure 3). The amino acid and nucleic acid sequences for the identified KRAB repressive mutants are shown in Table 3. When the score for each repressive mutant is shown compared to 0 for the wild-type sequence, a high score indicates enhanced repression of KRAB transcription.
[0132] The ZNF10 KRAB effector has three components: an A-box required for binding to KAP1 (Peng et al., 2009), a B-box thought to enhance binding to KAP1 (Peng et al., 2007), and an N-terminal extension found in nature on a separate exon upstream of the KRAB domain (Figure 3). Mutations at numerous locations within the A-box dramatically reduced repressive activity compared to the wild-type sequence (Figure 3). Some of these mutations have already been investigated by CAT recruitment assays in COS and 3T3 cells; these data correlated well with measurements obtained by deep mutation scanning in K562 cells (Figure 3). The complete absence of silencing function in KRAB A-box mutants was also examined individually (Figure 3). The periodic nature of the effects of mutations across the A-box suggests that the angles of these residues along the alpha-helix are functionally involved (Figure 3). These residues are designated as residues necessary for silencing (p<1×10 -5 On day 5, using the Wilcoxon rank-sum test to compare the distribution of all substitutions against the wild type, we found 12 required residues where mutations in the A box had a significant impact, and one residue where the impact in the B box was significant but weak (Figure 3).
[0133] These substitutions were mapped to an aligned mouse KRAB A-box structure (PDB:1v65: 55% identity and 69% similarity within the A-box [V13~Y54]; Figures 20C and 10), and it was found that the required residues had similar orientations in 3D space, suggesting a binding interface (Figures 3 and 20D; red; and Figure 10). These residues may be important for binding to KAP1, as it was actually shown in a previous recombinant protein binding assay using KRAB-O (Peng et al., 2009) to facilitate binding to KAP1, with 10 of these A-box residues being aligned to 12~71 (50% identity, 75% similarity) of ZNF10 KRAB within a region containing all 12 of the required residues (red KRAB-O residues; Figures 20C and 10). Of the remaining eight residues that have already been found to be unnecessary for binding, the eight residues were also unnecessary for repression in DMS (p<1×10). -4 Fischer's exact test; gray KRAB-O residues; Figures 20C and 10). When the DMS silencing scores on day 5 were examined for each single, double, and triple alanine substitution used in the binding assay, perfect agreement was found: mutations that detached binding also abolished silencing (Z score < -4 compared to wild-type distribution), and mutations that did not affect binding also did not affect silencing (|Z score| < 0.6) (p < 0.01; n=12 mutations; Fischer's exact test). This high degree of validation and their positions within the 3D structure suggest that, according to DMS, two of the remaining 12 required A-box residues (V41 and N45) may also be involved in binding to KAP1.
[0134] In contrast to the A-box mutations, the B-box mutations showed relatively small effects at the end of recruitment (day 5), with only one statistically significant site (P59) showing a consistent but weak effect. On the other hand, P59 and four other sites (K58, I62, L65, E66) showed significant effects on memory after doxycycline removal, as measured on day 9 (Figure 3). Individual validation of the four significant sites revealed that, as in the high-throughput experiment, the B-box mutants were strong gene silencers at 5 days after recruitment, but showed reduced memory after doxycycline release (Figures 3 and 20E and 10). To interpret these results, we considered the previously proposed gene silencing model (Bintu et al., 2016), which posits that silenced cells pass through a "reversible silent" state before entering an "irreversible silent" state. The memory reduction of the B-box mutant may result in a moderate reduction in the silencing rate, leading to a small number of cells committing to an irreversible silent state by day 5. Since reversible and irreversible silent cells are indistinguishable at day 5, the mutation's effect on the silencing rate may be masked. To investigate this possibility, the silencing time course was repeated with 1 / 100th of the normal dose of doxycycline by fine-tuning the recruitment intensity downwards. In this regime, the B-box mutation reduced the silencing rate before day 5 (Figures 20E and 20E). This result suggests that the B-box contributes partially to the KRAB silencing rate.
[0135] Finally, the N-terminus of KRAB contained residues where numerous substitutions consistently enhanced silencing compared to the wild type (Figure 3; blue; panel on day 13). In particular, almost all substitutions of tryptophan at position 8 resulted in an increased number of silencing cells compared to the wild type at day 13 (the point at which most dynamic ranges detect levels of silencing exceeding those of the wild type). This was the only significant site for enhancing silencing (Figure 3). Memory enhancement was individually examined for the two highest-ranked mutants (WSR8EEE and AW7EE) by high-doxycycline recruitment (Figures 3, 20E, and 10).
[0136] This enhanced silencing may be a result of increased KRAB protein expression levels. To explore the relationship between protein expression levels and KRAB silencing intensity, we examined high-throughput FLAG-tagged expression level measurements for a set of KAP1-binding KRAB domains and found a significant correlation between KRAB expression levels and silencing on day 13 (r 2=0.49, Figures 20F and 10). The low expression level of ZNF10 KRAB on day 13, compared to other KRAB domains that showed high levels of silencing, is strongly related to the results of deep mutation scanning, implying that this result can be improved through mutation. In particular, the N-terminus was extremely poorly preserved (Figure 3), and the fact that this was actually found uniquely within KRABs derived from ZNF10 by BLAST suggests that mutations that improve stability at the N-terminus are unlikely to interfere with KRAB function. In addition, high tryptophan (W) frequency was observed in domains negatively correlated with the expression level throughout domain expression, while high glutamate (E) frequency was observed in domains positively correlated with the expression level (Figures 20G and 10). This trend in amino acid composition suggests that removal of tryptophan from position 8 of KRAB by substitution enhances its effector function, and this enhancement was most pronounced when substituted with glutamate, further suggesting that enhancement of the N-terminal KRAB mutant may be due to improved expression levels. Western blotting of ZNF10 KRAB mutants confirmed that the N-terminal glutamate-substituted mutant was expressed more highly than the wild type (Figures 20H and 20H). In summary, these results support the use of deep mutation scanning to map sequences against function for human transcriptional repressors and to improve effector function by incorporating expression-enhancing substitutions into poorly conserved locations.
[0137] [Example 6] Homeodomain repression strength is collinear with the organization of Hox genes. In the screening, the second largest domain family containing repressor hits was the homeodomain family. Homeodomains are sequence-specific DNA-binding domains composed of three helices, with base contact occurring via helix 3 (Lynch et al., 2006). In some cases, homeodomains are also known to act as repressors (Holland et al., 2007; Schnabel and Abate-Shen, 1996). The library contained homeodomains derived from 216 human genes, of which 26% were repressor hits. Repressors were found in four of the 11 subclasses of homeodomains: PRD, NKL, HOXL, and LIM (Figure 13A). These recruitment assay results suggest that transcriptional repression is a potentially widespread, if not ubiquitous, function of homeodomain transcription factors.
[0138] Next, the results for the HOXL subclass were examined in more detail. This subclass contained Hox genes, a subset of 39 homeodomain transcription factors that are major regulators of cell fate and designate systemic regions along the anterior-posterior axis during embryonic development. These genes were found within four Hox paralog clusters (A-D) that were colinearly arranged from 3' to 5', corresponding to the temporal order and spatial patterning of their expression along the anterior-posterior axis (Gilbert, 1971). Interestingly, the repressive strength of their homeodomains was also colinear with their arrangement within the Hox clusters, with the gene homeodomains closer to the 5' side being stronger repressors (Spearman's ρ = 0.82; Figure 13B). This correlation suggested a possible link between homeodomain repressive function and the timing and spatial patterning of Hox gene expression along the anterior-posterior axis.
[0139] Multiple sequence alignments of Hox homeodomains revealed the RKKR (SEQ ID NO: 1330) motif present in the N-terminal arms of 11 of the strongest repression domains (Figure 13C). While the RKKR motif was present in the strongest repressors in the context of bases, lower-ranked domains, although lacking the RKKR motif, also contained some bases in their disordered N-terminal arms, resulting in a significant correlation between repression strength and the number of positively charged amino acids, arginine and lysine (R). 2 =0.85 (Figures 13C-13E).
[0140] Outside the Hox homeodomain, 99.5% of the repressor hits in the Pfam nuclear protein domain library did not contain the RKKR (SEQ ID NO: 1330) motif, whereas many of the domains without hits contained the RKKR motif. Furthermore, when the entire domain library was examined, no correlation was found between net domain charge and repression intensity on day 5 (R 2 (=0.04). In summary, these results suggest that the RKKR (SEQ ID NO: 1330) motif and charge contribute to the suppression of the Hox homeodomain in the recruitment assay, but are not sufficient for suppression when found in the context of other domains.
[0141] [Example 7] Discovery of transcription activators through high-throughput recruitment to minimal promoters Reporter K562 cell lines with a weak min CMV promoter were found to be activated when a fusion of rTetR and the activation domain was recruited (Figure 14A). To perform an activator screen, a nuclear Pfam domain library was delivered to these reporter cells using lentivirus, and rTetR-mediated recruitment was induced with doxycycline for 48 hours. The cells (Figure 21A) were magnetically separated, and the domains in the resulting two cell populations were sequenced. For each domain, the enrichment ratio of sequencing counts in the bead-bound population (on) to sequencing counts in the non-bound population (off) was calculated as a measure of transcriptional activation intensity, and domains exceeding two standard deviations of the mean of the poorly expressed negative control were considered hits (Figure 14B). The hits included FOXO-TAD derived from three already known transcriptional activation domain families present in the library: FOXO1 / 3 / 6, LMSTEN derived from Myb / Myb-A, and TORC_C derived from CRTC1 / 2 / 3. The measured activation strengths for the hits were highly reproducible among individually transduced biological replicates (r 2 =0.89, Figure 14B). This second screening using a short-chain nuclear domain library established that high-throughput recruitment can be used to measure activation or repression by altering the reporter promoter. The amino acid and nucleic acid sequences of the activators identified in the nuclear Pfam domain library are shown in Table 2, where lower scores indicate potent activators.
[0142] A total of 48 hits were found, derived from 26 domain families. Beyond the three known activating domain families mentioned above, the remaining families with activating hits had not yet been annotated as activating domains on Pfam (Figure 14C). Overall, fewer activators were found than repressors, which may simply be because activators are often disordered or low-complexity regions that are not annotated as Pfam domains (Liu et al., 2006). However, proteins containing activating domains were significantly enriched for gene ontology terms such as "positive regulation of transcription," with the strongest enrichment being for "signal transduction," which reflects that many of the source proteins for these terms are activators (Figure 21B). Furthermore, hits are a common characteristic within activating domains (Mitchell and Tjian, 1989; Staller et al., 2018), but they were significantly more acidic than those without hits (p ≤ 1 × 10⁻⁶). -5 Mann-Whitney test; Figure 14D).
[0143] Some hits did not originate from sequence-specific transcription factors where the classical activation domain was predicted, but rather from non-classical activators derived from activation cofactors and transcriptional mechanism proteins including Med9, TFIIEβ, and NCOA3. In particular, the Med9 domain (Takahashi et al., 2009), whose ortholog directly binds to other mediator complex components in yeast, was a strong activator despite its low expression level, with a mean log2 (off:on) of -5.5. Non-classical activators have already been reported to act individually in yeast (Gaudreau et al., 1999), but when recruited individually in mammalian cells, they only act weakly (Nevado et al., 1999). One exception is TATA-binding proteins (Dorris and Struhl, 2000). By screening more non-classical sequences, we found more exceptions to this concept.
[0144] For all test domains, doxycycline-dependent activation of the reporter gene was confirmed using both the 80-amino acid sequence extended from the library and the trimmed Pfam annotation domain (Figure 21C). The already annotated FOXO-TAD and LMSTEN were strong activators in both their extended and trimmed forms. The activating function of the QLQ domain, derived from the transcription factor EGR3 (DUF3446) and the SWI / SNF family's SMARCA2 protein (mostly uncharacterized), was also confirmed. Furthermore, the Dpy-30 motif domain, a DUF found within the Dpy-30 protein, was confirmed to be a weak activator. Dpy-30 is a core subunit of the histone methyltransferase complex (Hyun et al., 2017), denoted as H3K4me3, which is associated with the transcriptionally active chromatin region (Sims et al., 2003). When a total of 11 hit domains (including non-classical hits such as Nuc_rec_co-act derived from Med9 and NCOA3) were examined and the 80-amino acid sequences extended from the library were used, all were found to significantly activate the reporter. In summary, the screening and validation confirmed that the unbiased nucleoprotein domain library allows for productive re-screening to reveal domains with significantly different functions, and that a diverse set of domains beyond classical activation domains (and including DUFs) can activate transcription when recruited.
[0145] [Example 8] Discovery of the KRAB activation domain Surprisingly, the most potent activator in the library was the KRAB domain derived from ZNF473 (Figure 5B). The other three KRAB domains (derived from ZFP28, ZNF496, and ZNF597) were also activating hits, and all of them were stably expressed and not repressive. One of these domains, derived from ZNF496, had already been reported as an activator when individually recruited in HT1080 cells (Losson and Nielsen, 2010). Interestingly, ZFP28 contains two KRAB domains, with KRAB_1 being a repressor and KRAB_2 being an activator. Existing affinity purification / mass spectrometry performed on full-length ZFP28 identified significant interactions with both repressive and activating proteins (Schmitges et al., 2016). The activated KRAB domain was significantly more acidic than the inactivated KRAB (p=0.01, Mann-Whitney test, Figure 14D). Sequence analysis showed that the activated KRAB domains share homology to each other but are branched from the consensus KRAB sequence, forming a mutant KRAB subcluster (Figure 14E). Existing phylogenetic analyses have linked the mutant KRAB cluster to the lack of binding to KAP1 and its evolutionary age (Helleboid et al., 2019). More specifically, two of the activated KRAB source proteins (ZNF496 and ZNF597) have already been investigated by co-immunoprecipitation mass spectrometry, but were not found to interact with KAP1 (Helleboid et al., 2019).
[0146] Using the same 80-amino acid sequence centered on the KRAB domain used within the library, it was individually verified that KRAB derived from ZNF473 is a potent activator, and KRAB_2 derived from ZFP28 is a moderately potent activator (Figure 14F). Furthermore, while a 41-amino acid KRAB trimmed from ZNF473 was sufficient for potent activation, a 37-amino acid KRAB_2 trimmed from ZFP28 did not activate, suggesting that some of the surrounding sequence was required for activation (Figure 21C). Next, a thorough examination of the available ChiP-seq and ChIP-exo datasets (ENCODE Project Consortium et al., 2020; Imbeault et al., 2017; Najafabadi et al., 2015; Schmitges et al., 2016) revealed that, in contrast to the repressive ZNF10, ZNF473 colocalizes with the active chromatin marker, H3K27ac (Figure 14G). Manual examination revealed the most prominent ZNF473 peaks near the transcription start sites of genes (CASC3, STAT6, WASF2, ZKSCAN2) and the lncRNA (LINC00431). On the other hand, ZFP28 did not colocalize with H3K27ac, which likely indicates that its KAP1-binding repressive KRAB_1 domain was the primary effector, outweighing its activating KRAB_2 domain, which is generally of moderate intensity. When examining these individual KRAB proteins beyond their individual counterparts, zinc finger proteins containing repressive KRAB did not co-localize with H3K27ac, whereas the group of non-repressive KRAB proteins clearly contained a co-localization peak (Figure 14G). In summary, the results support the idea that mutant KRAB proteins are functionally diverse and, in some cases, function as transcription activators.
[0147] [Example 9] Tiling libraries reveal effector domains within non-annotation regions of nuclear proteins. Pfam annotation has provided a useful means of filtering the nuclear proteome to create a relatively compact library; however, Pfam currently has a high probability of missing many human effector domains. To discover effector domains within the non-annotated regions of proteins, a list of 238 proteins derived from the silencer complex was curated, and a tiling library was designed by tiling their sequences with 80 amino acids separated by a 10-amino acid tiling window (Figure 15A). High-throughput recruitment of a potent pEF reporter was performed, and the timing for measuring silencing was set after 5 days of doxycycline, and the timing for measuring epigenetic memory again was set on day 13 (8 days after doxycycline release) (Figure 22A). 4.3% of the tiles were rated as hits on day 5 (Figure 15B), and their suppression intensity measurements were reproducible (r 2=0.72, Figure 22B). In summary, the tiling screen identified short-chain repressive domains in 141 out of 238 proteins. Some of these hits included positive controls that overlapped with the annotation domain: for example, tiling ZNF57 and ZNF461 identified the KRAB domains of these transcription factors as repressive effectors, while the rest of the sequences were not identified as repressive effectors (Figure 22C). Similarly, the tiling strategy identified RYBP repressive domains annotated with Pfam, and individual validations showed that both the 80-amino acid tile and the 32-amino acid Pfam domain silenced with similar intensity and epigenetic memory (Figure 22D). Suppressors in REST (overlapping with the CoREST binding domain (Ballas et al., 2001)), DNMT3b (overlapping with the DNMT1 binding domain and the DNMT3a binding domain (Kim et al., 2002)), and CBX7 (overlapping with PcBox, which recruits PRC1 (Li et al., 2010)) were also identified and validated (Figures 22E-22G). Another category of tiling hits, although not annotated as domains within Pfam, was found in the literature to have previously reported inhibitory functions. For example, amino acids 121-220 of CTCF showed strong inhibitory function in the screen and, upon individual validation (Figures 15C and 15E), were consistent with previous recruitment studies in HeLa cells, HEK293 cells, and COS-7 cells (Drueppel et al., 2004). In summary, these results establish that high-throughput recruitment of protein tiles is an effective strategy for identifying true inhibitory domains. The amino acid sequences of the inhibitors identified in the tiling library are shown in Table 4, where high scores indicate a high degree of inhibition.
[0148] Novel, unannotated repressive domains were also discovered. For example, BAZ2A (also known as TIP5) is a component of the nuclear remodeling complex (NoRC) that mediates transcriptional silencing of certain rDNAs (Guetg et al., 2010), but it does not possess an annotation effector domain. Tiling data for BAZ2A showed a peak in repressive function within a glutamine-rich region, but it was individually validated as a moderately strong repressor (Figures 15D and 15E). Repressive tiles were found within the unannotated regions of three TETDNA demethylases (TET1 / 2 / 3). Unexpectedly, repressor tiles were also identified within the control protein DMD, which was validated by flow cytometry (Figure 22H).
[0149] It is thought that this protein represses transcription by binding to the genome in an E-box motif and recruiting the non-canonical Polycomb 1.6 complex (Blackledge et al., 2014; Jolma et al., 2013; Stielow et al., 2018). Within the MGA, tiling experiments revealed two domains with repressive function, located adjacent to two known DNA-binding domains, and referred to herein as repressor 1 and repressor 2 (Figure 15F). These repressive domains were individually examined, and remarkably different silencing dynamics and degrees of memory were observed; the first domain (amino acids 341-420) was characterized by slow silencing but strong memory, while the second domain (amino acids 2381-2460) was characterized by rapid silencing but weak memory accompanied by rapid reactivation (Figure 15G). These are thought to be the first effector domain isolated from the protein within the ncPRC1.6 silencing complex.
[0150] Next, we examined the overlaps within all tiles covering the protein regions exhibiting inhibitory function and attempted to identify the minimum sequence necessary for inhibitory function within each independent domain by determining which amino acid sequences are present within all inhibitory tiles (Figure 15H). Using this method, we generated two candidate minimizing effector domains for MGA: a 10-amino acid MGA sequence [381-390] and a 30-amino acid MGA sequence [2431-2460], both of which overlapped within the conserved region and included functionally exposed residues predicted by ConSurf. Individual validation experiments confirmed that both candidate minimizing sequences could efficiently silence the reporter (Figure 15I).
[0151] Materials and methods Cell lines and cell cultures All experiments were performed using K562 cells (ATCC CCL-243). Cells were cultured in RPMI 1640 (Gibco) medium supplemented with 10% FBS (Hyclone), penicillin (10,000 IU / mL), streptomycin (10,000 ug / mL), and L-glutamine (2 mM) at 37°C in a controlled CO2 humidified incubator. HEK293FT cells and HEK293T-LentiX cells were grown in DMEM (Gibco) medium supplemented with 10% FBS (Hyclone), penicillin (10,000 IU / mL), and streptomycin (10,000 ug / mL) and used to produce lentiviruses. Reporter cell lines were created by TALEN-mediated homology-directed repair as described below, and donor constructs were incorporated into the AAVS1 locus. 1.2 × 10 6K562 cells were electroporated in Amaxa solution (Lonza Nucleofector 2b, setting T0-16) containing 1000 ng of reporter donor plasmid and 500 ng each of TALEN-L (Addgene #35431) and TALEN-R (Addgene #35432) plasmids (targeting upstream and downstream of the intended DNA cleavage site, respectively). After 7 days, cells were treated with 1000 ng / mL of puromycin antibiotic for 5 days to select a population in which the donor was stably integrated into the intended locus, providing a promoter that expresses the PuroR resistance gene. Fluorescent reporter expression was measured by microscopy and flow cytometry (BD Accuri).
[0152] Nuclear protein Pfam domain library design We queried the UniProt database (UniProt Consortium, 2015) for human genes capable of nuclear localization. Intracellular location information in UniProt was determined from publicly available data, or by similarity if publicly available data only existed in similar genes (e.g., orthologs), and manually reviewed. Next, we searched for Pfam annotation domains using the ProDy searchPfam function (Bakan et al., 2011). Domains of 80 amino acids or less were filtered, and highly abundant and repetitive C2H2 zinc finger DNA-binding domains were excluded, as their function as transcriptional effectors was not expected. We searched for annotation domain sequences and extended them in any side chain until a total of 80 amino acids was reached. Duplicate sequences were removed, and then codon optimization was performed using human codons, removing BsmBI sites and limiting GC content to 20%–75% per 50-nucleotide window (performed using a DNA chisel (Zulkower and Rosser, 2020)). 499 random controls of 80 amino acids lacking a stop codon were computer-generated as controls. Since DMD was not thought to be a transcription regulator, 362 elements tiled with the DMD protein in an 80-amino acid tile with a 10-amino acid sliding window were also included as controls. In total, the library consists of 5,955 elements.
[0153] Silencer Tiling Library Design 216 proteins involved in transcriptional silencing were selected from a database of transcriptional regulators (Lambertf et al., 2018). 32 proteins potentially involved in transcriptional silencing were manually added, and then an unbiased protein tiling library was constructed. To do this, canonical transcripts for each gene were searched for in Ensembl BioMart (Kinsella et al., 2011) using the Python API. If a canonical transcript was not found, the longest transcript with a CDS was searched for. The coding sequences were divided into 80-amino acid tiles with a 10-amino acid sliding window between tiles. For each gene, the final tile, extending from 80 amino acids upstream of the last residue to its final residue, was included so that the C-terminal region was included in the library. Duplicate protein sequences were removed, codon optimization was performed using human codons, BsmBI sites were removed, and GC content was limited to 20%–75% per 50-nucleotide window (performed using a DNA chisel (Zulkower and Rosser, 2020)). Including 361 tiling negative controls, as in the previous library design, a total of 15,737 library elements were obtained.
[0154] KRAB Deep Mutation Scan Library Design A deep mutation scan of the ZNF10 KRAB domain sequence used in CRISPRi (Gilbert et al., 2014) was designed using all possible single substitutions and all consecutive double and triple substitutions of the same amino acid (e.g., substitutions by AAA). These amino acid sequences were backtranslated into DNA sequences using a stochastic codon optimization algorithm so that each DNA sequence contains some variation beyond the substituted residue, improving the ability to clearly align sequencing reads to unique library members. In addition, all Pfam-annotated KRAB domains derived from human KRAB found in InterPro were included, similar to those in the previous nuclear Pfam domain library. Tiling sequences designed in the previous tiling library were also included for the five KRAB zinc finger genes. 300 random control sequences and 200 tiles from the DMD gene were included as negative controls. During codon optimization, BsmBI sites were removed, and GC content was limited to 30%–70% per 80-nucleotide window (performed using a DNA chisel (Zulkower and Rosser, 2020)). The total library size was 5,731 elements.
[0155] Domain library cloning Oligonucleotides up to 300 nucleotides in length were synthesized as pooled libraries (Twist Biosciences) and then amplified by PCR. 6 × 50 μl of reaction were set up in a clean PCR hood to avoid amplification of contaminating DNA. For each reaction, 5 ng of template, 0.1 μl of 100 μM primers, 1 μl of Herculase II polymerase (Agilent), 1 μl of DMSO, 1 μl of 10 nM dNTPs, and 10 μl of 5 × Herculase buffer were used. The thermocycling protocol consisted of a final step cycle of 3 minutes at 98°C, followed by 20 seconds at 98°C, 20 seconds at 61°C, 30 seconds at 72°C, and 3 minutes at 72°C. The default number of cycles was 29, which was optimized for each library to find the fewest cycles that yielded a clean, visible product for gel extraction (in practice, 25 cycles was the minimum). After PCR, the obtained dsDNA library was loaded into ≥4 lanes of a 2% TBE gel, bands were excised at the expected length (around 300 bp), and gel extraction was performed using the QIAgen gel extraction kit. The library was cloned into the lentiviral recruitment vector pJT050 by digestion of 4 × 10 μl of GoldenGate (75 ng of gel-extracted skeleton plasmid before digestion, 5 ng of library, 0.13 μl of T4 DNA ligase (NEB, 20000 U / μl), 0.75 μl of Esp3I-HF (NEB), and 1 μl of 10 × T4 DNA ligase buffer) for 30 cycles at 37°C, followed by ligation at 16°C for 5 minutes each, then digestion for the last 5 minutes at 37°C, and finally thermal inactivation for 20 minutes at 70°C. Next, the reaction was pooled, purified using a MinElute column (QIAgen), and eluted with 6 μl of ddH2O. 2 μl per tube was transformed into 50 μl of electrocompetent cells (Lucigen DUO) in two tubes, according to the manufacturer's instructions. After harvesting, the cells were seeded into 3–7 large 10-inch × 10-inch LB plates containing carbenicillin.After overnight growth at 37°C, bacterial colonies were collected in collection bottles, and the plasmid pool was extracted using the HiSpeed Plasmid Maxiprep kit (QIAgen). Two to three small plates were prepared in parallel with diluted transformed cells to count the colonies, confirming that the transformation efficiency was sufficient to maintain at least 30× library coverage. To determine library quality, domains were amplified and sequenced from the plasmid pool and the original oligo pool by PCR using primers with extensions containing Illumina adapters. The PCR and sequencing protocols were the same as those described below for sequencing from genomic DNA, except that these PCRs used 10 ng of input DNA and 17 cycles. These sequencing data were analyzed as described below to determine the uniformity of coverage and the synthetic quality of the library. In addition, 20–30 colonies from transformation were Sanger sequenced (Quintara) to estimate cloning efficiency and the proportion of empty-skeleton plasmids in the pool.
[0156] High-throughput recruitment for measuring inhibitor activity Large-scale lentivirus production and spin-fection of K562 cells were performed. HEK293T cells were seeded into four 15 cm tissue culture plates to produce enough lentivirus to infect K562 cells with the library. In each plate, 9 × 10⁵ HEK293T cells were seeded in 30 mL of DMEM and grown overnight, then transfected with an equimolar mixture of 8 μg of three third-generation packaging plasmids and 8 μg of rTetR-domain library vector using 50 μl of polyethyleneimine (PEI, Polysciences #23966). Lentiviruses were harvested after 48 and 72 hours of incubation. The pooled lentiviruses were filtered through a 0.45 μm PVDF filter (Millipore) to remove any cell debris. For nuclear Pfam domain suppressor screening, 4.5 × 10⁵ cells were used.7 Individual K562 reporter cells were infected with a lentiviral library by spin-fection over 2 hours using two separate biological replicates of infection. Infected cells were grown for 3 days, and then selected with blastosidine (10 μg / mL, Sigma). Infection and selection efficiency was monitored daily using flow cytometry measuring mCherry (BD Accuri C6). Cells were stored in spinner flasks with a total of at least 1.5 × 10⁶ cells per replication. 8 While preserving individual cells, the cell concentration was reduced to the original 5 × 10⁻⁶ 5 Logarithmic growth conditions were maintained daily by diluting to individual cells / mL, resulting in the lowest maintenance coverage being >25,000 cells per library element (a very high coverage level that compensated for losses from incomplete blastosidine selection, library preparation, and library synthesis errors). On day 6 post-infection, recruitment was induced by treating cells with 1000 ng / ml doxycycline (Fisher Scientific) for 5 days, then the cells were spun down from doxycycline and blastosidine and maintained in untreated RPMI medium for a further 8 days, and counted up to day 13 from doxycycline addition. 2.5 × 10 8 Individual cells were collected for measurement at each time point (days 5, 9, and 13). The protocol was the same as for KRAB DMS, except that doxycycline was added on day 8 post-infection, and 2 × 10⁶ cells were measured with >12,500 × coverage. 8 ~2.2×10 8 Individual cells were collected at each time point. The protocol was the same as for the tiling screen, but 9.6 × 10⁶ cells were collected. 7 Individual cells were infected, and doxycycline was added on the 8th day after infection, at least 2 × 10⁻⁶ 8 Individual cells are maintained in each passage for >12,500× coverage, and 2×10 8 ~2.7×10 8 Individual cells were collected at each point in time.
[0157] High-throughput recruitment for measuring transcriptional activation activity For nuclear Pfam domain activator screening, lentiviruses for nuclear Pfam libraries in the rTetR(SE-G72P)-3XFLAG vector were generated for reporter screening, resulting in a size of 3.8 × 10⁻⁶. 7 K562-pDY32 minCMV reporter cells were infected with a lentiviral library by spin-fection over 2 hours using two separate biological replicates of infection. Infected cells were grown for 2 days, and then selected by blastosidine (10 μg / mL, Sigma). The efficiency of infection and selection was monitored daily using flow cytometry measuring mCherry (BD Accuri C6). Cells were stored in spinner flasks with at least 1 × 10⁶ cells per replication. 8 While leaving a total of 10 cells, the cell concentration returns to its original 5 × 10 5 Logarithmic growth conditions were maintained daily by diluting to individual cells / mL, resulting in the lowest maintenance coverage being >18,000 cells per library element. On day 7 post-infection, recruitment was induced by treating cells with 1000 ng / ml doxycycline (Fisher Scientific) for 2 days, after which the cells were spun down from doxycycline and blastosidine and maintained in untreated RPMI medium for a further 4 days. 2 × 10 8 Individual cells were collected for measurement on day 2. There was no evidence of activated memory on day 4 after doxycycline removal, as determined by the absence of citrin-positive cells by flow cytometry, and therefore no further collection was performed at any point in time.
[0158] Magnetic separation of reporter cells At each point in time, the cells were spun down at 300 × g for 5 minutes, and the culture medium was aspirated. The cells were then resuspended in the same volume of PBS (Gibco), and the spin-down and aspiration process was repeated to wash the cells and remove any IgG from the serum. Dynabeads® M-280 Protein G (ThermoFisher 10003D) was resuspended by vortexing for 30 seconds. 2 × 10 8 50 mL of blocking buffer per individual cell was prepared by adding 1 gram of biotin-free BSA (Sigma Aldrich) and 200 μl of 0.5 M pH 8.0 EDTA (ThemoFisher 15575020) to DPBS (Gibco), vacuum filtration through a 0.22 μm filter (Millipore), and holding on ice. 60 μl of beads were mixed by adding 1 mL of buffer per 200 μl of beads, vortexing for 5 seconds, sowing onto a magnetic tube rack (Eppendorf), waiting for 1 minute, removing the supernatant, and finally removing the beads from the magnet. The beads were then resuspended in 100-600 μl of blocking buffer per initial 60 μl of beads, resulting in all beads being 1 × 10⁶. 7 Individual cells were prepared. For KRAB DMS only, 30 μl of beads were added using the same method, all in 1 × 10⁶ units. 7 The preparation was done for individual cells. Beads were divided into 1 × 10⁶ units per 100 μl of resuspended beads. 7 The solution was added to cells in a sub-cell state, and then incubated at room temperature for 30 minutes while being shaken. 2 × 10 8 For samples containing individual cells, 1.2 mL of beads were used in a 15 mL Falcon tube and a large magnetic rack, and the cells were resuspended in 12 mL of blocking buffer. <5 × 10 7For samples containing individual cells, non-stacked Ambion 1.5 mL tubes and small magnetic racks were used. After incubation, the bead and cell mixture was placed on the magnetic rack for >2 minutes. The unbound supernatant was transferred to a new tube and again placed on the magnet for >2 minutes to remove any remaining beads, then the supernatant was transferred and stored as the unbound fraction. The beads were then resuspended in the same volume of blocking buffer, magnetically separated again, the supernatant was discarded, and the tube containing the beads was retained as the bound fraction. The bound fraction was resuspended in blocking buffer or PBS to dilute the cells (the unbound fraction was already diluted). Flow cytometry (BD Accuri) was performed using a small portion of each fraction to estimate the number of cells in each fraction (to ensure library coverage was maintained), and the separation was confirmed based on the citrin reporter level (the bound fraction should be >90% citrin positive, while the unbound fraction is more variable depending on the initial distribution of reporter levels). Finally, the sample was spun down, and the pellet was frozen at -20°C until genomic DNA extraction.
[0159] High-throughput measurement of domain fusion protein expression levels Expression levels were measured in K562-pDY32 cells (with citrin-off) infected with a 3×FLAG-tagged nuclear Pfam domain library. 1×10⁶ per biological repeat. 8 Individual cells were used 5 days after blastosidine selection (10 μg / mL, Sigma), which was 7 days after infection. 1 × 10 6One control K562-JT039 cell (citrinon, lentivirus-free) was added to each repeat. Fixation buffer I (BD Biosciences, BDB557870) was preheated at 37°C for 15 minutes, and permeabilization buffer III (BD Biosciences, BDB558050) and PBS (Gibco) with 10% FBS (Hyclone) were cooled on ice. A library of cells expressing the domain was collected, and the cell density was counted by flow cytometry (BD Accuri). For fixation, cells were resuspended in a volume of fixation buffer I (BD Biosciences, BDB557870) corresponding to the pellet volume at 20 μl per 1 million cells for 10–15 minutes at 37°C. Cells were washed with 1 mL of cold PBS containing 10% FBS, spun down at 500 × g for 5 minutes, and the supernatant was aspirated. Cells were permeabilized on ice for 30 minutes using Cold BD Permeabilization Buffer III (BD Biosciences, BDB558050), which was mixed by slowly adding 20 μl per 1 million cells and vortexing. The cells were then washed twice in 1 ml of PBS + 10% FBS as before, and the supernatant was aspirated. Antibody staining was performed using 5 μl / 1 × 10⁶ of α-FLAG-Alexa647 (RNDsystems, IC8529R). 6 Individual cells were used, protected from light, and the procedure was performed at room temperature for 1 hour. The cells were washed and placed in PBS + 10% FBS in 3 × 10⁶ solutions. 7Cells were resuspended at individual cell concentrations. After gateding for mCherry-positive viable cells, cells were sorted into two vials based on APC-A fluorescence (Sony SH800S) levels. A small number of unstained control cells were also analyzed in a sorter to confirm that staining was above the background. A citrin-positive cell boom was used to assess the background level of staining in cells known to lack the 3×FLAG tag, and gated for sorting was used to guide the cells above that level. After sorting, cell coverage ranged from 336 to 1,295 cells per library element across the sample. Sorted cells were spun down at 500×g for 5 minutes and then resuspended in PBS. Genomic DNA extraction was performed using the QIAgen Blood Maxi kit, >1×10⁻¹⁶, with one modification: proteinase K+AL buffer incubation was performed overnight at 56°C according to the manufacturer's instructions. 7 Used for samples containing individual cells, with the QIAamp DNA Mini Kit and up to 5 × 10⁶ 6 One column per individual cell, ≤1 × 10⁻⁶ 7 (Used for samples containing individual cells).
[0160] Library preparation and sequencing Genomic DNA, up to 1.25 × 10⁶ per column. 8Individual cells were extracted using the Blood & Tissue Kit (QIAgen) according to the manufacturer's instructions. To avoid subsequencing PCR inhibition, DNA was eluted with EB rather than AE. Domain sequences were amplified by PCR using primers containing Illumina adapters as extension. Test PCRs were performed using 5 μg of genomic DNA in 50 μl (half-size) reactions to verify that the PCR conditions produced visible bands of the expected size for each sample. Subsequently, 12–24 × 100 μl reactions were set on ice (in a clean PCR hood to avoid amplification of contaminating DNA), with the number of reactions depending on the amount of genomic DNA available in each experiment. 10 μg of genomic DNA, 0.5 μl each of 100 μM primers, and 50 μl of NEBnext 2 × Master Mix (NEB) were used in each reaction. The thermocycling protocol consisted of 32 cycles of preheating the thermocycler to 98°C, adding the sample over 3 minutes at 98°C, then a final step of 10 seconds at 98°C, 30 seconds at 63°C, 30 seconds at 72°C, and 2 minutes at 72°C. All subsequent steps were performed outside the PCR hood. The PCR reaction was pooled, and ≥140 μl was run over at least 1 hour in at least three lanes of a 2% TBE gel in parallel with a 100 bp ladder to cleave the library band around 395 bp. The DNA was purified by eluting 30 μl into a non-stick tube (Ambion) using the QIAquick Gel Extraction Kit (QIAgen). Confirmation gels were run to verify that small amounts of product had been removed. These libraries were then quantified using the Qubit HS kit (Thermo Fisher), pooled with 15% PhiX control (Illumina), and sequenced using single-ended forward reads (266 or 300 cycles) and 8-cycle index reads on Illumina NextSeq with a High Output kit.
[0161] Domain sequencing analysis Sequencing reads were demultiplexed using bcl2fastq (Illumina). Bowtie references were generated using library sequences designed by the script "makeIndices.py", and the reads were aligned to tolerate zero mismatches using the script "makeCounts.py". Enrichment for each domain between off-samples and on-samples (or FLAGhigh and FLAGlow) was calculated using the script "makeRhos.py". Domains with <5 reads in both samples for a given repeat were dropped from the repeat (assigned a count of 0), while domains with <5 reads in a single sample were adjusted to 5 to avoid a surge in enrichment values from low depths. For all nuclear domain screens, domains with ≤5 counts in both repeats under a given condition were filtered out from downstream analysis. For nuclear domain expression screens, well-expressed domains were those whose log2(FLAGhigh:FLAGlow) was ≥1 standard deviation above the median of the random controls. For the nuclear Pfam domain suppressor screen, hits were domains where log2(off:on) was ≥2 standard deviations above the mean of poorly expressed domains. For the nuclear domain activator screen, hits were domains where log2(off:on) was ≤2 standard deviations below the mean of poorly expressed domains. For the silencer tiling screen, tiles with ≤20 counts in both replicates of a given condition were filtered out, and hits were tiles where log2(off:on) was ≥2 standard deviations above the mean of the random control and DMD tiling control. Gene ontology analysis enrichment was calculated using the PantherDB web tool (www.pantherdb.org). The background set consisted of all proteins that were well-expressed and contained domains measured experimentally after count filtering was applied. P-values for statistical significance were calculated using Fisher's exact test, and the false detection rate (FDR) was calculated. Only the most significant results with FDR < 10% are shown.
[0162] Western blot and co-immunoprecipitation Cells transduced with a lentiviral vector containing rTetR-fusion-T2A-mCherry-BSD were selected using blastosidine (10 μg / mL) until mCherry content reached >80%. The cells were dissolved in lysis buffer (1% Triton X-100, 150 mM NaCl, 50 mM Tris pH 7.5, 1 mM EDTA, and a protease inhibitor cocktail). Protein levels were quantified using the DC Protein Assay Kit (Bio-Rad). Equal volumes were loaded onto gels and transferred to nitrocellulose or PVDF membranes. The membrane was explored using either GATA1 antibody (1:1000, rabbit, Cell Signaling Technologies catalog number 3535S) and GAPDH antibody (1:2000, mouse, ThermoFisher catalog number AM4300) or FLAG M2 monoclonal antibody (1:1000, mouse, Sigma-Aldrich, catalog number F1804) and histone 3 antibody (1:1000, mouse, Abcam catalog number AB1791) as primary antibodies. As secondary antibodies, donkey anti-rabbit IRDye 680 LT and goat anti-mouse IRDye 800CW (1:20,000 dilution, LI-COR Biosciences, catalog numbers 926-68023 and 926-32210, respectively) or goat anti-mouse IRDye 680 RD and goat anti-rabbit IRDye 800CW (1:20,000 dilution, LI-COR Biosciences, catalog numbers 926-68070 and 926-32211, respectively) were used.
[0163] The blots were imaged using LiCor Odyssey CLx. Band intensity was quantified using ImageJ.
[0164] Individual inhibitor recruitment assays Individual effector domains were cloned as fusions with rTetR or rTetR(SE-G72P) with or without a 3×FLAG tag upstream of the T2A-mCherry-BSD marker (see legend in figure) using GoldenGate cloning to the pJT050 or pJT126 scaffold. K562-pJT039-pEF-citrin reporter cells were then transduced with this lentiviral vector and selected by blastosidine (10 μg / mL) after 3 days until >80% of the cells were mCherry-positive (6-7 days). Cells were divided into separate wells of a 24-well plate and treated with or left untreated with doxycycline (Fisher Scientific). Five days after treatment, doxycycline was removed from the cells by spin-down, the medium was replaced with DPBS (Gibco) to dilute any remaining doxycycline, and then the cells were spun down again and transferred to fresh medium. Time points were measured every 2-3 days by flow cytometry analysis of >7,000 cells (either BD Accuri C6 or Beckman Coulter CytoFLEX). Data were analyzed using Cytoflow and a custom Python script. Events were gated for viability and for mCherry as a delivery marker. To calculate the off-cell fraction during doxycycline treatment, a two-component Gaussian mixture model was fitted to untreated rTetR-only negative control cells, then to both on-peak and background silenced off-cell subpopulations, and then a threshold was set two standard deviations below the on-peak mean to display silenced cells as off. Normalized percentages of the cell background were calculated using time-matched untreated controls. cell オフ、正規化 =cell オフ+ドキシサイクリン / (1-cell オフ、非処理We used two independently transduced physical replicates. A gene silencing model consisting of exponential decay during the doxycycline treatment phase (e.g., exponential decay subtracted from 1), and an increasing form of exponential decay during the doxycycline removal phase with additional parameters for lag time before silencing and reactivation initiation, was fitted to normalized data using SciPy.
[0165] Individual activator recruitment assays The domain was cloned as a fusion with rTetR (SE-G72P) upstream of the T2A-mCherry-BSD marker using GoldenGate cloning in the pJT126 skeleton. K562 pDY32 minCMV citrin reporter cells were then transduced with each lentiviral vector, and after 3 days, cells were selected by blastosidine (10 μg / mL) until >80% were mCherry positive (6-7 days). Cells were divided into separate wells of a 24-well plate and treated with doxycycline or left untreated. Time points were measured by flow cytometry analysis of >15,000 cells (Biorad ZE5). To calculate the on-cell fraction during doxycycline treatment, a Gaussian model was fitted to negative control cells with only untreated rTetR, then fitted to the off-peak, and a threshold was set two standard deviations above the mean of the off-peak to display cells activated as on. We used two independently transfected biological replicates.
[0166] Flow cytometry of FLAG-tagged protein levels Staining was performed at the FLAG-tagged fusion protein level. Specifically, K562 cells were transduced using lentivirus to express the fusion protein, selected by blastoscidin, and then fixed with Fixation Buffer I (BD Biosciences) at 37°C for 15 minutes. The cells were washed once with cold PBS containing 10% FBS and then permeabilized on ice for 30 minutes using Perm Buffer III (BD Biosciences). The cells were washed twice and then stained with anti-FLAG(XX) at 4°C for 1 hour. After the final round of washing, flow cytometry was performed using a CytoFLEX (Beckman Coulter) flow cytometer. The data were analyzed by CytoFlow by gatering the cells for mCherry expression, and then the FLAG-tagged protein levels in mCherry+ and non-transduced cells were plotted. A control of this method for variability in staining efficiency was mixed in the same sample as two cell groups.
[0167] Phylogenetic analysis and alignment analysis Using surrounding natural sequences, KRAB and homeodomain sequences were searched for and extracted from Pfam until 80AA was reached. Well-expressed domains were selected for alignment. Phylogenetic trees and sequence alignments were obtained using the alignment website Clustal Omega (McWilliam et al., 2013; Sievers et al., 2011) with default parameters, and 52 phylogenetic neighbor-joined trees without distance correction were constructed using the default parameters in Jalview (Waterhouse et al., 2009). Alignment visualization was performed in Jalview.
[0168] Analysis of amino acid residue conservation Protein sequences were submitted to the ConSurf web server and analyzed using the ConSeq method. Briefly, ConSeq selects up to 150 homologs for multiple string alignment by sampling from a list of homologs with 35–95% sequence identity. The phylogenetic tree is then reconstructed, and conservation is scored using Rate4Site. ConSurf provides normalized scores, resulting in a mean score of zero and a standard deviation of 1 for all residues. The conservation score calculated by ConSurf is a relative measure of evolutionary conservation of each residue in the protein, with the lowest score representing the most conserved position in the protein. The uniqueness of ZNF10 KRAB N-terminal elongation was determined by exploring protein BLAST for all human proteins and other zinc finger proteins in BLAST matches (Johnson et al., 2008).
[0169] ChiP-seq and ChiP-exo analysis External ChIP datasets were searched from multiple sources. ENCODE ChIP-seq data were processed using the uniform processing pipeline of ENCODE (ENCODE Project Consortium et al., 2020), and narrow peaks below an IDR threshold of 0.05 were searched. KRAB ZNF ChIP-exo data from tagged KRAB ZNF overexpression in HEK293 cells, and KAP1 ChIP-exo data from H1 hESCs were obtained from GEO accession GSE78099 (Imbeault et al., 2017). Reads were trimmed to a uniform length of 36 base pairs using Bowtie (version 1.0.1; (Langmead et al., 2009)) and mapped to the hg38 version of the human genome, allowing for the retention of a maximum of 2 mismatches and unique alignments. Peaks were called using MACS2 (version 2.1.0) (Feng et al., 2012) with the following settings: "-g hs-f BAM--keep-dup all--shift-75--extsize 150--nomodel". Browser tracks were generated using a Python script. For some KRAB ZNFs for which ChIP-exo data was not available, ChIP-seq data from overexpression of tagged KRAB ZNFs in HEK293 cells were obtained from GEO accessions GSE76496 (Schmitges et al., 2016) and GSE52523 (Najafabadi et al., 2015). KRAB ZNF peaks were defined as sole binding sites if other KRAB ZNFs in the dataset did not have peaks less than 250 base pairs apart. The ENCODE H3K27ac ChIP-seq dataset for H1 cells was processed using the ENCODE pipeline (ENCODE Project Consortium et al., 2020). Narrow peaks were called using MACS2, and peaks below an IDR threshold of 0.05 were searched for.
[0170] External dataset ChIP-seq and ChIP-exo data for KRAB ZNF, KAP1, and H3K27ac (ENCODE Project Consortium et al., 2020; Imbeault et al., 2017; Najafabadi et al., 2015; Schmitges et al., 2016), the evolutionary age of KRAB ZNF genes (Imbeault et al., 2017), co-immunoprecipitation / mass spectrometry data of KRAB ZNF proteins (Helleboid et al., 2019), and CAT assays for KRAB inhibitory activity (Margolin et al., 1994; Witzgall et al., 1994) were retrieved from previously published studies.
[0171] All references, including publications, patent applications, and patents, cited herein are incorporated by reference herein in their entirety to the same extent as if each reference were individually and specifically indicated to be incorporated by reference and as fully set forth herein.
[0172] Preferred embodiments of the present invention are described herein, including the best mode known to the inventors for carrying out the invention. Variations to these preferred embodiments will be apparent to those skilled in the art upon reading the foregoing description. The inventors expect those skilled in the art to employ such variations as appropriate, and the inventors intend for the invention to be practiced in a manner other than that specifically described herein. Accordingly, the invention includes all modifications and their equivalents to the objects recited in the claims appended hereto as permitted by applicable law, and further, unless otherwise indicated herein or otherwise indicated by context to the contrary, any combination of all possible variations of the elements described above is also included in the present invention.
[0173] [Table 1] TIFF0007869753000002.tif251166TIFF0007869753000003.tif250164TIFF0007869753000004.tif251167TIFF0007869753000005.tif250166TIFF0007869753000006.tif251167TIFF0007869753000007.tif250164TIFF0007869753000008.tif251166TIFF0007869753000009.tif251164TIFF0007869753000010.tif250162TIFF0007869753000011.tif250166TIFF0007869753000012.tif252166TIFF0007869753000013.tif251164TIFF0007869753000014.tif251164TIFF0007869753000015.tif251162TIFF0007869753000016.tif251166TIFF0007869753000017.tif249163TIFF0007869753000018.tif251163TIFF0007869753000019.tif250162TIFF0007869753000020.tif251163TIFF0007869753000021.tif250162TIFF0007869753000022.tif251164TIFF0007869753000023.tif250164TIFF0007869753000024.tif251164TIFF0007869753000025.tif249165TIFF0007869753000026.tif250164TIFF0007869753000027.tif250163TIFF0007869753000028.tif250163TIFF0007869753000029.tif250163TIFF0007869753000030.tif249164TIFF0007869753000031.tif249163TIFF0007869753000032.tif250163TIFF0007869753000033.tif250161TIFF0007869753000034.tif250162TIFF0007869753000035.tif251162TIFF0007869753000036.tif251167TIFF0007869753000037.tif251164TIFF0007869753000038.tif250164TIFF0007869753000039.tif250164TIFF0007869753000040.tif251162TIFF0007869753000041.tif250162TIFF0007869753000042.tif250161TIFF0007869753000043.tif249164TIFF0007869753000044.tif250158TIFF0007869753000045.tif249167TIFF0007869753000046.tif249162TIFF0007869753000047.tif250161.
[0174]
Table 2
[0175]
Table 3
[0176]
Table 4
[0177]
Table 5
Claims
1. A method for identifying transcriptional repression domains or transcriptional activation domains, a) A step of preparing a domain library comprising multiple nucleic acid sequences configured to express a fusion protein, each containing a protein domain linked to an inducible DNA-binding domain; b) A step of transforming reporter cells with a domain library, wherein the reporter cells contain a bipartite reporter gene comprising a surface marker and a fluorescent protein, A bipartite reporter gene is under the control of a promoter that confers a high transcription rate and can be silenced by a putative transcriptional repression domain after treatment with a first agent configured to induce an inducible DNA-binding domain, or A bipartite reporter gene is under the control of a promoter that confers a low transcription rate and can be activated by a putative transcriptional activation domain after treatment with a second agent configured to induce an inducible DNA-binding domain. Step; c) A step of treating reporter cells with a first agent for a period of time necessary for the degradation of intracellular proteins and mRNA, or treating reporter cells with a second agent for a period of time necessary for the production of intracellular proteins and mRNA; d) A step of isolating reporter cells based on the presence or absence of surface markers, fluorescent proteins, or combinations thereof; e) A step of sequencing protein domains from isolated reporter cells; f) For each protein domain sequence, calculate the ratio of sequencing counts from reporter cells without a surface marker, fluorescent protein, or combination thereof to sequencing counts from reporter cells with a surface marker, fluorescent protein, or combination thereof; and g) Step of identifying the protein domain as a transcription repressor or transcription activator. Includes, A method in which a protein domain is identified as a transcription repressor when the ratio of log2 is at least two standard deviations from the mean of a poorly expressed negative control, and a protein domain is identified as a transcription activator when the ratio of log2 is at least two standard deviations from the mean of a low-expression negative control.
2. The method according to claim 1, further comprising the step of stopping the treatment of reporter cells with a first or second agent, and repeating steps d to g one or more times.
3. The method according to claim 2, wherein steps d to g are repeated for at least 48 hours after the treatment of reporter cells with the first or second agent is stopped.
4. The method according to any one of claims 1 to 3, wherein each protein domain has 80 amino acids or less.
5. The method according to any one of claims 1 to 4, wherein the protein domain is derived from a nuclear localized protein.
6. The method according to any one of claims 1 to 5, wherein the protein domain comprises the amino acid sequence of a wild-type protein domain derived from a nuclear localized protein.
7. The method according to any one of claims 1 to 5, wherein the protein domain includes a mutant amino acid sequence of a protein domain derived from a nuclear localized protein.
8. The method according to any one of claims 1 to 7, wherein the inducible DNA-binding domain includes a tag.
9. The method according to any one of claims 1 to 8, further comprising the step of measuring the expression level of a protein domain.
10. The method according to claim 9, wherein the expression level is determined by measuring the relative presence or absence of a tag on the DNA-binding domain.
11. The method according to any one of claims 1 to 10, wherein reporter cells are treated with a first agent for at least 5 days.
12. The method according to any one of claims 1 to 11, wherein reporter cells are treated with a second agent for at least 24 hours.