Methods for parallel measurment of protein-DNA binding within cells

The method using nucleic acid molecule libraries with primer recognition sites and fusion proteins allows for high-throughput, parallel measurement of TF binding to DNA sequence variants, addressing the challenge of understanding motif interactions and genomic binding specificity, and identifying inhibitors.

WO2025158445A1PCT designated stage Publication Date: 2025-07-31YEDA RES & DEV CO LTD
View PDF 9 Cites 0 Cited by

Patent Information

Application Number
PCT/IL2025/050095
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-01-28
Filing Date
2025-01-28
Publication Date
2025-07-31

AI Technical Summary

Technical Problem

Current methods are limited in their ability to perform parallel measurements of transcription factor (TF) binding to thousands of DNA sequence variants within cells, which is necessary for understanding motif interactions and genomic binding specificity.

Method used

A method involving nucleic acid molecule libraries with 5' and 3' primer recognition sites, contacted with cells expressing fusion proteins, followed by incubation, amplification, and sequencing to determine protein association with nucleotide sequences, allowing for the identification of protein-DNA interactions and inhibitor detection.

Benefits of technology

Enables high-throughput, parallel measurement of TF binding to DNA sequence variants, revealing insights into motif interactions and genomic binding specificity, and identifying inhibitors of protein-DNA interactions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure IL2025050095_31072025_PF_FP_ABST
    Figure IL2025050095_31072025_PF_FP_ABST
Patent Text Reader

Abstract

Nucleic acid molecule libraries comprising 5' and 3' primer recognition sites and a plurality of molecules comprising a target sequence molecule and a plurality of variant sequence molecules are provided. Kits comprising the nucleic acid molecule libraries are also provided as are methods of determining protein association with a nucleotide sequence and identifying an inhibitor of protein-DNA interaction.
Need to check novelty before this filing date? Find Prior Art

Description

METHODS FOR PARALLEL MEASURMENT OF PROTEIN-DNA BINDINGWITHIN CELLSCROSS REFERENCE TO RELATED APPLICATIONS

[0001] This application claims the benefit of priority of Israeli Patent Application No. 310476 titled “METHODS FOR DETERMINING TRANSCRIPTION FACTOR TO MOTIF BINDING”, filed January 28, 2024, the contents of which are all incorporated herein by reference in their entirety.REFERENCE TO AN ELECTRONIC SEQUENCE LISTING

[0002] The contents of the electronic sequence listing (YEDA-P-031-PCT.xml; Size: 29,466 bytes; and Date of Creation: January 13, 2025) is herein incorporated by reference in its entirety.FIELD OF INVENTION

[0003] The present invention is in the field of transcription factor analysis.BACKGROUND OF THE INVENTION

[0004] To regulate gene expression, transcription factors (TFs) bind short DNA motifs in gene regulatory regions. In bacteria, binding motifs are long enough to provide precise addressing within the genome. However, in eukaryotes, genomes are longer and binding motifs are shorter, so that only a fraction of motif occurrences are regulatory relevant and TF-bound. TF binding locations in eukaryotes are therefore not determined solely by the presence of binding motifs but depend on additional properties that distinguish bound sites from unbound ones.

[0005] The prevalent model assumes that TF-bound motifs are distinguished by proximity to other nearby motifs. This model postulates that stable DNA binding is only possible when combinations of TFs co-bind at adjacent sites, which, in turn, is only possible in locations containing all respective motifs. In support of that, regulatory enhancers contain closely spaced motifs and bind multiple TFs. Yet binding of multiple TFs to the same regulatoryregion may serve purposes other than cooperative binding, including sensing multiple signals transduced by those different TFs, or cooperation of the bound TFs in the recruitment of various co-factors.

[0006] The use of motif combinations for increasing the low information content of a single motif is compelling, yet its broad applications for directing TF bindings is perhaps questioned by findings showing that the relative locations, distances, or orientations of those motifs diverge rapidly in evolution, contrasting the tight conservation of the composite motifs bound by obligate TF dimers. It should be noted in this context that composite motifs bound by such obligate heterodimers (e.g., of the bZIP and bHLH families) are still short and of low information content.

[0007] Experimentally, testing whether TFs can only bind at sites containing closely spaced motifs requires comparing TF binding in the presence or absence of the respective motif combinations. This was done in certain cases but has not yet been attempted at the genomic scale due to limited capacity for measuring TF binding to libraries of designed DNA sequences. There is a great need for a method for parallel measurements of TF binding to thousands of DNA sequence variants within cells which would allow for the measurement of motif interactions.SUMMARY OF THE INVENTION

[0008] The present invention provides nucleic acid molecule libraries comprising 5’ and 3’ primer recognition sites and a plurality of molecules comprising a target sequence molecule and a plurality of variant sequence molecules. Kits comprising the nucleic acid molecule libraries are also provided as are methods of determining protein association with a nucleotide sequence and identifying an inhibitor of protein-DNA interaction.

[0009] According to a first aspect, there is provided a method of determining protein association with a nucleotide sequence, the method comprising: a. receiving a nucleic acid molecule library comprising a plurality of nucleic acid molecules wherein each molecule comprises a 5’ primer recognition site and a 3’ primer recognition site which are common to all molecules of the plurality and wherein the plurality of molecules comprises: i. a target sequence molecule comprising a putative target sequence of the protein;ii. a plurality of variant sequences molecules wherein each variant sequence molecule comprises a variant target sequence comprising at least one single nucleotide change in the putative target sequence, such that each variant sequence differs from another variant sequence by a single nucleotide change and all possible mutations within the putative target sequence are made; b. contacting the library with a cell expressing a fusion protein, wherein the fusion protein comprises the protein fused to a nuclease; c. incubating the library in the cell for a period of time sufficient for cleavage of nucleic acid molecules bound by the fusion protein; d. amplifying nucleic acid molecules in the cell with the 5’ primer and the 3’ primer to produce amplification products comprising both the 5’ primer recognition site and the 3 ’ primer recognition site; and e. sequencing the amplification products to produce sequencing data, wherein variant target sequences depleted in the sequencing data as compared to a control sequence are sequences bound by the protein; thereby determining protein association with a nucleotide sequence.

[0010] By another aspect, there is provided a method of determining CRISPR associated protein (CAS) in complex with a guide RNA (gRNA) association with a nucleotide sequence, the method comprising: a. receiving a nucleic acid molecule library comprising a plurality of nucleic acid molecules wherein each molecule comprises a 5’ primer recognition site and a 3’ primer recognition site which are common to all molecules of the plurality, a guide RNA (gRNA) which is common to all molecules and wherein the plurality of molecules comprises: i. a target sequence molecule comprising a putative target sequence of the CAS-gRNA complex; ii. a plurality of variant sequences molecules wherein each variant sequence molecule comprises a variant target sequence comprising at least one single nucleotide change in the putative target sequence, such that each variant sequence differs from another variant sequence by a single nucleotide change and a plurality of mutations within the putative target sequence are made;b. contacting the library with a cell expressing a fusion protein, wherein the fusion protein comprises a dead CAS (dCAS) in which its nuclease activity is abolished fused to a nuclease; c. incubating the library in the cell for a period of time sufficient for cleavage of nucleic acid molecules bound by the fusion protein; d. amplifying nucleic acid molecules in the cell with the 5’ primer and the 3’ primer to produce amplification products comprising both the 5’ primer recognition site and the 3 ’ primer recognition site; and e. sequencing the amplification products to produce sequencing data, wherein variant target sequences depleted in the sequencing data as compared to a control sequence are sequences bound by the CAS-gRNA complex; thereby determining CAS-gRNA complex association with a nucleotide sequence.[Oi l] According to another aspect, there is provided a method of determining nucleotide sequence association with a peptide, the method comprising: a. receiving a nucleic acid molecule library comprising a plurality of nucleic acid molecules wherein each molecule comprises i. a 5’ primer recognition site and a 3’ primer recognition site which are common to all molecules of the plurality; ii. a coding sequence encoding a fusion protein, wherein the fusion protein comprises a peptide fused to a nuclease and wherein each peptide comprises at least one amino acid difference from every other peptide; and iii. a target sequence which is common to all molecules of the plurality; wherein at least one peptide of the library binds to the target sequence; b. contacting the library with a cell; c. incubating the library in the cell for a period of time sufficient for producing of the fusion protein and cleavage of nucleic acid molecules bound by the fusion protein; d. amplifying nucleic acid molecules in the cell with the 5’ primer and the 3’ primer to produce amplification products comprising both the 5’ primer recognition site, and the 3’ primer recognition site; ande. sequencing the amplification products to produce sequencing data, wherein peptides encoded by coding sequences depleted in the sequencing data as compared to a control sequence are peptides that bind the target sequence; thereby determining nucleotide sequence association with a peptide.

[0012] According to some embodiments, the nucleic acid molecules are DNA molecules.

[0013] According to some embodiments, the protein is a DNA binding protein and the nucleic acid molecules are DNA molecules.

[0014] According to some embodiments, the protein is a transcription factor.

[0015] According to some embodiments, the protein is a protein complex.

[0016] According to some embodiments, the protein complex is a CAS protein in complex with a guide RNA (gRNA).

[0017] According to some embodiments, the gRNA is reverse complementary to the target sequence.

[0018] According to some embodiments, the fusion protein comprises a CAS protein that is a dead CAS (dCAS) in which its nuclease activity is abolished, and wherein the dCAS is fused to a nuclease.

[0019] According to some embodiments, the target sequence or variant of the target sequence is between the 5’ primer recognition site and a 3’ primer recognition site and wherein the amplification products further comprise the target sequence or variant of the target sequence.

[0020] According to some embodiments, the method is a method of detecting off-target binding of a CAS-gRNA protein complex to variants of its target sequence, wherein variant target sequences depleted in the sequencing data as compared to control sequences are off- target sequences bound by the CAS-gRNA protein complex.

[0021] According to some embodiments, the plurality of mutations is all possible mutations within the putative target sequence.

[0022] According to some embodiments, the nuclease is micrococcal nuclease (MNase).

[0023] According to some embodiments, the sufficient time is at most 600 seconds.

[0024] According to some embodiments, the sufficient time is at most 300 seconds.

[0025] According to some embodiments, the sufficient time is at most 60 seconds.

[0026] According to some embodiments, the sequencing is next generation sequencing, high-throughput sequencing or massively parallel sequencing.

[0027] According to some embodiments, the relative depletion of a given sequence is proportional to the occupancy of the protein at the given sequence.

[0028] According to some embodiments, the relative depletion of a given sequence is proportional to the occupancy of the CAS at the given sequence.

[0029] According to some embodiments, the method further comprises contacting the cell comprising the library with a substrate that induces the nuclease activity of the fusion protein.

[0030] According to some embodiments, the substrate is selected from calcium, magnesium and sodium.

[0031] According to some embodiments, the substrate is calcium.

[0032] According to some embodiments, the control sequence is a sequence not bound by the protein.

[0033] According to some embodiments, the control sequence is a plurality of the other variant sequences.

[0034] According to some embodiments, the plurality of molecules each contain a transcription factor binding site adjacent to a 5’ end and a 3 ’end of the putative target sequence or variant target sequence.

[0035] According to some embodiments, the plurality of molecules each contain a Rebl binding site adjacent to a 5’ end and a 3 ’end of the putative target sequence or variant target sequence.

[0036] According to some embodiments, the transcription factor is Rebl.

[0037] According to some embodiments, the Rebl binding site comprises or consists of CGGGTAA or its reverse complement.

[0038] According to some embodiments, the coding sequence further encodes a selection protein and wherein the selection protein is separated from the fusion protein by a ribosomal skipping sequence, optionally wherein the ribosomal skipping sequence is selected from P2A, E2A, T2A and F2A.

[0039] According to some embodiments, the cell cannot grow in a first condition and the selection protein is a protein that allows growth in the first condition.

[0040] According to some embodiments, the first condition is absence of uracil and the selection protein is orotidine 5-phosphate decarboxylase (URA3).

[0041] According to another aspect, there is provided a method of identifying an inhibitor of protein-DNA interaction, the method comprising performing a method of the invention wherein the contacting and incubating is performed in the presence of a test molecule and in the absence of the test molecule, and comparing the sequencing data produced in the presence of the test molecule to sequencing data produced in the absence of the test molecule and wherein an increase in sequences of amplification products in the sequencing data produced in the presence of the test molecule as compared to sequences of amplification products in the sequencing data produced in the absence of the test molecule indicates the test molecule is an inhibitor of protein-DNA interaction, thereby identifying an inhibitor of protein-DNA interaction.

[0042] According to another aspect, there is provided a method of identifying an inhibitor of CAS-gRNA complex off-target binding, the method comprising performing a method of the invention wherein the contacting and incubating is performed in the presence of a test molecule and in the absence of the test molecule, and comparing the sequencing data produced in the presence of the test molecule to sequencing data produced in the absence of the test molecule and wherein an increase in sequences of amplification products in the sequencing data produced in the presence of the test molecule as compared to sequences of amplification products in the sequencing data produced in the absence of the test molecule indicates the test molecule is an inhibitor of CAS-gRNA complex off-target binding, thereby identifying an inhibitor.

[0043] According to some embodiments, the protein is a transcription factor.

[0044] Further embodiments and the full scope of applicability of the present invention will become apparent from the detailed description given hereinafter. However, it should be understood that the detailed description and specific examples, while indicating preferred embodiments of the invention, are given by way of illustration only, since various changes and modifications within the spirit and scope of the invention will become apparent to those skilled in the art from this detailed description.BRIEF DESCRIPTION OF THE DRAWINGS

[0045] Figures 1A-1G: MPBA allows the sensitive detection of TF binding to a library of DNA sequence variants within cells (1A) Selecting TFs and regulatory regions for testing motif interactions: To select candidates for cooperation, we searched for regulatory regions that contain multiple motifs bound by different TFs. For this, we used our lab dataset, which describes the genome-wide binding profiles of 145 (95%) of budding yeast TFs, collected using ChEC-seq. From those profiles, we focused on the 63 TFs showing a strong binding preference for their in-vitro motif (1), leaving out weak motif binders and the fungi- specific zinc-cluster TFs localizing to low-information motifs. For each of these TFs, we then selected the top-bound promoters and filtered them by the presence of the respective TF in vitro preferred motif showing increased signal at its proximity. We screened those regions manually to detect 164 bp segments that include not only the motif bound by the TF that directed us to this region, but also motifs associated with other TFs, showing enriched signal in their respective binding profiles (2). Each of the selected regulatory regions was used to generate a library of sequence variants, which includes the wild-type sequence and all combinations of mutated motifs, as demonstrated (3). For this, we defined for each motif the central position with the highest information content based on the TF in-vitro probability weight matrix (see the ‘Materials and methods’ section). For this position, we considered two variants: leaving it intact or generating a mutation, the latter expected to abolish the binding of the associated TF. (IB) Experimental procedure of MPBA: To measure the in- vivo binding of a TF of interest to a library of DNA sequence variants, we followed the steps shown. First, we fuse the TF of interest to a micrococcal nuclease (MNase). Cells containing this TF-MNase fusion are then transformed with a library of plasmids, each containing one DNA sequence variant (1). Next, cells are subjected to a short pulse of calcium, which activates the MNase (2). Sequences bound by the TF-MNase fusion are then cleaved, while the unbound variants remain intact. PCR amplification is used to enrich for uncleaved sequences, whose abundance is defined using high-throughput sequencing and compared to their abundance prior to MNase activation (3). Of note, multiple libraries based on distinct regulatory regions can be transformed into the same TF-MNase bearing strain and measured in a single experiment. (1C) Rebl binding at the AQR1 regulatory region depends on its in- vitro defined motif: The AQR1 regulatory sequence is shown on top, with all detected TF- bound motifs considered in our library design indicated in circles. The Rebl ChEC-Seq genomic binding signal at this region is shown in purple (see the ‘Materials and methods’ section). The distance of the selected region from the AQR1 transcription start site is also noted. A library containing all motif mutation combinations was transformed into a strain bearing a Rebl-MNase fusion. Shown in the lower panel are the analysis steps used to definethe Rebl occupancy at each sequence. First, the distribution of sequence reads was examined before and after triggering MNase based cleavage (left). Next, the post-cleavage abundance of each sequence was normalized by its pre-cleavage abundance (time 0) and then by the fully mutated sequence, converting normalized abundance to occupancy (middle; see the ‘Materials and methods’ section). Shown are the Rebl occupancy values at each of the library variants, comparing two MNase activation times (60 and 180 s). The intact sequence, containing all motifs, and the fully mutated sequence, lacking all motifs, are indicated (grey and black, respectively). As a further visualization of this data, the occupancy at sequences containing a single intact motif (top right; the variant in which only the Rebl motif is intact is marked) and the occupancy at sequences in which only a single motif is mutated (right bottom; the variant in which only the Rebl motif is mutated is marked) were plotted. Note that in the example shown, the single self-motif fully explained the high (>90%) Rebl occupancy of this region. (ID) Abfl binding to the MET 13 regulatory region depends on its own motif and increases over time: Shown on top is the MET 13 regulatory region, presentation as in Figure 1C, with the Abfl genomic signal presented in red. Shown on the left are the Abfl occupancies at sequences in which only one motif is intact (top) and ones in which only one motif is mutated (bottom). Shown on the right is the Abfl occupancy at the sequence containing only its motifs intact (light grey) and at the fully intact sequence (dark grey), as a function of MNase activation time. Error bars indicate the standard error of the mean (SEM) between experimental repeats. (IE) The GRFs bind promoters depleted of other TFs: The top bound promoters of each GRF were searched for bound GRF motifs found in proximity to another TF motif enriched in binding signal of the respective TF (distance <90 bps, Methods). Shown are the number of additional bound TF motifs found in proximity to all identified GRF bound motifs. Note that in most cases, there are no additional TFs bound at sites found in proximity to the GRF binding site. (IF) Regulatory regions selected for measuring motif-dependent GRF binding: We selected six regulatory regions bound by at least one of the three GRFs. Shown are the ChEC-seq binding signals across the selected regions, GRFs are indicated (Methods), where the signal of each GRF is normalized to its maximal signal received in this region. Distances of the selected regions from the TSS of the nearest genes are noted. The locations of all identified TF motifs are displayed as circles, while motifs corresponding to a GRF are shown. Based on these motifs, six libraries of combinatorial mutations were generated. The number of sequences in each library (32- 128) is also indicated. These libraries were tested for the binding of the relevant GRFs. (1G) GRFs reach high occupancy at individual motifs and exhibit high reproducibility: In each panel, shown on the left is the sequence abundance correlation for the indicated library boundby the indicated TF in each time point and for all technical repeats (3 per time point). The middle of each panel shows the occupancy of the GRF at each library sequence (dots) at the indicated MNase activation time -points. The number of intact GRF motifs in each library sequence is indicated. The dots representing the fully intact sequence (and its corresponding values) and the fully mutated variant are indicated in red and black respectively. Variants in which the motifs of the tested GRF are intact and those of other TFs are mutated, and variants in which the tested GRF motifs are mutated while other TF motifs are intact are both represented. On the right of each panel is the relevant promoter (top; presentation as in IE) and two heatmaps: The upper one shows the GRF occupancy at sequences in which all positions are mutated except for a single one, represented by each column, and the lower heatmap shows the GRF occupancy at sequences in which all positions are intact except for a single one.

[0046] Figures 2A-2K: The GRFs rarely cooperate to localize to their binding sites. (2A) TF mediated nucleosome eviction: A scheme describing the scenario of nucleosome eviction as a consequence of TF binding at its binding site. H3 histone was fused to an MNase and assayed using MPBA. (2B) H3 binding is increased when eliminating GRF binding: For each GRF library, shown is the fold change of H3 occupancy between the wildtype sequence and the sequence in which all variable positions are mutated. Error bars indicating SEM calculated for all repeats (3) in each experiment. The relevant GRFs are indicated for each library. (2C) The binding motifs of the GRFs Rebl and Abfl inhibit H3 histone binding to the Crcl promoter region: The binding at all individual sequences in the Crcl library is shown for Rebl (x-axis) and H3 (y-axis). The mode (intact / mutated) of the Abfl motif found in this region is shown and the number of Rebl intact motifs is indicated by dot size. Error bars represent the SEM calculated for all repeats. The wild-type and fully mutated sequences are marked. Note the inversed correlation between Rebl and H3 binding, reflecting the Rebl function in nucleosome eviction and the additional nucleosomedepleting effect attributed to Abfl binding (color). (2D-2E) The effect of the GRFs on nucleosome binding corresponds to the difference between in-vivo and in-vitro nucleosome occupancy: For each GRF, shown in (2D) is the mean in-vivo and in-vitro signal around the 200 most bound motif occurrences in promoters, measured by Kaplan et al (motifs are defined in Table 1; see the ‘Materials and methods’ section for promoter definition). (2E) The mean fold change between these signals is summarized. The fold change for each tested library is shown in dots and the mean fold change over the libraries relevant to each GRF is indicated. The line indicates the mean nucleosome-binding fold change between the wildtype and the fully mutated sequence for all MPBA-GRF tested libraries. (2F) Cross-motif dependencies: A scheme demonstrating three different modes of TF interactions. While recruitment (top) can occur independently of proximal motifs, cooperativity (middle) and inhibition (bottom) require an interaction between nearby binding sites. Only the two latter modes are able to increase TF genomic binding specificity. (2G-2H). The binding of the GRFs has little effect on TF binding to proximal motifs: (2G) TFs may rely on the GRFs for binding at their own motifs as the GRFs evict nucleosomes. To examine this, the binding of 10 additional TFs whose motifs are present within the designed libraries shown in Figure IF we measured. (2H) To measure the effect of the GRFs on the binding of those other TFs (or of other GRFs), the mean TF occupancies (at 180-s MNase activation time) at sequences containing (y-axis), or lacking (x-axis) the respective GRF motifs were compared. These comparisons reveal only a single case of recruitment, in which a Rebl motif allows Abfl occupancy at the CRC 1 regulatory region, and some cases of mild inhibition, including the inhibitory effect of the Rapl motifs found in the GPP1 regulatory region on the Abfl binding. Overall, the binding of all measured TFs appears to be independent of the GRF binding. (21) Defining motif-cooperativity score: motif cooperation refers to cases in which the presence of a motif pair stabilizes the binding of a TF, beyond the individual contributions of each motif. To measure cooperation, one first defines the expected occupancy for sequences containing a pair of motifs, using sequences with only one of the two motifs intact, assuming independent binding to each individual motif (see the ‘Materials and methods’ section). In the schematic example shown on the bottom left, sequences in which only one of the motifs is present while the other is mutated reach 50% occupancy. Accordingly, the expected occupancy of sequences in which both motifs are intact, assuming independence, is 75% (dashed line; see the ‘Materials and methods’ section). The measured occupancy of such sequences can exceed this expectation, suggesting that the motifs cooperate to stabilize the TF binding. Alternatively, the observed binding to the sequences containing both motifs can be lower than expected, indicating inhibition (bottom right). The cooperation score is defined as the difference between the measured and expected occupancy, with positive values corresponding to cooperative enhancement of binding while negative ones describe cases of inhibition (see the ‘Materials and methods’ section for mathematical definition). (2J) GRF binding is independent of motif cooperation: Shown are the averaged cooperation scores as a function of the mean observed occupancy over all sequence variants containing the two individual motifs intact. Plotted are all 63 motif pairs within the six GRF libraries that contain at least one GRF motif, measured for the respective GRF. Colors indicate the distance between the motifs, where distance 0 mean that the secondmotif starts one base-pair after the end of the first motif. The different shapes represent the different GRFs and self-motif pairs are marked in red. Error bars indicate the SEM (see the ‘Materials and methods’ section). (2K) Proximal Abfl motifs demonstrate inhibitory relations: The analysis in 2D demonstrates inhibition of Abfl binding when its two proximal motifs found on the GPP1 promoter are intact (left). This effective inhibitory relation is visualized by examining the Abfl occupancy of individual sequences. As shown in the library scheme (top left, presentation is as in Figure 1C), the GPP1 library contains a total of seven motifs. Therefore, the interaction between the two Abfl motifs can be measured within 32 different adjacent motif combinations (x-axis). Within each such motif combination, the occupancy of three mutation combinations was compared in the two Abfl motifs: Two that contain only one non-mutated Abfl motif (± and - / +) and one in which both motifs are intact (++). Shown on the bottom left are the occupancies of each of those three sequences within each context (y-axis), sorted by the observed occupancy when both motifs are intact. The respective expected occupancy of both motifs, assuming independence, is also shown (white). Fully intact and fully mutated motif combinations are marked. For comparison, the respective plot is also shown for the two Rebl motifs within the CRC1 regulatory region, measured with respect to Rebl binding, which shows no cooperation (right).

[0047] Figures 3A-3G: Cooperating motifs guide the binding of the interacting Teel and Stel2 TFs. (3A-3C) Selecting regulatory regions for testing motif-dependent binding of Teel and Stel2: To select promoter regions bound by Stel2 or Teel their respective genome-wide binding profiles were examined. (3A) The scatter plot compares the Teel and Stel2 genomic binding signal at each individual promoter, highlighting which motif combination is present in the promoter and the (closest) distances between each motif pair shown by size. The selected regions are found within the promoters marked in black. The TEC1 and PCL2 promoters, for example, contain both the Stel2 and Teel motifs at close distance, while KAR4 contains two proximal Stel2 motifs but no Teel motifs. (3B) When tested across the entire genome, binding locations of Stel2 and Teel are enriched not only in their own motif, but also in the motif of their partner, (see the ‘Materials and methods’ section), with Rebl serving as a control. (3C) The selected regulatory regions are shown in, indicating all motif locations mutated in our libraries and showing the genomic binding signal of both TFs in these regions. The signal of each TF is normalized to the maximum value received within each region. (3D-3F) High occupancy by Teel and Stel2 depends on cooperating motifs: (3D) Cooperation scores, as a function of the observed occupancy whenboth motifs are present (presentation as in Figure 2E). Included here are all pairs within the seven libraries covering all possible Teel and Stel2 motif combinations. (3E) A direct comparison of the cooperativity scores of Teel and Stel2 received for each motif pair reveals similarity in the dependent binding to the composite Stel2-Tecl motif. (3F) The pattern of motif pair cooperativity can be appreciated from the occupancy at each library sequence variant, as shown for the three indicated examples, with sequences colored by the mutation status of the two cooperating motifs, as indicated. (3G) Shown is the occupancy of Teel and Stel2 at each intact regulatory region, relative to the fully mutated sequence, as a function of the occupancy of the same region when only the motifs of the tested TF are intact. Note the high dependency of both Teel and Stel2 on motifs other than their own in both PCL2 and TEC1 promoter libraries, suggesting their cooperation.

[0048] Figures 4A-4S: Msn2 binding is explained by additive contributions of multiple, mostly self-preferred, proximal motifs. (4A-4B) Selecting regulatory regions for testing motif-dependent binding of Msn2 and Sok2: As candidates for motif-cooperation, we selected regulatory regions whose binding by Msn2 or Sok2 depends on amino-acid sequences found outside their DBDs. For this, the genome-wide binding dataset was used, in which the binding of the wild-type TFs as well as TF mutants containing only their DBDs was profiled. (4A) A scatter plot comparing the binding signal of Msn2 and its DBD-only mutant, as measured across each of the budding yeast promoters, with the number of Msn2- motifs (AGGGG) indicated. Regions found within the yellow-circled promoters were selected for further analysis and are shown in (4B) (presentation as in Figure 1C). A similar analysis and the libraries chosen for Sok2 is shown in Supplementary Figure 4J-4K. (4C- 4D) Msn2 binding reaches high occupancy and is guided by multiple, mostly canonical motifs: (4C) The occupancy of Msn2 at each intact regulatory region (relative to the fully mutated sequence), as a function of the occupancy of the same region mutated in all motifs other than the canonical Msn2 binding sites (AGGGG). (4D) Occupancy at all individual sequences is also shown for two libraries, with the mode (intact / mutated) of the single Msn2 motif found in these regions shown, and the number of non-canonical motifs contributing to occupancy indicated in dot-size (see the ‘Materials and methods’ section). Error bars represent the SEM calculated for all repeats in each experiment (3-4 repeats for each experiment). See a similar analysis of Sok2 in Supplementary Figure 4L-4M. (4E-4H) Msn2 binding triggers nucleosome evection around its genomic binding sites: (4E) The mean ATAC-seq and MNase-seq signals at the surrounding of the 200 most bound Msn2 motifs in all annotated promoters (see the ‘Materials and methods’ section), in wild-typecells (black) and cells deleted of Msn2 and its close paralog Msn4 (gray). (4F) The signals around the selected GSY2 regulatory region. Note the increase in nucleosome occupancy (MNase-seq) and the reduction in DNA accessibility (ATAC-seq) following the Msn2 / 4 double deletion. (4G) The occupancy at all individual sequences in the GSY2 library for Msn2 (x-axis) and H3 (y-axis). Msn2 binding is presented in % occupancy, as described above, while H3 binding was normalized to the genomic nucleosome occupancy (see the ‘Materials and methods’ section). The mode of the Msn2 motif (intact / mutated) in each sequence is shown, and dot sizes indicate the number of other intact TF motifs. (4H) A comparison of the Msn2 -dependent nucleosome eviction around the tested regulatory regions in MPBA and in the genome (ATAC-seq and MNase-seq). For MPBA, shown is the fold change between the mean binding at sequences containing intact Msn2 effective positions (occupancy > 10%; see the ‘Materials and methods’ section) to the mean binding at sequences in which these positions are mutated. For ATAC-seq and MNAse-seq, the fold change between the signals in wild-type cells and those deleted of Msn2 / 4 is shown. (41) The binding of Msn2 is independent of motif cooperation: Motif cooperation scores are shown as a function of the observed occupancy for all 97 motif pairs containing at least one Msn2 canonical motif within the Msn2-testing libraries following 180 s of MNAse activation. The distance between the tested motifs of each pair is indicated. A control that includes two variable positions found within the same Msn2 motif is shown. (4J-4K) Selection of regulatory regions: (4J) The sum of signals received on each promoter in the genome, comparing Sok2 to a mutant containing only its DNA Binding domain (DBD). The number of Sok2 motifs found in each promoter sequence is indicated by color. The chosen Sok2 bound regulatory regions are circled and shown in 4K (presentation as in IF). (4L- 4M) Sok2 depends mainly on its own motifs for binding: (4L) The occupancy of Sok2 at each intact regulatory region, relative to the fully mutated sequence, as a function of the occupancy of the same region when only the canonical Sok2 motifs are intact. The occupancy of each sequence of the library based on the OCA5 regulatory region is shown in 4M. The number of intact Sok2 motifs in each of the library sequences is indicated. (4N- 40) (4N) For each of the indicated libraries, is the H3 normalized occupancy (Methods), as a function of the occupancy of Msn2. Each dot represents a library variant, by the number of intact Msn2 motifs it contains. The size indicates the number of other TF intact motifs found in each sequence. Also indicated are the ranks of the Msn2 motifs found in each library, according to the collected genomic ChEC-seq signal. (40) The correlations of the scatterplots shown in 4N. (4P) Sok2 binding is independent of motif cooperation: Motif cooperation scores are shown as a function of the observed occupancy for all 67 motif pairscontaining at least one Sok2 motif within the Sok2-testing libraries. The distance between the tested motifs of each pair is indicated. (4Q) Rare cases of cooperativity and interruption of Msn2: Shown are the positive (left) and inhibitory (right) effects between two pairs of Msn2-Mig TFs motifs on the Msn2 occupancy in the indicated libraries (top, presentation as in Figure IF) for all combinations of non-pair motifs. For example, the HXK1 library (left) contains a total of 6 motifs, therefore, we can test the effect of the pair of interest in 16 different combinations of the remaining 4 motifs (x axis). Within each such context, the occupancy of three mutation combinations in the motif pair was compared (Msn2+-Mig3“; Msn2’-Mig3+; Msn2+-Mig3+). Shown are the occupancies of each of those three sequences within each context of other motif combination (y axis), sorted by the observed occupancy when both motifs are intact. The respective expected occupancy of both motifs, assuming independence, is also shown (white). Fully intact and fully mutated contexts are marked. (4R) Msn2 additive binding is independent of MNase activation time: Msn2 motif cooperation was tested along an MNase activation time course. Shown are the scores as a function of the observed occupancy of all motif pairs containing at least one Msn2 canonical motif within the Msn2-testing libraries, with timepoints indicated by color. Note that the occupancy increases with activation time, but, with the exception of the control, cooperation scores remain low. (4S) Msn2 binding at the tested regulatory regions depends on its own motif and increases over time: For each regulatory region shown are the Msn2 occupancies at the wild-type sequence and the sequence containing only its motifs intact, as a function of MNase activation time. Error bars indicate the standard error of the mean (SEM) between experimental repeats.

[0049] Figures 5A-5H: Msn2 shows no cooperativity across multiple DNA contexts. (5A-5B) Testing Msn2 libraries in different plasmid contexts: (5A) S scheme illustrating the embedding of the Msn2 libraries within five different plasmids, each containing a distinct 2 kb context. Four of the plasmids contain genomic contexts that differ in their chromatin landscape, as shown by the MNase-seq profile presented in 5B, and the fifth contains a synthetic context that is expected to be depleted of nucleosomes (see the ‘Materials and methods’ section). The number of TFs bound to the central promoter is mentioned for each context. (5C) Msn2 binding to the GSY2 promoter region depends on its binding motif in all measured contexts: Two repeats showing the binding of Msn2 to each sequence in the GSY2 library are presented for each context. The mode of the Msn2 motif (intact / mutated) and the number of other TF intact motifs are indicated. (5D) Msn2 binds its motifs independently in all tested contexts: For each context, motif cooperation scores are shownas a function of the observed occupancy at 180 s of MNase activation. The color indicates the different libraries measured within each context. The black-circled dots are controls that include two variable positions within the same Msn2 motif. (5E) Testing Msn2 libraries in the genomic context: A scheme illustrating CRISPR-Cas9-mediated library integration into the HO locus. (5F-5G) Msn2 binding in the genomic context depends on its binding sites: (5F) The binding of Msn2 to each sequence of the GSY2 genomically embedded library, presentation as in 5C. (5G) The occupancy at the wild-type sequence in each genomically embedded library. (5H) Msn2 binding is independent of other TF motifs also in the genomic context: Shown are the motif pair cooperation scores as a function of the observed occupancy at 180 s of MNase activation in the genomic context. The black-circled dots are the above- mentioned controls that include two variable positions within the same Msn2 motif.

[0050] Figures 6A-6H: Msn2 binding to non-canonical motifs is explained by direct recognition of its DBD. (6A-6B) Motifs contributing to Msn2 binding are similar in sequence to its canonical motif: (6A) The average occupancies of all Msn2 and Sok2 non- canonical motifs when all the other motifs are mutated (see the ‘Materials and methods’ section). While Msn2 depends on a few non-canonical motifs, Sok2 does not depend on motifs other than its own. (6B) The Msn2 occupancies at all motifs found in the tested libraries are plotted as a function of their Hamming distance from the Msn2 motif, i.e. the number of positions in which the respective motif differs from the canonical one (AGGGG). Purple lines indicate the median Msn2 dependent occupancy for each Hamming distance. (6C) Experimental scheme for distinction between direct and indirect binding of non- canonical Msn2 motifs: It was desired to distinguish whether Msn2 directly binds the non- canonical motifs contributing to its occupancy using its DBD (Option I, top left) or alternatively, is recruited to these motifs via interaction with another TF (Option II, top right). To this end, the occupancy of an Msn2 mutant containing only its DBD were measured. It was expected that this mutant would lack the ability to interact with other TFs (bottom). (6D) Regions outside the DBD stabilize Msn2 binding at regulatory regions: In addition to the wild-type Msn2, we tested the binding of the DBD-only mutant across the 11 Msn2 libraries presented above. Shown are the distributions of the occupancies measured for the wild-type Msn2 and its DBD-only mutant at sequence variants in which only one Msn2 canonical motif is intact and all other motifs found within the respective region mutated, over all tested libraries. As the binding of the DBD-only variant was low, we also tested its binding upon over-expression using the strongest budding yeast promoter (TDH3). (6E-6G) Msn2 motif occupancy is explained by preferences of its DBD: (6E) The singlemotif-dependent occupancies of the over-expressed Msn2 DBD-only mutant for all motifs found within the tested libraries. Sequences are classified based on the presence of a canonical Msn2 motif (AGGGG), a non-canonical motif that contributes to the binding of the wild-type Msn2 (Hamming distance > 0; as defined in a prior experiment shown in Figure 4A-4I, based on the wild-type Msn2 at occupancy at 180 s of activation), and non- canonical motifs showing no contribution to the occupancy of Msn2 (gray). (6F) The direct comparisons of the measured occupancies of the wild-type and DBD-only at each individual sequence for three selected libraries at 180 s of MNase activation (presentation as in Figure 4D). Error bars represent the SEM over repeats (5-7 repeats for each TF in each library). (6G) The distributions of correlations between occupancies at each sequence of the wildtype Msn2 and DBD-only mutant across all Msn2 libraries. The distributions of correlations between the wild-type Msn2 repeats and between Msn2 to unrelated TFs, measured across those same libraries, are also presented (see Figure 6, below). Note the tight correlation in the binding pattern between the wild-type Msn2 and its DBD-only variant. (6H) Time course showing high correspondence of non-canonical motif preference between Msn2 and an overexpressed mutant containing only its DBD: Shown are the occupancies (across all library sequences) of the wild-type Msn2 and of its over-expressed DBD-only mutant as a function of activation times. Sequences are classified based on the presence of a canonical Msn2 motif (AGGGG, blue), a non-canonical motif that contributes to the binding of Msn2 (Hamming distance > 1; as defined in a prior experiment shown in Figure 4A-4I, based on the wildtype Msn2 at occupancy at 180 seconds of activation), and non-canonical motifs showing no contribution to the occupancy of Msn2 (gray).

[0051] Figures 7A-7K: Screen for motif-dependent interactions reveals independent binding with rare cases of motif cooperation. (7A-7B) Motif-dependent TF occupancy across regulatory regions: (7A) TF occupancies at each regulatory sequence calculated using the abundance of intact versus fully mutated sequences. TFs are ordered by the median occupancies of their respective regulatory regions (gray line). (7B) The self-preferred motif dependent region occupancy, measured by comparing the average occupancy at sequences containing intact versus mutated canonical motifs of each TF (see the ‘Materials and methods’ section). Average region occupancies depending on non-canonical motifs are also plotted. (7C-7D) (7C) One sided recruitment is limited: TF occupancy may depend on both self-canonical motifs and motifs of other TFs, suggesting recruitment. This was estimated by the effect of non-canonical motifs on TF occupancy when mutating its canonical motifs (see the ‘Materials and methods’ section). Cases demonstrating such dependency are shown(>10% contribution to occupancy). The tested TFs are indicated in orange and the potential recruiters are indicated in purple. Medl5, known to be recruited by Msn2 (19), was tested as a control on Msn2 regulatory regions. The Hamming distance between canonical and non- canonical motifs is shown. (7D) As another measure for dependency, occupancy levels of TFs tested on the same regions were correlated, revealing that most TFs displayed distinct patterns (left). Correlating TF pairs (Pearson's r > 0.4) (right), where each dot corresponds to a library tested for both TFs are shown. The Euclidean distances between the respective TF DBDs preferred motifs (see the ‘Materials and methods’ section) is shown. TF pairs from the same DBD family are indicated in pink. (7E) Cooperation among motif pairs is rare: Shown are the cooperation scores amongst all 1917 motif pairs containing at least one canonical motif of the TFs measured in our screen, as a function of the observed occupancy over all variants containing the motif pairs intact. The distance between the two motifs within each pair is indicated. Highlighted are pairs with strong interaction (cooperation or inhibition), where the TF measured is on the left, and the associated second motif TF is on the right. Control cases carrying two variable positions found within the same motif are marked in red. Note the small fraction of cooperating motif pairs; only 36 of the 1917 (<2%) showed cooperativity scores higher than 10%. (7F) Summary scheme: Based on our screen we conclude that TF binding is well explained by the additive contributions of individual motifs. Thus, motif-dependent cooperativity is rare, refuting its general role in increasing TF genomic binding specificity. (7G-7I) TFs and regulatory regions selected for screening motif interactions: As candidates for cooperating motifs, the dataset of 145 TF binding profiles was searched for regulatory regions containing dense sets of bound motifs, as described in Figure 1A. Overall, 39 TFs, and 68 regulatory regions of the length of 164 bps were selected. (7G) The similarity (correlation) of the genomic binding profiles of the selected TFs, measured across the selected regions. Note the high binding similarity of the interacting factors Stel2-Tecl, as well as the clusters of TFs that bind at promoters overlapping with Msn2 or Sok2. (7H) The TF binding signals themselves, presented as the ratio between the total signal received for each TF on each region to the maximal signal received by the same TF on a similar-sized region when examining the entire genome (Methods). (71) The number of respective TF motifs in each selected region. The overall number of motifs within each regulatory region, as well as the overall number of TFs predicted to bind these same motifs are indicated (bottom stripes). (7 J) Cooperation among motif pairs is limited: Shown are the cooperation scores amongst all 1917 motif pairs containing at least one canonical motif of the TFs measured in our screen. Pairs are grouped by the tested TFs, with dots circled in black indicating a pair of motifs belonging to the tested TF. The maximal and median pairoccupancies measured for each TF are shown at the bottom. Controls including two variable positions found within the same motif are shown on the right panel. Note the small fraction of cooperating motif pairs. (7K) The majority of detected interactions are of highly proximal motifs: Shown are the distances for all motif pairs showing inhibition (left, score<-10) and cooperation (right, score>10) for pairs showing >10% contribution to the TF occupancy. The median values for each group are shown as black lines. The dashed line indicates the median distance between all bound motif pairs showing no interaction.

[0052] Figures 8A-8B: (8A) Schematic of a nucleic acid libraries (each marked by a different shape) for testing binding of dCas-9+gRNA to its target sequence and variants of that sequence is presented. (8B) Chart of the binding of each tested library variant by Cas9. Two repeats are plotted. The library to which the variant belongs is indicated by the same shape as in 8A. The number of mutations within each variant is shown.

[0053] Figures 9A-9B: (9A) Schematic of a nucleic acid libraries (each marked by a different shape) for testing binding of dCas-9+gRNA to its target sequence and variants of that sequence, when the target sequence is flanked by Rebl binding motifs. (9B) Chart of the binding of each tested library variant by Cas9. Two repeats are plotted. The library to which the variant belongs is indicated by the same shape as in 9A. The number of mutations within each variant is shown.

[0054] Figures 10A-10E: (10A) Shown is the binding occupancy of dCas9 at various sequences. A schematic of the sequences is shown on the right. From top to bottom the sequences are: The on-target sequence, PAM mutation, nucleotide scrambling and sliding mutations along the target sequence. The minimal mutation distance from the PAM is shown. (10B) Distributions of the binding signals obtained from dCas9 in the MPBA system for GUIDE-seq detected off targets and CIRCLE-seq off targets. Lower values equate to greater binding. (10C) A series of binding score thresholds were generated (x-axis) and for each threshold the fraction of variants passing is shown (y-axis). (10D) Dot plot comparing GUIDE-seq scores to MPBA scores. The on-target sequence is marked as are amplicon validated sequences. A suggested threshold for MPBA sensitivity is marked by a dashed line. (10E) A Venn diagram showing the overlap between MPBA and GUIDE-seq.

[0055] Figures 11A-11D: Peptide Massive parallel binding assay (pMPBA) using library of short peptides. (11A) pMPBA for peptide variants workflow: A target sequence of interest is inserted into the plasmid, then, a library of sequences coding for 64 AA peptides is integrated upstream to an MNase (1, top). The plasmid pool is then transformed intobacteria for propagation followed by plasmid extraction. A ribosome skipping inducing sequence (E2A) is located between the peptide-MNase and a URA3 gene (1, bottom). Following yeast transformation, only cells expressing an in-frame peptide will grow in uracil depleted media. A short calcium pulse activates the MNase, triggering the cleavage of peptide-bound plasmids (2). Next, a targeted PCR is performed, amplifying non-cleaved sequences covering the region from the peptide-coding sequence to the target promoter sequence. Amplicon sequencing before and after MNase activation provides a quantitative binding measurement for each variant (3). (11B) The zinc-finger domain of the DBD of the Msn2 TF was modified to generate a few variants: (1) Premature stop codon (E->*) (2) DNA contact residue modification to match its paralog Nrg2 (E->G), (3) Structural mutation predicted to disrupt proper 3D folding of the domain (H->Q), and (4) Combination of mutations (2) and (3). This library was integrated into three plasmids, each containing a different target promoter (HSP12, CDC6, and MGA1). (11C) Ribosome skipping peptide reduces incorrectly translated proteins: The fraction of library sequences containing a premature stop codon is shown before (grey) and after (orange) yeast transformation. (11D) pMPBA successfully quantifies variant binding: Library binding was tested over an MNase activation time course. The log2 fold-change from time-point zero (no MNase activation) is shown for the three tested promoters.DETAILED DESCRIPTION OF THE INVENTION

[0056] The present invention, in some embodiments, provides nucleic acid molecule libraries wherein each molecule comprises 5’ and 3’ primer recognition sites which are common to all molecules and a plurality of molecules comprising a target sequence molecule and a plurality of variant sequence molecules. Kits comprising the nucleic acid molecule libraries are also provided. Methods of determining protein association with a nucleotide sequence comprising receiving a nucleic acid molecule library; contacting the library with a cell expressing a fusion protein wherein the fusion protein comprises the protein fused to a nuclease; amplifying nucleic acid molecules with the 5’ and 3’ primer and sequencing the amplification products are provided. Methods of identifying an inhibitor of protein-DNA interaction are also provided.

[0057] By a first aspect, there is provided a library comprising a plurality of nucleic acid molecules.

[0058] In some embodiments, the library is a nucleic acid molecule library. In some embodiments, the library is a library of plasmids. In some embodiments, the library is a library of vectors. In some embodiments, the library is a library of DNA fragments. In some embodiment, the nucleic acid molecule is a plasmid. In some embodiments, the nucleic acid molecule is a vector. In some embodiments, the vector is an expression vector. In some embodiments, the expression vector is a eukaryotic expression vector. In some embodiments, the library is for use in a method of the invention.

[0059] In some embodiments, the nucleic acid molecule is a fragment. In some embodiments, the nucleic acid molecule is not circular. In some embodiments, the nucleic acid molecule does not comprise an open reading frame. In some embodiments, the nucleic acid molecule is at most 130, 140, 150, 160, 170, 175, 180, 190, 200, 225, 250, 275, 300, 325, 350, 375, 400, 425, 450, 475 or 500 nucleotides long. Each possibility represents a separate embodiment of the invention. In some embodiments, the nucleic acid molecule is at most 200 nucleotides long. In some embodiments, the nucleic acid molecule is at most 250 nucleotides long. In some embodiments, the nucleic acid molecule is at most 500 nucleotides long. In some embodiments, the nucleic acid molecule is about 130, 140, 150, 160, 170, 175, 180, 190, 200, 225, 250, 275, 300, 325, 350, 375, 400, 425, 450, 475 or 500 nucleotides long. Each possibility represents a separate embodiment of the invention. In some embodiments, the nucleic acid molecule is about 200 nucleotides long. In some embodiments, the nucleic acid molecule is 10-1000, 50-1000, 100-1000, 150-1000, 200- 1000, 10-500, 50-500, 100-500, 150-500, 200-500, 10-450, 50-450, 100-450, 150-450, 200- 450, 10-400, 50-400, 100-400, 150-400, 200-400, 10-350, 50-350, 100-350, 150-350, 200- 350, 10-300, 50-300, 100-300, 150-300, 200-300, 10-250, 50-250, 100-250, 150-250, 200- 250, 10-200, 50-200, 100-200, or 150-200 nucleotides long. Each possibility represents a separate embodiment of the invention. In some embodiments, the nucleic acid molecule is 100 to 300 nucleotides long. In some embodiments, the nucleic acid molecule is 150 to 250 nucleotides long.

[0060] The term "nucleic acid" is well known in the art. A "nucleic acid" as used herein will generally refer to a molecule (i.e., a strand) of DNA, RNA or a derivative or analog thereof, comprising a nucleobase. A nucleobase includes, for example, a naturally occurring purine or pyrimidine base found in DNA (e.g., an adenine "A," a guanine "G," a thymine "T" or a cytosine "C") or RNA (e.g., an A, a G, an uracil "U" or a C).

[0061] The terms “nucleic acid molecule” include but not limited to single-stranded RNA (ssRNA), double- stranded RNA (dsRNA), single- stranded DNA (ssDNA), double- strandedDNA (dsDNA), small RNA such as miRNA, siRNA and other short interfering nucleic acids, snoRNAs, snRNAs, tRNA, piRNA, tnRNA, small rRNA, hnRNA, circulating nucleic acids, fragments of genomic DNA or RNA, degraded nucleic acids, ribozymes, viral RNA or DNA, nucleic acids of infectious origin, amplification products, modified nucleic acids, plasmidical or organellar nucleic acids and artificial nucleic acids such as oligonucleotides. In some embodiments, the molecule is a DNA molecule. In some embodiments, the molecule is an RNA molecule. In some embodiments, the molecule is an isolated molecule. In some embodiments, the molecule is a purified molecule. In some embodiments, the molecule is a chimeric molecule. In some embodiments, the molecule is a not a naturally occurring molecule.

[0062] As used herein, the term "expression" refers to the biosynthesis of a gene product, including the transcription and / or translation of said gene product. Thus, expression of a nucleic acid molecule may refer to transcription of the nucleic acid fragment (e.g., transcription resulting in mRNA or other functional RNA) and / or translation of RNA into a precursor or mature protein (polypeptide). Expressing a gene / open reading frame within a cell is well known to one skilled in the art. It can be carried out by, among many methods, transfection, viral infection, or direct alteration of the cell’s genome. In some embodiments, the gene is in an expression vector such as plasmid or viral vector.

[0063] A vector nucleic acid sequence generally contains at least an origin of replication for propagation in a cell and optionally additional elements, such as a heterologous polynucleotide sequence, expression control element (e.g., a promoter, enhancer), selectable marker (e.g., antibiotic resistance), poly-Adenine sequence.

[0064] The vector may be a DNA plasmid delivered via non-viral methods or via viral methods. The viral vector may be a retroviral vector, a herpesviral vector, an adenoviral vector, an adeno-associated viral vector or a poxviral vector. The promoters may be active in mammalian cells. The promoter may be a viral promoter.

[0065] In some embodiments, the vector is introduced into the cell by standard methods including electroporation (e.g., as described in From et al., Proc. Natl. Acad. Sci. USA 82, 5824 (1985)), Heat shock, infection by viral vectors, high velocity ballistic penetration by small particles with the nucleic acid either within the matrix of small beads or particles, or on the surface (Klein et al., Nature 327. 70-73 (1987)), and / or the like.

[0066] In some embodiments, mammalian expression vectors include, but are not limited to, pcDNA3, pcDNA3.1 (±), pGL3, pZeoSV2(±), pSecTag2, pDisplay, pEF / myc / cyto,pCMV / myc / cyto, pCR3.1, pSinRep5, DH26S, DHBB, pNMTl, pNMT41, pNMT81, which are available from Invitrogen, pCI which is available from Promega, pMbac, pPbac, pBK- RSV and pBK-CMV which are available from Strategene, pTRES which is available from Clontech, and their derivatives.

[0067] In some embodiments, expression vectors containing regulatory elements from eukaryotic viruses such as retroviruses are used by the present invention. SV40 vectors include pSVT7 and pMT2. In some embodiments, vectors derived from bovine papilloma virus include pBV-lMTHA, and vectors derived from Epstein Bar virus include pHEBO, and p2O5. Other exemplary vectors include pMSG, pAV009 / A+, pMTO10 / A+, pMAMneo- 5, baculovirus pDSVE, and any other vector allowing expression of proteins under the direction of the SV-40 early promoter, SV-40 later promoter, metallo thionein promoter, murine mammary tumor virus promoter, Rous sarcoma virus promoter, polyhedrin promoter, or other promoters shown effective for expression in eukaryotic cells.

[0068] In some embodiments, recombinant viral vectors, which offer advantages such as lateral infection and targeting specificity, are used for in vivo expression. In one embodiment, lateral infection is inherent in the life cycle of, for example, retrovirus and is the process by which a single infected cell produces many progeny virions that bud off and infect neighboring cells. In one embodiment, the result is that a large area becomes rapidly infected, most of which was not initially infected by the original viral particles. In one embodiment, viral vectors are produced that are unable to spread laterally. In one embodiment, this characteristic can be useful if the desired purpose is to introduce a specified gene into only a localized number of targeted cells.

[0069] Various methods can be used to introduce the expression vector of the present invention into cells. Such methods are generally described in Sambrook et al., Molecular Cloning: A Laboratory Manual, Cold Springs Harbor Laboratory, New York (1989, 1992), in Ausubel et al., Current Protocols in Molecular Biology, John Wiley and Sons, Baltimore, Md. (1989), Chang et al., Somatic Gene Therapy, CRC Press, Ann Arbor, Mich. (1995), Vega et al., Gene Targeting, CRC Press, Ann Arbor Mich. (1995), Vectors: A Survey of Molecular Cloning Vectors and Their Uses, Butterworths, Boston Mass. (1988) and Gilboa et at. [Biotechniques 4 (6): 504-512, 1986] and include, for example, stable or transient transfection, lipofection, electroporation and infection with recombinant viral vectors. In addition, see U.S. Pat. Nos. 5,464,764 and 5,487,992 for positive-negative selection methods.

[0070] It will be appreciated that other than containing the necessary elements for the transcription and translation of the inserted coding sequence (encoding the polypeptide), the expression construct of the present invention can also include sequences engineered to optimize stability, production, purification, yield or activity of the expressed polypeptide.

[0071] In some embodiments, the library comprises a plurality of nucleic acid molecules. In some embodiments, each molecule of the plurality comprises two primer recognition sites. In some embodiments, primer recognition sites are primer binding sites. In some embodiments, each molecule of the plurality comprises a 5’ primer recognition site. In some embodiments, each molecule of the plurality comprises a 3’ primer recognition site. In some embodiments, the two primer recognition sites are common to all molecules. In some embodiments, the 5’ primer recognition site is common to all the molecules. In some embodiments, the 3’ recognition site is common to all the molecules. In some embodiments, all the molecules is all the molecules of the plurality. In some embodiments, common sites comprise the same sequences. In some embodiments, common sites are identical. It will be understood by a skilled artisan that use of the 5’ and 3’ primers for amplification will amplify all the molecules. However, when a molecule is cleaved between the primer recognition sites the cleaved molecule will not be amplified.

[0072] In some embodiments, each molecule of the plurality comprises a guide RNA (gRNA). In some embodiments, a gRNA is a single guide RNA (sgRNA). In some embodiments, the gRNA is common to all molecules of the plurality. In some embodiments, the gRNA is complementary to the target sequence. In some embodiments, complementary is reverse complementary. In some embodiments, the target sequence is the putative target sequence.

[0073] In some embodiments, the plurality of molecules comprises a target sequence molecule. In some embodiments, the library comprises a target sequence molecule. In some embodiments, the target sequence molecule comprises a target sequence. In some embodiments, the target sequence is between the 5’ primer recognition site and the 3’ primer recognition site. In some embodiments, the target sequence comprises at least one binding site. In some embodiments, the target molecule comprises at least one putative target sequence of the protein. In some embodiments, the target sequence is a putative target sequence. In some embodiments, the target sequence comprises at least one putative target sequence of the protein. In some embodiments, the binding site is the binding site of the protein. In some embodiments, the binding site is a binding site of a transcription factor. In some embodiments, the target sequence comprises at least 2, 3, 4, 5, 6, 7, 8, 9, or 10 bindingsites. Each possibility represents a separate embodiment of the invention. In some embodiments, the target sequence comprises at least 2 binding sites. In some embodiments, the target sequence comprises at least 3 binding sites. In some embodiments, the target sequence comprises at least 6 binding sites. In some embodiments, the target sequence comprises at least 7 binding sites. In some embodiments, the binding sites are different binding sites. In some embodiments, the binding sites are binding sites for different proteins. In some embodiments, the proteins are transcription factors. In some embodiments, the target sequence is common to all molecules of the plurality. In some embodiments, the target sequence is common to all molecules of the library. In some embodiments, the target sequence is a target of the protein. In some embodiments, the target sequence is a target of the peptide. In some embodiments, at least one peptide of the library binds to the target sequence.

[0074] In some embodiments, the plurality of molecules comprises a plurality of variant sequence molecules. In some embodiments, the library comprises a plurality of variant sequence molecules. In some embodiments, a variant sequence molecule comprises a variant target sequence. In some embodiments, the variant target sequence is between the 5’ primer recognition site and the 3’ primer recognition site. In some embodiments, a variant target sequence is a variant of the target sequence. In some embodiments, a variant target sequence comprises at least one mutation of the target sequence. In some embodiments, a variant target sequence comprises at least one mutation of a binding site of the target sequence. In some embodiments, all possible mutations of all nucleotides of the target sequence are present in the plurality of variant target sequences. In some embodiments, all possible mutations of all nucleotides of the binding site are present in the plurality of variant target sequences. In some embodiments, all possible mutations of all nucleotides of all binding sites are present in the plurality of variant target sequences. In some embodiments, each variant sequence differs from another variant sequence by a single nucleotide change. In some embodiments, all possible mutations within the binding sites are made in the plurality of variant sequences. It will be understood that the variant sequences vary by a single change such that a single base pair resolution on the effect on binding can be evaluated. Thus, not only are there variants with single nucleotide changes from the target sequence (1 change) but also there are all the variants with single nucleotide changes from those variants (2 changes) and then variants for those (3 changes) and so on until all nucleotides of the binding site have been changed to all possible permutations.

[0075] In some embodiments, each molecule contains a transcription factor binding site adjacent to a 5’ end of the target sequence. In some embodiments, each molecule contains a transcription factor binding site adjacent to a 5’ end of the putative target sequence or variant target sequence. In some embodiments, each molecule contains a transcription factor binding site adjacent to a 3’ end of the target sequence. In some embodiments, each molecule contains a transcription factor binding site adj acent to a 3 ’ end of the putative target sequence or variant target sequence. In some embodiments, the transcription factor binding site is adj acent to botthe 5’ end and the 3’ end. In some embodiments, adjacent is directly adjacent. In some embodiments, adjacent is within 1, 2, 3, 4, or 5 bases. Each possibility represents a separate embodiment of the invention. In some embodiments, adjacent is within 5 bases. In some embodiments, the transcription factor binding site is a site selected from those provided in Table 1. In some embodiments, the transcription factor is Rebl. In some embodiments, the Rebl binding site comprise or consists of CGGGTAA or its reverse complement. In some embodiments, the Rebl binding site comprise or consists of CGGGTAA.

[0076] In some embodiments, each molecule comprises a coding sequence. In some embodiments, the coding sequence is between the 5’ primer recognition site and the 3’ primer recognition stie. In some embodiments, the coding sequence encodes a fusion protein. In some embodiments, the fusion protein comprises the protein fused to a nuclease. In some embodiments, the fusion protein comprises a peptide fused to a nuclease. As used herein, the terms “protein” and “peptide” are used interchangeably. In some embodiments, the peptides are variant peptides. In some embodiments, each fusion protein comprises a different variant peptide. In some embodiments, a variant peptide comprises at least one mutation of the peptide amino acid sequence. In some embodiments, the peptide is a nucleic acid binding motif. In some embodiments, the peptide comprises a nucleic acid binding motif. In some embodiments, the motif is a DNA binding motif. In some embodiments, a motif is a domain. In some embodiments, each fusion protein comprises a variant of the motif. In some embodiments, a variant peptides comprises at least one mutation of an amino acid in a nucleic acid binding motif. In some embodiments, all possible mutations of all amino acids of the nucleic acid binding motif are present in the plurality of coding sequences. In some embodiments, all possible mutations of all amino acids of the nucleic acid binding motif are present in the plurality of variant coding sequences. In some embodiments, all possible mutations of all amino acids of all nucleic acid binding motifs are present in the library. In some embodiments, each variant peptide differs from another variant peptide by a single amino acid change. In some embodiments, all possible mutations within the binding motifare made in the plurality of variant coding sequences. It will be understood that the variant coding sequences vary by a single change resulting in a single amino acid alterations such that the effect of specific amino acid residues on binding can be evaluated. Thus, not only are there variants with single amino acid changes from the wild type peptide (1 change) but also there are all the variants with single amino acid changes from those variants (2 changes) and then variants for those (3 changes) and so on until all amino acids of the binding motif have been changed to all possible permutations. Alternatively, amino acids implicated in binding are mutated or amino acids implicated in the structure of the protein or motif are mutated. In some embodiments, all amino acids implicated in binding, structure or both are mutated.

[0077] In some embodiments, the protein is a DNA binding protein. In some embodiments, the protein is an RNA binding protein. In some embodiments, the protein is a transcription factor. In some embodiments, the protein is a eukaryotic protein. In some embodiments, the eukaryote is yeast. In some embodiments, the protein is a protein complex. In some embodiments, the protein is a genome editing protein. In some embodiments, the protein is a nuclease. In some embodiments, the protein is a sequence specific nuclease. Genome editing proteins are well known in the art and include for example, CRISPR-associated proteins (CASs), zinc finger nucleases (ZFNs), Transcription activator-like effector nucleases (TALENs) and meganucleases. In some embodiments, the protein is a CRISPR- associated protein (CAS). In some embodiments, the CAS is a CAS-CLOVER CAS. In some embodiments, the CAS is CAS9. In some embodiments, the protein complex is a CAS in complex with a guide RNA (gRNA). In some embodiments, the gRNA is a single guide RNA (sgRNA). In some embodiments, the binding site is the complement of the gRNA. In some embodiments, complementary is reverse complementary. In some embodiments, the binding site is at least 70, 75, 80, 85, 90, 92, 95, 97, 99 or 100% complementary to the gRNA. Each possibility represents a separate embodiment of the invention. In some embodiments, the binding site is at least 90% complementary to the gRNA. In some embodiments, the binding site is 100% reverse complementary to the binding site. In some embodiments, the molecule further comprises a PAM sequence. In some embodiments, the Cas is a dead Cas (dCAS) which no longer has the ability to cleave DNA. In some embodiments, the dCAS has abolished nuclease function. In some embodiments, the dCas is dCas9. In some embodiments, the method is a method of detecting off-target binding. In some embodiments, the off-target binding is off-target binding by the protein. In some embodiments, the off- target binding is off target CAS-gRNA protein complex binding. In some embodiments, off-target binding is binding to variants of the target sequence. In some embodiments, a variant target sequence depleted in the sequencing data is an off-target sequence bound by the protein. In some embodiments, the protein is selected from a CAS, a ZFN, a TALEN and a meganuclease.

[0078] In some embodiments, the target sequence molecule comprises a putative target sequence of the protein. In some embodiments, the target sequence molecule comprises a putative target sequence of the peptide. In some embodiments, at least one peptide of the plurality binds to the target sequence. In some embodiments, at least one peptide of the library binds to the target sequence. In some embodiments, one coding sequence encodes a wild-type version of the peptide. In some embodiments, the wild-type peptide binds the target sequence.

[0079] In some embodiments, the coding sequence is operably linked to at least one regulatory element. In some embodiments, the regulatory element is active in the cell. In some embodiments, the regulatory element is a promoter. In some embodiments, the regulatory element induces transcription of the coding sequence in the cell. In some embodiments, the regulatory element induces expression of the coding sequence in the cell. In some embodiments, contacting the library with a cell results in expression of the fusion protein in the cell. In some embodiments, contacting the library with a cell results in production of the fusion protein in the cell. In some embodiments, production is translation.

[0080] In some embodiments, the coding sequence further encodes a selection protein. In some embodiments, the cell cannot grow in a first condition and the selection protein is a protein that allows growth in the first condition. In some embodiments, the selection protein is a selection marker. Markers for selection, thus allowed for the detection of expression of a plasmid are well known in the art and include for example antibiotic resistance proteins (which allow growth in the condition of the antibiotic being present), protein for growth on specific sugars (which allow growth in the condition in which the normal sugar is missing and the other specific sugar is present) and proteins that allow for growth when an essential substance is missing. In some embodiments, the first condition is absence of an essential substance. In some embodiments, the essential substance is uracil and the selection protein is orotidine 5-phosphate decarboxylase (URA3). In some embodiments, URA3 is S. cerevisiae URA3. The S. cerevisiae URA3 gene can be found in the Entrez gene ID 856692. The URA3 mRNA sequence is provided in NM_001178836 and the protein sequence is provided in NP_010893 and Uniprot entry P03962.

[0081] In some embodiments, the selection protein is separated from the fusion protein by a sequence that results in the in-frame production of two proteins. In some embodiments, the sequence is a self-cleaving peptide. In some embodiments, the sequence is a ribosomal skipping sequence. The addition of the ribosomal skipping sequence followed by the selection marker ensures that a fusion protein is actually produced and protects against frame shifts, deletions / additions or stop codons. In such cases, the selection protein is not produced (out of frame or the mRNA is degraded) and cell growth is arrested when the cells are grown in the first condition. In some embodiments, the method comprises producing the library in cells growing in the first condition. In some embodiments, the cell is grown in the first condition after the contacting. In some embodiments, the ribosomal skipping sequence is selected from P2A, E2A, T2A and F2A. In some embodiments, the ribosomal skipping sequence is P2A. The ribosomal skipping peptide sequence and the nucleotide sequence encoding them are well-known in the art and may be integrated into the coding sequence between the fusion protein and the selection protein. In some embodiments, the sequence encoding the fusion protein is 5’ to the sequence encoding the selection protein.

[0082] By another aspect, there is provided a kit comprising a nucleic acid molecule library of the invention.

[0083] In some embodiments, the kit further comprises at least one fusion protein. As used herein, the term “fusion protein” refers to a single polypeptide chain that comprises parts of two different proteins, i.e., the protein of interest and the nuclease. It will therefore be understood that a nuclease (e.g., Cas9) on its own is not a fusion protein. In some embodiments, the kit further comprises a nucleic acid molecule encoding the at least one fusion protein. In some embodiments, the kit further comprises a cell comprising the fusion protein. In some embodiments, the cell comprises the nucleic acid molecule encoding the fusion protein. In some embodiments, the cell expresses the fusion protein. In some embodiments, the kit further comprises a cell comprising the gRNA. In some embodiments, the cell comprises the nucleic acid molecule encoding the gRNA. In some embodiments, the cell expresses the gRNA. In some embodiments, the fusion protein comprises the protein. In some embodiments, the protein is the protein that associates with the binding site. In some embodiments, the protein is the protein that associates with the target sequence. In some embodiments, associates with is binds to. In some embodiments, associates with is is recruited to. In some embodiments, the fusion protein comprises a transcription factor. In some embodiments, the fusion protein further comprises a nuclease. In some embodiments, fusion protein is a fusion of the protein and the nuclease. In some embodiments, the fusionis a direct fusion. In some embodiments, the fusion is fusion via a linker. In some embodiments, the linker is an amino acid linker.

[0084] In some embodiments, the nuclease is a DNA nuclease and the library is a DNA library. In some embodiments, the nuclease is an RNA nuclease and the library is an RNA library. In some embodiments, the nuclease is an activatable nuclease. In some embodiments, activatable is activatable by contact with a substrate. In some embodiments, the nuclease is an inducible nuclease. In some embodiments, inducible is inducible by contact with a substrate. In some embodiments, the nuclease is activatable by contact with a substrate. In some embodiments, the nuclease is substrate activated nuclease. In some embodiments, the nuclease is a sodium, calcium or magnesium activated nuclease. In some embodiments, the nuclease is a calcium or magnesium activated nuclease. In some embodiments, an activatable nuclease is a nuclease whose activity increases by at least 10X upon addition of the substrate. In some embodiments, the substrate is an activating molecule. In some embodiments, an activatable nuclease has no activity in the absence of the substrate. In some embodiments, no activity is no detectable activity. In some embodiments, no activity is essentially no activity. Examples of activatable nucleases include, but are not limited to micrococcal nuclease (MNase), cyclophilin A, cyclophilin B, cyclophilin C, 97 kDa endonuclease, membrane-associated nucleases in mycoplasma, Rpn (YhgA-like) E. coli nucleases, T5 Flap endonuclease, calcium / magnesium-dependent nucleases in lymphoid cell lines, calcium-dependent endonucleases in Pneumococci, Lucilia sericata nucleases and calcium and magnesium ion-dependent endonuclease 37 kDa. In some embodiments, the nuclease is micrococcal nuclease (MNase). In some embodiments, the nuclease is a calcium dependent nuclease. In some embodiments, the nuclease is a magnesium dependent nuclease. In some embodiments, the nuclease is a calcium and magnesium dependent nuclease.

[0085] In some embodiments, a nucleic acid molecule encoding the fusion protein is a vector. In some embodiments, the nucleic acid molecule is a nucleic acid molecule of the library of the invention. In some embodiments, the vector is an expression vector. In some embodiments, the expression vector is a eukaryotic expression vector. In some embodiments, the cell is a eukaryotic cell. In some embodiments, the expression vector is a yeast expression vector. In some embodiments, the cell is a yeast cell. In some embodiments, the cell is a S. cerevisiae cell. In some embodiments, the expression vector is a bacterial expression vector. In some embodiments, the cell is a bacterial cell. In some embodiments, the expression vector is a mammalian expression vector. In some embodiments, the cell is amammalian cell. In some embodiments, the mammal is human. In some embodiments, the nucleic acid molecule encoding the fusion protein comprises an open reading frame encoding the fusion protein. In some embodiments, the open reading frame is operably linked to at least one transcription regulatory element. In some embodiments, the regulatory element is a promoter. In some embodiments, the regulatory element is active in the cell. In some embodiments, the cell constitutively expresses the fusion protein. In some embodiments, a sequence encoding the fusion protein is integrated into the cell’s genome.

[0086] As used herein, the term “operably linked” is intended to mean that the open reading frame is linked to the regulatory element or elements in a manner that allows for expression of the nucleotide sequence (e.g., in an in vitro transcription / translation system or in a host cell when the vector is introduced into the host cell). In some embodiments, the molecule is configured to express the fusion protein in the cell. The term "promoter" as used herein refers to a group of transcriptional control modules that are clustered around the initiation site for an RNA polymerase i.e., RNA polymerase II. Promoters are composed of discrete functional modules, each consisting of approximately 7-20 bp of DNA, and containing one or more recognition sites for transcriptional activator or repressor proteins.

[0087] In some embodiments, nucleic acid sequences are transcribed by RNA polymerase II (RNAP II and Pol II). RNAP II is an enzyme found in eukaryotic cells. It catalyzes the transcription of DNA to synthesize precursors of mRNA and most snRNA and microRNA.

[0088] In some embodiments, mammalian expression vectors include, but are not limited to, pcDNA3, pcDNA3.1 (±), pGL3, pZeoSV2(±), pSecTag2, pDisplay, pEF / myc / cyto, pCMV / myc / cyto, pCR3.1, pSinRep5, DH26S, DHBB, pNMTl, pNMT41, pNMT81, which are available from Invitrogen, pCI which is available from Promega, pMbac, pPbac, pBK- RSV and pBK-CMV which are available from Strategene, pTRES which is available from Clontech, and their derivatives.

[0089] In some embodiments, expression vectors containing regulatory elements from eukaryotic viruses such as retroviruses are used by the present invention. SV40 vectors include pSVT7 and pMT2. In some embodiments, vectors derived from bovine papilloma virus include pBV-lMTHA, and vectors derived from Epstein Bar virus include pHEBO, and p2O5. Other exemplary vectors include pMSG, pAV009 / A+, pMTO10 / A+, pMAMneo- 5, baculovirus pDSVE, and any other vector allowing expression of proteins under the direction of the SV-40 early promoter, SV-40 later promoter, metallo thionein promoter,murine mammary tumor virus promoter, Rous sarcoma virus promoter, polyhedrin promoter, or other promoters shown effective for expression in eukaryotic cells.

[0090] In some embodiments, recombinant viral vectors, which offer advantages such as lateral infection and targeting specificity, are used for in vivo expression. In one embodiment, lateral infection is inherent in the life cycle of, for example, retrovirus and is the process by which a single infected cell produces many progeny virions that bud off and infect neighboring cells. In one embodiment, the result is that a large area becomes rapidly infected, most of which was not initially infected by the original viral particles. In one embodiment, viral vectors are produced that are unable to spread laterally. In one embodiment, this characteristic can be useful if the desired purpose is to introduce a specified gene into only a localized number of targeted cells.

[0091] Various methods can be used to introduce the expression vector of the present invention into cells. Such methods are generally described in Sambrook et al., Molecular Cloning: A Laboratory Manual, Cold Springs Harbor Laboratory, New York (1989, 1992), in Ausubel et al., Current Protocols in Molecular Biology, John Wiley and Sons, Baltimore, Md. (1989), Chang et al., Somatic Gene Therapy, CRC Press, Ann Arbor, Mich. (1995), Vega et al., Gene Targeting, CRC Press, Ann Arbor Mich. (1995), Vectors: A Survey of Molecular Cloning Vectors and Their Uses, Butterworths, Boston Mass. (1988) and Gilboa et at. [Biotechniques 4 (6): 504-512, 1986] and include, for example, stable or transient transfection, lipofection, electroporation and infection with recombinant viral vectors. In addition, see U.S. Pat. Nos. 5,464,764 and 5,487,992 for positive-negative selection methods.

[0092] In some embodiments, the kit comprises at least two different fusion proteins. In some embodiments, the different fusion proteins comprise different proteins. In some embodiments, the different fusion proteins comprise the same nuclease. In some embodiments, at least two is at least 3, 4, 5, 6, 7, 8, 9 or 10 fusion proteins. Each possibility represents a separate embodiment of the invention. In some embodiments, all the fusion proteins are different. In some embodiments, each fusion protein is expressed by a different cell. In some embodiments, the proteins of the fusion proteins are the proteins that associate with the binding sites in the target sequence. In some embodiments, the nuclease of each fusion protein is the same, and the protein of interest in each fusion protein is different.

[0093] By another aspect, there is provided a method of determining protein association with a nucleotide sequence, the method comprising:a. receiving a nucleic acid molecule library; b. contacting the library with a cell expressing a fusion protein; c. amplifying the nucleic acid molecules; and d. sequencing the amplified nucleic acid molecules; thereby determining protein association with a nucleotide sequence.

[0094] By another aspect, there is provided a method of determining nucleotide sequence association with a peptide, the method comprising: a. receiving a nucleic acid molecule library; b. contacting the library with a cell; c. amplifying the nucleic acid molecules; and d. sequencing the amplified nucleic acid molecules; thereby determining nucleotide sequence association with a peptide.

[0095] In some embodiments, the method is a diagnostic method. In some embodiments, the method is an in vitro method. In some embodiments, the method is an ex vivo method. In some embodiments, the method is performed in culture. In some embodiments, the method is a method of determining protein association. In some embodiments, determining is measuring. In some embodiments, determining is quantifying. In some embodiments, determining is testing. In some embodiments, the protein association is with a target sequence. In some embodiments, protein association is with variants of a target sequence. In some embodiments, the method is a method of determining off-target binding of the protein. In some embodiments, off-target is to a sequence other than the target sequence. In some embodiments, off-target is to a sequence other than the binding site. In some embodiments, the determining is determining protein association with a plurality of nucleotide sequences. In some embodiments, the determining is simultaneously determining association with a plurality of nucleotide sequences. In some embodiments, the determining is determining nucleotide sequence association with a plurality of peptides. In some embodiments, the determining is simultaneously determining association with a plurality of peptides. In some embodiments, the plurality is at least 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 35, 40, 45, 50, 60, 70, 80, 90 or 100. Each possibility represents a separate embodiment of the invention. In some embodiments, the plurality is at least 4 peptides. In some embodiments, the plurality is at least 20 nucleotide sequences. In some embodiments, protein association with anucleotide sequence is protein association with all nucleotide sequences. In some embodiments, nucleotide sequence association with a peptide is nucleotide sequence association with all peptides.

[0096] In some embodiments, association with is binding to. In some embodiments, association is occupancy at. In some embodiments, association with is being recruited to. In some embodiments, binding is direct binding. In some embodiments, binding is binding via an intermediate. In some embodiments, the intermediate is another protein. In some embodiments, the intermediate is a nucleic acid molecule.

[0097] In some embodiments, the nucleic acid molecule library is a nucleic acid molecule library of the invention. In some embodiments, all molecules of the library comprise 5’ and 3’ primer recognition sites. In some embodiments, the target sequence is between the 5’ and 3’ primer recognition sites. In some embodiments, a putative target sequence of the protein is between the 5’ and 3’ primer recognition sites.

[0098] In some embodiments, the library is contacted with a cell. In some embodiments, a cell is a plurality of cells. In some embodiments, the cell is a population of cells. In some embodiments, the cell expresses a fusion protein. In some embodiments, contacting comprises expression of the library in the cell. In some embodiments, contacting comprises infection of the cell with the library. In some embodiments, contacting comprises transfecting the cell with the library. In some embodiments, contacting comprises delivering the library into the cell. In some embodiments, into the cell is into the cytoplasm of the cell.

[0099] In some embodiments, the method further comprises incubating the cell and the library. In some embodiments, the incubating is after the contacting. In some embodiments, the method further comprises incubating the cell comprising the library. In some embodiments, the incubating is of the cell with the library. In some embodiments, the incubating is of the library in the cell. In some embodiments, the incubating is for a period of time sufficient for cleavage. In some embodiments, cleavage is cleavage of a nucleic acid molecule by the fusion protein. In some embodiments, cleavage is cleavage of the nucleic acid molecules of the library. In some embodiments, cleavage is cleavage of a nucleic acid molecule associated with the fusion protein. In some embodiments, cleavage is cleavage of a nucleic acid molecule bound by the fusion protein. In some embodiments, the incubating produces cleaved and non-cleaved nucleic acid molecules. In some embodiments, the sufficient time is at most 30, 45, 60, 90, 120, 150, 180, 210, 240, 270, 300, 330, 360, 390, 420, 450, 480, 510, 540, 570, 600, 630, 660, 690, 720, 750, 780, 810, 840, 870 or 900seconds. Each possibility represents a separate embodiment of the invention. In some embodiments, the sufficient time is at most 60 seconds. In some embodiments, the sufficient time is at most 300 seconds. In some embodiments, the sufficient time is at most 360 seconds. In some embodiments, the sufficient time is at most 600 seconds.

[0100] In some embodiments, the method further comprises amplifying the nucleic acid molecules. In some embodiments, the nucleic acid molecules of the cell are amplified. In some embodiments, the amplification produces amplification products. Methods of nucleic acid amplification are well known in the art and any method of amplification may be used, these include PCR, and isothermal amplifications such as RPA, RCA, LAMP and WGA. Any such method may be used. In some embodiments, the amplifying occurs in the cell. In some embodiments, the nucleic acid molecules from the cell are amplified. In some embodiments, the method further comprises lysing the cell. In some embodiments, the amplification occurs in the lysate. In some embodiments, the method further comprises isolating the nucleic acid molecules. In some embodiments, DNA is isolated. In some embodiments, RNA is isolated. In some embodiments, the amplification is within the isolated nucleic acid molecules. In some embodiments, the method is devoid of an isolation step.

[0101] In some embodiments, the amplification is with the 5’ primer. In some embodiments, the amplification is with the 3’ primer. In some embodiments, the amplification is with the 5’ and 3’ primer pair. In some embodiments, the amplification produces amplification products. In some embodiments, the amplification products comprise both the 5’ primer recognition site and the 3’ primer recognition site. In some embodiments, the amplification products further comprise the target sequence. In some embodiments, the amplification products further comprise the target sequence or variant of the target sequence. In some embodiments, the amplification products further comprise the target sequence and the coding sequence. In some embodiments, cleaved molecules are not amplified. In some embodiments, the isolation products are isolated. In some embodiments, isolation is by size isolation.

[0102] In some embodiments, the amplification products are sequenced. In some embodiments, the intact nucleic acid molecules are sequenced. In some embodiments, the non-cleaved nucleic acid molecules are sequenced. In some embodiments, the sequencing produces sequencing data. In some embodiments, the sequencing is next generation sequencing. In some embodiments, the sequencing is high-throughput sequencing. In someembodiments, the sequencing is massively parallel sequencing. It will be understood by a skilled artisan that any method that quantifies sequencing reads can be used.

[0103] In some embodiments, sequences depleted in the sequencing data are sequences associated with the protein. In some embodiments, sequences depleted in the sequencing data are sequences bound by the protein. In some embodiments, depleted is as compared to a control sequence. In some embodiments, the control sequence is a sequence which is not cleaved. In some embodiments, the library further comprises a control molecule. In some embodiments, the control molecule comprises a control sequence. In some embodiments, the control sequence is a target sequence in which the binding site has been completely mutated such that the protein does not bind. In some embodiments, the control sequence is a plurality of other variant sequences. In some embodiments, the plurality of other variant sequences are sequences showing the highest level of sequencing reads. In some embodiments, a control sequence is the sequence with the highest number of reads. In some embodiments, a control sequence is sequence with reads in the top percentage of total reads. In some embodiments, the top percentage is the top 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 12, 15, 17, 20, or 25% of molecules. Each possibility represents a separate embodiment of the invention.

[0104] In some embodiments, the depletion is at least a 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 92, 95, 97, 99 or 100% depletion. Each possibility represents a separate embodiment of the invention. In some embodiments, the depletion is at least a 20% depletion. In some embodiments, the depletion is relative depletion. In some embodiments, the relative depletion is proportional to the association. In some embodiments, the relative depletion of a given sequence is proportional to the association of the protein to the given sequence. In some embodiments, the relative depletion of a given sequence is proportional to the occupancy of the protein at the given sequence. It will be understood that the greater the depletion of the sequence the greater the occupancy / association of the protein with the binding site. In some embodiments, a given sequence is a given binding site.

[0105] In some embodiments, the method further comprises contacting the cell with a substrate. The use of a substrate to activate the nuclease is advantageous over constitutively active nucleases, as the timing of all cleavage can be synchronized, leading to more uniform and easier to interpret results. In some embodiments, the substrate is contacted after the library is contacted. In some embodiments, the substrate is contacted before the incubation. In some embodiments, the substrate is contacted at the beginning of the incubation. In some embodiments, incubation is incubation with the substrate. In some embodiments, the substrate is a substrate of the nuclease. In some embodiments, the substrate induces nucleaseactivity. In some embodiment, the substate induces the activity of the nuclease. In some embodiments, the substrate induces activity of the nuclease in the fusion protein. In some embodiments, the substrate induces nuclease activity of the fusion protein. In some embodiments, the substrate is specific to the nuclease. In some embodiments, the substrate is calcium. In some embodiments, the substrate is magnesium. In some embodiments, the substrate is sodium. In some embodiments, the substrate is both calcium and magnesium. In some embodiments, the substrate is selected from calcium and magnesium. In some embodiments, the nuclease is MNase and the substrate is calcium. In some embodiments, the nuclease is MNase and the substrate is magnesium. It is well known in the art that both calcium and magnesium can induce / activate MNase.

[0106] As used herein, the term "about" when combined with a value refers to plus and minus 10% of the reference value. For example, a length of about 1000 nanometers (nm) refers to a length of 1000 nm+- 100 nm.

[0107] It is noted that as used herein and in the appended claims, the singular forms "a," "an," and "the" include plural referents unless the context clearly dictates otherwise. Thus, for example, reference to "a polynucleotide" includes a plurality of such polynucleotides and reference to "the polypeptide" includes reference to one or more polypeptides and equivalents thereof known to those skilled in the art, and so forth. It is further noted that the claims may be drafted to exclude any optional element. As such, this statement is intended to serve as antecedent basis for use of such exclusive terminology as "solely," "only" and the like in connection with the recitation of claim elements, or use of a "negative" limitation.

[0108] In those instances where a convention analogous to "at least one of A, B, and C, etc." is used, in general such a construction is intended in the sense one having skill in the art would understand the convention (e.g., "a system having at least one of A, B, and C" would include but not be limited to systems that have A alone, B alone, C alone, A and B together, A and C together, B and C together, and / or A, B, and C together, etc.). It will be further understood by those within the art that virtually any disjunctive word and / or phrase presenting two or more alternative terms, whether in the description, claims, or drawings, should be understood to contemplate the possibilities of including one of the terms, either of the terms, or both terms. For example, the phrase "A or B" will be understood to include the possibilities of "A" or "B" or "A and B."

[0109] It is appreciated that certain features of the invention, which are, for clarity, described in the context of separate embodiments, may also be provided in combination in a singleembodiment. Conversely, various features of the invention, which are, for brevity, described in the context of a single embodiment, may also be provided separately or in any suitable sub-combination. All combinations of the embodiments pertaining to the invention are specifically embraced by the present invention and are disclosed herein just as if each and every combination was individually and explicitly disclosed. In addition, all subcombinations of the various embodiments and elements thereof are also specifically embraced by the present invention and are disclosed herein just as if each and every such sub-combination was individually and explicitly disclosed herein.

[0110] Additional objects, advantages, and novel features of the present invention will become apparent to one ordinarily skilled in the art upon examination of the following examples, which are not intended to be limiting. Additionally, each of the various embodiments and aspects of the present invention as delineated hereinabove and as claimed in the claims section below finds experimental support in the following examples.

[0111] Various embodiments and aspects of the present invention as delineated hereinabove and as claimed in the claims section below find experimental support in the following examples.EXAMPLES

[0112] Generally, the nomenclature used herein and the laboratory procedures utilized in the present invention include molecular, biochemical, microbiological and recombinant DNA techniques. Such techniques are thoroughly explained in the literature. See, for example, "Molecular Cloning: A laboratory Manual" Sambrook et al., (1989); "Current Protocols in Molecular Biology" Volumes I-III Ausubel, R. M., ed. (1994); Ausubel et al., "Current Protocols in Molecular Biology", John Wiley and Sons, Baltimore, Maryland (1989); Perbal, "A Practical Guide to Molecular Cloning", John Wiley & Sons, New York (1988); Watson et al., "Recombinant DNA", Scientific American Books, New York; Birren et al. (eds) "Genome Analysis: A Laboratory Manual Series", Vols. 1-4, Cold Spring Harbor Laboratory Press, New York (1998); methodologies as set forth in U.S. Pat. Nos. 4,666,828; 4,683,202; 4,801,531; 5,192,659 and 5,272,057; "Cell Biology: A Laboratory Handbook", Volumes I- III Cellis, J. E., ed. (1994); "Culture of Animal Cells - A Manual of Basic Technique" by Freshney, Wiley-Liss, N. Y. (1994), Third Edition; "Current Protocols in Immunology" Volumes I-III Coligan J. E., ed. (1994); Stites et al. (eds), "Basic and Clinical Immunology" (8th Edition), Appleton & Lange, Norwalk, CT (1994); Mishell and Shiigi (eds), "Strategiesfor Protein Purification and Characterization - A Laboratory Course Manual" CSHL Press (1996); all of which are incorporated by reference. Other general references are provided throughout this document.Materials and Methods

[0113] Strains: All used TF-MNase strains were built on the BY4741 strain and described in Brodsky et al. , “Intrinsically disordered regions direct transcription factor in vivo binding specificity”, Mol Cell, 2020 Aug 6;79(3):459-471.e4, the contents of which are hereby incorporated by reference herein in their entirety.

[0114] Library plasmid context: For changing the surrounding library context, four genomic loci bound by Msn2 were chosen based on their different nucleosome architecture and the total number of TFs bound. In each context, the central 120 bp promoter region containing the Msn2 binding sites was removed and replaced with Sgsl and Csil cutting sites. Next, the remaining 2 kb surrounding the binding site (1 kb upstream and 1 downstream) were introduced with point mutations to edit any existing Sgsl / Csil binding sites. In addition to these four genomic loci, another synthetic context was planned, in which each side of the cloning site has 200 bp with a low GC content (20%), aimed to evict nucleosomes. All contexts were ordered from IDT (gBlock) and integrated into the pRS316 plasmid between the Xhol and Spel cutting sites. The main plasmid sequence is provided in SEQ ID NO: 7.

[0115] Library cloning: Libraries were ordered from IDT (4 nmole Ultramer™ DNA Oligos) and resuspended in 40 pl of Tris pH 8 (to a final concentration of 100 pM). To amplify the libraries, we first calibrated the PCR reaction (Forward Primer: tgatcACCAGGTCGATGCGCATGCGTACGC; SEQ ID NO: 2 and Reverse primer: tgatcGGCGCGCCGCTCGGCAAGTCGTGTCC; SEQ ID NO: 3) using four dilutions for each library (reaching to final reaction concentrations of 44.4, 4.4, 1.8 and 0.9 nM) in order to avoid non-specific amplification products. The PCR reaction was done using Herculase II Fusion DNA Polymerase (Agilent, 600677) and was performed according to the recommendation of the manufacturer for 11 cycles. The concentration yielding the sharpest and most specific band was chosen for each library. This reaction was repeated twice more (a total of three reactions for each library to minimize PCR biases). Reactions were combined and cleaned using the QIAQuick PCR purification kit (QIAGEN, 28104). Concentrations were measured using nanodrop machine and ranged 20-120 ng / ul.

[0116] Digestion: Plasmids and libraries were digested using the restriction enzymes Sgsl and Csil (Thermofisher, FD1894 and FD2114, respectively). For the restriction, we used 180ng of the library and 2 ug of our 5 kb plasmid. Plasmids went through an additional step of 5' dephosphorylation using FastAP Thermosensitive Alkaline Phosphatase (Thermofisher, EF0654) to avoid self-ligation. Their cleavage was validated by comparing them to an uncut control plasmid using gel electrophoresis. Reactions were performed according to the recommendation of the manufacturer. Cleaved plasmids and oligos were then cleaned using the QIAQuick PCR purification kit and DNA concentrations were measured (2.5-10 ng / pl for libraries and 50-100 ng / pl for plasmids).

[0117] Ligation: Libraries and plasmids were ligated together using the fast-link DNA digestion kit (Biosearch technologies, LK0750h) in a respective ratio of 2:1, using 225 ng plasmid, and according to the manufacturer recommendation. Ligations were cleaned using the MinElute PCR Purification Kit (QIAGEN, 20-28004), with an additional wash with the PE buffer to avoid high salt concentration. Ligated plasmids were elutied twice to increase yield.

[0118] Bacterial transformation: E. cloni® Electrocompetent Cells (Biosearch technologies, LC601172) were used according to the provided protocol with some modifications. Cells were thawed over ice, and the entire process (until the recovery step) was performed in the cold (4°C). After the recovery step, cells were directly inoculated in 250 ml warm lysogeny broth (LB) media, and grown overnight (shaking, 37°C). A small amount of cells were plated for counting the total number of transformants (5-30M for transformants for each library); 16-20 colonies from each transformation were collected and library insertion was validated using PCR (Forward Primer: aggcgtatcacgaggccctttc; SEQ ID NO: 4 and Reverse Primer: accgtcatcaccgaaacgcg; SEQ ID NO: 5). The integration success rate was > 95%.

[0119] Plasmid extraction: Plasmids were extracted using the NucleoBond Xtra Maxi kit (Machery-Nagel, 740414.10), according to the provided protocol, yielding 1500 ng / pl on average.

[0120] Yeast transformation: To transform yeast with plasmids carrying the library variants the BY4741 strain, of genotype MATa his3-Al leu2-A0 lys2-A0 metl5-A0 ura3-A0, bearing a TF-micrococcal nuclease (TF-MNase) fusion, was transformed using the lithium acetate / single-stranded carrier DNA / polyethylene glycol (LiAc / SS DNA / PEG) method (22), with several modifications to increase transformant number and avoid library complexity issues. Briefly, the relevant yeast strain was freshly thawed from a frozen stock, plated on yeast extract peptone dextrose (YPD) plates and grown at 30°C. A single yeast colony was inoculated in fresh liquid YPD and grown to saturation overnight. The next morning, 250 plof the stationary culture was diluted into 12.5 ml YPD media and grown at 30°C for 4-5 h. The cells were washed with double distilled water (DDW) and then with LiAc 100 mM, and resuspended in a transformation mix, containing 33% PEG-3350, 100 mM LiAc, single stranded salmon sperm DNA, 10 ug of library plasmids and 700 ng of the plasmid containing the poorly bound sequence control. Of note, while multiplexing different libraries to the same transformation, concentrations were normalized according to the complexity of each library, based on the number of variable positions it contains. The cells were incubated at 30°C for 30 min followed by a 30-min heat shock at 42°C. After the heat shock, the cells were pelleted and resuspended in 1 ml SD-URA media; 1:500 of the resuspended cells were further diluted and plated for calculating the number of transformants. These transformations yield ranged between 200K and 4M. The rest of the cells were inoculated in fresh 25 ml SD- URA and were grown to stationary phase for 72 h (shaking at 30°C).

[0121] Genomic library integration: First, the HO coding sequence in the Msn2-MNase expressing strain was replaced by a short target sequence (GGTGTGGGTTTAGATGACAAGGG; SEQ ID NO: 6) using the CRISPR-Cas9 pbRA89, a gift from James Haber. Next, after testing positive using Sanger sequencing, the pbRA89 plasmid was lost from cells by growth in YPD and selection for colonies without bRA89- encoded Hygromycin resistance. Then, the libraries were PCR amplified using homologic HO locus sequences flanking the primers. The PCR-amplified libraries were then transformed together with the bRA89 plasmid, this time containing a gRNA directed to the short sequence introduced in the first transformation. Cells were grown in liquid YPD media containing 50 mg / ml Hygromycin to select for positive colonies containing the library embedded within the HO locus.

[0122] Experimental procedure: Library-transformed stationary culture was diluted into 10 ml of fresh SD-URA media (per sample) to reach OD600 of 4 after ~10 cell divisions at 30°C. Note that that this method obliges to compare the change of the variant pool following MNase activation. Therefore, at the beginning of the experiment, each sample was split into two 5 ml samples that were subjected to different treatments: the non-activated sample (time point 0) will be taken directly to DNA extraction (see below), while the other sample will be subjected to MNase activation (see below), until the Proteinase K digestion step, and only then is DNA extracted.

[0123] MNase activation: This procedure was performed as previously described. In brief, cultures (OD600 = 4) were pelleted at 1500 g and resuspended in 1 ml Buffer A (15 mM Tris pH 7.5, 80-mM KC1, 0.1 mM egtazic acid (EGTA), 0.2 mM spermine, 0.5 mMspermidine, lx Roche complete ethylenediaminetetraaceticacid (EDTA)-free mini protease inhibitors, 1 mM phenylmethylsulfonyl fluoride (PMSF)), and then transferred to deep well plate. Cells were washed twice more in 500 pl Buffer A, pelleted and resuspended in 150 pl Buffer A containing 0.1% digitonin. Then, cells were transferred to a 96-well plate (PCR- 96-FLT-C, Axygen) for permeabilization (30°C for 5 min). CaC12 was added to a final concentration of 2 mM. Mnase was activated for different times, as specified in the text. Next, 100 pl of stop buffer (400-mM NaCl, 20-mM EDTA, 4 mM EGTA and 1% SDS) were mixed with 100 pl of sample. Proteinase K was then added, and incubated at 55°C for 30 min.

[0124] DNA extraction: Plasmid DNA was extracted both for MNase-activated cells and non-activated cells using the MasterPure Yeast DNA Purification Kit (Lucigen Corporation, MPY80200), with some modifications. Briefly, 300 pl of lysis buffer supplemented with RNase A (as recommended) were added to each sample. For MNase-activated cells, 300 pl lysis buffer were added directly to the Proteinase K-treated samples. For non-activated samples, the OD600 4 was first pelleted in 1500 g to remove SD-URA media, then cells were resuspended in 300 pl lysis buffer. Following 15 min of incubation in 65°C, cells were moved into EoBind microcentrifuge tubes (Eppendorf®, 022431021) containing 0.5 mm zirconium oxide beads (ZrOB05, Next Advance) and blended using the Bullet Blender 24 (Next Advance) for a single cycle of 3 min in level 8. Samples were cooled over ice and lysates were moved to 1.5 ml tubes, vortexed and supplemented with 300 pl MPC buffer. Samples were vortexed vigorously and centrifuged for 10 min (17 000g). Supernatant was moved to 1.5 ml tubes containing 500 pl isopropanol. Samples were vortexed vigorously and DNA was precipitated (10 min, 17 000g). Isopropanol was removed and pellets were washed in 70% EtOH, then air dried and finally resuspended in TE buffer.

[0125] Library preparation: Samples were 1.5x SPRI (AMPure XP, A63881) cleaned and resuspended in 30 pl elution buffer (10-mM Tris-HCl, pH 8) and then diluted 1 :20; 1 pl was used as a template for PCR reaction for amplifying non-cleaved variants and sample barcoding (28 cycles) using the Phusion® High-Fidelity PCR Master Mix (NEB, M0531L). Samples were verified using gel electrophoresis and 0.5 reversed-SPRI cleaned. DNA concentrations were measured using Qubit™ Flex Fluorometer (Invitrogen) and samples were normalized to equal amounts, then pooled together; 1 ng was taken as a template for a PCR reaction using KAPA HiFi HotStart ReadyMix (Roche, KK2602) for adding Illumina indices and machine adapters. Samples were lx SPRI cleaned, then concentrations and library quality were validated using Qubit and tape station (Agilent).

[0126] Sequencing: Libraries were sequenced using the NovaSeq 6000 machine. Runs were performed with the SP200 kit (20040719); parameters: R1 - 150 cycles, Indexl - 8 cycles, Index2 - 8 cycles, R2 - 6 cycles. In each run, 5% PhiX DNA was added to increase complexity.

[0127] ATAC-seq

[0128] Experimental procedure: Cells were grown as in MPBA, to reach OD600= 4, and were then pelleted in 1.5 ml Eppendorf tubes at 1500 g for 60 s and washed twice with 200- pl Spheroplasting Buffer (SB) (40-mM HEPES [pH 7.5], 10-mM MgC12 1 M Sorbitol). The cells were next resuspended in 1 pl Lyticase lOku (Sigma, L2524) dissolved in 199 pl SB, transferred to a 96-well plate and incubated for 30 min at 30°C. The spheroplasted cells were pelleted at 1500 g for 60 s, washed twice with SB as in the previous step, resuspended in 2 pl of RNase-A (5 pg / pl) dissolved in 98 pl SB and incubated for 30 min at 37°C. Cells were then pelleted and washed twice as in the previous steps, resuspended on ice in 25 pl Transposase reaction mix (Tn5 transposase enzyme produced in house [6.5pM] loaded Tn5MEDS-A / B oligonucleotides, 2x Tn5 Buffer [10 mM Tris HCL PH 7.6, 10 mM MgC12, DMF in 1:4 ratio], DDW) and transferred to a 96-well plate for 60 min incubation at 37°C. Transposase reaction clean-up mix (5% SDS, 20mg / ml Proteinase K, clean-up buffer [NaCl 5M, EDTA 0.5M, DDW]) was added, and the plate was incubated for another 30 min at 40°C. Samples were 2x SPRI (AMPure XP, A63881) cleaned and resuspended in 22 pl elution buffer (10-mM Tris-HCl, pH 8).

[0129] Library preparation: The eluted samples were used as a template for a PCR reaction for amplifying Tn5 tagmented fragments and sample barcoding (2x Kappa HIFI [Roche], 5 pl nextera primers F: i5, R: i7; 12 cycles). Next, samples were 0.65 x reversed SPRI cleaned and resuspended in 22 pl elution buffer (10-mM Tris-HCl, pH 8), then concentrations and library quality were validated using Qubit and tape station (Agilent).

[0130] Sequencing: Libraries were sequenced using the NovaSeq 6000 machine. Runs were performed with the S 1 kit (20028319), parameters: R1 - 61 cycles, Indexl - 8 cycles, Index2 - 8 cycles, R2 - 61 cycles.

[0131] MPBA Pipeline: The pipeline was built using SnakeMake and can be found in Zenodo. Briefly, the forward library primers were removed using cutadapt. Then, AdapterRemoval was used to demultiplex each 32-sample pool into separate samples. Next, in order to demultiplex the different libraries in each sample, each read was assigned to the right library using cutadapt. Reads with improper 3' prime alignment were filtered out inorder to remove sequences containing indels. Then, each combination of variable positions within each library was counted. In addition, the number of reads in each sample aligning to the poorly bound sequence control was counted.

[0132] MPBA Data filtering: The read count of each sample was normalized to 10A6 to account for sequence depth. Next, technical repeats were averaged and log2 transformed. Then, each MNase-activated sample was normalized to the corresponding pre-activated sample. Lastly, biological repeats were averaged.

[0133] MPBA Data normalization: The read count of each sample was normalized to 106 to account for sequence depth. Next, technical repeats were averaged and log2 transformed. Then, each MNase-activated sample was normalized to the corresponding pre-activated sample. Lastly, biological repeats were averaged.

[0134] Occupancy measure: To calculate the relative occupancy, each time -point zero normalized sequence was normalized to the fully mutated sequences and converted to % using the following equation:(l-(0.5A-(value of tpO normalized Sequence - tpO normalized fully mutated sequence value))).

[0135] Normalized occupancy measure: To calculate histone occupancy compared to the genome, a parameter representing the highest nucleosome occupancy of library-sized promoter region (164 bp) was calculated. To this end, the MNase-seq (27) signal was summed over all library-sized regions promoter using a sliding window of 10 bp. The top 1% of this calculated distribution served as the value of a 100% nucleosome occupancy. The ratio between the MNase-seq signal over each library to this value was calculated and served as the genomic occupancy over each tested region. For calculating the nucleosome occupancy over each library sequence, as presented in Figure 4G, the genomic occupancy and the fold change between each sequence and the wild-type one was used, as follows:0.5A(value of sequence normalized to tpO & wild-type variant value) * (Library nucleosome occupancy).

[0136] Cooperativity score: For a given pair of motifs, the occupancy (in fraction) of each motif in the absence of the second motif was calculated in each context (all combinations of all non-pair motifs found in the respective library). Specifically, for each pair in each context, the log2 fold-change between the sequences containing only one intact motif and the sequence where both are mutated was calculated and converted to occupancy (as in the‘Occupancy measure’ section). Then, the expected cooperativity, assuming independence, of the motif pair was calculated for each context as follows:1 - (1-OccupancymotifA) * (l-OccupancymotifB).

[0137] The observed motif pair occupancy was calculated for each context using the log2 fold change between the sequence in which both motifs are intact to the one where both motifs are mutated. The cooperativity score is the average difference between the observed and expected occupancy over all contexts.

[0138] Genome-wide promoter binding: ChEC-seq Samples were processed as in Kumar et al., “Complementary strategies for directing in vivo transcription factor binding through DNA binding domains and intrinsically disordered regions”, Mol Cell, 2023 May 4;83(9): 1462-1473. e5, the contents of which are hereby incorporated by reference herein in their entirety. Promoters were defined as in Brodsky et al. Each promoter was assigned with the average sum of signal it received over all repeats.

[0139] Library region genomic binding: For the correlation heatmap presented in Figure 6A, the sum of signal received for each TF over each genomic region used to generate libraries was summed. Next, the sum of signal received for all these library regions was correlated between the different TFs. For the region binding strength heatmap presented in Figure 6B, the signal received over each genomic region representing a library was summed and divided by the highest signal received on a similar-sized region (164 bp) over the entire genome of the respective TF.

[0140] Motif enrichment: For the motif enrichment analysis, all possible 7-mer sequences were given a numerical index (forward and reverse complement forms of each 7 -mer were given the same index). Each nucleotide in the yeast genome was indexed according to the 7- mer that begins from it. To score each 7-mer occurrence, the signal around its mid -position was averaged (20 bp window). The averaged signal for each 7-mer was then calculated across all of its occurrences in all promoters and was assigned as its relative binding score. Next, the relative binding score of all 7-mers containing the relevant TF motif (Stel2 - TGAAAC, Teel - GAATG[CT], Rebl - CGGGTAA, Table 1) was averaged and divided by the average signal of all 7mer combinations.

[0141] Motif Euclidean distance: The Euclidean distances between the motifs of the TFs presented in Figure 6G were calculated using the Probability weight matrices (PWMs) from Jana et al., “Speed-specificity trade-offs in the transcription factors search for their genomic binding sites”, Trends Genet, 2021 May;37(5):421-432, which is the same as in Krieger etal., “Independent evolution of transcript abundance and gene regulatory dynamics”, Genome Res. 2020 Jul; 30(7): 1000-1011, the contents of which are hereby incorporated by reference herein in their entirety.

[0142] Non-cooperative recruitment: To measure the non-cooperative recruitment presented in Figures 5A and 6F, the average effect of each non-canonical motif in each library was calculated. In detail, after data normalization, the average log2 fold-change of all sequences in which the respective non-canonical motif is intact and the self-TF positions are mutated was normalized by the average log2 fold-change of all sequences in which the respective non-canonical motif is mutated, and the self-TF positions are mutated. Then, the relative position occupancy was calculated as in the ‘Occupancy measure’ section.

[0143] Self-motif-dependent region occupancy: To calculate the self-motif-dependent region occupancy presented in Figure 6E, after data normalization, the average log2 foldchange of all sequences in which all self-TF motifs are intact was normalized to the log2 fold-change of all sequences in which all self-TF motifs are mutated. Then, the relative occupancy was calculated as in the ‘Occupancy measure’ section.

[0144] Number of TFs bound to each context: For this aim, lab data containing the genomic binding signal of most yeast TFs was used. For each TF, the sum of normalized signal received on each promoter was calculated and z-score transformed. Then, the top number of TFs bound to the central promoter found in each context was calculated, asking how many TFs received a z-score > 3.

[0145] ATAC seq: As a first step in analyzing ATAC-seq read counts, the reads were aligned to the S. cerevisiae genome (SacerR64) using bowtie2 with the following parameters: ‘very sensitive’, ‘trim to 30’ and ‘dovetail’. Reads larger than 140 bp were discarded using genome coverage. The coverage was calculated on the 5' position of each read. The repetitive region in chromosome 12 (451600:489500) was removed and the reads of each sample were normalized to 10A7. The presented data was smoothed using 20 bp window size in all figures, except from Figure 4F in which 50 bp window was used.

[0146] In-vitro and in-vivo MNase-seq: The Tn Vitro’ and ‘YPD’ normalized data samples were taken from (Kaplan et al., “The DNA-encoded nucleosome organization of a eukaryotic genome”, Nature, 2009; 458:362-366, the contents of which are hereby incorporated by reference in their entirety). The genomic annotations were converted to the SacerR64 genome using CrossMap (Zhao, et al., “CrossMap: a versatile tool for coordinate conversionbetween genome assemblies”, Bioinformatics, 2014; 30: 1006-1007, the contents of which are hereby incorporated by reference in their entirety).Example 1: Massively Parallel Binding Assay (MPBA) for rapid screening of thousands of DNA variants

[0147] Analyzing motif interactions within regulatory regions requires comparing the in- vivo binding of a specific TF to a library of DNA variants, systematically perturbing all motif combinations (Fig. 1A). Herein there is provided Massively Parallel Binding Assay (MPBA), a sequencing-based method that enables measuring TF binding to thousands of designed DNA sequences within cells using a rapid, temporally controlled readout. MPBA was applied to measure motif interactions within dozens of 164 bp-long regulatory regions, each containing a total of 3-7 motifs bound by multiple TFs. Some cases of cooperative interactions are described herein, but those were found to be the exceptions. Rather, TF binding was explained primarily by the independent contributions of individual, mostly selfpreferred, motifs.

[0148] Across large genomes, a single short motif is found in hundreds of thousands of places, while the number of places containing a particular combination of motifs is significantly lower. Motif combinations could therefore define unique locations across the genome. In the case of TF binding, this can be implemented by interacting TFs that co-bind to adjacent motifs. Co-binding, by itself, however, is not sufficient for increasing binding specificity; for example, a TF that is recruited to a site bound by its interacting partner will still depend on the presence of a single motif, this time the motif of the recruiting partner, providing no advantage in restricting binding locations. Binding specificity, therefore, will only increase if both motifs are needed, that is, when the motif combination stabilizes the binding of the individual sites. We therefore designed our study to measure motif cooperativity, that is, define the extent to which proximal motifs cooperate in stabilizing TF binding.

[0149] Budding yeast offers a convenient platform for examining motif cooperation at the genomic scale, as its genome is compact and well-characterized, and its TFs are well annotated and of known binding motifs. Further, through different studies, our lab has accumulated genomic locations of -140 (95%) budding yeast TFs, all obtained under the same conditions and using a unified spatially resolved method (ChEC-seq; see Brodsky, et al., 2020, “Intrinsically disordered region direct transcription factor in vivo binding specificity”, Molecular Cell 79, 459-471; and Gera, et al., 2022, “Evolution of bindingpreferences among whole-genome duplicated transcription factors”, Elife, l l :e73225, the contents of which are hereby incorporated by reference in their entirety).

[0150] MPBA was developed for comparing the binding of a given TF to a library of designed DNA sequences. This method follows the conceptual design of massively parallel reporter assays (MPRA) but replaces the expression-based output of MPRA by a more direct readout of TF binding to DNA.

[0151] In MPBA, the endogenous TF of interest is fused to a MNase, which cleaves TF- bound DNA upon a short calcium pulse, as described in previous specific applications (see Schmid et at., “ChIC and ChEC; genomic mapping of chromatin proteins”, Mol. Cell, 2004, 16: 147-157, the content of which are hereby incorporated by reference in their entirety). This method was adapted to enable variant screening as follows (Fig. IB). First, the library variants are integrated into a specialized plasmid and the plasmids are transformed into cells carrying the TF-MNase fusion. Next, the MNase is activated to trigger the cleavage of TF- bound plasmids. Following this step, the library variants are PCR amplified, thereby capturing only the non-cleaved (unbound) sequences. Then, the frequency of each variant is defined using high-throughput sequencing. Finally, the abundance of each sequence is normalized to its abundance prior to MNase activation and then to that of the fully mutated sequence, lacking all motifs. This provides a measure for TF occupancy at each sequence within the library, namely the proportion of cells in which this sequence is bound.Example 2: MPBA captures motif binding and nucleosome eviction by the general regulatory factors

[0152] Plasmids provide a convenient platform for massive parallel analysis of sequence libraries but might be limiting for capturing properties of TF binding inside cells, given their uncertain chromatin environment. Budding yeast differs from higher eukaryotes in lacking large-scale closed chromatin domains. Still, nucleosomes might be more loosely packed on a plasmid, which would reduce cooperative effects or, conversely, might be denser, limiting TF binding. As a first test of the method, the nucleosome-organizing general regulatory factors (GRFs) Rebl, Rapl and Abfl were therefore tested, asking whether MPBA can capture their tight motif binding and their role in nucleosome eviction. Of note, when compared to other TFs, the GRFs bind a large fraction of their motif sites, the majority of which are present in promoters deprived of other TFs (Fig. IE).

[0153] Six 164 bp GRF-bound genomic loci were selected, biasing towards the limited fraction of regions bound by additional TFs (Fig. IF). Six respective sequence libraries werethen designed as follows. First, for each of the n TF-bound motifs present in the selected region, two variants were considered the native motif and a 1 bp motif mutation abolishing TF binding (Fig. 1A and Table 1). Each library then includes all combinations of these intact or mutated motifs, totaling 2n sequences. For the GRFs, the six libraries include 32-128 sequences corresponding to the 5-7 TF motifs present in the selected regions (Fig. IF).

[0154] The designed libraries were integrated into a specialized plasmid (SEQ ID NO: 7), multiplexed and transformed into cells carrying the respective GRF fused to an MNase (Fig. IB). As expected, MNase activation led to rapid and reproducible depletion of sequences containing intact GRF motifs, while sequences lacking the motif remained at high abundances (Fig. 1C and 1G). In the case of Rebl, the depletion at 60 s of MNase activation translated into estimated occupancies of 85-92% at the wild-type sequence containing all intact motifs, which increased to a maximum of 95% in 3 min. The binding of Rapl and Abfl was again dependent on their known motifs, reaching somewhat lower estimated occupancies at similar activation times (3 min, 40-80% at the fully intact sequence). Of note, the binding of Abfl and Rebl increased with MNase activation time, with different kinetics depending on the tested TF and regulatory region (Fig. ID and 1G). Given the rapid enzymatic activity of the MNase, this temporal increase likely captures the dynamic transition of each TF between binding sites, although it could also result from TF-specific differences in MNase catalytic rate. It was concluded that MPBA captures the strong GRF motif binding and further provides, for the first time, a lower bound on the fraction of cells in which this motif is bound.

[0155] It was next asked whether the plasmid-based approach can also capture the GRF role in nucleosome eviction. To quantify nucleosome occupancy, the binding of the histone H3 to the six libraries was measured using MPBA (Fig. 2A). As expected, in five of these libraries, H3 eviction increased continuously with increasing GRF binding (Fig. 2B-2C). In these libraries, the GRF-inflicted ~ 15-55% decrease in H3 found in the assay is compatible with the average GRF effects in the genome, as estimated by comparing in vivo and in vitro nucleosome occupancy at the GRF binding sites (Fig. 2D-2E). It was concluded that MPBA captures not only the high GRF motif occupancy but also their role in nucleosome eviction.

[0156] Table 1: List of TFs and motifs (the notation “{}” refers to the number of possible repeats; the notation “[]” refers to an optional sequence)Example 3: Cross-motif dependencies: recruitment versus cooperation

[0157] Within regulatory sequences, TFs mostly bind to their own motif but may also depend on motifs of other TFs. In the data, this was exemplified by the binding of Abfl to the CRC1 regulatory region, which was mostly explained by a Rebl motif rather than its own (Fig. 1G). Given the motivation of defining the role of motif interactions in guidingTF-binding preferences, such cross-motif dependencies were of primary interest, distinguishing between two modes: recruitment and cooperativity (Fig. 2F). In recruitment, a motif that is not directly bound by the tested TF still promotes binding but independently of other motifs. This would be the case, for example, when the tested TF is recruited by a second TF that directly binds at this recruiting motif. Those cases, while interesting, do not increase specificity as TF binding still depends on a single motif. Cooperative binding, by contrast, refers to cases where two proximal motifs stabilize binding beyond their individual contributions. In these cases, specificity increases, as two nearby motifs are significantly less likely to occur at random compared to one.

[0158] The GRFs bind a large fraction of their motifs, which may indicate independent motif binding. Through their role in organizing nucleosomes, however, they may promote the binding of other TFs (Fig. 2G). Such cases were examined by testing the binding of the 10 additional TFs localizing to the six GRF-bound regulatory regions using MPBA. These TFs reached occupancies that were significantly lower than those of the GRFs, and in none of those cases was TF binding significantly influenced by the presence of the GRF motif (Fig. 2H). In fact, the only exception was the dependence of Abfl on the Rebl motif at the CRC1 promoter (noted above).

[0159] Next, a rigorous measure of motif cooperation was defined, which is assigned for each pair of motifs relative to the binding of a specific TF. For this, the sequences in which one motif in the respective pair is intact and the second is mutated were considered. Assuming that the contribution of those motifs to TF binding is independent, the occupancy at sequences in which both motifs are intact can be predicted. The contribution of motif cooperation is, therefore, the difference between the measured occupancy when both motifs are intact and the predicted one (Fig. 21). It should be noted that this measure may depend on the state of other TF motifs (wild type or mutated) found within the tested region. Therefore, the cooperativity scores assigned to a given motif pair were averaged over all possible states (wild type or mutated) of the other TF motifs found in the library.

[0160] As expected, in the case of the GRFs, cooperation scores were low and contributed less than 10% to the total occupancy (Fig. 2J). The only exception from linearity was an apparent repression, whereby the binding of Abfl to its two adjacent motifs at the GPP1 promoter was lower than expected (Fig. 2J-2K).Example 4: Motif cooperation explains the binding of the interacting Stel2-Tecl TFs

[0161] To verify that the method can detect binding cooperativity, the two known interacting TFs, Stel2 and Teel were considered. The heterodimers Stel2-Tecl bind to filamentous- related genes, while Stel2-Stel2 homodimers bind to mating -related ones. Accordingly, Stel2 and Teel share common targets and, when tested across the genome, preferentially localize to the motifs of one another (Fig. 3A-B). Seven regulatory regions bound by either or both TFs were selected, the TF-bound motifs in those regions were defined and corresponding libraries containing all combinations of motif perturbations were generated, as above (Fig 3C). Measuring Teel and Stel2 across those libraries revealed high occupancies at the intact sequences, reaching ~80% at the strongly bound regions, including the PCL2 and TEC1 promoters bound by both TFs, and the GPA1 and KAR4 regions bound only by Stel2 (Fig. 3G). Regions showing lower binding in the genome (POP3, SIM1 and CDC6) showed correspondingly low occupancy in our system (Fig. 3A and 3G).

[0162] Cooperation scores were measured for the 35 motif pairs covering all Stel2 and Teel motif combinations. Cases of strong cooperation were readily identified (Fig. 3D-3E). For example, the regulatory regions of PCL2 and TEC1 were strongly bound by both TFs, but only when both adjacent Stel2 and Teel motifs were present, while mutating either motif led to a strong reduction in binding (Fig. 3F). Similarly, the high Stel2 occupancy at its unique KAR4 and GPA1 regions entailed two adjacent Stel2 motifs (Fig. 3D-3F). Therefore, the co-binding of the interacting Stel2-Tecl TFs is mirrored in their motif cooperation, which is readily detected by MPBA.Example 5: Msn2 and Sok2 occupancy at regulatory regions depends on multiple independent motifs with no apparent motif cooperativity

[0163] Motif cooperation is most likely to occur in regulatory enhancers that contain multiple TF binding motifs. Examining the lab compendium of 141 (>95%) TF binding profiles highlighted two groups of TFs that bind overlapping promoters, broadly corresponding to activators and repressors of stress genes (see Lupo et al., “The architecture of binding cooperativity between densely bound transcription factors”, Cell Syst., 2023; 14:732-745, the contents of which are hereby incorporated by reference). As representatives of those groups, Msn2 and Sok2 were selected, they appeared as promising candidates for participating in cooperative binding, given that both are guided to their genomic locations through (largely disordered) regions located outside their DNA-binding domains (DBDs) (Fig. 4A and 4J).

[0164] 18 regions were selected, each bound by Msn2 or Sok2 and multiple additional TFs, respective libraries mutating all motif combinations were designed, and MPBA was applied to measure library binding by Msn2 or Sok2 (Fig. 4B and 4K). In all libraries, the intact sequence was depleted, indicating ~40-80% and up to ~35% occupancies of Msn2 and Sok2 at their respective intact regions (Fig. 4C and 4L). Those occupancies were mostly explained by the known Msn2 or Sok2 motifs, as exemplified by EGO4 (Msn2) and OCA5 (Sok2) regions (Fig. 4D and 4M). In some cases, motifs that are better aligned with other TFs contributed to the occupancy, as was particularly notable in three of the tested Msn2 -bound regions (GSY2, BDH2 and HXT7-. Fig. 4C-4D).

[0165] Msn2 binding to the genome was reported to reposition or evict surrounding nucleosomes. To test this in the assay, MPBA was applied to measure the binding of H3 across the Msn2 libraries and further quantified nucleosome occupancy across the genome in the presence or absence of Msn2 and its close paralog Msn4 by applying ATAC-seq and analyzing existing MNase-seq data in cells growing under the tested conditions (see Lupo et al.) (OD600= 4, Fig. 4E-4F). In all 10 tested libraries, the H3 occupancy increased continuously with the decreasing Msn2 binding (Fig. 4F-4G and 4N-4O, see the ‘Materials and methods’ section). Notably, this Msn2 -triggered nucleosome eviction, measured on the plasmid, was comparable in strength to the eviction seen in the genome (Fig. 4H). Note that the measured evictions did differ in certain places between the plasmid-based measurements and the respective genomic analysis, which could indicate differences between chromatin environment on the plasmid and genome or could be due to methodological biases. It was concluded that MPBA simulates not only the binding of Msn2 to the tested regulatory regions but also its role in the organization of nucleosomes.

[0166] Next, searched for cooperating motif pairs was performed, considering those that contained at least one TF motif (97 for Msn2 and 67 for Sok2). Contrasting expectations, Msn2 and Sok2 occupancies were well explained by independent binding at individual motifs (Fig. 41 and 4P). The only exceptions were the HXK1 region, where motif cooperation increased somewhat an already strong Msn2 binding (Fig. 41 and 4Q), and a control sequence in which two variable positions were inserted within the same Msn2 motif, giving the expected high cooperating score (Fig. 41). To control for MNase activation, the analysis of Msn2 binding was repeated using a tight time course. However, despite the continuous increase in occupancies, cooperativity scores remained low (Fig. 4R-4S). Therefore, Msn2 and Sok2 binding at the selected regulatory sites appears independent of motif cooperation.Example 6: Msn2 retains independent motif binding across different plasmid and genomic contexts

[0167] Msn2 binding results in significant eviction of surrounding nucleosomes but appears independent of this eviction, as indicated by the lack of significant motif binding cooperation. Still, cooperative binding that propagates through chromatin might be sensitive to the broader chromatin context extending beyond the immediate motif- surrounding sequence. Since the plasmid provides an uncertain chromatin landscape that could differ from that found in the genome, it was asked whether motif cooperativity could be retrieved by changing the specific plasmid context within which were embedded the 164 bp tested regulatory regions (Fig. 5A).

[0168] To test this potential influence of the broader motif context on the binding of Msn2, five contexts of 2 kb lengths were selected, including a context that surrounds one of the tested regions (GSY2; SEQ ID NO: 8), three genomic contexts that display differential nucleosome occupancy (SEQ ID NO: 9-11) and a synthetic context that is expected to remain nucleosome-free (SEQ ID NO: 12) (see the ‘Materials and methods’ section; Fig. 5B). These contexts were used to construct five new MPBA plasmids, to which eight libraries of Msn2- bound regions were integrated and tested for Msn2 binding. Notably, while these contexts affected the overall motif binding occupancies, none have led to significant motif cooperation (Fig. 5C-5D).

[0169] To further test the effect of the surrounding context, a selected set of six libraries were introduced into the genome and MPBA was applied to measure Msn2 binding (Fig. 5E; see the ‘Materials and methods’ section). While Msn2 still showed a strong preference for its motif, the overall binding signals to these genomically embedded libraries were somewhat lower as compared to the plasmid (Fig. 5F-5G). Yet, also here, none have shown motif cooperativity of significant binding effect (Fig. 5H). It was concluded that Msn2 retains independent motif binding across a range of different sequence contexts.Example 7: Msn2 mutant containing only its DBD is dependent on both canonical and on similar, non-canonical motifs contributing to the binding of the full Msn2

[0170] Msn2 binding was independent of motif cooperation but was still dependent, in some cases, on motifs that better align with other TFs (Fig. 6A). Therefore, Msn2 might be recruited to these regions by TFs that bind these motifs. Alternatively, Msn2 may bind those motifs directly, as suggested by their limited divergence from its canonical motif (Fig. 6B). The Msn2 DBD is a well characterized member of the C2H2 family that binds DNAindependently and lacks any known dimerization domain or other interaction sites, suggesting that any potential Msn2 recruitment is likely mediated by regions outside the DBD. To distinguish between those possibilities, it was asked whether these motifs are also bound by an Msn2 mutant containing only its DBD, which is unlikely to interact with other TFs (Fig. 6C).

[0171] The DBD-only mutant displayed weak library binding when expressed under the endogenous Msn2 promoter, limiting the ability to measure motif occupancies reliably (Fig. 6D). Over-expression of the DBD overcame this apparent limited affinity, reaching high occupancies across the libraries that were comparable to the endogenously expressed Msn2 (Fig. 6D). Notably, the DBD motif-binding patterns within each library correlated well with those of Msn2, and it showed similar localization not only at Msn2 motifs but also to the non-canonical, Msn2 -recruiting ones (Fig. 6E-6H). Therefore, all motifs on which Msn2 depends for binding are directly recognized by its DBD.Example 8: Large-scale testing of motif interactions

[0172] The analysis so far questioned a general role for motif-dependent cooperation in guiding TF localization, as the motifs contributing to the binding of the GRFs, Msn2 and Sok2 appear to have independent effects. Still, cooperative motif interactions were clearly seen in the case of Teel and Stel2. To examine the prevalence of cooperating motifs more broadly, the analysis was extended to dozens of additional TFs (Fig. 7G-7I). For this, a total of 68 regions bound by multiple TFs were selected, corresponding libraries including all combinations of motif mutations were generated and the binding of respective TFs to those libraries was tested. Overall, 250 TF-library pairs were measured.

[0173] The estimated occupancies at intact regulatory regions varied between TFs. High occupancies that consistently reached ~60-90% were seen for Stpl, Mot3 and Cbfl, as well as for the GRFs discussed above (Fig. 7A). Further, in 33 / 39 of the TFs tested, the average occupancy at their own motif exceeded that of other TF motifs (Fig. 7B). In some cases, additional, non-canonical motifs contributed to binding, and those included sequence-similar motifs (e.g. The Met31 and the Cbfl) but also non-similar ones, which may suggest recruitment (e.g. Rgml by Sok2 or Gln3 by Stpl; Fig. 7C). To verify this, the correlation in motif effects of different TFs measured were compared across the same library. For this, Medl5, an Msn2-recruited coactivator was used as a control, its binding was measured across the Msn2 libraries (Fig. 7C-7D). Notably, only ~10% of the TF pairs showed similar motif effects (Pearson’s r > 0.4), and these included the Msn2 and Medl5 control, as well asparalogs whose DBDs indeed bind the same motif. Still, eight TF pairs localized to similar motifs despite having distinct DBDs, and those included the known Stel2-Tecl and Gln3- Stpl (noted above).

[0174] Next, the cooperation score was defined for all 1917 pairs, including at least one motif of the tested TF, available in the dataset (Fig. 7E and 7J). Apart from Stel2-Tecl discussed above, a few more cooperating motif pairs were detected. These include the Swi5- Sok2 motifs, whose cooperation increased Swi5 binding by 32% at the RSM7 promoter, and two Swi4 motifs, whose cooperation increased Swi4 binding by 31% at the CAP1 promoter. Notably, all of these pairs were located in high proximity to one another (Fig. 7E and 7K). Those cooperating motif pairs, however, were the exception. Out of 606 pairs showing > 10% contribution to TF occupancy, 589 (97%) were independent of motif cooperation. Therefore, at the global scale, TF localization to their genomic binding sites is largely invariant to the presence of nearby motifs (Fig. 7F).Example 9: Testing CRISPR-Cas9 off-target binding

[0175] One of the major bottlenecks for developing CRISPR -based therapeutic treatments is the risk of off-target CRISPR binding and DNA cleavage. Off-target binding could have catastrophic consequences for a subject receiving a CRISPR therapeutic. However, the method of the current invention allows for the screening of CRISPR complexes for their gRNA specificity and off-target effects. For a proof-of-principle, the golden standard EMX1 target sequence: GAGTCCGAGCAGAAGAAGAA (SEQ ID NO: l)-GGG(PAM) was used. dCas9 fused to an MNase was integrated into the yeast genome and a plasmid expressing a gRNA against SEQ ID NO: 1 was expressed in the yeast cells. A variant library of the target sequence was generated. A schematic of the library which contains 4 sublibraries is presented in Figure 8A. Each library contains 5 variable positions out of the 20 possible positions of the target sequence. Each variable position can contain the original target nucleotide or a mutated one (a total of 32 variants). Mutant bases were selected so as to minimize synthesis bias. In the first sublibrary (marked with a circle in Fig. 8A) positions -20, -19, -3, -2 and -1 from the PAM (not including the GGG spacer) were probed. That is, each of these positions either contained the correct base or a mutant base. The 15 middle bases were left unaltered. This led to 32 different options for the sublibrary. In the second sublibrary (marked with a square) positions -18, -17, -6, -5 and -4 relative to the PAM were probed. The other 15 bases were left unaltered which again produced 32 different variants. The third sublibrary (marked with a triangle) probed positions -16, -15, -9, -8 and -7 and thefourth sublibrary (marked with a pentagon) probed positions -14, -13, -12, -11 and -10. In total 128 molecules were analyzed.

[0176] The ability of the dCas9+gRNA to recognize each variant was tested, as binding would result in cleavage by the MNase and depletion of that sequence from the sequencing data. The results obtained corresponded well with what is known in the literature for the EMX1 target sequence. The dCas9 was found to be sensitive to mutations close to the 3’ end and less sensitive to one found in the 5’ end, but generally an increase in mutations lead to corresponding decrease in binding (Fig. 8B).

[0177] In the next step, a motif of the Rebl TF was added to each side of the target / variant sequence of the library (Rebl motif CGGGTAA on the 5’ side and its reverse complement TTACCCG on the 3’ side) (Fig. 9A). As Rebl is endogenous to yeast and strongly binds all genomic appearances of its motif it was hypothesized that the addition of this motif would increase the stringency of dCas9 binding. Indeed, that is what was observed as the system became much more stringent, with off-target binding largely abolished while preserving binding of dCas9 to its target sequence (Fig. 9B).Example 10: Comparison to existing technologies

[0178] In order to test the efficacy of the method of the invention to detect CRISPR off- target binding, it was compared to two preexisting technologies: GUIDE-seq (Tsai, et al., 2014 “GUIDE-seq enables genome-wide profiling of off-target cleavage by CRISPR-Cas nucleases”, Nature Biotechnology, 33, 187-197) and CIRCLE-Seq (Tsai, et al., 2017, “CIRCLE-seq: a highly sensitive in vitro screen for genome-wide CRISPR-Cas9 nuclease off-targets”, Nature Methods, 14, 607-614). GUIDE-seq is a cellular based method that detects binding by integrating a unique tag at DNA break sites; whereas CIRCLE-seq is an in vitro method performed without cells in which genomic DNA is circularized and exposed to CAS and linearized products are sequenced.

[0179] The VEGFA site 1 target sequence (GGGTGGGGGGAGTTTGCTCC (SEQ ID NO: 17)- TGG(PAM)) was again used as proof of principle. Approximately 500 potential off target sequences derived from these two methods were extracted and a library was generated containing these sequences and the control sequence. The library was integrated into plasmids containing the VEGFA site 1 gRNA, transformed into cells expressing dCas9- Mnase and the binding assay was performed. As expected, maximal binding was obtained for the on-target sequence and PAM mutation or target scrambling abolished Cas9 binding (Fig. 10A). dCas9 binding was found to be sensitive to off-targets with mutations closer tothe 3’ end and less sensitive to mutations in the 5’ end, which aligns with the previous results and what is known in the literature.

[0180] The binding signal was determined for each off-target sequence from GUIDE-seq and CIRCLE-seq (Fig. 10B) off-targets based on the instant method (Fig. 10C). Of the 24 sequences identified in cells by GUIDE-seq, 13 (54%) met the selected threshold ( — 0.2), while only 3% of the CIRCLE-Seq 483 sequences met the same threshold, suggesting that most of the in vitro identified sequences are artifacts and not relevant to Cas9 binding in cells. A direct comparison was made between GUIDE-seq scores and the binding occupancy measured by MPBA (Fig. 10D) and a high correlation was found (r=0.7). A threshold for method sensitivity was generated and can be used going forward. With this threshold 11 of the GUIDE-Seq sequences were found to also be off-target binding sites based on MPBA (Fig. 10E). These results indicate that MPBA is more sensitive to in-vivo defined off-targets and correlates well with GUIDE-seq scores.Example 11: Protein libraries

[0181] MPBA measures binding of thousands of designed nucleotide sequences to a given protein of interest. This concept was extended to compare the binding of thousands of designed peptides to a DNA sequence of interest.

[0182] In peptide MPBA (pMPBA), the tested peptide library and a DNA target sequence are integrated into plasmids and transformed into cells. The system is engineered such that the binding of the plasmid-carrying peptide to the tested DNA results in cleavage of the plasmid. In its current implementation, plasmid cleavage is measured within the region that includes the coding sequence of the tested peptides and can therefore be quantified by sequencing. To achieve this, pMPBA is implemented as follows: First, a plasmid is engineered to contain the target DNA adjacent to a library integration site (Fig. 11A). Second, a library of peptides is designed and cloned into the plasmid, ensuring that each plasmid carries a single variant. Critically, within the plasmid, the peptide is fused to an MNase and a nuclear localization signal (NLS) and is expressed under a strong promoter (e.g., TEF1). As an additional feature, each peptide was linked to a URA3 selection marker using the ribosome-skipping inducer (E2A) sequence, ensuring that only in-frame peptides are retained. Third, the plasmids are transformed into cells, and DNA cleavage is triggered by MNase activation through a short calcium pulse. Finally, PCR amplification of a region covering the cleaved site and the peptide sequence is performed and quantified by high- throughput sequencing. The changes in the abundance of each peptide sequence, relative tocells in which MNase was not activated, providing an estimate of the relative DNA occupancy compared to all tested peptides.

[0183] This approach was validated using a small peptide library of DNA-binding domain (DBD) mutants. The library included the native DBD of the Msn2 TF (SEQ ID NO: 13), and four additional variants: one inserting a premature stop codon, one replacing an essential structural residue to glutamine (H660Q) (SEQ ID NO: 14), one replacing a DNA-contacting residue to that of a related TF, Nrg2 (E655G, SEQ ID NO: 15; Fig. 11B) and one double mutant (H660Q and E655G, SEQ ID NO: 16). Applying pMPBA to measure library binding to three DNA sequences differing in the promoter region confirmed the expected results: variants including the premature stop-codon were eliminated following yeast transformation, validating the efficiency of the E2A selection for in-frame peptides (Fig. 11C). Further, plasmid carrying the WT DBD were depleted from the pool following MNase activation, with levels decreasing by ~6-fold relative to the DBD structural mutant at 360 second activation (Fig. 11D). Lastly, the contact residue mutant showed a partial decrease in signal, consistent with binding at reduced affinity (Fig. 11D). Of note, these measured occupancies were highly similar between the three tested promoters.

[0184] Although the invention has been described in conjunction with specific embodiments thereof, it is evident that many alternatives, modifications and variations will be apparent to those skilled in the art. Accordingly, it is intended to embrace all such alternatives, modifications and variations that fall within the spirit and broad scope of the appended claims.

Claims

CLAIMS:

1. A method of determining CRISPR associated protein (CAS) in complex with a guide RNA (gRNA) association with a nucleotide sequence, the method comprising: a. receiving a nucleic acid molecule library comprising a plurality of nucleic acid molecules wherein each molecule comprises a 5’ primer recognition site and a 3’ primer recognition site which are common to all molecules of said plurality, a guide RNA (gRNA) which is common to all molecules and wherein said plurality of molecules comprises: i. a target sequence molecule comprising a putative target sequence of said CAS-gRNA complex; ii. a plurality of variant sequences molecules wherein each variant sequence molecule comprises a variant target sequence comprising at least one single nucleotide change in said putative target sequence, such that each variant sequence differs from another variant sequence by a single nucleotide change and a plurality of mutations within the putative target sequence are made; b. contacting said library with a cell expressing a fusion protein, wherein said fusion protein comprises a dead CAS (dCAS) in which its nuclease activity is abolished fused to a nuclease; c. incubating said library in said cell for a period of time sufficient for cleavage of nucleic acid molecules bound by said fusion protein; d. amplifying nucleic acid molecules in said cell with said 5’ primer and said 3’ primer to produce amplification products comprising both said 5’ primer recognition site and said 3’ primer recognition site; and e. sequencing said amplification products to produce sequencing data, wherein variant target sequences depleted in said sequencing data as compared to a control sequence are sequences bound by said CAS-gRNA complex; thereby determining CAS-gRNA complex association with a nucleotide sequence.

2. The method of claim 1, wherein said nucleic acid molecules are DNA molecules.

3. The method of claim 1 or 2, wherein said target sequence or variant of said target sequence is between said 5’ primer recognition site and a 3’ primer recognition site and wherein said amplification products further comprise said target sequence or variant of said target sequence.

4. The method of any one of claims 1 to 3, wherein said gRNA is reverse complementary to said target sequence.

5. The method of any one of claims 1 to 4, wherein said method is a method of detecting off-target binding of a CAS-gRNA protein complex to variants of its target sequence, wherein variant target sequences depleted in said sequencing data as compared to control sequences are off-target sequences bound by said CAS-gRNA protein complex.

6. The method of any one of claims 1 to 5, wherein the CAS protein is CAS 9.

7. The method of any one of claims 1 to 6, wherein said plurality of mutations is all possible mutations within the putative target sequence.

8. The method of any one of claims 1 to 7, wherein said nuclease is micrococcal nuclease (MNase).

9. The method of any one of claims 1 to 8, wherein said sufficient time is at most 600 seconds.

10. The method of claim 9, wherein said sufficient time is at most 60 seconds.

11. The method of any one of claims 1 to 10, wherein said sequencing is next generation sequencing, high-throughput sequencing or massively parallel sequencing.

12. The method of any one of claims 1 to 11, wherein the relative depletion of a given sequence is proportional to the occupancy of said CAS at said given sequence.

13. The method of any one of claims 1 to 12, further comprising contacting said cell comprising said library with a substrate that induces the nuclease activity of said fusion protein.

14. The method of claim 13, wherein said substrate is selected from calcium, magnesium and sodium.

15. The method of claim 14, wherein said substrate is calcium.

16. The method of any one of claims 1 to 15, wherein the control sequence is a sequence not bound by the protein.

17. The method of any one of claims 1 to 16, wherein the control sequence is a plurality of the other variant sequences.

18. The method of any one of claims 1 to 17, wherein said plurality of molecules each contain a transcription factor binding site adjacent to a 5’ end and a 3 ’end of said putative target sequence or variant target sequence.

19. The method of claim 18, wherein said transcription factor binding site is a Rebl binding site comprises or consists of CGGGTAA or its reverse complement.

20. A method of identifying an inhibitor of CAS-gRNA complex off-target binding, the method comprising performing a method of any one of claims 1 to 19 wherein said contacting and incubating is performed in the presence of a test molecule and in the absence of said test molecule, and comparing said sequencing data produced in the presence of said test molecule to sequencing data produced in the absence of said test molecule and wherein an increase in sequences of amplification products in the sequencing data produced in the presence of said test molecule as compared to sequences of amplification products in the sequencing data produced in the absence of said test molecule indicates said test molecule is an inhibitor of CAS-gRNA complex off-target binding, thereby identifying an inhibitor.

21. A method of determining protein association with a nucleotide sequence, the method comprising: a. receiving a nucleic acid molecule library comprising a plurality of nucleic acid molecules wherein each molecule comprises a 5’ primer recognition site and a 3’ primer recognition site which are common to all molecules of said plurality and wherein said plurality of molecules comprises: i. a target sequence molecule comprising a putative target sequence of said protein; ii. a plurality of variant sequences molecules wherein each variant sequence molecule comprises a variant target sequence comprising at least one single nucleotide change in said putative target sequence, such that each variant sequence differs from another variant sequence by a single nucleotide change and all possible mutations within the putative target sequence are made; b. contacting said library with a cell expressing a fusion protein, wherein said fusion protein comprises said protein fused to a nuclease; c. incubating said library in said cell for a period of time sufficient for cleavage of nucleic acid molecules bound by said fusion protein; d. amplifying nucleic acid molecules in said cell with said 5’ primer and said 3’ primer to produce amplification products comprising both said 5’ primer recognition site and said 3’ primer recognition site; and e. sequencing said amplification products to produce sequencing data, wherein variant target sequences depleted in said sequencing data as compared to a control sequence are sequences bound by said protein;thereby determining protein association with a nucleotide sequence.

22. A method of determining nucleotide sequence association with a peptide, the method comprising: a. receiving a nucleic acid molecule library comprising a plurality of nucleic acid molecules wherein each molecule comprises i. a 5’ primer recognition site and a 3’ primer recognition site which are common to all molecules of said plurality; ii. a coding sequence encoding a fusion protein, wherein said fusion protein comprises a peptide fused to a nuclease and wherein each peptide comprises at least one amino acid difference from every other peptide ; and iii. a target sequence which is common to all molecules of said plurality; wherein at least one peptide of said library binds to said target sequence; b. contacting said library with a cell; c. incubating said library in said cell for a period of time sufficient for producing of said fusion protein and cleavage of nucleic acid molecules bound by said fusion protein; d. amplifying nucleic acid molecules in said cell with said 5’ primer and said 3’ primer to produce amplification products comprising both said 5’ primer recognition site, and said 3’ primer recognition site; and e. sequencing said amplification products to produce sequencing data, wherein peptides encoded by coding sequences depleted in said sequencing data as compared to a control sequence are peptides that bind said target sequence; thereby determining nucleotide sequence association with a peptide.

23. The method of claim 21 or 22, wherein said protein or peptide is a DNA binding protein and said nucleic acid molecules are DNA molecules.

24. The method of any one of claims 21 to 23, wherein said protein or peptide is a transcription factor.

25. The method of any one of claims 21 to 24, wherein said target sequence and said coding sequence are between said 5’ primer recognition site and a 3’ primer recognition site and wherein said amplification products further comprise said target sequence and said coding sequence.

26. The method of any one of claims 21 and 23 to 25, wherein said plurality of mutations is all possible mutations within the putative target sequence.

27. The method of any one of claims 21 to 26, wherein said nuclease is micrococcal nuclease (MNase).

28. The method of any one of claims 21 to 27, wherein said sufficient time is at most 600 seconds.

29. The method of claim 28, wherein said sufficient time is at most 300 seconds.

30. The method of any one of claims 21 to 29, wherein said sequencing is next generation sequencing, high-throughput sequencing or massively parallel sequencing.

31. The method of any one of claims 21 to 30, wherein the relative depletion of a given sequence is proportional to the occupancy of said protein or peptide at said given sequence.

32. The method of any one of claims 21 to 31, further comprising contacting said cell comprising said library with a substrate that induces the nuclease activity of said fusion protein.

33. The method of claim 32, wherein said substrate is selected from calcium, magnesium and sodium.

34. The method of claim 33, wherein said substrate is calcium.

35. The method of any one of claims 21 to 34, wherein the control sequence is a sequence not bound by the protein or peptide.

36. The method of any one of claims 22 to 35, wherein said coding sequence further encodes a selection protein and wherein said selection protein is separated from said fusion protein by a ribosomal skipping sequence, optionally wherein said ribosomal skipping sequence is selected from P2A, E2A, T2A and F2A.

37. The method of claim 36, wherein said cell cannot grow in a first condition and said selection protein is a protein that allows growth in said first condition.

38. The method of claim 37, wherein said first condition is absence of uracil and said selection protein is orotidine 5-phosphate decarboxylase (URA3).

Citation Information

Patent Citations

  • Test for Huntington's disease

    US4666828A

  • Process for amplifying nucleic acid sequences

    US4683202A

  • Apo AI / CIII genomic polymorphisms predictive of atherosclerosis

    US4801531A

  • Intron sequence analysis method for detection of adjacent and remote locus alleles as haplotypes

    US5192659A

  • Method of detecting a predisposition to cancer by the use of restriction fragment length polymorphism of the gene for human poly (ADP-ribose) polymerase

    US5272057A