Synthetic reporters for multiplexed detection of transcription factor activity

Optimized nucleic acid compositions with consensus TF binding sites and spacers enhance TF reporter sensitivity and specificity, addressing precision and scalability issues in TF activity detection, facilitating high-throughput analysis of physiological pathways.

WO2026024176A1PCT designated stage Publication Date: 2026-01-29THE TRUSTEES OF COLUMBIA UNIV IN THE CITY OF NEW YORK +1
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
PCT/NL2025/050355
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-07-22
Filing Date
2025-07-21
Publication Date
2026-01-29

AI Technical Summary

Technical Problem

Current methods for detecting transcription factor (TF) activities in parallel lack precision and scalability, as existing reporters are limited, rely on non-specific genomic elements, or have suboptimal synthetic sequences, and TF paralogs complicate design specificity.

Method used

Development of a systematically optimized nucleic acid composition with consensus TF binding sites, spacer sequences, and core promoters to create a library of highly sensitive and specific synthetic reporters for 86 TFs, enabling multiplexed detection using massively parallel reporter assays (MPRAs).

Benefits of technology

The optimized reporters provide accurate, high-throughput detection of TF activities, outperforming existing methods by >80% in sensitivity and specificity, allowing for the analysis of complex physiological pathways and signaling interdependencies.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure NL2025050355_29012026_PF_FP_ABST
    Figure NL2025050355_29012026_PF_FP_ABST
Patent Text Reader

Abstract

The invention covers an isolated nucleic acid composition that encodes a synthetic reporter capable of binding to a transcription factor, said composition comprising a) between four and eight copies of a consensus transcription factor binding site (TFBSs), b) between three and seven TFBSs-devoid spacer sequences between the TFBSs mentioned under a), c) one TFBS-devoid spacer sequence upstream of a core promoter, d) a core promoter and e) an unique barcode sequence and / or an open reading frame encoding a reporter protein. The invention also covers a library of isolated nucleic acid molecules that encode one or more synthetic reporters containing the designed spacer sequences and binding to one or more transcription factors, as well a kit containing all these elements and barcodes. The invention also covers a computer-implemented method for determining transcription factor (TF) reporter activities, as well as comparing the activities across different conditions. The invention also covers a computing apparatus and non-transitory computer-readable storage medium for performing the method.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Synthetic reporters for multiplexed detection of transcription factor activity Field of the invention The invention is related to the field of optimizing transcriptional reporters, in particular the use of novel massively parallel reporter assay (MPRA) to systematically optimize transcriptional reporters for all existing TFs, and in particular for a large subset of TFs involved in most physiologically-relevant signalling pathways in mammals and humans. Background Transcription factors present a growing area of possible therapeutic targets for novel drugs and treatments for a myriad of medical conditions. In any given cell type, dozens of transcription factors (TFs) act in concert to control the activity of the genome by binding to specific DNA sequences in regulatory elements. Intra- and extracellular signals intricately control the activity of dozens of interwoven signaling pathways, often TFs. TFs respond to these upstream signaling cascades and translate them to orchestrate the regulation of the genome. If we knew the activity of all TFs in any given cell type, we might be able to understand how TFs interpret incoming signals, how they drive the downstream changes in gene expression, and how cascades of TF activities progress over time. However, we currently have no reliable method to directly detect many TF activities in parallel. A variety of computational approaches have been developed to infer TF activities from genome-wide data such as TF binding data (chromatin immunoprecipitation (ChIP)-sequencing {Johnson, 2007, 17540862}), chromatin accessibility maps (assay for transposase-accessible chromatin (ATAC)-sequencing {Buenrostro, 2013, 24097267}) {Schep, 2017, 28825706; Baek, 2017, 28538187}, TF or target gene transcript abundance data (RNA-sequencing {Mortazavi, 2008, 18516045}) {Alvarez, 2016, 27322546; Müller-Dott, 2023, 37843125; Li, 2023, 37719147}, or a combination of these methods {Berest, 2019, 31801079; Garcia-Alonso, 2019, 31340985; Keenan, 2019, 31114921}. While these methods provide convenient tools to compute TF activities from well-established genomics assays, they do not directly measure the transcriptional activity of TFs and might therefore lack precision. For example, it is known that maps of TF binding poorly reflect TF activity {Lenstra, 2012, 22572961; Kang, 2022, 35666184}; ATAC-seq detects open chromatin regions which are not necessarily predictive of transcription activity {Gasperini, 2020, 31988385}; and inferring TF activity from mRNA-seq data requires assumptions regarding the distance over which each TF may be able to control gene activity. Traditional reporter assays, employing fluorescent or luminescent proteins expressed by TF response sequences, offer direct means to measure TF activities {Korinek, 1997, 9065401; Ting, 1996, 8947041; Turkson, 1998, 9566874; Fujino, 2004, 13130122; Chen, 2000, 10948206; Emmel, 1989, 2595372; Dennler, 1998, 9606191; Kastan, 1992, 1423616; Sladek, 1990, 2279702; Tamura, 1995, 8749394; Davidson, 1988, 2843293}. These assays have been used for decades and detect TF activity with great sensitivity. However, conventional reporter assays do not allow the detection of multiple TFs at once. A previous study circumvented this limitation and measured 58 TF activities in parallel from previously published TF reporters by utilizing RNA barcodes as reporters {O’Connell, 2016, 27211859}. This study also showed that TF reporter measurements can be more accurate than RNA-seq-inferred TF activities for a subset of TFs. Thus, directly measuring TF activities in a high-throughput fashion using barcoded reporters offers a direct and precise alternative to computational inference approaches. Despite the advantages of multiplexed TF reporter assays, there are still several challenges in achieving accurate high-throughput TF activity detection. First, reporters are only available for a limited number of TFs. Expanding the collection of TF reporters will be crucial to make multiplexed reporter assays more scalable. Second, most of the published TF reporters rely on either (i) TF response elements found in the genome {Turkson, 1998, 9566874; Fujino, 2004, 13130122; Kastan, 1992, 1423616; Sladek, 1990, 2279702}, which might lack specificity to the intended TF due to the presence of other TF binding sites (TFBSs), or (ii) poorly optimized synthetic TF response sequences {Korinek, 1997, 9065401; Ting, 1996, 8947041; Romanov, 2008, 18297081}, which could be suboptimal in terms of sensitivity and specificity. Hence, it is necessary to optimize TF reporters to obtain more reliable activity measurements. Finally, it is known that TFs within the same TF family, especially TF paralogs, can have highly similar DNA-binding domains and thus also TFBSs, which complicates the design of reporters that are specific for a single TF. Despite their considerable importance in determining cell identity and their pivotal role in numerous disorders, we currently lack simple tools to directly measure the activity of many TFs in parallel. Massively parallel reporter assays (MPRAs) allow the detection of TF activities in a multiplexed fashion; however, we lack basic understanding to rationally design sensitive reporters for many TFs. Summary of the invention In the invention hereunder, we used a novel massively parallel reporter assay (MPRA) to systematically optimize transcriptional reporters for all existing TFs, and in particular for a large subset of TFs involved in most physiologically-relevant signalling pathways in mammals and humans. We generated a set of reporters for 86 TFs and showed the specificity of all reporters across a wide array of TF perturbation conditions. We thus discovered critical TF reporter design features and obtained highly sensitive and specific reporters for all TFs, but in particular for 62 TFs, which based on our design logic outperform available reporters. The resulting collection of “prime” TF reporters can be used to uncover TF interaction networks and to illuminate signaling pathway interdependencies. The invention is summarized in the following embodiments: - Embodiment 1. An isolated nucleic acid composition that encodes a synthetic reporter capable of binding to a transcription factor, said composition comprising: a) between four and eight copies of a consensus transcription factor binding site (TFBSs), b) between three and seven TFBSs-devoid spacer sequences between the TFBSs mentioned under a), c) one TFBS-devoid spacer sequence upstream of a core promoter, d) a core promoter and e) an unique barcode sequence and / or an open reading frame encoding a reporter protein. - Embodiment 2. The isolated nucleic acid composition of embodiment 1 wherein the features b, c, and d, have been determined according to the formula: ^^^^^^^^^^^^^ ^^^^^^^^^^,^^^^^^^^^^ ~core promoter + promoter distance + spacer length: spacer sequence- Embodiment 3. The isolated nucleic acid composition according to embodiment 1 and 2 wherein the consensus TFBS are selected from the list of SEQ ID NO: 1 to SEQ ID NO: 86: - Embodiment 4. The isolated nucleic acid composition according to embodiments 1 to 3 wherein the length of the three to seven spacer sequences between the TFBSs is either 5 or 10 bp. - Embodiment 5. The isolated nucleic acid composition according to embodiments 1 to 4 wherein the spacer sequences between the TFBSs are selected from the list according to SEQ ID NO 87 to SEQ ID NO:146. - Embodiment 6. The isolated nucleic acid sequence composition according to any of the embodiments 1 to 5 wherein the length of the spacer sequence upstream of a core promoter is either 10 or 21 bp. - Embodiment 7. The isolated nucleic acid composition according to any of the embodiments 1 to 6 wherein each spacer length has one out of three designed sets of the TFBS-devoid spacer sequences selected from table 3. - Embodiment 8. The isolated nucleic acid composition from any of the preceding embodiments wherein the core promoter represents any promoter capable of initiating transcription of said nucleic acid composition, preferably the core promoter being selected from the group of: minP derived from pGL4, minCMV, minHBG, SCP1 (PMID 17124735), AdML, Hsp68 (2557196). - Embodiment 9. The isolated nucleic acid composition from any of the preceding embodiments wherein the reporter protein is a fluorescent or luminescent protein. - Embodiment 10. The isolated nucleic acid composition of any of the embodiments 1-9 wherein the synthetic reporter comprises a nucleic acid sequence with at least 95% identity according to SEQ ID NO: 147 to SEQ ID NO: 232, preferably at least 98% identity, more preferably at least 99% identity, most preferably 100% identical to said SEQ ID Nos - Embodiment 11. A library of isolated nucleic acid molecules that encode one or more synthetic reporters containing the designed spacer sequences from table 3 binding to one or more transcription factors according to embodiments 1-10. - Embodiment 12. The library of isolated nucleic acid molecules from embodiment 11, wherein the transcription factors represent all mammalian transcription factors, preferably all human transcription factors. - Embodiment 13. The library of isolated nucleic acid molecules from embodiment 12, wherein the transcription factors are selected from table 1. - Embodiment 14. A high-throughput method for simultaneously detecting the activity of two or more transcription factors in one or more mammalian cells, the method comprising the steps of: a) providing the library of synthetic transcription factor reporters according to any of the embodiments 11 to 13, b) transfecting mammalian cells with one or more vectors encoding the synthetic reporters, and c) detecting and measuring TF activity via a reporter assay selected from the group of: luminescence assay, fluorescence assay and / or measuring TF activity by estimating the abundance of the barcoded transcript by sequencing-based methods. - Embodiment 15. The method according to embodiment 14 wherein the mammalian cells are cultured in an organoid. - Embodiment 16. The method according to embodiment 15 wherein the mammalian cell is present in a transgenic animal model. - Embodiment 17. The method according to any of the embodiments 14 to 16 wherein the vectors are taken from the list of synthetic, bacterial, viral, or transposable element-based. - Embodiment 18. The method according to any of the embodiments 14 to 17, wherein the mammalian cells, organoids or transgenic animal models have been exposed to specific therapeutic molecules, toxins, chemicals or physical treatments prior to or during use of said method. - Embodiment 19. The method according to embodiments 14 to 18, wherein the synthetic reporter is capable of binding to transcription factors that are overexpressed or downregulated. - Embodiment 20. The method according to embodiment 14 to 19, wherein the synthetic reporter is capable of binding to mutant transcription factors. - Embodiment 21. The method according to any of the embodiments 14 to 20, combined with single- cell RNA sequencing methods for determining the activity of one or more transcription factors in single cells together with the expression status of any other gene in the cell. - Embodiment 22. The method according to any of the embodiments 14 to 21, combined with a (sc)ATAC-sequencing method. - Embodiment 23. A vector comprising the nucleic acid composition of any of the embodiments 1 to 10, wherein the vector is synthetic, bacterial, viral, or transposable element – based. - Embodiment 24. A mammalian cell, organoid or mammal containing one or more vectors from embodiment 23. - Embodiment 25. A mammal containing in its genome multiple barcoded synthetic TF reporters according to any of the embodiments 1 to 10. - Embodiment 26. The mammal from embodiment 25 wherein said mammal is a mouse - Embodiment 27. A kit of parts comprising: a) the isolated nucleic acid composition of embodiment 1 or the library of any of the embodiments 11 to 13, b) one or multiple vectors according to embodiment 23, c) a gene-specific primer used for reverse transcription, and d) a primer pair binding to the sequences up- and downstream of the barcode for amplification and quantification of barcodes. - Embodiment 28. A computer-implemented method for determining transcription factor (TF) reporter activities, the method comprising: a) receiving sequencing data for a plurality of TF reporters for at least one control sample; and b) receiving sequencing data for a plurality of TF reporters for at least one contrast sample; and c) Inputting the sequencing data in a) and the sequencing data in b) into a statistical model which outputs: i) a measure of TF reporter activities; and / or ii) a measure of changes in TF reporter activities between the control sample in a) and the contrast sample in b). - Embodiment 29. The computer-implemented method of embodiment 28, wherein plasmid DNA data or genomic DNA data is also input to the statistical model. - Embodiment 30. The computer-implemented method of embodiment 28 or embodiment 29, wherein at least one quality control measure and / or at least one exploratory plot is also generated. - Embodiment 31. A computing apparatus for determining transcription factor (TF) reporting activities, the apparatus comprising; one or more memory units; and one or more processors configured to execute instructions stored in the one or more memory units to perform the method of any of embodiments 28-30. - Embodiment 32. A non-transitory computer-readable storage medium including instructions that, when executed by one or more processors of a computer, configure the computer to perform the method of any one of embodiments 28-30. Methods and compositions are provided herein that are, inter alia, useful in therapeutic interrogation of complex physiologic pathways by massively parallel and permissive transcriptional screening. Thus, methods and compositions are provided herein that are useful for high-throughput functional analysis of complex, transcriptionally regulated physiological pathways. The methods and composition can be generalized and applied to any class of transcription factor or any class of gene product that can regulate the activity of transcription. Thus, for example, the methods and compositions provided herein are generally applicable to all known transcription factors and any gene encoded product that modulates said transcription factor activity. Moreover, data obtained through the methods provided herein are directly comparable thereby facilitating high-throughput functional analysis. In an aspect of the invention, a method is provided for identifying the activity of a transcription factor. The method includes transfecting (each of) a plurality of cells with the nucleic acid composition encoding a reporter protein that is capable of binding to a TF and acting as an indicator of TF activity (i.e. each of the plurality of cells are transfected with the same nucleic acid composition encoding the same reporter protein). The plurality of cells is transfected with a nucleic acid promoter sequence of known function (i.e. having a known functional characteristic) linked to a nucleic acid reporter sequence or a nucleic acid promoter sequence linked to a nucleic acid reporter sequence wherein the nucleic acid promoter sequence forms part of a family of a nucleic acid promoter sequences. Each of the plurality of cells is transfected with a single or multiple nucleic acid promoter sequences linked to a nucleic acid reporter sequence. Transcription of the nucleic acid reporter sequence in at least one of the plurality of reporter cells is detected thereby obtaining a fluorescent or luminescent readout of the nucleic acid promoter sequence interaction with the TF. Additionally, the TF activity can be detected by the quantification of expression levels of a unique barcode attached to each reporter. Each cell can therefore contain multiple reporters with a variety of different barcodes said reporters capable of simultaneously binding different TFs, allowing for the multiplexed detection of TF activity in a given cell. In another aspect, a kit is provided for identifying the activity of a transcription factor. The kit includes an isolated nucleic acid composition encoding a reporter capable of binding to a TF, or a library of multiple reporters, or one or multiple vectors comprising the reporters. In addition, the kit may contain a gene- specific primer used for reverse transcription, a primer pair binding to the sequences up- and downstream of the barcode for amplification and quantification of barcodes. The isolated nucleic acid composition encoding the reporter capable of binding a TF and the libraries of nucleic acid promoter sequences linked to a nucleic acid reporter sequence and libraries of nucleic acid composition encoding a reporter protein capable of binding a TF are described above in the description of methods of the present invention, and are equally applicable to the kits provided herein. Definitions Various terms relating to the methods, compositions, formulations, uses and other aspects of the present invention are used throughout the specification and claims. Such terms are to be given their ordinary meaning in the art to which the invention pertains, unless otherwise indicated. Other specifically defined terms are to be construed in a manner consistent with the definition provided herein. Although any methods and materials similar or equivalent to those described herein can be used in the practice for testing of the present invention, the preferred materials and methods are described herein. Methods of carrying out the conventional techniques used in methods of the invention will be evident to the skilled worker. The practice of conventional techniques in molecular biology, biochemistry, computational chemistry, cell culture, recombinant DNA, bioinformatics, genomics, sequencing and related fields are well-known to those of skill in the art and are discussed, for example, in the following literature references: Sambrook et al., Molecular Cloning. A Laboratory Manual, 2nd Edition, Cold Spring Harbor Laboratory Press, Cold Spring Harbor, N. Y., 1989; Ausubel et al., Current Protocols in Molecular Biology, John Wiley & Sons, New York, 1987 and periodic updates; and the series Methods in Enzymology, Academic Press, San Diego. “A,” “an,” and “the”: these singular form terms include plural referents unless the content clearly dictates otherwise. The indefinite article "a" or "an" thus usually means "at least one". Thus, for example, reference to “a cell” includes a combination of two or more cells, and the like. “About” and “approximately”: these terms, when referring to a measurable value such as an amount, a temporal duration, and the like, is meant to encompass variations of ±20% or ±10%, more preferably ±5%, even more preferably ±1%, and still more preferably ±0.1% from the specified value, as suchvariations are appropriate to perform the disclosed methods. Additionally, amounts, ratios, and othernumerical values are sometimes presented herein in a range format. It is to be understood that such range format is used for convenience and brevity and should be understood flexibly to include numerical values explicitly specified as limits of a range, but also to include all individual numerical values or sub-ranges encompassed within that range as if each numerical value and sub-range is explicitly specified. For example, a ratio in the range of about 1 to about 200 should be understood to include the explicitly recited limits of about 1 and about 200, but also to include individual ratios such as about 2, about 3, and about 4, and sub-ranges such as about 10 to about 50, about 20 to about 100, and so forth. “And / or”: The term “and / or” refers to a situation wherein one or more of the stated cases may occur, alone or in combination with at least one of the stated cases, up to with all of the stated cases. “Comprising”: this term is construed as being inclusive and open ended, and not exclusive. Specifically, the term and variations thereof mean the specified features, steps or components are included. These terms are not to be interpreted to exclude the presence of other features, steps or components. “Exemplary": this terms means "serving as an example, instance, or illustration," and should not be construed as excluding other configurations disclosed herein. A "nucleic acid reporter sequence" is a nucleic acid encoding one reporter gene that produces a detectable reporter protein. A reporter gene is a particularly useful component of some types of transgenes. A reporter gene comprises nucleotide sequences encoding a protein that will be expressed under the direction of a particular promoter of interest to which it is linked in a transgene, providing a measurable biochemical response of the promoter activity. A reporter gene is typically easy to detect or measure against the background of endogenous cellular proteins. Commonly used reporter genes include but are not limited to LacZ, green fluorescent protein, and luciferase, and other reporter genes, many of which are well-known to those skilled in the art. A tag or barcode in the form of a noncoding unique sequence may be encoded in a DNA construct so that the mRNA will have a unique non-coding region attached. This unique region serves as a as an identifier that may be detected by a wide variety of techniques well known in the art, including, but not limited to, RT- PCR, RNA-Seq, immunohistochemistry, or in situ hybridization. Such a barcode may be particularly advantageously designed to accommodate quantification and identification of nucleic acid sequences encoding for the transcription factor reporters (e.g. expressed and / or comprised in the genome) allowing the quantification of barcode expression in the RNA of the modified cells and correlating said expression with TF activity. A barcode sequence can be any nucleotide sequence that, based on its sequence, is uniquely distinguishable and allows identification of the encoded receptor. A barcode sequence may be between 5 and 20 nucleotides in length. It is understood that with a barcode sequence consisting of only 5 nucleotides, this allows for 625 (25 x 25) identifiers (without any redundancy). The skilled person is well capable of designing and incorporating a barcode sequence in the nucleic acid sequence, or, selecting suitable barcodes known in the art. An "isolated nucleic acid composition" is understood as a nucleic acid sequence encoding a reporter protein capable of binding a TF, also referred to herein as a "reporter". The nucleic acid composition is typically a DNA sequence. A "transcription factor," is a DNA binding protein that influences the transcription of a gene product from genomic material. Various transcription factors specifically influence (e.g., promote) transcription of particular gene products. A transcription factor can be a transcription modifying protein and it can be in some instances a "nuclear receptor" or "nuclear hormone receptor," which is a transcription modifying protein that activates or represses transcription of one or more genes in the nucleus (but can also have second messenger signalling actions), typically in conjunction with other transcription factors. In one embodiment the present invention refers to “any mammalian transcription factor”. This is to be understood as any transcription factor found in mammals such as the ones identified and characterized by Lambert et al, 2018, 29425488 or by Vorontsov et al, 2024 https: / / doi.org / 10.1093 / nar / gkad1077 In some embodiments of this invention, the nucleic acid promoter sequences can be naturally occurring sequences, can be modified or recombinant or mutated versions of naturally occurring sequences, or can be allelic variants or disease / medical condition specific variants. A plurality of reporters can be pooled into a library of reporters. A library of reporters can comprise at least 2 reporters. A library of reporters can, for example, comprise from 5 to a million reporters. A library of reporters can comprise at least a million reporters. It should be understood, that a library of reporters can comprise any number of reporters. A library of reporters can comprise reporters that have any combination of common elements and non-common or unique elements as compared to the other reporters within the pool. For example, a library of reporters can comprise common consensus TFBSs or common spacer sequences while also containing non-common or unique barcodes. Common elements can be shared by a plurality, majority, or all of the reporters within a library of reporters. Non-common elements can be shared by a plurality, minority, or sub-population of reporters within the library of reporters. Unique elements can be shared by a one, a few, or a sub-population of reporters within the library of reporters, such that it is able to identify or distinguish the one, few, or sub-population of reporters from the other reporters within the library of reporters. Such combinations of common and non-common are advantageous for multiplexing techniques as disclosed herein. A "nucleic acid promoter sequence" or "promoter" is a nucleic acid that facilitates transcription of a particular gene. Nucleic acid promoter sequence are typically regions of DNA located near the particular gene whose transcription is facilitated. In some embodiments, the nucleic acid promoter sequence is, or includes, a "transcription element," which is a regulatory DNA region that allows transcription of a gene product from a gene. A transcription element comprises specific nucleic acid sequences that are recognized by one or more transcription factors. Thus, in some embodiments, the nucleic acid promoter sequence is, or includes, a transcription factor-binding site or a response element. A “core promoter” as used in the inventions hereunder, is the minimal portion of any known promoter that can property initiate transcription of the consensus transcription factor binding sites of the invention hereunder. A core promoter is a specific region of DNA located immediately upstream of the transcription start site (TSS) of a gene. It serves as the primary binding site for the basic transcriptional machinery, including RNA polymerase II and various transcription factors, to initiate transcription. The core promoter typically spans about 30-100 base pairs and contains several conserved sequence elements that are crucial for the accurate initiation of transcription. Elements of such core promoters may include: a TATA Box, an Initiator (Inr) Element, a Transcription Factor IIB Recognition Element and / or a downstream promoter element. The combination and specific arrangement of these elements can vary, contributing to the regulation of gene expression by influencing the binding of transcription factors and the assembly of the transcriptional machinery. The term "test" in reference to an agent, compound, or method component (e.g. a nucleic acid promoter sequence, a transcription modulating agent, transcription modifying protein, an isolated nucleic acid composition encoding a reporter, etc.) means that the referenced agent, compound, or method component is to be analyzed (e.g. screened, assayed, identified or characterized) in one or more of the methods described herein. The agent, compound, or method component can exist as a single isolated compound or can be a member of library. The term "transfected" or "transfection" refers to the process of introducing nucleic acids into a cell by any appropriate method, including viral or non-viral means. Thus, as used herein, transfection includes transformation and transduction. Detailed description Despite the several challenges in achieving accurate high-throughput TF activity detection, such as i) reporters that are only available for a limited number of TFs and ii) most of the published TF reporters relying on either (a) TF response elements found in the genome, which might lack specificity to the intended TF due to the presence of other TF binding sites (TFBSs), or (b) poorly optimized synthetic TF response sequences, which could be suboptimal in terms of sensitivity and specificity and iii) it is known that TFs within the same TF family, especially TF paralogs, can have highly similar DNA-binding domains and thus also TFBSs, which complicates the design of reporters that are specific for a single TF, we have surprisingly found an method for optimizing TF reporters to obtain more reliable activity measurements against any individual TF. An important embodiment of the invention is that we have obtained consensus TFBSs that are recognized and bound by TFs with high affinity. These sequences are derived from the alignment of multiple binding sites recognized by a particular TF, revealing a common sequence pattern or motif that represents the optimal binding site. The consensus sequence represents the most frequently occurring nucleotides at each position in these binding sites. Also, the invention comprises the compositions of highly optimized reporters for every TF. We generated massively parallel reporter assays (MPRAs) with a systematically designed library of 35,500 different reporter designs for a set of 86 TFs, including TFs that respond to diverse signaling pathways and a variety of cell type-specific TFs. For each TF, we optimized the design of the reporter by varying the spacer sequences and spacer length between TFBSs, the distance to the core promoter and the core promoter itself. We evaluated the specificity and sensitivity of the generated TF reporters by probing the library across nine cell lines and almost 100 TF perturbation conditions. Detailed analysis of this rich dataset provided insights into the rules that determine the sensitivity and specificity of reporters for each TF, and yielded a collection of ‘prime’ reporters for 62 TFs, for many of which no reporters were available yet. Our reporters outperform published reporters in >80% of all comparisons. We demonstrate the utility of the identified prime reporter set by detecting signaling pathway interdependencies upon pluripotency-challenging perturbations in mouse embryonic stem cells (mESCs). An important embodiment of the invention are the spacer sequences between the consensus TFBSs, representing non-coding DNA segments located between TFBSs. The advantages of the spacer sequences between TFBSs of the invention are structural flexibility, optimal TF binding, regulatory complexity, prevention of steric hindrance and facilitation of DNA looping. The spacer sequences of the invention represent either 5 base pair or 10 base pair, preferably they are from the list according to SEQ ID NO 87 to SEQ ID NO:146, but also represents 5 base pair or 10 base pair sequences that are minimally 95%, or 98% identical to any of the SEQ ID NO 87 to SEQ ID NO:146. Provided herein are novel methods, including high-throughput methods, for functional analysis of complex physiologic pathways. The methods include analysis of promoter functionality, transcription modifying protein functionally, and / or transcription modulating agent functionality. Disclosed herein are methods that allow, for the first time, high throughput functional analysis of physiologic pathway components in cellular systems. In some embodiments, the analysis is performed in a reporter cell (i.e. the cellular system is a reporter cell) wherein the reporter cell provides a generic environment thereby allowing the functional studies, such as the interactions between physiologic pathway components. By providing a generic environment, the reporter cell enables the study, inter alia, of physiologic pathways derived from tissues exogenous to the tissue from which the reporter cell was derived. Moreover, it has been found that the methods provided herein allow the study of interactions between components of physiologic pathways derived from different tissues. An important part of the invention is the systematic design and identification of a large collection of optimized “prime” TF reporters. This collection encompasses reporters that significantly outperform currently available reporters (e.g., VDR, SOX2, PAX6), and reporters for TFs for which no reliable reporters were available yet (e.g., GATA4, TFCP2L1, KLF4). The sequences of the prime reporter for each TF are documented in Table S4, which can be used for various purposes. For instance, the prime reporters can be used individually in a conventional fluorescence / luminescence reporter assay to better characterize the role of single TFs in certain biological processes. Alternatively, the identified 62 prime reporters can be employed in a multiplexed fashion, where each TF drives a unique barcode. Signaling pathways could be challenged by an array of inhibitors or activators, similar to what has been done in this study, to unveil novel roles of TFs in signaling pathways. Likewise, TF responses can be tracked upon TF depletion to dissect TF-TF communications. Potentially, this can also be done in single cells and in time-course experiments to detect cascades of TF activities. Although the prime reporters are top-rated based on our performance criteria, there may be instances where other reporters with specific attributes are preferred for certain TFs (e.g., high cell type-specificity or responsiveness to perturbation of related TFs). Figure S6 can aid in identifying such cases (e.g., for identifying generic STAT reporters instead of STAT3-specific reporters). We have shown that our synthetic reporters outperform published reporters for >80% of all comparisons. This underscores that a subset of currently available reporters is suboptimal in terms of sensitivity (e.g., VDR, PAX6, NR1H4) or specificity (e.g., CLOCK, TP53, POU2F1). In comparison to published TF reporters, which rely on genomic response elements or unoptimized synthetic designs, the designed prime reporters exclusively contain TFBSs for the candidate TF, and are highly optimized to enable effective transcription. Through careful optimization of the spacer sequences between the TFBSs and the choice of the core promoter, we were able to achieve reporters with increased TF sensitivity and specificity. In some cases, this even enabled us to identify specific reporters for TFs with highly similar TFBSs (e.g., GATA1 / GATA4, TFCP2L1 / GRHL1). While we established prime reporters for 62 TFs, we still lack good reporters for many other TFs. For instance, our set of TFs did not include a large number important activator TFs that belong the basic domain or homeodomain TF superclass, many of which have non-unique binding motifs. These TFs can have crucial roles in development (e.g., HOX TFs), hence, generating reporters for these TFs would be important to dissect the roles of TFs during differentiation. Although it remains challenging to generate specific reporters for TFs with non-unique TFBSs, careful optimization of TFBS spacer sequences and thorough evaluation of reporter responses to a variety of target TF and off-target TF perturbations could offer solutions. An important embodiment of the invention is that multiplexed TF reporter measurements can complement indirect TF activity inference methods that rely on ATAC-seq, ChIP-seq, or RNA-seq data. Multiplexed (prime) TF reporter assays offer an orthogonal approach that provides functional evidence of TF activity with high specificity for the candidate TF. Use of luciferase constructs and their related assays are well known to those of skill in the art. See, e.g., Greer, et al., "Imaging of light emission from the expression of luciferases in living cells and organisms: a review," Luminescence, 2002, Jan.-Feb, 17(1):43- 74, Hutchens, et al, "Applications of bioluminescence imaging to the study of infectious diseases," Cellular Microbiology, 9:2315-2322, etc. In addition to luciferase, various embodiments of the current invention can also utilize other bioluminescent or biofluorescent reporter proteins in the promoter / transcription element - reporter gene constructs of the invention. For example, in addition to and / or alternative to luciferase, the invention can use, e.g., a fluorescent protein, a luminescent protein, a secretable reporter protein, a luciferase, a secretable luciferase, a green fluorescent protein, and a red fluorescent protein. Secretable luciferase (as well as other secretable reporter gene products) that can be used in various embodiments of the invention can be seen in, e.g., WO / 2008 / 073805 "Secretable Reporter System," filed December 6, 2007. Other bioluminescent and biofluorescent reporter proteins that can be used in various constructs of the invention will be familiar to those of skill in the art. See, e.g., Haugland, Handbook of Fluorescent Probes and Research Products, Molecular Probes, Inc., Eugene Oregon, 2005, and the references cited therein. The reporter protein expressed when the promoter is activated can be detected and quantified by any of a number of methods well known to those of skill in the art in addition to use of luciferase assays. These can include analytic biochemical methods such as electrophoresis, capillary electrophoresis, high performance liquid chromatography (HPLC), thin layer chromatography (TLC), hyperdiffusion chromatography, and the like, or various immunological methods such as fluid or gel precipitin reactions, immunodiffusion (single or double), immunoelectrophoresis, radioimmunoassay (RIA), enzyme-linked immunosorbent assays (ELISAs), immunofluorescent assays, western blotting, and the like. For example, an encoded polypeptide (e.g., luciferase) can be detected / quantified in an electrophoretic protein separation (e.g. a 1- or 2-dimensional electrophoresis). Means of detecting proteins using electrophoretic techniques are well known to those of skill in the art {see generally, R. Scopes (1982) Protein Purification, Springer- Verlag, N. Y.; Deutscher, (1990) Methods in Enzymology Vol.182: Guide to Protein Purification, Academic Press, Inc., N. Y.). Western blot (immunoblot) analysis can be used to detect and quantify the presence of an encoded reporter protein. The encoded reporter polypeptide can also be detected using an immunoassay. As used herein, an immunoassay is an assay that utilizes an antibody to specifically bind to the analyte (e.g., the target polypeptide(s)). The immunoassay is thus characterized by detection of specific binding of a reporter polypeptide to an antibody as opposed to the use of other physical or chemical properties to isolate, target, and quantify the analyte. The level of reporter polypeptide present can also be determined by an enzyme immunoassay (EIA) which utilizes, depending on the particular protocol employed, unlabeled or labeled (e.g., enzyme-labeled) derivatives of polyclonal or monoclonal antibodies or antibody fragments or single-chain antibodies that bind reporter polypeptide(s), either alone or in combination, hi the case where the antibody that binds the target polypeptide(s) is not labeled, a different detectable marker, for example, an enzyme-labeled antibody capable of binding to the monoclonal antibody which binds the target polypeptide, can be employed. Any of the known modifications of EIA, for example, enzyme-linked irnmunoabsorbent assay (ELISA), can also be employed. Changes in expression levels of a reporter gene (e.g., luciferase) can also be detected by measuring changes in mRNA and / or a nucleic acid derived from the mRNA (e.g. reverse-transcribed cDNA, etc.) that encodes a polypeptide of the gene product or a gene product of a nucleic acid that is under control of the transcription element in the reporter construct. Any of the methods provided herein are amenable to high throughput screening. Preferred assays detect increases or decreases in reporter (e.g., luciferase) transcription and / or translation, e.g., in response to the presence of a test transcription modulating agent (e.g. a test compound). Cells (or wells in an assay plate) utilized in the methods of this invention need not be contacted with a single test agent at a time. For example, to facilitate high-throughput screening, a single cell / well / etc, can be contacted by at least two, preferably by at least 5, more preferably by at least 10, and most preferably by at least 20 test compounds. If the cell / well scores positive, it can be subsequently tested with a subset of the test agents until the agents having the activity are identified. High throughput assays for various reporter gene products such as luciferase are well known to those of skill in the art. For example, multi-well fluorimeters are commercially available (e.g., from Perkin-Elmer). In addition, high throughput screening systems are commercially available {see, e.g., Zymark Corp., Hopkinton, MA; Air Technical Industries, Mentor, OH; Beckman Instruments, Inc. Fullerton, CA; Precision Systems, Inc., Natick, MA, etc.). These systems typically automate entire procedures including all sample and reagent pipetting, liquid dispensing, timed incubations, and final readings of the microplate in detector(s) appropriate for the assay. These configurable systems provide high throughput and rapid start up as well as a high degree of flexibility and customization. The manufacturers of such systems provide detailed protocols of the various high throughputs. Thus, for example, Zymark Corp. provides technical bulletins describing screening systems for detecting the modulation of gene transcription, ligand binding, and the like. High throughput screening formats are particularly useful in identifying activity of transcription factors. Generally in these methods, one or more biological sample that includes a transcription factor (i.e., in a driver construct and along with reporter constructs, etc.) is contacted, serially or in parallel, with a plurality of test compounds comprising putative modulators (e.g., the members of a modulator library). Binding to or modulation of the activity of the transcription factor by a test compound is detected, thereby identifying one or more modulator compound that binds to or modulates activity of the transcription factor. In one of the embodiments the reporters comprise a unique barcode sequence and / or an open reading frame encoding a reporter protein. Such a barcode may be particularly advantageously designed to accommodate quantification and identification of nucleic acid sequences encoding for the transcription factor reporters (e.g. expressed and / or comprised in the genome) allowing the quantification of barcode expression in the RNA of the modified cells and correlating said expression with TF activity. A barcode sequence can be any nucleotide sequence that, based on its sequence, is uniquely distinguishable and allows identification of the encoded receptor. A barcode sequence may be between 5 and 20 nucleotides in length. It is understood that with a barcode sequence consisting of only 5 nucleotides, this allows for 625 (25 x 25) identifiers (without any redundancy). The skilled person is well capable of designing and incorporating a barcode sequence in the nucleic acid sequence, or, selecting suitable barcodes known in the art. A “Mammal” or “mammalian” as used in relation to the invention hereunder means a class of vertebrate animals characterized by several distinctive features such as regulating their internal body temperature regardless of external conditions, female mammals have mammary glands that produce milk to feed their young and typically have larger and more complex brains relative to their body size compared to other animals, which enables higher order functions such as learning and problem-solving. Examples of mammals are humans, mice, rats, dogs, cats and rabbits. An important embodiment of the invention is the ability to detect the activity of multiple TFs simultaneously in a cell, an organoid, or an entire mammal. This may be achieved by assembling the TF reporters in a tandem array wherein TF reporters are assembled next to each other. Each individual TF reporter in the array may contain the following elements: a TF response element (containing four identical TFBSs and optimized spacer sequences; ~50-100 bp in total depending on the length of the TFBS), a minCMV core promoter (56 bp), a short transcription unit (~100 bp) containing a unique DNA barcode and primer binding sites to amplify the transcription unit including the DNA barcode, a strong transcription termination sequence (SV40 PolyA sequence; 100 bp). To prevent cross-talk, adjacent TF reporters are separated from each other with an insulator sequence (based on PMID 24098520; 237 bp). Individual TF reporters are ordered as DNA fragments and assembled into a plasmid backbone in one step using Golden Gate assembly. The array can then be integrated into a single safe-harbor locus in the genome of any cell line of interest using recombination-mediated cassette exchange or CRISPR-Cas9. As detailed above, essentially any available compound library, e.g., a peptide library, or any one or combination of compound libraries described herein, can be screened to identify putative modulators in a high-throughput format against a biological or biochemical sample.96-wells or 384-well drug screens may be used. In an embodiment, a computer-implemented method for determining transcription factor (TF) reporter activities is provided, the method comprising: a) Receiving sequencing data for a plurality of TF reporters for at least one control sample; and b) Receiving sequencing data for a plurality of TF reporters for at least one contrast sample; and c) Inputting the sequencing data in a) and the sequencing data in b) into a statistical model which outputs: i) a measure of TF reporter activities; and / or ii) a measure of changes in TF reporter activities between the control sample in a) and the contrast sample in b). The computer- implemented method provides a measure of TF reporter activities directly from sequencing data. The plurality of TF reporters as is described herein. The method compares at least two conditions at a time: at least one control or reference condition, and at least one contrast or treatment condition. There may be multiple contrast or treatment conditions. The control sample may be a first condition for example cells with no treatment or a different drug treatment to the contrast sample. The contrast sample may be a second condition for example cells with a drug treatment or a different drug treatment to the control sample. The input data may be: (1) high-throughput sequencing data for TF reporters in a first condition sample, (2) high-throughput sequencing data for TF reporters in a second condition sample, and optionally (3) the plasmid DNA (pDNA) counts or genomic DNA (gDNA counts). gDNA counts may be used for lentiviral data. pDNA sequencing data does not necessarily need to be generated for every single experiment. Previous pDNA sequencing data may be used if the same plasmid library was sequenced in an earlier experiment. The input data may be in fastq format. The statistical model may be based on barcode analysis using linear models (BCalm {Keukeleire, 2025, 39948460}). The model (BCalm) may calculate the activity of the TF reporters by deriving the ratio of transcribed RNA (the high-throughput sequencing data of (1) or (2)) to the input pDNA counts (high- throughput sequencing data of (3)). The model (BCalm) may perform differential analysis by fitting, for each individual barcode, a linear model on the input data. The model (BCalm) may perform statistical analysis to compute changes in reporter activities between two conditions by running a re-implementation of the Limma framework {Ritchie, 2015, 25605792}. P-values may be extracted from the statistical model. The p-values may be adjusted for multiple comparisons using any method known in the art, for example, Bonferroni correction or False Discovery Rate (FDR) correction. The method calculates the activity of each TF in each condition and / or performs a comparative analysis to statistically assess changes in TF activity between the two compared conditions. The method produces an output file stating the activity of the TF reporters for each condition, as well as the Log2Fold-Change values (comparing the activity of the reference condition and the contrast condition) and the raw and adjusted P- values. The method may use single comparisons or multiple comparisons. The at least one quality control measure may be a plasmid DNA bleedthrough estimation plot. These plots correlate the counts in the plasmid DNA with the counts derived from the RNA. This is done using a set of negative control reporters. The slope of the correlation is used to estimate the level of bleedthrough. The at least one quality control measure may be correlation plots between barcode counts to estimate the level of technical noise. The at least one quality control measure may be correlation plots between biological replicates to estimate the level of reproducibility. The at least one quality control measure may be an estimation of sequencing depth by plotting the number of reads per barcode per sample and the total number of reads per sample. The at least one quality control measure may be a plot of the reads normalized by sequencing depth of all barcodes and reporters per TF for all biological replicates per condition. These plots are very useful to determine if the perturbation has reproducible effects across all biological replicates and barcodes on each individual TF. The quality control measure may be one or more of the plots described above. Besides quality control measures, the method may also generate at least one exploratory plot guiding differential TF activity analysis. The exploratory plots may be useful for visualisation of the data and / or results. The at least one exploratory plot may include a lollipop plot representing the activity of all measured TFs, ordered by their activity level in a reference (e.g., control) condition. The lollipop plot may display TF activity in both the control and contrast conditions, with statistically significant differences indicated by color coding (e.g., red for upregulated TFs and blue for downregulated TFs), based on a p-value threshold produced by the statistical model. This plot may assist in identifying the most active TFs in the control condition and those showing significant regulation in the contrast condition. Alternatively or additionally, the at least one exploratory plot may be a volcano plot, which may display the Log2 fold-change in TF activity and corresponding adjusted p-values for each TF, thereby highlighting TFs with statistically significant changes in activity between the contrast and control conditions. The at least one exploratory plot may include a heatmap illustrating either the Log2 fold-change in TF activity or a transformed score, such as the product of the Log2 fold-change and the negative Log10 of the adjusted p-value, for all TFs across multiple contrast conditions. The at least one exploratory plot may include a dimensionality reduction plot, such as a principal component analysis (PCA), t-distributed stochastic neighbor embedding (t-SNE), or uniform manifold approximation and projection (UMAP) plot. These plots may visualize variance in TF reporter read counts across conditions and facilitate assessment of similarity or clustering among experimental conditions. The method may output at least one quality control measure and / or at least one exploratory plot. The method may output both (i) at least one quality measure and (ii) at least one exploratory plot. The methods of the present disclosure may be conducted on receipt of suitable computer readable instructions, which may be embodied within a computer program running on a processor. In another embodiment, a computing apparatus for determining transcription factor (TF) reporting activities is provided, the apparatus comprising; one or more memory units; and one or more processors configured to execute instructions stored in the one or more memory units to perform the method described herein. There may also be provided a non-transitory computer-readable storage medium including instructions that, when executed by one or more processors, cause the one or more processors to perform any of the methods described herein. The methods of the present disclosure may be implemented in hardware, or as software modules running on one or more processors. The methods may also be carried out according to the instructions of a computer program, and the present disclosure also provides a computer readable medium having stored thereon a program for carrying out any of the methods described herein. A computer program embodying the disclosure may be stored on a computer-readable medium, or it could, for example, be in the form of a signal such as a downloadable data signal provided from an Internet website, or it could be in any other form. Figure legends Systematic design and probing of TF reporters. (A) Design of the TF reporter library. Four copies of the TFBS are placed around variable spacer sequences and spacer lengths upstream of a core promoter and a unique barcode. The arrow in the core promoter indicates the transcription start site. (B) Correlation of all reporter activities (i.e., activation compared to TF-neg reporters) measured in biological replicate 1 compared to replicate 2 of reporters with mutated TFBSs (grey) and consensus TFBSs (red). Reporter activities in all nine cell lines are displayed together. (C) Activities of individual reporter designs per TF in mNPCs and mESCs. Highlighted in orange are mESC-specific TFs. Red line indicates median activity per TF. Figure S1: TF reporter library design and activity characterization. (A) Pairwise Pearson correlations coefficient (PCC) heatmap for all 86 motifs chosen for the TF reporter library design. Motifs clustering together are highlighted. (B) Schematic overview of the design of the motif-depleted spacer sequences. (C) Correlations between the reporter activities of individual barcodes. Displayed are 300 randomly sampled reporters. Figure panels on the diagonal show the density distribution of the reporter activities per barcode. Figure panels below the diagonal show the pairwise correlation plots, and panels above the diagonal indicate the PCCs. (D) Average PCC of all pairwise correlations between biological replicates per cell line. (E) Reporter activity distributions per cell line of reporters with mutated TFBSs (grey) and consensus TFBSs (red). (F) Comparison of reporter activities of TF reporters and the genomic Klf2 gene enhancer controls in mESCs. Red line indicates median reporter activity. Figure 2: Identification of design features causative for TF reporter activity. (A) Equation of the log-linear model. Reporter design features are indicated by color and reflect the reporter design cartoon underneath the equation. (B) Correlation between measured STAT3 reporter activities and predicted reporter activities by the log-linear model in mESCs. Color-coded are the two spacer lengths. (C) Weights of the individual coefficients of the STAT3 log-linear model. The color indicates the strength of the weight. (D) Coefficient weight heatmap for all TFs with a significant log-linear model fit and a total explained variance of > 50%. As in C, all weights are computed in contrast to the reference variables minP (core promoter), 10 bp (promoter distance), and #1 (spacer sequence). Weights of features that did not significantly contribute to the model (p ≥ 0.1) are set to 0 in this visualization. TFs highlighted in red display spacer sequence preferences, TFs in blue spacer length preferences. TFs indicated in bold are mentioned in the text. (E) Total variance explained by the individual design features for all TFs displayed in D. Red line indicates the median. The color of the dots indicates the total explained variance of the log-linear model. (F) Average difference between the weights of 10 bp spacer sequences in the log-linear model and the 5 bp spacer sequences, separately per monomeric and dimeric / multimeric TFBSs. Statistical significance of the difference in variance is estimated by Levene’s test. Figure S2: TF reporter activities across all probed cell lines. Reporter activities per TF in all nine probed cell lines. Each dot represents a unique reporter design. Figure 3: Investigating TF specificity of reporters. (A) Correlations between TF reporter activities and TF transcript abundances across the nine probed cell lines per TF. Only TFs with variable expression across the nine cell lines are included (see Methods). The black solid line indicates the mean PCC per TF. TFs highlighted in red are mentioned in the text. Dots highlighted with a red stroke are depicted in B-D. (B) Correlation between GATA4 transcript abundance and reporter activity for a highly (spacer sequence #6) and a poorly correlating reporter (spacer sequence #4). The two displayed reporters are identical except for the spacer sequence mentioned above the panels. Cell lines are color-coded. Solid line indicates linear regression, grey shade indicates standard deviation. nTPM = normalized TPM (see Methods). (C) Same as B, but for GATA1 and two promoter distances. (D) Same as B, but for TFCP2L1 and two spacer lengths. Log-linear models highlight importance of reporter design. (A) Distribution of the p-values of all log-linear models (one per TF) with measured data (blue) and randomized data (orange; measured activities were randomly assigned to reporter designs per TF). (B) HNF4A reporter activities per spacer length. The bar indicates the mean reporter activity per spacer length, the dots indicate activities of individual reporters. The red line denotes the standard deviation. Difference in activity between the groups is tested by student’s t-test; *** < 0.001. Response to signaling pathway perturbations. (A) Change in TF reporter activities upon signaling pathway perturbations. Shown are only the responses of the direct targets of the perturbations. Activating perturbations are shown in blue, repressing conditions in purple. Conditions that were selected as best perturbation condition for the TF (i.e., strongest average response of the tested perturbations for that TF, see Methods) are denoted by an asterisk. TFs highlighted in the text are indicated in bold. TFs depicted in B-G are indicated by letter. LIF = leukemia inhibitory factor. PMA = phorbol 12-myristate 13-acetate, HQ = hydroquinone, CDCA = chenodeoxycholic acid. (B-G) Response of reporters of six different TFs to TF- targeted pathway perturbation conditions. TFs for which reporters are displayed is denoted on top of each figure. Reporter activities (log2) in the basal condition are displayed on the x-axis and in the perturbation condition on the y-axis. Reporter design features are indicated by color, published reporters by shape. Reviewing TF specificity of reporters. (A) PCC p-value of GATA4 reporters per spacer sequence and spacer length. Mean is indicated by the bar, and individual dots represent individual reporters. Significant p-values (p < 0.1) are indicated by green color. (B) Same as (A) but for GATA1 reporters per promoter distance. (C) Same as (A) but for TFCP2L1 reporters. Difference in PCC p-value between the groups is tested by student’s t-test; * < 0.05, ** < 0.01, *** < 0.001. Response of reporters to direct TF perturbation. (A) Change in TF reporter activities upon direct TF perturbation. In some cases the target TF consists of two TFs (e.g., POU5F1::SOX2); the perturbed TF is then indicated in the x-axis labels. TF overexpression is shown in black, TF knockdown in purple, and TF degradation in green. Conditions that were selected as best perturbation condition for the TF are denoted by asterisk. TFs highlighted in the text are indicated in black. TFs depicted in B-H are indicated by letter. (B-G) Response of TF reporters to six different direct TF perturbation conditions. TFs for which reporters are displayed is denoted on top of each figure. Reporter activities (log2) in the basal condition are displayed on the x-axis and in the TF perturbation condition on the y-axis. (H) Response of GRHL1 reporters to GRHL1 (y-axis) and TFCP2L1 knockdown (x-axis). Reviewing TF specificity of reporters. (A) Changes in reporter activities upon perturbation of related TFs by either TF overexpression (black), knockdown (purple), or degradation (green). (B) Changes in reporter activity of NR1H4 reporters upon all perturbations. Only top changing conditions and conditions that perturb NR1H4 or TFs with similar TFBSs are indicated. Identification of TF-specific and sensitive reporters. (A) Reporter confidence levels are defined based on the four threshold criteria mentioned in the boxes. Response to known TF perturbation is given a higher weight due to its importance. (B) Reporter confidence scores of STAT3 reporters. Reporter activity, TF abundance correlation, or TF perturbation response meeting the threshold criteria outlined in A contribute to the reporter confidence level and are denoted by a plus or minus sign. (C) Overview of the confidence level of the best reporter per TF for TFs with both synthetic and published reporters probed. (D) Same as C but for TFs with only synthetic reporters probed. TP53 and NR3C1 are included in this list because their published reporters were not probed in TP53 / NR3C1 perturbation conditions, prohibiting comparisons between synthetic and published reporters. (E) Same as C and D but for TFs for which only published reporters were included in the reporter library design. (F) Reporter activity of the prime reporters with consensus TFBS (blue dot) and mutated TFBS (grey dot). Activities displayed are from the same conditions as used for the log-linear models. Example of reporter confidence level heatmaps for all TFs. Upper heatmap: confidence levels per reporter. Middle heatmap: Activities (first row), TF abundance correlation (second row), perturbation fold-change (third row), and off-target perturbation fold-change (fourth row) per reporter. Lower heatmap: Color-coding of the reporter design. Multiplexed detection of TF activities with prime reporters. (A) TF activities as measured by the prime reporters across all nine probed cell lines. Activities were scaled by dividing the reporter activities by the maximum activity per TF. (B) Changes in prime reporter TF activities upon various TF perturbations in mESCs. TF targets of perturbations are indicated by black rectangles and asterisks. Only TFs expressed in mESCs (nTPM > 4) and with a substantial perturbation response (fold-change > 2) in at least one condition are displayed. DEG = degradation. Identifying high-confidence TF reporters. Change in reporter activity upon TFBS mutation of the 47 prime reporters with matched mutated reporters compared to reporters for the same TFs with a lower confidence level (mean across all reporters with one confidence level lower than the prime reporter). Red line indicates a fold-change of 2. Perturbation responses of high-confidence TF reporters. (A) Activities of the prime reporters in the nine probed cell lines. (B) Data shown in Figure 7A visualized as UMAP. Color codes are based on clustering in Figure 7A. (C) Reporter activities (max normalized) of the synthetic HNF4A prime reporter (red) and the highest-ranking published HNF4A reporter per cell line. (D) Changes in TF activities upon KD of TFs in HEPG2 cells. KDs that did not reduce the activity of their target TF (log2-fc > -0.25) are not included in this visualization. Example mentioned in the text is highlighted in bold. Graphical representation of protocol for multiplexed TF activity detection using a barcoded plasmid library of optimized “prime” TF reporters. The protocol comprises library transfection, RNA processing for barcode sequencing, and a computational pipeline for analyzing differential TF activity, enabling high-throughput and quantitative TF profiling. Scheme of prime TF reporter library consisting of 100 distinct TF reporters. Each TF reporter carries four copies of TF binding sites located upstream of a core promoter (indicated by the arrow), a unique DNA barcode, and a GFP open reading frame. The plasmid library is transfected into cells after which the transcribed barcodes can be harvested from the RNA. Barcode cDNA synthesis and PCR. A) Scheme of barcode cDNA synthesis using a GFP-specific reverse transcription primer. B) Scheme of the PCR to amplify the barcodes from cDNA (upper panel) and pDNA (lower panel). C) Example of two PCR products loaded on an agarose gel, using either 18 or 20 cycles including -RT and H2O negative controls. The lower band (~100 bp) represents primer dimers and the upper band (~225 bp) indicates the expected PCR product of the amplified barcodes. Example of QC plots generated by the computational pipeline described herein. A) Correlation matrix between the five barcodes of a sample. Each row and column correspond to a barcode. In the lower- left corner, a scatterplot of the read counts depicts the correlation between the pairs of barcodes, with a dashed red line showing the diagonal. In the upper-right corner, we show the Pearson correlation values for each pair. On the diagonal, we show the distribution of read counts for each barcode. B) The same type of plot, illustrating the correlations of TF activity between biological replicates of a sample. C) Correlation plot between the pDNA read counts and cDNA read counts of five different samples to estimate the pDNA bleedthrough. The panel background colors represent the level of pDNA bleedthrough: blue for less than 10%, yellow for between 10% and 25%, and orange for greater than 25%. Exploratory table and plots generated by the computational pipeline described herein. A) Overview of the main output file. Results are stored in a tabular format which contains the activity of each TF for each condition (here: DMSO and calcitriol) and the results of the comparative analysis. Key metrics include LogFC values (Log2(Calcitriol / DMSO)), raw and adjusted p-values, and a summary column (“sig”) indicating significance: NS (non-significant, adjusted p-value ≥ 0.01), Downregulated (significant with higher activity in the reference condition), and Upregulated (significant with higher activity in the contrast condition). B) Volcano plot of the analysis, stored as “primetime_volcano.pdf”. This plot shows the adjusted p-values and LogFC for each TF, highlighting significant TFs (red for downregulated and blue for upregulated, non-significant TFs in gray). For a better visualization, this plot shows the names of only the most up-and downregulated TFs. C) Lollipop plot of the results, stored as “primetime_lollipop.pdf”. Black dots represent TF activity in the reference condition (here: DMSO), while colored dots indicate activity in the contrast condition (here: calcitriol). The significance of TFs follows the same color code as in the volcano plot. Read counts (in reads per million, RPM) for all barcodes associated with each TF across the tested samples. Each dot represents an individual barcode. TFs are sorted by the magnitude of differential activity, calculated as Log₂FoldChange multiplied by -Log₁₀(p-value). Displayed in the figure are only the top 20 changing TFs. Statistically up- and downregulated TFs are highlighted in red and blue, respectively. This plot enables assessment of fold-change consistency across replicates. Figure 15: TF activity measured from 12-wells (x-axis) or from 96-wells (y-axis). Measured activities almost perfectly correlate, indicating that TF reporter activities can be efficiently retrieved from 96-wells. Figure 16: TF activities measured after 1 hour of recovery from a 1-hour heat shock vs. immediately post- heat shock. A) shows measurements from total RNA; B) shows measurements from nascent RNA. In the nascent RNA, HSF1 activity decreases after recovery, indicating higher temporal resolution. In contrast, total RNA fails to capture this drop, as previously transcribed barcodes persist, masking dynamic changes in transcription factor activity. Examples Example 1 Systematic probing of a TF reporter library Selection of TFs. A main challenge in the design of specific TF reporters is the similarity between binding motifs of TFs. Therefore, to select TFs for which the generation of TF-specific reporters would be feasible, we manually examined all human TFs (n = 1,244) and reviewed their (i) TF motif quality (i.e., motif length and information content), (ii) the number of TFs with a similar motif, (iii) expression pattern across cell types, and (iv) stimulation and perturbation opportunities. Based on these criteria we selected a list of 86 TFs (Table S1). For each TF, we selected the best motif according to a previous motif curation {Lambert, 2018, 29425488}. We also included several heterodimeric motifs (e.g., POU5F1::SOX2), for which we carefully reviewed available motifs. Most of the selected TFs have unique motifs (i.e., no other TF has a similar motif, Figure S1A), and cover a large diversity of the human TF motif landscape. The selected 86 TFs include most well-known TFs downstream of generic signaling pathways such as MAPK, PI3K / AKT, TGFbeta, WNT, and JAK-STAT, as well as a diversity of nuclear receptors and tissue-specific and pluripotency- specific TFs (Table 1, Table S1). Table 1. Overview of the selected TFs and their primary associated cellular functions. Note that some TFs might have multiple functions. TFs for which only published reporters were included are displayed in brackets. TF Main cellular function AHR::ARNT, NR1I2, NR1I3 Xenobiotic stress response Example 2 Library design. We generated a library consisting of synthetic TF reporters for the selected 86 TFs. For each TF, we generated a consensus TFBS by choosing the most conserved base at each position of its motif (Table 2). We also included two sets of negative control TFBSs. First, for each TF we generated a matched mutated TFBS in which two or three conserved bases of the TFBSs were modified. Second, we designed three distinct 11 bp random sequences that are devoid of any TFBS (TF-neg) that served as generic negative controls and were used for normalization. To generate synthetic TF reporter sequences, we placed four copies of the TFBSs in front of a core promoter that drives the transcription of a unique 13-bp barcode sequence and a GFP open reading frame (Figure 1A). We chose to use four copies of TFBSs as this number was shown to achieve optimal activation for many TFs {Davis, 2020, 32603702; Sharon, 2012, 22609971; van Dijk, 2017, 27965290; Trauernicht, 2023, 37650627}. We then systematically varied several design parameters for each TFBS (Figure 1A). First, we designed three sets of spacer sequences around the four TFBSs of either 5 or 10 bp (i.e., TFBS1-spacer1-TFBS2-spacer2-TFBS3-spacer3-TFBS4) (Table 3). These spacer sequences were computationally designed to minimize occurrences of secondary activator TFBSs, even in the junctions between the spacer sequences and the TFBSs (Figure S1B). For each spacer length (5 and 10 bp), we then selected three distinct sets of spacer sequences. Second, we coupled the TFBSs to three different core promoters (minP (derived from pGL4 (Promega, Madison, WI, USA)), minCMV {Li, 2009, 18701910}, or for some TFs also minHBG {Collis, 1990, 2295312}). Third, we placed the core promoter at either 10 or 21 bp from the nearest TFBS. Together, the combination of these design parameters yielded 36 reporter designs for TFs with minHBG, and 24 for TFs without. Additionally, for comparison we also included previously established and published reporter sequences for 62 TFs from three different public sources {O’Connell, 2016, 27211859; Romanov, 2008, 18297081; Promega - see Table S1}. Another 120 genomic enhancer fragments of the Klf2 gene (previously shown to be active in MPRAs in mESCs {Martinez-Ara, 2022, 35594855}), and 86 reporters with a TFBS-devoid core promoter (one for each TF) were included as positive and negative controls, respectively (see Methods). Together, this yielded a collection of in total 5,530 unique reporter sequences. Finally, each of these sequences was coupled to 5-8 distinct barcodes to minimize biases caused by individual barcodes, yielding a library of 35,500 barcoded reporters. Table 2. selected TFBS AAACCGGTTT AGTTAATCATTAACT GGGGTCAAAGTCCAAT TATAATCGTTTT GGAACGTTCTAGAAG GGAAACGGAAACCGAAAC TACATGAATATTCATGTA SEQ ID NO: 26 KLF4 CCACGCCC SEQ ID NO: 27 MAF::NFE2 ATGACTCAGCAATTT SEQ ID NO: 28 MEF2A TCTAAAAATAGA SEQ ID NO: 46 NR4A2 GAGGTCATTGACCCC SEQ ID NO: 58 RARA AAGGTCATTTGAGGTCA AGGTCACGGAGAGGTCA TTCCCA GGAAATCCCC CGTTGCCATGGCAACG CAAAGGTCAAATTGAGGTCA CTAACCGCAAAAACCGCAAC GGGGTCAAAGGTCA SEQ ID NO: 66 SMAD2::3::4 CTGTCTGTCACCT SEQ ID NO: 67 SMAD4 TCTAGACA SEQ ID NO: 68 SOX9 TATCAATAACATTGATA SEQ ID NO: 74 TCF7 ACATCAAAG AGATCAAAGG CACATTCCAT TGCCCCCGGGCA CCGGTTCGAACCGG GTGACCTTATGAGGTCAC GTGACCTCAATGAGGTCAC GGACATGCCCGGGCATGT CAGGTCACCAGGTTCAC TGCGGGGGAGT SEQ ID NO: 84 XBP1 GATGACGTGGCATT

[0002] Table 3. spacer sets Spacer set 1 SEQ ID NO 87 TCTAT SEQ ID NO 88 GCTAT SEQ ID NO 89 GTCTT SEQ ID NO 90 TCGACACTCT SEQ ID NO 91 GATCGTTCAA SEQ ID NO 92 GGTCCACTAG SEQ ID NO 93 TCTAT SEQ ID NO 94 TCAGA SEQ ID NO 95 TCGCC SEQ ID NO 96 GAGCGGTGCA SEQ ID NO 97 CCGTACTCCT SEQ ID NO 98 GGAGAGTATA Spacer set 2 SEQ ID NO 99 ACTAT SEQ ID NO 100 GATAT SEQ ID NO 101 AGAAA SEQ ID NO 102 TCTAATATCT SEQ ID NO 103 TCGAGCTATC SEQ ID NO 104 AGGAGCTCGG SEQ ID NO 105 CGCTC SEQ ID NO 106 AGTAG SEQ ID NO 107 ACTCT SEQ ID NO 108 TCGATGGACG SEQ ID NO 109 TCGCATCTCG SEQ ID NO 110 GGACGAGTAC Spacer set 3 SEQ ID NO 111 GAGCG SEQ ID NO 112 CGATG SEQ ID NO 113 CCGAT SEQ ID NO 114 TAGTCTGAAG SEQ ID NO 115 CGACTATCTT SEQ ID NO 116 TAGTCCGATG SEQ ID NO 117 TTGGT SEQ ID NO 118 GGTCC SEQ ID NO 119 TGGTA SEQ ID NO 120 AGGATCCTAC SEQ ID NO 121 GAATCCGTCC SEQ ID NO 122 CAGGCCGCTT Promoter_spacer_10bp SEQ ID NO 123 GATCCTATAC SEQ ID NO 124 GATAGTACTC SEQ ID NO 125 TCTCGTATAC SEQ ID NO 126 GATCCTATAC SEQ ID NO 127 GATAGTACTC SEQ ID NO 128 TCTCGTATAC SEQ ID NO 129 CGACTCACTT SEQ ID NO 130 ATACTCCTCC SEQ ID NO 131 GATAGTACTC SEQ ID NO 132 CGACTCACTT SEQ ID NO 133 ATACTCCTCC SEQ ID NO 134 GATAGTACTC Promoter_spacer_21bp SEQ ID NO 135 TAACTCTACTGTGTCTATACC SEQ ID NO 136 SEQ ID NO 137 TCGTACAAACGCCTTTTTCAC SEQ ID NO 138 TAACTCTACTGTGTCTATACC SEQ ID NO 139 SEQ ID NO 140 TCGTACAAACGCCTTTTTCAC SEQ ID NO 141 SEQ ID NO 142 CGACGCTGTTTGTTACTCAGT SEQ ID NO 143 SEQ ID NO 144 SEQ ID NO 145 CGACGCTGTTTGTTACTCAGT SEQ ID NO 146

[0003] Table 4. synthetic reporter sequences NO.168 HOMEZ GCCATCTATTGCTTACATTTGCTTCT

[0004] NO.170 CGGGCTGGGCATAAAAGTCAGGGCAGAGCCATCTATTGCTTACATTTGCTTCT SEQ ID 1 IRX TACATGAATATTCATGTATCTAATATCTTACATGAATATTCATGTATCGAGCTATCTACATGAATATTCATGTAAGGAGCTCGGTACATGAATATTCATGTAGGGTTCTAACGAT NO.17 3 ATGTGAAAGGGCTGGGCATAAAAGTCAGGGCAGAGCCATCTATTGCTTACATTTGCTTCT SEQ ID KLF4 CCACGCCCTCTATCCACGCCCGCTATCCACGCCCGTCTTCCACGCCCTAACTCTACTGTGTCTATACCGGCGTTTACTATGGGAGGTCTATATAAGCAGAGCTCGTTTAGTG NO.172 AACCGTCAGATC SEQ ID MAF::NF ATGACTCAGCAATTTAGGATCCTACATGACTCAGCAATTTGAATCCGTCCATGACTCAGCAATTTCAGGCCGCTTATGACTCAGCAATTTGATAGTACTCGGCGTTTACTATG NO.173 E2 GGAGGTCTATATAAGCAGAGCTCGTTTAGTGAACCGTCAGATC SEQ ID TCTAAAAATAGAGAGCGGTGCATCTAAAAATAGACCGTACTCCTTCTAAAAATAGAGGAGAGTATATCTAAAAATAGAGCTCGAAATCTAGTGGGTTTGGGTTAGCGATCCAA NO.174 MEF2A TTCAGCTAGATTTTAAGC SEQ ID MTF1 AAGGCCGTGTGCAAAAGTCTATAAGGCCGTGTGCAAAAGTCAGAAAGGCCGTGTGCAAAAGTCGCCAAGGCCGTGTGCAAAAGGCTCGAAATCTAGTGGGTTTGGGCGTT NO.175 TACTATGGGAGGTCTATATAAGCAGAGCTCGTTTAGTGAACCGTCAGATC SEQ ID MYBL2 TTAACGGTTATTCGATGGACGTTAACGGTTATTCGCATCTCGTTAACGGTTATGGACGAGTACTTAACGGTTATATACTCCTCCTAGAGGGTATATAATGGAAGCTCGACTTC SEQ ID NR3C2 GGGAACACAATGTTCCCGAGCGGTGCAGGGAACACAATGTTCCCCCGTACTCCTGGGAACACAATGTTCCCGGAGAGTATAGGGAACACAATGTTCCCGCTCGAAATCTAG NO.190 TGGGTTTGGGTTAGCGATCCAATTCAGCTAGATTTTAAGC SEQ ID N TAAAGGTCACTCTATTAAAGGTCACTCAGATAAAGGTCACTCGCCTAAAGGTCACGCTCGAAATCTAGTGGGTTTGGGCGTTTACTATGGGAGGTCTATATAAGCAGAGCTC NO.191 R4A1 GTTTAGTGAACCGTCAGATC SEQ ID GAGGTCATTGACCCCGAGCGGTGCAGAGGTCATTGACCCCCCGTACTCCTGAGGTCATTGACCCCGGAGAGTATAGAGGTCATTGACCCCGCTCGAAATCTAGTGG 192 NR GTTT NO. 4A2 GGGCGTTTACTATGGGAGGTCTATATAAGCAGAGCTCGTTTAGTGAACCGTCAGATC

[0005] NO.213 TTGCTTCT SEQ ID GAACAATGGTCTATGAACAATGGTCAGAGAACAATGGTCGCCGAACAATGGGCTCGAAATCTAGTGGGTTTGGGCGTTTACTATGGGAGGTCTATATAAGCAGAGCTCGTTT NO.214 SOX2 AGTGAACCGTCAGATC SEQ ID S TATCAATAACATTGATAGAGCGGTGCATATCAATAACATTGATACCGTACTCCTTATCAATAACATTGATAGGAGAGTATATATCAATAACATTGATAGCTCGAAATCTAGTGG NO.215 OX9 GTTTGGGTTAGCGATCCAATTCAGCTAGATTTTAAGC SEQ ID SP CCCCGCCCCCTCGACACTCTCCCCGCCCCCGATCGTTCAACCCCGCCCCCGGTCCACTAGCCCCGCCCCCGATCCTATACGGGCTGGGCATAAAAGTCAGGGCAGAGCC NO.216 1 ATCTATTGCTTACATTTGCTTCT

[0006] SEQ ID SRF GCCATATATGGTTCGATGGACGGCCATATATGGTTCGCATCTCGGCCATATATGGTGGACGAGTACGCCATATATGGTATACTCCTCCGGCGTTTACTATGGGAGGTCTATA NO.217 SEQ ID NO.218 GAGGTCTATATAAGCAGAGCTCGTTTAGTGAACCGTCAGATC SEQ ID NO.219 SEQ ID ACATCAAAGGAGCGGTGCAACATCAAAGCCGTACTCCTACATCAAAGGGAGAGTATAACATCAAAGGCTCGAAATCTAGTGGGTTTGGGCGTTTACTATGGGAGGTCTATAT NO.220 TCF7 AAGCAGAGCTCGTTTAGTGAACCGTCAGATC SEQ ID TCF7L AGATCAAAGGACTATAGATCAAAGGGATATAGATCAAAGGAGAAAAGATCAAAGGGGGTTCTAACGATATGTGAAAGGGCTGGGCATAAAAGTCAGGGCAGAGCCATCTATT NO.221 2 GCTTACATTTGCTTCT SEQ ID CACATTCCATTCGATGGACGCACATTCCATTCGCATCTCGCACATTCCATGGACGAGTACCACATTCCATATACTCCTCCGGCGTTTACTATGGGAGGTCTATATAAGC O.222 T AGAG N EAD1 CTCGTTTAGTGAACCGTCAGATC SEQ ID NO.223 SEQ ID NO.224 1 GTTTACTATGGGAGGTCTATATAAGCAGAGCTCGTTTAGTGAACCGTCAGATC SEQ ID NO.225 SEQ ID GTGACCTCAATGAGGTCACTCTATGTGACCTCAATGAGGTCACTCAGAGTGACCTCAATGAGGTCACTCGCCGTGACCTCAATGAGGTCACCGACTCACT 6 THR TTAGAGGGTATA NO.22 B TAATGGAAGCTCGACTTCCAG SEQ ID TP53 GGACATGCCCGGGCATGTTCTATGGACATGCCCGGGCATGTGCTATGGACATGCCCGGGCATGTGTCTTGGACATGCCCGGGCATGTGATCCTATACGGCGTTTACTATG NO.227 GGAGGTCTATATAAGCAGAGCTCGTTTAGTGAACCGTCAGATC SEQ ID VD NO.228 R TTACTCAGTGGCGTTTACTATGGGAGGTCTATATAAGCAGAGCTCGTTTAGTGAACCGTCAGATC SEQ ID NO.229 WT1 SEQ ID NO.230 XBP1 TTTACTATGGGAGGTCTATATAAGCAGAGCTCGTTTAGTGAACCGTCAGATC SEQ ID NO.231 SEQ ID NO.232 ZFX AGGCCTGAGCGAGGCCTCGATGAGGCCTCCGATAGGCCTTCTCGTATACGGCGTTTACTATGGGAGGTCTATATAAGCAGAGCTCGTTTAGTGAACCGTCAGATC

[0007] For the design of the reporters the design logic of table 5 was used. Table 5: Design logic of TF reporters. Summary of which reporter design features are important for reporter activity per TF. 14 FOXO1 NA 16 GATA4 Spacer sequence = SEQ ID NO 120-122 18 GLI1 Spacer sequence = SEQ ID NO 99-101 20 HNF1A minCMV 26 KLF4 Spacer sequence 10bp spacer length NA 5bp spacer length 34 NFE2L2 minHBG / minCMV & 5bp spacer length STAT1::2 NA minHBG / minCMV & & = SEQ ID NO 84 XBP1 minCMV & 10bp spacer length 86 ZFX 5bp spacer length Example 3. Systematic testing of TF activities. Among the 86 included TFs are many tissue-specific TFs. We therefore probed the reporter library in nine different cell lines from distinct tissues. Since TF binding specificities are highly conserved between human and mouse {Jolma, 2013, 23332764; Vorontsov, 2024, 37971293}, we tested the library in cell lines derived from both human (n = 7) and mouse (n = 2). Furthermore, we extensively perturbed TF activities by (i) activating or inhibiting upstream signaling pathways (n = 25), or (ii) changing the TF abundance by overexpressing, knocking down or degrading individual TFs (n = 73). We thus queried all 5,530 reporters across 98 TF-perturbation conditions. For each tested condition or cell line, we first normalized the barcode counts in the mRNA to the barcode counts in the input plasmid DNA. Activities were then computed from the plasmid DNA-normalized counts by calculating the induction over the median counts of the collection of the TF-neg reporters. This was done separately per core promoter. Reporter activities between individual barcodes correlated highly (Pearson's correlation coefficient (PCC) range 0.84 – 0.87, Figure S1C) and were averaged. We probed the reporter library per cell line in at least three (HEK293, K562) and up to 11 (mESCs) biological replicates, yielding per cell line an average PCC between replicates of 0.77 - 0.94 (Figure S1D). For downstream analyses we averaged the reporter activities of the replicates. As expected, reporters with consensus TFBSs were more active than reporters with minimal mutations in the TFBS in all nine tested cell lines (Figure 1B, S1F). Furthermore, the synthetic TF reporters reached activities as high as the genomic enhancer element reporters, showing that four copies of the same TFBS are as potent as highly active native enhancer elements of approximately the same length (Figure S1F). Example 4. Reporter activities depend on cell type. SEQ ID NO: We first characterized activities for all TFs and their 24-36 reporter designs across the nine probed cell lines. We found that known generic TFs displayed activities in all cell lines (e.g., ELK1, FOS::JUN), while known cell type-specific TFs were predominantly detected in a subset of cell lines (e.g., HNF1A or HNF4A in HEPG2), and some were not active in any cell type (e.g., VDR, see below; Figure S2). Next, to explore these cell type-specific activities in more detail, we focused on two different cell lines: mESCs and mESC-derived neural precursor cells (mNPCs). As expected, the reporter activities differed for many TFs between the two different cell lines (Figure 1C). TFs that displayed substantially higher activity in mESCs compared to mNPCs included POU5F1::SOX2, TFCP2L1, STAT3, KLF4, SOX2, and TCF7 (Figure 1C, highlighted in orange), which are known activating TFs of the mESC pluripotency network {Dunn, 2014, 24904165; Hackett, 2014, 25280218}. Interestingly, for several TFs the reporter designs showed substantial differences in activity, despite having identical TFBS. For instance, in mESCs some STAT3 reporter designs were as inactive as the TF-neg control reporters, while others were up to 25-fold more active than those controls (Figure 1C, highlighted in bold and orange). This indicates that the design of the reporter can have substantial effects on its activity. Example 5 Relation between reporter design and reporter activity. To investigate the relation between reporter design and reporter activity, we fitted for each TF a log-linear model using the reporter design features (core promoter identity, promoter distance, spacer sequence and length) as categorical input variables (Figure 2A). This analysis enabled us to extract which reporter design features contribute to the variation in reporter activity. For example, for STAT3 the model accurately reflected the measured reporter activities (adjusted R2 = 0.95) (Figure 2B), and indicated that spacer length and sequence were crucial to achieve high transcriptional activity, while promoter identity contributed moderately, and promoter distance was largely irrelevant (Figure 2C). This suggested that STAT3 is more active with TFBSs spaced by 10 bp, and that the spacer sequence can strongly impact reporter activity, even though the spacers were designed not to contain any known TF motif. Next, we designed active TF reporters according to universal reporter design rules, or whether each TF requires its own specific rules. We applied the log-linear model analysis to each of the 86 probed TFs, focusing on the cell line and culture condition in which the TF is most active (see Methods). For 67 out of 86 TFs (78%) the models reached statistical significance (adjusted p-value < 0.05; Figure S3A) and explained >50% of the variance in reporter activity. For these models we then extracted the underlying weights of the individual reporter design features (Figure 2D). This analysis revealed several important insights. First, for almost all tested TFs, reporters were more active when having a minCMV or minHBG promoter compared to a minP promoter. Note that the fitted activities are normalized to the background activity of the core promoter (as described in the ‘Data overview’ section), meaning that reporter activities reflect the TF-induced activity change compared to the promoter-only activity. Thus, minCMV and minHBG promoters allow for stronger induction, regardless of the TF. Second, although the promoter distance explained the least variance compared to all other investigated features (Figure 2D, E), the majority of TFs had a slightly decreased activity when the core promoter was placed 21 bp away from the first TFBS instead of 10 bp. This shows that for many TFs placing the TFBS closer to the TSS can subtly increase transcription activity. Example 6. TFBS spacer length affects activity. Besides the generic role of the core promoter and the core promoter distance, we found a striking TF- specific role for the spacer length between the TFBSs. For ten TFs, all three 10 bp spacer sequences consistently increased activity compared to the 5 bp spacer sequences (Figure 2D, TFs highlighted in dark blue). A readily interpretable example is HNF4A, for which > 90% of all variance in the reporter activity was caused by changing the spacer length from 5 to 10 bp (Figure 2D, E); this increased reporter activity by roughly 8-fold on average (Figure S3B). Conversely, six TFs had significant negative weights for all three 10 bp spacer sequences, and hence favored the shorter 5 bp spacer length (Figure 2D, highlighted in light blue). We then examined in greater detail which TFs exhibited these spacer length-preferences. Interestingly, we observed that TFs that bind DNA as monomers tended to be unaffected by changes in spacer length, while dimeric or multimeric TFs had significantly stronger spacer length-preferences (Figure 2F). In fact, 15 out of 16 TFs with consistent spacer length-preferences were dimeric or multimeric TFs. Possibly, dimeric or multimeric TF assemblies have more complex DNA interactions and might therefore need precise relative positioning to be able to activate efficiently from adjacent TFBSs. Several TFs benefit from specific spacer sequences. Besides TFs that clearly require certain spacer lengths to effectively activate, several TFs showed strong preferences for individual spacer sequences (Figure 2D, highlighted in red). For GATA4, for instance, only spacer sequence #6 (spacer length of 10 bp) significantly contributed to reporter activity, while for TEAD1 spacer sequence #1 (spacer length of 5 bp) was the only spacer sequence with strong activation. Surprisingly, our log-linear model analysis revealed that TF reporter design can be optimized regardless of the TF through the choice and positioning of the core promoter. Nevertheless, many TFs require TF-specific spacer lengths or spacer sequences for efficient activation, underscoring the importance of systematic reporter design optimization. Example 7 Correlating reporter activities with TF abundance across cell lines. After identifying the features facilitating high reporter activity, we showed characterization of the TF specificity of each reporter. One line of evidence for such specificity would be if the activity of a reporter correlates positively with the abundance of the corresponding TF across the nine tested cell lines. Therefore, we generated mRNA-seq data for mNPCs, mESCs, and HEPG2 cells and collected publicly available mRNA-seq data for the other six probed cell lines (see Methods). We then conducted a transcript abundance correlation analysis for 37 TFs that showed sufficient variation in expression level across the cell lines (Figure 3A). For some TFs (e.g., POU5F1::SOX2, HNF1A) the activities of all reporters significantly correlated with TF transcript abundance. For other TFs (e.g., IRX3), none of the reporters had a significant correlation. While this could indicate that the reporters for these TFs lack specificity, it is also possible that the activity of those TFs is controlled primarily by intracellular signaling or by certain co-factors; alternatively, their protein abundance is not reliably predicted by their mRNA level. Example 8. Using expression correlation to identify optimal reporters. For most TFs only a subset of reporters significantly correlated with TF transcript abundance (e.g., GATA4, TFCP2L1, GATA1; Figure 3A, highlighted in red). For example, one GATA4 reporter design with spacer sequence #4 was not active in any cell type, but the same design (i.e., the same core promoter, promoter distance and spacer length) with spacer sequence #6 was more active in cell types where GATA4 is expressed (HEPG2, mESC; Figure 3B). Indeed, GATA4 reporters with spacer sequence #6 almost exclusively displayed activities that significantly correlated with GATA4 transcript abundance (Figure S4A), suggesting that this spacer sequence renders GATA4 reporters GATA4-specific. In line with these findings, spacer sequence #6 was also identified as the most important feature in the log-linear model for GATA4 (note that this model was fit in HEPG2, Figure 2D). Additional examples of design-dependent TF specificity. GATA1 reporters were more GATA1-specific (i.e., activity only in K562) with a 10 bp rather than a 21 bp promoter distance (Figure 3C, S4B, 2D). The latter displayed activity in GATA1-lacking cell types, possibly because these reporters respond to other GATAs (e.g., GATA3 in MCF7 or GATA4 in HEPG2). TFCP2L1 reporters give another example of design- dependent TF specificity. We found that a TFCP2L1 reporter with a 10 bp spacer length (spacer sequence #4) was predominantly active in the cell line where TFCP2L1 is highly expressed (mESC), while the same reporter with a 5 bp spacer length (spacer sequence #1) was also highly active in other cell types (Figure 3D). Indeed, all TFCP2L1 reporters with spacer sequence #4 and #5 (both 10 bp) displayed activities that significantly correlated with TFCP2L1 transcript abundance (Figure S4C). Activities of TFCP2L1 reporters in TFCP2L1-lacking cell types might be explained by response to GRHL1, which is a TF with a highly similar binding motif (Figure S1A), but a distinct expression pattern (GRHL1 is lowly expressed in all nine cell lines). Together, these findings highlight that fine-tuning the reporter design can substantially improve the specificity, even for TFs with highly similar TFBSs. Example 9. Experimental design of pathway perturbations. Many TFs are known to depend on specific stimuli or upstream signaling events for their activity. To further test the responsiveness of the reporters, we therefore applied a total of 23 different pathway inhibitors, ligands, drugs and culture conditions that are known to influence the activity of at least one of the TFs (Figure 4A, Table S3). For each perturbation we chose one cell type that was most likely responsive to this stimulus. Altogether, we expected these perturbations to activate 27 TFs and suppress 9 TFs within our set of 86 TFs. Examples of strong responses. For some of those TFs (e.g., HSF1 upon heat shock, TCF7 upon CHIR- 99021 (WNT activator) removal), we saw robust responses across almost all reporter designs (Figure 4A). The most potent TF-stimulating condition was activation of VDR (vitamin D receptor) reporters by its ligand calcitriol. In U2OS cells this yielded activation levels up to 180-fold (Figure 4A, B). Other strong reporter responses were also achieved by stimulating the heat shock-responsive HSF1 at 43^C (Figure 4C); the oxidative stress response factor NFE2L2 by treatment with hydroquinone (Figure 4D); the bile acid receptor NR1H4 by the bile acid CDCA (Figure 4E); the c-AMP responsive TF CREB1 by c-AMP activator forskolin (Figure 4F); and STAT3 by removal of JAK-STAT activator LIF (Figure 4G). Example 10. Variation in responses between reporter designs. Overall, there was a marked variation in the strength of the response between reporters of the same TF. The strength of the responses in the examples above strongly depended on the core promoter (VDR, HSF1, NR1H4), or the spacer sequences (NFE2L2, STAT3), which is in line with the findings of the log-linear model (Figure 2D). For other TFs (e.g., AHR::ARNT, NR4A2), only a few the reporters showed a clear response (fold-change > 2). Example 11. Altered TF expression: experimental design and interpretation. Finally, as a more direct method of perturbing TF activity, we tested the response of all reporters to transient knockdown (KD), protein degradation, or overexpression of individual TFs. Among our set of 86 TFs, we knocked down 16 TFs in mESCs and 28 TFs in HEPG2 cells by RNA interference. For SOX2 and POU5F1 we additionally used degron-mediated depletion in mESCs {Maresca, 2023, 37691488}. Moreover, to evaluate specificity and off-target responses of the TF reporters, we also included nine KDs in mESCs and 13 KDs in HEPG2 cells of related TFs that have similar TFBSs as our candidate TFs. Finally, we overexpressed four TFs that are not naturally expressed in mESCs. The scale of these experiments prohibited the verification of the KD or overexpression efficiency for each individual TF by Western blotting or mass-spectrometry. For this reason, a lack of a response of reporters to the perturbation of their cognate TF does not necessarily imply that the reporters lack specificity; it is possible that we simply failed to alter the level of the TF sufficiently. Conversely, however, a strong response of reporters to the perturbation of the cognate TF can be regarded as evidence of specificity. The results of these experiments are summarized in Figure 5A. Approximately one-third of all KD-targeted TFs showed a strong decrease in reporter activity (fold-change > 2) across the majority of reporters, although for most of these TFs the strength of the response varied substantially between reporters. Protein degradation of SOX2 strongly reduced activities of all POU5F1::SOX2 reporters, and a subset of SOX2 reporters. Similarly, POU5F1 degradation decreased activity of a subset of POU5F1 reporters and all POU5F1::SOX2 reporters. Overexpression of FOXA1 significantly increased the majority of the FOXA1 reporters, while FOSL1 overexpression only led to an increase in FOS::JUN, but not FOSL1 reporter activities. GATA1 and NR4A2 overexpression did not increase activities of their target reporters. Example 12. Perturbation response depends on reporter design. Conformingly, we found the reporter responses were often dependent on the precise design. While all POU5F1::SOX2 reporters strongly reduced their activity upon POU5F1 degradation (Figure 5B), there was a marked difference in response to SOX2 degradation, with POU5F1::SOX2 reporters with a 10 bp spacer length showing stronger responses (Figure 5C). Similarly, we found that PAX6 reporters with reduced activity upon PAX6 KD mostly had 10 bp spacers, while the published PAX6 reporters did not show any response (Figure 5D). Other examples of design-dependent responses are highlighted in Figures 5E-G. Overall, of the 44 TFs that were targeted by KD, 34 had at least one reporter with a more than two-fold reduction in activity (Figure 5A). Example 13. Probing reporter cross-reactivity. Many TFs belong to families that share highly similar binding motifs. Therefore, to test for off-target responses we also evaluated responses upon perturbations of other members within the same TF family. In total, we investigated 50 pathway perturbations and 87 TF perturbations that could potentially result in cross-reactivity due to TFBS similarity of the target TF and another TF. Of these, reporters for around 20 TFs showed substantial off-target responses (Figure S5A). A striking example of high selectivity is NR1H4 reporters, which have a TFBS that is highly similar to other nuclear receptor TFBSs (Figure S1A); nevertheless, they strongly responded only to bile acid stimulation (CDCA) and not to any other nuclear receptor stimulation (Figure S5B). Off-target responses often varied in magnitude depending on the reporter design. For example, all GRHL1 reporters had a reduced activity upon KD of GRHL1, while only GRHL1 reporters with spacer sequences #1-4 additionally responded to TFCP2L1 KD. We found that CLOCK reporters, for which we only probed published reporter designs, reduced their activity by approximately twofold upon removal of LIF; these reporters carry a repeat sequence that significantly matches the STAT3 motif, possibly explaining the erroneous response to LIF. Example 14. Assigning confidence levels to TF reporters. Using the abundance of the cell type-specific activities and the perturbation data described above, we aimed to integrate all data to identify the most optimal reporters for each TF. To do so, we assigned confidence levels to each individual reporter, ranging from 0 (low confidence) to 4 (very high confidence), based on the criteria summarized in Figure 6A. For level 4, we required reporters to be responsive to a relevant stimulus, display activities that correlate with the abundance of the TF across the tested cell lines, and show a substantial response to depletion or overexpression of the TF, without responding to off-target perturbations. Figure 6B illustrates how each of the confidence level criteria contributes to the confidence scores of all STAT3 reporters. Out of 51 reporters, 21 had a confidence level of 0 because they did not display any significant activity, and also did not respond to LIF removal. Only six reporters were assigned level 4 because they displayed high activity in basal conditions, correlated with STAT3 abundance, strongly responded to LIF removal, and did not show an off-target response to STAT1 KD (Figure S5A). As established previously (Figure 2B-D, 4G), these high-confidence reporters are characterized by a 10 bp spacer sequence #6, but also include published reporters. We generated similar reporter confidence heatmaps for all 86 TFs (Figure S6). Finally, for TFs with reporters with a confidence level of 2 or higher, we selected a single "prime" reporter, based on the confidence scores and – in case of ties – additional performance criteria (Table S4; see Methods). For a total of 62 TFs, this yielded a prime reporter with confidence level 4 (11 TFs), 3 (33 TFs), or 2 (18 TFs). We emphasize that level 2 means that the reporter is significantly active and that there is evidence for TF specificity, and thus such a reporter is likely to provide meaningful information. While most prime reporters feature a minCMV or minHBG core promoter (46 / 62), the spacer sequences are distributed relatively evenly across prime reporters (#1 [5 bp]: 11, #2 [5 bp]: 4, #3 [5 bp]: 8, #4 [10 bp]: 13, #5 [10 bp]: 9, #6 [10 bp]: 6), highlighting their TF-specific nature. This underscores the necessity for TF-specific spacer sequence optimization. Furthermore, the set of 62 prime reporters consists of 51 synthetic reporters and 11 published reporters. Notably, of the 38 TFs in the prime reporter set for which we probed both synthetic and published reporters, synthetic reporters outperformed the published reporters for 32 TFs (84%), while published reporters outperformed the synthetic reporters for only 6 TFs (Figure 6C). For 19 TFs, the synthetic prime reporters even scored at least one confidence level higher than the published reporters. This demonstrates the value of systematic optimization. Additionally, the prime set includes 19 TFs for which we did not test published reporters, primarily because they were not available, (Figure 6D), and five published reporters for which we did not test synthetic designs (Figure 6E, Table S4). Example 15. Prime reporters typically require high-affinity BSs. As a final characterization of the synthetic prime reporters, we checked whether their activities are dependent on full integrity of the respective TF motifs. Indeed, of the 49 synthetic prime reporters for which we had matched mutated controls, 39 decreased their activity upon mutation of a two to three nucleotides in the TFBS by at least 2-fold, and up to 450-fold (Figure 6F). Prime reporters also had a significantly increased sensitivity to these mutations compared to reporters of the same TF with a lower confidence level (Figure S7). These strong responses to minimal alterations in the TFBS reaffirm the TF specificity of the identified prime reporters. We note, that the remaining 10 reporters (of which four are confidence level 4, and three are confidence level 3) should not be rejected based on this result, because some TFs might be able to activate a promoter stronger through low- or medium-affinity TFBSs than through high-affinity TFBSs {Trauernicht, 2023, 37650627}. Example 16. Specific TF activity detection across nine cell lines. Having identified the prime reporters for 62 TFs, we reassessed the activities of those TFs across all tested conditions. We first focused on the steady-state activities across the nine probed cell lines (Figure S8A). To be able to compare reporters of different strengths with each other, we rescaled the reporter activities separately per TF. This allowed us to identify cell type-specificities of TFs and to identify clusters of TFs with similar activity patterns (Figure 7A, S8B). We found a large number of TFs displaying distinct cell type- specific activities, which match their known biological functions (e.g., HNF4A in HEPG2, ESR1 in MCF7, or SOX2 in mESC; Figure 7A, S8C). The prime reporters also discriminate TFs with highly similar TFBSs, like GATA1 / GATA4, TFCP2L1 / GRHL1, EGR1 / KLF4, or a variety of nuclear receptor TFs. Thus, our set of 62 prime reporters can identify TF activity differences between cell types, and highlight functional similarities between TFs. Exploring TF-TF communications. Besides steady-state activities, we can use the prime reporters to monitor 62 TF activities across all tested 98 TF perturbation conditions. This offers an immense resource to uncover novel roles for TFs and interactions between TFs. To showcase the responses of prime reporter activities to the direct perturbations of TFs, we quantified prime reporter responses upon all KDs in HEPG2 cells with a strong effect on their direct target (n = 21). We found a large number of TFs that change their activity upon downregulation of another TF (e.g., PAX6 activation upon HNF1A KD, Figure S8D). These data offer a large resource to explore potential TF-TF communications. We then focused our analysis on perturbations in mESCs that challenge the pluripotency network (Figure 7B). Interestingly, besides altering the activity of its cognate TF, most perturbations led to strong secondary TF activity changes. For instance, we found that degradation of key pluripotency factors POU5F1 and SOX2 substantially reduced the activity of other pluripotency TFs like STAT3, TFCP2L1 or KLF4, highlighting their core function in the pluripotency network {Dunn, 2014, 24904165}. Furthermore, removal of JAK-STAT activator LIF led to strong inactivation of its target STAT3, but also decreased the activity of WNT target TCF7 as well as many other pluripotency TFs like SOX2 or KLF4 (Figure 7B). This suggests that LIF is needed to maintain pluripotency, potentially through crosstalk with the WNT signaling pathway. Similarly, we found that MEK-ERK inhibitor PD (PD0325901) crosstalks with WNT signaling, and WNT activator CH (CHIR-99021) with MEK-ERK signaling, suggesting that these signaling pathways reinforce each other and have redundant targets, as has been discussed before {Dunn, 2014, 24904165; Hackett, 2014, 25280218}. Besides this, we found that addition of serum increased the activity of pluripotency TFs such as POU5F1::SOX2, reinforcing the pluripotency network. Together, this analysis shows that multiplexed TF activity detection using prime reporters has the potential to link targeted signaling pathway perturbations to functional changes in TF activity to discover signaling pathway interdependencies. Example 17 TF reporter tandem arrays To be able to detect activities of multiple TFs simultaneously (at least 5 and up to 20), we have successfully integrated a tandem array containing reporters for 5 different TFs into the genome of mouse embryonic stem cells (mESCs). Reporter activities retrieved from this array correlate well with activities measured from individual TF reporter plasmids, and target TF reporters are responding to stimuli (e.g., STAT3 response to LIF) without affecting the activity of other TFs in the array. We are currently assembling a tandem array carrying 20 TF reporters. Using the engineered TF reporter tandem array cells, activities of many TFs (at least 5, likely up to 20) can be tracked simultaneously from the same cell. TF reporter tandem arrays could be used to track TF activities during differentiation (e.g., gastrulation of mESCs) to track TF activities over time from single cells. This may help to understand how TFs are responding to stimuli and how TFs orchestrate changes in gene regulation during differentiation. TF activities from engineered TF reporter tandem arrays can also be tracked along with the transcriptome (using scRNA-seq). mESCs carrying a TF reporter tandem array may also be used for blastocyst injection to generate transgenic mice. This can be combined with a wide variety of disease models (cancer, obesity, brain disorders, genetic disorders, etc.) to reveal how the activity of TFs is altered in the diseased tissues. An example of this is a “nuclear receptor reporter mouse” could be generated that carries in its genome ~10 different barcoded reporters for TFs of the nuclear hormone receptor class. This mouse could be used to determine the efficacy and specificity of new (and known) steroid- and related drugs throughout all tissues. Example 18: High-throughput screens using prime TF reporters TF reporter activities may be measured from cells cultured in 96- or 384-well plate formats, provided that the TF reporter library is efficiently delivered into the cellular model of interest. The reporter library may be introduced via any suitable vector system, including but not limited to synthetic vectors, bacterial plasmids, viral vectors (e.g., lentiviral, AAV), or transposon-based systems. For high-throughput applications, a reduced-complexity library comprising only the prime TF reporters may be employed. Each prime TF reporter may be represented by five distinct barcodes, along with a set of negative control reporters, resulting in a total library complexity of approximately 655 barcodes. This reduced complexity enables the recovery of high-quality data from small cell numbers, such as approximately 8,000 cells per well in 384- well formats. Following delivery of the TF reporters, cells may be seeded into 96- or 384-well plates and subjected to high-throughput perturbagen screening. Perturbations may include treatment with small molecule libraries, genetic modifications (e.g., CRISPR-based screens), or environmental stressors. The inventors have demonstrated that TF activity profiles recovered from U2OS cells in 96-well format show high correlation with profiles obtained from larger well formats such as 12-well plates (see Figure 15). These results establish that prime TF reporter libraries enable robust and reproducible measurement of TF activity at high throughput and in miniaturized assay formats, supporting their application in automated screening workflows. Example 19: Simultaneous measurement of TF reporter activity and the transcriptome The method enables concurrent measurement of TF reporter activity and global transcriptome expression from the same cellular sample. Since the quantification of TF reporter activity is based on sequencing barcoded transcripts derived from total RNA, the extracted RNA may also be used to generate standard whole-transcriptome RNA sequencing (RNA-seq) libraries. This dual measurement may be achieved by dividing the extracted total RNA into two separate aliquots. One aliquot may be subjected to reverse transcription and amplification of the reporter-derived barcodes to quantify TF activity. The second aliquot may be processed using conventional RNA-seq library preparation protocols to capture the endogenous transcriptome. The inventors have routinely applied this approach, enabling direct comparison of TF activity readouts with corresponding gene expression profiles from the same population of cells. This simultaneous measurement supports integrative analyses of transcriptional regulation and downstream gene expression responses within a unified experimental workflow. Example 20: Measuring nascent TF reporter activities The inventors have adapted the method to measure nascent TF reporter activities, enabling the detection of ongoing, active transcription rather than accumulated transcript levels. This provides a higher temporal resolution in monitoring transcriptional responses. To achieve this, cells are labeled with a nucleotide analog, 4-thiouridine (4sU), for a short pulse duration (e.g., 15 minutes) prior to harvesting. The 4sU is incorporated into newly synthesized RNA transcripts during the labeling window. Following RNA extraction, 4sU-labeled transcripts are selectively biotinylated and captured via streptavidin pulldown. This enrichment step isolates nascent RNA, which is then subjected to reverse transcription and amplification of the barcodes corresponding to TF reporters. The inventors demonstrated in HEK293 cells that this nascent RNA protocol enables robust detection of TF reporter activity and provides enhanced temporal resolution. For example, it allowed the observation of a transient decrease in HSF1 activity during the early recovery phase after heat shock, a change that was not detectable using total RNA due to the persistence of previously transcribed barcodes (Figure 16). Example 21: Computational pipeline provides quantification of TF activities directly from sequencing data The inventors developed a novel computational pipeline for easy quantification of TF activities directly from sequencing data which also provides a statistical framework and exploratory visualizations to guidecomparative analyses between tested conditions. The computational pipeline calculates the activity of theTF reporters and performs differential analysis to statistically test the effect of the perturbation on the activity of the reporters. The pipeline is based on an MPRA analysis framework, enabling sensitive reporter activity computation and powerful statistical analysis. Additionally, we include several quality control (QC) plots that provide in-depth information about the data quality. The pipeline also performs statistical analyses to compute changes in reporter activities between two conditions. As an example, we compare the activity of the reporters on U2OS cells under calcitriol treatment versus DSMO treatment using the pipeline. The pipeline compares two conditions at a time: one reference / control condition, and one contrast / treatment condition. Multiple tests can be made by inputting multiple contrast / treatment conditions. The required input parameters for the pipeline are the following fastq files: (1) high-throughput sequencing data for TF reporters in a control sample, (2) high-throughput sequencing data for TF reporters in a contrast / perturbed sample, and optionally (3) the pDNA counts or gDNA counts. gDNA counts may be used for lentiviral data. pDNA sequencing data does not necessarily need to be generated for every single experiment. Previous pDNA sequencing data can be used if the same plasmid library was sequenced in an earlier experiment. The input data parameter is as follows: ^ INPUT_DATA - Considering that the input data information is configured in the “yaml” format with the fields: INPUT_DATA: <sample_name>: is_pdna: <True / False> condition: <condition_of_this_sample> fastq: ^ <path_to_replicate1_fastq_file> ^ <path_to_replicate2_fastq_file> All of the information in the INPUT_DATA parameter is transformed into a “design” dataframe, which is used mainly for the differential activity analysis, but also for other steps of the pipeline, such as for QC plots. Example dataframe: Sample Replicate pDNA Treatment <sample_name> <sample_name>_1 <True / False> <condition_of_this_sample> <sample_name> <sample_name>_2 <True / False> <condition_of_this_sample> Further pre-processing of the data includes: • The <sample_name> is processed as a string factor for the output plots. • The value of “is_pdna” is processed as a boolean value, informing the pipeline whether this sample is pDNA or not. In some steps of the pipeline, pDNA and cDNA samples are treated differently (e.g., for calling BCalm for the differential activity analysis, for estimating the bleedthrough, and to calculate the cDNA / pDNA ratio). The Boolean value of this parameter is used to guide the pipeline on which type of processing should be done for this sample. • The value of “condition” is processed as a string factor for generating the output plots and output results. Further steps include: • Extracting the barcode counts from the fastq sequencing files • Clustering the barcodes based on sequence similarity using Starcode. • Matching the barcode sequences to TFs using a barcode annotation file. The parameters are as follows: ^ PVALUE_THRESHOLD: 0.01 - This parameter sets the threshold for the adjusted p-value to assign the TFs as responsive to the treatment. ^ BARCODE_DOWNSTREAM_SEQUENCE: CATCGTCGCATCCAAGAGGCTAGCTAACTA - This parameter sets the sequence downstream of the barcode. To identify the barcode of the plasmid, the pipeline first identifies this sequence (allowing by default 3 mismatches, see “MAX_MISMATCH_DOWNSTREAM_SEQ” parameter) that is known to be right after the barcode in the plasmid map. Then, it assumes that the 12 nucleotides (by default, see “BARCODE_LENGTH” parameter) upstream of this sequence represent the barcode of the plasmid. The default sequence was defined because the default plasmids of the library contain this sequence downstream of the barcode. Some implementations may change this sequence, so therefore this is a parameter. It’s also important for us to have this as a parameter because we are also use lentiviral constructs; for them, the downstream sequence is different from the default. For the lentiviral backbone, the downstream sequence is: CATCGTCGCATCCAAGAGTGCCACCATGCC. ^ BARCODE_LENGTH: 12 - This parameter sets the expected length of the barcode, for a correct barcode identification (see explanation above). This value was defined based on the standard plasmids. Some implementations with different constructs might use bigger or smaller barcodes, therefore this as a parameter. ^ MAX_MISMATCH_DOWNSTREAM_SEQ: 3 - This parameter sets the maximum mismatch that is allowed during the identification of the downstream sequence. ^ BARCODE_ANNOTATION_FILE: misc / bc_annotation_prime.csv - This parameter points to the file containing the information of the barcode, i.e., which TF each barcode corresponds to. This information is necessary for connecting the barcode counts to the TF activity in the pipeline. Since this information is provided by us as an external file, we decided to have this as a parameter. ^ EXPECTED_PDNA_COUNTS: misc / expected_pDNA_counts.txt - This parameter points to the file containing the pDNA counts that we, in our lab, get in our transfections of the library. This information is used as a quality-check step, where the user can see if their pDNA sequencing is similar to what we get in our lab. Since this information is provided by us as an external file, we decided to have this as a parameter. The pipeline outputs QC plots, computes the activity of the TF reporters, and performs comparative analyses between conditions. To validate that the measured TF activities are robust, the pipeline generates several QC plots. A key aspect of this validation is estimating the technical noise of the experiment, which can be assessed by reviewing the correlations between the five barcodes within a single experiment. An example of this is shown in Figure 12A. A Pearson correlation of ≥ 0.9 is typically targeted. To evaluate reproducibility between biological replicates, it is crucial to assess the strength of correlations in TF activity. Figure 12B provides an example of this. Again, a Pearson correlation of ≥ 0.9 is typically targeted. In plasmid-based MPRAs, a portion of cDNA reads can originate unintentionally from the plasmids during cDNA amplification. This phenomenon, which we call pDNA bleedthrough, can skew cDNA counts and, consequently, distort the measured activity of the reporters. The pipeline performs a bleedthrough estimation step as another quality control by correlating the pDNA counts with the cDNA counts for every sample, as shown in Figure 12C. Ideally, we aim at less than 10% of bleedthrough but accept up to 25%. We consider a bleedthrough higher than 25% as too excessive pDNA contamination in the cDNA counts Next, the pipeline calculates the activity of each TF in each condition and performs a comparative analysis to statistically assess changes in TF activity between the two compared conditions. The model performs differential analyses by fitting, for each individual barcode, a linear model on the input data. The model performs statistical analyses to compute changes in reporter activities between two conditions by running a re-implementation of the Limma framework {Ritchie, 2015, 25605792}. P-values are extracted from the statistical model and corrected for multiple comparisons using any method known in the art, for example, Bonferroni correction or False Discovery Rate (FDR) correction. The main results are stored in a tabular file (Figure 13A), containing both the activity of the TFs and the results of the comparative analysis. By default, the pipeline applies a threshold of 0.01 on the adjusted p-values to call for significant changes. The pipeline also offers two different ways of visualizing the results: a volcano plot showing the adjusted p- values and the fold-change of each TF (Figure 13B), and a lollipop plot showing the activity of the TFs in both conditions, highlighting the significant changes (Figure 13C). To assess the robustness of fold-changes across replicates, the pipeline generates plots of normalized read counts (RPM) for all barcodes linked to each transcription factor (TF) across all samples and replicates (Figure 14). Statistically up- or downregulated TFs are highlighted, helping to evaluate the consistency of fold-changes across both replicates and barcodes. The pipeline generates critical quality check plots to assess the quality of the data. Without this insight, it is impossible to assess if the experiments have succeeded. These plots are: • Plasmid DNA bleedthrough estimation plots. In these plots, we correlate the counts in the plasmid DNA with the counts of the counts derived from the RNA. This is done using our set of negative control reporters. We use the slope of the correlation to estimate the level of bleedthrough. This step is very important to assess the quality of the data. As far as we know, no other computational tools implement this QC step. • Correlation plots between barcode counts to estimate the level of technical noise. • Correlation plots between biological replicates to estimate the level of reproducibility. • Estimation of sequencing depth by plotting the number of reads per barcode per sample and the total number of reads per sample. • Plotting the activity of all reporters per TF for all biological replicates. These plots are very useful to determine if the perturbation has reproducible effects across all biological replicates and barcodes on each individual TF. • All these QC plots are done using barcode count estimations and reporter activity estimations (RNA / pDNA). These plots may be generated using data from any part of the pipeline described above and may be generated independently from the statistical model. MATERIALS & METHODS TF reporter library design The 86 TFs were chosen based on motif quality, motif uniqueness, expression patterns, and perturbation opportunities (Table S1). For each TF, consensus TFBSs and mutated TFBSs were created by mutating two to three conserved bases (Table S1). In addition to the mutated TFBSs, three random TFBS-devoid (TF-neg) 11 bp sequences were included as negative controls. The absence of TFBSs was confirmed in the mutated and random sequences using FIMO (p-value threshold 1e-4) {Grant, 2011, 21330290}. Synthetic TF reporters were then created by placing four adjacent copies of the consensus, mutated, or negative TFBS. The four TFBSs were separated by in silico-designed TFBS-devoid spacer sequences with lengths of 5 or 10 bp. In total, three different spacer sequences were generated per spacer length. To do so, random sequences with a GC content of 40-60% were generated (sim.DNAseq function in R from package SimRAD (version 0.96)). These sequences were combined with 3 bp of the left and right side of all TFBSs and then scanned using FIMO (Figure S1C). For the two spacer lengths (5 and 10 bp), nine sequences with the fewest predicted significant TFBSs were selected and placed in between the TFBSs (three different spacer sequences per reporter, times the three spacer sequences). A similar approach was taken to generate three 10 or 21 bp spacer sequences in front of the core promoter. One of three core promoter sequences, minCMV {Li, 2009, 18701910}, minHBG {Collis, 1990, 2295312}, or minP (derived from pGL4 (Promega, Madison, WI, USA)), was placed downstream of the TFBSs and spacer sequences, followed by a S1 Illumina adapter sequence and a unique 12-13 bp random barcode sequence (each unique construct was linked to five to eight different barcodes). All generated random barcodes had a Levenshtein distance of at least three with respect to one another and barcodes with an unbalanced GC ratio were removed (create.dnabarcodes function from the R package DNABarcodes (version 1.2.2) {Buschmann, 2013, 25638815}). For 64 TFs we also included published reporter sequences. The response element sequences were retrieved from three different sources {O’Connell, 2016, 27211859; Romanov, 2008, 18297081; Promega pGL4.XX – see Table S1}. For some TFs, multiple TF response elements were included (see Table S1 for all included published TF response elements). Again, each published response element was placed 10 or 21 bp upstream of a minP or minCMV core promoter. The same spacer sequence as for the synthetic TF reporters was used upstream of the core promoter. Several other controls were included in the design. First, to estimate the effect of the TFBSs alone, TF reporters with a TFBS-devoid core promoter were designed. This promoter was previously shown to be inactive {Trauernicht, 2023, 37650627}. For each TF, this TFBS-devoid core promoter was attached to one reporter design only (background #4, promoter distance 21 bp). Second, two different positive controls were included to benchmark the expression levels of the synthetic TF reporters: 1) a 183-bp region of the hPGK promoter, and 2) 120 (40 for each of the three core promoters minP, minCMV, and minHBG) 100-bp regions of Klf2 gene enhancers with known activity in reporter assays {Martinez-Ara, 2022, 35594855}. Each of these control reporters were also linked to five to eight different barcodes. All reporter sequences were completed with 18 bp primer adapter sequences (that were also scanned using FIMO) in both flanks for cloning purposes. The resulting sequence pool had a total length of on average 202 bp (at least 148 bp up to 297 bp) and was ordered as oligonucleotide library from Twist Biosciences. Cloning of the TF reporter library The vector backbone was constructed as mentioned previously {Trauernicht, 2023, 37650627}. The oligonucleotide library was resuspended in TE buffer (Invitrogen) to a final concentration of 20 ng / µl.10 ng of the oligonucleotide library was then PCR amplified (1’ 95°C, 6x(15’’ 95°C, 15’’ 57°C, 15’’ 72°C), 1’ 72°C) by MyTaq Red mix (Bioline) using primers that add overhangs with EcoRI (MT024, Table S5) or NheI (MT025) restriction enzyme sites. The PCR product was then purified using CleanPCR beads (#CPCR, CleanNA) at 1.8:1 beads:sample ratio, digested with EcoRI-HF (#R3101, NEB) and NheI-HF (#3131, NEB) by incubating the PCR product at 37°C for 1 h, and then again bead purified as before.1 µg of the entry vector was also digested with EcoRI-HF and NheI-HF and the linearized product was purified from a 2% agarose gel using PCR Isolate II PCR and Gel Kit (Bioline). The digested and purified reporter pool was then ligated into 80 ng of the linearized entry vector using Takara ligation kit v1.0 (#6021; Takara) at a 1:3 (vector:insert) ratio. The ligation mix was then bead purified as before and transformed into MegaX DH10B T1R Electrocomp™ Cells (Invitrogen) using 1 µl of the ligation mix. The library complexity was estimated from plated serial dilutions of the transformed cells to be ~300,000 colony forming units. Transformed cells were transferred to 200 ml standard Luria Broth (LB) plus kanamycin (50μg / ml), grown overnight and purified using a Maxi plasmid purification kit (#12162; Qiagen). Cell culture MCF7 (#HTB-22, ATCC), HEK293 (#CRL-1573, ATCC), and A549 (#CCL-185, ATCC) cells were cultured in DMEM medium (#41966029, Gibco), K562 (#CCL-243, ATCC) in RPMI 1640 medium (#11875093, Gibco), U2OS (#HTB-96, ATCC) and HCT116 (#CCL-247, ATCC) in McCoy's 5a medium (#26600023, Gibco) and HEPG2 (#HB-8065, ATCC) in MEM (#11095080, Gibco). All media were supplemented with 10% fetal bovine serum (FBS, Sigma). mESC (E14TG2a, #CRL-1821, ATCC) were cultured in 2i+LIF culturing media according to the 4DN protocol (https: / / data.4dnucleome.org / protocols / cb03c0c6-4ba6- 4bbe-9210-c430ee4fdb2c / ). The reagents used were neurobasal medium (#21103-049, Gibco), DMEM- F12 medium (#11320-033, Gibco), BSA (#15260-037, Gibco), N27 (#17504-044, Gibco), B2 (#17502-048, Gibco), LIF (#ESG1107, Sigma-Aldrich), CHIR-99021 (#HY-10182; MedChemExpress) and PD0325901 (#HY-10254, MedChemExpress), monothioglycerol (#M6145-25ML, Sigma) and L-Glutamine (#25030- 081, Gibco). The mNPCs used in this study were differentiated from E14TG2a mESCs and cultured in mNPC medium as mentioned previously {Peric-Hupkes, 2010, 20513434}. HEK293T (#CRL-3216, ATCC) cells used for lentivirus production were cultured in DMEM-F12 (#11320-033, Gibco) supplemented with FBS (Sigma) and L-glutamine (#25030-081, Gibco). All cells used in this study were routinely tested for mycoplasm. Reporter library transfection and pathway perturbations All cell lines except for K562 were transfected using lipofection. Per lipofection condition, 1.5x105cells were seeded in a 12-well and transfected 8 hours later by adding 1 µg of TF reporter plasmid library with 3 µl of Lipofectamine 3000 (#L3000150, ThermoFisher) in 100 µl Opti-MEM (#31985070, Gibco). mESCs were plated directly before lipofection instead of 8 hours prior and transfected using Lipofectamine 2000 (#11668027, ThermoFisher). K562 cells were electroporated using an Amaxa 2D Nucleofector. Per transfection, 1x106K562 cells were resuspended in transfection buffer (100 mM KH2PO4, 15 mM NaHCO3, 12 mM MgCl2, 8 mM ATP, 2 mM glucose (pH 7.4)) supplied with 1 µg of plasmid library and electroporated using program T-003. After nucleofection, cells were resuspended in 2 mL complete medium and plated in 6-well plates. For the signaling pathway perturbation conditions, inhibitors or activators were added to the cells directly after transfections. All inhibitors and activators used in this study are mentioned in Table S3. 24 hours after transfection, cells were harvested and resuspended in 800 µl TRIsure (#BIO-38032; Bioline) and stored at -80 °C until further use. Transfections were done at least in biological duplicates on separate days. siRNA TF knockdown experiments The TF knockdown experiments were performed in HEPG2 and mESCs. For HEPG2 cells, reverse siRNA transfections were done by mixing 20 nM siRNA with 1.5 µl Lipofectamine RNAiMAX transfection reagent (#13778075, ThermoFisher) in 100 µl Opti-MEM in 24-wells. Then, 7.5x104HEPG2 cells were added to the wells. The list of siGENOME SMARTpool siRNAs (Dharmacon) used in the screen can be found in Table S3.24h after siRNA transfection, 0.5 µg of the TF reporter plasmid library was transfected by mixing the library with 1.5 µl Lipofectamine 3000 in 50 µl Opti-MEM and adding the mix directly to the cells. For mESCs, 1.5x105cells were reverse lipofected in 12-wells using 40 nM siRNA and 3 µl Lipofectamine RNAiMAX transfection reagent (#13778075, ThermoFisher) in 200 µl Opti-MEM. All used ON-TARGETplus siRNAs (Dharmacon) are listed Table S3.24h after siRNA transfection, 1 µg of the TF reporter plasmid library was mixed with 3 µl Lipofectamine 2000 in 100 µl Opti-MEM and plated in new 12-wells. The siRNA-transfected mESCs were then collected and added to new 12-wells with the TF reporter plasmid library lipofection mix. Knockdown efficiency was evaluated by killing controls using siRNAs targeting PLK1 (#L-003290 (human), #L-040566 (mouse), Dharmacon). Non-targeting siRNAs were used as negative controls (#D-001210-01, Dharmacon).24 hours after TF reporter library plasmid transfection and 48 hours after siRNA transfection the cells were harvested as mentioned in the ‘Reporter library transfection and pathway perturbations’ section. TF overexpression experiments Lentiviral plasmids carrying doxycycline-inducible open reading frames for GATA1, FOSL1, FOXA1, NR4A2 or RFX1 and a puromycin selection cassette were a kind gift from Bart Deplancke (EPFL, Lausanne, Switzerland). To generate lentivirus, 5x105HEK293T cells were plated in 6-well plates per condition. At ~75% confluency, 1.5 µg TF ORF lentiviral plasmid was mixed with 1.125 µg psPAX2 (#12260, Addgene), 0.375 µg pMD2.G (#12259, Addgene) and 5 µl Lipofectamine 2000 in 250 µl Opti-MEM and added to the 6-wells. The medium was refreshed after 12 hours and lentivirus was collected after 48 hours from the supernatant. To transduce cells with the lentivirus, 1x105mESCs were plated in 12-wells in 500 µl 2i / LIF medium supplemented with 8.5 µg polybrene (#TR-1003, Sigma). Then, 500 µl of lentiviral supernatant was added to the cells. Medium was changed to fresh 2i / LIF medium 24 hours later and to puromycin- containing (2 µg / ml) 2i / LIF medium after 48 hours. Puromycin-resistant cells were grown and used for the subsequent TF reporter plasmid library transfection experiments. To transfect the TF reporter plasmid library, the TF ORF-carrying mESCs were pretreated for 24 hours with 2 µg / ml doxycycline (#D9891, Sigma) and then lipofected as mentioned in the ‘Reporter library transfection and pathway perturbations’ section. TF degradation experiments mESCs with FKBP-tagged POU5F1 (genetic background: V6.5 {Boija, 2018, 30449618}), SOX2 (IB10), or NANOG (E14tg2a) were generated as described in {Maresca, 2023, 37691488} and were a kindly provided by Elzo de Wit (Netherlands Cancer Institute). TF degradation was induced directly after TF reporter library transfections using 500 nM dTAG-13 (#SML2601, Sigma). Cells were harvested for RNA extraction 24h after library transfection and degradation induction. RNA extraction, reverse transcription and barcode amplification RNA extraction was done using the standard procedure according to the TRIsure protocol. After RNA extraction, 1 µg of RNA was treated with DNase I for 30 minutes (#04716728001; Roche) and subsequently treated with 1 μl 25 mM EDTA at 70 °C for 10 minutes to inactivate DNase I. cDNA synthesis was primed by addition of 1 μl gene-specific primer targeting the GFP ORF (10 µM, MT165) and 1 μl dNTPs (10 mM each) followed by incubation at 65 °C for 5 minutes. Then, the reverse transcription reaction was set up by adding 20 units RiboLock RNase inhibitor (#EO0381; ThermoFisher Scientific), 200 units of Maxima reverse transcriptase (#EP0743; ThermoFisher Scientific, 4 μl of 5x Maxima reverse transcriptase buffer and 2.5 μl of nuclease-free water. The reaction was then incubated for 30 minutes at 50 °C followed by heat-inactivation at 85 °C for 5 minutes.20 μl of cDNA were then PCR amplified (1′ 96 °C, 20x(15″ 96 °C, 15″ 60 °C, 15″ 72 °C)) in a 100 μl reaction using MyTaq Red mix and primers containing the Illumina S1 and p5 adapter (MT397) and the Illumina S2 and p7 adapter (MT164). To generate input plasmid DNA (pDNA) barcode counts that serve as normalization control, the plasmid library that was used for the transfections was linearized using EcoRI-HF and subsequently 1 ng of linearized vector was PCR amplified as before using 8 cycles. PCR products were pooled and purified by double-sided CleanPCR bead purification using beads:sample ratios of 0.6:1 followed by 1.2:1 on the supernatant. The sequencing library was then sequenced using a 75 bp single-read NextSeq High Output kit (Illumina), yielding on average ~8.8x106reads per sample, and thus on average ~248 reads per barcode. RNA-seq data generation and analysis RNA-seq data was generated for mNPCs as following.1x106mNPCs were collected on two separate days and resuspended in 600 µl RLT buffer (#79216, Qiagen). RNA was isolated using RNeasy column purification (#74104, Qiagen). Sequencing libraries were prepared using TruSeq polyA stranded mRNA library prep kit (#20020595, Illumina) and sequenced on a NovaSeq 6000 with 51 bp paired-end reads yielding 20x106reads per sample. RNA-seq data for mESCs was retrieved from {Joshi, 2015, 26637943}. Data for all other cell lines was collected from the Human Protein Atlas (https: / / www.proteinatlas.org / about / download, #25 - RNA HPA cell line gene data, The Human Protein Atlas version 23.0, Ensembl version 109). For all cell lines and all genes, transcripts per million (TPM) were calculated and then normalized to nTPM using Trimmed mean of M values {Robinson, 2010, 20196867} to allow for between-sample comparisons. To compute correlations between TF reporter activity and TF expression, only TFs with differences in expression across cell lines were included (nTPM > 8 in at least one cell line, nTPM < 1 in at least one cell line). Additionally, TFs that were not active in any cell line (reporter activity (log2) < 0.75) were excluded. Several TFs were included in the analysis even though they did not pass these filters (STAT3, SP1, TEAD1, NFKB1, ZFX, NR4A1). In case of heterodimeric TFs (e.g., POU5F1::SOX2), we considered in each cell line the nTPM value of the TF with the lowest abundance, since this TF is the limiting factor of the heterodimer. Reporter activity computation and normalizations Raw barcode counts were clustered using starcode {Zorita, 2015, 25638815} using a maximum Levenshtein distance of 1. Next, clustered barcode counts were normalized by library size. To be more precise, the clustered barcode counts were divided by the total sum of all barcode counts per sample per million. From these normalized barcode counts activities were computed by dividing the cDNA barcode counts by the plasmid DNA barcode counts. The activities were normalized by dividing the activities by the median of the activities of the scrambled BS reporters per core promoter and sample. Normalized activities were then averaged over the different barcodes and finally over the independent replicates. Log-linear model of reporter activities To explore the impact of the reporter design on the reporter activity, a log-linear model was fit using the following equation. ^^^^^^^^^^^^^ ^^^^^^^^^^,^^^^^^^^^^ ~ core promoter + promoter distance + spacer length: spacer sequenceThe reporter activities were fit for each TF in three different conditions where the TF is a) expressed highest, or b) stimulated or overexpressed (if data available). We reasoned that these conditions would represent the most TF-specific conditions. The condition with the best model performance was chosen as representative model for the TF and is displayed in Figure 2. See Table S2 for chosen reference conditions. All input features in the model were used as categorical variables. Models were fit using the lm function in R from the stats package (version 3.6.2). Reporter confidence level and reporter score computation To evaluate the performance of each individual TF reporter, reporter confidence levels were computed as mentioned in the Results section. In case more than one perturbation condition was tested for a TF, the perturbation with the strongest average reporter activity fold-change was selected. TF abundance correlation was only taken into consideration for TFs that were included in the TF abundance correlation analysis (see 3A, “RNA-seq data generation and analysis” section). Moreover, to rank reporters within a confidence level, a reporter quality score was computed as follows. where =>?@^^7refers to the correlation of the reporter activities with the TF transcript abundance across the nine tested cell lines. TF tandem arrays TF reporters are assembled next to each other in a tandem array (Figure 8). Each individual TF reporter in this array contains the following elements: a TF response element (containing four identical TFBSs and optimized spacer sequences; ~50-100 bp in total depending on the length of the TFBS), a minCMV core promoter (56 bp), a short transcription unit (~100 bp) containing a unique DNA barcode and primer binding sites to amplify the transcription unit including the DNA barcode, a strong transcription termination sequence (SV40 PolyA sequence; 100 bp). To prevent cross-talk, adjacent TF reporters are separated from each other with an insulator sequence (based on PMID 24098520; 237 bp). Individual TF reporters are ordered as DNA fragments and assembled into a plasmid backbone in one step using Golden Gate assembly. The array can then be integrated into a single safe-harbor locus in the genome of any cell line of interest using recombination-mediated cassette exchange or CRISPR-Cas9.

[0008] Software and algorithms for the computer-implemented method Software and algorithms BioPython Open https: / / biopython Bioinformatics .org / Foundation ggplot2 v3.4.1 Wickham et al. https: / / ggplot2.ti {Wickham, dyverse.org / 2009} stringr v1.5.1 Wickham et al. https: / / stringr.tid {Wickham, yverse.org 2009} ggpubr v0.6.0 Alboukadel https: / / rpkgs.dat Kassambara anovia.com / ggp ubr / ggbeeswarm v0.7.2 Erik Clarke https: / / github.co m / eclarke / ggbee swarm ggnewscale v0.5.0 Elio Campitelli https: / / eliocamp. github.io / ggnew scale / ggrepel v0.9.6 Kamil https: / / ggrepel.sl Slowikowski owkow.com dplyr v2.5.0 Wickham et al. https: / / dplyr.tidy {Wickham, verse.org / 2019} optparse v1.7.5 Trevor L. Davis https: / / trevorldav is.com / R / optpar se / Starcode Zorita et al. http: / / doi.org / 10. {Zorita, 2015, 1093 / bioinforma 25638815} tics / btv053 BCalm Keukeleire et al. https: / / doi.org / 1 {Keukeleire, 0.1186 / s12859- 2025, 025-06065-9 39948460} BioRender BioRender https: / / biorender .com Image Lab Software Bio-Rad https: / / www.bio- rad.com / en- nl / product / image -lab-software

Claims

Claims 1. An isolated nucleic acid composition that encodes a synthetic reporter capable of binding to a transcription factor, said composition comprising a. Between four and eight copies of a consensus transcription factor binding site (TFBSs). b. Between three and seven TFBSs-devoid spacer sequences between the TFBSs mentioned under a). c. One TFBS-devoid spacer sequence upstream of a core promoter d. A core promoter. e. A unique barcode sequence and / or an open reading frame encoding a reporter protein.

2. The isolated nucleic acid composition of claim 1 wherein the features b, c, and d, have been determined according to the formula: ^^^^^^^^^^^^^ ^^^^^^^^^^,^^^^^^^^^^ ~ core promoter + promoter distance + spacer length: spacer sequence3. The isolated nucleic acid composition according to claim 1 and 2 wherein the consensus TFBS are selected from the list of SEQ ID NO: 1 to SEQ ID NO: 86:

4. The isolated nucleic acid composition according to claims 1 to 3 wherein the length of the three to seven spacer sequences between the TFBSs is either 5 or 10 bp.

5. The isolated nucleic acid composition according to claims 1 to 4 wherein the spacer sequences between the TFBSs are selected from the list according to SEQ ID NO 87 to SEQ ID NO:

146.

6. The isolated nucleic acid sequence composition according to any of the claims 1 to 5 wherein the length of the spacer sequence upstream of a core promoter is either 10 or 21 bp.

7. The isolated nucleic acid composition according to any of the claims 1 to 6 wherein each spacer length has one out of three designed sets of the TFBS-devoid spacer sequences selected from table 3.

8. The isolated nucleic acid composition from any of the preceding claims wherein the core promoter represents any promoter capable of initiating transcription of said nucleic acid composition, preferably the core promoter being selected from the group of: minP derived from pGL4, minCMV, minHBG, SCP1 (PMID 17124735), AdML, Hsp68 (2557196).

9. The isolated nucleic acid composition from any of the preceding claims wherein the reporter protein is a fluorescent or luminescent protein.

10. The isolated nucleic acid composition of any of the claims 1-9 wherein the synthetic reporter comprises a nucleic acid sequence with at least 95% identity according to SEQ ID NO: 147 to SEQ ID NO: 232, preferably at least 98% identity, more preferably at least 99% identity, most preferably 100% identical to said SEQ ID NOs 11. A library of isolated nucleic acid molecules that encode one or more synthetic reporters containing the designed spacer sequences from table 3 binding to one or more transcription factors according to claims 1-10.

12. The library of isolated nucleic acid molecules from claim 11, wherein the transcription factors represent all mammalian transcription factors, preferably all human transcription factors.

13. The library of isolated nucleic acid molecules from claim 12, wherein the transcription factors are selected from table 1.

14. A high-throughput method for simultaneously detecting the activity of two or more transcription factors in one or more mammalian cells, the method comprising the steps of: a. Providing the library of synthetic transcription factor reporters according to any of the claims 11 to 13. b. Transfecting mammalian cells with one or more vectors encoding the synthetic reporters. c. Detecting and measuring TF activity via a reporter assay selected from the group of: luminescence assay, fluorescence assay and / or measuring TF activity by estimating the abundance of the barcoded transcript by sequencing-based methods.

15. The method according to claim 14 wherein the mammalian cells are cultured in an organoid.

16. The method according to claim 15 wherein the mammalian cell is present in a transgenic animal model.

17. The method according to any of the claims 14 to 16 wherein the vectors are taken from the list of synthetic, bacterial, viral, or transposable element-based.

18. The method according to any of the claims 14 to 17, wherein the mammalian cells, organoids or transgenic animal models have been exposed to specific therapeutic molecules, toxins, chemicals or physical treatments prior to or during use of said method.

19. The method according to claims 14 to 18, wherein the synthetic reporter is capable of binding to transcription factors that are overexpressed or downregulated.

20. The method according to claim 14 to 19, wherein the synthetic reporter is capable of binding to mutant transcription factors.

21. The method according to any of the claims 14 to 20, combined with single-cell RNA sequencing methods for determining the activity of one or more transcription factors in single cells together with the expression status of any other gene in the cell.

22. The method according to any of the claims 14 to 21, combined with a (sc)ATAC-sequencing method.

23. A vector comprising the nucleic acid composition of any of the claims 1 to 10, wherein the vector is synthetic, bacterial, viral, or transposable element – based.

24. A mammalian cell, organoid or mammal containing one or more vectors from claim 23.

25. A mammal containing in its genome multiple barcoded synthetic TF reporters according to any of the claims 1 to 10 26. The mammal from claim 25 wherein said mammal is a mouse 27. A kit of parts comprising: a. The isolated nucleic acid composition of claim 1 or b. The library of any of the claims 11 to 13 c. One or multiple vectors according to claim 23 d. A gene-specific primer used for reverse transcription. e. A primer pair binding to the sequences up- and downstream of the barcode for amplification and quantification of barcodes.

28. A computer-implemented method for determining transcription factor (TF) reporter activities, the method comprising: a) receiving sequencing data for a plurality of TF reporters for at least one control sample; and b) receiving sequencing data for a plurality of TF reporters for at least one contrast sample; and c) Inputting the sequencing data in a) and the sequencing data in b) into a statistical model which outputs: (i) a measure of TF reporter activities; and / or (ii) a measure of changes in TF reporter activities between the control sample in a) and the contrast sample in b).

29. The computer-implemented method of claim 28, wherein plasmid DNA data or genomic DNA data is also input to the statistical model.

30. The computer-implemented method of claim 28 or claim 29, wherein at least one quality control measure and / or at least one exploratory plot is generated.

31. A computing apparatus for determining transcription factor (TF) reporting activities, the apparatus comprising; one or more memory units; and one or more processors configured to execute instructions stored in the one or more memory units to perform the method of any of claims 28-30.

32. A non-transitory computer-readable storage medium including instructions that, when executed by one or more processors of a computer, configure the computer to perform the method of any one of claims 28-30.

Citation Information

Patent Citations

  • Secretable reporter system

    WO2008073805A1

  • Multiplexing transcription factor reporter protein assay process and system

    US10544472B2

  • Nucleic acid constructs and vectors for podocyte specific expression

    WO2023213738A1

  • Synthetic cancer-specific promoters

    WO2025019712A1