Synthetic transcription factors

By using synthetic transcription factors derived from the mitochondrial DNA binding protein MTERF1, the problems of immune response and endogenous interference of non-human transcription factors in the human body were solved, and efficient and safe gene expression control in human cells was achieved.

CN120752252APending Publication Date: 2025-10-03ETH ZURICH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202380094145.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-08-25
Filing Date
2023-12-21
Publication Date
2025-10-03

AI Technical Summary

Technical Problem

In the existing technology, transcription factors of non-human origin are prone to trigger immune responses and interfere with endogenous gene regulation in the human body, resulting in reduced safety and effectiveness of gene therapy and difficulty in achieving tissue-specific and reliable gene expression.

Method used

A synthetic transcription factor containing a DNA binding domain derived from the mitochondrial DNA binding protein MTERF1 and a human transcriptional regulatory domain is used, designed as a fusion protein to ensure high orthogonality in human cells and avoid interference with endogenous gene regulation.

Benefits of technology

Highly orthogonal gene expression control in human cells is achieved, ensuring that genes of interest are expressed only in specific cells, reducing the risk of side effects and improving the safety and efficacy of gene therapy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure BDA0005548411470000391
    Figure BDA0005548411470000391
  • Figure BDA0005548411470000401
    Figure BDA0005548411470000401
  • Figure BDA0005548411470000411
    Figure BDA0005548411470000411
Patent Text Reader

Abstract

The present invention is generally in the field of synthetic biology, gene therapy and cell therapy. In particular, the present invention relates to synthetic transcription factors comprising a DNA binding domain derived from a mitochondrial DNA binding protein (e.g., MTERF1) and a transcriptional regulatory domain derived from one or more other proteins; nucleic acids or combinations of nucleic acids encoding the synthetic transcription factors of the invention; a DNA construct comprising an MTERF1 binding site and a minimum promoter; a system comprising a synthetic transcription factor of the invention and a DNA construct of the invention; and uses thereof, such as medical uses.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] The present invention generally relates to the fields of synthetic biology, gene therapy, and cell therapy. In particular, the present invention relates to synthetic transcription factors comprising a DNA binding domain derived from a mitochondrial DNA binding protein (e.g., MTERF1) and a transcriptional regulatory domain derived from one or more other proteins; nucleic acids or combinations of nucleic acids encoding the synthetic transcription factors of the invention; DNA constructs comprising an MTERF1 binding site and a minimal promoter; systems comprising the synthetic transcription factors of the invention and the DNA constructs of the invention; and uses thereof, such as medical uses. In addition, the present invention also relates to libraries of DNA constructs and their use in methods for optimizing promoters for binding to transcription factors.

[0002] In gene and cell therapy (GCT), artificial gene sequences are at the heart of therapeutic formulations. In gene therapy, a delivery vehicle carrying the therapeutic gene sequence ("payload") is administered directly to the patient, delivering the therapeutic payload to multiple cells within the patient's body. In cell therapy, the payload is introduced ex vivo into cells, such as patient-derived cells, which are then (re)infused into the patient [1].

[0003] The therapeutic transgene encoded on the gene payload of a GCT is typically driven by a constitutive or tissue-specific promoter. However, the spectrum of naturally occurring promoters for specific expression of transgenes is limited [2]. Alternatively, tissue or cell type specificity can be achieved with the help of an engineered gene network (also fully encoded on the gene payload) that processes multiple cellular inputs in a programmable logic manner (i.e., a "biological computational circuit"). For example, a cancer cell classifier circuit was designed to restrict the expression of pro-apoptotic genes to cancer cells while sparing healthy cells [3]. In the context of chimeric antigen receptor (CAR) T cell therapy, an AND-gate was constructed, resulting in increased specificity for cancer cells that co-express two specific antigens on their surface [4]. Such circuits can lead to safer and more effective GCTs. However, in addition to the therapeutic transgene, the circuit usually encodes protein components ("accessory proteins"). These protein components help to execute the logic control process of precisely activating the therapeutic transgene when multiple conditions are met in the target cell.

[0004] Accessory proteins, such as transcription factors, can increase toxicity and reduce the efficacy of biocomputation-based therapies in human patients, particularly if these proteins are of non-human origin. The use of accessory proteins of non-human origin is driven by the so-called “orthogonality requirement”: these proteins should not interfere with endogenous gene regulation in human cells, and they or the processes they control should not be alterable or perturbed by endogenous human factors. A priori, there is a concern that fully human proteins used as accessory proteins in gene therapy payloads will bind to endogenous components of human cells and cause side effects and toxicity. To achieve orthogonality in accessory proteins, particularly when the protein is a transcription factor, the prior art employs non-human DNA binding domains (DBDs) that have no binding sites in the regulatory sequences of human genes. Non-human DBDs are typically derived from prokaryotes and are further fused to transcriptional activation domains of viral or human origin [5]. While meeting the orthogonality requirement, non-human proteins have been found in animal models and human clinical trials to elicit immune responses and lead to clearance of cells expressing such proteins [6][7]. This results in reduced therapeutic efficacy of GCT products. Thus, the orthogonality perspective favors the use of proteins of non-human origin, while the low immunogenicity requirement suggests the opposite, i.e., the use of human or mostly human proteins, which in turn can lead to reduced orthogonality. Reconciling these two opposing requirements is therefore challenging.

[0005] There are currently two general approaches to addressing the immunogenicity of transgenes, and in particular, the immunogenicity of proteins of non-human origin.

[0006] The first approach aims to eliminate the immune response while simultaneously accepting the potentially immunogenic protein sequence. This often requires systemic drugs to suppress the immune system. However, this strategy can increase the risk of infection in patients. Other strategies borrow from or replicate viral strategies that antagonize host defenses and include blocking proteasomal degradation and epitope production [9], downregulating MHC class I

[10] , surface enzymes that degrade antibodies

[11] , and combinations thereof. Disadvantages of these strategies are difficulties in fine-tuning and reliability, as well as the often large DNA cargo in the case of immune-evading viral proteins encoded on the delivery vector. A more targeted approach is to introduce antigen-presenting cell (APC)-specific miRNA target sites into the 5' or 3' UTR of the transgene

[12] . This prevents transgene expression in APCs, a key step in inducing an immune response. While this approach has been shown to be effective in some conditions, it has failed in others, hindering its application in a range of diseases

[13] .

[0007] For some delivery vectors, such as AAV, inhibition of AAV interaction with TLR9 in dendritic cells was shown to reduce the increase in adaptive immunity against the viral capsid and transgene product

[14] .

[0008] Another approach is to try to reduce the immunogenicity of the protein itself by modifying or cleverly selecting the protein sequence. One strategy is to “humanize” the transgene, that is, to replace the non-human protein domain with a human-derived domain that the immune system cannot recognize as foreign / non-self. For synthetic transcription factors, this strategy is suitable for transcriptional activation domains because these domains are generally not expected to cause crosstalk with their own endogenous processes. However, it is difficult to apply this strategy to DNA binding domains (DBDs) because human-derived DBDs do pose the risk of interfering with endogenous gene expression. In one previously proposed “humanized” DBD solution, zinc finger (ZNF) protein domains derived from human proteins (each domain has a DNA binding specificity of 2 to 3 DNA base pairs) are fused to each other to recognize longer DNA sequences that do not naturally occur in the human genome, thereby meeting the orthogonality criterion

[15]

[16] . However, the fusion of several ZNF domains produces multiple domain junction regions of non-human sequence, which can themselves be immunogenic. Therefore, such artificial DBDs are largely non-human. Especially considering the large patient population with a diverse MHC, TCR and antibody repertoire, immune responses to junction-derived peptides may still occur. Another approach to humanizing the DBD of a synthetic transcription factor (TF) is to utilize a naturally occurring human DBD. This reduces the number of new (non-human) junctions to one, namely the junction between the human-derived DBD and the human-derived TAD, thereby minimizing the potential for immunogenicity. Typically, one uses a DBD of a human TF that is not expressed in the cell type expressing the synthetic TF

[17] . However, there is a risk that the chimeric protein constructed according to this strategy will bind to its cognate response element in the human genome and thus lead to unwanted expression of endogenous human genes, which violates the requirement of orthogonality.

[0009] Therefore, there remains a need for improved means and methods to regulate gene expression, particularly in humans.

[0010] The invention relates to the embodiments as characterized in the claims and described hereinafter.

[0011] Thus, the present invention relates to synthetic transcription factors comprising (i) a DNA binding domain (DBD) derived from a mitochondrial DNA binding protein and (ii) a transcriptional regulatory domain derived from one or more other proteins.

[0012] The present invention is based, at least in part, on the surprising discovery that synthetic transcription factors comprising a DNA binding domain (DBD) from a mitochondrial DNA binding protein, such as MTERF1, are capable of regulating transcription in cells, particularly in the nucleus of the cell.

[0013] As illustrated in the accompanying Examples, synthetic transcription factors, in particular fusion proteins (referred to as "MTF"), have been discovered that comprise (i) a C-terminal fragment of the human MTERF1 protein containing a DNA binding domain but lacking a mitochondrial transfer peptide (MTP), and (ii) a human transcriptional activation domain, RelA, capable of promoting transcription of a gene of interest (GoI) in a cell, in particular in the nucleus of a cell. 430-551 Furthermore, in this context, DNA constructs, ie gene expression constructs, have been developed which comprise a promoter containing a response element (RE) for binding the synthetic transcription factor and a minimal promoter, wherein the promoter can be operably linked to the gene of interest.

[0014] As further illustrated in the accompanying examples, it was surprisingly found that the synthetic transcription factors of the present invention, such as MTF proteins, and the transcription systems of the present invention (e.g., comprising the synthetic transcription factors and the DNA constructs) are highly orthogonal in cells of interest (e.g., human cells); see, e.g., Example 3.

[0015] In particular, the inventors could demonstrate that the synthetic transcription factor, MTF protein, diffuses within the cell, including the nucleus (see, e.g. Figure 7 A). This functionality is in stark contrast to the WT MTERF1 protein, which localizes only to mitochondria. Furthermore, the ability of the synthetic transcription factor to promote transcription of the gene of interest from the synthetic DNA construct is not affected in cells when the WT MTERF1 protein is overexpressed (see, e.g. Figure 7 C).

[0016] Using RNA-sequencing, the inventors further found that MTF does not retain the gene regulatory functionality of WT MTERF1. Surprisingly, only very few genes were differentially expressed after transfection of the MTF construct, which showed a high degree of orthogonality with the endogenous gene regulatory process in human cells (see, e.g., Figure 8 ).

[0017] Thus, the present inventors have developed synthetic transcription factors comprising human DNA binding domains that unexpectedly exhibit extremely high orthogonality in human cells. Specifically, (i) endogenous human gene expression is not substantially altered by the synthetic transcription factors of the present invention; and (ii) endogenously expressed wild-type MTERF1 does not regulate (i.e., does not interfere with) expression of a gene of interest in human cells driven by a corresponding promoter containing a response element (i.e., a DNA construct of the present invention). In other words, the inventive methods of the present invention, as illustrated in the accompanying Examples, do not substantially interfere with endogenous gene regulation in human cells, and transcription of the gene of interest is not interfered with by the endogenous human factor (i.e., WT MTERF1).

[0018] Thus, the present invention provides, inter alia, improved components for improved transcription systems, in particular improved synthetic transcription factors and corresponding gene expression constructs, which function in a substantially orthogonal manner in cells of a particular species (eg, in humans).

[0019] In cells (e.g., in human cells), a high degree of orthogonality is conducive to ensuring reliable control of the expression of a gene of interest, e.g., a gene encoding a pro-apoptotic protein of interest. In particular, a high degree of orthogonality ensures that the gene of interest is expressed only in those cells in which it should be expressed, and only when it should be expressed. In addition, the high degree of orthogonality of synthetic transcription systems ensures that endogenous gene expression is not disturbed or destroyed in an undesirable manner, and therefore reduces the risk of side effects or toxicity.

[0020] In particular, in the context of engineered gene networks that process multiple cellular inputs in a programmable logic manner (i.e., “biocomputing circuits”), such as in cancer cell classifier circuits, a high degree of orthogonality is highly advantageous.

[0021] It is expected that gene therapy products that require engineered transcription factors as part of their mechanism of action, such as gene therapy products that otherwise operate as multi-component networks (known as "biocomputational gene circuits"), will have a favorable safety profile when using synthetic transcription factors according to the present invention compared to alternatives.

[0022] Thus, the present invention further provides more effective and / or safer methods described herein for gene and / or cell therapy, e.g., therapy employing biological computational circuits.

[0023] As used herein, gene refers to a nucleotide sequence in DNA that is transcribed to produce functional RNA. A gene can be a protein-coding gene or a non-coding gene. In addition, a gene is typically associated with at least one regulatory sequence, such as a promoter and an optional enhancer, that can participate in the transcription of the gene. A promoter that can participate in gene transcription can further be considered to be operably connected to the gene. Typically, a regulatory sequence (such as a promoter) comprises at least one response element, and the response element comprises at least one binding site for a transcription factor, as further described herein.

[0024] As used herein, a transcription factor refers to a protein that regulates the transcription of one or more genes. In particular, transcription factors control the transcription rate of a gene, i.e., the rate at which the gene information is transcribed from DNA into messenger RNA, by binding to a specific DNA sequence, a response element (RE), such as in a promoter. A "response element" may also be referred to as a "transcription factor binding site," or may contain at least one transcription factor binding site.

[0025] The defining characteristic of transcription factors is that they contain at least one DNA binding domain (DBD), which binds (i.e., attaches to) specific DNA sequences adjacent to or at a distance from the genes they regulate (i.e., regulatory sequences). In particular, the DBD binds to response elements contained within regulatory sequences (e.g., promoters or enhancers).

[0026] In addition, transcription factors contain transcriptional regulatory domains, such as activation domains or repression domains, which typically contain interaction sites for other proteins (such as transcriptional co-regulators). Activation domains may also be referred to as "transactivation domains (TADs)" or "transcriptional activation domains." Similarly, repression domains may also be referred to as "transcriptional repression domains."

[0027] Optionally, a transcription factor may further comprise a signal sensing domain (SSD) (eg, a ligand binding domain) that senses external signals and, in response, transmits these signals to the rest of the transcription complex, resulting in upregulation or downregulation of gene expression.

[0028] Although transcription factors can work alone, they typically work with other proteins in complexes by promoting (as activators) or inhibiting (as repressors) the recruitment of RNA polymerase (the enzyme that carries out the transcription of genetic information from DNA into RNA) to specific genes (e.g., the genes of interest described herein).

[0029] Thus, the DNA binding domain of a transcription factor directs the transcription factor to the regulatory sequence of a gene, particularly a response element contained in a promoter or enhancer associated with the gene, and the transcription regulatory domain typically acts in concert with an endogenous transcription regulatory factor (i.e., a co-regulator) that binds to or interacts with the transcription regulatory domain to promote or inhibit transcription of the gene, as further described below. In particular, a transcription factor can stimulate transcription initiation, particularly when it has an activation domain, or more precisely, a transcription factor can hinder transcription initiation, particularly when it has an inhibition domain. For example, a transcription factor can help RNA polymerase bind to DNA, which in turn promotes transcription, particularly transcription initiation, or a transcription factor can hinder RNA polymerase binding to DNA, which in turn inhibits transcription, particularly transcription initiation.

[0030] As used herein and in the context of the present invention, the term "synthetic" refers specifically to a compound, such as a transcription factor, comprising at least two parts that do not occur together in this manner in nature. In particular, a synthetic protein, such as a synthetic transcription factor, comprises one part derived from a specific protein and at least one other part derived from at least one other protein.

[0031] In particular, herein, the synthetic transcription factors of the present invention include (i) a DNA binding domain (DBD) derived from a certain protein (i.e., mitochondrial DNA binding protein) and (ii) a transcriptional regulatory domain derived from one or more other proteins.

[0032] In a preferred embodiment of the present invention, the synthetic transcription factor is a fusion protein.

[0033] As used herein and in the context of the present invention, a fusion protein refers to a synthetic protein, such as a synthetic transcription factor, in which at least two or all parts of a protein, particularly parts that do not occur together in a single polypeptide in nature, are contained in one amino acid chain, i.e. one polypeptide.

[0034] Therefore, in a preferred embodiment, the synthetic transcription factor of the present invention is a fusion protein comprising a DNA binding domain according to the present invention and a transcription regulatory domain according to the present invention. In particular, the two domains are linked in the fusion protein via a peptide bond, either directly or via a peptide linker. In other words, in a preferred embodiment, the DNA binding domain according to the present invention and the transcription regulatory domain according to the present invention are contained in a single amino acid chain, i.e., a single polypeptide. More preferably, in the context of these embodiments, substantially all parts of the synthetic transcription factor of the present invention are contained in a single polypeptide.

[0035] In a further embodiment of the present invention, the synthetic transcription factor comprises or consists of a first polypeptide and a second polypeptide, wherein the first polypeptide comprises a DNA binding domain according to the present invention and the second polypeptide comprises a transcriptional regulatory domain according to the present invention. In particular, the first polypeptide and the second polypeptide each comprise a multimerization domain as described herein, wherein the multimerization domains of the first polypeptide and the second polypeptide are capable of binding to and / or interacting with each other.

[0036] In the context of this document and the present invention, a mitochondrial DNA binding protein refers to a DNA binding protein that is typically localized to mitochondria. In particular, a mitochondrial DNA binding protein binds to mitochondrial DNA in a sequence-specific manner. For example, a mitochondrial DNA binding protein can be a mitochondrial transcription factor or a mitochondrial transcription termination factor.

[0037] Preferably, herein and in the context of the present invention, the mitochondrial DNA binding protein is from a mammalian species. Preferably, herein, the mammalian species is human.

[0038] In the context of the present invention, it has been further discovered that MTERF1 has a large recognition site that is rare or absent in gene regulatory sequences in the nuclear genome of human cells, and that MTERF1 does not have any perfect binding sites in the human nuclear genome. This is particularly advantageous for the high degree of orthogonality (e.g., in human cells), as described herein.

[0039] Therefore, in the context of this document and the present invention, the mitochondrial DNA binding protein is preferably MTERF1, preferably human MTERF1, namely Uniprot Q99551. In particular, human MTERF1, namely wild-type (WT) MTERF1 has the amino acid sequence shown in SEQ ID NO:108.

[0040] However, MTERF1 can also be derived from other species, such as mouse or dog. In particular, mouse MTERF1 has the amino acid sequence set forth in SEQ ID NO:111 or 113. Furthermore, canine MTERF1 has the amino acid sequence set forth in SEQ ID NO:115. Furthermore, orthologous sequences of MTERF1 from other species, such as mammalian species, are readily available to those skilled in the art and can also be used herein and in the context of the present invention.

[0041] In the context of this document and the present invention, the term "derived" should be interpreted in a technically meaningful way. For example, a domain derived from a certain protein refers to a part or fragment of said protein, which may further comprise at least one modification, in particular at least one amino acid substitution, deletion and / or insertion. The degree of modification can be defined by sequence identity with a reference sequence. Furthermore, a domain derived from a certain protein particularly has qualitatively similar functionality to the domain in said protein (although the functionality may be enhanced or diminished to a certain extent).

[0042] In particular, a DNA binding domain derived from a certain DNA binding protein, such as MTERF1, has the ability to bind to DNA in a sequence-specific manner, in particular to the response element of the DNA binding protein, such as MTERF1.

[0043] Similarly, a transcriptional regulatory domain (e.g., an activation domain) derived from a certain transcriptional activator protein (e.g., a transcription factor such as RELA) has, in particular, the ability to regulate (e.g., promote) the transcription of a gene when it is bound to or located near a regulatory sequence (e.g., a response element) of the gene.

[0044] In general, herein and in the context of the present invention, an amino acid sequence (e.g., the amino acid sequence of a certain protein domain or motif) or a DNA sequence (e.g., the DNA sequence of a binding site or minimal protein) can be defined by a certain % sequence identity to a reference sequence. Thus, as used herein, the term "sequence identity" is particularly used to describe the sequence relationships between two or more amino acid sequences, proteins (or fragments thereof) or polypeptides (or fragments thereof). In particular, a sequence can have at least n% sequence identity to a reference sequence, where n is an integer between 60 and 100, for example, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98 or 99. For example, an amino acid sequence may have at least 60%, 70%, 80% or 90%, preferably at least 80%, 85%, 90% or 95%, more preferably at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98% or 99% sequence identity with an amino acid sequence as shown in a SEQ ID NO. The same applies to nucleic acid sequences, such as DNA sequences, mutatis mutandis.

[0045] In general, the higher the % sequence identity, the more preferred the sequence. However, further preferred % identities are described directly in the context of certain embodiments herein. It should further be noted that the present invention is in no way limited to high or preferred sequence identities, but rather any % sequence identity is contemplated, such as those immediately described above.

[0046] As used herein and in the context of the present invention, the term "sequence identity" has substantially the same meaning as that commonly used and understood by those skilled in the art. The degree of sequence identity can be determined according to methods well known in the art, preferably using a suitable computer algorithm, such as CLUSTAL.

[0047] When using the Clustal analysis method to determine whether a particular sequence is, for example, at least 60% identical to a reference sequence, the default settings can be used.

[0048] In a preferred embodiment, Clustal Omega (Madeira F, Park YM, Lee J, et al., The EMBL-EBI search and sequence analysis tools APIs in 2019. Nucleic Acids Research. 2019 Jul; 47(W1): W636-W641. DOI: 10.1093 / nar / gkz268. PMID: 30976793; PMCID: PMC6602479) is used to compare amino acid sequences. In the case of pairwise comparisons / alignments, the following default settings are preferably selected: Program: clustalo; Version: 1.2.4; Input Parameters: Output guide tree: True; Output distance matrix: False; Unalign input sequences: False; mBed-like clustering guide tree: True; mBed-like clustering iterations: True; Iteration number: 0; Maximum guide tree iterations: -1; Maximum HMM iterations: -1; Output alignment format: clustal_num; Output order: alignment; Sequence type: protein. Preferably, the degree of identity is calculated over the entire length of the sequence.

[0049] Further, those skilled in the art can identify the amino acid residue at a position corresponding to the position in the reference sequence by methods known in the art. The comparison can be performed using means and methods known to those skilled in the art, for example, by using a known computer algorithm, such as the Lipman-Pearson method (Science 227 (1985), 1435) or the CLUSTAL algorithm. In such comparisons, it is preferred that maximum homology be assigned to the conserved amino acid residues present in the amino acid sequence.

[0050] In a preferred embodiment, Clustal Omega is used for amino acid sequence comparison. In the case of pairwise comparisons / alignments, the following default settings are preferably selected: Program: clustalo; Version: 1.2.4; Input Parameters: Output guide tree: True; Output distance matrix: False; Unalign input sequences: False; mBed-like clustering guide tree: True; mBed-like clustering iterations: True; Iteration number: 0; Maximum guide tree iterations: -1; Maximum HMM iterations: -1; Output alignment format: clustal_num; Output order: Alignment; Sequence type: Protein.

[0051] In the context of the present invention, "amino acid substitution" means that each amino acid residue at a specified position can be replaced by any other possible amino acid residue, such as a naturally occurring amino acid or a non-naturally occurring amino acid (Brustad and Arnold, Curr. Opin. Chem. Biol. 15 (2011), 201-210).

[0052] Generally speaking, herein and in the context of the present invention, a feature "having" a certain sequence may "comprise" the sequence, may be "defined by" the sequence, or may "consist of" the sequence.

[0053] In the context of this invention, the DNA binding domain according to the invention may have at least 60%, preferably at least 70%, more preferably at least 80% sequence identity with SEQ ID NO: 1. In certain preferred embodiments, the DNA binding domain has the sequence shown in SEQ ID NO: 1.

[0054] As illustrated in the accompanying examples, it was further surprisingly found in the context of the present invention that the DNA binding domain of MTERF1 can be truncated, for example at the N-terminus, and still fully retain its functionality, i.e. bind to its response element and provide functionality for synthetic transcription factors (see, e.g. Figure 3 Shorter DNA binding domains can be advantageous because they have a reduced DNA footprint / gene payload, for example for viral transduction. Further, adjusting the length of the MTERF1 binding domain can alter the expression level of a gene of interest.

[0055] Therefore, the DNA binding domain according to the present invention may comprise MTERF1 subdomain B having the sequence shown in SEQ ID NO:7 or a sequence having at least 70%, preferably at least 80%, more preferably at least 90% sequence identity thereto.

[0056] Surprisingly, it has further been found that one of the MTERF1 motifs in the MTERF1 DNA binding domain can be omitted while substantially retaining the functionality of the DNA binding domain.

[0057] Therefore, the DNA binding domain according to the present invention may comprise MTERF1 subdomain A (i.e., the shorter subdomain), which has the sequence shown in SEQ ID NO: 9 or a sequence having at least 80%, preferably at least 90%, more preferably at least 95% sequence identity with SEQ ID NO: 9.

[0058] Thus, the DNA binding domain according to the present invention may comprise: (i) a first MTERF1 motif having the sequence shown in SEQ ID NO: 104 or a sequence having at least 80%, preferably at least 90%, and more preferably at least 95% sequence identity thereto; and / or (ii) a second MTERF1 motif having the sequence shown in SEQ ID NO: 106 or a sequence having at least 80%, preferably at least 90%, and more preferably at least 95% sequence identity thereto. Preferably, the first MTERF1 motif is located N-terminally to the second MTERF1 motif. Furthermore, the first and second MTERF1 motifs may be directly adjacent to each other (preferably via a peptide bond) or connected via a linker (particularly via a peptide linker).

[0059] Preferably, the DNA binding domain comprising the first and / or second MTERF1 motif further comprises (iii) a MTERF1 C-terminal domain having the sequence shown in SEQ ID NO: 11 or a sequence having at least 70%, preferably at least 80%, more preferably at least 90% sequence identity with SEQ ID NO: 11.

[0060] In particular, the first MTERF1 motif and / or the second MTERF1 motif described herein are comprised in the subdomain A described herein. Furthermore, the MTERF1 subdomain A described herein is particularly comprised in the MTERF1 subdomain B described herein.

[0061] Furthermore, according to the present invention, several amino acids, for example, about 1 to 30 or about 1 to 10 amino acids, can be deleted from the C-terminus of the MTERF1-derived DNA-binding domain or the MTERF1-derived DNA-binding subdomain. According to the present invention, it is also possible to delete the entire C-terminal subdomain of the MTERF1-derived DNA-binding domain or the MTERF1-derived DNA-binding subdomain.

[0062] Preferably, herein and in the context of the present invention, the inventive MTERF1 derived DNA binding domain exhibits an arginine (R) at the position corresponding to position 387 in the sequence of SEQ ID NO: 108 (eg at position 330 in the sequence of SEQ ID NO: 1).

[0063] Preferably, in the context of this invention and herein, the synthetic transcription factor does not comprise a mitochondrial transfer peptide having the sequence shown in SEQ ID NO: 37 or a sequence having at least 90% sequence identity to SEQ ID NO: 37. Preferably, the synthetic transcription factor is completely free of a functional mitochondrial transfer peptide. In particular, the synthetic transcription factor of the invention does not comprise a mitochondrial transfer peptide at its N-terminus.

[0064] Herein, “mitochondrial transfer peptide” may also be referred to as “mitochondrial targeting signal.” In particular, the mitochondrial transfer peptide directs proteins to mitochondria, allowing the proteins to enter and / or be localized in mitochondria.

[0065] However, the synthetic transcription factors of the present invention are preferably unable to enter or localize to mitochondria.

[0066] In the context of this document and the present invention, it is highly preferred that the synthetic transcription factors of the present invention are able to enter and / or localize to the nucleus. This is particularly advantageous for achieving the high degree of orthogonality described herein.

[0067] Thus, the synthetic transcription factors of the present invention may have the ability to localize more efficiently to the nucleus of a cell than to the mitochondria of the cell. This ability can be determined by separately measuring the amount of the synthetic transcription factor in the nucleus and mitochondria of the same cell. This can be accomplished by conventional methods in the art, such as immunostaining, separation of nuclei and mitochondria followed by Western blotting or ELISA.

[0068] When there is no functional mitochondrial transfer peptide described herein (e.g., as shown in SEQ ID NO: 37), the synthetic transcription factor of the present invention can enter and / or be localized to the cell nucleus. In addition, when the synthetic transcription factor of the present invention comprises a nuclear localization signal, it can enter and / or be localized to the cell nucleus.

[0069] Therefore, the synthetic transcription factors of the present invention may comprise a nuclear localization signal (NLS). A nuclear localization signal is also referred to as a "nuclear localization peptide". NLS sequences are well known in the art, and in principle any of them may be employed.

[0070] Herein, the term "gene expression" always encompasses the term "gene transcription" or "transcription of a gene," but in some cases, it may also include post-transcriptional mechanisms. As used herein, the term "gene transcription" or "transcription of a gene" refers to gene expression in a more specific manner. However, unless otherwise specifically stated, the term "gene expression" herein may be replaced by the term "gene transcription" or "transcription of a gene," as the present invention particularly relates to means and methods for regulating gene transcription.

[0071] In particular, herein and in the context of the present invention, the synthetic transcription factor is capable of regulating the transcription of at least one gene of interest in a cell. Preferably, the synthetic transcription factor is capable of regulating the transcription of at least one gene of interest in the nucleus of a cell.

[0072] The synthetic transcription factors of the present invention can regulate the transcription of a gene of interest in the same manner as generally described herein in the context of transcription factors. In particular, the synthetic transcription factors of the present invention can control the transcription rate of a gene of interest. In addition, the synthetic transcription factors of the present invention can induce or initiate transcription of a gene of interest.

[0073] As used herein and in the context of the present invention, gene of interest is not limited to any gene.Herein and in the context of the present invention, gene of interest preferably encodes, for example, pro-cell death proteins such as hBAX or HSV-TK, immunostimulatory cytokines such as IL-2 or IL-12 or antigen receptors such as CAR or TCR.

[0074] hBAX refers to a pro-apoptotic protein that can be used in a cell sorter to kill cells, for example, in a cancer cell sorter to kill cancer cells. HSV-TK refers to a protein that metabolizes ganciclovir into toxic metabolites. It can be used in a cell sorter to kill cells, for example, in a cancer cell sorter to kill cancer cells. IL-2 or IL-12 are immunostimulatory cytokines that can be used in, for example, a cell sorter (such as a cancer cell sorter) to attract and induce T cell proliferation. CAR and TCR refer to antigen receptors that recognize cells with corresponding surface antigens. They can be used, for example, in combination with the SynNotch system; Morsut (2016), Cell 164 (4): 780-91.

[0075] Preferably, herein and in the context of the present invention, the synthetic transcription factors regulate the transcription of a gene of interest in a cell by (i) promoting the transcription of the gene of interest, or by (ii) inhibiting the transcription of the gene of interest. As described herein, the synthetic transcription factors of the present invention preferably promote or inhibit the transcription of the gene of interest in the nucleus of the cell. This is particularly advantageous for achieving the high degree of orthogonality described herein.

[0076] In particular, the synthetic transcription factors of the present invention can promote gene transcription by increasing the transcription rate. In addition, the synthetic transcription factors of the present invention can inhibit gene transcription by decreasing the transcription rate.

[0077] As used herein, promoting gene transcription can include, for example, inducing, initiating and / or enhancing gene transcription. Additionally, inhibiting gene transcription can include, for example, blocking or inhibiting gene transcription, such as blocking or inhibiting the induction or initiation of gene transcription.

[0078] Preferably, in the context of this invention and herein, the synthetic transcription factor is capable of binding to a response element in a cell, preferably in the nucleus of a cell. In particular, the DNA binding domain comprised in the synthetic transcription factor of the invention is capable of binding to a response element in a cell, preferably in the nucleus of a cell.

[0079] In the context of this document and the present invention, the response element may comprise an MTERF1 binding site having the sequence shown in SEQ ID NO: 42 or a sequence having at least 60%, preferably at least 80%, more preferably at least 90% sequence identity to SEQ ID NO: 42. Preferably, the MTERF1 binding site according to the present invention consists of the sequence shown in SEQ ID NO: 42 or a sequence having at least 60%, preferably at least 80%, more preferably at least 90% sequence identity to SEQ ID NO: 42.

[0080] In particular, the synthetic transcription factors and / or DNA binding domains according to the present invention are capable of binding to at least one MTERF1 binding site in a response element, as described herein.

[0081] It has further been found that the expression level of a gene of interest can be modulated by varying the design of the promoter containing the response element and the amount of synthetic transcription factor expressed in the cell. In addition, it has been observed that the copy number of REs and the spacing between them also affect GoI expression at a constant MTF level (see, e.g., Figure 1 ). It is worth noting that Figure 1 In the present invention, the term "response element" (RE) corresponds to the term "binding site" used herein, in particular the MTERF1 binding site defined by SEQ ID NO:42.

[0082] Also in the context of the DNA constructs of the invention, the response element according to the invention may comprise multiple copies of an MTERF1 binding site, as described herein, for example 2 to 50, 2 to 25, or 2 to 15 copies, preferably 2 to 5. Furthermore, the individual copies of the binding site need not be identical to each other, but may comprise variations.

[0083] Furthermore, the copies (e.g., two or more or all copies) of the binding site in the response element may be directly adjacent to each other or separated by one or more (e.g., 1 to 100, 1 to 50, 1 to 20, 1 to 10, such as 6 or 10) nucleotides.

[0084] The synthetic transcription factors of the present invention are particularly capable of binding to a promoter in a cell (preferably in the nucleus of a cell) that comprises a response element as described herein. The promoter may further comprise a minimal promoter, for example, a minimal TATA box preferably having the sequence set forth in SEQ ID NO: 103, or a minimal CMV promoter preferably having the sequence set forth in SEQ ID NO: 136. Preferably, the minimal promoter is located 3' to the response element. For example, the promoter has the sequence set forth in SEQ ID NO: 110.

[0085] Further promoters, response elements and minimal promoters are described herein in the context of the DNA constructs of the invention and may also be contemplated and employed in this specific context herein.

[0086] Preferably, herein and in the context of the present invention, the promoter is operably linked to a gene of interest as described herein, preferably in the nucleus as described herein.

[0087] In particular, herein and in the context of the present invention, the binding of a synthetic transcript of the invention (in particular a synthetic transcript of a DNA binding domain) to a promoter of the invention (in particular to a response element of the invention in said promoter) regulates (preferably in the nucleus of a cell) the transcription of a gene of interest operably linked to said promoter as described herein.

[0088] The orthogonality of the synthetic transcription factors of the invention or the transcription systems of the invention described herein is particularly assessed in cells from the same species (e.g., the same mammalian species) as the mitochondrial DNA binding protein of the invention (and therefore the DNA binding domain contained in the synthetic transcription factors of the invention). Such cells are also referred to herein as "homologous cells."

[0089] As described herein, the synthetic transcription factors of the present invention and the transcription systems of the present invention can act on endogenous processes in cells (ie, cells of the same species), particularly gene regulatory processes, in a highly orthogonal manner.

[0090] Thus, herein and in the context of the present invention, the cell, in particular the cell in which the synthetic transcription factor of the present invention regulates transcription, is from the same species, e.g., the same mammalian species, as the mitochondrial DNA binding protein employed in the context of the present invention (and thus the DNA binding domain in the synthetic transcription factor of the present invention), i.e., is a homologous cell as described herein. Preferably, the DNA binding domain is derived from a human mitochondrial DNA binding protein (e.g., human MTERF1), and thus, the cell (i.e., the homologous cell) is preferably a human cell.

[0091] Preferably, herein, the synthetic transcription factors of the present invention do not substantially alter transcription other than that of the gene of interest, ie, the transcriptome, in the same cell.

[0092] Furthermore, the synthetic transcription factors of the present invention may not specifically bind to substantially any endogenous DNA sequence in the same cell. Preferably, the synthetic transcription factors of the present invention do not specifically bind to substantially any DNA sequence in the cell other than the promoter (particularly a promoter comprising a response element according to the present invention).

[0093] Furthermore, a DNA binding domain according to the invention (ie a DNA binding domain comprised in a synthetic transcription factor of the invention) preferably does not specifically bind to substantially any endogenous DNA sequence in the nucleus of a cell of the same species.

[0094] It is further possible that the synthetic transcription factors of the invention do not substantially compete with the mitochondrial DNA binding protein from which the DBD according to the invention is derived for sequence-specific DNA binding in the same cell.

[0095] Furthermore, the synthetic transcription factors of the present invention may not substantially interfere with the function of the mitochondrial DNA binding protein from which the DBD according to the present invention is derived.

[0096] It is also possible that the synthetic transcription factors of the present invention do not substantially interfere with the function of the protein from which the transcriptional regulatory domain according to the present invention is derived.

[0097] As described herein, transcriptional regulatory domains are generally involved in promoting or inhibiting transcription. Typically, a transcriptional regulatory domain comprises an interaction site with other proteins such as transcriptional co-regulators. Transcriptional co-regulators are proteins that interact with transcription factors to promote or inhibit transcription of a particular gene. The transcriptional co-regulators that activate gene transcription are called coactivators, while the transcriptional co-regulators that inhibit gene transcription are called co-repressors. As used herein, an activation domain can be associated with a coactivator rather than a co-repressor, while the inhibition domain used herein can be associated with a co-repressor instead.

[0098] The primary mechanism of action of transcriptional coregulators is to modify chromatin structure and, thereby, render the associated DNA more or less susceptible to transcription. In humans, dozens to hundreds of coregulators are known, depending on the confidence level achieved by characterizing the proteins acting as coregulators. For example, one class of transcriptional coregulators modifies chromatin structure by covalently modifying histones, while a second class, which is ATP-dependent, modifies the conformation of chromatin.

[0099] Typical coactivators include, inter alia, the pre-initiation complex (comprising, for example, transcription factor IID (TFIID)), the mediator complex, histone acetyltransferases, and chromatin remodeling complexes. Typical co-repressors include, inter alia, the polycomb repressive complex (e.g., PRC1 or PRC2), histone deacetylases, and histone methyltransferases.

[0100] In the context of this document and the present invention, a transcriptional regulatory domain is particularly capable of regulating the transcription of a gene, in particular when the transcriptional regulatory domain is part of, binds to or interacts with a DNA binding protein, wherein the DNA binding protein is capable of binding to or interacting with a regulatory sequence of the gene, such as a promoter or enhancer.

[0101] In the context of the present invention, since the transcription regulatory domain is contained in the synthetic transcription factor of the present invention together with the DNA binding domain according to the present invention, the transcription regulatory domain is particularly capable of interacting with the DNA binding domain and can therefore target gene regulatory sequences, such as the promoters described herein, and regulate the transcription of the corresponding gene.

[0102] In particular, the transcriptional regulatory domain according to the present invention may be capable of binding to and / or interacting with an RNA polymerase, preferably RNA polymerase II; at least one other transcription factor, such as a general transcription factor (e.g. TFIID); and / or at least one transcriptional co-regulator, such as a transcriptional co-activator (e.g. a Mediator complex and / or a histone acetyltransferase) and / or a transcriptional co-repressor (e.g. a Polycomb repressive complex or a histone deacetylase), as described herein.

[0103] In the context of this document and the present invention, a transcriptional regulatory domain may be (i) an activation domain or (ii) an inhibition domain, as described herein. Preferably, the transcriptional regulatory domain according to the present invention is an activation domain.

[0104] In particular, a transcriptional regulatory domain is capable of (i) promoting transcription of a gene (particularly when defined as an activation domain), or (ii) inhibiting transcription of a gene (particularly when defined as an inhibition domain). Preferably, a transcriptional regulatory domain according to the present invention is capable of (i) promoting transcription of a gene.

[0105] Therefore, the transcriptional regulatory domain according to the present invention can be (i) an activation domain that binds to and / or interacts with at least one coactivator to promote the transcription of a gene, or (ii) an inhibition domain that binds to and / or interacts with at least one co-repressor to inhibit the transcription of a gene; preferably, it is an activation domain as described in (i).

[0106] The present invention is not particularly limited with respect to the transcriptional regulatory domain, and many suitable activation domains and repression domains are readily available and can be used in the context of the present invention. The functionality of the transcriptional regulatory domain in the context of the synthetic transcription factors of the invention provided herein can be determined by conventional methods (e.g., as described in the accompanying examples and as, for example, Figure 4 For example, to assess the functionality of the activation domain in the context of the present invention, the following simple assay can be performed:

[0107] A reporter gene DNA construct (eg, a plasmid) comprising the promoter shown in SEQ ID NO: 110 operably linked to a gene encoding a detectable protein (eg, a fluorescent protein) is introduced (eg, transfected) into suitable cells.

[0108] In addition, a DNA construct encoding the transcriptional regulatory domain to be tested is introduced (e.g., transfected) into the same cells, wherein the transcriptional regulatory domain is fused to the human MTERF1 binding domain shown in SEQ ID NO: 1 under the control of a constitutive promoter such as EF1a (SEQ ID NO: 137), CMV (SEQ ID NO: 138) or UbC (SEQ ID NO: 139). Further, a construct encoding an additional protein that is constitutively expressed is introduced into the same cells, and the additional protein is detectable independently of the detectable protein encoded in the reporter gene construct. The amounts of the two detectable proteins are then quantitatively analyzed (e.g., by flow cytometry, microscopy, ELISA, or Western blotting). "Signal" is defined as the ratio between the reporter gene and the constitutively expressed detectable protein. As a negative control, the same assay is performed, but the construct encoding the transcriptional regulatory domain-MTERF1 fusion protein is not introduced into the cells.

[0109] The signal for the experimental condition is then compared to the signal for the negative control. If the signal is higher than the negative control, then the activation domain is determined to be functional in the context of the synthetic transcription factor of the invention (ie, it is able to promote transcription of the gene of interest).

[0110] Where the transcriptional regulatory domain is an inhibitory domain, the following assays can be performed to assess its functionality:

[0111] In a reporter gene DNA construct (e.g., a plasmid), a promoter to be repressed, such as EF1a (SEQ ID NO: 137), CMV (SEQ ID NO: 138), or UbC (SEQ ID NO: 139), additionally comprises a binding site for a transcription factor of the present invention within the promoter sequence or within 0 to 2000 bases adjacent to the 5' or 3' end of the promoter, and is operably linked to a gene encoding a detectable protein. This reporter gene DNA construct is introduced (e.g., transfected) into a suitable cell.

[0112] In addition, a DNA construct encoding the transcriptional regulatory domain to be tested is introduced (e.g., transfected) into the same cells, wherein the transcriptional regulatory domain is fused to the human MTERF1 binding domain (as defined in SEQ ID NO: 1) under the control of a constitutive promoter such as EF1a (SEQ ID NO: 137), CMV (SEQ ID NO: 138) or UbC (SEQ ID NO: 139). Further, a construct encoding an additional protein that is constitutively expressed is introduced into the same cells, wherein the additional protein is detectable independently of the detectable protein encoded in the reporter gene construct. The amounts of the two detectable proteins are then quantified (e.g., by flow cytometry, microscopy, ELISA, or Western blotting). "Signal" is defined as the ratio between the reporter gene and the constitutively expressed detectable protein. As a negative control, the same assay is performed, but the construct encoding the transcriptional regulatory domain-MTERF1 fusion protein is not introduced into the cells.

[0113] The signal under the experimental conditions is then compared to the signal of the negative control. If the signal is lower than the negative control, the repression domain is determined to be functional in the context of the synthetic transcription factor of the invention (i.e., it is able to block transcription of the gene of interest).

[0114] Furthermore, in cases where a transcriptional regulatory domain (e.g., an activation domain) is found to have little functionality on its own, it may provide good functionality when included multiple times in synthetic transcription factors and / or when combined with further transcriptional regulatory domains. For example, as described herein and illustrated in the accompanying examples (see, e.g., Figure 4 and Figure 9 ), the use of two FOXO domains strongly enhanced the transcriptional activation activity compared with a single FOXO domain.

[0115] The transcriptional regulatory domain of the present invention can be an activation domain comprising at least one transcriptional activation domain independently selected from the group consisting of: a RELA domain (e.g., RelA 430-551 ,RelA 342-551 ,RelA 361-551 or RelA 521-551 (ie, "TA1"), WW domain (WWC1 2-81 ), KRAB domain (ZNF473 5-48 ), NucRecCoAct domain (NCOA3 1045-1092 ), LMSTEN domain (MYB 251-330 ) and FoxoTAD (FOXO3 604-644 ).

[0116] In the context of this document and the present invention, a transcriptional regulatory domain, in particular an activation domain, may have at least 60%, preferably at least 70%, more preferably at least 80% sequence identity to a sequence selected from the group consisting of SEQ ID NO: 3, SEQ ID NO: 33, SEQ ID NO: 52, SEQ ID NO: 27, SEQ ID NO: 29, SEQ ID NO: 54, SEQ ID NO: 56, SEQ ID NO: 58, SEQ ID NO: 17, SEQ ID NO: 19, SEQ ID NO: 60, SEQ ID NO: 62, SEQ ID NO: 13, SEQ ID NO: 64, SEQ ID NO: 66 and SEQ ID NO: 68; or a transcriptional regulatory domain, in particular an activation domain, may comprise at least one sequence having at least 60%, preferably at least 70%, more preferably at least 80% sequence identity to a sequence selected from the group consisting of SEQ ID NO: 3, SEQ ID NO: 33, SEQ ID NO: 52, SEQ ID NO: 27, SEQ ID NO: 29, SEQ ID NO: 54, SEQ ID NO: 56, SEQ ID NO: 58, NO:56, SEQ ID NO:58, SEQ ID NO:17, SEQ ID NO:19, SEQ ID NO:60, SEQ ID NO:62, SEQ ID NO:13, SEQ ID NO:64, SEQ ID NO:66 and SEQ ID NO:68.

[0117] In some preferred embodiments of the present invention, the transcriptional regulatory domain, in particular the activation domain, has at least 60%, preferably at least 70%, and more preferably at least 80% sequence identity with a sequence selected from the group consisting of SEQ ID NO:3, SEQ ID NO:33, SEQ ID NO:52, SEQ ID NO:27, SEQ ID NO:29, SEQ ID NO:54, SEQ ID NO:56, and SEQ ID NO:58.

[0118] In a more preferred embodiment of the present invention, the transcriptional regulatory domain, in particular the activation domain, has at least 70%, preferably at least 80%, more preferably at least 90% sequence identity with a sequence selected from the group consisting of: SEQ ID NO: 3, SEQ ID NO: 33 and SEQ ID NO: 52.

[0119] In some preferred embodiments, the transcriptional regulatory domain of the present invention comprises a first, second, and / or third RELA transcriptional activation domain; wherein the first RELA transcriptional activation domain (TA1) has the sequence set forth in SEQ ID NO: 31, or a sequence having at least 80% sequence identity thereto; wherein the second RELA transcriptional activation domain has the sequence set forth in SEQ ID NO: 132, or a sequence having at least 80% sequence identity thereto; and wherein the third RELA transcriptional activation domain has the sequence set forth in SEQ ID NO: 134, or a sequence having at least 80% sequence identity thereto. Preferably, the transcriptional regulatory domain comprises at least a first RELA transcriptional activation domain, as described herein.

[0120] In addition, the transcriptional regulatory domain may comprise multiple copies, such as two or three copies, of the first, second and / or third RELA domains, preferably the first RELA domain (TA1). As illustrated in the accompanying examples, this may increase the expression level of the gene of interest (see, e.g. Figure 4 and 9 ).

[0121] Furthermore, for example in the context of these embodiments, the transcriptional regulatory domain according to the present invention may comprise a RELA subdomain A having a sequence as shown in SEQ ID NO: 3 or a sequence having at least 70%, preferably at least 80%, more preferably at least 90% sequence identity with SEQ ID NO: 3.

[0122] It has further been found in the context of the present invention that the expression level of a gene of interest can be increased by using a subunit with a larger RELA transcriptional activation domain (see, e.g. Figure 4 and 10 C).

[0123] Thus, for example, in the context of these "RELA" embodiments, the transcriptional regulatory domain according to the invention may comprise a RELA subdomain B having the sequence shown in SEQ ID NO:27 or a sequence having at least 70%, preferably at least 80%, more preferably at least 90% sequence identity to SEQ ID NO:27.

[0124] Furthermore, for example in the context of these embodiments, the transcriptional regulatory domain according to the present invention may comprise a RELA subdomain C having a sequence as shown in SEQ ID NO: 29 or a sequence having at least 60%, preferably at least 70%, more preferably at least 80% sequence identity with SEQ ID NO: 29.

[0125] In particular, the first, second and / or third RELA motif, the RELA subdomain A and / or the RELA subdomain B as described herein are comprised in the RELA subdomain C. Furthermore, the RELA subdomain A is specifically comprised in the RELA subdomain B.

[0126] In addition, the transcriptional regulatory domain of the present invention may comprise a FOXO3 transcriptional activation domain (FOXO TAD) having the sequence shown in SEQ ID NO: 17 or a sequence having at least 80% sequence identity with SEQ ID NO: 17.

[0127] It was further surprisingly found in the context of the present invention that the use of two copies of the FOXO TAD provides high expression levels of the gene of interest while maintaining a small size (and thus low gene payload); (see, e.g. Figure 4 and 9 ).

[0128] Therefore, in a further preferred embodiment, the transcriptional regulatory domain of the present invention may comprise multiple copies, for example, two or three copies of the FOXO3 transcriptional activation domain (FOXO TAD). Preferably, the transcriptional regulatory domain has the sequence shown in SEQ ID NO: 33 or a sequence having at least 70%, preferably at least 80%, and more preferably at least 90% sequence identity to SEQ ID NO: 33.

[0129] In addition, the transcriptional regulatory domain of the present invention may comprise a MYB transcriptional activation domain (LMSTEN) having the sequence shown in SEQ ID NO: 17 or a sequence having at least 80% sequence identity with SEQ ID NO: 17.

[0130] In some embodiments, the transcriptional regulatory domain comprises multiple copies, such as two or three copies of the MYB transcriptional activation domain (LMSTEN). Preferably, the transcriptional regulatory domain has a sequence as shown in SEQ ID NO: 196 or a sequence having at least 70%, preferably at least 80%, and more preferably at least 90% sequence identity to SEQ ID NO: 196.

[0131] In some embodiments, the transcriptional regulatory domain comprises two copies of (i) a first RELA transcriptional activation domain (TA1) having the sequence set forth in SEQ ID NO:31 or a sequence having at least 80% sequence identity to SEQ ID NO:31;

[0132] (ii) a FOXO3 transcriptional activation domain (FOXO TAD) having the sequence set forth in SEQ ID NO: 33 or a sequence having at least 80% sequence identity to SEQ ID NO: 33; or

[0133] (iii) a MYB transcriptional activation domain (LMSTEN) having the sequence shown in SEQ ID NO: 196 or a sequence having at least 80% sequence identity to SEQ ID NO: 196.

[0134] Furthermore, the transcriptional regulatory domain according to the present invention may comprise three copies of the first RELA transcriptional activation domain (TA1) having the sequence shown in SEQ ID NO: 31 or a sequence having at least 80% sequence identity with SEQ ID NO: 31.

[0135] Thus, in some embodiments, the transcriptional regulatory domain has a sequence as shown in SEQ ID NO:33, SEQ ID NO:35, SEQ ID NO:194 or SEQ ID NO:196, or a sequence having at least 70%, preferably at least 80%, and more preferably at least 90% sequence identity to any of these sequences.

[0136] It was further surprisingly found in the context of the present invention that two copies of the FOXO TAD enhance the transcriptional activity of synthetic TFs more than two copies of TA1 or two copies of LMSTEN (see Figure 9 ).

[0137] Thus, in the context of the present invention, a transcriptional regulatory domain comprising multiple copies of the same transcriptional activation domain preferably comprises at least two copies of the FOXO3 transcriptional activation domain (FOXO TAD) as represented by SEQ ID NO: 33 or a sequence having at least 70%, preferably at least 80%, more preferably at least 90% sequence identity to SEQ ID NO: 33.

[0138] In some preferred embodiments, the transcriptional regulatory domain of the present invention comprises at least two domains independently selected from the group consisting of: TA1 described herein, FOXO TAD described herein, and LMSTEN described herein. Preferably, the transcriptional regulatory domain comprises at least FOXO TAD. More preferably, the transcriptional regulatory domain comprises at least FOXO TAD, TA1, and LMSTEN described herein.

[0139] In some preferred embodiments, the transcriptional regulatory domain comprises (i) two FOXO TADs, (ii) a FOXO TAD and a LMSTEN, (iii) a LMSTEN and a TA1, (iv) a FOXO TAD and a TA1, or (v) a FOXO TAD, a LMSTEN, and a TA1, as described herein. More preferably, the transcriptional regulatory domain comprises a FOXO TAD, particularly option (i), (ii), (iv), or (v).

[0140] In a further preferred embodiment of the present invention, the transcriptional regulatory domain has a sequence shown in SEQ ID NO: 52, SEQ ID NO: 54, SEQ ID NO: 56 or SEQ ID NO: 58, or a sequence having at least 70%, preferably at least 80%, and more preferably at least 90% sequence identity with any of these sequences.

[0141] It has further been found in the context of the present invention that certain transcriptional regulatory domains, e.g. comprising a FOXO-TAD and further domains, in particular further FOXO-TAD, TA1 and / or LMSTEN domains, confer high transcriptional activity to synthetic transcription factors; see e.g. Figure 9 C.

[0142] Furthermore, it has been found that such combined transcriptional regulatory domains (e.g. comprising a FOXO-TAD and further domains, in particular a further FOXO-TAD, a TA1 and / or a LMSTEN domain) compensate or overcompensate for the N-terminally truncated MTERF1 DNA binding domain (e.g. MTERF1 104-399 (SEQ ID NO: 9)) slight loss of transcriptional activity; see, e.g. Figure 9 F.

[0143] Therefore, in a further preferred embodiment of the present invention, the transcriptional regulatory domain has a sequence as shown in SEQ ID NO: 52, SEQ ID NO: 33, SEQ ID NO: 54 or SEQ ID NO: 58, or a sequence having at least 70%, preferably at least 80%, and more preferably at least 90% sequence identity with any of these sequences. In addition, the DNA binding domain according to the present invention may comprise an MTERF1 subdomain A having a sequence as shown in SEQ ID NO: 9, or a sequence having at least 80%, preferably at least 90%, and more preferably at least 95% sequence identity with SEQ ID NO: 9.

[0144] In some most preferred embodiments of the present invention, the transcriptional regulatory domain has a sequence as shown in SEQ ID NO: 52, or a sequence having at least 70%, preferably at least 80%, and more preferably at least 90% sequence identity to SEQ ID NO: 52. In addition, the DNA binding domain according to the present invention may comprise an MTERF1 subdomain A having a sequence as shown in SEQ ID NO: 9, or a sequence having at least 80%, preferably at least 90%, and more preferably at least 95% sequence identity to SEQ ID NO: 9.

[0145] Furthermore, the transcriptional regulatory domain, in particular the activation domain, of the present invention may comprise a sequence having at least 60%, preferably at least 70%, more preferably at least 80% sequence identity to a sequence selected from the group consisting of SEQ ID NO: 60, SEQ ID NO: 62, SEQ ID NO: 13, SEQ ID NO: 64, SEQ ID NO: 66 and SEQ ID NO: 68, preferably in addition to the first, second and / or third RELA motif, such as the TA1, RELA subdomain A, B or C, FOXO TAD and / or LMSTEN, as described herein.

[0146] : 94, SEQ ID NO: 96 and SEQ ID NO: 98; or wherein the transcriptional regulatory domain, in particular the repression domain, of the present invention may comprise at least one sequence having at least 60%, preferably at least 70%, more preferably at least 80% sequence identity to a sequence selected from the group consisting of SEQ ID NO: 70, SEQ ID NO: 72, SEQ ID NO: 74, SEQ ID NO: 76, SEQ ID NO: 78, SEQ ID NO: 80, SEQ ID NO: 82, SEQ ID NO: 84, SEQ ID NO: 86, SEQ ID NO: 88, SEQ ID NO: 90, SEQ ID NO: 92, SEQ ID NO: 94, SEQ ID NO: 96 and SEQ ID NO: 98. ID NO:84, SEQ ID NO:86, SEQ ID NO:88, SEQ ID NO:90, SEQ ID NO:92, SEQ ID NO:94, SEQ ID NO:96 and SEQ ID NO:98.

[0147] Preferably, the transcriptional regulatory domain, in particular the repression domain, has at least 60%, preferably at least 70%, more preferably at least 80% sequence identity with a sequence selected from the group consisting of: preferably SEQ ID NO: 70, SEQ ID NO: 72, SEQ ID NO: 74 or SEQ ID NO: 76.

[0148] More preferably, the transcriptional regulatory domain, in particular the repression domain, has a sequence identity of at least 60%, preferably at least 70%, more preferably to SEQ ID NO:70.

[0149] In addition, the synthetic transcription factor of the present invention may further comprise a controllable domain, preferably a controllable destabilization domain or a controllable localization domain.

[0150] In particular, the controllable domains used herein and in the context of the present invention are controllable by a compound or light (e.g., infrared light, visible light, and / or UV light). Preferably, the compound is a small molecule. Preferably, the light is light of a specific wavelength or a specific range of wavelengths. Exposure of a synthetic transcription factor comprising a controllable domain to a compound or light can modulate, for example, promote or inhibit the transcriptional activity of the transcription factor.

[0151] Synthetic transcription factors comprising the controllable domains described herein may also be referred to as "inducible transcription factors" because their activity (or inactivation) can be induced by an external stimulus, particularly by a compound described herein or light. Many suitable controllable domains are known in the art, and any of them can be used in the context of the present invention. In addition, the induction of the activity (or inactivation) of a transcription factor is not limited to a particular mechanism. For example, certain controllable domains (e.g., the NS3 domain derived from the hepatitis C virus) act as destabilizing domains that destabilize proteins to which they are fused, where the fusion protein can be stabilized by small molecules that bind to the destabilizing domain. Other controllable domains act in the opposite manner, where the fusion protein can be destabilized by small molecules that bind to the controllable domain.

[0152] Again, other controllable domains (e.g., ERT2 derived from the human estrogen receptor) control the localization of proteins to which they are fused, where localization can be altered by small molecules that bind to the controllable domains. Synthetic transcription factors in the context of the present invention play a role, in particular, in the nucleus of the cell. Thus, when a transcription factor is localized from outside the nucleus (e.g., the cytoplasm) into the nucleus, its activity can be induced.

[0153] Some controllable domains, such as the FRB and FKBP domains, bind to each other in the presence of a compound or light, thereby bringing together the proteins they are fused to. If the functional protein consists of two parts, its activity is induced upon dimerization in the presence of a compound or light.

[0154] Other controllable domains known in the art or to be developed and methods of inducing the synthetic transcription factors of the invention may be used in the context of the present invention.

[0155] In some embodiments, the synthetic transcription factors of the present invention comprise a controllable destabilization domain and are stabilized or destabilized (preferably stabilized) by a compound or light. Preferably, the controllable destabilization domain comprises an NS3 domain having a sequence as set forth in SEQ ID NO: 158 or a sequence having at least 80%, preferably at least 90%, and more preferably at least 95% sequence identity to SEQ ID NO: 158. In particular, the synthetic transcription factors comprising the NS3 domain according to the present invention are stabilized by the small molecule grazoprevir.

[0156] In some embodiments, the synthetic transcription factor comprises a controllable localization domain and is localized to the nucleus or cytoplasm of a cell, preferably to the nucleus, by the compound or light. Preferably, the controllable localization domain comprises an ERT2 domain having the sequence set forth in SEQ ID NO:152, or a sequence having at least 80%, preferably at least 90%, and more preferably at least 95% sequence identity to SEQ ID NO:152. In particular, the synthetic transcription factor comprising the ERT2 domain is localized to the nucleus of a cell by 4-hydroxytamoxifen.

[0157] In a further embodiment, in particular when the DNA binding domain and the transcriptional regulatory domain according to the present invention are comprised in separate polypeptides, the controllable domain comprises a FRB domain and a FKBP domain, wherein the FRB domain has the sequence shown in SEQ ID NO: 140 or a sequence having at least 80%, preferably at least 90%, more preferably at least 95% sequence identity thereto, and / or wherein the FKBP domain has the sequence shown in SEQ ID NO: 142 or a sequence having at least 80%, preferably at least 90%, more preferably at least 95% sequence identity thereto. In particular, the FRB domain and the FKBP domain bind to each other in the presence of C16-(S)-7-methylindorapamycin.

[0158] In some embodiments, the synthetic transcription factor of the present invention comprises a synNotch core having a sequence as shown in SEQ ID NO: 160 or a sequence having at least 80%, preferably at least 90%, more preferably at least 95% sequence identity to SEQ ID NO: 160.

[0159] The synNotch core is a surface receptor and is fused to a single-chain variable fragment (scFv) at its N-terminus. Once the scFc::synNotch core::transcription factor fusion protein binds to its target (via the scFv), the transcription factor is cleaved, localized to the nucleus, and regulates (e.g., activates) gene expression; see Morsut (2016), Cell, 164(4).

[0160] In the context of this context and the present invention, it is further desirable that the immunogenicity of the synthetic transcription factor is low in the organism in which it is intended to be employed, such as in humans.

[0161] Furthermore, it was surprisingly found within the context of the present invention that human-derived synthetic transcription factors can be assembled which retain a high degree of orthogonality as described herein in human cells. In particular, in the human-derived synthetic transcription factors according to the invention, both the DNA binding domain according to the invention and the transcriptional regulatory domain according to the invention are derived from human proteins, preferably from human MTERF1 and at least one other human protein, respectively.

[0162] In the context of synthetic transcription factors comprising an activation domain, these synthetic transcription factors are also referred to herein as "human-derived transcriptional activator proteins" or "HumTAPs" for short.

[0163] Preferably, herein and in the context of the present invention, human-derived synthetic transcription factors such as HumTAP consist of only human protein-derived domains, more preferably they consist of only human protein domains.

[0164] What is generally described herein in the context of orthogonality, particularly in the context of humans (or human cells), also applies to the human-derived synthetic transcription factors described herein, at least not because orthogonality was tested with HumTAP in the accompanying Examples (see, e.g., Example 3).

[0165] To assess the immunogenicity of proteins, the inventors employed an assay to study the immune response to peptides derived from proteins of interest using primary PBMCs from normal human donors, as illustrated in the accompanying examples. It was found that the response of PBMCs to HumTAP (i.e., MTF)-derived peptides (specifically corresponding to the domain junction of HumTAP) was more similar to the response to self-peptides (i.e., non-immunogenic control peptides) than to immunogenic positive control peptides (see, e.g., Example 2). This demonstrates the low immunogenicity of the human-derived synthetic transcription factors of the present invention, such as HumTAP, in humans.

[0166] Therefore, the inventors have further developed a class of synthetic transcription factors made entirely of human protein subunits. The data shown in the accompanying examples show the favorable properties of GoI transcriptional activation, immunogenicity and orthogonality. As illustrated in the accompanying examples, the immunogenic potential of the synthetic transcription factors generated by using two completely human protein domains (i.e., a human DNA binding domain and a human transcriptional regulatory domain) is minimized. Therefore, it is expected that less or no local or systemic immunosuppression is required, and less or no additional genetic elements are required to offset the patient's immune response to gene or cell therapy (GCT) payloads (which contain genes encoding human-derived synthetic transcription factors (e.g., HumTAP proteins)). This can further lead to improved clinical performance of GCTs compared to the alternatives of the prior art.

[0167] Thus, the human-derived synthetic transcription factors according to the present invention successfully reconcile the requirements of high orthogonality (eg in human cells) and low immunogenicity (eg in humans).

[0168] Therefore, the development of human-derived synthetic transcription factors (such as HumTAP) with high orthogonality in human cells is a particularly great (and unexpected) achievement of the present inventors. In addition, the present invention is particularly valuable for synthetic gene circuits that can be used for therapies (such as gene therapy or cell therapy as described herein).

[0169] Thus, in the context of this document and the present invention, the one or more proteins from which the transcriptional regulatory domain is derived are preferably from the same species, e.g. the same mammalian species, as the mitochondrial DNA binding protein according to the present invention. Furthermore, the synthetic transcription factors of the present invention may consist essentially or completely of protein moieties from the same species, e.g. the same mammalian species.

[0170] Preferably, in the context of this document and the present invention, the mitochondrial DNA binding protein according to the invention (e.g., MTERF1) and one or more other proteins according to the invention from which the transcriptional regulatory domain is derived (e.g., RELA, FOXO3, and / or MYB) are of human origin. Furthermore, the synthetic transcription factors of the invention may consist essentially or entirely of human protein moieties.

[0171] In certain embodiments, the transcriptional regulatory domain of the present invention is derived from a single human protein. Preferably, the transcriptional regulatory domain, particularly the activation domain, has at least 90%, preferably at least 95%, and more preferably at least 99% sequence identity to SEQ ID NO: 3, SEQ ID NO: 27, or SEQ ID NO: 29, preferably to SEQ ID NO: 3.

[0172] Alternatively, the transcriptional regulatory domain, in particular the repression domain, of the present invention may have at least 90%, preferably at least 95%, more preferably at least 99% sequence identity to SEQ ID NO:70, SEQ ID NO:72, SEQ ID NO:74, SEQ ID NO:76, SEQ ID NO:78, SEQ ID NO:80, SEQ ID NO:82, SEQ ID NO:84, SEQ ID NO:86, SEQ ID NO:88, SEQ ID NO:90, SEQ ID NO:92, SEQ ID NO:94, SEQ ID NO:96 or SEQ ID NO:98, preferably to SEQ ID NO:70.

[0173] The synthetic transcription factors of the present invention may be substantially non-immunogenic in the mammalian species from which the mitochondrial DNA binding protein is derived. In the context of this document and the present invention, preferably, the synthetic transcription factors are substantially non-immunogenic in humans.

[0174] As described herein above, in some preferred embodiments, the synthetic transcription factor of the present invention is a fusion protein comprising a DNA binding domain according to the present invention and a transcription regulatory domain according to the present invention. As described herein above, the DNA binding domain and the transcription regulatory domain are linked to each other in the fusion protein directly via a direct peptide bond or via a peptide linker (and two corresponding peptide bonds at either end of the linker).

[0175] Therefore, the synthetic transcription factor fusion protein of the present invention may further comprise a peptide linker between the DNA binding domain and the transcriptional regulatory domain of the synthetic transcription factor.

[0176] Many suitable peptide linkers are known in the art, any of which can be employed in the context of the present invention, for example, in the context of the fusion proteins of the present invention. Suitable peptide linkers are, in particular, the two amino acid linker "GS", G4S as shown in SEQ ID NO: 118, AP6 as shown in SEQ ID NO: 120, cMycNLS as shown in SEQ ID NO: 122, "EAAAK" as shown in SEQ ID NO: 124, and the SV40 linker as shown in SEQ ID NO: 126. In addition, the cMycNLS and the SV40 linker can also function as nuclear localization sequences (NLS) and can therefore also be employed for the purposes herein and in the context of the present invention.

[0177] Furthermore, the present invention relates to a nucleic acid encoding a fusion protein of the invention comprising a DNA binding domain according to the invention and a transcriptional regulatory domain according to the invention (ie a fusion protein comprising a synthetic transcription factor of the invention).

[0178] Nucleic acid of the present invention can be DNA or RNA.In addition, nucleic acid can be single-stranded or double-stranded, for example dsDNA, ssRNA, ssDNA or dsRNA.Preferably, nucleic acid of the present invention comprises coding strand (i.e., sense strand).However, nucleic acid of the present invention also can refer to antisense strand (or even be made up of antisense strand), and therefore, it is characterized in that reverse complementary sequence.This is equally applicable to DNA construct of the present invention.

[0179] Furthermore, the nucleic acid of the present invention may be mRNA, such as mRNA contained in a lipid nanoparticle.

[0180] It was further found in the context of the present invention that DNA sequences codon-optimized for humans provide for higher expression of the gene of interest (see, e.g. Figure 2 ).

[0181] Therefore, in a preferred embodiment, the nucleic acid of the present invention comprises the DNA sequence shown in SEQ ID NO: 5, or a DNA sequence having at least 60%, preferably at least 70%, more preferably at least 80% sequence identity to SEQ ID NO: 5, wherein the DNA sequence encodes a DNA binding domain having at least 60%, preferably at least 70%, more preferably at least 80% sequence identity to SEQ ID NO: 1. Preferably, the DNA binding domain has the sequence shown in SEQ ID NO: 1.

[0182] Furthermore, the nucleic acid of the present invention may comprise a DNA sequence as shown in SEQ ID NO: 6, or a DNA sequence having at least 60%, preferably at least 70%, more preferably at least 80% sequence identity to SEQ ID NO: 6, wherein the DNA sequence encodes a transcriptional regulatory domain having at least 60%, preferably at least 70%, more preferably at least 80% sequence identity to SEQ ID NO: 3. Preferably, the transcriptional regulatory domain has the sequence shown in SEQ ID NO: 3.

[0183] Furthermore, the present invention relates to a DNA plasmid comprising the nucleic acid of the present invention. Preferably, the plasmid is suitable for expressing a synthetic transcription factor in a cell.

[0184] Furthermore, the present invention relates to a viral vector comprising a nucleic acid according to the invention or a plasmid according to the invention.As described herein, a nucleic acid may also be defined by the corresponding reverse complement sequence.

[0185] Suitable viral vectors employed in the context of the present invention include, inter alia, adeno-associated virus (AAV) vectors, lentiviral vectors, adenoviral vectors, herpes simplex virus vectors and VSV vectors. Preferably, the viral vector is an adeno-associated virus vector or a lentiviral vector.

[0186] Furthermore, the present invention relates to a cell comprising the nucleic acid of the invention, the plasmid of the invention and / or the viral vector of the invention.

[0187] In general, any cell can be used in the context of this article and the present invention. Preferably, the cell is a mammalian cell, preferably a human cell. For example, the cell can be an immune cell, such as a T cell, a B cell, or a NK cell. In an in vivo context, the cell can also preferably be a cancer cell or a tumor cell.

[0188] As mentioned above, in some embodiments, the synthetic transcription factor of the present invention comprises or consists of a first polypeptide and a second polypeptide, wherein the first polypeptide comprises a DNA binding domain according to the present invention and the second polypeptide comprises a transcriptional regulatory domain according to the present invention. In other words, the synthetic transcription factor can be a multimeric protein comprising a first polypeptide (i.e., a second amino acid chain) and a second polypeptide (i.e., a second amino acid chain), wherein the first polypeptide comprises a DNA binding domain according to the present invention and the second polypeptide comprises a transcriptional regulatory domain according to the present invention. In particular, the first polypeptide and the second polypeptide are capable of binding to and / or interacting with each other. The binding or interaction can be reversible and / or inducible, for example, by a compound as described herein or light.

[0189] Preferably, when the synthetic transcript is a multimeric protein as described above, it comprises a multimerization domain, wherein the multimerization domains of the first polypeptide and the second polypeptide are capable of binding to and / or interacting with each other. Preferably, the multimerization domain is a dimerization domain.

[0190] In some embodiments, the multimerization domain is a homodimerization domain. In that case, the multimerization domains of the first polypeptide and the second polypeptide are substantially identical to each other.

[0191] In a preferred embodiment, the multimerization domain is a heterodimerization domain. In that case, the multimerization domains of the first polypeptide and the second polypeptide are different from each other.

[0192] In some embodiments, (i) the multimerization domain of the first polypeptide comprises or consists of a SYNZIP1 domain, and the multimerization domain of the second polypeptide comprises or consists of a SYNZIP2 domain; or (ii) the multimerization domain of the first polypeptide comprises or consists of a SYNZIP2 domain, and the multimerization domain of the second polypeptide comprises or consists of a SYNZIP1 domain. In particular, the SYNZIP1 domain has the sequence set forth in SEQ ID NO: 154, or a sequence having at least 80%, preferably at least 90%, more preferably at least 95% sequence identity thereto; and the SYNZIP2 domain has the sequence set forth in SEQ ID NO: 156, or a sequence having at least 80%, preferably at least 90%, more preferably at least 95% sequence identity thereto.

[0193] In some embodiments, the multimerization domain is a controllable domain as described herein. In particular, the multimerization domain can be a dimerization domain, wherein a compound as described herein or light controls the dimerization of the multimerization domains of the first and second polypeptides, i.e., a controllable dimerization domain. Preferably, the first and second polypeptides bind and / or interact with each other in the presence of the small molecule or light.

[0194] In a preferred embodiment, the multimerization domain of the first polypeptide comprises or consists of an FKBP domain, and the multimerization domain of the second polypeptide comprises or consists of a FRB domain; or the multimerization domain of the first polypeptide comprises or consists of an FRB domain, and the multimerization domain of the second polypeptide comprises or consists of an FKBP domain. Preferably, the multimerization domain of the first polypeptide comprises or consists of an FKBP domain, and the multimerization domain of the second polypeptide comprises or consists of a FRB domain. In the presence of C16-(S)-7-methylindorapamycin, the first polypeptide and the second polypeptide may bind to and / or interact with each other, in particular, wherein C16-(S)-7-methylindorapamycin induces heterodimerization of the FKBP domain and the FRB domain.

[0195] In particular, the FKBP domain herein has the sequence shown in SEQ ID NO: 142, or a sequence having at least 80%, preferably at least 90%, and more preferably at least 95% sequence identity thereto. Furthermore, the FRB domain particularly has the sequence shown in SEQ ID NO: 140, or a sequence having at least 80%, preferably at least 90%, and more preferably at least 95% sequence identity thereto. Preferably, the DNA binding domain is located at the N-terminus of the FKBP domain in the first polypeptide, and / or the transcriptional regulatory domain is located at the C-terminus of the FRB domain in the second polypeptide.

[0196] In addition, the DNA binding domain and the FKBP domain may be connected to each other via a first peptide linker, and / or the transcriptional regulatory domain and the FRB domain may be connected to each other via a second peptide linker. The first and second peptide linkers may be independently selected from the group consisting of: a cMyc NLS linker as shown in SEQ ID NO: 122, a 6AP (AP6) linker as shown in SEQ ID NO: 120, an AP8 linker as shown in SEQ ID NO: 144, a G4S linker as shown in SEQ ID NO: 118, an EAAAK3 linker as shown in SEQ ID NO: 146, an EAAAK2 linker as shown in SEQ ID NO: 148, and a G4S4 linker as shown in SEQ ID NO: 150. It is worth noting that the terms "6AP" and "AP6" are used interchangeably herein.

[0197] In a preferred embodiment, the first peptide linker is the cMyc NLS linker shown in SEQ ID NO: 122 or the 6AP (AP6) linker shown in SEQ ID NO: 120, and / or the second peptide linker is the 6AP (AP6) linker shown in SEQ ID NO: 120.

[0198] In a further preferred embodiment, for example in the context of an FKBP domain and a FRB domain, the transcriptional regulatory domain has a sequence as set forth in SEQ ID NO: 52, SEQ ID NO: 29, SEQ ID NO: 33, SEQ ID NO: 54 or SEQ ID NO: 58, or a sequence having at least 70%, preferably at least 80%, and more preferably at least 90% sequence identity to any of these sequences. More preferably, the transcriptional regulatory domain has a sequence as set forth in SEQ ID NO: 52 or SEQ ID NO: 29, or a sequence having at least 70%, preferably at least 80%, and more preferably at least 90% sequence identity to SEQ ID NO: 52 or SEQ ID NO: 29.

[0199] In some preferred embodiments, the first polypeptide of the synthetic transcription factor comprises the sequence shown in SEQ ID NO: 174, 176, 178, 180, 182, 184, 186, 188 or 192, or a sequence having at least 70%, preferably at least 80%, more preferably at least 90% sequence identity to SEQ ID NO: 174, 176, 178, 180, 182, 184, 186, 188 or 192; and / or the second polypeptide of the synthetic transcription factor comprises the sequence shown in SEQ ID NO: 162, 164, 166, 168, 170, 172 or 190, or a sequence having at least 70%, preferably at least 80%, more preferably at least 90% sequence identity to SEQ ID NO: 162, 164, 166, 168, 170, 172 or 190.

[0200] Furthermore, the present invention relates to a combination of nucleic acids encoding a synthetic transcription factor according to the present invention, wherein one nucleic acid encodes a DNA binding domain according to the present invention and another nucleic acid encodes a transcription regulatory domain according to the present invention. In particular, the nucleic acid encoding the DNA binding domain according to the present invention corresponds to a first nucleic acid encoding a first polypeptide according to the present invention, as described herein. Furthermore, the nucleic acid encoding the transcription regulatory domain according to the present invention particularly corresponds to a second nucleic acid encoding a second polypeptide according to the present invention, as described herein.

[0201] Therefore, the present invention further relates to a combination of nucleic acids encoding the synthetic transcription factors of the present invention, which comprises a first polypeptide and a second polypeptide as described herein, wherein the combination of nucleic acids comprises a first nucleic acid and a second nucleic acid, wherein the first nucleic acid encodes the first polypeptide and the second nucleic acid encodes the second polypeptide.

[0202] Furthermore, the nucleic acids of the invention may be contained in a plurality of plasmids, viral vectors, cells or kits, as described herein. In certain embodiments, the combination of nucleic acids according to the invention refers to a kit comprising the combination of nucleic acids.

[0203] In certain embodiments of the combination of nucleic acids, one nucleic acid has the DNA sequence set forth in SEQ ID NO: 5, or a DNA sequence having at least 60%, preferably at least 70%, and more preferably at least 80% sequence identity to SEQ ID NO: 5, wherein the DNA sequence encodes a DNA binding domain having at least 60%, preferably at least 70%, and more preferably at least 80% sequence identity to SEQ ID NO: 1. Preferably, the DNA binding domain has the sequence set forth in SEQ ID NO: 1.

[0204] Furthermore, the additional nucleic acid in the combination may have the DNA sequence shown in SEQ ID NO: 6, or a DNA sequence having at least 60%, preferably at least 70%, more preferably at least 80% sequence identity to SEQ ID NO: 6, wherein the DNA sequence encodes a transcriptional regulatory domain having at least 60%, preferably at least 70%, more preferably at least 80% sequence identity to SEQ ID NO: 3. Preferably, the transcriptional regulatory domain has the sequence shown in SEQ ID NO: 3.

[0205] Furthermore, the present invention relates to a DNA construct comprising a promoter (P), said promoter comprising a response element and a minimal promoter, characterized in that said response element comprises an MTERF1 binding site, said binding site having (in particular comprising or consisting of) the sequence shown in SEQ ID NO: 42 or SEQ ID NO: 200, or a sequence having at least 60%, preferably at least 80%, more preferably at least 90% sequence identity with SEQ ID NO: 42 or SEQ ID NO: 200.

[0206] SEQ ID NO:200 corresponds to the reverse complementary (ie, antisense) sequence of SEQ ID NO:42.

[0207] In the context of this document and the present invention, sequences are read from the 5' to 3' end, i.e., in the sense read. Thus, a DNA construct comprising an MTERF1 binding site in the sense direction relative to the minimal promoter, for example, comprises the sequence shown in SEQ ID NO:42, or a sequence having at least 60%, preferably at least 80%, and more preferably at least 90% sequence identity to SEQ ID NO:42. Following the same logic, a DNA construct comprising an MTERF1 binding site in the antisense direction relative to the minimal promoter, in particular, comprises the sequence shown in SEQ ID NO:200, or a sequence having at least 60%, preferably at least 80%, and more preferably at least 90% sequence identity to SEQ ID NO:200.

[0208] Nevertheless, in particular in the context of a dsDNA construct according to the present invention, SEQ ID NO:42 or a sequence having at least 60%, preferably at least 80%, more preferably at least 90% sequence identity to SEQ ID NO:42 may be included in one strand of the dsDNA construct in the 5' to 3' direction and the sequence of the minimal promoter described herein may be included in the other strand of the dsDNA construct in the 5' to 3' direction.

[0209] Preferably, herein, in particular in the context of the DNA constructs of the invention provided herein, the MTERF1 binding site comprises or consists of the sequence shown in SEQ ID NO: 42 or a sequence having at least 60%, preferably at least 80%, more preferably at least 90% sequence identity thereto, in particular MTERF1 binding in the positive sense relative to the minimal promoter.

[0210] Preferably, herein, for example, in the context of a DNA construct of the present invention, the response element of the present invention consists of: (i) one or more copies, for example 2 to 50 copies (preferably 2 to 5 copies) of the MTERF1 binding site, wherein the multiple copies are directly adjacent to each other; or (ii) multiple copies, for example 2 to 50 copies (preferably 2 to 5 copies) of the MTERF1 binding site, and a BS-BS spacer between at least two, preferably all consecutive copies of the binding site. In particular, the spacer has a length of 1 to 1000 nucleotides, preferably 1 to 100 nucleotides, more preferably 1 to 10 nucleotides, for example 1 to 6 nucleotides. Preferably, the BS-BS spacer consists of 1 to 10 nucleotides at the 5' end of the sequence shown in SEQ ID NO: 198, i.e., the first 1, 2, 3, 4, 5, 6, 7, 8 or 9 or all nucleotides at the 5' end of the sequence shown in SEQ ID NO: 198. Furthermore, in cases where the MTERF1 binding site is in the sense orientation relative to the minimal promoter, as described herein, the BS-BS spacer may consist of 1, 4, 5 or 8 nucleotides, preferably the first 1, 4, 5 or 8 nucleotides at the 5' end of the sequence shown in SEQ ID NO: 198.

[0211] Furthermore, instead of the exemplary 2 to 50 copies, other ranges may be considered, such as 2 to 25 copies, 2 to 15 copies, 2 to 10 copies, or preferably 2 to 5 copies.

[0212] In some preferred embodiments, the response element consists of multiple copies, such as 2 to 15 copies, of the MTERF1 binding site that are directly adjacent to each other.

[0213] Furthermore, the DNA constructs of the present invention may, for example, have a length of up to about 10 6 , preferably up to about 10 5 More preferably, up to about 10,000 nucleotides.

[0214] Preferably, herein, for example, in the context of a DNA construct of the invention, the response element of the invention and the minimal promoter are separated from each other by at most about 2000 nucleotides (i.e., separated by a RE-minP spacer of at most about 2000 nucleotides in length), preferably at most about 200 nucleotides, more preferably at most about 20 nucleotides, such as about 6 or 8 nucleotides. Preferably, the RE-minP spacer consists of 1 to 10 nucleotides at the 5' end of the sequence shown in SEQ ID NO: 199; and preferably, the MTERF1 binding site comprises or consists of the sequence shown in SEQ ID NO: 42 or a sequence having at least 60%, preferably at least 80%, more preferably at least 90% sequence identity thereto, in particular in the positive sense relative to the minimal promoter.

[0215] Furthermore, herein, for example in the context of the DNA construct of the present invention, the minimal promoter may be located 3' or 5' to the response element. Preferably, the minimal promoter is located 3' to the response element.

[0216] In the context of this document and the present invention, a minimal promoter may be a minimal TATA box having a sequence identity of at least 60%, preferably at least 80%, more preferably at least 90% to SEQ ID NO: 103.

[0217] Furthermore, the minimal promoter may be a minimal CMV promoter having at least 60%, preferably at least 80%, more preferably at least 90% sequence identity to SEQ ID NO:136.

[0218] In particular, in the context of this text and the present invention, the promoter (P) according to the invention is capable of binding to the synthetic transcription factor according to the invention. In particular, the response element according to the invention is capable of binding to the DNA binding domain according to the invention.

[0219] In particular, in the present context and in the context of the present invention, the DNA binding domain according to the invention binds to the response element according to the invention in a sequence-specific manner.

[0220] Preferably, in the context of this document and the present invention, the DNA construct of the present invention further comprises at least one gene of interest. Preferably, the at least one gene of interest is located 3' to the minimal promoter.

[0221] In preferred embodiments, the gene of interest encodes a pro-cell death protein such as hBAX or HSV-TK, an immunostimulatory cytokine such as IL-2 or IL-12, and / or an antigen receptor such as CAR or TCR, as described herein.

[0222] Preferably, in the context of this invention and in the present invention, the promoter (P) according to the invention is operably linked to at least one gene of interest, as described herein.

[0223] In particular, when the promoter (P), in particular the response element of the invention, is bound by the synthetic transcription factor of the invention in a cell, for example in the nucleus of a human cell, at least one of the genes of interest is transcribed.

[0224] In certain embodiments, the DNA construct of the present invention does not comprise the sequence shown in SEQ ID NO:117 or a sequence having at least 90% sequence identity to SEQ ID NO:117.

[0225] As illustrated in the accompanying examples, the strength of the promoter (P) contained in the DNA construct of the present invention can be adjusted as desired, i.e., stronger or weaker promoter variants can be used; see Figure 14 and 15 and SEQ ID NOs: 201 to 1190.

[0226] Therefore, the DNA construct of the present invention may comprise a sequence selected from the group consisting of SEQ ID NOs: 201 to 1190. In addition, the spacer sequence therein, i.e., the consecutive nucleotides (particularly 1-10 nucleotides in length) that do not belong to the MTERF1 binding site (SEQ ID NO: 42) or the minimal promoter sequence (SEQ ID NO: 103), may be replaced by other corresponding spacer sequences of the same length.

[0227] like Figure 15 As shown in FIG, and as reflected in SEQ ID NO: 300, a particular promoter variant of particular interest has two MTERF1 binding sites in positive sense relative to the minimal promoter, and an 8-nucleotide spacer between the response element (or the 3'-most MTERF1 binding site) and the minimal promoter sequence, the two MTERF1 binding sites being directly adjacent to each other. This promoter variant is relatively strong while having a relatively small size.

[0228] Thus, in some preferred embodiments, the response element comprises or consists of two copies of an MTERF1 binding site, each copy having a sequence as shown in SEQ ID NO:42 or a sequence having at least 60%, preferably at least 80%, and more preferably at least 90% sequence identity to SEQ ID NO:42; wherein the two copies of the MTERF1 binding site are directly adjacent to each other; and wherein the response element and the minimal promoter are separated by 8 nucleotides from each other.

[0229] In some preferred embodiments, the DNA construct of the present invention comprises the sequence shown in SEQ ID NO:300.

[0230] like Figure 15 A specific promoter variant of further particular interest, as shown in FIG, and reflected in SEQ ID NO: 693, has five MTERF1 binding sites in a positive orientation relative to the minimal promoter, and a 10-nucleotide spacer between the response element (or 3'-most MTERF1 binding site) and the minimal promoter sequence, wherein the five MTERF1 binding sites are separated from each other by an 8-nucleotide spacer. This promoter variant is particularly strong.

[0231] Therefore, in other preferred embodiments, the response element consists of 5 copies of the MTERF1 binding site, which are separated from each other by 8 nucleotides, and wherein each MTERF1 binding site has the sequence shown in SEQ ID NO:42 or a sequence having at least 60%, preferably at least 80%, more preferably at least 90% sequence identity with SEQ ID NO:42; and wherein the response element and the minimal promoter are separated from each other by 10 nucleotides.

[0232] In some preferred embodiments, the DNA construct of the present invention comprises the sequence shown in SEQ ID NO:693.

[0233] In addition, the DNA constructs of the present invention can be single-stranded or double-stranded, i.e., dsDNA, ssDNA, or dsRNA. In addition, the present invention relates to RNA corresponding to the DNA constructs of the present invention (i.e., with the exception of uracil relative to thymine, having the same sequence). The corresponding RNA can also be single-stranded or double-stranded, i.e., ssRNA or dsRNA. In addition, the RNA can be mRNA, for example, contained in lipid nanoparticles, as described herein.

[0234] Preferably, DNA construct of the present invention (or corresponding RNA) comprises coding strand (i.e. sense strand). However, DNA construct of the present invention (or corresponding RNA) also can refer to antisense strand (even being made up of antisense strand), and therefore, it is characterized in that reverse complementary sequence.

[0235] Therefore, the present invention further relates to a single-stranded or double-stranded nucleic acid comprising the sense strand of a DNA construct of the present invention and / or the antisense strand of a DNA construct of the present invention.

[0236] Furthermore, the present invention relates to a single-stranded or double-stranded nucleic acid, such as DNA or RNA, comprising a sequence corresponding to the sense strand of a DNA construct of the invention and / or a sequence corresponding to the antisense strand of a DNA construct of the invention.

[0237] Furthermore, the present invention relates to a plasmid comprising the DNA construct of the present invention.

[0238] The present invention also relates to a viral vector comprising the DNA construct, the corresponding RNA or the corresponding plasmid according to the invention.

[0239] Furthermore, the present invention relates to a cell comprising the DNA construct of the invention.Regarding the cell, the same applies as disclosed herein in the context of the nucleic acid of the invention encoding the fusion protein of the invention (ie the synthetic transcription factor).

[0240] Furthermore, the present invention relates to a system (particularly a transcription system) comprising: (i) a synthetic transcription factor according to the invention, a corresponding nucleic acid according to the invention, a corresponding DNA plasmid according to the invention and / or a corresponding viral vector according to the invention; and (ii) a DNA construct according to the invention, a DNA plasmid according to the invention and / or a viral vector according to the invention. Preferably, the system comprises a synthetic transcription factor according to the invention and a DNA construct according to the invention.

[0241] Furthermore, the system can be an engineered genetic network, such as a biological computational circuit.

[0242] In particular, the present system is suitable for regulating the transcription of at least one gene of interest and can be used for this purpose. Preferably, the gene of interest is contained in a DNA construct of the present invention, as described herein. Preferably, the gene of interest encodes a cell death-promoting protein such as hBAX or HSV-TK, an immunostimulatory cytokine such as IL-2 or IL-12, and / or an antigen receptor such as CAR or TCR, as described herein.

[0243] In addition, the present invention relates to cells comprising the synthetic transcription factor of the present invention and the DNA construct of the present invention. Preferably, the cell is a mammalian cell, preferably a human cell. For example, when the gene of interest encodes an antigen receptor such as a CAR or TCR as described herein, the cell according to the present invention can be an immune cell, such as a T cell, a B cell or a NK cell. In the in vivo context, the cell may also preferably be a cancer cell or a tumor cell.

[0244] Furthermore, the present invention relates to a kit comprising a synthetic transcription factor of the invention, a corresponding nucleic acid of the invention, a corresponding DNA plasmid of the invention, a corresponding viral vector of the invention, a combination of nucleic acids of the invention, a DNA construct of the invention, a single-stranded or double-stranded nucleic acid of the invention, a DNA plasmid of the invention, a viral vector of the invention and / or a system of the invention.

[0245] In certain embodiments, the kit comprises (i) a nucleic acid of the invention (i.e., encoding a fusion protein / synthetic transcription factor of the invention), a corresponding DNA plasmid of the invention, or a corresponding viral vector of the invention, and (ii) a DNA construct of the invention, a corresponding DNA plasmid of the invention, or a corresponding viral vector of the invention.

[0246] Furthermore, the present invention relates to a pharmaceutical composition comprising a synthetic transcription factor of the invention, a corresponding nucleic acid of the invention, a corresponding DNA plasmid of the invention, a corresponding viral vector of the invention, a combination of nucleic acids of the invention, a DNA construct of the invention, a single-stranded or double-stranded nucleic acid of the invention, a corresponding DNA plasmid of the invention (i.e. corresponding to a DNA construct of the invention), a corresponding viral vector of the invention, a system of the invention or any cell of the invention.

[0247] In certain embodiments, the pharmaceutical composition comprises (i) a nucleic acid of the invention (corresponding to a fusion protein / synthetic transcription factor of the invention), a corresponding DNA plasmid of the invention, or a corresponding viral vector of the invention, and (ii) a DNA construct of the invention, a corresponding DNA plasmid of the invention, or a corresponding viral vector of the invention.

[0248] In certain embodiments, a pharmaceutical composition comprises a cell of the invention comprising a system of the invention.

[0249] In addition, the pharmaceutical composition of the present invention may further comprise a pharmaceutically acceptable excipient.

[0250] In addition, the pharmaceutical compositions of the present invention can be used to treat diseases in which target cells are killed and / or manipulated. In particular, the treatment may involve a cancer cell classifier circuit as described and / or mentioned herein. Preferably, in this context, at least one gene of interest may encode a pro-cell death protein such as hBAX or HSV-TK, an immunostimulatory cytokine such as IL-2 or IL-12, and / or an antigen receptor such as CAR or TCR.

[0251] In addition, the pharmaceutical compositions of the present invention can be used in methods for treating tumors or cancer. In particular, the treatment may involve the cancer cell classifier circuit described herein. Preferably, in this context, at least one gene of interest encodes a pro-cell death protein such as hBAX or HSV-TK, an immunostimulatory cytokine such as IL-2 or IL-12, and / or an antigen receptor such as CAR or TCR.

[0252] Furthermore, the nucleic acids of the invention (corresponding to the fusion proteins / synthetic transcription factors of the invention), the corresponding DNA plasmids of the invention, the corresponding viral vectors of the invention, combinations of nucleic acids of the invention, DNA constructs of the invention, the corresponding single-stranded or double-stranded nucleic acids of the invention, the corresponding DNA plasmids of the invention, the corresponding viral vectors of the invention or the systems of the invention can be used for gene therapy.

[0253] In addition, the cells of the present invention can be used for cell therapy. Preferably, the cells are T cells, such as CART cells, and the cell therapy is T cell therapy, such as CAR T cell therapy.

[0254] Furthermore, the nucleic acids of the invention (corresponding to the fusion proteins / synthetic transcription factors of the invention), the corresponding DNA plasmids of the invention, the corresponding viral vectors of the invention, combinations of nucleic acids of the invention, DNA constructs of the invention, the corresponding single-stranded or double-stranded nucleic acids of the invention, the corresponding DNA plasmids of the invention, the corresponding viral vectors of the invention, or the systems of the invention can be used as part of or in combination with engineered gene networks, in particular biological computational circuits.

[0255] Furthermore, the nucleic acids of the invention (corresponding to the fusion proteins / synthetic transcription factors of the invention), the corresponding DNA plasmids of the invention, the corresponding viral vectors of the invention, combinations of nucleic acids of the invention, DNA constructs of the invention, the corresponding single-stranded or double-stranded nucleic acids of the invention, the corresponding DNA plasmids of the invention, the corresponding viral vectors of the invention or the systems of the invention can be used to transcribe a gene of interest in vitro or in vivo (e.g., in cells in vitro or in vivo).

[0256] As illustrated in the accompanying examples, the inventors further generated libraries of promoter variants and developed methods for screening promoters optimized for binding to transcription factors; see Examples 6 and Figures 11 to 15 . The inventors have found, in particular, promoter variants that are relatively small in size and provide relatively high transcriptional activity, i.e., they are relatively strong (e.g., SEQ ID NO: 300). In addition, the inventors have found, in particular, particularly strong promoter variants (e.g., SEQ ID NO: 693). In addition, the inventors have surprisingly found that, in addition to the number of transcription factor (TF) binding sites and the orientation of the transcription factor binding sites relative to the minimal promoter, the presence or length of the spacer between the TF binding sites and the spacer between the 3' TF binding site and the minimal promoter (in particular, the spacer between the 3' TF binding site and the minimal promoter) affects the strength of the promoter.

[0257] Thus, the present invention further relates to a library of DNA constructs comprising at least about 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 150, 200, 250, 300, 350, 400, 450, 500, 550, 600, 650, 700, 750, 800, 850, 900, 950, 1000, 2000, 3000, 4000, 5000, 6000, 7000, 8000, 9000, 10000, 50000, 100000, 500000 or 10000000 different DNA constructs, preferably at least about 100, 500, 900, 950, 990 or 1000 different DNA constructs, wherein each DNA construct in the library comprises

[0258] (i) a promoter (P) consisting of a response element (RE), a minimal promoter (minP) located 3' to the response element, and an optional RE-minP spacer between the response element and the minimal promoter; wherein each response element consists of one or more copies of a transcription factor binding site (BS) and an optional BS-BS spacer between at least two, preferably all consecutive copies of the binding site;

[0259] (ii) an export sequence (preferably located 3' to the minimal promoter); and

[0260] Wherein all DNA constructs in the library differ from each other in their promoter (P) sequence. Optionally, each DNA construct in the library comprises a unique barcode sequence that distinguishes all DNA constructs in the library from each other. Means and methods for barcoding are well known in the art.

[0261] The DNA constructs in the library, particularly the promoters, response elements, transcription factor binding sites, minimal promoters and spacers can be designed as described herein in the context of the DNA constructs of the invention.

[0262] Therefore, the transcription factor binding site is preferably an MTERF1 binding site, which comprises or consists of the sequence shown in SEQ ID NO: 42 or SEQ ID NO: 200 (preferably SEQ ID NO: 42), or a sequence having at least 60%, preferably at least 80%, and more preferably at least 90% sequence identity thereto. Preferably, the MTERF1 binding site consists of the sequence shown in SEQ ID NO: 42.

[0263] The DNA constructs in the library can be identical to each other except for the promoter sequence and the optional barcode sequence. In addition, the minimal promoters in different DNA constructs (particularly in different promoters) can be identical to each other. In addition, the transcription factor binding sites in different DNA constructs (particularly in different promoters) can have the same sequence in sense or antisense orientation relative to the minimal promoter, preferably the sequence shown in SEQ ID NO: 42 or SEQ ID NO: 200, respectively.

[0264] In some embodiments, at least 20%, 30%, 40% or 50% of the promoters of the DNA constructs differ from each other in (i) binding site copy number and / or (ii) presence or length of the RE-minP spacer; and / or at least 80%, at least 90% or all of the promoters of the DNA constructs differ from each other in at least one parameter selected from the group consisting of: (i) binding site copy number, (ii) presence or length of the RE-minP spacer, (iii) presence or length of the BS-BS spacer, and (iv) orientation of the binding site sequence relative to the minimal promoter in sense or antisense.

[0265] In a preferred embodiment of the library, the minimal promoter is a minimal TATA box having a sequence identity of at least 60%, preferably at least 80%, more preferably at least 90% to SEQ ID NO:103.

[0266] In some embodiments of the library, the minimal promoter is a minimal CMV promoter having at least 60%, preferably at least 80%, more preferably at least 90% sequence identity to SEQ ID NO:136.

[0267] In particular, the promoter (P) and in particular the response element in the DNA constructs of the library are capable of (or suspected of being capable of) binding to the synthetic transcription factors of the present invention.

[0268] In a further aspect, the present invention relates to a method for optimizing a promoter for binding to a transcription factor, the method comprising the steps of:

[0269] a) preparing a library of DNA constructs according to the invention,

[0270] b) combining the library of DNA constructs with said transcription factors in cells or in an in vitro transcription system, preferably in cells,

[0271] c) determining the transcriptional activity of the promoter of each DNA construct in the library, preferably by determining the amount of mRNA produced by each DNA construct in the library, in particular, wherein said mRNA comprises a sequence corresponding to the output sequence of the DNA construct, and

[0272] d) Selecting a promoter based on its transcriptional activity, thereby obtaining a promoter optimized for binding to said transcription factor.

[0273] Preferably, the transcriptional activity of the promoter of each DNA construct in the library is determined by RNA sequencing, preferably by next generation RNA sequencing, for example as explained in Example 6. More preferably, the method comprises a massively parallel reporter gene assay, for example as described in Example 6. In a preferred embodiment of the method, the transcription factor is a synthetic transcription factor according to the present invention.

[0274] Sequence Listing

[0275] The following sequences and SEQ ID NOs refer to the SEQ ID NOs described herein and in the context of the present invention:

[0276]

[0277]

[0278]

[0279]

[0280]

[0281]

[0282]

[0283]

[0284]

[0285]

[0286]

[0287]

[0288]

[0289]

[0290]

[0291]

[0292]

[0293]

[0294]

[0295]

[0296]

[0297]

[0298]

[0299]

[0300]

[0301] SEQ ID NO Amino acid sequence SEQ ID NO: 144, DNA sequence SEQ ID NO: 145 Construct AP8 illustrate connector organism Synthetic constructs Amino acid sequence APAPAPAPAPAPAPAP DNA sequence GCACCAGCACCAGCTCCTGCCCCGGCACCAGCGCCAGCACCAGCACCA

[0302] SEQ ID NO Amino acid sequence SEQ ID NO: 146, DNA sequence SEQ ID NO: 147 Construct EAAAK3 illustrate connector organism Synthetic constructs Amino acid sequence EAAAKEAAAKEAAAK DNA sequence GAAGCTGCGGCAAAAGAAGCCGCTGCGAAGGAAGCGGCAGCAAAA

[0303] SEQ ID NO Amino acid sequence SEQ ID NO: 148, DNA sequence SEQ ID NO: 149 Construct EAAAK2 illustrate connector organism Synthetic constructs Amino acid sequence EAAAKEAAAK DNA sequence GAAGCAGCGGCAAAAGAGGCAGCGGCAAAA

[0304] SEQ ID NO Amino acid sequence SEQ ID NO: 150, DNA sequence SEQ ID NO: 151 Construct G4S4 illustrate connector organism Synthetic constructs Amino acid sequence GGGGSGGGGSGGGGSGGGGS DNA sequence See the attached sequence listing according to WIPOSt.26

[0305]

[0306] SEQ ID NO Amino acid sequence SEQ ID NO: 154, DNA sequence SEQ ID NO: 155 Construct SYNZIP1 organism Synthetic constructs Amino acid sequence NLVAQLENEVASLENENETLKKKNLHKKDLIAYLEKEIANLRKKIEE DNA sequence See the attached sequence listing according to WIPOSt.26

[0307] SEQ ID NO Amino acid sequence SEQ ID NO: 156, DNA sequence SEQ ID NO: 157 Construct SYNZIP2 organism Synthetic constructs Amino acid sequence ARNAYLRKKIARLKKDNLQLERDEQNLEKIIANLRDEIARLENEVASHEQ DNA sequence See the attached sequence listing according to WIPOSt.26

[0308]

[0309]

[0310]

[0311]

[0312]

[0313]

[0314]

[0315]

[0316] The sequences SEQ ID NO: 1-200 are also shown in the attached sequence listing according to WIPOS t.26.

[0317] The following promoter (sensor) sequences, SEQ ID NOs: 201-1190, are shown in the attached sequence listing according to WIPOSit. 26. These promoter sequences have the following symbols: (a) orientation of the binding site (BS) relative to the minimal promoter (TATA) - (b) number of BSs - (c) BS-BS distance - (d) BS-TATA distance.

[0318]

[0319]

[0320]

[0321]

[0322]

[0323]

[0324]

[0325]

[0326]

[0327]

[0328]

[0329] The invention is further characterized by the following figures, illustrations and the following non-limiting examples.

[0330] BRIEF DESCRIPTION OF THE DRAWINGS

[0331] Figure 1 MTF activity against various reporter gene constructs in human cell lines.

[0332] (A) DNA sequence of the response element (RE) repeat and spacer region, which is repeated n times. Figure 1In Figure 2, the term "response element" (RE) corresponds to the term "binding site" as used herein, in particular the MTERF1 binding site defined by SEQ ID NO:42. (B) Dependence of gene of interest (GoI) expression on the number of REs in the reporter construct in the HeLa human cell line. Microscope images of HeLa cells showing expression of the mCerulean GoI (top) and expression of the reporter control mCherry transfected (bottom). The names of the fluorescent channels are indicated on the left. The number of RE repeats in each promoter construct is indicated above each pair of images. (C) Quantitative analysis of mCerulean levels (normalized to mCherry levels) in HeLa cells transfected with reporter constructs with different numbers of RE-spacer repeats as shown in Figure A. The analysis is based on flow cytometry data. The number of RE / MTERF1 binding sites (SEQ ID NO:42) in each reporter construct is shown on the X-axis. (D) Dependence of GoI expression on the number of REs in the reporter construct in the HEK293 human cell line. Microscope images of HEK293 cells show expression of mCerulean GoI (top) and expression of the reporter gene control mCherry (bottom). The names of the fluorescence channels are indicated on the left. The number of RE repeats in each promoter construct is indicated above each pair of images. (E) Quantitative analysis of mCerulean levels in HEK293 cells shown in Figure D, which were normalized to mCherry levels. This analysis was based on flow cytometry data. The number of RE / MTERF1 binding sites in each reporter gene construct is indicated on the X-axis. (F) DNA sequences of 5x RE repeats were spaced by 10, 6, or 0 base pairs (bp). (G) Effect of the spacer length between REs in the reporter gene construct on GoI expression in HeLa cells. Schematic diagram of the promoter region and the position of the variable spacer is shown above. Microscope images of HeLa cells show expression of mCerulean GoI (top) and expression of the reporter gene control mCherry (bottom). The names of the fluorescence channels are indicated on the left. The length of the spacer between REs in each promoter construct is indicated above each pair of images. (H) Quantitative analysis of mCerulean levels in HeLa cells transfected with the reporter construct containing the spacer shown in Figure F, which was normalized to mCherry levels. The analysis was based on flow cytometry data. The length of the spacer between REs in each reporter construct is shown on the X-axis. (I) Effect of the length of the spacer between REs in the reporter construct on GoI expression in HEK293 cells. A schematic diagram of the promoter region and the position of the variable spacer is shown above. Microscope images of HEK293 cells show expression of mCerulean GoI (top) and expression of the transfected reporter control mCherry (bottom).Fluorescence channel names are indicated on the left. The length of the spacer between REs in each promoter construct is shown above each pair of images. (J) Quantitative analysis of mCerulean levels in HeLa cells transfected with reporter constructs containing the spacers shown in Figure F, normalized to mCherry levels. This analysis is based on flow cytometry data. The length of the spacer between REs in each reporter construct is shown on the X-axis. (K) DNA sequence of the repeats of the scrambled (Scr) RE repeated five times in the negative control promoter (SEQ ID NO:51). (L) Microscope images show HeLa cells transfected with a negative control (Neg.ctrl) reporter design containing five scrambled (Scr) REs spaced 10 bp apart. (M) Flow cytometric measurements of the cells shown in Figure L. (N) Microscope images show HEK293 cells transfected with a reporter design containing five scrambled REs spaced 10 bp apart. (O) Flow cytometric measurements of the cells shown in Figure N.

[0333] The scale bar in the microscopy images indicates 600 μm. Panels B, G, and L: mCerulean, 100 ms exposure, LUT range 0-26,000; mCherry, 75 ms exposure, LUT range 0-35,000. Panels D, I, and N: mCerulean, 75 ms exposure, LUT range 0-65 k; mCherry, 50 ms exposure, LUT range 0-40,000. All micrographs are at 100x magnification.

[0334] RE, response element; bp, base pair; Rel., relative; Scr, scrambled sequence; Neg.ctrl, negative control

[0335] Figure 2 .Codon optimization of the MTF coding sequence (CDS).

[0336] (A) The six amino acid sequences of an exemplary MTERF1 and its genetic code in wild-type (WT) and codon-optimized (CO) forms. The different DNA bases are shown on a light gray background. Each triplet is underlined, and the encoded amino acid is shown below. (B) Micrographs of HeLa cells transfected with a 5xRE, 6bp spacer reporter gene (SEQ ID NO: 44) construct, MTF encoded by wild-type or codon-optimized DNA sequences, and a transfection control construct encoding constitutively expressed mCitrine. The upper image shows mCerulean, and the lower image shows mCitrine expression. Fluorescence channels are indicated on the left. (C) Quantitative analysis of mCerulean levels in the cells shown in Figure B, normalized to mCitrine levels. The analysis is based on flow cytometry data. (D) Microscope images show mCerulean fluorescence levels (top) and mCitrine fluorescence levels (bottom) in HEK293 cells. The names of the fluorescence channels are indicated on the left. (E) Quantification of mCerulean levels in cells shown in panel D, normalized to mCitrine levels. This analysis is based on flow cytometry data.

[0337] The scale bar in B and D represents 600 μm. Panel B: mCerulean, 2-second exposure, LUT range 5000-30,000; mCitrine, 500-ms exposure, LUT range 0-65,000. Panel D: mCerulean, 500-ms exposure, LUT range; mCitrine, 300-ms exposure, LUT range: 0-65,000. All micrographs are at 10x magnification. WT, wild type; CO, codon-optimized; CDS, coding sequence.

[0338] Figure 3 . Expression levels of reporter protein as a function of the amount of MTF construct transfected.

[0339] The y-axis represents mCerulean expression relative to the control mCherry transfection. Error bars indicate the standard deviation of three replicates of HEK293 cells transfected with the MTF construct, with the amount of MTF construct indicated in nanograms on the x-axis. The line represents the dose-response curve, which was fit to the Hill equation with n=1. This analysis is based on flow cytometry data.

[0340] Figure 4 Comparison of HumTAP variants with different transcriptional activation domains.

[0341] (A) Relative reporter gene expression in HEK293 cells following co-transfection of the indicated HumTAP constructs (names indicated on the X-axis use the HUGO gene nomenclature, and amino acid numbering is according to UniProt) with a 5xRE, 0 bp-spacer (SEQ ID NO: 48) promoter driving mCerulean GoI, normalized to the expression of the transfection control mCherry. (B) DNA length required to encode each TAD, in base pairs (bp). The dashed lines in both figures indicate the expression level of mCerulean GoI or RelA. 430-551 DNA size of TADs. Rel., relative; bp, base pairs.

[0342] Figure 5 . Construction and testing of HumTAP with reduced genetic footprint.

[0343] (A) Schematic representation of the different domains of MTERF1. The numbers above the horizontal bars indicate the amino acid number, starting from 1 at the N-terminus. The width of the bar is proportional to its amino acid length, except that the dashed bar indicates a stretch of amino acids that is not shown. (B) Quantitative analysis of mCerulean levels in cells transfected with HumTAP constructs constructed with the DNA binding domains shown in Figure A, normalized to mCherry levels. Bars of different grayscale indicate different amounts of transfected HumTAP constructs. The MTERF1 domains used are indicated on the x-axis. This analysis is based on flow cytometry data. (C) The length of DNA required to encode the different DNA binding domains, in base pairs (bp). The peptide range of MTERF1 is indicated on the x-axis. The dashed line indicates the baseline MTERF1 58-399 Domain length. MTP, mitochondrial transfer peptide; WT, wild type; Rel., relative; bp, base pairs.

[0344] Figure 6 . Experimental evaluation of immunogenicity.

[0345] (A) Scheme of the experimental procedure. Peripheral blood mononuclear cells (PBMCs) are shown as various shapes in the diagram of the culture wells, and the peptides are plotted as sticks. Dimethyl sulfoxide (DMSO) is the solvent for the peptides. (B) P values ​​from a t-test comparing the number of spot-forming units (SFU) in cells primed and recalled with the same peptide pool and in the same PBMCs primed but not recalled. The x-axis shows the protein from which the peptide pools used for priming and recall were derived. Light gray indicates a p-value above 0.05, in which the donor was classified as a non-responder. Dark gray indicates a region with a p-value below 0.05. Donor-derived samples whose p-values ​​fell within this region were considered responders. The dotted line indicates a p-value of 0.05. Different dot characters indicate different PBMC donors. DMSO, dimethyl sulfoxide; ELISPot, enzyme-linked immunospot assay; H0: null hypothesis; SFU, spot-forming unit.

[0346] Figure 7 .MTF orthogonality evaluation.

[0347] A) Confocal microscopy of HeLa cells transfected with Flag-tagged wild-type MTERF1 or MTF. Scale bar indicates 10 μm. MitoRed is a dye that stains mitochondria, anti-Flag Ab is used to stain Flag-tagged transfected proteins, and Hoechst 33342 stains DNA. B) Graph showing the expected interactions of plasmids and their encoded proteins. C) Reporter fluorescence levels in HEK293 cells transfected with the MTF construct and a reporter gene construct containing the 5xRE (SEQ ID NO:48) without a spacer in its promoter, in the presence or absence of EF1a-driven WT MTERF1.

[0348] Figure 8 Volcano plot of differential gene expression analysis of transfected HEK293 cells.

[0349] (A) Volcano plot showing differentially regulated genes between cells transfected with WT MTERF1 and a useless plasmid. The Y-axis represents the log10 of the false detection rate (FDR). The X-axis represents the log2 of the fold change (FC). Light gray dots indicate genes with no significant differential expression, medium gray dots indicate downregulated genes, and darker colors represent upregulated genes. Labels indicate the ENSEMBL symbol of the gene corresponding to the nearest dot, or “-” for transcripts not associated with the designated gene. (B) Volcano plot comparing gene expression between cells transfected with an MTF construct and cells that received a useless DNA plasmid. The Y-axis represents the log10 of the false detection rate (FDR). The X-axis represents the log2 of the fold change (FC). Light gray dots indicate genes with no significant differential expression, medium gray dots indicate downregulated genes, and darker colors represent upregulated genes. Labels indicate the ENSEMBL symbol of the gene corresponding to the nearest dot, or “-” for transcripts not associated with the designated gene. WT, wild type; FDR, false detection rate; FC, fold change.

[0350] Figure 9 Optimization of transcriptional activation strength and protein size.

[0351] A) Overview of the TAD domain. Subscript numbers indicate the amino acid ranges comprising the TAD domain used in this work. B, C, E, F) Transcriptional activation mediated by a single transcriptional activation domain. The number of base pairs required to encode HumTAP variants is shown in the upper panel. The lower panel shows flow cytometric analysis of HEK293 cells co-transfected with plasmids encoding pEF1a-driven mCherry, six reporter gene designs, and equimolar amounts of plasmids encoding pEF1a-driven MTERF1 fused to the TAD domain. 58-399 The plasmids are indicated below each column. The symbols below the columns correspond to those in panels A and D. The columns correspond to the average of three replicates, and each dot indicates one replicate. The horizontal dashed line indicates the cells transfected with the baseline MTERF1 58-399 ::RelA 430-551 rel.mCerulean levels in cells expressing a HumTAP variant, "MTF" (SEQ ID NO: 100). D) Schematic diagram of the MTERF1 structure and the amino acid ranges conserved across the different variants. Amino acid sequences not shown in the schematic are indicated by dashed lines, while vertical dashed lines indicate the starting amino acid position of each variant. TAD, transcriptional activation domain; DBD, DNA binding domain; MTP, mitochondrial transfer peptide; bp, base pairs.

[0352] Figure 10 .Construction of a HumTAP-based gene expression system inducible by rapamycin analogs.

[0353] A) MTERF1 mediated by A / C heterodimers 58-399 ::FKBP (SEQ ID NO: 192) and FRB::TAD dimerization schematic. B) Flow cytometric analysis of HEK293 cells co-transfected with six reporter plasmid designs, a plasmid encoding pEF1a-driven mCherry, and varying amounts of each plasmid encoding the expression of each pEF1a-driven dimerization protein. Axis units refer to molar equivalents of plasmid amount, while relative mCerulean is calculated by normalizing the mCerulean signal to the mCherry signal. C) Relative mCerulean expression levels in HEK293 cells co-transfected with six reporter plasmid designs, a plasmid encoding pEF1a-driven mCherry expression, and a plasmid encoding each of the two pEF1a-driven dimerization proteins. Dots represent mCerulean levels normalized to mCherry levels, and the dashed line represents the logistic model fit. The color indicates which TAD was used in constructing the FRB::TAD fusion protein transfected into the cells from which the data were obtained. D) Relative mCerulean expression levels in HEK293 cells co-transfected with six reporter plasmid designs, a plasmid for pEF1a-driven mCherry expression, and a plasmid encoding each of the two pEF1a-driven dimerization proteins. Points represent mCerulean levels normalized to mCherry levels, and a line was fitted to the data using a logistic model. Colors indicate the expression of mCerulean in the construct transfected into the cells from which the data were obtained. 58-399 Which linker is used in the ::FKBP fusion protein (Linker 1) and the FRB::TAD fusion protein (Linker 2). TAD, transcription activation domain; Rv1, RelA. 430-551 (SEQ ID NO:3)

[0354] Figure 11 Schematic diagram of a massively parallel reporter gene assay to establish principles for cognate promoter design.

[0355] A) Schematic diagram of the components and various characteristics of the plasmid library. B) Overview of the experimental process. BS, binding site; bp, base pair; BC, barcode; RT, reverse transcription; PCR, polymerase chain reaction

[0356] Figure 12 .Quality control of promoter design library.

[0357] A) The chimerism rate indicates the ratio of reads containing an unexpected promoter design coupled to a given barcode to the total number of reads containing that barcode. The barcode read count indicates the total number of reads in which that barcode was present. The distribution of points on each axis is shown at the top and left edge. R2 represents the square of the Pearson correlation coefficient. B) The distribution of design parameters is shown at the top of each stack. The percentage indicates the frequency of each given parameter identified across all reads, and the boxes have the corresponding height. Dir., direction.

[0358] Figure 13 .Quality control of sequencing results.

[0359] A) Histogram showing the frequency of reads for each barcode in a plasmid DNA sample. B) The height of the stacked bar indicates the distribution of barcodes from any library member or any UbC control for each sample. C) Dot plot showing the activity score for each barcode in two replicates of cells transfected with an amount of MTF plasmid corresponding to the EC90. The marginal density plot indicates the distribution of barcode activity scores for replicate 2. D) Histogram of promoter-level activity scores corresponding to MTF levels at EC90. E) Dot plot showing Citrine fluorescence levels normalized to mCherry expression on the y-axis in HEK293 cells co-transfected with individually selected plasmids, a plasmid encoding pEF1a-driven MTF at an amount corresponding to the EC90, and a plasmid encoding pEF1a-driven mCherry. The x-axis indicates the activity score for the design corresponding to the individually selected plasmids. R2 indicates the square of the Pearson correlation coefficient, p indicates the Spearman correlation coefficient, and the linear regression fit is shown as a dashed line. F) Activity score traces obtained from RNA samples of cells co-transfected with a plasmid library, a plasmid encoding UbC-driven Citrine, and varying amounts of a plasmid encoding pEF1a-driven MTF. EC, effective concentration

[0360] Figure 14 . Promoter activity at the MTF level corresponding to its EC90.

[0361] A) Schematic representation of the potential gene regulatory cascades under the measured promoter design activity scores. B) Ranks assigned to promoters in descending order according to their activity scores. The marginal density plot represents the distribution of promoter activity scores. C) Heat map of all activity scores for all designs. For designs with one BS, the "BS-BS distance" is not applicable and is set to 0. D) Distribution of activity scores. The violin plot contains all designs that share the indicated design parameters, and the black dots represent the median activity scores. P values ​​were calculated using analysis of variance (ANOVA). E,F) Distribution of activity scores. The violin plot contains all designs that share the indicated design parameters, where the binding site is oriented in the sense (E) or antisense (F) orientation relative to the minimal promoter. The black dots represent the median activity scores. P values ​​were calculated using analysis of variance (ANOVA). BS, binding site; bp, base pairs

[0362] Figure 15 .Characterization of individual promoter variants.

[0363] A) Variants were randomly selected or specifically chosen for high activity scores and small size. Dark dots indicate promoter design variants that were selected and transfected individually, and dots with different fill patterns represent designs indicated by the same patterns in Figures B, C, and D. B, D) Average relative mCitrine levels of three or six replicates in cells co-transfected with MTF-encoding plasmids at levels corresponding to the EC90 were measured by flow cytometry. C) Relative mCitrine levels in cells transfected with the amounts of MTF-encoding plasmids indicated by the axis labels. D) White dots indicate individual replicates. Example

[0364] The methods and materials used in the present disclosure are described herein; other suitable methods and materials known in the art can also be used. These materials, methods, and examples are illustrative only and are not intended to be limiting.

[0365] Example 1. Engineering of a synthetic transcriptional activation system based on MTERF1

[0366] An exemplary HumTAP transcriptional activation system was designed by engineering a "human-derived transcriptional activator protein" (HumTAP) transcriptional activator and a promoter containing a response element / MTERF1 binding site operably linked to a gene of interest. 58-399 (SEQ ID NO: 1), Uniprot Q99551) and a peptide consisting of amino acids 430 to 551 of the RELA protein (RelA 430-551(SEQ ID NO: 3) and Uniprot Q04206) to prepare the prototype HumTAP protein. This fusion protein is hereafter referred to as MTERF1 58-399 ::RelA 430-551 (or "MTF"; SEQ ID NO: 100). A DNA construct was prepared that carries the strong EF1A promoter driving constitutive expression of MTF (SEQ ID NO: 101) in human cells ("MTF construct").

[0367] To construct a promoter driving the response element / MTERF1 binding site of a gene of interest, the inventors placed the MTERF1 binding site (SEQ ID NO: 42) upstream of a minimal TATA box (SEQ ID NO: 103). The mCerulean fluorescent reporter gene representing the gene of interest (GoI) was placed downstream of the TATA box. Each of these is referred to as a reporter construct. Five different reporter constructs were prepared, containing 3, 5, 7, 9, and 11 MTERF1 binding sites, respectively, 6 base pairs upstream of the TATA box ( Figure 1 A; SEQ ID NO: 41, 44, 45, 46 and 47, respectively). They were co-transfected with the MTF construct into Hela cells. Microscopic analysis showed that the maximum mCerulean signal for the 11 repeated MTERF1 binding sites ( Figure 1 B). Relative mCerulean levels in HeLa cells were measured using flow cytometry and calculated by normalization to transfection control ( Figure 1 C). Examination using a microscope ( Figure 1 D) and flow cytometry ( Figure 1 E), The same transfection conditions in HEK293 cells revealed that the reporter gene fluorescence level increased with increasing number of RE / MTERF1 binding sites in the promoter design. Next, MTF constructs were transfected with different spacer lengths (0, 6 bp, and 10 bp, respectively; SEQ ID NOs: 48, 44, and 49) between a fixed number of REs (i.e., 5 RE / MTERF1 binding sites) ( Figure 1 F) Reporter gene plasmid. In HeLa cells, microscopic examination revealed a large increase in reporter gene expression when there was no spacer at all, and similar amounts of mCerulean when there were 6 bp and 10 bp spacers ( Figure 1 G). This was confirmed using flow cytometry ( Figure 1 H). A different trend was observed in HEK293 cells, where microscopy was used ( Figure 1 I) and flow cytometry ( Figure 1J) It was observed that the reporter amount was highest for 6 bp intervals and the mCerulean signal was very small for 10 bp intervals. Sequence specificity was confirmed by microscopic examination ( Figure 1 L) and flow cytometry ( Figure 1 M) evaluation, a promoter having five scrambled DNA binding sites (SEQ ID NO: 51) spaced 10 bp apart ( Figure 1 K) reporter gene constructs did not result in mCerulean expression after co-transfection with the MTF construct in HeLa cells. HEK293 cells were examined under a microscope ( Figure 1 N) and flow cytometry ( Figure 1 O) shows the same behavior.

[0368] Next, the inventors optimized the DNA sequence encoding MTF (SEQ ID NO: 102; Figure 2 A). Microscopic examination ( Figure 2 B) and flow cytometry ( Figure 2 C) analysis, using a reporter gene construct (SEQ ID NO: 44) with 5xRE / MTERF1 binding sites (separated by a 6 bp spacer) in its promoter, this optimization increased reporter gene expression in Hela cells by about 3-fold. In HEK cells, the expression of the reporter gene was observed by micrographs ( Figure 2 D) and flow cytometry analysis ( Figure 2 As depicted in E), this modification had no significant effect on reporter gene fluorescence levels.

[0369] By transfecting different amounts of codon-optimized MTF constructs into HEK293 cells, the inventors established a dose-dependent effect of mCerulean reporter gene (GoI) expression on MTF expression levels ( Figure 3 HEK293 cells in 24-well plates were transfected with 0.75, 3, 6, 12, 24, 48, 96, 192, and 384 ng of MTF construct, 100 ng of reporter construct, and 50 ng of constitutive mCherry transfection control construct. Two days after transfection, cells were analyzed by flow cytometry. The dose-response curve was modeled as a Hill equation, assuming a Hill coefficient of 1. EC10 and EC50 were reached at 1.25 ng and 11.24 ng of transfected MTF construct, respectively. Although a slight decrease in relative reporter fluorescence levels was observed for higher amounts of transfected MTF construct, the theoretical EC90 was reached at 101 ng.

[0370] To design a stronger and potentially shorter HumTAP, the inventors inserted MTERF1 into the 58-399The peptide (SEQ ID NO: 1) was fused to various transcriptional activation domains

[18] , and HEK293 cells were transfected with these constructs and a reporter gene construct (SEQ ID NO: 48) containing 5xRE / MTERF1 binding sites with no spacer between them.

[0371] First, MTERF1 58-399 (SEQ ID NO: 1) fused to a different subsequence of RelA. 361-551 (SEQ ID NO: 27) corresponds to the sequence used to construct the synthetic TF PIT2 of bacterial origin

[19] , while RelA 342-551 (SEQ ID NO: 29) contains all annotated transcriptional activation domains according to the domain annotation under Uniprot Q04206 (https: / / www.uniprot.org / uniprotkb / Q04206 / entry). Both resulted in stronger reporter gene expression ( Figure 4 A), but with RelA 430-551 (SEQ ID NO: 3; Figure 4 B) also requires a larger gene footprint. 521-551 , SEQ ID NO: 31) domain

[20] did not significantly activate reporter gene expression. We also 58-399 (SEQ ID NO: 1) was fused to a set of transcriptional activation domains previously found to be strong activators

[18] with a small gene footprint ( Figure 4 B). WW domain of WWC1 (SEQ ID NO: 21; WWC1 2-81 , Uniprot Q8IX03), the KRAB domain of ZNF473 (ZNF473 5-48 , Uniprot M0R032, SEQ ID NO: 23) and the Nuc_rec_co-act domain of NCOA3 (SEQ ID NO: 25; NCOA3 1045-1092 , Uniprot Q9Y6Q9) did not result in high reporter gene expression. In contrast, the LMSTEN domain of MYB (SEQ ID NO: 19; MYB 251-330 , Uniprot P10242) and the FOXO-TAD domain of FOXO3 (SEQ ID NO: 17; FOXO3 604-644, Uniprot O43524) mediated significant mCerulean expression. VP64 (which contains three repeats of the transcriptional activation domain of HHV11 (SEQ ID NO: 15) (SEQ ID NO: 13; HHV11) connected by a GS linker 437-447 , Uniprot P06492)), a commonly used strong virus-derived transcriptional activation domain, mediated reporter gene expression below the benchmark RelA used 430-551 Interestingly, two copies of FOXO3 604-644 (SEQ ID NO: 33) or RelA 521-551 Fusion of (SEQ ID NO: 35) resulted in 12 and 201 times the fluorescence conferred by only one copy of each individual TAD ( Figure 4 A) These results suggest that the DBD of MTERF1 can serve as a versatile building block to regulate gene expression when fused to a TAD that may or may not be derived from RelA.

[0372] To further reduce the DNA footprint of HumTAP, the inventors investigated to what extent the DNA binding domain could be reduced. To this end, the inventors created three versions of HumTAP fused to RelA. 430-551 MTERF1 subsequence (SEQ ID NO: 3). 73-399 (SEQ ID NO: 7), MTERF1 104-399 (SEQ ID NO:9) and MTERF1 135-399 (SEQ ID NO: 11) is a subunit of MTERF1 without its mitochondrial transfer peptide (MTERF1 1-57 ; SEQ ID NO: 37), which lacks only the first 14 N-terminal amino acids (MTERF 58-72 ; SEQ ID NO: 39), or in addition lacking the first (MTERF1 73-98 ; SEQ ID NO: 104) or the first and second (MTERF1 104-134 ; SEQ ID NO: 106) mterf motif ( Figure 5 A)

[21] . Removal of the N-terminal domain reduces the genetic footprint of HumTAP ( Figure 5 B) These constructs were co-transfected into HEK293 cells with a reporter gene construct containing 5xRE without a spacer in its promoter (SEQ ID NO: 48). 73-399 (SEQ ID NO: 7) and MTERF1 104-399(SEQ ID NO: 9), it was found that the output expression was reduced but not eliminated. However, further truncation to amino acid 135 (MTERF1 135-399 (SEQ ID NO: 11)) results in a complete loss of transcriptional activation ( Figure 5 B) This demonstrates that the genetic footprint of synthetic TFs can be reduced while retaining activity.

[0373] Example 2. Evaluation of immunogenicity

[0374] T cells that strongly bind to self-peptide-MHC complexes are deleted in the thymus during maturation, which allows tolerance to self-peptides

[22] . Therefore, since HumTAP is composed entirely of human protein subunits, it is expected to have more favorable immunogenic characteristics compared to proteins of non-human origin.

[0375] The inventors evaluated whether protein subsequences could elicit an immune response. Using NetMHCpan4.0

[23] , the inventors predicted which peptides derived from the MTERF-RelA junction were most likely to bind strongly to the most common MHC I allele, HLA-A02:01, in the human population. Pools of five peptides predicted to bind most strongly were synthesized. In addition, for each of those “core” binders, the inventors synthesized 15-mers containing the flanking sequences, generating 10 peptides. As a positive control, a pool of peptides spanning NY-ESO1, an immunogenic cancer-associated protein, was used. Peptides derived from self-portions (i.e., RelA and MTERF, but not the junction) served as negative controls. Peripheral blood mononuclear cells (PBMCs) were isolated from normal blood donors, transferred to 24-well plates, and each peptide pool was added to each of three replicate wells. The cells were then incubated for 14 days in the presence of IL-2 and IL-7. At this stage, T cells containing T cell receptors (TCRs) that recognize peptides bound to MHC class I or class II on the surface of PBMCs (particularly antigen-presenting cells) will be stimulated and begin to proliferate (a process called "priming"). Next, the PBMCs in each well are restimulated overnight with the same peptide pool and subjected to IFNγ-ELISPot in 4 replicates to generate and count spots ( Figure 6A). Spots are induced by T cells that have been previously primed with one of the peptides in the priming peptide mixture, so more spots correspond to the more immunogenic peptide mixture in the priming step. Donors were classified as "responders" if significantly more spot-forming units (SFU) (p<0.05) were observed for restimulated cells compared to their respective primed but not restimulated background controls, as suggested by Moodie (2010), Cancer Immunology, Immunotherapy, Vol. 59, pp. 1489-1501. Some donors did not show a significant response to the NY-ESO1 peptide pool and were omitted from the analysis. Seven PBMC samples showed a response to NY-ESO1 (positive control) and were included in the analysis. Of these samples, one produced a significant response to the junction-derived peptide pool, while three PBMC samples were considered reactive to the negative control peptide ( Figure 6 B) This shows that the behavior of the MTF junction-derived peptide pool is much more similar to the negative control consisting of self-peptides than to the known immunogenicity positive control, and points to low immunogenicity in humans.

[0376] Example 3. Orthogonality evaluation

[0377] The inventors examined whether the lack of HumTAP's mitochondrial transfer peptide (MTP) would result in the orthogonality of HumTAP to human host cells, and vice versa, that is, whether (i) endogenous human gene expression is not substantially altered by HumTAP, and (ii) endogenously expressed wild-type MTERF1 (i.e., including the MTP set forth in SEQ ID NO: 37) lacks the ability to regulate the expression of a gene of interest driven by a promoter containing a response element.

[0378] First, the inventors evaluated intracellular localization by transfecting HeLa cells with a plasmid constitutively expressing Flag-tagged MTF (SEQ ID NO: 128) or a plasmid constitutively expressing Flag-tagged wild-type MTERF1 (SEQ ID NO: 130). Cells were stained with anti-Flag antibody and imaged by confocal microscopy. Indeed, it was found that WT MTERF1 was localized to mitochondria, while MTF was distributed throughout the cell, including the nucleus, without obvious localization to mitochondria ( Figure 7 A).

[0379] To assess the orthogonality of the host response element driven gene of interest (GoI), the inventors transfected HumTAP representative MTF (SEQ ID NO: 100) with the response element driven (SEQ ID NO: 48) mCerulean fluorescent reporter gene in the presence and absence of a construct encoding wild-type MTERF1 to assess whether it could interfere with the induction of GoI by MTF ( Figure 7 B). The inventors found that there was no difference in GoI fluorescence levels with or without WT MTERF1 expression ( Figure 7 C).

[0380] Alignment of the MTERF1 binding sites with the human genome using BLAST showed that there was no perfect match in the nuclear genome, but the presence of imperfect binding sites to which MTF could bind could not be ruled out. Therefore, the inventors used RNA-sequencing to investigate whether MTF (SEQ ID NO: 100) activates the expression of endogenous human genes. The inventors also transfected WT MTERF1 (SEQ ID NO: 108) to compare the gene regulatory effects of MTERF1 and MTF. Differential gene expression analysis was performed on cells transfected with "junk" plasmids using the edgeR package for R

[24] . Genes were considered differentially expressed (DE) if their expression was found to be significantly different (FDR < 0.5) with a fold change of at least 2. For WT MTERF1, 7 genes were found to be downregulated, and 3 genes were found to be upregulated (excluding MTERF1 expressed from the transfected plasmid) ( Figure 8 A). For MTF, 24 genes were down-regulated and 4 genes were up-regulated ( Figure 8 B). Notably, none of the genes downregulated by WT MTERF1 were found to be affected by MTF. Only one gene, HSPA7, was upregulated under both conditions (Table 1). A few genes were significantly downregulated by MTF, but none of them contained sequences similar to MTERF1 binding sites in their genomic vicinity. Similarly, none of the upregulated genes in MTF-transfected cells contained typical MTERF1 binding motifs. The inventors concluded from this experiment that MTF does not retain the gene regulatory function of the wild-type MTERF1 protein and does not exert extensive off-target gene activation. The lack of MTERF1 binding sites in the genomic vicinity of downregulated genes may suggest the existence of a mechanism independent of typical MTF binding. Future studies will assess the dose-dependence of differentially expressed (DE) genes on the amount of MTF and may help to elucidate the origin of DE genes.

[0381] Table 1. List of all DE genes of MTF (upper panel) and WT MTERF1 (lower panel).

[0382] MTF

[0383]

[0384]

[0385] WT MTERF1

[0386]

[0387]

[0388] Example 4. Optimization of humTAP

[0389] This example is an extension of Example 1 herein above, and describes in part the same data and results.

[0390] For many applications, strong transcriptional activation of the gene of interest by synthetic transcription factors (synTFs) is required, while viral delivery methods require components that are as small as possible

[25] . Therefore, the inventors investigated the size and transcriptional activation strength of various human and engineered chimeric transcriptional activation domains (TADs) and used RelA 430-551 The peptide (SEQ ID NO: 3) was used as a benchmark.

[0391] In addition to RelA 430-551 In addition to the RelA subsequence (referred to as "v1"), the inventors tested three alternative subsequences: v2) RelA 342-551 (SEQ ID NO: 29), which contains all the transcriptional activation domains annotated on UniProt Q04206; v3) RelA 361-551 (SEQ ID NO: 27), which is the same peptide used in the construction of PIT2 [2]; and v4) RelA 521-551 (SEQ ID NO: 31), which only constitutes the transcriptional activation domain 1 (TA1)

[26] . The protein domains WW (SEQ ID NO: 21), KRAB of ZNF473 (SEQ ID NO: 23), Nuc-rec-co-Act (SEQ ID NO: 25), LMSTEN (SEQ ID NO: 19) and FoxoTAD (SEQ ID NO: 17) have recently been identified as strong transcriptional activators

[27] . To evaluate their functionality in the HumTAP context, the inventors created their respective MTERF1 58-399 (SEQ ID NO: 1) Figure 9 A).

[0392] When HEK293 cells were co-transfected with plasmids encoding pEF1a-driven fusion HumTAP expression, response element-driven mCerulean reporter protein, and an mCherry transfection control, all peptides derived from RelA, except TA1, mediated high levels of reporter gene expression. RelA protein containing a larger portion resulted in higher reporter gene expression. In contrast, the core transcription activation domains TA1, WW, ZNF473 KRAB, and NucRecCoAct mediated little to no reporter gene expression. LMSTEN and FoxoTAD domains generated some mCerulean expression, although the levels were 11.3-fold and 8.4-fold lower than the benchmark RelA430-551 TAD, respectively ( Figure 9 B).

[0393] Interestingly, fusion of two TA1 domains (SEQ ID NO: 35) resulted in reporter gene expression, and addition of a third domain (SEQ ID NO: 194) further increased output expression ( Figure 9 C). Similarly, two copies of FoxoTAD (SEQ ID NO: 33) mediated 20-fold higher mCerulean expression than only one copy of FoxoTAD and 20-fold higher mCerulean expression than the benchmark RelA. 430-551 This effect was not evident for the LMSTEN domain, for which two copies (SEQ ID NO: 196) mediated ∼2.7-fold higher reporter gene expression than one copy. Fusion of FoxoTAD to LMSTEN (SEQ ID NO: 54), LMSTEN to TA1 (SEQ ID NO: 56), or FoxoTAD to TA1 (SEQ ID NO: 58) produced ∼2.7-fold higher reporter gene expression than the benchmark RelA. 430-551 A TAD that is smaller and stronger. The chimeric TAD composed of FoxoTAD, LMSTEN, and TA1 (SEQ ID NO: 52) produces a reporter gene expression level that is 3.7 times higher than that of RelA430-551 ( Figure 9 C) Therefore, it is possible to design a RelA that is more efficient than the commonly used one without introducing non-human protein domains. 430-551 Stronger and smaller transcriptional activation domains. Given the synergy observed between homotypic and heterotypic transcriptional activation domains, further work exploring this effect will likely result in human-based chimeric TADs that mediate even higher levels of transgene expression without increasing protein domain size.

[0394] To further reduce the size of the HumTAP protein, the inventors deleted the N-terminus (SEQ ID NO: 39), the optional Mterf motif 1 (SEQ ID NO: 104), and the optional Mterf motif 2 (SEQ ID NO: 106) in addition to the MTP of MTERF1 (SEQ ID NO: 37)

[28] ( Figure 9 D). Given the abrogating effect of the R387A mutation

[28] (which is located only 12 amino acids from the C-terminus of MTERF1), the inventors did not delete any C-terminal domains. 73-399 ; SEQ ID NO: 7) and omitting both the N-terminus and the Mterf motif 1 (MTERF1 104-399 ; SEQ ID NO: 9) produced a ~1.5-fold decrease in reporter gene expression, but the DNA size was reduced by 45 and 135 bases, respectively. Also omitting the Mterf motif 2 (shown in SEQ ID NO: 11) resulted in a non-functional synTF Figure 9 E).

[0395] To examine whether the reduced transcriptional activation activity of HumTAP with a smaller DBD could be compensated, we fused the novel chimeric TAD to MTERF1. 104-399 In this case, the FoxoTAD::TA1 chimeric TAD (SEQ ID NO:58) produced a synTF that was approximately as potent as MTF (SEQ ID NO:100), but required 288 fewer DNA bases to encode on the vector. 104-399 Combinations of FoxoTAD::FoxoTAD (SEQ ID NO: 33) and FoxoTAD::LMSTEN (SEQ ID NO: 54) resulted in synTFs that were ~1.5-fold more potent than MTF (SEQ ID NO: 100), but were 96 and 86 amino acids smaller, respectively. In this case, using FoxoTAD::LMSTEN::TA1 (SEQ ID NO: 52) resulted in a synTF similar in size to MTF, but mediated 3.5-fold more reporter gene expression ( Figure 9 F).

[0396] Taken together, this dataset supports the notion that MTERF1 retains DNA binding functionality even with the deletion of its N-terminal amino acids. The reduced reporter gene levels could be explained by a reduced affinity for DNA or by a more unstable protein. Regardless, high reporter gene expression levels could be restored by using an optimized TAD.

[0397] Example 5. Engineering of an MTF-based inducible gene expression system

[0398] Regulation of output expression intensity by small molecules has been recognized as an important feature of synthetic transcriptional activation systems

[29] . Therefore, the inventors attempted to design an inducible HumTAP system based on dimerization.

[0399] A well-characterized small molecule-based dimerization system is based on the rapamycin analog A / C (C16-(S)-7-methylindorapamycin), which induces heterodimerization between the human protein FK506 binding protein 12 (FKBP; SEQ ID NO: 142) and the FKBP12-rapamycin binding domain (FRB) mutant FRBT2098L (SEQ ID NO: 140)

[30] ,

[31] . To achieve gene expression inducible by the A / C heterodimerizer, the inventors fused FKBP (SEQ ID NO: 142) to MTERF1. 58-399 (SEQ ID NO: 1) and fused FRB (SEQ ID NO: 140) to the C-terminus of different transcriptional activation domains ( Figure 10 A). When the concentration of A / C heterodimer (i.e. C16-(S)-7-methylindorapamycin) was kept constant at 100 nM, the inventors found that compared with MTERF1 58-399 ::FKBP (SEQ ID NO: 192) plasmid, 4-fold molar excess of FRB::RelA 430-551 (SEQ ID NO: 170) plasmid, resulting in the highest mCerulean levels in HEK293 cells co-transfected with an inducible TF component and a response element-driven mCerulean reporter protein ( Figure 10 B) MTERF1 was co-transfected 58-399 In HEK cells expressing FKBP and a fusion protein between FRB and the newly developed strongest TAD (see Example 4), mCerulean levels increased in a dose-dependent manner with increasing inducer concentrations ( Figure 10 C) According to the above results, the maximum output expression level mediated by the chimeric transcriptional activation domain is higher than that of RelA 430-551 .

[0400] Following ingestion, physiological blood concentrations of rapamycin (from which the A / C dimerizer is derived) reach ∼22 nM

[32] , whereas the present system exhibits an EC50 between 100 and 120 nM. Previously reported inducible gene expression systems based on A / C-inducible FRB and FKBP dimerization exhibit EC50 values ​​of ∼1 nM

[31] ,

[33] ,

[34] , and are therefore inducible by pharmacologically achievable inducer concentrations. The inventors reasoned that in the present system, dimerization could be hindered by a fusion partner for either FRB or FKBP, and investigated whether the addition of a linker between the protein domains affected the EC50 values ​​( Figure 10 D). The combination of the cMyc nuclear localization signal (NLS) (SEQ ID NO: 122) linking MTERF158-399 and FKBP and the rigid linker AP6 (SEQ ID NO: 120) linking FRB and RelA430-551 resulted in a decrease in EC50 to ∼60 nM, while also increasing the maximum output by ∼3-fold compared to variants without a linker and with RelA430-551 as the TAD ( Figure 10 D).

[0401] Example 6 Engineering of HumTAP-homologous promoter (sensor) collections

[0402] Multiple factors interact in a complex manner to influence the strength of synthetic promoters. The inventors performed massively parallel reporter assays (MPRAs)

[35]

[37] to quantify the effects of four promoter design parameters on the strength of transcriptional activation: the number of MTERF1 binding sites (BS) (SEQ ID NO: 42), their orientation relative to the minimal promoter, the number of intervening bp between BSs, and the number of intervening bp between the minimal promoter and the proximal BS ( Figure 11 A).

[0403] The inventors investigated a range of 1 to 5 copies of the BS (SEQ ID NO:42), each oriented in either sense (SEQ ID NO:42) or antisense (SEQ ID NO:200) orientation relative to the yb_TATA minimal promoter (SEQ ID NO:103), with distances from the minimal promoter ranging from 0 to 10 bp (for a 10 bp spacer, see SEQ ID NO:199; the shorter spacer is truncated at the 3' end of SEQ ID NO:199), and spacing between the BSs ranging from 0 to 10 bp (for a 10 bp spacer, see SEQ ID NO:198; the shorter spacer is truncated at the 3' end of SEQ ID NO:198). For each of the 990 possible parameter combinations, 10 unique 11 bp barcodes were assigned. The resulting library of 9910 different DNA sequences was cloned in such a way that each encoded promoter variant drove expression of mCitrine (with the associated barcode in the 3' UTR). Next generation sequencing was used to measure the relative frequency of each barcode in the plasmid library. To control for differences in transfection efficiency and expression levels under different conditions, the inventors cloned 10 plasmids encoding the constitutive UbC promoter (SEQ ID NO: 139) driving mCitrine, each with a unique barcode in the 3'UTR. Then, the MTF (SEQ ID NO: 100) plasmids (see Examples 1 and 2) were added to the mCitrine library at the corresponding EC10, EC50, and EC90 levels. Figure 3 ) were co-transfected into HEK293 cells with the promoter library and 10 UbC control plasmids in triplicate. To quantify the expression level of each barcode, RNA was extracted from the cells, reverse transcribed, amplified, and sequenced. For each barcode, the score for each condition and replicate was calculated as the barcode count from the RNA sample (normalized to the barcode frequency in the plasmid library) and the median of the 10 barcodes associated with the UbC control. The "barcode level activity score" was the average of the barcode scores of the three replicates. For each promoter design, the final activity score was calculated as the median of the 10 associated barcode level activity scores ( Figure 11 B).

[0404] We performed quality control on our plasmid libraries using nanopore long-read sequencing, which was analyzed with an algorithm that extracted the promoter design parameters and barcode for each read. A common problem in MPRA library construction is the decoupling of barcodes from their assigned variants due to the formation of chimeric DNA sequences during PCR amplification of the DNA oligonucleotide pool

[38] ,

[39] . In our library, this effect occurred at an average rate of 17%, meaning that on average 17% of the plasmids encoding a given barcode were coupled to a promoter design different from the assigned promoter design. The lack of correlation between barcode read counts and chimera rate suggests that our analysis does not underestimate the true chimera rate ( Figure 12 A). Although chimeras may contribute noise in the screen, the inventors do not expect it to be high enough to hamper their ability to draw conclusions. Furthermore, analysis of the design parameter distributions did not show a strong bias for any particular parameter, except for a bias towards promoters containing BS in an antisense orientation ( Figure 12 B) Therefore, the inventors chose to use a plasmid library for screening.

[0405] Then, the inventors performed Figure 11 The entire workflow is outlined in B. Analysis of barcode frequencies in the plasmid library showed complete coverage of the library, with no barcode being read less than 10 times, an average of ˜27,000 reads per barcode, and tailing towards a maximum of 224,959 reads ( Figure 13 A). As the amount of MTF increases, the ratio of library-derived reads compared to UbC control-derived reads should also increase. In fact, the inventors found that for 0 MTF and EC90 levels, between 75% and 97.6% of the reads were derived from the plasmid library ( Figure 13 B). After calculating the barcode-level activity score for each sample, bimodality in the score distribution becomes apparent: barcodes fall into either the "low" group with scores below ~0.001 or the "high" group with scores above this threshold. Comparison of the two replicates shows that some barcodes fall into the "low" group in one replicate and the "high" group in the other, and vice versa. ( Figure 13 C). After averaging the replicates and calculating the promoter-level activity score, values ​​below 0.001 were no longer observed under EC90 conditions ( Figure 13 D) The inventors then isolated 39 individual plasmids from the library, co-transfected each plasmid with an amount of MTF and mCherry transfection control plasmid corresponding to the EC90 into HEK293 cells, and measured fluorescence levels using flow cytometry. mCitrine fluorescence levels correlated well with the promoter activity scores obtained from the screening, with R 2 is 0.77, and the Spearman correlation coefficient (not assuming a linear relationship) is 0.9 ( Figure 13E). Confirming the responsiveness of the promoter to MTF, the activity scores of most designs increased with increasing MTF levels, while no consistent increase was observed for the negative control design based on scrambled BS ( Figure 13 F).

[0406] The inventors concluded that the promoter design-level activity score reflects the reporter gene expression level mediated by the promoter variant under EC90 conditions. Therefore, the inventors analyzed the promoter-level activity score to identify the impact of design parameters on the promoter activity score at the MTF level corresponding to EC90 ( Figure 14 A). At the MTF level corresponding to EC90, the highest promoter score was ∼32-fold higher than the lowest score ( Figure 14 B). Although most promoters score in the lower half of the overall activity range, there is a tail of high scoring promoters ( Figure 14 B), where the differences and trends are shown in the heatmap ( Figure 14 C). Promoters constructed with binding sites oriented in the sense relative to the minimal promoter generated significantly higher scores than designs oriented in the opposite manner (i.e., binding sites oriented in the antisense relative to the minimal promoter) ( Figure 14 D). As previously observed in other contexts

[35] , each additional copy of the binding site leads to a higher score, independent of its orientation ( Figure 14 E). However, this effect seems to reach saturation for sense-oriented BS designs. This means that more than five binding sites provide only a relatively small increase in activity. For example, for sense-oriented BS, the median of all 2-BS-based designs is 1.69 times higher than the median of 1-BS-based designs, but going from 4 to 5 BSs only increases the median activity score by 1.09 times. Analysis of variance (ANOVA) showed that the distance between BSs had no significant effect on the activity scores of both sense- and antisense-based designs. Nevertheless, the highest scores were achieved by promoters with sense-oriented BSs spaced 1, 4, 5, and 8 bp apart ( Figure 14 E). On the other hand, the effect on the activity score of the base pair spacing between the minimal promoter and the proximal BS was highly significant. For the design based on the sense orientation, there was a positive effect for spacings greater than 5 bp, with a peak at 8 bp ( Figure 14 E). Notably, for designs in which the BS was in an antisense orientation relative to the minimal promoter, a greater distance between the proximal BS and the minimal promoter was associated with a lower activity score ( Figure 14 F).

[0407] The inventors next sought to facilitate potential future applications and created a series of promoters with different strengths to meet the needs of potential future applications. The 20 randomly selected variants mostly produced weak promoters. To maximize the diversity of size and strength, the inventors selected a further 19 promoters based on their high activity scores or small gene size. In summary, the inventors created a promoter collection across a wide range of sizes and strengths ( Figure 15 A, B). After co-transfection of each variant (with or without an amount of MTF plasmid corresponding to EC90) and an mCherry transfection control plasmid into HEK293 cells, relative mCitrine levels were measured using flow cytometry. As a result of repeated screening, the number of base pairs separating the promoter components had a strong effect on the expression level of the reporter gene - this allowed the inventors to find small-sized promoters that mediated high expression levels. For example, a promoter with a sense-2-0-8 structure (orientation - number of BSs - distance between BSs - distance to the minimum promoter (SEQ ID NO: 300) mediated as much mCitrine expression as the sense-5-9-8 design (SEQ ID NO: 672), but required 102 bp less to encode. In contrast, the two promoter designs sense-5-9-8 (SEQ ID NO: 672) and sense-5-8-10 (SEQ ID NO: 693) differed in length by only 2 bp, but the latter mediated twice as much mCitrine expression as the former ( Figure 15 B, D). Confirming responsiveness to MTF, reporter gene expression was observed only in cells co-transfected with the MTF-encoding plasmid but not in cells lacking this plasmid ( Figure 15 C).

[0408] While most attention has been directed to proteins in mammalian synTF engineering, the inventors hereby investigate and highlight the significant impact of homologous promoter design parameters. In addition to the importance of the number of binding sites, the spacing between promoter components has a profound impact on output expression strength. These surprising findings allowed the inventors to construct a series of promoters with a range of output strengths and find variants that mediate high expression levels, yet require only 52 bases to encode, regardless of the minimal promoter (i.e., sense-2-0-8; SEQ ID NO: 300).

[0409] References The following references refer to the reference numbers in parentheses, i.e., [1] to

[24] , as disclosed herein (including the accompanying Examples).

[0410] [1]L.Naldini,“Gene therapy returns to centre stage,”Nature,vol.526,no.7573,pp.351–360,Oct.2015,doi:10.1038 / nature15818.

[0411] [2]S.Merlin and A.Follenzi,“Transcriptional Targeting and MicroRNARegulation of Lentiviral Vectors,”Mol.Ther.-Methods Clin.Dev.,vol.12,pp.223–232,Mar.2019,doi:10.1016 / j.omtm.2018.12.013.

[0412] [3]Z.Xie,L.Wroblewska,L.Prochazka,R.Weiss,and Y.Benenson,“Multi-InputRNAi-Based Logic Circuit for Identification of Specific Cancer Cells,”Science,vol.333,no.6047,pp.1307–1311,Sep.2011,doi:10.1126 / science.1205527.

[0413] [4]L.Morsut et al.,“Engineering Customized Cell Sensing and ResponseBehaviors Using Synthetic Notch Receptors,”Cell,vol.164,no.4,pp.780–791,Feb.2016,doi:10.1016 / j.cell.2016.01.012.

[0414] [5]M.Gossen and H.Bujard,“Tight control of gene expression inmammalian cells by tetracycline-responsive promoters.,”Proc.Natl.Acad.Sci.,vol.89,no.12,pp.5547–5551,Jun.1992,doi:10.1073 / pnas.89.12.5547.

[0415] [6]M.Latta-Mahieu et al.,“Gene Transfer of a Chimeric Trans-ActivatorIs Immunogenic and Results in Short-Lived Transgene Expression,”Hum.GeneTher.,vol.13,no.13,pp.1611–1620,Sep.2002,doi:10.1089 / 10430340260201707.

[0416] [7]L.E.Mays and J.M.Wilson,“The Complex and Evolving Story of T cellActivation to AAV Vector-encoded Transgene Products,”Mol.Ther.,vol.19,no.1,pp.16–27,Jan.2011,doi:10.1038 / mt.2010.250.

[0417] [9]S.A.Hoyng et al.,“Developing a potentially immunologicallyinerttetracycline-regulatable viral vector for gene therapy in the peripheralnerve,”Gene Ther.,vol.21,no.6,pp.549–557,Jun.2014,doi:10.1038 / gt.2014.22.

[0418]

[10] X.Wang,F.G.Cabrera,K.L.Sharp,D.M.Spencer,A.E.Foster,andJ.H.Bayle,“Engineering Tolerance toward Allogeneic CAR-T Cells by Regulationof MHC Surface Expression with Human Herpes Virus-8 Proteins,”Mol.Ther.,Oct.2020,doi:10.1016 / j.ymthe.2020.10.019.

[0419]

[11] L.Peraro et al.,“Incorporation of bacterial immunoevasins toprotect cell therapies from host antibody-mediated immune rejection,”Mol.Ther.,Jul.2021,doi:10.1016 / j.ymthe.2021.06.022.

[0420]

[12] A.Annoni,B.D.Brown,A.Cantore,L.S.Sergi,L.Naldini,and M.-G.Roncarolo,“In vivo delivery of a microRNA-regulated transgene inducesantigen-specific regulatory T cells and promotes immunologic tolerance,”Blood,vol.114,no.25,pp.5152–5161,Dec.2009,doi:10.1182 / blood-2009-04-214569.

[0421]

[13] F.Boisgerault et al.,“Prolonged Gene Expression in Muscle IsAchieved Without Active Immune Tolerance Using MicrorRNA 142.3p-RegulatedrAAV Gene Transfer,”Hum.Gene Ther.,vol.24,no.4,pp.393–405,Feb.2013,doi:10.1089 / hum.2012.208.

[0422]

[14] Y.K.Chan et al.,“Engineering adeno-associated viral vectors toevade innate immune and inflammatory responses,”Sci.Transl.Med.,vol.13,no.580,p.eabd3438,Feb.2021,doi:10.1126 / scitranslmed.abd3438.

[0423]

[15] V.M.Rivera et al.,“A humanized system for pharmacologic controlof gene expression,”Nat.Med.,vol.2,no.9,pp.1028–1032,Sep.1996,doi:10.1038 / nm0996-1028.

[0424]

[16] I.Zhu et al.,“Modular design of synthetic receptors forprogrammed gene regulation in cell therapies,”Cell,vol.185,no.8,pp.1431-1443.e16,Apr.2022,doi:10.1016 / j.cell.2022.03.023.

[0425]

[17] I.Zhu et al.,“Design and modular assembly of syntheticintramembrane proteolysis receptors for custom gene regulation in therapeuticcells,”bioRxiv,p.2021.05.21.445218,May 2021,doi:10.1101 / 2021.05.21.445218.

[0426]

[18] J.Tycko et al.,“High-Throughput Discovery and Characterization ofHuman Transcriptional Effectors,”Cell,vol.0,no.0,Dec.2020,doi:10.1016 / j.cell.2020.11.024.

[0427]

[19] M.Fussenegger et al.,“Streptogramin-based gene regulation systemsfor mammalian cells,”Nat.Biotechnol.,vol.18,no.11,Art.no.11,Nov.2000,doi:10.1038 / 81208.

[0428]

[20] J.M.O’shea and N.D.Perkins,“Regulation of the RelA(p65)transactivation domain,”Biochem.Soc.Trans.,vol.36,no.4,pp.603–608,Jul.2008,doi:10.1042 / BST0360603.

[0429]

[21] E.Yakubovskaya,E.Mejia,J.Byrnes,E.Hambardjieva,and M.Garcia-Diaz,“Helix Unwinding and Base Flipping Enable Human MTERF1 to TerminateMitochondrial Transcription,”Cell,vol.141,no.6,pp.982–993,Jun.2010,doi:10.1016 / j.cell.2010.05.018.

[0430]

[22] L.Klein,B.Kyewski,P.M.Allen,and K.A.Hogquist,“Positive andnegative selection of the T cell repertoire:what thymocytes see(and don’tsee),”Nat.Rev.Immunol.,vol.14,no.6,pp.377–391,Jun.2014,doi:10.1038 / nri3667.

[0431]

[23] V.Jurtz,S.Paul,M.Andreatta,P.Marcatili,B.Peters,and M.Nielsen,“NetMHCpan-4.0:Improved Peptide–MHC Class I Interaction PredictionsIntegrating Eluted Ligand and Peptide Binding Affinity Data,”J.Immunol.,vol.199,no.9,pp.3360–3368,Nov.2017,doi:10.4049 / jimmunol.1700893.

[0432]

[24] M.D.Robinson,D.J.McCarthy,and G.K.Smyth,“edgeR:a Bioconductorpackage for differential expression analysis of digital gene expressiondata,”Bioinformatics,vol.26,no.1,pp.139–140,Jan.2010,doi:10.1093 / bioinformatics / btp616.

[0433]

[25] J.T.Bulcha,Y.Wang,H.Ma,P.W.L.Tai,and G.Gao,“Viral vectorplatforms within the gene therapy landscape,”Sig Transduct Target Ther,vol.6,no.1,Art.no.1,Feb.2021,doi:10.1038 / s41392-021-00487-6.

[0434]

[26] J.M.O’shea and N.D.Perkins,“Regulation of the RelA(p65)transactivation domain,”Biochemical Society Transactions,vol.36,no.4,pp.603–608,Jul.2008,doi:10.1042 / BST0360603.

[0435]

[27] J.Tycko et al.,“High-Throughput Discovery and Characterization ofHuman Transcriptional Effectors,”Cell,vol.0,no.0,Dec.2020,doi:10.1016 / j.cell.2020.11.024.

[0436]

[28] E.Yakubovskaya,E.Mejia,J.Byrnes,E.Hambardjieva,and M.Garcia-Diaz,“Helix Unwinding and Base Flipping Enable Human MTERF1 to TerminateMitochondrial Transcription,”Cell,vol.141,no.6,pp.982–993,Jun.2010,doi:10.1016 / j.cell.2010.05.018.

[0437]

[29] H.-S.Li et al.,“Multidimensional control of therapeutic humancell function with synthetic gene circuits,”Science,Dec.2022,doi:10.1126 / science.ade0156.

[0438]

[30] J.H.Bayle,J.S.Grimley,K.Stankunas,J.E.Gestwicki,T.J.Wandless,andG.R.Crabtree,“Rapamycin Analogs with Differential Binding Specificity PermitOrthogonal Control of Protein Activity,”Chemistry&Biology,vol.13,no.1,pp.99–107,Jan.2006,doi:10.1016 / j.chembiol.2005.10.017.

[0439]

[31] V.M.Rivera et al.,“A humanized system for pharmacologic controlof gene expression,”Nat Med,vol.2,no.9,pp.1028–1032,Sep.1996,doi:10.1038 / nm0996-1028.

[0440]

[32] M.J.Drummond et al.,“Rapamycin administration in humans blocksthe contraction-induced increase in skeletal muscle protein synthesis,”TheJournal of Physiology,vol.587,no.7,pp.1535–1546,2009,doi:10.1113 / jphysiol.2008.163816.

[0441]

[33] A.Mazéand Y.Benenson,“Artificial signaling in mammalian cellsenabled by prokaryotic two-component system,”Nat Chem Biol,vol.16,no.2,pp.179–187,Feb.2020,doi:10.1038 / s41589-019-0429-9.

[0442]

[34] P.S.Donahue,J.W.Draut,J.J.Muldoon,H.I.Edelstein,N.Bagheri,andJ.N.Leonard,“The COMET toolkit for composing customizable genetic programs inmammalian cells,”Nat Commun,vol.11,no.1,Art.no.1,Feb.2020,doi:10.1038 / s41467-019-14147-5.

[0443]

[35] R.P.Smith et al.,“Massively parallel decoding of mammalianregulatory sequences supports a flexible organizational model,”Nat Genet,vol.45,no.9,Art.no.9,Sep.2013,doi:10.1038 / ng.2713.

[0444]

[36] J.E.Davis,K.D.Insigne,E.M.Jones,Q.B.Hastings,and S.Kosuri,“Multiplexed dissection of a model human transcription factor binding sitearchitecture.”bioRxiv,p.625434,May 13,2019.doi:10.1101 / 625434.

[0445]

[37] E.Sharon et al.,“Inferring gene regulatory logic from high-throughput measurements of thousands of systematically designed promoters,”Nat Biotechnol,vol.30,no.6,Art.no.6,Jun.2012,doi:10.1038 / nbt.2205.

[0446]

[38] M.Hegde,C.Strand,R.E.Hanna,and J.G.Doench,“Uncoupling of sgRNAsfrom their associated barcodes during PCR amplification of combinatorialCRISPR screens,”PLOS ONE,vol.13,no.5,p.e0197547,May 2018,doi:10.1371 / journal.pone.0197547.

[0447]

[39] J.Liu et al.,“Extensive Recombination Due to HeteroduplexesGenerates Large Amounts of Artificial Gene Fragments during PCR,”PLOS ONE,vol.9,no.9,p.e106658,Sep.2014,doi:10.1371 / journal.pone.0106658.

Claims

1. A synthetic transcription factor comprising (i) a DNA binding domain (DBD) derived from a mitochondrial DNA binding protein and (ii) a transcriptional regulatory domain derived from one or more other proteins.

2. The synthetic transcription factor of claim 1, wherein the mitochondrial DNA binding protein is from a mammalian species.

3. The synthetic transcription factor of claim 2, wherein the mammalian species is human.

4. The synthetic transcription factor according to any one of claims 1 to 3, wherein the mitochondrial DNA binding protein is MTERF1.

5. The synthetic transcription factor according to any one of claims 1 to 4, wherein the DNA binding domain comprises: (i) a first MTERF1 motif having a sequence as shown in SEQ ID NO: 104 or a sequence having at least 80%, preferably at least 90%, more preferably at least 95% sequence identity thereto; and / or (ii) a second MTERF1 motif having a sequence as shown in SEQ ID NO: 106 or a sequence having at least 80%, preferably at least 90%, more preferably at least 95% sequence identity thereto; And preferably, the first MTERF1 motif is located at the N-terminus of the second MTERF1 motif.

6. The synthetic transcription factor according to claim 5, wherein the DNA binding domain further comprises (iii) an MTERF1 C-terminal domain having the sequence shown in SEQ ID NO: 11 or a sequence having at least 70%, preferably at least 80%, more preferably at least 90% sequence identity with SEQ ID NO:

11.

7. The synthetic transcription factor according to any one of claims 1 to 6, wherein the DNA binding domain comprises MTERF1 subdomain A having the sequence shown in SEQ ID NO: 9 or a sequence having at least 80%, preferably at least 90%, more preferably at least 95% sequence identity with SEQ ID NO: 9; In particular, wherein the first MTERF1 motif and / or the second MTERF1 motif is comprised in the subdomain A.

8. The synthetic transcription factor according to any one of claims 1 to 7, wherein the DNA binding domain comprises MTERF1 subdomain B, which has the sequence shown in SEQ ID NO: 7 or a sequence having at least 70%, preferably at least 80%, more preferably at least 90% sequence identity with SEQ ID NO: 7; in particular wherein the MTERF1 subdomain A is comprised in the MTERF1 subdomain B.

9. The synthetic transcription factor according to any one of claims 1 to 7, wherein the DNA binding domain has a sequence identity of at least 60%, preferably at least 70%, more preferably at least 80% with SEQ ID NO: 1; preferably, wherein the DNA binding domain has the sequence shown in SEQ ID NO:

1.

10. The synthetic transcription factor according to any one of claims 1 to 9, which does not comprise a mitochondrial transfer peptide having the sequence shown in SEQ ID NO: 37 or a sequence having at least 90% sequence identity to SEQ ID NO:

37.

11. The synthetic transcription factor according to any one of claims 1 to 10, which does not have a functional mitochondrial transfer peptide.

12. The synthetic transcription factor according to any one of claims 1 to 11, comprising a nuclear localization signal.

13. A synthetic transcription factor according to any one of claims 1 to 12, which is capable of regulating the transcription of at least one gene of interest in a cell, preferably in the nucleus of a cell; and preferably, wherein the gene of interest encodes a pro-cell death protein such as hBAX or HSV-TK, an immunostimulatory cytokine such as IL-2 or IL-12 and / or an antigen receptor such as CAR or TCR.

14. The synthetic transcription factor according to claim 13, wherein the synthetic transcription factor regulates the transcription of the gene of interest in a cell, preferably in the cell nucleus, by (i) promoting the transcription of the gene of interest or by (ii) inhibiting the transcription of the gene of interest.

15. The synthetic transcription factor according to any one of claims 1 to 14, wherein the synthetic transcription factor and / or the DNA binding domain is capable of binding to a response element in a cell, preferably in the nucleus of a cell.

16. The synthetic transcription factor according to claim 15, wherein the response element comprises an MTERF1 binding site having the sequence shown in SEQ ID NO: 42 or a sequence having at least 60%, preferably at least 80%, more preferably at least 90% sequence identity to SEQ ID NO:

42.

17. The synthetic transcription factor according to claim 16, wherein the response element comprises multiple copies of the binding site, such as 2 to 15 copies; preferably 2 to 5 copies.

18. A synthetic transcription factor according to claim 17, wherein two or more or all copies of the binding site in the response element are directly adjacent to each other or separated by one or more, such as 6 or 10 nucleotides.

19. The synthetic transcription factor according to any one of claims 15 to 18, which is capable of binding to a promoter comprising said response element in a cell, preferably in the nucleus of a cell.

20. The synthetic transcription factor according to claim 19, wherein the promoter further comprises a minimal promoter, for example, a minimal TATA box preferably having the sequence shown in SEQ ID NO: 103, or a minimal CMV promoter preferably having the sequence shown in SEQ ID NO: 136; and preferably, wherein the minimal promoter is located 3' of the response element.

21. The synthetic transcription factor according to any one of claims 19 to 20, wherein the promoter has the sequence shown in SEQ ID NO:

110.

22. The synthetic transcription factor according to any one of claims 19 to 21, wherein the promoter is operably linked to a gene of interest, preferably in the nucleus of the cell.

23. The synthetic transcription factor according to claim 22, wherein the binding of the synthetic transcription factor to the promoter, in particular the binding of the DNA binding domain to the response element within the promoter, regulates the transcription of a gene of interest operably linked to the promoter, e.g. in the nucleus of a cell; in particular, wherein the binding regulates the transcription of a gene of interest in the nucleus of a cell as defined in claim 13 or 14.

24. The synthetic transcription factor according to any one of claims 1 to 23, which is capable of entering the nucleus of a cell.

25. The synthetic transcription factor according to any one of claims 1 to 24, which has the ability to localize more efficiently to the nucleus of a cell than to the mitochondria of said cell; in particular, wherein said ability is determined by measuring the amount of the synthetic transcription factor in the nucleus and mitochondria, respectively, of the same cell.

26. The synthetic transcription factor according to any one of claims 1 to 25, which is unable to enter mitochondria.

27. The synthetic transcription factor according to any one of claims 13 to 26, wherein the cell and the mitochondrial DNA binding protein are from the same species, such as the same mammalian species.

28. A synthetic transcription factor according to claim 27 which does not substantially alter transcription, i.e. the transcriptome, in the cell, other than the transcription of the gene of interest.

29. The synthetic transcription factor according to claim 27 or 28, which does not specifically bind to substantially any endogenous DNA sequence in the cell.

30. The synthetic transcription factor of any one of claims 27 to 29, wherein the DNA binding domain does not specifically bind to substantially any endogenous DNA sequence in the nucleus of the cell.

31. The synthetic transcription factor of any one of claims 27 to 30 which does not substantially compete with the mitochondrial DNA binding protein from which the DBD is derived for sequence-specific DNA binding in the cell.

32. The synthetic transcription factor of any one of claims 27 to 31 which does not substantially interfere with the function of the mitochondrial DNA binding protein from which the DBD is derived.

33. A synthetic transcription factor according to any one of claims 1 to 32, wherein the transcriptional regulatory domain is capable of regulating the transcription of a gene, in particular when the transcriptional regulatory domain is part of, binds to or interacts with a DNA binding protein that is capable of binding to or interacting with a regulatory sequence of the gene, such as a promoter or enhancer.

34. A synthetic transcription factor according to any one of claims 1 to 33, wherein the transcriptional regulatory domain is capable of binding to and / or interacting with: an RNA polymerase, preferably RNA polymerase II; at least one other transcription factor, e.g. a general transcription factor; and / or at least one transcriptional co-regulator, such as a transcriptional co-activator and / or a transcriptional co-repressor.

35. The synthetic transcription factor of any one of claims 1 to 34, wherein the transcriptional regulatory domain is (i) an activation domain or (ii) an inhibition domain.

36. The synthetic transcription factor according to any one of claims 1 to 34, wherein the transcriptional regulatory domain is capable of (i) promoting the transcription of a gene, in particular when it is defined as an activation domain, or (ii) inhibiting the transcription of a gene, in particular when it is defined as an inhibition domain.

37. A synthetic transcription factor according to any one of claims 1 to 36, wherein the transcriptional regulatory domain, in particular the activation domain, has at least 60%, preferably at least 70%, more preferably at least 80% sequence identity with a sequence selected from the group consisting of SEQ ID NO: 3, SEQ ID NO: 33, SEQ ID NO: 52, SEQ ID NO: 27, SEQ ID NO: 29, SEQ ID NO: 54, SEQ ID NO: 56, SEQ ID NO: 58, SEQ ID NO: 17, SEQ ID NO: 19, SEQ ID NO: 60, SEQ ID NO: 62, SEQ ID NO: 13, SEQ ID NO: 64, SEQ ID NO: 66 and SEQ ID NO: 68; or wherein the transcriptional regulatory domain, in particular the activation domain, comprises at least one sequence having at least 60%, preferably at least 70%, more preferably at least 80% sequence identity with a sequence selected from the group consisting of SEQ ID NO: 3, SEQ ID NO: 33, SEQ ID NO: 52, SEQ ID NO: 27, SEQ ID NO: 29, SEQ ID NO: 54, SEQ ID NO: 56, SEQ ID NO: 58, NO:54, SEQ ID NO:56, SEQ ID NO:58, SEQ ID NO:17, SEQ ID NO:19, SEQ ID NO:60, SEQ ID NO:62, SEQ ID NO:13, SEQ ID NO:64, SEQ ID NO:66 and SEQ ID NO:

68.

38. The synthetic transcription factor of any one of claims 1 to 37, wherein the transcriptional regulatory domain comprises a first, a second and / or a third RELA transcriptional activation domain; wherein the first RELA transcriptional activation domain (TA1) has the sequence shown in SEQ ID NO: 31, or a sequence having at least 80% sequence identity to SEQ ID NO: 31; wherein the second RELA transcriptional activation domain has the sequence shown in SEQ ID NO: 132, or a sequence having at least 80% sequence identity to SEQ ID NO: 132; and wherein the third RELA transcriptional activation domain has the sequence shown in SEQ ID NO: 134, or a sequence having at least 80% sequence identity to SEQ ID NO: 134; and preferably, wherein the transcriptional regulatory domain comprises at least a first RELA transcriptional activation domain.

39. The synthetic transcription factor according to claim 38, wherein the transcriptional regulatory domain comprises multiple copies, such as two or three copies, of the first, second and / or third RELA domain, preferably the first RELA domain (TA1).

40. The synthetic transcription factor according to any one of claims 1 to 39, wherein the transcriptional regulatory domain comprises a RELA subdomain A having a sequence as shown in SEQ ID NO: 3 or a sequence having at least 70%, preferably at least 80%, more preferably at least 90% sequence identity thereto.

41. A synthetic transcription factor according to any one of claims 1 to 40, wherein the transcriptional regulatory domain comprises a RELA subdomain B having a sequence as shown in SEQ ID NO: 27 or a sequence having at least 70%, preferably at least 80%, more preferably at least 90% sequence identity with SEQ ID NO: 27; in particular, wherein the RELA subdomain A is contained in the RELA subdomain B.

42. A synthetic transcription factor according to any one of claims 1 to 41, wherein the transcriptional regulatory domain comprises a RELA subdomain C having a sequence as shown in SEQ ID NO: 29 or a sequence having at least 60%, preferably at least 70%, more preferably at least 80% sequence identity with SEQ ID NO: 29; in particular, wherein the first, second and / or third RELA motif, the RELA subdomain A and / or the RELA subdomain B are contained in the RELA subdomain C.

43. The synthetic transcription factor of any one of claims 1 to 42, wherein the transcriptional regulatory domain comprises a FOXO3 transcriptional activation domain (FOXO TAD) having the sequence shown in SEQ ID NO: 17 or a sequence having at least 80% sequence identity to SEQ ID NO:

17.

44. A synthetic transcription factor according to claim 43, wherein the transcriptional regulatory domain comprises multiple copies, such as two or three copies of the FOXO3 transcriptional activation domain (FOXO TAD); preferably, wherein the transcriptional regulatory domain has the sequence shown in SEQ ID NO: 33 or a sequence having at least 70%, preferably at least 80%, and more preferably at least 90% sequence identity with SEQ ID NO:

33.

45. The synthetic transcription factor of any one of claims 1 to 44, wherein the transcriptional regulatory domain comprises a MYB transcriptional activation domain (LMSTEN) having the sequence shown in SEQ ID NO: 17 or a sequence having at least 80% sequence identity to SEQ ID NO:

17.

46. ​​A synthetic transcription factor according to claim 45, wherein the transcriptional regulatory domain comprises multiple copies, such as two or three copies of the MYB transcriptional activation domain (LMSTEN); preferably, wherein the transcriptional regulatory domain has the sequence shown in SEQ ID NO: 196 or a sequence having at least 70%, preferably at least 80%, and more preferably at least 90% sequence identity with SEQ ID NO:

196.

47. The synthetic transcription factor of any one of claims 1 to 46, wherein the transcriptional regulatory domain comprises two copies of: (i) a first RELA transcriptional activation domain (TA1) having the sequence set forth in SEQ ID NO: 31 or a sequence having at least 80% sequence identity to SEQ ID NO: 31; (ii) a FOXO3 transcriptional activation domain (FOXO TAD) having the sequence shown in SEQ ID NO: 33 or a sequence having at least 80% sequence identity to SEQ ID NO: 33; or (iii) a MYB transcriptional activation domain (LMSTEN) having the sequence set forth in SEQ ID NO: 196 or a sequence having at least 80% sequence identity to SEQ ID NO: 196; Preferably, two copies of the FOXO3 transcriptional activation domain (FOXO TAD) are represented by SEQ ID NO: 33 or a sequence having at least 80% sequence identity to SEQ ID NO:

33.

48. The synthetic transcription factor of any one of claims 1 to 47, wherein the transcriptional regulatory domain comprises three copies of a first RELA transcriptional activation domain (TA1) having the sequence shown in SEQ ID NO: 31 or a sequence having at least 80% sequence identity to SEQ ID NO:

31.

49. The synthetic transcription factor according to any one of claims 1 to 48, wherein the transcriptional regulatory domain has a sequence as shown in SEQ ID NO: 33, SEQ ID NO: 35, SEQ ID NO: 194 or SEQ ID NO: 196, or a sequence having at least 70%, preferably at least 80%, more preferably at least 90% sequence identity to any of these sequences.

50. The synthetic transcription factor according to any one of claims 1 to 48, wherein the transcriptional regulatory domain has the sequence shown in SEQ ID NO: 33 or a sequence having at least 70%, preferably at least 80%, more preferably at least 90% sequence identity to SEQ ID NO:

33.

51. The synthetic transcription factor according to any one of claims 38 to 50, wherein the transcriptional regulatory domain comprises at least two domains independently selected from the group consisting of: the TA1, the FOXO TAD and the LMSTEN; and preferably, wherein the transcriptional regulatory domain comprises at least the FOXO TAD; and more preferably, wherein the transcriptional regulatory domain comprises at least the FOXO TAD, TA1 and the LMSTEN.

52. A synthetic transcription factor according to claim 51, wherein the transcriptional regulatory domain comprises (i) two FOXOTADs, (ii) FOXO TAD and LMSTEN, (iii) LMSTEN and TA1, (iv) FOXO TAD and TA1, or (v) FOXO TAD, LMSTEN and TA1; and preferably comprises FOXO TAD, in particular option (i), (ii), (iv) or (v).

53. The synthetic transcription factor according to any one of claims 1 to 51, wherein the transcriptional regulatory domain has a sequence as shown in SEQ ID NO: 52, SEQ ID NO: 54, SEQ ID NO: 56 or SEQ ID NO: 58, or a sequence having at least 70%, preferably at least 80%, more preferably at least 90% sequence identity to any of these sequences.

54. The synthetic transcription factor according to any one of claims 1 to 52, wherein the transcriptional regulatory domain has a sequence as shown in SEQ ID NO: 52, SEQ ID NO: 33, SEQ ID NO: 54 or SEQ ID NO: 58, or a sequence having at least 70%, preferably at least 80%, more preferably at least 90% sequence identity to any of these sequences.

55. The synthetic transcription factor according to any one of claims 51 to 54, wherein the transcriptional regulatory domain has the sequence shown in SEQ ID NO: 52 or a sequence having at least 70%, preferably at least 80%, more preferably at least 90% sequence identity with SEQ ID NO:

52.

56. The synthetic transcription factor according to any one of claims 51 to 55, wherein the DNA binding domain comprises or consists of the MTERF1 subdomain A having the sequence shown in SEQ ID NO: 9 or a sequence having at least 80%, preferably at least 90%, more preferably at least 95% sequence identity thereto; and preferably, wherein the transcriptional regulatory domain comprises or consists of a sequence as defined in claim 54 or 55.

57. A synthetic transcription factor according to any one of claims 37 to 45, 51 and 53, wherein the transcriptional regulatory domain has at least 60%, preferably at least 70%, more preferably at least 80% sequence identity with a sequence selected from the group consisting of: SEQ ID NO: 3, SEQ ID NO: 33, SEQ ID NO: 52, SEQ ID NO: 27, SEQ ID NO: 29, SEQ ID NO: 54, SEQ ID NO: 56 and SEQ ID NO:

58.

58. A synthetic transcription factor according to any one of claims 37 to 45, 51, 53 and 57, wherein the transcriptional regulatory domain has at least 70%, preferably at least 80%, more preferably at least 90% sequence identity with a sequence selected from the group consisting of: SEQ ID NO: 3, SEQ ID NO: 33 and SEQ ID NO:

52.

59. A synthetic transcription factor according to any one of claims 1 to 58, wherein the transcriptional regulatory domain comprises a sequence having at least 60%, preferably at least 70%, more preferably at least 80% sequence identity to a sequence selected from the group consisting of: SEQ ID NO: 60, SEQ ID NO: 62, SEQ ID NO: 13, SEQ ID NO: 64, SEQ ID NO: 66 and SEQ ID NO: 68, preferably in addition to the first, second and / or third RELA motif, such as the TA1, the RELA subdomain A, B or C, the FOXO TAD and / or the LMSTEN.

60. The synthetic transcription factor according to any one of claims 1 to 36, wherein the transcriptional regulatory domain, in particular the repression domain, has at least 60%, preferably at least 70%, more preferably at least 80% sequence identity with a sequence selected from the group consisting of SEQ ID NO: 70, SEQ ID NO: 72, SEQ ID NO: 74, SEQ ID NO: 76, SEQ ID NO: 78, SEQ ID NO: 80, SEQ ID NO: 82, SEQ ID NO: 84, SEQ ID NO: 86, SEQ ID NO: 88, SEQ ID NO: 90, SEQ ID NO: 92, SEQ ID NO: 94, SEQ ID NO: 96 and SEQ ID NO: 98; or wherein the transcriptional regulatory domain, in particular the repression domain, comprises at least one sequence having at least 60%, preferably at least 70%, more preferably at least 80% sequence identity with a sequence selected from the group consisting of SEQ ID NO: 70, SEQ ID NO: 72, SEQ ID NO: 74, SEQ ID NO: 76, SEQ ID NO: 78, SEQ ID NO: 80, SEQ ID NO: 82, SEQ ID NO: 84, SEQ ID NO: 86, SEQ ID NO: 88, SEQ ID NO: 90, SEQ ID NO: 92, SEQ ID NO: 94, SEQ ID NO: 96 and SEQ ID NO:

98. ID NO:82, SEQ ID NO:84, SEQ ID NO:86, SEQ ID NO:88, SEQ ID NO:90, SEQ ID NO:92, SEQ ID NO:94, SEQ ID NO:96 and SEQ ID NO:

98.

61. A synthetic transcription factor according to claim 60, wherein the transcriptional regulatory domain has at least 60%, preferably at least 70%, more preferably at least 80% sequence identity with a sequence selected from the group consisting of: preferably SEQ ID NO: 70, SEQ ID NO: 72, SEQ ID NO: 74 or SEQ ID NO:

76.

62. The synthetic transcription factor according to claim 60 or 61, wherein the transcriptional regulatory domain has at least 60%, preferably at least 70%, more preferably sequence identity to SEQ ID NO:

70.

63. The synthetic transcription factor of any one of claims 1 to 62, wherein the one or more proteins from which the transcriptional regulatory domain is derived are from the same species as the mitochondrial DNA binding protein, such as the same mammalian species.

64. A synthetic transcription factor according to any one of claims 1 to 63, which essentially comprises protein parts from the same species, such as the same mammalian species.

65. A synthetic transcription factor according to any one of claims 1 to 58 and 60 to 64, wherein the mitochondrial DNA binding protein, such as MTERF1, and the one or more other proteins from which the transcriptional regulatory domain is derived, such as RELA, FOXO3 and / or MYB, are of human origin.

66. A synthetic transcription factor according to any one of claims 1 to 58 and 60 to 65, which essentially comprises part of a human protein.

67. A synthetic transcription factor according to claim 65 or 66, wherein the transcriptional regulatory domain is derived from a single human protein.

68. A synthetic transcription factor according to any one of claims 65 to 67, wherein the transcriptional regulatory domain, in particular the activation domain, has a sequence identity of at least 90%, preferably at least 95%, more preferably at least 99% to SEQ ID NO: 3, SEQ ID NO: 27 or SEQ ID NO: 29, preferably to SEQ ID NO:

3.

69. A synthetic transcription factor according to any one of claims 65 to 67, wherein the transcriptional regulatory domain, in particular the repression domain, has at least 90%, preferably at least 95%, more preferably at least 99% sequence identity to SEQ ID NO: 70, SEQ ID NO: 72, SEQ ID NO: 74, SEQ ID NO: 76, SEQ ID NO: 78, SEQ ID NO: 80, SEQ ID NO: 82, SEQ ID NO: 84, SEQ ID NO: 86, SEQ ID NO: 88, SEQ ID NO: 90, SEQ ID NO: 92, SEQ ID NO: 94, SEQ ID NO: 96 or SEQ ID NO: 98, preferably to SEQ ID NO:

70.

70. The synthetic transcription factor of any one of claims 1 to 69, which is substantially non-immunogenic in humans.

71. The synthetic transcription factor of any one of claims 1 to 69, which is substantially non-immunogenic in the mammalian species from which the mitochondrial DNA binding protein is derived.

72. The synthetic transcription factor of any one of claims 65 to 69 which is substantially non-immunogenic in humans.

73. The synthetic transcription factor according to any one of claims 1 to 71, further comprising a controllable domain, preferably a controllable destabilization domain or a controllable localization domain.

74. The synthetic transcription factor according to claim 73, wherein the controllable domain is controllable (i) by a chemical compound, preferably by a small molecule, or (ii) by light, preferably by a specific wavelength or a specific range of wavelengths.

75. The synthetic transcription factor according to claim 73 or 74, wherein the synthetic transcription factor comprising the controllable destabilization domain is stabilized or destabilized by the compound or light; and / or the synthetic transcription factor comprising the controllable localization domain is localized to the nucleus or cytoplasm of the cell by the compound or light, preferably to the nucleus of the cell.

76. A synthetic transcription factor according to any one of claims 73 to 75, wherein the controllable destabilization domain comprises an NS3 domain having the sequence shown in SEQ ID NO: 158 or a sequence having at least 80%, preferably at least 90%, more preferably at least 95% sequence identity to SEQ ID NO:

158.

77. The synthetic transcription factor of claim 76, wherein the synthetic transcription factor comprising the NS3 domain is stabilized by grazoprevir.

78. The synthetic transcription factor according to any one of claims 73 to 77, wherein the controllable localization domain comprises an ERT2 domain having the sequence shown in SEQ ID NO: 152 or a sequence having at least 80%, preferably at least 90%, more preferably at least 95% sequence identity thereto.

79. The synthetic transcription factor of claim 78, wherein the synthetic transcription factor comprising the ERT2 domain is localized to the nucleus of a cell by 4-hydroxytamoxifen.

80. The synthetic transcription factor according to any one of claims 73 to 80, wherein the controllable domain comprises a FRB domain and a FKBP domain, wherein the FRB domain has a sequence as shown in SEQ ID NO: 140 or a sequence having at least 80%, preferably at least 90%, more preferably at least 95% sequence identity thereto, and / or wherein the FKBP domain has a sequence as shown in SEQ ID NO: 142 or a sequence having at least 80%, preferably at least 90%, more preferably at least 95% sequence identity thereto.

81. The synthetic transcription factor of claim 80, wherein the FRB domain and the FKBP domain bind to each other in the presence of C16-(S)-7-methylindorapamycin.

82. The synthetic transcription factor according to any one of claims 1 to 81, further comprising a synNotch core having the sequence shown in SEQ ID NO: 160 or a sequence having at least 80%, preferably at least 90%, more preferably at least 95% sequence identity thereto.

83. The synthetic transcription factor according to any one of claims 1 to 82, which is a fusion protein comprising the DNA binding domain and the transcriptional regulatory domain.

84. The synthetic transcription factor of claim 83, wherein the DNA binding domain is located N-terminally to the transcriptional regulatory domain.

85. The synthetic transcription factor according to claim 83 or 84, wherein the fusion protein further comprises a controllable destabilization domain as defined in any one of claims 73 to 77 or a controllable localization domain as defined in any one of claims 73 to 75, 78 and 79.

86. A synthetic transcription factor according to claim 83 or 84, wherein the fusion protein further comprises a single-chain variable fragment (scFv) and a synNotch core as defined in claim 82; and preferably, wherein the domains in the fusion protein are in the following order from N-terminus to C-terminus: (i) scFv, (ii) synNotch, and (iii) a DNA binding domain and a transcriptional regulatory domain, wherein the DNA binding domain may be located at the N-terminus or C-terminus of the transcriptional regulatory domain.

87. A nucleic acid encoding the synthetic transcription factor of any one of claims 83 to 86.

88. The nucleic acid of claim 87, which is DNA or RNA.

89. The nucleic acid of claim 87 or 88, which is mRNA, such as mRNA contained in a lipid nanoparticle.

90. The nucleic acid according to claim 87, comprising the DNA sequence shown in SEQ ID NO:5 or a DNA sequence having at least 60%, preferably at least 70%, more preferably at least 80% sequence identity with SEQ ID NO:5, and wherein the DNA sequence encodes a DNA binding domain having at least 60%, preferably at least 70%, more preferably at least 80% sequence identity with SEQ ID NO:1; preferably, wherein the DNA binding domain has the sequence shown in SEQ ID NO:

1.

91. A DNA plasmid comprising the nucleic acid of any one of claims 87 to 90, preferably wherein the plasmid is suitable for expressing a synthetic transcription factor in a cell.

92. A viral vector comprising the nucleic acid of any one of claims 87 to 90, the plasmid of claim 91, or a nucleic acid having a complementary sequence to the nucleic acid of any one of claims 87 to 90.

93. A cell comprising the nucleic acid of any one of claims 87 to 90, the plasmid of claim 91, or the viral vector of claim 92.

94. The synthetic transcription factor of any one of claims 1 to 82, comprising or consisting of a first polypeptide and a second polypeptide, wherein the first polypeptide comprises the DNA binding domain and the second polypeptide comprises the transcriptional regulatory domain.

95. The synthetic transcription factor of claim 94, wherein the first polypeptide and the second polypeptide each comprise a multimerization domain, and wherein the multimerization domains of the first polypeptide and the second polypeptide are capable of binding to and / or interacting with each other.

96. The synthetic transcription factor of claim 95, wherein the multimerization domain is a dimerization domain.

97. The synthetic transcription factor according to claim 95 or 96, wherein the multimerization domain is a homodimerization domain, in particular, wherein the multimerization domains of the first polypeptide and the second polypeptide are identical to each other.

98. The synthetic transcription factor according to claim 95 or 96, wherein the multimerization domain is a heterodimerization domain, in particular, wherein the multimerization domains of the first polypeptide and the second polypeptide are different from each other.

99. The synthetic transcription factor of claim 98, wherein (i) the multimerization domain of the first polypeptide comprises or consists of a SYNZIP1 domain, and the multimerization domain of the second polypeptide comprises or consists of a SYNZIP2 domain; or (ii) the multimerization domain of the first polypeptide comprises or consists of a SYNZIP2 domain, and The multimerization domain of the second polypeptide comprises or consists of a SYNZIP1 domain; wherein the SYNZIP1 domain has the sequence shown in SEQ ID NO: 154 or a sequence having at least 80%, preferably at least 90%, more preferably at least 95% sequence identity to SEQ ID NO: 154; and The SYNZIP2 domain has the sequence shown in SEQ ID NO: 156 or a sequence having at least 80%, preferably at least 90%, more preferably at least 95% sequence identity with SEQ ID NO:

156.

100. The synthetic transcription factor according to any one of claims 95 to 99, wherein the multimerization domain is a controllable domain, in particular, wherein the multimerization domain is controllable (i) by a compound, preferably by a small molecule, or (ii) by light, preferably by a specific wavelength or a specific range of wavelengths.

101. A synthetic transcription factor according to claim 100, wherein the multimerization domain is a dimerization domain, and wherein the compound or light controls the dimerization of the multimerization domain of the first polypeptide and the second polypeptide; and preferably, wherein the first polypeptide and the second polypeptide bind to and / or interact with each other in the presence of the small molecule or light.

102. The synthetic transcription factor of any one of claims 94 to 96, 98, 100 and 101, wherein (i) the multimerization domain of the first polypeptide comprises or consists of an FKBP domain, and the multimerization domain of the second polypeptide comprises or consists of an FRB domain; or (ii) the multimerization domain of the first polypeptide comprises or consists of a FRB domain, and The multimerization domain of the second polypeptide comprises or consists of an FKBP domain; wherein the FKBP domain has the sequence shown in SEQ ID NO: 142 or a sequence having at least 80%, preferably at least 90%, more preferably at least 95% sequence identity to SEQ ID NO: 142; and wherein the FRB domain has the sequence shown in SEQ ID NO: 140 or a sequence having at least 80%, preferably at least 90%, more preferably at least 95% sequence identity to SEQ ID NO: 140; And preferably, the multimerization domain of the first polypeptide comprises or consists of the FKBP domain, and the multimerization domain of the second polypeptide comprises or consists of the FRB domain.

103. The synthetic transcription factor of claim 102, wherein the DNA binding domain is located N-terminally to the FKBP domain in the first polypeptide, and / or wherein the transcriptional regulatory domain is located C-terminally to the FRB domain in the second polypeptide.

104. The synthetic transcription factor according to claim 102 or 103, wherein the first polypeptide and the second polypeptide bind to and / or interact with each other in the presence of C16-(S)-7-methylindorapamycin, in particular, wherein C16-(S)-7-methylindorapamycin induces heterodimerization of the FKBP domain and the FRB domain.

105. The synthetic transcription factor according to any one of claims 102 to 104, wherein (i) the DNA binding domain and the FKBP domain are linked to each other via a first peptide linker, and / or (ii) the transcriptional regulatory domain and the FRB domain are linked to each other via a second peptide linker.

106. A synthetic transcription factor according to claim 105, wherein the first peptide linker and the second peptide linker are independently selected from the group consisting of: a cMyc NLS linker shown in SEQ ID NO: 122, a 6AP (AP6) linker shown in SEQ ID NO: 120, an AP8 linker shown in SEQ ID NO: 144, a G4S linker shown in SEQ ID NO: 118, an EAAAK3 linker shown in SEQ ID NO: 146, an EAAAK2 linker shown in SEQ ID NO: 148, and a G4S4 linker shown in SEQ ID NO:

150.

107. The synthetic transcription factor of claim 105, wherein the first peptide linker is a cMyc NLS linker as shown in SEQ ID NO: 122 or a 6AP (AP6) linker as shown in SEQ ID NO: 120; and / or wherein the second peptide linker is a 6AP (AP6) linker as shown in SEQ ID NO:

120.

108. The synthetic transcription factor according to any one of claims 102 to 107, wherein the transcriptional regulatory domain is as defined in any one of claims 51 to 58.

109. The synthetic transcription factor according to any one of claims 102 to 108, wherein the transcriptional regulatory domain has a sequence as shown in SEQ ID NO: 52, SEQ ID NO: 29, SEQ ID NO: 33, SEQ ID NO: 54 or SEQ ID NO: 58, or a sequence having at least 70%, preferably at least 80%, more preferably at least 90% sequence identity to any of these sequences.

110. The synthetic transcription factor according to any one of claims 102 to 108, wherein the transcriptional regulatory domain has the sequence shown in SEQ ID NO: 52 or SEQ ID NO: 29, or a sequence having at least 70%, preferably at least 80%, more preferably at least 90% sequence identity thereto.

111. The synthetic transcription factor according to any one of claims 102 to 110, wherein the first polypeptide comprises the sequence shown in SEQ ID NO: 174, 176, 178, 180, 182, 184, 186, 188 or 192, or a sequence having at least 70%, preferably at least 80%, more preferably at least 90% sequence identity thereto.

112. The synthetic transcription factor according to any one of claims 102 to 111, wherein the second polypeptide comprises the sequence shown in SEQ ID NO: 162, 164, 166, 168, 170, 172 or 190, or a sequence having at least 70%, preferably at least 80%, more preferably at least 90% sequence identity to SEQ ID NO: 162, 164, 166, 168, 170, 172 or 190.

113. A combination of nucleic acids encoding the synthetic transcription factor of any one of claims 1 to 82, wherein one nucleic acid encodes the DNA binding domain and another nucleic acid encodes the transcriptional regulatory domain.

114. A nucleic acid combination encoding the synthetic transcription factor of any one of claims 94 to 112, comprising a first nucleic acid and a second nucleic acid, wherein the first nucleic acid encodes the first polypeptide and the second nucleic acid encodes the second polypeptide.

115. The combination according to claim 113 or 114, wherein one of the nucleic acids, in particular the first nucleic acid, has the DNA sequence shown in SEQ ID NO:5 or a DNA sequence having at least 60%, preferably at least 70%, more preferably at least 80% sequence identity with SEQ ID NO:5, and wherein the DNA sequence encodes a DNA binding domain having at least 60%, preferably at least 70%, more preferably at least 80% sequence identity with SEQ ID NO:1; preferably, wherein the DNA binding domain has the sequence shown in SEQ ID NO:

1.

116. The combination of claim 113 or 114, wherein the nucleic acid is a DNA or RNA molecule.

117. The combination of claim 116, wherein the nucleic acid is an mRNA molecule, such as an mRNA contained in a lipid nanoparticle.

118. A combination of DNA plasmids comprising a combination of nucleic acids according to any one of claims 113 to 115, preferably wherein the plasmids are suitable for expressing a synthetic transcription factor in a cell.

119. A viral vector combination comprising the nucleic acid of any one of claims 113 to 115, the plasmid of claim 119, or a nucleic acid having a complementary sequence to the nucleic acid of any one of claims 113 to 115.

120. A cell comprising the combination of any one of claims 113 to 119.

121. A DNA construct comprising a promoter (P), said promoter (P) comprising a response element and a minimal promoter, said DNA construct being characterized in that said response element comprises an MTERF1 binding site, said MTERF1 binding site comprising or consisting of a sequence as shown in SEQ ID NO: 42 or SEQ ID NO: 200, or a sequence having at least 60%, preferably at least 80%, more preferably at least 90% sequence identity thereto.

122. The DNA construct according to claim 121, wherein the MTERF1 binding site comprises or consists of the sequence shown in SEQ ID NO: 42 or a sequence having at least 60%, preferably at least 80%, more preferably at least 90% sequence identity to SEQ ID NO:

42.

123. The DNA construct of claim 121 or 122, wherein the response element consists of: (i) one or more copies, such as 2 to 50 copies, preferably 2 to 5 copies, of the MTERF1 binding site, wherein the multiple copies are directly adjacent to each other, or (ii) multiple copies, such as 2 to 50 copies, preferably 2 to 5 copies, of the MTERF1 binding site, and a BS-BS spacer between at least two, preferably all consecutive copies of the binding site, wherein the BS-BS spacer has a length of 1 to 1000 nucleotides, preferably 1 to 100 nucleotides, more preferably 1 to 10 nucleotides, such as 1 to 6 nucleotides; and preferably, wherein the BS-BS spacer consists of 1 to 10 nucleotides at the 5' end of the sequence shown in SEQ ID NO:

198.

124. The DNA construct according to any one of claims 121 to 123, wherein the response element consists of multiple copies, such as 2 to 15 copies, of the MTERF1 binding site, which copies are directly adjacent to each other.

125. according to the DNA construct described in any one in claim 121 to 124, the length that wherein said DNA construct has is at the most about 10 6 preferably at most 10 5 More preferably, up to about 10,000 nucleotides.

126. The DNA construct according to any one of claims 121 to 125, wherein the response element and the minimal promoter are separated from each other by a RE-minP spacer having a length of up to about 2000 nucleotides, preferably up to about 200 nucleotides, more preferably up to about 20 nucleotides, most preferably about 6 to 10 nucleotides, for example about 6 or 8 nucleotides; preferably, wherein the RE-minP spacer consists of 1 to 10 nucleotides at the 5' end of the sequence shown in SEQ ID NO: 199; and preferably, the MTERF1 binding site comprises or consists of the sequence shown in SEQ ID NO: 42 or a sequence having at least 60%, preferably at least 80%, more preferably at least 90% sequence identity to SEQ ID NO:

42.

127. The DNA construct according to any one of claims 121 to 126, wherein the minimal promoter is located 3' or 5', preferably 3', of the response element.

128. A DNA construct according to any one of claims 121 to 127, wherein the minimal promoter is a minimal TATA box having at least 60%, preferably at least 80%, more preferably at least 90% sequence identity to SEQ ID NO:

103.

129. A DNA construct according to any one of claims 121 to 127, wherein the minimal promoter is a minimal CMV promoter having a sequence identity of at least 60%, preferably at least 80%, more preferably at least 90% to SEQ ID NO:

136.

130. The DNA construct according to any one of claims 121 to 129, wherein the promoter (P), in particular the response element, is capable of binding to the synthetic transcription factor according to any one of claims 1 to 86 and 94 to 112.

131. A DNA construct according to any one of claims 121 to 130, further comprising at least one gene of interest, preferably encoding a pro-cell death protein such as hBAX or HSV-TK, an immunostimulatory cytokine such as IL-2 or IL-12, and / or an antigen receptor such as CAR or TCR.

132. The DNA construct of claim 131, wherein the promoter (P) is operably linked to at least one gene of interest; preferably, wherein the at least one gene of interest is located 3' to a minimal promoter.

133. A DNA construct according to any one of claims 130 to 132, wherein at least one of the genes of interest is transcribed when the promoter (P), in particular the response element, is bound by the synthetic transcription factor according to any one of claims 1 to 86 and 94 to 112 in a cell.

134. A DNA construct according to any one of claims 130 to 133, wherein at least one of the genes of interest is transcribed when the promoter (P), in particular the response element, is bound by the synthetic transcription factor according to any one of claims 1 to 86 and 94 to 112 in the nucleus of a human cell.

135. The DNA construct of any one of claims 121 to 134, which does not comprise the sequence shown in SEQ ID NO: 117 or a sequence having at least 90% sequence identity to SEQ ID NO:

117.

136. A DNA construct according to any one of claims 121 to 135, comprising a sequence selected from the group consisting of SEQ ID NOs: 201 to 1190.

137. The DNA construct of any one of claims 121 to 135, wherein the response element comprises or consists of two copies of an MTERF1 binding site, each MTERF1 binding site having a sequence as shown in SEQ ID NO: 42 or a sequence having at least 60%, preferably at least 80%, more preferably at least 90% sequence identity thereto; wherein the two copies of the MTERF1 binding site are directly adjacent to each other; and wherein the response element and the minimal promoter are separated by 8 nucleotides from each other.

138. A DNA construct according to any one of claims 121 to 137, comprising the sequence shown in SEQ ID NO:

300.

139. The DNA construct of any one of claims 121 to 135, wherein the response element consists of 5 copies of an MTERF1 binding site, the copies being separated from each other by 8 nucleotides, and wherein each MTERF1 binding site has the sequence shown in SEQ ID NO: 42, or a sequence having at least 60%, preferably at least 80%, more preferably at least 90% sequence identity to SEQ ID NO: 42; and wherein the response element and the minimal promoter are separated from each other by 10 nucleotides.

140. A DNA construct according to any one of claims 121 to 135 or 139, comprising the sequence shown in SEQ ID NO:

693.

141. A single-stranded or double-stranded nucleic acid comprising the sense strand of the DNA construct of any one of claims 121 to 140 and / or the antisense strand of the DNA construct of any one of claims 121 to 140.

142. The single-stranded or double-stranded nucleic acid of claim 141, which is DNA or RNA.

143. A DNA plasmid comprising the DNA construct of any one of claims 121 to 140.

144. A viral vector comprising the DNA construct of any one of claims 121 to 140, the single-stranded or double-stranded nucleic acid of claim 141 or 142, or the DNA plasmid of claim 143.

145. A cell comprising the DNA construct of any one of claims 121 to 140.

146. A system comprising: (i) the synthetic transcription factor of any one of claims 1 to 86 and 94 to 112, the nucleic acid of any one of claims 87 to 93, the DNA plasmid of claim 94, the viral vector of claim 95, or the combination of any one of claims 113 to 119; and (ii) the DNA construct of any one of claims 121 to 140, the DNA plasmid of claim 143, and / or the viral vector of claim 144; preferably, wherein the system comprises the synthetic transcription factor of any one of claims 1 to 86 and 94 to 112 and the DNA construct of any one of claims 121 to 140; and preferably, wherein the system is an engineered gene network, preferably a biocomputing circuit.

147. A system according to claim 146 for regulating the transcription of at least one gene of interest, preferably, wherein the gene of interest is contained in the DNA construct, preferably located 3' to a minimal promoter; and preferably, wherein the gene of interest encodes a pro-cell death protein such as hBAX or HSV-TK, an immunostimulatory cytokine such as IL-2 or IL-12, and / or an antigen receptor such as CAR or TCR.

148. A cell comprising the synthetic transcription factor of any one of claims 1 to 86 and 94 to 112 and the DNA construct of any one of claims 121 to 140.

149. A kit comprising the synthetic transcription factor of any one of claims 1 to 86 and 94 to 112, the nucleic acid of any one of claims 87 to 93, the DNA plasmid of claim 94, the viral vector of claim 95, the combination of any one of claims 113 to 119, the DNA construct of any one of claims 121 to 140, the DNA plasmid of claim 143, the viral vector of claim 144 and / or the system of claim 146 or 147.

150. The kit of claim 149, comprising (i) the nucleic acid of any one of claims 87 to 93, the DNA plasmid of claim 94, the viral vector of claim 95, or the combination of any one of claims 113 to 119, and (ii) the DNA construct of any one of claims 121 to 140, the DNA plasmid of claim 143, or the viral vector of claim 144.

151. A pharmaceutical composition comprising the synthetic transcription factor of any one of claims 1 to 86 and 94 to 112, the nucleic acid of any one of claims 87 to 93, the DNA plasmid of claim 94, the viral vector of claim 95, the combination of any one of claims 113 to 119, the DNA construct of any one of claims 121 to 140, the single-stranded or double-stranded nucleic acid of claim 141 or 142, the DNA plasmid of claim 143, the viral vector of claim 144, the system of claim 146 or 147, or the cell of any one of claims 93, 120, 145 and 148.

152. The pharmaceutical composition of claim 151 , comprising: (i) the nucleic acid of any one of claims 87 to 93, the DNA plasmid of claim 94, the viral vector of claim 95, or the combination of any one of claims 113 to 119; and (ii) the DNA construct of any one of claims 121 to 140, the DNA plasmid of claim 143, or the viral vector of claim 144.

153. The pharmaceutical composition of claim 151 or 152, comprising the cells of claim 148.

154. The pharmaceutical composition of any one of claims 151 to 153, further comprising a pharmaceutically acceptable excipient.

155. A pharmaceutical composition according to any one of claims 151 to 154 for use in treating a disease, wherein target cells are killed and / or manipulated, in particular wherein the treatment involves a cancer cell classifier circuit; preferably, wherein at least one gene of interest encodes a pro-cell death protein such as hBAX or HSV-TK, an immunostimulatory cytokine such as IL-2 or IL-12, and / or an antigen receptor such as CAR or TCR.

156. A pharmaceutical composition according to any one of claims 151 to 154 for use in a method of treating a tumor or cancer, in particular, wherein the treatment involves a cancer cell classifier circuit; preferably, wherein at least one gene of interest encodes a pro-cell death protein such as hBAX or HSV-TK, an immunostimulatory cytokine such as IL-2 or IL-12, and / or an antigen receptor such as CAR or TCR.

157. The nucleic acid of any one of claims 87 to 93, the DNA plasmid of claim 94, the viral vector of claim 95, the combination of any one of claims 113 to 119, the DNA construct of any one of claims 121 to 140, the single-stranded or double-stranded nucleic acid of claim 141 or 142, the DNA plasmid of claim 143, the viral vector of claim 144, or the system of claim 146 or 147, for use in gene therapy.

158. The cell of any one of claims 93, 120, 145 and 148, for use in cell therapy.

159. The cell of claim 158, wherein the cell is a T cell, such as a CAR T cell, and the cell therapy is a T cell therapy, such as a CAR T cell therapy.

160. Use of the synthetic transcription factor of any one of claims 1 to 86 and 94 to 112, the nucleic acid of any one of claims 87 to 93, the DNA plasmid of claim 94, the viral vector of claim 95, the combination of any one of claims 113 to 119, the DNA construct of any one of claims 121 to 140, the single-stranded or double-stranded nucleic acid of claim 141 or 142, the DNA plasmid of claim 143, the viral vector of claim 144, the system of claim 146 or 147, or the cell of any one of claims 93, 120, 145 and 148 as part of or in combination with an engineered gene network, in particular a biological computational circuit.

161. Use of the synthetic transcription factor of any one of claims 1 to 86 and 94 to 112, the nucleic acid of any one of claims 87 to 93, the DNA plasmid of claim 94, the viral vector of claim 95, the combination of any one of claims 113 to 119, the DNA construct of any one of claims 121 to 140, the single-stranded or double-stranded nucleic acid of claim 141 or 142, the DNA plasmid of claim 143, the viral vector of claim 144, the system of claim 146 or 147, or the cell of any one of claims 93, 120, 145 and 148 for transcribing a gene of interest in a cell in vitro or in vivo, e.g., in vitro or in vivo.

162. A library of DNA constructs comprising at least about 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 150, 200, 250, 300, 350, 400, 450, 500, 550, 600, 650, 700, 750, 800, 850, 900, 950, 1000, 2000, 3000, 4000, 5000, 6000, 7000, 8000, 9000, 10000, 50000, 100000, 500000, or 10000000, preferably at least about 100, 500, 900, 950, 990, or 1000 different DNA constructs, wherein each DNA construct in the library comprises (i) a promoter (P) consisting of a response element (RE), a minimal promoter (minP) located 3' to the response element, and an optional RE-minP spacer between the response element and the minimal promoter; wherein each response element consists of one or more copies of a transcription factor binding site (BS) and an optional BS-BS spacer located between at least two, preferably all consecutive copies of the binding site; (ii) an export sequence, preferably located 3' to the minimal promoter; wherein all DNA constructs in the library differ from each other in their promoter (P) sequences; and Optionally, each DNA construct in the library comprises a unique barcode sequence that distinguishes all DNA constructs in the library from one another.

163. The library according to claim 162, wherein the transcription factor binding site is a MTERF1 binding site, and the MTERF1 binding site comprises the sequence shown in SEQ ID NO:42 or SEQ ID NO:200, or a sequence having at least 60%, preferably at least 80%, more preferably at least 90% sequence identity with SEQ ID NO:42 or SEQ ID NO:200, or consists of the same; preferably, the MTERF1 binding site consists of the sequence shown in SEQ ID NO:

42.

164. The library of claim 162 or 163, wherein the DNA constructs are identical to each other except for the promoter sequence and optional barcode sequence.

165. A library according to any one of claims 162 to 164, wherein the minimal promoters in the different DNA constructs, in particular the different promoters, are identical to each other.

166. The library according to any one of claims 162 to 165, wherein the transcription factor binding sites in different DNA constructs, in particular in different promoters, have the same sequence in sense orientation or antisense orientation relative to the minimal promoter, preferably SEQ ID NO: 42 or SEQ ID NO: 200, respectively.

167. A library according to any one of claims 162 to 166, wherein the promoters of at least 20%, 30%, 40% or 50% of the DNA constructs differ from each other in (i) binding site copy number and / or (ii) the presence or length of the RE-minP spacer.

168. The library of any one of claims 162 to 167, wherein at least 80%, at least 90% or all of the promoters of the DNA constructs differ from each other in at least one parameter selected from the group consisting of: (i) binding site copy number, (ii) presence or length of the RE-minP spacer, (iii) presence or length of the BS-BS spacer and (iv) sense or antisense orientation of the binding site sequence relative to the minimal promoter.

169. The library according to any one of claims 162 to 168, wherein the minimal promoter is a minimal TATA box having at least 60%, preferably at least 80%, more preferably at least 90% sequence identity to SEQ ID NO:

103.

170. The library according to any one of claims 162 to 168, wherein the minimal promoter is a minimal CMV promoter having at least 60%, preferably at least 80%, more preferably at least 90% sequence identity to SEQ ID NO:

136.

171. The library according to any one of claims 162 to 170, wherein the promoter (P), in particular the response element, is capable of binding to the synthetic transcription factor according to any one of claims 1 to 86 and 94 to 112.

172. A method for optimizing a promoter for binding to a transcription factor, the method comprising the steps of: a) preparing a library of DNA constructs as defined in any one of claims 161 to 171, b) combining the library of DNA constructs with said transcription factors in cells or in an in vitro transcription system, preferably in cells, c) determining the transcriptional activity of the promoter of each DNA construct in the library, preferably by determining the amount of mRNA produced from each DNA construct in the library, in particular, wherein said mRNA comprises a sequence corresponding to said output sequence, and d) Selecting a promoter based on its transcriptional activity, thereby obtaining a promoter optimized for binding to the transcription factor.

173. The method of claim 172, wherein the transcriptional activity of the promoter of each DNA construct in the library is determined by RNA sequencing, preferably by next generation RNA sequencing.

174. The method of claim 172 or 173, wherein the method is a massively parallel reporter gene assay.

175. The method of any one of claims 172 to 174, wherein the transcription factor is as defined in any one of claims 1 to 86 or 94 to 112.