Optimised promoter design for industrial application
Synthetic promoters with modified TFBSs in yeast systems enhance protein expression by replacing repressor sites with activator sites, achieving substantial yield improvements for recombinant proteins.
Patent Information
- Application Number
- PCT/EP2025/073942
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-08-26
- Filing Date
- 2025-08-21
- Publication Date
- 2026-03-05
AI Technical Summary
Existing endogenous promoters in yeast-based expression systems exhibit low yields and variable activity, limiting the production of recombinant proteins and hindering efficient and scalable protein production.
Development of synthetic promoters (synPs) with modified core and upstream regulatory domains by replacing repressor TFBSs with activator TFBSs, specifically targeting MIG1, MIG2, MIG3, R0X1, RFX1, MATALPHA2, RDR1, and Y0X1 repressors with activators like Cat8, Gal4, TATA box-binding protein, GATA box-binding protein, and Sip4, resulting in enhanced protein expression.
The synPs demonstrate significantly increased protein yields, including a 4.1 to 15.5-fold improvement in extracellular protein titers, showcasing their potential for industrial applications without strain engineering.
Smart Images

Figure EP2025073942_05032026_PF_FP_ABST
Abstract
Description
[0001] Optimised promoter design for industrial application
[0002] FIELD OF THE INVENTION
[0003] The invention relates to synthetic promoters for use in heterologous protein expression, as well as methods for identifying and generating novel synthetic promoters.
[0004] BACKGROUND TO THE INVENTION
[0005] Over the last decades, there has been an increasing focus on improving the optimization of production yeast strains, with the primary objective of enhancing productivity. In response to the growing demand for heterologous protein production in industry, many technologies have been developed to optimize the process. Among the various strategies employed, the expression strength of the heterologous gene has been identified as a key bottleneck in efforts to increase the titres of proteins of interest.
[0006] Promoters play a pivotal role in regulating the transcriptional activity of genes and the expression levels of recombinant proteins in yeast. Different types of promoters may have distinct strengths, regulation patterns, and other characteristics, which can impact the efficiency and specificity of protein production.
[0007] Although endogenous promoters are often used due to their inherent advantages, they have some limitations. Studies have shown that endogenous promoters may result in low yields of recombinant proteins and their activity can vary depending on the specific yeast strain and growth conditions, which can impact the production of the desired gene products.
[0008] To address these limitations, synthetic promoters (synPs) may be required to achieve higher levels of protein expression in yeast-based expression systems. Indeed, synPs have become valuable tools for fine-tuning gene expression levels in yeast expression systems. In particular, the use of synPs offers greater flexibility in gene expression regulation and improves the scalability and efficiency of protein production, leading to more sustainable and cost-effective bioprocessing.
[0009] However, there exists a need to develop and improve existing synPs, and in particular synPs that can achieve high levels of target gene expression. SUMAMRY OF THE INVENTION
[0010] We describe the creation of a set of synPs for yeast-based expression systems by removing repressor binding sites and adding activator binding sites. These synPs showed significantly increased activity and improved protein yields compared to endogenous promoters. Using a small-scale screening method, we were able to enhance the production of two target proteins - an intracellular protein and an extracellular protein. Our results showed that the synPs of the invention can improve the production of different proteins when compared to the endogenous promoters; this is proof of principle that the synPs of the invention have the remarkable capability to regulate the expression of any target protein can be used to increase the expression of any target protein. Bioreactor cultivations of the top performing strains also showed significant improvements in protein expression without any strain engineering, demonstrating also the significant potential industrial application of this invention.
[0011] Accordingly, in one aspect of the invention, there is provided a synP comprising a core promoter domain and an upstream regulatory domain, wherein the regulatory domain comprises or consists of one or more repressor and wherein at least one repressor has been or is replaced by an, or a plurality of, activator transcription factor binding site (TFBS(s)).
[0012] In one embodiment, the repressor TFBS binds a transcriptional repressor, wherein the transcriptional repressor is selected from MIG1, MIG2, MIG3, R0X1, RFX1, MATALPHA2, RDR1 and Y0X1. Preferably, the at least one repressor TFBS is selected from one of SEQ ID NO: 6 to 49 or 167 to 169 or 171 to 194 or a variant thereof.
[0013] In another embodiment, the activator TFBS binds to a transcriptional activator, wherein the transcriptional activator is selected from Cat8, Gal4, TATA box-binding protein, GATA box-binding protein, Adri and Sip4. Preferably, the at least one activator TFBS is selected from one of SEQ ID NO: 50 to 157 or a variant thereof.
[0014] In a further embodiment additionally at least one repressor TFBS is deleted (and not replaced by an activator TFBS). In one embodiment at least two, preferably at least four, more preferably between six and ten repressor TFBSs are replaced with at least two, preferably at least four, more preferably between six and ten activator TFBSs.
[0015] The promoter can exhibit enhanced performance when exposed to various carbon sources, such as ethanol or methanol. In one embodiment, the promoter may be a methanol-inducible promoter. In one embodiment, the promoter is selected from the alcohol dehydrogenase 2 gene (PAD ), alcohol oxidase 1 (PAOXI), or glyceradehyde-3- phosphate dehydrogenase (PGAP).
[0016] In one embodiment, the promoter PADH2 comprises or consists of a sequence defined in SEQ ID NO: 4.
[0017] In one embodiment, the promoter PAOXI comprises or consists of a sequence defined in SEQ ID NO: 5.
[0018] In one embodiment, the promoter PG P comprises or consists of a sequence defined in SEQ ID NO: 170.
[0019] In one embodiment, the promoter does not comprise any additional TFBSs as a consequence of replacement of the at least one repressor TFBS with at least one activator TFBS.
[0020] In one embodiment, the promoter comprises a sequence selected from SEQ ID NO: 1 , 2 or 3, or a variant thereof, wherein the variant has at least 60% overall sequence identity to one of SEQ ID NO: 1 , 2 or 3. Accordingly, the promoter may comprise or consist of SEQ ID NO: 1 or a variant thereof, wherein the variant has at least 60% overall sequence identity to SEQ ID NO: 1. Alternatively, the promoter may comprise or consist of SEQ ID NO: 2 or a variant thereof, wherein the variant has at least 60% overall sequence identity to SEQ ID NO: 2. Alternatively, the promoter may comprise or consist of SEQ ID NO: 3 or a variant thereof, wherein the variant has at least 60% overall sequence identity to SEQ ID NO: 3.
[0021] In another aspect of the invention, there is provided a nucleic acid construct comprising the synP described herein and a nucleic acid sequence encoding a target protein, wherein the synP is operably linked to the nucleic acid sequence encoding a target protein. The target protein may be any exogenous protein.
[0022] In one embodiment, the nucleic acid construct further comprises at least one nucleic acid sequence encoding a transcription factor (TF), wherein preferably the TF is selected from Cat8, Gal4, TATA box-binding protein, GATA box-binding protein, Adri and Sip4. The nucleic acid construct may also further comprise at least one helper factor.
[0023] In another aspect of the invention, there is provided a cell comprising the nucleic acid construct of the invention.
[0024] In another aspect of the invention, there is provided an organism comprising, and preferably capable of expressing or expressing the nucleic acid construct described herein, wherein preferably, the organism is a yeast. For example, the yeast is a methylotrophic yeast, preferably selected from the Pichia genus, most preferably selected from Pichia pastoris, Hansenula polymorpha, Candida boidinii or Pichia methanolica.
[0025] In another aspect of the invention, there is provided a method of heterologous protein production, the method comprising culturing the yeast described herein under conditions that allow expression of the nucleic acid encoding a target protein. The method may further comprise screening the yeast described herein and selecting yeast with the greatest expression of the target protein. Expression of the target protein may be compared to the level of expression in yeast that do not comprise the nucleic acid construct described herein. The method may further comprise extracting the exogenous protein.
[0026] In another aspect of the invention there is provided a method for producing a synP, the method comprising
[0027] (a) identifying a plurality of TFBSs within the upstream regulatory sequence of a promoter;
[0028] (b) determining whether said TFBS is a repressor TFBS or an activator TFBS; and
[0029] (c) replacing at least one TFBS that binds a transcriptional repressor with at least one TFBS that binds a transcriptional activator. Again, the promoter can exhibit enhanced performance when exposed to various carbon sources, such as ethanol or methanol. In one embodiment, the promoter may be a methanol-inducible promoter. In one embodiment, the promoter is selected fromP iDH2, or PGAP.
[0030] In one embodiment, the repressor TFBS binds a transcriptional repressor, wherein the transcriptional repressor is selected from MIG1, MIG2, MIG3, R0X1, RFX1, MATALPHA2, RDR1 and Y0X1. Preferably, the at least one repressor TFBS is selected from one of SEQ ID NO: 6 to 49 or 167 to 169 or 171 to 194 or a variant thereof.
[0031] In another embodiment, the activator TFBS binds to a transcriptional activator, wherein the transcriptional activator is selected from Cat8, Gal4, TATA box-binding protein, GATA box-binding protein, Adri and Sip4. Preferably, the activator TFBS is selected from one of SEQ ID NO: 50 to 157 or a variant thereof.
[0032] In another aspect of the invention there is provided a synP obtained or obtainable by the method described herein.
[0033] DESCRIPTION OF THE FIGURES
[0034] Features of examples of the present disclosure will become apparent by reference to the following detailed description and drawings, in which like reference numerals correspond to similar, though perhaps not identical, components. For the sake of brevity, reference numerals or features having a previously described function may or may not be described in connection with other drawings in which they appear.
[0035] Figure 1 : Graphical representation of the Matinspector prediction of the TFBSs of the endogenous and synthetic promoters.
[0036] Figure 2: The level of Psyn2, Psyn3, and Psyn8, relative expression of an enhanced green fluorescent protein_eGFP by FACS analysis. 24 individual clones were selected per promoter in a small-scale screening. The average expression level of each promoter was normalized using the expression of eGFP under the control of PGAP (as reference. PGAP is set to 1 . Figure 3: Comparative expression of target A protein across various promoters in topperforming strains (based on small screening data in 24DWP).
[0037] Figure 4: Comparative analysis of target A protein titers driven by synPs Psyn2, Psyn3, Psyn8 and PGAP in bioreactor cultivations.
[0038] Figure 5: The gene expression levels and their gene copy numbers of seven different colonies were tested. The expression levels of each clone were normalized to the expression of eGFP of colonies 1 (C1), which is set as 1 .0.
[0039] Figure 6: General overview of the small-scale screening protocol.
[0040] DETAILED DESCRIPTION
[0041] In the following passages, different aspects of the invention are defined in more detail. Each aspect so defined may be combined with any other aspect or aspects unless clearly indicated to the contrary. In particular, any feature indicated as being preferred or advantageous may be combined with any other feature or features indicated as being preferred or advantageous.
[0042] Abbreviations: transcription factor binding site (TFBS), transcription factor (TF), synthetic promoters (synPs), enhanced green fluorescent protein (eGFP), transcription starting signal (TSS), upstream activator sequences (UAS), upstream repressive sequences (URS).
[0043] In one aspect of the invention there is provided a synthetic promoter comprising a core promoter domain and an upstream regulatory domain, wherein the regulatory domain comprises one or more, more likely a plurality of repressor TFBSs and wherein at least one, preferably a plurality of, repressor TFBS(s) is replaced by an, preferably a plurality of, activator TFBS(s).
[0044] In another aspect of the invention, there is provided a method for producing a synP, the method comprising
[0045] (d) identifying a plurality of TFBSs within the upstream regulatory sequence of a promoter;
[0046] (e) determining whether said TFBS is a repressor TFBS or an activator TFBS; and replacing at least one TFBS that binds a transcriptional repressor with at least one TFBS that binds a transcriptional activator.
[0047] The term "promoter" typically refers to a nucleic acid control sequence located upstream from the transcriptional start of a gene and which is involved in the binding of RNA polymerase and other proteins, thereby directing transcription of an operably linked nucleic acid (e.g. a target sequence). Encompassed by the aforementioned terms are transcriptional regulatory sequences derived from a classical eukaryotic genomic gene (including the TATA box which is required for accurate transcription initiation, with or without a CCAAT box sequence) and additional regulatory elements (i.e. upstream activator sequences, enhancers and silencers) which alter gene expression in response to developmental and / or external stimuli. Also included within the term is a transcriptional regulatory sequence of a classical prokaryotic gene, in which case it may include a -35 box sequence and / or -10 box transcriptional regulatory sequences.
[0048] In yeast, the minimal modular yeast promoter consists of several basic parts in order, namely, hybrid regulatory domain (which may comprise an upstream activator sequence (UAS) or an upstream repressor sequence (URS), neutral AT-rich spacer, TATA-box, N30 core promoter, and transcriptional starting site (TSS). The TATA box may be a consensus TATA box (TATAWAW; SEQ ID NO: 158), or a TATA box with one or two mismatches to the consensus sequence. Within the upstream regulatory domain, there may be a carbon source responsive element (CSRE). CSREs mediate transcriptional activation of the gluconeogenic genes during growth of the yeast on non-fermentable carbon sources.
[0049] The core promoter may be considered to be the minimal stretch of contiguous DNA sequence that is sufficient to direct initiation of transcription by the RNA polymerase II machinery. Preserving the 150bp region from the 3'-UTR of endogenous promoters is our preference. For Psyn3 the core promoter would be the endogenous core promoter of the PADH2- For Psyn2 and Psyn8 the core promoter would be the endogenous core promoter of PGAP, except for a single modification. In this modification, the TFBS Y0X1 is replaced with TATA box. Accordingly, in one embodiment, the core promoter is selected from the core promoter sequence of PD SV, PD S2, PFLDI, PFDH2, PAOX2, PPEXS, PFGHI, PDAKI, PPEXS, PFBA2, PTPH, PFBPI, PPGU. In one embodiment the core promoter sequence is the core promoter sequence of PADH2 and PGAP. In a further embodiment, the core promoter sequence comprises or consists of a sequence selected from SEQ ID NO: 159 or 160 or a functional variant thereof.
[0050] By “synthetic” promoter is meant a promoter that has been modified from a wild-type promoter. That is, the synP does not exist naturally.
[0051] By “upstream regulatory domain” is meant a DNA sequence found upstream of the core promoter and adjacent gene, and typically comprising one or more binding sites for a TF(s). In eukaryotic cells, including yeast, multiple interactions between TFs bound to the upstream regulatory domain of a promoter influence transcription. Within the upstream regulatory domain, upstream activator sequences (UAS) and / or upstream repressive sequences (URS) that bind activating or repressing TFs respectively may be found. In yeast, binding of a TF to a UAS is necessary for transcription of a gene positioned downstream. UAS in yeast are found adjacently upstream to a TATA box and TSS, and work in either orientation and at variable distance with respect to the TATA box and TSS.
[0052] A TF may be defined as a protein comprising at least one DNA-binding domain, and which binds to a specific sequence of DNA and that regulates expression of at least one gene. The TF may comprise an activator or repressor domain that mediates a variety of protein-protein interactions to alter the specific level of gene expression, i.e. to increase or decrease gene transcription, respectively.
[0053] By “repressor transcription factor” or “transcriptional repressor” is meant a protein, which binds (i.e. hybridises) to a DNA sequence and mediates protein-protein interactions that decrease or inhibit transcription (via its repressor domain). A repressor TF may impede transcription through a number of mechanisms, including but not limited to: binding to co-repressors to form a repressor complex that blocks RNA polymerase binding or binding to RNA polymerase and disrupting the initiation of transcription.
[0054] By “repressor transcription factor binding sequence” or “transcription repressor binding sequence” is meant the sequence of DNA that a repressor TF binds to exert its repressive effects on transcription. Such sequences are found upstream of the proteincoding gene, within an upstream regulatory domain, and may also be referred to as a repressor site or upstream repressor sequence (URS). Such terms may be used interchangeably.
[0055] Accordingly, a transcription repressor may bind to a repressor TFBSs. In one embodiment, the transcription repressor may be selected from MIG1, MIG2, MIG3 R0X1, RFX1, MATALPHA2, RDR1 and / or Y0X1. In one embodiment, the sequence of MIG1, MIG2, MIG3 R0X1, RFX1, MATALPHA2, RDR1 and Y0X1 can be found by reference to http: / / www.yeastract.com / consensuslist.php. In another embodiment, the repressor TFBS may be selected from one of the following sequences in Table 1. In another embodiment, the repressor TFBS may be selected from one of the sequences shown in Table 3 or 4.
[0056] Table 1: Repressor transcription factor binding sequences
[0057]
[0058] By an “activator transcription factor” or “transcriptional activator” is meant a protein which binds to a DNA sequence and mediates protein-protein interactions that increase transcription (via an activator domain). An activator TF may promote transcription through a number of mechanisms, including but not limited to: promoting a conformational change in the pre-initiation complex that promotes transcription, or recruiting RNA polymerase and / or co-activators.
[0059] By “activator transcription factor binding sequence” or “transcription activator binding sequence” is meant the sequence of DNA that an activator TF binds (i.e. hybridises) to exert a promotive effect on transcription. Such sequences are found upstream of the protein-coding gene, within an upstream regulatory domain, and may also be referred to as activator sites or upstream activator sequences (UAS) or an activator-binding site. Such terms may be used interchangeably. An activating sequence is usually within a few hundred bp of a promoter. For example, most activating sequences are within about 200 to 400 bp of the promoter that is enhanced. Accordingly, the activating TFBS may bind to a transcriptional activator, wherein the transcriptional activator is selected from Cat8, Gal4, a TATA box-binding protein, a GATA box-binding protein, Sip4 and Adri. In one embodiment, the sequence of Cat8, Gal4, TATA box, GATA box, Sip4 and Adri can be found by reference to http: / / www.yeastract.com / consensuslist.php. In one embodiment, the sequence of Cat8, Gal4, TATA box, GATA box, Sip4 and Adri can be found in Table 6. In a preferred embodiment, the activating TFBS binds to one or more of Cat8, Gal4, TATA box and GATA box.
[0060] By “binds” means that the activator or repressor TFBS binds to an activator or repressor TF respectively. The level of binding can be measured using any technique known in the art, for example, by using a DNA electrophoretic mobility shift assay (EMSA). In one embodiment, the activator or repressor TFBS binds only the activator or repressor TF (and no other TF).
[0061] In one embodiment, the activating TF may bind to a carbon source responsive element in the promoter. In this case, the carbon source responsive element acts as a TFBS.
[0062] In another embodiment, the activator TFBS may be selected from one of the following sequences in Table 2.
[0063] Table 2: Activator transcription factor binding sequences
[0064]
[0065]
[0066] Preferably, the TFBS is specific for the TF. Techniques to experimentally identify and validate a sequence as a TFBS are known to those skilled in the art, and have been described in the literature (for example, Geertz and Maerkl 2010). Alternatively, the Matinspector programme (K. Cartharius, et al., 2005) or the Transfac programme (https: / / qenexplain.com / transfac / ) can be used.
[0067] An appropriate method to identify a location of a TFBS is a protein binding microarray. Micro-arrayed double- stranded DNA oligos are subject to a binding reaction by adding TFs to these microarrays. If a TFBS is present within the microarray, the DNA-binding domain of TF will bind to it. Following a wash step, bound TFs can be immunodetected by a TFs-specific, fluorescent antibody. The sequence within the well will be known, and so the sequence of the TFBS is determined.
[0068] Systematic evolution of ligands by exponential enrichment (SELEX) may also be used to identify a TFBS (Geertz and Maerkl 2010). This technique is based on incubating a purified TF with a pool of random DNA oligos. In vitro selection of TFs binding sites consists of several rounds of binding and amplification of captured double- stranded DNA targets. Captured targets are either analysed individually by cloning and sequencing or in bulk by deep sequencing approaches.
[0069] Immunoprecipitation based approaches to identify and validate a TF consist of crosslinking a TF to genomic loci in vivo (using ChlP-seq) or in vitro (DNA immunoprecipitation, DIP), followed by shearing of DNA and precipitation with a TF specific antibody. Enriched DNA fragments are analysed after reversal of cross-linking by microarray or deep sequencing. A multiplicity of experimental and computational methods may be employed to characterise a TF as an activator or repressor. A widely accepted standard for measuring the transcriptional activity of a TF is a reporter assay. In this assay, in vitro cells are transfected with TF reporter plasmids bearing regulatory sequences containing TFBSs of the target TF. The interaction between the TF and TFBS influences the transcription efficiency of a reporter gene. Thus, measuring the changes in reporter gene expression allows for direct measurement and comparison of the gene transcription activity of the targeted TF.
[0070] Suitable reporter genes are well-known to those in the art and may include:
[0071] • A fluorescent reporter, for example green fluorescent protein (GFP)
[0072] • A luminescent reporter gene, for example luciferase.
[0073] • A chemical reporter gene, for example p-galactosidase, alkaline phosphatase or beta-lactamase and the like.
[0074] The skilled person would expect the expression (and therefore the output measurement) of the reporter gene to decrease if the TF was a repressor (compared to the level of expression without the TF). The skilled person would expect the expression (and therefore the output measurement) of the reporter gene to increase if the TF was an activator (compared to the level of expression without the TF).
[0075] The transcriptional activity of a TF may also be characterised using a Northern blot. After overexpressing the putative TF within a cell, it is possible to measure the change in transcription of a target gene, by preparing a northern blot. Northern Blotting measures the mRNA of a particular gene by separating mRNA of the cell according to size (using gel electrophoresis) and subsequently quantifying mRNA of interest using a hybridization probe with a base sequence complementary to all, or a part, of the sequence of the target mRNA.
[0076] By “replaced” is meant the deletion (i.e. removal) of at least one repressor TFBS and the insertion of at least one activator TFBS at the same loci. The activator TFBS that is inserted at the same loci as the deleted repressor site may or may not be of an equivalent sequence length to the repressor TFBS (that has been removed), and so the length of the upstream regulatory region may be altered. In one embodiment, at least two or more repressor sites are deleted. In another embodiment, where a sequence contains a plurality of repressor sites the entire region comprising (all of) the repressor sites may be deleted including any surrounding nucleotides that did not contain activator sites. Accordingly, deletion of at least one repressor TFBS may comprise the deletion of one or more 5’ and / or 3’ nucleotides around the repressor TFBS, but only where these 5’ and / or 3’ nucleotides do not contain an activator TFBS. By eliminating unnecessary sequences (i.e. those that do not bind an activator TF, and so do not act as activating TFBS) we created shorter but stronger promoter variants (e.g. Psyn2, Psyn3 and Psyn8 are only around 700bp long). This is advantageous since shorter promoter sequences offer greater precision, efficiency, predictability and are easier to use in genetic engineering applications.
[0077] Accordingly, in one embodiment, the length of the synP is between 600 and 900bp, more preferably between 650 and 850bp, even more preferably around 700bp. In one embodiment, the length of the synP is or is around 690, 708 or 711bp.
[0078] Methods using routine experimentation to identify and validate a TFBS have been discussed above and have been described in the literature. Using the techniques described above and those known in the art, identifying and validating a sequence as a TFBS, and characterising it as a repressor or activator site, the skilled person would be able to determine the sequence of the repressor that needs to be deleted, and optionally the activator binding sequence which needs to be isolated for insertion into a yeast cell, as per the methods described below.
[0079] The skilled person will appreciate that replacing a (repressor) TFBS with an alternative (activator) TFBS may result inadvertently in the creation of new TFBSs, i.e. at the insertion sites between the insert and the endogenous DNA. This can be checked by performing an additional in silico analysis using, for example, Matinspector or Transfac, which as described elsewhere, can be used to predict putative TFBSs. These new sequences may bind to repressor TFs and negatively affect the performance of the synP (by decreasing expression of the target gene). Accordingly, in one embodiment, the synP may also not comprise any additional or new (i.e. not present in the endogenous promoter sequence) TFBSs as a consequence of replacement of the at least one repressor TFBS with at least one activator TFBS. In an alternative or an additional embodiment, the at least one repressor TFBS is deleted. That is, and as shown in Table 2, one or more repressor sites are deleted without being replaced by an activator site. This may be in addition to the replacement of one or more repressor site(s) with an activator site(s), as described above.
[0080] In one embodiment, the deletion of the at least one repressor TFBS and / or the replacement of the at least one repressor TFBS by at least one activator TFBS is achieved using targeted genome modification, such as, for example, Zinc Finger Nuclease (ZNFs), TALENs or CRISPR / Cas9 (or Cpf1).
[0081] The synP may be characterised by having multiple modifications (i.e. multiple replacements or multiple replacements and / or deletions). For example, the synP may comprise at least two, at least three, at least four, at least five, at least six, at least seven, at least eight, at least nine, at least ten, at least eleven, at least twelve, at least thirteen, at least fourteen or at least fifteen replacements. Where the multiple modifications are replacements, the multiple replacements may be to the same activator sites or to one or more different activator sites. For example at least two, at least three, at least four, at least five, at least six, at least seven, at least eight, at least nine, at least ten, at least eleven, at least twelve, at least thirteen, at least fourteen or at least fifteen repressor sites may be replaced with at least two, at least three, at least four, at least five, at least six, at least seven, at least eight, at least nine, at least ten, at least eleven, at least twelve, at least thirteen, at least fourteen or at least fifteen same or different activator sites.
[0082] In one embodiment, the promoter is selected from the promoter of: Acetyl-CoA hydrolase (P CW), ACC Synthase 1 (P CSV), ADY2 (PADY2), ATO3 (PATOS), Fructose-Bisphosphatase 1 (PFBPV), isocitrate lyase (P / c ), JEN1 PJENI), Malate Dehydrogenase 2 (PMDH2), MLS1 (P / WLSV), and Proline permease (PPUT4). Preferably, the promoter is selected from PADH2 or PGAP. In one embodiment, the promoter is PADH2- In another embodiment, the promoter is PGAP.
[0083] The promoter may be an inducible promoter. In one embodiment, the inducible promoter is an ethanol-inducible promoter. In another embodiment, the synthetic promoter is PADH2- Preferably, the sequence of the synthetic PADH2 promoter comprises or consists of SEQ ID NO: 2 or a functional fragment or variant thereof. In another embodiment, the synthetic promoter is PGAP. Preferably, the sequence of the synthetic promoters based on PGAP promoter comprises or consists of SEQ ID NO: 1 or 3 or a functional fragment or variant thereof.
[0084] In another embodiment, the inducible promoter is a methanol-inducible promoter. This may be particularly advantageous where the yeast is methylotrophic. As such, the addition of methanol to the culture medium will activate the promoter and induce expression of the target protein.
[0085] Heterologous gene expression systems driven by strong methanol-inducible promoters have a number of benefits in heterologous yeast protein production. Such benefits include (i) a cheap synthetic salt-based media for growing the yeast and (ii) strong and tightly regulated promoters induced by methanol.
[0086] As used in any aspect of the invention described throughout a “variant” or a “functional variant” has at least 25%, 26%, 27%, 28%, 29%, 30%, 31 %, 32%, 33%, 34%, 35%, 36%, 37%, 38%, 39%, 40%, 41 %, 42%, 43%, 44%, 45%, 46%, 47%, 48%, 49%, 50%, 51 %,
[0087] 52%, 53%, 54%, 55%, 56%, 57%, 58%, 59%, 60%, 61 %, 62%, 63%, 64%, 65%, 66%,
[0088] 67%, 68%, 69%, 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81 %,
[0089] 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91 %, 92%, 93%, 94%, 95%, 96%,
[0090] 97%, 98%, or at least 99% overall sequence identity to the non-variant nucleic acid or amino acid sequence. A “functional” variant is still able to act as a promoter - that is to cause expression of the target gene.
[0091] In another aspect of the invention, there is provided a nucleic acid construct comprising the synP of the invention, and a nucleic acid sequencing encoding a target protein. The synP is preferably operably linked to the nucleic acid sequence encoding a target protein.
[0092] The term "operably linked" as used throughout refers to a functional linkage between the regulatory sequence (for example, a promoter) and the gene of interest, such that the regulatory sequence is able to initiate transcription of the gene of interest (for example, by initiating transcription).
[0093] The nature of the target protein and the relevant genetic material encoding the target material is not significant to the practice of this invention. The target protein may typically be a secreted protein; however, the production of non-secreted proteins is also within the scope of this invention. Examples of suitable exogenous proteins include food proteins, enzymes, antibodies, cytokines, insulin, interferon, growth hormones, vaccines, plasma proteins and hormones. In one specific example, the exogenous protein is a dairy protein, egg protein or collagen. The terms “target protein” and “exogenous protein” or “heterologous protein” can be used interchangeably herein.
[0094] A further approach to enhance the strength of the synPs of this invention, is to additionally overexpress at least one TF activator. Accordingly, in a further embodiment, the nucleic acid construct further comprises at least one nucleic acid sequence encoding a TF. The TF may be selected from Cat8, Gal4, TATA box-binding protein, GATA box binding protein, Adri and Sip4. In one embodiment, the TF may be selected from SEQ ID NO: 161 to 166 or a functional variant thereof, shown in Table 6. In one embodiment, the TATA box-binding protein may be SPLT15. By “functional” variant is meant herein a variant that is able to act as a TF. Preferably, the at least one TF is also operably linked to the synP of the invention (in addition to the target gene). Alternatively, the at least one TF is operably linked to a different regulatory sequence, such as a constitutive or strong promoter.
[0095] Table 6: Transcription Factor Sequences
[0096] It may also be useful to optimize the processing of the target proteins, as this can have a significant impact on final protein yield in yeast expression systems. Protein processing includes all of the steps from translation to post-translational modifications and folding, and any step that is suboptimal can lead to reduced yield or poor protein quality. One approach to optimize protein processing is to engineer the yeast expression system to further include helper factors, such as chaperones, which can facilitate protein folding and reduce aggregation. Such sequences may be constitutively or inducibly expressed in the yeast host cell, using vectors, markers, and the like as known in the art. Preferably the sequences, including transcriptional regulatory elements sufficient for the desired pattern of expression, are stably integrated in the yeast genome through a targeted methodology.
[0097] Chaperones are a class of proteins that play a crucial role in protein folding and stability. They are able to bind to unfolded or partially folded proteins and guide them towards the correct conformation, preventing misfolding and aggregation. There are several types of chaperones that are used in yeast expression systems, including Hsp70, Hsp60. These chaperones have been shown to improve protein yield and quality in a variety of expression systems.
[0098] Accordingly, in a further embodiment, the nucleic acid construct may further express at least one helper factor. The helper factor may be a chaperone. In one embodiment, the chaperone is Hsp60 or Hsp70, Kar2 (BiP). The helper factor may be a co-chaperone. In one embodiment, the co-chaperone is a nucleotide exchange factor, such as Lhs1 , Sill , Erol , or a J-protein, for example Jem1 or Scj1 or Sec63. The helper factor is preferably also operably linked to the synP of the invention. Alternatively, the at least one TF is operably linked to a different regulatory sequence, such as a constitutive or strong promoter.
[0099] In a further aspect of the invention, there is provided a cell obtained or obtainable by the above method. The cell is preferably a yeast cell.
[0100] In another aspect of the invention there is provided the use of the recombinant cell, preferably a yeast cell of the invention for exogenous protein production.
[0101] The yeast may be of any suitable species. Brewer’s yeast (Saccharomyces sp.) is commonly used for exogenous protein production; the yeast may be S. cerevisiae. In preferred embodiments of the invention, however, the yeast may be Pichia sp., preferably P. pastoris. (Note that some authors have reassigned P. pastoris to Komagataella sp., or as K. phaffii, K. pastoris, and K. pseudopastoris). In one embodiment, the strain of P. pastoris may be selected from CBS2612 or CBS7435. Other suitable yeasts may include Schizosaccharomyces (for example, S. pombe), Kluyveromyces lactis, and Yarrowia lipolytica.
[0102] The yeast may preferably be a methylotrophic yeast.
[0103] The recombinant yeast of the invention may alternatively be produced using genome editing. Genome editing of yeast includes the use of the CRISPR / Cas system to selectively and rapidly engineer yeast strains, without the need for marker-based selection. Optionally, several gRNAs can be combined to target multiple sites simultaneously. CRISPR / Cas can be particularly effective in yeast, since off-target effects are unlikely due to the relatively small genome size. Tools for designing gRNA in yeast include CRISPy, CRISPy-web, CRISPR-ERA, Yeastriction and CHOPCHOP v2. The most commonly used Cas enzyme for CRISPR / Cas in yeast is Cas9, and particularly Cas9 from Streptococcus pyogenes, which may be yeast codon-optimised.
[0104] The use of synPs to regulate expression of enhanced green fluorescent protein (eGFP) and Target A resulted in notable increases in protein production. Bioreactor cultivations of the top strains that produced the highest protein levels using synPs control demonstrated substantial improvements in protein expression compared to the endogenous promoter. These findings show the enormous potential synPs have in industrial applications.
[0105] Furthermore, the approach used to design synPs in this study is also unique. Unlike the random mutation approach, which is unpredictable and time-consuming, or the computational modelling and simulation approach, which requires significant computational work, synPs were designed herein by combining knowledge on the functions of individual TFs with an analysis of the presence or absence of their binding sites in individual promoters. Thus, unique new combinations of TFBS with increasing numbers were designed. Additionally, this study incorporated multiple mutations simultaneously, unlike previous studies where single or a few modifications were made to the same constructs. This method allows for the creation of synPs customized to meet specific requirements and achieve high levels of gene expression. These findings highlight the potential of synPs in improving the scalability and efficiency of protein production, leading to more sustainable and cost-effective bioprocessing. The invention is now described in the following non-limiting examples:
[0106] EXAMPLE 1: In-silico analysis
[0107] In this study, Matinspector was performed to analyse the TFBSs. Table 3 provides the predicted putative repressors for PADH2 and PGAP. Using this information, synPs were designed to target these TFBS repressors. This approach enables the control of gene expression, which could facilitate the precise regulation of synPs and the production of desired gene products.
[0108] Table 3: Predicted putative TFBS repressors on PADH2 and PGAP by Matinspector. The protein names for TFs that bind to the predicted binding motifs were based on their homologs in S. cerevisiae. Sequence ID for TFBSs shown on the left hand side.
[0109] EXAMPLE 2: Design of synPs through iterative modification of TFBSs
[0110] The main objective of this study was to design strong synPs for industrial applications. In contrast to prior studies where only single or a few modifications were made to the same constructs, we designed synPs by incorporating multiple mutations at once. In order to eliminate internal TFBSs, either the entire binding site was removed or replaced with activators, with a particular emphasis on Cat8, Gal4, TATA box-binding proteins and GATA box-binding proteins. A summary of the newly constructed promoters Psyn3, Psyn2 and Psyn8 is given in Table 4. To ensure that no additional TFBSs were created, we performed an additional in silico analysis by Matinspector for predicting putative TFBSs after designing the promoter variants, as changes in the promoter sequence can lead to the emergence of new TFBSs.
[0111] Table 4: List of repressor factors on PADH2 and their replacing activator factors on synthetic promoters. Sequence Identifier numbers can be found to the left of each sequence (Sequences in PADH2 and PGAP) or to the right of each sequence (sequence in Psyn).
[0112] A schematic analysis was carried out on the promoters of Psyn3, Psyn2, and Psyn8, as well as their corresponding endogenous promoters, to highlight the differences between synPs and their corresponding native sequences. The analysis involved using coloured boxes to represent the different TFs, with boxes of the same colour belonging to the same family. Binding predictions on the positive and negative strands were indicated by boxes on the upper and lower sides, respectively (Fig. 1).
[0113] The effect of modifying binding sites on promoter activity was dependent on their specific location within the promoter. This suggests that by arranging these TFBSs in particular combinations, it is possible to generate synPs with distinct properties. This methodology was then used to develop additional synPs from the same native sequence. To accomplish this, gBIocks were acquired from IDT (integration DNA technologies company) or Twist Bioscience.
[0114] EXAMPLE 3: The expression of enhanced green fluorescent protein (eGFP) was controlled by synthetic promoters.
[0115] In this study, we assessed the promoter activity of three synthetic promoters (Psyn3, Psyn2 and Psyns) that were developed using the methodology described earlier. We compared their performance to that of the endogenous PGAP and P ADH2 in P. pastoris. To do this, we conducted an experiment involving the screening of individual single clones for each promoter in a small-scale cultivation setup. The expression levels of these promoters were normalized using eGFP expression under the control of the PGAP as reference, with PGAP set as 1 (Fig. 2).
[0116] In figure 2, we observed varying average eGFP expression levels under the control of different promoters: PADH2 exhibited an expression level of 0.4, Psyn2 demonstrated a significantly higher level at 3, while Psyn3 and Psyn8 had expression levels of 2.2 and 4, respectively. Notably, when compared to PGAP as the control, the t-test results revealed significant differences with the following p-values: Psyn2 had a p-value of 2.4E-08, Psyn3 showed a p-value of 1.8E-08, Psyn8 had the lowest p-value of 9.7E-11 , and PADH2 had a p-value of 2.0E-06. These findings suggest that Psyn3, Psyn2, and Psyns are promising promoters for the expression of heterologous proteins in P. pastoris. Psyn3, Psyn2 and Psyns had higher promoter activity than the endogenous promoter.
[0117] EXAMPLE 4: Evaluation of protein expression driven by synthetic promoters
[0118] After identifying these promising synPs, the next step was to investigate their ability to drive the expression of specific proteins. As eGFP is an intracellular protein, we next focused on an extracellular protein - referred to herein as target A, we evaluated the expression levels of target A under the control of various synPs (Psyn2, Psyn3 and Psyns). To quantify the productivity of target A, we used an automated microcapillary electrophoresis method (LabChip GX).
[0119] Figure 3 shows that the synPs are stronger than the current state-of-the-art endogenous promoters, including PGAP, PADH2, PAOXI. The T arget-A titre was found to be 119 mg / L with clone #39.3.09 under the control of Psyn2, and 106 mg / L with clone #39.4.14 under control of Psyn8. The highest expression level was observed for clone #38.4.06 under the control of Psyn3, which reached a titer of 130 mg / L. This titer is 4.1 , 4.2, and 15.5-fold higher than the highest titer observed for the clones under the control of PGAP, PADH2 and PAOXI, respectively.
[0120] Based on these results, the top three strains that produced Target A under the control of Psyn2, Psyn3 and Psyn8 were selected. Subsequently, these strains were evaluated in bioreactor cultivations.
[0121] EXAMPLE 5: Bioreactor cultivations
[0122] Figure 4 shows the titres of Target A (mg / L) produced using different synPs in a bioreactor cultivation setting. The promoters used for this study are PG P, Psyn2, Psyn3, and Psyn8, with four different sampling time points, which are batch end (BE), feed end (FE), and two intermediate-time points during the feed (F1 and F2). The aim of the study was to evaluate the performance of synP-driven expression strains in industrially relevant protein production conditions compared to the positive control PGAP. The results show that synPs have a significant impact on protein expression, with Psyn3 exhibiting the highest expression levels for T arget A. The use of synPs led to a significant increase in protein yield compared to the positive control PGAP. At the time point BE (batch end), the titres of Target A produced using Psyn2, Psyn3, and Psyns were 25, 24 and 24 mg / L, respectively, while the titer was only 11 mg / L under the regulation of PGAP. Moreover, at subsequent time points (F1 , F2, and FE), the synPs consistently outperformed PGAP.
[0123] At time point F1 , the titres of Target A produced using synPs were between 22-30 times higher than the titre produced using PG P. The titres of Target A produced using Psyn2, Psyn3, and Psyn8 were 177, 231 and 152 mg / L, respectively, while the PGAP driven reference produced only 8 mg / L.
[0124] At time point F2, the fold change in titre between the synPs and PGAP increased up to 35- fold. The titres of Target A produced using Psyn2, Psyn3, and Psyn8 were 261 , 334 and 237 mg / L, respectively, while PGAP produced only 10 mg / L.
[0125] Finally, at the FE time point, the fold change in titre between the synPs and PGAP increased up to 4.1 -fold. The highest titre of Target A produced was 866 mg / L under control of Psyn3.
[0126] This data suggests that the synPs, particularly Psyn3, are much stronger than the naturally occurring PG P for producing Target A.
[0127] EXAMPLE 6: Gene copy numbers
[0128] To identify the correlation between gene copy number and protein expression driven by the same promoter (PGAP) in different colonies, the gene copy numbers were checked. Colonies with varying levels of protein expression driven by PGAP were selected, and their gene copy numbers were determined using qPCR.
[0129] There was a correlation between the change of gene copy numbers and gene expression. In numerous instances, increasing the gene copy number leads to increases in the product titres (Fig.5). EXAMPLE 7: Material and Methods
[0130] Strains
[0131] The fully sequenced Pichia pastoris (Komagataella phaffiP) CBS2612 strain (= NRRL Y- 7556; Type strain of Komagataella phaffiP) (Kurtzmann) was used as basis for all protein producing yeast strains.
[0132] For all cloning tasks, the fully sequenced Escherichia coli strain DH10B (genotype: F- mcrA A(mrr-hsdRMS-mcrBC) <t>80dlacZAM15 AlacX74 endA1 recA1 deoR A(ara,leu)7697 araD139 galll galK nupG rpsL A-) (Durfee et al.) was used.
[0133] In silico TFBS analysis
[0134] The study examined promoter variants for putative TFBSs and analysed them using Matinspector (Cartharius, Freeh et al. 2005), library version 11.1 , with the fungi matrix as a reference. The default parameters were used, and matches with similarity scores greater than 0.75 were considered as possible matches.
[0135] Cloning
[0136] For all model protein cloning applications, the newly developed GoldenP / CS (Pichia Cloning System) was used (Prielhofer, Barrero et al. 2017), which is a bespoke implementation of a Golden-Gate / Gateway modular cloning system for Pichia pastoris. The modularity of this technique provides the advantage of reliable high-quality assembly of different promoter, gene of interest and terminator combinations in (compared to conventional cloning techniques) relatively short time.
[0137] P. pastoris Transformation
[0138] To transform P. pastoris, each plasmid construct (2 pg per construct) was linearized at the genome integration locus and then enzyme inactivated. Electroporation (BioRad Gene Pulser, 2000 V, 25 mF, and 200 V) was used for P. pastoris transformation. Regeneration of the transformed cells was achieved by incubating them in YPD medium at 30°C for about 3 h, followed by plating on YPD plates containing an antibiotic selection marker. After 48 h at 28°C, some transformants were randomly chosen and re-streaked on selective YPD plates, which were then incubated for an additional 48 h at 28°C or 72 h at 25°C. Small scale screening cultivation (24DWP)
[0139] The screening cultivation of the protein producing P. pastoris strains was performed according to a standardized and optimized high-throughput small-scale cultivation protocol in 24-deep well plates (DWP) (Fig.6). Cultures are grown in a volume of 2 mL, the basic carbon supply is based on the enzymatic release system EnPump 200 (Enpresso GmbH, Germany), where glucose is released from a polymer solution. Therefore, the basic media is prepared in a 2x-concentration and for the cultivation diluted to the normal concentration with a 2x-enzyme / polymer-solution mix.
[0140] Bioreactor cultivation All bioreactor cultivation experiments are performed in DASGIP® Parallel Bioreactor systems (Eppendorf, Germany; #76DG04MBBB) in 1.0 L bioreactor vessels (#76SR07000DLS). Process parameters like temperature (T), pH and dissolved oxygen (DO) are controlled using the DASware® control software (#76DGCS4). Experiments were carried out according to the following standard protocol:
[0141] Table 5: bioreactor process settings of fed-batch cultivation processes batch volume (V) 300 mL inoculation density 0.5-2.0 OD600 (corresponds approx.
[0142] 0.125-0.5 g / L YDM) temperature (T) 25°C pH 5.5 (batch phase) 5.0 (feed phase) dissolved oxygen (DO) 20% automated control by DO-cascade (stirrer, in-gas flow, in-gas x(O2)) fed batch 72h exponential feed of 50% glucose with set to 0.05 / h stirrer speed (rpm) 200 - 1200 rpm ingas flow rate (F) 6 - 50 sL / h ingas oxygen concentration x(O2) 21 - 100%
[0143] NH3 (12.5 or 25%) by pump B 5 - 15 mL / h off-gas off-gas flow, xco2, X02 feed media balance values
[0144] Quantification of eGFP by flow cytometry
[0145] Flow cytometer (CytoFLEX S, Beckman Coulter) was used for the determination of eGFP. For flow cytometer analysis, cells were diluted in 96-well-plates to an ODeoo of 0.4 in PBS. Relative eGFP expression levels were calculated compared to eGFP expression under endogenous versions (PGAP or PAD ).
[0146] Microcapillary electrophoresis
[0147] The protein product quality and quantity are analysed from small scale screening or bioreactor cultivation supernatant samples, which are frozen to -20°C until analysis. Right before analysis they are thawed at room temperature (RT).
[0148] The LabChip GXII Touch HT Protein Characterization System (CLS138160; PerkinElmer) was applied with ProteinExpress Chips and Reagents (760499 + CLS960008; PerkinElmer) according to the standard analysis protocol.
[0149] Determination of gene copy numbers by quantitative PCR
[0150] Quantitaive PCR was used for relative gene copy number determination. Genomic DNAs were isolated according to the protocol from the Wizard® Genomic DNA purification kit (Promega Corp., USA). Relative gene copy number was determined from relative concentrations of the gene of interest (GOI) normalized to the internal control, the housekeeping gene ACT1.
[0151] REFERENCES
[0152] Geertz M, Maerkl SJ. Experimental strategies for studying transcription factor-DNA binding specificities. Brief Funct Genomics. 2010 Dec;9(5-6):362-73.
[0153] K. Cartharius, K. Freeh, K. Grote, B. Klocke, M. Haltmeier, A. Klingenhoff, M. Frisch, M. Bayerlein, T. Werner, Matinspector and beyond: promoter analysis based on transcription factor binding sites, Bioinformatics, Volume 21 , Issue 13, July 2005, Pages 2933-2942, SEQUENCE LISTING
[0154] SEQ ID NO: 1 (PSyn2)
[0155] TCACAGTTTTCCAGTAGTTCCGATCAAATTACCATCGAAATGGTCCCATAAACGGA
[0156] CATTTGACATCCGTTCCTGAATTATAGTCTTCCACCGTGGATCATGGTGTTCCTTTT
[0157] TTTCCCAAAGAATATCAGCATCCCTTAACTACGTTAGGTCAGGACATGCGTCCGCC
[0158] CTCTTTTTCTTTCATCGGCACATTTCAGCCTCACATGCGACTATTATCGATCAATGA
[0159] AATCCATCAAGATTGAAATCTTAAAATTGCCCCTTTCACTTGACAGGATCCTTTTTT
[0160] GTAGAAATGTCTTGGTGTCCTCGTCCAATCAGGTAGCCATCTCTGAAATATCTGTT
[0161] CCGTTCGTCCGAGGAACCAGAAACGTCTCTTCCCTTCTCTCTCCTTCCACCGCCC
[0162] GTTACCGTCCCTAGGAAATTTTACTCTGCTGGAGAGCTTCTTCTACGGCCCCCTTG
[0163] CAGCAATGCTCTTCCCAGCATTACGTTCCGTTCGTCCGAACGGAGGTCGTGTACC
[0164] CGACCTAGCAGCCCAGGGATGGAAAAGTCCCGGCCGTCGCTGGCAATAATAGCG
[0165] GGCGGACGCATGTCATGATTCCGTTCGTCCGACCAGAATCGAATATAAAAGGCGA
[0166] ACACCTTTCCCAATTTTGGTTTCTCCTGACCCAAAGACTTTACTATATAAAACAAAA
[0167] GTCCCTATTTCAATCAATTGAACAACTAT
[0168] SEQ ID NO: 2 (PSyn3)
[0169] >
[0170] AATCAGTATCACCGATTTTCCGTTCGTCCGATCAACGGTCCCTCATCCTTTCCGTT
[0171] CGTCCGAATGGCAGTTAGCATTGGTGCACTGACTGACTGCCCAACCTTAAACCCA
[0172] AATTTCTTAGAAGGGGCCCATCTAGTTAGCGAGGGGTGAAAAATTCCTCCATCGG
[0173] AGATGTATTGACCGTAAGTTGCTGCTTAAAAAAAATCAGTTCAGATAGCGAGACTT
[0174] TTGACTGATTAGATTAGAGTGCCTGTTCCATTCGATTGCAATTCTCACCCCTTCTG
[0175] CCCAGTCCTGCCAATTGCCCATGAATCTGCTAATTTCGTTGATTTTCCGTTCGTCC
[0176] GACCACAAATTGTCCAATCTCGTTTTCCATTTGGGAGAATCTGCATGTCGACTACA
[0177] TAAAGCGACCGGTGTCCGAAAAGATCTGTGTAGTTTTCAACATTTTGTGTTCCGTT
[0178] CGTCCGAACGGGGGTGAGCGCTCTCCGGGGTGCGAATTCGTGCCCAATTCCTTT
[0179] CACCCTTTCCGTTCGTCCGAGTCAACCCGCATCTGGTGCGAATATAGCGCACCCC
[0180] CAATGATCATCCCCAACGATTGCATTGGGGATCCACCCCTCCCCAATCTCTAATAT
[0181] TCACAATTCACCTCACTATAAATACCCCTGTCCTGCTCCCAAATTCTTTTTTCCTTC
[0182] TTCCATCAGCTACTAGCTTTTATCTTATTTACTTTACGAA
[0183] SEQ ID NO: 3 (Psyns)
[0184] TCACAGTTTTCCAGTAGTTCCGATCAAATTACCATCGAAATGGTCCCATAAACGGA
[0185] CATTTGACATCCGTTCCTGAATTATAGTCTTCCACCGTGGATCATGGTGTTCCTTTT TTTCCCAAAGAATATCAGCATCCCTTAACTACGTTAGGTCAGGACATGCGTCCGCC
[0186] CTCCTTTCATCCGGTTTCAGCCTCACATGCGACTATTATCGATCTCCATTCGTCCG
[0187] GCGGAGGACAGTCCTCCGATTGCCCCTTTCACTTGACAGGATCCTTTTTTGTAGA
[0188] AATGTCTTGGTGTCCTCGTCCAATCAGGTAGCCATCTCTGAAATATCTGTTCCGTT
[0189] CGTCCGAGGAATCCATTGATCCGACTTCTCTCTCCTTCCACCGCCCGTTACCGTC
[0190] CCTAGGAAATCCCCAACGATTGCATTGGGGATCTTCTACGGCCCCCTTGCAGCAA
[0191] TGCTCTTCCCAGCATTACGTTCCGTTCGTCCGAACGGAGGTCGTGTACCCGACCT
[0192] AGCAGCCCAGGGATGGAAAAGTCCCGGCCGTCGCTGGCAATAATAGCGGGCGG
[0193] ACGCATGTCATGATTCCGTTCGTCCGACCAGAATCGAATATAAAAGGCGAACACC
[0194] TTTCCCAATTTTGGTTTCTCCTGACCCAAAGACTTTACTATATAAAACAAAAGTCCC
[0195] TATTTCAATCAATTGAACAACTAT
[0196] SEQ ID NO: 4 PADH2
[0197] TGATTATGTAAGAAGAGGGGGGTGATTCGGCCGGCTATCGAACTCTAACAACTAG
[0198] GGGGGTGAACAATGCCCAGCAGTCCTCCCCACTCTTTGACAAATCAGTATCACCG
[0199] ATTAACACCCCAAATCTTATTCTCAACGGTCCCTCATCCTTGCACCCCTCTTTGGA
[0200] CAAATGGCAGTTAGCATTGGTGCACTGACTGACTGCCCAACCTTAAACCCAAATTT
[0201] CTTAGAAGGGGCCCATCTAGTTAGCGAGGGGTGAAAAATTCCTCCATCGGAGATG
[0202] TATTGACCGTAAGTTGCTGCTTAAAAAAAATCAGTTCAGATAGCGAGACTTTTTTGA
[0203] TTTCGCAACGGGAGTGCCTGTTCCATTCGATTGCAATTCTCACCCCTTCTGCCCGT
[0204] CCTGCCAATTGCCCATGAATCTGCTAATTTCGTTGATTCCCACCCCCCTTTCCAAC
[0205] TCCACAAATTGTCCAATCTCGTTTTCCATTTGGGAGAATCTGCATGTCGACTACAT
[0206] AAAGCGACCGGTGTCCGAAAAGATCTGTGTAGTTTTCAACATTTTGTGCTCCCCCC
[0207] GCTGTTTGAAAACGGGGGTGAGCGCTCTCCGGGGTGCGAATTCGTGCCCAATTC
[0208] CTTTCACCCTGCCTATTGTAGACGTCAACCCGCATCTGGTGCGAATATAGCGCAC
[0209] CCCCAATGATCACACCAACAATTGGTCCACCCCTCCCCAATCTCTAATATTCACAA
[0210] TTCACCTCACTATAAATACCCCTGTCCTGCTCCCAAATTCTTTTTTCCTTCTTCCAT
[0211] CAGCTACTAGCTTTTATCTTATTTACTTTACGAAA
[0212] SEQ ID NO: 5 PAOXI
[0213] AGACGAAAGGTTGAATGAAACCTTTTTGCCATCCGACATCCACAGGTCCATTCTCA
[0214] CACATAAGTGCCAAACGCAACAGGAGGGGATACACTAGCAGCAGACCGTTGCAA
[0215] ACGCAGGACCTCCACTCCTCTTCTCCTCAACACCCACTTTTGCCATCGAAAAACCA
[0216] GCCCAGTTATTGGGCTTGATTGGAGCTCGCTCATTCCAATTCCTTCTACTAGGCTA
[0217] CTAACACCGTGACTTTATTAGCCTGTCTATCCTGGCCCCCCTGGCGAGGTTCATG TTTGTTTATTTCCGAATGCAACAAGCTCCGCATTACACCCGAACATCACTCCAGAT
[0218] GAGGGCTTTCTGAGTGTGGGGTCAAATAGTTTCATGTTCCCCAAATGGCCCAAAA
[0219] CTGACAGTTTAAACGCTGTCTTGGAACCTAATATGACAAAAGCGTGATCTCATCCA
[0220] AGATGAACTAAGTTTGGTTCGTTGAAATGCTAACGGCCAGTTGGTCAAAAAGAAAT
[0221] TCCAAAAGTCGGCATACCGTTTGTCTTGTTTGGTATTGATTGACGAATGCTCAAAA
[0222] ATAATCTCATTAATGCTTAGCGCAGTCTCTCTATCGCTTCTGAACCCCGGTGCACC
[0223] TGTGCCGAAACGCAAATGGGGAAACACCCGCTTTTTGGATGATTATGCATTGTCT
[0224] CCACATTGTATGCTTCCAAGATTCTGGTGGGAATACTGCTGATAGCCTAACGTTCA
[0225] TGATCAAAATTTAACTGTTCTAACCCCTACTTGACAGCAATATATAAACAGAAGGAA
[0226] GCTGCCCTGTCTTAAACCTTTTTTTTTATCATCATTATTAGCTTACTTTCATAATTGC
[0227] GACTGGTTCCAATTGACAAGCTTTTGATTTTAACGACTTTTAACGACAACTTGAGAA
[0228] GATCAAAAAACAACTAATTATTCGAAAC
[0229] SEQ ID NO: 170 PGAP
[0230] TCACAGTTTTCCAGTAGTTCCGATCAAATTACCATCGAAATGGTCCCATAAACGGA
[0231] CATTTGACATCCGTTCCTGAATTATAGTCTTCCACCGTGGATCATGGTGTTCCTTTT
[0232] TTTCCCAAAGAATATCAGCATCCCTTAACTACGTTAGGTCAGTGATGACAATGGAC
[0233] CAAATTGTTGCAAGGTTTTTCTTTTTCTTTCATCGGCACATTTCAGCCTCACATGCG
[0234] ACTATTATCGATCAATGAAATCCATCAAGATTGAAATCTTAAAATTGCCCCTTTCAC
[0235] TTGACAGGATCCTTTTTTGTAGAAATGTCTTGGTGTCCTCGTCCAATCAGGTAGCC
[0236] ATCTCTGAAATATCTGGCTCCGTTGCAACTCCGAACGACCTGCTGGCAACGTAAA
[0237] ATTCTCCGGGGTAAAACTTAAATGTGGAGTAATGGAACCAGAAACGTCTCTTCCCT
[0238] TCTCTCTCCTTCCACCGCCCGTTACCGTCCCTAGGAAATTTTACTCTGCTGGAGAG
[0239] CTTCTTCTACGGCCCCCTTGCAGCAATGCTCTTCCCAGCATTACGTTGCGGGTAA
[0240] AACGGAGGTCGTGTACCCGACCTAGCAGCCCAGGGATGGAAAAGTCCCGGCCGT
[0241] CGCTGGCAATAATAGCGGGCGGACGCATGTCATGAGATTATTGGAAACCACCAGA
[0242] ATCGAATATAAAAGGCGAACACCTTTCCCAATTTTGGTTTCTCCTGACCCAAAGAC
[0243] TTTAAATTTAATTTATTTGTCCCTATTTCAATCAATTGAACAACTAT
[0244] SEQ ID NO: 159 Core Psyn3
[0245] TGATCATCCCCAACGATTGCATTGGGGATCCACCCCTCCCCAATCTCTAATATTCA
[0246] CAATTCACCTCACTATAAATACCCCTGTCCTGCTCCCAAATTCTTTTTTCCTTCTTC
[0247] CATCAGCTACTAGCTTTTATCTTATTTACTTTACGAA
[0248] SEQ ID NO: 160 Core Psyn2 and Psyn8 ATAATAGCGGGCGGACGCATGTCATGATTCCGTTCGTCCGACCAGAATCGAATAT
[0249] AAAAGGCGAACACCTTTCCCAATTTTGGTTTCTCCTGACCCAAAGACTTTACTATAT
[0250] AAAACAAAAGTCCCTATTTCAATCAATTGAACAACTAT
[0251] SEQ ID NO: 171. MIG1 TFBS (shown in Table 4)
[0252] T gtcaaagagtggggagga
[0253] SEQ ID NO: 172. RFX1 TFBS (shown in Table 4) ttgatttcgcaacgg
[0254] SEQ ID NO: 173. MIG1 TFBS (shown in Table 4) gagttggaaaggggggtggg
[0255] SEQ ID NO: 174. Matalpha-2 TFBS (shown in Table 4) gcctaTTGTagac
[0256] SEQ ID NO: 175. ROX1 TFBS (shown in Table 3) tccaTTGTcatca
[0257] SEQ ID NO: 176. Matalpha-2 TFBS (shown in Table 3) gtccaTTGTcatc
[0258] SEQ ID NO: 177. Matalpha-2 TFBS ((shown in Table 3) ccaaaTTGTtgca
[0259] SEQ ID NO: 178. RFX1 TFBS (shown in Table 3) aaaaccttGCAAcaa
[0260] SEQ ID NO: 179. RFX1 TFBS (shown in Table 3) gctccgttGCAActc
[0261] SEQ ID NO: 180. RFX1 TFBS (shown in Table 3) tcggagttGCAAcgg
[0262] SEQ ID NO: 181. RFX1 TFBS (shown in Table 3) acctgctgGCAAcgt
[0263] SEQ ID NO: 182. MIG1 TFBS (shown in Table 3) taaaattctccggggtaaa
[0264] SEQ ID NO: 183. MIG1 TFBS (shown in Table 3) aacttaaatgtGGAGtaat
[0265] SEQ ID NO: 184. RDR1 TFBS (shown in Table 3) ttgcgggtaaa
[0266] SEQ ID NO: 185. RFX1 TFBS (shown in Table 3) gattattggaaacca
[0267] SEQ ID NO: 186. YOX1 TFBS (shown in Table 3) aaataaattaaattt SEQ ID NO: 187. YOX1 TFBS (shown in Table 3) aatttAATTtatttg
[0268] SEQ ID NO: 188. RFX1 TFBS (shown in Table 4) GCTCCGTTGCAAC
[0269] SEQ ID NO: 189. RFX1 TFBS (shown in Table 4)
[0270] GATTATTGGAAACC
[0271] SEQ ID NO: 190. Yox1 TFBS (shown in Table 4) AATTTAATTTATTT
[0272] SEQ ID NO: 191. (shown in Table 4)
[0273] AATGAAATCCATCAAGATTGAAATCTTAAA
[0274] SEQ ID NO: 192. (shown in Table 4)
[0275] TTTACTCTGCTGGAGAGCT
[0276] SEQ ID NO: 193. MIG-1 TFBS (shown in Table 4)
[0277] GCGGGTAA
[0278] SEQ ID NO: 194. MIG-1 TFBS (shown in Table 4)
[0279] GATTATTGGAAACC
[0280] SEQ ID NO: 195. Modified sequence in Psyn3. (shown in Table 4) TTCCGTTCGTCCGA
[0281] SEQ ID NO: 196. Modified sequence in Psyn3. (shown in Table 4) GACTGATTAGATTA
[0282] SEQ ID NO: 197. Modified sequence in Psyn3. (shown in Table 4) TTCCGTTCGTCCGAC
[0283] SEQ ID NO: 198. Modified sequence in Psyn3. (shown in Table 4) TTCCGTTCGTCCG
[0284] SEQ ID NO: 199. Modified sequence in Psyn3. (shown in Table 4) TTTCCGTTCGTCCGA
[0285] SEQ ID NO: 200. Modified sequence in Psyn3. (shown in Table 4) TCCCCAACGATTGCATTGGGGA
[0286] SEQ ID NO: 201. Modified sequence in Psyn3. (shown in Table 4) GACATGCGTCCGCCC
[0287] SEQ ID NO: 202. Modified sequence in Psyn2. (shown in Table 4) TTCCGTTCG
[0288] SEQ ID NO: 203. Modified sequence in Psyn2. (shown in Table 4) TTCCGTTGTCCG SEQ ID NO: 204. Modified sequence in Psyn2. (shown in Table 4) TTCCGTTCGTCCG
[0289] SEQ ID NO: 205. Modified sequence in Psyn2. (shown in Table 4) CTATATAAAACAAAA
[0290] SEQ ID NO: 206. Modified sequence in Psyn8. (shown in Table 4) GACATGCGTCCGCCCTCCTTTCATCCGG
[0291] SEQ ID NO: 207. Modified sequence in Psyn8. (shown in Table 4) TCCATTCGTCCGGCGGAGGACAGTCCTCCG
[0292] SEQ ID NO: 208. Modified sequence in Psyn8. (shown in Table 4) TTCCGTTCGTCCGAGGAATCCATTGATCCGA
[0293] SEQ ID NO: 209. Modified sequence in Psyn8. (shown in Table 4) CCCCAACGATTGCATTGGGGA
[0294] SEQ ID NO: 210. Modified sequence in Psyn8. (shown in Table 4) CCGTTCGTCCG
[0295] SEQ ID NO: 211. Modified sequence in Psyn8. (shown in Table 4) TTCCGTTCGTCCG
Claims
38CLAIMS:
1. A synthetic promoter comprising a core promoter domain and an upstream regulatory domain, wherein the regulatory domain comprises one or more repressor TFBSs and wherein at least one repressor TFBS has been replaced by an activator TFBS.
2. The synP of claim 1 , wherein the repressor TFBS binds a transcriptional repressor, wherein the transcriptional repressor is selected from MIG1, MIG2, MIG3, R0X1, RFX1, Matalpha 2, RDR1 and Y0X1.
3. The synP of claim 1 or 2, wherein the at least one repressor TFBS is selected from SEQ ID NO: 6 to 49 or 167 to 169 or 171 to 194 or a variant thereof.
4. The synP of any preceding claim, wherein the activator TFBS binds to a transcriptional activator, wherein the transcriptional activator is selected from Cat8, Gal4, TATA box-binding protein, GATA box-binding protein, Adri and Sip4.
5. The synP of claim 4, wherein the at least one activator TFBS is selected from SEQ ID NO: 50 to 157 or a variant thereof.
6. The synP of any preceding claim, wherein additionally at least one repressor TFBS is deleted.
7. The synP of any preceding claim, wherein at least two, preferably at least four, more preferably between six and ten repressor TFBSs are replaced with at least two, preferably at least four, more preferably between six and ten activator TFBSs.
8. The synP of any preceding claim, wherein the promoter is a methanol-inducible promoter.
9. The synP of any preceding claim, wherein the promoter is selected from PADH2, alcohol oxidase 1 (P^oxi), or is PGAP.
10. The synP of any preceding claim, wherein the promoter does not comprise any additional TFBSs as a consequence of replacement of the at least one repressor TFBS being replaced by at least one activator TFBS.3911. The synP of any preceding claim, wherein the promoter comprises a sequence selected from SEQ ID NO: 1 to 3, or a variant thereof, wherein the variant has at least 60% overall sequence identity to one of SEQ ID NO: 1 to 3.
12. A nucleic acid construct comprising the synP of any of claims 1 to 11 and a nucleic acid sequence encoding a target protein, wherein the synP is operably linked to the nucleic acid sequence encoding a target protein.
13. The nucleic acid construct of claim 12, wherein the target protein is an exogenous protein.
14. The nucleic acid construct of any of claims 12 to 13, wherein the construct further comprises at least one nucleic acid sequence encoding a TF, wherein preferably the TF is selected from Cat8, Gal4, TATA box-binding protein, GATA box-binding protein, Adri and Sip4.
15. The nucleic acid construct of any of claims 12 to 14, wherein the construct further comprises at least one helper factor.
16. A cell comprising the nucleic acid construct of any of claims 12 to 15.
17. An organism comprising the nucleic acid construct of any of claims 12 to 15, wherein preferably, the organism is a yeast.
18. The organism of claim 17, wherein the yeast is a methylotrophic yeast, preferably selected from the Pichia genus, most preferably selected from Pichia pastoris, Hansenula polymorpha, Candida boidinii or Pichia methanolica.
19. A method of heterologous protein production, the method comprising culturing the yeast of claim 17 or 18 under conditions that allow expression of the nucleic acid encoding a target protein.
20. The method of claim 19, wherein the method further comprises screening the yeast of claim 17 or 18 and selecting yeast with the greatest expression of the target protein.4021. The method of claim 20, wherein the method further comprises extracting the exogenous protein.
22. A method for producing a synP, the method comprising a. identifying a plurality of TFBSs within the upstream regulatory sequence of a promoter; b. determining whether said TFBS is a repressor TFBS or an activator TFBS; and c. replacing at least one TFBS that binds a transcriptional repressor with at least one TFS that binds a transcriptional activator.
23. The method of claim 22, wherein the repressor TFBS binds a transcriptional repressor, wherein the transcriptional repressor is selected from MIG1, MIG2, MIG3, R0X1, RFX1, Matalpha 2, RDR1 and Y0X1.
24. The method of claim 23, wherein the at least one repressor TFBS is selected from SEQ ID NO: 6 to 49 or 167 to 169 or 171 to 194 or a variant thereof.
25. The method of claim 22, wherein the activator TFBS binds to a transcriptional activator, wherein the transcriptional activator is selected from Cat8, Gal4, TATA box-binding protein, GATA box-binding protein, Adri and Sip4.
26. The method of claim 25, wherein the at least one activator TFBS is selected from SEQ ID NO: 50 to 157 or a variant thereof.
27. The method of any of claims 22 to 26, wherein the promoter is a methanol- regulated promoter, preferably selected from PADH2, alcohol oxidase 1 (PAOXI), or PGAP.
28. A synthetic promoter obtained or obtainable by the method of any of claims 22 to
Citation Information
Patent Citations
Promoter variants
WO2017021541A1
Design of alcohol dehydrogenase 2 (ADH2) promoter variants by promoter engineering
WO2020068019A2