Non-viral transcription activation domains and related methods and uses

Plant-derived, modified transcription activation domains address the limitations of viral-based systems by ensuring stable and efficient gene expression across species, enhancing the use of artificial expression systems in diverse applications.

JP7763756B2Active Publication Date: 2025-11-04TEKNOLOGIAN TUTKIMUSKESKUS VTT OY
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2022527828
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2019-11-19
Filing Date
2020-11-18
Publication Date
2025-11-04
Estimated Expiration
2040-11-18

AI Technical Summary

Technical Problem

Current gene expression systems, particularly those using viral or cancer-associated transcription activation domains, face challenges in achieving stable and predictable gene expression across various species and are limited by regulatory and customer acceptance issues, necessitating the development of non-viral alternatives that maintain functionality and stability.

Method used

Development of plant-derived non-viral transcription activation domains, specifically modified from transcription factors in edible plant species like Arabidopsis thaliana and Brassica napus, which are engineered to retain high activity and stability in diverse eukaryotic organisms, replacing existing viral-based domains.

Benefits of technology

These modified plant-derived domains provide reliable and stable gene expression across multiple species, enabling efficient production of target compounds and broadening the applicability of artificial expression systems in industries like food and pharmaceuticals.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007763756000010
    Figure 0007763756000010
  • Figure 0007763756000011
    Figure 0007763756000011
  • Figure 0007763756000012
    Figure 0007763756000012
Patent Text Reader

Abstract

The present invention relates to the fields of life sciences, genetics, and regulation of gene expression. Specifically, the present invention relates to a non-viral transcription activation domain for a eukaryotic host. The present invention also relates to a polypeptide or artificial transcription factor comprising the transcription activation domain of the present invention. Furthermore, the present invention relates to a polynucleotide, an expression cassette, an expression system, and / or a eukaryotic host. Furthermore, the present invention relates to a method for producing a desired protein product in a eukaryotic host of the present invention, or a method for preparing the non-viral transcription activation domain of the present invention or a polynucleotide encoding this non-viral transcription activation domain. Still further, the present invention relates to the use of the transcription activation domain, polypeptide, artificial transcription factor, polynucleotide, expression cassette, expression system, or eukaryotic host of the present invention for metabolic engineering and / or production of a desired protein product.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to the fields of life sciences, genetics, and regulation of gene expression. Specifically, the present invention relates to a non-viral transcription activation domain for a eukaryotic host. The present invention also relates to a polypeptide or artificial transcription factor comprising the transcription activation domain of the present invention. Furthermore, the present invention relates to a polynucleotide, an expression cassette, an expression system, and / or a eukaryotic host. Furthermore, the present invention relates to a method for producing a desired protein product in a eukaryotic host of the present invention, or a method for preparing the non-viral transcription activation domain of the present invention or a polynucleotide encoding this non-viral transcription activation domain. Still further, the present invention relates to the use of the transcription activation domain, polypeptide, artificial transcription factor, polynucleotide, expression cassette, expression system, or eukaryotic host of the present invention for metabolic engineering and / or production of a desired protein product. [Background technology]

[0002] Controlled and predictable gene expression is very difficult to achieve even in established hosts, especially with regard to stable expression under various culture conditions or growth stages. In addition, for many potentially interesting industrial hosts, the spectrum of tools and / or methods for achieving heterologous gene expression or controlling endogenous gene expression is very limited (or even non-existent). In many cases, this prevents the use of such interesting industrial hosts (which are often very promising hosts) in industrial applications.

[0003] Transcription factors play a major role in the regulation of gene expression. Transcription factors usually contain at least two domains: the DNA-binding domain (DBD), which binds to the promoter of the target gene, and the activation domain (AD), which is involved in activating transcription by interacting with the transcription machinery. There have been numerous previous attempts to introduce new transcription factors or their domains suitable for tight control of gene expression in genetically engineered biological systems.

[0004] In artificial gene expression systems, the use of viral transcriptional activation domains (e.g., VP16 or VP64) is currently the most common solution for high-level expression. Other components derived from viral proteins or proteins associated with cancer development may also be used in efficient artificial expression systems. For example, Chavez et al. described an improved transcriptional regulator obtained through the rational design of a tripartite activator, VP64-p65-Rta (VPR), fused to a nuclease-deficient Cas9. In this transcriptional regulator, VP64 is derived from human herpes simplex virus, p65 is a human protein associated with multiple types of cancer, and Rta is derived from Epstein-Barr virus (Chavez et al., 2015, Nat Methods, 12(4), 326-328).

[0005] The use of plant (Arabidopsis thaliana) native transcription factors for the regulation of gene expression in yeast has been described by Naseri G et al. (2017, ACS Synthetic Biology, 6, 1742-1756). In their study, Naseri G et al. focused on the use of fusion transcription factors containing additional activation domains in their structure, specifically the virus-based VP16 activation domain, the GAL4 activation domain from Saccharomyces cerevisiae (budding yeast), and the EDLL motif from Arabidopsis thaliana.

[0006] Although expression systems containing viral or cancer-associated transcription activation domains are highly efficient, their use in many biotechnology applications, particularly in food or pharmaceutical manufacturing, can be problematic due to current regulations and customer and / or patient acceptance. Therefore, novel transcription activation domains are needed to replace currently used viral-based domains. Furthermore, new types of activation domains must provide a sufficient level of functionality in gene expression systems to achieve similar or better production of target compounds. Additionally, efficient non-viral transcription activation domains and gene expression systems based on them must provide reliable and stable gene expression in several different species and genera of production organisms. [Prior art documents] [Non-patent literature]

[0007] [Non-Patent Document 1] Chavez A et al., 2015, Nat Methods, 12(4), 326-328 [Non-patent document 2] Naseri G et al., 2017, ACS Synthetic Biology, 6, 1742-1756 Summary of the Invention [Problem to be solved by the invention]

[0008] The objective of the present invention, i.e., a novel, efficient transcription activation domain and related tools and methods, can be used to functionally replace virus-based activation domains without compromising the performance of gene expression systems. Expression systems containing the novel transcription activation domain provide reliable and stable expression, a wide range of expression levels, and can be used in several different species and genera. This is achieved by utilizing a transcription activation domain derived from a transcription factor found in plant species, such as edible plant species. [Means for solving the problem]

[0009] Indeed, it has now surprisingly been found that engineering of plant-derived transcription activation domains results in novel activation domains that are highly active and, importantly, retain high activity in diverse eukaryotic organisms. These novel activation domains are plant-derived, non-viral transcription activation domains that can be used to regulate gene expression in expression systems, e.g., eukaryotic organisms.

[0010] The present invention can be used to overcome deficiencies in the prior art, including but not limited to the use of viral DNA elements in artificial expression systems, which lack efficient activation domains and expression systems that are functional across a wide variety of species and are also acceptable or suitable for all technical fields and industries that utilize gene expression, including food and pharmaceuticals.

[0011] Surprisingly, the present inventors have been able to develop specific activation domains derived from plant species that can be used in a variety of overexpression systems, for example, to replace currently used activation domains. Indeed, the activation domains of the present invention can be incorporated into expression systems based on artificial (synthetic) transcription factors without compromising the function of the system, and all previously demonstrated advantages of artificial transcription systems can be retained or improved.

[0012] The present invention allows for the efficient transfer of engineered metabolic pathways and their testing simultaneously in several potential production hosts for, for example, functional evaluation. Furthermore, the present invention provides tools for orthogonal gene expression, thus providing benefits to the scientific community studying, for example, eukaryotes.

[0013] Furthermore, the present invention makes it possible to broaden the use of artificial expression systems in applications where the use of potentially problematic (viral) DNA elements is undesirable.

[0014] The present invention relates to a non-viral transcription activation domain for a eukaryotic host or an artificial expression system in a eukaryotic host, wherein the transcription activation domain is derived from a plant or a plant transcription factor, for example derived from or found in an edible plant.

[0015] The present invention also relates to a polypeptide comprising a non-viral transcription activation domain for a eukaryotic host or an artificial expression system in a eukaryotic host, wherein the transcription activation domain is derived from a plant or a plant transcription factor.

[0016] The present invention also relates to an artificial transcription factor comprising a non-viral transcription activation domain, a DNA binding domain and a nuclear localization signal for a eukaryotic host or an artificial expression system in a eukaryotic host, wherein the transcription activation domain is derived from a plant or a plant transcription factor.

[0017] Furthermore, the present invention relates to polynucleotides encoding the transcription activation domains, polypeptides or artificial transcription factors of the present invention.

[0018] Furthermore, the present invention relates to an expression cassette or expression system comprising a polynucleotide encoding a transcription activation domain, a polypeptide or an artificial transcription factor of the present invention.

[0019] Still further, the present invention relates to a eukaryotic host comprising a transcriptional activation domain, a polypeptide, an artificial transcription factor, a polynucleotide, an expression cassette or an expression system of the present invention.

[0020] Still further, the present invention relates to a method for producing a desired protein product in a eukaryotic host, comprising culturing a host of the invention under suitable culture conditions.

[0021] Still further, the present invention relates to the use of a transcriptional activation domain, polypeptide, artificial transcription factor, polynucleotide, expression cassette, expression system or eukaryotic host of the invention for metabolic engineering and / or production of a desired protein product.

[0022] Furthermore, the present invention relates to a method for preparing a non-viral transcription activation domain of the present invention or a polynucleotide encoding the non-viral transcription activation domain, the method comprising the steps of obtaining a transcription activation domain polypeptide derived from a plant transcription factor or obtaining a polynucleotide encoding said transcription activation domain polypeptide derived from a plant transcription factor, and modifying the obtained transcription activation domain polypeptide or polynucleotide.

[0023] Other objects, details and advantages of the present invention will become apparent from the following drawings, detailed description and examples. [Brief explanation of the drawings]

[0024] [Figure 1]Figure 1 shows an example of a scheme for an expression system including a transcription activation domain of the present invention. Indeed, Figure 1 shows an example of a scheme for an expression system for testing the production of a transcription activation domain and a protein product of interest in a eukaryote or microorganism, as exemplified by the evaluation of the production of, for example, a red fluorescent protein, mCherry, in Trichoderma reesei (Examples 1 and 8). Accordingly, this scheme also shows an expression system used for heterologous protein production in, for example, Trichoderma reesei (Example 3), Myceliophthora thermophila (Example 5), and / or Aspergillus oryzae (Example 7). This expression system is constructed as a single DNA molecule and includes or consists of a target gene expression cassette, an sTF expression cassette, a selectable marker (SM) expression cassette, and genomic integration DNA regions (flanking regions), exemplified herein by genomic DNA sequences from Trichoderma reesei located upstream of the egl1 gene (EGL1-5') and downstream of the egl1 gene (EGL1-3'). In one embodiment, Figure 1 shows a synthetic expression system for use in filamentous fungi, such as T. reesei, M. thermophila, and / or Aspergillus oryzae. The target gene expression cassette can include, or includes, multiple sTF-specific binding sites, exemplified herein by eight sTF-specific binding sites (8BS) located upstream of a core promoter, exemplified herein by An_201cp (SEQ ID NO: 23) from Aspergillus niger (Aspergillus niger). The eight sTF-specific binding sites and the core promoter form a synthetic promoter that potently activates transcription of the target gene in the presence of synthetic transcription factors (sTFs).The target gene may be any DNA sequence encoding a protein product of interest, as exemplified herein by a DNA sequence encoding mCherry (see Examples 1, 2, and 8), a DNA sequence encoding a xylanase enzyme (see Examples 3 and 5), or a DNA sequence encoding bovine β-lactoglobulin B (see Example 7). Transcription of the target gene can be terminated on a transcription termination sequence, as exemplified herein by the Trichoderma reesei pdc1 terminator (Tr_PDC1t). The synthetic transcription factor (sTF) expression cassette contains a core promoter (Tr_hfb2cp; SEQ ID NO: 25), an sTF coding sequence, and a terminator. The core promoter provides constitutive low expression of sTF. sTF binds to the sTF-dependent synthetic promoter in the target gene expression cassette to promote transcription of the target gene. sTFs contain or consist of a DNA-binding domain (BDB) consisting of a bacterial DNA-binding protein and a nuclear localization signal, such as the SV40 NLS, and a transcription activation domain (AD). The AD can be any transcription activation domain of plant origin, exemplified herein by 10 examples based on or derived from transcription factors found in Arabidopsis thaliana, Brassica napus, and Spinacia oleracea (spinach). A control AD ​​is VP16, derived from herpes simplex virus. Transcription of the sTF gene can be terminated on a transcription termination sequence, exemplified herein by the Trichoderma reesei tef1 terminator (Tr_TEF1t). A selectable marker (SM) expression cassette is any expression cassette that allows the production of a specific protein in a host organism and provides the host organism with the means to grow under selective conditions, such as in the presence of an antibiotic compound or the absence of an essential metabolite.SM cassettes are exemplified herein by expression cassettes that enable expression of the pyr4 gene (encoding the enzyme orotidine 5'-phosphate decarboxylase) in Trichoderma reesei strains (Examples 1, 3, and 8), or by expression cassettes that enable expression of the hygR gene (encoding hygromycin-B 4-O-kinase) in Myceliophthora thermophila (Example 5), or by expression cassettes that enable expression of the pyrG gene (encoding the enzyme orotidine 5'-phosphate decarboxylase) in Aspergillus oryzae strains (Example 7). [Figure 2]Figure 2 shows an example of a schematic diagram of an expression system containing the transcription activation domain of the present invention. Indeed, Figure 2 shows an example of a schematic diagram of an expression system for testing the production of a transcription activation domain and a protein product of interest in eukaryotes or microorganisms, as exemplified by the evaluation of the production of a heterologous protein, such as a bacterial phytase enzyme, in Pichia pastoris (Example 4). This expression system can include, or be constructed as, two separate DNA molecules: the first DNA includes or consists of an sTF expression cassette, a selectable marker (SM) expression cassette, and a genome-integrated DNA region (flanking regions); the second DNA includes or consists of a target gene expression cassette, a selectable marker (SM) expression cassette, and a genome-integrated DNA region (flanking regions). Each cassette integrates into a separate locus in the host genome, and together they form a functional gene expression system. In one embodiment, Figure 2 shows a synthetic expression system used in Pichia pastoris. The sTF expression cassette can comprise (or consist of) a core promoter (An_008cp, SEQ ID NO: 22), an sTF coding sequence, and a terminator. The sTF comprises (or consists of) a DNA-binding domain (BDB) consisting of a bacterial DNA-binding protein, exemplified herein by the Bm3R1 repressor (Example 4), and a nuclear localization signal, such as the SV40 NLS, and a transcription activation domain (AD). The AD can be any transcription activation domain of plant origin, exemplified herein by five examples based on or derived from transcription factors found in Arabidopsis thaliana, Brassica napus, and Spinacia oleracea, selected based on the analysis performed in Example 1 (Figure 4). The control AD ​​can be, for example, VP16 from herpes simplex virus. Transcription of the sTF gene can be terminated on a transcription termination sequence, exemplified herein by the Trichoderma reesei tef1 terminator (Tr_TEF1t).An SM cassette is exemplified herein by an expression cassette enabling expression of the kanR gene (encoding an aminoglycoside phosphotransferase enzyme) in a Pichia pastoris strain (Example 4). Genomic integrated DNA regions (flanking regions) are exemplified herein by genomic DNA sequences from Pichia pastoris located upstream of the URA3 gene (URA3-5') and downstream of the URA3 gene (URA3-3'). A target gene expression cassette can contain or include multiple sTF-specific binding sites, exemplified herein by eight Bm3R1-specific binding sites (8BS) located upstream of a core promoter, exemplified herein by An_201cp (SEQ ID NO: 23) from Aspergillus niger. A target gene may be any DNA sequence encoding a protein product of interest, exemplified herein by a DNA sequence encoding a phytase enzyme (see Example 4). Transcription of the target gene can be terminated on a transcription termination sequence, exemplified herein by the Saccharomyces cerevisiae ADH1 terminator (Sc_ADH1t). The SM cassette is exemplified herein by an expression cassette (Example 4) that enables expression of the Pichia pastoris URA3 gene (encoding the orotidine 5'-phosphate decarboxylase enzyme) in Pichia pastoris. Genomic integrated DNA regions (flanking regions) are exemplified herein by genomic DNA sequences from Pichia pastoris located upstream of the AOX2 gene (AOX2-5') and downstream of the AOX2 gene (AOX2-3'). [Figure 3]Figure 3 shows an example of a scheme for an expression system comprising the transcription activation domain of the present invention. Indeed, Figure 3 shows an example of a scheme for an expression system for testing the production of a transcription activation domain and a protein product of interest in eukaryotes or microorganisms, as exemplified by the evaluation of the production of, for example, a red fluorescent protein, mCherry, in, for example, CHO cells (Chinese hamster (Cricetulus griseus)) (Example 6). This expression system is constructed as a single DNA molecule and includes or consists of a target gene expression cassette, an sTF expression cassette, and a selectable marker (SM) expression cassette. More specifically, Figure 3 shows a synthetic expression system for use in CHO cells. The target gene expression cassette can include or contains multiple sTF-specific binding sites, exemplified herein by eight sTF-specific binding sites (8BS) located upstream of a core promoter (CP1), exemplified herein by either Mm_Atp5Bcp (SEQ ID NO: 26), Mm_Eef2cp (SEQ ID NO: 27), or Mm_Rpl4cp (SEQ ID NO: 28) from Mus musculus. The target gene may be any DNA sequence encoding a protein product of interest, exemplified herein by a DNA sequence encoding mCherry (see Example 6). Transcription of the target gene can be terminated on a transcription termination sequence (term1), exemplified herein by either the SV40 terminator from simian virus 40 or the FTH1 terminator from Mus musculus (Table 1F; sequences highlighted in gray and italics). The sTF expression cassette can include a core promoter (CP2), an sTF coding sequence, and a terminator. CP2 is exemplified herein by either Mm_Atp5Bcp (SEQ ID NO: 26) or Mm_Eef2cp (SEQ ID NO: 27) or Mm_Rpl4cp (SEQ ID NO: 28) of house mouse origin (Example 6).The sTF comprises or consists of a DNA-binding domain (BDB) containing or consisting of a bacterial DNA-binding protein and nuclear localization signal, such as the SV40 NLS, as exemplified herein by the PhlF repressor from Pseudomonas protegens or the McbR repressor from Corynebacterium species (Example 6). The AD can be any transcription activation domain of plant origin, as exemplified herein by two examples based on transcription factors found in Brassica napus and Spinachia oleracea (So-NAC102M - SEQ ID NO: 10 and Bn-TAF1M - SEQ ID NO: 11), selected based on analyses performed in fungal hosts (Examples 3, 4, and 5). The control AD ​​is VP64 (SEQ ID NO: 30) from herpes simplex virus. Transcription of the sTF gene can be terminated on a transcription termination sequence (term 2), exemplified herein by either the SV40 terminator of simian virus 40 origin or the FTH1 terminator of house mouse origin (Table 1F; sequences shown in gray, italics). The SM cassette is exemplified herein by an expression cassette (Example 6) that allows expression of the pac gene (encoding the puromycin N-acetyltransferase enzyme) in CHO cells. [Figure 4]Figure 4 depicts an example of an analysis of the red fluorescent protein mCherry expressed in a Trichoderma reesei strain transformed with the expression system shown in Figure 1. The purpose of this experiment was to evaluate the performance of a plant-based transcription activation domain compared to the virus-based VP16 activation domain (Examples 1 and 2). A set of 11 T. reesei strains, each containing the expression system (egl1 gene replaced by the expression system) with the indicated AD integrated into the genome at the egl1 locus, was grown in YE-glucose medium for 24 hours before analysis. Quantitative analysis was performed by fluorimetric measurement of mycelial suspensions using a Varioskan instrument (Thermo Electron Corporation). The graph shows the fluorescence intensity (mCherry) normalized by the optical density of the mycelial suspension used for fluorimetric analysis. Bars represent the average value from at least three experimental replicates, and error bars represent the standard deviation. Five activation domains (marked with arrows in the graph) were selected for further testing. [Figure 5]Figure 5 depicts an SDS-PAGE analysis (Coomassie-stained gel) of xylanase protein (Xyn) produced by Trichoderma reesei strains using expression systems containing various transcription activation domains (24-well plates, see Example 3). A set of eight T. reesei strains, each containing the expression system with the indicated AD integrated into the genome at the egl1 locus (egl1 gene replaced by the expression system), was grown in 4 mL of YE-glc medium for 3 days before analysis. A 10 μL aliquot of culture supernatant from each culture was loaded onto a gel (4–20% gradient), and proteins were separated in an electric field (PowerPac HC; BioRad). The gel was stained with colloidal Coomassie (PageBlue Protein Staining Solution; Thermo Fisher Scientific) and visualized using an Odyssey CLx Imaging System (LI-COR Biosciences). The xylanase protein (Xyn) is indicated by an arrow. Three strains were selected for bioreactor cultivation: strains with expression systems containing the So-NAC102M (SEQ ID NO: 10) and Bn-TAF1M (SEQ ID NO: 11) activation domains, and a control strain with VP16 AD (SEQ ID NO: 1). [Figure 6]Figure 6 depicts an SDS-PAGE analysis (Coomassie-stained gel) of the xylanase protein (Xyn) produced by Trichoderma reesei strains in a 1 L bioreactor (see Example 3). A set of three T. reesei strains was cultured for 6 days in YE-glucose medium with continuous glucose feeding. Two microliters of culture supernatant from each culture at different time points was loaded onto a gel (4-20% gradient), and proteins were separated in an electric field (PowerPac HC; BioRad). The gel was stained with colloidal Coomassie (PageBlue Protein Staining Solution; Thermo Fisher Scientific) and visualized using an Odyssey CLx Imaging System (LI-COR Biosciences). The xylanase protein (Xyn) is indicated by an arrow. Cultures from the 5th and 6th day were analyzed for specific xylanase activity (Figure 7). [Figure 7] Figure 7 depicts xylanase activity analysis in the culture supernatant of a Trichoderma reesei strain cultivated in a 1 L bioreactor (see Example 3). Culture supernatants from days 5 and 6 (diluted in 50 mM Tris·HCl, pH 8.0) were assayed for xylanase activity using the EnzCheck® Ultra Xylanase Assay Kit (Invitrogen). Activity is expressed in arbitrary units per mL of culture supernatant (AU / mL). The negative control (NC) represents the culture supernatant of a 1 L bioreactor culture (day 6) of a Trichoderma reesei strain that does not produce xylanase. Bars represent the mean values ​​from at least three technical replicates, and error bars represent the standard deviation. [Figure 8]Figure 8 shows an SDS-PAGE analysis (Coomassie-stained gel) of phytase protein (Appa) produced by Pichia pastoris strains using expression systems containing various transcription activation domains (24-well plate, Figure 2, Example 4). A set of five P. pastoris strains was grown in duplicate (duplicate) in 4 mL of BMG medium for 3 days prior to analysis. Each strain contained the indicated AD, an sTF expression cassette integrated into the genome at the ura3 locus (the ura3 gene replaced by the sTF expression cassette), and a target gene cassette integrated into the aox2 locus (the aox2 gene replaced by the target gene expression cassette). A 10 μL aliquot of culture supernatant from each culture was loaded onto a gel (4–20% gradient), and proteins were separated using an electric field (PowerPac HC; BioRad). The gel was stained with colloidal Coomassie (PageBlue Protein Staining Solution; Thermo Fisher Scientific) and visualized using an Odyssey CLx Imaging System (LI-COR Biosciences). Phytase (AppA) is indicated by an arrow. Three strains were selected for bioreactor cultivation: strains with expression systems containing the So-NAC102M (SEQ ID NO: 10) and Bn-TAF1M (SEQ ID NO: 11) activation domains, and a control strain with VP16 AD (SEQ ID NO: 1) (Figure 9). [Figure 9]Figure 9 depicts an SDS-PAGE analysis (Coomassie-stained gel) of the phytase protein (AppA) produced by Pichia pastoris strains in a 1 L bioreactor (see Example 4). A set of three P. pastoris strains was cultured for 6 days in BMG medium with continuous glucose feeding. Two microliters of culture supernatant from each culture at different time points was loaded onto a gel (4–20% gradient), and proteins were separated using an electric field (PowerPac HC; BioRad). The gel was stained with colloidal Coomassie (PageBlue Protein Staining Solution; Thermo Fisher Scientific) and visualized using an Odyssey CLx Imaging System (LI-COR Biosciences). The phytase protein (AppA) is indicated by an arrow. [Figure 10] Figure 10 depicts the phytase (AppA) activity analysis in the culture supernatant of Pichia pastoris strains cultivated in 1 L bioreactors (see Example 4). 1 mL samples of culture supernatant from days 4 and 6 were diluted with 100 mM Na acetate solution (pH 4.7) and processed by gravity gel filtration (PD-10 desalting column; BioRad). Phytase activity was assayed using a Phytase Assay Kit (MyBioSource). Activity is expressed in arbitrary units per mL of culture supernatant (AU / mL). The negative control (NC) represents the culture supernatant of a 1 L bioreactor culture of a Pichia pastoris strain that does not produce phytase. Bars represent the mean values ​​from three technical replicates, and error bars represent the standard deviation. [Figure 11]Figure 11 depicts an SDS-PAGE analysis (Coomassie-stained gel) of the xylanase protein (Xyn) produced by Myceliophthora thermophila strains using expression systems containing three selected transcription activation domains (24-well plates, Figure 1, Example 5). A set of four M. thermophila clones from each transformation was analyzed. Each clone contained the expression system with the indicated AD integrated into the genome in a random manner (one or more integration events at unknown genomic loci). These strains were grown in 4 mL of BMG medium for 3 days before analysis. A 10 μL aliquot of culture supernatant from each culture was loaded onto a gel (4–20% gradient). The gel was stained with colloidal Coomassie (PageBlue Protein Staining Solution; Thermo Fisher Scientific) and visualized using an Odyssey CLx Imaging System (LI-COR Biosciences). The xylanase protein (Xyn) is indicated by an arrow. All cultures were analyzed for specific xylanase activity (Figure 12). [Figure 12] Figure 12 depicts xylanase activity analysis in the culture supernatant of Myceliophthora thermophila strains cultured in 4 mL of BMG medium for 3 days (24-well plate, Figure 11, Example 5). The culture supernatant was diluted with 50 mM Tris·HCl (pH 8.0) and assayed for xylanase activity using the EnzCheck® Ultra Xylanase Assay Kit (Invitrogen). Activity is expressed in arbitrary units per mL of culture supernatant (AU / mL). The negative control (NC) represents the culture supernatant from the parent Myceliophthora thermophila strain cultured in BMG medium. Bars represent the mean values ​​from at least three technical replicates, and error bars represent the standard deviation. [Figure 13]Figure 13 depicts an SDS-PAGE analysis (Coomassie-stained gel) of bovine β-lactoglobulin B protein (LGB) produced by an Aspergillus oryzae strain using an expression system containing the Bn-TAF1M (SEQ ID NO: 11) transcription activation domain (24-well plate culture, expression system scheme shown in Figure 1; details are described in Example 7). A set of four A. oryzae clones was analyzed. The clones contained the expression system integrated into the genome at two selected loci (see Example 7). These strains were grown in 4 mL of BMG medium for up to 4 days before analysis. A 10 μL equivalent of culture supernatant from each culture was loaded onto the gel (4–20% gradient), and commercially available pure bovine β-lactoglobulin B protein was loaded as a positive control. The gel was stained with colloidal Coomassie (PageBlue Protein Staining Solution; Thermo Fisher Scientific) and visualized using an Odyssey CLx Imaging System (LI-COR Biosciences). The β-lactoglobulin B protein (LGB) is indicated by an arrow. [Figure 14]Figure 14 shows an example of a scheme for an expression system comprising a transcription activation domain of the present invention. Indeed, Figure 14 shows an example of a scheme for an expression system for testing the production of a transcription activation domain and a protein product of interest in a eukaryote or microorganism, as exemplified by the evaluation of regulated production of, e.g., a red fluorescent protein, mCherry, in, e.g., Pichia pastoris or Yarrowia lipolytica (Example 8), or the evaluation of constitutive production of, e.g., a red fluorescent protein, mCherry, in, e.g., Yarrowia lipolytica or Cutaneotrichosporon oleaginosus (Example 9). The expression system is constructed as a single DNA molecule and comprises or consists of a target gene expression cassette, an sTF expression cassette, a selectable marker (SM) expression cassette, and a genome-integrated DNA region (flanking region), exemplified herein by genomic DNA sequences from P. pastoris located upstream (5') and downstream (3') of the ADE1 gene, or sequences from Y. lipolytica located upstream (5') and downstream (3') of the ANT1 gene. In one embodiment, Figure 14 shows a synthetic expression system for use in yeast species, such as P. pastoris, Y. lipolytica, and / or C. oleaginosus. The target gene expression cassette can include, or includes, multiple sTF-specific binding sites, exemplified herein by eight sTF-specific binding sites (8BS) located upstream of a core promoter (cp1), exemplified in Example 8 by An_201cp (SEQ ID NO: 23) from Aspergillus niger, or exemplified by Yl_565cp (SEQ ID NO: 32) from Yarrowia lipolytica, or by other core promoters in Example 9. The eight sTF-binding sites and the core promoter form a synthetic promoter that potently activates transcription of the target gene in the presence of synthetic transcription factors (sTFs).The target gene may be any DNA sequence encoding a protein product of interest, exemplified herein by the DNA sequence encoding mCherry (see Examples 8 and 9). Transcription of the target gene can be terminated on a transcription termination sequence, exemplified herein by the Saccharomyces cerevisiae ADH1 terminator (term1). A synthetic transcription factor (sTF) expression cassette contains a core promoter (cp2), exemplified in Example 8 by An_008cp (SEQ ID NO: 22) or Yl_242cp (SEQ ID NO: 33), or other core promoters in Example 9. The expression cassette further contains an sTF coding sequence and a terminator. The core promoter provides constitutive low expression of the sTF. The sTF comprises or consists of a DNA-binding domain (BDB) consisting of a bacterial DNA-binding protein, e.g., Bm3R1 or TetR, and a nuclear localization signal, e.g., SV40 NLS, and a transcription activation domain, exemplified herein by Bn_TAF1M (SEQ ID NO: 11). sTF binds to an sTF-dependent synthetic promoter in a target gene expression cassette to promote transcription of the target gene. In Example 8, where TetR was used as the DBD of sTF, binding occurs in the absence of doxycycline, and the presence of increasing amounts of doxycycline results in inhibition of binding. Transcription of the sTF gene can be terminated on a transcription termination sequence, exemplified herein by the Trichoderma reesei tef1 terminator (term2). A selectable marker (SM) expression cassette is any expression cassette that allows for the production of a specific protein in a host organism and provides the host organism with a means to grow under selective conditions, such as in the presence of an antibiotic compound or the absence of an essential metabolite.SM cassettes are exemplified herein by expression cassettes that allow expression of the kanR gene (encoding an aminoglycoside phosphotransferase enzyme) in a Pichia pastoris strain (Example 8), or expression cassettes that allow expression of the NAT gene (encoding nourseothricin N-acetyltransferase) in Yarrowia lipolytica (Examples 8 and 9) or Cutaneotrichosporon oleaginosus (Example 9). [Figure 15] Figure 15 depicts an example of an analysis of the red fluorescent protein mCherry expressed in a Trichoderma reesei strain transformed with the expression system shown in Figure 1 (the TetR-based sTF version) and in Pichia pastoris and Yarrowia lipolytica strains transformed with the expression systems shown in Figure 14. The purpose of this experiment was to demonstrate the feasibility of using a plant-based transcriptional activation domain (exemplified herein by Bn_TAF1M) in a doxycycline-regulated Tet-OFF-like expression system (Example 8). A set of strains, each containing the genomically integrated expression system, was grown in BMG medium for 24 hours before analysis. Doxycycline-dependent inhibition of reporter gene expression was assessed using BMG medium without doxycycline (DOX-free) and with 1 mg / L or 3 mg / L doxycycline (DOX). Quantitative analysis was performed by fluorimetric measurement of mycelial or cell suspensions using a Varioskan instrument (Thermo Electron Corporation). The graph shows the fluorescence intensity (mCherry) normalized by the optical density of the mycelium / cell suspension used for fluorimetric analysis. Bars represent the average value from three experimental replicates (three individual clones tested for each species), and error bars represent the standard deviation. [Figure 16]Figure 16 depicts an example of an analysis of the red fluorescent protein mCherry expressed in Yarrowia lipolytica and Cutaneotrichosporon oleaginosus strains transformed with the expression system shown in Figure 14. The purpose of this experiment was to demonstrate the use of a plant-based transcription activation domain (exemplified herein by Bn_TAF1M) in an industrially relevant yeast production host (Example 9). A set of strains, each containing the expression system integrated into their genomes, was grown in YPD medium for 24 hours before analysis. Quantitative analysis was performed by fluorimetric measurement of the cell suspension using a Varioskan instrument (Thermo Electron Corporation). The graph shows the fluorescence intensity (mCherry) normalized by the optical density of the cell suspension used for fluorimetric analysis. Bars represent the average value from three experimental replicates, and error bars represent the standard deviation.

[0025] Sequence Listing SEQ ID NO: 1 VP16 SEQ ID NO: 2 At_NAC102 SEQ ID NO: 3 So_NAC102 SEQ ID NO: 4 At_TAF1 SEQ ID NO: 5 So_NAC72 SEQ ID NO: 6 Bn_TAF1 Sequence number 7 At_JUB1 Sequence number 8 So_JUB1 Sequence number 9 Bn_JUB1 Sequence number 10 So_NAC102M Sequence number 11 Bn_TAF1M SEQ ID NO: 12 At_NAC102 (including nuclear localization signal) SEQ ID NO: 13 So_NAC102 (including nuclear localization signal) SEQ ID NO: 14 At_TAF1 (including nuclear localization signal) SEQ ID NO: 15 So_NAC72 (including nuclear localization signal) SEQ ID NO: 16 Bn_TAF1 (including nuclear localization signal) SEQ ID NO: 17 At_JUB1 (including nuclear localization signal) SEQ ID NO: 18 So_JUB1 (including nuclear localization signal) SEQ ID NO: 19 Bn_JUB1 (including nuclear localization signal) SEQ ID NO: 20 So_NAC102M (including nuclear localization signal) SEQ ID NO: 21 Bn_TAF1M (including nuclear localization signal) Sequence number 22 An_008cp Sequence number 23 An_201cp SEQ ID NO: 24 Phytase enzyme, thermostable variant AppA_K24E Sequence number 25: Tr_hfb2cp SEQ ID NO: 26 Mm_Atp5Bcp Sequence number 27 Mm_Eef2cp SEQ ID NO: 28 Mm_Rpl4cp SEQ ID NO: 29 Bovine β-lactoglobulin B protein SEQ ID NO: 30 VP64 SEQ ID NO: 31 Alkaline xylanase, thermostable mutant xynHB_N188A SEQ ID NO: 32 Yl_565cp SEQ ID NO: 33 Yl_242cp SEQ ID NO: 34 Yl_205cp SEQ ID NO: 35 Yl_TEF1cp SEQ ID NO: 36 Yl_137cp SEQ ID NO: 37 Yl_113cp SEQ ID NO: 38 Yl_697cp Sequence number 39 Cc_RAScp Sequence number 40 Cc_MFScp SEQ ID NO: 41 Cc_HSP9cp SEQ ID NO: 42 Cc_GSTcp SEQ ID NO: 43 Cc_AKRcp SEQ ID NO: 44 Cc_FbPcp DETAILED DESCRIPTION OF THE INVENTION

[0026] The transcription factors studied by Naseri G et al. (2017, ACS Synthetic Biology, 6, 1742-1756) were from the NAC family of Arabidopsis thaliana transcription factors. Several of the transcription factors tested, namely JUB1 and ATAF1, were shown to activate transcription in Saccharomyces cerevisiae, even without fusions to other activation domains.

[0027] The NAC (i.e., NAM, ATAF, and CUC) family of transcription factors is a large protein family containing functionally and structurally distinct proteins (Olsen, Ernst, et al., 2015, Trends Plant Sci 10(2):79-87). NAC transcription factors share a high degree of homology in the DNA-binding domain (NAC domain), but often share very little homology in the transcription activation domain.

[0028] The present inventors have now been able to identify transcription activation domains of transcription factors (e.g., NAC family) from, for example, Arabidopsis thaliana, Brassica napus, and Spinachia oleracea. The latter two species are common edible plant species: oilseed rape and spinach, respectively. Although a high degree of sequence identity was present within the NAC domain, large variations in sequence homology were found between corresponding activation domains. For example, the amino acid sequence identity between the TAF1 activation domain from Arabidopsis thaliana and that from Brassica napus was approximately 77%, whereas the amino acid sequence identity between the JUB1 activation domain from Arabidopsis thaliana and that from Spinachia oleracea was only approximately 23%.

[0029] Furthermore, the level of functionality of activation domains in expression systems implemented in various fungal hosts varied greatly. For example, the TAF1 activation domain from Arabidopsis thaliana was highly active in Trichoderma reesei but nearly inactive in Pichia pastoris (Figures 4 and 8).

[0030] In addition, EDLL motifs previously successfully used in S. cerevisiae by Naseri G et al. and in Arabidopsis thaliana by Tiwari, Belachew et al. (2012, The Plant Journal 70(5):855-865) proved completely inactive when tested in Trichoderma reesei (data not shown). Therefore, the present observations point to the unpredictable function of some plant activation domains in diverse host organisms.

[0031] We noticed that some of the plant-derived activation domains we tested, particularly the TAF1 activation domain from Brassica napus (Bn-TAF1 - SEQ ID NO: 6) and the NAC102 activation domain from Spinachia oleracea (So-NAC102 - SEQ ID NO: 3), contained amino acid compositions similar to those of typical acidic activation domains, rich in acidic amino acids (e.g., glutamic acid and / or aspartic acid) and hydrophobic amino acids (e.g., leucine, isoleucine, and / or phenylalanine). However, the native forms of these activation domains also contained several basic amino acids (e.g., lysine, among others) that were hypothesized to limit the activity of the activation domain. We modified the sequences of the two mentioned activation domains by replacing unfavorable amino acids (e.g., lysine) in the structures with amino acids that more closely matched the typical acidic activation domain sequence (e.g., leucine and / or glutamic acid). Surprising results were found with the modified domains.

[0032] Indeed, the inventors of the present disclosure have been able to create effective transcription activation domains modified from natural plant transcription activation domains, resulting in very potent domains that can be successfully used, for example, to replace current viral or other domains in artificial expression systems.

[0033] Indeed, the present invention relates to modified non-viral transcription activation domains, i.e., mutants (variants) of non-viral transcription activation domains. As used herein, "modified domain" or "modified transcription activation domain" refers to any non-naturally occurring domain or non-naturally occurring transcription activation domain, respectively, that contains different material (e.g., different or modified amino acids) compared to the corresponding unmodified (i.e., native or wild-type) domain. By way of example, a modified domain may include a deletion, substitution, disruption, or insertion of one or more amino acids or portions of the domain, or an insertion of one or more modified amino acids, compared to the corresponding (native or wild-type) domain without said modification.

[0034] The modification of the domain may be obtained, for example, by modifying the polynucleotide encoding said domain by any genetic method. Methods for making genetic modifications are generally well known and are described in various practical manuals describing laboratory molecular techniques. Some examples of general procedures and specific embodiments are described in the Examples section. In one specific embodiment of the present invention, the modified non-viral transcription activation domain was obtained by rational mutagenesis (mutagenesis) or random mutagenesis of the polynucleotide encoding said transcription activation domain.

[0035] In one embodiment of the invention, the transcription activation domain comprises one or more alterations (modifications) and / or mutations compared to the corresponding wild-type transcription activation domain (amino acid) sequence. In certain embodiments, the transcription activation domain comprises one or more amino acid modifications or mutations compared to the corresponding wild-type (i.e., naturally occurring) transcription activation domain sequence.

[0036] In one embodiment, the modified transcription activation domain is a transcription activation domain variant that contains an increased acidic and / or hydrophobic amino acid content compared to the native (i.e., unmodified) transcription activation domain. Acidic amino acids include aspartic acid and glutamic acid. Hydrophobic amino acids include alanine, valine, leucine, isoleucine, proline, phenylalanine, cysteine, and methionine. In certain embodiments, the modified transcription activation domain or transcription activation domain variant contains more aspartic acid, glutamic acid, leucine, isoleucine, and / or phenylalanine amino acids compared to the native (i.e., unmodified) transcription activation domain.

[0037] In one embodiment, the transcription activation domain is a recombinant, synthetic, or artificial transcription activation domain. As used herein, a "recombinant activation domain" refers to an activation domain obtained by genetic modification of genetic material; i.e., the domain may be produced by recombinant DNA technology. In one embodiment, a polynucleotide encoding the "recombinant activation domain" contains mutations compared to the corresponding wild-type polynucleotide (e.g., a deletion, substitution, disruption, or insertion of one or more nucleic acids, including the entire gene or a portion thereof, compared to the domain before modification). In one embodiment, a "recombinant activation domain" comprises or is a polypeptide encoded by a polynucleotide that has been cloned into a system that supports expression of the polynucleotide and further translation of the polypeptide. Indeed, a (genetically) modified polynucleotide can encode a variant (mutant) polypeptide. As used herein, a "synthetic domain" refers to a domain produced by linking multiple amino acids via amide bonds. Polypeptide synthesis can be performed by methods including, but not limited to, classical solution-phase and solid-phase techniques. Also, in some embodiments, "synthetic" can be considered synonymous with "recombinant," as defined above. An "artificial domain" refers to a domain that is non-naturally occurring, i.e., not produced by nature or not occurring in nature, or, for example, a wild-type domain when used in a non-natural context.

[0038] The transcription activation domain (e.g., modified transcription activation domain) of the present invention is derived from a plant or plant transcription factor (e.g., an edible plant). As used herein, "derived from a plant or plant transcription factor," i.e., "originating from a plant or plant transcription factor," or "derived from a plant or plant transcription factor," refers to the situation where the transcription activation domain is a protein or polypeptide, typically a transcription factor, present in a plant. Indeed, in one embodiment of the present invention, the amino acid sequence of the plant activation domain or the nucleotide sequence encoding the plant activation domain is modified. In one particular embodiment, the transcription activation domain is derived from an edible plant or plant species, or a food-grade plant or plant species. As used herein, a "food-grade plant" refers to a non-toxic plant that is safe for consumption and of sufficient quality to be used, for example, for purposes of food production, food storage, or food preparation.

[0039] In one embodiment, the transcription activation domain is derived from Spinacia, Brassica, Ocimum, or Arabidopsis, or from Spinacia oleracea, Brassica napus, Ocimum basilicum, or Arabidopsis thaliana, Ocimum basilicum (sweet basil), or Arabidopsis thaliana. The transcription activation domain may be any transcription activation domain of plant origin, and is exemplified herein by ten examples based on or derived from transcription factors found in Arabidopsis thaliana, Brassica napus, and Spinacia oleracea.

[0040] Many view the use of viral activation domains or viral transcription factors as problematic in synthetic expression systems. Thus, there is a significant need for highly functional activation domains derived from acceptable sources (e.g., as judged by the public or industry). The present invention provides non-viral transcription activation domains derived from plants, i.e., transcription activation domains that do not contain any viral components. Such non-viral transcription activation domains can provide the same or improved efficiency as current viral-based transcription activation domains.

[0041] In one embodiment, the transcription activation domain is selected from the group consisting of a transcription activation domain from a plant NAC family transcription factor (e.g., a TAF (e.g., TAF1) transcription activation domain, a JUB (e.g., JUB1) transcription activation domain), or any fragment thereof. A JUB transcription activation domain refers to the transcription activation domain of a JUNGBRUNNEN factor. For example, JUB1 acts, among other effects, as a negative regulator of senescence and a positive regulator of tolerance to heat and salt stress in plants.

[0042] New activation domains can be incorporated into existing synthetic expression systems, particularly into the construction of synthetic transcription factors of the expression system, where they can replace current activation domains without impairing the function of the system. In one embodiment, the transcription activation domains of the invention are used in the construction of artificial transcription factors, or the transcription activation domains are for synthetic expression systems.

[0043] In one embodiment of the invention, the transcription activation domain is functional across multiple species. If the transcription activation domain is for a synthetic expression system, the synthetic expression system is functional across multiple species.

[0044] The activation domain of the present invention can be of any length, preferably less than 500 amino acids in length. In one embodiment, the transcription activation domain is 20 to 300 amino acids, specifically 30 to 250 amino acids, or more specifically 40 to 200 amino acids, for example, 20 to 30 amino acids, 31 to 40 amino acids, 41 to 50 amino acids, 51 to 60 amino acids, 61 to 70 amino acids, 71 to 80 amino acids, 81 to 90 amino acids, 91 to 100 amino acids, 101 to 110 amino acids, 111 to 120 amino acids, 121 to 130 amino acids, or 131 to 140 amino acids. , 141 to 150 amino acids, 151 to 160 amino acids, 161 to 170 amino acids, 171 to 180 amino acids, 181 to 190 amino acids, 191 to 200 amino acids, 201 to 210 amino acids, 211 to 220 amino acids, 221 to 230 amino acids, 231 to 240 amino acids, 241 to 250 amino acids, 251 to 260 amino acids, 261 to 270 amino acids, 271 to 280 amino acids, 281 to 290 amino acids, and 291 to 300 amino acids in length.

[0045] In certain embodiments, the transcription activation domain has 70 to 100%, 75 to 100%, 80 to 100, 85 to 100%, 90 to 100%, or 95 to 100% sequence identity to the amino acid sequence of SEQ ID NO: 2, 3, 4, 5, 6, 7, 8, 9, 10, or 11 (without the nuclear localization signal contained within the sequence), for example, SEQ ID NO: 3, 5, 6, 8, 9, 10, or 11. , for example, comprising or consisting of an amino acid sequence having at least 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity.

[0046] In one embodiment, the transcription activation domain has 60 to 100%, 65 to 100%, 70 to 100%, 75 to 100%, 80 to 100, 85 to 100%, 90 to 100%, or 95 to 100% sequence identity, for example, at least 100%, to the amino acid sequence of SEQ ID NO: 12, 13, 14, 15, 16, 17, 18, 19, 20, or 21 (a nuclear localization signal contained in the sequence), for example, SEQ ID NO: 13, 15, 16, 18, 19, 20, or 21. and / or any of the above amino acid sequences, each of which comprises or consists of an amino acid sequence having 61%, 62%, 63%, 64%, 65%, 66%, 67%, 68%, 69%, 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity with the other of the above amino acid sequences.

[0047] In very specific embodiments, said transcription activation domain belongs to the group of i) acidic domains (also called "acid blobs" or "negative noodles", rich in D and E amino acids), ii) glutamine-rich domains (comprising multiple repeats, e.g. "QQQXXXQQQ" type repeats), iii) proline-rich domains (comprising "PPPXXXPPP" like repeats) or iv) isoleucine-rich domains (comprising repeats, e.g. "IIXXII").

[0048] The present invention also relates to polypeptides comprising the modified non-viral, plant-based transcriptional activation domain of the present invention and a nuclear localization signal.

[0049] In one embodiment, the modified activation domain of the present invention is for an artificial transcription factor. The present invention also relates to artificial transcription factors. Generally, a transcription factor refers to a protein that binds to a specific DNA sequence present in an upstream activating sequence (UAS), thereby controlling the rate of transcription carried out by RNA II polymerase. Transcription factors perform this function alone or together with other proteins in a complex by promoting (as an activator) or blocking (as a repressor) the recruitment of RNA polymerase to the core promoter of a gene. An artificial or synthetic transcription factor (sTF) refers to a protein that functions as a transcription factor but is not a native protein of the host organism. The artificial transcription factor of the present invention comprises a transcription activation domain of the present invention, a DNA-binding domain, and a nuclear localization signal. In one embodiment, the DNA-binding protein of the artificial transcription factor is of prokaryotic origin. In one embodiment, the artificial transcription factor comprises a transcription activation domain of the present invention, a DNA-binding protein derived from a prokaryotic, typically bacterial, source, and a nuclear localization signal such as the SV40 NLS.

[0050] In the polypeptide or artificial transcription factor of the present invention, the nuclear localization signal can be any suitable localization signal known to those skilled in the art, such as the SV40 nuclear localization signal, or the nuclear localization signal can have an amino acid sequence comprising or consisting of the PKKKRKV.

[0051] A DNA binding domain refers to a region of a protein, typically a specific protein domain, that is involved in the interaction (binding) of the protein with a specific DNA sequence, such as the promoter of a target gene.

[0052] The modified transcription activation domain, polypeptide or artificial transcription factor of the present invention can be obtained from a polynucleotide that encodes the modified transcription activation domain, polypeptide or artificial transcription factor, or from a polynucleotide that has been modified to encode the modified transcription activation domain, polypeptide or artificial transcription factor.

[0053] The present invention also relates to polynucleotides encoding the transcription activation domains, polypeptides or artificial transcription factors of the invention.

[0054] A polynucleotide encoding a transcription activation domain, polypeptide, or artificial transcription factor of the present invention may be operably linked to any suitable promoter or regulatory sequence, including, but not limited to, a core promoter sequence, such as any one of those set forth in SEQ ID NOs: 22, 23, 25, 26, 27, 28, or SEQ ID NOs: 32-44, or any combination thereof.

[0055] As used herein, "polynucleotide" refers to any polynucleotide, such as, for example, single- or double-stranded DNA (synthetic, genomic, or cDNA) or RNA, that comprises a nucleic acid sequence that encodes a polymer of amino acids or a polypeptide of interest.

[0056] A codon is a trinucleotide unit that encodes a single amino acid in a protein-encoding gene. Codons that encode an amino acid may differ in any of their three nucleotides. Different organisms have different frequencies of codons in their genomes, which affects the efficiency of mRNA translation and protein production.

[0057] A coding sequence refers to a DNA sequence that encodes a specific RNA or polypeptide (i.e., a specific amino acid sequence). A coding sequence can optionally contain introns (i.e., additional sequences that interrupt the reading frame and are removed during maturation of the RNA molecule in a process called RNA splicing). If the coding sequence encodes a polypeptide, then the sequence includes a reading frame.

[0058] A reading frame is defined by a start codon (AUG in RNA, corresponding to ATG in DNA), and is a sequence of consecutive codons that encodes a polypeptide (protein). A reading frame ends with a stop codon (one of three: UAG, UGA, and UAA in RNA, corresponding to TAG, TGA, and TAA in DNA). Those skilled in the art can predict the location of open reading frames by using publicly available computer programs and databases.

[0059] The terms "polypeptide" and "protein" are used interchangeably herein to refer to polymers of amino acids of any length.

[0060] Variations or modifications of any one of the sequences or subsequences described and claimed herein remain within the scope of the invention, provided they can be used in the present invention or as an activation domain or polynucleotide encoding this activation domain for manipulation of gene expression.

[0061] The identity of any sequence or fragment thereof compared to a sequence of the present disclosure refers to the identity of that sequence compared to the entire sequence of the present invention. As used herein, the percent identity between two sequences is a function of the number of identical positions shared by the sequences, taking into account the number of gaps and the length of each gap that need to be introduced for optimal alignment of the two sequences (e.g., % identity = number of identical positions / total number of positions × 100). Sequence comparison and determination of the percent identity between two sequences can be accomplished using mathematical algorithms available in the art. This applies to both amino acid and nucleic acid sequences. As an example, sequence identity may be determined by using BLAST (Basic Local Alignment Search Tools) or FASTA (FAST-AII). Searches typically use the default setting parameters "gap penalties" and "matrix."

[0062] The expression cassette or expression system of the present invention comprises a polynucleotide encoding a transcription activation domain, polypeptide, or artificial transcription factor of the present invention. In one embodiment, the expression cassette further comprises a polynucleotide sequence encoding a desired product.

[0063] In one embodiment, a polynucleotide encoding a modified activation domain of the invention is for an expression cassette or expression system, or a modified activation domain of the invention is for an expression cassette or expression system.

[0064] In one embodiment, the expression system comprises one or more expression cassettes, optionally at least one expression cassette further comprising a polynucleotide sequence encoding a desired product.

[0065] The expression system of the present invention can be an orthogonal expression system, i.e., a system comprising or consisting of a heterologous (non-native) core promoter, transcription factor(s), and transcription factor-specific binding sites. Typically, an orthogonal expression system is functional (transducible) in a variety of eukaryotic organisms, such as eukaryotic microorganisms.

[0066] In one embodiment, the expression system comprises a target gene expression cassette and / or an artificial transcription factor expression cassette comprising the activation domain of the present invention. Furthermore, the expression system may further comprise, for example, one or more selectable marker (SM) expression cassettes and, optionally, genome-integrated DNA regions (flanking regions). In one embodiment, the expression system is constructed as a single DNA molecule or as two separate DNA molecules.

[0067] Figures 1, 2, 3 and 14 show exemplary schemes of expression systems or expression cassettes comprising activation domains of the invention, eg for heterologous protein production.

[0068] In one embodiment, a target gene expression cassette refers to a cassette comprising a target gene coding sequence and a sequence that controls expression (see Figures 1 to 3 and 14). In one embodiment, the expression cassette comprises a promoter sequence and / or a 3' untranslated region that optionally includes a polyadenylation site. The sequence controlling the expression of the target gene can include, but is not limited to, a promoter (e.g., a core promoter, e.g., An_201cp from Aspergillus niger, as exemplified in Figure 1 or Figure 2, or CP1 (e.g., Mm_Atp5Bcp, Mm_Eef2cp, or Mm_Rpl4cp from House mouse, or An_201cp from Aspergillus niger, or Yl_565cp from Yarrowia lipolytica) as exemplified in Figure 3 or Figure 14), and one or more sTF-specific binding sites (e.g., exemplified by sTF-specific binding sites (BS) in Figure 1, Figure 2, Figure 3, or Figure 14), which can be located, for example, upstream of the core promoter.

[0069] In one embodiment, the target gene expression cassette comprises a synthetic promoter, usually 1-10, typically 1, 2, 4 or 8 sTF binding sites, separated by 0-20, typically 5-15 random nucleotides, and a core promoter (CP), the target gene, and a terminator.

[0070] The target gene can be any DNA sequence (e.g., native or heterologous) that encodes a polypeptide or protein product of interest (see, e.g., Examples 1, 4, 6, 8, and 9, Figures 1-3, and 14). In one embodiment, transcription of the target gene is terminated on a transcription termination sequence (e.g., exemplified by the Trichoderma reesei pdc1 terminator (Tr_PDC1t) in Figure 1, the Saccharomyces cerevisiae ADH1 terminator (Sc_ADH1t) in Figure 2, either the SV40 terminator of simian virus 40 origin or the FTH1 terminator of Mus musculus origin in Figure 3, and the Saccharomyces cerevisiae ADH1 terminator in Figure 14).

[0071] In one embodiment, an artificial transcription factor (sTF) expression cassette comprises a core promoter (e.g., exemplified as Tr_hfb2cp in FIG. 1 , or An_008cp in FIG. 2 , or CP2 in FIG. 3 (Mm_Atp5Bcp, Mm_Eef2cp, or Mm_Rpl4cp of house mouse origin), or CP2 in FIG. 14 (e.g., An_008cp or Yl_242cp)), an sTF coding sequence, and a terminator (see FIGS. 1-3 and 14 ). The core promoter provides constitutive low expression of the sTF. The sTF binds to an sTF-dependent synthetic promoter in the target gene expression cassette, facilitating transcription of the target gene. The sTF comprises or consists of a DNA binding domain (BDB) that optionally includes or consists of a bacterial DNA binding protein (e.g., the Bm3R1 transcriptional regulator from Bacillus megaterium in Example 1, the PhlF transcriptional regulator from Pseudomonas protegens in Example 6, the McbR transcriptional regulator from Corynebacterium species in Example 6, or the TetR transcriptional regulator from Escherichia coli in Example 8) and / or a nuclear localization signal such as the SV40 NLS, and a transcription activation domain (AD). Transcription of the sTF gene can be terminated on a transcription termination sequence (e.g., as exemplified by the Trichoderma reesei tef1 terminator (Tr_tef1t) in Figure 1 or Figure 2, or by either the SV40 terminator of simian virus 40 origin or the FTH1 terminator of house mouse origin in Figure 3, or the Trichoderma reesei tef1 terminator in Figure 14).

[0072] In certain embodiments, the expression system comprises at least two individual expression cassettes, e.g., expression cassettes formed as one or more (e.g., two or more) DNA molecules: (a) a target gene expression cassette comprising a synthetic promoter containing a variable number of sTF binding sites, usually 1 to 10, typically 1, 2, 4 or 8, separated by 0 to 20, typically 5 to 15 random nucleotides, and a CP, a target gene, and a terminator; and (b) an artificial transcription factor cassette containing a CP that controls the expression of a gene encoding a fusion protein (artificial transcription factor, sTF), the artificial transcription factor itself (sTF), and a terminator; Includes.

[0073] A selectable marker (SM) expression cassette is any expression cassette that allows for the production of a particular protein in a host organism, providing the host or organism with the means to grow under selective conditions, such as in the presence of an antibiotic compound or the absence of an essential metabolite. In one embodiment of the present invention, the SM cassette is selected from the group consisting of the pyr4 gene (encoding the orotidine 5'-phosphate decarboxylase enzyme) in, for example, a Trichoderma reesei strain (see, e.g., Examples 1 and 3), the pyrG gene (encoding the orotidine 5'-phosphate decarboxylase enzyme) in, for example, an Aspergillus oryzae strain (see, e.g., Example 7), the hygromycin-B gene (encoding the orotidine 5'-phosphate decarboxylase enzyme) in, for example, a Myceliophthora thermophila strain, and the SM cassette is selected from the group consisting of the pyr4 gene (encoding the orotidine 5'-phosphate decarboxylase enzyme) in, for example, a Trichoderma reesei strain (see, e.g., Examples 1 and 3), the pyrG gene (encoding the orotidine 5'-phosphate decarboxylase enzyme) in, for example, an Aspergillus oryzae strain (see, e.g., Example 7), the SM cassette is selected from the group consisting of the pyr4 gene (encoding the orotidine 5'-phosphate decarboxylase enzyme) in, for example, a Myceliophthora thermophila strain (see, e.g., Example 8), the SM cassette is selected from the group consisting of the pyrG gene (encoding the orotidine 5'-phosphate decarboxylase enzyme) in, for example, a Myceliophthora thermophila strain (see, e.g., Example 9), the SM cassette is selected from the group consisting of the pyrG gene (encoding the orotidine 5'-phosphate decarboxylase enzyme) in, for example, a Myceliophthora thermophila strain (see The expression cassette may be an expression cassette allowing the expression of a hygR gene (encoding the 4-O-kinase) (see e.g. Example 5), a URA3 gene (encoding the orotidine 5'-phosphate decarboxylase enzyme) e.g. in a Pichia pastoris strain (see e.g. Example 4), a A gene (encoding the aminoglycoside phosphotransferase enzyme) e.g. in a Pichia pastoris strain (see e.g. Example 4), a pac gene (encoding the puromycin N-acetyltransferase enzyme) e.g. in a CHO cell (see e.g. Example 6), a kanR gene (encoding the aminoglycoside phosphotransferase enzyme) e.g. in a Pichia pastoris strain (see e.g. Example 8), and / or a NAT gene (encoding the nourseothricin N-acetyltransferase) e.g. in a Yarrowia lipolytica strain or a Cutaneotrichosporon oleaginosus strain (see e.g. Examples 8 and 9),

[0074] When the expression system is constructed as two separate DNA molecules, the first DNA can comprise or consist of an artificial transcription factor expression cassette containing the activation domain of the present invention, and optionally a selection marker (SM) expression cassette and / or a genome-integrated DNA region (flanking region), and the second DNA can comprise or consist of a target gene expression cassette, and optionally a selection marker (SM) expression cassette and / or a genome-integrated DNA region (flanking region). Each cassette can be integrated into a separate locus in the host genome and together form a functional gene expression system.

[0075] The genome-integrated DNA region (flanking region) used in the present invention can be any genomic locus present in the production host, such as genomic DNA sequences from Trichoderma reesei located upstream of the egl1 gene (EGL1-5') and downstream of the egl1 gene (EGL1-3') (see, e.g., Example 5), genomic DNA sequences from Pichia pastoris located upstream of the URA3 gene (URA3-5') and downstream of the URA3 gene (URA3-3') (see, e.g., Example 4), and genomic DNA sequences from Pichia pastoris located upstream of the AOX2 gene (AOX2-5') and downstream of the AOX2 gene (AOX2-3'). The target sequence can be selected from genomic DNA sequences (see, e.g., Example 4), or, for example, genomic DNA sequences derived from Aspergillus oryzae located upstream of the gaaC gene (gaaC-5') and downstream of the gaaC gene (gaaC-3') (see, e.g., Example 7), and genomic DNA sequences derived from Aspergillus oryzae located upstream of the gluC gene (gluC-5') and downstream of the gluC gene (gluC-3') (see, e.g., Example 7), or, for example, genomic DNA sequences for targeting the ADE1 gene of Pichia pastoris or the ant1 gene of Y. lipolytica (Examples 8 and 9).

[0076] In one particular embodiment of the invention, the expression system, e.g., for a eukaryotic or microbial host, comprises: (a) an expression cassette comprising a core promoter, which is the sole "promoter" controlling the expression of a DNA sequence encoding an activation domain or artificial transcription factor (sTF) of the invention; and (b) one or more expression cassettes comprising a target gene sequence encoding a desired protein product operably linked to a synthetic promoter comprising the same core promoter as (a) or a different core promoter, and an activation domain or sTF-specific binding site, respectively, upstream of the core promoter.

[0077] A eukaryotic promoter is a region of DNA required for initiating transcription of a gene. A eukaryotic promoter is located upstream of a DNA sequence (coding sequence) that encodes a specific RNA or polypeptide. A eukaryotic promoter includes an upstream activating sequence (UAS) and a core promoter. Those skilled in the art can predict the location of a promoter by using publicly available computer programs and databases.

[0078] A core promoter (CP) is the part of a (eukaryotic) promoter that is the region of DNA immediately upstream of the coding sequence encoding a polypeptide (5' upstream region), as defined by the start codon. A core promoter contains all the general transcriptional regulatory motifs required for transcription initiation, such as the TATA box, but does not contain any specific regulatory motifs, such as UAS sequences (natural activator and repressor binding sites).

[0079] The selection of a CP can be based on the expression level of a gene in a selected organism containing the candidate CP in its promoter. Another selection criterion can be the presence of a TATA box in the candidate CP. In one embodiment, screening for functional CPs used in the present invention is advantageously carried out by in vivo assembly of the candidate CP with an sTF-dependent reporter cassette expressed in an organism that constitutively expresses sTF, such as an S. cerevisiae strain. The resulting strain is tested for the level of reporter, preferably fluorescence, and these levels are compared with a control strain.

[0080] A core promoter (CP) typically comprises a DNA sequence comprising the 5' upstream region of a eukaryotic gene, starting 10 to 50 bp upstream of the TATA box and ending 9 bp upstream of the ATG start codon. In one embodiment, the distance between the TATA box and the start codon is 80 bp or more and 180 bp or less. A core promoter also typically comprises a DNA sequence comprising 1 to 20 bp of random sequence at its 3' end. In one embodiment, a core promoter comprises a DNA sequence having at least 90% sequence identity to the 5' upstream region of a eukaryotic gene and a DNA sequence comprising 1 to 20 bp of random sequence at its 3' end.

[0081] In one embodiment, the core promoter is a DNA sequence comprising: 1) a 5' upstream region of a highly expressed gene starting 10 to 50 bp upstream of the TATA box and ending 9 bp upstream of the start codon, wherein the distance between the TATA box and the start codon is 80 bp or more and 180 bp or less; and 2) a random 1 to 20 bp, typically 5 to 15 bp or 6 to 10 bp, positioned in place of the 9 bp of the DNA region (1) immediately upstream of the start codon; or a DNA sequence comprising: 1) a DNA sequence having at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity to the 5' upstream region; and 2) a random 1 to 20 bp, typically 5 to 15 bp or 6 to 10 bp, positioned in place of the 9 bp of the DNA region (1) immediately upstream of the start codon.

[0082] As used in the above sections, a "highly expressed gene" in an organism is a gene in that organism that has been shown to be expressed in the top 3% or 5% of all genes in any studied condition as determined by transcriptomic analysis, or, in organisms where transcriptomic analysis has not been performed, the gene that is the closest sequence homolog of a highly expressed gene.

[0083] A TATA box is a DNA sequence (TATA) upstream of the start codon that is at least 80 bp but not more than 180 bp from the start codon. If multiple sequences meet the above description, the TATA box is defined as the TATA sequence that is the closest to the start codon.

[0084] The core promoters (CPs) used in the expression system or one or several expression cassettes of the present invention may be different from each other or identical, for example the first one, CP1, can be the same as the second one, CP2 (or the third one, CP3, or the fourth one, CP4, in an expression system composed of multiple expression cassettes), or the first one, CP1, can be different from the second one, CP2.

[0085] In one embodiment, one or more CPs are universal core promoters functional in a variety of eukaryotic organisms. For example, Tr_hfb2cp (SEQ ID NO: 25), An_008cp (SEQ ID NO: 22), or Yl_242cp (SEQ ID NO: 33) can be used to control expression of sTFs in several organisms, such as Trichoderma reesei (see, e.g., Examples 1, 3, and 8), Aspergillus oryzae (see, e.g., Example 7), Myceliophthora thermophila strains (see, e.g., Example 5), Pichia pastoris (see, e.g., Example 8), or Yarrowia lipolytica (see, e.g., Example 8). In another embodiment of the invention, for example, An_201cp (SEQ ID NO: 23) can be used to control expression of a target gene in conjunction with an upstream sTF binding site in several organisms, such as Pichia pastoris (see, e.g., Examples 4 and 8), Trichoderma reesei (see, e.g., Examples 1 and 3 and 8), Aspergillus oryzae (see, e.g., Example 7), Myceliophthora thermophila strains (see, e.g., Example 5), or Yarrowia lipolytica (see, e.g., Example 8). Other CPs suitable for the present invention include, but are not limited to, An_008cp (SEQ ID NO: 22) (e.g., in Pichia pastoris; see Example 4), Mm_Atp5Bcp (SEQ ID NO: 26) (e.g., in Trichoderma reesei or CHO cells; see Examples 1 and 6), Mm_Eef2cp (SEQ ID NO: 27) (e.g., in Trichoderma reesei or CHO cells; see Examples 1 and 6), Mm_Rpl4cp (SEQ ID NO: 28), any CP of SEQ ID NOs: 32 to 44, or any combination thereof.

[0086] The sTF binding sites and core promoter (e.g., eight Bm3R1-specific binding sites and An_201cp; Figures 1 and 2) can form a synthetic promoter that strongly activates transcription of a target gene in the presence of an artificial transcription factor. In specific applications where the target gene is a native (homologous) gene of a host organism, the synthetic promoter can be inserted immediately upstream of the target gene coding region in the genome of the host organism, potentially replacing the original (native) promoter of the target gene.

[0087] A synthetic promoter refers to a region of DNA that functions as a eukaryotic promoter but is not a naturally occurring promoter in the host organism. A synthetic promoter contains an upstream activating sequence (UAS) and a core promoter, where either the UAS or the core promoter, or both elements, are not native to the host organism. In one embodiment of the present invention, the synthetic promoter contains (usually 1-10, typically 1, 2, 4, or 8) sTF-specific binding sites (synthetic UAS - sUAS) linked to the core promoter. In one embodiment of the present invention, the sTF binding sites and the core promoter form a synthetic promoter that strongly activates transcription of a target gene in the presence of an artificial transcription factor that can bind to the sTF binding sites. It is also possible to construct multiple synthetic promoters with different numbers of binding sites (usually 1-10, typically 1, 2, 4, or 8, separated by 0-20, typically 5-15 random nucleotides) to simultaneously control different target genes with a single sTF. This can result in a set of differentially expressed genes that form, for example, a metabolic pathway.

[0088] Two or more expression cassettes can be introduced into a eukaryotic host (typically integrated into the genome) either as two or more individual DNA molecules, or as one DNA molecule in which two or more expression cassettes are linked (fused) to form a single DNA molecule.

[0089] In one embodiment, the present invention provides tools for expression systems that are independent of the endogenous transcriptional regulation of the expression host.

[0090] Tuning of the expression system for different expression levels of at least the target gene and / or transcription factor can be performed in a host organism in which multiple options can be tested, including the choice of CP, sTF, different numbers of BS, and target gene.

[0091] The present invention relates to a non-viral transcription activation domain that can be used in a eukaryotic host. In one embodiment, the polypeptide, artificial transcription factor, polynucleotide, expression cassette, or expression system of the present invention is for a eukaryotic host. The eukaryotic host of the present invention comprises the transcription activation domain, polypeptide, artificial transcription factor, polynucleotide, expression cassette, or expression system of the present invention.

[0092] Eukaryotic (production) hosts suitable for the present invention can be selected from the group consisting of: 1) the class Saccharomycetales, which includes, but is not limited to, the species Saccharomyces cerevisiae, Kluyveromyces lactis, Candida krusei (Pichia kudriavzevii), Pichia pastoris (Komagataella pastoris), Pichia kudriavzevii, Eremothecium gossypii, Kazachstania exigua, Yarrowia lipolytica, Zygosaccharomyces lentus, and the like, or Schizosaccharomyces pombe yeasts such as the class of Schizosaccharomycetes, e.g., Aspergillus pombe; the kingdom Fungi, including filamentous fungi such as the class of Eurotiomycetes, including but not limited to the species Aspergillus niger, Aspergillus nidulans, Aspergillus oryzae, Penicillium chrysogenum, etc.; the kingdom Sordariomycetes, including but not limited to the species Trichoderma reesei, Myceliophthora thermophila, etc.; or the kingdom Mucorales, including filamentous fungi such as the class of Mucor indicus, e.g., Mucor indicus; 2) Mammals (Mammalia) and cells thereof, including but not limited to species such as Mus musculus (mouse), Chinese hamster (hamster), and Homo sapiens (human); and the Animalia kingdom, including but not limited to insects, including but not limited to species such as Mamestra brassicae, Spodoptera frugiperda, Trichoplusia ni, and Drosophila melanogaster.

[0093] In one embodiment, the eukaryotic host is selected from the group consisting of cells of fungal species, including yeast and filamentous fungi, and cells of animal species, including mammals (e.g., non-human mammals), or cells of Trichoderma, Trichoderma reesei, Pichia, Pichia pastoris, Pichia kudriabzevi, Aspergillus, Aspergillus oryzae, Aspergillus niger, Myceliophthora, Myceliophthora thermophila, Saccharomyces, Saccharomyces cerevisiae, Yarrowia, Yarrowia lipolytica, Cutaneotrichosporon, Cutaneotrichosporon oleaginosus, Cryptococcus curvatus, or the like. curvatus), Zygosaccharomyces, Chinese hamster ovary (CHO) cells, and Chinese hamster cells.

[0094] Methods for producing a desired protein product in a eukaryotic host include culturing the host under suitable culture conditions. "Suitable culture conditions" means any conditions that allow the survival or growth of the host organism and / or the production of a desired product in the host organism. The desired product can be a product of a target polynucleotide (i.e., a polypeptide or protein), or a compound produced by a polypeptide or protein or by a metabolic pathway. In this context, the desired product is typically a protein product.

[0095] The present invention also relates to the use of transcription activation domains, polypeptides, artificial transcription factors, polynucleotides, expression cassettes, expression systems, or eukaryotic hosts for metabolic engineering and / or production of desired protein products. As used herein, "metabolic engineering" refers to the control or optimization of genetic or regulatory processes within a cell. Metabolic engineering, for example, allows for the altered production of a desired protein product in a cell.

[0096] The tools of the present invention accelerate the process of industrial host development and enable the use of novel hosts that have high potential for specific purposes but have only a very limited range of tools for genetic engineering.

[0097] The present invention also relates to a method for preparing a non-viral transcription activation domain of the present invention or a polynucleotide encoding the non-viral transcription activation domain, the method comprising the steps of obtaining a transcription activation domain polypeptide derived from a plant transcription factor or obtaining a polynucleotide encoding the transcription activation domain polypeptide derived from a plant transcription factor, and modifying the obtained transcription activation domain polypeptide or polynucleotide. Methods for modifying polypeptides are well known to those skilled in the art and include, but are not limited to, methods that result in the deletion, substitution, disruption, or insertion of one or more amino acids or portions of a polypeptide, or the insertion of one or more modified amino acids. Methods for modifying polynucleotides are also well known to those skilled in the art and include, but are not limited to, methods that result in the deletion, substitution, disruption, or insertion of one or more nucleic acids or portions of a polynucleotide, or the insertion of one or more modified nucleic acids. Polypeptide modifications can be obtained, for example, by modifying a polynucleotide encoding the polypeptide by any genetic method. Methods for making genetic modifications are generally well known and are described in various practical manuals describing laboratory molecular techniques. General procedures and some examples of specific embodiments are described in the Examples section. In one particular embodiment of the present invention, the modified non-viral transcriptional activation domain was obtained by rational or random mutagenesis of a polynucleotide encoding said transcriptional activation domain.

[0098] It will be obvious to those skilled in the art that as technology advances, the concept of the present invention can be implemented in various ways. The present invention and its embodiments are not limited to the examples described below, but may vary within the scope of the claims. [Example]

[0099] Example 1 Testing transcription activation domains derived from plant transcription factors for heterologous gene expression in Trichoderma reesei (Figure 1, Figure 4) Reporter expression systems for testing different transcription activation domains were constructed as single DNA molecules (plasmids) (Figure 1). All plasmids contained T. reesei genome integration flanking regions (JGI122081; https: / / genome.jgi.doe.gov / Trire2 / Trire2.home.html) that allowed integration of the constructs into the egl1 locus. The egl1 integration flanking regions contained DNA sequences corresponding to DNA regions outside the egl1 coding region. EGL1-5' was the sequence 811 to 1811 bp upstream of the start codon. EGL1-3' was the sequence 2 to 1001 bp downstream of the stop codon. In addition, the plasmids contained the pyr4 selectable marker (SM) gene with an appropriate promoter and terminator. In addition, the plasmids contained the regions necessary for propagation of the plasmids in E. coli (not shown in Figure 1). This plasmid also contained a target gene cassette consisting of eight Bm3R1 binding sites (BS; sequence shown in Tables 1A and 1B); the An_201 core promoter (An_201cp; sequence shown in Tables 1A and 1B); DNA encoding mCherry (target gene; sequence shown in Tables 1A and 1B); and the Trichoderma reesei pdc1 terminator (Tr_PDC1t). This plasmid also contained a synthetic transcription factor (sTF) expression cassette consisting of the Trichoderma reesei hfb2 core promoter (Tr_hfb2cp; sequence shown in Tables 1A and 1B); the sTF coding region; and the Trichoderma reesei tef1 terminator (Tr_TEF1t).

[0100] The sTF coding regions of all plasmids contained the same DNA-binding domain (DBD; Bm3R1 transcriptional regulator from Bacillus megaterium; NCBI reference sequence: WP_013083972.1; coding DNA codon optimized for Aspergillus niger; sequence shown in Tables 1A and 1B) and SV40 NLS. The transcription activation domain (AD) was selected from plant transcription factors available in public databases, and the corresponding protein-coding DNA was codon-optimized for T. reesei. The following protein sequences were selected and used: At_NAC102-AD (SEQ ID NO: 2) = Region of amino acid sequence 126-215 from the AT5G63790 protein of Arabidopsis thaliana (GenBank: BAH57132.1) So_NAC102-AD (SEQ ID NO: 3) = Region of amino acid sequence 173-303 from Spinachia oleracea NAC domain-containing protein 2 (NCBI reference sequence: XP_021863783.1) At_TAF1-AD (SEQ ID NO: 4) = Region of amino acid sequence 129-229 from the ATAF1 protein of Arabidopsis thaliana (GenBank: CAA52771.1) So_NAC72-AD (SEQ ID NO: 5) = Region of amino acid sequence 185-369 from Spinachia oleracea NAC domain-containing protein 72 (NCBI Reference Sequence: XP_021840466.1) Bn_TAF1-AD (SEQ ID NO: 6) = Region of amino acids 186 to 286 from Brassica napus NAC domain-containing protein 2 (NCBI Reference Sequence: NP_001302866.1) At_JUB1-AD (SEQ ID NO: 7) = Region of amino acids 106 to 197 from Arabidopsis thaliana NAC domain-containing protein 42 (NCBI Reference Sequence: NP_001324496.1) So_JUB1-AD (SEQ ID NO: 8) = Region of amino acid sequence 227-357 from the JUNGBRUNNEN1-like protein of Spinachia oleracea (NCBI Reference Sequence: XP_021854333.1) Bn_JUB1-AD (SEQ ID NO: 9) = Region of amino acids 189-279 from the JUNGBRUNNEN1 protein of Brassica napus (NCBI Reference Sequence: XP_013670411.1) VP16-AD (SEQ ID NO: 1) was used as the transcription activation domain in the control construct.

[0101] Trichoderma reesei strain M1909 (VTT culture collection) was used as the parent strain. This strain is a mutagenized version of strain QM9414, containing additional deletions, including a deletion of the pyr4 gene, making the strain auxotrophic for uracil. A reporter expression system (Figure 1) was integrated into the egl1 locus (replacing the native coding region) using the corresponding flanking regions for homologous recombination. Transformation was performed using a CRISPR-Cas9-protein transformation protocol. Isolated T. reesei protoplasts were suspended in 1500 μL of STC solution (1.33 M sorbitol, 10 mM Tris-HCl, 50 mM CaCl2, pH 8.0). For each transformation, 100 μL of protoplast suspension was mixed with 2 μg of donor DNA (linear fragment corresponding to the construct shown in Figure 1) and 50 μL of EGL1-targeting RNP solution (1 μM Cas9 protein (IDT), 1 μM synthetic crRNA (IDT), and 1 μM tracrRNA (IDT)) and 100 μL of transformation solution (25% PEG6000, 50 mM CaCl2, 10 mM Tris-HCl, pH 7.5). This mixture was incubated on ice for 20 min. 2 mL of transformation solution was added, and the mixture was incubated at room temperature for 5 min. Four milliliters of STC was added, followed by 7 milliliters of molten (50°C) top agar (200 g / L D-sorbitol, 6.7 g / L Yeast Nitrogen Base (YNB, Becton, Dickinson and Company), synthetic complete amino acids without uracil, 20 g / L D-glucose, and 20 g / L agar). The mixture was poured onto selective plates (200 g / L D-sorbitol, 6.7 g / L Yeast Nitrogen Base (YNB, Becton, Dickinson and Company), synthetic complete amino acids without uracil, 20 g / L D-glucose, and 20 g / L agar).Cultures were grown at 28°C for 5 or 7 days, and colonies were picked and re-cultured on SCD-URA plates (6.7 g / L Yeast Nitrogen Base (YNB, Becton, Dickinson and Company), synthetic complete amino acids without uracil, 20 g / L D-glucose, and 20 g / L agar).

[0102] Correct strains were selected by qPCR of genomic DNA from each transformant. The qPCR signal of the mCherry gene was compared with the qPCR signal of the native sequence specific to each host. In addition, the correct deletion of the egl1 gene was confirmed by the absence of a qPCR signal for the egl1 target. Selected strains were sporulated on PDA agar plates (39 g / L BD-Difco potato dextrose agar). Spores (conidia) were collected from the PDA plates and used as inoculum in liquid culture for fluorescence analysis.

[0103] For quantitative fluorometric analysis of mCherry production in mycelia of the tested strains (Figure 4), precultures of Trichoderma reesei strains (inoculated with conidia) were grown for 24 h in YPG medium (20 g / L bacto peptone, 10 g / L yeast extract, and 30 g / L gelatin). Four milliliters of YE-glc medium (20 g / L glucose, 10 g / L yeast extract, 15 g / L KH2PO4, 5 g / L (NH4)2SO4, 1 mL / L trace elements (3.7 mg / L CoCl2, 5 mg / L FeSO4·7H2O, 1.4 mg / L ZnSO4·7H2O, 1.6 mg / L MnSO4·7H2O), 2.4 mM MgSO4, and 4.1 mM CaCl2, pH adjusted to 4.8) in a 24-well culture plate was inoculated with the mycelium suspension to an OD600 of 0.5. The culture was grown at 800 rpm (Infors HT Microtron) and 28°C for 24 hours, centrifuged, and the pellet was washed with water and resuspended in 0.2 mL of sterile water. 200 μL of each mycelium suspension was analyzed in a black 96-well plate (Black Cliniplate; Thermo Scientific) using a Varioskan (Thermo Electron Corporation) fluorometer. The mCherry settings were 587 nm (excitation) and 610 nm (emission). To normalize the fluorescence results, the analyzed mycelium suspension was diluted 100-fold, and the OD600 was measured in a clear 96-well microtiter plate (NUNC) using a Varioskan (Thermo Electron Corporation). The results of the analysis are shown in Figure 4.

[0104] [Table 1(1)] [Table 1(2)] [Table 1(3)] [Table 1(4)] [Table 1(5)]

[0105] Example 2. Mutagenesis of selected activation domains to improve activity To enhance the activity of plant-based transcription activation domains, we performed rational mutagenesis on two selected activation domains derived from transcription factors found in the edible plant species spinach (Spinach aurea) and rapeseed / canola (Brassica napus). So_NAC102-AD and Bn_TAF1-AD (Example 1) contain significant amounts of acidic (glutamic acid and aspartic acid) and hydrophobic (leucine, isoleucine, phenylalanine) amino acids, indicating that they may belong to the group of acidic / hydrophobic transcription activation domains, which are typically rich in these types of amino acids. However, several basic amino acids (lysine and arginine) are present in the native sequences of these activation domains. We modified the sequences of these selected activation domains by mutating some of these amino acids (and introducing other changes) to obtain a more pronounced acid / hydrophobic pattern. Two novel activation domains were designed. So_NAC102M (SEQ ID NO: 10)-AD = So_NAC102-AD with the following amino acid changes: removal (deletion) of amino acids 1 to 3, and mutations K18L, K44L, R58D, C59L, K78L, K85L, and K91D. Bn_TAF1M (SEQ ID NO: 11)-AD = Bn_TAF1-AD with the following amino acid changes: K25D, K51L, K53D, K62D.

[0106] The new activation domains were tested using the same setup and following the same procedures as in Example 1. These domains were tested in a reporter expression system (Figure 1), and the fluorescence of T. reesei strains containing the corresponding reporter expression systems was analyzed, as shown in Figure 4. It was demonstrated that the modifications introduced into So_NAC102-AD and Bn_TAF1-AD resulted in significantly more active activation domains, So_NAC102M-AD and Bn_TAF1M-AD.

[0107] Example 3. Production of a prokaryotic xylanase in Trichoderma reesei by a synthetic expression system containing a plant-derived activation domain The five best-performing expression systems containing plant-based activation domains (indicated by arrows) as well as the expression systems with So_NAC102-AD and Bn_TAF1-AD were compared with an expression system containing VP16-AD (as a benchmark control) based on the results shown in Figure 4. The comparison was performed in experiments in which an exemplary heterologous protein product was produced by Trichoderma reesei (secreted into the culture medium). The expression systems described in Examples 1 and 2 were modified by replacing the mCherry coding sequence with a DNA sequence encoding an alkaline xylanase (thermostable mutant xynHB_N188A, SEQ ID NO: 31) from Bacillus pumilus, which had previously been produced in Pichia pastoris (Lu, Y. et al., 2016, Scientific Reports, Vol. 6, Paper No. 37869). The DNA encoding the xylanase was codon-optimized for Trichoderma reesei, and an appropriate secretory signal sequence (SS) with a Kex2 recognition site was added in-frame to its 5' end. This resulted in DNA encoding a fusion protein (SS-Kex2-xynHB_N188A; target gene in Figure 1). This fusion protein can be efficiently processed and secreted into the culture medium by T. reesei.

[0108] The xylanase expression cassette was transformed into T. reesei using the protocol described in Example 1. Trichoderma reesei strain M1909 was used as the parent strain, and DNA was transformed into T. reesei protoplasts using the CRISPR-Cas9 protein transformation protocol. Selection of transformed colonies and analysis of strains were performed as described above (Experiment 1), except that the xynHB_N188A gene was targeted in qPCR analysis instead of the mCherry gene.

[0109] Xylanase production was tested in small-scale liquid cultures and analyzed in the culture supernatants by SDS-PAGE (Figure 5). Four milliliters of YE-glc medium (20 g / L glucose, 10 g / L yeast extract, 15 g / L KH2PO4, 5 g / L (NH4)2SO4, 1 mL / L trace elements (3.7 mg / L CoCl2, 5 mg / L FeSO4·7H2O, 1.4 mg / L ZnSO4·7H2O, 1.6 mg / L MnSO4·7H2O), 2.4 mM MgSO4, and 4.1 mM CaCl2, pH adjusted to 4.8, in 24-well culture plates was inoculated with conidia of selected clones collected from PDA plates. The cultures were incubated at 28°C and 800 rpm (Infors HT Microtron) for 3 days and then centrifuged to pellet the mycelium. One hundred microliters of each culture supernatant was mixed with 50 μL of 4x SDS loading buffer (400 mL / L glycerol; 240 mM Tris·HCl, pH 6.8; 80 g / L SDS; 0.4 g / L bromophenol blue; and 50 mL / L β-mercaptoethanol) and incubated at 95°C for 4 minutes. 15 μL of this mixture was loaded onto a 4-20% SDS-PAGE gradient gel next to molecular weight standards. After complete protein separation in an electric field (PowerPac HC; BioRad), the gel was stained with colloidal Coomassie stain (PageBlue Protein Staining Solution; Thermo Fisher Scientific) according to the manufacturer's protocol. Visualization of the stained gel was performed using an Odyssey CLx Imaging System (LI-COR Biosciences). A scan of the stained gel is shown in Figure 5. The relative amounts of xylanase produced corresponded somewhat to the mCherry fluorescence levels shown in Figure 4. The best-performing expression systems with plant-based activation domains were the So_NAC102M- and Bn_TAF1M-containing systems. These two corresponding strains, as well as the strain producing xylanase in the VP16-AD-containing expression system, were tested in a 1 L bioreactor setup for evaluation of xylanase production.

[0110] 1 L bioreactor cultivation was performed in a Sartorius Stedim BioStat Q Plus Fermentor Bioreactor System. The preculture (inoculated with conidia) was grown in 100 mL of YE-glc medium for 24 hours to generate a sufficient amount of mycelium for bioreactor inoculation. Bioreactor cultivation was initiated by inoculating 800 mL of YE-glucose medium (10 g / L glucose, 20 g / L yeast extract, 5 g / L KH2PO4, 5 g / L NH4SO4, 1 mL / L trace elements, 2.4 mM MgSO4, and 4.1 mM CaCl2, 1 mL / L Antifoam J647, pH 4.8) with 80 mL of preculture. These cultures were continuously supplied with 500 g / L glucose (using a Watson Marlow 120 U / DV peristaltic pump at a flow rate of 0.3–0.7 rpm), an airflow of 0.5 slpm (0.4–0.6 vvm), and agitated at 900–1200 rpm. Cultures were grown for 6 days, with samples taken daily. A subset of the culture supernatants was analyzed by SDS-PAGE (Figure 6) and for xylanase activity (Figure 7).

[0111] Two microliters of culture supernatant from each culture at different time points was loaded onto a gel (4–20% gradient), and proteins were separated using an electric field (PowerPac HC; BioRad). The gel was stained with colloidal Coomassie (PageBlue Protein Staining Solution; Thermo Fisher Scientific) and visualized using an Odyssey CLx Imaging System (LI-COR Biosciences). A scan of the stained gel is shown in Figure 6. Xylanase production appeared equally well in all three strains, demonstrating the utility of the selected plant-based activation domain in potentially replacing the viral-based VP16 activation domain for heterologous protein production in Trichoderma reesei.

[0112] Culture supernatants from xylanase-producing bioreactor cultures (days 5 and 6) and from bioreactor cultures run under the same conditions using a T. reesei strain without a xylanase-producing expression system (day 6, negative control—NC in Figure 7) were serially diluted in 50 mM Tris·HCl (pH 8.0). Xylanase activity was assayed using the EnzCheck® Ultra Xylanase Assay Kit (Invitrogen). Fifty μL of the diluted culture supernatants was mixed with 50 μL of a 50 μg / mL solution of xylanase substrate (component A of the kit) in 50 mM Tris·HCl (pH 8.0) in a black 96-well plate (Black Cliniplate; Thermo Scientific). The reaction was incubated in the dark at room temperature for 25 minutes. Fluorescence of the xylanase reaction product (released from the substrate by the action of xylanase) was measured using a Varioskan (Thermo Electron Corporation) fluorometer. The measurement settings were 358 nm (excitation) and 455 nm (emission). Activity was calculated and expressed as arbitrary units per mL of culture supernatant (AU / mL). The resulting xylanase activity is shown in Figure 7. These results again clearly demonstrate that the selected plant-based activation domain can be successfully used in place of the virus-based VP16 AD for heterologous gene expression without loss of expression level. In fact, the xylanase activity in the supernatant from cultures using strains containing the plant-based AD in this expression system appears to be higher than the corresponding activity from the VP16 control (day 5, Figure 7). Additionally, these results clearly demonstrate that the xylanase protein produced in Trichoderma reesei is a functional, catalytically active enzyme.

[0113] Example 4. Production of Prokaryotic Phytase in Pichia pastoris by a Synthetic Expression System Containing a Plant-Derived Activation Domain To construct a synthetic expression system for Pichia pastoris, we selected the five best-performing plant-based activation domains (indicated by arrows) and VP16-AD (as a benchmark control) based on the results presented in Figure 4. Comparison of these gene constructs (transcription activation domains) was performed in experiments in which an exemplary heterologous protein product was produced (secreted into the culture medium) by Pichia pastoris. The expression system (Figure 2) was constructed as two separate DNA molecules (plasmids).

[0114] The first DNA consisted of 1) an sTF expression cassette; 2) a selectable marker (SM) expression cassette; 3) a genome-integrated DNA region (flanking regions); and 4) regions required for propagation of this plasmid in Escherichia coli (E. coli). The sTF expression cassette consisted of a core promoter (An_008cp, SEQ ID NO: 22), an sTF coding sequence, and a terminator (see Tables 1C and 1D for exemplary sequences of sTF expression cassettes used in Pichia pastoris). The sTF gene encoded a fusion protein (synthetic transcription factor) consisting of the bacterial DNA-binding protein Bm3R1 (the coding DNA sequence was codon-optimized for Saccharomyces cerevisiae), the SV40 nuclear localization signal (NLS), a short peptide linker, and a transcription activation domain (AD). The DNA sequence encoding the activation domain was codon-optimized for Pichia pastoris. The control AD ​​was VP16-AD. The terminator was the Trichoderma reesei tef1 terminator (Tr_TEF1t). The SM cassette was an expression cassette that enabled expression of the kanR gene (encoding the aminoglycoside phosphotransferase enzyme) in Pichia pastoris using an appropriate promoter and terminator. The above genomic integration DNA regions (flanking regions) were used to enable integration of the above construct into the URA3 locus of P. pastoris (JGI38543; https: / / genome.jgi.doe.gov / Picpa1 / Picpa1.home.html). The URA3 integration flanking regions contained DNA sequences corresponding to DNA regions outside the URA3 coding region. URA3-5' was the sequence 500 to 1 bp upstream of the start codon. URA3-3' was the sequence 1 to 499 bp downstream of the stop codon.

[0115] The second DNA consisted of 1) a target gene expression cassette; 2) a selectable marker (SM) expression cassette; 3) a genome-integrated DNA region (flanking regions); and 4) a region required for propagation of this plasmid in E. coli. The target gene expression cassette contained eight Bm3R1 binding sites (BS; sequence shown in Tables 1A and 1B); the An_201 core promoter (An_201cp, SEQ ID NO: 23; sequence shown in Tables 1A and 1B); DNA encoding the target gene (target gene); and the Saccharomyces cerevisiae ADH1 terminator (Sc_ADH1t). The target gene was a DNA sequence encoding a thermostable mutant phytase enzyme (AppA_K24E, amino acid sequence SEQ ID NO: 24) of E. coli origin, which had previously been produced in Pichia pastoris (Zhang J. et al., 2016, Biosci. Biotech. Res. Comm. 9(3):357-365). The phytase-encoding DNA was codon-optimized for Pichia pastoris, and an appropriate secretory signal sequence (SS) with a Kex2 recognition site was added in-frame to its 5' end. This resulted in DNA encoding a fusion protein (SS-Kex2-AppA_K24E; target gene in Figure 2). This fusion protein could be efficiently processed and secreted into the culture medium by P. pastoris. The SM cassette was an expression cassette that enabled expression of the URA3 gene (encoding the orotidine 5'-phosphate decarboxylase enzyme) in Pichia pastoris using an appropriate promoter and terminator. The genomic integration DNA region (flanking region) described above was used to enable integration of the construct into the AOX2 locus of P. pastoris (JGI39494; https: / / genome.jgi.doe.gov / Picpa1 / Picpa1.home.html). The AOX2 integration flanking region contained DNA sequences corresponding to DNA regions within and outside the AOX2 coding region. AOX2-5' was the sequence 504 to 6 bp upstream of the start codon, and AOX2-3' was the sequence starting at bp 1806 of the coding region and ending at bp 313 after the stop codon.

[0116] Each cassette was integrated into a separate locus in the P. pastoris genome. The transformations were performed sequentially. First, the sTF expression cassette-containing construct was integrated into the P. pastoris parent strain to form the sTF background strain. Then, the target gene expression cassette-containing construct was integrated into the sTF background strain to form the final production strain.

[0117] Pichia pastoris strain Y-11430 (now also known as Komagataella phafii, obtained from the NRRL Culture Collection) was used as the parent strain. The sTF expression cassette-containing construct (Figure 2) was integrated into the URA3 locus (replacing the native coding region) using the corresponding flanking regions for homologous recombination. Transformation was performed using the CRISPR-Cas9 protein transformation protocol. Isolated P. pastoris protoplasts were suspended in 600 μL of STC solution (1.33 M sorbitol, 10 mM Tris-HCl, 50 mM CaCl2, pH 8.0). For each transformation, 100 μL of protoplast suspension was mixed with 5 μg of donor DNA (linear fragment corresponding to the construct shown in Figure 2) and 50 μL of URA3-targeting RNP solution (1 μM Cas9 protein (IDT), 1 μM synthetic crRNA (IDT), and 1 μM tracrRNA (IDT)) and 100 μL of transformation solution (25% PEG6000, 50 mM CaCl2, 10 mM Tris-HCl, pH 7.5). This mixture was incubated on ice for 20 min. 2 mL of transformation solution was added, and the mixture was incubated at room temperature for 5 min. Four milliliters of STC was added, followed by 7 milliliters of molten (50°C) top agar (200 g / L D-sorbitol, 20 g / L bacto peptone, 10 g / L yeast extract, 1 g / L uracil, 20 g / L D-glucose, 500 mg / L G418, and 20 g / L agar). The mixture was poured onto selective plates (200 g / L D-sorbitol, 20 g / L bacto peptone, 10 g / L yeast extract, 1 g / L uracil, 20 g / L D-glucose, 500 mg / L G418, and 20 g / L agar). The plates were incubated at 30°C for 5 or 7 days until colonies appeared. Colonies were picked and re-cultured on YPD-G418 selection plates (20 g / L bacto peptone, 10 g / L yeast extract, 1 g / L uracil, 20 g / L D-glucose, 500 mg / L G418, and 20 g / L agar).

[0118] Transformed clones were first tested for growth in the absence of uracil, and those unable to grow were analyzed by qPCR. Genomic DNA from each selected strain was isolated and used as template DNA in qPCR reactions. The qPCR signal of the sTF gene (Bm3R1) was compared with the qPCR signal of the unique native sequence in each strain. In addition, precise deletion of the URA3 gene was confirmed by the absence of a qPCR signal for the URA3 target. Strains with the precise URA3 deletion and a single-copy sTF cassette integrated into the genome (sTF background strains) were selected for a second round of transformation.

[0119] The second transformation was performed using the lithium acetate protocol. The sTF background strain was grown in YPD+URA medium (20 g / L bacto peptone, 10 g / L yeast extract, 1 g / L uracil, 20 g / L D-glucose) to reach an OD600 of 0.6–1.0. 50 mL of each culture was centrifuged, and the cell pellet was washed with water and then with LiAc / TE solution (100 mM lithium acetate; 10 mM Tris HCl (pH 7.5); 1 mM EDTA). The washed cell pellet was resuspended in 0.5 mL of LiAc / TE solution. A 50 μL cell suspension was mixed with 10 μg of AppA expression construct DNA (a linear AppA target gene expression cassette fragment corresponding to the construct shown in Figure 2) and 400 μL of LiAc transformation solution (40% polyethylene glycol 4000 (PEG-4000); 100 mM lithium acetate; 10 mM Tris·HCl (pH = 7.5); 1 mM EDTA; 400 μg / mL herring sperm-derived DNA). The mixture was incubated at 30°C for 30 min and then at 42°C for 20 min. The transformation mixture was centrifuged, and the cell pellet was resuspended in 200 μL of water and plated on SCD-URA plates (6.7 g / L Yeast Nitrogen Base (YNB, Becton, Dickinson and Company), synthetic complete amino acids without uracil, 20 g / L D-glucose, and 20 g / L agar). The plates were cultured at 30°C for 3 or 5 days until colonies appeared. Colonies were picked and re-cultured on SCD-URA plates.

[0120] Genomic DNA from each selected clone was isolated and used as template DNA in qPCR reactions. The qPCR signal of the target gene (AppA) was compared with that of the native sequence unique to each strain. Strains with a single-copy target gene cassette integrated into their genome were used for phytase production experiments.

[0121] Phytase production was tested in small-scale liquid cultures and analyzed in the culture supernatants by SDS-PAGE (Figure 8). Cells of selected clones were inoculated into 4 mL of BMG medium (20 g / L glucose, 10 g / L yeast extract, 20 g / L bacto peptone, 13.4 g / L YNB, 0.4 mg / L biotin, and 100 mM KH2PO4, pH = 6.0) in 24-well culture plates. The cultures were incubated at 28 °C and 800 rpm (Infors HT Microtron) for 2 days and then centrifuged to pellet the cells. 100 μL of each culture supernatant was mixed with 50 μL of 4x SDS supplement (400 mL / L glycerol; 240 mM Tris·HCl, pH 6.8; 80 g / L SDS; 0.4 g / L bromophenol blue; and 50 mL / L β-mercaptoethanol) and incubated at 95°C for 4 minutes. 15 μL of the mixture was loaded onto a 4-20% SDS-PAGE gradient gel next to molecular weight standards. After complete protein separation using an electric field (PowerPac HC; BioRad), the gel was stained with colloidal Coomassie stain (PageBlue Protein Staining Solution; Thermo Fisher Scientific) according to the manufacturer's protocol. Visualization of the stained gel was performed using an Odyssey CLx Imaging System (LI-COR Biosciences). A scan of the stained gel is shown in Figure 8. Based on these results, the best-performing expression systems with plant-based activation domains appeared to be the So_NAC102M- and Bn_TAF1M-containing systems. These two corresponding strains, as well as the strain producing phytase in the VP16-AD-containing expression system, were tested in a 1 L bioreactor setup for evaluation of phytase production.

[0122] 1 L bioreactor cultures were performed in a Sartorius Stedim BioStat Q Plus Fermentor Bioreactor System. The preculture was grown in 100 mL of BMG medium for 24 hours to generate sufficient biomass for bioreactor inoculation. Bioreactor cultures were initiated by inoculating 800 mL of BMG medium containing 1 mL / L Antifoam J647 with 80 mL of the preculture. These cultures were continuously supplied with 500 g / L glucose (using a Watson Marlow 120U / DV peristaltic pump at a flow rate of 0.3–0.7 rpm), an airflow of 0.5 slpm (0.4–0.6 vvm), and agitated at 900–1200 rpm. Cultures were grown for 6 days, with samples taken daily. Culture supernatants were analyzed by SDS-PAGE (Figure 9) and for phytase activity (Figure 10).

[0123] Two microliters of culture supernatant from each culture at different time points was loaded onto a gel (4–20% gradient), and proteins were separated using an electric field (PowerPac HC; BioRad). The gel was stained with colloidal Coomassie (PageBlue Protein Staining Solution; Thermo Fisher Scientific) and visualized using an Odyssey CLx Imaging System (LI-COR Biosciences). A scan of the stained gel is shown in Figure 9. AppA_K24E phytase appeared to be produced equally well in all three strains, demonstrating the utility of the selected plant-based activation domain in potentially replacing the viral-based VP16 activation domain for heterologous protein production in Pichia pastoris.

[0124] Culture supernatants from phytase-producing bioreactor cultures (days 4 and 6) and from bioreactor cultures run under the same conditions using a P. pastoris strain without a phytase-producing expression system (negative control—NC in Figure 10) were subjected to gel filtration to remove phosphate, which interferes with the phytase assay. Gel filtration was performed on a PD-10 desalting column (BioRad) using 100 mM Na acetate (pH 4.7). The eluate from the gel filtration was assayed for phytase activity using a Phytase Assay Kit (MyBioSource). 14 μL of the eluate diluted with phytase reaction buffer was combined with 56 μL of substrate solution (containing phytic acid; Reagent No. 1 in the kit) in a clear 96-well plate (Thermo Scientific) and incubated at 37°C for 30 minutes. 70 μL of reaction termination solution (Reagent No. 2 in the kit) was added, followed by 70 μL of color development solution. The solution was mixed and incubated at room temperature for 10 minutes. The absorbance of the phosphomolybdate complex (a phytase reaction product released from phytic acid conjugated to molybdate by the action of phytase) was measured using a Varioskan (Thermo Electron Corporation) instrument. The absorbance of the solution was measured at 700 nm. Activity was calculated and expressed as arbitrary units per mL of culture supernatant (AU / mL). The resulting phytase activity is shown in Figure 10. These results clearly demonstrate that the selected plant-based activation domain can be successfully used in place of the virus-based VP16 AD for heterologous gene expression in Pichia pastoris without loss of expression level. Additionally, these results clearly demonstrate that the produced phytase protein is a functional, catalytically active enzyme.

[0125] Example 5. Production of prokaryotic xylanase in Myceliophthora thermophila using a synthetic expression system containing a plant-derived activation domain The two best-performing plant-based activation domains (So_NAC102M and Bn_TAF1M), with results presented in Figures 5, 6, 7, 8, and 9, were compared with VP16-AD in experiments in which exemplary heterologous protein products were produced (secreted into the culture medium) by Myceliophthora thermophila. The expression systems described in Example 3, containing the xylanase expression cassette So_NAC102M-AD, Bn_TAF1M-AD, or VP16-AD, were modified by replacing the pyr4 selection marker (SM) expression cassette with the hygR selection marker (SM) expression cassette, enabling expression of the hygR gene (encoding hygromycin-B 4-O-kinase) in Myceliophthora thermophila.

[0126] Myceliophthora thermophila strain D-76003 (also known as Thielavia heterothallica, VTT culture collection) was used as the parent strain, and DNA was transformed into M. thermophila protoplasts using the PEG transformation protocol. Isolated M. thermophila protoplasts were suspended in 400 μL of STC solution (1.33 M sorbitol, 10 mM Tris-HCl, 50 mM CaCl2, pH 8.0). For each transformation, 100 μL of the protoplast suspension was mixed with 30 μg of expression construct DNA (linear fragment corresponding to the construct shown in Figure 1) dissolved in less than 100 μL of solution and 100 μL of transformation solution (25% PEG 6000, 50 mM CaCl2, 10 mM Tris-HCl, pH 7.5). The mixture was incubated on ice for 20 min. Two milliliters of transformation solution was added, and the mixture was incubated at room temperature for 5 minutes. Four milliliters of STC was added, followed by 7 milliliters of molten (50°C) top agar (200 g / L D-sorbitol, 20 g / L D-glucose, 20 g / L bacto peptone, 10 g / L yeast extract, 200 mg / L hygromycin-B, and 20 g / L agar). The mixture was poured onto selective plates (200 g / L D-sorbitol, 20 g / L D-glucose, 20 g / L bacto peptone, 10 g / L yeast extract, 200 mg / L hygromycin-B, and 20 g / L agar). Cultures were grown at 35°C for 4 to 7 days, and colonies were picked and recultured on YPD-HYG plates (20 g / L D-glucose, 20 g / L bacto peptone, 10 g / L yeast extract, 200 mg / L hygromycin-B, and 20 g / L agar).

[0127] Four clones from each transformation were selected for small-scale liquid culture and analysis of the culture supernatant by SDS-PAGE (Figure 8). Four milliliters of BMG medium (20 g / L glucose, 10 g / L yeast extract, 20 g / L bacto peptone, 13.4 g / L YNB, 0.4 mg / L biotin, and 100 mM KH2PO4, pH = 6.0) in 24-well culture plates was inoculated with a mixture of mycelium and conidia collected from clones growing on YPD-HYG plates. The cultures were incubated at 35°C and 800 rpm (Infors HT Microtron) for 3 days and then centrifuged to pellet the mycelium. 100 μL of each culture supernatant was mixed with 50 μL of 4x SDS supplement (400 mL / L glycerol; 240 mM Tris·HCl, pH 6.8; 80 g / L SDS; 0.4 g / L bromophenol blue; and 50 mL / L β-mercaptoethanol) and incubated at 95°C for 4 minutes. 15 μL of the mixture was loaded onto a 4-20% SDS-PAGE gradient gel next to molecular weight standards. After complete protein separation using an electric field (PowerPac HC; BioRad), the gel was stained with colloidal Coomassie stain (PageBlue Protein Staining Solution; Thermo Fisher Scientific) according to the manufacturer's protocol. Visualization of the stained gel was performed using an Odyssey CLx Imaging System (LI-COR Biosciences). A scan of the stained gel is shown in Figure 11. There was a large variation in xylanase production levels among individual clones, which is the result of random DNA integration (the transformed DNA is not targeted to a specific genomic locus). In this type of transformation, the expression cassette is typically integrated into various unknown genomic loci in one or more integration events. However, the range of xylanase production levels obtained, particularly the maximum xylanase production in certain clones, indicates that the plant-based activation domains (So_NAC102M and Bn_TAF1M) can provide similar or even higher levels of heterologous gene expression than the virus-based VP16 AD.It is therefore clear that this plant-based activation domain can be successfully used in place of a virus-based activation domain for the production of recombinant proteins in Myceliophthora thermophila.

[0128] Culture supernatants from cultures of M. thermophila strains transformed with the xylanase expression constructs and from cultures transformed under the same conditions as the parent M. thermophila strain (NC in Figure 12) were serially diluted in 50 mM Tris·HCl (pH 8.0) and assayed for xylanase activity using the EnzCheck® Ultra Xylanase Assay Kit (Invitrogen). Fifty microliters of the diluted culture supernatants were mixed with 50 μL of a 50 μg / mL solution of xylanase substrate (component A of the kit) in 50 mM Tris·HCl (pH 8.0) in a black 96-well plate (Black Cliniplate; Thermo Scientific). The reactions were incubated in the dark at room temperature for 25 minutes. Fluorescence of the xylanase reaction product (released from the substrate by the action of xylanase) was measured using a Varioskan (Thermo Electron Corporation) fluorometer. The measurement settings were 358 nm (excitation) and 455 nm (emission), respectively. Activity was calculated and expressed in arbitrary units per mL of culture supernatant (AU / mL). The obtained xylanase activity is shown in Figure 12. These results closely correlate with those presented in Figure 11 and clearly demonstrate that the xylanase protein produced in Myceliophthora thermophila is a functional, catalytically active enzyme.

[0129] Example 6 Testing selected plant-derived activation domains in CHO (Chinese Hamster Ovary) cells The two best plant-based activation domains based on fungal experiments, So_NAC102M and Bn_TAF1M, were used to construct an artificial expression system for Chinese hamster ovarian (CHO) cells (see Tables 1E and 1F for exemplary sequences of expression cassettes for CHO cells). The CHO K1 cell line was transformed with a plasmid containing eight sTF-specific binding sites (8BS) located upstream of the core promoter Mm_Atp5Bcp (SEQ ID NO: 26). The target gene mCherry was located immediately after the core promoter. Transcription of mCherry was terminated by the SV40 terminator. Adjacent to the mCherry expression cassette, in the opposite direction, was an sTF expression cassette consisting of the core promoter Mm_Eef2cp (SEQ ID NO: 27), the PhlF repressor, the nuclear localization signal, the SV40 NLS, and a transcription activation domain (AD) of plant origin. Transcription of the sTF gene terminated at the transcription termination sequence FTH1 terminator of Mus musculus origin. This plasmid also contains the pac gene, which encodes the puromycin N-acetyltransferase enzyme, which confers resistance to the antibiotic puromycin. The performance of these expression systems is compared to an expression system that uses the CMV (cytomegalovirus) promoter for expression of mCherry, and an artificial expression system in which the VP64 activation domain (of herpes simplex virus origin) (SEQ ID NO: 30) is used instead of the plant-based AD.

[0130] CHO-K1 cells are maintained in RPMI medium (Thermo Fischer) supplemented with 2 mM L-glutamine, 10% fetal bovine serum, and penicillin-streptomycin solution to a final concentration of 100 units of penicillin and 0.1 g / L streptomycin. Cells are grown at 37°C in the presence of 5% CO2. The day before transfection, 70-80% confluent CHO cells are washed with PBS, pH approximately 7.4, and then transfected with 250 mL of 75 cm 2Trypsinize the culture in the flask by adding 2 mL of trypsin and incubating at 37 °C for 2-4 minutes until the cells dissociate. Add 8 mL of fresh RPMI medium containing the above supplements to the flask. Pipette 100 µL of the above cell solution into each well of a 24-well plate containing 400 µL of RPMI medium (1 / 5 dilution) supplemented with 2 mM L-glutamine, 10% fetal bovine serum, and penicillin-streptomycin solution to a final concentration of 100 units of penicillin and 0.1 g / L streptomycin. The next day, remove the medium by pipetting and immediately replace it with 400 µL of fresh RPMI medium without antibiotic supplements. Incubate the cells at 37 °C with 5% CO2 for 20 minutes. For each transfection, combine 2 μL of Lipofectamine LTX (Thermo Fischer) with 25 μL of Opti-MEM medium (Thermo Fischer), and combine 0.5–1 μg of plasmid DNA with 0.5 μL of Plus Reagent (provided with the Lipofectamine LTX Reagent) and 25 μL of Opti-MEM medium. The Opti-MEM-diluted DNA is then mixed with the diluted Lipofectamine® LTX Reagent and incubated at room temperature for 5 minutes. The DNA-lipid complex is immediately added to the CHO cells by gently pipetting it onto the top of each culture. The cells are incubated at 37°C in the presence of 5% CO2 for 1–2 days. mCherry expression can be visualized and analyzed by fluorescence microscopy or flow cytometry. For selection of stably transfected cells, replace the medium with RPMI medium supplemented with puromycin (1–10 μg / mL) 2–4 days after transfection.

[0131] Example 7 Production of bovine β-lactoglobulin B protein (LGB) in Aspergillus oryzae using a synthetic expression system containing a plant-derived activation domain An expression system containing one exemplary plant-based activation domain, Bn_TAF1M-AD (SEQ ID NO: 11), was constructed and tested in Aspergillus oryzae for the production of an exemplary heterologous protein product secreted into the culture medium. The expression system described in Example 2 (and its schematic shown in Figure 1) containing Bn_TAF1M-AD was modified by replacing the mCherry coding sequence with a DNA sequence encoding bovine β-lactoglobulin B protein (LGB, SEQ ID NO: 29). The DNA encoding LGB was extended with an appropriate secretory signal sequence (SS) containing a Kex2 recognition site added in-frame to its 5' end. This resulted in DNA encoding a fusion protein (SS-Kex2-LGB; target gene in Figure 1). This fusion protein could be efficiently processed and secreted into the culture medium by A. oryzae. This expression system was further modified by providing an A. oryzae-specific selectable marker (SM in Figure 1) and genomic integration DNA regions (shown as EGL1-5' and EGL1-3' in Figure 1) to target the selected A. oryzae genomic locus. The selectable marker was the A. oryzae pyrG gene with appropriate promoter and terminator regions. The genomic integration DNA region was selected to allow integration of the above construct into the A. oryzae gaaC locus (AO090011000868) (https: / / fungi.ensembl.org / ). The gaaC integration flanking regions contained DNA sequences corresponding to DNA regions outside the gaaC coding region in the genome. gaaC-5' spanned a sequence from 600 bp upstream to 15 bp downstream of the start codon. gaaC-3' spanned a sequence from 1 to 600 bp downstream of the stop codon. Another set of genomic integration DNA regions was selected to allow integration of the above construct into the gluC locus of A. oryzae (AO090701000403 (https: / / fungi.ensembl.org / )). The gluC integration flanking regions contained DNA sequences corresponding to DNA regions outside the gluC coding region in the genome. gluC-5' was the sequence 600 to 29 bp upstream of the start codon.gluC-3' was the sequence 1 to 600 bp downstream of the stop codon. Therefore, two LGB expression cassettes were constructed: one targeted to the gaaC locus of A. oryzae and the other to the gluC locus.

[0132] Aspergillus oryzae strain D-171652 (VTT culture collection) was used as the parent strain. This strain was first modified by deleting two genes: the AO090011000868 gene (https: / / fungi.ensembl.org / ), which encodes the orotidine 5'-phosphate decarboxylase (pyrG) enzyme, and the AO090120000322 gene (https: / / fungi.ensembl.org / ), which encodes a homolog of the NHEJ complex subunit (lig4) protein. The resulting strain (referred to herein as A. oryzae pyrGΔ / lig4Δ) is unable to grow in the absence of uracil and is defective in the non-homologous end-joining DNA repair pathway.

[0133] The two LGB expression cassettes were transformed into protoplasts prepared from the A. oryzae pyrGΔ / lig4Δ strain using the PEG transformation protocol. The isolated A. oryzae pyrGΔ / lig4Δ protoplasts were suspended in 400 μL of STC solution (1.33 M sorbitol, 10 mM Tris-HCl, 50 mM CaCl, pH 8.0). For transformation, 100 μL of the protoplast suspension was mixed with 20 μg of an LGB expression construct with gaaC genomic integration flanking regions (a linear fragment corresponding to the construct shown in Figure 1, in which the EGL1-5' and EGL1-3' regions are replaced with the gaaC-5' and gaaC-3' regions) dissolved in 50 μL of solution, 20 μg of an LGB expression construct with gluC genomic integration flanking regions (a linear fragment corresponding to the construct shown in Figure 1, in which the EGL1-5' and EGL1-3' regions are replaced with the gluC-5' and gluC-3' regions) dissolved in 50 μL of solution, and 100 μL of transformation solution (25% PEG 6000, 50 mM CaCl2, 10 mM Tris-HCl, pH 7.5). This mixture was incubated on ice for 20 min. 2 mL of transformation solution was added, and the mixture was incubated at room temperature for 5 min. Four milliliters of STC was added, followed by 7 milliliters of molten (50°C) top agar (200 g / L D-sorbitol, 6.7 g / L Yeast Nitrogen Base (YNB, Becton, Dickinson and Company), synthetic complete amino acids without uracil, and 20 g / L agar). The mixture was poured onto selective plates (200 g / L D-sorbitol, 20 g / L D-glucose, 6.7 g / L Yeast Nitrogen Base (YNB, Becton, Dickinson and Company), synthetic complete amino acids without uracil, and 20 g / L agar). Incubation was continued at 28°C for 4 to 7 days. Colonies were picked and re-plated onto SDC-URA plates (6.7 g / L Yeast Nitrogen Base (YNB, Becton, Dickinson and Company), synthetic complete amino acids without uracil, 20 g / L D-glucose, and 20 g / L agar).

[0134] Transformants were tested by qPCR of genomic DNA isolated from the strains. The qPCR signals of the LGB genes were compared with the qPCR signals of the unique native sequences in each strain. In addition, precise simultaneous deletion of the gaaC and gluC genes was confirmed by the absence of qPCR signals of the gaaC and gluC targets. Four correct selected strains were sporulated on PDA agar plates (39 g / L BD-Difco potato dextrose agar). Spores (conidia) were collected from the PDA plates and used as inoculum in liquid culture for LBG production experiments.

[0135] Four selected clones were tested in small-scale liquid cultures, and analysis of culture supernatants by SDS-PAGE was performed on days 2, 3, and 4 (Figure 13). Four mL of BMG medium (20 g / L glucose, 10 g / L yeast extract, 20 g / L bacto peptone, 13.4 g / L YNB, 0.4 mg / L biotin, and 100 mM KH2PO4, pH = 6.0) in 24-well culture plates was inoculated with conidia collected from PDA plates. Cultures were incubated at 28 °C and 800 rpm (Infors HT Microtron) and centrifuged on the indicated days to partially pellet the mycelium. Fifty microliters of each culture supernatant was mixed with 25 μL of 4x SDS supplement (400 mL / L glycerol; 240 mM Tris·HCl, pH 6.8; 80 g / L SDS; 0.4 g / L bromophenol blue; and 50 mL / L β-mercaptoethanol) and incubated at 95°C for 4 minutes. 15 μL of the mixture was loaded onto a 4-20% SDS-PAGE gradient gel next to a molecular weight standard, commercially available pure β-lactoglobulin B from bovine milk. After complete protein separation using an electric field (PowerPac HC; BioRad), the gel was stained with colloidal Coomassie stain (PageBlue Protein Staining Solution; Thermo Fisher Scientific) according to the manufacturer's protocol. Visualization of the stained gel was performed using an Odyssey CLx Imaging System (LI-COR Biosciences). A scan of the stained gel is shown in Figure 13. In all tested strains, there was clear and consistent production of protein (identical to pure LGB as determined by molecular weight) in the culture supernatant. High-level production of LGB in all four tested clones was achieved using an expression system containing the Bn_TAF1M activation domain. Therefore, it is clear that plant-based activation domains can be successfully used for recombinant protein production in Aspergillus oryzae.

[0136] Example 8 Testing the Transcriptional Activation Domain Bn-TAF1M as Part of a Doxycycline-Regulated Synthetic Expression System in Trichoderma reesei, Pichia pastoris, and Yarrowia lipolytica A reporter expression system for testing doxycycline-dependent expression in Trichoderma reesei was constructed as a single DNA molecule (plasmid) (Figure 1, Table 2A). This plasmid contained the same parts as those described in Example 1, except for the sTF DNA-binding domain and sTF-dependent binding site (Table 2A). Reporter expression systems for testing doxycycline-dependent expression in Pichia pastoris (Table 2B) and Yarrowia lipolytica (Table 2C) were constructed as single DNA molecules (plasmids) (Figure 14).

[0137] In all three expression cassettes, the DNA-binding domain (DBD) was TetR (a transcriptional regulator from Escherichia coli, GenBank: EFK45326.1) extended with an SV40 NLS. The DNA encoding the DBD was codon-optimized for Saccharomyces cerevisiae in the construct used in Pichia pastoris (Table 2B) or for Aspergillus niger in the construct used in Trichoderma reesei (Table 2A) and Yarrowia lipolytica (Table 2C).

[0138] The transcription activation domain (AD) was Bn-TAF1M (SEQ ID NO: 11) in all expression cassettes. The DNA encoding the AD was codon-optimized for Aspergillus niger in the constructs used in Trichoderma reesei and Yarrowia lipolytica (Tables 2A and 2B) or for Pichia pastoris in the construct used in Pichia pastoris (Table 2C).

[0139] The expression cassettes contained eight TetR binding sites (BS; sequences shown in Tables 2A, 2B, and 2C); the Aspergillus niger 201 core promoter (An_201cp; sequences shown in Tables 2A and 2B) or the Yarrowia lipolytica 565 core promoter (Yl_565cp; sequences shown in Table 2C); mCherry-encoding DNA (target gene; sequences shown in Tables 2A, 2B, and 2C); and a target gene cassette consisting of the Trichoderma reesei pdc1 terminator (Tr_PDC1t; Table 2A) or the Saccharomyces cerevisiae ADH1 terminator (Sc_ADH1t; Tables 2B and 2C). These plasmids further contained synthetic transcription factor (sTF) expression cassettes consisting of the Trichoderma reesei hfb2 core promoter (Tr_hfb2cp; sequence shown in Table 2A), or the Aspergillus niger 008 core promoter (An_008cp; Table 2B), or the Yarrowia lipolytica 242 core promoter (Yl_242cp; Table 2C); the sTF coding region; and the Trichoderma reesei tef1 terminator (Tr_TEF1t; Tables 2A, 2B, and 2C).

[0140] The expression cassette for Pichia pastoris also contained a selectable marker that allowed expression of the kanR gene and flanking DNA regions for genomic integration to target the ADE1 gene. The expression cassette for Yarrowia lipolytica also contained a selectable marker that allowed expression of the NAT gene and flanking DNA regions for genomic integration to target the ant1 gene.

[0141] Trichoderma reesei strain M1909 (VTT culture collection), Pichia pastoris strain Y-11430, and Yarrowia lipolytica strain C-00365 (VTT culture collection) were used as parent strains. The expression system (Figure 1, Table 2A) was transformed into T. reesei using the PEG transformation protocol (described in Example 5). The expression system (Figure 14, Tables 2B and 2C) was transformed into P. pastoris or Y. lipolytica, respectively, using the lithium acetate protocol (described in Example 4). T. reesei transformants were selected for growth on medium lacking uracil, P. pastoris transformants were selected on medium containing 500 mg / L G418, and Y. lipolytica transformants were selected on medium containing 150 mg / L nourseothricin.

[0142] Three randomly selected colonies from each transformation were analyzed for mCherry fluorescence in liquid culture in the absence and presence of 1 mg / L or 3 mg / L doxycycline (DOX) (Figure 15).

[0143] For quantitative fluorometric analysis of mCherry production in mycelia of T. reesei strains or cells of P. pastoris and Y. lipolytica strains (Figure 15), 4 mL of BMG medium (20 g / L glucose, 10 g / L yeast extract, 20 g / L bacto peptone, 13.4 g / L YNB, 0.4 mg / L biotin, and 100 mM KH2PO4, pH = 6.0) without doxycycline or containing 1 mg / L or 3 mg / L doxycycline in 24-well culture plates was inoculated with spores / cells of selected clones to an OD600 of 0.1. The cultures were grown at 800 rpm (Infors HT Microtron) and 28 °C for 24 h, centrifuged, and the pellets were washed with water and resuspended in 0.5 mL of sterile water. 200 μL of each mycelium / cell suspension was analyzed in a black 96-well plate (Black Cliniplate; Thermo Scientific) using a Varioskan (Thermo Electron Corporation) fluorometer. The mCherry settings were 587 nm (excitation) and 610 nm (emission), respectively. To normalize the fluorescence results, the analyzed mycelium / cell suspension was diluted 100-fold and the OD600 was measured using a Varioskan (Thermo Electron Corporation) in a clear 96-well microtiter plate (NUNC). The results of the analysis are shown in Figure 15. These results clearly demonstrate that the selected plant-based activation domains can be successfully used in a doxycycline-dependent expression system (TET-OFF) for the controlled expression of heterologous genes in diverse fungal species.

[0144] Example 9. Development of a synthetic expression system based on plant-derived activation domains for high-level gene expression in Yarrowia lipolytica and Kutaneotrichosporon oleaginosus Microbial lipid production is becoming an increasingly attractive topic in biotechnology, including food applications. Several promising production hosts have been identified, and some of them have been established in bioprocesses for the production of diverse lipid compounds. However, further development of production hosts is often hindered by the limited amount of robust gene expression tools available for genetic manipulation, such as heterologous gene expression. A synthetic expression system based on sTF-containing plant-derived activation domains was tested and optimized for two yeast species known for high-level lipid production: Yarrowia lipolytica and Cutaneotrichosporon oleaginosus.

[0145] Bn_TAF1M, one of the best-performing plant-based activation domains identified and extensively tested in previous examples, was selected as the activation domain for the development of expression systems for Yarrowia lipolytica and Cutaneotrichosporon oleaginosus. This expression system was constructed as a single DNA molecule (Figure 14), with the DBD being Bm3R1 and the target gene being the reporter mCherry. The terminators used in this cassette were the S. cerevisiae ADH1 terminator (term 1 in Figure 14) and the T. reesei tef1 terminator (term 2 in Figure 14). This construct also contained a selectable marker (SM in Figure 14) that enabled expression of the NAT gene and flanking DNA regions for genome integration (5' and 3' in Figure 14) for targeting the ant1 gene in Y. lipolytica. A control expression system containing a viral-based VP16 activation domain instead of Bn_TAF1M-AD shown in Figure 14 was also constructed and tested.

[0146] For Y. lipolytica, the expression systems (Figures 14 and 16) contained different combinations of two core promoters (cp): one upstream of the target gene (cp1 in the target gene cassette in Figure 14) and the other upstream of the sTF (cp2 in the sTF cassette in Figure 14). The following cp1-core promoters were tested: An_201cp (SEQ ID NO: 23), Yl_205cp (SEQ ID NO: 34), Yl_565cp (SEQ ID NO: 32), Yl_137cp (SEQ ID NO: 36), Yl_113cp (SEQ ID NO: 37), and Yl_697cp (SEQ ID NO: 38). The following cp2-core promoters were tested: An_008cp (SEQ ID NO: 22), Yl_TEF1cp (SEQ ID NO: 35), Yl_242cp (SEQ ID NO: 33), and Cc_MFScp (SEQ ID NO: 40). Bm3R1 (DBD in Figure 14) was codon-optimized for Aspergillus niger.

[0147] For C. oleaginosus, the expression system (Figures 14 and 16) contained different combinations of two core promoters (cp): one upstream of the target gene (cp1 in the target gene cassette in Figure 14) and the other upstream of the sTF (cp2 in the sTF cassette in Figure 14). The following cp1-core promoters were tested: An_201cp (SEQ ID NO: 23), Cc_RAScp (SEQ ID NO: 39), Cc_GSTcp (SEQ ID NO: 42), Cc_AKRcp (SEQ ID NO: 43), and Cc_FbPcp (SEQ ID NO: 44). The following cp2-core promoters were tested: An_008cp (SEQ ID NO: 22), Cc_HSP9cp (SEQ ID NO: 41), and Cc_MFScp (SEQ ID NO: 40). Bm3R1 (DBD in Figure 14) was codon-optimized for C. oleaginosus. The DNA sequences of exemplary expression systems containing Cc_FbPcp and Cc_MFScp are shown in Table 2D.

[0148] Yarrowia lipolytica strain C-00365 (VTT culture collection) and Cutaneotrichosporon oleaginosus (formerly known as Trichosporon oleaginosus, Cryptococcus curvatus, Apiotrichum curvatum, or Candida curvata) strain ATCC20509 were used as parent strains. This expression system was transformed into Y. lipolytica using the lithium acetate protocol (described in Example 4). This expression system was transformed into C. oleaginosus by electroporation (the following protocol is for one transformation): a 20 mL liquid culture grown in YPD to an OD of ∼1.0 was centrifuged briefly (4000 rpm / 1 min) to pellet the cells. The cells were washed with 10 mL of ice-cold sterile EB solution (10 mM Tris, pH 7.5; 270 mM sucrose; 1 mM MgCl2) and resuspended in 5 mL of IB solution (25 mM DTT; 20 mM HEPES, pH 8.0; in YPD). The cell suspension was incubated at 30°C with shaking at 22 rpm for 30 minutes, followed by a brief centrifugation (4000 rpm / 1 minute) to pellet the cells. The cells were washed with 20 mL of EB solution, and the cell pellet after centrifugation (4000 rpm / 1 minute) was resuspended in 500 μL of EB solution to prepare transformation-competent cells. 400 μL of this cell suspension was mixed with 5–10 μg of DNA (expression system DNA cassette) in an electroporation cuvette (4 mm gap) and incubated on ice for 15 minutes. Two sequential electroporations were performed (BioRad GenePulser; 1800V; 1000Ω; 25 μF). The transformation mixture was diluted with 1 mL of YPD and incubated at 30°C with shaking at 220 rpm for 4 hours, after which the cells were spread onto selective agar plates.

[0149] Transformants of Y. lipolytica and C. oleaginosus were selected for growth on medium containing 150 mg / L nourseothricin (YPD agar). Three colonies from each transformation were analyzed for mCherry fluorescence in liquid culture.

[0150] For quantitative fluorometric analysis of mCherry production in P. pastoris cells (Figure 16), 4 mL of YPD medium in a 24-well culture plate was inoculated with cells of the selected clones to an OD of 0.1. The cultures were grown at 800 rpm (Infors HT Microtron) and 28°C for 24 hours, centrifuged, and the pellets were washed with water and resuspended in 0.5 mL of sterile water. 200 μL of each cell suspension was analyzed in a black 96-well plate (Black Cliniplate; Thermo Scientific) using a Varioskan (Thermo Electron Corporation) fluorometer. The mCherry settings were 587 nm (excitation) and 610 nm (emission), respectively. For normalization of fluorescence results, the analyzed cell suspensions were diluted 100-fold, and the OD was measured using a Varioskan (Thermo Electron Corporation) in a clear 96-well microtiter plate (NUNC). The results of the analysis are shown in Figure 16. These results clearly demonstrate that selected plant-based (e.g., edible plant-based) activation domains can be successfully used in place of the virus-based VP16 AD for high-level expression of heterologous genes in Y. lipolytica and C. oleaginosus. A control system with VP16 AD was also tested in C. oleaginosus, but no fluorescence was detected in the transformed cells (data not shown). However, the lack of mCherry expression could be due to the nonfunctional core promoters An_201cp and An_008 rather than the nonfunctional VP16 AD in C. oleaginosus.

[0151] [Table 2(1)] [Table 2(2)] [Table 2(3)] [Table 2(4)]

[0152] References Chavez A et al. (2015). “Highly efficient Cas9-mediated transcriptional programming”, Nat Methods, 12(4), 326-328. Lu, Y. et al. (2016). "High-level expression of improved thermo-stable alkaline xylanase variant in Pichia Pastoris through codon optimization, multiple gene insertion and high-density fermentation", Scientific Reports, Vol. 6, Article No.: 37869 Naseri G et al. (2017). "Plant-derived transcription factors for orthologous regulation of gene expression in the yeast Saccharomyces cerevisiae", ACS Synthetic Biology, 6, 1742-1756. Olsen, AN, HAErnst et al. (2005). “NAC transcription factors: structurally distinct, functionally diverse”, Trends Plant Sci 10(2):79-87. Tiwari,S.B.、A.Belachewら(2012). 「The EDLL motif: a potent plant transcriptional activation domain from AP2 / ERF transcription factors」、The Plant Journal 70(5):855-865. Zhang,J.ら(2016). 「Site-directed mutagenesis and thermal stability analysis of phytase from Escherichia coli」、Biosci.Biotech.Res.Comm. 9(3):357-365.

Claims

1. 1. A non-viral transcription activation domain for artificial expression systems in eukaryotic hosts, the transcription activation domain being derived from a transcription factor found in a plant species, The transcription activation domain comprises an amino acid sequence having 90 to 100% sequence identity with SEQ ID NO: 10 or 11.

2. The transcription activation domain of claim 1, wherein the transcription activation domain is obtained by rational mutagenesis of a polynucleotide encoding the transcription activation domain.

3. 3. The transcription activation domain of claim 1 or 2, wherein the transcription activation domain is a recombinant or synthetic transcription activation domain.

4. 3. The transcription activation domain of claim 1 or 2, wherein the transcription activation domain is used in the construction of an artificial transcription factor.

5. The transcription activation domain of claim 1 or 2, wherein the transcription activation domain is functional across multiple species.

6. A polypeptide comprising the transcription activation domain of claim 1 or 2.

7. 3. An artificial transcription factor comprising the transcription activation domain of claim 1 or 2, a DNA binding domain, and a nuclear localization signal.

8. A polynucleotide encoding a transcription activation domain, polypeptide or artificial transcription factor according to any one of claims 1 to 7.

9. An expression cassette or expression system comprising a polynucleotide encoding a transcription activation domain, polypeptide or artificial transcription factor according to any one of claims 1 to 8.

10. 10. The expression cassette of claim 9, wherein the expression cassette further comprises a polynucleotide sequence encoding a desired product.

11. 10. The expression system of claim 9, wherein the expression system comprises one or more expression cassettes, optionally at least one expression cassette further comprising a polynucleotide sequence encoding a desired product.

12. A polypeptide, artificial transcription factor, polynucleotide, expression cassette or expression system according to any one of claims 1 to 11 for a eukaryotic host.

13. A eukaryotic host comprising a transcription activation domain, polypeptide, artificial transcription factor, polynucleotide, expression cassette or expression system according to any one of claims 1 to 12.

14. The eukaryotic host is selected from the group consisting of cells of fungal species, including yeast and filamentous fungi, and cells of animal species, including non-human mammals, or selected from the group consisting of Trichoderma, Trichoderma reesei, Pichia, Pichia pastoris, Pichia kudriabzevi, Aspergillus, Aspergillus niger, Aspergillus oryzae, Myceliophthora, Myceliophthora thermophila, Saccharomyces, Saccharomyces cerevisiae, Yarrowia, Yarrowia lipolytica, Kut 14. The transcription activation domain, polypeptide, artificial transcription factor, polynucleotide, expression cassette or expression system according to any one of claims 1 to 13, or a eukaryotic host selected from the group consisting of Neotrichosporon, Cutaneotrichosporon oleaginosus (Trichosporon oleaginosus, Cryptococcus curbatus), Zygosaccharomyces, Chinese hamster ovary (CHO) cells, and Chinese hamster ovary (CHO) cells.

15. 15. A method for producing a desired protein product in a eukaryotic host, comprising culturing a host according to claim 13 or claim 14 under suitable culture conditions.

16. 16. Use of a transcription activation domain, polypeptide, artificial transcription factor, polynucleotide, expression cassette, expression system or eukaryotic host according to any one of claims 1 to 15 for metabolic engineering and / or production of a desired protein product.

17. A method for preparing a non-viral transcription activation domain or a polynucleotide encoding said non-viral transcription activation domain described in any one of claims 1 to 5, comprising the steps of obtaining a transcription activation domain polypeptide derived from a plant transcription factor or obtaining a polynucleotide encoding said transcription activation domain polypeptide derived from a plant transcription factor, and modifying the obtained transcription activation domain polypeptide or polynucleotide.

Citation Information

Patent Citations

  • Expression Systems for Eukaryotes

    JP2019505230A