Means and methods for increasing protein expression using transcription factors
Overexpressing transcription factors like Msn4/2 in eukaryotic host cells enhances recombinant protein yield and secretion by addressing transcription and translation limitations, achieving improved protein production in host cells.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- BOEHRINGER INGELHEIM RCV GMBH & CO KG
- Filing Date
- 2024-11-08
- Publication Date
- 2026-07-06
AI Technical Summary
Existing methods struggle to achieve high yields of recombinant proteins in eukaryotic host cells due to limitations in transcription and translation processes, as well as challenges in protein folding, secretion, and cellular fitness, despite the use of strong promoters and increased gene copy numbers.
Overexpressing at least one polynucleotide encoding a transcription factor, such as Msn4/2, in eukaryotic host cells to enhance protein yield, utilizing a DNA binding domain and an activation domain, with sequences having at least 60% sequence identity to SEQ ID NO: 1 or 87, and optionally incorporating endoplasmic reticulum accessory proteins.
Significantly increases the yield and titer of recombinant proteins, particularly in fungal host cells like Pichia pastoris, by improving protein production and secretion, while maintaining cellular fitness.
Smart Images

Figure 0007885305000012 
Figure 0007885305000013 
Figure 0007885305000014
Abstract
Description
[Technical Field]
[0001] Cross-reference of related applications This application claims priority to European Patent Application No. 18180164.8, filed on 27 June 2018, the contents of which are hereby incorporated by reference in their entirety for all purposes.
[0002] Field of Invention The present invention relates to the field of recombinant biotechnology, particularly the field of protein expression. The present invention generally relates to a method for increasing the yield of a protein of interest (POI) in eukaryotic host cells, preferably yeast, by overexpressing at least one transcription factor, preferably at least one polynucleotide encoding Msn4 / 2. The present invention further relates to recombinant eukaryotic host cells for producing POI (the host cells hereby engineered to overexpress at least one polynucleotide encoding at least one transcription factor), and to the use of host cells for producing POI.
[0003] Background of the Invention Successful production of point of interest (POIs) can be achieved using either prokaryotic or eukaryotic cell hosts. Most notable examples include bacteria such as Escherichia coli, yeasts such as Saccharomyces cerevisiae, Pichia pastoris, or Hansenula polymorpha, filamentous fungi such as Aspergillus awamori or Trichoderma reesei, or mammalian cells such as CHO cells. While yields of some proteins are easily achieved at high ratios, many other proteins are produced only at relatively low levels.
[0004] Generally, the synthesis of heterologous proteins can be restricted at various levels. Possible restrictions include transcription and translation, protein folding and, where applicable, secretion, disulfide bridge formation and glycosylation, as well as aggregation and degradation of target proteins. Transcription can be enhanced by utilizing strong promoters or by increasing the copy number of heterologous genes. However, these measures clearly reach a plateau, indicating that other obstacles downstream of transcription limit expression.
[0005] High levels of protein yield in host cells can also be limited by one or more different processes, such as folding, disulfide bond formation, glycosylation, intracellular transport, or release from the cell. Many of the mechanisms involved are still not fully understood and cannot be predicted based on current state-of-the-art technology, even when the DNA sequence of the entire host organism's genome is available. Furthermore, the phenotype of cells producing recombinant proteins in high yield may include reduced growth rate, decreased biomass formation, and overall reduced cellular fitness.
[0006] In this field, various attempts have been made to improve the production of proteins of interest, such as the overexpression of chaperones that promote protein folding and the supplementation of amino acids from external sources.
[0007] However, there is still a need for methods to improve the host cell's ability to produce and / or secrete proteins of interest. The technical problem underlying this invention is to address this need.
[0008] The solution to the technical problems is to provide means such as engineered host cells, methods for applying means to increase the yield of recombinant protein of interest in a host cell by overexpressing at least one polynucleotide encoding at least one transcription factor in a eukaryotic host cell, and uses for such application. These means, methods, and uses are described in detail herein, explained in the claims, illustrated in the examples, and illustrated in the drawings.
[0009] Therefore, the present invention provides a novel method and use for increasing the yield of recombinant proteins in host cells, which is simple, efficient, and suitable for use in industrial methods. The present invention also provides host cells for achieving this objective.
[0010] In this specification, the singular forms “a,” “an,” and “the” include plural nouns unless the context clearly indicates otherwise, and vice versa. Therefore, for example, a reference to “one host cell” or “one method” includes one or more such host cells or methods, and a reference to “the method” includes equivalent steps and methods that are known to those skilled in the art, are modifiable, or can be replaced. Similarly, for example, a reference to “methods” or “host cells” includes “one host cell” or “one method,” respectively.
[0011] Unless otherwise indicated, the term “at least” preceding a set of elements should be understood to refer to all elements within that set. Those skilled in the art will be able to identify or verify, by mere ordinary experimentation, many equivalents to the specific embodiments of the invention described herein. Such equivalents shall be encompassed by the invention.
[0012] Whenever the term "and / or" is used herein, it includes the meanings of "and," "or," and "all elements or any other combination of elements connected by the term." For example, A, B and / or C means A, B, C, A+B, A+C, B+C, and A+B+C.
[0013] As used herein, the terms “about” or “approximately” mean within 20%, preferably within 10%, and more preferably within 5% of a given number or range. It also includes specific numbers; for example, “about 20” includes 20.
[0014] The terms "less than," "more than," and "greater than" include specific numbers. For example, "less than 20" means 20 or less, and "more than 20" means 20 or more.
[0015] Throughout this specification and any claim or item, unless the context requires otherwise, the word “comprises” and its variations, e.g., “comprises” and “comprising,” will be understood to mean including the integer (or process) or set of integers (or sets of processes) described. It does not exclude any other integer (or process) or set of integers (or sets of processes). Where used herein, the term “comprises” may be replaced with “contains,” “composed of,” “including,” “has,” or “possesses,” and vice versa; for example, the term “has” may be replaced with the term “comprising.” Where used herein, “consists of” excludes any integer or process not explicitly stated in the claim / item. Where used herein, “essentially consists of” does not exclude a set of integers or processes that do not substantially affect the basic and novel features of the claim / item.
[0016] Furthermore, in describing each embodiment of the present invention, this specification may present a specific sequence of steps for the method and / or process of the present invention. However, to the extent that the method or process does not rely on the specific sequence of steps shown herein, the method or process should not be limited to the specific sequence of steps described herein. Other sequences of steps may be possible, as those skilled in the art will recognize. Therefore, the specific sequence of steps shown herein should not be considered a limitation on the claims. Furthermore, claims directed to the method and / or process of the present invention should not be limited to the execution of those steps in the order described, and those skilled in the art will readily recognize that the order may be changed and still remain within the spirit and scope of the invention.
[0017] It should be understood that the present invention is not limited to the specific methods, protocols, materials, reagents, and substances described herein. The terms used herein are for the sole purpose of describing specific embodiments and are not intended to limit the scope of the invention, and the scope of the invention is defined solely by the claims / subjects.
[0018] All publications and patents cited throughout this specification, whether above or below (including all patents, patent applications, scientific literature, manufacturer specifications, instructions, etc.), are incorporated by reference in their entirety. Nothing in this specification should be construed as acknowledging that the present invention is not entitled to precede such disclosure for the sake of prior art. This specification will be substituted by any such material to the extent that the material incorporated by reference is inconsistent or contradictory to this specification.
[0019] summary The inventors' findings are quite surprising, because, to the best of our knowledge, the transcription factors of the present invention have never been associated with increasing the yield of the protein of interest in eukaryotic host cells, particularly in fungal host cells.
[0020] The present invention includes a method for increasing the yield of a recombinant protein of interest in a eukaryotic host cell, the method comprising overexpressing in the eukaryotic host cell at least one polynucleotide encoding at least one transcription factor, whereby increasing the yield of the recombinant protein of interest compared to a host cell that does not overexpress the polynucleotide encoding the transcription factor, wherein the transcription factor comprises at least: a) a DNA binding domain which comprises i) an amino acid sequence as set forth in SEQ ID NO: 1, or ii) a functional homolog of the amino acid sequence as set forth in SEQ ID NO: 1 having at least 60% sequence identity to the amino acid sequence as set forth in SEQ ID NO: 1 and / or having at least 60% sequence identity to the amino acid sequence as set forth in SEQ ID NO: 87, and b) an activation domain.
[0021] The method of the present invention i) at least: a) a DNA binding domain which a1) an amino acid sequence as set forth in SEQ ID NO: 1, or a2) a functional homolog of the amino acid sequence as set forth in SEQ ID NO: 1 having at least 60% sequence identity to the amino acid sequence as set forth in SEQ ID NO: 1 and / or having at least 60% sequence identity to the amino acid sequence as set forth in SEQ ID NO: 87, and b) an activation domain and engineering the host cell to overexpress at least one polynucleotide encoding at least one transcription factor comprising the same, ii) engineering the host cell to contain a polynucleotide encoding the protein of interest, iii) overexpressing at least one polynucleotide encoding at least one transcription factor and culturing the host cell under conditions suitable for overexpressing the protein of interest, optionally iv) isolating the protein of interest from the cell culture, and optionally v) The process of purifying the protein of interest. It may include.
[0022] Furthermore, the present invention is i) The step of preparing host cells that have been engineered to overexpress at least one polynucleotide encoding at least one transcription factor. (The host cell here further contains polynucleotides encoding the protein of interest, and the transcription factors here are at least: a) DNA binding domain (this is, a1) An amino acid sequence as shown in Sequence ID No. 1, or a2) Including a functional homolog of the amino acid sequence shown in SEQ ID NO: 1, which has at least 60% sequence identity with the amino acid sequence shown in SEQ ID NO: 1, and / or has at least 60% sequence identity with the amino acid sequence shown in SEQ ID NO: 87, and b) Activation domain (including), ii) A step of overexpressing at least one polynucleotide encoding at least one transcription factor and culturing the host cells under conditions suitable for overexpressing the protein of interest, optionally iii) A step of isolating the protein of interest from the cell culture medium, and, if applicable iv) A step to purify the protein of interest, and, if applicable, v) A step of modifying the protein of interest, and, if applicable vi) The process of formulating the protein of interest. We envision a method for producing recombinant proteins of interest using eukaryotic host cells, including [specific example].
[0023] The method of the present invention may include the overexpression of the transcription factor increasing the yield of the model proteins scFv (SEQ ID NO: 13) and / or vHH (SEQ ID NO: 14) compared to host cells before engineering.
[0024] Furthermore, the present invention may include a method in which a polynucleotide encoding at least one transcription factor is incorporated into the genome of the host cell or contained in a vector or plasmid that is not incorporated into the genome of the host cell.
[0025] The present invention relates to eukaryotic host cells, and fungal host cells, preferably Pichia pastoris (synonym Komagataella), Hansenula polymorpha (synonym H. angusta), Trichoderma reesei, Aspergillus niger, Saccharomyces cerevisiae, Kluyveromyces lactis, Yarrowia lipolytica, Pichia methanolica, Candida boidinii, Komagataella, and Schizosaccharomyces pombe. The present invention may encompass yeast host cells selected from the group consisting of pombe. Hanzenula polymorpha has been reclassified into the genus Ogataea (Yamada et al. 1994. Biosci Biotechnol Biochem. 58(7):1245-57). Ogataea angusta, Ogataea polymorpha, and Ogataea parapolymorpha are closely related species that have recently been separated from each other (Kurtzman et al. 2011. Antonie Van Leeuwenhoek. 100(3):455-62).
[0026] The present invention may envision a method in which the recombinant protein of interest is an enzyme, a therapeutic protein, a food additive, or a feed additive.
[0027] Furthermore, the present invention may include a method of the present invention that further comprises the steps of overexpressing in a host cell at least one polynucleotide encoding at least one endoplasmic reticulum accessory protein, or engineering the host cell to overexpress it.
[0028] Preferably, the endoplasmic reticulum accessory protein has an amino acid sequence as shown in SEQ ID NO: 28, or a functional homolog thereof having at least 70% sequence identity to the amino acid sequence shown in SEQ ID NO: 28.
[0029] The present invention may further include a method comprising the steps of overexpressing at least two polynucleotides encoding at least two endoplasmic reticulum accessory proteins in a host cell, or engineering the host cell to overexpress them.
[0030] Preferably, the first endoplasmic reticulum accessory protein has an amino acid sequence as shown in SEQ ID NO: 28, or a functional homolog thereof having at least 70% sequence identity to the amino acid sequence shown in SEQ ID NO: 28, and the second endoplasmic reticulum accessory protein is i) an amino acid sequence as shown in SEQ ID NO: 37, or a functional homolog thereof having at least 25% sequence identity with respect to the amino acid sequence shown in SEQ ID NO: 37, ii) An amino acid sequence as shown in SEQ ID NO: 47, or its homologue (wherein a homologue has at least 20% sequence identity with the amino acid sequence shown in SEQ ID NO: 47). It may have. In some cases, the third endoplasmic reticulum accessory protein may have an amino acid sequence as shown in SEQ ID NO: 55, or a functional homolog thereof having at least 25% sequence identity to the amino acid sequence shown in SEQ ID NO: 55.
[0031] Furthermore, the present invention may include a method of the present invention that further comprises the steps of overexpressing in a host cell at least one polynucleotide encoding one additional transcription factor, or engineering the host cell to overexpress it.
[0032] Preferably, additional transcription factors include at least: a) DNA binding domain (this is, i) an amino acid sequence as shown in Sequence ID No. 65, or ii) including functional homologs of the amino acid sequence shown in SEQ ID NO: 65 that have at least 50% sequence identity with the amino acid sequence shown in SEQ ID NO: 65, and b) Activation domain Includes.
[0033] The present invention also comprises recombinant eukaryotic host cells for producing a protein of interest, wherein the host cells are engineered to overexpress at least one polynucleotide encoding at least one transcription factor, and the transcription factor is at least a) DNA binding domain (this is, i) an amino acid sequence as shown in Sequence ID No. 1, or ii) including a functional homolog of the amino acid sequence shown in SEQ ID NO: 1, which has at least 60% sequence identity with the amino acid sequence shown in SEQ ID NO: 1, and / or has at least 60% sequence identity with the amino acid sequence shown in SEQ ID NO: 87, and b) Activation domain Includes.
[0034] Another possibility associated with the present invention is the use of recombinant eukaryotic host cells, as described above, for producing recombinant proteins of interest. [Brief explanation of the drawing]
[0035] [Figure 1]Improvement of vHH secretion (titer and yield) in small-scale screening cultures. An overview of overexpressed genes or gene combinations that increase vHH secretion in Pichia pastris in small-scale screening. Plasmids or groups of plasmids used to overexpress these genes or gene combinations by engineering host cells are shown below the gene or gene combination in parentheses. The change factor values for small-scale screening are the arithmetic mean of 20 or fewer clones / transformers. [Figure 2] Improvement of vHH secretion (titer and yield) in fed-batch bioreactor cultures. An overview of overexpressed genes or gene combinations that increase vHH secretion in Pichia pastris in fed-batch cultures. Plasmids or groups of plasmids used to overexpress these genes or gene combinations by engineering host cells are shown below the gene or gene combination in parentheses. The numerical values for the change in fed-batch cultures are for one selected clone. [Figure 3] Improvement of scFv secretion (titer and yield) in small-scale screening cultures. An overview of overexpressed genes or gene combinations that increase scFv secretion in Pichia pastris in small-scale screening. Plasmids or groups of plasmids used to overexpress these genes or gene combinations by engineering host cells are shown below the gene or gene combination in parentheses. The change factor values for small-scale screening are the arithmetic mean of 20 or fewer clones / transformers. [Figure 4] Improvement of scFv secretion (titer and yield) in fed-batch bioreactor cultures. An overview of overexpressed genes or gene combinations that increase scFv secretion in Pichia pastris in fed-batch cultures. Plasmids or groups of plasmids used to overexpress these genes or gene combinations by engineering host cells are shown below the gene or gene combination in parentheses. The numerical values for the change in fed-batch cultures are for one selected clone. [Figure 5]Improvement of scFv secretion (potency and yield) by overexpression of MSN2 / 4 homologs derived from other species in fed-batch bioreactor cultures. [Figure 6] An overview of the alignment of Msn4p transcription factors from different origins. The zinc finger of the protein structural motif clearly shows strong conservation (box in Figure 6), which is known as the DNA-binding domain of the Msn4p and Msn2p (ScMsn4 / 2) transcription factors well-characterized in Saccharomyces cerevisiae. [Figure 7-1] Common amino acid sequence of the Msn4-like C2H2 zinc finger DNA-binding domain. [Figure 7-2] Common amino acid sequence of the Msn4-like C2H2 zinc finger DNA-binding domain. [Figure 8] Sequence alignment of MSN4 / 2 in Pichia pastris. Pairwise sequence similarity / identity between the full-length Msn4p of Pichia pastris and its homologs in other organisms was evaluated using global pairwise sequence alignment with an embossed needle algorithm. Pairwise sequence similarity / identity of the DNA-binding domain of Pichia pastris Msn4p and the DNA-binding domains of its homologs in other organisms were also investigated. [Figure 9] Sequence identity rate of Pichia pastris to KAR2. Sequence identity rate was evaluated using BLASTp. [Figure 10] Sequence identity rate of Pichia pastris to LHS1. Sequence identity rate was evaluated using BLASTp. [Figure 11] Sequence identity rate of Pichia pastris to SIL1. Sequence identity rate was evaluated using BLASTp. [Figure 12] Sequence identity rate of Pichia pastris to ERJ5. Sequence identity rate was evaluated using BLASTp. [Figure 13]Sequence alignment of HAC1 in Pichia pastris. Pairwise sequence similarity / identity between the full-length Hac1p of Pichia pastris and its homologs in other organisms was evaluated using global pairwise sequence alignment with an embossed needle algorithm. Pairwise sequence similarity / identity of the DNA-binding domain of Hac1p in Pichia pastris and the DNA-binding domains of its homologs in other organisms were also investigated. [Figure 14] Sequence similarity rate of the DNA-binding domain of MSN4 / 2 relative to the common sequence. The pairwise sequence similarity / identity rate between the common sequence of the DNA-binding domain (DBD) of Msn4p / Msn2p and the DNA-binding domains of each homolog of other organisms was investigated using global pairwise sequence alignment with an embossed needle algorithm.
[0036] Detailed description of the invention The present invention is partly based on the surprising finding that overexpression of at least one transcription factor, as described herein, has been found to increase the yield of the recombinant protein of interest. In particular, the present invention includes a method for increasing the yield of the recombinant protein of interest in a eukaryotic host cell, comprising the step of overexpressing at least one polynucleotide encoding at least one transcription factor of the present invention in the eukaryotic host cell, thereby increasing the yield of the recombinant protein of interest compared to host cells that do not overexpress the polynucleotide encoding the transcription factor.
[0037] The term "increasing the yield of recombinant protein of interest in host cells" means that the yield of the protein of interest is increased compared to cells expressing the same protein of interest (POI) under the same culture conditions, but without overexpression of the transcription factor-encoding polynucleotide, or without being engineered to overexpress the transcription factor-encoding polynucleotide.
[0038] In this context, the term “yield” refers to the amount of the protein of interest or model protein(s) described herein, particularly scFv, i.e., single-strand variable fragment (SEQ ID NO: 13), and vHH (or VHHV), single-domain antibody fragment (SEQ ID NO: 14), respectively, collected from engineered host cells, where the increased yield may result from increased production within the host cell or increased secretion of the protein of interest by the host cell. The term “yield” also refers to the amount of the protein of interest or model protein(s) described herein per cell, and may be expressed in mg of the protein of interest per gram of host cell biomass (measured as dry or wet cell weight). The term “titer” as used herein similarly refers to the amount of the protein of interest or model protein produced, expressed as mg of the protein of interest per liter of culture supernatant or whole cell broth. The present invention may also include a method for increasing the titer of a recombinant protein of interest, in which the transcription factor of the present invention is overexpressed in eukaryotic host cells. An increase in yield can be determined by comparing the yield obtained from engineered host cells with the yield obtained from unengineered host cells, i.e., from unengineered host cells. Preferably, as used herein in the context of model proteins as described herein, “yield” is determined as described in Examples 3, 4, and 5. For example, the term “yield” may refer to the amount of protein of interest produced by a certain amount of biomass over the entire immersion culture period. Among these recombinant protein of interest, it may be produced and accumulated intracellularly or secreted into the culture supernatant. The term “increasing the yield of recombinant protein of interest in host cells” refers to increasing the amount of protein of interest produced intracellularly or by cells, and / or increasing the amount of protein of interest secreted by cells.
[0039] As will be recognized by those skilled in the art, overexpression of the transcription factors of the present invention has been shown to increase not only the yield of the protein of interest, particularly recombinant protein of interest, but also its titer.
[0040] As used herein, the term “protein of interest” (POI) generally refers to any protein, but preferably to a “heterogeneous protein” or “recombinant protein,” and preferably to the model proteins scFv (SEQ ID NO: 13) and / or vHH (SEQ ID NO: 14). Specific examples of the POI of the present invention are shown elsewhere herein. As used herein, “recombinant” refers to the modification of genetic material by human intervention. Typically, recombinant refers to the manipulation of DNA or RNA in viruses, cells, plasmids, or vectors by molecular biological (recombinant DNA technology) methods, including cloning and recombination. Recombinant proteins may typically be described in terms of how they differ from their natural counterparts ("wild-type"). Preferably, the recombinant POI of interest expressed by the eukaryotic host cells of the present invention originates from a different organism. The POI of interest is preferably not a transcription factor; i.e., the transcription factor and the POI of interest are not identical. The recombinant protein may also be an homogeneous protein. In this case, one or more copies of the polynucleotide encoding the homogeneous protein are introduced into the host cell by genetic engineering.
[0041] The term "expressing a polynucleotide" means that the polynucleotide is transcribed into mRNA, and that mRNA is translated into a polypeptide. The term "overexpressing" generally refers to any amount higher than the expression level indicated by a reference standard (e.g., the same host cells under the same culture conditions that have not been engineered to overexpress a protein-coding polynucleotide). In this invention, the terms "overexpressing," "overexpressed," and "overexpressed" refer to the expression of a gene product or polypeptide at a level higher than the expression of the gene product or polypeptide in a host cell before genetic modification, or in an equivalent host that has not been genetically modified under specified conditions. In this invention, a transcription factor comprising an amino acid sequence or a functional homolog thereof, such as that shown in any one of SEQ ID NOs: 15-27, is overexpressed. If a host cell does not contain a given gene product, it is possible to introduce the gene product into the host cell for expression; in this case, any detectable expression is encompassed by the term "overexpression." In preferred embodiments, "overexpressing" means "engineered to overexpress" as described below. Such preferred embodiments are possible for any embodiment relating to “overexpression” or “being overexpressed” as described herein.
[0042] As used herein, “polynucleotide” refers to a polymeric, unbranched nucleotide, ribonucleotide, or deoxyribonucleotide of any length, or a combination thereof. Preferably, polynucleotide refers to a polymeric, unbranched deoxyribonucleotide of any length. Here, a nucleotide consists of a pentose sugar (deoxyribose), a nitrogen-containing base (adenine, guanine, cytosine, or thymine), and a phosphate group. The terms “polynucleotide (group)” and “nucleic acid sequence (group)” are used synonymously herein.
[0043] As used herein, the term "at least one polynucleotide encoding at least one transcription factor" refers to one polynucleotide encoding one transcription factor, two polynucleotides encoding two transcription factors, three polynucleotides encoding three transcription factors, four polynucleotides encoding four transcription factors, and so on. Preferably, the present invention includes one polynucleotide encoding one transcription factor. More preferably, the present invention includes one polynucleotide encoding one transcription factor and one polynucleotide encoding one additional transcription factor.
[0044] The term "transcription factor" refers to a protein that controls the rate of transcription of genetic information from DNA to messenger RNA by binding to a specific DNA sequence, preferably with its DNA-binding domain. Their function is to regulate and / or activate genes to ensure that genes are expressed in the correct cells, at the correct time, and in the correct amount. For example, transcription factors can initiate transcription of a specific gene(s) in response to stimuli such as starvation or heat shock. In the present invention, the Msn4p transcription factor refers to a transcription factor comprising SEQ ID NOs: 15-27, which include a DNA-binding domain, and a functional homolog of the amino acid sequence shown in SEQ ID NO: 1, or having at least 60% sequence identity to the amino acid sequence shown in SEQ ID NO: 1 and / or having at least 60% sequence identity to the amino acid sequence shown in SEQ ID NO: 87 as described herein, and an optional activation domain (e.g., a synthetic domain, viral domain, or activation domain of the transcription factor of the present invention or any other transcription factor of any species described elsewhere herein), preferably an activation domain as seen in SEQ ID NO: 83. The alignment of the DNA-binding domain and any activation domain of the transcription factor of the present invention as described herein may be carried out according to the knowledge of those skilled in the art, and may be carried out in any order. The DNA-binding domain of the transcription factor of the present invention may be aligned to the C-terminus or N-terminus, preferably to the C-terminus, by those skilled in the art. In further embodiments, synthetic forms of the transcription factor of the present invention (e.g., synMSN4) may also be used in the present invention (e.g., SEQ ID NO: 27). Synthetic forms of transcription factors may include a synthetic DNA-binding domain (e.g., SEQ ID NO: 12). Furthermore, synthetic forms of transcription factors of the present invention may include an activation domain (a synthetic domain, viral domain, or activation domain of the transcription factor of the present invention or any other transcription factor of any kind described elsewhere herein), preferably an activation domain as seen in SEQ ID NO: 84. Again, the alignment of the DNA-binding domain and any activation domain of the transcription factor of the present invention as described herein may be carried out according to the knowledge of those skilled in the art, and may be carried out in any order.The DNA-binding domain of the synthetic transcription factor of the present invention can be arranged at the C-terminus or N-terminus, preferably at the C-terminus, by those skilled in the art.
[0045] In this invention, the transcription factor refers to the Msn4 / 2 protein (Msn4 / 2p or MSN4 / 2). Msn4p is a homolog of Msn2p in yeasts and their relatives, such as Saccharomyces cerevisiae, that have undergone a whole-genome duplication event. Most other yeast and fungal species contain only Msn-type transcription factors, and these transcription factors in these species cannot be reasonably distinguished. Due to this functional redundancy, these transcription factors can be called Msn2, Msn4, or Msn4 / 2. Due to high homology, Msn4p and Msn2p are highly likely to be interchangeable; i.e., the transcription factors are redundant. There is no fundamental difference between Msn2-dependent expression and Msn4-dependent expression, and the structures of Msn4p and Msn2p are also very similar. Pichia pastris has the only homolog, named Msn4p. In some other yeasts, there is only one homolog for Msn4 / 2, but it may have a different name. In Aspergillus niger, the homolog of Msn4 / 2 is called Seb1. In Saccharomyces cerevisiae, the homolog of Msn4 / 2 is called Com2.
[0046] MSN4 (e.g., MSN2) encodes a transcription factor that modulates general stress responses. In Saccharomyces cerevisiae, Msn4p (e.g., Msn2p) modulates approximately 200 genes in response to several stresses, including heat shock, osmotic shock, oxidative stress, low pH, glucose depletion, sorbic acid, and high ethanol concentrations, by C-terminal binding of the Msn4p (e.g., Msn2p) zinc finger-binding domain to the STRE element 5'-CCCCT-3' located in the promoters of these genes. At its N-terminus, Msn4p (e.g., Msn2p) contains a transcriptional activation domain and a nuclear export sequence. Furthermore, Msn4p (e.g., Msn2p) contains a nuclear localization signal, which is repressed by phosphorylation of protein kinase A and activated by dephosphorylation by protein phosphatase 1. Under stress-free conditions, Msn4p (e.g., Msn2p) is located in the cytoplasm. Cytoplasmic localization is regulated in part by TOR signaling. Under stress, Msn4p (e.g., Msn2p) is hyperphosphorylated, relocalizes to the nucleus, and subsequently exhibits periodic nuclear-cytoplasmic cyclic behavior.
[0047] Preferably, the transcription factor of the present invention comprises an amino acid sequence such as those shown in SEQ ID NOs: 15-27.
[0048] To date, it has not been known anywhere that the transcription factor Msn4p is involved in increasing the yield / titer of recombinant target protein of interest, or in general, in the secretion of recombinant target protein of interest by eukaryotic host cells. Therefore, it was surprising that overexpression of Msn4p in eukaryotic host cells increased the yield / titer of the recombinant target protein of interest of the present invention.
[0049] The transcription factors in this invention were initially isolated from Pichia pastris (Comagataera fafi) strain CBS7435 (CBS-KNAW culture collection). It is assumed that the transcription factors can be overexpressed in a wide variety of host cells. Therefore, instead of using sequences that are native to a species or genus, transcription factor sequences may also be obtained or derived from other prokaryotes or eukaryotes, preferably from fungal host cells, more preferably from yeast host cells, such as Pichia pastris (synonymous Comagataera), Hansenula polymorpha (synonymous Hansenula angusta), Trichoderma liesei, Aspergillus niger, Saccharomyces cerevisiae, Kluiveromyces lactis, Yarowia lipopolitica, Pichia metanorica, Candida boidini, Comagataera, and Schizosaccharomyces pombe. Preferably, the transcription factor is derived from Pichia pastris (Komagataella species), Saccharomyces cerevisiae, Yarowia liporitica, or Aspergillus niger, more preferably from Pichia pastris (Komagataella species). Furthermore, synthetic forms of the transcription factor of the present invention may be used. The Komagataella species used herein include all species of the genus Komagataella. In preferred embodiments, the transcription factor is derived from Komagataella pastris, Komagataella pseudopastoris, or Komagataella fafi. In even more preferred embodiments, the transcription factor is derived from Komagataella pastris or Komagataella fafi.
[0050] Preferably, the method, recombinant host cell, and transcription factor used in the use of the recombinant host cell of the present invention include at least one DNA-binding domain (DNA-binding domain of Pichia pastris, particularly Komagataera fafi or Komagataera pastris Msn4p) containing an amino acid sequence as shown in Sequence ID No. 1, and an activation domain. Therefore, the method, recombinant host cell, and use of the present invention preferably involve overexpression of a transcription factor containing at least one DNA-binding domain and an activation domain containing an amino acid sequence as shown in Sequence ID No. 1 in Pichia pastris (Komagataera species). Overexpression of the transcription factor, which includes at least one DNA-binding domain and an activation domain containing an amino acid sequence as shown in SEQ ID NO: 1, is also preferred in Hanzenula polymorpha, Trichoderma liesei, Aspergillus niger, Saccharomyces cerevisiae, Kluiveromyces lactis, Yarowia liporitica, Pichia metanorica, Candida boidini, Komagataella species, and Schizosaccharomyces pombe.
[0051] The method, recombinant host cell, and transcription factor used in the use of the recombinant host cell of the present invention include at least one DNA-binding domain and an activation domain, which include a functional homolog of the amino acid sequence shown in SEQ ID NO: 1 (the Msn4p DNA-binding domain of Pichia pastris), having at least 60% sequence identity with the amino acid sequence shown in SEQ ID NO: 1. Furthermore, the present invention also includes a method, recombinant host cell, and transcription factor used in the use of the recombinant host cell, which include at least one DNA-binding domain and an activation domain, which include a functional homolog of the amino acid sequence shown in SEQ ID NO: 1 (the Msn4p DNA-binding domain of Pichia pastris), having at least 60% sequence identity with the amino acid sequence shown in SEQ ID NO: 87. Preferably, the method, recombinant host cell, and transcription factor used in the use of the recombinant host cell of the present invention comprises at least one DNA-binding domain and an activation domain, which includes a functional homolog of the amino acid sequence shown in SEQ ID NO: 1 (the DNA-binding domain of Pichia pastris Msn4p), having at least 60% sequence identity to the amino acid sequence shown in SEQ ID NO: 1 and / or at least 60% sequence identity to the amino acid sequence shown in SEQ ID NO: 87. Accordingly, the method, recombinant host cell, and use of the present invention may further include overexpression in Pichia pastris of a transcription factor comprising at least one DNA-binding domain and an activation domain, which includes a functional homolog of the amino acid sequence shown in SEQ ID NO: 1, having at least 60% sequence identity to the amino acid sequence shown in SEQ ID NO: 1 and / or at least 60% sequence identity to the amino acid sequence shown in SEQ ID NO: 87.Accordingly, the method, recombinant host cells, and use of the present invention may further involve overexpression of a transcription factor comprising at least one DNA-binding domain and an activation domain, which is a functional homolog of the amino acid sequence shown in SEQ ID NO: 1, having at least 60% sequence identity to the amino acid sequence shown in SEQ ID NO: 1, and / or having at least 60% sequence identity to the amino acid sequence shown in SEQ ID NO: 87, in Hanzenula polymorpha, Trichoderma liesei, Aspergillus niger, Saccharomyces cerevisiae, Kluiveromyces lactis, Yarowia liporitica, Pichia metanorica, Candida boidini, Komagataella species, or Schizosaccharomyces pombe.
[0052] Preferably, a functional homolog of the amino acid sequence shown in SEQ ID NO: 1, having at least 60% sequence identity with the amino acid sequence shown in SEQ ID NO: 1, and / or having at least 60% sequence identity with the amino acid sequence shown in SEQ ID NO: 87, has the amino acid sequences shown in SEQ ID NOs: 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, and 12.
[0053] Therefore, the method, recombinant host cells, and use of the present invention may further involve overexpressing a transcription factor comprising at least one DNA-binding domain and an activation domain, which include an amino acid sequence as shown in SEQ ID NOs: 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, and 12.
[0054] Furthermore, the methods, recombinant host cells, and uses of the present invention may further include overexpression in Pichia pastris of a transcription factor comprising at least one DNA-binding domain and an activation domain, comprising an amino acid sequence as shown in SEQ ID NOs: 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, and 12. Accordingly, the methods, recombinant host cells, and uses of the present invention may further include overexpression in Hanzenula polymorpha, Trichoderma liesei, Aspergillus niger, Saccharomyces cerevisiae, Kluiveromyces lactis, Yarowia lipopolitica, Pichia metanorica, Candida boidini, Komagataella species, or Schizosaccharomyces pombe of the present invention of a transcription factor comprising at least one DNA-binding domain and an activation domain, comprising an amino acid sequence as shown in SEQ ID NOs: 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, and 12.
[0055] As used herein, "DNA-binding domain" or "binding domain" refers to a domain of a transcription factor that binds to the DNA of the gene being regulated. Preferably, the DNA-binding domain of the present invention is selected from the group consisting of functional homologs of the amino acid sequence shown in SEQ ID NO: 1 (e.g., SEQ ID NOs: 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, and 12) that have at least 60% sequence identity to the amino acid sequence shown in SEQ ID NO: 1, and / or have at least 60% sequence identity to the amino acid sequence shown in SEQ ID NO: 87. Most preferably, the DNA-binding domain is as shown in SEQ ID NO: 1. Therefore, the present invention may also include synthetic DNA-binding domains, as can be seen from SEQ ID NO: 12.
[0056] In this specification, Sequence ID No. 87 refers to the common sequence of the MSN4 / 2-like C2H2-type zinc finger DNA-binding domain (see Figure 6). Alignment of MSN4 / 2 transcription factors from different origins was performed using the CLC Main Workbench software (Qiagen Bioinformatics) as described in Example 6. Here, the known DNA-binding domains of Msn4p / Msn2p in Saccharomyces cerevisiae, a model organism frequently used in experiments and having undergone whole-genome duplication (WGD) and thus possessing two homologs, Msn4p and Msn2p, are used to induce the same function in other organisms. The zinc finger in Msn2 / 4 of Saccharomyces cerevisiae is X2-CX 2,4 -CX 12 -HX 3,4,5 It has a C2H2-like folding pattern with an -H amino acid sequence motif (see Figure 7). The common sequence of the Msn4 / 2DNA binding domain (SEQ ID NO: 87) is as follows: [Table 1] (In the formula, The 10th card, K, can be interchangeable with R; The 11th card, R, can be interchangeable with K; The 15th ranked Xaa can be Q or S; The 19th card, K, can be interchangeable with R; Xaa, ranked 22nd, can be any natural amino acid; The 25th ranked Xaa can be V or L; S, ranked 27th, can be interchangeable with T; Xaa, ranked 28th, can be any natural amino acid; The 30th card, K, can be interchangeable with R; Xaa, ranked 33rd, can be any natural amino acid; The Xaa at positions 35-36 can be any natural amino acid; Xaa, ranked 38th, can be any natural amino acid; The 40th card, K, can be interchangeable with R; S, ranked 44th, can be interchangeable with T; Xaa, ranked 48th, can be any natural amino acid; (R, ranked 52nd, may be interchangeable with K.) It has. Bold letters are highly preserved, and underlined letters are part of the C2H2 type zinc finger.
[0057] As used herein, the term “homologue” or “homolog” of a transcription factor or transcription factor binding domain of the present invention means that a protein has the same or conserved residues at corresponding positions in its primary, secondary, or tertiary structure. The term is also extended to two or more nucleotide sequences encoding the same polypeptide. When the function as a transcription factor or as a transcription factor binding domain is demonstrated using such a homolog, the homolog is called a “functional homolog.” A functional homolog performs the same or substantially the same function as the transcription factor or transcription factor binding domain from which it is derived. In the case of a nucleotide sequence, a “functional homolog” preferably means a nucleotide sequence that has a different sequence from the original nucleotide sequence but still encodes the same amino acid sequence due to the use of a degenerate genetic code. Functional homologs of proteins, particularly transcription factors or transcription factor binding domains, are obtained by substituting one or more amino acids in the protein, particularly the transcription factor or transcription factor binding domain, and such substitutions preserve the function of the protein, particularly the transcription factor or transcription factor binding domain.In particular, functional homologs of amino acid sequences such as the one shown in SEQ ID NO: 1 are at least approximately 60% of the amino acid sequence shown in SEQ ID NO: 1 (the DNA-binding domain of Msn4p in Pichia pastris), for example, at least 61%, 62%, 63%, 64%, 65%, 66%, 67%, 68%, 69%, 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, Or, furthermore, 100% amino acid sequence identity and / or at least about 60% of the amino acid sequence (common sequence) as shown in Sequence ID No. 87, for example, at least 61%, 62%, 63%, 64%, 65%, 66%, 67%, 68%, 69%, 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or furthermore, 100% amino acid sequence identity. In some embodiments, a functional homolog of an amino acid sequence such as the one shown in SEQ ID NO: 1 has at least about 60% amino acid sequence identity with respect to the amino acid sequence shown in SEQ ID NO: 1 (the DNA-binding domain of Msn4p of Pichia pastris), and at least about 60% amino acid sequence identity with respect to the amino acid sequence such as the one shown in SEQ ID NO: 87 (the common sequence), for example, at least 61%, 62%, 63%, 64%, 65%, 66%, 67%, 68%, 69%, 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or even 100% amino acid sequence identity.In some embodiments, a functional homolog of an amino acid sequence such as the one shown in SEQ ID NO: 1 has at least about 61% amino acid sequence identity with respect to the amino acid sequence such as the one shown in SEQ ID NO: 1 (the DNA-binding domain of Msn4p of Pichia pastris), and at least about 60% amino acid sequence identity with respect to the amino acid sequence such as the one shown in SEQ ID NO: 87 (the common sequence), for example, at least 61%, 62%, 63%, 64%, 65%, 66%, 67%, 68%, 69%, 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or even 100% amino acid sequence identity. In some embodiments, a functional homolog of an amino acid sequence such as the one shown in SEQ ID NO: 1 has at least about 62% amino acid sequence identity with respect to the amino acid sequence such as the one shown in SEQ ID NO: 1 (the DNA-binding domain of Msn4p of Pichia pastris), and at least about 60% amino acid sequence identity with respect to the amino acid sequence such as the one shown in SEQ ID NO: 87 (the common sequence), for example, at least 61%, 62%, 63%, 64%, 65%, 66%, 67%, 68%, 69%, 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or even 100% amino acid sequence identity.In some embodiments, a functional homolog of an amino acid sequence such as the one shown in SEQ ID NO: 1 has at least about 63% amino acid sequence identity with respect to the amino acid sequence such as the one shown in SEQ ID NO: 1 (the DNA-binding domain of Msn4p of Pichia pastris), and at least about 60% amino acid sequence identity with respect to the amino acid sequence such as the one shown in SEQ ID NO: 87 (the common sequence), for example, at least 61%, 62%, 63%, 64%, 65%, 66%, 67%, 68%, 69%, 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or even 100% amino acid sequence identity. In some embodiments, a functional homolog of an amino acid sequence such as the one shown in SEQ ID NO: 1 has at least about 64% amino acid sequence identity with respect to the amino acid sequence such as the one shown in SEQ ID NO: 1 (the DNA-binding domain of Msn4p of Pichia pastris), and at least about 60% amino acid sequence identity with respect to the amino acid sequence such as the one shown in SEQ ID NO: 87 (the common sequence), for example, at least 61%, 62%, 63%, 64%, 65%, 66%, 67%, 68%, 69%, 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or even 100% amino acid sequence identity.In some embodiments, a functional homolog of an amino acid sequence such as the one shown in SEQ ID NO: 1 has at least about 65% amino acid sequence identity with respect to the amino acid sequence such as the one shown in SEQ ID NO: 1 (the DNA-binding domain of Msn4p of Pichia pastris), and at least about 60% amino acid sequence identity with respect to the amino acid sequence such as the one shown in SEQ ID NO: 87 (the common sequence), for example, at least 61%, 62%, 63%, 64%, 65%, 66%, 67%, 68%, 69%, 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or even 100% amino acid sequence identity. In some embodiments, a functional homolog of an amino acid sequence such as the one shown in SEQ ID NO: 1 has at least about 66% amino acid sequence identity with respect to the amino acid sequence such as the one shown in SEQ ID NO: 1 (the DNA-binding domain of Msn4p of Pichia pastris), and at least about 60% amino acid sequence identity with respect to the amino acid sequence such as the one shown in SEQ ID NO: 87 (the common sequence), for example, at least 61%, 62%, 63%, 64%, 65%, 66%, 67%, 68%, 69%, 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or even 100% amino acid sequence identity.In some embodiments, a functional homolog of an amino acid sequence such as the one shown in SEQ ID NO: 1 has at least about 67% amino acid sequence identity with respect to the amino acid sequence such as the one shown in SEQ ID NO: 1 (the DNA-binding domain of Msn4p of Pichia pastris), and at least about 60% amino acid sequence identity with respect to the amino acid sequence such as the one shown in SEQ ID NO: 87 (the common sequence), for example, at least 61%, 62%, 63%, 64%, 65%, 66%, 67%, 68%, 69%, 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or even 100% amino acid sequence identity. In some embodiments, a functional homolog of an amino acid sequence such as the one shown in SEQ ID NO: 1 has at least about 68% amino acid sequence identity with respect to the amino acid sequence such as the one shown in SEQ ID NO: 1 (the DNA-binding domain of Msn4p of Pichia pastris), and at least about 60% amino acid sequence identity with respect to the amino acid sequence such as the one shown in SEQ ID NO: 87 (the common sequence), for example, at least 61%, 62%, 63%, 64%, 65%, 66%, 67%, 68%, 69%, 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or even 100% amino acid sequence identity.In some embodiments, a functional homolog of an amino acid sequence such as the one shown in SEQ ID NO: 1 has at least about 69% amino acid sequence identity with respect to the amino acid sequence such as the one shown in SEQ ID NO: 1 (the DNA-binding domain of Msn4p of Pichia pastris), and at least about 60% amino acid sequence identity with respect to the amino acid sequence such as the one shown in SEQ ID NO: 87 (the common sequence), for example, at least 61%, 62%, 63%, 64%, 65%, 66%, 67%, 68%, 69%, 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or even 100% amino acid sequence identity. In some embodiments, a functional homolog of an amino acid sequence such as the one shown in SEQ ID NO: 1 has at least about 70% amino acid sequence identity with respect to the amino acid sequence shown in SEQ ID NO: 1 (the DNA-binding domain of Msn4p of Pichia pastris), and at least about 60% amino acid sequence identity with respect to the amino acid sequence such as the one shown in SEQ ID NO: 87 (the common sequence), for example, at least 61%, 62%, 63%, 64%, 65%, 66%, 67%, 68%, 69%, 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or even 100% amino acid sequence identity. In some embodiments, a functional homolog of an amino acid sequence such as the one shown in SEQ ID NO: 1 has at least about 71% amino acid sequence identity with respect to the amino acid sequence such as the one shown in SEQ ID NO: 1 (the DNA-binding domain of Msn4p of Pichia pastris), and at least about 60% amino acid sequence identity with respect to the amino acid sequence such as the one shown in SEQ ID NO: 87 (the common sequence), for example, at least 61%, 62%, 63%, 64%, 65%, 66%, 67%, 68%, 69%, 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or even 100% amino acid sequence identity. In some embodiments, a functional homolog of an amino acid sequence such as the one shown in SEQ ID NO: 1 has at least about 72% amino acid sequence identity with respect to the amino acid sequence such as the one shown in SEQ ID NO: 1 (the DNA-binding domain of Msn4p of Pichia pastris), and at least about 60% amino acid sequence identity with respect to the amino acid sequence such as the one shown in SEQ ID NO: 87 (the common sequence), for example, at least 61%, 62%, 63%, 64%, 65%, 66%, 67%, 68%, 69%, 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or even 100% amino acid sequence identity.In some embodiments, a functional homolog of an amino acid sequence such as the one shown in SEQ ID NO: 1 has at least about 73% amino acid sequence identity with respect to the amino acid sequence such as the one shown in SEQ ID NO: 1 (the DNA-binding domain of Msn4p of Pichia pastris), and at least about 60% amino acid sequence identity with respect to the amino acid sequence such as the one shown in SEQ ID NO: 87 (the common sequence), for example, at least 61%, 62%, 63%, 64%, 65%, 66%, 67%, 68%, 69%, 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or even 100% amino acid sequence identity. In some embodiments, a functional homolog of an amino acid sequence such as the one shown in SEQ ID NO: 1 has at least about 74% amino acid sequence identity with respect to the amino acid sequence such as the one shown in SEQ ID NO: 1 (the DNA-binding domain of Msn4p of Pichia pastris), and at least about 60% amino acid sequence identity with respect to the amino acid sequence such as the one shown in SEQ ID NO: 87 (the common sequence), for example, at least 61%, 62%, 63%, 64%, 65%, 66%, 67%, 68%, 69%, 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or even 100% amino acid sequence identity.In some embodiments, a functional homolog of an amino acid sequence such as the one shown in SEQ ID NO: 1 has at least about 75% amino acid sequence identity with respect to the amino acid sequence such as the one shown in SEQ ID NO: 1 (the DNA-binding domain of Msn4p of Pichia pastris), and at least about 60% amino acid sequence identity with respect to the amino acid sequence such as the one shown in SEQ ID NO: 87 (the common sequence), for example, at least 61%, 62%, 63%, 64%, 65%, 66%, 67%, 68%, 69%, 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or even 100% amino acid sequence identity. In some embodiments, a functional homolog of an amino acid sequence such as the one shown in SEQ ID NO: 1 has at least about 76% amino acid sequence identity with respect to the amino acid sequence such as the one shown in SEQ ID NO: 1 (the DNA-binding domain of Msn4p in Pichia pastris), and at least about 60% amino acid sequence identity with respect to the amino acid sequence such as the one shown in SEQ ID NO: 87 (the common sequence), for example, at least 61%, 62%, 63%, 64%, 65%, 66%, 67%, 68%, 69%, 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or even 100% amino acid sequence identity.In some embodiments, a functional homolog of an amino acid sequence such as the one shown in SEQ ID NO: 1 has at least about 77% amino acid sequence identity with respect to the amino acid sequence such as the one shown in SEQ ID NO: 1 (the DNA-binding domain of Msn4p of Pichia pastris), and at least about 60% amino acid sequence identity with respect to the amino acid sequence such as the one shown in SEQ ID NO: 87 (the common sequence), for example, at least 61%, 62%, 63%, 64%, 65%, 66%, 67%, 68%, 69%, 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or even 100% amino acid sequence identity. In some embodiments, a functional homolog of an amino acid sequence such as the one shown in SEQ ID NO: 1 has at least about 78% amino acid sequence identity with respect to the amino acid sequence shown in SEQ ID NO: 1 (the DNA-binding domain of Msn4p of Pichia pastris), and at least about 60% amino acid sequence identity with respect to the amino acid sequence such as the one shown in SEQ ID NO: 87 (the common sequence), for example, at least 61%, 62%, 63%, 64%, 65%, 66%, 67%, 68%, 69%, 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or even 100% amino acid sequence identity.In some embodiments, a functional homolog of an amino acid sequence such as the one shown in SEQ ID NO: 1 has at least about 79% amino acid sequence identity with respect to the amino acid sequence such as the one shown in SEQ ID NO: 1 (the DNA-binding domain of Msn4p of Pichia pastris), and at least about 60% amino acid sequence identity with respect to the amino acid sequence such as the one shown in SEQ ID NO: 87 (the common sequence), for example, at least 61%, 62%, 63%, 64%, 65%, 66%, 67%, 68%, 69%, 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or even 100% amino acid sequence identity. In some embodiments, a functional homolog of an amino acid sequence such as the one shown in SEQ ID NO: 1 has at least about 80% amino acid sequence identity with respect to the amino acid sequence such as the one shown in SEQ ID NO: 1 (the DNA-binding domain of Msn4p of Pichia pastris), and at least about 60% amino acid sequence identity with respect to the amino acid sequence such as the one shown in SEQ ID NO: 87 (the common sequence), for example, at least 61%, 62%, 63%, 64%, 65%, 66%, 67%, 68%, 69%, 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or even 100% amino acid sequence identity.In some embodiments, a functional homolog of an amino acid sequence such as the one shown in SEQ ID NO: 1 has at least about 81% amino acid sequence identity with respect to the amino acid sequence shown in SEQ ID NO: 1 (the DNA-binding domain of Msn4p of Pichia pastris), and at least about 60% amino acid sequence identity with respect to the amino acid sequence such as the one shown in SEQ ID NO: 87 (the common sequence), for example, at least 61%, 62%, 63%, 64%, 65%, 66%, 67%, 68%, 69%, 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or even 100% amino acid sequence identity. In some embodiments, a functional homolog of an amino acid sequence such as the one shown in SEQ ID NO: 1 has at least about 82% amino acid sequence identity with respect to the amino acid sequence shown in SEQ ID NO: 1 (the DNA-binding domain of Msn4p of Pichia pastris), and at least about 60% amino acid sequence identity with respect to the amino acid sequence such as the one shown in SEQ ID NO: 87 (the common sequence), for example, at least 61%, 62%, 63%, 64%, 65%, 66%, 67%, 68%, 69%, 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or even 100% amino acid sequence identity.In some embodiments, a functional homolog of an amino acid sequence such as the one shown in SEQ ID NO: 1 has at least about 83% amino acid sequence identity with respect to the amino acid sequence shown in SEQ ID NO: 1 (the DNA-binding domain of Msn4p in Pichia pastris), and at least about 60% amino acid sequence identity with respect to the amino acid sequence such as the one shown in SEQ ID NO: 87 (the common sequence), for example, at least 61%, 62%, 63%, 64%, 65%, 66%, 67%, 68%, 69%, 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or even 100% amino acid sequence identity. In some embodiments, a functional homolog of an amino acid sequence such as the one shown in SEQ ID NO: 1 has at least about 84% amino acid sequence identity with respect to the amino acid sequence such as the one shown in SEQ ID NO: 1 (the DNA-binding domain of Msn4p of Pichia pastris), and at least about 60% amino acid sequence identity with respect to the amino acid sequence such as the one shown in SEQ ID NO: 87 (the common sequence), for example, at least 61%, 62%, 63%, 64%, 65%, 66%, 67%, 68%, 69%, 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or even 100% amino acid sequence identity. In some embodiments, a functional homolog of an amino acid sequence such as the one shown in SEQ ID NO: 1 has at least about 85% amino acid sequence identity with respect to the amino acid sequence such as the one shown in SEQ ID NO: 1 (the DNA-binding domain of Msn4p of Pichia pastris), and at least about 60% amino acid sequence identity with respect to the amino acid sequence such as the one shown in SEQ ID NO: 87 (the common sequence), for example, at least 61%, 62%, 63%, 64%, 65%, 66%, 67%, 68%, 69%, 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or even 100% amino acid sequence identity. In some embodiments, a functional homolog of an amino acid sequence such as the one shown in SEQ ID NO: 1 has at least about 86% amino acid sequence identity with respect to the amino acid sequence shown in SEQ ID NO: 1 (the DNA-binding domain of Msn4p of Pichia pastris), and at least about 60% amino acid sequence identity with respect to the amino acid sequence such as the one shown in SEQ ID NO: 87 (the common sequence), for example, at least 61%, 62%, 63%, 64%, 65%, 66%, 67%, 68%, 69%, 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or even 100% amino acid sequence identity.In some embodiments, a functional homolog of an amino acid sequence such as the one shown in SEQ ID NO: 1 has at least about 87% amino acid sequence identity with respect to the amino acid sequence such as the one shown in SEQ ID NO: 1 (the DNA-binding domain of Msn4p of Pichia pastris), and at least about 60% amino acid sequence identity with respect to the amino acid sequence such as the one shown in SEQ ID NO: 87 (the common sequence), for example, at least 61%, 62%, 63%, 64%, 65%, 66%, 67%, 68%, 69%, 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or even 100% amino acid sequence identity. In some embodiments, a functional homolog of an amino acid sequence such as the one shown in SEQ ID NO: 1 has at least about 88% amino acid sequence identity with respect to the amino acid sequence such as the one shown in SEQ ID NO: 1 (the DNA-binding domain of Msn4p of Pichia pastris), and at least about 60% amino acid sequence identity with respect to the amino acid sequence such as the one shown in SEQ ID NO: 87 (the common sequence), for example, at least 61%, 62%, 63%, 64%, 65%, 66%, 67%, 68%, 69%, 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or even 100% amino acid sequence identity.In some embodiments, a functional homolog of an amino acid sequence such as the one shown in SEQ ID NO: 1 has at least about 89% amino acid sequence identity with respect to the amino acid sequence shown in SEQ ID NO: 1 (the DNA-binding domain of Msn4p of Pichia pastris), and at least about 60% amino acid sequence identity with respect to the amino acid sequence such as the one shown in SEQ ID NO: 87 (the common sequence), for example, at least 61%, 62%, 63%, 64%, 65%, 66%, 67%, 68%, 69%, 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or even 100% amino acid sequence identity. In some embodiments, a functional homolog of an amino acid sequence such as the one shown in SEQ ID NO: 1 has at least about 90% amino acid sequence identity with respect to the amino acid sequence such as the one shown in SEQ ID NO: 1 (the DNA-binding domain of Msn4p of Pichia pastris), and at least about 60% amino acid sequence identity with respect to the amino acid sequence such as the one shown in SEQ ID NO: 87 (the common sequence), for example, at least 61%, 62%, 63%, 64%, 65%, 66%, 67%, 68%, 69%, 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or even 100% amino acid sequence identity.In some embodiments, a functional homolog of an amino acid sequence such as the one shown in SEQ ID NO: 1 has at least about 91% amino acid sequence identity with respect to the amino acid sequence such as the one shown in SEQ ID NO: 1 (the DNA-binding domain of Msn4p of Pichia pastris), and at least about 60% amino acid sequence identity with respect to the amino acid sequence such as the one shown in SEQ ID NO: 87 (the common sequence), for example, at least 61%, 62%, 63%, 64%, 65%, 66%, 67%, 68%, 69%, 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or even 100% amino acid sequence identity. In some embodiments, a functional homolog of an amino acid sequence such as the one shown in SEQ ID NO: 1 has at least about 92% amino acid sequence identity with respect to the amino acid sequence such as the one shown in SEQ ID NO: 1 (the DNA-binding domain of Msn4p of Pichia pastris), and at least about 60% amino acid sequence identity with respect to the amino acid sequence such as the one shown in SEQ ID NO: 87 (common sequence), for example, at least 61%, 62%, 63%, 64%, 65%, 66%, 67%, 68%, 69%, 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or even 100% amino acid sequence identity.In some embodiments, a functional homolog of an amino acid sequence such as the one shown in SEQ ID NO: 1 has at least about 93% amino acid sequence identity with respect to the amino acid sequence such as the one shown in SEQ ID NO: 1 (the DNA-binding domain of Msn4p of Pichia pastris), and at least about 60% amino acid sequence identity with respect to the amino acid sequence such as the one shown in SEQ ID NO: 87 (the common sequence), for example, at least 61%, 62%, 63%, 64%, 65%, 66%, 67%, 68%, 69%, 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or even 100% amino acid sequence identity. In some embodiments, a functional homolog of an amino acid sequence such as the one shown in SEQ ID NO: 1 has at least about 94% amino acid sequence identity with respect to the amino acid sequence such as the one shown in SEQ ID NO: 1 (the DNA-binding domain of Msn4p of Pichia pastris), and at least about 60% amino acid sequence identity with respect to the amino acid sequence such as the one shown in SEQ ID NO: 87 (common sequence), for example, at least 61%, 62%, 63%, 64%, 65%, 66%, 67%, 68%, 69%, 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or even 100% amino acid sequence identity.In some embodiments, a functional homolog of an amino acid sequence such as the one shown in SEQ ID NO: 1 has at least about 95% amino acid sequence identity with respect to the amino acid sequence such as the one shown in SEQ ID NO: 1 (the DNA-binding domain of Msn4p in Pichia pastris), and at least about 60% amino acid sequence identity with respect to the amino acid sequence such as the one shown in SEQ ID NO: 87 (the common sequence), for example, at least 61%, 62%, 63%, 64%, 65%, 66%, 67%, 68%, 69%, 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or even 100% amino acid sequence identity. In some embodiments, a functional homolog of an amino acid sequence such as the one shown in SEQ ID NO: 1 has at least about 96% amino acid sequence identity with respect to the amino acid sequence such as the one shown in SEQ ID NO: 1 (the DNA-binding domain of Msn4p of Pichia pastris), and at least about 60% amino acid sequence identity with respect to the amino acid sequence such as the one shown in SEQ ID NO: 87 (the common sequence), for example, at least 61%, 62%, 63%, 64%, 65%, 66%, 67%, 68%, 69%, 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or even 100% amino acid sequence identity.In some embodiments, a functional homolog of an amino acid sequence such as the one shown in SEQ ID NO: 1 has at least about 97% amino acid sequence identity with respect to the amino acid sequence such as the one shown in SEQ ID NO: 1 (the DNA-binding domain of Msn4p of Pichia pastris), and at least about 60% amino acid sequence identity with respect to the amino acid sequence such as the one shown in SEQ ID NO: 87 (the common sequence), for example, at least 61%, 62%, 63%, 64%, 65%, 66%, 67%, 68%, 69%, 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or even 100% amino acid sequence identity. In some embodiments, a functional homolog of an amino acid sequence such as the one shown in SEQ ID NO: 1 has at least about 98% amino acid sequence identity with respect to the amino acid sequence such as the one shown in SEQ ID NO: 1 (the DNA-binding domain of Msn4p of Pichia pastris), and at least about 60% amino acid sequence identity with respect to the amino acid sequence such as the one shown in SEQ ID NO: 87 (the common sequence), for example, at least 61%, 62%, 63%, 64%, 65%, 66%, 67%, 68%, 69%, 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or even 100% amino acid sequence identity. In some embodiments, a functional homolog of an amino acid sequence such as the one shown in SEQ ID NO: 1 has at least about 99% amino acid sequence identity with respect to the amino acid sequence such as the one shown in SEQ ID NO: 1 (the DNA-binding domain of Msn4p in Pichia pastris), and at least about 60% amino acid sequence identity with respect to the amino acid sequence such as the one shown in SEQ ID NO: 87 (common sequence), for example, at least 61%, 62%, 63%, 64%, 65%, 66%, 67%, 68%, 69%, 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or even 100% amino acid sequence identity. In some embodiments, a functional homolog of an amino acid sequence such as the one shown in SEQ ID NO: 1 has at least approximately 100% amino acid sequence identity to the amino acid sequence shown in SEQ ID NO: 1 (the DNA-binding domain of Msn4p in Pichia pastris), and at least approximately 60% to the amino acid sequence such as the one shown in SEQ ID NO: 87 (the common sequence), for example, at least 61%, 62%, 6%. It has amino acid sequence identity rates of 3%, 64%, 65%, 66%, 67%, 68%, 69%, 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or even 100%.
[0058] Generally, homologs can be prepared using any mutagenesis procedure known in the art, such as site-directed mutagenesis, synthetic gene construction, semi-synthetic gene construction, random mutagenesis, or shuffling. Site-directed mutagenesis is a technique in which one or more (e.g., several) mutations are introduced at one or more designated sites within a parental polynucleotide. Site-directed mutagenesis can be achieved in vitro by PCR, which involves the use of oligonucleotide primers containing the desired mutation. Site-directed mutagenesis can also be performed in vitro by cassette mutagenesis, which involves restriction enzyme cleavage at a site within a plasmid containing the parental polynucleotide, followed by ligation of an oligonucleotide containing the mutation within the polypeptide.
[0059] Typically, the restriction enzymes that digest plasmids and oligonucleotides are the same, allowing the sticky ends of the plasmid and insertion fragment to ligate to each other. See, for example, Scherer and Davis, 1979, Proc. Natl. Acad. Sci. USA 76: 4949-4955; and Barton et al., 1990, Nucleic Acids Res. 18: 7349-4966. Site-directed mutagenesis can also be achieved in vivo by methods known in the art. See, for example, U.S. Patent Application Publication No. 2004 / 0171154; Storici et ai, 2001, Nature Biotechnol. 19: 773-776; Kren et ai, 1998, Nat. Med. 4: 285-290; and Calissano and Macino, 1996, Fungal Genet. Newslett. 43: 15-16. Gene construction by synthesis involves the in vitro synthesis of polynucleotide molecules designed to encode polypeptides of interest. Gene synthesis can be carried out using a number of techniques, such as the multi-microchip-based technique described by Tian et al. (2004, Nature 432: 1050-1054), and similar techniques for synthesizing and assembling oligonucleotides on photoprogrammable microfluidic chips. Single or multiple amino acid substitutions, deletions, and / or insertions may be prepared and tested using known mutagenesis, recombination, and / or shuffling methods, followed by relevant screening procedures, e.g., Reidhaar-Olson and Sauer, 1988, Science 241:53-57; Bowie and Sauer, 1989, Proc. Natl. Acad. Sci. USA 86: 2152-2156; International Publication No. 95 / 17413; or International Publication No. 95 / 22625.Other methods that may be used include error-prone PCR, phage display (e.g., Lowman et al., 1991, Biochemistry 30: 10832-10837; U.S. Patent No. 5,223,409; International Publication No. 92 / 06204), and region-specific mutagenesis (Derbyshire et al., 1986, Gene 46: 145; Ner et al., 1988, DNA 7:127). By combining mutagenesis / shuffling methods with high-throughput automated screening methods, the activity of cloned and mutated polypeptides expressed by host cells can be detected (Ness et a / ., 1999, Nature Biotechnology 17: 893-896). Mutant DNA molecules encoding active polypeptides can be recovered from host cells and rapidly sequenced using standard methods known in the art. These methods allow for rapid determination of the importance of individual amino acid residues within the polypeptide. Semi-synthetic gene construction is achieved by combining synthetic gene construction and / or site-directed mutagenesis and / or random mutagenesis and / or shuffling. Semi-synthetic construction is typically represented by a process utilizing polynucleotide fragments synthesized in combination with PCR techniques. Thus, a defined region of the gene may be newly synthesized, while other regions may be amplified using site-directed mutagenesis primers, and yet other regions may be amplified by error-prone PCR or non-error-prone PCR. Subsequently, the polynucleotide subsequence may be shuffled.Alternatively, homologs can be obtained, for example, from natural sources, by screening cDNA libraries of other organisms, or by searching for homologs in nucleic acid databases, preferably closely related or related organisms, such as *Chamaeira pastoris*, *Chamaeira pseudopastoris*, or *Chamaeira fafi*, *Chamaeira* species, *Hanzenula polymorpha*, *Trichoderma liesei*, *Aspergillus niger*, *Saccharomyces cerevisiae*, *Cluyveromyces lactis*, *Yarowia liporitica*, *Pichia metanorica*, *Candida boidini*, *Chamaeira* species, or *Schizosaccharomyces pombe* homologs. Thus, SEQ ID NOs. 2-12 are functional homologs of the binding domain of transcription factors as shown in SEQ ID NO. 1, and SEQ ID NOs. 16-27 are functional homologs of transcription factors as shown in SEQ ID NO. 15.
[0060] Functional homologues of the amino acid sequence of the DNA-binding domain shown in SEQ ID NO: 1, having at least 60% sequence identity to the amino acid sequence shown in SEQ ID NO: 1 (e.g., SEQ ID NOs: 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, and 12) and / or having at least 60% sequence identity to the amino acid sequence shown in SEQ ID NO: 87, or functional homologues of the amino acid sequence of the transcription factor shown in SEQ ID NO: 15 (e.g., SEQ ID NOs: 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27), having at least 11% sequence identity to the amino acid sequence shown in SEQ ID NO: 15, or functional homologues of the amino acid sequence of the DNA-binding domain of an additional transcription factor shown in SEQ ID NO: 65 (e.g., SEQ ID NOs: 66-73), having at least 50% sequence identity to the amino acid sequence shown in SEQ ID NO: 65, or at least 20% sequence identity to the amino acid sequence shown in SEQ ID NO: 74 The function of an amino acid sequence homolog of an additional transcription factor having sequence identity, such as the one shown in SEQ ID NO: 74 (e.g., SEQ ID NOs: 75, 76, 77, 78, 79, 80, 81, 82), can be tested by preparing an expression cassette into which an amino acid sequence homolog of a DNA-binding domain, an activation domain (e.g., SEQ ID NO: 83 or 84), and a nuclear localization signal (NLS) (e.g., SEQ ID NO: 85 or 86), as shown in SEQ ID NO: 1, or an additional transcription factor, such as the one shown in SEQ ID NO: 65, an amino acid sequence homolog of a transcription factor, as shown in SEQ ID NO: 15, or an amino acid sequence homolog of a transcription factor, as shown in SEQ ID NO: 74, is inserted; transforming host cells having a sequence encoding a test protein, such as one of the model proteins or another of interest proteins used in the Examples section, and determining the difference in yield of the model protein or the interest protein under identical conditions.
[0061] The term "amino acid" refers to natural and synthetic amino acids, as well as amino acid analogs and amino acid mimes that function in a similar manner to natural amino acids. Natural amino acids are those encoded by the genetic code, as well as later modified amino acids, such as hydroxyproline, γ-carboxyglutamate, and O-phosphoserine. Amino acid analogs refer to compounds that have the same basic chemical structure as natural amino acids, i.e., hydrogen, a carboxyl group, an amino group, and a carbon atom bonded to an R group, such as homoserine, norleucine, methionine sulfoxide, and methionine methylsulfonium. Such analogs have a modified R group (e.g., norleucine) or a modified peptide skeleton, but retain the same basic chemical structure as natural amino acids. Amino acid mimes refer to compounds that have a different structure from the general chemical structure of amino acids, but function in a similar manner to natural amino acids.
[0062] "Sequence identity" or "identity %" refers to the percentage of residues that match between at least two polypeptide sequences or polynucleotide sequences aligned using a standardized algorithm. Such algorithms can insert gaps into the sequences being compared to optimize the alignment between the two sequences in a standardized and reproducible manner, thus achieving a more meaningful comparison of the two sequences. As used in this invention, sequence identity refers to the ratio of identical amino acids between at least two polypeptide sequences (amino acid sequences). Sequence similarity, as enumerated in this invention, refers to the ratio of amino acid groups that are similar in their side chains and charge between at least two polypeptide sequences (amino acid sequences). For the purposes of this invention, the sequence identity between two amino acid sequences or nucleotide sequences is determined using the NCBI BLAST program version 2.2.29 (January 6, 2014) (Altschul et al., Nucleic Acids Res. (1997) 25:3389-3402). The sequence identity rate of two amino acid sequences can be determined using blastp set with the following parameters: Matrix: BLOSUM62, Word size: 3; Expected value: 10; Gap cost: Existence = 11, Extension = 1; Filter = Inactivated low-complexity region; Composition adjustment: Score matrix adjustment according to conditional components. For the purposes of this invention, the sequence identity rate between two nucleotide sequences is determined using the NCBI BLAST program version 2.2.29 (January 6, 2014) with blastn set with the following exemplary parameters: Word size: 28; Expected value: 10; Gap cost: Linear; Filter = Activated low-complexity region; Match / Mismatch score: 1, -2. For the purposes of this invention, the sequence identity rate between two amino acid sequences or nucleotide sequences is further determined using BLAST and the Embossed Needle algorithm. The sequence identity rate of DNA-binding domains was evaluated by the global pairwise sequence alignment using the Embossed Needle algorithm.The Emboss Needle web server (https: / / www.ebi.ac.uk / Tools / psa / emboss_needle / ) was used for pairwise protein sequence alignment using default settings (matrix: BLOSUM62; gap open: 10; gap elongation: 0.5; end gap penalty: false; end gap open: 10; end gap elongation: 0.5). Emboss Needle decodes two input sequences and exports their optimal global sequence alignment to a file. It uses the Needleman-Bunsch alignment algorithm to find the optimal alignment (including gaps) of the two sequences along their entire length. Sequence identity rates for Pichia pastris KAR2, LHS1, SIL1, and ERJ5 were determined by BLAST.
[0063] As used herein, the term “activating domain” refers to any domain capable of activating transcription. In the present invention, each activating domain derived from any transcription factor of any organism known to those skilled in the art may be used as an activating domain. Preferably, for the transcription factor of the present invention, any activating domain of the transcription factor of any specified species herein, preferably an activating domain such as that shown in SEQ ID NO: 83, may be used. For additional transcription factors, any activating domain of additional transcription factors of any specified species herein may also be used. In further embodiments, synthetic (e.g., SEQ ID NO: 84) or viral (e.g., VP64) activating domains may also be used in the present invention for the transcription factor of the present invention or for additional transcription factors. The function of the activating domain can be measured by methods known in the art, namely by yeast two-hybrid (Y2H) techniques that enable the detection of interacting proteins in living yeast cells. Therefore, the method, recombinant host cell, and transcription factor used in the present invention each include at least one DNA-binding domain and an activating domain. Activating domains such as those shown in SEQ ID NO: 83 or SEQ ID NO: 84 may be preferred. It is also conceivable that activating domains derived from functional homologs may be used. The activation domain specific to MSN4 in Pichia pastris may be part of sequence number 83.
[0064] The present invention further provides a method for increasing the yield of recombinant protein of interest in host cells, comprising the steps of: i) engineering host cells to overexpress at least one polynucleotide encoding at least one transcription factor of the present invention, comprising at least one DNA-binding domain and an activation domain; ii) engineering the host cells to contain a polynucleotide encoding a protein of interest; iii) culturing the host cells under conditions suitable for overexpressing at least one polynucleotide encoding at least one transcription factor and for overexpressing the protein of interest; optionally iv) isolating the protein of interest from the cell culture medium; and optionally v) purifying the protein of interest.
[0065] It should be noted that the steps listed in (i) and (ii) do not need to be performed in the order they are listed. The steps listed in (ii) can be performed first, followed by the steps listed in (i). In step (i), the engineer can be manipulated to overexpress at least one polynucleotide encoding at least one transcription factor of the present invention, which includes a DNA-binding domain containing a functional homolog of the amino acid sequence shown in SEQ ID NO: 1, having at least 60% sequence identity to the amino acid sequence shown in SEQ ID NO: 1, and / or having at least 60% sequence identity to the amino acid sequence shown in SEQ ID NO: 87.
[0066] When a host cell is “engineered to overexpress” a given protein, the host cell has the ability to express, preferably overexpress, the transcription factor of the present invention or its functional homologue, thereby the host cell is engineered to increase the expression of a given protein, such as a protein of interest or a model protein, compared to a host cell under the same conditions before the engineering. In one embodiment, “engineered to overexpress” means that a genetic modification of the host cell is made to increase the expression of the protein, i.e., the cell is genetically engineered to (intentionally) overexpress such a protein.
[0067] When used in the context of host cells of the present invention, “pre-engineered” or “before manipulation” means that such host cells have not been engineered using polynucleotides encoding the transcription factor or its functional homologue of the present invention. Accordingly, the term also means that the host cells do not overexpress polynucleotides encoding the transcription factor or its functional homologue of the present invention, or have not been engineered to overexpress polynucleotides encoding the transcription factor or its functional homologue of the present invention. Accordingly, “pre-engineered host cells” or “before manipulation host cells” or “host cells that do not overexpress polynucleotides encoding the transcription factor” is a host cell that does not overexpress polynucleotides encoding the transcription factor or its functional homologue of the present invention, or a host cell that has not been engineered to overexpress polynucleotides encoding the transcription factor or its functional homologue of the present invention. Furthermore, “host cells before engineering,” “host cells before manipulation,” or “host cells that do not overexpress polynucleotides encoding the transcription factor” are the same host cells from which the increase in yield of the recombinant protein of interest is compared, but which do not overexpress polynucleotides encoding the transcription factor of the present invention or its functional homologue, or have not been engineered to overexpress polynucleotides encoding the transcription factor of the present invention or its functional homologue.
[0068] As used herein, the term "engineering the host cell to contain the polynucleotide encoding the protein of interest" means that the host cell of the present invention is equipped with the polynucleotide encoding the protein of interest, i.e., the host cell of the present invention has been engineered to contain the polynucleotide encoding the protein of interest. This can be achieved, for example, by transformation or transfection, or by any other suitable technique known in the art for introducing polynucleotides into host cells.
[0069] For example, procedures used to manipulate polynucleotide sequences encoding transcription factors and / or proteins of interest, promoters, enhancers, readers, etc., are well known to those skilled in the art and are described, for example, in J. Sambrook et al., Molecular Cloning: A Laboratory Manual (3rd edition), Cold Spring Harbor Laboratory, Cold Spring Harbor Laboratory Press, New York (2001).
[0070] Foreign or targeted polynucleotides, such as polynucleotides encoding overexpressed transcription factors or proteins of interest, can be inserted into chromosomes by various means, for example, homologous recombination or by using hybrid recombinases that specifically target the sequence at the integration site. The foreign or targeted polynucleotides are typically present in vectors ("insertion vectors"). These vectors are typically circular and are chained before use for homologous recombination. Alternatively, the foreign or targeted polynucleotides may be DNA fragments linked by fusion PCR or constructed synthetically, which are then recombined in host cells. In addition to homologous arms, vectors may also contain markers, origins of replication, and other elements suitable for selection or screening. Heterogeneous recombination, resulting in random or untargeted integration, is also possible. Heterogeneous recombination refers to the recombination between DNA molecules with significantly different sequences. Recombination methods are well known in the art and are described, for example, in Boer et al., Appl Microbiol Biotechnol (2007) 77:513-523. For genetic manipulation of yeast cells, see also Principles of Gene Manipulation and Genomics by Primrose and Twyman (7th edition, Blackwell Publishing 2006).
[0071] Polynucleotides encoding overexpressed transcription factors and / or proteins of interest may be present on expression vectors. Such vectors are known in the art. In expression vectors, the promoter is located upstream of the gene encoding the heterologous protein and regulates gene expression. Multicloning vectors are particularly useful due to their multicloning sites. For expression, the promoter is generally located upstream of the multicloning site. Vectors for incorporating polynucleotides encoding transcription factors and / or proteins of interest may be constructed either by first preparing a DNA construct containing the entire DNA sequence encoding the transcription factor and / or protein of interest, and then inserting this construct into a suitable expression vector, or by sequentially inserting DNA fragments containing genetic information for individual elements such as DNA-binding domains and activation domains, and then ligating them. As an alternative to restriction enzyme cleavage and ligation of fragments, DNA sequences may be inserted into vectors using recombination methods based on attachment sites (ATT) and recombinant enzymes. Such methods are described, for example, by Landy (1989) Ann. Rev. Biochem. 58:913-949 and are known to those skilled in the art.
[0072] The host cells according to the present invention can be obtained by introducing a vector or plasmid containing a target polynucleotide sequence into cells. Techniques for transfecting or transforming eukaryotic cells, or for transforming prokaryotic cells, are well known in the art. These include lipid vesicle-mediated uptake, heat shock-mediated uptake, calcium phosphate-mediated transfection (coprecipitation of calcium phosphate / DNA), viral infection, particularly viral infection using modified viruses such as modified adenoviruses, microinjection, and electroporation. Techniques for transforming prokaryotes include heat shock-mediated uptake, fusion of intact cells and bacterial protoplasts, microinjection, and electroporation. Techniques for transforming plants include Agrobacterium-mediated introduction, such as by Agrobacterium tumefaciens, rapidly sprayed tungsten or gold particle guns, electroporation, microinjection, and polyethylene glycol-mediated uptake. DNA may be single-stranded or double-stranded, linear or circular, unwound or supercoiled. For various techniques for transfecting mammalian cells, see, for example, Keown et al. (1990) Processes in Enzymology 185:527-537.
[0073] The phrase "culturing host cells under conditions suitable for overexpressing at least one polynucleotide encoding at least one transcription factor, and for overexpressing a protein of interest" means maintaining and / or growing eukaryotic host cells under conditions (e.g., temperature, pressure, pH, induction, growth rate, culture medium, duration, etc.) that are suitable or sufficient for obtaining the production of the desired compound (protein of interest), or for obtaining or overexpressing the transcription factor of the present invention.
[0074] Host cells obtained by transformation using transcription factor genes(s) and / or genes(s) of a protein of interest according to the present invention can preferably be cultured first under conditions that allow for efficient proliferation to a large number of cells without the burden of expressing recombinant proteins. When cells are prepared for the expression of a protein of interest, appropriate culture conditions are selected and optimized to produce the protein of interest.
[0075] For example, different promoters and / or copy and / or integration sites can be used for transcription factors and proteins of interest to control the timing and intensity of induction of the expression of the proteins of interest. For instance, transcription factors may be expressed first before the induction of the expression of the proteins of interest. This has the advantage that the transcription factors are already present at the start of translation of the proteins of interest. Alternatively, transcription factors and proteins of interest may be induced simultaneously.
[0076] An inductive promoter can be used that, upon application of an inductive stimulus, immediately becomes activated and directs the transcription of the gene under its control. Under growth conditions including an inductive stimulus, cells typically grow more slowly than under normal conditions, but since the culture medium has already grown to a large number of cells in the preceding stage, the overall culture system produces a large amount of recombinant protein. The inductive stimulus is preferably the addition of a suitable agent (e.g., methanol for the alcohol oxidase promoter) or the depletion of a suitable nutrient (e.g., methionine for the MET3 promoter). In addition, the addition of ethanol, methylamine, cadmium, or copper, as well as heat or osmotic pressure increasing agents, can induce expression depending on the promoter operably linked to the transcription factor and the protein(s) of interest.
[0077] It is preferable to culture the host cells(s) according to the present invention in a bioreactor under optimized growth conditions to obtain a cell density of at least 1 g / L, preferably at least 10 g / L, and more preferably at least 50 g / L. Achieving such biomolecule production yields is advantageous not only on a laboratory scale but also on an experimental or industrial scale.
[0078] According to the present invention, the protein of interest can be obtained in high yield, even when biomass is kept low, by overexpression of at least one transcription factor. Therefore, high specific yields, measured in mg of the protein of interest per gram of dry biomass, can be and are achievable in the range of 1 to 200, e.g., 50 to 200, e.g., 100 to 200, at laboratory, experimental, and industrial scales. The specific yield of production host cells according to the present invention preferably gives at least 1.1 times, more preferably at least 1.2 times, at least 1.3 times, or at least 1.4 times, compared to the expression of the product without overexpression of at least one transcription factor, and in some cases, an increase of more than 2 times may be shown.
[0079] The host cells according to the present invention can be tested for expression / secretion capacity or yield by measuring the titer of the protein of interest in the supernatant of the cell culture medium or in the cellular homogenate of the cells after homogenization, using standard tests such as ELISA, activity assay, HPLC, surface plasmon resonance (viacore), Western blotting, capillary electrophoresis (caliper), or SDS-PAGE.
[0080] Preferably, the host cells are cultured in a minimal medium containing a suitable carbon source, which further simplifies the isolation process. For example, the minimal medium contains salts containing available carbon sources (e.g., glucose, glycerol, ethanol, or methanol), macroelements (potassium, magnesium, calcium, ammonium, chlorides, sulfates, phosphates), and trace elements (salts of copper, iodine, manganese, molybdenum, cobalt, zinc, and iron, and boric acid).
[0081] In the case of yeast cells, the cells can be transformed using one or more of the expression vectors described above, crossed to form diploid strains, and cultured in a standard nutrient medium modified to be suitable for promoter induction, selection of transformants, or amplification of genes encoding the desired sequence. Many minimal media suitable for yeast growth are known in the art. Any of these media may be supplemented as needed with salts (e.g., sodium chloride, calcium, magnesium, and phosphates), buffers (e.g., HEPES, citrate, and phosphate buffer), nucleosides (e.g., adenosine and thymidine), antibiotics, trace elements, vitamins, and glucose, or equivalent energy sources. Any other necessary auxiliary substances, known to those skilled in the art, may also be included in appropriate concentrations. Culture conditions, such as temperature and pH, are known to those skilled in the art and have been previously used with host cells selected for expression. Cell culture conditions for other types of host cells are also known and can be readily determined by those skilled in the art. Descriptions of various culture media for microorganisms can be found, for example, in the American Society for Microbiology's handbook, "Manual of Methods for General Bacteriology" (Washington, D.C., USA, 1981).
[0082] Host cells can be cultured (e.g., maintained and / or proliferated) in a liquid medium, preferably continuously or intermittently by conventional culture methods, such as static culture, test tube culture, shaking culture (e.g., rotational shaking culture, shaking flask culture, etc.), aerated spinner culture, or fermentation. In some embodiments, cells are cultured in shaking flasks or deep-bottom well plates. In yet other embodiments, cells are cultured in a bioreactor (e.g., a bioreactor culture process). Culture processes include, but are not limited to, batch culture, fed-batch culture, and continuous culture. The terms “batch process” and “batch culture” refer to a closed system in which the composition of the medium, nutrients, auxiliary additives, etc., is set at the start of the culture and does not change during the culture; however, attempts may be made to control such factors, such as pH and oxygen concentration, to prevent excessive acidification of the medium and / or cell death. The terms “fed-batch process” and “fed-batch culture” refer to batch culture except that one or more substrates or auxiliary substances are added (e.g., gradually or continuously) as the culture progresses. The terms "continuous process" and "continuous culture" refer to a system in which a predetermined culture medium is continuously added to a bioreactor, and an equal amount of used or "acclimatized" medium is simultaneously removed, for example, to recover the desired product. A wide variety of such processes have been developed and are well known in the art.
[0083] In some embodiments, host cells are cultured for approximately 12–24 hours, while in other embodiments, host cells are cultured for approximately 24–36 hours, 36–48 hours, 48–72 hours, 72–96 hours, 96–120 hours, 120–144 hours, or more than 144 hours. In yet another embodiment, the culture is continued for a sufficient time to reach the desired production yield of the protein of interest.
[0084] The above method may further include a step of isolating the expressed protein of interest. If the protein of interest is secreted from the cells, it can be isolated and purified from the culture medium using state-of-the-art techniques. Secretion of the protein of interest from cells is generally preferred because the product is recovered from the culture supernatant rather than from a complex mixture of proteins produced when cells are destroyed and release intracellular proteins. Protease inhibitors such as phenylmethylsulfonyl fluoride (PMSF) may be useful in suppressing proteolysis during purification, and antibiotics may be included to prevent the growth of foreign contaminants. The composition may be concentrated, filtered, dialyzed, etc., using methods known in the art. Cells can be separated from the culture supernatant by centrifugation of the cell culture medium after fermentation / culture using a centrifuge or centrifuge tube. The supernatant can then be filtered from the concentrate using tangential flow filtration. Alternatively, cultured host cells may be destroyed sonically or mechanically (e.g., by high-pressure homogenization), enzymatically or chemically to obtain a cell extract containing the desired protein of interest, from which the protein of interest may be isolated and purified.
[0085] Isolation and purification methods for obtaining proteins of interest can be based on methods utilizing differences in solubility, such as salting out, solvent precipitation, and thermal precipitation; methods utilizing differences in molecular weight, such as size exclusion chromatography, ultrafiltration, and gel electrophoresis; methods utilizing differences in charge, such as ion exchange chromatography; methods utilizing specific affinity, such as affinity chromatography; methods utilizing differences in hydrophobicity, such as hydrophobic interaction chromatography and reversed-phase high-performance liquid chromatography; methods utilizing differences in isoelectric point (for example, isoelectric focusing electrophoresis may be used); and methods utilizing specific amino acids, such as IMAC (fixed metal ion affinity chromatography). If the protein of interest is expressed as an inactive and soluble inclusion body, the solubilized inclusion body needs to be refolded.
[0086] The isolated and purified protein of interest can be identified by conventional methods, such as Western blotting or specific assays for the activity of the protein of interest. The structure of the purified protein of interest can be determined by amino acid analysis, amino-terminal peptide sequencing, primary structure analysis, such as mass spectrometry, RP-HPLC, ion-exchange HPLC, or ELISA. It is preferable that the protein of interest can be obtained in large quantities and with high purity, and therefore meet the necessary requirements for use as an active ingredient in pharmaceutical compositions or as a feed additive or food additive.
[0087] As used herein, the term “isolated” means a substance in form or environment that does not exist in nature. Non-limiting examples of isolated substances include: (1) any non-natural substance; (2) any substance, including but not limited to any enzyme, mutant, nucleic acid, protein, peptide, or cofactor, that is at least partially isolated from one or more or all of the naturally occurring components that are associated with it in nature; (3) any substance that has been modified by human intervention compared to a naturally occurring substance, e.g., cDNA made from mRNA; or (4) any substance that has been modified by increasing the amount of the substance compared to other components that are associated with it in nature (e.g., recombinant production in a host cell; multiple copies of the gene encoding the substance; and the use of a promoter stronger than the promoter naturally associated with the gene encoding the substance).
[0088] The present invention further provides a method for producing a recombinant protein of interest using eukaryotic host cells, comprising the steps of (i) preparing host cells that have been engineered to overexpress at least one polynucleotide encoding at least one transcription factor (wherein the host cells further comprising a polynucleotide encoding a protein of interest, the transcription factor of the present invention comprising at least one DNA-binding domain and an activation domain); (ii) culturing the host cells under conditions suitable for overexpressing at least one polynucleotide encoding at least one transcription factor or a functional homolog thereof, and for overexpressing a protein of interest; and optionally (iii) isolating the protein of interest from the cell culture medium; optionally (iv) purifying the protein of interest; optionally (v) modifying the protein of interest; and optionally (vi) formulating the protein of interest.
[0089] Preferably, in step (i), the host cell is engineered to overexpress at least one polynucleotide encoding at least one transcription factor of the present invention, which includes a DNA-binding domain containing an amino acid as shown in SEQ ID NO: 1, or a functional homolog of the amino acid sequence shown in SEQ ID NO: 1 having at least 60% sequence identity with respect to the amino acid sequence shown in SEQ ID NO: 87.
[0090] In this context, the term "producing a recombinant protein of interest by / in a eukaryotic host cell" as used herein means that the recombinant protein of interest may be produced by using a eukaryotic host cell for the formation of a recombinant host cell. This allows the eukaryotic host cell to produce the recombinant protein of interest intracellularly and maintain the recombinant protein of interest inside the cell (intracellularly), or to secrete the recombinant protein of interest into the culture medium (extracellularly) (where the host cell is cultured). Thus, the protein of interest can be isolated from the culture medium (supernatant of the cell culture medium) or from the cellular homogenate of the cells after cell homogenization.
[0091] In this context, the term "modifying the protein of interest" means that the protein of interest is chemically modified. Many protein modification methods are known in the art. Proteins may be conjugated to carbohydrates or lipids. Proteins of interest may be PEGylated (chemically conjugated to polyethylene glycol) or HESlated (chemically conjugated to hydroxyethyl starch) for extension of their half-life. Proteins of interest may also be conjugated to other moieties, such as affinity domains for human serum albumin, for extension of their half-life. Proteins of interest may also be treated with proteases or under hydrolytic conditions for cleavage to form active components from a precursor sequence, or for removal of tags such as affinity tags for purification. Proteins of interest may also be conjugated to other moieties, such as toxins, radioactive moieties, or any other moieties. Proteins of interest may be further treated under conditions to form dimers, trimers, etc.
[0092] Furthermore, the term "formulating a protein of interest" refers to bringing the protein of interest into conditions that allow for longer storage. Many different methods known in the art are available for stabilizing proteins. By changing the buffer in which the protein of interest is present after purification and / or modification, the protein of interest can be brought into more stable conditions. Various buffers and additives known in the art, such as sucrose, mild surfactants, and stabilizers, can be used. The protein of interest may also be stabilized by lyophilization. Formulation for some proteins of interest can be carried out by forming a complex of the protein of interest with a lipid or lipoprotein, such as a polyplex. Some proteins may be co-formulated with other proteins.
[0093] Overexpression of the Msn4p transcription factors (see SEQ ID NOs. 15-27) of the present invention used in the method, recombinant host cells, and the present invention may increase the yield of the model protein scFv (SEQ ID NOs. 13) and / or vHH (SEQ ID NOs. 14) compared to host cells before engineering. The yield of the above model protein(s) may increase by at least 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 100%, 110%, 120%, 130%, 140%, 150%, 160%, 170%, 180%, 190%, 200%, 210%, 220%, 230%, 240%, 2 It may increase by 50%, 260%, 270%, 280%, 290%, 300%, 310%, 320%, 330%, 340%, 350%, 360%, 370%, 380%, 390%, 400%, 410%, 420%, 430%, 440%, 450%, 460%, 470%, 480%, 490%, or 500%. In this specification, the terms "0%, 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 100%, 200%, 300%, 400%, 500%, 600%, etc." refer to multiples of 1, 1.1, 1.2, 1.3, 1.4, 1.5, 1.6, 1.7, 1.8, 1.9, 2, 3, 4, 5, 6, etc. The prefix "times" indicates a multiple. "1x" means the whole, "2x" means twice that, and "3x" means three times that. Overexpression of the Pichia pastris native transcription factor Msn4p according to the present invention increases the yield of the model protein, preferably scFv (SEQ ID NO: 13), by at least 10%, for example, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 100%, 110%, 120%, 130%, 140%, 150%, 160%, 170%, 180%, compared to host cells before engineering manipulation. It can be increased by 190%, 200%, 210%, 220%, 230%, 240%, 250%, 260%, 270%, 280%, 290%, 300%, 310%, 320%, 330%, 340%, 350%, 360%, 370%, 380%, 390%, 400%, 410%, 420%, 430%, 440%, 450%, 460%, 470%, 480%, 490%, or 500%.Overexpression of the synthetic transcription factor synMsn4p of the present invention increases the yield of the model protein, preferably vHH (SEQ ID NO: 14), by at least 10%, for example, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 100%, 110%, 120%, 130%, 140%, 150%, 160%, 170%, 180%, 190%, compared to host cells before engineering. It can be increased by 200%, 210%, 220%, 230%, 240%, 250%, 260%, 270%, 280%, 290%, 300%, 310%, 320%, 330%, 340%, 350%, 360%, 370%, 380%, 390%, 400%, 410%, 420%, 430%, 440%, 450%, 460%, 470%, 480%, 490%, or 500%.
[0094] The method of the present invention, the recombinant host cell, and the polynucleotides encoding the transcription factors(s) and / or the protein of interest used are preferably integrated into the genome of the host cell. The term “genome” generally refers to the entire genetic information of an organism encoded in DNA (or RNA in some virus species). It can reside in chromosomes, on plasmids, on vectors, or both. Preferably, the polynucleotides encoding the transcription factors are integrated into the chromosomes of the cell.
[0095] Polynucleotides encoding transcription factors and proteins of interest can be recombined in host cells by ligating the relevant genes into separate vectors. It is possible to construct a single vector containing the genes, or two separate vectors, one containing the transcription factor gene and the other the protein of interest gene. These genes can be incorporated into the host cell genome by transforming the host cell using such vectors or vector sets. In some embodiments, the genes encoding the protein of interest are incorporated into the genome, while the genes encoding the transcription factors are incorporated into plasmids or vectors. In some embodiments, the genes encoding the transcription factors are incorporated into the genome, while the genes encoding the protein of interest are incorporated into plasmids or vectors. In some embodiments, the genes encoding both the protein of interest and the transcription factor are incorporated into the genome. In some embodiments, the genes encoding both the protein of interest and the transcription factor are incorporated into plasmids or vectors. When multiple genes encoding the protein of interest are used, some of the genes encoding the protein of interest may be incorporated into the genome, while the others may be incorporated into the same or different plasmids or vectors. When multiple genes encoding transcription factors are used, some of the transcription factor-encoding genes may be integrated into the genome, while others may be integrated into the same or different plasmids or vectors.
[0096] A polynucleotide encoding a transcription factor or its functional homolog may be incorporated into its native locus. “Natural locus” means the specific chromosomal location where the transcription factor-encoding polynucleotide is located, for example, the native locus of the gene encoding the transcription factor according to the present invention. However, in another embodiment, the transcription factor-encoding polynucleotide is located in a location other than its native locus within the host cell's genome and is ectopically incorporated. The term “ectopic integration” means the insertion of a nucleic acid into a location in the microbial genome other than its usual chromosomal locus, i.e., predetermined or random integration. Alternatively, the transcription factor or its functional homolog may be incorporated into its native locus and ectopically.
[0097] In yeast cells, polynucleotides encoding transcription factors and / or polynucleotides encoding proteins of interest can be inserted into desired gene loci, such as AOX1, GAP, ENO1, TEF, HIS4 (Zamir et al., Proc. NatL Acad. Sci. USA (1981) 78(6):3496-3500), HO (Voth et al. Nucleic Acids Res. 2001 June 15; 29(12): e59), TYR1 (Mirisola et al., Yeast 2007; 24: 761-766), His3, Leu2, Ura3 (Taxis et al., BioTechniques (2006) 40:73-78), Lys2, ADE2, TRP1, GAL1, ADH1, RGI1, etc., or into ribosomal RNA loci.
[0098] In other embodiments, polynucleotides encoding at least one transcription factor and / or polynucleotides encoding a protein of interest may be incorporated into the plasmid or vector. The terms “plasmid” and “vector” include autonomously replicating nucleotide sequences and genome-integrated nucleotide sequences. Those skilled in the art can use an appropriate plasmid or vector depending on the host cell being used.
[0099] Preferably, the plasmid is a eukaryotic cell expression vector, preferably a yeast expression vector.
[0100] Plasmids can be used for the transcription of cloned recombinant nucleotide sequences, i.e., the transcription of recombinant genes, and the translation of their mRNAs in a suitable host organism. Plasmids can also be used to integrate target polynucleotides into the host cell genome by methods known in the art, for example, as described by J. Sambrook et al., Molecular Cloning: A Laboratory Manual (3rd edition), Cold Spring Harbor Laboratory, Cold Spring Harbor Laboratory Press, New York (2001). A "plasmid" typically contains an origin of autonomous replication, a selection marker, several restriction enzyme cleavage sites, a suitable promoter sequence, and a transcription termination factor, whose components are operably linked to one another. The polypeptide coding sequence of interest is operably linked to transcriptional and translational regulatory sequences that result in polypeptide expression in the host cell.
[0101] Nucleic acids are "operably linked" when they are positioned to have a functional relationship with another nucleic acid sequence on the same nucleic acid molecule. For example, a promoter is operably linked to the coding sequence of a recombinant gene when it can influence the expression of that coding sequence.
[0102] Most plasmids exist in only one copy per bacterial cell. However, some plasmids exist in multiple copies. For example, plasmid ColE1 typically exists in 10 to 20 plasmid copies per chromosome in Escherichia coli (E. coli). When the nucleotide sequence of the present invention is contained within a plasmid, the plasmid may have 1 to 10, 10 to 20, 20 to 30, 30 to 100, or more copies per host cell. Multiple copy number plasmids allow cells to overexpress transcription factors.
[0103] Many suitable plasmids or vectors are known to those skilled in the art, and many are commercially available. Examples of suitable vectors are provided in Sambrook et al, eds., Molecular Cloning: A Laboratory Manual (2nd Ed.), Vols. 1-3, Cold Spring Harbor Laboratory (1989) and Ausubel et al, eds., Current Protocols in Molecular Biology, John Wiley & Sons, Inc., New York (1997).
[0104] The vector or plasmid of the present invention refers to a DNA construct that includes a yeast artificial chromosome, which may be genetically modified to contain a telomere sequence, a centromere sequence, and an origin of replication sequence, and to contain a heterologous DNA sequence (for example, a DNA sequence as large as 3000kb).
[0105] The vector or plasmid of the present invention also includes a bacterial artificial chromosome (BAC), which is a DNA construct that may contain an origin of replication sequence (Ori), one or more helicases (e.g., parA, parB, and parC), and may be genetically modified to contain a heterologous DNA sequence (e.g., a DNA sequence as large as 300kb).
[0106] Examples of plasmids that use yeast as a host include YIp vectors, YEp vectors, YRp vectors, YCp vectors (Yxp vectors are described, for example, in Romanos et al. 1992, Yeast. 8(6):423-488), pGPD-2 (described in Bitter et al., 1984, Gene, 32:263-274), pYES, pAO815, pGAPZ, pGAPZα, pHIL-D2, pHIL-S1, pPIC3.5K, pPIC9K, pPICZ, pPICZα, pPIC3K, pPINK-HC, pPINK-LC (all available from Thermo Fisher Scientific / Invitrogen), and pHWO10 (Waterham et al., 1997, Gene, (as described in 186:37-44), pPZeoR, pPKanR, pPUZZLE and pPUZZLE derivatives, e.g., pPM2d, pPM2aK21 or pPM2eH21 (Stadlmayr et al., 2010, J Biotechnol. 150(4):519-29.; Marx et al. 2009, FEMS Yeast Res. Examples include: 9(8):1260-70; the Golden PiCS system (consisting of the BB1, BB2, and BB3aK / BB3eH / BB3rN skeletons); pJ vectors (e.g., pJAN, pJAG, pJAZ and their derivatives; all available from Biogrammatics); pJexpress vectors; pD902, pD905, pD915, pD912 and their derivatives; pD12xx, pJ12xx (all available from ATUM / DNA2.0); pRG plasmids (Gnugge et al., 2016, Yeast 33:83-98); and 2 μm plasmids (e.g., Ludwig et al., 1993, Gene 132(1):33-40). Such vectors are publicly known and are described, for example, in Cregg et al., 2000, Mol Biotechnol. 16(1):23-52 or Ahmad et al. 2014., Appl Microbiol Biotechnol. 98(12):5301-17.Further suitable vectors can be readily prepared by applied molecular cloning techniques, as described, for example, by Lee et al. 2015, ACS Synth Biol. 4(9):975-986; Agmon et al. 2015, ACS Synth. Biol., 4(7):853-859; or Wagner and Alper, 2016, Fungal Genet Biol. 89:126-136. Furthermore, these and other suitable vectors are also available from Adgene, Inc. (Cambridge, MA, USA).
[0107] Preferably, the gene fragments of the transcription factor of the present invention are introduced using a BB1 plasmid in the Golden PiCS system by the use of specific restriction enzymes (Table 1). The constructed BB1 having the respective coding sequences can then be further processed in the Golden PiCS system to produce the required BB3 integration plasmid as described in Prielhofer et al. 2017.
[0108] The method, recombinant host cells, and polynucleotides encoding at least one transcription factor used in the present invention may encode heterologous or homologous transcription factors.
[0109] As used herein, the term “heterogeneous” means derived from cells or organisms (preferably yeast) or synthetic sequences having different genomic backgrounds. Therefore, “heterogeneous transcription factors” are derived from foreign origins (or species, e.g., Msn4p or synMsn4p of Saccharomyces cerevisiae) and are used for origins other than foreign origins (or species, e.g., Pichia pastris). The term “homogeneous” means derived from the same cells or organisms having the same genomic background. Therefore, “homogeneous transcription factors” are derived from the same origins (or species, e.g., Msn4p of Pichia pastris) and are used for the same origins (or species, e.g., Pichia pastris).
[0110] In general, overexpression can be achieved by any method known to those skilled in the art, as will be detailed later. Overexpression can be achieved by increasing the transcription / translation of a gene, for example, by increasing the copy number of the gene, or by altering or modifying a regulatory sequence. For example, overexpression can be achieved by introducing one or more copies of a polynucleotide encoding a transcription factor or functional homolog, operably ligated to a regulatory sequence (e.g., a promoter). For example, a gene can be operably ligated to a strong constitutive promoter to reach a high level of expression. Such a promoter may be an endogenous promoter or a recombinant promoter. Alternatively, a regulatory sequence can be removed to achieve constitutive expression. The native promoter of a given gene may be replaced with a heterologous promoter that increases gene expression or results in constitutive expression of the gene. For example, transcription factors may be overexpressed by host cells by more than 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 100%, 200%, or 300% compared to host cells cultured under the same conditions before engineering. Furthermore, overexpression can also be achieved by modifying proteins involved in the transcription of a gene and / or the translation of a gene product (e.g., regulatory proteins, suppressors, enhancers, transcription activators, etc.) by, for example, modifying the chromosomal location of a particular gene, altering nucleic acid sequences adjacent to a particular gene, such as ribosome binding sites or transcription termination factors, or by using antisense nucleic acid molecules to block the expression of repressor proteins, or by any other conventional means of deregulating the expression of a particular gene that is conventional in the art, including but not limited to deletions or mutations in the genes of transcription factors that normally repress the expression of the gene to be overexpressed. Extending the lifetime of mRNA can also improve expression levels.For example, the half-life of mRNA can be extended using specific termination region regions (Yamanishi et al., Biosci. Biotechnol. Biochem. (2011) 75:2234 and U.S. Patent No. 2013 / 0244243). If a gene contains multiple copies, the gene may be located in a plasmid with a variable copy number or may be incorporated and amplified within a chromosome. If the host cell does not contain the gene encoding the transcription factor, the gene can be introduced into the host cell for expression. In this case, "overexpression" means expressing the gene product using any method known to those skilled in the art.
[0111] Those skilled in the art will particularly refer to Martin et al. (Bio / Technology 5, 137-146 (1987)), Guerrero et al. (Gene 138, 35-41 (1994)), Tsuchiya and Morinaga (Bio / Technology 6, 428-430 (1988)), Eikmanns et al. (Gene 102, 93-98 (1991)), European Patent No. 0472869, U.S. Patent No. 4,601,893, Schwarzer and Puhler (Bio / Technology 9, 84-87 (1991)), Reinscheid et al. (Applied and Environmental Microbiology 60, 126-132 (1994)), LaBarre et al. (Journal of Bacteriology 175, 1001- Relevant explanations can be found in 1007 (1993), International Publication No. 96 / 15246, Malumbres et al. (Gene 134, 15-24 (1993)), Japanese Patent Publication No. 10-229891, Jensen and Hammer (Biotechnology and Bioengineering 58, 191-195 (1998)), and Makrides (Microbiological Reviews 60, 512-538 (1996)), as well as in well-known texts on genetics and molecular biology.
[0112] Therefore, the methods of the present invention, recombinant host cells, and the overexpression of polynucleotides encoding heterologous transcription factors used in their application can be achieved by replacing or modifying regulatory elements operably linked to the polynucleotide encoding the heterologous transcription factor. In this context, a “regulatory element” is a segment of a nucleic acid molecule that can increase or decrease the expression of a particular gene in an organism. Positive regulatory elements can increase expression, while negative regulatory elements can decrease expression. Regulatory sequences (elements) include, for example, promoters, enhancers, silencers, polyadenylation signals, transcription termination factors (termination factor sequences), coding sequences, and internal ribosome entry sites (IRESs). Positive regulatory sequences may include, but are not limited to, enhancers. Negative regulatory sequences may include, but are not limited to, silencers. In this context, the exchange of regulatory sequences means exchanging the native termination factor sequence of the heterologous transcription factor for a more efficient termination factor sequence, or exchanging the coding sequence of the heterologous transcription factor for a codon-optimized coding sequence (codon optimization is performed according to the codon usage frequency of the host cell), or exchanging the native positive regulatory element of the heterologous transcription factor for a more efficient regulatory element.
[0113] The method of the present invention, recombinant host cells, and the overexpression of polynucleotides encoding heterologous transcription factors used in the present invention can further be achieved by introducing one or more copies of the polynucleotide encoding heterologous transcription factors into host cells under promoter control.
[0114] As used herein, the term “promoter” refers to a region that promotes the transcription of a particular gene. A promoter typically increases the amount of recombinant product expressed from a nucleotide sequence compared to the amount expressed in the absence of a promoter. A promoter derived from one organism can be used to enhance the expression of recombinant products from sequences derived from another organism. Promoters can be incorporated into the chromosomes of host cells by homologous recombination using methods known in the art (e.g., Datsenko et al, Proc. Natl. Acad. Sci. USA, 97(12): 6640-6645 (2000)). Furthermore, a single promoter element can increase the amount of product expressed from multiple sequences attached in series. Thus, a single promoter element can enhance the expression of one or more recombinant products. The activity of a promoter can be evaluated by its transcription efficiency. This can be determined directly by measuring the amount of mRNA transcription from the promoter, for example by Northern blotting or quantitative PCR, or indirectly by measuring the amount of gene product expressed from the promoter.
[0115] A promoter may be either an "inducible promoter" or a "constitutive promoter." An "inducible promoter" is a promoter that can be induced by the presence or absence of a specific factor, while a "constitutive promoter" is a promoter that is always active regardless of the inducer, and therefore enables the continuous transcription of the associated gene or group of genes.
[0116] In a preferred embodiment, the transcription of both the transcription factor and the nucleotide sequence encoding the protein of interest is driven by an inductive promoter. In another preferred embodiment, the transcription of both the transcription factor and the nucleotide sequence encoding the protein of interest is driven by a constitutive promoter. In yet another preferred embodiment, the transcription of the nucleotide sequence encoding the transcription factor is driven by a constitutive promoter, and the transcription of the nucleotide sequence encoding the protein of interest is driven by an inductive promoter. In yet another preferred embodiment, the transcription of the nucleotide sequence encoding the transcription factor is driven by an inductive promoter, and the transcription of the nucleotide sequence encoding the protein of interest is driven by a constitutive promoter. As an example, the transcription of the nucleotide sequence encoding the transcription factor may be driven by a constitutive GAP promoter, and the transcription of the nucleotide sequence encoding the protein of interest may be driven by an inductive AOX promoter. In one embodiment, the transcription of the transcription factor and the nucleotide sequence encoding the protein of interest are driven by the same or similar promoters in terms of promoter activity, promoter regulation, and / or expression behavior. In another embodiment, the transcription of nucleotide sequences encoding transcription factors and proteins of interest is driven by different promoters in terms of promoter activity, promoter regulation, and / or expression behavior.
[0117] Suitable promoter sequences for use in yeast host cells are described in Mattanovich et al., Methods Mol. Biol. (2012). As described in 824:329-58, this includes promoters and variants of glycolytic enzymes such as triose phosphate isomerase (TPI), 3-phosphoglycerate kinase (PGK), glucose-6-phosphate isomerase (PGI), glyceraldehyde-3-phosphate dehydrogenase (GAPDH or GAP), promoters of lactase (LAC) and galactosidase (GAL), translation elongation factor promoter (PTEF), and promoters of Pichia pastris enolase 1 (ENO1), triose phosphate isomerase (TPI), ribosomal subunit proteins (RPS2, RPS7, RPS31, RPL1), alcohol oxidase promoter (AOX) or its modified variants, formaldehyde dehydrogenase promoter (FLD), isocitrate lyase promoter (ICL), α - This includes the promoters for ketoisocaproate decarboxylase (THI), heat shock protein family members (SSA1, HSP90, KAR2), 6-phosphogluconate dehydrogenase (GND1), phosphoglycerate mutase (GPM1), transketolase (TKL1), phosphatidylinositol synthase (PIS1), iron dioxide oxidoreductase (FET3), high affinity iron permease (FTR1), inhibitory alkaline phosphatase (PHO8), N-myristoyltransferase (NMT1), pheromone-responsive transcription factor (MCM1), ubiquitin (UBI4), single-stranded DNA endonuclease (RAD2), the promoter for the major ADP / ATP carrier of the inner mitochondrial membrane (PET9) (International Publication No. 2008 / 128701), and formate dehydrogenase (FDH) promoter.Further suitable promoters are described by Prielhofer et al. 2017 (BMC Syst Biol. 11(1):123.), Gasser et al. 2015 (Microb Cell Fact. 14:196.), Portela et al. 2017 (ACS Synth Biol. 6(3):471-484), or Vogl et al. 2016 (ACS Synth Biol. 5(2):172-86). The AOX promoter can be induced by methanol and repressed by glucose, for example.
[0118] Further examples of suitable promoters include those for Saccharomyces cerevisiae enolase (ENOI-1), galactokinase (GAL1), alcohol dehydrogenase / glyceraldehyde-3-phosphate dehydrogenase (ADH1, ADH2 / GAP), triose phosphate isomerase (TPI), metallothionein (CUP1), 3-phosphoglycerate kinase (PGK), and the maltase gene (MAL) promoter.
[0119] Other useful promoters for yeast host cells are described by Romanos et al, 1992, Yeast 8:423-488.
[0120] Each coding sequence of the heterologous transcription factor of the present invention (e.g., synMsn4p) can be incorporated into an integration plasmid, preferably BB3, together with the GAP promoter.
[0121] The methods of the present invention, recombinant host cells, and the overexpression of polynucleotides encoding allogeneic transcription factors used in the present invention can be achieved by using a promoter that drives the expression of the polynucleotide encoding the allogeneic transcription factor. Higher expression levels can be achieved by replacing the endogenous / native promoter, which is operably ligated to the endogenous allogeneic transcription factor, with another, more potent promoter. Such promoters may be inductive or constitutive. Modification and / or replacement of the endogenous promoter can be carried out by mutation or homologous recombination using methods known in the art.
[0122] Each coding sequence of the allogeneic transcription factor of the present invention (for example, the natural Msn4p of P. pastris when expressed in Pichia pastris) can be incorporated into an integration plasmid, such as BB3, together with a strong constitutive or inducible promoter, such as the GAP promoter, pTHI11, pSBH17, or pPOR1.
[0123] Overexpression of polynucleotides encoding transcription factors can be achieved by genetically modifying their endogenous regulatory regions, as described by other methods known in the art, for example, by Marx et al., 2008 (Marx, H., Mattanovich, D. and Sauer, M. Microb Cell Fact 7 (2008): 23) and Pan et al., 2011 (Pan et al., FEMS Yeast Res. (2011) May; (3):292-8), such methods include, for example, the incorporation of recombinant promoters that increase the expression of transcription factors. Transformation is described in Cregg et al. (1985) Mol. Cell. Biol. 5:3376-3385.
[0124] Accordingly, the present invention may include the method of the present invention, recombinant host cells, and the overexpression of a polynucleotide encoding an allogeneic transcription factor used in use, which can be further achieved by replacing or modifying a regulatory sequence operably linked to the polynucleotide encoding the allogeneic transcription factor.
[0125] In this context, replacing regulatory sequences means, for example, replacing the native termination factor sequence of the allogeneic transcription factor with a more efficient termination factor sequence, or replacing the coding sequence of the allogeneic transcription factor with a codon-optimized coding sequence (codon optimization is performed according to the codon usage frequency of the host cell), or replacing the native positive regulatory element of the allogeneic transcription factor with a more efficient positive regulatory element.
[0126] In this context, the term “modifying a regulatory sequence” as used herein means the addition of another positive regulatory sequence or the deletion of a negative regulatory sequence. Therefore, modifying a regulatory sequence means introducing / adding another positive regulatory sequence that is not present in the native expression cassette of the homogeneous / heterogeneous transcription factor (element), or deleting a negative regulatory sequence (element) that is normally present in the native expression cassette of the homogeneous / heterogeneous transcription factor. A native expression cassette means a protein-coding sequence, including its 5' flanking and 3' flanking sequences, such as promoters, termination factors, and polyadenylation signals, that is naturally present in cells and not artificially created by humans using recombinant gene technology, and that is involved in the positive or negative regulation of the protein's expression. Heterogeneous and homogeneous native expression cassettes can exist. If an expression cassette from one species is introduced into another species and the expression of the protein encoded by that native expression cassette still occurs, then this native expression cassette is considered a heterogeneous native expression cassette.
[0127] The method of the present invention, recombinant host cells, and the overexpression of polynucleotides encoding allogeneic transcription factors used in the present invention can further be achieved by introducing one or more copies of polynucleotides encoding allogeneic transcription factors into host cells under the control of a promoter.
[0128] The method, recombinant host cells, and overexpression of a polynucleotide encoding at least one transcription factor used in the present invention comprises: i) replacing the native promoter of the allogeneic transcription factor with a different promoter operably linked to the polynucleotide encoding the allogeneic transcription factor, e.g., a stronger promoter; ii) replacing the native termination factor sequence of the heterogeneic and / or allogeneic transcription factor with a more efficient termination factor sequence; and iii) replacing the coding sequence of the heterogeneic and / or allogeneic transcription factor with a codon-optimized coding sequence (e.g., optimized for mRNA stability or half-life, or the most... This can be achieved by: iv) replacing a codon with a codon that is more frequently used (codon optimization here is performed according to the codon usage frequency of the host cell), for example; iv) replacing a native positive regulatory element of the heterogeneous and / or homogeneous transcription factor with a more efficient regulatory element; v) introducing another positive regulatory element that is not present in the native expression cassette of the homogeneous transcription factor; vi) deleting a negative regulatory element that is normally present in the native expression cassette of the homogeneous transcription factor; or vii) introducing one or more copies of a polynucleotide encoding a heterogeneous and / or homogeneous transcription factor or a combination thereof.
[0129] The present invention may further include a method, recombinant host cell, and transcription factor(s) used in use, comprising an amino acid sequence as shown in SEQ ID NOs: 15-27, or a functional homolog of an amino acid sequence as shown in SEQ ID NOs: 15 having at least 11% sequence identity to the amino acid sequence as shown in SEQ ID NOs: 15. In further embodiments, the present invention may further include a method, recombinant host cell, and transcription factor(s) used in use, comprising an amino acid sequence as shown in SEQ ID NOs: 15-27, or a functional homolog of an amino acid sequence as shown in SEQ ID NOs: 15 having at least 11% sequence identity to the amino acid sequence as shown in SEQ ID NOs: 15, for example, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 98%, or even 100% sequence identity to the amino acid sequence as shown in SEQ ID NOs: 15.
[0130] The methods, recombinant host cells, and transcription factors(s) used in the present invention may further include any nuclear localization signal (NLS). Therefore, the transcription factors of the present invention may include a DNA-binding domain as otherwise described herein, an activation domain as otherwise described herein, and any NLS. Any NLS in this specific context may include a synthetic NLS (e.g., SEQ ID NO: 86), a viral NLS, or an NLS of the transcription factors of the present invention or any other protein of any species described herein. An NLS is an amino acid sequence that "tags" a protein for translocation into the cell nucleus via nuclear transport. Typically, an NLS consists of one or more short sequences of positively charged lysine or arginine exposed on the protein surface. Amino acid sequences such as those shown in SEQ ID NO: 85 (Predicted NLS of Pichia pastrix Msn4p: EPRKKETKQRKRAK; Best prediction by SeqNLS (score > 0.89); http: / / mleg.cse.sc.edu / seqNLS / MainProcess.cgi) or SEQ ID NO: 86 (NLS of synMsn4p: PKKKRKV) are preferred as NLS in the present invention.
[0131] The nuclear localization signal may be from a homogeneous or heterogeneous NLS. In this context, the term “heterogeneous NLS” refers to an NLS derived from an exotic origin (or species, e.g., NLS from Saccharomyces cerevisiae or human NLS; see also Weninger et al. 2015. FEMS Yeast Res. 15:7) or a synthetic sequence, and is used for origins other than exotic origins (or species, e.g., Pichia pastris). “Homogeneous NLS” refers to a sequence derived from the same origin (or species, e.g., NLS from Pichia pastris), and is used for the same origin (or species, e.g., Pichia pastris).
[0132] The present invention may further comprise the method, recombinant host cells, and transcription factors(s) used in the use of the transcription factors(s) that do not stimulate the promoter used for the expression of the protein of interest. This means that the transcription factors of the present invention have no effect whatsoever on the promoter of the protein of interest. Rather, they act on the promoters of different proteins other than the protein of interest. In this context, the terms “non-stimulating” or “unstimulating” mean having no effect whatsoever on the promoter of the protein of interest, or having only a minor effect on the promoter of the protein of interest, and thus resulting in a small increase in the yield of the protein of interest of about 10% or less, for example, an increase in the yield of the protein of interest of 1%, 2%, 3%, 4%, 5%, 6%, 7%, 8%, 9%, or 10%.
[0133] The methods, recombinant host cells, and uses of the present invention use eukaryotic cells as host cells. As used herein, “host cell” refers to a cell capable of expressing a protein and, optionally, secreting a protein. Such a host cell is to be applied to the methods of the present invention. For this purpose, in order for the host cell to overexpress at least one polynucleotide encoding at least one transcription factor, the polynucleotide sequence encoding the transcription factor is either present in the cell or introduced into the cell. Examples of eukaryotic cells include, but are not limited to, vertebrate cells, mammalian cells, human cells, animal cells, invertebrate cells, plant cells, nematode cells, insect cells, stem cells, fungal cells, or yeast cells.
[0134] Preferably, the eukaryotic host cell is a fungal cell. More preferably, it is a yeast host cell. Examples of yeast cells include the genera Saccharomyces (e.g., Saccharomyces cerevisiae, Saccharomyces kluyveri, Saccharomyces uvarum), Komagataella (Komagataella pastoris, Komagataella pseudopastoris, or Komagataella fafi), Kluyveromyces (e.g., Kluyveromyces lactis, Kluyveromyces marxianus), Candida (e.g., Candida utilis, Candida cacaoi), and Geotrium (e.g., Geotrium fermentus). Examples include, but are not limited to, *Fermentans*, as well as *Hanzenula polymorpha* and *Yarrowia liporitica*.
[0135] In a preferred embodiment, Pichia is of particular interest. Pichia includes many species, including species such as Pichia pastoris, Pichia methanolica, Pichia kluyveri and Pichia angusta. Most preferred is the Pichia pastoris species.
[0136] The species that was formerly Pichia pastoris has been classified and renamed Komagataella pastoris, Komagataella phaffii and Komagataella pseudopastoris. Therefore, Pichia pastoris is synonymous with both Komagataella pastoris, Komagataella phaffii and Komagataella pseudopastoris.
[0137] Examples of Pichia pastoris strains useful in the present invention are X33 and its subtypes GS115, KM71, KM71H; CBS7435(mut+) and its subtype CBS7435mut S , CBS7435mut S ΔArg, CBS7435mut S ΔHis, CBS7435mut S ΔArgΔHis, CBS7435mut S PDI + , CBS704(=NRRL Y-1603=DSMZ70382), CBS2612(=NRRL Y-7556), CBS9173-9189 and DSMZ 70877 and mutants thereof. These yeast strains are available from industrial product suppliers, or cell depositories, such as the American Type Culture Collection (ATCC), the "Deutsche Sammlung von Mikroorganismen und Zellkulturen" (DSMZ) in Braunschweig, Germany, or the "Centraalbureau voor Schimmelcultures" (CBS) in Utrecht, the Netherlands.
[0138] In a further preferred embodiment, the yeast host cells are selected from the group consisting of Pichia pastris (Chomagataera species), Hanzenula polymorpha, Trichoderma liesei, Aspergillus niger, Saccharomyces cerevisiae, Kluiveromyces lactis, Yarowia liporitica, Pichia metanolica, Candida boidini, Chomagataera species, and Schizosaccharomyces pombe. These yeast strains are available from cell collections, such as the United States Cell Lineage Preservation Center (ATCC), the German Microbial Cell Culture Collection (DSMZ) in Braunschweig, Germany, or the Netherlands Westerdijk Institute for Mycological Diversity (CBS) in Utrecht, Netherlands.
[0139] The present invention further includes the fact that the method, recombinant host cells, and recombinant protein of interest used in the use of the present invention may be enzymes. Preferred enzymes can be used for industrial applications, for example, in the manufacture of detergents, starches, fuels, textiles, pulp and paper, oils, personal care products, or, for example, in bread making, organic synthesis, etc. (See Kirk et al., Current Opinion in Biotechnology (2002) 13:345-351).
[0140] The present invention further includes the fact that the recombinant protein of interest may be a therapeutic protein. The protein of interest may be, but is not limited to, a biopharmaceutical substance such as an antigen-binding protein, an antibody or antibody fragment, or an antibody-derived backbone, single-domain antibodies and their derivatives, an affinity backbone not derived from other antibodies, such as an antibody mimetic, a growth factor, a hormone, or a vaccine, as described in more detail herein.
[0141] Examples of such therapeutic proteins include, but are not limited to, insulin, insulin-like growth factor, human growth hormone, tissue plasminogen activator, cytokines such as interleukins, such as IL-1, IL-2, IL-3, IL-4, IL-5, IL-6, IL-7, IL-8, IL-9, IL-10, IL-11, IL-12, IL-13, IL-14, IL-15, IL-16, IL-17, IL-18, interferon (IFN)α, IFNβ, IFNγ, IFNω, or IFNτ, tumor necrosis factor (TNF), TNFα, and TNFβ, TRAIL; granulocyte colony-stimulating factor (G-CSF), granulocyte-macrophage colony-stimulating factor (GM-CSF), macrophage colony-stimulating factor (M-CSF), monocyte chemotactic protein 1, and vascular endothelial growth factor.
[0142] Further examples of therapeutic proteins include blood coagulation factors (VII, VIII, IX), Fusarium-derived alkaline proteases, calcitonin, CD4 receptor darbepoetin, deoxyribonuclease (pustular fibrosis), erythropoietin, eutropin (human growth hormone derivative), follicle-stimulating hormone (follitropin), gelatin, glucagon, glucocerebrosidase (Gaucher disease), Aspergillus niger-derived glucoamylase, Aspergillus niger-derived glucose oxidase, gonadotropins, growth factors (G-CSF, GM-CSF), growth hormone (somatotropin), hepatitis B vaccine, hirudin, human antibody fragments, human apolipoprotein A1, human calcitonin precursor, human collagenase IV, and human Examples include epidermal growth factor, human insulin-like growth factor, human interleukin-6, human laminin, human proapolipoprotein A1, human serum albumin, insulin, insulin and mutein, insulin, interferon-α and mutein, interferon-β, interferon-γ (mutein), interleukin-2, luteinizing hormone, monoclonal antibody 5T4, mouse collagen, OP-1 (osteogenesis and neuroprotection factor), oprelbequin (interleukin-11 agonist), organic phosphohydrolases, platelet-derived growth factor agonists, phytase, platelet-derived growth factor (PDGF), recombinant plasminogen activator G, staphylokinase, stem cell factors, tetanus toxin fragment C, tissue plasminogen activator, and tumor necrosis factor (see Schmidt, Appl Microbiol Biotechnol (2004) 65:363-372).
[0143] Preferably, the therapeutic protein is an antigen-binding protein. More preferably, the therapeutic protein comprises an antibody, an antibody fragment, or an antibody mimetic. Even more preferably, the therapeutic protein is an antibody or an antibody fragment.
[0144] In preferred embodiments, the protein is an antibody fragment. The term “antibody” is intended to include any polypeptide chain-containing molecular structure having a specific shape that fits and recognizes an epitope, where one or more non-covalent interactions stabilize the complex between the molecular structure and the epitope. Typical antibody molecules are immunoglobulins, and all types of immunoglobulins from all origins, e.g., humans, rodents, rabbits, cattle, sheep, pigs, dogs, other mammals, chickens, other birds, etc., i.e., IgG, IgM, IgA, IgE, IgD, IgY, etc., are considered “antibodies.” For example, antibody fragments may include, but are not limited to, Fv (a molecule containing VL and VH), single-chain Fv (scFv) (a molecule containing VL and VH linked by a peptide linker), Fab, Fab', F(ab')2, single-domain antibodies (sdAb) (a molecule containing a single variable domain and three CDRs), and their multivalent presenters. Antibodies or fragments thereof may be mouse antibodies, human antibodies, humanized antibodies, or chimeric antibodies, or fragments thereof. Examples of therapeutic proteins include antibodies, polyclonal antibodies, monoclonal antibodies, recombinant antibodies, antibody fragments, e.g., Fab', F(ab')2, Fv, scFv, di-scFv, bi-scFv, serial scFv, bi-specific serial scFv, sdAb, nanobody, V H and V L Examples include human antibodies, humanized antibodies, chimeric antibodies, IgA antibodies, IgD antibodies, IgE antibodies, IgG antibodies, IgM antibodies, intracellularly expressed antibodies, diabodies, tetrabodies, minibodies, or monobodies. Preferably, the antibody fragment is scFv (SEQ ID NO: 13) and / or vHH (SEQ ID NO: 14). Antibody mimetic refers to an organic compound that binds to an antigen but is not structurally related to an antibody. Such antibody mimetic refers to an artificial peptide or protein having a molecular weight of about 3 to 20 kDa, such as affibody molecules, affilins, affimers, affitins, alphabodies, antikalin, avimers, artificial ankyrin repeat proteins (DARPins), monobodies, and nanoCLAMPs, as known in the prior art.
[0145] The proteins of interest may also be food additives. Food additives are proteins used as nutritional supplements, dietary supplements, or digestive aids in foods, feed, or cosmetics, etc. Foods may be, for example, bouillon, desserts, cereal bars, sweets, sports drinks, diet products, or other nutritional products. "Food" means natural or artificial diets, or components of such diets that are intended or suitable to be eaten, ingested, or digested by humans.
[0146] The proteins of interest may also be feed additives. Examples of enzymes that can be used as feed additives include phytase, xylanase, and β-glucanase.
[0147] The methods, recombinant host cells, and uses of the present invention may further include the steps of overexpressing, or engineering, a host cell to overexpress, at least one polynucleotide encoding at least one endoplasmic reticulum (ER) accessory protein. In this context, the term "ER" refers to the "endoplasmic reticulum." Preferably, by adding and overexpressing at least one polynucleotide encoding at least one endoplasmic reticulum accessory protein in the host cell, the yield of the recombinant protein of interest is increased compared to host cells that overexpress at least one polynucleotide encoding at least one transcription factor, but do not overexpress at least one polynucleotide encoding at least one endoplasmic reticulum accessory protein.
[0148] As used herein, the term "at least one polynucleotide encoding at least one endoplasmic reticulum (ER) accessory protein" means one polynucleotide encoding one ER accessory protein, two polynucleotides encoding at least two ER accessory proteins, three polynucleotides encoding three ER accessory proteins, and so on.
[0149] The term “endoplasmic reticulum accessory proteins” refers to chaperones, co-chaperones, and / or nucleotide exchange factors. As used herein, the term “chaperone” refers to polypeptides that assist in the folding, unfolding, assembly, or disassembly of other polypeptides. Chaperones are proteins involved in the correct folding or unfolding and transport of newly translated eukaryotic intracellular and secretory proteins. There are many different families of chaperones, each acting to assist in protein folding in a different way. Endoplasmic chaperones and cytoplasmic chaperones exist.
[0150] Cytoplasmic chaperones in yeast cells include, but are not limited to, Ssa1p, Ssa2p, Ssa3p, Ssa4p, Ssb1p, Ssb2p, Sse1p, and Sse2p, and refer to the Hsp70 system. Ssa1-4p are involved in the folding of newly synthesized proteins, as well as the transport of intermediate proteins to the endoplasmic reticulum and mitochondria. Ssb1p and Ssb2p are involved in the folding of nascent chains bound to ribosomes, and Sse1p and Sse2p act as nucleotide exchange factors for Ssap and Ssbp. Ydj1p and Sis1p belong to the Hsp40 system in yeast and interact with non-native polypeptides as co-chaperones, triggering ATP hydrolysis by Ssa1-4p and participating in transmembrane protein transport. Snl1p, Fes1p, and Cns1p are other co-chaperones of Ssa1-4p (Chang et al., Cell 128 (2007)). In this context, the term “co-chaperone” refers to a protein that assists a chaperone in protein folding and other functions. Co-chaperones are non-client binding molecules that assist in protein folding mediated by Hsp70 and Hsp90.
[0151] In yeast cells, endoplasmic reticulum chaperones include, but are not limited to, Kar2p, which refers to the Hsp70 system or Pdi1p. Kar2p binds to unassembled / misfolded endoplasmic reticulum protein subunits and is involved in protein translocation to the endoplasmic reticulum by regulating the unfolded protein response (UPR). It interacts with its co-chaperones, such as Lhs1p, Sil1p, Erj5p, Sec63p, Scj1p, Jem1p, or others known in the art. Lhs1p and Sil1p refer to nucleotide exchange factors of Kar2p and belong to the Hsp70 system (Chang et al., Cell 128 (2007)). In this context, the term "nucleotide exchange factor" refers to a protein that stimulates the exchange (substitution) of nucleoside diphosphate (ADP, GDP) with nucleoside triphosphate (ATP, GTP) bound to other proteins (preferably chaperones). Erj5p, Sec63, and Scj1 belong to the Hsp40 type protein group. Erj5p is, for example, a type I membrane protein with a J domain; it is required to maintain the folding ability of the endoplasmic reticulum; and deletion of the non-essential ERJ5 gene results in a constitutively induced endoplasmic reticulum stress response (Mehnert et al., Molecular biology of the cell, 26 (2014)).
[0152] At least one endoplasmic reticulum (ER) accessory protein may be obtained from Pichia pastris (Komagataera pastris or Komagataera fafi), Hanzenula polymorpha, Trichoderma liesei, Saccharomyces cerevisiae, Kluiveromyces lactis, Yarowia liporitica, Candida boidini, Aspergillus niger, preferably Pichia pastris (Komagataera pastris or Komagataera fafi), for additional overexpression or for engineering host cells to induce additional overexpression. The closest homologs from other eukaryotic species may also be obtained for at least one ER accessory protein.
[0153] Preferably, the endoplasmic reticulum accessory protein of the present invention, which is additionally overexpressed in the host cell, has a functional homolog thereof having at least 70%, for example, at least 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or even 100% sequence identity with the amino acid sequence shown in SEQ ID NO: 28 (Kar2p of Pichia pastris). Preferably, the functional homologs of SEQ ID NO: 28 are SEQ ID NOs: 29-36. Therefore, the endoplasmic reticulum accessory protein of the present invention, which is additionally overexpressed in the host cell, has an amino acid sequence as shown in SEQ ID NOs. 28-36. An endoplasmic reticulum accessory protein having an amino acid sequence as shown in SEQ ID NO. 28 is preferred. Preferably, the accessory protein is not identical to the transcription factor of the present invention as described above, nor is it identical to the protein of interest.
[0154] When a polynucleotide encoding at least one transcription factor is introduced under the control of a promoter by a vector or plasmid, the polynucleotide encoding an additional endoplasmic reticulum (ER) accessory protein may be incorporated under the control of the same promoter on the same vector or plasmid, or under the control of a different promoter (Msn4p under one promoter, Kar2p under a different promoter). When a polynucleotide encoding at least one transcription factor is introduced under the control of a promoter by a vector or plasmid, the polynucleotide encoding an additional ER accessory protein may be incorporated simultaneously or sequentially (one at a time) on different vectors or plasmids. If both the polynucleotide encoding at least one transcription factor and the polynucleotide encoding an additional ER accessory protein can be introduced on different vectors or plasmids, it is preferable to use one plasmid containing only the at least one transcription factor and another plasmid containing an overexpression cassette for at least one additional ER accessory protein.
[0155] When one or more copies of a polynucleotide encoding at least one transcription factor are introduced under the control of a promoter by a vector or plasmid, additional polynucleotides encoding endoplasmic reticulum (ER) accessory proteins may be incorporated into the same vector or plasmid under the control of the same promoter or different promoters (one or more copies of Msn4p under the control of one promoter, and one or more copies of Kar2p under the control of different promoters). When one or more copies of a polynucleotide encoding at least one transcription factor are introduced under the control of a promoter by a vector or plasmid, additional polynucleotides encoding ER accessory proteins may be incorporated into different vectors or plasmids simultaneously or sequentially (one at a time).
[0156] It is hypothesized that overexpression of additional endoplasmic reticulum accessory proteins could ensure that the protein of interest folds correctly within the endoplasmic reticulum, thereby further increasing the yield of the protein of interest.
[0157] Overexpression of the Msn4p transcription factor(s) and the first Kar2p accessory protein(s) of the present invention increased the yield of the model protein by at least 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 100%, 110%, 120%, 130%, 140%, 150%, 160%, 170%, 180%, and 19% compared to host cells before engineering. It can be increased by 0%, 200%, 210%, 220%, 230%, 240%, 250%, 260%, 270%, 280%, 290%, 300%, 310%, 320%, 330%, 340%, 350%, 360%, 370%, 380%, 390%, 400%, 410%, 420%, 430%, 440%, 450%, 460%, 470%, 480%, 490%, or 500%. Overexpression of the Pichia pastris native (homologous) transcription factor Msn4p and the Pichia pastris first endoplasmic reticulum accessory protein Kar2p of the present invention increases the yield of the model protein, preferably vHH (SEQ ID NO: 14), by at least 40%, for example, 50%, 60%, 70%, 80%, 90%, 100%, 110%, 120%, 130%, 140%, 150%, 16% compared to host cells before engineering. It can be increased by 0%, 170%, 180%, 190%, 200%, 210%, 220%, 230%, 240%, 250%, 260%, 270%, 280%, 290%, 300%, 310%, 320%, 330%, 340%, 350%, 360%, 370%, 380%, 390%, 400%, 410%, 420%, 430%, 440%, 450%, 460%, 470%, 480%, 490%, or 500%. Overexpression of the synthetic transcription factor synMsn4p of the present invention and the first endoplasmic reticulum accessory protein Kar2p of Pichia pastris can increase the yield of the model protein, preferably vHH (SEQ ID NO: 14), by at least 30%, for example, 40%, 50%, 60%, 70%, 80%, 90%, 100%, 120%, 130%, 140%, 150%, 160%, 170%, 180%, 190%, 200%, 250%, 300%, 350%, 400%, or 500% compared to host cells before engineering.
[0158] The methods, recombinant host cells, and uses of the present invention may further include overexpressing, or engineering, the host cells to overexpress, at least two polynucleotides encoding at least two endoplasmic reticulum accessory proteins in the host cells.
[0159] When the present invention refers to two additional endoplasmic reticulum (ER) auxiliary proteins, this means the “first ER auxiliary protein” and the “second ER auxiliary protein.” When the present invention refers to three additional ER auxiliary proteins, this means the “first ER auxiliary protein,” the “second ER auxiliary protein,” and the “third ER auxiliary protein.” Preferably, by adding and overexpressing at least two polynucleotides encoding at least two ER auxiliary proteins in the host cell, the yield of the recombinant protein of interest is increased compared to host cells that overexpress at least one polynucleotide encoding at least one transcription factor, but do not overexpress at least two polynucleotides encoding at least two ER auxiliary proteins. By adding and overexpressing at least two polynucleotides encoding at least two endoplasmic reticulum (ER) accessory proteins in the host cells, the yield of the recombinant protein of interest is increased compared to host cells that overexpress at least one polynucleotide encoding at least one transcription factor and at least one polynucleotide encoding at least one additional ER accessory protein, but do not overexpress at least two polynucleotides encoding at least two ER accessory proteins.
[0160] Preferably, the first endoplasmic reticulum accessory protein has a functional homologue having at least 70%, for example 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or even 100% sequence identity with the amino acid sequence shown in SEQ ID NO: 28 (Kar2p of Pichia pastris). Preferably, the functional homologues of SEQ ID NO: 28 as the first endoplasmic reticulum accessory protein that are overexpressed in addition to the transcription factor are SEQ ID NOs: 29-36. Therefore, the first endoplasmic reticulum accessory protein of the present invention, which is additionally overexpressed in the host cell, has an amino acid sequence as shown in SEQ ID NOs. 28 is preferred as the first endoplasmic reticulum accessory protein.
[0161] Preferably, the second endoplasmic reticulum accessory protein has an amino acid sequence as shown in SEQ ID NO: 37, or at least 25% of the amino acid sequence as shown in SEQ ID NO: 37 (Lhs1p of Pichia pastris), for example, 26%, 27%, 28%, 29%, 30%, 31%, 32%, 33%, 34%, 35%, 36%, 37%, 38%, 39%, 40%, 41%, 42%, 43%, 44%, 45%, 46%, 47%, 48%, 49%, 50%, 51%, 52%, 53%, 54% It has functional homologs having sequence identity rates of 55%, 56%, 57%, 58%, 59%, 60%, 61%, 62%, 63%, 64%, 65%, 66%, 67%, 68%, 69%, 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or even 100%. Accordingly, the present invention includes the overexpression of a combination of the transcription factor of the present invention and a first accessory protein (Pichia pastris Kar2p) or its functional homologue described in SEQ ID NO: 28 and a second endoplasmic reticulum accessory protein (Pichia pastris Lhs1p) or its functional homologue described in SEQ ID NO: 37. Preferably, the functional homologues of SEQ ID NO: 37 as the second endoplasmic reticulum accessory protein, which are overexpressed in addition to the transcription factor and the first endoplasmic reticulum accessory protein, are SEQ ID NOs: 38 to 46.
[0162] A second endoplasmic reticulum accessory protein having an amino acid sequence or functional homolog thereof, such as that shown in Sequence ID No. 37, may be obtained from Pichia pastris (Chomagataera pastris or Chomagataera fafi), Hanzenula polymorpha, Trichoderma liesei, Saccharomyces cerevisiae, Kluiveromyces lactis, Yarowia liporitica, Candida boidini, Schizosaccharomyces pombe, Aspergillus niger, preferably Pichia pastris (Chomagataera pastris or Chomagataera fafi), for additional overexpression or for engineering host cells to induce additional overexpression.
[0163] Overexpression of the Msn4p transcription factor(s) and the first Kar2p accessory protein(s) and the second Lhs1p accessory protein(s) of the present invention increases the yield of model proteins, preferably scFv (SEQ ID NO: 13) and / or vHH (SEQ ID NO: 14), by at least 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 100%, 110%, 120%, and 130% compared to host cells before engineering. It can be increased by %, 140%, 150%, 160%, 170%, 180%, 190%, 200%, 210%, 220%, 230%, 240%, 250%, 260%, 270%, 280%, 290%, 300%, 310%, 320%, 330%, 340%, 350%, 360%, 370%, 380%, 390%, 400%, 410%, 420%, 430%, 440%, 450%, 460%, 470%, 480%, 490%, or 500%. Overexpression of the Pichia pastris native transcription factor Msn4p, the first endoplasmic reticulum accessory protein Kar2p, and the second accessory protein Lhs1p of Pichia pastris according to the present invention increases the yield of the model protein, preferably vHH (SEQ ID NO: 14), by at least 60%, for example, 70%, 80%, 90%, 100%, 110%, 120%, 130%, or 140%, compared to host cells before engineering. It can be increased by 150%, 160%, 170%, 180%, 190%, 200%, 210%, 220%, 230%, 240%, 250%, 260%, 270%, 280%, 290%, 300%, 310%, 320%, 330%, 340%, 350%, 360%, 370%, 380%, 390%, 400%, 410%, 420%, 430%, 440%, 450%, 460%, 470%, 480%, 490%, or 500%.Overexpression of the synthetic transcription factor synMsn4p, the first endoplasmic reticulum accessory protein Kar2p of Pichia pastris, and the second accessory protein Lhs1p of Pichia pastris increases the yield of the model protein, preferably scFv (SEQ ID NO: 13), by at least 80%, for example, 90%, 100%, 110%, 120%, 130%, 140%, 150%, 1% compared to host cells before engineering. It can be increased by 60%, 170%, 180%, 190%, 200%, 210%, 220%, 230%, 240%, 250%, 260%, 270%, 280%, 290%, 300%, 310%, 320%, 330%, 340%, 350%, 360%, 370%, 380%, 390%, 400%, 410%, 420%, 430%, 440%, 450%, 460%, 470%, 480%, 490%, or 500%.
[0164] The present invention includes another overexpression of a combination of the transcription factor of the present invention and a first accessory protein or its functional homologue described in SEQ ID NO: 28 and another second endoplasmic reticulum accessory protein or its functional homologue described in SEQ ID NO: 47.
[0165] Preferably, the other second endoplasmic reticulum accessory protein has an amino acid sequence as shown in SEQ ID NO: 47, or a homolog thereof, wherein the homolog is at least 20% of the amino acid sequence shown in SEQ ID NO: 47 (Sil1p of Pichia pastris), for example, 21%, 22%, 23%, 24%, 25%, 26%, 27%, 28%, 29%, 30%, 31%, 32%, 33%, 34%, 35%, 36%, 37%, 38%, 39%, 40%, 41%, 42%, 43%, 44%, 45%, 46%, 47%, 4 The sequences have a sequence identity rate of 8%, 49%, 50%, 51%, 52%, 53%, 54%, 55%, 56%, 57%, 58%, 59%, 60%, 61%, 62%, 63%, 64%, 65%, 66%, 67%, 68%, 69%, 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or even 100%. Preferably, functional homologs of SEQ ID NO: 47 as other second endoplasmic reticulum accessory proteins that are overexpressed in addition to the transcription factor and the first endoplasmic reticulum accessory protein are SEQ ID NOs: 48-54.
[0166] A second endoplasmic reticulum (ER) accessory protein having an amino acid sequence or functional homolog thereof as shown in SEQ ID NO: 47 may be obtained from Pichia pastris (Comagataera pastris or Comagataera fafi), Hanzenula polymorpha, Trichoderma liesei, Saccharomyces cerevisiae, Kluiveromyces lactis, Yarowia liporitica, Candida boidini, preferably Pichia pastris (Comagataera pastris or Comagataera fafi), for additional overexpression or for engineering host cells to induce additional overexpression. The closest homologs from other eukaryotic species may also be obtained for at least one ER accessory protein having an amino acid sequence or functional homolog thereof as shown in SEQ ID NO: 47.
[0167] Overexpression of the aforementioned Msn4p transcription factor(s), the aforementioned first Kar2p accessory protein(s), and the aforementioned second Sil1p accessory protein(s) of the present invention increases the yield of model proteins, preferably scFv (SEQ ID NO: 13) and / or vHH (SEQ ID NO: 14), by at least 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 100%, 110%, 120%, and 130% compared to host cells before engineering. It can be increased by %, 140%, 150%, 160%, 170%, 180%, 190%, 200%, 210%, 220%, 230%, 240%, 250%, 260%, 270%, 280%, 290%, 300%, 310%, 320%, 330%, 340%, 350%, 360%, 370%, 380%, 390%, 400%, 410%, 420%, 430%, 440%, 450%, 460%, 470%, 480%, 490%, or 500%.
[0168] When a polynucleotide encoding at least one transcription factor is introduced under the control of a promoter by a vector or plasmid, the polynucleotides encoding two additional endoplasmic reticulum accessory proteins are incorporated into the same vector or plasmid under the control of the same promoter, or under the control of different promoters ((a) Msn4p under the control of one promoter, Kar2p under the control of a different promoter, and Lhs1p or Sil1p under the control of yet another different promoter; or b) Msn4p and Kar2p under the control of the same promoter, and Lhs1p or Sil1p under the control of different promoters; or c) Msn4p under the control of one promoter, and Kar2p and Lhs1p or Sil1p under the control of yet another promoter). When introducing a polynucleotide encoding at least one transcription factor under promoter control via a vector or plasmid, polynucleotides encoding two additional endoplasmic reticulum (ER) accessory proteins (one polynucleotide encoding the first ER accessory protein and the other polynucleotide encoding the other second ER accessory protein) can be incorporated simultaneously or sequentially (one at a time) onto separate vectors or plasmids (one vector / plasmid containing the polynucleotide encoding at least one transcription factor, and the other vector / plasmid containing the polynucleotides encoding the first and second ER accessory proteins). For example, if both the polynucleotide encoding at least one transcription factor and the polynucleotides encoding at least two additional ER accessory proteins can be introduced onto separate vectors or plasmids, one integration plasmid BB3 having only the at least one transcription factor under promoter control, and another integration plasmid BB3 having two additional ER accessory proteins (e.g., Kar2p under promoter control and Lhs1p or Sil1p under the control of another promoter) can be used.
[0169] When introducing one or more copies of a polynucleotide encoding at least one transcription factor into the control of a promoter via a vector or plasmid, polynucleotides encoding at least two additional endoplasmic reticulum accessory proteins may be incorporated into the same vector or plasmid under the control of the same promoter or different promoters(s): (a) one or more copies of Msn4p under the control of one promoter, one or more copies of Kar2p under the control of a different promoter, and one or more copies of Lhs1p or Sil1p under the control of another different promoter; or (b) one or more copies of Msn4p and Kar2p under the control of the same promoter, and one or more copies of Lhs1p or Sil1p under the control of different promoters; or (c) one or more copies of Msn4p under the control of one promoter, and one or more copies of Kar2p and Lhs1p or Sil1p under the control of another promoter. When one or more copies of a polynucleotide encoding at least one transcription factor are introduced under the control of a promoter by a vector or plasmid, one or more copies of polynucleotides encoding two additional endoplasmic reticulum (ER) accessory proteins (one polynucleotide encoding the first ER accessory protein and the other polynucleotide encoding the other second ER accessory protein) are incorporated simultaneously or sequentially (one at a time) onto different vectors or plasmids (one vector / plasmid containing the polynucleotide encoding at least one transcription factor, and the other vector / plasmid containing the polynucleotides encoding the first and second ER accessory proteins).
[0170] Overexpression of two additional endoplasmic reticulum (ER) accessory proteins (Kar2p and Lhs1p or Kar2p and Sil1p) can ensure that the protein of interest folds correctly within the ER, thereby further increasing the yield / titer of the protein of interest. In this embodiment, the second accessory protein (e.g., Lhs1p or Sil1p) may interact with the first ER accessory protein (e.g., Kar2p) as a co-chaperone when folding the protein of interest.
[0171] Overexpression of the aforementioned additional endoplasmic reticulum accessory proteins (e.g., Kar2p, Lhs1p, or Sil1p) or engineering host cells to overexpress them can be accomplished by any method known to those skilled in the art, which has been previously described herein for the allogeneic transcription factors of the present invention or for the heterogeneous transcription factors of the present invention.
[0172] The present invention includes another overexpression of a combination of the transcription factor of the present invention with the first endoplasmic reticulum accessory protein or its functional homologue described in SEQ ID NO: 28, another second endoplasmic reticulum accessory protein or its functional homologue described in SEQ ID NO: 37 / SEQ ID NO: 47, and optionally the third endoplasmic reticulum accessory protein or its functional homologue described in SEQ ID NO: 55.
[0173] Preferably, the third endoplasmic reticulum accessory protein has an amino acid sequence or homologue thereof as shown in SEQ ID NO: 55, wherein the homologue is at least 25%, for example, 26%, 27%, 28%, 29%, 30%, 31%, 32%, 33%, 34%, 35%, 36%, 37%, 38%, 39%, 40%, 41%, 42%, 43%, 44%, 45%, 46%, 47%, 48%, 49%, 50%, 5% relative to the amino acid sequence shown in SEQ ID NO: 55 (Erj5p of Pichia pastris). The sequences have a sequence identity rate of 1%, 52%, 53%, 54%, 55%, 56%, 57%, 58%, 59%, 60%, 61%, 62%, 63%, 64%, 65%, 66%, 67%, 68%, 69%, 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or even 100%. Preferably, the functional homologs of SEQ ID NO: 55 as a third endoplasmic reticulum accessory protein overexpressed in addition to the transcription factor, the first endoplasmic reticulum accessory protein, and the second endoplasmic reticulum accessory protein are SEQ ID NOs: 56-64.
[0174] A third endoplasmic reticulum accessory protein having an amino acid sequence or functional homolog thereof as shown in Sequence ID No. 55 is obtained from Pichia pastris (Chomagataera pastris or Chomagataera fafi), Hanzenula polymorpha, Trichoderma liesei, Saccharomyces cerevisiae, Kluiveromyces lactis, Yarowia liporitica, Candida boidini, Schizosaccharomyces pombe, and Aspergillus niger, preferably from Pichia pastris (Chomagataera pastris or Chomagataera fafi).
[0175] When a polynucleotide encoding at least one transcription factor is introduced under the control of a promoter by a vector or plasmid, three additional polynucleotides encoding endoplasmic reticulum (ER) accessory proteins are incorporated into the same vector or plasmid under the control of the same promoter or different promoters. When a polynucleotide encoding at least one transcription factor is introduced under the control of a promoter by a vector or plasmid, three additional polynucleotides encoding ER accessory proteins (one polynucleotide encoding the first ER accessory protein, another polynucleotide encoding the other second ER accessory protein, and another polynucleotide encoding the other third ER accessory protein) are incorporated into different vectors or plasmids simultaneously or sequentially (one at a time) (one vector / plasmid contains a polynucleotide encoding at least one transcription factor, and the other vector / plasmid contains polynucleotides encoding the first, second, and third ER accessory proteins). For example, if both a polynucleotide encoding at least one transcription factor and a polynucleotide encoding three additional endoplasmic reticulum (ER) accessory proteins can be introduced onto different vectors or plasmids, one embedded plasmid BB3 having only at least one transcription factor under promoter control and another embedded plasmid BB3 having three additional ER accessory proteins (e.g., Kar2p under promoter control, Lhs1p or Sil1p under another promoter control, and Erj5p under yet another promoter control) can be used.
[0176] When one or more copies of a polynucleotide encoding at least one transcription factor are introduced under the control of a promoter by a vector or plasmid, polynucleotides encoding one or more additional three endoplasmic reticulum (ER) accessory proteins are incorporated into the same vector or plasmid under the control of the same promoter or under the control of different promoters. When one or more copies of a polynucleotide encoding at least one (homogeneous and / or heterogeneous) transcription factor are introduced under the control of a promoter by a vector or plasmid, one or more copies of polynucleotides encoding three additional ER accessory proteins (one polynucleotide encoding the first ER accessory protein, another polynucleotide encoding the other second ER accessory protein, and another polynucleotide encoding the third ER accessory protein) are incorporated into different vectors or plasmids simultaneously or sequentially (one at a time) (one vector / plasmid contains the polynucleotide encoding at least one transcription factor, and the other vector / plasmid contains the polynucleotides encoding the first, second, and third ER accessory proteins).
[0177] Overexpression of the Msn4p transcription factor(s) of the present invention, the first Kar2p accessory protein(s), the second Lhs1p accessory protein(s), and the third Erj5p accessory protein(s) increases the yield of model proteins, preferably scFv (SEQ ID NO: 13) and / or vHH (SEQ ID NO: 14), by at least 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 100%, 11% compared to host cells before engineering. It can be increased by 0%, 120%, 130%, 140%, 150%, 160%, 170%, 180%, 190%, 200%, 210%, 220%, 230%, 240%, 250%, 260%, 270%, 280%, 290%, 300%, 310%, 320%, 330%, 340%, 350%, 360%, 370%, 380%, 390%, 400%, 410%, 420%, 430%, 440%, 450%, 460%, 470%, 480%, 490%, or 500%. Overexpression of the Pichia pastris native transcription factor Msn4p, the first endoplasmic reticulum accessory protein Kar2p, the second endoplasmic reticulum accessory protein Lhs1p, and the third accessory protein Erj5p of the Pichia pastris according to the present invention increases the yield of the model protein, preferably vHH (SEQ ID NO: 14), by at least 110%, 120%, and 130% compared to host cells before engineering. It can be increased by %, 140%, 150%, 160%, 170%, 180%, 190%, 200%, 210%, 220%, 230%, 240%, 250%, 260%, 270%, 280%, 290%, 300%, 310%, 320%, 330%, 340%, 350%, 360%, 370%, 380%, 390%, 400%, 410%, 420%, 430%, 440%, 450%, 460%, 470%, 480%, 490%, or 500%.Overexpression of the synthetic transcription factor synMsn4p of the present invention, and the first endoplasmic reticulum accessory protein Kar2p of Pichia pastris, the second endoplasmic reticulum accessory protein Lhs1p of Pichia pastris, and the third endoplasmic reticulum accessory protein Erj5p of Pichia pastris increases the yield of the model protein, preferably vHH (SEQ ID NO: 14), by at least 70%, e.g., 80%, 90%, 100%, 110%, compared to host cells before engineering. It can be increased by 120%, 130%, 140%, 150%, 160%, 170%, 180%, 190%, 200%, 210%, 220%, 230%, 240%, 250%, 260%, 270%, 280%, 290%, 300%, 310%, 320%, 330%, 340%, 350%, 360%, 370%, 380%, 390%, 400%, 410%, 420%, 430%, 440%, 450%, 460%, 470%, 480%, 490%, or 500%.
[0178] Overexpression of the aforementioned Msn4p transcription factor(s), the aforementioned first Kar2p accessory protein(s), the aforementioned second Sil1p accessory protein(s), and the aforementioned third Erj5p accessory protein(s) of the present invention increases the yield of the model protein scFv (SEQ ID NO: 13) and / or vHH (SEQ ID NO: 14) by at least 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 100%, 110% compared to host cells before engineering. It can be increased by %, 120%, 130%, 140%, 150%, 160%, 170%, 180%, 190%, 200%, 210%, 220%, 230%, 240%, 250%, 260%, 270%, 280%, 290%, 300%, 310%, 320%, 330%, 340%, 350%, 360%, 370%, 380%, 390%, 400%, 410%, 420%, 430%, 440%, 450%, 460%, 470%, 480%, 490%, or 500%.
[0179] The methods, recombinant host cells, and uses of the present invention may further include overexpressing, or engineering, the host cells to overexpress, at least one polynucleotide encoding one additional transcription factor. Thus, the host cells overexpress at least one transcription factor and at least one polynucleotide encoding one additional transcription factor. Preferably, by overexpressing at least one polynucleotide encoding at least one additional transcription factor in the host cells, the yield of the recombinant protein of interest is increased compared to host cells that overexpress at least one polynucleotide encoding at least one transcription factor but not at least one polynucleotide encoding at least one additional transcription factor.
[0180] The additional transcription factors were initially isolated from Pichia pastris (Chomagataera fafi) strain CBS7435 (CBS-KNAW culture collection). The transcription factors(s) are assumed to be able to be overexpressed in a wide variety of host cells. Therefore, instead of using sequences that are native to a species or genus, the transcription factor sequences(s) may also be obtained from or derived from other prokaryotes or eukaryotes. Preferably, transcription factors(s) are employed from Pichia pastris (Chomagataera pastris or Chomagataera fafi), Hanzenula polymorpha, Trichoderma liesei, Saccharomyces cerevisiae, Kluiveromyces lactis, Yarowia liporitica, Candida boidini, and Aspergillus niger for additional overexpression, or for engineering host cells to be additionally overexpressed.
[0181] In the present invention, additional Hac1 transcription factors refer to SEQ ID NOs: 74-82, which include a DNA-binding domain containing an amino acid sequence such as that shown in SEQ ID NO: 65, or a functional homolog of the amino acid sequence such as that shown in SEQ ID NO: 65 having at least 50% sequence identity with the amino acid sequence such as that shown in SEQ ID NO: 65 as described herein, and an optional activation domain (a synthetic domain, viral domain, or activation domain of any type of additional transcription factor as described elsewhere herein). Alignment of the DNA-binding domain and the optional activation domain of the additional transcription factor as described herein may be carried out according to the knowledge of those skilled in the art and may be carried out in any order.
[0182] Preferably, the additional transcription factor comprises at least one DNA-binding domain and an activation domain, wherein the DNA-binding domain comprises an amino acid sequence such as that shown in SEQ ID NO: 65 (the DNA-binding domain of Hac1p in Pichia pastris).
[0183] Preferably, the additional transcription factor comprises at least one DNA-binding domain and an activation domain, wherein the DNA-binding domain comprises at least 50% of the amino acid sequence as shown in SEQ ID NO: 65, for example, at least 51%, 52%, 53%, 54%, 55%, 56%, 57%, 58%, 59%, 60%, 61%, 62%, 63%, 64%, 65%, 66%, 67%, 68%, 69% This includes functional homologs of amino acid sequences such as the one shown in SEQ ID NO: 65, having sequence identity rates of %, 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or even 100%.
[0184] Preferably, functional homologs of the amino acid sequence shown in SEQ ID NO: 65, having at least 50% sequence identity with the amino acid sequence shown in SEQ ID NO: 65, are SEQ ID NOs: 66-73.
[0185] Therefore, the method, recombinant host cells, and use of the present invention may further include overexpression of an additional transcription factor comprising at least one DNA-binding domain and an activation domain having an amino acid sequence as shown in SEQ ID NOs. 65-73.
[0186] HAC1 encodes a transcription factor of the basic leucine zipper (bZIP) family involved in the endoplasmic reticulum (ER) stress response (Mori K et al., Genes Cells 1(9):803-17, 1996 and Cox JS and Water P, Cell 87(3):391-404, 1996). Heat stress, drug treatment, mutations in secretory proteins, or overexpression of wild-type secretory proteins can cause unfolded proteins to accumulate in the ER, triggering the ER stress response (UPR). HAC1 is not essential under normal growth conditions but is essential under conditions that trigger the ER stress response. Hac1p binds to a DNA sequence called the ER stress response element (UPRE) within the promoters of genes regulated by the ER stress response, such as KAR2, PDI1, EUG1, and FKB2. The amount of Hac1p is regulated by splicing of HAC1 mRNA. Spliced HAC1 mRNA is translated with much higher efficiency than unspliced transcripts. Hac1p induces transcription of genes encoding endoplasmic reticulum chaperones, such as Kar2p, which is involved in the endoplasmic reticulum stress response. Increased transcription of genes encoding soluble endoplasmic reticulum proteins, including endoplasmic reticulum chaperones, is a key feature of the endoplasmic reticulum stress response. Furthermore, Hac1p increases the synthesis of proteins present in the endoplasmic reticulum that are necessary for protein folding.
[0187] When a polynucleotide encoding at least one transcription factor is introduced under the control of a promoter by a vector or plasmid, polynucleotides encoding additional transcription factors are incorporated into the same vector or plasmid under the control of the same promoter or under the control of different promoters (Msn4p under the control of one promoter, Hac1p under the control of a different promoter). When both a polynucleotide encoding at least one transcription factor and a polynucleotide encoding an additional transcription factor can be introduced into the same vector or plasmid, recombinant plasmid B33 is preferably used, where the polynucleotide encoding at least one transcription factor is under the control of a promoter, and the polynucleotide encoding at least one additional transcription factor is under the control of a different promoter. When a polynucleotide encoding at least one transcription factor is introduced under the control of a promoter by a vector or plasmid, polynucleotides encoding additional transcription factors are incorporated into different vectors or plasmids simultaneously or sequentially (one at a time). For example, if both a polynucleotide encoding at least one transcription factor and a polynucleotide encoding an additional transcription factor can be introduced onto different vectors or plasmids, then one embedded plasmid BB3 having only at least one transcription factor and another embedded plasmid BB3 having only at least one additional transcription factor can be used.
[0188] When one or more copies of a polynucleotide encoding at least one transcription factor are introduced under the control of a promoter by a vector or plasmid, one or more copies of the polynucleotide encoding additional transcription factors are incorporated into the same vector or plasmid under the control of the same promoter or under the control of different promoters (one or more copies of Msn4p under the control of one promoter, and one or more copies of Hac1p under the control of different promoters). When one or more copies of a polynucleotide encoding at least one transcription factor are introduced under the control of a promoter by a vector or plasmid, one or more copies of the polynucleotide encoding additional transcription factors are incorporated into different vectors or plasmids simultaneously or sequentially (one at a time).
[0189] Overexpression of additional transcription factors can lead to overexpression of endoplasmic reticulum chaperones, such as Kar2p, which are a key feature of the endoplasmic reticulum stress response, thereby further increasing the yield of the protein of interest.
[0190] Overexpression of the Msn4p transcription factor(s) and the additional Hac1p transcription factor(s) of the present invention increases the yield of the model protein scFv (SEQ ID NO: 13) and / or vHH (SEQ ID NO: 14) by at least 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 100%, 110%, 120%, 130%, 140%, 150%, and 160% compared to host cells before engineering. It can be increased by 170%, 180%, 190%, 200%, 210%, 220%, 230%, 240%, 250%, 260%, 270%, 280%, 290%, 300%, 310%, 320%, 330%, 340%, 350%, 360%, 370%, 380%, 390%, 400%, 410%, 420%, 430%, 440%, 450%, 460%, 470%, 480%, 490%, or 500%. Overexpression of the Pichia pastris native transcription factor Msn4p and the additional transcription factor Hac1p of Pichia pastris according to the present invention increases the yield of the model protein, preferably vHH (SEQ ID NO: 14), by at least 60%, for example, 70%, 80%, 90%, 100%, 110%, 120%, 130%, 140%, 150%, 160%, 170%, 1% compared to host cells before engineering. It can be increased by 80%, 190%, 200%, 210%, 220%, 230%, 240%, 250%, 260%, 270%, 280%, 290%, 300%, 310%, 320%, 330%, 340%, 350%, 360%, 370%, 380%, 390%, 400%, 410%, 420%, 430%, 440%, 450%, 460%, 470%, 480%, 490%, or 500%.Overexpression of the synthetic transcription factor synMsn4p of the present invention and the additional transcription factor Hac1p of Pichia pastris increases the yield of the model protein, preferably vHH (SEQ ID NO: 14), by at least 80%, for example, 90%, 100%, 110%, 120%, 130%, 140%, 150%, 160%, 170%, 180%, 190%, compared to host cells before engineering. It can be increased by 200%, 210%, 220%, 230%, 240%, 250%, 260%, 270%, 280%, 290%, 300%, 310%, 320%, 330%, 340%, 350%, 360%, 370%, 380%, 390%, 400%, 410%, 420%, 430%, 440%, 450%, 460%, 470%, 480%, 490%, or 500%.
[0191] The at least one polynucleotide encoding at least one additional transcription factor encodes a heterogeneous or homogeneous additional transcription factor. Overexpression of the additional transcription factor (Hac1p), or engineering host cells to overexpress it, can be achieved for the homogeneous or heterogeneous transcription factor of the present invention as discussed above.
[0192] The methods, recombinant host cells, and additional transcription factors(s) used in the present invention may include amino acid sequences as shown in SEQ ID NOs. 74-82, or functional homologs of amino acid sequences as shown in SEQ ID NOs. 74, having at least 20% sequence identity to the amino acid sequences shown in SEQ ID NOs. 74. In further embodiments, the methods, recombinant host cells, and additional transcription factors(s) used in the present invention may include amino acid sequences as shown in SEQ ID NOs. 74-82, or functional homologs of amino acid sequences as shown in SEQ ID NOs. 74, having at least 20% sequence identity, for example, 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 98%, or even 100% sequence identity to the amino acid sequences shown in SEQ ID NOs. 74. The additional transcription factors(s) may further include nuclear localization signals (NLS).
[0193] The present invention further envisions a method for increasing the secretion of recombinant protein of interest by a eukaryotic host cell, comprising the step of overexpressing at least one polynucleotide encoding at least one transcription factor in the host cell, thereby increasing the yield of the recombinant protein of interest compared to host cells that do not overexpress the polynucleotide encoding the transcription factor, wherein the transcription factor comprises at least one DNA-binding domain and an activation domain having an amino acid sequence as shown in SEQ ID NO: 1.
[0194] Furthermore, the present invention further envisions a method for increasing the secretion of recombinant protein of interest by eukaryotic host cells, comprising the step of overexpressing at least one polynucleotide encoding at least one transcription factor in eukaryotic host cells, thereby increasing the yield of recombinant protein of interest compared to host cells that do not overexpress the polynucleotide encoding the transcription factor, wherein the transcription factor comprises at least one DNA-binding domain and an activation domain, including a functional homolog of the amino acid sequence shown in SEQ ID NO: 1, having at least 60% sequence identity with the amino acid sequence shown in SEQ ID NO: 87, and / or having at least 60% sequence identity with the amino acid sequence shown in SEQ ID NO: 87.
[0195] The present invention also provides recombinant eukaryotic host cells for producing a protein of interest, wherein the host cells are engineered to overexpress at least one polynucleotide encoding at least one transcription factor.
[0196] Preferably, the present invention provides recombinant eukaryotic host cells for producing a protein of interest, wherein the host cells are engineered to overexpress at least one polynucleotide encoding at least one transcription factor, wherein the transcription factor comprises at least one DNA-binding domain and an activation domain, wherein the DNA-binding domain comprises an amino acid sequence such as that shown in Sequence ID No. 1.
[0197] Furthermore, the present invention provides recombinant eukaryotic host cells for producing a protein of interest, wherein the host cells are engineered to overexpress at least one polynucleotide encoding at least one transcription factor, wherein the transcription factor comprises at least one DNA-binding domain and an activation domain, including a functional homolog of the amino acid sequence shown in SEQ ID NO: 1, having at least 60% sequence identity to the amino acid sequence shown in SEQ ID NO: 87, and / or having at least 60% sequence identity to the amino acid sequence shown in SEQ ID NO: 87.
[0198] "Recombinant cell" or "recombinant host cell" refers to a cell or host cell that has been genetically modified to contain nucleic acid sequences that were not naturally occurring in the cell.
[0199] The present invention further encompasses the use of recombinant eukaryotic host cells for producing recombinant proteins of interest. Host cells may be advantageously used to introduce polypeptides encoding one or more proteins of interest and then cultured under conditions suitable for expressing the proteins of interest.
[0200] Examples The following examples are provided to the art to provide a complete disclosure and explanation of the method of preparation and use of the invention and are not intended to limit the scope of the invention as defined in the claims. While efforts have been made to ensure accuracy with respect to the figures used (e.g., quantity, temperature, concentration, etc.), some experimental error and deviation should be permitted. Unless otherwise noted, parts are by weight, molecular weight is average molecular weight, temperature is Celsius, and pressure is atmospheric pressure or near atmospheric pressure.
[0201] The following examples will demonstrate that newly identified accessory proteins increase the titer (product per volume (mg / L)) and yield (product per biomass (mg / g), where biomass is measured as dry or wet cell weight) of recombinant proteins, respectively, when they are overexpressed. As an example, the yield of single-strand variable fragments (scFv, vHH) of recombinant antibodies in Pichia pastris yeast increases. Positive effects were shown in shaking cultures (conducted in shaking flasks or deep-bottom well plates) and in laboratory-scale fed-batch cultures.
[0202] Example 1: Preparation and selection of Pichia pastris strains secreting antibody fragments scFv and vHH
[0203] Pichia pastrius CBS7435mut s A mutant (genome sequenced by Sturmberger et al. 2016) was used as the host strain. The pPM2d_pGAP and pPM2d_pAOX expression plasmids are derivatives of the pPuzzle_ZeoR plasmid backbone described in International Publication No. 2008 / 128701A2, consisting of a pUC19 bacterial origin of replication and a zeosin antibiotic resistance cassette. Heterogenetic expression was mediated by the Pichia pastris glyceraldehyde-3-phosphate dehydrogenase (GAP) promoter or alcohol oxidase (AOX) promoter, and the Saccharomyces cerevisiae CYC1 transcription termination factor, respectively. The plasmids already contained a Saccharomyces cerevisiae α junction factor pre-pro-reader sequence at the N-terminus. The scFv and vHH genes were codon-optimized using DNA2.0 and obtained as synthetic DNA. His6-tags were fused to the C-terminus of the genes for detection. After restriction enzyme digestion using XhoI and BamHI (for scR) or EcoRV (for vHH), each gene was ligated to both the plasmids pPM2d_pGAP and pPM2d_pAOX, which were digested with XhoI and BamHI or EcoRV.
[0204] Plasmids were chained using either AvrII restriction enzyme (for pPM2d_pGAP) or Pmel restriction enzyme (for pPM2d_pAOX), and then electroporated into Pichia pastris (using a standard transformation protocol as described in Gasser et al. 2013. Future Microbiol. 8(2):191-208). Selection of positive transformants was performed on YPD plates containing 100 μg / ml zeosin (per liter: 10 g yeast extract, 20 g peptone, 20 g glucose, 20 g agar-agar).
[0205] One colony from each transformation approach (approximately 120 in total) was picked from the transformation plate and placed in one well of a 96-depth plate. After biomass generation during the initial growth phase, expression from the AOX1 promoter was induced by supplementation with a methanol-containing medium composition (a total of four times). 72 hours after the initial methanol induction, all depth-bottom plates were centrifuged, and the supernatant from all wells was collected in a stock microtiter plate for subsequent analysis. Expression from the GAP promoter was continued at predetermined time points after the initial growth phase (i.e., twice per day over two days) by glucose supplementation. A total of 110 hours after the initial inoculation, the culture medium was collected as described above.
[0206] Clones with the highest productivity in small-scale screening (Example 3) and in fed-batch culture media (Example 4) were selected to serve as the basic production strain for further engineering operations. Clone CBS7435mut s pAOXscR 4E3 was selected as the basic producing strain for scFv secretion. Clone CBS7435mut s pAOX vHH 14G8 was selected as the basic producing strain for vHH secretion.
[0207] Example 2: Creation of an engineered strain overexpressing accessory genes To investigate the positive effects of scFv and vHH secretion, candidate accessory genes were introduced into two basic production strains: CBS7435mut. s pAOXscR(scFv)4E3 and CBS7435mut s It was overexpressed in pAOX vHH(vHH)14G8 (see Example 1 for preparation).
[0208] a) General procedure for amplification and cloning of selected candidate secretion-supporting genes The genes selected for overexpression were amplified from start codon to stop codon by PCR (Q5® high-fidelity DNA polymerase, New England Biolabs) or split into two or more fragments. The Golden PiCS system (Prielhofer et al. 2017. BMC Systems Biol. doi: 10.1186 / s12918-017-0492-3) requires the introduction of silent mutations into several coding sequences. This was done by amplifying several fragments from a single coding sequence. Alternatively, gBlocks or synthetic codon-optimized genes were obtained from suppliers (including Integrated DNA Technology IDT, Geneart, and ATUM). The amplified coding sequences were cloned into either the pPUZZLE-type expression plasmid pPM2aK21 or pPM2eH21, or the Golden PiCS system (consisting of the BB1, BB2, and BB3aK / BB3eH / BB3rN skeletons). The gene fragments listed in Table 1 were introduced into BB1 of the Golden PiCS system using BsaI restriction enzymes. All promoters and termination factors used to construct expression cassettes within the BB2 or BB3 skeleton are described in Prielhofer et al. 2017 (BMC Systems Biol. doi: 10.1186 / s12918-017-0492-3). pPM2aK21 and BB3aK allow integration into the 3'-AOX1 genomic region and contain KanMX selection marker cassettes for selection in Escherichia coli (E. coli) and yeast. pPM2eH21 and BB3eH contain a 5'-ENO1 genome integration region and an HphMX selection marker cassette for selection with hygromycin. BB3rN contains a 5'-RGI1 genome integration region and a NatMX selection marker cassette for selection with notheotricin. All plasmids contain a replication origin for E. coli (pUC19). Pichia pastris strain CBS7435mut sThe derived genomic DNA or gBlocks (Integrated DNA Technologies) served as a PCR template.
[0209] Table 1 lists the gene fragments required to introduce these fragments into BB1 in the Golden PiCS system using the restriction enzyme BsaI. The constructed BB1 cells, each containing the respective coding sequence, were then further processed in the Golden PiCS system to produce the required BB3 integration plasmids, as described in Prielhofer et al. 2017. Underlined nucleotides indicate the first forward and last reverse primers required to produce Golden PiCS-compatible gene fragments, while start and stop codons are shown in bold.
[0210] [Table 2] TIFF0007885305000003.tif251165 TIFF0007885305000004.tif241165 TIFF0007885305000005.tif254165 TIFF0007885305000006.tif244165 TIFF0007885305000007.tif245165 TIFF0007885305000008.tif251165 TIFF0007885305000009.tif251165 TIFF0007885305000010.tif82165
[0211] b) Creation of natural and synthetic MSN4 overexpression strains A single silent mutation was introduced into the natural coding sequence of Pichia pastris MSN4 to remove the BsaI restriction enzyme site. This coding sequence was introduced into BB1 of the Golden PiCS system. The synthetic MSN4 coding sequence was constructed by fusing the transcriptional activation domain sequence (VP64) and nuclear localization sequence (SV40) with the natural DNA-binding domain of MSN4 from nucleotides 883 to 1071. The DNA-binding domain was identified by sequence homology to the amino acid sequence published in Nicholls et al. 2004 (Eukaryot Cell. doi: 10.1128 / EC.3.5.1111-1123.2004). This synthetic coding sequence (synMSN4) was introduced into BB1 of the Golden PiCS system. MSN2 of Saccharomyces cerevisiae, MSN4 of Saccharomyces cerevisiae, Seb1 (an MSN4 homolog of Aspergillus niger), and MSN4 homolog of Yarowia liporitica were amplified from the genomic DNA of Saccharomyces cerevisiae CEN.PK, Aspergillus niger CBS513.88, and Yarowia liporitica DSMZ, respectively, and introduced into BB1.
[0212] Each MSN4 coding sequence was combined with the glyceraldehyde-3-phosphate dehydrogenase (GAP) promoter and the CYC1 transcription termination factor of Saccharomyces cerevisiae in the integrated plasmid BB3rN (e.g., 189_BB3rN or 142_BB3eH for the MSN4 of natural Pichia pastris). The MSN4 of Pichia pastris was also combined with the THI11 promoter and IDP1 termination factor (253_BB3eH), or the POR1 promoter and IDP1 termination factor (254_BB3eH). The synMSN4 coding sequence was further combined with the THI11 promoter (Landes et al. 2016. Biotechnol Bioeng. doi: 10.1002 / bit.26041) and the IDP1 transcription termination factor (258_BB3eH), or the SBH17 promoter and the TDH3 termination factor (191_BB3aK). The synMSN4 coding sequence was also combined with the GAP promoter and the TDH3 transcription termination factor in the integrated plasmid 208_BB3aK. All integrated plasmids were chained using the restriction enzyme AscI, and then applied to transform the basic producing strains. The titer and yield (titer per wet cell weight) of clones overexpressing MSN4 or synthetic MSN4 were determined by small-scale screening and compared to their parental basic producing strains (Example 3).
[0213] c) Creation of a (synthetic) MSN4+KAR2 overexpressing strain An overexpression cassette containing only KAR2 was constructed in the embedded plasmid BB3eH (219_BB3eH). This plasmid was obtained by combining the BB1 plasmid with the KAR2 coding sequence, the GAP promoter, and the RPS3 termination factor.
[0214] Based on the yield of the product determined in a small-scale screening (Example 3), the best clone overexpressing MSN4 or synthetic MSN4 was selected after transformation using the respective plasmids of Example 2b, and this clone was further transformed with the SmaI-chained KAR2-integrated plasmid 219_BB3eH. This ultimately yielded clones with two different overexpression cassettes, introduced by two sequential transformations using two different integration plasmids.
[0215] d) Creation of (synthetic) MSN4+HAC(i) overexpression strains The introduced (i) version of the HAC(i) coding sequence was created by removing alternative introns from nucleotide numbers 857–1178, according to Guerfal et al. 2010 (Microb Cell Fact. doi: 10.1186 / 1475-2859-9-49). The coding sequence was introduced into BB1. Furthermore, the codon-optimized HAC(i) sequence was used for overexpression of Hac1(i). This was further combined with the FDH1 promoter and RPL2A termination factor in the BB2 plasmid. Other BB2 constructs contained HAC1 under the control of the MDH3 promoter and RPL2A termination factor, or the ADH2 promoter and RPL2A termination factor.
[0216] The embedded plasmids 243_BB3eH, 253_BB3eH, 254_BB3eH, and 257_BB3eH, which possess the MSN4+HAC1(i) combination under the control of different promoters, were prepared by combining BB2 from Example 2d with the BB2 plasmid (Example 2b) containing an expression cassette for MSN4. The same combinations were also prepared by sequential transformation using the embedded plasmid BB3rN (189_BB3rN) which possesses only MSN4, and the embedded plasmid BB3eH (234_BB3eH) which possesses only HAC1(i) along with the FDH1 promoter and RPL2A termination factor. In the integrated plasmid (258_BB3eH), the plasmid containing the synMSN+HAC1(i) combination was combined with the BB2 plasmid from Example 2d. This BB2 plasmid was obtained from the BB1 plasmid (Example 2b) containing synMSN4, which was combined with the THI11 promoter and the IDP1 transcription termination factor. Both integrated plasmids were chained using the restriction enzyme SmaI, and then applied to transform the basic production strain.
[0217] e) Creation of (synthetic) MSN4+KAR2 and / or LHS1, (synthetic) MSN4+KAR2 and / or SIL, (synthetic) MSN4+KAR2+LHS1 or SIL1 and ERJ5 overexpression strains The coding sequences for KAR2 (requiring 7 silent mutations), LHS1 (requiring 1 silent mutation), SIL1 (no mutations), and ERJ5 (requiring 1 silent mutation) were introduced into BB1 of the Golden PiCS system. The integrated plasmid 219_BB3eH contains KAR2 along with the GAP promoter and RPS3 transcription termination factor. Overexpression of KAR2 combined with LHS1 was constructed in integrated plasmid 174_BB3eH, which was obtained from two BB2 molecules (one containing KAR2 along with the GAP promoter and RPS3 transcription termination factor, and the other containing LHS1 along with the POR1 promoter and IDP1 transcription termination factor). Overexpression of KAR2 in combination with SIL1 was constructed in the integrated plasmid 078_BB3eH, which was obtained from two BB2 molecules (one containing KAR2 along with the GAP promoter and RPS3 transcription termination factor, and the other containing SIL1 along with the POR1 promoter and IDP1 transcription termination factor). Overexpression of KAR2 in combination with LHS1 and ERJ5 was constructed in the integrated plasmid 052_BB3eH, which was obtained from three BB2 molecules (the first containing KAR2 along with the GAP promoter and Saccharomyces cerevisiae CYC1 transcription termination factor, the second containing LHS1 along with the POR1 promoter and IDP1 transcription termination factor, and the third containing ERJ5 along with the MDH3 promoter and TDH1 transcription termination factor).
[0218] Based on the yield (titer per biomass) determined in a small-scale screening (Example 3), the best clone was selected after transformation using each plasmid from Example 2b, and further transformed using the BB3eH-integrated plasmids chained with each of the aforementioned SmaI plasmids. This ultimately yielded clones with two different overexpression cassettes introduced by two sequential transformations using two different integration plasmids.
[0219] Example 3: Screening for increased scFv or vHH secretion In small-scale screening, transformants of up to 20 overexpression combinations were tested after transformation. Transformants were evaluated by comparing their scFv or vHH titers in the supernatant, their wet cell weight (biomass after centrifugation and supernatant removal), and their scFv or vHH yield (titer per wet cell weight) to the respective parental basal-producing strains. For each overexpression combination, the mean multipliers of change in titer, yield, and wet cell weight were determined to assess secretion improvement. The mean multipliers of change in titer, yield, and wet cell weight were calculated by dividing the arithmetic mean values of titer, yield, and wet cell weight for all transformants by the arithmetic mean values of titer, yield, and wet cell weight for four biological replicas of the basal-producing strain cultured on the same deep-well plate.
[0220] a) Small-scale screening culture of scFv or vHH-producing strains One colony of a Pichia pastris clone was inoculated into 2 mL of YP medium (10 g / L yeast extract, 20 g / L peptone) containing 10 g / L glucose and 50 μg / mL zeosin (basic production strain) or 50 μg / mL zeosin and 500 μg / mL G418 and / or 200 μg / mL hygromycin and / or 100 μg / mL notheotricin (depending on the incorporated plasmid of the engineering strain), and grown overnight at 25°C. These cultures were transferred to 2 mL of synthetic screening medium M2 or ASMv6 (medium composition shown below), supplemented with glucose feed tablets (Kuhner, Switzerland; lot number SMFB63319) or x% enzyme (m2p medium development kit), and incubated in a 24-depth well plate at 25°C at 280 rpm for 1–25 hours. Aliquots of these cultures (final OD) were prepared. 600The cells (corresponding to 4 or 8) were transferred to 2 mL of synthetic screening medium M2 or ASMv6 (in the case of ASMv6, in a fresh 24-depth well plate using the m2p medium development kit). 0.5% by volume of pure methanol was added first, followed by 1% by volume of pure methanol at 19, 27, and 43 hours. After 48 hours, the cells were collected by centrifugation at 2,500 × g for 10 minutes at room temperature and prepared for analysis. Biomass was determined by measuring the cell weight of 1 mL of cell suspension, and the determination of recombinant proteins secreted into the supernatant is described in Examples 3b-3c below.
[0221] Synthetic screening medium M2 contained the following per liter: 22.0 g of citric acid monohydrate, 3.15 g of (NH4)2HPO4, 0.49 g of MgSO4·7H2O, 0.80 g of KCl, 0.0268 g of CaCl2·2H2O, 1.47 mL of PTM1 trace metals, and 4 mg of biotin; pH was set to 5 using KOH (solid).
[0222] The synthetic screening medium ASMv6 contained the following per liter: 44.0 g of citric acid monohydrate, 12.60 g of (NH4)2HPO4, 0.98 g of MgSO4·7H2O, 5.28 g of KCl, 0.1070 g of CaCl2·2H2O, 2.94 mL of PTM1 trace metals, and 8 mg of biotin; the pH was set to 6.5 using KOH (solid).
[0223] b) SDS-PAGE and Western blot analysis For protein gel analysis, the NuPAGE® Novex® Bis-Tris system was used, either with a 12% Bis-Tris gel with MOPS electrophoresis buffer or a 4-12% Bis-Tris gel with MES electrophoresis buffer (all manufactured by Invitrogen). After electrophoresis, proteins were visualized by colloidal Coomassie staining or transferred to a nitrocellulose membrane for Western blot analysis. Therefore, proteins were electroblotted onto the nitrocellulose membrane (for 7 minutes) using Bio-Rad's TransBlot® Turbo® transfer system with ready-to-use membranes and filter paper for minigels and the Turbo program. After blocking, the Western blots were probed using the following antibodies. His-tagged scFv and vHH were detected using the following antibody: Anti-polyhistidine-peroxidase antibody (A7058, Sigma-A) diluted 1:2,000. Detection of horseradish peroxidase conjugate was performed using the chemiluminescent agent Super Signal West chemiluminescent substrate (Thermo Scientific).
[0224] c) Quantitative analysis by microfluidic capillary electrophoresis (mCE) The "LabChip GX / GXII system" (PerkinElmer) was used for the quantitative analysis of protein titer secreted into the culture supernatant. The consumables "Protein Express Lab Chip" (760499, PerkinElmer) and "Protein Express Reagent Kit" (CLS960008, PerkinElmer) were used. In short, several μL of the total culture supernatant were fluorescently labeled and analyzed according to protein size using a microfluidic-based electrophoresis system. Internal standards allowed for the assignment of the approximate size (kDa) and approximate concentration of the detected signals.
[0225] Example 4: Fed Batch Culture Clones of the engineered strain (Example 2) were selected after small-scale screening culture (Example 3). The selected clones were further evaluated in larger culture volumes using fed-batch bioreactor culture. The improvement in secretion observed in small-scale screening was also present and confirmed in fed-batch bioreactor culture.
[0226] a) Fed-batch bioreactor culture procedure Each strain was inoculated into a 300 mL wide-mouthed, baffled, lidded shaking flask filled with 50 mL of YPhyG and shaken overnight at 110 rpm at 28°C (pre-culture 1). Pre-culture 2 (100 mL of YPhyG in a 1000 mL wide-mouthed, baffled, lidded shaking flask) was inoculated with OD 600 Pre-culture 1 was inoculated so that the absorbance (measured at 600 nm) reached approximately 20 in the evening (measured relative to YPhyG medium) (doubling time: approximately 2 hours). Pre-culture 2 was incubated similarly at 110 rpm and 28°C.
[0227] Fed batches were carried out in a 0.8 L working-capacity bioreactor (Minifors, Infors, Switzerland). Pre-culture 2 was individually inoculated into all bioreactors (filled with 400 mL of BSM medium at approximately pH 5.5) until the OD600 reached 2.0. Generally, Pichia pastris was grown on glycerol to generate biomass, and the culture was then subjected to glycerol feeding followed by methanol feeding.
[0228] In the initial batch phase, the temperature was set to 28°C. Over the hour prior to starting the production phase, the temperature was lowered to 24°C and maintained at this level throughout the rest of the process, during which time the pH was reduced to 5.0 and maintained at this level. Oxygen saturation was set to 30% throughout the entire process (cascade control: stirrer, flow rate, oxygen supply). Stirring was applied at 700-1200 rpm, and a flow rate range of 1.0-2.0 L / min (air) was selected. pH 5.0 control was achieved using 25% ammonium. Foaming was controlled by adding the defoamer Glanapon 2000 as needed.
[0229] During the batch phase, biomass was generated until the wet cell weight (WCW) reached approximately 110–120 g / L (μ = approximately 0.30 per hour). The classic batch phase (biomass generation) would last approximately 14 hours. Glycerol was fed at a rate defined by the equation 2.6 + 0.3 × t (g / h), so a total of 30 g of glycerol (60%) would be supplied within 8 hours. The first sampling point was selected to occur at 20 hours (0 hours being the induction time).
[0230] For the following 18 hours (process time from 20 to 38 hours), a glycerol / methanol mixed feed was applied: 66 g of glycerol (60%) was supplied at a glycerol feed rate defined by formula: 2.5 + 0.13 × t (g / h), and 21 g of methanol was added at a methanol feed rate defined by formula: 0.72 + 0.05 × t (g / h).
[0231] During the following 72–74 hours (process time from 38 hours to 110–112 hours), methanol was fed at a feed rate defined by equation 2.2 + 0.016 × t (g / L).
[0232] The YPhyG pre-culture medium (per liter) contained: 20 g of phyton-peptone, 10 g of Bacto yeast extract, and 20 g of glycerol.
[0233] Batch medium: Modified basal salt medium (BSM) (per liter) contained the following: 13.5 mL of H3PO4 (85%), 0.5 g of CaCl·2H2O, 7.5 g of MgSO4·7H2O, 9 g of K2SO4, 2 g of KOH, 40 g of glycerol, 0.25 g of NaCl, 4.35 mL of PTM1, and 0.1 mL of Glanapon 2000 (antifoaming agent).
[0234] The trace elements (per liter) of PTM1 contained the following: 0.2 g biotin, 6.0 g CuSO4·5H2O, 0.09 g KI, 3.00 g MnSO4·H2O, 0.2 g Na2MoO4·2H2O, 0.02 g H3BO3, 0.5 g CoCl2, 42.2 g ZnSO4·7H2O, 65.0 g FeSO4·7H2O, and 5.0 mL H2SO4 (95%~98%).
[0235] The feed solution glycerol (per 1 kg) contained the following: 600 g of glycerol and 12 mL of PTM1.
[0236] The feed solution methanol contained the following: pure methanol.
[0237] b) Analysis of samples from fed-batch bioreactor cultures Samples were obtained at various points in time using the following procedure. The first 3 mL of sampled culture broth (using a syringe) was discarded. 1 mL of freshly obtained sample (3–5 mL) was transferred to a 1.5 mL centrifuge tube and centrifuged at 13,200 rpm (16,100 g) for 5 minutes. The supernatant was carefully transferred to separate vials and stored at 4°C or frozen until analysis.
[0238] One mL of culture broth in a weighed Eppendorf vial was centrifuged at 13,200 rpm (16,100 g) for 5 minutes, and the resulting supernatant was accurately removed. The vial was weighed (accuracy 0.1 mg), and the weight of the empty vial was subtracted to obtain the infused cell weight.
[0239] The supernatant from each bioreactor culture at individual sampling points was analyzed against bovine serum albumin or purified standards (for scR-GG-6xHIS and vHH-GG-6xHIS) using mCE (microfluidic capillary electrophoresis, GXII, PerkinElmer).
[0240] Example 5: Improvement of recombinant protein production and secretion by overexpression of transcription factors and accessory genes. The improvement in secretion is measured by the change in potency and yield ratios, referring to each unengineered basic production strain (Example 1).
[0241] a) Improvement in vHH protein secretion yield by overexpression of transcription factors alone or in combination with accessory genes (groups) - results from a small-scale screening.
[0242] Figure 1 lists the overexpressed genes or gene combinations that increased vHH secretion in Pichia pastris in a small-scale screening (Example 3). The change ratio values for the small-scale screening are the arithmetic mean of up to 20 clones / transformers (see Example 3).
[0243] vHH secretion is increased by overexpression of the transcription factor Msn4 (Figure 1). Both natural and synthetic Msn4 mutants increase vHH titer and yield to similar levels. Surprisingly, overexpression of chaperone Kar2 alone or in combination with co-chaperone Lhs1 did not increase vHH secretion. Increased vHH titer and yield were observed only when these were co-overexpressed with the transcription factors Msn4 or synMsn4. Furthermore, co-expression of Hsp40 proteins such as Erj5 further increased vHH secretion.
[0244] Furthermore, co-expression of Hac1 and Msn4 or synMsn4 enhanced vHH secretion, superior to overexpression of single Hac1. Similar levels of enhancement were achieved regardless of whether the two transcription factors were expressed from the same vector or from two separate vectors. There was also no significant difference when different promoter pairs were used for the expression of the two transcription factors.
[0245] b) Improvement of vHH protein secretion yield by overexpression of transcription factors alone or in combination with auxiliary gene(s) - Results from fed-batch bioreactor cultures.
[0246] Figure 2 lists the overexpressed genes or gene combinations that increase the secretion of vHH in Pichia pastoris in fed-batch culture (Example 4). The fold change values for the fed-batch culture are those of one selected clone.
[0247] The positive effect of overexpressing the transcription factor Msn4 on recombinant protein production observed in the screening was also confirmed in controlled bioreactor cultures (Figure 2). Similar to the screening, overexpression of Msn4 or synMsn4 in combination with chaperones or other transcription factors significantly outperformed the performance of strains overexpressing only the latter factors. No clear difference was seen between overexpression of the natural Msn4 and the synthetic form of Msn4 with respect to the beneficial effect on vHH secretion.
[0248] c) Improvement of scFv protein secretion yield by overexpression of transcription factors alone or in combination with auxiliary gene(s) - Results from small-scale screening. Figure 3 lists the overexpressed genes or gene combinations that increase the secretion of scFv in Pichia pastoris in small-scale screening (Example 3). The fold change values for the small-scale screening are the arithmetic mean values of up to 20 clones / transformants (see Example 3).
[0249] Overexpression of Msn4 also enhanced the secretion level of scFv, which is representative of another model protein of interest (Figure 3). The secretion yield and titer of vHH were further enhanced by combining overexpression of a chaperone such as Kar2 alone or in combination with Lhs1 with overexpression of Msn4 or synMsn4, exceeding the improvement obtained by overexpression of Kar2 and Lhs1 without Msn4. Also, the combination of overexpression of Hac1 with Msn4 or synMsn4 had a positive effect on the secretion of scFv.
[0250] d) Improvement of scFv protein secretion yield by overexpression of a transcription factor alone or in combination with an auxiliary gene(s) - Results from fed-batch bioreactor cultures. Figure 4 lists the overexpressed genes or gene combinations that increase the secretion of vHH in Pichia pastoris in fed-batch cultures (Example 4). The fold change values for the fed-batch cultures are those of one selected clone.
[0251] For the second recombinant model protein, the results obtained in screening were also confirmed under bioreactor conditions similar to a controlled process (Figure 4). Overexpression of only Msn4 improved the titer and yield of scFv compared to the wild-type producing strain (parent). Co-overexpression of a chaperone or other transcription factor (e.g., Hac1) with Msn4 stimulated the secretion of scFv compared to overexpression of the chaperone or Hac1 alone.
[0252] e) Improvement of scFv secretion (titer and yield) by overexpression of MSN2 / 4 homologs from other species in fed-batch bioreactor cultures. Figure 5 lists the overexpressed MSN2 / 4 homologs that increase the secretion of scFv in Pichia pastoris in fed-batch cultures (Example 4). The fold change values for the fed-batch cultures are those of one selected clone.
[0253] Overexpression of two Msn4 homologs from Saccharomyces cerevisiae had a positive effect on scFv secretion (Figure 5), confirming that homologs from other species also have a positive effect on protein secretion in Pichia pastris. Combined with results from natural and synthetic Msn4 mutants in Pichia pastris, this also points to a conserved effect in other production hosts where targeted Msn4 overexpression improves recombinant protein production, highlighting the versatile applicability of our approach.
[0254] Example 6: Alignment of MSN4 and sequence identity for PpMSN4 Knowledge regarding the function of MSN2 / 4 is derived from Saccharomyces cerevisiae, as it is the most important model organism for eukaryotic cells. In this context, it is important to note that Saccharomyces cerevisiae underwent whole-genome duplication (WGD). This results in the genome of Saccharomyces cerevisiae having many very similar copies of its genes. The redundant transcription factors Msn2p and Msn4p are examples of this. Due to this functional redundancy, these transcription factors are commonly referred to as MSN2 / 4. Descriptions of the function of proteins in other yeasts are derived from experiments using the model organism Saccharomyces cerevisiae. Pichia pastris, for example, did not undergo whole-genome duplication and therefore has only one homolog, namely Msn4p. Since there is essentially no clear functional difference between Msn2p and Msn4p in Saccharomyces cerevisiae, these transcription factors in other yeasts cannot be reasonably distinguished.
[0255] Alignment was performed using the CLC Main Workbench software (Qiagen Bioinformatics), and can be seen in Figure 6. The only strongly conserved region, highlighted in the dotted box in Figure 6, consists of the zinc finger protein structure motif. This is the known DNA-binding domain of Msn4p and Msn2p (ScMSN4 / 2), transcription factors well-characterized in Saccharomyces cerevisiae, and can be similarly used to obtain the same function in other organisms (Nicholls et al. 2004).
[0256] The zinc finger in MSN2 / 4 of Saccharomyces cerevisiae has a C2H2-like folding. The amino acid sequence motif is X2-CX 2,4 -CX 12 -HX 3,4,5 It is -H, which is also shown in Figure 7. By gradually zooming in on the image of the strongly conserved region of the sequence alignment (Figure 7) (the black dotted box in Figure 6), this motif can be clearly observed.
[0257] The common sequence of the C2H2 type zinc finger DNA binding domain in MSN4 is highlighted in gray. The C2H2 motif is represented by a black asterisk. * It is marked with ). The common sequence is, [Table 3] That is the case.
[0258] Furthermore, pairwise sequence similarity / identity between the full-length Msn4p of Pichia pastris and its homologs in other organisms was investigated using global pairwise sequence alignment with an embossed needle algorithm. Pairwise sequence similarity / identity was also examined for the DNA-binding domain of Pichia pastris Msn4p and the DNA-binding domain of its homologs in other organisms. The EmbossNeedle web server ((https: / / www.ebi.ac.uk / Tools / psa / emboss_needle / ) was used for pairwise protein sequence alignment with default settings (Matrix: BLOSUM62; Gap Open: 10; Gap Extension: 0.5; End Gap Penalty: False; End Gap Open: 10; End Gap Extension: 0.5). EmbossNeedle decodes two input sequences and exports their optimal global sequence alignment to a file. It uses the Needleman-Bunsch alignment algorithm to find the optimal alignment (including gaps) of the two sequences along their entire length.
[0259] The results for the degree of identity are listed in Figure 8. As expected, the global sequence identity rate of full-length Msn4 shows a much lower degree of conservation than that of the DNA-binding domain alone.
[0260] The pairwise sequence similarity / identity between the common DNA-binding domain (DBD) sequences of Msn4p / Msn2p and the common DNA-binding domain sequences of each homolog of other organisms was investigated using global pairwise sequence alignment with the same embossed needle algorithm (see Figure 14).
[0261] Example 7: Sequence similarity between HAC1 alignment and PpHAC1 Alignment was performed using the CLC Main Workbench software (Qiagen Bioinformatics).
[0262] We investigated the pairwise sequence similarity / identity between the full-length Hac1p or its DNA-binding domain of Pichia pastris and its homologs in other organisms. Global similarity / identity was assessed by global pairwise sequence alignment using an embossed needle algorithm (Figure 13).
Claims
1. A method for increasing the yield of a recombinant protein of interest in a eukaryotic host cell, comprising the step of overexpressing at least one polynucleotide encoding at least one transcription factor in the eukaryotic host cell, thereby increasing the yield of the recombinant protein of interest compared to host cells that do not overexpress the polynucleotide encoding the transcription factor, wherein the transcription factor is at least a) i) the amino acid sequence shown in SEQ ID NO: 5 or 6, or ii) a DNA-binding domain containing a functional homolog of the amino acid sequence shown in SEQ ID NO: 5 or 6 having at least 90% sequence identity with the amino acid sequence shown in SEQ ID NO: 5 or 6, and b) Transcriptional activation domain, and c) Nuclear localization signals Includes, The eukaryotic host cell is a yeast host cell selected from the group consisting of Pichia pastoris, Hansenula polymorpha, Trichoderma reesei, Aspergillus niger, Saccharomyces cerevisiae, Kluyveromyces lactis, Yarrowia lipolytica, Pichia methanolica, Candida boidinii, Komagataella species, and Schizosaccharomyces pombe, according to the method.
2. i) At least: a) a1) a DNA-binding domain containing a functional homolog of the amino acid sequence shown in SEQ ID NO: 5 or 6, which has at least 90% sequence identity with the amino acid sequence shown in SEQ ID NO: 5 or 6, and b) Transcriptional activation domain, and c) Nuclear localization signals A step of engineering a host cell to overexpress at least one polynucleotide encoding at least one transcription factor, including ii) A step of engineering the host cell to contain a polynucleotide encoding a protein of interest, iii) A step of overexpressing at least one polynucleotide encoding at least one transcription factor and culturing the host cells under conditions suitable for overexpressing the protein of interest, optionally iv) A step of isolating the protein of interest from the cell culture medium, and, if applicable v) The process of purifying the protein of interest. The method according to claim 1, including the method described in claim 1.
3. i) A step of preparing a host cell engineered to overexpress at least one polynucleotide encoding at least one transcription factor, wherein the host cell further comprises a polynucleotide encoding a protein of interest, wherein the transcription factor is at least a) a1) a DNA-binding domain comprising a functional homolog of the amino acid sequence shown in SEQ ID NO: 5 or 6, having at least 90% sequence identity with the amino acid sequence shown in SEQ ID NO: 5 or 6, b) Transcriptional activation domain, and c) Nuclear localization signals including, ii) A step of overexpressing at least one polynucleotide encoding at least one transcription factor and culturing the host cells under conditions suitable for overexpressing the protein of interest, optionally iii) A step of isolating the protein of interest from the cell culture medium, and, if applicable iv) A step to purify the protein of interest, and, if applicable v) A step of modifying the protein of interest, and, if applicable vi) Process of formulating the protein of interest Includes, A method for producing recombinant proteins of interest using eukaryotic host cells, wherein the host cells are yeast host cells selected from the group consisting of Pichia pastoris, Hansenula polymorpha, Trichoderma reesei, Aspergillus niger, Saccharomyces cerevisiae, Kluyveromyces lactis, Yarrowia lipolytica, Pichia methanolica, Candida boidinii, Komagataella species, and Schizosaccharomyces pombe.
4. The method according to any one of claims 1 to 3, wherein overexpression of the transcription factor increases the yield of the model protein scFv (SEQ ID NO: 13) and / or vHH (SEQ ID NO: 14) compared to host cells before engineering.
5. The method according to any one of claims 1 to 4, wherein a polynucleotide encoding at least one transcription factor is incorporated into the genome of a host cell or contained in a vector or plasmid that is not incorporated into the genome of a host cell.
6. The method according to any one of claims 1 to 5, wherein the polynucleotide encoding at least one transcription factor encodes a heterogeneous or homogeneous transcription factor.
7. Overexpression of polynucleotides encoding heterologous transcription factors i) A step of replacing or modifying a regulatory sequence operably linked to the polynucleotide encoding a heterologous transcription factor, ii) The step of introducing one or more copies of a polynucleotide encoding a heterologous transcription factor into a host cell under the control of a promoter. The method according to claim 6, which is achieved by...
8. Overexpression of polynucleotides encoding the same transcription factor i) A step of using a promoter that drives the expression of the polynucleotide encoding an allogeneic transcription factor, ii) A step of replacing or modifying a regulatory sequence operably linked to the polynucleotide encoding an allogeneic transcription factor, iii) The step of introducing one or more copies of the polynucleotide encoding an allogeneic transcription factor into a host cell under the control of a promoter. The method according to claim 6, which is achieved by...
9. Overexpression of polynucleotides i) A step of replacing the native promoter of an allogeneic transcription factor with a different promoter operably linked to a polynucleotide encoding the allogeneic transcription factor. ii) A step of replacing the native termination factor sequences of heterologous and / or homologous transcription factors with more efficient termination factor sequences. iii) A step of replacing the coding sequences of heterologous and / or homologous transcription factors with codon-optimized coding sequences, wherein codon optimization is performed according to the codon usage frequency of the host cell. iv) A process of replacing the natural positive regulatory element of an allogeneic transcription factor with a more efficient regulatory element. v) A step of introducing another positive regulatory element that is not present in the natural expression cassette of the same transcription factor. vi) A step of deleting a negative regulatory element that is normally present in the natural expression cassette of an allogeneic transcription factor, vii) A step of introducing one or more copies or combinations thereof of polynucleotides encoding heterogeneous and / or homogeneous transcription factors, The method according to any one of claims 1 to 8, achieved by...
10. The method according to any one of claims 1 to 9, wherein the transcription factor comprises the amino acid sequence shown in SEQ ID NO: 20, 21, or 26.
11. The method according to any one of claims 1 to 10, wherein the nuclear localization signal is a homologous or heterologous nuclear localization signal.
12. The method according to any one of claims 1 to 11, wherein the transcription factor does not stimulate a promoter used for the expression of the protein of interest.
13. The method according to any one of claims 1 to 12, wherein the recombinant protein of interest is an enzyme, a therapeutic protein, a food additive, or a feed additive.
14. The method according to claim 13, wherein the therapeutic protein is an antigen-binding protein.
15. The method according to any one of claims 1 to 14, further comprising the steps of overexpressing in a host cell at least one polynucleotide encoding at least one endoplasmic reticulum accessory protein, or engineering a host cell to overexpress it.
16. The method according to claim 15, wherein the endoplasmic reticulum accessory protein has the amino acid sequence shown in SEQ ID NO: 28, or a functional homolog thereof having at least 90% sequence identity with respect to the amino acid sequence shown in SEQ ID NO:
28.
17. The method according to any one of claims 1 to 14, further comprising the steps of overexpressing at least two polynucleotides encoding at least two endoplasmic reticulum accessory proteins in a host cell, or engineering a host cell to overexpress them.
18. a) The first endoplasmic reticulum accessory protein has the amino acid sequence shown in SEQ ID NO: 28, or a functional homolog thereof having at least 90% sequence identity with respect to the amino acid sequence shown in SEQ ID NO: 28, and b) The second endoplasmic reticulum accessory protein, i) The amino acid sequence shown in SEQ ID NO: 37, or a functional homolog thereof having at least 90% sequence identity with respect to the amino acid sequence shown in SEQ ID NO: 37, ii) The amino acid sequence shown in SEQ ID NO: 47, or its homologue, wherein the homologue has at least 90% sequence identity with respect to the amino acid sequence shown in SEQ ID NO:
47. It has, and in some cases, c) The third endoplasmic reticulum accessory protein is i) The amino acid sequence shown in SEQ ID NO: 55, or a functional homolog thereof having at least 90% sequence identity with respect to the amino acid sequence shown in SEQ ID NO:
55. The method according to claim 17, comprising:
19. The method according to any one of claims 1 to 14, further comprising the steps of overexpressing in a host cell at least one polynucleotide encoding one additional transcription factor, or engineering a host cell to overexpress it.
20. Additional transcription factors, at least: a) i) the amino acid sequence shown in SEQ ID NO: 65, or ii) a DNA-binding domain containing a functional homolog of the amino acid sequence shown in SEQ ID NO: 65 having at least 90% sequence identity with respect to the amino acid sequence shown in SEQ ID NO: 65, and b) Transcriptional activation domain The method according to claim 19, including the method described in claim 19.
21. The method according to claim 20, wherein the additional transcription factor comprises the amino acid sequences shown in SEQ ID NOs. 74-82.
22. The method according to any one of claims 19 to 21, wherein the additional transcription factor does not stimulate a promoter used for the expression of the protein of interest.
23. Recombinant eukaryotic host cells for producing recombinant proteins of interest, wherein the host cells are engineered to overexpress at least one polynucleotide encoding at least one transcription factor, the transcription factor being at least a) i) the amino acid sequence shown in SEQ ID NO: 5 or 6, or ii) a DNA-binding domain containing a functional homolog of the amino acid sequence shown in SEQ ID NO: 5 or 6 having at least 90% sequence identity with the amino acid sequence shown in SEQ ID NO: 5 or 6, and b) Transcriptional activation domain, and c) Nuclear localization signals Includes, Recombinant eukaryotic host cell, wherein the eukaryotic host cell further comprises a polynucleotide encoding a recombinant protein of interest, and the host cell is a yeast host cell selected from the group consisting of Pichia pastoris, Hansenula polymorpha, Trichoderma reesei, Aspergillus niger, Saccharomyces cerevisiae, Kluyveromyces lactis, Yarrowia lipolytica, Pichia methanolica, Candida boidinii, Komagataella, and Schizosaccharomyces pombe.
24. Recombinant eukaryotic host cell according to claim 23, wherein overexpression of the transcription factor increases the yield of the model protein scFv (SEQ ID NO: 13) and / or vHH (SEQ ID NO: 14) compared to host cells before engineering.
25. Recombinant eukaryotic host cell according to claim 23 or 24, wherein a polynucleotide encoding at least one transcription factor is incorporated into the genome of the host cell or contained in a vector or plasmid that is not incorporated into the genome of the host cell.
26. The recombinant eukaryotic host cell according to any one of claims 23 to 25, wherein the polynucleotide encoding at least one transcription factor encodes a heterogeneous or homogeneous transcription factor.
27. Overexpression of polynucleotides encoding heterologous transcription factors (i) A step of replacing or modifying a regulatory sequence operably linked to the polynucleotide encoding a heterologous transcription factor, (ii) The process of introducing one or more copies of a polynucleotide encoding a heterologous transcription factor into a host cell under the control of a promoter. The recombinant eukaryotic host cell according to claim 26, achieved by...
28. Overexpression of polynucleotides encoding the same transcription factor (i) A step of using a promoter that drives the expression of the polynucleotide encoding an allogeneic transcription factor, (ii) A step of replacing or modifying a regulatory sequence operably linked to the polynucleotide encoding an allogeneic transcription factor, (iii) The process of introducing one or more copies of a polynucleotide encoding an allogeneic transcription factor into a host cell under the control of a promoter. The recombinant eukaryotic host cell according to claim 26, achieved by...
29. Overexpression of polynucleotides i) A step of replacing the native promoter of a heterogeneous and / or homogeneous transcription factor with a different promoter operably linked to a polynucleotide encoding the homogeneous transcription factor. ii) A step of replacing the native termination factor sequences of heterologous and / or homologous transcription factors with more efficient termination factor sequences. iii) A step of replacing the coding sequences of heterologous and / or homologous transcription factors with codon-optimized coding sequences, wherein codon optimization is performed according to the codon usage frequency of the host cell. iv) A process of replacing a native positive regulatory element of a heterogeneous and / or homogeneous transcription factor with a more efficient regulatory element. v) The step of introducing another positive regulatory element that is not present in the natural expression cassette of heterogeneous and / or homogeneous transcription factors. vi) A step of deleting negative regulatory elements that are normally present in the natural expression cassettes of heterologous and / or homologous transcription factors, vii) The process of introducing one or more copies or combinations thereof of polynucleotides encoding heterologous and / or homologous transcription factors. A recombinant eukaryotic host cell according to any one of claims 23 to 28, achieved by the above.
30. Recombinant eukaryotic host cell according to any one of claims 23 to 29, wherein the transcription factor comprises the amino acid sequence shown in SEQ ID NO: 20, 21, or 26.
31. The recombinant eukaryotic host cell according to any one of claims 23 to 30, wherein the nuclear localization signal is a homologous or heterologous nuclear localization signal.
32. Recombinant eukaryotic host cell according to any one of claims 23 to 31, wherein the recombinant protein of interest is an enzyme, a therapeutic protein, a food additive, or a feed additive.
33. Recombinant eukaryotic host cell according to claim 32, wherein the therapeutic protein is an antigen-binding protein.
34. The recombinant eukaryotic host cell according to any one of claims 23 to 33, wherein the host cell is further engineered to overexpress at least one polynucleotide encoding at least one endoplasmic reticulum accessory protein.
35. The recombinant eukaryotic host cell according to claim 34, wherein the auxiliary protein has the amino acid sequence shown in SEQ ID NO: 28, or a functional homolog thereof having at least 90% sequence identity with respect to the amino acid sequence shown in SEQ ID NO:
28.
36. Recombinant eukaryotic host cell according to any one of claims 23 to 33, wherein the host cell is further engineered to overexpress at least two polynucleotides encoding at least two endoplasmic reticulum accessory proteins.
37. a) The first endoplasmic reticulum accessory protein has the amino acid sequence shown in SEQ ID NO: 28, or a functional homolog thereof having at least 90% sequence identity with respect to the amino acid sequence shown in SEQ ID NO: 28, and b) The second endoplasmic reticulum accessory protein, i) The amino acid sequence shown in SEQ ID NO: 37, or a functional homolog thereof having at least 90% sequence identity with respect to the amino acid sequence shown in SEQ ID NO: 37, ii) The amino acid sequence shown in SEQ ID NO: 47, or its homologue, wherein the homologue has at least 90% sequence identity with respect to the amino acid sequence shown in SEQ ID NO:
47. It has and / or c) The third endoplasmic reticulum accessory protein is i) The amino acid sequence shown in SEQ ID NO: 55, or a functional homolog thereof having at least 90% sequence identity with respect to the amino acid sequence shown in SEQ ID NO:
55. A recombinant eukaryotic host cell according to claim 36, having the characteristics of the following:
38. The recombinant eukaryotic host cell according to any one of claims 23 to 33, wherein the host cell is further engineered to overexpress at least one polynucleotide encoding one additional transcription factor.
39. Additional transcription factors, at least: a) i) the amino acid sequence shown in SEQ ID NO: 65, or ii) a DNA-binding domain containing a functional homolog of the amino acid sequence shown in SEQ ID NO: 65 having at least 90% sequence identity with respect to the amino acid sequence shown in SEQ ID NO: 65, and b) Transcriptional activation domain Recombinant eukaryotic host cells according to claim 38, including the above.
40. Recombinant eukaryotic host cell according to any one of claims 38 to 39, wherein the additional transcription factor comprises the amino acid sequence shown in SEQ ID NOs. 74 to 82.
41. Use of recombinant eukaryotic host cells according to any one of claims 23 to 40 for producing recombinant proteins of interest.