Engineered eukaryotic cell
Patent Information
- Application Number
- EP2024707272
- Authority / Receiving Office
- EP · EP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-02-13
- Filing Date
- 2024-02-13
- Publication Date
- 2025-12-24
AI Technical Summary
Current methods for improving recombinant protein production in Saccharomyces cerevisiae are labor-intensive, costly, and inefficient, often requiring multiple strain optimizations for each product due to the complexity of yeast metabolism and the limited scope of genetic changes explored, leading to sub-optimal production yields and high costs for biopharmaceuticals.
Identification and modification of specific genomic regions and genes, such as those on Saccharomyces cerevisiae chromosome XII 706198 - 717026, including UBP3 and BUD27, through quantitative trait loci analysis to create engineered eukaryotic cells with improved recombinant protein production capabilities, utilizing non-naturally occurring alleles and single nucleotide polymorphisms to enhance protein secretion and yield.
This approach enables faster, less expensive, and less labor-intensive strain improvement, resulting in increased recombinant protein production and secretion, making biopharmaceutical production more accessible and cost-effective by optimizing key genetic elements associated with protein production.
Smart Images

Figure GB2024050387_22082024_PF_FP
Abstract
Description
[0001] Engineered Eukaryotic Cell
[0002] The invention relates to recombinant or engineered eukaryotic cells, and particularly, although not exclusively, to recombinant or engineered yeast cells, such as Saccharomyces cerevisiae, exhibiting improved recombinant protein production. The invention is especially concerned with generating eukaryotic cells exhibiting increased protein secretion. The invention also extends to the use of recombinant or engineered eukaryotic cells for producing recombinant proteins, as well as the recombinant protein products obtained from such recombinant or engineered eukaryotic cells.
[0003] For over 40 years, microorganisms have been used to produce recombinant proteins. However, the cost of producing high-quality, biologically compatible, and commercially relevant quantities is often prohibitive. The baker's yeast, Saccharomyces cerevisiae, is well-known for producing high-quality, correctly folded heterologous proteins, often at lower costs and at levels greater than mammalian cells. Alternative yeasts, such as Komagataella species, K. phaffii, K. pastoris, and K. pseudopastoris, including industrial strains also known as Pichia pastoris, are also used to make recombinant products, such as biopharmaceutical proteins at competitive costs of goods. However, Pichia pastoris, is often associated with a requirement for methanol and oxygen for large-scale manufacturing, nonexchangeable plasmid systems requiring selection with toxic zeocin, and posttranslational quality issues for biopharmaceuticals. While these disadvantages are not generally associated with Saccharomyces cerevisiae, S. cerevisiae tends to have lower productivity. Therefore, it would be advantageous to increase the yields of recombinant proteins from S. cerevisiae, which will act to reduce the cost of goods sold (COGS) during large-scale manufacture, whilst retaining all the positive features of this expression host.
[0004] Alternative eukaryotic production hosts, such as mammalian cells, tend to be significantly more expensive for manufacturing biopharmaceuticals than yeasts. The media and growth requirements are also more expensive for mammalian cells, and processes are not free from animal-derived components.
[0005] Alternative microbial production hosts, including prokaryotic / bacterial systems like E. coli, lack eukaryotic cellular machinery, e.g., for protein folding and secretion, frequently resulting in insoluble inclusion bodies and endotoxin contamination issues affecting downstream purification. Previous work to improve product yields from S. cerevisiae has succeeded in developing manufacturing processes for a range of biopharmaceuticals, such as insulins, vaccines (e.g. virus-like particles), albumin and albumin fusion proteins. However, the production hosts are often sub-optimal, and the cost of goods sold limits access to biopharmaceutical products for all those who need them, thereby hindering the eradication of treatable human diseases. The development of improved production yeast, especially for S. cerevisiae, which has a long, safe history of human use and is most commonly used for the manufacture of yeast- derived biologies approved by the FDA, would be advantageous. Improving product yields from S. cerevisiae would be especially beneficial.
[0006] Existing methods for improving product yields include optimising the expression construct, e.g., the promoter, leader sequence (for secretory products), the coding sequence, terminator sequences, and the copy number and stability of the expression construct. However, improvements to the yeast genome can also lead to substantial improvements in productivity, product quality and other valuable bioprocess phenotypes. Nevertheless, the existing methods for improving the genome of the production yeast are sub-optimal, slow, expensive and labour- intensive. These methods have also focused on engineering the genome of a single strain or a family of closely related strains derived from commonly available laboratory strains, e.g. CEN.PK or relatives of W303. Therefore, the optimal starting strain was not used. The underlying biology is complex and not fully understood, especially the bottlenecks and limitations in metabolism affecting product yields and quality. Each protein product has different requirements, and many different limitations or bottlenecks can be encountered during production strain development. A production host improved for one product is seldom optimal for all products. Furthermore, methods for improvement have so far explored only a relatively small number of genetic changes. For example, random mutagenesis approaches tend to make limited types of genetic changes and frequently result in loss of function. There is also a risk of introducing unwanted mutations. Unless the beneficial mutation is then identified and re-engineered into a clean genetic background of the progenitor strain, there is a tendency for the accumulation of these undesirable mutations. Therefore, multiple rounds of random mutagenesis tend to be limited in the scope of improvements possible and can generate adverse phenotypes. On the other hand, rational genome engineering is unlikely to identify all options for strain improvement, some of which are not obvious targets for improvement, e.g. UBC4, M0T2, GHS1. The engineering of the yeast secretion system is also complicated by the involvement of many cross-reacting factors. The tight interdependence of each of these factors makes genetic modification difficult. Many attempts in strain engineering also fail to be transferrable from the laboratory to an industrial process. Each different recombinant protein product presents different challenges for production strain optimisation and generates a different burden on host cell metabolism. Therefore, the changes made to a single yeast strain improved for one product are unlikely to be optimal and may even be detrimental to the production of a different product. Consequently, bespoke production strain improvement is frequently required for each new product, which can also be sub-optimal, slow, expensive and labour-intensive. For example, each recombinant product might require an optimal combination of chaperone proteins for maximal secretion of the correctly folded protein. This might require multiple chaperones to be overexpressed, each requiring their expression level to be fine-tuned, thereby requiring the generation of multiple different strains and slow, expensive testing in fermenters.
[0007] It would, therefore, be advantageous to have faster, less expensive and less labour-intensive methods for strain improvement, which explore greater genetic diversity and make it easier to select strains improved for each recombinant product instead of relying solely on traditional mutagenesis and strain engineering methods. Improved strains for recombinant protein manufacture, which are genetically diverse to common laboratory strains, would also be valuable for investigating the manufacture of products that are difficult to express in existing systems. Furthermore, it would be beneficial to be able to identify the key regions of the yeast genome responsible for improvements in recombinant protein production, whether for intracellular products or secreted products.
[0008] Accordingly, in a first aspect of the invention, there is provided a recombinant or engineered eukaryotic cell exhibiting improved recombinant protein production compared to a corresponding wild-type, progenitor, or common laboratory strain cell, the recombinant or engineered eukaryotic cell comprising a non-naturally occurring combination of alleles associated with improved recombinant protein production, wherein the at least one allele is for a gene selected from a group consisting of: a gene present in Saccharomyces cerevisiae chromosome XII 706198 - 717026 inclusive (SEQ ID No: 11), UBP3 (YER151C), and BUD27 (YFL023W), or a homologue, orthologue or paralogue thereof, wherein the at least one gene is associated with recombinant protein production.
[0009] In a second aspect of the invention, there is provided a recombinant or engineered eukaryotic cell exhibiting improved recombinant protein production compared to a corresponding wild-type, progenitor, or common laboratory strain cell, the recombinant or engineered eukaryotic cell comprising at least one gene with a non- naturally occurring combination of single nucleotide polymorphisms (SNPs), wherein the at least one gene is selected from a group consisting of: a gene present in Saccharomyces cerevisiae chromosome XII 706198 - 717026 inclusive (SEQ ID No: 11), UBP3 (YER151C), and BUD27 (YFL023W), or a homologue, orthologue or paralogue thereof, wherein the at least one gene is associated with recombinant protein production.
[0010] As discussed in the Examples, using quantitative trait loci (QTL) analysis, the inventors have identified 16 genomic regions comprising approximately 3.3% of the total Saccharomyces cerevisiae genome containing genes, and different alleles of genes, responsible for the differential expression of the recombinant amylase- mCherry protein (see Table 1). In particular, the inventors have identified genes present in Saccharomyces cerevisiae chromosome XII 706198 - 717026 (e.g. MEC3 (YLR288C) and YLR287C), UBP3 (YER151C), and BUD27 (YFL023W), as key genes.
[0011] Advantageously, by identifying these genomic regions, and more specifically the SNPs within the genes, improved strains for recombinant protein manufacture can be provided, resulting in improvements in recombinant protein production, whether for intracellular products or secreted products.
[0012] From this QTL analysis, the inventors were able to identify genes in the QTLs that are associated with higher or lower levels of recombinant protein production. As such, the inventors believe that by modifying the expression of such genes (e.g. through knock-outs, overexpression, protein engineering or other methods to modify the genome or the expression levels of the protein), they can provide further improved strains exhibiting increased recombinant protein production. Accordingly, in a third aspect of the invention, there is provided a recombinant or engineered eukaryotic cell exhibiting improved recombinant protein production compared to a corresponding wild-type, progenitor, or common laboratory strain cell, the recombinant or engineered eukaryotic cell comprising at least one modified gene, wherein the at least one gene is selected from a group consisting of: a gene present in Saccharomyces cerevisiae chromosome XII 706198 - 717026 inclusive (SEQ ID No: 11), UBP3 (YER151C), and BUD27 (YFL023W), or a homologue, orthologue or paralogue thereof, wherein the at least one gene is associated with recombinant protein production.
[0013] Thus, preferably the recombinant or engineered eukaryotic cell of the invention exhibits improved recombinant protein secretion.
[0014] In one embodiment, the recombinant or engineered eukaryotic cell comprises a non-naturally occurring combination of alleles associated with improved recombinant protein production, wherein the at least one allele is for a gene selected from a group consisting of: a gene present in Saccharomyces cerevisiae chromosome XII 706198 - 717026 inclusive (SEQ ID No: 11), UBP3 (YER151C), and BUD27 (YFL023W), or a homologue, orthologue or paralogue thereof, and the recombinant or engineered eukaryotic cell comprises at least one modified gene, wherein the at least one gene is selected from a group consisting of: a gene present in Saccharomyces cerevisiae chromosome XII 706198 - 717026 inclusive (SEQ ID No: 11), UBP3 (YER151C), and BUD27 (YFL023W), or a homologue, orthologue or paralogue thereof.
[0015] In another embodiment, the recombinant or engineered eukaryotic cell comprises at least one gene with a non-naturally occurring combination of single nucleotide polymorphisms (SNPs), and at least one modified gene, wherein the at least one gene is selected from a group consisting of: a gene present in Saccharomyces cerevisiae chromosome XII 706198 - 717026 inclusive (SEQ ID No: 11), UBP3 (YER151C), and BUD27 (YFL023W), or a homologue, orthologue or paralogue thereof.
[0016] In one embodiment, there is provided a recombinant or engineered eukaryotic cell exhibiting improved recombinant protein production compared to a corresponding wild-type, progenitor, or common laboratory strain cell, the recombinant or engineered eukaryotic cell comprising at least one gene with a non-naturally occurring combination of single nucleotide polymorphisms (SNPs) associated with improved recombinant protein production, wherein the at least one gene is present within SEQ ID No: 11, or a homologue, orthologue or paralogue thereof.
[0017] In another embodiment, there is provided a recombinant or engineered eukaryotic cell exhibiting improved recombinant protein production compared to a corresponding wild-type, progenitor, or common laboratory strain cell, the recombinant or engineered eukaryotic cell comprising at least one gene with a non- naturally occurring combination of SNPs associated with improved recombinant protein production, wherein the at least one gene is selected from Table 7 and is present within SEQ ID No: 11 or a homologue, orthologue or paralogue thereof. In one embodiment, the non-naturally occurring combination of SNPs associated with improved recombinant protein production is selected from Table 7.
[0018] One embodiment of the nucleotide sequence of Saccharomyces cerevisiae chromosome XII 706198 - 717026 inclusive (from the S288c reference genome), is provided herein as SEQ ID No: 11, as follows:
[0019] TCATAAACGATGCTTCCTTTGTTTCGAGCCCAGCTGCCTAAATCTTTTTAAGGGCTCTCCATCTACCCAATACTTGAG AGATTCATTTACTTCCACTGAGTTAGCCTTATTGAATGCATCGATGTGGTTCGATTTCAGCAATTTTTTCATCCCTAA GCAACTGGGCAGGTATAGCCCTTTCACTTTCTCCCTCAATTCTTCTAAGACCTTTGCATTGAACGCTTCAGCGTTTGA AGATGGCATGTTAAAATTCTTGCTTATAAATCCGTTCTCGCACATTATATCGTACTTGAATGGTTTGTTGAACATGAG GCATTCATACGTCGTATTTGTGCCAAACTTCAATGGCAAAGAGACCGTTGTACCACCTTCGGTAATTAGTCCTAAGTT AGCAAAGGGGTATAGCAAATAAACCTTGTCATTTATACTGTACACAATGTCACATAACGCTACCAGTGCCGCGCTCAA CCCTATTGCTGGTCCATTCAAACAGCAAATTAAAACTTTGGAATGCTTGATGAAGGCATCAGTGACATAAACATTTCT AGCGACAAAATTTGACACCCACTTGCTTGTTTCCGAAGGATATTTATTGGTATCATCCCCTTGGGCTTTTGCAATACC CTTGAAATCAGCACCACTGGAAAAAAATCTACCACTGCTTTGTATAATTGTAAAATATACATCACGATTTCTGTCCGC TAGTTCTAGTAACTCTCCTAAATAAATATAGTCTTCACCTTCTAGTGCATTCAAATTGTCAGGGTTCATTAAGTGAAT AATGAAGAATGGTCCTTCAATACGATAACTGATTTTCTCATTTTGCCTAATTTCTTGCGACATATTGTTCTTCCTTCC TTTACTGTGCGAGCAATTTGTGCCATACACTCCAATTATCTATCCATATGTTCATCTTTCAGTTCTTATATATAGTGC TTCACTTTTATTAACCGCTCTTGGGGTAAAAGAGAAACGGTCACTGCTACAGCCTTTCGGAATAATATTTCTTATGCC CGTCGCCAATTTCAGATTCTTTCCTTCGTGCCGCATTTATTTTATTTTTCAAAATGACGAGAAGTTAAAATCTAAAAT CAGGAAAAAGAACAGTTGCAGATGACGAAATACTTCGAGCATTATCGCAATAAATAAGGCGTGTAAACGTATGTCAGA TATAGAATCATTGGGAGAGGCAGCAGGTTTATTTGAGGAGCCAGAGGATTTTCTTCCTCCTCCACCAAAACCTCATTT TGCAGAATATCAAAGGTCACACATCACAAAAGAGTCCAAATCAGATGTTAAAGACATCAAACTCCGTCTCGTTGGAAC GTCACCACTTTGGGGTCATCTCTTGTGGAATGCAGGAATATACACCGCAAATCACTTGGATTCTCATCCAGAATTAAT AAAGGGGAAGACTGTTTTAGAATTGGGTGCTGCTGCTGCTTTACCTTCCGTTATTTGTGCTTTGAATGGGGCTCAAAT GGTTGTTTCAACTGATTATCCAGATCCTGATTTAATGCAGAACATCGATTATAATATAAAGTCTAACGTTCCTGAAGA TTTTAATAATGTCAGTACGGAGGGTTATATTTGGGGAAACGATTATTCTCCATTGTTGGCACATATCGAAAAAATAGG TAACAATAATGGAAAGTTTGACTTAATCATTTTAAGTGATTTAGTCTTCAATCATACGGAACATCACAAATTACTTCA AACAACAAAGGATTTATTAGCTGAGAAAGGTCAGGCGTTGGTAGTATTTTCACCGCATAGACCAAAATTGTTGGAGAA GGATTTAGAATTCTTCGAATTGGCTAAGAACGAGTTCCATTTGGTTCCTCAGCTAATTGAAATGGTTAATTGGAAACC GATGTTTGACGAAGACGAGGAAACAATCGAAGTCAGATCTCGTGTTTATGCGTACTATTTAACACATGAAAAGTAGTA AGTGGGCGGAATTGGCATTCTCAACAATCATTTTTATGTTTGTAAACCAGTGTCCTTTGATATGTAACAATTAAATAA AATAAGTCATGATGGGGCACGAATTCGCAAAATGGTAGAAACACAGTAAATCGAAGAAAATTAAATAATTCGTTATAG ATATAAAGTAGTCGATTATAAAAACGTAAAAACATGAAGGCAGGGTACCTTGACGAAATAAAAGAGAGGCGGTGAGAT GAGGTGAAAAATAAAGAATATTGCAAAAAAAAATATTCGGACTCTATGAATCAATCTCTAATAACTTTAAAAGTAATT GCTTTCCAAATAAGAGAAATTACATTGGGTATAGACTGAGTCGCCGGAGTCATAAGCATAACATGTGGTTCCAGAAGC ACATTCCATGTAAACCCAAGCGCTATGATCACAAACGGCGAACTTCCCATCAGCAGAGCATGCAATTTCACCTTCAGT ACAGGTAGATTTACCGTTCAATTTACCAGCCGCATATTGAGCATTCAATTCTTTAGCCAATGTACGAGCTGTACTGTC TGAGCTTGTACTACCTGAGCTTGTACTACCTGAGCTTGTACTGGTCGTTGTTGGGGAAAGCGTACTAGTAGTAGCTGT CTGTAGGGAAACGACAGAAGAACTCTTCGTTGCTGGCGAAAGAGTACTAGTGATAGCTGTTTGAATTGGGGCCGAAGA AACTATGCTAGTAGTAGTTGTTTGAGGGGTCACCAAGGCAGCACTGGTTATTTGGGAAGATAGAGTAGTTTTCATACT TGTGATAGCAACTGAATTTAAAGTGCTCTCTGTTGTGGTGGTACCTAGACTAGATTTTGTCTTGGTGCTACTCGTCAA
[0020] TGTTTTTGTAGTTTGAGTAATTGATGTTTTGATAGCGCTGCTTGCAGTTGGAGATAAAGTAACTTTGCTTTTACTTTG
[0021] TGTAGATGTCGTAGATTGTGTGGTCTTTTTCTGAGAAGTTGAAGCAGATGAAGTTGAAGCAGATGAAGTTGAGGCTGC
[0022] TGAGGTTTTTGAGGTGGCAACTGTAGTAGTGGCGGTCTGGCTAGCACTTGTTAGCAAATTCTTCAAAATCTCAACATA
[0023] TGGTTCACCATTTAGCTCGTTGGAAAAGGCTTGAGATGCATCCCATAACGCAATACCACCAAAAGAACTTGAAGAGGC
[0024] AATATCTGCAATAGTTGATTCCAATAAAGAAGTGTCAGAAATATAACCAGAGCCAGCAGCAGAAGCAGAACCAGGTAA
[0025] ACCTAAGAACAGTTTGATATTTTTATTTGGGGATACAGTTTGAGCATAGGTTAACCAAGTATCCCAATTGAATTGACC
[0026] ACTCACACTGCAGTAATTATTGTAAAATTGGATGAACGCAAAATCAATGTCTGCATTTTCCAACAAGTCACCAACAGA
[0027] AGCATCCGGGTATGGACATTGTGGTGCGGCAGAAAGGTAATATTGCTTTGTACCTTCGGCAAACAAAGTTCTTAACTT
[0028] GGTAGCTAACGCACTATAGCCTACTTCGTTGTTGTTTTCAATATCAAAATCAAAACCATCAACGACTGCTGAGTCAAA
[0029] TGGTCTCTCACTGGCACCTGTACCTTCACCGAAAGTATCCCATAAAGTTTGTGCAAAAGTTTCCGCTTGAGAATCATC
[0030] TGAAAAGAGGTAGCTACCAGATGCACCACCTAATGATAATAGAACTTTCTTTCCTAGGGACTGGCAAGTTTCAATATC
[0031] TTCAGCAATCTGGGTGCAGTGAAGTAAGCCATCAGAAAAAGTATCAGAGCATGCGTTGGCAAAGTTCAAACCAAGGGT
[0032] TGGAAATTGGTTCAAGAAAGATAATAGGAAAATATCAGCATCAGAAGATTCACAGTAAGTAGCTAAGGATTCTTGCGT
[0033] TCCTGCTGAGTTTTGACCCCAATAAACAGCAATATTTGTGTTAGCAGACCTATCAAAGGCATCGGTTGGCAGTAGTAA
[0034] GAATTGTGTGAATAGAAGAATGATGTAAAGGAGTGACATTCTATTATTAATTATTTTATATTTAAATTAGAATTTCAA
[0035] TGTATTGGAAAAAGAGTGGTTTTAAAAAAGGTAGGTTGTGAAACGAGCGACAGTTATTATTTTTGGATATAAAAAGGT
[0036] TCAAAGGAATGACAGGTTTATTTTTTTTATGTTATTGAGTCTTAAATGAGTGAAAGATCTTCAATGTTATGAAAAAAC
[0037] TTGACTGATATAATCTTAATGTAATGTATTGTTGTTGTCCATTCTATAAATTATACAAGGAAAGGTAAGTGAGACGTT
[0038] TTTCCTCCATCCAAGGACCTAATACCTGCATCTTTTTATATCTTTGACCAATGCCTATGAAGCCAAATACTTCCGTCC
[0039] ATATTAGTTTTTTTTTGTTCCTTGAATCCTGCACCGAACCATCGAGGGGCTTAGTTCTTACACCAATGCTAATGTTTA
[0040] CTCAACATCTGAAACACATAAACAAATAAACAAACACCAAACATTGCACGTTAGAAACTCATATTTACTCGCACATAT
[0041] CTGTTAAGGCGAGGCTGGTTATTTTTTCAAGGGACCAGCATTTCCAATTTTTCACCAGCGGCAAAGAAAAATACCAAC
[0042] AAGCAGGAATCACACCCACGCATCTCATTTTGCGTTACTATGACAACTACAAAATTTGAATTTGTTAACGAGAATGAC
[0043] GTATATAAGTTCGATTTTGTTGAGGATATCCCACATATTCCCACCAACCAACCACTATATTAATGGACATGGTTTTTG
[0044] CCGTGACCGTGGATCTCGCAGCTACACAATGATATCTCTTGTATCCGTTTTAGCATGAATTTCTAAACAAGACTGATT
[0045] ATATACCATGCACTGTAACATTCAGAGAAAATATACTAAAACAATTCATTGAATTTTATCTCCCATGTGTCAATCCAT
[0046] TTCTTTCTTTTTGGATCACCTTCATATAGTTTACTCATTTGGATAAGAACTTTTTTAGTTACTGCAACTAATGCAGCC
[0047] TGTTCTTCTTTTACCTCCTCATTGGTAAAGTTCTTCGGCTCAAATTGTACAGTGGATACCACTTCATCCAAAAGTAAC
[0048] TGTATTTCCTTCAGATATGTATGTATACTGTCTAAAGTTTCTGCTTGATTCCTCTTAGGAGTAAAATCTTTTGACACC
[0049] AGAGTTTTCCTGAAAGTTGATACAAGTAATTTGATTAGCTTTATCTTCCTCGTAAACCCATCGAAAAACAACTTTATA
[0050] TTTTCATAGACCTTTTCCTGGTCAAATTCCTCTTTCTGTGAATCAGATTCGGATTCTGAATCCTCAAAATCCAAAAAT
[0051] ATATCGTCCGAGTTTGCTGAAAAATCGGGTTCCTCAAGCCACTCTTTTATTTCATTCATTGTATCGTCCATTATGGCC
[0052] ACATTATCTTTTAAGATGTTTGCTAGGATCCCATAAGGTCCAGCCTTGGAACAATTAGAGAGAGAATCGCACGCATTG
[0053] AATATTTTACCAACTGAGGTTAACCTTTCTTTGTCTAGAGATGCATTTTCGTCATTTTTCAATCTTTCTTGTAATTCG
[0054] GCAATGAAATCTCTCAAGCCATCGAGCAATTGTAGCGTACTTTCATCCAGTTGGTCTGTAAAGTATTTGGGGCAATCC
[0055] TTGTTATTGTAAAACAAAGGGAATAAACTTAGTAAGTAGAATAATGGCCTACTAAAATTTTGGATTTCGGTAATCACA
[0056] ACTTTATGATTGTTATCAAAAGTTCCTGGCTTGCAAACAATACCTATTTTTGTGCAATGTGCCTTCAATACAGAGGCT
[0057] AATTTGTCTAGTTCCTTAGTTGGGGTGGAACCTTGTAGTTTGGTAGTAGAGGAGATTTTACGCAAGTCTTCCGGCTTT
[0058] TTGTATGGCACTAAAAACTGTTCATCTATAGAATTCAGCAGTTCCAATAATTTGACGTCATCCTTTCTATCACTGCTT
[0059] CCAGTAGACATTATTGTATAACTGGTATTTCTCTTATAAAGTTTATAATAGGCGTCTCTGGTGATGCTTGTTGACCTC
[0060] AATGGATTTTGTTAAAAGGAGGCTTCATGTCATCGTCAAGATCATTGCAGTTTTTTTCTAGTTTTTTTTTTTTCATTG
[0061] CCACATTATGAAGCGCCCATTCAAAAACAACTAGAAAAACGTTTGCCAACCCTAGGTACGGCAAGATCCTCCTTATTA
[0062] ATAAAAGCATCACACCAAGACAGCAAAGAAGTAAATCGATAATATGAATAGCATAATTTTGCGAAGTTAAAGTCCTAA
[0063] AGAAATAATTTATAGTTAACACTACTTACATAATTTCAATACTGCTTTATATTTAAAGTAAAAAACGCATAACGGAAA
[0064] TACAAATACAAGTCAAAAACTATCTTGAAGTAAAGATGGCATGACTGCTATTTTGTTTCTTGTTACTTTTAAGCTCCT
[0065] TTTCCACAACTGTTAATTTTCTTATTGGACGGATGGACCTGGGTTCATTCTTCTCTTACCGTTAACCAAGGTAACGTT
[0066] AACGAATCTTCTGGTGTACAATAATCTTTTGTAAGCACGACCCTTAGGCTTCTTTGGCTTTTCGGTTTTTTCAACCTT
[0067] TGGTGTTTGAGACTTGACTTTACCAGCACGAGCTAGAGAACCATGAACTTTAGCCTAAATAAATAATATTAATATTAA
[0068] AAATTAATGACCATGTTAGTATAGTGAATTTTTGAATTATTTGGCGTATCAGCAAGGATGCTAAAAATAAACTTCACT
[0069] CTGCTATTTTCTCAAAGCCTTTTTAACTCTCCGTCAAAAAATGTTGTTATTAAACTCGACATTATGTTAAGTCCACGC
[0070] TGTTGCCCTCGAAAAATCTGTGTCAAGCAATTACGTTTTTGATGTAACAGAACTCAAATTTAAAAAGTTTCCTATCAT
[0071] ATTCTCAGGGTGGAGCCAGTTTAAAGATATATTGCGTTTTATTTCTTTTTCTGCGGTATATATAGATATGAAACACCA
[0072] GTAATATATTCAAATTAGCTATTTCTTGTTATTCATTCCTCGCTGAAACGTCCATTTTTTTCATATTTTAAAATCCTA
[0073] ATGAGGCATTACGTACCATATTTGCGTAGTTTTTGTATGGGGATATGGAAGATTATTACAGTGGCAAATTTAATGTGC
[0074] AGTTCATGTAGCCTTTAAAATGCAACATACCGTTCGTAAGGCTGTTTTGTTCTATCTATTTTTTCTGAGGTTTCCGAC
[0075] AACCTCACGACATTTTTTTTTTTCGAACGGAAAAAAGAAAAGCAGTTAGTATGTAAAGCACGGATTTTTTTATACATA
[0076] TGTACGGATGTGATAGTCTGCAACTCTGCAACTCTGCAGTTCTCAGCCATGTCAGATTTAAAGAGTTTCCTTAGCAAC
[0077] GTAGCAAAGAAATGTACCGCTGTAGGGTTTACAAGCCCTTCGATCTTGCTATATAATATATGATTTGTCCTCTTTCCC
[0078] TTGGTTTTTCCACATCTTCGCTATCTTCCAAGCTACCACGATCCAACGAACAGTGGAATACGCAGGATTCGTCGTGAG
[0079] AAATCGCCAAAACAACTTCTTCAAATGCAGCGTATAACTTGGAACACACCTTCCAATCTTTGCAACGGATGATTACTT
[0080] CATGTGTGGACGAACTTTCCTGTTCAGCCTTTTCCACCATAACGGATATGTCATTAAATTCAGTATCACCGCTAGTAT
[0081] CAGCTGTGTAAATGTTTCCCCGCGTATCTGCGATCGAGCTATCCTCAATTCTTAATAAATCTTCATCGTAGCGGATAT
[0082] TTTCTTCCATCTCTCGATCTCTAGTATTGGTATATAGTGAAGACATCGGTTTATCCGCTTCGATAATCGGAAGAGATC
[0083] CTTCCTCCTGCCGGCCGTCTGTGTCGATGTGCTGGTTTTGGGAAGGATTGTCAGTGAGCCCTTCTTGGCGTTGTATCA CAGAGTCTAAGGGTCCATTCCAGCATATTTCCAAATGCCAATCTAATTCATTCACAATTATCTTAAGTTCTACATCAT CACCTTCATTTCCATGCTCCTTTTTTTTGACTCCCATTAAATGAATGTGGTTGACATTACTGTACCGTTCAACACGTC TAATAAACCCGTGGAAGGCGGAGCCAAACTCGCCCGATATTGGTGGCAGCTTGTACATCATCAGTTGAATATAGTTAA TCATTGGCTCTTGTATTCGCGTATCTTGTGCTCGGAATAATAGTTTGACAGGTACTTTGAACGAATGCATGATAACCT TATTGCTTGCGAGTAGATTTCCTGTGCCTACTGTGGTTGGTAAACCATTGTGCTCATCCACTCCCCCATTCATAACTA TAGCGTCATTTGGGCCGCTAGTATGTACAATTTCTTCGAATGTTATACCTAGAGCACAAATGGGGTTCGGTTTTGAGG TTGTATCTACGCCTCCTGCTGTTCCTCCTGAAAGTGTTCCGTTATTTGTATTCCATTCTGGCATTGACTGTAGTTTTA TAGTCATATTAGAGGAAGATCCTTGGTTCATTACTCGATCATACCTTTTAAACACACTCAATAAAGAATCACAATTAC ACTCCATTGTTATTGTATTGAGCTCGCGAGCTGATATAACTGTATATAATCTGAATACATCATGAGGAATGGTACACC AAAGCTGACCAGTATCCCCTCGTAATATTGTACCGTTGTTACTGCTGTTGAGTGATGATTTTGGAGTGGATATTATTG TCAATCTTTCACTATTAAATCTTAAGATAGCCGTCTTTCGTAGCGAAGCAACTGTATTGATAGTAGTTCTTAGCAATT TATAATCATCAGGTGCTTCACAACCATTTACTATCAATTTTAATTTCATTTAACTGAATTAAGACACACCTTTTGTCT TCTTTTTTCTCTCATCATCTCCGTATGTTTATGTTGCTATTTTGATGTAAATAAAAAAGTTGAATAATAGACGAGGGC AAGTATAACTCGCCTATATTTGTAGCCGCAACCATTGAAAAAGAGCCATGAATATGGGAAACTAGTTGCACATAAAAA TGCTGAAATTTAGAATTAGGCCAGTGAGACATATACGGTGTTATAAACGACACGCATATTTCTTACGATATAACCATA CGACTACCCCTGCACAGAAGTTACAAGCACAGATCGAGCAAATACCTCTCGAAAATTACAGAAATTTCTCTATAGTTG CCCATGTTGACCATGGGAAGTCAACCTTAAGTGACAGACTGCTGGAAATAACGCATGTCATCGATCCCAATGCGAGAA ATAAACAAGTTTTGGATAAATTGGAAGTCGAAAGAGAAAGAGGTATTACTATAAAGGCGCAAACATGTTCGATGTTTT ATAAAGATAAGAGGACCGGAAAAAACTATCTTTTACATTTAATTGACACGCCAGGACATGTGGACTTCAGAGGTGAAG TTTCACGGTCATATGCGTCTTGTGGGGGAGCAATTCTTTTGGTTGATGCATCACAAGGCATACAAGCACAGACGGTTG CTAATTTTTATTTAGCCTTCAGTTTAGGATTGAAATTAATTCCAGTAATAAACAAAATTGATCTAAATTTTACAGATG TTAAACAGGTAAAGGATCAGATAGTGAATAACTTTGAGCTCCCCGAGGAAGATATAATCGGAGTAAGTGCTAAAACAG GATTAAATGTAGAGGAACTGTTACTACCGGCTATAATTGATCGTATACCACCACCAACGGGGAGGCCTGATAAACCCT TCAGAGCATTATTAGTGGATTCTTGGTACGACGCATACTTAGGAGCGGTTCTTCTAGTGAATATTGTTGATGGTTCTG TACGTAAAAATGACAAGGTTATTTGTGCTCAGACAAAAGAAAAATACGAAGTCAAAGATATTGGAATCATGTATCCTG ACAGAACCTCTACAGGTACGCTAAAGACAGGACAGGTTGGCTATCTAGTGCTGGGAATGAAGGATTCTAAAGAAGCAA AAATTGGAGATACTATAATGCATTTAAGTAAAGTAAATGAAACGGAAGTACTTCCCGGATTTGAAGAACAAAAACCCA TGGTATTTGTGGGTGCTTTCCCGGCTGATGGGATTGAATTCAAAGCCATGGATGATGATATGAGTAGACTTGTCCTCA ACGATAGGTCAGTTACTTTGGAACGTGAGACCTCCAATGCTTTGGGTCAAGGTTGGAGATTGGGCTTTTTAGGATCTT TACATGCATCTGTTTTTCGTGAACGACTAGAAAAAGAGTATGGTTCGAAATTGATCATTACTCAACCCACAGTTCCTT ATTTGGTGGAGTTTACCGATGGTAAGAAAAAGCTTATAACAAATCCGGATGAGTTTCCAGACGGAGCAACAAAGAGGG TGAACGTTGCTGCTTTCCATGAACCGTTTATAGAGGCAGTTATGACATTGCCCCAGGAATATTTAGGTAGTGTCATAC GCTTATGCGATAGTAATAGAGGAGAACAAATTGATATAACATACCTAAACACCAATGGACAAGTGATGTTAAAATATT ACCTTCCGCTATCGCATCTAGTTGATGACTTTTTTGGTAAATTAAAATCGGTGTCCAGAGGATTTGCCTCTTTAGATT ATGAGGATGCTGGCTATAGAATTTCTGATGTTGTAAAACTGCAACTCTTGGTTAATGGAAATGCGATTGATGCCTTGT CAAGAGTACTTCATAAATCGGAAGTAGAGAGAGTGGGTAGAGAATGGGTGAAGAAGTTTAAAGAGTATGTTAAATCAC AATTATATGAGGTCGTTATACAGGCCCGAGCTAATAACAAGATAATAGCTAGAGAAACAATTAAGGCAAGAAGAAAAG ATGTTCTCCAAAAGCTGCATGCTTCTGATGTCTCACGAAGGAAAAAACTTTTGGCGAAACAGAAAGAGGGTAAAAAGC ATATGAAAACTGTAGGTAATATTCAAATCAACCAAGAGGCATATCAGGCTTTTTTGCGCCGTTAG
[0084] [SEQ ID No: 11]
[0085] Accordingly, in one embodiment, Saccharomyces cerevisiae chromosome XII 706198 - 717026 inclusive comprises a nucleotide sequence as set out in SEQ ID No: 11, or a fragment or variant thereof.
[0086] In one embodiment, Saccharomyces cerevisiae chromosome XII 706198 - 717026 inclusive comprises a nucleotide sequence as set out in SEQ ID No: 11, or a fragment or variant thereof, comprising at least one SNP as defined in Table 7.
[0087] The skilled person would appreciate that the term "Saccharomyces cerevisiae chromosome XII 706198 -717026 inclusive" includes genes or coding sequences from the beginning of the YLR284C coding sequence to the end of the YLR289W coding sequence from the S288c reference genome. Thus, the skilled person would appreciate that the term "Saccharomyces cerevisiae chromosome XII 706198 - 717026 inclusive" includes the following coding sequences: YLR284C, YLR285W, YLR285C-A, YLR286C, YLR286W-A, YLR287C, YLR287C-A, MEC3 (YLR288C), and YLR289W, and the sequences between these coding regions, including any intergenic regions. It will be appreciated that these genes are part of a single QTL.
[0088] Accordingly, in one embodiment, the at least one gene is selected from a group consisting of: YLR284C, YLR285W, YLR285C-A, YLR286C, YLR286W-A, YLR287C, YLR287C-A, YLR288C, and YLR289W, or a homologue, orthologue or paralogue thereof.
[0089] In one embodiment, the at least one gene is MEC3 (YLR288C) and / or YLR287, or a homologue, orthologue or paralogue thereof. In one embodiment, the at least one gene is MEC3 (YLR288C), or a homologue, orthologue or paralogue thereof. In one embodiment, the at least one gene is YLR287C, or a homologue, orthologue or paralogue thereof.
[0090] In one embodiment, the at least one gene is selected from a group consisting of: NNT1 (YLR285W), CTS1 (YLR286C), YLR287C, RPS30A (YLR287C-A), MEC3 (YLR288C) and GUF1 (YLR289W), or a homologue, orthologue or paralogue thereof.
[0091] In one embodiment, the at least one gene is CTS1 (YLR286C), or a homologue, orthologue or paralogue thereof. In one embodiment, the at least one gene is RPS30A (YLR287C-A), or a homologue, orthologue or paralogue thereof. Alternatively, in another embodiment, the at least one gene is UBP3 (YER151C), or a homologue, orthologue or paralogue thereof. In another embodiment, the at least one gene is BUD27 (YFL023W), or a homologue, orthologue or paralogue thereof.
[0092] In one embodiment, the at least one gene may be at least two, at least three, at least four, at least five, or at least six genes selected from a group consisting of:
[0093] YLR284C, YLR285W, YLR285C-A, YLR286C, YLR286W-A, YLR287C, YLR287C-A, YLR288C, YLR289W, UBP3 (YER151C), and BUD27 (YFL023W), or a homologue, orthologue or paralogue thereof. In one embodiment, the at least one gene may be at least seven, at least eight, at least nine, at least ten, or at least eleven genes selected from a group consisting of: YLR284C, YLR285W, YLR285C-A, YLR286C, YLR286W-A, YLR287C, YLR287C-A, YLR288C, YLR289W, UBP3 (YER151C), and BUD27 (YFL023W),or a homologue, orthologue or paralogue thereof.
[0094] In one embodiment, the at least one gene may be at least two, at least three, or at least four genes selected from a group consisting of: MEC3 (YLR288C), YLR287C, UBP3 (YER151C), and BUD27 (YFL023W), or a homologue, orthologue, or paralogue thereof.
[0095] In one embodiment, the at least two genes may be MEC3 (YLR288C) and YLR287C, or a homologue, orthologue or paralogue thereof. In one embodiment, the at least two genes may be MEC3 (YLR288C) and UBP3 (YER151C), or a homologue, orthologue or paralogue thereof. In one embodiment, the at least two genes may be MEC3 (YLR288C) and BUD27 (YFL023W), or a homologue, orthologue or paralogue thereof. In one embodiment, the at least two genes may be YLR287C and UBP3 (YER151C), or a homologue, orthologue or paralogue thereof. In one embodiment, the at least two genes may be YLR287C and BUD27 (YFL023W), or a homologue, orthologue or paralogue thereof. In one embodiment, the at least two genes may be UBP3 (YER151C) and BUD27 (YFL023W), or a homologue, orthologue or paralogue thereof.
[0096] In one embodiment, the at least three genes may be MEC3 (YLR288C), YLR287C and UBP3 (YER151C), or a homologue, orthologue or paralogue thereof. In one embodiment, the at least three genes may be MEC3 (YLR288C), YLR287C and BUD27 (YFL023W), or a homologue, orthologue or paralogue thereof. In one embodiment, the at least three genes may be MEC3 (YLR288C), UBP3 (YER151C) and BUD27 (YFL023W), or a homologue, orthologue or paralogue thereof. In one embodiment, the at least three genes may be YLR287C, UBP3 (YER151C), and BUD27 (YFL023W), or a homologue, orthologue or paralogue thereof. In another embodiment, the at least four genes are MEC3 (YLR288C), YLR287C, UBP3 (YER151C), and BUD27 (YFL023W), or a homologue, orthologue, or paralogue thereof.
[0097] In one embodiment, the recombinant or engineered eukaryotic cell according to the first aspect comprises at least one additional non-naturally occurring allele associated with improved recombinant protein production, wherein the at least one additional allele is for at least one additional gene selected from a list of genes in Table 1, or a homologue, orthologue or paralogue thereof, wherein the at least one additional gene is associated with recombinant protein production.
[0098] In one embodiment, the recombinant or engineered eukaryotic cell according to the second aspect comprises at least one additional gene with a non-naturally occurring combination of single nucleotide polymorphisms (SNPs), wherein the at least one additional gene is selected from a list of genes in Table 1, or a homologue, orthologue or paralogue thereof, wherein the at least one additional gene is associated with recombinant protein production.
[0099] In one embodiment, the recombinant or engineered eukaryotic cell according to the third aspect comprises at least one additional modified gene, wherein the at least one additional modified gene is selected from a list of genes in Table 1, or a homologue, orthologue or paralogue thereof, wherein the at least one additional gene is associated with recombinant protein production.
[0100] In one embodiment, the at least one additional gene may be at least two, at least three, at least four, or at least five genes selected from the list of genes in Table 1, or a homologue, orthologue, or paralogue thereof. In another embodiment, the at least one additional gene may be at least six, at least seven, at least eight, at least nine or at least ten genes selected from the list of genes in Table 1, or a homologue, orthologue, or paralogue thereof.
[0101] Alternatively, in another embodiment, the at least one additional gene may be at least eleven, at least twelve, at least thirteen, at least fourteen or at least fifteen genes selected from the list of genes in Table 1, or a homologue, orthologue, or paralogue thereof. Alternatively, in another embodiment, the at least one additional gene may be at least sixteen, at least seventeen, at least eighteen, at least nineteen or at least twenty genes selected from the list of genes in Table 1, or a homologue, orthologue, or paralogue thereof.
[0102] Alternatively, in another embodiment, the at least one additional gene may be at least twenty-two, at least twenty-four, at least twenty-six, at least twenty-eight or at least thirty genes selected from the list of genes in Table 1, or a homologue, orthologue, or paralogue thereof.
[0103] Preferably, the at least one additional gene is selected from a group consisting of: GTB1, GCD7, SLM 1, PIP2, YIL089W, ZTA1, UTP25, KTR7, MFB1, TRM7, PUF4, CHS7, SHQ1, CRF1, PRP6, C0G3, DPH1, SPT2, XBP1, PRK1, REB1, RCI37, YBR053C, YHR131C, BMT5, CNM1, C0Q11, YIL092W, DFP4, DSE2, FMC1, VMR1, RSO55, MAM33, CRP1, AIR1, MYG1, NDD1, YCK1, ARO9, SDS3, ECM2, CBP2, MOB1, ADRI, PDR1, and BEM2.
[0104] In one preferred embodiment, the at least one additional gene is selected from a group consisting of: GTB1, GCD7, SLM1, PIP2, YIL089W, ZTA1, UTP25, KTR7, MFB1, TRM7, PUF4, CHS7, SHQ1, CRF1, PRP6, COG3, DPH1, SPT2, XBP1, PRK1, and REB1.
[0105] More preferably, the at least one additional gene is selected from a group consisting of: GTB1, GCD7, SLM1, and PIP2. In another embodiment, the at least one additional gene is selected from a group consisting of: YIL089W, ZTA1, UTP25, KTR7, MFB1, TRM7, PUF4, CHS7, SHQ1, CRF1, PRP6, COG3, DPH1, SPT2, XBP1, PRK1, and REB1.
[0106] In another embodiment, the at least one additional gene is selected from a group consisting of: RCI37, YBR053C, YHR131C, BMT5, CNM 1, COQ11, YIL092W, DFP4, DSE2, FMC1, VMR1, RSO55, MAM33, CRP1, AIR1, MYG1, NDD1, YCK1, ARO9, SDS3, ECM2, CBP2, MOB1, ADRI, PDR1, and BEM2.
[0107] A gene or allele of this invention includes upstream and downstream regions from the protein coding sequence, which may be at least lOObp, at least 200bp, at least 300bp, at least 4OObp, at least 500bp, at least 600bp, at least 700bp, at least 800bp, at least 900bp, at least lOOObp either side of the coding sequence, or up to the next adjacent gene or coding sequence. SNPs present in these regions are also claimed in this invention, as are SNPs within protein coding sequences which are synonymous and do not result in amino acid changes.
[0108] MEC3 (YLR288C)
[0109] In one embodiment, when the at least one gene is MEC3, most preferably the gene comprises the SNP T1362G. In this embodiment, a guanine replaces a thymine at position 1362 of the nucleotide sequence. Preferably, this SNP results in the amino acid substitution Asp454Glu, i.e. glutamine replaces aspartic acid at position 454 of the amino acid sequence. Alternatively, in a next preferred embodiment, the gene comprises an adenine or a cytosine at position 1362 of the nucleotide sequence.
[0110] Least preferably, the gene comprises a thymine at position 1362 of the nucleotide sequence. In another embodiment, when the at least one gene is MEC3, most preferably the gene comprises the SNP C359T. In this embodiment, a thymine replaces a cytosine at position 359 of the nucleotide sequence. Preferably, this SNP results in the amino acid substitution Thrl20Ile, i.e. isoleucine replaces threonine at position 120 of the amino acid sequence. Alternatively, in a next preferred embodiment, the gene comprises an adenine or a guanine at position 359 of the nucleotide sequence.
[0111] Least preferably, the gene comprises a cytosine at position 359 of the nucleotide sequence.
[0112] YLR287C In one embodiment, when the at least one gene is YLR287C, most preferably the gene comprises the SNP G19A. In this embodiment, an adenine replaces a guanine at position 19 of the nucleotide sequence. Preferably, this SNP results in the amino acid substitution Asp7Asn, i.e. asparagine replaces aspartic acid at position 7 of the amino acid sequence. Alternatively, in a next preferred embodiment, the gene comprises a thymine or a cytosine at position 19 of the nucleotide sequence. Least preferably, the gene comprises a guanine at position 19 of the nucleotide sequence.
[0113] In one embodiment, when the at least one gene is YLR287C, the gene comprises the SNP A1001G. In this embodiment, a guanine replaces an adenine at position 1001 of the nucleotide sequence. Preferably, this SNP results in the amino acid substitution Tyr334Cys, i.e. cysteine replaces tyrosine at position 334 of the amino acid sequence. Alternatively, in a next preferred embodiment, the gene comprises a thymine or a cytosine at position 1001 of the nucleotide sequence. Least preferably, the gene comprises an adenine at position 1001 of the nucleotide sequence.
[0114] In one embodiment, when the at least one gene is YLR287C, the gene comprises the SNP G992A. In this embodiment, an adenine replaces a guanine at position 992 of the nucleotide sequence. Preferably, this SNP results in the amino acid substitution Ser331Asn, i.e. asparagine replaces serine at position 331 of the amino acid sequence. Alternatively, in a next preferred embodiment, the gene comprises a thymine or a cytosine at position 992 of the nucleotide sequence. Least preferably, the gene comprises a guanine at position 992 of the nucleotide sequence. In one embodiment, when the at least one gene is YLR287C, the gene comprises the SNP T722C. In this embodiment, a cytosine replaces a thymine at position 722 of the nucleotide sequence. Preferably, this SNP results in the amino acid substitution Leu241Ser, i.e. serine replaces leucine at position 241 of the amino acid sequence. Alternatively, in a next preferred embodiment, the gene comprises a guanine or an adenine at position 722 of the nucleotide sequence. Least preferably, the gene comprises a thymine at position 722 of the nucleotide sequence. RPS30A (YLR287C-A)
[0115] In one embodiment, when the at least one gene is RPS30A, most preferably the gene comprises the SNP G148A. In this embodiment, an adenine replaces a guanine at position 148 of the nucleotide sequence. Preferably, this SNP results in the amino acid substitution Val50Ile, i.e. isoleucine replaces valine at position 50 of the amino acid sequence. Alternatively, in a next preferred embodiment, the gene comprises a cytosine or thymine at position 148 of the nucleotide sequence. Least preferably, the gene comprises a guanine at position 148 of the nucleotide sequence. CTS1 (YLR286C)
[0116] In one embodiment, when the at least one gene is CTS1, most preferably the gene comprises the SNP T47C. In this embodiment, a cytosine replaces a thymine at position 47 of the nucleotide sequence. Preferably, this SNP results in the amino acid substitution Leul6Pro, i.e. proline replaces leucine at position 16 of the amino acid sequence. Alternatively, in a next preferred embodiment, the gene comprises an adenine or a guanine at position 47 of the nucleotide sequence. Least preferably, the gene comprises a thymine at position 47 of the nucleotide sequence.
[0117] In another embodiment, when the at least one gene is CTS1, most preferably the gene comprises the SNP C962G. In this embodiment, a guanine replaces a cytosine at position 962 of the nucleotide sequence. Preferably, this SNP results in the amino acid substitution Thr321Ser, i.e. serine replaces threonine at position 321 of the amino acid sequence. Alternatively, in a next preferred embodiment, the gene comprises an adenine or a thymine at position 962 of the nucleotide sequence. Least preferably, the gene comprises a cytosine at position 962 of the nucleotide sequence. In another embodiment, when the at least one gene is CTS1, the gene comprises the SNP G1570A. In this embodiment, an adenine replaces a guanine at position 1570 of the nucleotide sequence. Preferably, this SNP results in the amino acid substitution Asp524Asn, i.e. asparagine replaces aspartic acid at position 524 of the amino acid sequence. Alternatively, in a next preferred embodiment, the gene comprises a cytosine or a thymine at position 1570 of the nucleotide sequence.
[0118] Least preferably, the gene comprises a guanine at position 1570 of the nucleotide sequence. NNT1 (YLR285W)
[0119] In one embodiment, when the at least one gene is NNT1, most preferably the gene comprises the SNP G413C. In this embodiment, a cytosine replaces a guanine at position 413 of the nucleotide sequence. Preferably, this SNP results in the amino acid substitution Serl38Thr, i.e. threonine replaces serine at position 138 of the amino acid sequence. Alternatively, in a next preferred embodiment, the gene comprises an adenine or a thymine at position 413 of the nucleotide sequence.
[0120] Least preferably, the gene comprises a guanine at position 413 of the nucleotide sequence. GUF1 (YLR289W)
[0121] In one embodiment, when the at least one gene is GUF1, most preferably the gene comprises the SNP C779T. In this embodiment, a thymine replaces a cytosine at position 779 of the nucleotide sequence. Preferably, this SNP results in the amino acid substitution Ser260Phe, i.e. phenylalanine replaces serine at position 260 of the amino acid sequence. Alternatively, in a next preferred embodiment, the gene comprises an adenine or a guanine at position 779 of the nucleotide sequence.
[0122] Least preferably, the gene comprises a cytosine at position 779 of the nucleotide sequence. BUD27 (YFL023W)
[0123] In one embodiment, when the at least one gene is BUD27, most preferably the gene comprises the SNP A1268G. In this embodiment, a guanine replaces an adenine at position 1268 of the nucleotide sequence. Preferably, this SNP results in the amino acid substitution Glu423Gly, i.e. glycine replaces glutamic acid at position 423 of the amino acid sequence. Alternatively, in a next preferred embodiment, the gene comprises a thymine or a cytosine at position 1268 of the nucleotide sequence. Least preferably, the gene comprises an adenine at position 1268 of the nucleotide sequence.
[0124] In another embodiment, when the at least one gene is BUD27, most preferably the gene comprises the SNP T1638G. In this embodiment, a guanine replaces a thymine at position 1638 of the nucleotide sequence. Preferably, this SNP results in the amino acid substitution Phe546Leu, i.e. leucine replaces phenylalanine at position 546 of the amino acid sequence. Alternatively, in a next preferred embodiment, the gene comprises an adenine or a cytosine at position 1638 of the nucleotide sequence. Least preferably, the gene comprises a thymine at position 1638 of the nucleotide sequence.
[0125] UBP3 (YER151O
[0126] In one embodiment, when the at least one gene is UBP3, most preferably the gene comprises the SNP G770A. In this embodiment, an adenine replaces a guanine at position 770 of the nucleotide sequence. Preferably, this SNP results in the amino acid substitution Ser257Asn, i.e. asparagine replaces serine at position 257 of the amino acid sequence. Alternatively, in a next preferred embodiment, the gene comprises a thymine or a cytosine at position 770 of the nucleotide sequence. Least preferably, the gene comprises a guanine at position 770 of the nucleotide sequence.
[0127] In another embodiment, when the at least one gene is UBP3, most preferably the gene comprises the SNP G643A. In this embodiment, an adenine replaces a guanine at position 643 of the nucleotide sequence. Preferably, this SNP results in the amino acid substitution Ala215Thr, i.e. threonine replaces alanine at position 215 of the amino acid sequence. Alternatively, in a next preferred embodiment, the gene comprises a thymine or a cytosine at position 643 of the nucleotide sequence. Least preferably, the gene comprises a guanine at position 643 of the nucleotide sequence.
[0128] In another embodiment, when the at least one gene is UBP3, most preferably the gene comprises the SNP A230C. In this embodiment, a cytosine replaces an adenine at position 230 of the nucleotide sequence. Preferably, this SNP results in the amino acid substitution His77Pro, i.e. proline replaces histidine at position 77 of the amino acid sequence. Alternatively, in a next preferred embodiment, the gene comprises a thymine or a guanine at position 230 of the nucleotide sequence. Least preferably, the gene comprises an adenine at position 230 of the nucleotide sequence.
[0129] GTB1 (YDR221W) In one embodiment, when the at least one additional gene is GTB1, most preferably the gene comprises the SNP T955C. In this embodiment, a cytosine replaces a thymine at position 955 of the nucleotide sequence. Preferably, this SNP results in the amino acid substitution Tyr319His, i.e. histidine replaces tyrosine at position 319 of the amino acid sequence. Alternatively, in a next preferred embodiment, the gene comprises an adenine or a guanine at position 955 of the nucleotide sequence.
[0130] Least preferably, the gene comprises a thymine at position 955 of the nucleotide sequence.
[0131] In another embodiment, when the at least one additional gene is GTB1, most preferably the gene comprises the SNP G1633A. In this embodiment, an adenine replaces a guanine at position 1633 of the nucleotide sequence. Preferably, this SNP corresponds to the amino acid substitution Asp545Asn, i.e. asparagine replaces aspartic acid at position 545 of the amino acid sequence. Alternatively, in a next preferred embodiment, the gene comprises a thymine or a cytosine at position 545 of the nucleotide sequence. Least preferably, the gene comprises a guanine at position 545 of the nucleotide sequence.
[0132] GCD7 (YLR291C)
[0133] In one embodiment, when the at least one additional gene is GCD7, most preferably the gene comprises the SNP A427G. In this embodiment, a guanine replaces an arginine at position 427 of the nucleotide sequence. Preferably, this SNP results in the amino acid substitution Serl43Gly, i.e. glycine replaces serine at position 143 of the amino acid sequence. Alternatively, in a next preferred embodiment, the gene comprises a thymine or a cytosine at position 427 of the nucleotide sequence. Least preferably, the gene comprises an adenine at position 427 of the nucleotide sequence.
[0134] PIP2 (YOR363C)
[0135] In one embodiment, when the at least one additional gene is PIP2, most preferably the gene comprises the SNP T2618C. In this embodiment, a cytosine replaces a thymine at position 2618 of the nucleotide sequence. Preferably, this SNP results in the amino acid substitution Leu873Ser, i.e. serine replaces leucine at position 873 of the amino acid sequence. Alternatively, in a next preferred embodiment, the gene comprises an adenine or a guanine at position 2618 of the nucleotide sequence. Least preferably, the gene comprises a thymine at position 2618 of the nucleotide sequence.
[0136] KTR7 (YILQ85O
[0137] In one embodiment, when the at least one additional gene is KTR7, most preferably the gene comprises the SNP G877C. In this embodiment, a cytosine replaces a guanine at position 877 of the nucleotide sequence. Preferably, this SNP results in the amino acid substitution Glu293Gln, i.e. glutamine replaces glutamic acid at position 293 of the amino acid sequence. Alternatively, in a next preferred embodiment, the gene comprises a thymine or an adenine at position 877 of the nucleotide sequence. Least preferably, the gene comprises a guanine at position 877 of the nucleotide sequence.
[0138] In another embodiment, when the at least one additional gene is KTR7, most preferably the gene comprises the SNP C437G. In this embodiment, a guanine replaces a cytosine at position 437 of the nucleotide sequence. Preferably, this SNP results in the amino acid substitution Thrl46Ser, i.e. serine replaces threonine at position 146 of the amino acid sequence. Alternatively, in a next preferred embodiment, the gene comprises a thymine or an adenine at position 437 of the nucleotide sequence. Least preferably, the gene comprises a cytosine at position 437 of the nucleotide sequence. In another embodiment, when the at least one additional gene is KTR7, most preferably the gene comprises the SNP A230G. In this embodiment, a guanine replaces an adenine at position 230 of the nucleotide sequence. Preferably, this SNP results in the amino acid substitution Lys77Arg, i.e. arginine replaces lysine at position 77 of the amino acid sequence. Alternatively, in a next preferred embodiment, the gene comprises a thymine or a cytosine at position 230 of the nucleotide sequence. Least preferably, the gene comprises an adenine at position 230 of the nucleotide sequence.
[0139] YIL089W In one embodiment, when the at least one additional gene is YIL089W, most preferably the gene comprises the SNP T469A. In this embodiment, an adenine replaces a thymine at position 469 of the nucleotide sequence. Preferably, this SNP results in the amino acid substitution Leul57Ile, i.e. isoleucine replaces leucine at position 157 of the amino acid sequence. Alternatively, in a next preferred embodiment, the gene comprises a guanine or a cytosine at position 469 of the nucleotide sequence. Least preferably, the gene comprises a thymine at position 469 of the nucleotide sequence.
[0140] PUF4 (YGL014W)
[0141] In one embodiment, when the at least one additional gene is PUF4, most preferably the gene comprises the SNP C671T. In this embodiment, a thymine replaces a cytosine at position 671 of the nucleotide sequence. Preferably, this SNP results in the amino acid substitution Ser224Leu, i.e. leucine replaces serine at position 224 of the amino acid sequence. Alternatively, in a next preferred embodiment, the gene comprises an adenine or a guanine at position 671 of the nucleotide sequence. Least preferably, the gene comprises a cytosine at position 671 of the nucleotide sequence.
[0142] In another embodiment, when the at least one additional gene is PUF4, most preferably the gene comprises the SNP A2570G. In this embodiment, a guanine replaces an adenine at position 2570 of the nucleotide sequence. Preferably, this SNP results in the amino acid substitution Lys857Arg, i.e. arginine replaces lysine at position 857 of the amino acid sequence. Alternatively, in a next preferred embodiment, the gene comprises a thymine or a cytosine at position 857 of the nucleotide sequence. Least preferably, the gene comprises an adenine at position 857 of the nucleotide sequence.
[0143] COG3 (YER157W)
[0144] In one embodiment, when the at least one additional gene is COG3, most preferably the gene comprises the SNP G593A. In this embodiment, an adenine replaces a guanine at position 593 of the nucleotide sequence. Preferably, this SNP results in the amino acid substitution Serl98Asn, i.e. asparagine replaces serine at position 198 of the amino acid sequence. Alternatively, in a next preferred embodiment, the gene comprises a thymine or a cytosine at position 593 of the nucleotide sequence. Least preferably, the gene comprises a guanine at position 593 of the nucleotide sequence.
[0145] SHO1 (YIL104C) In one embodiment, when the at least one additional gene is SHQ1, most preferably the gene comprises the SNP C889T. In this embodiment, a thymine replaces a cytosine at position 889 of the nucleotide sequence. Preferably, this SNP results in the amino acid substitution Pro297Ser, i.e. serine replaces proline at position 297 of the amino acid sequence. Alternatively, in a next preferred embodiment, the gene comprises an adenine or a guanine at position 889 of the nucleotide sequence. Least preferably, the gene comprises a cytosine at position 889 of the nucleotide sequence. REB1 (YBR049C)
[0146] In one embodiment, when the at least one additional gene is REB1, most preferably the gene comprises the SNP G1727A. In this embodiment, an adenine replaces a guanine at position 1727 of the nucleotide sequence. Preferably, this SNP results in the amino acid substitution Arg576His, i.e. histidine replaces arginine at position 576 of the amino acid sequence. Alternatively, in a next preferred embodiment, the gene comprises a thymine or a cytosine at position 1727 of the nucleotide sequence. Least preferably, the gene comprises a guanine at position 1727 of the nucleotide sequence. PRP6 (YBR055Q
[0147] In one embodiment, when the at least one additional gene is PRP6, most preferably the gene comprises the SNP A2229T. In this embodiment, a thymine replaces an adenine at position 2229 of the nucleotide sequence. Preferably, this SNP results in the amino acid substitution Lys743Asn, i.e. asparagine replaces lysine at position 743 of the amino acid sequence. Alternatively, in a next preferred embodiment, the gene comprises a guanine or a cytosine at position 2229 of the nucleotide sequence. Least preferably, the gene comprises an adenine at position 2229 of the nucleotide sequence. In another embodiment, when the at least one additional gene is PRP6, most preferably the gene comprises the SNP T916G. In this embodiment, a guanine replaces a thymine at position 916 of the nucleotide sequence. Preferably, this SNP results in the amino acid substitution Ser306Ala, i.e. alanine replaces serine at position 306 of the amino acid sequence. Alternatively, in a next preferred embodiment, the gene comprises an adenine or a cytosine at position 916 of the nucleotide sequence. Least preferably, the gene comprises a thymine at position 916 of the nucleotide sequence. In another embodiment, when the at least one additional gene is PRP6, most preferably the gene comprises the SNP C743T. In this embodiment, a thymine replaces a cytosine at position 743 of the nucleotide sequence. Preferably, this SNP results in the amino acid substitution Ala248Val, i.e. valine replaces alanine at position 248 of the amino acid sequence. Alternatively, in a next preferred embodiment, the gene comprises an adenine or a guanine at position 743 of the nucleotide sequence. Least preferably, the gene comprises a cytosine at position 743 of the nucleotide sequence.
[0148] In another embodiment, when the at least one additional gene is PRP6, most preferably the gene comprises the SNP A110G. In this embodiment, a guanine replaces an adenine at position 110 of the nucleotide sequence. Preferably, this SNP results in the amino acid substitution Asp37Gly, i.e. glycine replaces aspartic acid at position 37 of the amino acid sequence. Alternatively, in a next preferred embodiment, the gene comprises a thymine or a cytosine at position 110 of the nucleotide sequence. Least preferably, the gene comprises an adenine at position 110 of the nucleotide sequence. TRM7 (YBR061Q
[0149] In one embodiment, when the at least one additional gene is TRM7, most preferably the gene comprises the SNP C228A. In this embodiment, an adenine replaces a cytosine at position 228 of the nucleotide sequence. Preferably, this SNP results in the amino acid substitution Asp76Glu, i.e. glutamic acid replaces aspartic acid at position 76 of the amino acid sequence. Alternatively, in a next preferred embodiment, the gene comprises a thymine or a guanine at position 228 of the nucleotide sequence. Least preferably, the gene comprises a cytosine at position 228 of the nucleotide sequence. MFB1 (YDR219C)
[0150] In one embodiment, when the at least one additional gene is MFB1, most preferably the gene comprises the SNP C592A. In this embodiment, an adenine replaces a cytosine at position 592 of the nucleotide sequence. Preferably, this SNP results in the amino acid substitution Hisl98Asn, i.e. asparagine replaces histidine at position 198 of the amino acid sequence. Alternatively, in a next preferred embodiment, the gene comprises a thymine or a guanine at position 592 of the nucleotide sequence. Least preferably, the gene comprises a cytosine at position 592 of the nucleotide sequence.
[0151] CRF1 (YDR223W) In one embodiment, when the at least one additional gene is CRF1, most preferably the gene comprises the SNP T91C. In this embodiment, a cytosine replaces a thymine at position 91 of the nucleotide sequence. Preferably, this SNP results in the amino acid substitution Ser31Pro, i.e. proline replaces serine at position 31 of the amino acid sequence. Alternatively, in a next preferred embodiment, the gene comprises an adenine or a guanine at position 91 of the nucleotide sequence. Least preferably, the gene comprises a thymine at position 91 of the nucleotide sequence.
[0152] In another embodiment, when the at least one additional gene is CRF1, most preferably the gene comprises the SNP C422T. In this embodiment, a thymine replaces a cytosine at position 422 of the nucleotide sequence. Preferably, this SNP results in the amino acid substitution Serl41Leu, i.e. leucine replaces serine at position 141 of the amino acid sequence. Alternatively, in a next preferred embodiment, the gene comprises an adenine or a guanine at position 422 of the nucleotide sequence. Least preferably, the gene comprises a cytosine at position 422 of the nucleotide sequence.
[0153] In another embodiment, when the at least one additional gene is CRF1, most preferably the gene comprises the SNP C740T. In this embodiment, a thymine replaces a cytosine at position 740 of the nucleotide sequence. Preferably, this SNP results in the amino acid substitution Pro247Leu, i.e. leucine replaces proline at position 247 of the amino acid sequence. Alternatively, in a next preferred embodiment, the gene comprises an adenine or a guanine at position 740 of the nucleotide sequence. Least preferably, the gene comprises a cytosine at position 740 of the nucleotide sequence.
[0154] In another embodiment, when the at least one additional gene is CRF1, most preferably the gene comprises the SNP A1234G. In this embodiment, a guanine replaces an adenine at position 1234 of the nucleotide sequence. Preferably, this SNP results in the amino acid substitution Ile412Val, i.e. valine replaces isoleucine at position 412 of the amino acid sequence. Alternatively, in a next preferred embodiment, the gene comprises a thymine or a cytosine at position 1234 of the nucleotide sequence. Least preferably, the gene comprises an adenine at position 1234 of the nucleotide sequence.
[0155] In another embodiment, when the at least one additional gene is CRF1, most preferably the gene comprises the SNP G1396A. In this embodiment, an adenine replaces a guanine at position 1396 of the nucleotide sequence. Preferably, this SNP results in the amino acid substitution Gly466Ser, i.e. serine replaces glycine at position 466 of the amino acid sequence. Alternatively, in a next preferred embodiment, the gene comprises a thymine or a cytosine at position 1396 of the nucleotide sequence. Least preferably, the gene comprises a guanine at position 1396 of the nucleotide sequence.
[0156] SPT2 (YER161O
[0157] In one embodiment, when the at least one additional gene is SPT2, most preferably the gene comprises the SNP A419T. In this embodiment, a thymine replaces an adenine at position 419 of the nucleotide sequence. Preferably, this SNP results in the amino acid substitution Glnl40Leu, i.e. leucine replaces glutamine at position 140 of the amino acid sequence. Alternatively, in a next preferred embodiment, the gene comprises a guanine or a cytosine at position 419 of the nucleotide sequence. Least preferably, the gene comprises an adenine at position 419 of the nucleotide sequence.
[0158] UTP25 (YILQ91O
[0159] In one embodiment, when the at least one additional gene is UTP25, most preferably the gene comprises the SNP G1561A. In this embodiment, an adenine replaces a guanine at position 1561 of the nucleotide sequence. Preferably, this SNP results in the amino acid substitution Val521Ile, i.e. isoleucine replaces valine at position 521 of the amino acid sequence. Alternatively, in a next preferred embodiment, the gene comprises a thymine or a cytosine at position 1561 of the nucleotide sequence. Least preferably, the gene comprises a guanine at position 1561 of the nucleotide sequence.
[0160] In another embodiment, when the at least one additional gene is UTP25, most preferably the gene comprises the SNP G443A. In this embodiment, an adenine replaces a guanine at position 443 of the nucleotide sequence. Preferably, this SNP results in the amino acid substitution Glyl48Glu, i.e. glutamic acid replaces glycine at position 148 of the amino acid sequence. Alternatively, in a next preferred embodiment, the gene comprises a thymine or a cytosine at position 443 of the nucleotide sequence. Least preferably, the gene comprises a guanine at position 443 of the nucleotide sequence. PRK1 (YIL095W)
[0161] In one embodiment, when the at least one additional gene is PRK1, most preferably the gene comprises the SNP A7G. In this embodiment, a guanine replaces an adenine at position 7 of the nucleotide sequence. Preferably, this SNP results in the amino acid substitution Thr3Ala, i.e. alanine replaces threonine at position 3 of the amino acid sequence. Alternatively, in a next preferred embodiment, the gene comprises a thymine or a cytosine at position 7 of the nucleotide sequence. Least preferably, the gene comprises an adenine at position 7 of the nucleotide sequence.
[0162] PDR1 (YGLQ13O In one embodiment, when the at least one additional gene is PDR1, most preferably the gene comprises the SNP A280G. In this embodiment, a guanine replaces an adenine at position 280 of the nucleotide sequence. Preferably, this SNP results in the amino acid substitution Thr94Ala, i.e. alanine replaces threonine at position 94 of the amino acid sequence. Alternatively, in a next preferred embodiment, the gene comprises a thymine or a cytosine at position 280 of the nucleotide sequence. Least preferably, the gene comprises an adenine at position 280 of the nucleotide sequence.
[0163] RSO55 (YLR281O In one embodiment, when the at least one additional gene is RSO55, most preferably the gene comprises the SNP A52C. In this embodiment, a cytosine replaces an adenine at position 52 of the nucleotide sequence. Preferably, this SNP results in the amino acid substitution Ilel8Leu, i.e. leucine replaces isoleucine at position 18 of the amino acid sequence. Alternatively, in a next preferred embodiment, the gene comprises a thymine or a guanine at position 52 of the nucleotide sequence. Least preferably, the gene comprises an adenine at position 52 of the nucleotide sequence.
[0164] COO11 (YLR290C) In one embodiment, when the at least one additional gene is COQ11, most preferably the gene comprises the SNP C82A. In this embodiment, an adenine replaces a cytosine at position 82 of the nucleotide sequence. Preferably, this SNP results in the amino acid substitution Gln28Lys, i.e. lysine replaces glutamine at position 28 of the amino acid sequence. Alternatively, in a next preferred embodiment, the gene comprises a thymine or a guanine at position 82 of the nucleotide sequence. Least preferably, the gene comprises a cytosine at position 82 of the nucleotide sequence.
[0165] YBR053C
[0166] In one embodiment, when the at least one additional gene is YBR053C, most preferably the gene comprises the SNP G473A. In this embodiment, an adenine replaces a guanine at position 473 of the nucleotide sequence. Preferably, this SNP results in the amino acid substitution Glyl58Glu, i.e. glutamic acid replaces glycine at position 158 of the amino acid sequence. Alternatively, in a next preferred embodiment, the gene comprises a thymine or a cytosine at position 473 of the nucleotide sequence. Least preferably, the gene comprises a guanine at position 473 of the nucleotide sequence.
[0167] In another embodiment, when the at least one additional gene is YBR053C, most preferably the gene comprises the SNP G218C. In this embodiment, a cytosine replaces a guanine at position 218 of the nucleotide sequence. Preferably, this SNP results in the amino acid substitution Gly73Ala, i.e. alanine replaces glycine at position 73 of the amino acid sequence. Alternatively, in a next preferred embodiment, the gene comprises a thymine or an adenine at position 218 of the nucleotide sequence. Least preferably, the gene comprises a guanine at position 218 of the nucleotide sequence.
[0168] CNM1 (YBR063C)
[0169] In one embodiment, when the at least one additional gene is CNM1, most preferably the gene comprises the SNP G396T. In this embodiment, a thymine replaces a guanine at position 396 of the nucleotide sequence. Preferably, this SNP results in the amino acid substitution Argl32Ser, i.e. serine replaces arginine at position 132 of the amino acid sequence. Alternatively, in a next preferred embodiment, the gene comprises an adenine or a cytosine at position 396 of the nucleotide sequence. Least preferably, the gene comprises a guanine at position 396 of the nucleotide sequence.
[0170] In another embodiment, when the at least one additional gene is CNM1, most preferably the gene comprises the SNP G232A. In this embodiment, an adenine replaces a guanine at position 232 of the nucleotide sequence. Preferably, this SNP results in the amino acid substitution Ala78Thr, i.e. threonine replaces alanine at position 78 of the amino acid sequence. Alternatively, in a next preferred embodiment, the gene comprises a thymine or a cytosine at position 232 of the nucleotide sequence. Least preferably, the gene comprises a guanine at position 232 of the nucleotide sequence.
[0171] ECM2 (YBR065Q
[0172] In one embodiment, when the at least one additional gene is ECM2, most preferably the gene comprises the SNP T984A. In this embodiment, an adenine replaces a thymine at position 984 of the nucleotide sequence. Preferably, this SNP results in the amino acid substitution Phe328Leu, i.e. leucine replaces phenylalanine at position 328 of the amino acid sequence. Alternatively, in a next preferred embodiment, the gene comprises a guanine or a cytosine at position 984 of the nucleotide sequence. Least preferably, the gene comprises a thymine at position 984 of the nucleotide sequence.
[0173] In another embodiment, when the at least one additional gene is ECM2, most preferably the gene comprises the SNP G295A. In this embodiment, an adenine replaces a guanine at position 295 of the nucleotide sequence. Preferably, this SNP results in the amino acid substitution Val99Ile, i.e. isoleucine replaces valine at position 99 of the amino acid sequence. Alternatively, in a next preferred embodiment, the gene comprises a thymine or a cytosine at position 295 of the nucleotide sequence. Least preferably, the gene comprises a guanine at position 295 of the nucleotide sequence.
[0174] ADRI (YDR216W)
[0175] In one embodiment, when the at least one additional gene is ADRI, most preferably the gene comprises the SNP T935C. In this embodiment, a cytosine replaces a thymine at position 935 of the nucleotide sequence. Preferably, this SNP results in the amino acid substitution Leu312Ser, i.e. serine replaces leucine at position 312 of the amino acid sequence. Alternatively, in a next preferred embodiment, the gene comprises an adenine or a guanine at position 935 of the nucleotide sequence. Least preferably, the gene comprises a thymine at position 935 of the nucleotide sequence. In another embodiment, when the at least one additional gene is ADRI, most preferably the gene comprises the SNP C959T. In this embodiment, a thymine replaces a cytosine at position 959 of the nucleotide sequence. Preferably, this SNP results in the amino acid substitution Thr320Ile, i.e. isoleucine replaces threonine at position 320 of the amino acid sequence. Alternatively, in a next preferred embodiment, the gene comprises an adenine or a guanine at position 959 of the nucleotide sequence. Least preferably, the gene comprises a cytosine at position 959 of the nucleotide sequence. In another embodiment, when the at least one additional gene is ADRI, most preferably the gene comprises the SNP C1498A. In this embodiment, an adenine replaces a cytosine at position 1498 of the nucleotide sequence. Preferably, this SNP results in the amino acid substitution Leu500Ile, i.e. isoleucine replaces leucine at position 500 of the amino acid sequence. Alternatively, in a next preferred embodiment, the gene comprises a thymine or a guanine at position 1498 of the nucleotide sequence. Least preferably, the gene comprises a cytosine at position 1498 of the nucleotide sequence.
[0176] In another embodiment, when the at least one additional gene is ADRI, most preferably the gene comprises the SNP C1691A. In this embodiment, an adenine replaces a cytosine at position 1691 of the nucleotide sequence. Preferably, this SNP results in the amino acid substitution Thr564Lys, i.e. lysine replaces threonine at position 564 of the amino acid sequence. Alternatively, in a next preferred embodiment, the gene comprises a thymine or a guanine at position 1691 of the nucleotide sequence. Least preferably, the gene comprises a cytosine at position 1691 of the nucleotide sequence.
[0177] In another embodiment, when the at least one additional gene is ADRI, most preferably the gene comprises the SNP A2282T. In this embodiment, a thymine replaces an adenine at position 2282 of the nucleotide sequence. Preferably, this SNP results in the amino acid substitution Tyr761Phe, i.e. phenylalanine replaces tyrosine at position 761 of the amino acid sequence. Alternatively, in a next preferred embodiment, the gene comprises a guanine or a cytosine at position 2282 of the nucleotide sequence. Least preferably, the gene comprises an adenine at position 2282 of the nucleotide sequence. In another embodiment, when the at least one additional gene is ADRI, most preferably the gene comprises the SNP G3157A. In this embodiment, an adenine replaces a guanine at position 3157 of the nucleotide sequence. Preferably, this SNP results in the amino acid substitution GlulO53Lys, i.e. lysine replaces glutamic acid at position 1053 of the amino acid sequence. Alternatively, in a next preferred embodiment, the gene comprises a thymine or a cytosine at position 427 of the nucleotide sequence. Least preferably, the gene comprises an adenine at position 427 of the nucleotide sequence. BEM2 (YER155O
[0178] In one embodiment, when the at least one additional gene is BEM2, most preferably the gene comprises the SNP G5937T. In this embodiment, a thymine replaces a guanine at position 5937 of the nucleotide sequence. Preferably, this SNP results in the amino acid substitution Leul979Phe, i.e. phenylalanine replaces leucine at position 1979 of the amino acid sequence. Alternatively, in a next preferred embodiment, the gene comprises an adenine or a cytosine at position 5937 of the nucleotide sequence. Least preferably, the gene comprises a guanine at position 5937 of the nucleotide sequence. In another embodiment, when the at least one additional gene is BEM2, most preferably the gene comprises the SNP G2584A. In this embodiment, an adenine replaces a guanine at position 2584 of the nucleotide sequence. Preferably, this SNP results in the amino acid substitution Asp862Asn, i.e. asparagine replaces aspartic acid at position 862 of the amino acid sequence. Alternatively, in a next preferred embodiment, the gene comprises a thymine or a cytosine at position 2584 of the nucleotide sequence. Least preferably, the gene comprises a guanine at position 2584 of the nucleotide sequence.
[0179] In another embodiment, when the at least one additional gene is BEM2, most preferably the gene comprises the SNP C926T. In this embodiment, a thymine replaces a cytosine at position 926 of the nucleotide sequence. Preferably, this SNP results in the amino acid substitution Ala309Val, i.e. valine replaces alanine at position 309 of the amino acid sequence. Alternatively, in a next preferred embodiment, the gene comprises an adenine or a guanine at position 926 of the nucleotide sequence. Least preferably, the gene comprises a cytosine at position 926 of the nucleotide sequence. In another embodiment, when the at least one additional gene is BEM2, most preferably the gene comprises the SNP C533T. In this embodiment, a thymine replaces a cytosine at position 533 of the nucleotide sequence. Preferably, this SNP results in the amino acid substitution Thrl78Ile, i.e. isoleucine replaces threonine at position 178 of the amino acid sequence. Alternatively, in a next preferred embodiment, the gene comprises an adenine or a guanine at position 533 of the nucleotide sequence. Least preferably, the gene comprises a cytosine at position 533 of the nucleotide sequence. MYG1 (YER156O
[0180] In one embodiment, when the at least one additional gene is MYG1, most preferably the gene comprises the SNP A439G. In this embodiment, a guanine replaces an adenine at position 439 of the nucleotide sequence. Preferably, this SNP results in the amino acid substitution Thrl47Ala, i.e. alanine replaces threonine at position 147 of the amino acid sequence. Alternatively, in a next preferred embodiment, the gene comprises a thymine or a cytosine at position 439 of the nucleotide sequence. Least preferably, the gene comprises an adenine at position 439 of the nucleotide sequence. YHR131C
[0181] In one embodiment, when the at least one additional gene is YHR131C, most preferably the gene comprises the SNP G1924A. In this embodiment, an adenine replaces a guanine at position 1924 of the nucleotide sequence. Preferably, this SNP results in the amino acid substitution Gly642Ser, i.e. serine replaces glycine at position 642 of the amino acid sequence. Alternatively, in a next preferred embodiment, the gene comprises a thymine or a cytosine at position 1924 of the nucleotide sequence. Least preferably, the gene comprises a guanine at position 1924 of the nucleotide sequence. In another embodiment, when the at least one additional gene is YHR131C, most preferably the gene comprises the SNP C1633G. In this embodiment, a guanine replaces a cytosine at position 1633 of the nucleotide sequence. Preferably, this SNP results in the amino acid substitution Pro545Ala, i.e. alanine replaces proline at position 545 of the amino acid sequence. Alternatively, in a next preferred embodiment, the gene comprises a thymine or an adenine at position 1633 of the nucleotide sequence. Least preferably, the gene comprises a cytosine at position 1633 of the nucleotide sequence. In another embodiment, when the at least one additional gene is YHR131C, most preferably the gene comprises the SNP A328G. In this embodiment, a guanine replaces an adenine at position 328 of the nucleotide sequence. Preferably, this SNP results in the amino acid substitution IlellOVal, i.e. valine replaces isoleucine at position 110 of the amino acid sequence. Alternatively, in a next preferred embodiment, the gene comprises a thymine or a cytosine at position 328 of the nucleotide sequence. Least preferably, the gene comprises an adenine at position 328 of the nucleotide sequence.
[0182] ARO9 (YHR137W)
[0183] In one embodiment, when the at least one additional gene is ARO9, most preferably the gene comprises the SNP C1121G. In this embodiment, a guanine replaces a cytosine at position 1121 of the nucleotide sequence. Preferably, this SNP results in the amino acid substitution Ser374Cys, i.e. cysteine replaces serine at position 374 of the amino acid sequence. Alternatively, in a next preferred embodiment, the gene comprises a thymine or an adenine at position 1121 of the nucleotide sequence. Least preferably, the gene comprises a cytosine at position 1121 of the nucleotide sequence.
[0184] DSE2 (YHR143W)
[0185] In one embodiment, when the at least one additional gene is DSE2, most preferably the gene comprises the SNP A71G. In this embodiment, a guanine replaces an adenine at position 71 of the nucleotide sequence. Preferably, this SNP results in the amino acid substitution Asn24Ser, i.e. serine replaces asparagine at position 24 of the amino acid sequence. Alternatively, in a next preferred embodiment, the gene comprises a thymine or a cytosine at position 71 of the nucleotide sequence. Least preferably, the gene comprises an adenine at position 71 of the nucleotide sequence.
[0186] CRP1 (YHR146W)
[0187] In one embodiment, when the at least one additional gene is CRP1, most preferably the gene comprises the SNP G859A. In this embodiment, an adenine replaces a guanine at position 859 of the nucleotide sequence. Preferably, this SNP results in the amino acid substitution Val287Met, i.e. methionine replaces valine at position 287 of the amino acid sequence. Alternatively, in a next preferred embodiment, the gene comprises a thymine or a cytosine at position 859 of the nucleotide sequence. Least preferably, the gene comprises a guanine at position 859 of the nucleotide sequence.
[0188] In another embodiment, when the at least one additional gene is CRP1, most preferably the gene comprises the SNP T1043C. In this embodiment, a cytosine replaces a thymine at position 1043 of the nucleotide sequence. Preferably, this SNP results in the amino acid substitution Leu348Ser, i.e. serine replacing leucine at position 348 of the amino acid sequence. Alternatively, in a next preferred embodiment, the gene comprises an adenine or a guanine at position 1043 of the nucleotide sequence. Least preferably, the gene comprises a thymine at position 1043 of the nucleotide sequence.
[0189] RCI37 (YIL077C)
[0190] In one embodiment, when the at least one additional gene is RCI37, most preferably the gene comprises the SNP C697T. In this embodiment, a thymine replaces a cytosine at position 697 of the nucleotide sequence. Preferably, this SNP results in the amino acid substitution Pro233Ser, i.e. serine replacing proline at position 233 of the amino acid sequence. Alternatively, in a next preferred embodiment, the gene comprises an adenine or a guanine at position 697 of the nucleotide sequence. Least preferably, the gene comprises a cytosine at position 697 of the nucleotide sequence.
[0191] AIR1 (YILQ79O
[0192] In one embodiment, when the at least one additional gene is AIR1, most preferably the gene comprises the SNP G487A. In this embodiment, an adenine replaces a guanine at position 487 of the nucleotide sequence. Preferably, this SNP results in the amino acid substitution Alal63Thr, i.e. threonine replacing alanine at position 163 of the amino acid sequence. Alternatively, in a next preferred embodiment, the gene comprises a thymine or a cytosine at position 487 of the nucleotide sequence. Least preferably, the gene comprises a guanine at position 487 of the nucleotide sequence.
[0193] In another embodiment, when the at least one additional gene is AIR1, most preferably the gene comprises the SNP A827G. In this embodiment, a guanine replaces an adenine at position 827 of the nucleotide sequence. Preferably, this SNP results in the amino acid substitution Lys276Arg, i.e. an arginine replacing a lysine at position 276. Alternatively, in a next preferred embodiment, the gene comprises a thymine or a cytosine at position 827 of the nucleotide sequence. Least preferably, the gene comprises an adenine at position 827 of the nucleotide sequence. SDS3 (YILQ84O
[0194] In one embodiment, when the at least one additional gene is SDS3, most preferably the gene comprises the SNP T627A. In this embodiment, an adenine replaces a thymine at position 627 of the nucleotide sequence. Preferably, this SNP results in the amino acid substitution Asp209Glu, i.e. glutamic acid replacing asparagine at position 209 of the amino acid sequence. Alternatively, in a next preferred embodiment, the gene comprises a guanine or a cytosine at position 627 of the nucleotide sequence. Least preferably, the gene comprises a thymine at position 627 of the nucleotide sequence. YIL092W
[0195] In one embodiment, when the at least one additional gene is YIL092W, most preferably the gene comprises the SNP G130A. In this embodiment, an adenine replaces a guanine at position 130 of the nucleotide sequence. Preferably, this SNP results in the amino acid substitution Glu44Lys, i.e. lysine replacing glutamic acid at position 44 of the amino acid sequence. Alternatively, in a next preferred embodiment, the gene comprises a thymine or a cytosine at position 130 of the nucleotide sequence. Least preferably, the gene comprises a guanine at position 130 of the nucleotide sequence. In another embodiment, when the at least one additional gene is YIL092W, most preferably the gene comprises the SNP T137C. In this embodiment, a cytosine replaces a thymine at position 137 of the nucleotide sequence. Preferably, this SNP results in the amino acid substitution Leu46Ser, i.e. serine replacing leucine at position 46 of the amino acid sequence. Alternatively, in a next preferred embodiment, the gene comprises an adenine or a guanine at position 137 of the nucleotide sequence. Least preferably, the gene comprises a thymine at position 137 of the nucleotide sequence.
[0196] In another embodiment, when the at least one additional gene is YIL092W, most preferably the gene comprises the SNP A257G. In this embodiment, a guanine replaces an adenine at position 257 of the nucleotide sequence. Preferably, this SNP results in the amino acid substitution Asn86Ser, i.e. serine replacing asparagine at position 86 of the amino acid sequence. Alternatively, in a next preferred embodiment, the gene comprises a thymine or a cytosine at position 257 of the nucleotide sequence. Least preferably, the gene comprises an adenine at position 257 of the nucleotide sequence.
[0197] In another embodiment, when the at least one additional gene is YIL092W, most preferably the gene comprises the SNP A778G. In this embodiment, a guanine replaces an adenine at position 778 of the nucleotide sequence. Preferably, this SNP results in the amino acid substitution Ser260Gly, i.e. glycine replacing serine at position 260 of the amino acid sequence. Alternatively, in a next preferred embodiment, the gene comprises a thymine or a cytosine at position 778 of the nucleotide sequence. Least preferably, the gene comprises an adenine at position 778 of the nucleotide sequence. In another embodiment, when the at least one additional gene is YIL092W, most preferably the gene comprises the SNP G955A. In this embodiment, an adenine replaces a guanine at position 955 of the nucleotide sequence. Preferably, this SNP results in the amino acid substitution Glu319Lys, i.e. lysine replacing glutamic acid at position 319 of the amino acid sequence. Alternatively, in a next preferred embodiment, the gene comprises a thymine or a cytosine at position 955 of the nucleotide sequence. Least preferably, the gene comprises a guanine at position 955 of the nucleotide sequence.
[0198] In another embodiment, when the at least one additional gene is YIL092W, most preferably the gene comprises the SNP G1055A. In this embodiment, an adenine replaces a guanine at position 1055 of the nucleotide sequence. Preferably, this SNP results in the amino acid substitution Ser352Asn, i.e. asparagine replacing serine at position 352 of the amino acid sequence. Alternatively, in a next preferred embodiment, the gene comprises a thymine or a cytosine at position 1055 of the nucleotide sequence. Least preferably, the gene comprises a guanine at position 1055 of the nucleotide sequence.
[0199] In another embodiment, when the at least one additional gene is YIL092W, most preferably the gene comprises the SNP A1086G. In this embodiment, a guanine replaces an adenine at position 1086 of the nucleotide sequence. Preferably, this SNP results in the amino acid substitution Ile362Met, i.e. methionine replacing isoleucine at position 362 of the amino acid sequence. Alternatively, in a next preferred embodiment, the gene comprises a thymine or a cytosine at position 1086 of the nucleotide sequence. Least preferably, the gene comprises an adenine at position 1086 of the nucleotide sequence. In another embodiment, when the at least one additional gene is YIL092W, most preferably the gene comprises the SNP C1661T. In this embodiment, a thymine replaces a cytosine at position 1661 of the nucleotide sequence. Preferably, this SNP results in the amino acid substitution Ala554Val, i.e. valine replacing alanine at position 554 of the nucleotide sequence. Alternatively, in a next preferred embodiment, the gene comprises an adenine or a guanine at position 1661 of the nucleotide sequence. Least preferably, the gene comprises a cytosine at position 1661 of the nucleotide sequence.
[0200] In another embodiment, when the at least one additional gene is YIL092W, most preferably the gene comprises the SNP C1730G. In this embodiment, a guanine replaces a cytosine at position 1730 of the nucleotide sequence. Preferably, this SNP results in the amino acid substitution Pro577Arg, i.e. arginine replacing proline at position 577 of the amino acid sequence. Alternatively, in a next preferred embodiment, the gene comprises a thymine or an adenine at position 1730 of the nucleotide sequence. Least preferably, the gene comprises a cytosine at position 1730 of the nucleotide sequence.
[0201] In another embodiment, when the at least one additional gene is YIL092W, most preferably the gene comprises the SNP G1802T. In this embodiment, a thymine replaces a guanine at position 1802 of the nucleotide sequence. Preferably, this SNP results in the amino acid substitution Gly601Val, i.e. valine replacing glycine at position 601 of the amino acid sequence. Alternatively, in a next preferred embodiment, the gene comprises an adenine or a cytosine at position 1802 of the nucleotide sequence. Least preferably, the gene comprises a guanine at position 1802 of the nucleotide sequence.
[0202] BMT5 (YIL096Q
[0203] In one embodiment, when the at least one additional gene is BMT5, most preferably the gene comprises the SNP A64C. In this embodiment, a cytosine replaces an adenine at position 64 of the nucleotide sequence. Preferably, this SNP results in the amino acid substitution Lys22Gln, i.e. glutamine replacing lysine at position 22 of the amino acid sequence. Alternatively, in a next preferred embodiment, the gene comprises a thymine or a guanine at position 64 of the nucleotide sequence.
[0204] Least preferably, the gene comprises an adenine at position 64 of the nucleotide sequence. FMC1 (YILQ98O
[0205] In one embodiment, when the at least one additional gene is FMC1, most preferably the gene comprises the SNP A112G. In this embodiment, a guanine replaces an adenine at position 112 of the nucleotide sequence. Preferably, this SNP results in the amino acid substitution Thr38Ala, i.e. alanine replacing threonine at position 38 of the amino acid sequence. Alternatively, in a next preferred embodiment, the gene comprises a thymine or a cytosine at position 112 of the nucleotide sequence. Least preferably, the gene comprises an adenine at position 112 of the nucleotide sequence. CBP2 (YHL038Q
[0206] In one embodiment, when the at least one additional gene is CBP2, most preferably the gene comprises the SNP A1094C. In this embodiment, a cytosine replaces an adenine at position 1094 of the nucleotide sequence. Preferably, this SNP results in the amino acid substitution Gln365Pro, i.e. proline replacing glutamine at position 365 of the amino acid sequence. Alternatively, in a next preferred embodiment, the gene comprises a thymine or a guanine at position 1094 of the nucleotide sequence. Least preferably, the gene comprises an adenine at position 1094 of the nucleotide sequence. In another embodiment, when the at least one additional gene is CBP2, most preferably the gene comprises the SNP G372T. In this embodiment, a thymine replaces a guanine at position 372 of the nucleotide sequence. Preferably, this SNP results in the amino acid substitution Glul24Asp, i.e. asparagine replacing glutamine at position 124 of the amino acid sequence. Alternatively, in a next preferred embodiment, the gene comprises an adenine or a cytosine at position 372 of the nucleotide sequence. Least preferably, the gene comprises a guanine at position 372 of the nucleotide sequence.
[0207] YER152C In one embodiment, when the at least one additional gene is YER152C, most preferably the gene comprises the SNP G1030A. In this embodiment, an adenine replaces a guanine at position 1030 of the nucleotide sequence. Preferably, this SNP results in the amino acid substitution Ala344Thr, i.e. threonine replaces alanine at position 344 of the amino acid sequence. Alternatively, in a next preferred embodiment, the gene comprises a thymine or a cytosine at position 1030 of the nucleotide sequence. Least preferably, the gene comprises a guanine at position 1030 of the nucleotide sequence.
[0208] In another embodiment, the at least one additional gene may comprise any one of the SNPs (nucleotide substitutions) referred to in Table 2. Accordingly, the at least one additional gene may comprise any one of the amino acid substitutions referred to in Table 2.
[0209] When performing the QTL analysis, the inventors also discovered that particular SNPs in certain genes are associated with poor recombinant protein production. Accordingly, in one embodiment, the recombinant or engineered eukaryotic strain does not comprise a combination of SNPs associated with poor recombinant protein production.
[0210] SLM1 (YIL105C)
[0211] In one embodiment, when the at least one additional gene is SLM1, most preferably the gene does not comprise the SNP A1873G (resulting in the amino acid substitution Met625Val). Accordingly, in this embodiment, adenine is present at position 1873 of the nucleotide sequence. Preferably, therefore, methionine is present at position 625 of the amino acid sequence. Alternatively, in a next preferred embodiment, the gene comprises a thymine or a cytosine at position 1873 of the nucleotide sequence. Least preferably, the gene comprises a guanine at position 1873 of the nucleotide sequence.
[0212] In another embodiment, when the at least one additional gene is SLM1, most preferably the gene does not comprise the SNP T1295A (resulting in the amino acid substitution Phe432Tyr). Accordingly, in this embodiment, thymine is present at position 1295 of the nucleotide sequence. Preferably, therefore, phenylalanine is present at position 432 of the amino acid sequence. Alternatively, in a next preferred embodiment, the gene comprises a guanine or a cytosine at position 1295 of the nucleotide sequence. Least preferably, the gene comprises an adenine at position 1295 of the nucleotide sequence. In another embodiment, when the at least one additional gene is SLM1, most preferably the gene does not comprise the SNP A1103G (resulting in the amino acid substitution Gln368Arg). Accordingly, in this embodiment, adenine is present at position 1103 of the nucleotide sequence. Preferably, therefore, glutamine is present at position 368 of the amino acid sequence. Alternatively, in a next preferred embodiment, the gene comprises a thymine or a cytosine at position 1103 of the nucleotide sequence. Least preferably, the gene comprises a guanine at position 1103 of the nucleotide sequence. In another embodiment, when the at least one additional gene is SLM1, most preferably the gene does not comprise the SNP G524A (resulting in the amino acid substitution Argl75Lys). Accordingly, in this embodiment, guanine is present at position 524 of the nucleotide sequence. Preferably, therefore, arginine is present at position 175 of the amino acid sequence. Alternatively, in a next preferred embodiment, the gene comprises a thymine or a cytosine at position 524 of the nucleotide sequence. Least preferably, the gene comprises an adenine at position 524 of the nucleotide sequence.
[0213] In another embodiment, when the at least one additional gene is SLM1, most preferably the gene does not comprise the SNP G417T (resulting in the amino acid substitution Glnl39His). Accordingly, in this embodiment, guanine is present at position 417 of the nucleotide sequence. Preferably, therefore, glutamine is present at position 139 of the amino acid sequence. Alternatively, in a next preferred embodiment, the gene comprises an adenine or a cytosine at position 417 of the nucleotide sequence. Least preferably, the gene comprises a thymine at position 417 of the nucleotide sequence.
[0214] In another embodiment, when the at least one additional gene is SLM1, most preferably the gene does not comprise the SNP T386C (resulting in the amino acid substitution Vall29Ala). Accordingly, in this embodiment, thymine is present at position 386 of the nucleotide sequence. Preferably, therefore, valine is present at position 129 of the amino acid sequence. Alternatively, in a next preferred embodiment, the gene comprises an adenine or a guanine at position 386 of the nucleotide sequence. Least preferably, the gene comprises a cytosine at position 386 of the nucleotide sequence. In another embodiment, when the at least one additional gene is SLM1, most preferably the gene does not comprise the SNP G197C (resulting in the amino acid substitution Ser66Thr). Accordingly, in this embodiment, guanine is present at position 197 of the nucleotide sequence. Preferably, therefore, serine is present at position 66 of the amino acid sequence. Alternatively, in a next preferred embodiment, the gene comprises a thymine or an adenine at position 197 of the nucleotide sequence. Least preferably, the gene comprises a cytosine at position 197 of the nucleotide sequence. In another embodiment, when the at least one additional gene is SLM1, most preferably the gene does not comprise the SNP G98A (resulting in the amino acid substitution Arg33Lys). Accordingly, in this embodiment, guanine is present at position 98 of the nucleotide sequence. Preferably, therefore, arginine is present at position 33 of the amino acid sequence. Alternatively, in a next preferred embodiment, the gene comprises a thymine or a cytosine at position 98 of the nucleotide sequence. Least preferably, the gene comprises an adenine at position 98 of the nucleotide sequence.
[0215] BUD27 (YFL023W) In one embodiment, when the at least one gene is BUD27, most preferably the gene does not comprise the SNP G941T (resulting in the amino acid substitution Gly314Val). Accordingly, in this embodiment, guanine is present at position 941 of the nucleotide sequence. Preferably, therefore, glycine is present at position 314 of the amino acid sequence. Alternatively, in a next preferred embodiment, the gene comprises an adenine or a cytosine at position 941 of the nucleotide sequence.
[0216] Least preferably, the gene comprises a thymine at position 941 of the nucleotide sequence.
[0217] ZTA1 (YBR046C) In one embodiment, when the at least one additional gene is ZTA1, most preferably the gene does not comprise the SNP G259A (resulting in the amino acid substitution Val87Ile). Accordingly, in this embodiment, guanine is present at position 259 of the nucleotide sequence. Preferably, therefore, valine is present at position 87 of the amino acid sequence. Alternatively, in a next preferred embodiment, the gene comprises a thymine or a cytosine at position 259 of the nucleotide sequence. Least preferably, the gene comprises an adenine at position 259 of the nucleotide sequence. C0G3 (YER157W)
[0218] In one embodiment, when the at least one additional gene is COG3, most preferably the gene does not comprise the SNP A148G (resulting in the amino acid substitution Ser50Gly). Accordingly, in this embodiment, adenine is present at position 148 of the nucleotide sequence. Preferably, therefore, serine is present at position 50 of the amino acid sequence. Alternatively, in a next preferred embodiment, the gene comprises a thymine or a cytosine at position 148 of the nucleotide sequence. Least preferably, the gene comprises a guanine at position 148 of the nucleotide sequence.
[0219] SHO1 (YIL104C)
[0220] In one embodiment, when the at least one additional gene is SHQ1, most preferably the gene does not comprise the SNP A1360G (resulting in the amino acid substitution Met454Val). Accordingly, in this embodiment, adenine is present at position 1360 of the nucleotide sequence. Preferably, therefore, methionine is present at position 454 of the amino acid sequence. Alternatively, in a next preferred embodiment, the gene comprises a thymine or a cytosine at position 1360 of the nucleotide sequence. Least preferably, the gene comprises a guanine at position 1360 of the nucleotide sequence.
[0221] In another embodiment, when the at least one additional gene is SHQ1, most preferably the gene does not comprise the SNP A352G (resulting in the amino acid substitution Thrll8Ala). Accordingly, in this embodiment, adenine is present at position 352 of the nucleotide sequence. Preferably, therefore, threonine is present at position 118 of the amino acid sequence. Alternatively, in a next preferred embodiment, the gene comprises a thymine or a cytosine at position 352 of the nucleotide sequence. Least preferably, the gene comprises a guanine at position 352 of the nucleotide sequence.
[0222] CHS7 (YHR142W)
[0223] In one embodiment, when the at least one additional gene is CHS7, most preferably the gene does not comprise the SNP T536C (resulting in the amino acid substitution Phel79Ser). Accordingly, in this embodiment, thymine is present at position 536 of the nucleotide sequence. Preferably, therefore, phenylalanine is present at position 179 of the amino acid sequence. Alternatively, in a next preferred embodiment, the gene comprises a guanine or an adenine at position 536 of the nucleotide sequence. Least preferably, the gene comprises a cytosine at position 536 of the nucleotide sequence.
[0224] PRK1 (YIL095W) In one embodiment, when the at least one additional gene is PRK1, most preferably the gene does not comprise the SNP T649A (resulting in the amino acid substitution Tyr217Asn). Accordingly, in this embodiment, thymine is present at position 649 of the nucleotide sequence. Preferably, therefore, tyrosine is present at position 217 of the amino acid sequence. Alternatively, in a next preferred embodiment, the gene comprises a guanine or a cytosine at position 649 of the nucleotide sequence. Least preferably, the gene comprises an adenine at position 649 of the nucleotide sequence.
[0225] In another embodiment, when the at least one additional gene is PRK1, most preferably the gene does not comprise the SNP C989A (resulting in the amino acid substitution Thr330Asn). Accordingly, in this embodiment, cytosine is present at position 989 of the nucleotide sequence. Preferably, therefore, threonine is present at position 330 of the amino acid sequence. Alternatively, in a next preferred embodiment, the gene comprises a thymine or a guanine at position 989 of the nucleotide sequence. Least preferably, the gene comprises an adenine at position 989 of the nucleotide sequence.
[0226] In another embodiment, when the at least one additional gene is PRK1, most preferably the gene does not comprise the SNP T1077G (resulting in the amino acid substitution Ser359Arg). Accordingly, in this embodiment, thymine is present at position 1077 of the nucleotide sequence. Preferably, therefore, serine is present at position 359 of the amino acid sequence. Alternatively, in a next preferred embodiment, the gene comprises an adenine or a cytosine at position 1077 of the nucleotide sequence. Least preferably, the gene comprises a guanine at position 1077 of the nucleotide sequence.
[0227] In another embodiment, when the at least one additional gene is PRK1, most preferably the gene does not comprise the SNP T1216A (resulting in the amino acid substitution Ser406Thr). Accordingly, in this embodiment, thymine is present at position 1216 of the nucleotide sequence. Preferably, therefore, serine is present at position 406 of the amino acid sequence. Alternatively, in a next preferred embodiment, the gene comprises a guanine or a cytosine at position 1216 of the nucleotide sequence. Least preferably, the gene comprises an adenine at position 1216 of the nucleotide sequence.
[0228] In another embodiment, when the at least one additional gene is PRK1, most preferably the gene does not comprise the SNP G2117A (resulting in the amino acid substitution Gly706Asp). Accordingly, in this embodiment, guanine is present at position 2117 of the nucleotide sequence. Preferably, therefore, glycine is present at position 706 of the amino acid sequence. Alternatively, in a next preferred embodiment, the gene comprises a thymine or a cytosine at position 2117 of the nucleotide sequence. Least preferably, the gene comprises an adenine at position 2117 of the nucleotide sequence.
[0229] XBP1 (YIL101C)
[0230] In one embodiment, when the at least one additional gene is XBP1, most preferably the gene does not comprise the SNP T1883C (resulting in the amino acid substitution Phe628Ser). Accordingly, in this embodiment, thymine is present at position 1883 of the nucleotide sequence. Preferably, therefore, phenylalanine is present at position 628 of the amino acid sequence. Alternatively, in a next preferred embodiment, the gene comprises an adenine or a guanine at position 1883 of the nucleotide sequence. Least preferably, the gene comprises a cytosine at position 1883 of the nucleotide sequence.
[0231] In another embodiment, when the at least one additional gene is XBP1, most preferably the gene does not comprise the SNP C1877G (resulting in the amino acid substitution Thr626Ser). Accordingly, in this embodiment, cytosine is present at position 1877 of the nucleotide sequence. Preferably, therefore, threonine is present at position 626 of the amino acid sequence. Alternatively, in a next preferred embodiment, the gene comprises a thymine or an adenine at position 1877 of the nucleotide sequence. Least preferably, the gene comprises a guanine at position 1877 of the nucleotide sequence.
[0232] In another embodiment, when the at least one additional gene is XBP1, most preferably the gene does not comprise the SNP A1466G (resulting in the amino acid substitution Lys489Arg). Accordingly, in this embodiment, adenine is present at position 1466 of the nucleotide sequence. Preferably, therefore, lysine is present at position 489 of the amino acid sequence. Alternatively, in a next preferred embodiment, the gene comprises a thymine or a cytosine at position 1466 of the nucleotide sequence. Least preferably, the gene comprises a guanine at position 1466 of the nucleotide sequence.
[0233] In another embodiment, when the at least one additional gene is XBP1, most preferably the gene does not comprise the SNP C872T (resulting in the amino acid substitution Ser291Leu). Accordingly, in this embodiment, cytosine is present at position 872 of the nucleotide sequence. Preferably, therefore, serine is present at position 291 of the amino acid sequence. Alternatively, in a next preferred embodiment, the gene comprises an adenine or a guanine at position 872 of the nucleotide sequence. Least preferably, the gene comprises a thymine at position 872 of the nucleotide sequence.
[0234] In another embodiment, when the at least one additional gene is XBP1, most preferably the gene does not comprise the SNP G839A (resulting in the amino acid substitution Gly280Glu). Accordingly, in this embodiment, guanine is present at position 839 of the nucleotide sequence. Preferably, therefore, glycine is present at position 280 of the amino acid sequence. Alternatively, in a next preferred embodiment, the gene comprises a thymine or a cytosine at position 839 of the nucleotide sequence. Least preferably, the gene comprises an adenine at position 839 of the nucleotide sequence.
[0235] DPH1 (YIL103W)
[0236] In one embodiment, when the at least one additional gene is DPH1, most preferably the gene does not comprise the SNP A56G (resulting in the amino acid substitution Lysl9Arg). Accordingly, in this embodiment, adenine is present at position 56 of the nucleotide sequence. Preferably, therefore, lysine is present at position 19 of the amino acid sequence. Alternatively, in a next preferred embodiment, the gene comprises a thymine or a cytosine at position 56 of the nucleotide sequence. Least preferably, the gene comprises a guanine at position 56 of the nucleotide sequence.
[0237] In another embodiment, when the at least one additional gene is DPH1, most preferably the gene does not comprise the SNP A737G (resulting in the amino acid substitution Lys246Arg). Accordingly, in this embodiment, adenine is present at position 737 of the nucleotide sequence. Preferably, therefore, lysine is present at position 246 of the amino acid sequence. Alternatively, in a next preferred embodiment, the gene comprises a thymine or a cytosine at position 737 of the nucleotide sequence. Least preferably, the gene comprises a guanine at position 737 of the nucleotide sequence.
[0238] VMR1 (YHL035Q In one embodiment, when the at least one additional gene is VMR1, most preferably the gene does not comprise the SNP A3616G (resulting in the amino acid substitution Asnl206Asp). Accordingly, in this embodiment, adenine is present at position 3616 of the nucleotide sequence. Preferably, therefore, asparagine is present at position 1206 of the amino acid sequence. Alternatively, in a next preferred embodiment, the gene comprises a thymine or a cytosine at position 3616 of the nucleotide sequence. Least preferably, the gene comprises a guanine at position 3616 of the nucleotide sequence.
[0239] DFP4 (YHL044W) In one embodiment, when the at least one additional gene is DFP4, most preferably the gene does not comprise the SNP G308C (resulting in the amino acid substitution ArglO3Thr). Accordingly, in this embodiment, guanine is present at position 308 of the nucleotide sequence. Preferably, therefore, arginine is present at position 103 of the amino acid sequence. Alternatively, in a next preferred embodiment, the gene comprises a thymine or an adenine at position 308 of the nucleotide sequence. Least preferably, the gene comprises a cytosine at position 308 of the nucleotide sequence.
[0240] BEM2 (YER155O In another embodiment, when the at least one additional gene is BEM2, most preferably the gene does not comprise the SNP A1109G (resulting in the amino acid substitution Asn370Ser). Accordingly, in this embodiment, adenine is present at position 1109 of the nucleotide sequence. Preferably, therefore, asparagine is present at position 370 of the amino acid sequence. Alternatively, in a next preferred embodiment, the gene comprises a thymine or a cytosine at position 1109 of the nucleotide sequence. Least preferably, the gene comprises a guanine at position 1109 of the nucleotide sequence.
[0241] MYG1 (YER156O In one embodiment, when the at least one additional gene is MYG1, most preferably the gene does not comprise the SNP G614A (resulting in the amino acid substitution Arg205Lys). Accordingly, in this embodiment, guanine is present at position 614 of the nucleotide sequence. Preferably, therefore, arginine is present at position 205 of the amino acid sequence. Alternatively, in a next preferred embodiment, the gene comprises a thymine or a cytosine at position 614 of the nucleotide sequence. Least preferably, the gene comprises an adenine at position 614 of the nucleotide sequence.
[0242] In another embodiment, when the at least one additional gene is MYG1, most preferably the gene does not comprise the SNP A340G (resulting in the amino acid substitution Asnll4Asp). Accordingly, in this embodiment, adenine is present at position 340 of the nucleotide sequence. Preferably, therefore, asparagine is present at position 114 of the amino acid sequence. Alternatively, in a next preferred embodiment, the gene comprises a thymine or a cytosine at position 340 of the nucleotide sequence. Least preferably, the gene comprises a guanine at position 340 of the nucleotide sequence.
[0243] In another embodiment, when the at least one additional gene is MYG1, most preferably the gene does not comprise the SNP A36T (resulting in the amino acid substitution Lysl2Asn). Accordingly, in this embodiment, adenine is present at position 36 of the nucleotide sequence. Preferably, therefore, lysine is present at position 12 of the amino acid sequence. Alternatively, in a next preferred embodiment, the gene comprises a guanine or a cytosine at position 36 of the nucleotide sequence. Least preferably, the gene comprises a thymine at position 36 of the nucleotide sequence. YCK1 (YHR135O
[0244] In one embodiment, when the at least one additional gene is YCK1, most preferably the gene does not comprise the SNP G1606T (resulting in the amino acid substitution Gly536Cys).
[0245] Accordingly, in this embodiment, guanine is present at position 1606 of the nucleotide sequence. Preferably, therefore, glycine is present at position 536 of the amino acid sequence. Alternatively, in a next preferred embodiment, the gene comprises an adenine or a cytosine at position 1606 of the nucleotide sequence. Least preferably, the gene comprises a thymine at position 1606 of the nucleotide sequence.
[0246] ARO9 (YHR137W) In one embodiment, when the at least one additional gene is ARO9, most preferably the gene does not comprise the SNP C17T (resulting in the amino acid substitution Ala6Val). Accordingly, in this embodiment, cytosine is present at position 17 of the nucleotide sequence. Preferably, therefore, alanine is present at position 6 of the amino acid sequence. Alternatively, in a next preferred embodiment, the gene comprises an adenine or a guanine at position 17 of the nucleotide sequence. Least preferably, the gene comprises a thymine at position 17 of the nucleotide sequence. DSE2 (YHR143W)
[0247] In one embodiment, when the at least one additional gene is DSE2, most preferably the gene does not comprise the SNP T320C (resulting in the amino acid substitution VallO7Ala). Accordingly, in this embodiment, thymine is present at position 320 of the nucleotide sequence. Preferably, therefore, valine is present at position 107 of the amino acid sequence. Alternatively, in a next preferred embodiment, the gene comprises an adenine or a guanine at position 320 of the nucleotide sequence.
[0248] Least preferably, the gene comprises a cytosine at position 320 of the nucleotide sequence. CRP1 (YHR146W)
[0249] In one embodiment, when the at least one additional gene is CRP1, most preferably the gene does not comprise the SNP A643G (resulting in the amino acid substitution Ile215Val). Accordingly, in this embodiment, adenine is present at position 643 of the nucleotide sequence. Preferably, therefore, isoleucine is present at position 215 of the amino acid sequence. Alternatively, in a next preferred embodiment, the gene comprises a thymine or a cytosine at position 643 of the nucleotide sequence. Least preferably, the gene comprises a guanine at position 643 of the nucleotide sequence. In another embodiment, when the at least one additional gene is CRP1, most preferably the gene does not comprise the SNP A952G (resulting in the amino acid substitution Ile318Val). Accordingly, in this embodiment, adenine is present at position 952 of the nucleotide sequence. Preferably, therefore, isoleucine is present at position 318 of the amino acid sequence. Alternatively, in a next preferred embodiment, the gene comprises a thymine or a cytosine at position 952 of the nucleotide sequence. Least preferably, the gene comprises a guanine at position 952 of the nucleotide sequence. MAM33 (YIL070Q
[0250] In one embodiment, when the at least one additional gene is MAM33, most preferably the gene does not comprise the SNP G80T (resulting in the amino acid substitution Trp27Leu). Accordingly, in this embodiment, guanine is present at position 80 of the nucleotide sequence. Preferably, therefore, tryptophan is present at position 27 of the amino acid sequence. Alternatively, in a next preferred embodiment, the gene comprises a cytosine or an adenine at position 427 of the nucleotide sequence. Least preferably, the gene comprises a thymine at position 427 of the nucleotide sequence.
[0251] BMT5 (YIL096Q
[0252] In another embodiment, when the at least one additional gene is BMT5, most preferably the gene does not comprise the SNP C793T (resulting in the amino acid substitution Leu265Phe). Accordingly, in this embodiment, cytosine is present at position 793 of the nucleotide sequence. Preferably, therefore, leucine is present at position 265 of the amino acid sequence. Alternatively, in a next preferred embodiment, the gene comprises an adenine or a guanine at position 793 of the nucleotide sequence. Least preferably, the gene comprises a thymine at position 793 of the nucleotide sequence.
[0253] MOB1 (YIL106W)
[0254] In one embodiment, when the at least one additional gene is MOB1, most preferably the gene does not comprise the SNP G85A (resulting in the amino acid substitution Ala29Thr). Accordingly, in this embodiment, guanine is present at position 85 of the nucleotide sequence. Preferably, therefore, alanine is present at position 29 of the amino acid sequence. Alternatively, in a next preferred embodiment, the gene comprises a thymine or a cytosine at position 85 of the nucleotide sequence. Least preferably, the gene comprises an adenine at position 85 of the nucleotide sequence.
[0255] NDD1 (YOR372Q
[0256] In one embodiment, when the at least one additional gene is NDD1, most preferably the gene does not comprise the SNP A1510G (resulting in the amino acid substitution Ser504Gly). Accordingly, in this embodiment, adenine is present at position 1510 of the nucleotide sequence. Preferably, therefore, serine is present at position 504 of the amino acid sequence. Alternatively, in a next preferred embodiment, the gene comprises a thymine or a cytosine at position 1510 of the nucleotide sequence. Least preferably, the gene comprises a guanine at position 1510 of the nucleotide sequence. In another aspect of the invention, there is provided a recombinant or engineered eukaryotic cell exhibiting improved recombinant protein production compared to a corresponding wild-type, progenitor, or common laboratory strain cell, the recombinant or engineered eukaryotic cell comprising a non-naturally occurring combination of alleles associated with improved recombinant protein production, wherein the at least one allele is for a gene selected from a list of genes in Table 1, or a homologue, orthologue or paralogue thereof, wherein the at least one gene is associated with recombinant protein production.
[0257] In another aspect of the invention, there is provided a recombinant or engineered eukaryotic cell exhibiting improved recombinant protein production compared to a corresponding wild-type, progenitor, or common laboratory strain cell, the recombinant or engineered eukaryotic cell comprising at least one gene with a non- naturally occurring combination of single nucleotide polymorphisms (SNPs), wherein the at least one gene is selected from a list of genes in Table 1, or a homologue, orthologue or paralogue thereof, wherein the at least one gene is associated with recombinant protein production.
[0258] In another aspect of the invention, there is provided a recombinant or engineered eukaryotic cell exhibiting improved recombinant protein production compared to a corresponding wild-type, progenitor, or common laboratory strain cell, the recombinant or engineered eukaryotic cell comprising at least one modified gene, wherein the at least one gene is selected from a list of genes in Table 1, or a homologue, orthologue or paralogue thereof, wherein the at least one gene is associated with recombinant protein production.
[0259] In one embodiment, the recombinant or engineered eukaryotic cell comprises a non-naturally occurring combination of alleles associated with improved recombinant protein production, wherein the at least one allele is for a gene selected from a list of genes in Table 1, or a homologue, orthologue or paralogue thereof, and the recombinant or engineered eukaryotic cell comprises at least one modified gene, wherein the at least one gene is selected from a list of genes in Table 1, or a homologue, orthologue or paralogue thereof. In another embodiment, the recombinant or engineered eukaryotic cell comprises at least one gene with a non-naturally occurring combination of single nucleotide polymorphisms (SNPs), and at least one modified gene, wherein the at least one gene is selected from a list of genes in Table 1, or a homologue, orthologue or para log ue thereof.
[0260] In one preferred embodiment, the eukaryotic cell is a fungal cell. The fungal cell may be unicellular or have a hyphal (filamentous) morphology or exist as propagules. Most preferably, the eukaryotic cell is a yeast cell.
[0261] In one embodiment, the yeast cell is Pichia pastoris. Preferably, the Pichia pastoris is a Komagataella species (such as K. phaffii, K. pastoris, and K. pseudopastoris), Hansenula polymorpha (also known as Ogataea polymorpha), Kluyveromyces lactis, a Yarrowia species (such as Yarrowia lipolytica), or Schizosaccharomyces pombe. Alternatively, in a preferred embodiment, the yeast cell is a Saccharomyces species yeast, such as Saccharomyces cerevisiae. Most preferably, the yeast cell is Saccharomyces cerevisiae. In one embodiment, the yeast cell is a Wine / European (WE), West African (WA), North American (NA) or Sake (SA) strain. In one embodiment, when the at least one gene is MEC3 (YLR288C), the yeast cell is a Wine / European (WE) strain. In one embodiment, when the at least one gene is YLR287C, the yeast cell is a Wine / European (WE) strain. In one embodiment, when the at least one gene is GAT1, the yeast cell is a West African (WA) strain. In one embodiment, when the at least one gene is UBP14, the yeast cell is a West African (WA) strain.
[0262] In another embodiment, the eukaryotic cell is a fungal cell such as an Aspergillus species, including Aspergillus oryzae and Aspergillus niger, or a Trichoderma species, or Myceliophthora thermophila.
[0263] In another embodiment, the eukaryotic cell is an insect cell. Examples of insect cells include cell lines Sf9 and Sf21 from Spodoptera frugiperda cells, Hi-5 from Trichoplusia ni cells, and Schneider 2 cells and Schneider 3 cells from Drosophila melanogaster cells. Alternatively, in a preferred embodiment, the eukaryotic cell is from Excavata, such as Leishmania tarentolae. In another embodiment the eukaryotic cell is a mammalian cell type, such as a Chinese hamster ovary (CHO) cell, a Mouse myeloma lymphoblastoid, e.g. an NSO cell, a Human embryonic kidney cell, e.g. a HEK-293 cell, a Human embryonic retinal cell, e.g. a Crucell's Per.C6 cell, or a Human amniocyte cell, e.g. Glycotope or CEVEC.
[0264] A "common laboratory strain cell" will be well-known to the skilled person and may include those defined in Louis, E.J. 20161. A common laboratory strain may include yeast strains listed on the Saccharomyces Genome Database (SGD). A common laboratory strain may include one of the following yeast strains: S288C (Reference Genome: GenBank GCF_000146045.2); W303 (GenBank: JRIUOOOOOOOO.l);
[0265] CEN.PK; JRY188; AH22; S150-2B; and CB11 / 63. A common laboratory strain is not a natural strain, and therefore, may contain its own non-naturally occurring combination of alleles.
[0266] In a case where a yeast strain has been developed through multiple steps, e.g., to improve bioprocessing phenotypes, a progenitor strain of this invention is the original strain used in a strain development program, i.e. the closest relative, or least genetically diverse, compared to the wild-type. The progenitor strain may also be any strain, such as an intermediate strain in a multi-strain lineage, which has been further improved, e.g. for recombinant protein production, to give a final production strain derived from it.
[0267] Preferably, the term "recombinant" when referring to a "recombinant eukaryotic cell", will be understood to mean a eukaryotic cell into which recombinant DNA has been introduced, i.e. a cell containing genetically engineered DNA.
[0268] In one embodiment, the term "recombinant protein" may be any protein not naturally produced by the expression host, including fusion proteins, tagged proteins, muteins, analogues, derivatives, domains, precursors and fragments of any protein or polypeptide, including, but not limited to, the following proteins (or other polypeptides) of interest. Proteins (or other polypeptides) of interest include albumin, transferrin, lactoferrin, immunoglobulin (such as an Fab fragment or single-chain antibody, including, ScFvs, VHHs and VNARs), (haemo)globin, leghaemoglobin, myoglobin, blood clotting factors (such as factors II, VII, VIII, IX), von Willebrand's factor, tick anticoagulant peptide, endostatin, angiostatin, icestructuring proteins, hydrophobins interferons, interleukins, alpha-l-antitrypsin, insulin, GLP-1, glucagon, calcitonin, cell surface receptors, fibronectin, prourokinase, (pre-pro)-chymosin, antigens for vaccines (including virus-like particles), t-PA, urokinase, prourokinase, hirudin, tumour necrosis factor, G-CSF, GM-CSF, Kunitz domain proteins, CNTF, growth hormone, transforming growth factors, fibroblast growth factors, nerve growth factors, serum cholinesterase, aprotinin, amyloid precursor protein, inter-alpha trypsin inhibitor, antithrombin III, apolipoproteins, bone morphogenic proteins, MIC-1, leptin, erythropoietin (EPO), thrombopoietin (TPO), parathyroid hormone, platelet-derived endothelial cell growth factor, platelet-derived growth factor, Protein C, Protein S, keratins, collagens, antimicrobial peptides, defensins, chymosins, casein, amylases, and enzymes generally, such as glucose oxidase and superoxide dismutase. The protein may be a viral, microbial, fungal, plant or animal protein, for example, a mammalian protein. In one embodiment, it is a human protein. The recombinant protein may be a protein endogenous to the host, such as an enzyme, for which production has been improved by strain engineering in a recombinant eukaryotic cell.
[0269] Preferably, the term "engineered" when referring to an "engineered eukaryotic cell", will be understood to mean a eukaryotic cell whose genome comprises an addition, deletion or modification of genetic sequences, i.e. a cell containing genetically engineered DNA. Such engineering may be achieved by standard breeding or crossing technologies, mutagenesis or targeted genome or plasmid engineering methods. The term "allele" refers to the alternative forms of a gene that are found at the same place on a chromosome. For example, genetically diverse strains used for breeding may contain different alleles for a particular gene which have arisen from mutations within the gene. Novel alleles may also be generated during breeding, e.g. through recombination between alleles from different parents.
[0270] It is well-known that a SNP is a substitution of a single nucleotide at a specific position within the genome. In a preferred embodiment, the SNP is a non- synonymous SNP. A non-synonymous SNP is a SNP that results in a single amino acid substitution in the protein sequence encoded by the nucleotide sequence. In another preferred embodiment, the SNP introduces or removes a stop codon. The SNP may also be synonymous within a coding sequence or within an intergenic region, which may also result in phenotypic changes. It will be appreciated that a SNP within one gene or an intergenic region might impact the expression of adjacent genes in the locus, thereby influencing a phenotype indirectly. For example, a SNP in an intergenic region might affect the promoter function of one or both adjacent genes. Similarly, a SNP in one gene, especially if it results in a stop codon or significantly alters DNA structural features, might prevent, reduce or increase transcription of that gene, leading to altered expression of an adjacent gene.
[0271] The term "homologue" will be well understood by the skilled person to mean a gene or genetic region that is similar in sequence, structure or evolutionary origin to a gene or genetic region in another species or organism. For example, homologous genes may be derived from a single common ancestral gene present in the common ancestor of different organisms. Homologous genes will encode proteins with the same or similar function in different species, and may also be referred to as an "ortholog ue".
[0272] The sequence identity between two genes, nucleotide sequences or protein sequences can be defined by a sequence alignment algorithm, such as Smith- Waterman or Needleman-Wunsch for local and global alignments, respectively. These algorithms compare the sequences and calculate a score based on the similarity between the nucleotides or residues at each position. The resulting score is used to infer the degree of identity between the sequences. Other popular algorithms for measuring sequence identity include BLAST (Basic Local Alignment Search Tool) and FASTA (Fast Alignment Search Tool for Assessment of similarity). These algorithms are widely used for large-scale sequence comparisons and are commonly available in bioinformatics software packages. For example, the algorithm is in CloneManager 11 software (Sci Ed Software LLC), selecting the Align > Compare Two Sequences > Global options, using standard Scoring Matrix and other parameters for either amino acid or DNA sequences.
[0273] The level of identity is preferably at least 30%, at least 40%, at least 50%, more preferably at least 60%, at least 70%, at least 80%, or at least 90% for amino acid sequences. Preferably, the level of identity is at least 30%, at least 40%, at least 50%, more preferably at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 99% or at least 99.9% for nucleotide sequences. The term "orthologue" will be well understood by the skilled person to mean one of two or more homologous gene sequences found in different species. The term "paralogue" will be well understood by the skilled person to mean a gene which has evolved by a gene duplication event within a genome. For example, gene duplication within a single species may involve one copy of the gene receiving a mutation that gives rise to a new gene. Paralogous genes code for a protein with similar, but not necessarily identical functions.
[0274] A gene "associated with recombinant protein production" may be one in which different alleles or modification of the gene, result in a modulation in recombinant protein production. A cell exhibiting "improved recombinant protein production" may also be defined as a cell exhibiting increased recombinant protein production. Preferably, improved recombinant protein production comprises improved secretion. In one embodiment, the recombinant or engineered eukaryotic cell may exhibit an increase in recombinant protein production or secretion compared to a eukaryotic cell that does not comprise at least one gene selected from a group consisting of: a gene present in Saccharomyces cerevisiae chromosome XII 706198 - 717026 inclusive (SEQ ID No: 11), UBP3 (YER151C), and BUD27 (YFL023W), or at least one gene selected from Table 1, with a non-naturally occurring combination of single nucleotide polymorphisms (SNPs), or a non-natural combination of alleles of genes selected from a group consisting of: a gene present in Saccharomyces cerevisiae chromosome XII 706198 - 717026 inclusive (SEQ ID No: 11), UBP3 (YER151C), and BUD27 (YFL023W), or of genes listed in Table 1, or has a gene selected from a group consisting of: a gene present in Saccharomyces cerevisiae chromosome XII 706198 - 717026 inclusive (SEQ ID No: 11), UBP3 (YER151C), and BUD27 (YFL023W), or from Table 1 modified to improve recombinant protein production.
[0275] For example, in a preferred embodiment, the recombinant or engineered eukaryotic cell may exhibit an increase in recombinant protein production or secretion compared to a wild-type, progenitor, or common laboratory strain cell of at least 5%, at least 10%, at least 20%, at least 30%, at least 40%, at least 50%, at least
[0276] 60%, at least 70%, at least 80%, at least 90%, or at least 100%. Alternatively, the recombinant or engineered eukaryotic cell may exhibit an increase in recombinant protein production or secretion compared to a wild-type, progenitor, or common laboratory strain cell, of at least 2-fold, at least 3-fold, at least 4-fold, at least 5- fold, at least 6-fold, at least 7-fold, at least 8-fold, at least 9-fold or at least 10- fold. Alternatively, the recombinant or engineered eukaryotic cell may exhibit an increase in recombinant protein production or secretion compared to a wild-type, progenitor, or common laboratory strain cell, with total soluble intracellular protein or culture supernatant yields of at least 1 g / L, at least 2 g / L, at least 3 g / L, at least 4 g / L, at least 5 g / L, at least 6 g / L, at least 7 g / L, at least 8 g / L, at least 9 g / L, at least 10 g / L, at least 20g / L, at least 30 g / L, at least 40 g / L, at least 50 g / L, at least 60 g / L, at least 70 g / L, at least 80 g / L, at least 90 g / L, at least 100 g / L, at least 200g / L, or at least 300g / L.
[0277] In a preferred embodiment, the at least one modified gene may be a gene that has been over-expressed, a gene that has been knocked-down or a gene that has been knocked-out. In another embodiment, the at least one modified gene may be engineered to alter the protein function, e.g. by protein engineering to change the amino acid sequence. This may alter the activity of the protein, including its interaction with other cellular components, such as metabolites, substrates, cofactors, nucleic acids, proteins, lipids and carbohydrates. In another embodiment, e.g. if the gene encodes for a nucleic acid component such as a tRNA, the modified gene may be engineered to alter its structure and binding properties. This engineering may also affect the stability and cellular levels of the gene product, e.g. by modulating its degradation. Undesirable proteolysis often occurs within eukaryotic cells when producing recombinant protein products, which can affect the final product and have significant cost and yield implications. Accordingly, in a preferred embodiment, the engineered or recombinant eukaryotic cell (preferably, Saccharomyces cerevisiae) exhibits reduced proteolysis.
[0278] The inventors have identified that proteinase A gene PEP4) disruption was important for reducing the overall protease levels in Saccharomyces cerevisiae strains that are genetically diverse for protease metabolism. Accordingly, in a preferred embodiment, the engineered or recombinant eukaryotic cell has a modified or disrupted PEP4 gene, or a homologue, orthologue or paralogue thereof. More preferably, the engineered or recombinant Saccharomyces cerevisiae cell has a modified or disrupted PEP4 gene. It will be appreciated that the recombinant or engineered eukaryotic cell according to the invention, can be used for the production of recombinant proteins. Accordingly, in a fourth aspect, there is provided use of the recombinant or engineered eukaryotic cell according to the first, second or third aspect, for producing a recombinant protein.
[0279] In a fifth aspect, there is provided a method for producing a recombinant protein, the method comprising:
[0280] (i) transforming the recombinant or engineered eukaryotic cell or obtaining a transformed recombinant or engineered eukaryotic cell according to the first, second or third aspect with an expression vector encoding at least one recombinant protein; and (ii) culturing the cell in a medium under conditions to produce the recombinant protein.
[0281] The method of culturing the cell under conditions to produce the recombinant protein will be well-known to the skilled person, and will depend on the nature of the cell and the recombinant protein. For example, the cell may be cultured in conventional nutrient medium well-known to the skilled person for culturing eukaryotic cells, and preferably yeast cells. This may be a minimal media containing Yeast Nitrogen Base, e.g. without amino acids, or Yeast Nitrogen Base without amino acids or ammonium sulphate, with ammonium sulphate added separately, which is buffered with sodium phosphate / citrate pH 6.0 (e.g. made using citric acid and di-sodium hydrogen orthophosphate) and containing 2% w / v glucose, which may or may not be supplemented with amino acids, vitamins and other nutrients, including Adenine, L-Arginine, L-Aspartic acid, L-Histidine, L-Isoleucine, L-Leucine, L-Lysine, L-Methionine, L-Phenylalanine, L-Threonine, L-Tryptophan, L-Tyrosine, Uracil or Valine. Media are preferably made with animal-free components. Rich media, e.g. YPD media with 1% yeast extract, 2% peptone, 2% glucose, may also be used comprising yeast extract, peptone and glucose. Alternative carbon sources can be used, such as sucrose, galactose or glycerol. For stable maintenance of the expression plasmid selective media are preferred. For example, for plasmids containing the LEU2 selectable marker, media lacking leucine are preferred. For large-scale culture in stirred tank bioreactors additional components such as antifoams may be preferred with media and feed-regimes. In one preferred embodiment, the expression vector is an episomal partial-2-micron plasmid, or a CEN vector, or a yeast artificial chromosome, encoding at least one recombinant protein, or integrating an expression cassette encoding at least one recombinant protein into the genome. In one preferred embodiment, the expression vector is a whole-2-micron family plasmid. The whole-2-micron family plasmid may be an expression vector comprising at least 50%, at least 60%, at least 70%, at least 80%, preferably at least 90%, or most preferably 100% of the sequence of a natural yeast 2-micron plasmid. Alternatively, the whole-2-micron family plasmid may be a whole-2- micron-family plasmid from another species as described in Sleep et al. 20052(W02005061719A1).
[0282] Preferably, the method comprises transforming at least one yeast strain with a whole -2-micron family plasmid. The whole -2-micron plasmid of Saccharomyces cerevisiae is a small circular, multicopy DNA element that resides in the yeast nucleus at a copy number of about 40-60 per haploid cell. Examples of whole 2- micron family plasmids include Scpl, Scp2 and Scp3, or those described by Strope et a / . 20153. More preferably, the expression vector is an engineered whole-2-micron family plasmid. Preferably, an engineered whole-2-micron family plasmid is a whole-2- micron family plasmid which has been engineered for recombinant protein production (a whole-2-micron expression plasmid). Preferably, recombinant protein production is inactive, repressed or uninduced during breeding.
[0283] Alternatively, the expression plasmid is a stable partial-2-micron plasmid. A stable partial-2-micron plasmid may be an expression vector comprising the 2-micron origin of replication which is dependent on a whole-2-micron family plasmid, a whole-2-micron expression plasmid or functions provided from these plasmids for stable replication and maintenance. Alternatively, the expression plasmid may be an integrative plasmid (e.g. a plasmid that integrates into the yeast genome).
[0284] Alternatively, the expression plasmid may be a centromeric plasmid (e.g. containing a centromeric sequence and / or an autonomous replicating sequence). Alternatively, the expression plasmid may be an artificial chromosome (e.g. a yeast artificial chromosome or YAC). In one embodiment, the cell is cultured for between 12 and 24 hours, between 24 and 36 hours, between 36 and 48 hours, between 48 and 72 hours, between 72 and 96 hours, between 96 and 120 hours, between 120 and 144 hours, or for a duration greater than 144 hours. In another embodiment, the cell is cultured for a time sufficient to reach the desirable production yields of the recombinant protein.
[0285] In one embodiment, the cell is cultured at between 20°C and 40°C, between 25°C and 35°C, between 26°C and 34°C, between 27°C and 33°C, and between 28°C and 32°C. Most preferably, the cell is cultured at 30°C. The method according to the fifth aspect may also comprise the step of isolating and / or purifying the recombinant protein. Such processes for isolating and purifying the recombinant protein will be well-known to the skilled person, and may include, for example, precipitation, ultrafiltration, gel electrophoresis, and chromatography. In a sixth aspect, there is provided a recombinant protein obtained from the recombinant or engineered eukaryotic cell according to the first, second or third aspect.
[0286] Preferably, the recombinant protein is purified.
[0287] It will be appreciated that the invention extends to any nucleic acid or peptide or variant, derivative or analogue thereof, which comprises substantially the amino acid or nucleic acid sequences of any of the sequences referred to herein, including variants or fragments thereof. The terms "substantially the amino acid / nucleotide / peptide sequence", "variant" and "fragment", can be a sequence that has at least 40% sequence identity with the amino acid / nucleotide / peptide sequences of any one of the sequences referred to herein, for example 40% identity with the sequence identified as SEQ ID No: 1-11, and so on. Amino acid / polynucleotide / polypeptide sequences with a sequence identity which is greater than 65%, in some embodiments, greater than 70%, in some embodiments, greater than 75%, and in some embodiments, greater than 80% sequence identity to any of the sequences referred to are also envisaged. In some embodiments, the amino acid / polynucleotide / polypeptide sequence has at least 85% identity with any of the sequences referred to, in some embodiments at least 90% identity, in some embodiments at least 92% identity, in some embodiments at least 95% identity, in some embodiments at least 97% identity, in some embodiments at least 98% identity and, in some embodiments at least 99% identity with any of the sequences referred to herein.
[0288] The skilled technician will appreciate how to calculate the percentage identity between two amino acid / polynucleotide / polypeptide sequences. In order to calculate the percentage identity between two amino acid / polynucleotide / polypeptide sequences, an alignment of the two sequences must first be prepared, followed by calculation of the sequence identity value. The percentage identity for two sequences may take different values depending on:- (i) the method used to align the sequences, for example, ClustalW, BLAST, FASTA,
[0289] Smith-Waterman (implemented in different programs), or structural alignment from 3D comparison; and (ii) the parameters used by the alignment method, for example, local vs global alignment, the pair-score matrix used (e.g. BLOSUM62, PAM250, Gonnet etc.), and gap-penalty, e.g., functional form and constants.
[0290] Having made the alignment, there are many different ways of calculating percentage identity between the two sequences. For example, one may divide the number of identities by: (i) the length of shortest sequence; (ii) the length of alignment; (iii) the mean length of sequence; (iv) the number of non-gap positions; or (iv) the number of equivalenced positions excluding overhangs. Furthermore, it will be appreciated that percentage identity is also strongly length dependent. Therefore, the shorter a pair of sequences is, the higher the sequence identity one may expect to occur by chance. Hence, it will be appreciated that the accurate alignment of protein or DNA sequences is a complex process. The popular multiple alignment program ClustalW (Thompson et al., 1994, Nucleic Acids Research, 22, 4673-4680; Thompson et al. , 1997, Nucleic Acids Research, 24, 4876-4882) is one way for generating multiple alignments of proteins or DNA in accordance with the invention. Suitable parameters for ClustalW may be as follows: For DNA alignments: Gap Open Penalty = 15.0, Gap Extension Penalty = 6.66, and Matrix = Identity. For protein alignments: Gap Open Penalty = 10.0, Gap Extension Penalty = 0.2, and Matrix = Gonnet. For DNA and Protein alignments: ENDGAP = -1, and GAPDIST = 4. Those skilled in the art will be aware that it may be necessary to vary these and other parameters for optimal sequence alignment. In some embodiments, calculation of percentage identities between two amino acid / polynucleotide / polypeptide sequences may then be calculated from such an alignment as (N / T)*100, where N is the number of positions at which the sequences share an identical residue, and T is the total number of positions compared including gaps but excluding overhangs. In some embodiments, overhangs are included in the calculation. Hence, one method for calculating percentage identity between two sequences comprises (i) preparing a sequence alignment using the ClustalW program using a suitable set of parameters, for example, as set out above; and (ii) inserting the values of N and T into the following formula:- Sequence Identity = (N / T)*100.
[0291] Alternative methods for identifying similar sequences will be known to those skilled in the art. For example, a substantially similar nucleotide sequence will be encoded by a sequence which hybridizes to DNA sequences or their complements under stringent conditions. By stringent conditions, we mean the nucleotide hybridizes to filter-bound DNA or RNA in 3x sodium chloride / sodium citrate (SSC) at approximately 45°C followed by at least one wash in 0.2x SSC / 0.1% SDS at approximately 20-65°C. Alternatively, a substantially similar polypeptide may differ by at least 1, but less than 5, 10, 20, 50 or 100 amino acids from the sequences shown in, for example, SEQ ID Nos: 1-11.
[0292] Due to the degeneracy of the genetic code, it is clear that any nucleic acid sequence described herein could be varied or changed without substantially affecting the sequence of the protein encoded thereby, to provide a functional variant thereof. Suitable nucleotide variants are those having a sequence altered by the substitution of different codons that encode the same amino acid within the sequence, thus producing a silent change. Other suitable variants are those having homologous nucleotide sequences but comprising all, or portions of, sequence, which are altered by the substitution of different codons that encode an amino acid with a side chain of similar biophysical properties to the amino acid it substitutes, to produce a conservative change. For example small non-polar, hydrophobic amino acids include glycine, alanine, leucine, isoleucine, valine, proline, and methionine. Large non-polar, hydrophobic amino acids include phenylalanine, tryptophan and tyrosine. The polar neutral amino acids include serine, threonine, cysteine, asparagine and glutamine. The positively charged (basic) amino acids include lysine, arginine and histidine. The negatively charged (acidic) amino acids include aspartic acid and glutamic acid. It will therefore be appreciated which amino acids may be replaced with an amino acid having similar biophysical properties, and the skilled technician will know the nucleotide sequences encoding these amino acids.
[0293] All of the features described herein (including any accompanying claims, abstract and drawings), and / or all of the steps of any method or process so disclosed, may be combined with any of the above aspects in any combination, except combinations where at least some of such features and / or steps are mutually exclusive.
[0294] For a better understanding of the invention, and to show how embodiments of the same may be carried into effect, reference will now be made, by way of example, to the accompanying Figures, in which:-
[0295] Figure 1 shows an expression construct (DNA) for amylase-mCherry expression and its protein product (on the left), with secreted amylase activity assay shown on a yeast agar plate containing starch (on the right). MET17 promoter = MET17p, SUC2 leader = Pre, ADH1 terminator = ADHlt.
[0296] Figure 2 shows specific mCherry activities for 41 Progeny Yeast Strains grown at 30°C. The average across four replicates is displayed, and error bars show 1 SD from 4 replicates.
[0297] Figure 3 shows mCherry vs amylase for yeast strains grown in four microtitre plates (Plates 1 to 4) at 30°C. Figure 4 shows different embodiments of an expression construct for amylase- mCherry expression which was inserted into a whole-2-micron plasmid and the protein derived from it. Figure 4a is the construct, pHRIK, which encodes the entire 2-micron plasmid from yeast strain YLF185, and the S. cerevisiae S288C LEU2 gene. Figure 4b is the construct, pHR2A, which encodes the pHRIK Notl-Hpal fragment, the pHRIK Notl-Adl fragment and the pUC57-Amp Eco J-Narl fragment. Figure 4c is the construct, pEV7, which encodes a repressible MET17 promoter driving expression of the alpha-amylase (AA) mCherry fusion protein without any N-linked glycosylation sites. Figure 4d is the construct, pHR!K-pEV7, which comprises DNA from the pHRIK and pEV7 constructs, and is the final whole- 2-micron expression in the cell. Table 1 shows the 16 genomic regions identified using QTL analysis, which contain genes, and thereby alleles, responsible for the differential expression of the recombinant amylase-mCherry protein. The table defines the 16 QTL regions from data at 30°C, either based on the LOD (Logarithm Of the Odds) scores (i.e. the local maximum and 1.5 LOD drop), or + / - 5Kb or 10Kb from the peak itself. The four rows with grey backgrounds have the maximum scores for the replicates or mean LOD.
[0298] Table 2 defines the non-synonymous SNPs in the genes listed in Table 1, which are different in the best two strains and worst two strains. The columns indicate which are the most preferred, preferred / neutral and least preferred bases at these positions.
[0299] Table 3 defines the prioritised list of genes which are associated with the phenotype of increased recombinant protein production. This table indicates both the gene name and the systematic name.
[0300] Table 4 shows the results of the reciprocal hemizygosity analysis, measuring foldchange improvements in OD normalised mCherry fluorescence in strains with the alleles of interest, compared to isogenic strains without the alleles of interest. The strain identified for carrying the allele of interest (MEC3 (YLR288C), YLR287C, GAT1 (YFL021W) and UBP14 (YBR058C)) is highlighted bold and underlined. Strains disrupted in their relevant allele are denoted by the suffix -K. Strains transformed to introduce the expression vector pHRlK-pEV7 are denoted by the suffix -L. Table 5 shows the plasmids for expression of recombinant proteins, with a description of the expression cassettes they contain (including any detection tags and linkers) and whether they are designed for intracellular or secreted production.
[0301] Table 6 shows a comparison of one of the best two strains for amylase-mCherry secretion (2-A2) with one of the worst two strains for amylase-mCherry secretion (1-C12), for expression of multiple different recombinant proteins. Table 7 shows preferred SNPs in Saccharomyces cerevisiae chromosome XII 706198 - 717026 inclusive (SEQ ID No: 11), including the most preferred non- synonymous SNPs in coding sequences, synonymous SNPs in coding sequences and SNPs outside coding sequences.
[0302] Examples
[0303] The inventors set out to identify the key regions of the yeast genome responsible for improvements in recombinant protein production, whether for intracellular products or secreted products. In order to do this, the inventors performed quantitative trait loci (QTL) analysis, which identifies regions in the genome associated with improving the phenotype analysed. As described below, the inventors used two detectable markers which are secreted in order to measure and quantify recombinant protein production, i.e. the enzyme, amylase, and the fluorophore, mCherry from Anaplasma marginale (UniProt X5DSL3). These two readily detectable markers are secreted and acted as a measure of recombinant protein production, and facilitated screening and selection of final yeast strains harbouring desired alleles within the defined QTLs.
[0304] Materials and Methods
[0305] Strain Engineering, Breeding and Selection
[0306] The original strains described by Liti, G., Carter, D., Moses, A. et al. Population genomics of domestic and wild yeasts. Nature 458, 337-341 (2009), from the Saccharomyces Genome resequencing Project (SGRP), were made genetically tractable by Cubillos et al. 20094and Louvel et al. 20145, using standard methods of genome engineering. Representatives of the Wine / European (WE), West African (WA), North American (NA) and Sake (SA) clean lineages were selected, with derivatives YFL185, YFL187, YFL190, and YFL191 described in Louvel et al. 20145, being further modified in this study to create strains Q416, Q413, Q426 and Q427, respectively, with genotypes shown in Table 8 below. Plasmid curing is described by Rose and Broach 19906.
[0307] Table 8 - Yeast strains used in this work
[0308] Plasmid Construction pHRIK was constructed from two PCR fragments encoding the entire 2-micron plasmid from strain YLF185, which were cloned into the EcoRV site of pUC57-Kan (GeneWiz / Azenta) with a 2.1Kb PCR fragment encoding the S. cerevisiae S288C LEU2 gene (flanked by two Pad-sites introduced on the PCR primers), using NEBuilder® HiFi DNA Assembly Cloning Kit.
[0309] This whole 2-micron-family plasmid contained the LEU2 selectable marker and pUC57-Kan integrated at the SnaBI-site downstream of the 2-micron D gene (see GenBank: J01347.1 sequence for Sna BI site location). Sbfl and Notl sites were introduced on PCR primers for the directional insertion of expression constructs downstream of the LEU2 selectable marker. See the pHRIK map in Figure 4a for details. Standard methods were used for E. coli transformation and plasmid preparation.
[0310] Oligonucleotide PCR primers used to construct pHRIK are provided below, with F (forward) and R (reverse) primers binding to the regions PCR1-4 shown in the pHRIK map corresponding to the primer names (Figure 4a).
[0311] Table 9 - Primer sequences Regions of primer sequence binding are marked on the pHRIK map as PCR1 to PCR4. Amplification was performed using a Q5 HiFi PCR kit (NEB) and / or restriction enzyme digestion to generate: a) pUC57-Kan (from GeneWiz) cut with EtoRV (buffer 3.1) or uncut plasmid can be used as PCR template DNA. b) The 2-micron fragment from the SnaBI-site near STB to just before the second inverted repeat, which contains the first inverted repeat, FLP and R.EP2 of the 2-micron B-form. PCR primers contain 20-30bp homology to pUC57-Kan and the second 2-micron fragment, respectively. This fragment is around 3.1Kb in length. c) The 2-micron fragment containing the second inverted repeat, R.EP1 and D. PCR primers contain 20-30bp homology with the first 2-micron fragment and the
[0312] LEU2 fragment. This fragment is around 3.4Kb in length. d) The LEU2 fragment flanked by Pad-sites. PCR primers contain 20-30bp homology with the second 2-micron fragment (encoding R.EP1 and D) and the pUC57-Kan fragment. This fragment is around 2.1Kb in length and can be amplified from S288C genomic DNA.
[0313] PCR and DNA cloning methods are well known to those skilled in the art of molecular biology. Co-transformation of a cir° yeast strain, such as YLF185 cir°, with all four fragments followed by selection of leucine prototrophs, extraction of total DNA, transformation of f. coli with kanamycin selection and then extraction and sequencing of plasmid DNA for alignment to the expected pHRl-K sequence may be performed. Alternatively, seamless cloning can be performed in vitro (e.g. using a NEBuilder® HiFi DNA assembly kit) before the yeast and / or E. coli transformations. pHRIK plasmid DNA can be isolated from E. coli, and its identity confirmed by diagnostic restriction enzyme digests and / or DNA sequencing.
[0314] Alternatively, disintegration vectors could be used to construct expression plasmids as shown by Chinery et al. 19897.
[0315] Referring to Figure 4b, pHR2A was constructed in a 3-way ligation of the 2.6kp pHRIK Notl-Hpal fragment, the 1.2kp pHRIK Notl-AcII fragment and the 2.5kp pUC57-Amp EcoR -Narl fragment (also digested with NEB Antarctic phosphatase). Ligation used NEB high concentration T4 ligase with transformation into NEB10P competent cells for growth on LB plates containing ampicillin. A synthetic expression cassette for amylase-mCherry (GeneWiz / Azenta) was cloned into the Sbfl and Notl sites to give pEV7, as shown in Figure 4c. It was preferred in this case for the recombinant product to lack N-linked glycosylation, which might affect protein activities independently of product yield. Options include the expression of an alpha-amylase lacking potential N-linked glycosylation motifs (-N-X-S / T-) or expressing an alpha-amylase analogue modified to remove existing N-linked glycosylation motifs, e.g. Aspergillus oryzae alpha-amylase (UniProt P0C1B3) with serine 199 in the mature protein changed to an alanine residue. The coding sequence for the fluorophore, mCherry from Anaplasma marginale (UniProt X5DSL3), was genetically fused to the alpha-amylase C-terminal coding sequence to facilitate screening and selection of final yeast strains and to provide an additional phenotype for evaluating productivity and differential proteolysis. Alternatively, fluorescent proteins such as yEGFP, ymUkGl, mOrange and mNeonGreen and others, such as described by Thorn, 20178and Kaishima at al. 20169., could be used. For all codon sequences, unless specified otherwise, codons were selected for rapid mRNA translation in S. cerevisiae as described by Chu et al. 201410. pEV7 contains an expression construct encoding a repressible MET17 promoter as described by Solow et al., 200511, which is driving expression of the alpha-amylase (AA) mCherry fusion protein without any N-linked glycosylation sites (see the DNA map in Figure 1). The repressible MET17 promoter can be switched off during breeding (by adding methionine to the culture media at 20mM or above. See Solow et al., 200511for details of methionine concentrations needed for repression and expression) and switched on for the production of amylase-mCherry in selected strains by growth in media lacking methionine. Secretion was directed by the secretory leader sequence from the S. cerevisiae SUC2 (invertase) gene. This preleader sequence is removed by signal peptidase during translocation into the endoplasmic reticulum to give the mature amylase-mCherry protein, which is then secreted into the culture media. The S. cerevisiae ADH1 terminator ADHlt) was used for transcriptional termination.
[0316] Expression cassettes for additional recombinant proteins were designed similarly to pEV7, as Sbfl-Notl fragments, which were cloned in place of the amylase-mCherry expression cassette for transcription in the same direction as the LEU2 gene in the final whole-2-micron vectors. The plasmids for expression of the additional recombinant proteins are described below and in Table 5, with a description of the expression cassettes they contain (including any detection tags and linkers) and whether they are designed for intracellular or secreted production. pEVl contains an expression cassette for intracellular expression of mCherry. The mCherry coding sequence is essentially the same in all constructs of this invention. This expression cassette contains the MET17 promoter and ADH1 terminator described above for pEV7. pEV51 & pEV52 contain intracellular expression cassettes for the virus-like particle proteins HPV16(Ll)-mCherry and mCherry-GSl-HBsAg, respectively, where GS1 is a linker with sequence GSGGSGGSGPVTN (SEQ ID No: 9). HPV16(L1) encodes human papillomavirus type 16 major capsid protein LI (UniProt: P03101 VL1_HPV16). HBsAg encodes Hepatitis B virus S protein (GenBank: AIJ50189.1). These expression cassettes contain the MET17 promoter and ADH1 terminator described above for pEV7. pEV3 contains an expression cassette for the secretion of rHA-mCherry, where rHA encodes recombinant human albumin with the mature albumin sequence from UniProt P02768. This expression cassette contains the MET17 promoter, SUC2 leader and ADH1 terminator described above for pEV7. pEV388 contains an expression cassette for the secretion of rHA (without an mCherry tag) with transcription from the Saccharomyces cerevisiae proteinase B promoter PRBlp). Secretion is directed by the modified fusion leader sequence (mFL). The DNA coding sequence for mFL-rHA in pEV388 is the same as the open reading for mFL-rHA in SEQ ID 19 of W02004009819. pEV275 contains an expression cassette for a VHH domain antibody for prostatespecific membrane antigen, PSMA (Chatalic et al., 201520). This expression cassette contains the MET17 promoter, SUC2 leader and ADH1 terminator described above for pEV7. pEV299 contains the same expression cassette as pEV275, except with transcription from the Saccharomyces cerevisiae proteinase B promoter
[0317] (PRBlp). pEV298 contains the same expression cassette as pEV299, except for secretion of a VHH-mCherry fusion protein. pEV395 contains an expression cassette for secretion of a GLP1 (9-37) analogue precursor-GS-HiBit fusion protein with amino acid sequence EGTFTSDVSSYLEGQAAKEFIAWLVRGRGGGGGSGGGGSVSGWRLFKKIS (SEQ ID No: 10) with a Saccharomyces cerevisiae Mating Factor-alpha-derived leader sequence and Saccharomyces cerevisiae BCY1 -derived terminator.
[0318] Yeast Strain Transformation
[0319] S. cerevisiae Q427 (SA lineage) was transformed to leucine prototrophy using a lithium acetate method (Sigma Aldrich Yeast Transformation Kit YEAST1) with DNA fragments from pHRIK and pEV7 (described below) for in vivo assembly of the final expression plasmid pHRlK-pEV7 by homologous recombination, which is shown in Figure 4d. Prototrophic transformants were selected on media without leucine, and cryopreserved stocks were prepared with 25% glycerol after plating a single transformant to provide single colonies. The expression plasmid, pHRlK-pEV7, shown in Figure 4d, was subsequently transferred to all progeny during multigeneration breeding.
[0320] Approximately lOOng each of gel purified 7.3Kb pHRIK Bst I-Notl fragment and pEV7 digested with SwaI+Acc65I were co-transformed into Q427 cir° to give strain Q427 [pHRlK-pEV7] following homologous recombination (gap-repair) of the plasmid DNA fragments. Gap-repair transformation and other yeast methods are described in Andersen et al., 201212, and Finnis et al. 201813(WO2018234349A1). Equivalent yeast transformations were performed as required for the introduction of additional expression plasmids with homologous recombination (gap-repair) between DNA fragments comprising the whole-2-micron plasmid sequence and fragments comprising the expression construct with homologous flanking sequences from the different pEV-plasmids described above.
[0321] Breeding
[0322] Breeding methods are described in Cubillos et al. 201314. Breeding was performed for 12 generations before strain selection. Cpl2 populations were generated from Q416, Q413, Q426 and Q427 with two pairwise crosses first and then mixed 4-way crosses with selection for Ura + Lys+ in between cycles, in the presence of methionine repression. Cpl2 diploid libraries comprise approximately 25% PEP4: :PEP4 homozygotes, 50% PEP4: :pep4 heterozygotes, and 25% pep4: :pep4 homozygotes. Selection from Cpl2 cir° libraries, or Cpl2 libraries containing a whole-2-micron expression plasmid, such as pHR!K-pEV7 introduced by using Q427 [pHR!K-pEV7] as a parental stain in place of Q427 cir°, was performed using G418. G418 is an aminoglycoside antibiotic similar in structure to gentamicin Bl. , which blocks polypeptide synthesis by inhibiting the elongation step as described by Goldstein et al. 199915. Strains containing the KanMX gene are resistant to G418. A homozygous pep4::pep4 diploid population ("pure diploids", PD) was prepared by allowing germination of genetically diverse Cpl2 progeny in the presence of G418 to select for pep4: : KanMX spores only, followed by mating to form pep4::pep4 diploids. Alternatively, a mixed population of heterozygous pep4: :PEP4 and homozygous pep4: : pep4 diploids ("mixed diploids", MD was prepared by allowing germination and mating of pep4: : KanMX and PEP4 haploids followed by selection against PEP4:PEP4 homozygous diploids with G418. Therefore, multigenerational libraries can be produced with a range of proteinase A genotypes, including an absence of proteinase A gene, despite this protease being essential for the breeding process.
[0323] Strain Selection Spores with the pep4 genotype were selected with G418 after germination and mated to give homozygote pep4::pep4 diploids as described above. 48 strains were selected with a range of mCherry levels in the supernatant following microtitre plate culture. Cryopreserved stocks were made with 25% final glycerol concentration for storage at below -70°C. Genotypic and phenotypic data were obtained for all strains, of which seven strains were excluded, e.g., for poor growth and / or low- quality sequence data, and the remaining 41 strains were used for QTL analysis. Strains were typically grown on BMMD media as defined by Evans et al., 201016, with appropriate supplements, e.g., CSM-Leu (Formedium Ltd). For expression studies, strains were typically cultured in 0.5mL BMMD+(CSM-Leu-Met) media in 48-well microtitre plates (MTP) shaken at 30°C, 280rpm, 2.5cm orbit in a humidity chamber. Details of methods are provided in Schelde et al., 201917, Ramaiya et al. 201718(WO2017112847A1) and Finnis et al. 201813(WO2018234349A1). Phenotyping
[0324] Secreted amylase activity from the expression construct was confirmed by growing the yeast on agar plates containing starch, which was degraded by the secreted amylase to give a visible zone around the cultured yeast (Figure 1). Protease activities affecting either the amylase domain or the mCherry domain can influence the results from amylase assays or mCherry signal detection.
[0325] For accurate detection of specific mCherry levels in the culture supernatant for use as phenotypic data in QTL analysis, yeast stocks were grown up at 500pL scale in 48-well MTPs, using repressive media containing methionine for three days before cultures were transferred into non-repressible minimal growth media to allow recombinant protein expression. Cultures were set up in replicates of four. Growth was monitored, and cultures were harvested at 3.5 days, at which point OD620 and mCherry fluorescence readings were recorded using a TECAN Spark with the following settings: OD readings were taken at 620nm using 10 flashes and a settle time of 50 ms. mCherry fluorescence readings were taken using excitation of 540nm and emission of 614nm. Whole culture readings were corrected using a gain of 100, and supernatant-only values were corrected using a gain of 150. Cultures were then centrifuged, and 300pL supernatant was removed before an additional centrifugation step. 250pL cleared supernatant was removed, and OD620 and mCherry fluorescence was recorded using the settings above.
[0326] To generate amylase activity data, 50pL cleared supernatant was analysed using the EnzChek ® Ultra Amylase Assay Kit (E33651) in 48-well MTPs. Fluorescence intensity at 505 / 512nm was measured at various points during the assay using the TECAN Spark with a gain of 30, for 15 minutes.
[0327] DNA Preparation and Sequencing Genotypic data were obtained by the following methods: Genomic DNA was retrieved from a 5mL overnight culture of each strain grown in YPD media (1% yeast extract, 2% peptone, 2% glucose) using the Promega Wizard DNA extraction kit. gDNA was re-precipitated in lOmM Tris-HCI, pH 8.5 and DNA quality was assessed using a Nanodrop 2000 Spectrophotometer and visualised using agarose gel electrophoresis before genomic DNA sequencing. LITE (Low Input, Transposase Enabled) Library preparation was carried out on each DNA sample before fortyeight diploid genomes were sequenced on the Illumina NovaSeq 6000 SP lane, generating 150bp PE reads. Sequencing was carried out by the Earlham Enterprises Ltd, Norwich, UK. The average read number per sample was 10,556,161, which represents an average estimated genome coverage of 263 times. OTL Analysis
[0328] Quantitative Trait Locus (QTL) analysis is a statistical method that links both phenotypic and genotypic data to explain the genetic basis of variation in complex traits as described by Miles et al., 200819. This method was performed on a range of S. cerevisiae strains selected with a range of expression levels for the recombinant protein amylase-mCherry. This analysis identifies regions of the genome containing alleles (also containing SNPs, single nucleotide polymorphisms, or QTNs, quantitative trait nucleotides) causing a phenotype, e.g. increased levels of the recombinant protein amylase-mCherry. The QTL analysis identifies "regions" (also called "intervals") in the genome associated with improving the phenotype analysed.
[0329] Short reads were first assessed for sequencing quality using fastqc, before each read was aligned against the reference genome of S288C (R64-2-1), using bwa. Alignments were indexed and sorted using samtools, and duplicate reads marked and removed using picard tools. Variants were then called using freebayes. The parameters were set for the minimum mapping quality to 20 and ploidy to diploid (- -min-mapping-quality 20 -min-base-quality 20 -p 2), then this output was subjected to a set of filters to use as genetic markers with SNP sites for the samples.
[0330] The following filters were applied: a. The variant calling quality is more than 20; b. The observation of the variant is 100% of the samples in the calling set; c. Allele frequency (REF / (REF + ALT in (0.1, .09)); and d. Calling positions are the known bi-allele variant sites for SGRP founders.
[0331] The reproducibility of strain measurements across each plate was assessed using R. QTL (Quantitative Trait Loci) analysis was applied to find the association between genotypes and phenotype measurements for each plate as well as the average records. Specific activity values (mCherry fluorescence I OD620) from each culture (obtained using the methods described above) were used as phenotype inputs for the QTL analysis. LOD (Logarithm Of the Odds) score was calculated for each locus. The selected candidate QTL intervals are listed if 1) the LOD score is > 3 for each separate replicate plate analysis and 2) the max LOD score is > 5 for the max score when all replicate plates are considered. The intervals are summarised by the local maximum and 1.5 LOD drop. In addition, 5k and 10k flanking regions are also summarised for each of the selected peak markers. Variants in QTLs resulting in nonsynonymous mutation were further annotated. To narrow down and identify candidate causative genes, only the sites which appear in the top 2 performing strains and alternatives present in the bottom 2 performing strains are considered. Additionally, subsets of these genes with nonsynonymous mutations were selected based on gene function and position within each interval.
[0332] Plasmid Curing
[0333] Progeny strains are leucine auxotrophs with a non-functional genomic Ieu2 allele.
[0334] Growth in synthetic media lacking leucine is restored when strains are transformed with expression plasmids containing a functional copy of the LEU2 gene.
[0335] To cure yeast strains of the whole 2-micron expression plasmids for recombinant protein expression, such as original amylase-mCherry secretion plasmid, serial passaging was conducted in small (500pl) liquid cultures with a non-selective, synthetic drop-out medium containing leucine, e.g. BMMD+(CSM-Leu-Met) media in 48-well microtitre plates (MTP) shaken at 30°C, 280rpm, 2.5cm orbit in a humidity chamber. After 5 passages with growth to late-log or stationary phase and approximately 1-2% inocula, single colonies of each strain were isolated on YPD agar plates, patched onto YPD agar plates, and replica plated onto synthetic dropout plates lacking leucine. Colonies exhibiting no growth on the synthetic dropout medium indicated potential loss of the expression plasmid. To confirm plasmid loss, PCR of Leu- colonies was conducted with primers annealing to the 2p origin of the expression plasmid (Strope et al., 20153). Samples were analysed by gel electrophoresis, with the absence of a DNA amplicon indicating successful plasmid loss.
[0336] Reciprocal Hemizvaositv
[0337] Reciprocal hemizygosity is the standard method to validate the effect of one allele over the other in an Fl hybrid, as taught by Mackay et al. 200921. The principle of reciprocal hemizygosity is to assay the phenotype of interest in isogenic diploids that differ only in which allele of a candidate gene of interest is present using the methods described in Cubillos, et al. 201122, Liti and Louis 201224. In general, this involves a pair of haploid strains bearing different alleles at the locus of interest being mated to produce a diploid. Two diploids are made, each with one or the other allele deleted. Therefore, pairwise crossings of up to three haploid parental strains disrupted of the wild-type alleles, respectively, and the parental strain harbouring a novel allele of interest, were produced. Reciprocal crossings of the relevant haploid strain, disrupted for the novel allele, and up to three parental strains, harbouring the wildtype allele, were produced. For each gene, the allele of interest can only be found in one of the four parental backgrounds that were previously described in Table 8.
[0338] The expression cassette pHRlK-pEV7 was introduced to the diploid strains via homologous (gap-repair) transformation of the haploid cell of each above crossing that was not disrupted for any allele. Transformation was conducted as described above in 'Yeast Strain Transformation'.
[0339] All allele disruptions were coding sequence deletions, using standard methods of genome engineering known in the art. The deletion of each gene in the yeast genome was achieved by insertion of an antibiotic resistance cassette, KanMX, conferring resistance to G418. PCR-generated deletion cassettes can be made either from genomic DNA from the relevant strain from the deletion collection or from long oligos with homology to the flanking regions of the coding sequence of the gene of interest and homology to the KanMX cassette used for replacing the open reading frame (Cubillos, et al. 201122; Parts et al. 201123; Liti and Louis 201224; Cubillos et al. 201314).
[0340] Supernatant mCherry fluorescence values (FU) were measured for each diploid strain after growth using the protocols previously described for culturing and measurement. The values were normalised for cell growth (OD620). A ratio of these values for each pairwise crossing was calculated, where these values denote folddifference relative to the diploid with wild-type allele. An unpaired two-tailed t-test was conducted to determine if the result was statistically significant.
[0341] If different alleles of the gene alter amylase-mCherry expression in the isogenic diploids, the gene is associated with recombinant protein production. This also identifies the preferred allele for improving or increasing recombinant protein production. Results
[0342] Example 1 - mCherrv and amylase activities An amylase-mCherry fusion protein was expressed from pHRlK-pEV7 shown in Figure 4d in the genetically diverse yeast population. Strains were selected from this genetically diverse population with a range of mCherry levels. In the absence of proteolysis to degrade the secreted amylase-mCherry product, the different progeny selected would have different levels of amylase-mCherry polypeptide production with secretion into the extracellular media, e.g. based on their different productivities. If the full-length protein product is highly stable during both the secretion process and in the culture media (e.g. without significant proteolysis acting on the amylase-mCherry polypeptide), the ratio of amylase to mCherry activity is expected to remain relatively constant.
[0343] Figure 2 shows the mCherry activities from the 41 strains selected for growth at 30°C and used for the QTL analysis. The supernatant mCherry levels were corrected for growth differences (OD620) for the strains used in the QTL analysis. Significant diversity in the mCherry phenotype is observed, with more than a 10- fold difference observed between the highest and lowest producers. A 10.8 fold difference in mCherry levels in the supernatant for the highest producer (2A2) was obtained compared to the lowest producer (3D9).
[0344] Figure 3 shows plots of mCherry against amylase specific activities for each strain grown at 30°C, with strains grown in four MTPs. While the mCherry levels range widely from low to high, the amylase levels tend not to fall so far towards zero. This indicates that proteolysis has affected the mCherry and amylase domains differently. The mCherry domain appears to be more protease sensitive than the amylase domain.
[0345] Example 2 - QTL Analysis
[0346] QTL analysis identified 16 genomic regions comprising approximately 3.3% of the total Saccharomyces cerevisiae genome containing genes and alleles responsible for the differential expression of the recombinant amylase-mCherry protein (see Table 1). Successive statistical and bioinformatic filters and rational selection methods were used to shortlist sequences, e.g. QTLs, genes and QTNs, responsible for the increased recombinant protein production. 16 QTL intervals were identified using the specific mCherry activity values (mCherry fluorescence I OD620) at 30°C. Of these, 4 QTL intervals contained the maximum LOD scores for the mean LOD scores and the four replicates. These are QTLs 6, 7, 9 and 13 highlighted in grey in Table 1.
[0347] Table 1 shows the 16 genomic regions identified using QTL analysis. Each of these QTL regions lists the genes (i.e. interval genes), containing SNPs responsible for the differential expression of the recombinant amylase-mCherry protein. As such, the inventors identified these "interval genes" as being associated with recombinant protein production. The table defines the 16 QTL regions based on their chromosome number, and their start and end bp position.
[0348] Table 2 shows the non-synonymous SNPs identified in the genes listed in Table 1, which were found to differ amongst the best two strains at recombinant protein production and the worst two strains at recombinant protein production. In other words, the inventors identified these specific SNPs as being associated with good or poor recombinant protein production. The table identifies the SNP (i.e. the nucleotide substitution), as well as the resulting amino acid substitution.
[0349] The table then indicates whether this SNP is associated with improved recombinant protein production, by the presence of a "1 / 1" in the column "Good Performance Genotype". For these SNPs, the preferred nucleotide is the substituted base after the ">" symbol, and the preferred amino acid is the second one listed. For example, for YER151C, preferably the gene has an adenine at position 770 of the nucleotide sequence, and an asparagine at position 257 of the amino acid sequence.
[0350] Alternatively, the Table indicates if a SNP is associated with poor recombinant protein production, by the presence of a "1 / 1" in the column "Poor Performance Genotype". For these SNPs, the preferred nucleotide is the base before the ">" symbol, and the preferred amino acid is the first one listed. For example, for YIL105C, preferably the gene has an adenine at position 1873 of the nucleotide sequence, and a methionine at position 625 of the amino acid sequence. Additionally, Table 2 indicates which is the most preferred base at this position of the respective gene, which bases are preferred / neutral, and which base is the least preferred at this position, for improved recombinant protein production. Table 3 shows the preferred list of genes which are associated with the phenotype of increased recombinant protein production. The genes have been split into three groups, RPP 1, RPP 2, and RPP 3. RPP 1 are the most preferred genes associated with improved recombinant protein production, whilst RPP 2 is the second most preferred set of genes, and RPP 3 is the next preferred set of genes after that.
[0351] Example 3 - Reciprocal Hemizvaositv Analysis
[0352] As alleles of interest, MEC3 (YLR288C), YLR287C, GAT1 (YFL021W) and UBP14 (YBR058C) were chosen, harbouring all applicable SNPs to the gene, respectively, as described in Table 2.
[0353] For each gene, the allele of interest can only be found in one of the four parental backgrounds that are described in Table 8. The strain identified for carrying the allele of interest is highlighted bold and underlined in Table 4. In Table 4, strains disrupted in their relevant allele are denoted by the suffix -K. Strains transformed with the expression vector pEV7 are denoted by the suffix -L.
[0354] In some cases, haploid disruption of the relevant gene resulted in inviable strains, which could not be included in the dataset. In some cases, mating of the haploid strains would fail and the resulting diploid could therefore not be included in the dataset. Consequentially, Table 4 does not show the reciprocal pairings of WA / WE for YLR287C and SA / WA, WE / WA for UBP14. In Table 4, if the two isogenic diploid strains differing only by the allele expressed differ in outcome, then this validates that the variation at the gene of interest causes a difference in the phenotypic outcome.
[0355] As can be seen in Table 4, alleles of interest resulted in statistically significant fold- change improvements (values >1) in OD normalised mCherry fluorescence compared to isogenic strains without the allele of interest (** denotes p>0.01, **** denotes p>0.0001).
[0356] The skilled person will appreciate that reciprocal hemizygosity experiments validated the effectiveness of the alleles of interest to increase recombinant protein production and / or reduce proteolysis levels. Furthermore, the skilled person will appreciate that the phenotypic effects of the allele of interest that were measured for MEC3 (YLR288C), YLR287C, GAT1 (YFL021W) are independent of strain background and the positive properties of alleles of interest are therefore not a strain-specific invention. Example 4 - Improved Production of Multiple Protein Types
[0357] A comparison of one of the best two strains for amylase-mCherry secretion (2-A2) with one of the worst two strains for amylase-mCherry secretion (1-C12) was performed for multiple different recombinant proteins. Nine additional recombinant proteins were expressed, which were diverse in structure, size, and other physiochemical properties (Table 5). In all cases except one, the best strain for amylase mCherry production gave higher production for the other recombinant proteins (Table 6).
[0358] For amylase-mCherry control, 2-A2 gave approximately 8.6 times more amylase- mCherry based on mCherry fluorescence than 1-C12, which is statistically consistent with the results in Figure 2. For the other proteins, the fold increase was between approximately 4.2 and 1.2, with one protein giving approximately equal productivity between the two strains. In this case, HPV16(Ll)-mCherry production is very different to amylase-mCherry, so it is not unexpected that alleles beneficial to amylase-mCherry production that are present in strain 2-A2 would improve HPV16(L1) production as significantly as for the other recombinant proteins because the HPV16(Ll)-mCherry was expressed intracellularly for accumulation and VLP formation in the nucleus, whereas amylase-mCherry was expressed for secretion into the extracellular media. For all the other proteins, which had a range of detection tags and assay methods and were either secreted or expressed intracellularly for cytosolic accumulation, the alleles in 2-A2 were beneficial for improved recombinant protein production, e.g. through increased productivity and / or reduced proteolysis. The recombinant proteins expressed have a diverse range of sizes, folding, domain structures and other physiochemical structures, indicating that strain 2-A2 is also generally improved to produce many other recombinant proteins of interest. Multiple alleles beneficial for recombinant protein production and / or reduced proteolysis originating from the different parental strains have been combined in strain 2-A2. While this combination of alleles was originally selected for improved production of amylase-mCherry, clearly, many of these alleles and other combinations of these alleles and the SNPs within them are also beneficial for the production of multiple other recombinant proteins. For expression of multiple different types of recombinant protein, Strains 2-A2 and 1-C12 were cured of the whole-2-micron expression vector for amylase-mCherry (pHR!K-pEV7) by the method described above and retransformed for expression from whole-2-micron plasmids equivalent to pHRIK of multiple different recombinant proteins comprising the expression cassettes described in Table 5. All final whole-2-micron expression plasmids contain a LEU2 gene for leucine selection. Transformants were isolated as colonies on synthetic drop-out agar lacking leucine, e.g. BMMD+(CSM-Leu+Met). Three transformants were selected for each strain / plasmid combination for expression studies. Controls were strains 2-A2 and 1-C12 secreting amylase-mCherry from the pEV7 expression cassette in pHRlK- pEV7.
[0359] Inoculum cultures for three transformants of 2-A2 and 1-C12 for each plasmid (and triplicate pHRlK-pEV7 controls) were started by picking cells from patches on solid media and transferring to 500pL liquid cultures in clear, 48 well microtiter plates.
[0360] Buffered synthetic drop-out media with 2%(w / v) dextrose, lacking leucine and containing 3 g / L methionine was used to maintain plasmids and to repress expression from constructs utilising the MET17 promoter. Inoculum cultures were incubated at 30°C for 2 days after which 20pL of each inoculum culture was passaged into 500pL synthetic dropout media with 2%(w / v) dextrose, lacking leucine in new 48 well microtiter plates, e.g. BMMD+(CSM-Leu-Met), in triplicate, to inoculate expression cultures. Expression cultures were incubated in shaking humidity chambers at 30°C over 4 days before harvesting. Upon harvest, culture OD was measured in wells using a TECAN Spark plate reader (Tecan, Switzerland). Culture supernatants were isolated by centrifugation at 1800 RCF and analysed for secreted product, where applicable.
[0361] For strains transformed with pEVl, pEV51 and pEV52 for the expression of intracellular mCherry and mCherry-tagged recombinant proteins, mCherry fluorescence was measured directly from the expression cultures in 48 well clear microtiter plates upon harvest at Aex540nm; Aem614nm, gain 100 on a TECAN Spark plate reader (TECAN, Switzerland).
[0362] For strains transformed with pEV3, pEV7 and pEV298 expression constructs for secreted expression mCherry-tagged recombinant proteins, 200pL culture supernatant was isolated as described previously and transferred to new, clear, 48 well microtiter plates. mCherry fluorescence of supernatants was measured at Aex540nm; Aem614nm, gain 100 on a TECAN Spark plate reader. OD was also measured to check for any accidental transfer of the cell pellet.
[0363] For strains transformed with pEV388 for the secreted expression of recombinant human albumin (rHA), titres were quantified using the Albumin Blue Fluorescence Assay Kit (Active Motif, Belgium) with a modified protocol for high-throughput detection in 384 plates. Briefly, 12.5pL culture supernatant was transferred to wells in a black, clear bottomed, non-treated 96 well assay plate. 75pL assay reagent comprising Ipl Albumin Blue dye and 74pL Buffer A from the kit was added to each well and mixed by pipetting. Plates were incubated for 5 minutes at room temperature, then fluorescence at Aex560nm; Aem620nm was measured on a TECAN Spark plate reader. Fluorescence signals for each sample were averaged from three technical replicates and converted to relative levels using a standard curve of rHA prepared in expression medium.
[0364] For strains transformed with pEV395 for the secreted expression of HiBit-tagged GLP1 analogue precursor, titres were quantified using the Hi-Bit (HiBit) extracellular detection kit (Promega, US), according to the manufacturer's instructions. A standard curve was made from a HiBit-tagged control protein (Promega, US) of known concentrations prepared in expression medium. Supernatant samples were diluted 1 / 104to generate a signal within the linear range of the standard curve. Reactions were conducted in 20pL final volumes (lOpL sample, lOpL assay mix) in white, 384 well low-volume assay plates. Luminescence was measured on a BMG FLUOstar Omega plate reader (BMG Labtech, Germany). Luminescence signals for each sample were averaged from 3 technical replicates and converted to relative levels using the HiBit control protein standard curve.
[0365] For strains transformed with pEV299 and pEV275 for the secreted expression of untagged VHH, titres were quantified by SDS PAGE. Supernatant samples were run on NuPAGE 4-12% Bis-Tris gels (Thermo Fisher Scientific, US) according to the manufacturer's instructions, alongside 3 prepared samples of a purified VHH standard at known concentrations. Gels were Coomassie stained and imaged on an Amersham ImageQuant 800 (Cytiva, US), and densitometry analysis of bands corresponding to the VHH samples was conducted using ImageQuantTL (Cytiva, US). Band intensity values were converted to estimated relative levels using values obtained from the VHH reference standard. All data was for triplicates corrected for growth / biomass (ODeoo or OD620).
[0366] Accordingly, the inventors have demonstrated that yeast cells comprising the alleles and SNPs of interest are able to increase recombinant protein production of multiple different types of protein, including both secreted and intracellular proteins.
[0367] Conclusions
[0368] Using QTL analysis, the inventors have identified 16 genomic regions comprising approximately 3.3% of the total Saccharomyces cerevisiae genome containing genes and alleles responsible for the differential expression of the recombinant amylase-mCherry protein. Advantageously, by identifying these genomic regions, the inventors have identified specific genes, and SNPs within these genes, that are associated with improved recombinant protein production. As such, improved strains for recombinant protein manufacture can be provided, resulting in improvements in recombinant protein production, whether for intracellular products or secreted products.
[0369] References 1. Louis, E. J. Historical Evolution of Laboratory Strains of Saccharomyces cerevisiae.
[0370] Cold Spring Harb. Protoc. 2016, (2016).
[0371] 2. Sleep, D. & Finnis, C. 2-MICRON FAMILY PLASMID AND USE THEREOF, WO 2005 / 061719 Al. (2005).
[0372] 3. Strope, P. K. et al. 2p plasmid in Saccharomyces species and in Saccharomyces cerevisiae. FEMS Yeast Res. 15, (2015).
[0373] 4. Cubillos, F. A., Louis, E. J. & Liti, G. Generation of a large set of genetically tractable haploid and diploid Saccharomyces strains. FEMS Yeast Res. 9, 1217-1225 (2009).
[0374] 5. Louvel, H., Gillet-Markowska, A., Liti, G. & Fischer, G. A set of genetically diverged Saccharomyces cerevisiae strains with markerless deletions of multiple auxotrophic genes. Yeast 31, 91-101 (2014).
[0375] 6. Rose, A. B. & Broach, J. R. Propagation and expression of cloned genes in yeast: 2- microns circle-based vectors. Methods Enzymol. 185, 234-279 (1990).
[0376] 7. Chinery, S. A. 8i Hi nchliffe, E. A novel class of vector for yeast transformation. Curr. Genet. 16, 21-25 (1989). 8. Thorn, K. Genetically encoded fluorescent tags. Mol. Biol. Cell 28, 848-857 (2017).
[0377] 9. Kaishima, M., Ishii, J., Matsuno, T., Fukuda, N. 8i Kondo, A. Expression of varied GFPs in Saccharomyces cerevisiae: codon optimization yields stronger than expected expression and fluorescence intensity. Sci. Rep. 6, 35932 (2016).
[0378] 10. Chu, D. et al. Translation elongation can control translation initiation on eukaryotic mRNAs. EMBO J. 33, 21-34 (2014).
[0379] 11. Solow, S. P., Sengbusch, J. 8i Laird, M. W. Heterologous protein production from the inducible MET25 promoter in Saccharomyces cerevisiae. Biotechnol. Prog. 21, 617-620 (2005).
[0380] 12. Andersen, J. T. et al. Structure-based mutagenesis reveals the albumin-binding site of the neonatal Fc receptor. Nat. Commun. 3, 610 (2012).
[0381] 13. Finnis, C., Nordeide, P. & McLaughlan, J. IMPROVED PROTEIN EXPRESSION STRAINS, WO 2018 / 234349 Al. (2018).
[0382] 14. Cubillos, F. A. et al. High-resolution mapping of complex traits with a four-parent advanced intercross yeast population. Genetics 195, 1141-1155 (2013). 15. Goldstein, A. L. & McCusker, J. H. Three new dominant drug resistance cassettes for gene disruption in Saccharomyces cerevisiae. Yeast 15, 1541-1553 (1999).
[0383] 16. Evans, L. et al. The production, characterisation and enhanced pharmacokinetics of scFv-albumin fusions expressed in Saccharomyces cerevisiae. Protein Expr. Purif. 73, 113- 124 (2010).
[0384] 17. Schelde, K. K. et al. A new class of recombinant human albumin with multiple surface thiols exhibits stable conjugation and enhanced FcRn binding and blood circulation. J. Biol. Chem. 294, 3735-3743 (2019).
[0385] 18. Ramaiya, P., Finnis, C., McLaughlan, J. 8i Nordeide, P. IMPROVED PROTEIN EXPRESSION STRAINS, WO 2017 / 112847 Al. (2017).
[0386] 19. Miles, C.; Wayne, M. Quantitative Trait Locus (QTL) Analysis. Nat. Educ. 1, 208 (2008).
[0387] 20. Chatalic KL, Veldhoven-Zweistra J, Bolkestein M, Hoeben S, Koning GA, Boerman OC, de Jong M, van Weerden WM. A Novel11 :LIn-Labeled Anti-Prostate-Specific Membrane Antigen Nanobody for Targeted SPECT / CT Imaging of Prostate Cancer. J Nucl Med. 2015 Jul;56(7) : 1094-9 (2015).
[0388] 21. Mackay TF, Stone EA, Ayroles JF (2009) The genetics of quantitative traits: challenges and prospects. Nat Rev Genet 10: 565-577.
[0389] 22. Cubillos, Francisco A, Billi, Eleonora, Zbrgb, Enikb, Parts, Leopold, Fargier, Patrick, Omholt, Stig, Blomberg, Anders, Warringer, Jonas, Louis, Edward J and Liti, Gianni, 2011. Assessing the complex architecture of polygenic traits in diverged yeast populations. Molecular Ecology 20: 1401-1413.
[0390] 23. Leopold Parts, Francisco A. Cubillos, Jonas Warringer, Kanika Jain, Francisco Salinas, Suzannah J. Bumpstead, Mikael Molln, Amin Zia, Jared Simpson, Michael A. Quail, Alan Moses, Edward J. Louis, Richard Durbin and Gianni Liti. 2011. Revealing the genetic structure of a trait by sequencing a population under selection. Genome Research 21: 1131-1138. Gianni Liti and Edward J Louis. 2012. Advances in Quantitative Trait Analysis in Yeast. PLoS Genetics 8(8): el002912
Claims
Claims1. A recombinant or engineered eukaryotic cell exhibiting improved recombinant protein production compared to a corresponding wild-type, progenitor, or common laboratory strain cell, the recombinant or engineered eukaryotic cell comprising a non-naturally occurring combination of alleles associated with improved recombinant protein production, wherein the at least one allele is for a gene selected from a group consisting of: a gene present in Saccharomyces cerevisiae chromosome XII 706198 - 717026 inclusive (SEQ ID No: 11), UBP3 (YER151C), and BUD27 (YFL023W), or a homologue, orthologue or paralogue thereof, wherein the at least one gene is associated with recombinant protein production.
2. A recombinant or engineered eukaryotic cell exhibiting improved recombinant protein production compared to a corresponding wild-type, progenitor, or common laboratory strain cell, the recombinant or engineered eukaryotic cell comprising at least one gene with a non-naturally occurring combination of single nucleotide polymorphisms (SNPs), wherein the at least one gene is selected from a group consisting of: a gene present in Saccharomyces cerevisiae chromosome XII 706198 - 717026 inclusive (SEQ ID No: 11), UBP3 (YER151C), and BUD27(YFL023W), or a homologue, orthologue or paralogue thereof, wherein the at least one gene is associated with recombinant protein production.
3. A recombinant or engineered eukaryotic cell exhibiting improved recombinant protein production compared to a corresponding wild-type, progenitor, or common laboratory strain cell, the recombinant or engineered eukaryotic cell comprising at least one modified gene, wherein the at least one gene is selected from a group consisting of: a gene present in Saccharomyces cerevisiae chromosome XII 706198 - 717026 inclusive (SEQ ID No: 11), UBP3 (YER151C), and BUD27 (YFL023W), or a homologue, orthologue or paralogue thereof, wherein the at least one gene is associated with recombinant protein production.
4. A recombinant or engineered eukaryotic cell according to any one of the preceding claims, wherein the at least one gene is selected from a group consisting of: YLR284C, YLR285W, YLR285C-A, YLR286C, YLR286W-A, YLR287C, YLR287C-A, YLR288C, and YLR289W, or a homologue, orthologue or paralogue thereof.
5. A recombinant or engineered eukaryotic cell according to any one of the preceding claims, wherein the at least one gene is MEC3 (YLR288C) and / or YLR287C, or a homologue, orthologue or paralogue thereof.
6. A recombinant or engineered eukaryotic cell according to any one of claims 1 to 4, wherein the at least one gene is selected from a group consisting of: NNT1 (YLR285W), CTS1 (YLR286C), YLR287C, RPS30A (YLR287C-A), MEC3 (YLR288C) and GUF1 (YLR289W), or a homologue, orthologue or paralogue thereof.
7. A recombinant or engineered eukaryotic cell according to any one of claims1 to 4 or 6, wherein the at least one gene is CTS1 (YLR286C) and / or RPS30A (YLR287C-A), or a homologue, orthologue or paralogue thereof.
8. A recombinant or engineered eukaryotic cell according to any one of the preceding claims, wherein the at least one gene is at least two, at least three, or at least four genes selected from a group consisting of: MEC3 (YLR288C), YLR287C, UBP3 (YER151C), and BUD27 (YFL023W), or a homologue, orthologue, or paralogue thereof.
9. A recombinant or engineered eukaryotic cell according to any one of the preceding claims, wherein:(i) when the at least one gene is MEC3, the gene comprises the SNP T1362G;(ii) when the at least one gene is MEC3, the gene comprises the SNP C359T; (iii) when the at least one gene is YLR287C, the gene comprises the SNPG19A;(iv) when the at least one gene is YLR287C, the gene comprises the SNP A1001G;(v) when the at least one gene is YLR287C, the gene comprises the SNP G992A;(vi) when the at least one gene is YLR287C, the gene comprises the SNP T722C;(vii) when the at least one gene is RPS30A, the gene comprises the SNP G148A; (viii) when the at least one gene is CTS1, the gene comprises the SNP T47C;(ix) when the at least one gene is CTS1, the gene comprises the SNP C962G;(x) when the at least one gene is CTS1, the gene comprises the SNP G1570A;(xi) when the at least one gene is NNT1, the gene comprises the SNP G413C; (xii) when the at least one gene is GUF1, the gene comprises the SNPC779T;(xiii) when the at least one gene is BUD27, the gene comprises the SNP A1268G;(xiv) when the at least one gene is BUD27, the gene comprises the SNP T1638G;(xv) when the at least one gene is UBP3, the gene comprises the SNP G770A;(xvi) when the at least one gene is UBP3, the gene comprises the SNP G643A; and / or (xvii) when the at least one gene is UBP3, most preferably the gene comprises the SNP A230C.
10. A recombinant or engineered eukaryotic cell according to any one of the preceding claims, wherein the recombinant or engineered eukaryotic cell comprises: (i) at least one additional non-naturally occurring allele for at least one additional gene; (ii) at least one additional gene with a non-naturally occurring combination of single nucleotide polymorphisms (SNPs); and / or (iii) at least one additional modified gene, wherein the at least one additional gene is selected from a list of genes in Table 1, or a homologue, orthologue or paralogue thereof, wherein the at least one additional gene is associated with recombinant protein production.
11. A recombinant or engineered eukaryotic cell according to claim 10, wherein the at least one additional gene is at least two, at least three, at least four, or at least five genes selected from the list of genes in Table 1, or a homologue, orthologue, or paralogue thereof.
12. A recombinant or engineered eukaryotic cell according to claim 10 or claim 11, wherein the at least one additional gene is at least six, at least seven, at least eight, at least nine or at least ten genes selected from the list of genes in Table 1, or a homologue, orthologue, or paralogue thereof.
13. A recombinant or engineered eukaryotic cell according to any one of claims 10-12, wherein the at least one additional gene is selected from a group consisting of: GTB1, GCD7, SLM1, PIP2, YIL089W, ZTA1, UTP25, KTR7, MFB1, TRM7, PUF4, CHS7, SHQ1, CRF1, PRP6, COG3, DPH1, SPT2, XBP1, PRK1, REB1, RCI37, YBR053C, YHR131C, BMT5, CNM1, COQ11, YIL092W, DFP4, DSE2, FMC1, VMR1, RSO55, MAM33, CRP1, AIR1, MYG1, NDD1, YCK1, ARO9, SDS3, ECM2, CBP2, MOB1, ADRI, PDR1, and BEM2.
14. A recombinant or engineered eukaryotic cell according to any one of claims 10-13, wherein the at least one additional gene is selected from a group consisting of: GTB1, GCD7, SLM1, and PIP2.
15. A recombinant or engineered eukaryotic cell according to any one of claims 10-14, wherein the at least one additional gene is selected from a group consisting of: YIL089W, ZTA1, UTP25, KTR7, MFB1, TRM7, PUF4, CHS7, SHQ1, CRF1, PRP6,COG3, DPH1, SPT2, XBP1, PRK1, and REB1.
16. A recombinant or engineered eukaryotic cell according to any one of claims 10-15, wherein the at least one additional gene is selected from a group consisting of: RCI37, YBR053C, YHR131C, BMT5, CNM1, COQ11, YIL092W, DFP4, DSE2,FMC1, VMR1, RSO55, MAM33, CRP1, AIR1, MYG1, NDD1, YCK1, ARO9, SDS3, ECM2, CBP2, MOB1, ADRI, PDR1, and BEM2.
17. A recombinant or engineered eukaryotic cell according to any one of claims 10-16, wherein :(i) when the at least one additional gene is GTB1, the gene comprises the SNP T955C;(ii) when the at least one additional gene is GTB1, the gene comprises the SNP G1633A; (iii) when the at least one additional gene is GCD7, the gene comprises theSNP A427G; and / or(iv) when the at least one additional gene is PIP2, the gene comprises the SNP T2618C.
18. A recombinant or engineered eukaryotic cell according to any one of claims10-17, wherein :(i) when the at least one additional gene is YIL089W, the gene comprises the SNP T469A;(ii) when the at least one additional gene is UTP25, the gene comprises the SNP G1561A; (iii) when the at least one additional gene is UTP25, the gene comprises theSNP G443A;(iv) when the at least one additional gene is KTR7, the gene comprises the SNP G877C;(v) when the at least one additional gene is KTR7, the gene comprises the SNP C437G;(vi) when the at least one additional gene is KTR7, the gene comprises the SNP A230G;(vii) when the at least one additional gene is MFB1, the gene comprises the SNP C592A; (viii) when the at least one additional gene is TRM7, the gene comprises theSNP C228A;(ix) when the at least one additional gene is PUF4, the gene comprises the SNP C671T;(x) when the at least one additional gene is PUF4, the gene comprises the SNP A2570G;(xi) when the at least one additional gene is SHQ1, the gene comprises the SNP C889T;(xii) when the at least one additional gene is CRF1, the gene comprises the SNP T91C; (xiii) when the at least one additional gene is CRF1, the gene comprises theSNP C422T;(xiv) when the at least one additional gene is CRF1, the gene comprises the SNP C740T;(xv) when the at least one additional gene is CRF1, the gene comprises the SNP A1234G;(xvi) when the at least one additional gene is CRF1, the gene comprises the SNP G1396A;(xvii) when the at least one additional gene is PRP6, the gene comprises the SNP A2229T; (xviii) when the at least one additional gene is PRP6, the gene comprises theSNP T916G;(xix) when the at least one additional gene is PRP6, the gene comprises the SNP C743T;(xx) when the at least one additional gene is PRP6, the gene comprises the SNP A110G; (xxi) when the at least one additional gene is COG3, the gene comprises theSNP G593A;(xxii) when the at least one additional gene is SPT2, the gene comprises the SNP A419T;(xxiii) when the at least one additional gene is PRK1, the gene comprises the SNP A7G; and / or(xxix) when the at least one additional gene is REB1, the gene comprises the SNP G1727A.
19. A recombinant or engineered eukaryotic cell according to any one of claims 10-18, wherein:(i) when the at least one additional gene is RCI37, the gene comprises the SNP C697T;(ii) when the at least one additional gene is YBR053C, the gene comprises the SNP G473A; (iii) when the at least one additional gene is YBR053C, the gene comprises the SNP G218C;(iv) when the at least one additional gene is YHR131C, the gene comprises the SNP G1924A;(v) when the at least one additional gene is YHR131C, the gene comprises the SNP C1633G;(vi) when the at least one additional gene is YHR131C, the gene comprises the SNP A328G;(vii) when the at least one additional gene is BMT5, the gene comprises the SNP A64C; (viii) when the at least one additional gene is CNM1, the gene comprises theSNP G396T;(ix) when the at least one additional gene is CNM1, the gene comprises the SNP G232A;(x) when the at least one additional gene is COQ11, the gene comprises the SNP C82A;(xi) when the at least one additional gene is YIL092W, the gene comprises the SNP G130A;(xii) when the at least one additional gene is YIL092W, the gene comprises the SNP T137C;(xiii) when the at least one additional gene is YIL092W, the gene comprises the SNP A257G; (xiv) when the at least one additional gene is YIL092W, the gene comprises the SNP A778G;(xv) when the at least one additional gene is YIL092W, the gene comprises the SNP G955A;(xvi) when the at least one additional gene is YIL092W, the gene comprises the SNP G1055A;(xvii) when the at least one additional gene is YIL092W, the gene comprises the SNP A1086G;(xviii) when the at least one additional gene is YIL092W, the gene comprises the SNP C1661T; (xix) when the at least one additional gene is YIL092W, the gene comprises the SNP C1730G;(xx) when the at least one additional gene is YIL092W, the gene comprises the SNP G1802T;(xxi) when the at least one additional gene is DSE2, the gene comprises the SNP A71G;(xxii) when the at least one additional gene is FMC1, the gene comprises the SNP A112G;(xxiii) when the at least one additional gene is RSO55, the gene comprises the SNP A52C; (xxiv) when the at least one additional gene is CRP1, the gene comprises theSNP G859;(xxv) when the at least one additional gene is CRP1, the gene comprises the SNP T1043C;(xxvi) when the at least one additional gene is AIR1, the gene comprises the SNP G487A;(xxvii) when the at least one additional gene is AIR1, the gene comprises the SNP A827G;(xxviii) when the at least one additional gene is MYG1, the gene comprises the SNP A439G; (xxix) when the at least one additional gene is ARO9, the gene comprises the SNP C1121G;(xxx) when the at least one additional gene is SDS3, the gene comprises the SNP T627A;(xxxi) when the at least one additional gene is ECM2, the gene comprises the SNP T984A; (xxxii) when the at least one additional gene is ECM2, the gene comprises the SNP G295A;(xxxiii) when the at least one additional gene is CBP2, the gene comprises the SNP A1094C;(xxxiv) when the at least one additional gene is CBP2, the gene comprises the SNP G372T;(xxxv) when the at least one additional gene is ADRI, the gene comprises the SNP T935C;(xxxvi) when the at least one additional gene is ADRI, the gene comprises the SNP C959T; (xxxvii) when the at least one additional gene is ADRI, the gene comprises the SNP C1498A;(xxxviii) when the at least one additional gene is ADRI, the gene comprises the SNP C1691A;(xxxix) when the at least one additional gene is ADRI, the gene comprises the SNP A2282T;(xxxx) when the at least one additional gene is ADRI, the gene comprises the SNP G3157A;(xxxxi) when the at least one additional gene is PDR1, the gene comprises the SNP A280G; (xxxxii) when the at least one additional gene is BEM2, the gene comprises the SNP G5937T;(xxxxiii) when the at least one additional gene is BEM2, the gene comprises the SNP G2584A;(xxxxiv) when the at least one additional gene is BEM2, the gene comprises the SNP C926T; and / or(xxxxv) when the at least one additional gene is BEM2, the gene comprises the SNP C533T.
20. A recombinant or engineered eukaryotic cell according to any preceding claim, wherein the recombinant or engineered eukaryotic strain does not comprise a combination of SNPs associated with poor recombinant protein production.
21. A recombinant or engineered eukaryotic cell according to any preceding claim, wherein the eukaryotic cell is a fungal cell, or a yeast cell, optionally wherein the engineered or recombinant eukaryotic cell has a modified or disrupted PEP4 gene, or a homologue, orthologue or paralogue thereof.
22. A recombinant or engineered eukaryotic cell according to claim 21, wherein the yeast cell is Pichia pastoris, optionally a Komagataella species, or Hansenula polymorpha, Kluyveromyces lactis, a Yarrowia species, or Schizosaccharomyces pombe.
23. A recombinant or engineered eukaryotic cell according to claim 21, wherein the yeast cell is a Saccharomyces species yeast, preferably Saccharomyces cerevisiae.
24. A recombinant or engineered eukaryotic cell according to claim 21, wherein the fungal cell is an Aspergillus species, optionally Aspergillus oryzae or Aspergillus niger, or a Trichoderma species, or Myceliophthora thermophila.
25. A recombinant or engineered eukaryotic cell according to any one of claims 1 to 20, wherein:(i) the eukaryotic cell is an insect cell, optionally a Sf9 or Sf21 cell line from Spodoptera frugiperda cells, Hi-5 from Trichoplusia ni cells, or Schneider 2 cells or Schneider 3 cells from Drosophila melanogaster cells; or(ii) the eukaryotic cell is from Excavata, optionally Leishmania tarentolae.
26. A recombinant or engineered eukaryotic cell according to any one of claims 1 to 20, wherein the eukaryotic cell is a mammalian cell type, optionally a Chinese hamster ovary (CHO) cell, a Mouse myeloma lymphoblastoid, a Human embryonic kidney cell, a Human embryonic retinal cell, or a Human amniocyte cell.
27. A recombinant or engineered eukaryotic cell according to any preceding claim, wherein the recombinant or engineered eukaryotic cell exhibits an increase in recombinant protein production or secretion compared to a wild-type, progenitor, or common laboratory strain cell, of at least 10%, at least 20%, at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, at least 80%, at least 90% and preferably at least 100%.
28. Use of the recombinant or engineered eukaryotic cell according to any one of the preceding claims, for producing a recombinant protein.
29. A method for producing a recombinant protein, the method comprising: (i) transforming the recombinant or engineered eukaryotic cell or obtaining a transformed recombinant or engineered eukaryotic cell according to any one of claims 1 to 27 with an expression vector encoding at least one recombinant protein; and(ii) culturing the cell in a medium under conditions to produce the recombinant protein.
30. A recombinant protein obtained from the recombinant or engineered eukaryotic cell according to any one of claims 1 to 27.