Engineered eukaryotic cell

EP4665746A1Pending Publication Date: 2025-12-24PHENOTYPECA LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
EP2024707273
Authority / Receiving Office
EP · EP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-02-13
Filing Date
2024-02-13
Publication Date
2025-12-24

AI Technical Summary

Technical Problem

Current methods for improving recombinant protein yields in Saccharomyces cerevisiae are sub-optimal, labor-intensive, and costly, with challenges in controlling proteolysis and identifying optimal genetic modifications for specific products, leading to inefficient production and high costs for biopharmaceuticals.

Method used

Development of recombinant or engineered eukaryotic cells with specific combinations of alleles and single nucleotide polymorphisms (SNPs) associated with reduced proteolysis, particularly targeting genes like GAT1, UBP14, and MEC3, to enhance protein production and reduce proteolytic degradation.

Benefits of technology

The engineered cells exhibit reduced proteolysis and improved recombinant protein production, leading to increased yields and cost-effectiveness in biopharmaceutical manufacturing while maintaining the advantages of Saccharomyces cerevisiae as a production host.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure GB2024050388_22082024_PF_FP
    Figure GB2024050388_22082024_PF_FP
Patent Text Reader

Abstract

The invention relates to recombinant or engineered eukaryotic cells, and particularly, although not exclusively, to recombinant or engineered yeast cells, such as Saccharomyces cerevisiae, exhibiting reduced proteolysis. The invention also extends to the use of recombinant or engineered eukaryotic cells for producing recombinant proteins, as well as the recombinant protein products obtained from such recombinant or engineered eukaryotic cells.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Engineered Eukaryotic Cell

[0002] The invention relates to recombinant or engineered eukaryotic cells, and particularly, although not exclusively, to recombinant or engineered yeast cells, such as Saccharomyces cerevisiae, exhibiting reduced proteolysis. The invention also extends to the use of recombinant or engineered eukaryotic cells for producing recombinant proteins, as well as the recombinant protein products obtained from such recombinant or engineered eukaryotic cells.

[0003] For over 40 years, microorganisms have been used to produce recombinant proteins. However, the cost of producing high-quality, biologically compatible, and commercially relevant quantities is often prohibitive. The baker's yeast, Saccharomyces cerevisiae, is well-known for producing high-quality, correctly folded heterologous proteins, often at lower costs and at levels greater than mammalian cells. Alternative yeasts, such as Komagataella species, K. phaffii, K. pastoris, and K. pseudopastoris, including industrial strains also known as Pichia pastoris, are also used to make recombinant products, such as biopharmaceutical proteins at competitive costs of goods. However, Pichia pastoris, is often associated with a requirement for methanol and oxygen for large-scale manufacturing, nonexchangeable plasmid systems requiring selection with toxic zeocin, and posttranslational quality issues for biopharmaceuticals. While these disadvantages are not generally associated with Saccharomyces cerevisiae, S. cerevisiae tends to have lower productivity. Therefore, it would be advantageous to increase the yields of recombinant proteins from S. cerevisiae, which will act to reduce the cost of goods sold (COGS) during large-scale manufacture, whilst retaining all the positive features of this expression host.

[0004] Alternative eukaryotic production hosts, such as mammalian cells, tend to be significantly more expensive for manufacturing biopharmaceuticals than yeasts. The media and growth requirements are also more expensive for mammalian cells, and processes are not free from animal-derived components.

[0005] Alternative microbial production hosts, including prokaryotic / bacterial systems like E. coli, lack eukaryotic cellular machinery, e.g., for protein folding and secretion, frequently resulting in insoluble inclusion bodies and endotoxin contamination issues affecting downstream purification. Previous work to improve product yields from S. cerevisiae has succeeded in developing manufacturing processes for a range of biopharmaceuticals, such as insulins, vaccines (e.g. virus-like particles), albumin and albumin fusion proteins. However, the production hosts are often sub-optimal, and the cost of goods sold limits access to biopharmaceutical products for all those who need them, thereby hindering the eradication of treatable human diseases. The development of improved production yeast, especially for S. cerevisiae, which has a long, safe history of human use and is most commonly used for the manufacture of yeast- derived biologies approved by the FDA, would be advantageous. Improving product yields from S. cerevisiae would be especially beneficial.

[0006] Existing methods for improving product yields include optimising the expression construct, e.g., the promoter, leader sequence (for secretory products), the coding sequence, terminator sequences, and the copy number and stability of the expression construct. However, improvements to the yeast genome can also lead to substantial improvements in productivity, product quality and other valuable bioprocess phenotypes. Nevertheless, the existing methods for improving the genome of the production yeast are sub-optimal, slow, expensive and labour- intensive. These methods have also focused on engineering the genome of a single strain or a family of closely related strains derived from commonly available laboratory strains, e.g. CEN.PK or relatives of W303. Therefore, the optimal starting strain was not used. The underlying biology is complex and not fully understood, especially the bottlenecks and limitations in metabolism affecting product yields and quality. Each protein product has different requirements, and many different limitations or bottlenecks can be encountered during production strain development. A production host improved for one product is seldom optimal for all products.

[0007] Furthermore, methods for improvement have so far explored only a relatively small number of genetic changes. For example, random mutagenesis approaches tend to make limited types of genetic changes and frequently result in loss of function. There is also a risk of introducing unwanted mutations. Unless the beneficial mutation is then identified and re-engineered into a clean genetic background of the progenitor strain, there is a tendency for the accumulation of these undesirable mutations. Therefore, multiple rounds of random mutagenesis tend to be limited in the scope of improvements possible and can generate adverse phenotypes. On the other hand, rational genome engineering is unlikely to identify all options for strain improvement, some of which are not obvious targets for improvement, e.g. UBC4, M0T2, GHS1. The engineering of the yeast secretion system is also complicated by the involvement of many cross-reacting factors. The tight interdependence of each of these factors makes genetic modification difficult. Many attempts in strain engineering also fail to be transferrable from the laboratory to an industrial process. Each different recombinant protein product presents different challenges for production strain optimisation and generates a different burden on host cell metabolism. Therefore, the changes made to a single yeast strain improved for one product are unlikely to be optimal and may even be detrimental to the production of a different product. Consequently, bespoke production strain improvement is frequently required for each new product, which can also be sub-optimal, slow, expensive and labour-intensive. For example, each recombinant product might require an optimal combination of chaperone proteins for maximal secretion of the correctly folded protein. This might require multiple chaperones to be overexpressed, each requiring their expression level to be fine-tuned, thereby requiring the generation of multiple different strains and slow, expensive testing in fermenters.

[0008] A particular problem exists in controlling the proteolysis of the final recombinant product. This undesirable proteolysis can occur at any point in the secretory pathway within the cell after the polypeptide chain has been synthesised by the ribosome, in the extracellular media before harvesting, or even during downstream processing and purification, for example, where a particular protease is brought into contact with the product protein under conditions allowing for proteolysis. Furthermore, it can be especially difficult to remove a "product-related impurity", such as a proteolytic fragment of the desired protein from its full-length form, because the proteolytic fragment will retain some of the physiochemical properties generally exploited during downstream processing to efficiently obtain the desired full-length product. Therefore, even a small amount of proteolysis affecting the final product can have significant cost implications in its removal, not to mention the fact that some of the product has been lost due to proteolytic degradation reducing overall yields. Therefore, it is common practice to disrupt non-essential genes in the production strain which are known to be responsible for product proteolysis. However, this can result in undesirable phenotypes, such as poor growth of the final production strain. Furthermore, the expression of proteases by S. cerevisiae is extremely complex, with each recombinant protein product being affected differently by the nearly 200 diverse peptidases reportedly encoded in its genome. Consequently, it can be extremely difficult to identify which protease(s) is / are degrading a particular protein product. This is especially true when the recombinant product is a substrate for multiple S. cerevisiae proteases because even if the recombinant protein is expressed in a strain disrupted for one of the proteases which degrade it, it will still be degraded by the other(s), and so it is difficult to know if an improvement has been made. Additionally, a limit is quickly reached for the number of protease genes which can be disrupted in any S. cerevisiae production host before it becomes unsuitable for industrial manufacturing due to adverse phenotypes introduced. Therefore, the strategy of identifying and disrupting proteases to solve this problem is both laborious and extremely limited. It is impractical to have strains disrupted for all combinations of S. cerevisiae proteases. This means that the manufacture of desirable biopharmaceutical products is often not commercially viable, despite other obvious benefits of this well-established yeast, and medical innovations may fail to reach the market.

[0009] It would, therefore, be highly advantageous to have a method of controlling proteolysis, without these difficulties in identifying the culprit proteases or the adverse effect of multiple gene disruptions. Improved strains for recombinant protein manufacture that exhibit reduced proteolysis, which are genetically diverse to common laboratory strains, would also be valuable for investigating the manufacture of products that are difficult to express in existing systems. Furthermore, it would be beneficial to be able to identify the key regions of the yeast genome responsible for reducing proteolysis, thus providing improved strains for recombinant protein manufacture with improved product yields.

[0010] Accordingly, in a first aspect of the invention, there is provided a recombinant or engineered eukaryotic cell exhibiting reduced proteolysis compared to a corresponding wild-type, progenitor, or common laboratory strain cell, the recombinant or engineered eukaryotic cell comprising a non-naturally occurring combination of alleles associated with reduced proteolysis, wherein the at least one allele is for a gene selected from a group consisting of: GAT1 (YFL021W), UBP14 (YBR058C), and a gene present in Saccharomyces cerevisiae chromosome XII 706198 - 717026 inclusive (SEQ ID No: 11), or a homologue, orthologue or paralogue thereof, wherein the at least one gene is associated with proteolysis. In a second aspect of the invention, there is provided a recombinant or engineered eukaryotic cell exhibiting reduced proteolysis compared to a corresponding wildtype, progenitor, or common laboratory strain cell, the recombinant or engineered eukaryotic cell comprising at least one gene with a non-naturally occurring combination of single nucleotide polymorphisms (SNPs), wherein the at least one gene is selected from a group consisting of: GAT1 (YFL021W), UBP14 (YBR058C), and a gene present in Saccharomyces cerevisiae chromosome XII 706198 - 717026 inclusive (SEQ ID No: 11), or a homologue, orthologue or paralogue thereof, wherein the at least one gene is associated with proteolysis. As discussed in the Examples, using quantitative trait loci (QTL) analysis, the inventors have identified 16 genomic regions comprising approximately 3.3% of the total Saccharomyces cerevisiae genome containing genes, and different alleles of genes, responsible for the differential expression of the recombinant amylase- mCherry protein (see Table 1). In particular, the inventors have identified GAT1 (YFL021W), UBP14 (YBR058C), and a gene present in Saccharomyces cerevisiae chromosome XII 706198 - 717026 inclusive (SEQ ID No: 11), e.g. MEC3 (YLR288C), and YLR287C, as key genes. Advantageously, by identifying these genomic regions, and more specifically the SNPs within the genes, improved strains exhibiting reduced proteolysis of the recombinant protein product can be provided, resulting in improvements in recombinant protein production.

[0011] From this QTL analysis, the inventors were able to identify genes in the QTLs that are associated with higher or lower levels of proteolysis. As such, the inventors believe that by modifying the expression of such genes (e.g. through knock-outs, overexpression, protein engineering or other methods to modify the genome or the expression levels of the protein), they can provide further improved strains exhibiting reduced proteolysis.

[0012] Accordingly, in a third aspect of the invention, there is provided a recombinant or engineered eukaryotic cell exhibiting reduced proteolysis compared to a corresponding wild-type, progenitor, or common laboratory strain cell, the recombinant or engineered eukaryotic cell comprising at least one modified gene, wherein the at least one gene is selected from a group consisting of: GAT1 (YFL021W), UBP14 (YBR058C), and a gene present in Saccharomyces cerevisiae chromosome XII 706198 - 717026 inclusive (SEQ ID No: 11), or a homologue, orthologue or paralogue thereof, wherein the at least one gene is associated with proteolysis.

[0013] In one embodiment, the recombinant or engineered eukaryotic cell comprises a non-naturally occurring combination of alleles associated with reduced proteolysis, wherein the at least one allele is for gene selected from a group consisting of: GAT1 (YFL021W), UBP14 (YBR058C), and a gene present in Saccharomyces cerevisiae chromosome XII 706198 - 717026 inclusive (SEQ ID No: 11), or a homologue, orthologue or paralogue thereof, and the recombinant or engineered eukaryotic cell comprises at least one modified gene, wherein the at least one gene is selected from a group consisting of: GAT1 (YFL021W), UBP14 (YBR058C), and a gene present in Saccharomyces cerevisiae chromosome XII 706198 - 717026 inclusive (SEQ ID No: 11), or a homologue, orthologue or paralogue thereof.

[0014] In another embodiment, the recombinant or engineered eukaryotic cell comprises at least one gene with a non-naturally occurring combination of single nucleotide polymorphisms (SNPs), and at least one modified gene, wherein the at least one gene is selected from a group consisting of: GAT1 (YFL021W), UBP14 (YBR058C), and a gene present in Saccharomyces cerevisiae chromosome XII 706198 - 717026 inclusive (SEQ ID No: 11), or a homologue, orthologue or paralogue thereof. In one embodiment, there is provided a recombinant or engineered eukaryotic cell exhibiting reduced proteolysis compared to a corresponding wild-type, progenitor, or common laboratory strain cell, the recombinant or engineered eukaryotic cell comprising at least one gene with a non-naturally occurring combination of single nucleotide polymorphisms (SNPs) associated with reduced proteolysis, wherein the at least one gene is present within SEQ ID No: 11, or a homologue, orthologue or paralogue thereof.

[0015] In another embodiment, there is provided a recombinant or engineered eukaryotic cell exhibiting reduced proteolysis compared to a corresponding wild-type, progenitor, or common laboratory strain cell, the recombinant or engineered eukaryotic cell comprising at least one gene with a non-naturally occurring combination of SNPs associated with reduced proteolysis, wherein the at least one gene is selected from Table 7 and is present within SEQ ID No: 11 or a homologue, orthologue or paralogue thereof. In one embodiment, the non-naturally occurring combination of SNPs associated with reduced proteolysis is selected from Table 7.

[0016] One embodiment of the nucleotide sequence of Saccharomyces cerevisiae chromosome XII 706198 - 717026 inclusive (from the S288c reference genome), is provided herein as SEQ ID No: 11, as follows:

[0017] TCATAAACGATGCTTCCTTTGTTTCGAGCCCAGCTGCCTAAATCTTTTTAAGGGCTCTCCATCTACCCAATACTTGAG AGATTCATTTACTTCCACTGAGTTAGCCTTATTGAATGCATCGATGTGGTTCGATTTCAGCAATTTTTTCATCCCTAA GCAACTGGGCAGGTATAGCCCTTTCACTTTCTCCCTCAATTCTTCTAAGACCTTTGCATTGAACGCTTCAGCGTTTGA AGATGGCATGTTAAAATTCTTGCTTATAAATCCGTTCTCGCACATTATATCGTACTTGAATGGTTTGTTGAACATGAG GCATTCATACGTCGTATTTGTGCCAAACTTCAATGGCAAAGAGACCGTTGTACCACCTTCGGTAATTAGTCCTAAGTT AGCAAAGGGGTATAGCAAATAAACCTTGTCATTTATACTGTACACAATGTCACATAACGCTACCAGTGCCGCGCTCAA CCCTATTGCTGGTCCATTCAAACAGCAAATTAAAACTTTGGAATGCTTGATGAAGGCATCAGTGACATAAACATTTCT AGCGACAAAATTTGACACCCACTTGCTTGTTTCCGAAGGATATTTATTGGTATCATCCCCTTGGGCTTTTGCAATACC CTTGAAATCAGCACCACTGGAAAAAAATCTACCACTGCTTTGTATAATTGTAAAATATACATCACGATTTCTGTCCGC TAGTTCTAGTAACTCTCCTAAATAAATATAGTCTTCACCTTCTAGTGCATTCAAATTGTCAGGGTTCATTAAGTGAAT AATGAAGAATGGTCCTTCAATACGATAACTGATTTTCTCATTTTGCCTAATTTCTTGCGACATATTGTTCTTCCTTCC TTTACTGTGCGAGCAATTTGTGCCATACACTCCAATTATCTATCCATATGTTCATCTTTCAGTTCTTATATATAGTGC TTCACTTTTATTAACCGCTCTTGGGGTAAAAGAGAAACGGTCACTGCTACAGCCTTTCGGAATAATATTTCTTATGCC CGTCGCCAATTTCAGATTCTTTCCTTCGTGCCGCATTTATTTTATTTTTCAAAATGACGAGAAGTTAAAATCTAAAAT CAGGAAAAAGAACAGTTGCAGATGACGAAATACTTCGAGCATTATCGCAATAAATAAGGCGTGTAAACGTATGTCAGA TATAGAATCATTGGGAGAGGCAGCAGGTTTATTTGAGGAGCCAGAGGATTTTCTTCCTCCTCCACCAAAACCTCATTT TGCAGAATATCAAAGGTCACACATCACAAAAGAGTCCAAATCAGATGTTAAAGACATCAAACTCCGTCTCGTTGGAAC GTCACCACTTTGGGGTCATCTCTTGTGGAATGCAGGAATATACACCGCAAATCACTTGGATTCTCATCCAGAATTAAT AAAGGGGAAGACTGTTTTAGAATTGGGTGCTGCTGCTGCTTTACCTTCCGTTATTTGTGCTTTGAATGGGGCTCAAAT GGTTGTTTCAACTGATTATCCAGATCCTGATTTAATGCAGAACATCGATTATAATATAAAGTCTAACGTTCCTGAAGA TTTTAATAATGTCAGTACGGAGGGTTATATTTGGGGAAACGATTATTCTCCATTGTTGGCACATATCGAAAAAATAGG TAACAATAATGGAAAGTTTGACTTAATCATTTTAAGTGATTTAGTCTTCAATCATACGGAACATCACAAATTACTTCA AACAACAAAGGATTTATTAGCTGAGAAAGGTCAGGCGTTGGTAGTATTTTCACCGCATAGACCAAAATTGTTGGAGAA GGATTTAGAATTCTTCGAATTGGCTAAGAACGAGTTCCATTTGGTTCCTCAGCTAATTGAAATGGTTAATTGGAAACC GATGTTTGACGAAGACGAGGAAACAATCGAAGTCAGATCTCGTGTTTATGCGTACTATTTAACACATGAAAAGTAGTA AGTGGGCGGAATTGGCATTCTCAACAATCATTTTTATGTTTGTAAACCAGTGTCCTTTGATATGTAACAATTAAATAA AATAAGTCATGATGGGGCACGAATTCGCAAAATGGTAGAAACACAGTAAATCGAAGAAAATTAAATAATTCGTTATAG ATATAAAGTAGTCGATTATAAAAACGTAAAAACATGAAGGCAGGGTACCTTGACGAAATAAAAGAGAGGCGGTGAGAT GAGGTGAAAAATAAAGAATATTGCAAAAAAAAATATTCGGACTCTATGAATCAATCTCTAATAACTTTAAAAGTAATT GCTTTCCAAATAAGAGAAATTACATTGGGTATAGACTGAGTCGCCGGAGTCATAAGCATAACATGTGGTTCCAGAAGC ACATTCCATGTAAACCCAAGCGCTATGATCACAAACGGCGAACTTCCCATCAGCAGAGCATGCAATTTCACCTTCAGT ACAGGTAGATTTACCGTTCAATTTACCAGCCGCATATTGAGCATTCAATTCTTTAGCCAATGTACGAGCTGTACTGTC TGAGCTTGTACTACCTGAGCTTGTACTACCTGAGCTTGTACTGGTCGTTGTTGGGGAAAGCGTACTAGTAGTAGCTGT CTGTAGGGAAACGACAGAAGAACTCTTCGTTGCTGGCGAAAGAGTACTAGTGATAGCTGTTTGAATTGGGGCCGAAGA AACTATGCTAGTAGTAGTTGTTTGAGGGGTCACCAAGGCAGCACTGGTTATTTGGGAAGATAGAGTAGTTTTCATACT TGTGATAGCAACTGAATTTAAAGTGCTCTCTGTTGTGGTGGTACCTAGACTAGATTTTGTCTTGGTGCTACTCGTCAA TGTTTTTGTAGTTTGAGTAATTGATGTTTTGATAGCGCTGCTTGCAGTTGGAGATAAAGTAACTTTGCTTTTACTTTG TGTAGATGTCGTAGATTGTGTGGTCTTTTTCTGAGAAGTTGAAGCAGATGAAGTTGAAGCAGATGAAGTTGAGGCTGC TGAGGTTTTTGAGGTGGCAACTGTAGTAGTGGCGGTCTGGCTAGCACTTGTTAGCAAATTCTTCAAAATCTCAACATA TGGTTCACCATTTAGCTCGTTGGAAAAGGCTTGAGATGCATCCCATAACGCAATACCACCAAAAGAACTTGAAGAGGC AATATCTGCAATAGTTGATTCCAATAAAGAAGTGTCAGAAATATAACCAGAGCCAGCAGCAGAAGCAGAACCAGGTAA ACCTAAGAACAGTTTGATATTTTTATTTGGGGATACAGTTTGAGCATAGGTTAACCAAGTATCCCAATTGAATTGACC ACTCACACTGCAGTAATTATTGTAAAATTGGATGAACGCAAAATCAATGTCTGCATTTTCCAACAAGTCACCAACAGA AGCATCCGGGTATGGACATTGTGGTGCGGCAGAAAGGTAATATTGCTTTGTACCTTCGGCAAACAAAGTTCTTAACTT GGTAGCTAACGCACTATAGCCTACTTCGTTGTTGTTTTCAATATCAAAATCAAAACCATCAACGACTGCTGAGTCAAA TGGTCTCTCACTGGCACCTGTACCTTCACCGAAAGTATCCCATAAAGTTTGTGCAAAAGTTTCCGCTTGAGAATCATC TGAAAAGAGGTAGCTACCAGATGCACCACCTAATGATAATAGAACTTTCTTTCCTAGGGACTGGCAAGTTTCAATATC TTCAGCAATCTGGGTGCAGTGAAGTAAGCCATCAGAAAAAGTATCAGAGCATGCGTTGGCAAAGTTCAAACCAAGGGT TGGAAATTGGTTCAAGAAAGATAATAGGAAAATATCAGCATCAGAAGATTCACAGTAAGTAGCTAAGGATTCTTGCGT TCCTGCTGAGTTTTGACCCCAATAAACAGCAATATTTGTGTTAGCAGACCTATCAAAGGCATCGGTTGGCAGTAGTAA GAATTGTGTGAATAGAAGAATGATGTAAAGGAGTGACATTCTATTATTAATTATTTTATATTTAAATTAGAATTTCAA TGTATTGGAAAAAGAGTGGTTTTAAAAAAGGTAGGTTGTGAAACGAGCGACAGTTATTATTTTTGGATATAAAAAGGT TCAAAGGAATGACAGGTTTATTTTTTTTATGTTATTGAGTCTTAAATGAGTGAAAGATCTTCAATGTTATGAAAAAAC TTGACTGATATAATCTTAATGTAATGTATTGTTGTTGTCCATTCTATAAATTATACAAGGAAAGGTAAGTGAGACGTT

[0018] TTTCCTCCATCCAAGGACCTAATACCTGCATCTTTTTATATCTTTGACCAATGCCTATGAAGCCAAATACTTCCGTCC

[0019] ATATTAGTTTTTTTTTGTTCCTTGAATCCTGCACCGAACCATCGAGGGGCTTAGTTCTTACACCAATGCTAATGTTTA

[0020] CTCAACATCTGAAACACATAAACAAATAAACAAACACCAAACATTGCACGTTAGAAACTCATATTTACTCGCACATAT

[0021] CTGTTAAGGCGAGGCTGGTTATTTTTTCAAGGGACCAGCATTTCCAATTTTTCACCAGCGGCAAAGAAAAATACCAAC

[0022] AAGCAGGAATCACACCCACGCATCTCATTTTGCGTTACTATGACAACTACAAAATTTGAATTTGTTAACGAGAATGAC

[0023] GTATATAAGTTCGATTTTGTTGAGGATATCCCACATATTCCCACCAACCAACCACTATATTAATGGACATGGTTTTTG

[0024] CCGTGACCGTGGATCTCGCAGCTACACAATGATATCTCTTGTATCCGTTTTAGCATGAATTTCTAAACAAGACTGATT

[0025] ATATACCATGCACTGTAACATTCAGAGAAAATATACTAAAACAATTCATTGAATTTTATCTCCCATGTGTCAATCCAT

[0026] TTCTTTCTTTTTGGATCACCTTCATATAGTTTACTCATTTGGATAAGAACTTTTTTAGTTACTGCAACTAATGCAGCC

[0027] TGTTCTTCTTTTACCTCCTCATTGGTAAAGTTCTTCGGCTCAAATTGTACAGTGGATACCACTTCATCCAAAAGTAAC

[0028] TGTATTTCCTTCAGATATGTATGTATACTGTCTAAAGTTTCTGCTTGATTCCTCTTAGGAGTAAAATCTTTTGACACC

[0029] AGAGTTTTCCTGAAAGTTGATACAAGTAATTTGATTAGCTTTATCTTCCTCGTAAACCCATCGAAAAACAACTTTATA

[0030] TTTTCATAGACCTTTTCCTGGTCAAATTCCTCTTTCTGTGAATCAGATTCGGATTCTGAATCCTCAAAATCCAAAAAT

[0031] ATATCGTCCGAGTTTGCTGAAAAATCGGGTTCCTCAAGCCACTCTTTTATTTCATTCATTGTATCGTCCATTATGGCC

[0032] ACATTATCTTTTAAGATGTTTGCTAGGATCCCATAAGGTCCAGCCTTGGAACAATTAGAGAGAGAATCGCACGCATTG

[0033] AATATTTTACCAACTGAGGTTAACCTTTCTTTGTCTAGAGATGCATTTTCGTCATTTTTCAATCTTTCTTGTAATTCG

[0034] GCAATGAAATCTCTCAAGCCATCGAGCAATTGTAGCGTACTTTCATCCAGTTGGTCTGTAAAGTATTTGGGGCAATCC

[0035] TTGTTATTGTAAAACAAAGGGAATAAACTTAGTAAGTAGAATAATGGCCTACTAAAATTTTGGATTTCGGTAATCACA

[0036] ACTTTATGATTGTTATCAAAAGTTCCTGGCTTGCAAACAATACCTATTTTTGTGCAATGTGCCTTCAATACAGAGGCT

[0037] AATTTGTCTAGTTCCTTAGTTGGGGTGGAACCTTGTAGTTTGGTAGTAGAGGAGATTTTACGCAAGTCTTCCGGCTTT

[0038] TTGTATGGCACTAAAAACTGTTCATCTATAGAATTCAGCAGTTCCAATAATTTGACGTCATCCTTTCTATCACTGCTT

[0039] CCAGTAGACATTATTGTATAACTGGTATTTCTCTTATAAAGTTTATAATAGGCGTCTCTGGTGATGCTTGTTGACCTC

[0040] AATGGATTTTGTTAAAAGGAGGCTTCATGTCATCGTCAAGATCATTGCAGTTTTTTTCTAGTTTTTTTTTTTTCATTG

[0041] CCACATTATGAAGCGCCCATTCAAAAACAACTAGAAAAACGTTTGCCAACCCTAGGTACGGCAAGATCCTCCTTATTA

[0042] ATAAAAGCATCACACCAAGACAGCAAAGAAGTAAATCGATAATATGAATAGCATAATTTTGCGAAGTTAAAGTCCTAA

[0043] AGAAATAATTTATAGTTAACACTACTTACATAATTTCAATACTGCTTTATATTTAAAGTAAAAAACGCATAACGGAAA

[0044] TACAAATACAAGTCAAAAACTATCTTGAAGTAAAGATGGCATGACTGCTATTTTGTTTCTTGTTACTTTTAAGCTCCT

[0045] TTTCCACAACTGTTAATTTTCTTATTGGACGGATGGACCTGGGTTCATTCTTCTCTTACCGTTAACCAAGGTAACGTT

[0046] AACGAATCTTCTGGTGTACAATAATCTTTTGTAAGCACGACCCTTAGGCTTCTTTGGCTTTTCGGTTTTTTCAACCTT

[0047] TGGTGTTTGAGACTTGACTTTACCAGCACGAGCTAGAGAACCATGAACTTTAGCCTAAATAAATAATATTAATATTAA

[0048] AAATTAATGACCATGTTAGTATAGTGAATTTTTGAATTATTTGGCGTATCAGCAAGGATGCTAAAAATAAACTTCACT

[0049] CTGCTATTTTCTCAAAGCCTTTTTAACTCTCCGTCAAAAAATGTTGTTATTAAACTCGACATTATGTTAAGTCCACGC

[0050] TGTTGCCCTCGAAAAATCTGTGTCAAGCAATTACGTTTTTGATGTAACAGAACTCAAATTTAAAAAGTTTCCTATCAT

[0051] ATTCTCAGGGTGGAGCCAGTTTAAAGATATATTGCGTTTTATTTCTTTTTCTGCGGTATATATAGATATGAAACACCA

[0052] GTAATATATTCAAATTAGCTATTTCTTGTTATTCATTCCTCGCTGAAACGTCCATTTTTTTCATATTTTAAAATCCTA

[0053] ATGAGGCATTACGTACCATATTTGCGTAGTTTTTGTATGGGGATATGGAAGATTATTACAGTGGCAAATTTAATGTGC

[0054] AGTTCATGTAGCCTTTAAAATGCAACATACCGTTCGTAAGGCTGTTTTGTTCTATCTATTTTTTCTGAGGTTTCCGAC

[0055] AACCTCACGACATTTTTTTTTTTCGAACGGAAAAAAGAAAAGCAGTTAGTATGTAAAGCACGGATTTTTTTATACATA

[0056] TGTACGGATGTGATAGTCTGCAACTCTGCAACTCTGCAGTTCTCAGCCATGTCAGATTTAAAGAGTTTCCTTAGCAAC

[0057] GTAGCAAAGAAATGTACCGCTGTAGGGTTTACAAGCCCTTCGATCTTGCTATATAATATATGATTTGTCCTCTTTCCC

[0058] TTGGTTTTTCCACATCTTCGCTATCTTCCAAGCTACCACGATCCAACGAACAGTGGAATACGCAGGATTCGTCGTGAG

[0059] AAATCGCCAAAACAACTTCTTCAAATGCAGCGTATAACTTGGAACACACCTTCCAATCTTTGCAACGGATGATTACTT

[0060] CATGTGTGGACGAACTTTCCTGTTCAGCCTTTTCCACCATAACGGATATGTCATTAAATTCAGTATCACCGCTAGTAT

[0061] CAGCTGTGTAAATGTTTCCCCGCGTATCTGCGATCGAGCTATCCTCAATTCTTAATAAATCTTCATCGTAGCGGATAT

[0062] TTTCTTCCATCTCTCGATCTCTAGTATTGGTATATAGTGAAGACATCGGTTTATCCGCTTCGATAATCGGAAGAGATC

[0063] CTTCCTCCTGCCGGCCGTCTGTGTCGATGTGCTGGTTTTGGGAAGGATTGTCAGTGAGCCCTTCTTGGCGTTGTATCA

[0064] CAGAGTCTAAGGGTCCATTCCAGCATATTTCCAAATGCCAATCTAATTCATTCACAATTATCTTAAGTTCTACATCAT

[0065] CACCTTCATTTCCATGCTCCTTTTTTTTGACTCCCATTAAATGAATGTGGTTGACATTACTGTACCGTTCAACACGTC

[0066] TAATAAACCCGTGGAAGGCGGAGCCAAACTCGCCCGATATTGGTGGCAGCTTGTACATCATCAGTTGAATATAGTTAA

[0067] TCATTGGCTCTTGTATTCGCGTATCTTGTGCTCGGAATAATAGTTTGACAGGTACTTTGAACGAATGCATGATAACCT

[0068] TATTGCTTGCGAGTAGATTTCCTGTGCCTACTGTGGTTGGTAAACCATTGTGCTCATCCACTCCCCCATTCATAACTA

[0069] TAGCGTCATTTGGGCCGCTAGTATGTACAATTTCTTCGAATGTTATACCTAGAGCACAAATGGGGTTCGGTTTTGAGG

[0070] TTGTATCTACGCCTCCTGCTGTTCCTCCTGAAAGTGTTCCGTTATTTGTATTCCATTCTGGCATTGACTGTAGTTTTA

[0071] TAGTCATATTAGAGGAAGATCCTTGGTTCATTACTCGATCATACCTTTTAAACACACTCAATAAAGAATCACAATTAC

[0072] ACTCCATTGTTATTGTATTGAGCTCGCGAGCTGATATAACTGTATATAATCTGAATACATCATGAGGAATGGTACACC

[0073] AAAGCTGACCAGTATCCCCTCGTAATATTGTACCGTTGTTACTGCTGTTGAGTGATGATTTTGGAGTGGATATTATTG

[0074] TCAATCTTTCACTATTAAATCTTAAGATAGCCGTCTTTCGTAGCGAAGCAACTGTATTGATAGTAGTTCTTAGCAATT

[0075] TATAATCATCAGGTGCTTCACAACCATTTACTATCAATTTTAATTTCATTTAACTGAATTAAGACACACCTTTTGTCT

[0076] TCTTTTTTCTCTCATCATCTCCGTATGTTTATGTTGCTATTTTGATGTAAATAAAAAAGTTGAATAATAGACGAGGGC

[0077] AAGTATAACTCGCCTATATTTGTAGCCGCAACCATTGAAAAAGAGCCATGAATATGGGAAACTAGTTGCACATAAAAA

[0078] TGCTGAAATTTAGAATTAGGCCAGTGAGACATATACGGTGTTATAAACGACACGCATATTTCTTACGATATAACCATA

[0079] CGACTACCCCTGCACAGAAGTTACAAGCACAGATCGAGCAAATACCTCTCGAAAATTACAGAAATTTCTCTATAGTTG

[0080] CCCATGTTGACCATGGGAAGTCAACCTTAAGTGACAGACTGCTGGAAATAACGCATGTCATCGATCCCAATGCGAGAA

[0081] ATAAACAAGTTTTGGATAAATTGGAAGTCGAAAGAGAAAGAGGTATTACTATAAAGGCGCAAACATGTTCGATGTTTT ATAAAGATAAGAGGACCGGAAAAAACTATCTTTTACATTTAATTGACACGCCAGGACATGTGGACTTCAGAGGTGAAG

[0082] TTTCACGGTCATATGCGTCTTGTGGGGGAGCAATTCTTTTGGTTGATGCATCACAAGGCATACAAGCACAGACGGTTG CTAATTTTTATTTAGCCTTCAGTTTAGGATTGAAATTAATTCCAGTAATAAACAAAATTGATCTAAATTTTACAGATG TTAAACAGGTAAAGGATCAGATAGTGAATAACTTTGAGCTCCCCGAGGAAGATATAATCGGAGTAAGTGCTAAAACAG GATTAAATGTAGAGGAACTGTTACTACCGGCTATAATTGATCGTATACCACCACCAACGGGGAGGCCTGATAAACCCT

[0083] TCAGAGCATTATTAGTGGATTCTTGGTACGACGCATACTTAGGAGCGGTTCTTCTAGTGAATATTGTTGATGGTTCTG

[0084] TACGTAAAAATGACAAGGTTATTTGTGCTCAGACAAAAGAAAAATACGAAGTCAAAGATATTGGAATCATGTATCCTG

[0085] ACAGAACCTCTACAGGTACGCTAAAGACAGGACAGGTTGGCTATCTAGTGCTGGGAATGAAGGATTCTAAAGAAGCAA

[0086] AAATTGGAGATACTATAATGCATTTAAGTAAAGTAAATGAAACGGAAGTACTTCCCGGATTTGAAGAACAAAAACCCA TGGTATTTGTGGGTGCTTTCCCGGCTGATGGGATTGAATTCAAAGCCATGGATGATGATATGAGTAGACTTGTCCTCA

[0087] ACGATAGGTCAGTTACTTTGGAACGTGAGACCTCCAATGCTTTGGGTCAAGGTTGGAGATTGGGCTTTTTAGGATCTT

[0088] TACATGCATCTGTTTTTCGTGAACGACTAGAAAAAGAGTATGGTTCGAAATTGATCATTACTCAACCCACAGTTCCTT

[0089] ATTTGGTGGAGTTTACCGATGGTAAGAAAAAGCTTATAACAAATCCGGATGAGTTTCCAGACGGAGCAACAAAGAGGG

[0090] TGAACGTTGCTGCTTTCCATGAACCGTTTATAGAGGCAGTTATGACATTGCCCCAGGAATATTTAGGTAGTGTCATAC GCTTATGCGATAGTAATAGAGGAGAACAAATTGATATAACATACCTAAACACCAATGGACAAGTGATGTTAAAATATT

[0091] ACCTTCCGCTATCGCATCTAGTTGATGACTTTTTTGGTAAATTAAAATCGGTGTCCAGAGGATTTGCCTCTTTAGATT

[0092] ATGAGGATGCTGGCTATAGAATTTCTGATGTTGTAAAACTGCAACTCTTGGTTAATGGAAATGCGATTGATGCCTTGT CAAGAGTACTTCATAAATCGGAAGTAGAGAGAGTGGGTAGAGAATGGGTGAAGAAGTTTAAAGAGTATGTTAAATCAC AATTATATGAGGTCGTTATACAGGCCCGAGCTAATAACAAGATAATAGCTAGAGAAACAATTAAGGCAAGAAGAAAAG ATGTTCTCCAAAAGCTGCATGCTTCTGATGTCTCACGAAGGAAAAAACTTTTGGCGAAACAGAAAGAGGGTAAAAAGC ATATGAAAACTGTAGGTAATATTCAAATCAACCAAGAGGCATATCAGGCTTTTTTGCGCCGTTAG

[0093] [SEQ ID No: 11]

[0094] Accordingly, in one embodiment, Saccharomyces cerevisiae chromosome XII 706198 - 717026 inclusive comprises a nucleotide sequence as set out in SEQ ID No: 11, or a fragment or variant thereof.

[0095] In one embodiment, Saccharomyces cerevisiae chromosome XII 706198 - 717026 inclusive comprises a nucleotide sequence as set out in SEQ ID No: 11, or a fragment or variant thereof, comprising at least one SNP as defined in Table 7.

[0096] The skilled person would appreciate that the term "Saccharomyces cerevisiae chromosome XII 706198 -717026 inclusive" includes genes or coding sequences from the beginning of the YLR284C coding sequence to the end of the YLR289W coding sequence from the S288c reference genome. Thus, the skilled person would appreciate that the term "Saccharomyces cerevisiae chromosome XII 706198 - 717026 inclusive" includes the following coding sequences: YLR284C, YLR285W, YLR285C-A, YLR286C, YLR286W-A, YLR287C, YLR287C-A, MEC3 (YLR288C), and YLR289W, and the sequences between these coding regions, including any intergenic regions. It will be appreciated that these genes are part of a single QTL.

[0097] Accordingly, in one embodiment, the at least one gene is selected from a group consisting of: YLR284C, YLR285W, YLR285C-A, YLR286C, YLR286W-A, YLR287C, YLR287C-A, YLR288C, and YLR289W, or a homologue, orthologue or paralogue thereof. In one embodiment, the at least one gene is MEC3 (YLR288C) and / or YLR287, or a homologue, orthologue or paralogue thereof. In one embodiment, the at least one gene is MEC3 (YLR288C), or a homologue, orthologue or paralogue thereof. In one embodiment, the at least one gene is YLR287C, or a homologue, orthologue or paralogue thereof.

[0098] In one embodiment, the at least one gene is selected from a group consisting of: NNT1 (YLR285W), CTS1 (YLR286C), YLR287C, RPS30A (YLR287C-A), MEC3 (YLR288C) and GUF1 (YLR289W), or a homologue, orthologue or paralogue thereof.

[0099] In one embodiment, the at least one gene is CTS1 (YLR286C), or a homologue, orthologue or paralogue thereof. In one embodiment, the at least one gene is RPS30A (YLR287C-A), or a homologue, orthologue or paralogue thereof. Alternatively, in another embodiment, the at least one gene is GAT1 (YFL021W), or a homologue, orthologue or paralogue thereof. In another embodiment, the at least one gene is UBP14 (YBR058C), or a homologue, orthologue or paralogue thereof.

[0100] In one embodiment, the at least one gene may be at least two, at least three, at least four, at least five, or at least six genes selected from a group consisting of:

[0101] GAT1 (YFL021W), UBP14 (YBR058C), YLR284C, YLR285W, YLR285C-A, YLR286C, YLR286W-A, YLR287C, YLR287C-A, YLR288C, and YLR289W, or a homologue, orthologue or paralogue thereof. In one embodiment, the at least one gene may be at least seven, at least eight, at least nine, at least ten, or at least eleven genes selected from a group consisting of: GAT1 (YFL021W), UBP14 (YBR058C), YLR284C, YLR285W, YLR285C-A, YLR286C, YLR286W-A, YLR287C, YLR287C-A, YLR288C, and YLR289W, or a homologue, orthologue or paralogue thereof.

[0102] In one embodiment, the at least one gene may be at least two, at least three, or at least four genes selected from a group consisting of: GAT1 (YFL021W), UBP14 (YBR058C), MEC3 (YLR288C), and YLR287C, or a homologue, orthologue, or paralogue thereof.

[0103] In one embodiment, the at least two genes may be GAT1 (YFL021W) and UBP14 (YBR058C), or a homologue, orthologue or paralogue thereof. In one embodiment, the at least two genes may be GAT1 (YFL021W) and MEC3 (YLR288C), or a homologue, orthologue or paralogue thereof. In one embodiment, the at least two genes may be GAT1 (YFL021W) and YLR287C, or a homologue, orthologue or paralogue thereof. In one embodiment, the at least two genes may be UBP14 (YBR058C) and MEC3 (YLR288C), or a homologue, orthologue or paralogue thereof.

[0104] In one embodiment, the at least two genes may be UBP14 (YBR058C) and YLR287C, or a homologue, orthologue or paralogue thereof. In one embodiment, the at least two genes may be MEC3 (YLR288C) and YLR287C, or a homologue, orthologue or paralogue thereof.

[0105] In one embodiment, the at least three genes may be GAT1 (YFL021W), UBP14 (YBR058C) and MEC3 (YLR288C), or a homologue, orthologue or paralogue thereof. In one embodiment, the at least three genes may be GAT1 (YFL021W), UBP14 (YBR058C) and YLR287C, or a homologue, orthologue or paralogue thereof. In one embodiment, the at least three genes may be GAT1 (YFL021W), MEC3 (YLR288C) and YLR287C, or a homologue, orthologue or paralogue thereof. In one embodiment, the at least three genes may be UBP14 (YBR058C), MEC3 (YLR288C) and YLR287C, or a homologue, orthologue or paralogue thereof. In another embodiment, the at least four genes are GAT1 (YFL021W), UBP14 (YBR058C), MEC3 (YLR288C), and YLR287C, or a homologue, orthologue, or paralogue thereof.

[0106] In one embodiment, the recombinant or engineered eukaryotic cell according to the first aspect comprises at least one additional non-naturally occurring allele associated with reduced proteolysis, wherein the at least one additional allele is for at least one additional gene selected from a list of genes in Table 1, or a homologue, orthologue or paralogue thereof, wherein the at least one additional gene is associated with proteolysis.

[0107] In one embodiment, the recombinant or engineered eukaryotic cell according to the second aspect comprises at least one additional gene with a non-naturally occurring combination of single nucleotide polymorphisms (SNPs), wherein the at least one additional gene is selected from a list of genes in Table 1, or a homologue, orthologue or paralogue thereof, wherein the at least one additional gene is associated with proteolysis. In one embodiment, the recombinant or engineered eukaryotic cell according to the third aspect comprises at least one additional modified gene, wherein the at least one additional modified gene is selected from a list of genes in Table 1, or a homologue, orthologue or paralogue thereof, wherein the at least one additional gene is associated with proteolysis.

[0108] In one embodiment, the at least one additional gene may be at least two, at least three, at least four, or at least five genes selected from the list of genes in Table 1, or a homologue, orthologue, or paralogue thereof. In another embodiment, the at least one additional gene may be at least six, at least seven, at least eight, at least nine or at least ten genes selected from the list of genes in Table 1, or a homologue, orthologue, or paralogue thereof.

[0109] Alternatively, in another embodiment, the at least one additional gene may be at least eleven, at least twelve, at least thirteen, at least fourteen or at least fifteen genes selected from the list of genes in Table 1, or a homologue, orthologue, or paralogue thereof. Alternatively, in another embodiment, the at least one additional gene may be at least sixteen, at least seventeen, at least eighteen, at least nineteen or at least twenty genes selected from the list of genes in Table 1, or a homologue, orthologue, or paralogue thereof.

[0110] Alternatively, in another embodiment, the at least one additional gene may be at least twenty-two, at least twenty-four, at least twenty-six, at least twenty-eight or at least thirty genes selected from the list of genes in Table 1, or a homologue, orthologue, or paralogue thereof.

[0111] Preferably, the at least one additional gene is selected from a group consisting of: UBP3, SPL2, AVT7, YIL108W, FYV10, WSS1, GPB1, RPN2, and RAD4. Most preferably, the at least one gene is selected from a group consisting of: UBP3, SPL2 and AVT7.

[0112] In another preferred embodiment, the at least one gene is selected from a group consisting of: YIL108W, FYV10, WSS1, GPB1, RPN2, and RAD4.

[0113] / gene or allele of this invention includes upstream and downstream regions from the protein coding sequence, which may be at least lOObp, at least 200bp, at least 300bp, at least 400bp, at least 500bp, at least 600bp, at least 700bp, at least 800bp, at least 900bp, at least lOOObp either side of the coding sequence, or up to the next adjacent gene or coding sequence. SNPs present in these regions are also claimed in this invention, as are SNPs within protein coding sequences which are synonymous and do not result in amino acid changes.

[0114] MEC3 (YLR288C)

[0115] In one embodiment, when the at least one gene is MEC3, most preferably the gene comprises the SNP T1362G. In this embodiment, a guanine replaces a thymine at position 1362 of the nucleotide sequence. Preferably, this SNP results in the amino acid substitution Asp454Glu, i.e. glutamine replaces aspartic acid at position 454 of the amino acid sequence. Alternatively, in a next preferred embodiment, the gene comprises an adenine or a cytosine at position 1362 of the nucleotide sequence.

[0116] Least preferably, the gene comprises a thymine at position 1362 of the nucleotide sequence.

[0117] In another embodiment, when the at least one gene is MEC3, most preferably the gene comprises the SNP C359T. In this embodiment, a thymine replaces a cytosine at position 359 of the nucleotide sequence. Preferably, this SNP results in the amino acid substitution Thrl20Ile, i.e. isoleucine replaces threonine at position 120 of the amino acid sequence. Alternatively, in a next preferred embodiment, the gene comprises an adenine or a guanine at position 359 of the nucleotide sequence.

[0118] Least preferably, the gene comprises a cytosine at position 359 of the nucleotide sequence.

[0119] YLR287C

[0120] In one embodiment, when the at least one gene is YLR287C, most preferably the gene comprises the SNP G19A. In this embodiment, an adenine replaces a guanine at position 19 of the nucleotide sequence. Preferably, this SNP results in the amino acid substitution Asp7Asn, i.e. asparagine replaces aspartic acid at position 7 of the amino acid sequence. Alternatively, in a next preferred embodiment, the gene comprises a thymine or a cytosine at position 19 of the nucleotide sequence. Least preferably, the gene comprises a guanine at position 19 of the nucleotide sequence. In one embodiment, when the at least one gene is YLR287C, the gene comprises the SNP A1001G. In this embodiment, a guanine replaces an adenine at position 1001 of the nucleotide sequence. Preferably, this SNP results in the amino acid substitution Tyr334Cys, i.e. cysteine replaces tyrosine at position 334 of the amino acid sequence. Alternatively, in a next preferred embodiment, the gene comprises a thymine or a cytosine at position 1001 of the nucleotide sequence. Least preferably, the gene comprises an adenine at position 1001 of the nucleotide sequence.

[0121] In one embodiment, when the at least one gene is YLR287C, the gene comprises the SNP G992A. In this embodiment, an adenine replaces a guanine at position 992 of the nucleotide sequence. Preferably, this SNP results in the amino acid substitution Ser331Asn, i.e. asparagine replaces serine at position 331 of the amino acid sequence. Alternatively, in a next preferred embodiment, the gene comprises a thymine or a cytosine at position 992 of the nucleotide sequence. Least preferably, the gene comprises a guanine at position 992 of the nucleotide sequence.

[0122] In one embodiment, when the at least one gene is YLR287C, the gene comprises the SNP T722C. In this embodiment, a cytosine replaces a thymine at position 722 of the nucleotide sequence. Preferably, this SNP results in the amino acid substitution Leu241Ser, i.e. serine replaces leucine at position 241 of the amino acid sequence. Alternatively, in a next preferred embodiment, the gene comprises a guanine or an adenine at position 722 of the nucleotide sequence. Least preferably, the gene comprises a thymine at position 722 of the nucleotide sequence.

[0123] RPS30A (YLR287C-A)

[0124] In one embodiment, when the at least one gene is RPS30A, most preferably the gene comprises the SNP G148A. In this embodiment, an adenine replaces a guanine at position 148 of the nucleotide sequence. Preferably, this SNP results in the amino acid substitution Val50Ile, i.e. isoleucine replaces valine at position 50 of the amino acid sequence. Alternatively, in a next preferred embodiment, the gene comprises a cytosine or thymine at position 148 of the nucleotide sequence. Least preferably, the gene comprises a guanine at position 148 of the nucleotide sequence.

[0125] CTS1 (YLR286C)

[0126] In one embodiment, when the at least one gene is CTS1, most preferably the gene comprises the SNP T47C. In this embodiment, a cytosine replaces a thymine at position 47 of the nucleotide sequence. Preferably, this SNP results in the amino acid substitution Leul6Pro, i.e. proline replaces leucine at position 16 of the amino acid sequence. Alternatively, in a next preferred embodiment, the gene comprises an adenine or a guanine at position 47 of the nucleotide sequence. Least preferably, the gene comprises a thymine at position 47 of the nucleotide sequence.

[0127] In another embodiment, when the at least one gene is CTS1, most preferably the gene comprises the SNP C962G. In this embodiment, a guanine replaces a cytosine at position 962 of the nucleotide sequence. Preferably, this SNP results in the amino acid substitution Thr321Ser, i.e. serine replaces threonine at position 321 of the amino acid sequence. Alternatively, in a next preferred embodiment, the gene comprises an adenine or a thymine at position 962 of the nucleotide sequence. Least preferably, the gene comprises a cytosine at position 962 of the nucleotide sequence.

[0128] In another embodiment, when the at least one gene is CTS1, the gene comprises the SNP G1570A. In this embodiment, an adenine replaces a guanine at position 1570 of the nucleotide sequence. Preferably, this SNP results in the amino acid substitution Asp524Asn, i.e. asparagine replaces aspartic acid at position 524 of the amino acid sequence. Alternatively, in a next preferred embodiment, the gene comprises a cytosine or a thymine at position 1570 of the nucleotide sequence.

[0129] Least preferably, the gene comprises a guanine at position 1570 of the nucleotide sequence.

[0130] NNT1 (YLR285W)

[0131] In one embodiment, when the at least one gene is NNT1, most preferably the gene comprises the SNP G413C. In this embodiment, a cytosine replaces a guanine at position 413 of the nucleotide sequence. Preferably, this SNP results in the amino acid substitution Serl38Thr, i.e. threonine replaces serine at position 138 of the amino acid sequence. Alternatively, in a next preferred embodiment, the gene comprises an adenine or a thymine at position 413 of the nucleotide sequence.

[0132] Least preferably, the gene comprises a guanine at position 413 of the nucleotide sequence.

[0133] GUF1 (YLR289W)

[0134] In one embodiment, when the at least one gene is GUF1, most preferably the gene comprises the SNP C779T. In this embodiment, a thymine replaces a cytosine at position 779 of the nucleotide sequence. Preferably, this SNP results in the amino acid substitution Ser260Phe, i.e. phenylalanine replaces serine at position 260 of the amino acid sequence. Alternatively, in a next preferred embodiment, the gene comprises an adenine or a guanine at position 779 of the nucleotide sequence. Least preferably, the gene comprises a cytosine at position 779 of the nucleotide sequence. UBP14 (YBRQ58O

[0135] In one embodiment, when the at least one gene is UBP14, most preferably the gene comprises the SNP C2244A. In this embodiment, an adenine replaces a cytosine at position 2244 of the nucleotide sequence. Preferably, this SNP results in the amino acid substitution Asp748Glu, i.e. glutamic acid replaces aspartic acid at position 748 of the amino acid sequence. Alternatively, in a next preferred embodiment, the gene comprises a thymine or a guanine at position 2244 of the nucleotide sequence. Least preferably, the gene comprises a cytosine at position 2244 of the nucleotide sequence. In another embodiment, when the at least one gene is UBP14, most preferably the gene comprises the SNP A2129G. In this embodiment, a guanine replaces an adenine at position 2129 of the nucleotide sequence. Preferably, this SNP results in the amino acid substitution Lys710Arg, i.e. arginine replaces lysine at position 710 of the amino acid sequence. Alternatively, in a next preferred embodiment, the gene comprises a cytosine or a thymine at position 2129 of the nucleotide sequence. Least preferably, the gene comprises an adenine at position 2129 of the nucleotide sequence.

[0136] In another embodiment, when the at least one gene is UBP14, most preferably the gene comprises the SNP T2086C. In this embodiment, a cytosine replaces a thymine at position 2086 of the nucleotide sequence. Preferably, this SNP results in the amino acid substitution Phe696Leu, i.e. leucine replaces phenylalanine at position 696 of the amino acid sequence. Alternatively, in a next preferred embodiment, the gene comprises an adenine or a guanine at position 2086 of the nucleotide sequence. Least preferably, the gene comprises a thymine at position 2086 of the nucleotide sequence.

[0137] In another embodiment, when the at least one gene is UBP14, most preferably the gene comprises the SNP A1591G. In this embodiment, a guanine replaces an adenine at position 1591 of the nucleotide sequence. Preferably, this SNP results in the amino acid substitution Arg531Gly, i.e. glycine replaces arginine at position 531 of the amino acid sequence. Alternatively, in a next preferred embodiment, the gene comprises a cytosine or a thymine at position 1591 of the nucleotide sequence. Least preferably, the gene comprises an adenine at position 1591 of the nucleotide sequence. In another embodiment, when the at least one gene is UBP14, most preferably the gene comprises the SNP G1298A. In this embodiment, an adenine replaces a guanine at position 1298 of the nucleotide sequence. Preferably, this SNP results in the amino acid substitution Ser433Asn, i.e. asparagine replaces serine at position 433 of the amino acid sequence. Alternatively, in a next preferred embodiment, the gene comprises a cytosine or a thymine at position 1298 of the nucleotide sequence. Least preferably, the gene comprises a guanine at position 1298 of the nucleotide sequence.

[0138] In another embodiment, when the at least one gene is UBP14, most preferably the gene comprises the SNP T217C. In this embodiment, a cytosine replaces a thymine at position 217 of the nucleotide sequence. Preferably, this SNP results in the amino acid substitution Ser73Pro, i.e. proline replaces serine at position 73 of the amino acid sequence. Alternatively, in a next preferred embodiment, the gene comprises a guanine or an adenine at position 217 of the nucleotide sequence. Least preferably, the gene comprises a thymine at position 217 of the nucleotide sequence.

[0139] GAT1 (YFL021W)

[0140] In one embodiment, when the at least one gene is GAT1, most preferably the gene comprises the SNP G647A. In this embodiment, an adenine replaces a guanine at position 647 of the nucleotide sequence. Preferably, this SNP results in the amino acid substitution Ser216Asn, i.e. asparagine replaces serine at position 216 of the amino acid sequence. Alternatively, in a next preferred embodiment, the gene comprises a cytosine or a thymine at position 647 of the nucleotide sequence. Least preferably, the gene comprises a guanine at position 647 of the nucleotide sequence.

[0141] UBP3 (YER151O

[0142] In one embodiment, when the at least one additional gene is UBP3, most preferably the gene comprises the SNP G770A. In this embodiment, an adenine replaces a guanine at position 770 of the nucleotide sequence. Preferably, this SNP results in the amino acid substitution Ser257Asn, i.e. asparagine replaces serine at position 257 of the amino acid sequence. Alternatively, in a next preferred embodiment, the gene comprises a thymine or a cytosine at position 770 of the nucleotide sequence.

[0143] Least preferably, the gene comprises a guanine at position 770 of the nucleotide sequence. In another embodiment, when the at least one additional gene is UBP3, most preferably the gene comprises the SNP G643A. In this embodiment, an adenine replaces a guanine at position 643 of the nucleotide sequence. Preferably, this SNP results in the amino acid substitution Ala215Thr, i.e. threonine replaces alanine at position 215 of the amino acid sequence. Alternatively, in a next preferred embodiment, the gene comprises a thymine or a cytosine at position 643 of the nucleotide sequence. Least preferably, the gene comprises a guanine at position 643 of the nucleotide sequence.

[0144] In another embodiment, when the at least one additional gene is UBP3, most preferably the gene comprises the SNP A230C. In this embodiment, a cytosine replaces an adenine at position 230 of the nucleotide sequence. Preferably, this SNP results in the amino acid substitution His77Pro, i.e. proline replaces histidine at position 77 of the amino acid sequence. Alternatively, in a next preferred embodiment, the gene comprises a thymine or a guanine at position 230 of the nucleotide sequence. Least preferably, the gene comprises an adenine at position 230 of the nucleotide sequence.

[0145] SPL2 (YHR136O

[0146] In one embodiment, when the at least one additional gene is SPL2, most preferably the gene comprises the SNP G98C. In this embodiment, a cytosine replaces a guanine at position 98 of the nucleotide sequence. Preferably, this SNP results in the amino acid substitution Gly33Ala, i.e. alanine replaces glycine at position 33 of the amino acid sequence. Alternatively, in a next preferred embodiment, the gene comprises an adenine or a thymine at position 98 of the nucleotide sequence. Least preferably, the gene comprises a guanine at position 98 of the nucleotide sequence.

[0147] AVT7 (YILQ88O

[0148] In one embodiment, when the at least one additional gene is AVT7, most preferably the gene comprises the SNP A1148T. In this embodiment, a thymine replaces an adenine at position 1148 of the nucleotide sequence. Preferably, this SNP results in the amino acid substitution Glu383Val, i.e. valine replaces glutamic acid at position 383 of the amino acid sequence. Alternatively, in a next preferred embodiment, the gene comprises a cytosine or a guanine at position 1148 of the nucleotide sequence. Least preferably, the gene comprises an adenine at position 1148 of the nucleotide sequence. GPB1 (YOR371C)

[0149] In one embodiment, when the at least one additional gene is GPB1, most preferably the gene comprises the SNP C881T. In this embodiment, a thymine replaces a cytosine at position 881 of the nucleotide sequence. Preferably, this SNP results in the amino acid substitution Ser294Leu, i.e. leucine replaces serine at position 294 of the amino acid sequence. Alternatively, in a next preferred embodiment, the gene comprises an adenine or a guanine at position 881 of the nucleotide sequence. Least preferably, the gene comprises a cytosine at position 881 of the nucleotide sequence. WSS1 (YHR134W)

[0150] In one embodiment, when the at least one additional gene is WSS1, most preferably the gene comprises the SNP A85G. In this embodiment, a guanine replaces an adenine at position 85 of the nucleotide sequence. Preferably, this SNP results in the amino acid substitution Ile29Val, i.e. valine replaces isoleucine at position 29 of the amino acid sequence. Alternatively, in a next preferred embodiment, the gene comprises a cytosine or a thymine at position 85 of the nucleotide sequence. Least preferably, the gene comprises an adenine at position 85 of the nucleotide sequence. In another embodiment, when the at least one additional gene is WSS1, most preferably the gene comprises the SNP G250A. In this embodiment, an adenine replaces a guanine at position 250 of the nucleotide sequence. Preferably, this SNP results in the amino acid substitution Val84Ile, i.e. isoleucine replaces valine at position 84 of the amino acid sequence. Alternatively, in a next preferred embodiment, the gene comprises a cytosine or a thymine at position 250 of the nucleotide sequence. Least preferably, the gene comprises a guanine at position 250 of the nucleotide sequence.

[0151] In another embodiment, the at least one additional gene may comprise any one of the SNPs (nucleotide substitutions) referred to in Table 2. Accordingly, the at least one additional gene may comprise any one of the amino acid substitutions referred to in Table 2. When performing the QTL analysis, the inventors also discovered that particular SNPs in certain genes are associated with increased proteolysis of the recombinant protein product. Accordingly, in one embodiment, the recombinant or engineered eukaryotic strain does not comprise a combination of SNPs associated with increased proteolysis.

[0152] YIL108W

[0153] In one embodiment, when the at least one additional gene is YIL108W, most preferably the gene does not comprise the SNP C1862T (resulting in the amino acid substitution Ala621Val). Accordingly, in this embodiment, cytosine is present at position 1862 of the nucleotide sequence. Preferably, therefore, alanine is present at position 621 of the amino acid sequence. Alternatively, in a next preferred embodiment, the gene comprises an adenine or a guanine at position 1862 of the nucleotide sequence. Least preferably, the gene comprises a thymine at position 1862 of the nucleotide sequence.

[0154] RAD4 (YER162C)

[0155] In one embodiment, when the at least one additional gene is RAD4, most preferably the gene does not comprise the SNP C2112A (resulting in the amino acid substitution His704Gln). Accordingly, in this embodiment, cytosine is present at position 2112 of the nucleotide sequence. Preferably, therefore, histidine is present at position 704 of the amino acid sequence. Alternatively, in a next preferred embodiment, the gene comprises a thymine or a guanine at position 2112 of the nucleotide sequence. Least preferably, the gene comprises an adenine at position 2112 of the nucleotide sequence.

[0156] RPN2 (YILQ75O

[0157] In one embodiment, when the at least one additional gene is RPN2, most preferably the gene does not comprise the SNP G2357A (resulting in the amino acid substitution Arg786Lys). Accordingly, in this embodiment, guanine is present at position 2357 of the nucleotide sequence. Preferably, therefore, arginine is present at position 786 of the amino acid sequence. Alternatively, in a next preferred embodiment, the gene comprises a cytosine or a thymine at position 2357 of the nucleotide sequence. Least preferably, the gene comprises an adenine at position 2357 of the nucleotide sequence. In another embodiment, when the at least one additional gene is RPN2, most preferably the gene does not comprise the SNP C850T (resulting in the amino acid substitution Pro284Ser). Accordingly, in this embodiment, cytosine is present at position 850 of the nucleotide sequence. Preferably, therefore, proline is present at position 284 of the amino acid sequence. Alternatively, in a next preferred embodiment, the gene comprises an adenine or a guanine at position 850 of the nucleotide sequence. Least preferably, the gene comprises a thymine at position 850 of the nucleotide sequence. FYV10 (YIL097W)

[0158] In one embodiment, when the at least one additional gene is FYV10, most preferably the gene does not comprise the SNP T139A (resulting in the amino acid substitution Ser47Thr). Accordingly, in this embodiment, thymine is present at position 139 of the nucleotide sequence. Preferably, therefore, serine is present at position 47 of the amino acid sequence. Alternatively, in a next preferred embodiment, the gene comprises a cytosine or a guanine at position 139 of the nucleotide sequence. Least preferably, the gene comprises an adenine at position 139 of the nucleotide sequence. In another embodiment, when the at least one additional gene is FYV10, most preferably the gene does not comprise the SNP C247T (resulting in the amino acid substitution His83Tyr). Accordingly, in this embodiment, cytosine is present at position 247 of the nucleotide sequence. Preferably, therefore, histidine is present at position 83 of the amino acid sequence. Alternatively, in a next preferred embodiment, the gene comprises an adenine or a guanine at position 247 of the nucleotide sequence. Least preferably, the gene comprises a thymine at position 247 of the nucleotide sequence.

[0159] In another embodiment, when the at least one additional gene is FYV10, most preferably the gene does not comprise the SNP G1025C (resulting in the amino acid substitution Gly342Ala). Accordingly, in this embodiment, guanine is present at position 1025 of the nucleotide sequence. Preferably, therefore, glycine is present at position 342 of the amino acid sequence. Alternatively, in a next preferred embodiment, the gene comprises an adenine or a thymine at position 1025 of the nucleotide sequence. Least preferably, the gene comprises a cytosine at position 1025 of the nucleotide sequence. In another aspect of the invention, there is provided a recombinant or engineered eukaryotic cell exhibiting reduced proteolysis compared to a corresponding wildtype, progenitor, or common laboratory strain cell, the recombinant or engineered eukaryotic cell comprising a non-naturally occurring combination of alleles associated with reduced proteolysis, wherein the at least one allele is for a gene selected from a list of genes in Table 1, or a homologue, orthologue or paralogue thereof, wherein the at least one gene is associated with proteolysis.

[0160] In another aspect of the invention, there is provided a recombinant or engineered eukaryotic cell exhibiting reduced proteolysis compared to a corresponding wildtype, progenitor, or common laboratory strain cell, the recombinant or engineered eukaryotic cell comprising at least one gene with a non-naturally occurring combination of single nucleotide polymorphisms (SNPs), wherein the at least one gene is selected from a list of genes in Table 1, or a homologue, orthologue or paralogue thereof, wherein the at least one gene is associated with proteolysis.

[0161] In another aspect of the invention, there is provided a recombinant or engineered eukaryotic cell exhibiting reduced proteolysis compared to a corresponding wildtype, progenitor, or common laboratory strain cell, the recombinant or engineered eukaryotic cell comprising at least one modified gene, wherein the at least one gene is selected from a list of genes in Table 1, or a homologue, orthologue or paralogue thereof, wherein the at least one gene is associated with proteolysis.

[0162] In one embodiment, the recombinant or engineered eukaryotic cell comprises a non-naturally occurring combination of alleles associated with reduced proteolysis, wherein the at least one allele is for gene selected from a list of genes in Table 1, or a homologue, orthologue or paralogue thereof, and the recombinant or engineered eukaryotic cell comprises at least one modified gene, wherein the at least one gene is selected from a list of genes in Table 1, or a homologue, orthologue or paralogue thereof.

[0163] In another embodiment, the recombinant or engineered eukaryotic cell comprises at least one gene with a non-naturally occurring combination of single nucleotide polymorphisms (SNPs), and at least one modified gene, wherein the at least one gene is selected from a list of genes in Table 1, or a homologue, orthologue or paralogue thereof. In one preferred embodiment, the eukaryotic cell is a fungal cell. The fungal cell may be unicellular or have a hyphal (filamentous) morphology or exist as propagules. Most preferably, the eukaryotic cell is a yeast cell. In one embodiment, the yeast cell is Pichia pastoris. Preferably, the Pichia pastoris is a Komagataella species (such as K. phaffii, K. pastoris, and K. pseudopastoris), Hansenula polymorpha (also known as Ogataea polymorpha), Kluyveromyces lactis, a Yarrowia species (such as Yarrowia lipolytica), or Schizosaccharomyces pombe. Alternatively, in a preferred embodiment, the yeast cell is a Saccharomyces species yeast, such as Saccharomyces cerevisiae. Most preferably, the yeast cell is Saccharomyces cerevisiae.

[0164] In one embodiment, the yeast cell is a Wine / European (WE), West African (WA), North American (NA) or Sake (SA) strain. In one embodiment, when the at least one gene is MEC3 (YLR288C), the yeast cell is a Wine / European (WE) strain. In one embodiment, when the at least one gene is YLR287C, the yeast cell is a Wine / European (WE) strain. In one embodiment, when the at least one gene is GAT1, the yeast cell is a West African (WA) strain. In one embodiment, when the at least one gene is UBP14, the yeast cell is a West African (WA) strain.

[0165] In another embodiment, the eukaryotic cell is a fungal cell such as an Aspergillus species, including Aspergillus oryzae and Aspergillus niger, or a Trichoderma species, or Myceliophthora thermophila. In another embodiment, the eukaryotic cell is an insect cell. Examples of insect cells include cell lines Sf9 and Sf21 from Spodoptera frugiperda cells, Hi-5 from Trichoplusia ni cells, and Schneider 2 cells and Schneider 3 cells from Drosophila melanogaster cells. Alternatively, in a preferred embodiment, the eukaryotic cell is from Excavata, such as Leishmania tarentolae.

[0166] In another embodiment the eukaryotic cell is a mammalian cell type, such as a Chinese hamster ovary (CHO) cell, a Mouse myeloma lymphoblastoid, e.g. an NSO cell, a Human embryonic kidney cell, e.g. a HEK-293 cell, a Human embryonic retinal cell, e.g. a Crucell's Per.C6 cell, or a Human amniocyte cell, e.g. Glycotope or CEVEC. A "common laboratory strain cell" will be well-known to the skilled person and may include those defined in Louis, E.J. 20161. A common laboratory strain may include yeast strains listed on the Saccharomyces Genome Database (SGD). A common laboratory strain may include one of the following yeast strains: S288C (Reference Genome: GenBank GCF_000146045.2); W303 (GenBank: JRIUOOOOOOOO.l);

[0167] CEN.PK; JRY188; AH22; S150-2B; and CB11 / 63. A common laboratory strain is not a natural strain, and therefore, may contain its own non-naturally occurring combination of alleles. In a case where a yeast strain has been developed through multiple steps, e.g., to improve bioprocessing phenotypes, a progenitor strain of this invention is the original strain used in a strain development program, i.e. the closest relative, or least genetically diverse, compared to the wild-type. The progenitor strain may also be any strain, such as an intermediate strain in a multi-strain lineage, which has been further improved, e.g. for recombinant protein production, to give a final production strain derived from it.

[0168] Preferably, the term "recombinant" when referring to a "recombinant eukaryotic cell", will be understood to mean a eukaryotic cell into which recombinant DNA has been introduced, i.e. a cell containing genetically engineered DNA.

[0169] In one embodiment, the term "recombinant protein" may be any protein not naturally produced by the expression host, including fusion proteins, tagged proteins, muteins, analogues, derivatives, domains, precursors and fragments of any protein or polypeptide, including, but not limited to, the following proteins (or other polypeptides) of interest. Proteins (or other polypeptides) of interest include albumin, transferrin, lactoferrin, immunoglobulin (such as an Fab fragment or single-chain antibody, including, ScFvs, VHHs and VNARs), (haemo)globin, leghaemoglobin, myoglobin, blood clotting factors (such as factors II, VII, VIII, IX), von Willebrand's factor, tick anticoagulant peptide, endostatin, angiostatin, icestructuring proteins, hydrophobins interferons, interleukins, alpha-l-antitrypsin, insulin, GLP-1, glucagon, calcitonin, cell surface receptors, fibronectin, prourokinase, (pre-pro)-chymosin, antigens for vaccines (including virus-like particles), t-PA, urokinase, prourokinase, hirudin, tumour necrosis factor, G-CSF, GM-CSF, Kunitz domain proteins, CNTF, growth hormone, transforming growth factors, fibroblast growth factors, nerve growth factors, serum cholinesterase, aprotinin, amyloid precursor protein, inter-alpha trypsin inhibitor, antithrombin III, apo- lipoproteins, bone morphogenic proteins, MIC-1, leptin, erythropoietin (EPO), thrombopoietin (TPO), parathyroid hormone, platelet-derived endothelial cell growth factor, platelet-derived growth factor, Protein C, Protein S, keratins, collagens, antimicrobial peptides, defensins, chymosins, casein, amylases, and enzymes generally, such as glucose oxidase and superoxide dismutase. The protein may be a viral, microbial, fungal, plant or animal protein, for example, a mammalian protein. In one embodiment, it is a human protein. The recombinant protein may be a protein endogenous to the host, such as an enzyme, for which production has been improved by strain engineering in a recombinant eukaryotic cell.

[0170] Preferably, the term "engineered" when referring to an "engineered eukaryotic cell", will be understood to mean a eukaryotic cell whose genome comprises an addition, deletion or modification of genetic sequences, i.e. a cell containing genetically engineered DNA. Such engineering may be achieved by standard breeding or crossing technologies, mutagenesis or targeted genome or plasmid engineering methods.

[0171] The term "allele" refers to the alternative forms of a gene that are found at the same place on a chromosome. For example, genetically diverse strains used for breeding may contain different alleles for a particular gene which have arisen from mutations within the gene. Novel alleles may also be generated during breeding, e.g. through recombination between alleles from different parents. It is well-known that a SNP is a substitution of a single nucleotide at a specific position within the genome. In a preferred embodiment, the SNP is a non- synonymous SNP. A non-synonymous SNP is a SNP that results in a single amino acid substitution in the protein sequence encoded by the nucleotide sequence. In another preferred embodiment, the SNP introduces or removes a stop codon.

[0172] The SNP may also be synonymous within a coding sequence or within an intergenic region, which may also result in phenotypic changes. It will be appreciated that a SNP within one gene or an intergenic region might impact the expression of adjacent genes in the locus, thereby influencing a phenotype indirectly. For example, a SNP in an intergenic region might affect the promoter function of one or both adjacent genes. Similarly, a SNP in one gene, especially if it results in a stop codon or significantly alters DNA structural features, might prevent, reduce or increase transcription of that gene, leading to altered expression of an adjacent gene.

[0173] The term "homologue" will be well understood by the skilled person to mean a gene or genetic region that is similar in sequence, structure or evolutionary origin to a gene or genetic region in another species or organism. For example, homologous genes may be derived from a single common ancestral gene present in the common ancestor of different organisms. Homologous genes will encode proteins with the same or similar function in different species, and may also be referred to as an "orthologue".

[0174] The sequence identity between two genes, nucleotide sequences or protein sequences can be defined by a sequence alignment algorithm, such as Smith- Waterman or Needleman-Wunsch for local and global alignments, respectively. These algorithms compare the sequences and calculate a score based on the similarity between the nucleotides or residues at each position. The resulting score is used to infer the degree of identity between the sequences. Other popular algorithms for measuring sequence identity include BLAST (Basic Local Alignment Search Tool) and FASTA (Fast Alignment Search Tool for Assessment of similarity). These algorithms are widely used for large-scale sequence comparisons and are commonly available in bioinformatics software packages. For example, the algorithm is in CloneManager 11 software (Sci Ed Software LLC), selecting the Align > Compare Two Sequences > Global options, using standard Scoring Matrix and other parameters for either amino acid or DNA sequences.

[0175] The level of identity is preferably at least 30%, at least 40%, at least 50%, more preferably at least 60%, at least 70%, at least 80%, or at least 90% for amino acid sequences. Preferably, the level of identity is at least 30%, at least 40%, at least 50%, more preferably at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 99% or at least 99.9% for nucleotide sequences.

[0176] The term "orthologue" will be well understood by the skilled person to mean one of two or more homologous gene sequences found in different species. The term "paralogue" will be well understood by the skilled person to mean a gene which has evolved by a gene duplication event within a genome. For example, gene duplication within a single species may involve one copy of the gene receiving a mutation that gives rise to a new gene. Paralogous genes code for a protein with similar, but not necessarily identical functions.

[0177] A gene "associated with proteolysis" may be one in which different alleles or modification of the gene, result in a modulation in proteolysis.

[0178] The recombinant or engineered eukaryotic cell according to the invention exhibits reduced proteolysis of a recombinant protein product. In one embodiment, the recombinant or engineered eukaryotic cell may exhibit reduced proteolysis compared to a eukaryotic cell that does not comprise at least one gene selected from a group consisting of: GAT1 (YFL021W), UBP14 (YBR058C), and a gene present in Saccharomyces cerevisiae chromosome XII 706198 - 717026 inclusive (SEQ ID No: 11), or at least one gene selected from Table 1 with a non-naturally occurring combination of single nucleotide polymorphisms (SNPs), or a non-natural combination of alleles of genes selected from a group consisting of: GAT1 (YFL021W), UBP14 (YBR058C), and a gene present in Saccharomyces cerevisiae chromosome XII 706198 - 717026 inclusive (SEQ ID No: 11), or from the genes listed in Table 1, or does not have a gene from Table 1 modified to reduce proteolysis.

[0179] For example, in a preferred embodiment, the recombinant or engineered eukaryotic cell may exhibit a reduction in proteolysis compared to a wild-type, progenitor, or common laboratory strain cell of at least 5%, at least 10%, at least 20%, at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, or at least 100%.

[0180] Preferably, the recombinant or engineered eukaryotic cell exhibits a reduction in proteolysis at any point within the secretory pathway. For example, in one embodiment, the recombinant or engineered eukaryotic cell may exhibit a reduction in proteolysis after a polypeptide chain of a recombinant protein product has been synthesised by the ribosome. In another embodiment, the recombinant or engineered eukaryotic cell may exhibit a reduction in proteolysis in the extracellular media before harvesting of the recombinant protein product. Alternatively, in another embodiment, the recombinant or engineered eukaryotic cell may exhibit a reduction in proteolysis during downstream processing and / or purification of the recombinant protein product. The reduction in proteolysis may be determined by comparing samples prepared from different yeast strains or populations of yeast strains using proteolytic assays. Samples may comprise secreted proteins in culture supernatants or in a plate assay (e.g. to give clearing zones of a protease substrate, such as casein), or intracellular proteins, which may be fractionated, e.g. soluble or insoluble fractions, or subcellular fractions, which may or may not be treated with different reagents, e.g. solubilising agents or protease inhibitors. Proteolytic assays may be in solution, on solid agar, or in a gel system, e.g. zymography. A wide variety of proteolytic assays may be used. Generally, these are enzymatic with a proteolytic substrate being degraded by one or more proteases. Protease assays may include:

[0181] Fluorogenic assays: These assays use a synthetic substrate that contains a fluorophore and a quencher. When the substrate is cleaved by the protease, the fluorophore and quencher separate, leading to an increase in fluorescence that is proportional to the amount of protease activity.

[0182] Colourimetric assays: These assays use a synthetic substrate that produces a coloured product when cleaved by the protease. The amount of colour produced is proportional to the amount of protease activity. For example, the APE (Alkaline Protease Enzyme) test.

[0183] Gel-based assays: In these assays, the biological sample is separated by gel electrophoresis, and the protease activity is visualised by staining for cleaved products or by zymography, which uses a gel matrix containing a substrate for the protease.

[0184] Spot assay: This assay is similar to zymography, but the substrate is spotted on a solid surface, such as a nitrocellulose or PVDF membrane. The yeast protease is then added, and the protease activity is visualised as a clear spot on the membrane.

[0185] Mass spectrometry-based assays: These assays use mass spectrometry to identify and quantify cleaved peptides, providing a highly sensitive and specific measurement of protease activity. For example, Activity-based protein profiling (ABPP). Radioactivity-based assays: These assays use a labelled substrate, such as a radiolabelled peptide, that is cleaved by the protease, releasing a radioactive product that can be quantified. Antibody or HTRF (Homogeneous Time-Resolved Fluorescence)-based assays: These use highly specific antibodies to detect a specific protease or its substrate.

[0186] In situ hybridisation: This assay allows the visualisation of protease activity within a yeast cell, and can be performed using fluorescently labelled substrate or antibodybased detection of a cleaved substrate.

[0187] Each of these assay types has its own strengths and limitations, and the choice of the assay will depend on the specific goals of the experiment and the properties of the protease being studied, which are well-known to those skilled in the art.

[0188] In a preferred embodiment, the at least one modified gene may be a gene that has been over-expressed, a gene that has been knocked-down or a gene that has been knocked-out. In another embodiment, the at least one modified gene may be engineered to alter the protein function, e.g. by protein engineering to change the amino acid sequence. This may alter the activity of the protein, including its interaction with other cellular components, such as metabolites, substrates, cofactors, nucleic acids, proteins, lipids and carbohydrates. In another embodiment, e.g. if the gene encodes for a nucleic acid component such as a tRNA, the modified gene may be engineered to alter its structure and binding properties. This engineering may also affect the stability and cellular levels of the gene product, e.g. by modulating its degradation.

[0189] The inventors have identified that proteinase A gene PEP4 disruption was important for reducing the overall protease levels in Saccharomyces cerevisiae strains that are genetically diverse for protease metabolism. Accordingly, in a preferred embodiment, the engineered or recombinant eukaryotic cell has a modified or disrupted PEP4 gene, or a homologue, orthologue or paralogue thereof. More preferably, the engineered or recombinant Saccharomyces cerevisiae cell has a modified or disrupted PEP4 gene. It will be appreciated that the recombinant or engineered eukaryotic cell according to the invention, can be used for the production of recombinant proteins.

[0190] Accordingly, in a fourth aspect, there is provided use of the recombinant or engineered eukaryotic cell according to the first, second or third aspect, for producing a recombinant protein.

[0191] In a fifth aspect, there is provided a method for producing a recombinant protein, the method comprising: (i) transforming the recombinant or engineered eukaryotic cell or obtaining a transformed recombinant or engineered eukaryotic cell according to the first, second or third aspect with an expression vector encoding at least one recombinant protein; and

[0192] (ii) culturing the cell in a medium under conditions to produce the recombinant protein.

[0193] The method of culturing the cell under conditions to produce the recombinant protein will be well-known to the skilled person, and will depend on the nature of the cell and the recombinant protein. For example, the cell may be cultured in conventional nutrient medium well-known to the skilled person for culturing eukaryotic cells, and preferably yeast cells. This may be a minimal media containing Yeast Nitrogen Base, e.g. without amino acids, or Yeast Nitrogen Base without amino acids or ammonium sulphate, with ammonium sulphate added separately, which is buffered with sodium phosphate / citrate pH 6.0 (e.g. made using citric acid and di-sodium hydrogen orthophosphate) and containing 2% w / v glucose, which may or may not be supplemented with amino acids, vitamins and other nutrients, including Adenine, L-Arginine, L-Aspartic acid, L-Histidine, L-Isoleucine, L-Leucine, L-Lysine, L-Methionine, L-Phenylalanine, L-Threonine, L-Tryptophan, L-Tyrosine, Uracil or Valine. Media are preferably made with animal-free components. Rich media, e.g. YPD media with 1% yeast extract, 2% peptone, 2% glucose, may also be used comprising yeast extract, peptone and glucose. Alternative carbon sources can be used, such as sucrose, galactose or glycerol.

[0194] For stable maintenance of the expression plasmid selective media are preferred. For example, for plasmids containing the LEU2 selectable marker, media lacking leucine are preferred. For large-scale culture in stirred tank bioreactors additional components such as antifoams may be preferred with media and feed-regimes.

[0195] In one preferred embodiment, the expression vector is an episomal partial-2-micron plasmid, or a CEN vector, or a yeast artificial chromosome, encoding at least one recombinant protein, or integrating an expression cassette encoding at least one recombinant protein into the genome.

[0196] In one preferred embodiment, the expression vector is a whole-2-micron family plasmid. The whole-2-micron family plasmid may be an expression vector comprising at least 50%, at least 60%, at least 70%, at least 80%, preferably at least 90%, or most preferably 100% of the sequence of a natural yeast 2-micron plasmid. Alternatively, the whole-2-micron family plasmid may be a whole-2- micron-family plasmid from another species as described in Sleep et al. 20052(W02005061719A1).

[0197] Preferably, the method comprises transforming at least one yeast strain with a whole-2-micron family plasmid. The whole-2-micron plasmid of Saccharomyces cerevisiae is a small circular, multicopy DNA element that resides in the yeast nucleus at a copy number of about 40-60 per haploid cell. Examples of whole 2- micron family plasmids include Scpl, Scp2 and Scp3, or those described by Strope et a / . 20153.

[0198] More preferably, the expression vector is an engineered whole-2-micron family plasmid. Preferably, an engineered whole-2-micron family plasmid is a whole-2- micron family plasmid which has been engineered for recombinant protein production (a whole-2-micron expression plasmid). Preferably, recombinant protein production is inactive, repressed or uninduced during breeding. Alternatively, the expression plasmid is a stable partial-2-micron plasmid. A stable partial-2-micron plasmid may be an expression vector comprising the 2-micron origin of replication which is dependent on a whole-2-micron family plasmid, a whole-2-micron expression plasmid or functions provided from these plasmids for stable replication and maintenance. Alternatively, the expression plasmid may be an integrative plasmid (e.g. a plasmid that integrates into the yeast genome). Alternatively, the expression plasmid may be a centromeric plasmid (e.g. containing a centromeric sequence and / or an autonomous replicating sequence). Alternatively, the expression plasmid may be an artificial chromosome (e.g. a yeast artificial chromosome or YAC).

[0199] In one embodiment, the cell is cultured for between 12 and 24 hours, between 24 and 36 hours, between 36 and 48 hours, between 48 and 72 hours, between 72 and 96 hours, between 96 and 120 hours, between 120 and 144 hours, or for a duration greater than 144 hours. In another embodiment, the cell is cultured for a time sufficient to reach the desirable production yields of the recombinant protein. In one embodiment, the cell is cultured at between 20°C and 40°C, between 25°C and 35°C, between 26°C and 34°C, between 27°C and 33°C, and between 28°C and 32°C. Most preferably, the cell is cultured at 30°C.

[0200] The method according to the fifth aspect may also comprise the step of isolating and / or purifying the recombinant protein. Such processes for isolating and purifying the recombinant protein will be well-known to the skilled person, and may include, for example, precipitation, ultrafiltration, gel electrophoresis, and chromatography.

[0201] In a sixth aspect, there is provided a recombinant protein obtained from the recombinant or engineered eukaryotic cell according to the first, second or third aspect.

[0202] Preferably, the recombinant protein is purified. It will be appreciated that the invention extends to any nucleic acid or peptide or variant, derivative or analogue thereof, which comprises substantially the amino acid or nucleic acid sequences of any of the sequences referred to herein, including variants or fragments thereof. The terms "substantially the amino acid / nucleotide / peptide sequence", "variant" and "fragment", can be a sequence that has at least 40% sequence identity with the amino acid / nucleotide / peptide sequences of any one of the sequences referred to herein, for example 40% identity with the sequence identified as SEQ ID No: 1-11, and so on.

[0203] Amino acid / polynucleotide / polypeptide sequences with a sequence identity which is greater than 65%, in some embodiments, greater than 70%, in some embodiments, greater than 75%, and in some embodiments, greater than 80% sequence identity to any of the sequences referred to are also envisaged. In some embodiments, the amino acid / polynucleotide / polypeptide sequence has at least 85% identity with any of the sequences referred to, in some embodiments at least 90% identity, in some embodiments at least 92% identity, in some embodiments at least 95% identity, in some embodiments at least 97% identity, in some embodiments at least 98% identity and, in some embodiments at least 99% identity with any of the sequences referred to herein.

[0204] The skilled technician will appreciate how to calculate the percentage identity between two amino acid / polynucleotide / polypeptide sequences. In order to calculate the percentage identity between two amino acid / polynucleotide / polypeptide sequences, an alignment of the two sequences must first be prepared, followed by calculation of the sequence identity value. The percentage identity for two sequences may take different values depending on:- (i) the method used to align the sequences, for example, ClustalW, BLAST, FASTA, Smith-Waterman (implemented in different programs), or structural alignment from 3D comparison; and (ii) the parameters used by the alignment method, for example, local vs global alignment, the pair-score matrix used (e.g. BLOSUM62, PAM250, Gonnet etc.), and gap-penalty, e.g., functional form and constants. Having made the alignment, there are many different ways of calculating percentage identity between the two sequences. For example, one may divide the number of identities by: (i) the length of shortest sequence; (ii) the length of alignment; (iii) the mean length of sequence; (iv) the number of non-gap positions; or (iv) the number of equivalenced positions excluding overhangs. Furthermore, it will be appreciated that percentage identity is also strongly length dependent.

[0205] Therefore, the shorter a pair of sequences is, the higher the sequence identity one may expect to occur by chance.

[0206] Hence, it will be appreciated that the accurate alignment of protein or DNA sequences is a complex process. The popular multiple alignment program ClustalW (Thompson et al., 1994, Nucleic Acids Research, 22, 4673-4680; Thompson et al. , 1997, Nucleic Acids Research, 24, 4876-4882) is one way for generating multiple alignments of proteins or DNA in accordance with the invention. Suitable parameters for ClustalW may be as follows: For DNA alignments: Gap Open Penalty = 15.0, Gap Extension Penalty = 6.66, and Matrix = Identity. For protein alignments: Gap Open Penalty = 10.0, Gap Extension Penalty = 0.2, and Matrix = Gonnet. For DNA and Protein alignments: ENDGAP = -1, and GAPDIST = 4. Those skilled in the art will be aware that it may be necessary to vary these and other parameters for optimal sequence alignment.

[0207] In some embodiments, calculation of percentage identities between two amino acid / polynucleotide / polypeptide sequences may then be calculated from such an alignment as (N / T)*100, where N is the number of positions at which the sequences share an identical residue, and T is the total number of positions compared including gaps but excluding overhangs. In some embodiments, overhangs are included in the calculation. Hence, one method for calculating percentage identity between two sequences comprises (i) preparing a sequence alignment using the ClustalW program using a suitable set of parameters, for example, as set out above; and (ii) inserting the values of N and T into the following formula:- Sequence Identity = (N / T)*100.

[0208] Alternative methods for identifying similar sequences will be known to those skilled in the art. For example, a substantially similar nucleotide sequence will be encoded by a sequence which hybridizes to DNA sequences or their complements under stringent conditions. By stringent conditions, we mean the nucleotide hybridizes to filter-bound DNA or RNA in 3x sodium chloride / sodium citrate (SSC) at approximately 45°C followed by at least one wash in 0.2x SSC / 0.1% SDS at approximately 20- 65°C. Alternatively, a substantially similar polypeptide may differ by at least 1, but less than 5, 10, 20, 50 or 100 amino acids from the sequences shown in, for example, SEQ ID Nos: 1-11.

[0209] Due to the degeneracy of the genetic code, it is clear that any nucleic acid sequence described herein could be varied or changed without substantially affecting the sequence of the protein encoded thereby, to provide a functional variant thereof. Suitable nucleotide variants are those having a sequence altered by the substitution of different codons that encode the same amino acid within the sequence, thus producing a silent change. Other suitable variants are those having homologous nucleotide sequences but comprising all, or portions of, sequence, which are altered by the substitution of different codons that encode an amino acid with a side chain of similar biophysical properties to the amino acid it substitutes, to produce a conservative change. For example small non-polar, hydrophobic amino acids include glycine, alanine, leucine, isoleucine, valine, proline, and methionine. Large non-polar, hydrophobic amino acids include phenylalanine, tryptophan and tyrosine. The polar neutral amino acids include serine, threonine, cysteine, asparagine and glutamine. The positively charged (basic) amino acids include lysine, arginine and histidine. The negatively charged (acidic) amino acids include aspartic acid and glutamic acid. It will therefore be appreciated which amino acids may be replaced with an amino acid having similar biophysical properties, and the skilled technician will know the nucleotide sequences encoding these amino acids.

[0210] All of the features described herein (including any accompanying claims, abstract and drawings), and / or all of the steps of any method or process so disclosed, may be combined with any of the above aspects in any combination, except combinations where at least some of such features and / or steps are mutually exclusive.

[0211] For a better understanding of the invention, and to show how embodiments of the same may be carried into effect, reference will now be made, by way of example, to the accompanying Figures, in which:-

[0212] Figure 1 shows an expression construct (DNA) for amylase-mCherry expression and its protein product (on the left), with secreted amylase activity assay shown on a yeast agar plate containing starch (on the right). MET17 promoter = MET17p, SUC2 leader = Pre, ADH1 terminator = ADHlt.

[0213] Figure 2 shows specific mCherry activities for 41 Progeny Yeast Strains grown at 30°C. The average across four replicates is displayed, and error bars show 1 SD from 4 replicates. Figure 3 shows mCherry vs amylase for yeast strains grown in four microtitre plates (Plates 1 to 4) at 30°C.

[0214] Figure 4 shows different embodiments of an expression construct for amylase- mCherry expression which was inserted into a whole-2-micron plasmid and the protein derived from it. Figure 4a is the construct, pHRIK, which encodes the entire 2-micron plasmid from yeast strain YLF185, and the S. cerevisiae S288C LEU2 gene. Figure 4b is the construct, pHR2A, which encodes the pHRIK Notl-Hpal fragment, the pHRIK Notl-Adl fragment and the pUC57-Amp Eco J-Narl fragment. Figure 4c is the construct, pEV7, which encodes a repressible MET17 promoter driving expression of the alpha-amylase (AA) mCherry fusion protein without any N-linked glycosylation sites. Figure 4d is the construct, pHR!K-pEV7, which comprises DNA from the pHRIK and pEV7 constructs, and is the final whole- 2-micron expression in the cell.

[0215] Table 1 shows the 16 genomic regions identified using (quantitative trait locus) QTL analysis, which contain genes, and thereby alleles, responsible for the differential expression of the recombinant amylase-mCherry protein. The table defines the 16 QTL regions from data at 30°C, either based on the LOD (Logarithm Of the Odds) scores (i.e. the local maximum and 1.5 LOD drop), or + / - 5Kb or 10Kb from the peak itself. The four rows with grey backgrounds have the maximum scores for the replicates or mean LOD.

[0216] Table 2 defines the non-synonymous SNPs in the genes listed in Table 1, which are different in the best two strains and worst two strains. The columns indicate which are the most preferred, preferred / neutral and least preferred bases at these positions.

[0217] Table 3 defines the prioritised list of genes which are associated with the phenotype of reduced proteolysis. This table indicates both the gene name and the systematic name.

[0218] Table 4 shows the results of the reciprocal hemizygosity analysis, measuring foldchange improvements in OD normalised mCherry fluorescence in strains with the alleles of interest, compared to isogenic strains without the alleles of interest. The strain identified for carrying the allele of interest (MEC3 (YLR288C), YLR287C, GAT1 (YFL021W) and UBP14 (YBR058C)) is highlighted bold and underlined. Strains disrupted in their relevant allele are denoted by the suffix -K. Strains transformed with the expression vector pEV7 are denoted by the suffix -L.

[0219] Table 5 shows the plasmids for expression of recombinant proteins, with a description of the expression cassettes they contain (including any detection tags and linkers) and whether they are designed for intracellular or secreted production.

[0220] Table 6 shows a comparison of one of the best two strains for amylase-mCherry secretion (2-A2) with one of the worst two strains for amylase-mCherry secretion (1-C12), for expression of multiple different recombinant proteins. Table 7 shows preferred SNPs in Saccharomyces cerevisiae chromosome XII 706198 - 717026 inclusive (SEQ ID No: 11), including the most preferred non- synonymous SNPs in coding sequences, synonymous SNPs in coding sequences and SNPs outside coding sequences.

[0221] Examples

[0222] The inventors set out to identify the key regions of the yeast genome responsible for reducing proteolysis of a recombinant protein product. In order to do this, the inventors performed quantitative trait loci (QTL) analysis, which identifies regions in the genome associated with improving the phenotype analysed. As described below, the inventors used two detectable markers which are secreted in order to measure and quantify recombinant protein production, i.e. the enzyme, amylase, and the fluorophore, mCherry from Anaplasma marginale (UniProt X5DSL3). These two readily detectable markers are secreted and acted as a measure of recombinant protein production, and facilitated screening and selection of final yeast strains harbouring desired alleles within the defined QTLs.

[0223] Materials and Methods Strain Engineering, Breeding and Selection

[0224] The original strains described by Liti, G., Carter, D., Moses, A. et al. Population genomics of domestic and wild yeasts. Nature 458, 337-341 (2009), from the Saccharomyces Genome resequencing Project (SGRP), were made genetically tractable by Cubillos et al. 20094and Louvel et al. 20145, using standard methods of genome engineering. Representatives of the Wine / European (WE), West African (WA), North American (NA) and Sake (SA) clean lineages were selected, with derivatives YFL185, YFL187, YFL190, and YFL191 described in Louvel et al. 20145, being further modified in this study to create strains Q416, Q413, Q426 and Q427, respectively, with genotypes shown in Table 8 below. Plasmid curing is described by Rose and Broach 19906.

[0225] Table 8 - Yeast strains used in this work

[0226] Plasmid Construction pHRIK was constructed from two PCR fragments encoding the entire 2-micron plasmid from strain YLF185, which were cloned into the EcoRV site of pUC57-Kan (GeneWiz / Azenta) with a 2.1Kb PCR fragment encoding the S. cerevisiae S288C LEU2 gene (flanked by two Pad-sites introduced on the PCR primers), using NEBuilder® HiFi DNA Assembly Cloning Kit.

[0227] This whole 2-micron-family plasmid contained the LEU2 selectable marker and pUC57-Kan integrated at the SnaBI-site downstream of the 2-micron D gene (see GenBank: J01347.1 sequence for Sna BI site location). Sbfl and Notl sites were introduced on PCR primers for the directional insertion of expression constructs downstream of the LEU2 selectable marker. See the pHRIK map in Figure 4a for details. Standard methods were used for E. coli transformation and plasmid preparation.

[0228] Oligonucleotide PCR primers used to construct pHRIK are provided below, with F (forward) and R (reverse) primers binding to the regions PCR1-4 shown in the pHRIK map corresponding to the primer names (Figure 4a).

[0229] Table 9 - Primer sequences Regions of primer sequence binding are marked on the pHRIK map as PCR1 to PCR4. Amplification was performed using a Q5 HiFi PCR kit (NEB) and / or restriction enzyme digestion to generate: a) pUC57-Kan (from GeneWiz) cut with EtoRV (buffer 3.1) or uncut plasmid can be used as PCR template DNA. This is the 1F-2R fragment b) The 2-micron fragment from the SnaBI-site near STB to just before the second inverted repeat, which contains the first inverted repeat, FLP and R.EP2 of the 2-micron B-form. PCR primers contain 20-30bp homology to pUC57-Kan and the second 2-micron fragment, respectively. This fragment is around 3.1Kb in length. This is the 2F-3R fragment. c) The 2-micron fragment containing the second inverted repeat, R.EP1 and D. PCR primers contain 20-30bp homology with the first 2-micron fragment and the LEU2 fragment. This fragment is around 3.4Kb in length. This is the 3F-4R fragment d) The LEU2 fragment flanked by Pad-sites. PCR primers contain 20-30bp homology with the second 2-micron fragment (encoding R.EP1 and D) and the pUC57-Kan fragment. This fragment is around 2.1Kb in length and can be amplified from S288C genomic DNA. This is the 4F-1R fragment

[0230] PCR and DNA cloning methods are well known to those skilled in the art of molecular biology. Co-transformation of a cir° yeast strain, such as YLF185 cir°, with all four fragments followed by selection of leucine prototrophs, extraction of total DNA, transformation of f. coli with kanamycin selection and then extraction and sequencing of plasmid DNA for alignment to the expected pHRl-K sequence may be performed. Alternatively, seamless cloning can be performed in vitro (e.g. using a NEBuilder® HiFi DNA assembly kit) before the yeast and / or E. coli transformations. pHRIK plasmid DNA can be isolated from E. coli, and its identity confirmed by diagnostic restriction enzyme digests and / or DNA sequencing.

[0231] Alternatively, disintegration vectors could be used to construct expression plasmids as shown by Chinery et al. 19897. Referring to Figure 4b, pHR2A was constructed in a 3-way ligation of the 2.6kp pHRIK Notl-Hpal fragment, the 1.2kp pHRIK Notl-AcII fragment and the 2.5kp pUC57-Amp Eco M-Narl fragment (also digested with NEB Antarctic phosphatase). Ligation used NEB high concentration T4 ligase with transformation into NEB10P competent cells for growth on LB plates containing ampicillin. A synthetic expression cassette for amylase-mCherry (GeneWiz / Azenta) was cloned into the Sbfl and Notl sites to give pEV7, as shown in Figure 4c. It was preferred in this case for the recombinant product to lack N-linked glycosylation, which might affect protein activities independently of product yield. Options include the expression of an alpha-amylase lacking potential N-linked glycosylation motifs (-N-X-S / T-) or expressing an alpha-amylase analogue modified to remove existing N-linked glycosylation motifs, e.g. Aspergillus oryzae alpha-amylase (UniProt P0C1B3) with serine 199 in the mature protein changed to an alanine residue. The coding sequence for the fluorophore, mCherry from Anaplasma marginale (UniProt X5DSL3), was genetically fused to the alpha-amylase C-terminal coding sequence to facilitate screening and selection of final yeast strains and to provide an additional phenotype for evaluating productivity and differential proteolysis. Alternatively, fluorescent proteins such as yEGFP, ymUkGl, mOrange and mNeonGreen and others, such as described by Thorn 20178and Kaishima at al. 20169, could be used. For all codon sequences, unless specified otherwise, codons were selected for rapid mRNA translation in S. cerevisiae as described by Chu et al. 201410. pEV7 contains an expression construct encoding a repressible MET17 promoter as described by Solow et al. 200511, which is driving expression of the alpha-amylase

[0232] (AA) mCherry fusion protein without any N-linked glycosylation sites (see the DNA map in Figure 1). The repressible MET17 promoter can be switched off during breeding (by adding methionine to the culture media at 20mM or above. See Solow et al. 200511for details of methionine concentrations needed for repression and expression) and switched on for the production of amylase-mCherry in selected strains by growth in media lacking methionine. Secretion was directed by the secretory leader sequence from the S. cerevisiae SUC2 (invertase) gene. This preleader sequence is removed by signal peptidase during translocation into the endoplasmic reticulum to give the mature amylase-mCherry protein, which is then secreted into the culture media. The S. cerevisiae ADH1 terminator ADHlt) was used for transcriptional termination. Expression cassettes for additional recombinant proteins were designed similarly to pEV7, as Sbfl-Notl fragments, which were cloned in place of the amylase-mCherry expression cassette for transcription in the same direction as the LEU2 gene in the final whole-2-micron vectors. The plasmids for expression of the additional recombinant proteins are described below and in Table 5, with a description of the expression cassettes they contain (including any detection tags and linkers) and whether they are designed for intracellular or secreted production. pEVl contains an expression cassette for intracellular expression of mCherry. The mCherry coding sequence is essentially the same in all constructs of this invention.

[0233] This expression cassette contains the MET17 promoter and ADH1 terminator described above for pEV7. pEV51 & pEV52 contain intracellular expression cassettes for the virus-like particle proteins HPV16(Ll)-mCherry and mCherry-GSl-HBsAg, respectively, where GS1 is a linker with sequence GSGGSGGSGPVTN (SEQ ID No: 9). HPV16(L1) encodes human papillomavirus type 16 major capsid protein LI (UniProt: P03101 VL1_HPV16). HBsAg encodes Hepatitis B virus S protein (GenBank: AIJ50189.1). These expression cassettes contain the MET17 promoter and ADH1 terminator described above for pEV7. pEV3 contains an expression cassette for the secretion of rHA-mCherry, where rHA encodes recombinant human albumin with the mature albumin sequence from UniProt P02768. This expression cassette contains the MET17 promoter, SUC2 leader and ADH1 terminator described above for pEV7. pEV388 contains an expression cassette for the secretion of rHA (without an mCherry tag) with transcription from the Saccharomyces cerevisiae proteinase B promoter PRBlp). Secretion is directed by the modified fusion leader sequence (mFL). The DNA coding sequence for mFL-rHA in pEV388 is the same as the open reading for mFL-rHA in SEQ ID 19 of W02004009819. pEV275 contains an expression cassette for a VHH domain antibody for prostatespecific membrane antigen, PSMA (Chatalic et al., 201520). This expression cassette contains the MET17 promoter, SUC2 leader and ADH1 terminator described above for pEV7. pEV299 contains the same expression cassette as pEV275, except with transcription from the Saccharomyces cerevisiae proteinase B promoter (PRBlp). pEV298 contains the same expression cassette as pEV299, except for secretion of a VHH-mCherry fusion protein. pEV395 contains an expression cassette for secretion of a GLP1 (9-37) analogue precursor-GS-HiBit fusion protein with amino acid sequence EGTFTSDVSSYLEGQAAKEFIAWLVRGRGGGGGSGGGGSVSGWRLFKKIS (SEQ ID No: 10) with a Saccharomyces cerevisiae Mating Factor-alpha-derived leader sequence and Saccharomyces cerevisiae BCY1 -derived terminator. Yeast Strain Transformation

[0234] S. cerevisiae Q427 (SA lineage) was transformed to leucine prototrophy using a lithium acetate method (Sigma Aldrich Yeast Transformation Kit YEAST1) with DNA fragments from pHRIK and pEV7 (described below) for in vivo assembly of the final expression plasmid pHRlK-pEV7 by homologous recombination, which is shown in Figure 4d. Prototrophic transformants were selected on media without leucine, and cryopreserved stocks were prepared with 25% glycerol after plating a single transformant to provide single colonies. The expression plasmid, pHRlK-pEV7, shown in Figure 4d, was subsequently transferred to all progeny during multigeneration breeding.

[0235] Approximately lOOng each of gel purified 7.3Kb pHRIK Bst' I-Notl fragment and pEV7 digested with SwaI+Acc65I were co-transformed into Q427 cir° to give strain Q427 [pHRlK-pEV7] following homologous recombination (gap-repair) of the plasmid DNA fragments. Gap-repair transformation and other yeast methods are described in Andersen et al. 201212, and Finnis et al. 201813(WO2018234349A1).

[0236] Equivalent yeast transformations were performed as required for the introduction of additional expression plasmids with homologous recombination (gap-repair) between DNA fragments comprising the whole-2-micron plasmid sequence and fragments comprising the expression construct with homologous flanking sequences from the different pEV-plasmids described above.

[0237] Breeding

[0238] Breeding methods are described in Cubillos et al. 201314. Breeding was performed for 12 generations before strain selection. Cpl2 populations were generated from

[0239] Q416, Q413, Q426 and Q427 with two pairwise crosses first and then mixed 4-way crosses with selection for Ura + Lys+ in between cycles, in the presence of methionine repression.

[0240] Cpl2 diploid libraries comprise approximately 25% PEP4: :PEP4 homozygotes, 50% PEP4: :pep4 heterozygotes, and 25% pep4: :pep4 homozygotes. Selection from

[0241] Cpl2 cir° libraries, or Cpl2 libraries containing a whole-2-micron expression plasmid, such as pHR!K-pEV7 introduced by using Q427 [pHR!K-pEV7] as a parental stain in place of Q427 cir°, was performed using G418. G418 is an aminoglycoside antibiotic similar in structure to gentamicin Bl, which blocks polypeptide synthesis by inhibiting the elongation step as described by Goldstein et al. 199915. Strains containing the KanMX gene are resistant to G418.

[0242] A homozygous pep4::pep4 diploid population ("pure diploids", PD) was prepared by allowing germination of genetically diverse Cpl2 progeny in the presence of G418 to select for pep4: : KanMX spores only, followed by mating to form pep4::pep4 diploids.

[0243] Alternatively, a mixed population of heterozygous pep4: :PEP4 and homozygous pep4::pep4 diploids ("mixed diploids", MD was prepared by allowing germination and mating of pep4: -.KanMX and PEP4 haploids followed by selection against PEP4:PEP4 homozygous diploids with G418.

[0244] Therefore, multigenerational libraries can be produced with a range of proteinase A genotypes, including an absence of proteinase A gene, despite this protease being essential for the breeding process.

[0245] Strain Selection

[0246] Spores with the pep4 genotype were selected with G418 after germination and mated to give homozygote pep4::pep4 diploids as described above. 48 strains were selected with a range of mCherry levels in the supernatant following microtitre plate culture. Cryopreserved stocks were made with 25% final glycerol concentration for storage at below -70°C. Genotypic and phenotypic data were obtained for all strains, of which seven strains were excluded, e.g., for poor growth and / or low- quality sequence data, and the remaining 41 strains were used for QTL analysis. Strains were typically grown on BMMD media as defined by Evans et al. 201016, with appropriate supplements, e.g., CSM-Leu (Formedium Ltd). For expression studies, strains were typically cultured in 0.5mL BMMD+(CSM-Leu-Met) media in 48-well microtitre plates (MTP) shaken at 30°C, 280rpm, 2.5cm orbit in a humidity chamber. Details of methods are provided in Schelde et al. 201917, Ramaiya et al. 201718(WO2017112847A1) and Finnis et al. 201813(WO2018234349A1).

[0247] Phenotvoing Secreted amylase activity from the expression construct was confirmed by growing the yeast on agar plates containing starch, which was degraded by the secreted amylase to give a visible zone around the cultured yeast (Figure 1). Protease activities affecting either the amylase domain or the mCherry domain can influence the results from amylase assays or mCherry signal detection.

[0248] For accurate detection of specific mCherry levels in the culture supernatant for use as phenotypic data in QTL analysis, yeast stocks were grown up at 500pL scale in 48-well MTPs, using repressive media containing methionine for three days before cultures were transferred into non-repressible minimal growth media to allow recombinant protein expression. Cultures were set up in replicates of four. Growth was monitored, and cultures were harvested at 3.5 days, at which point OD620 and mCherry fluorescence readings were recorded using a TECAN Spark with the following settings: OD readings were taken at 620nm using 10 flashes and a settle time of 50 ms. mCherry fluorescence readings were taken using excitation of 540nm and emission of 614nm. Whole culture readings were corrected using a gain of 100, and supernatant-only values were corrected using a gain of 150. Cultures were then centrifuged, and 300pL supernatant was removed before an additional centrifugation step. 250pL cleared supernatant was removed, and OD620 and mCherry fluorescence was recorded using the settings above.

[0249] To generate amylase activity data, 50pL cleared supernatant was analysed using the EnzChek ® Ultra Amylase Assay Kit (E33651) in 48-well MTPs. Fluorescence intensity at 505 / 512nm was measured at various points during the assay using the TECAN Spark with a gain of 30, for 15 minutes.

[0250] DNA Preparation and Sequencing

[0251] Genotypic data were obtained by the following methods: Genomic DNA was retrieved from a 5mL overnight culture of each strain grown in YPD media (1% yeast extract, 2% peptone, 2% glucose) using the Promega Wizard DNA extraction kit. gDNA was re-precipitated in lOmM Tris-HCI, pH 8.5 and DNA quality was assessed using a Nanodrop 2000 Spectrophotometer and visualised using agarose gel electrophoresis before genomic DNA sequencing. LITE (Low Input, Transposase Enabled) Library preparation was carried out on each DNA sample before fortyeight diploid genomes were sequenced on the Illumina NovaSeq 6000 SP lane, generating 150bp PE reads. Sequencing was carried out by the Earlham Enterprises Ltd, Norwich, UK. The average read number per sample was 10,556,161, which represents an average estimated genome coverage of 263 times.

[0252] OTL Analysis

[0253] Quantitative Trait Locus (QTL) analysis is a statistical method that links both phenotypic and genotypic data to explain the genetic basis of variation in complex traits as described by Miles et al. 200819. This method was performed on a range of

[0254] S. cerevisiae strains selected with a range of expression levels for the recombinant protein amylase-mCherry. This analysis identifies regions of the genome containing alleles (also containing SNPs, single nucleotide polymorphisms, or QTNs, quantitative trait nucleotides) causing a phenotype, e.g. reduced levels of proteolysis. The QTL analysis identifies "regions" (also called "intervals") in the genome associated with improving the phenotype analysed.

[0255] Short reads were first assessed for sequencing quality using fastqc, before each read was aligned against the reference genome of S288C (R64-2-1), using bwa. Alignments were indexed and sorted using samtools, and duplicate reads marked and removed using picard tools. Variants were then called using freebayes. The parameters were set for the minimum mapping quality to 20 and ploidy to diploid (- -min-mapping-quality 20 -min-base-quality 20 -p 2), then this output was subjected to a set of filters to use as genetic markers with SNP sites for the samples.

[0256] The following filters were applied: a. The variant calling quality is more than 20; b. The observation of the variant is 100% of the samples in the calling set; c. Allele frequency (REF / (REF + ALT in (0.1, .09)); and d. Calling positions are the known bi-allele variant sites for SGRP founders.

[0257] The reproducibility of strain measurements across each plate was assessed using R.

[0258] QTL (Quantitative Trait Loci) analysis was applied to find the association between genotypes and phenotype measurements for each plate as well as the average records. Specific activity values (mCherry fluorescence I OD620) from each culture (obtained using the methods described above) were used as phenotype inputs for the QTL analysis. LOD (Logarithm Of the Odds) score was calculated for each locus. The selected candidate QTL intervals are listed if 1) the LOD score is > 3 for each separate replicate plate analysis and 2) the max LOD score is > 5 for the max score when all replicate plates are considered. The intervals are summarised by the local maximum and 1.5 LOD drop. In addition, 5k and 10k flanking regions are also summarised for each of the selected peak markers. Variants in QTLs resulting in nonsynonymous mutation were further annotated. To narrow down and identify candidate causative genes, only the sites which appear in the top 2 performing strains and alternatives present in the bottom 2 performing strains are considered.

[0259] Additionally, subsets of these genes with nonsynonymous mutations were selected based on gene function and position within each interval.

[0260] Plasmid Curing Progeny strains are leucine auxotrophs with a non-functional genomic Ieu2 allele. Growth in synthetic media lacking leucine is restored when strains are transformed with expression plasmids containing a functional copy of the LEU2 gene.

[0261] To cure yeast strains of the whole 2-micron expression plasmids for recombinant protein expression, such as original amylase-mCherry secretion plasmid, serial passaging was conducted in small (500pl) liquid cultures with a non-selective, synthetic drop-out medium containing leucine, e.g. BMMD+(CSM-Leu-Met) media in 48-well microtitre plates (MTP) shaken at 30°C, 280rpm, 2.5cm orbit in a humidity chamber. After 5 passages with growth to late-log or stationary phase and approximately 1-2% inocula, single colonies of each strain were isolated on YPD agar plates, patched onto YPD agar plates, and replica plated onto synthetic dropout plates lacking leucine. Colonies exhibiting no growth on the synthetic dropout medium indicated potential loss of the expression plasmid. To confirm plasmid loss, PCR of Leu- colonies was conducted with primers annealing to the 2p origin of the expression plasmid (Strope et al., 20153). Samples were analysed by gel electrophoresis, with the absence of a DNA amplicon indicating successful plasmid loss.

[0262] Reciprocal Hemizvaositv

[0263] Reciprocal hemizygosity is the standard method to validate the effect of one allele over the other in an Fl hybrid, as taught by Mackay et al. 200921. The principle of reciprocal hemizygosity is to assay the phenotype of interest in isogenic diploids that differ only in which allele of a candidate gene of interest is present using the methods described in Cubillos, et al. 201122, Liti and Louis 201224.

[0264] In general, this involves a pair of haploid strains bearing different alleles at the locus of interest being mated to produce a diploid. Two diploids are made, each with one or the other allele deleted.

[0265] Therefore, pairwise crossings of up to three haploid parental strains disrupted of the wildtype alleles, respectively, and the parental strain harbouring a novel allele of interest, were produced. Reciprocal crossings of the relevant haploid strain, disrupted for the novel allele, and up to three parental strains, harbouring the wildtype allele, were produced. For each gene, the allele of interest can only be found in one of the four parental backgrounds that were previously described in Table 8. The expression cassette pHRlK-pEV7 was introduced to the diploid strains via homologous (gap-repair) transformation of the haploid cell of each above crossing that was not disrupted for any allele. Transformation was conducted as described above in 'Yeast Strain Transformation'. All allele disruptions were coding sequence deletions, using standard methods of genome engineering known in the art. The deletion of each gene in the yeast genome was achieved by insertion of an antibiotic resistance cassette, KanMX, conferring resistance to G418. PCR-generated deletion cassettes can be made either from genomic DNA from the relevant strain from the deletion collection or from long oligos with homology to the flanking regions of the coding sequence of the gene of interest and homology to the KanMX cassette used for replacing the open reading frame (Cubillos, et al. 201122; Parts et al. 201123; Liti and Louis 201224; Cubillos et al. 201314). Supernatant mCherry fluorescence values (FU) were measured for each diploid strain after growth using the protocols previously described for culturing and measurement. The values were normalised for cell growth (OD620). A ratio of these values for each pairwise crossing was calculated, where these values denote folddifference relative to the diploid with wildtype allele. An unpaired two-tailed t-test was conducted to determine if the result was statistically significant. If different alleles of the gene alter amylase-mCherry expression in the isogenic diploids, the gene is associated with proteolysis. This also identifies the preferred allele for reducing proteolysis. Results

[0266] Example 1 - mCherrv and amylase activities

[0267] An amylase-mCherry fusion protein was expressed from pHRlK-pEV7 shown in Figure 4d in the genetically diverse yeast population. Strains were selected from this genetically diverse population with a range of mCherry levels. In the absence of proteolysis to degrade the secreted amylase-mCherry product, the different progeny selected would have different levels of amylase-mCherry polypeptide production with secretion into the extracellular media, e.g. based on their different productivities. If the full-length protein product is highly stable during both the secretion process and in the culture media (e.g. without significant proteolysis acting on the amylase-mCherry polypeptide), the ratio of amylase to mCherry activity is expected to remain relatively constant.

[0268] Figure 2 shows the mCherry activities from the 41 strains selected for growth at 30°C and used for the QTL analysis. The supernatant mCherry levels were corrected for growth differences (OD620) for the strains used in the QTL analysis. Significant diversity in the mCherry phenotype is observed, with more than a 10- fold difference observed between the highest and lowest producers. A 10.8 fold difference in mCherry levels in the supernatant for the highest producer (2A2) was obtained compared to the lowest producer (3D9).

[0269] Figure 3 shows plots of mCherry against amylase specific activities for each strain grown at 30°C, with strains grown in four MTPs. While the mCherry levels range widely from low to high, the amylase levels tend not to fall so far towards zero. This indicates that proteolysis has affected the mCherry and amylase domains differently. The mCherry domain appears to be more protease sensitive than the amylase domain.

[0270] Example 2 - QTL Analysis QTL analysis identified 16 genomic regions comprising approximately 3.3% of the total Saccharomyces cerevisiae genome containing genes and alleles responsible for the differential expression of the recombinant amylase-mCherry protein (see Table 1). Successive statistical and bioinformatic filters and rational selection methods were used to shortlist sequences, e.g. QTLs, genes and QTNs, responsible for reduced proteolysis. 16 QTL intervals were identified using the specific mCherry activity values (mCherry fluorescence I OD620) at 30°C. Of these, 4 QTL intervals contained the maximum LOD scores for the mean LOD scores and the four replicates. These are QTLs 6, 7, 9 and 13 highlighted in grey in Table 1. Table 1 shows the 16 genomic regions identified using QTL analysis. Each of these QTL regions lists the genes (i.e. interval genes), containing SNPs responsible for the differential expression of the recombinant amylase-mCherry protein. As such, the inventors identified these "interval genes" as being associated with reduced proteolysis. The table defines the 16 QTL regions based on their chromosome number, and their start and end bp position.

[0271] Table 2 shows the non-synonymous SNPs identified in the genes listed in Table 1, which were found to differ amongst the best two strains at recombinant protein production and the worst two strains at recombinant protein production. In other words, the inventors identified these specific SNPs as being associated with increased or decreased proteolysis. The table identifies the SNP (i.e. the nucleotide substitution), as well as the resulting amino acid substitution.

[0272] The table then indicates whether this SNP is associated with reduced proteolysis, by the presence of a "1 / 1" in the column "Good Performance Genotype". For these

[0273] SNPs, the preferred nucleotide is the substituted base after the ">" symbol, and the preferred amino acid is the second one listed. For example, for UBP3, preferably the gene has an adenine at position 770 of the nucleotide sequence, and an asparagine at position 257 of the amino acid sequence.

[0274] Alternatively, the Table indicates if a SNP is associated with increased proteolysis, by the presence of a "1 / 1" in the column "Poor Performance Genotype". For these SNPs, the preferred nucleotide is the base before the ">" symbol, and the preferred amino acid is the first one listed. For example, for YIL108W, preferably the gene has a cytosine at position 1862 of the nucleotide sequence, and an alanine at position 621 of the amino acid sequence. Additionally, Table 2 indicates which is the most preferred base at this position of the respective gene, which bases are preferred / neutral, and which base is the least preferred at this position, for improved recombinant protein production. Table 3 shows the preferred list of genes which are associated with the phenotype of reduced proteolysis.

[0275] Example 3 - Reciprocal Hemizvaositv Analysis

[0276] As alleles of interest, MEC3 (YLR288C), YLR287C, GAT1 (YFL021W) and UBP14 (YBR058C) were chosen, harbouring all applicable SNPs to the gene, respectively, as described in Table 2.

[0277] For each gene, the allele of interest can only be found in one of the four parental backgrounds that are described in Table 8. The strain identified for carrying the allele of interest is highlighted bold and underlined in Table 4. In Table 4, strains disrupted in their relevant allele are denoted by the suffix -K. Strains transformed with the expression vector pEV7 are denoted by the suffix -L.

[0278] In some cases, haploid disruption of the relevant gene resulted in inviable strains, which could not be included in the dataset. In some cases, mating of the haploid strains would fail and the resulting diploid could therefore not be included in the dataset. Consequentially, Table 4 does not show the reciprocal pairings of WA / WE for YLR287C and SA / WA, WE / WA for UBP14. In Table 4, if the two isogenic diploid strains differing only by the allele expressed differ in outcome, then this validates that the variation at the gene of interest causes a difference in the phenotypic outcome.

[0279] As can be seen in Table 4, alleles of interest resulted in statistically significant fold- change improvements (values >1) in OD normalised mCherry fluorescence compared to isogenic strains without the allele of interest (** denotes p>0.01, **** denotes p>0.0001).

[0280] The skilled person will appreciate that reciprocal hemizygosity experiments validated the effectiveness of the alleles of interest to increase recombinant protein production and / or reduce proteolysis levels. Furthermore, the skilled person will appreciate that the phenotypic effects of the allele of interest that were measured for MEC3 (YLR288C), YLR287C, GAT1 (YFL021W) are independent of strain background and the positive properties of alleles of interest are therefore not a strain-specific invention. Example 4 - Improved Production of Multiple Protein Types

[0281] A comparison of one of the best two strains for amylase-mCherry secretion (2-A2) with one of the worst two strains for amylase-mCherry secretion (1-C12) was performed for multiple different recombinant proteins. Nine additional recombinant proteins were expressed, which were diverse in structure, size, and other physiochemical properties (Table 5). In all cases except one, the best strain for amylase mCherry production gave higher production for the other recombinant proteins (Table 6).

[0282] For amylase-mCherry control, 2-A2 gave approximately 8.6 times more amylase- mCherry based on mCherry fluorescence than 1-C12, which is statistically consistent with the results in Figure 2. For the other proteins, the fold increase was between approximately 4.2 and 1.2, with one protein giving approximately equal productivity between the two strains. In this case, HPV16(Ll)-mCherry production is very different to amylase-mCherry, so it is not unexpected that alleles beneficial to amylase-mCherry production that are present in strain 2-A2 would improve HPV16(L1) production as significantly as for the other recombinant proteins because the HPV16(Ll)-mCherry was expressed intracellularly for accumulation and VLP formation in the nucleus, whereas amylase-mCherry was expressed for secretion into the extracellular media. For all the other proteins, which had a range of detection tags and assay methods and were either secreted or expressed intracellularly for cytosolic accumulation, the alleles in 2-A2 were beneficial for increased productivity and / or reduced proteolysis. The recombinant proteins expressed have a diverse range of sizes, folding, domain structures and other physiochemical structures, indicating that strain 2-A2 is also generally improved to produce many other recombinant proteins of interest. Multiple alleles beneficial for recombinant protein production and / or reduced proteolysis originating from the different parental strains have been combined in strain 2-A2. While this combination of alleles was originally selected for improved production of amylase- mCherry, clearly many of these alleles and other combinations of these alleles and the SNPs within them are also beneficial for the production of multiple other recombinant proteins. For expression of multiple different types of recombinant protein, Strains 2-A2 and 1-C12 were cured of the whole-2-micron expression vector for amylase-mCherry (pHR!K-pEV7) by the method described above and retransformed for expression from whole-2-micron plasmids equivalent to pHRIK of multiple different recombinant proteins comprising the expression cassettes described in Table 5. All final whole-2-micron expression plasmids contain a LEU2 gene for leucine selection. Transformants were isolated as colonies on synthetic drop-out agar lacking leucine, e.g. BMMD+(CSM-Leu+Met). Three transformants were selected for each strain / plasmid combination for expression studies. Controls were strains 2-A2 and 1-C12 secreting amylase-mCherry from the pEV7 expression cassette in pHRlK- pEV7.

[0283] Inoculum cultures for three transformants of 2-A2 and 1-C12 for each plasmid (and triplicate pHRlK-pEV7 controls) were started by picking cells from patches on solid media and transferring to 500pL liquid cultures in clear, 48 well microtiter plates.

[0284] Buffered synthetic drop-out media with 2%(w / v) dextrose, lacking leucine and containing 3g / L methionine was used to maintain plasmids and to repress expression from constructs utilising the MET17 promoter. Inoculum cultures were incubated at 30°C for 2 days after which 20pL of each inoculum culture was passaged into 500pL synthetic dropout media with 2%(w / v) dextrose, lacking leucine in new 48 well microtiter plates, e.g. BMMD+(CSM-Leu-Met), in triplicate, to inoculate expression cultures. Expression cultures were incubated in shaking humidity chambers at 30°C over 4 days before harvesting. Upon harvest, culture OD was measured in wells using a TECAN Spark plate reader (Tecan, Switzerland). Culture supernatants were isolated by centrifugation at 1800 RCF and analysed for secreted product, where applicable.

[0285] For strains transformed with pEVl, pEV51 and pEV52 for the expression of intracellular mCherry and mCherry-tagged recombinant proteins, mCherry fluorescence was measured directly from the expression cultures in 48 well clear microtiter plates upon harvest at Aex540nm; Aem614nm, gain 100 on a TECAN Spark plate reader (TECAN, Switzerland).

[0286] For strains transformed with pEV3, pEV7 and pEV298 expression constructs for secreted expression mCherry-tagged recombinant proteins, 200pL culture supernatant was isolated as described previously and transferred to new, clear, 48 well microtiter plates. mCherry fluorescence of supernatants was measured at Aex540nm; Aem614nm, gain 100 on a TECAN Spark plate reader. OD was also measured to check for any accidental transfer of the cell pellet.

[0287] For strains transformed with pEV388 for the secreted expression of recombinant human albumin (rHA), titres were quantified using the Albumin Blue Fluorescence Assay Kit (Active Motif, Belgium) with a modified protocol for high-throughput detection in 384 plates. Briefly, 12.5pL culture supernatant was transferred to wells in a black, clear bottomed, non-treated 96 well assay plate. 75pL assay reagent comprising Ipl Albumin Blue dye and 74pL Buffer A from the kit was added to each well and mixed by pipetting. Plates were incubated for 5 minutes at room temperature, then fluorescence at Aex560nm; Aem620nm was measured on a TECAN Spark plate reader. Fluorescence signals for each sample were averaged from three technical replicates and converted to relative levels using a standard curve of rHA prepared in expression medium.

[0288] For strains transformed with pEV395 for the secreted expression of HiBit-tagged GLP1 analogue precursor, titres were quantified using the Hi-Bit (HiBit) extracellular detection kit (Promega, US), according to the manufacturer's instructions. A standard curve was made from a HiBit-tagged control protein (Promega, US) of known concentrations prepared in expression medium. Supernatant samples were diluted 1 / 104to generate a signal within the linear range of the standard curve. Reactions were conducted in 20pL final volumes (lOpL sample, lOpL assay mix) in white, 384 well low-volume assay plates. Luminescence was measured on a BMG FLUOstar Omega plate reader (BMG Labtech, Germany). Luminescence signals for each sample were averaged from 3 technical replicates and converted to relative levels using the HiBit control protein standard curve.

[0289] For strains transformed with pEV299 and pEV275 for the secreted expression of untagged VHH, titres were quantified by SDS PAGE. Supernatant samples were run on NuPAGE 4-12% Bis-Tris gels (Thermo Fisher Scientific, US) according to the manufacturer's instructions, alongside 3 prepared samples of a purified VHH standard at known concentrations. Gels were Coomassie stained and imaged on an Amersham ImageQuant 800 (Cytiva, US), and densitometry analysis of bands corresponding to the VHH samples was conducted using ImageQuantTL (Cytiva, US). Band intensity values were converted to estimated relative levels using values obtained from the VHH reference standard. All data was for triplicates corrected for growth / biomass (ODeoo or OD620).

[0290] Accordingly, the inventors have demonstrated that yeast cells comprising the alleles and SNPs of interest are able to reduce proteolysis of multiple different types of protein, including both secreted and intracellular proteins.

[0291] Conclusions

[0292] Using QTL analysis, the inventors have identified 16 genomic regions comprising approximately 3.3% of the total Saccharomyces cerevisiae genome containing genes and alleles responsible for the differential expression of the recombinant amylase-mCherry protein. Advantageously, by identifying these genomic regions, the inventors have identified specific genes, and SNPs within these genes, that are associated with modified proteolysis. As such, improved strains for recombinant protein manufacture can be provided, exhibiting reduced proteolysis, resulting in improved product yields.

[0293] References

[0294] 1. Louis, E. J. Historical Evolution of Laboratory Strains of Saccharomyces cerevisiae. Cold Spring Harb. Protoc. 2016, (2016).

[0295] 2. Sleep, D. & Finnis, C. 2-MICRON FAMILY PLASMID AND USE THEREOF, WO 2005 / 061719 Al. (2005).

[0296] 3. Strope, P. K. et al. 2p plasmid in Saccharomyces species and in Saccharomyces cerevisiae. FEMS Yeast Res. 15, (2015). 4. Cubillos, F. A., Louis, E. J. & Liti, G. Generation of a large set of genetically tractable haploid and diploid Saccharomyces strains. FEMS Yeast Res. 9, 1217-1225 (2009).

[0297] 5. Louvel, H., Gillet-Markowska, A., Liti, G. & Fischer, G. A set of genetically diverged Saccharomyces cerevisiae strains with markerless deletions of multiple auxotrophic genes. Yeast 31, 91-101 (2014). 6. Rose, A. B. & Broach, J. R. Propagation and expression of cloned genes in yeast: 2- microns circle-based vectors. Methods Enzymol. 185, 234-279 (1990).

[0298] 7. Chinery, S. A. 8i Hi nchliffe, E. A novel class of vector for yeast transformation. Curr. Genet. 16, 21-25 (1989).

[0299] 8. Thorn, K. Genetically encoded fluorescent tags. Mol. Biol. Cell 28, 848-857 (2017). 9. Kaishima, M., Ishii, J., Matsuno, T., Fukuda, N. 8i Kondo, A. Expression of varied

[0300] GFPs in Saccharomyces cerevisiae: codon optimization yields stronger than expected expression and fluorescence intensity. Sci. Rep. 6, 35932 (2016).

[0301] 10. Chu, D. et al. Translation elongation can control translation initiation on eukaryotic mRNAs. EMBO J. 33, 21-34 (2014). 11. Solow, S. P., Sengbusch, J. 8i Laird, M. W. Heterologous protein production from the inducible MET25 promoter in Saccharomyces cerevisiae. Biotechnol. Prog. 21, 617-620 (2005).

[0302] 12. Andersen, J. T. et al. Structure-based mutagenesis reveals the albumin-binding site of the neonatal Fc receptor. Nat. Commun. 3, 610 (2012). 13. Finnis, C., Nordeide, P. & McLaughlan, J. IMPROVED PROTEIN EXPRESSION STRAINS,

[0303] WO 2018 / 234349 Al. (2018).

[0304] 14. Cubillos, F. A. et al. High-resolution mapping of complex traits with a four-parent advanced intercross yeast population. Genetics 195, 1141-1155 (2013).

[0305] 15. Goldstein, A. L. 8i McCusker, J. H. Three new dominant drug resistance cassettes for gene disruption in Saccharomyces cerevisiae. Yeast 15, 1541-1553 (1999). 16. Evans, L. et al. The production, characterisation and enhanced pharmacokinetics of scFv-albumin fusions expressed in Saccharomyces cerevisiae. Protein Expr. Purif. 73, 113- 124 (2010).

[0306] 17. Schelde, K. K. et al. A new class of recombinant human albumin with multiple surface thiols exhibits stable conjugation and enhanced FcRn binding and blood circulation. J. Biol.

[0307] Chem. 294, 3735-3743 (2019).

[0308] 18. Ramaiya, P., Finnis, C., McLaughlan, J. 8i Nordeide, P. IMPROVED PROTEIN EXPRESSION STRAINS, WO 2017 / 112847 Al. (2017).

[0309] 19. Miles, C.; Wayne, M. Quantitative Trait Locus (QTL) Analysis. Nat. Educ. 1, 208 (2008).

[0310] 20. Chatalic KL, Veldhoven-Zweistra J, Bolkestein M, Hoeben S, Koning GA, Boerman OC, de Jong M, van Weerden WM. A Novel11 :LIn-Labeled Anti-Prostate-Specific Membrane Antigen Nanobody for Targeted SPECT / CT Imaging of Prostate Cancer. J Nucl Med. 2015 Jul;56(7) : 1094-9 (2015). 21. Mackay TF, Stone EA, Ayroles JF (2009) The genetics of quantitative traits: challenges and prospects. Nat Rev Genet 10: 565-577.

[0311] 22. Cubillos, Francisco A, Billi, Eleonora, Zbrgb, Enikb, Parts, Leopold, Fargier, Patrick, Omholt, Stig, Blomberg, Anders, Warringer, Jonas, Louis, Edward J and Liti, Gianni, 2011. Assessing the complex architecture of polygenic traits in diverged yeast populations. Molecular Ecology 20: 1401-1413.

[0312] 23. Leopold Parts, Francisco A. Cubillos, Jonas Warringer, Kanika Jain, Francisco Salinas, Suzannah J. Bumpstead, Mikael Molin, Amin Zia, Jared Simpson, Michael A. Quail, Alan Moses, Edward J. Louis, Richard Durbin and Gianni Liti. 2011. Revealing the genetic structure of a trait by sequencing a population under selection. Genome Research 21: 1131-1138. 24. Gianni Liti and Edward J Louis. 2012. Advances in Quantitative Trait Analysis in

[0313] Yeast. PLoS Genetics 8(8) : el002912

Claims

Claims1. A recombinant or engineered eukaryotic cell exhibiting reduced proteolysis compared to a corresponding wild-type, progenitor, or common laboratory strain cell, the recombinant or engineered eukaryotic cell comprising a non-naturally occurring combination of alleles associated with reduced proteolysis, wherein the at least one allele is for a gene selected from a group consisting of: GAT1 (YFL021W), UBP14 (YBR058C), and a gene present in Saccharomyces cerevisiae chromosome XII 706198 - 717026 inclusive (SEQ ID No: 11), or a homologue, orthologue or paralogue thereof, wherein the at least one gene is associated with proteolysis.

2. A recombinant or engineered eukaryotic cell exhibiting reduced proteolysis compared to a corresponding wild-type, progenitor, or common laboratory strain cell, the recombinant or engineered eukaryotic cell comprising at least one gene with a non-naturally occurring combination of single nucleotide polymorphisms (SNPs), wherein the at least one gene is selected from a group consisting of: GAT1 (YFL021W), UBP14 (YBR058C), and a gene present in Saccharomyces cerevisiae chromosome XII 706198 - 717026 inclusive (SEQ ID No: 11), or a homologue, orthologue or paralogue thereof, wherein the at least one gene is associated with proteolysis.

3. A recombinant or engineered eukaryotic cell exhibiting reduced proteolysis compared to a corresponding wild-type, progenitor, or common laboratory strain cell, the recombinant or engineered eukaryotic cell comprising at least one modified gene, wherein the at least one gene is selected from a group consisting of: GAT1 (YFL021W), UBP14 (YBR058C), and a gene present in Saccharomyces cerevisiae chromosome XII 706198 - 717026 inclusive (SEQ ID No: 11), or a homologue, orthologue or paralogue thereof, wherein the at least one gene is associated with proteolysis.

4. A recombinant or engineered eukaryotic cell according to any one of the preceding claims, wherein the at least one gene is GAT1 (YFL021W) and / or UBP14 (YBR058C).

5. A recombinant or engineered eukaryotic cell according to any one of claims1 to 3, wherein the at least one gene is selected from a group consisting of:YLR284C, YLR285W, YLR285C-A, YLR286C, YLR286W-A, YLR287C, YLR287C-A, YLR288C, and YLR289W, or a homologue, orthologue or paralogue thereof.

6. A recombinant or engineered eukaryotic cell according to any one of claims 1 to 3, wherein the at least one gene is MEC3 (YLR288C) and / or YLR287C, or a homologue, orthologue or paralogue thereof.

7. A recombinant or engineered eukaryotic cell according to any one of claims1 to 3, wherein the at least one gene is selected from a group consisting of: NNT1 (YLR285W), CTS1 (YLR286C), YLR287C, RPS30A (YLR287C-A), MEC3 (YLR288C) and GUF1 (YLR289W), or a homologue, orthologue or paralogue thereof.

8. A recombinant or engineered eukaryotic cell according to any one of claims1 to 3 or 7, wherein the at least one gene is CTS1 (YLR286C) and / or RPS30A (YLR287C-A), or a homologue, orthologue or paralogue thereof.

9. A recombinant or engineered eukaryotic cell according to any one of the preceding claims, wherein the at least one gene is at least two, at least three, or at least four genes selected from a group consisting of: GAT1 (YFL021W), UBP14 (YBR058C), MEC3 (YLR288C), and YLR287C, or a homologue, orthologue, or paralogue thereof.

10. A recombinant or engineered eukaryotic cell according to any one of the preceding claims, wherein: (i) when the at least one gene is GAT1, the gene comprises the SNP G647A;(ii) when the at least one gene is UBP14, the gene comprises the SNP C2244A;(iii) when the at least one gene is UBP14, the gene comprises the SNP A2129G; (iv) when the at least one gene is UBP14, the gene comprises the SNPT2086C;(v) when the at least one gene is UBP14, the gene comprises the SNP A1591G;(vi) when the at least one gene is UBP14, the gene comprises the SNP G1298A;(vii) when the at least one gene is UBP14, the gene comprises the SNPT217C;(viii) when the at least one gene is MEC3, the gene comprises the SNP T1362G;(ix) when the at least one gene is MEC3, the gene comprises the SNP C359T; (x) when the at least one gene is YLR287C, the gene comprises the SNPG19A;(xi) when the at least one gene is YLR287C, the gene comprises the SNP A1001G;(xii) when the at least one gene is YLR287C, the gene comprises the SNP G992A;(xiii) when the at least one gene is YLR287C, the gene comprises the SNP T722C;(xiv) when the at least one gene is RPS30A, the gene comprises the SNP G148A; (xv) when the at least one gene is CTS1, the gene comprises the SNP T47C;(xvi) when the at least one gene is CTS1, the gene comprises the SNPC962G;(xvii) when the at least one gene is CTS1, the gene comprises the SNP G1570A; (xviii) when the at least one gene is NNT1, the gene comprises the SNPG413C; and / or(xix) when the at least one gene is GUF1, most preferably the gene comprises the SNP C779T.

11. A recombinant or engineered eukaryotic cell according to any one of the preceding claims, wherein the recombinant or engineered eukaryotic cell comprises: (i) at least one additional non-naturally occurring allele for at least one additional gene; (ii) at least one additional gene with a non-naturally occurring combination of single nucleotide polymorphisms (SNPs); and / or (iii) at least one additional modified gene, wherein the at least one additional gene is selected from a list of genes in Table 1, or a homologue, orthologue or paralogue thereof, wherein the at least one additional gene is associated with proteolysis.

12. A recombinant or engineered eukaryotic cell according to claim 11, wherein the at least one additional gene is at least two, at least three, at least four, or at least five genes selected from the list of genes in Table 1, or a homologue, orthologue, or paralogue thereof.

13. A recombinant or engineered eukaryotic cell according to claim 11 or claim12, wherein the at least one additional gene is at least six, at least seven, at least eight, at least nine or at least ten genes selected from the list of genes in Table 1, or a homologue, orthologue, or paralogue thereof.

14. A recombinant or engineered eukaryotic cell according to any one of claims 11-13, wherein the at least one additional gene is selected from a group consisting of: UBP3, SPL2, AVT7, YIL108W, FYV10, WSS1, GPB1, RPN2, and RAD4.

15. A recombinant or engineered eukaryotic cell according to any one of claims 11-14, wherein the at least one additional gene is selected from a group consisting of: UBP3, SPL2 and AVT7.

16. A recombinant or engineered eukaryotic cell according to any one of claims11-15, wherein the at least one additional gene is selected from a group consisting of: YIL108W, FYV10, WSS1, GPB1, RPN2, and RAD4.

17. A recombinant or engineered eukaryotic cell according to any one of claims 14-16, wherein:(i) when the at least one additional gene is UBP3, the gene comprises the SNP G770A;(ii) when the at least one additional gene is UBP3, the gene comprises the SNP G643; (iii) when the at least one additional gene is UBP3, the gene comprises theSNP A230C;(iv) when the at least one additional gene is SPL2, the gene comprises theSNP G98C; and / or(v) when the at least one additional gene is AVT7, the gene comprises the SNP A1148T.

18. A recombinant or engineered eukaryotic cell according to any one of claims 14-17, wherein:(i) when the at least one additional gene is WSS1, the gene comprises the SNP A85G;(ii) when the at least one additional gene is WSS1, the gene comprises theSNP G250A; and / or(iii) when the at least one additional gene is GPB1, the gene comprises the SNP C881T.

19. A recombinant or engineered eukaryotic cell according to any preceding claim, wherein the recombinant or engineered eukaryotic strain does not comprise a combination of SNPs associated with increased proteolysis.

20. A recombinant or engineered eukaryotic cell according to any preceding claim, wherein the eukaryotic cell is a fungal cell, or a yeast cell, optionally wherein the engineered or recombinant eukaryotic cell has a modified or disrupted PEP4 gene, or a homologue, orthologue or paralogue thereof.

21. A recombinant or engineered eukaryotic cell according to claim 20, wherein the yeast cell is Pichia pastoris, optionally a Komagataella species, or Hansenula polymorpha, Kluyveromyces lactis, a Yarrowia species, or Schizosaccharomyces pombe.

22. A recombinant or engineered eukaryotic cell according to claim 20, wherein the yeast cell is a Saccharomyces species yeast, preferably Saccharomyces cerevisiae.

23. A recombinant or engineered eukaryotic cell according to claim 20, wherein the fungal cell is an Aspergillus species, optionally Aspergillus oryzae or Aspergillus niger, or a Trichoderma species, or Myceliophthora thermophila.

24. A recombinant or engineered eukaryotic cell according to any one of claims1 to 19, wherein:(i) the eukaryotic cell is an insect cell, optionally a Sf9 or Sf21 cell line from Spodoptera frugiperda cells, Hi-5 from Trichoplusia ni cells, or Schneider 2 cells or Schneider 3 cells from Drosophila melanogaster cells; or(ii) the eukaryotic cell is from Excavata, optionally Leishmania tarentolae.

25. A recombinant or engineered eukaryotic cell according to any one of claims1 to 19, wherein the eukaryotic cell is a mammalian cell type, optionally a Chinese hamster ovary (CHO) cell, a Mouse myeloma lymphoblastoid, a Human embryonic kidney cell, a Human embryonic retinal cell, or a Human amniocyte cell.

26. A recombinant or engineered eukaryotic cell according to any preceding claim, wherein the recombinant or engineered eukaryotic cell exhibits a reduction in proteolysis compared to a wild-type, progenitor, or common laboratory strain cell of at least 5%, at least 10%, at least 20%, at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, or at least 100%.

27. Use of the recombinant or engineered eukaryotic cell according to any one of claims 1 to 26, for producing a recombinant protein.

28. A method for producing a recombinant protein, the method comprising:(i) transforming the recombinant or engineered eukaryotic cell or obtaining a transformed recombinant or engineered eukaryotic cell according to any one of claims 1 to 26 with an expression vector encoding at least one recombinant protein; and (ii) culturing the cell in a medium under conditions to produce the recombinant protein.

29. A recombinant protein obtained from the recombinant or engineered eukaryotic cell according to any one of claims 1 to 26.