Production of peptides using tandem repetitive recombination strategy

WO2025137687A3PCT designated stage expired Publication Date: 2025-08-07CONAGEN INC
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
PCT/US2024/061682
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-12-22
Filing Date
2024-12-23
Publication Date
2025-08-07

AI Technical Summary

Technical Problem

Current methods for producing short peptides are inefficient and lack the ability to produce bioactive peptides with high yield and purity, particularly in genetically engineered organisms like Pichia and E. coli.

Method used

The development of a tandem repetitive recombinant protein strategy using codon-optimized nucleic acid sequences, optimized gene spacers, and signal peptides, combined with co-expression of chaperones and optimized protease cleavage conditions, to enhance the production of short peptides in Pichia pastoris and E. coli.

Benefits of technology

This approach significantly improves the production efficiency and yield of bioactive short peptides, making it suitable for applications in flavor and cosmetics industries, and demonstrates high efficiency in genetically engineered organisms.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2024061682_07082025_PF_FP_ABST
    Figure US2024061682_07082025_PF_FP_ABST
Patent Text Reader

Abstract

A strategy to produce short peptides through tandem repetitive recombinant protein has been developed as disclosed herein. Furthermore, the production of tandem repetitive recombinant proteins in Pichia pastoris and optimized protease cleavage conditions for purification of bioactive peptides has been improved using techniques disclosed herein. Provided herein is the production of flavor and cosmetics peptides, and the technology is applied to production of various short peptides biologically. These techniques incorporate novel applications of signal peptides, design spacers for generating tandem repetitive recombinant protein and increasing copy number, co-expression of chaperones and optimization of protease cleavage conditions. Specifically, provided herein are methods using designed tandem repeated genes, optimized gene spacers and signal peptides useful for improved production of short peptide in Pichia and E. coli. These novel applications afford improvements in short peptide production and high efficiency biological production of bioactive peptides.
Need to check novelty before this filing date? Find Prior Art

Description

PRODUCTION OF PEPTIDES USING TANDEM REPETITIVERECOMBINATION STRATEGYCROSS REFERENCE TO RELATED APPLICATIONS

[0001] This application claims the benefit of U.S. Provisional Application No. 63 / 613,846, filed December 22, 2023, the contents of which is incorporated by reference herein in its entirety.INCORPORATION OF SEQUENCE LISTING XML

[0002] A computer readable form of the Sequence Listing XML containing the file named "CONAG3520.WO Sequence Listing. xml," which is 61,429 bytes in size (as measured in MICROSOFT WINDOWS® EXPLORER) and was created on December 17, 2024, is provided herein and is herein incorporated by reference. This Sequence Listing consists of SEQ ID NOs: 1-42.FIELD OF INVENTION

[0003] The present disclosure provides constructs and methods of synthetic biology and genetic engineering, and a strategy to produce short peptides through tandem repetitive recombinant protein. Furthermore, the production of tandem repetitive recombinant proteins in Pichia pastoris and optimized protease cleavage conditions for purification of bioactive peptides has been improved using techniques disclosed herein. Provided herein is the production of flavor and cosmetics peptides, and the technology is applied to production of various short peptides biologically. These techniques incorporate novel applications of signal peptides, design spacers for generating tandem repetitive recombinant protein and increasing copy number, coexpression of chaperones and optimization of protease cleavage conditions. Specifically, provided herein are methods using designed tandem repeated genes, optimized gene spacers and signal peptides useful for improved production of short peptide in Pichia and E. coli. These novel applications afford improvements in short peptide production and high efficiency biological production of bioactive peptides. Such improvements are applicable in genetically engineered organisms with specific application in Pichia strains.BACKGROUND OF INVENTION

[0004] Proteins containing internal repetitions are found in nature. Such protein repeats possess regular secondary structures and form multi-repeat assemblies in three dimensions of diverse sizes and functions. Repeat proteins are formed by the repeating of modular units of protein substructures. The structure of the repeat proteins as a whole is dictated by the internal geometry of the protein and the local packing of the repeat modular units. These features are generated by underlying patterns of amino acid sequences that are repetitive.

[0005] Repeat proteins function in nature as macromolecular binding and scaffolding domains, enzymes, and building blocks for the assembly of fibrous materials. The structure and identity of these repeat proteins are highly diverse, ranging from extended, super-helical folds that bind peptide, DNA, and RNA partners, to closed and compact conformations with internal cavities suitable for small molecule binding and catalysis.

[0006] Tandem repeat protein (“TRPs”) are proteins containing multiple repeated peptide sequences. TRPs form highly similar folded topologies that selfassociate to form closed circular or open linear protein scaffolds. Naturally occurring TRPs can bind to a variety of protein and nucleic acid partners. TRPs can be redesigned and often display several favorable biophysical properties, including high thermal stability and solubility, and ease of expression. TRPs have been used as scaffolds to host various protein domains. These TRPs can be expressed at very high levels, can be thermostable and soluble, and form structures that corresponded closely to their original designs.

[0007] A DNA assembly method that incorporates design of tandem repeat proteins can provide improved methods for expressing the desired proteins in various cell lines and improved methods are needed to increase expression and efficiency of production of TRPs.SUMMARY OF INVENTION

[0008] The present disclosure provides a codon optimized nucleic acid sequence encoding a tandem repeat amino acid sequence corresponding to an identified peptide.

[0009] Another aspect of the disclose is a plasmid comprising a nucleic acid sequence encoding a tandem repeat amino acid sequence corresponding to an identified peptide.

[0010] Yet another aspect of the disclosure is a plasmid comprising a nucleic acid sequence of any one or more of SEQ ID NO: 20, 22, 24, 26, 28, 30, 32, 34, 36, 38, 40, or 42 or a nucleotide sequence with at least 85%, 90%, 95%, 99%, or 99.9% sequence homology to any one or more of SEQ ID NO: 20, 22, 24, 26, 28, 30, 32, 34, 36, 38, 40, or 42.

[0011] A further aspect of the disclosure is a recombinantly produced tandem repeat amino acid sequence corresponding to an identified peptide.

[0012] Yet another aspect is an E. coli strain comprising one or more of the plasmids described herein, one or more of the recombinantly produced tandem repeat amino acid sequences described herein, or a combination thereof.

[0013] A further aspect is a Pichia pastoris strain comprising one or more of the plasmids described herein, one or more of the recombinantly produced tandem repeat amino acid sequences described herein, or a combination thereof.

[0014] Another aspect of the disclosure is a method of DNA assembly comprising identifying a peptide and forming a tandem repeat amino acid sequence corresponding to the identified peptide; optimizing the codon for expressing the tandem repeat amino acid sequence for expression in E. coli or Pichia pastoris to form the encoding fragment; forming a middle fragment by adding Bsal restriction sites to the 5' and 3' ends of the final sequence with four base overhangs corresponding to the last two bases and first two bases of the encoding fragment; generating a start fragment by amplifying the middle fragment with a 5' forward primer by adding a Bsal site for the last four bases of a mating factor alpha signal peptide; and generating an end fragment by using a reverse primer that added a stop codon and four bases after the stop codon, thereby forming a vector comprising a start fragment, an encoding fragment, a middle fragment, and an end fragment.

[0015] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains. Although methods and materials similar or equivalent to those described herein can be used to practice the invention, suitable methods and materials are described below. All publications, patent applications, patents, and other references mentioned herein are incorporated by reference in their entirety. In case ofconflict, the present specification, including definitions, will control. In addition, the materials, methods, and examples are illustrative only and not intended to be limiting.

[0016] The details of one or more embodiments of the invention are set forth in the accompanying drawings and the description below. Other features, objects, and advantages of the invention will be apparent from the description and drawings, and from the claims.BRIEF DESCRIPTION OF THE DRAWINGS

[0017] Figure 1 depicts SDS-PAGE analysis of medium samples from induced culture. M: standard ladder (17, 11 kDa from top to bottom); 1-9: PP9TR13; 10-14: PP9TR25. Arrows show GHK tandem repeat bands.

[0018] Figure 2 depicts SDS-PAGE analysis of medium samples from induced culture. 1 : PP9TR25x2, 2: PP9TR25xl, 3 PP9TR25x4.

[0019] Figure 3A-B depict LC MS analysis. Figure 3A depicts PP9 standard. Figure 3B depicts trypsin treated media sample from pHKA-PP9TR25x4 strain.

[0020] Figure 4 depicts SDS-PAGE analysis of medium samples from induced culture. 1 : PP22TR18xl, 2: PP22TR18x2, M: Protein Standard (250,180, 130, 95,72,55,43,34,26,17,11k Da from top to bottom).

[0021] Figures 5A-D depict HPLC analysis. Figure 5 A depicts 0.1 g / L PP22 standard. Figure 5B depicts uncut PP22TR18 (0 minutes digestion). Figure 5C depicts partially digested PP22TR18. Figure 5D depicts fully digested PP22TR18.

[0022] Figure 6A-6B depict LC MS analysis. Figure 6A depicts PP22 standard. Figure 6B depicts chymotrypsin treated media sample from pHKA-PP22TRl 8x4 strain.

[0023] Figure 7 depicts SDS-PAGE analysis of medium samples from induced culture. 1 : PP35TR20x4, M: Protein Standard (130, 95,72,55,43,34,26,17,11 kDa from top to bottom).

[0024] Figure 8A-B depict LC MS analysis. Figure 8 A depicts PP35 standard. Figure 8B depicts chymotrypsin treated media sample from pHKA-PP35TR20x4 strain.

[0025] Figure 9 depicts SDS-PAGE analysis of medium samples from induced culture. 1 : PP36TR12x4, M: Protein Standard (26.6, 17, 14.2, 6.5, 3.5, 1.06 kDa from top to bottom). Arrow shows expressed PP36 tandem repetitive recombinant protein.

[0026] Figure 10A-B depicts LC MS analysis. Figure 10A depicts PP36 standard. Figure 10B depicts trypsin treated media sample from pHKA-PP36TR12x4 strain.

[0027] Figure 11 depicts co-expression of PpPDIl and PpKAR2 in PP22TR18x4 strain increased PP22TR18 production. pHK-PP22TR18x4: parent strain, PDI1-KAR2 1, 2, 3: individual colonies from transformation of pPICZ- PpPDIl / PpKAR2 into PP22TR18x4 strain.

[0028] Figure 12 depicts fermentation samples diluted 1 :30 and run on SDS- PAGE. Lane 1 : Oh induction, 2: 21.5 h, 3: 44.5 h, 4: 71.5 h, M: Protein Standard (250,180, 130, 95,72,55,43,34,26,17,llkDa from top to bottom).

[0029] Figure 13 depicts flavor enhancing profile of PP22 (Salatide). The circle indicates significant differences at 95% Cl.DETAILED DESCRIPTION OF INVENTION

[0030] This disclosure describes various peptides, methods for production of the peptides biologically, and uses of the peptides for flavor and in cosmetics. For example, short peptides are often chemically synthesized. In order to generate technology for production of native bioactive peptides with synthetic biology and genetic engineering techniques, a strategy to produce short peptides through tandem repetitive recombinant protein has been developed as disclosed herein. Furthermore, the production of tandem repetitive recombinant proteins in Pichia pastoris and optimized protease cleavage conditions for purification of bioactive peptides has been improved using techniques disclosed herein. Provided herein is the production of flavor and cosmetics peptides, and the technology is applied to production of various short peptides biologically. These techniques incorporate novel applications of signal peptides, design spacers for generating tandem repetitive recombinant protein and increasing copy number, co-expression of chaperones and optimization of protease cleavage conditions.

[0031] Specifically, provided herein are methods using designed tandem repeated genes, optimized gene spacers and signal peptides useful for improved production of short peptides in Pichia and E. coli. These novel applications afford improvements in short peptide production and high efficiency biological production of bioactive peptides. Such improvements are applicable in genetically engineered organisms with specific application in Pichia strains.

[0032] Another aspect of the disclosure is a codon optimized nucleic acid sequence encoding a tandem repeat amino acid sequence corresponding to an identified peptide.

[0033] For example, the codon optimized nucleic acid sequences described herein can be optimized for E. coli or Pichia pastoris expression.

[0034] Also, the codon optimized nucleic acid sequences can comprise SEQ ID NOs: 2, 6, 10,14, or 18 or a nucleic acid sequence with at least 85%, 90%, 95%, 99%, or 99.9% sequence homology to any one of SEQ ID NOs: 2, 6, 10, 14, or 18.

[0035] Preferably, the nucleic acid sequences can comprise SEQ ID NO: 2 or a nucleic acid sequence with at least 85%, 90%, 95%, 99%, or 99.9% sequence homology to SEQ ID NO: 2.

[0036] Also, the nucleic acid sequences disclosed herein can comprise SEQ ID NO: 6 or a nucleic acid sequence with at least 85%, 90%, 95%, 99%, or 99.9% sequence homology to SEQ ID NO: 6.

[0037] Further, the nucleic acid sequences can comprise SEQ ID NO: 10 or a nucleic acid sequence with at least 85%, 90%, 95%, 99%, or 99.9% sequence homology to SEQ ID NO: 10.

[0038] The nucleic acid sequences described herein can comprise SEQ ID NO: 14 or a nucleic acid sequence with at least 85%, 90%, 95%, 99%, or 99.9% sequence homology to SEQ ID NO: 14.

[0039] Yet other nucleic acid sequences disclosed herein can comprise SEQ ID NO: 18 or a nucleic acid sequence with at least 85%, 90%, 95%, 99%, or 99.9% sequence homology to SEQ ID NO: 18.

[0040] Additionally, the identified peptide can comprise an amino acid sequence corresponding to GHK or to any one of SEQ ID NOs: 3, 7, 11, or 15 or an amino acid sequence with at least 85%, 90%, 95%, 99%, or 99.9% sequence homology to SEQ ID NOs: 3, 7, 11, or 15.

[0041] Another aspect of the disclosure is a plasmid comprising a nucleic acid sequence encoding a tandem repeat amino acid sequence corresponding to an identified peptide.

[0042] For example, the plasmids disclosed herein can be expressed in E. coli or Pichia pastoris.

[0043] Also, the plasmids can have the nucleic acid sequence comprise any one of SEQ ID NOs: 2, 6, 10,14, or 18 or a nucleic acid sequence with at least 85%, 90%, 95%, 99%, or 99.9% sequence homology to any one of SEQ ID NOs: 2, 6, 10, 14, or 18.

[0044] Further, the plasmids described herein can have the nucleic acid sequence comprise SEQ ID NO: 2 or a nucleic acid sequence with at least 85%, 90%, 95%, 99%, or 99.9% sequence homology to SEQ ID NO: 2.

[0045] The plasmids can also have the nucleic acid sequence comprise SEQ ID NO: 6 or a nucleic acid sequence with at least 85%, 90%, 95%, 99%, or 99.9% sequence homology to SEQ ID NO: 6.

[0046] Additionally, the plasmids can have the nucleic acid sequence comprise SEQ ID NO: 10 or a nucleic acid sequence with at least 85%, 90%, 95%, 99%, or 99.9% sequence homology to SEQ ID NO: 10.

[0047] The plasmids described herein can have the nucleic acid sequence comprises SEQ ID NO: 14 or a nucleic acid sequence with at least 85%, 90%, 95%, 99%, or 99.9% sequence homology to SEQ ID NO: 14.

[0048] The plasmids can have the nucleic acid sequence comprises SEQ ID NO: 18 or a nucleic acid sequence with at least 85%, 90%, 95%, 99%, or 99.9% sequence homology to SEQ ID NO: 18.

[0049] Also, the plasmids can further comprise the nucleic acid sequence of any one of SEQ ID NOs: 20, 22, 24, 26, 28, 30, 32, 34, 36, 38, 40, or 42 or a nucleic acid sequence with at least 85%, 90%, 95%, 99%, or 99.9% sequence homology to any one of SEQ ID NOs: 20, 22, 24, 26, 28, 30, 32, 34, 36, 38, 40, or 42.

[0050] Preferably, the plasmids further comprise the nucleic acid sequence of any one of SEQ ID NOs: 20, 22, 24, 26, 30, 32, 34, or 36 or a nucleic acid sequence with at least 85%, 90%, 95%, 99%, or 99.9% sequence homology to any one of SEQ ID NOs: 20, 22, 24, 26, 30, 32, 34, or 36.

[0051] The plasmids more preferably further comprise the nucleic acid sequence of any one of SEQ ID NOs: 20 or 36 or a nucleic acid sequence with at least 85%, 90%, 95%, 99%, or 99.9% sequence homology to any one of SEQ ID NOs: 20 or 36.

[0052] Yet another aspect of the disclosure is a plasmid comprising a nucleic acid sequence of any one or more of SEQ ID NO: 20, 22, 24, 26, 28, 30, 32, 34, 36, 38, 40, or 42 or a nucleotide sequence with at least 85%, 90%, 95%, 99%, or 99.9% sequence homology to any one or more of SEQ ID NO: 20, 22, 24, 26, 28, 30, 32, 34, 36, 38, 40, or 42.

[0053] A further aspect of the disclosure is a recombinantly produced tandem repeat amino acid sequence corresponding to an identified peptide.

[0054] The recombinantly produced tandem repeat amino acid sequence described herein can have the recombinantly produced tandem repeat amino acid sequence corresponding to an identified peptide comprises an amino acid sequence of any one of SEQ ID NOs: 1, 5, 9, 13, and 17 or an amino acid sequence with at least 85%, 90%, 95%, 99%, or 99.9% sequence homology to any one of SEQ ID NOs: 1, 5, 9, 13, and 17.

[0055] For the recombinantly produced tandem repeat amino acid sequences described herein, the recombinantly produced tandem repeat amino acid sequence can correspond to an identified peptide comprises an amino acid sequence of SEQ ID NO: 1 or an amino acid sequence with at least 85%, 90%, 95%, 99%, or 99.9% sequence homology to SEQ ID NO: 1.

[0056] Also, the recombinantly produced tandem repeat amino acid sequences described herein can comprise an amino acid sequence of SEQ ID NO: 5 or an amino acid sequence with at least 85%, 90%, 95%, 99%, or 99.9% sequence homology to SEQ ID NO: 5.

[0057] For the recombinantly produced tandem repeat amino acid sequences described, the recombinantly produced tandem repeat amino acid sequence corresponding to an identified peptide can comprise an amino acid sequence of SEQ ID NO: 9 or an amino acid sequence with at least 85%, 90%, 95%, 99%, or 99.9% sequence homology to SEQ ID NO: 9.

[0058] Further, the recombinantly produced tandem repeat amino acid sequences described herein can comprise an amino acid sequence of SEQ ID NO: 13 or an amino acid sequence with at least 85%, 90%, 95%, 99%, or 99.9% sequence homology to SEQ ID NO: 13.

[0059] For the recombinantly produced tandem repeat amino acid sequences described, the recombinantly produced tandem repeat amino acid sequence corresponding to an identified peptide can comprise an amino acid sequence of SEQ ID NO: 17 or an amino acid sequence with at least 85%, 90%, 95%, 99%, or 99.9% sequence homology to SEQ ID NO: 17.

[0060] Yet another aspect is an E. coli strain comprising one or more of the plasmids described herein, one or more of the recombinantly produced tandem repeat amino acid sequences described herein, or a combination thereof.

[0061] A further aspect is a Pichia pastoris strain comprising one or more of the plasmids described herein, one or more of the recombinantly produced tandem repeat amino acid sequences described herein, or a combination thereof.

[0062] Another aspect of the disclosure is a method of DNA assembly comprising identifying a peptide and forming a tandem repeat amino acid sequence corresponding to the identified peptide; optimizing the codon for expressing the tandem repeat amino acid sequence for expression in E. coli or Pichia pastoris to form the encoding fragment; forming a middle fragment by adding Bsal restriction sites to the 5' and 3' ends of the final sequence with four base overhangs corresponding to the last two bases and first two bases of the encoding fragment; generating a start fragment by amplifying the middle fragment with a 5' forward primer by adding a Bsal site for the last four bases of a mating factor alpha signal peptide; and generating an end fragment by using a reverse primer that added a stop codon and four bases after the stop codon, thereby forming a vector comprising a start fragment, an encoding fragment, a middle fragment, and an end fragment.

[0063] The method of DNA assembly described herein can have the tandem repeat amino acid sequences were selected to have a protease cut the peptide between each tandem repeat amino acid sequence but not within the tandem repeat amino acid sequence.

[0064] Further, the method of DNA assembly described herein can be used when no protease able to cut the peptide was available, a spacer amino acid was added to create a unique endopeptidase cut site and allow removal of the spacer amino acide from the C-terminus of the peptide by a carboxypeptidase.

[0065] The method of DNA assembly described herein can have the number of tandem repeat amino acid sequences in the DNA sequence were adjusted until the tandem repeat amino acid sequence could be synthesized by a commercial DNA vendor.

[0066] The method of DNA assembly described herein can have a pHK Pichia expression vector prepared by this method was modified to remove all Bsal sites in the backbone and a GFP flanked by Bsal sites between the mating factor alpha signal peptide and the Notl site.

[0067] The method of DNA assembly described herein can have a full-length mating factor alpha signal peptide was modified to replace the first 19 amino acids with the 22 amino acids signal peptide from S. cerevisiae Ostl.

[0068] The method of DNA assembly described herein can have the start fragment, the middle fragment, and the encoding fragment were ligated for up to 15 cycles.

[0069] The method of DNA assembly described herein can have the middle fragment, the end fragment, and the encoding fragment were ligated for up to 15 cycles.

[0070] The present disclosure provides a codon optimized nucleic acid sequence encoding a tandem repeat amino acid sequence corresponding to an identified peptide. Codon optimization refers to editing the codons of the nucleic acid sequence of a gene to encode for the same amino acids (i.e. synonymous codon changes) but use nucleic acid codons that are most efficient in the expression system of interest to improve protein expression. The sequence can be codon optimized for E. coli expression, Pichia pastoris expression, or any other bacterial or yeast expression system known in the art.

[0071] As used in this application, including the appended claims, the singular forms "a," "an," and "the" include plural references unless the content clearly dictates otherwise, and are used interchangeably with "at least one" and "one or more."

[0072] The invention will be further described in the following examples, which do not limit the scope of the invention described in the claims.EXAMPLES

[0073] It is to be understood that while the invention has been described in conjunction with the detailed description thereof, the preceding description is intended to illustrate and not limit the scope of the invention, which is defined by the scope of the appended claims. Other aspects, advantages, and modifications are within the scope of the following claims.EXAMPLE 1 : EXPRESSION OF TANDEM REPETITIVE RECOMBINANT PROTEINS IN PICHIA PASTORIS AND E. COLI

[0074] Tandem repeat amino acid sequences were designed for each peptide of interest, then codon optimized for expression in E. coli or Pichia pastoris. Amino acid sequences were checked on PeptideCutter (https: / / web.expasy.org / peptide_cutter / ) to ensure that a protease is predicted to cut between each tandem repeat but nowhere else in the mature peptide monomer. For peptides where no protease was available, a spacer amino acid was added that both created a unique endopeptidase cut site and can beremoved from the C-terminus of the mature peptide by a carboxypeptidase. The number of tandem repeats in the DNA sequences were adjusted until they had low enough complexity to be synthesized by a commercial DNA vendor. Bsal restriction sites were added to the 5’ and 3’ ends of the final sequence with 4 base overhangs that corresponded to the last two and first two bases of the DNA fragment encoding the tandem repeat polypeptide. This fragment was termed the “middle” fragment for each assembly. The DNA fragments were then amplified with a 5’ forward primer that added a Bsal site for the last four bases of the mating factor alpha signal peptide to generate a “start” fragment. The “end” fragment was generated by using a reverse primer that added a stop codon and the four bases after the stop codon in the destination vector. For this assembly, the pHK Pichia expression vector was modified to remove all Bsal sites in the backbone and a GFP flanked by Bsal sites between the mating factor alpha signal peptide and the Notl site. In some instances, the full-length mating factor alpha signal peptide was modified to replace the first 19 amino acids with the 22 amino acids signal peptide from S. cerevisiae Ostl. The “start,” “middle,” and vector backbone and in another tube the “middle” “end” and vector backbone were ligated for 15 cycles each. The 2 reactions were pooled and ligated for another 15 cycles before heat deactivation. The mixture was used to transform E. coli competent cells. Resulting colonies were screened by colony PCR using a forward primer binding to the start of the tandem repeat polypeptide DNA and a reverse primer binding to the A0X1 terminator. Colonies that produced multiple bands or a high molecular weight smear were chosen to be miniprepped and sequence confirmed by Sanger sequencing.

[0075] Identified plasmids were linearized at the HIS4 gene in the backbone with a BspEI digestion. The linearized expression plasmid was transformed into Pichia pastoris (GS 115) cells using known methods and the expression cassette was integrated into the His 4 locus of Pichia genome. After screening, the positive strains were identified, as summarized in Table 1.Table 1. List of tandem repeat peptide targets, property, plasmids, and protease used for peptide monomer releaseEXAMPLE 2: PRODUCTION OF PP9

[0076] The DNA assembly method described in Example 1 was used on PP9. This process yielded constructs with 13, 25 and 60 tandem repeats of PP9 secreted by the hybrid Ostl -mating factor alpha signal peptide. To demonstrate tandem repeat peptide production, the following experiment was conducted to confirm tandem repeat polypeptide expression and secretion by GS115 Pichia strains listed in Table 1. Single colonies of the Pichia pastoris strains were inoculated in BMGY medium in a 24 wells plate or baffled flask and grown at 28-30°C in a shaking incubator (250-300 rpm) until the culture reached an ODeoo of 2-6 (log-phase growth). The cells were harvested by centrifuging and resuspended to an ODeoo of 1.0 in BMM / BMMY medium to induce expression. 100% methanol was added to the BMMY medium to a final concentrationof 1% methanol every 24 hours to maintain induction of expression. Strains that secreted tandem repeat peptide were identified by subjecting media samples to electrophoresis on a 10-20% SDS-PAGE gel. As shown in Figure 1, there are expected molecular weight bands in expressing media samples, indicating secreted tandem repeat peptide production in the engineered Pichia strain. Furthermore, the correct peptide monomer can be produced by cleavage of suitable proteases (Example 3). Figure 1 shows SDS-PAGE from GS115 strains expressing PP9 expressed as 13 or 25 tandem repeats in the polypeptide chain.

[0077] The PP9TR25 plasmid was multimerized in vitro to generate constructs containing 2 and 4 copies of the gene. To generate multiple copies of the expression cassette in vitro, the above plasmid was digested with BspEI and Bglll or BspEI and BamHI. The fragments containing the peptide coding sequence were gel-purified then ligated together. Resulting E. coli colonies were screened by digestion with Bglll and BamHI to find colonies with an insert that is double the size of the signal expression cassette, which are plasmids containing 2 expression cassettes. This procedure was repeated on the 2 copies plasmid to generate pHKA Pichia expression plasmids harboring 4 copies of identical expression cassettes.

[0078] These plasmids were linearized and used to transform P. pastoris GS115. Strains were induced as above and compared to GS115 expressing one copy of the tandem repeat peptide gene. Figure 2 demonstrates that increasing the copy number leads to strains that secrete more tandem repeat protein into the culture media.

[0079] To prove PP9 monomer can be produced from the secreted tandem repeat polypeptide, culture supernatant was mixed with 500 mM ammonium bicarbonate. Trypsin was added at and digested at 37 °C for 2 hours. The digested media was run on LCMS alongside a chemically synthesized PP9 standard confirm the elution time and mass match (Figure 3A-B).EXAMPLE 3: PP22 PRODUCTION

[0080] For the PP22 target, a construct with 18 tandem repeats of the PP22 monomer (SEQ ID NO: 4) was obtained using the assembly method described in Example 1. For this target, the hybrid Ostl / mating factor alpha signal peptide was used. Once expression was proven for the one copy plasmid, it was subjected to in vitro multimerization to generate a two-copy plasmid. After GS115 transformation and induction, improved expression of PP22 with multiple integrated copies was observedon SDS-PAGE shown in Figure 4. In vitro multimerization of the PP22TR18x2 construct was performed once more to generate a plasmid with 4 copies of the expression cassette.

[0081] Pichia expressing four copies of the PP22 gene was used to generate samples to prove the release of PP22 monomer from the tandem repeat polypeptide. Ammonium sulfate was added to fermentation media and rotated at 4 °C for 1 hour. After centrifugation, the protein pellet was resuspended in water. Digestion was performed with chymotrypsin in 0.1M TRIS buffer pH 8 with 10 mM CaCh added. Digestion proceeded at room temperature for up to 4 hours and samples were run on HPLC. HPLC analysis was performed using a hypersil gold C18 HPLC column. A linear gradient increased from 10% - 35% 0.1% TFA in water: 0.1 % TFA in acetonitrile over 6 minutes then dropped down to 10% for an additional 4 minutes at a flow rate of 0.6 mL / min. An intact PP22TR18 tandem repeat polypeptide elutes as a single peak at 4.5 minutes when no chymotrypsin was added. In samples where the tandem repeat is partially digested, a series of peaks is observed, each corresponding to tandem repeats of differing length. In a fully digested sample, a single peak eluting at 2.2 minutes is observed and matches that of a chemically synthesized PP22 standard.

[0082] Samples were analyzed by LC-MS using a Waters Acquity UPLC BEH C18 column (1.7 pm, 2.1 x 100mm). Mobile phase A was 0.1% formic acid in water, and mobile phase B was 0.1% formic acid in acetonitrile. The flow rate was 0.2 ml / minute. Mass spectrometry analysis of the samples was done on the Q Exactive Hybrid Quadrupole-Orbitrap Mass Spectrometer (Thermo Fisher Scientific) with an optimized method in positive ion mode. The PP22 monomer released from the PP22TR18 polypeptide by chymotrypsin digestion has similar elution time to the PP22 standard (6.15-6.25 minutes). The produced PP22 peptide has same mass ([M+H]+: 1203 m / z ) as the chemically synthesized PP22 standard. These results provided the evidence supporting PP22 production in engineered Pichia strain.EXAMPLE 4: PRODUCTION OF PP35

[0083] The DNA assembly method described in Example 1 was used on PP35 (SEQ ID NO: 8). This process yielded constructs with 20 tandem repeats of PP35 secreted by the hybrid Ostl -mating factor alpha. In vitro multimerization was performed twice on this plasmid, yielding pHK-PP35TR20x4. To demonstrate tandem repeat peptide production, the following experiment was conducted to confirm tandemrepeat polypeptide expression and secretion by GS115 Pi chia strains listed in Table 1. Single colonies of the Pichia pastoris strains were inoculated in BMGY medium in a 24 wells plate or baffled flask and grown at 28-30°C in a shaking incubator (250-300 rpm) until the culture reached an ODeoo of 2-6 (log-phase growth). The cells were harvested by centrifuging and resuspended to an ODeoo of 1.0 in BMM / BMMY medium to induce expression. 100% methanol was added to the BMMY medium to a final concentration of 1% methanol every 24 hours to maintain induction of expression.

[0084] Strains that secreted tandem repeat peptide were identified by subjecting media samples to electrophoresis on a 10-20% SDS-PAGE gel. Figure 7 shows SDS- PAGE from GS115 strains expressing PP35 expressed as 20 tandem repeats in the polypeptide chain.

[0085] To prove monomer release, media containing PP35TR20 was precipitated by adding ammonium sulfate. The resuspended precipitate was subject to digestion with chymotrypsin in 0.1 M TRIS buffer pH 8 with 10 mM CaCh added. After 8 hours digestion, samples were run on LCMS to confirm that the released PP35 monomer ([M+H]+: 1147 m / z) matches a chemically synthesized standard (Figure 8A- B).EXAMPLE 5: PRODUCTION OF PP36

[0086] The DNA assembly method described in Example 1 was used on PP36 (SEQ ID NO: 12). This process yielded constructs with 8 and 12 tandem repeats of PP36 secreted by the hybrid Ostl -mating factor alpha or full-length mating factor alpha signal peptide. To demonstrate tandem repeat peptide production, the following experiment was conducted to confirm tandem repeat polypeptide expression and secretion by GS115 Pichia strains listed in Table 1. Single colonies of the Pichia pastoris strains were inoculated in BMGY medium in a 24 wells plate or baffled flask and grown at 28-30°C in a shaking incubator (250-300 rpm) until the culture reached an ODeoo of 2-6 (log-phase growth). The cells were harvested by centrifuging and resuspended to an ODeoo of 1.0 in BMM / BMMY medium to induce expression. 100% methanol was added to the BMMY medium to a final concentration of 1% methanol every 24 hours to maintain induction of expression. Strains that secreted tandem repeat peptide were identified by subjecting media samples to electrophoresis on a 10-20% SDS-PAGE gel. Figure 9 shows SDS-PAGE from GS115 strains expressing PP36 expressed as 12 tandem repeats in the polypeptide chain.

[0087] To prove monomer release, media containing PP36TR12 was precipitated by adding ammonium sulfate to 80% saturation. The resuspended precipitate was subject to digestion with 100 mg / mL trypsin in 0.1 M TRIS buffer pH 8. After 4 hours digestion at 37 °C, samples were run on LC MS to confirm that the released PP36 monomer ([M+H]+: 1576 m / z) matches a chemically synthesized standard (Figure 10A-B).EXAMPLE 6: IMPROVEMENT OF TANDEM REPEAT PEPTIDE SECRETION AND PRODUCTION IN PICHIA PASTORIS

[0088] To improve the amount of tandem repeat peptide secreted by above identified Pichia strains, a series of chaperone and proteins related to protein expression, secretion, and folding were co-expressed. These chaperones and proteins were selected from P. pastoris or plant to be heterologously expressed with a tandem repeat peptide in Pichia (Table 1).

[0089] Protein disulfide isomerase (PDI) is a chaperone localized primarily in the endoplasmic reticulum that aids in forming disulfide bonds between cysteine residues. Overexpressing PDI in P. pastoris has also been shown to improve the expression of certain non-disulfide bond containing proteins. ER oxidoreductin (ERO) proteins work in tandem with PDI by donating oxidating equivalents for disulfide bond formation. ERVs are a family of sulfhydryl oxidases that play a similar role as EROs but may also directly catalyze disulfide bond formation. HAC1 is a transcriptional regulator of the unfolded protein response in P. pastoris. GPX1 is a cytosolic peroxidase that is involved in cellular redox balancing. KAR2 codes for the ER chaperone BiP that aids in proper folding and directs misfolded proteins to be degraded. The genes SLY1 and SEC1 regulate vesicle traffic from the ER to the Golgi and from the Golgi to the extracellular membrane respectively. All selected transcription regulator, chaperones and disulfide bond formation related proteins are listed in Table 2.Table 2. Selection of candidates for co-expression in PP22 production strain.

[0090] Chaperone genes were cloned into a modified pPICZ vector that has an Ndel site after the AOX1 promoter in place of the EcoRI site. To construct vectors with multiple chaperones or multiple copies of the same chaperone, the vector containing chaperone to be added was digested with Bglll and BamHI to release the expression cassette. The expression cassette was then ligated into a second chaperone vector linearized with BamHI. A Bglll and BamHI digest was performed on the resulting plasmids to confirm the insertion of the chaperone in the proper orientation. Generated vectors were linearized at the A0X1 promoter with SacI restriction enzyme and used to transform GS115 P. pastoris with tandem repeat expression cassettes integrated at the HIS4 locus. Colonies with chaperone integration were selected on YPD-Zeocin and confirmed by colony PCR.

[0091] As the chaperones selected often work in tandem with other enzymes to improve secretion or disulfide bonding, expression vectors with multiple genes were generated. To ensure integration at the A0X1 locus, the promoter from the added genes needed to be replaced. The GAP1 and CAT1 promoters were amplified from GS115 genomic DNA with primers that added Bglll and Ndel sites to the 5’ and 3’ ends respectively. The AOX1 promoter was then excised from the pPICZ vectors with a Bglll / Ndel digest and replaced with the GAP1 and CAT1 promoters. Variouscombinations of different expression cassettes were generated by ligating a chaperone with GAP1 or CAT1 promoter excised from its expression vector with a BamHI / Bglll digest into a pPICZ A0X1 chaperone vector linearized with BamHI. All plasmids are listed in Table 3.Table 3: Summary of selected plasmids of single and combination of different chaperones for co-expression

[0092] To demonstrate improvement of tandem repeat peptide production with chaperone co-expression, colonies of Pichia expressing PP22TR18x4 from integration at the HIS4 locus and PpPDIl and PpKAR2 from the AOX1 locus were grown overnight in 0.5 mL BMGY media in 96 well plates. The next day the cells were resuspended in 0.5 mL BMMY media to an OD600 of 1.0. The cultures were induced at 30 °C for 48 hours with the additional feeding of 1% methanol to each well twice daily. The cells were harvested, spun down, and the supernatant was analyzed by SDS- PAGE and HPLC. As shown in Figure 11, co-expression of PpPDIl and PpKAR2 can increase the amount of PP22TR18 tandem repeat polypeptide when compared to the parent strain.EXAMPLE 7: FERMENTATION, PURIFICATION AND IDENTIFICATION OF PEPTIDES

[0093] PP22 production strain with PP22TR18x4 construct integrated was cultured in 3 L fermenter for tandem repeat peptide production. Seed cultures were inoculated into defined media with glycerol and methanol continually fed into the medium for PP22TR18 induction after glycerol was fully consumed. Medium samples were collected at different time points and analyzed by SDS-PAGE and HPLC. As shown in Figure 12, no PP22 recombinant protein was detected before methanol feeding (Figure 12. 1); After methanol feeding, PP22 recombinant protein production is increased along with the induction time through the fermentation time (Figure 12. 2-4).

[0094] PP22TR18 protein was precipitated from fermentation media by adding ammonium sulfate to 50% saturation and stirring at 4 °C for 2 hours. Precipitated protein was collected by centrifugation at 15,000xg at 4 C for 20 minutes. Protein pellet was resuspended in 0.1 M Tris pH 8.0 with 10 mM calcium chloride for protease cleavage. PP22 monomer production was monitored by HPLC analysis.

[0095] In order to purify PP22 monomer, ethanol was added into PP22 cleavage mixture. After centrifugation, the supernatant was passed through ultrafiltration system to remove large and small molecular weight impurity. The retentate was adjusted to pH 3.5 and subjected to pretreatment activated charcoal column and resin column. The column was gradient eluted with 0.1% formic acid of water and 0.1% formic acid of methanol. The fractions were collected and detected by HPLC. The fractions with more than 95% HPLC purity were pooled and desalted by membrane dialysis. The desaltedPP22 solution was lyophilized to dry. The final PP22 dry product is more than 95% purity.EXAMPLE 8: TASTING EVALUATION OF FLAVOR PEPTIDES

[0096] A 2AFC test (Hartman, Hallagan and FEMA Science Committee Sensory Data Task Force I November 2013, Volume 67, No. 11, updated January 2022).) was conducted to verify that PP22 (Salatide) does not produce an inherent salty taste at 500ppm. 0.25 NaCl in distilled water (Control) and 500ppm PP22 in distilled water (sample) were tested by 30 participants with 2 repetitions. Of 60 panelist evaluations, 60 selected controls show saltier than pp22 samples, indicating a statistically significant different from PP22 at 500ppm at 95% confidence.

[0097] A descriptive analysis profile was captured to understand the potential flavor modulation enhancement properties of PP22. The level chosen were from the range shown to enhance overall impact in a commercial no-sodium added chicken broth (plus one additional lower point of PP22 (350ppm and 500ppm). Swanson’s tetrapack commercial no sodium added chicken broth with 0.25% NaCl as control. 350ppm and 500ppm PP22 was added in the control broth separately for flavor enhancement test. Evaluations were conducted using a modified Spectrum-style defined interval scaling approach, by an expert panel of 6 panelists, in 3 repetitions / sessions for a total of 18 responses. Approximately 3-ounces of each sample were served each time at room temperature in plastic souffle cups labeled with 3 -digit codes with a sample of Control used as the reference. A 1-hour orientation was conducted at the start to allow the panel to align on vocabulary and ratings for the Control, and identify references. Consensus scores for the Control were discussed and agreed upon during training, and provided during testing to help calibrate the participants. A 15 pt. scale was used for ratings, with 1 -decimal allowance. Data was collected on CompusenseCloud and analyzed by XLSTAT using ANOVA and Fisher’s LSD at 95% CI. As shown in Figure 13, total impact is significantly higher for the samples containing PP22 (Salatide) at both 350ppm and 500ppm as compared to control (no PP22). No significant differences were found for PP22 and umami. Based on these results, PP22 (Salatide) can be used as an overall flavor enhancer, without inherent salty basic taste, at a maximum use rate of 500ppm.EXAMPLE 9: SEQUENCES

[0098] PP9 Amino Acid Sequence

[0099] GHK

[0100] PP9 DNA Sequence

[0101] GGCCACAAG

[0102] SEQ ID NO: 1 PP9TR25 Amino Acid Sequence

[0103] GHI<GHI<GHI<GHI<GHI<GHI<GHI<GHI<GHI<GHI<GHI<GHKGHKGHKGHKGHKGHKGHKGHKGHKGHKGHKGHKGHKGHK

[0104] SEQ ID NO: 2 PP9TR25 DNA Sequence

[0105] GGCCACAAGGGCCACAAGGGCCACAAGGGCCACAAGGGCCACAAGGGCCACAAGGGCCACAAGGGCCACAAGGGCCACAAGGGCCACAAGGGCCACAAAGGCCACAAAGGGCATAAAGGGCACAAAGGTCACAAGGGCCACAAGGGTCACAAAGGCCACAAGGGCCACAAAGGCCACAAAGGGCATAAAGGGCACAAAGGTCACAAGGGCCACAAGGGTCACAAA

[0106] SEQ ID NO: 3 PP22 Amino Acid Sequence

[0107] EDEGEQPRPF

[0108] SEQ ID NO: 4 PP22 DNA Sequence

[0109] GAAGATGAAGGTGAACAGCCAAGACCCTTT

[0110] SEQ ID NO: 5 PP22TR18 Amino Acid Sequence

[0111] EDEGEQPRPFEDEGEQPRPFEDEGEQPRPFEDEGEQPRPFEDEGEQPRPFEDEGEQPRPFEDEGEQPRPFEDEGEQPRPFEDEGEQPRPFEDEGEQPRPFEDEGEQPRPFEDEGEQPRPFEDEGEQPRPFEDEGEQPRPFEDEGEQPR PFEDEGEQPRPFEDEGEQPRPFEDEGEQPRPF

[0112] SEQ ID NO: 6 PP22TR18 DNA Sequence

[0113] GAAGATGAAGGTGAAC AGCC AAGACCCTTTGAAGATGAGGGCGAGCAGCCTAGACCATTTGAGGACGAAGGCGAACAACCTCGCCCATTCGAAGATGAAGGTGAGCAACCCAGACCATTTGAAGATGAAGGTGAACAGCCAAGACCCTTTGAAGATGAGGGCGAGCAGCCTAGACCATTTGAGGACGAAGGCGAACAACCTCGCCCATTCGAAGATGAAGGTGAGCAACCCAGACCATTTGAAGATGAAGGTGAACAGCCAAGACCCTTTGAAGATGAGGGCGAGCAGCCTAGACCATTTGAGGACGAAGGCGAACAACCTCGCCCATTCGAAGATGAAGGTGAGCAACCCAGACCATTTGAAGATGAAGGTGAACAGCCAAGACCCTTTGAAGATGAGGGCGAGCAGCCTAGACCATTTGAGGACGAAGGCGAACAACCTCGCCCATTCGAAGATGAAGGTGAGCAACCCAGACCATTTGAAGATGAAGGTGAACAGCCAAGACCCTTTGAAGATGAGGGCGAGCAGCCT AGACCATTTTAA

[0114] SEQ ID NO: 7 PP35 Amino Acid Sequence

[0115] VGPDDDEKSW

[0116] SEQ ID NO: 8 PP35 DNA Sequence

[0117] GTGGGACC AGATGACGATGAGAAGTCCTGG

[0118] SEQ ID NO: 9 PP3TR20 Amino Acid Sequence

[0119] VGPDDDEKSWVGPDDDEKSWVGPDDDEKSWVGPDDDEKSWVGPDDDEKSWVGPDDDEKSWVGPDDDEKSWVGPDDDEKSWVGPDDDEKSWVGPDDDEKSWVGPDDDEKSWVGPDDDEKSWVGPDDDEKSWVGPDDDEKSWVGPDDDEKSWVGPDDDEKSWVGPDDDEKSWVGPDDDEKSWVGPDDDEKSWVGPDDDEKSWVGPDDDEKSWVGPDDDEKSW

[0120] SEQ ID NO: 10 PP35TR20 DNA Sequence

[0121] GTGGGACCAGATGACGATGAGAAGTCCTGGGTAGGTCCCGACGATGATGAGAAATCTTGGGTTGGTCCTGACGATGATGAGAAATCATGGGTGGGTCCAGACGATGATGAAAAATCCTGGGTGGGACCAGATGACGATGAGAAGTCCTGGGTAGGTCCCGACGATGATGAGAAATCTTGGGTTGGTCCTGACGATGATGAGAAATCATGGGTGGGTCCAGACGATGATGAAAAATCCTGGGTGGGACCAGATGACGATGAGAAGTCCTGGGTAGGTCCCGACGATGATGAGAAATCTTGGGTTGGTCCTGACGATGATGAGAAATCATGGGTGGGTCCAGACGATGATGAAAAATCCTGGGTGGGACCAGATGACGATGAGAAGTCCTGGGTAGGTCCCGACGATGATGAGAAATCTTGGGTTGGTCCTGACGATGATGAGAAATCATGGGTGGGTCCAGACGATGATGAAAAATCCTGGGTGGGACCAGATGACGATGAGAAGTCCTGGGTAGGTCCCGACGATGATGAAAAATCCTGGGTGGGACCAGATGACGATGAGAAGTCCTGGGTAGGTCCCGACGATGATGAGAAATCTTGGGTTGGTCCTGACGATGATGAGAAATCATGGGTGGGTCCAGACGATGATGAAAAATCCTGG

[0122] SEQ ID NO: 11 PP36 Amino Acid Sequence

[0123] GENEEEDSGAIVTVK

[0124] SEQ ID NO: 12 PP36 DNA Sequence

[0125] GGAGAAAACGAGGAGGAAGATTCAGGTGCCATCGTAACTGTGAAA

[0126] SEQ ID NO: 13 PP36TR12 Amino Acid Sequence

[0127] GENEEEDSGAIVTVKGENEEEDSGAIVTVKGENEEEDSGAIVTVKGENEEEDSGAIVTVKGENEEEDSGAIVTVKGENEEEDSGAIVTVKGENEEEDSGAIVTVKGENEEEDSGAIVTVKGENEEEDSGAIVTVKGENEEEDSGAIVTVKGENEEEDSGAIVTVKGENEEEDSGAIVTVK

[0128] SEQ ID NO: 14 PP36TR12 DNA Sequence

[0129] GGAGAAAACGAGGAGGAAGATTCAGGTGCCATCGTAACTGTGAAAGGTGAGAACGAAGAAGAAGATTCAGGCGCAATAGTTACTGTAAAGGGCGAGAATGAAGAGGAAGATAGTGGTGCCATAGTAACAGTTAAAGGTGAAAATGAGGAGGAGGATTCAGGCGCCATTGTGACTGTTAAAGGAGAAAACGAGGAGGAAGATTCAGGTGCCATCGTAACTGTGAAAGGTGAGAACGAAGAAGAAGATTCAGGCGCAATAGTTACTGTAAAGGGCGAGAATGAAGAGGAAGATAGTGGTGCCATAGTAACAGTTAAAGGTGAAAATGAGGAGGAGGATTCAGGCGCCATTGTGACTGTTAAAGGAGAAAACGAGGAGGAAGATTCAGGTGCCATCGTAACTGTGAAAGGTGAGAACGAAGAAGAAGATTCAGGCGCAATAGTTACTGTAAAGGGCGAGAATGAAGAGGAAGATAGTGGTGCCATAGTAACAGTTAAAGGTGAAAATGAGGAGGAGGATTCAGGCGCCATTGTGACTGTTAAA

[0130] SEQ ID NO: 15 PP29 Amino Acid Sequence

[0131] PLWR

[0132] SEQ ID NO: 16 PP29 DNA Sequence

[0133] CCTCTGTGGCGG

[0134] SEQ ID NO: 17 PP29TR18 Amino Acid Sequence

[0135] PLWRPLWRPLWRPLWRPLWRPLWRPLWRPLWRPLWRPLWRPLWRPLWRPLWRPLWRPLWRPLWRPLWRPLWR

[0136] SEQ ID NO: 18 PP29TR18 DNA Sequence

[0137] CCTCTGTGGCGGCCCTTGTGGCGTCCCCTTTGGAGGCCCTTATGGAGGCCATTATGGAGGCCTCTATGGCGTCCTCTGTGGCGGCCCTTGTGGCGTCCCCTTTGGAGGCCCTTATGGAGGCCATTATGGAGGCCTCTATGGCGTCCTCTGTGGCGGCCCTTGTGGCGTCCCCTTTGGAGGCCCTTATGGAGGCCATTATGGAGGCCTCTATGGCGT

[0138] SEQ ID NO: 19 AA (PpPDIl)

[0139] MQFNWDIKTVASILSALTLAQASDQEAIAPEDSHVVKLTEATFESFITSNPHVLAEFFAPWCGHCKKLGPELVSAAEILKDNEQVKIAQIDCTEEKELCQGYEIKGYPTLKVFHGEVEVPSDYQGQRQSQSIVSYMLKQSLPPVSEINATKDLDDTIAEAKEPVIVQVLPEDASNLESNTTFYGVAGTLREKFTFVSTKSTDYAKKYTSDSTPAYLLVRPGEEPSVYSGEELDETHLVHWIDIESKPLFGDIDGSTFKSYAEANIPLAYYFYENEEQRAAAADIIKPFAKEQRGKINFVGLDAVKFGKHAKNLNMDEEKLPLFVIHDLVSNKKFGVPQDQELTNKDVTELIEKFIAGEAEPIVKSEPIPEIQEEKVFKLVGKAHDEVVFDESKDVLVKYYAPWCGHCKRMAPAYEELATLYANDEDASSKVVIAKLDHTLNDVDNVDIQGYPTLILYPAGDKSNPQLYDGSRDLESLAEFVI<ERGTHI<VDALALRPVEEEI<EAEEEAESEADAHDEL

[0140] SEQ ID NO: 20 DNA (PpPDIl)

[0141] ATGCAATTCAACTGGGATATTAAAACTGTGGCAAGTATTTTGTCCGCTCTCACACTAGCACAAGCAAGTGATCAGGAGGCTATTGCTCCAGAGGACTCTCATGTCGTCAAATTGACTGAAGCCACTTTTGAGTCTTTCATCACCAGTAATCCTCACGTTTTGGCAGAGTTTTTTGCCCCTTGGTGTGGTCACTGTAAGAAGTTGGGCCCTGAACTTGTTTCTGCTGCCGAGATTTTAAAGGACAATGAGCAGGTTAAGATTGCTCAAATTGATTGTACGGAGGAGAAGGAATTATGTCAAGGCTACGAAATTAAAGGGTATCCTACTTTGAAGGTGTTCCATGGTGAGGTTGAGGTCCCAAGTGACTATCAAGGTCAAAGACAGAGCCAAAGCATTGTCAGCTATATGCTAAAGCAGAGTTTACCCCCTGTCAGTGAAATCAATGCAACCAAAGATTTAGACGACACAATCGCCGAGGCAAAAGAGCCCGTGATTGTGCAAGTACTACCGGAAGATGCATCCAACTTGGAATCTAACACCACATTTTACGGAGTTGCCGGTACTCTCAGAGAGAAATTCACTTTTGTCTCCACTAAGTCTACTGATTATGCCAAAAAATACACTAGCGACTCGACTCCTGCCTATTTGCTTGTCAGACCTGGCGAGGAACCTAGTGTTTACTCTGGTGAGGAGTTAGATGAGACTCATTTGGTGCACTGGATTGATATTGAGTCCAAACCTCTATTTGGAGACATTGACGGATCTACCTTCAAATCATACGCTGAAGCTAACATCCCTTTAGCCTACTATTTCTATGAGAACGAAGAACAACGTGCTGCTGCTGCCGATATTATTAAACCTTTTGCTAAAGAGCAACGTGGCAAAATTAACTTTGTTGGCTTAGATGCCGTTAAATTCGGTAAGCATGCCAAGAACTTAAACATGGATGAAGAGAAACTCCCTCTATTTGTCATTCATGATTTGGTGAGCAACAAGAAGTTTGGAGTTCCTCAAGACCAAGAATTGACGAACAAAGATGTGACCGAGCTGATTGAGAAATTCATCGCAGGAGAGGCAGAACCAATTGTGAAATCAGAGCCAATTCCAGAAATTCAAGAAGAGAAAGTCTTCAAGCTAGTCGGAAAGGCCCACGATGAAGTTGTCTTCGATGAATCTAAAGATGTTCTAGTCAAGTACTACGCCCCTTGGTGTGGTCACTGTAAGAGAATGGCTCCTGCTTATGAGGAATTGGCTACTCTTTACGCCAATGATGAGGATGCCTCTTCAAAGGTTGTGATTGCAAAACTTGATCACACTTTGAACGATGTTGACAACGTTGATATTCAAGGTTATCCTACTTTGATCCTTTATCCAGCTGGTGATAAATCCAATCCTCAACTGTATGATGGATCTCGTGACCTAGAATCATTGGCTGAGTTTGTAAAGGAGAGAGGAACCCACAAAGTGGATGCCCTAGCACTCAGACCAGTCGAGGAAGAAAAGGAAGCTGAAGAAGAAGCTGAAAGTGAGGCAGACGCTCACGACGAG CTTTAA

[0142] SEQ ID NO: 21 AA (AtPDIl)

[0143] MASSSTSISLLLFVSFILLLVNSRAENASSGSDLDEELAFLAAEESKEQSHGGGSYHEEEHDHQHRDFENYDDLEQGGGEFHHGDHGYEEEPLPPVDEKDVAVLTKDNFTEFVGNNSFAMVEFYAPWCGACQALTPEYAAAATELKGLAALAKIDATEEGDLAQKYEIQGFPTVFLFVDGEMRKTYEGERTKDGIVTWLKKKASPSIHNITTKEEAERVLSAEPKLVFGFLNSLVGSESEELAAASRLEDDLSFYQTASPDIAKLFEIETQVKRPALVLLKKEEEKLARFDGNFTKTAIAEFVSANKVPLVINFTREGASLIFESSVKNQLILFAKANESEKHLPTLREVAKSFKGKFVFVYVQMDNEDYGEAVSGFFGVTGAAPKVLVYTGNEDMRKFILDGELTVNNIKTLAEDFLADKLKPFYKSDPLPENNDGDVKVIVGNNFDEIVLDESKDVLLEIYAPWCGHCQSFEPIYNKLGKYLKGIDSLVVAKMDGTSNEHPRAKADGFPTILFFPGGNKSFDPIAVDVDRTVVELYKFLKKHASIPFKLEKPATPEPVISTMKSDE KIEGDSSKDEL

[0144] SEQ ID NO: 22 DNA (AtPDIl)

[0145] ATGGCTTCTTCTTCTACTTCTATTTCTTTGTTGTTGTTCGTTTCTTTCATCTTGTTGTTGGTTAATTCTAGAGCTGAAAACGCTTCTTCTGGTTCTGATTTGGATGAAGAATTGGCTTTTCTTGCTGCTGAAGAATCTAAAGAACAATCTCATGGTGGTGGTTCTTATCATGAAGAAGAACATGATCATCAACATAGAGATTTTGAAAACTACGATGATTTGGAACAAGGTGGTGGTGAATTTCATCATGGTGACCATGGTTACGAAGAAGAACCATTGCCACCAGTTGATGAAAAAGATGTTGCTGTTTTGACTAAAGATAACTTCACTGAATTTGTCGGTAATAACTCTTTCGCTATGGTTGAATTTTACGCTCCATGGTGTGGTGCTTGTCAAGCTTTGACTCCTGAATATGCTGCTGCTGCTACTGAATTGAAAGGTTTGGCTGCTTTGGCTAAGATTGATGCTACTGAAGAAGGTGACTTGGCTCAAAAGTATGAAATTCAAGGTTTTCCTACTGTTTTCTTGTTTGTTGATGGTGAAATGAGAAAGACTTATGAAGGTGAAAGAACTAAGGATGGTATTGTTACTTGGTTGAAAAAGAAAGCTTCTCCTTCTATTCATAACATTACTACTAAGGAAGAGGCTGAAAGAGTTTTGTCTGCTGAACCAAAGTTGGTTTTTGGTTTTCTTAACTCTTTGGTTGGTTCTGAATCTGAAGAATTGGCCGCTGCTTCTAGATTGGAAGATGATTTGTCTTTTTACCAAACTGCTTCTCCTGATATTGCTAAATTGTTCGAAATTGAAACCCAAGTTAAGCGTCCTGCTTTGGTTTTGTTGAAAAAGGAAGAAGAAAAGTTGGCTAGATTTGATGGTAATTTTACTAAGACTGCTATCGCTGAATTTGTTTCTGCTAATAAGGTTCCATTGGTTATTAATTTCACCAGAGAAGGTGCTTCTTTGATTTTCGAATCTTCTGTTAAGAACCAATTGATTTTGTTCGCTAAAGCTAATGAATCTGAAAAGCATTTGCCTACTTTGAGAGAAGTTGCTAAGTCTTTCAAAGGTAAATTCGTTTTCGTTTACGTTCAAATGGATAATGAAGATTACGGTGAAGCTGTTTCTGGTTTCTTTGGTGTTACTGGTGCTGCTCCAAAGGTTTTGGTTTATACTGGTAACGAAGATATGAGAAAGTTCATTTTGGATGGTGAATTGACTGTTAACAATATTAAGACTCTGGCTGAAGATTTTCTTGCTGATAAGTTGAAACCATTCTACAAGTCTGATCCATTGCCTGAAAACAACGATGGTGACGTTAAGGTTATTGTTGGTAACAACTTCGATGAAATTGTTTTGGATGAATCTAAGGATGTTTTGTTGGAAATCTATGCTCCATGGTGCGGTCATTGTCAATCTTTTGAACCAATCTATAACAAGTTGGGTAAATACTTGAAGGGTATTGATTCTTTGGTTGTTGCTAAAATGGATGGTACTTCTAACGAACATCCAAGAGCTAAAGCTGATGGTTTTCCTACCATTTTGTTTTTCCCTGGTGGTAATAAGTCTTTCGATCCTATTGCTGTTGATGTTGATAGAACTGTTGTTGAATTGTATAAGTTCTTGAAGAAGCATGCTTCTATTCCTTTCAAGTTGGAAAAGCCAGCTACTCCAGAACCTGTTATTTCTACTATGAAGTCTGATGAAAAGATCGAAGGTGACTCTTCTAAGGATGAATTGTAA

[0146] SEQ ID NO: 23 AA (HAC1)

[0147] MPVDSSHKTASPLPPRKRAKTEEEKEQRRVERILRNRRAAHASREKKRRHVEFLENHVVDLESALQESAKATNKLKEIQDIIVSRLEALGGTVSDLDLTVPEVDFPKSSDLEPMSDLSTSSKSEKASTSTRRSLTEDLDEDDVAEYDDEEEDEELPRKMKVLNDKNKSTSIKQEKLNELPSPLSSDFSDVDEEKSTLTHLKLQQQQQQPVDNYVSTPLSLPEDSVDFINPGNLKIESDENFLLSSNTLQIKHENDTDYITTAPSGSINDFFNSYDISESNRLHHPAVMTDSSLHITAGSIGFFSLIGGGESSVAGRRSSVGTYQLTCIAIR

[0148] SEQ ID NO: 24 DNA (HAC1)

[0149] ATGCCCGTAGATTCTTCTCATAAGACAGCTAGCCCACTTCCACCTCGTAAAAGAGCAAAGACGGAAGAAGAAAAGGAGCAGCGTCGAGTGGAACGTATCCTACGTAATAGGAGAGCGGCCCATGCTTCCAGAGAGAAGAAACGAAGACACGTTGAATTTCTGGAAAACCACGTCGTCGACCTGGAATCTGCACTTCAAGAATCAGCCAAAGCCACTAACAAGTTGAAAGAAATACAAGATATCATTGTTTCAAGGTTGGAAGCCTTAGGTGGTACCGTCTCAGATTTGGATTTAACAGTTCCGGAAGTCGATTTTCCCAAATCTTCTGATTTGGAACCCATGTCTGATCTCTCAACTTCTTCGAAATCGGAGAAAGCATCTACATCCACTCGCAGATCTTTGACTGAGGATCTGGACGAAGATGACGTCGCTGAATATGACGACGAAGAAGAGGACGAAGAGTTACCCAGGAAAATGAAAGTCTTAAACGACAAAAACAAGAGCACATCTATCAAGCAGGAGAAGTTGAATGAACTTCCATCTCCTTTGTCATCCGATTTTTCAGACGTAGATGAAGAAAAGTCAACTCTCACACATTTAAAGTTGCAACAGCAACAACAACAACCAGTAGACAATTATGTTTCTACTCCTTTGAGTCTTCCGGAGGATTCAGTTGATTTTATTAACCCAGGTAACTTAAAAATAGAGTCCGATGAGAACTTCTTGTTGAGTTCAAATACTTTACAAATAAAACACGAAAATGACACCGACTACATTACTACAGCTCCATCAGGTTCCATCAATGATTTTTTTAATTCTTATGACATTAGCGAGTCGAATCGGTTGCATCATCCAGCAGTGATGACGGATTCATCTTTACACATTACAGCAGGCTCCATCGGCTTTTTCTCTTTGATTGGGGGGGGGGAAAGTTCTGTAGCAGGGAGGCGCAGTTCAGTTGGCACATATCAGTTGACATGCATAGCGATCAGG

[0150] SEQ ID NO: 25 AA (AtEROl)

[0151] MGKGAIKEEESEKKRKTWRWPLATLVVVFLAVAVSSRTNSNVGFFFSDRNSCSCSLQKTGKYKGMIEDCCCDYETVDNLNTEVLNPLLQDLVTTPFFRYYKVKLWCDCPFWPDDGMCRLRDCSVCECPENEFPEPFKKPFVPGLPSDDLKCQEGKPQGAVDRTIDNRAFRGWVETKNPWTHDDDTDSGEMSYVNLQLNPERYTGYTGPSARRIWDSIYSENCPKYSSGETCPEKKVLYKLISGLHSSISMHIAADYLLDESRNQWGQNIELMYDRILRHPDRVRNMYFTYLFVLRAVTKATAYLEQAEYDTGNHAEDLKTQSLIKQLLYSPKLQTACPVPFDEAKLWQGQSGPELKQQIQKQFRNISALMDCVGCEKCRLWGKLQVQGLGTALKILFSVGN QDIGDQTLQLQRNEVIALVNLLNRLSESVKMVHDMSPDVERLMEDQIAKVSA KPARLRRIWDLAVSFW

[0152] SEQ ID NO: 26 DNA (AtEROl)

[0153] ATGGGTAAAGGTGCTATTAAGGAAGAAGAATCTGAAAAGAAGAGAAAAACTTGGAGATGGCCTTTGGCTACTTTGGTTGTTGTTTTCTTGGCTGTTGCTGTTTCTTCTAGAACTAACTCTAACGTTGGTTTCTTTTTCTCTGATAGAAATTCTTGTTCCTGTTCTTTGCAAAAAACTGGTAAATACAAGGGTATGATTGAAGATTGTTGTTGTGATTATGAGACTGTTGATAACTTGAATACTGAAGTTTTGAACCCTTTGTTGCAAGATTTGGTTACTACTCCATTTTTCAGATACTACAAAGTTAAGTTGTGGTGTGATTGTCCATTCTGGCCAGATGATGGTATGTGTAGATTGAGAGATTGTTCTGTTTGTGAATGTCCAGAAAACGAATTTCCTGAACCATTCAAAAAGCCTTTCGTTCCTGGTTTGCCATCTGATGATTTGAAATGTCAAGAAGGTAAACCACAAGGTGCTGTTGATAGAACTATTGATAACAGAGCTTTTAGAGGTTGGGTTGAAACTAAAAACCCTTGGACTCATGATGATGATACTGATTCTGGTGAAATGTCTTATGTTAATTTGCAATTGAACCCAGAAAGATACACTGGTTACACTGGTCCTTCTGCTAGAAGAATTTGGGATTCTATCTATTCTGAAAACTGTCCAAAGTACTCTTCTGGTGAAACTTGTCCAGAAAAGAAAGTTTTGTATAAGTTGATCTCCGGTTTGCATTCTTCTATTTCTATGCATATTGCTGCTGATTATTTGTTGGATGAATCTAGAAATCAGTGGGGTCAAAACATTGAATTGATGTATGATAGAATCCTGAGACATCCAGATAGAGTTAGAAATATGTATTTCACTTACCTGTTCGTTTTGAGAGCTGTTACTAAAGCTACTGCTTATTTGGAACAAGCTGAATACGATACTGGTAACCATGCTGAAGATTTGAAAACTCAATCTTTGATTAAGCAGTTGTTGTATTCTCCTAAATTGCAAACTGCTTGTCCAGTTCCTTTTGATGAAGCTAAGTTGTGGCAAGGTCAATCTGGTCCAGAATTGAAACAACAAATTCAAAAACAGTTCAGAAACATCTCTGCTTTGATGGATTGTGTTGGTTGTGAAAAGTGTAGATTGTGGGGTAAATTGCAAGTTCAAGGTTTGGGTACTGCTTTGAAAATTTTGTTTTCTGTTGGTAACCAGGATATCGGTGACCAAACTTTGCAATTGCAAAGAAACGAAGTTATTGCTTTGGTTAATTTGTTGAACAGATTGTCTGAATCTGTTAAGATGGTTCATGATATGTCTCCAGATGTTGAAAGATTGATGGAAGATCAAATTGCTAAAGTTTCTGCTAAACCTGCTAGATTGAGAAGAATTTGGGACTTGGCTGTTTCTTTCTGGT AA

[0154] SEQ ID NO: 27 AA (AtERO2)

[0155] MAETDVGSVKGKEKGSGKRWILLIGAIAAVLLAVVVAVFLNTQNSSISEFTGKICNCRQAEQQKYIGIVEDCCCDYETVNRLNTEVLNPLLQDLVKTPFYRYFKVKLWCDCPFWPDDGMCRLRDCSVCECPESEFPEVFKKPLSQYNPVCQEGKPQATVDRTLDTRAFRGWTVTDNPWTSDDETDNDEMTYVNLRLNPERYTGYIGPSARRIWEAIYSENCPKHTSEGSCQEEKILYKLVSGLHSSISVHIASDYLLDEATNLWGQNLTLLYDRVLRYPDRVQNLYFTFLFVLRAVTKAEDYLGEAEYETGNVIEDLKTKSLVKQVVSDPKTKAACPVPFDEAKLWKGQRGPELKQQLEKQFRNISAIMDCVGCEKCRLWGKLQILGLGTALKILFTVNGEDNLR HNLELQRNEVIALMNLLHRLSESVKYVHDMSPAAERIAGGHASSGNSFWQRI VTSIAQSKAVSGKRS

[0156] SEQ ID NO: 28 DNA (AtERO2)

[0157] ATGGCTGAAACTGATGTTGGTTCTGTTAAGGGTAAAGAAAAGGGTTCTGGTAAAAGATGGATTTTGTTGATTGGTGCTATTGCTGCTGTTTTGTTGGCTGTTGTTGTTGCTGTTTTCTTGAACACTCAAAACTCTTCTATTTCTGAGTTTACTGGTAAAATCTGTAACTGTAGACAAGCTGAACAACAAAAGTACATTGGTATTGTTGAAGATTGTTGTTGTGATTATGAGACTGTTAACAGATTGAACACTGAAGTTTTGAACCCATTGTTGCAAGATTTGGTTAAGACTCCATTCTACAGATACTTTAAGGTTAAGTTGTGGTGTGATTGTCCTTTCTGGCCAGATGATGGTATGTGTAGATTGAGAGATTGTTCTGTTTGTGAATGTCCAGAATCTGAATTTCCTGAAGTTTTCAAGAAACCTTTGTCTCAATATAACCCAGTTTGTCAAGAAGGTAAACCACAAGCTACTGTTGATAGAACTTTGGATACTAGAGCTTTCAGAGGTTGGACTGTTACTGATAATCCTTGGACTTCTGATGATGAAACTGATAACGATGAAATGACTTATGTTAACTTGAGATTGAACCCAGAAAGATACACTGGTTATATTGGTCCATCTGCTAGAAGAATTTGGGAAGCTATCTATTCTGAAAATTGTCCAAAACATACCTCTGAAGGTTCTTGTCAAGAAGAAAAGATTTTGTATAAGCTGGTTTCTGGTTTGCATTCTTCTATTTCCGTTCATATTGCTTCTGATTACTTGTTGGATGAAGCTACTAACTTGTGGGGTCAAAACTTGACTTTGTTGTATGATAGAGTTTTGAGATACCCAGATAGAGTTCAAAACTTGTACTTTACTTTCTTGTTCGTTTTGAGAGCTGTTACTAAAGCTGAAGATTACTTGGGTGAAGCTGAATACGAAACTGGTAACGTTATTGAAGATTTGAAAACTAAATCTCTGGTCAAGCAAGTTGTTTCTGATCCAAAAACTAAGGCTGCTTGTCCAGTTCCATTTGATGAAGCTAAGTTGTGGAAGGGTCAAAGAGGTCCAGAATTGAAGCAACAATTGGAAAAGCAATTTCGTAACATTTCTGCTATTATGGATTGTGTTGGTTGTGAAAAATGTAGATTGTGGGGTAAATTGCAAATTTTGGGTTTGGGTACTGCTTTGAAAATTTTGTTTACTGTTAACGGTGAGGATAATTTGAGACATAACTTGGAATTGCAAAGAAACGAAGTTATTGCTTTGATGAATTTGTTGCATAGATTGTCTGAATCTGTTAAATACGTTCATGATATGTCTCCTGCTGCTGAAAGAATTGCTGGTGGTCATGCTTCTTCTGGTAATTCTTTTTGGCAAAGAATTGTTACTTCCATTGCTCAATCTAAAGCTGTTTCTGGTA AAAGATCCTAA

[0158] SEQ ID NO: 29 AA (AtERVl)

[0159] MGEKPWQPLLQSFEKLSNCVQTHLSNFIGIKNTPPSSQSTIQNPIISLDSSPPIATNSSSLQKLPLKDKSTGPVTKEDLGRATWTFLHTLAAQYPEKPTRQQKKDVKELMTILSRMYPCRECADHFKEILRSNPAQAGSQEEFSQWL CHVHNTVNRSLGKLVFPCERVDARWGKLECEQKSCDLHGTSMDF

[0160] SEQ ID NO: 30 DNA (AtERVl)

[0161] ATGGGTGAAAAACCATGGCAACCATTGTTGCAATCTTTCGAAAAGTTGTCTAATTGTGTTCAAACTCATTTGTCTAACTTCATTGGTATTAAGAACACTCCACCATCTTCTCAATCTACTATTCAAAACCCTATTATCT CTTTGGATTCTTCTCCACCAATTGCTACTAATTCTTCTTCTTTGCAAAAGTT GCCTTTGAAGGATAAGTCTACTGGTCCAGTTACTAAGGAAGATTTGGGTA GAGCTACTTGGACTTTTCTTCATACTTTGGCTGCTCAATACCCTGAAAAACCTACTAGACAACAAAAGAAAGATGTTAAGGAATTGATGACTATCTTGTCTAGAATGTATCCATGTAGAGAATGTGCTGATCATTTCAAAGAAATTTTGAGATCCAACCCTGCTCAAGCTGGTTCTCAAGAAGAATTTTCTCAATGGTTGTG TCATGTTCATAACACTGTTAATAGATCCTTGGGTAAATTGGTTTTCCCTTGT GAAAGAGTTGATGCTAGATGGGGTAAATTGGAATGTGAACAAAAATCTTG TGACTTGCATGGTACTTCTATGGATTTTTAA

[0162] SEQ ID NO: 31 AA (PpERV2)

[0163] MIKFNKRVATLTATLLSFIVLYTLFNSGARFANQLDQPVPLKTPELIIPNQSTKNDAPLPFMPKMANETLKAELGNASWKLFHTILARYPESP SENQKSTLNDYIYLFAQVYPCGDCARHFNLLLQKYPPQLSSRQVAAVWGCHI HNQVNKRLEKPQYDCSNILEDYDCGCGSDEKEVDDTLNNETMEHLQSIKITE KENEQFGR

[0164] SEQ ID NO: 32 DNA (PpERV2)

[0165] ATGATAACATTCAACAAACGAATAGCAACATTAGCGGCAACGTTATTTTCATTCATTGTGCTTTATACTCTCTTTAACAGTGGTGCTCAATTTTCCAACCAACTAGATCAGCCTGTTCCCCTCAAAACTCCAGAACTCAT CATACCGAATCAGAGTACTGAGAATGATCCCCCTCTTCCATTCATGCCAAA AATGGCTAACGAAACTTTGAAAGCAGAACTTGGAAATGCTTCCTGGAAAC TCTTTCACACTATTCTTGCTAGATATCCTGAATCCCCATCGGAGAATCAAAAATCAACCTTAAATGACTACATTTATTTGTTTGCACAGGTTTATCCATGTG GAGACTGTGCAAGACATTTCAATTTATTGCTGCAGAAATACCCTCCACAAT TGTCCTCAAGACAGGTGGCTGCAGTGTGGGGATGTCATATTCACAATCAG GTCAATAAGAGATTGGAGAAACCACAATACGACTGCTCCAATATTCTAGAGGATTACGATTGTGGATGTGGCTCTGATGAAAAGGAAGTAGATGACACTC TGAATAACGAAACAATAGAACACTTGCAAAGTATCAAAATTACTGAAAAA GAGAGTGAACAATTTGGTCGA

[0166] SEQ ID NO: 33 AA (PpEROl)

[0167] MRIVRSLAVTITCYCITALANPQIPFDGNYTEITVPDTEVNIGQIVDINHEIKPKLVELVNTDFFKYYKLNLWKPCPFWNGDEGFCKYKDCS VDFITDWSQVPDIWQPDQLGKLGDNTVHKDKGQDENELSSNDYCALDKDDD EDLVYVNLIDNPERFTGYGGQQSESIWTAVYDENCFQPNEGSQLGQVEDLCL EKQIFYRLVSGLHSSISTHLTNEYLNLKNGEYEPNLKQFMIKVGYFTERIQNLHLNYVLVLKSLIKLQEYNVIENLPLDDSLKAGLSGLISQGAQNINQTDDYLFNE KVLFQNDQNDDLKNEFRDKFRNVTRLMDCVHCERCKLWGKLQTTGYGTAL KILFDLKNPNDSINLKRVELVALVNTFHRLSKSVESIENFEKLYKIQPPTQDHP SPSSESLDVFDNEDEQNFFDSFSVDQTVTSSKEPPEEIKSKPVGKAEYKKTNSCPSSGSKSIKEAFHEELYAFIDAIGFILNSYRTLPKLLYTLFLVKSSELWDIFIGTQ RHRDSTYRVDL

[0168] SEQ ID NO: 34 DNA (PpEROl)

[0169] ATGAGGATAGTAAGGAGCGTAGCTATCGCAATAGCCTGTCATTGTATAACAGCGTTAGCAAACCCTCAAATCCCTTTTGACGGCAACTACACCGAGATCATCGTGCCAGATACCGAAGTTAACATCGGACAGATTGTA GATATTAACCACGAAATAAAACCCAAACTGGTGGAACTGGTCAACACAGA CTTCTTCAAATATTACAAATTAAACCTATGGAAACCATGTCCGTTTTGGAA TGGTGATGAGGGATTCTGCAAGTATAAGGATTGCTCTGTTGACTTTATCACTGATTGGTCCCAGGTGCCTGATATCTGGCAACCAGACCAATTGGGTAAGCTTGGAGATAACACGGTACATAAGGATAAGGGCCAAGATGAAAATGAGCTGTCCTCAAATGATTATTGCGCTTTGGATAAAGACGACGATGAAGATTTAG TATATGTCAATTTGATTGATAACCCTGAAAGATTCACCGGTTATGGTGGTC AGCAATCTGAATCTATTTGGACTGCGGTCTATGATGAGAACTGTTTCCAGC CGAATGAAGGATCACAATTGGGTCAAGTTGAAGACCTCTGTTTGGAGAAACAAATCTTTTACCGATTGGTTTCTGGTTTGCATTCTAGTATCTCCACCCACCTCACAAACGAATATCTGAATTTGAAAAATGGAGCATACGAACCAAATTTG AAACAGTTCATGATCAAAGTTGGGTATTTTACTGAAAGAATCCAAAACTT ACATCTCAATTATGTCCTTGTATTGAAGTCACTAATAAAGCTACAAGAATA CAATGTTATCGACAATCTACCTCTCGATGACTCTTTGAAAGCTGGTCTTAGCGGTTTAATATCTCAAGGAGCACAGGGTATTAACCAGAGTTCTGATGATTATCTATTTAACGAGAAGGTTCTTTTCCAAAATGACCAAAATGATGATTTGAAAAATGAATTTCGTGACAAATTCCGCAACGTGACTAGATTAATGGATTGT GTCCATTGCGAGAGATGCAAATTATGGGGAAAATTGCAAACTACAGGGTACGGGACTGCATTGAAGATTCTATTTGATTTGAAGAATCCTAATGACTCCAT CAATTTAAAGAGAGTTGAGTTAGTTGCTCTAGTCAACACATTCCATAGATT GTCCAAATCTGTTGAAAGCATTGAAAACTTTGAAAAACTATATAAGATTC AACCGCCAACGCAGGATCGTGCATCAGCGTCGTCCGAATCCTTAGGCCTT TTCGATAACGAAGATGAACAAAATCTCCTCAACTCGTTTTCGGTTGATCAG GCAGTCATTTCATCGAAAGAGGCACCAGAAGAAATCAAAAGCAAACCTGT TGGAAAAGCCGCATATAAACAAAACAGTTGTCCATCATTGGGTTCAAAAT CTATCAAAGAAGCATTCCATGAAGAACTTCACGCATTTATTGATGCAATTG GATTTATATTGAACTCTTACAGGACTTTGCCCAAGCTGTTGTACACACTTT TCCTCGTTAAATCATCTGAATTATGGGACATTTTCATTGGCACTCAAAGGC ACCGAGATACCACATATAGAGTAGACTTGTAAGCGGCCGCCAGCTT

[0170] SEQ ID NO: 35 AA (PpKAR2)

[0171] MLSLKPSWLTLAALMYAMLLVVVPFAKPVRADDVESYGTVIGIDLGTTYSCVGVMKSGRVEILANDQGNRITPSYVSFTEDERLVGDAAK NLAASNPKNTIFDIKRLIGMKYDAPEVQRDLKRLPYTVKSKNGQPVVSVEYK GEEKSFTPEEISAMVLGKMKLIAEDYLGKKVTHAVVTVPAYFNDAQRQATK DAGLIAGLTVLRIVNEPTAAALAYGLDKTGEERQIIVYDLGGGTFDVSLLSIEG GAFEVLATAGDTHLGGEDFDYRVVRHFVKIFKKKHNIDISNNDKALGKLKRE VEKAKRTLSSQMTTRIEIDSFVDGIDFSEQLSRAKFEEINIELFKKTLKPVEQVL KDAGVKKSEIDDIVLVGGSTRIPKVQQLLEDYFDGKKASKGINPDEAVAYGA AVQAGVLSGEEGVDDIVLLDVNPLTLGIETTGGVMTTLINRNTAIPTKKSQIFS TAADNQPTVLIQVYEGERALAKDNNLLGKFELTGIPPAPRGTPQVEVTFVLDA NGILKVSATDKGTGKSESITINNDRGRLSKEEVDRMVEEAEKYAAEDAALRE KIEARNALENYAHSLRNQVTDDSETGLGSKLDEDDKETLTDAIKDTLEFLEDN FDTATKEELDEQREKLSKIAYPITSKLYGAPEGGTPPGGQGFDDDDGDFDYDY DYDHDEL

[0172] SEQ ID NO: 36 DNA (PpKAR2)

[0173] ATGCTGTCGTTAAAACCATCTTGGCTGACTTTGGCGGCATTAATGTATGCCATGCTATTGGTCGTAGTGCCATTTGCTAAACCTGTTAG AGCTGACGATGTCGAATCTTATGGAACAGTGATTGGTATCGATTTGGGTA CCACGTACTCTTGTGTCGGTGTGATGAAGTCGGGTCGTGTAGAAATTCTTG CTAATGACCAAGGTAACAGAATCACTCCTTCCTACGTTAGTTTCACTGAAGACGAGAGACTGGTTGGTGATGCTGCTAAGAACTTAGCTGCTTCTAACCCAAAAAACACCATCTTTGATATTAAGAGATTGATCGGTATGAAGTATGATGCCCCAGAGGTCCAAAGAGACTTGAAGCGTCTTCCTTACACTGTCAAGAGCAAGAACGGCCAACCTGTCGTTTCTGTCGAGTACAAGGGTGAGGAGAAGTCTTTCACTCCTGAGGAGATTTCCGCCATGGTCTTGGGTAAGATGAAGTTGATCGCTGAGGACTACTTAGGAAAGAAAGTCACTCATGCTGTCGTTACCGTTCCAGCCTACTTCAACGACGCTCAACGTCAAGCCACTAAGGATGCCGGTCTGATCGCCGGTTTGACTGTTCTGAGAATTGTGAACGAGCCTACCGCCGCTGCCCTTGCTTACGGTTTGGACAAGACTGGTGAGGAAAGACAGATCATCGTCTACGACTTGGGTGGAGGAACCTTCGATGTTTCTCTGCTTTCTATTGAGGGTGGTGCTTTCGAGGTTCTTGCTACCGCCGGTGACACCCACTTGGGTGGTGAGGACTTTGACTACAGAGTTGTTCGCCACTTCGTTAAGATTTTCAAGAAGAAGCATAACATTGACATCAGCAACAATGATAAGGCTTTAGGTAAGCTGAAGAGAGAGGTCGAAAAGGCCAAGCGTACTTTGTCTTCCCAGATGACTACCAGAATTGAGATTGACTCTTTCGTTGACGGTATCGACTTCTCTGAGCAACTGTCTAGAGCTAAGTTTGAGGAGATCAACATTGAATTATTCAAGAAGACACTGAAACCAGTTGAACAAGTCCTCAAAGACGCTGGTGTCAAGAAATCTGAAATTGATGACATTGTCTTGGTTGGTGGTTCTACCAGAATCCCAAAGGTTCAACAATTATTGGAGGATTACTTTGACGGAAAGAAGGCTTCTAAGGGAATTAACCCAGATGAAGCTGTCGCATACGGTGCTGCTGTTCAGGCTGGTGTTTTGTCTGGTGAGGAAGGTGTCGATGACATCGTCTTGCTTGATGTGAACCCCCTAACTCTGGGTATCGAGACTACTGGTGGCGTTATGACTACCTTAATCAACAGAAACACTGCTATCCCAACTAAGAAATCTCAAATTTTCTCCACTGCTGCTGACAACCAGCCAACTGTGTTGATTCAAGTTTATGAGGGTGAGAGAGCCTTGGCTAAGGACAACAACTTGCTTGGTAAATTCGAGCTGACTGGTATTCCACCAGCTCCAAGAGGTACTCCTCAAGTTGAGGTTACTTTTGTTTTAGACGCTAACGGAATTTTGAAGGTTTCTGCCACCGATAAGGGAACTGGAAAATCCGAGTCCATCACCATCAACAATGATCGTGGTAGATTGTCCAAGGAGGAGGTTGACCGTATGGTTGAAGAGGCCGAGAAGTACGCCGCTGAGGATGCTGCACTAAGAGAAAAGATTGAGGCTAGAAACGCTCTGGAGAACTACGCTCATTCCCTTAGGAACCAAGTTACTGATGACTCTGAAACCGGGCTTGGTTCTAAATTGGACGAGGACGACAAAGAGACATTGACAGATGCCATCAAAGATACCCTAGAGTTCTTGGAAGACAACTTCGACACCGCAACCAAGGAAGAATTAGACGAACAAAGAGAAAAGCTTTCCAAGATTGCTTACCCAATCACTTCTAAGCTATACGGTGCTCCAGAGGGTGGTACTCCACCTGGTGGTCAAGGTTTTGACGATGATGATGGAGACTTTGACTACGACTATGACTATGATCATGATGAGTTGTAA

[0174] SEQ ID NO: 37 AA (PpSECl)

[0175] MDLVKVGQSYVDKIVTDTGIKVLLLDDITSSIISLVSTQSELLNHQVYLIDKLENENRDTIKQLDCVCFLSVSEKTINLLVEELGAPKYKSYKLYFNNVVPNSFLERLAERDDLEMVDKVMELFLDYDILNKNLFSFKQLNIFNSIDAWNQQQFLLTLASLKSLCFSLQTNPIIRYESNSRMCSKLASDLSYEFGQSSKIMEKFPVNDIPPVLLILDRKNDPITPLLNPWTYQSMVHELLGIFNNTVDLTGTPSDLPPDLIKLVLNPSQDPFYAQSLYLNFGDLSDSIKTYVNEYKEKTVKHNSNELTDLNDMKHFLESFPEFKKLSNNISKHMGLITELDRKINENHLWQVSELEQSIAVNDNHNADLQELEKLLTSQEFKIANNLKVKLVCLYAIRYELHPNNQLPKMLSILLQQGVPEFEINTVNRMLKYSGSTKRLNDDSESSIFNQATNNLLQGFKQSHENDNIYMQHIPRLERVISKLVKNKLPTAHYPTLINDFLKKQRPVSDLNGARLQDIIIFFVGGVTYEEARIINNFNLVNKSTRIVIGGTTVHNTNSFMTQVLELE

[0176] SEQ ID NO: 38 DNA (PpSECl)

[0177] ATGGACTTGGTTAAGGTTGGACAATCCTACGTGGATAAAATTGTCACAGACACAGGCATTAAGGTTCTTTTATTGGATGATATCACTTCTTCCATAATTTCCCTAGTGAGCACCCAATCAGAATTGTTGAACCATCAGGTGTATTTGATCGACAAGTTGGAGAACGAGAATAGAGATACGATAAAGCAATTGGATTGTGTGTGTTTCCTATCAGTATCAGAAAAAACTATAAACTTGCTTGTTGAGGAATTAGGTGCTCCCAAATACAAATCCTACAAGCTCTACTTCAATAATGTAGTTCCCAACTCATTCTTAGAGAGGTTGGCGGAGAGGGACGATTTGGAAATGGTCGATAAGGTCATGGAATTGTTCCTAGATTACGACATTTTGAACAAGAACTTGTTTTCCTTCAAACAACTGAATATTTTCAATTCAATTGATGCTTGGAATCAGCAACAGTTTCTCTTGACTTTAGCAAGCTTGAAATCACTCTGCTTCTCCTTGCAAACGAATCCTATAATCAGGTATGAATCTAATAGTCGAATGTGTTCTAAGCTAGCTTCCGATTTGTCATACGAATTTGGGCAAAGTTCTAAAATTATGGAAAAGTTCCCGGTGAATGATATCCCTCCTGTCCTGTTAATTCTTGACCGAAAAAACGACCCAATCACTCCATTATTAAATCCTTGGACTTATCAATCTATGGTACACGAGCTTTTAGGAATTTTCAATAATACGGTGGATTTAACGGGAACTCCTTCTGATCTGCCCCCAGACCTAATCAAACTGGTATTGAATCCCTCTCAAGATCCATTTTATGCTCAGTCTCTATATTTGAATTTCGGAGACTTGTCCGATAGTATAAAAACATACGTAAACGAGTACAAAGAAAAAACCGTCAAACACAATTCTAATGAATTGACAGATTTGAATGATATGAAACACTTTCTGGAATCTTTTCCAGAGTTCAAAAAACTTTCAAACAACATTTCCAAACACATGGGCTTGATTACAGAATTAGATAGAAAAATCAACGAAAATCACTTATGGCAAGTGAGTGAATTGGAACAATCCATAGCTGTTAATGACAATCATAATGCTGACCTTCAAGAACTAGAAAAGCTGTTGACATCTCAAGAGTTCAAGATTGCCAACAACTTAAAAGTTAAATTAGTATGTTTGTATGCCATACGATATGAACTTCATCCCAACAACCAGCTTCCAAAAATGTTGTCAATACTTTTACAGCAGGGGGTGCCAGAGTTTGAAATAAATACAGTCAACAGGATGTTGAAATACTCGGGAAGTACCAAACGATTGAATGATGACTCTGAATCTTCGATATTTAACCAGGCAACAAATAATCTACTGCAGGGGTTCAAACAAAGTCATGAAAACGACAATATTTATATGCAGCATATTCCAAGGTTGGAAAGAGTTATCAGCAAGTTAGTGAAAAATAAGCTACCCACAGCGCATTATCCGACTTTAATCAATGATTTTTTGAAGAAGCAACGCCCTGTTTCTGATCTAAATGGAGCCAGGCTGCAAGATATTATTATTTTCTTTGTTGGTGGAGTCACTTATGAAGAGGCCCGAATAATTAACAATTTCAATCTGGTGAACAAGTCTACGAGGATAGTTATAGGGGGAACTACAGTACACAACACGAATAGTTTTATGACTCAAGTTCTAGAATTGGAG TAA

[0178] SEQ ID NO: 39 AA (PpSLYl)

[0179] MSFTTSLPSLRDRQIATLEKMLHLNEPIVDNGSDIQAELTWKVLILDSRSTAIVSSVLRVNDLLSSGITMHSNIRSKRAALPDVPVIYFVEPNAENINFIIDDLERDQYAHFYINFTSSLNRDLLEEFAKKVATIGKSYKIKQVYDQYLDYIVTEPNLFSLDLVNIYSQLNNPNSLEDEINKVADKISNGIFAAILTMNGIPTIRCCRGGPAELIASKLDQKLRDHVINTKSSASFTNSKLVLILLDRNIDLASMFAHSWIYQCMVSDVFELKRNTIKIPSQKPNESTKEYDIDPKDFFWAANNSLPFPDAVENVENELSRYKADAAELTRKTGVSSLQDIDPNAITDTTDIQLAVKSLPELAFRKSILDMHMKVLASLLQELESKSLDSYFEIEQNYKDPKNQKQFISILNNGNEHTLNDKLRTYIMLYLLTDLPGSFVEECEEYFKKNSAELGSLSYIKRAKEVIKLSNYELSMSIDASHSTTSGLVNEAQKSALFQGLSSKLYGLTDGGSRLTEGVGSLITGLKNLLPDKKQLPITNIVESIMEPSLATQESIKLTDDYLYFDPISTRGVHSKPPKRQQYNNSIVFVVGGGNYLEYQNLQEWVTKTNTSNVNGTKSVIYGSTSIVTANEFLKECSLLGAEAK

[0180] SEQ ID NO: 40 DNA (PpSLYl)

[0181] ATGCTTCATTTGAATGAGCCCATTGTGGATAATGGTTCAGATATACAAGCGGAGTTAACATGGAAGGTACTGATTCTGGATAGTAGGAGTACTGCAATTGTTTCTTCTGTTCTGCGAGTTAATGACCTGCTTTCTTCTGGCATCACTATGCATAGCAATATCAGATCCAAGAGAGCGGCTTTGCCAGATGTTCCTGTCATTTACTTTGTTGAACCTAATGCGGAAAATATCAACTTTATCATTGATGACTTGGAAAGAGATCAGTACGCTCATTTTTATATCAACTTCACTTCCAGTCTAAATAGGGACCTTTTGGAGGAGTTTGCTAAGAAAGTGGCTACGATTGGTAAGTCCTACAAGATTAAACAGGTTTATGATCAGTACCTCGATTACATTGTCACTGAACCCAACCTGTTCTCTTTGGACTTGGTTAACATTTACTCGCAGCTAAATAACCCTAACTCACTGGAAGATGAAATCAATAAAGTTGCTGACAAGATTTCCAATGGTATATTCGCAGCAATCCTAACTATGAATGGTATCCCTACTATTAGATGTTGCAGAGGAGGTCCAGCAGAACTAATAGCGTCCAAACTAGATCAGAAGCTACGTGATCATGTTATCAATACAAAGTCATCTGCCTCTTTCACTAACAGTAAATTAGTGCTTATCCTGCTGGATAGAAACATTGATTTGGCTTCCATGTTTGCTCATTCATGGATTTATCAATGTATGGTGAGTGATGTTTTTGAGTTGAAAAGAAATACAATCAAAATTCCCTCTCAAAAGCCCAATGAATCTACGAAAGAATATGATATCGACCCAAAGGATTTTTTTTGGGCAGCCAACAACAGTTTGCCCTTCCCTGATGCTGTAGAAAATGTGGAGAACGAACTTTCTAGATACAAAGCGGATGCTGCAGAGCTAACTAGAAAGACTGGGGTTTCTTCTCTTCAAGATATTGATCCCAATGCAATTACTGACACCACAGATATACAGCTTGCTGTGAAGTCTTTACCTGAATTGGCTTTTAGAAAAAGCATCCTTGATATGCACATGAAAGTACTTGCGTCTTTGCTGCAAGAACTGGAATCAAAGTCATTGGATTCATACTTTGAAATTGAACAAAACTACAAAGATCCCAAAAACCAGAAGCAGTTTATCAGTATCCTCAACAACGGGAATGAGCATACCTTGAACGACAAACTGAGAACCTACATCATGTTGTATCTGTTAACAGACCTCCCAGGGTCGTTCGTTGAAGAATGTGAAGAGTATTTCAAAAAGAACTCCGCTGAGCTTGGTTCGTTGAGTTATATCAAGCGGGCAAAAGAGGTGATCAAGTTGTCTAATTATGAGTTGTCCATGTCAATTGATGCTAGCCACTCGACCACTAGTGGATTGGTGAATGAAGCTCAAAAGTCTGCTTTGTTCCAAGGATTGTCGTCCAAGCTATATGGATTAACAGATGGTGGTAGTAGGCTTACAGAGGGGGTGGGGTCATTAATTACTGGGTTGAAAAACTTGCTACCCGACAAGAAACAACTGCCTATTACCAATATTGTTGAATCGATAATGGAACCAAGTCTGGCCACTCAAGAGTCGATAAAACTAACGGACGATTACCTATATTTTGACCCTATTAGCACAAGAGGAGTTCACTCCAAACCACCCAAAAGACAGCAATACAACAATTCTATTGTGTTTGTTGTAGGAGGGGGCAACTATTTGGAGTACCAAAATTTGCAAGAATGGGTTACGAAGACCAATACTAGCAACGTCAATGGCACTAAGTCTGTAATCTACGGTAGTACCAGTATCGTGACCGCGAACGAGTTCTTGAAGGAGTGCTCCTTGCTCGGTGCCGAAGCAAAATAA

[0182] SEQ ID NO: 41 AA (PpGPXl)

[0183] MSSFYDLAPLDKKGEPFPFEQLKGKVVLIVNVASKCGFTPQYTELEKLYKDHKDEGLTIVGFPCNQFGHQEPGNDEEIGQFCQLNFGVTFPILKKIDVNGSEADPVYEFLKSKKSGLLGFKGIKWNFEKFLIDKQGNVIERYSSL TKPS SIESKIEELLKK

[0184] SEQ ID NO: 42 DNA (PpGPXl)

[0185] ATGTCTTCATTTTATGATCTGGCCCCATTAGATAAGAAAGGCGAACCTTTTCCTTTCGAACAATTAAAAGGCAAAGTGGTGTTGATTGTGAATGTTGCTTCTAAGTGTGGGTTTACTCCACAATATACCGAGTTGGAAAAGCTCTACAAAGACCACAAGGACGAGGGATTGACTATTGTCGGATTTCCCTGTAACCAGTTTGGTCATCAGGAACCAGGAAATGATGAAGAAATTGGACAGTTTTGCCAGTTGAATTTTGGTGTAACTTTCCCAATTCTAAAAAAGATTGATGTCAACGGTTCGGAAGCTGATCCTGTTTACGAATTTCTCAAGTCAAAAAAGTCTGGTCTGCTCGGATTCAAAGGTATTAAGTGGAACTTTGAAAAATTCTTGATCGATAAGCAAGGAAACGTTATTGAGAGATATTCGTCCTTGACTAAGCCCTCATCGATCGAGTCCAAGATTGAAGAACTATTAAAGAAATAA

Claims

WHAT IS CLAIMED IS:

1. A codon optimized nucleic acid sequence encoding a tandem repeat amino acid sequence corresponding to an identified peptide.

2. The codon optimized nucleic acid sequence of claim 1, wherein the nucleic acid sequence is optimized for E. coli or Pichia pastoris expression.

3. The nucleic acid sequence of claim 1 or 2, wherein the nucleic acid sequence comprises SEQ ID NOs: 2, 6, 10,14, or 18 or a nucleic acid sequence with at least 85%, 90%, 95%, 99%, or 99.9% sequence homology to any one of SEQ ID NOs: 2, 6, 10, 14, or 18.

4. The nucleic acid sequence of claim 1 or 2, wherein the nucleic acid sequence comprises SEQ ID NO: 2 or a nucleic acid sequence with at least 85%, 90%, 95%, 99%, or 99.9% sequence homology to SEQ ID NO: 2.

5. The nucleic acid sequence of claim 1 or 2, wherein the nucleic acid sequence comprises SEQ ID NO: 6 or a nucleic acid sequence with at least 85%, 90%, 95%, 99%, or 99.9% sequence homology to SEQ ID NO: 6.

6. The nucleic acid sequence of claim 1 or 2, wherein the nucleic acid sequence comprises SEQ ID NO: 10 or a nucleic acid sequence with at least 85%, 90%, 95%, 99%, or 99.9% sequence homology to SEQ ID NO: 10.

7. The nucleic acid sequence of claim 1 or 2, wherein the nucleic acid sequence comprises SEQ ID NO: 14 or a nucleic acid sequence with at least 85%, 90%, 95%, 99%, or 99.9% sequence homology to SEQ ID NO: 14.

8. The nucleic acid sequence of claim 1 or 2, wherein the nucleic acid sequence comprises SEQ ID NO: 18 or a nucleic acid sequence with at least 85%, 90%, 95%, 99%, or 99.9% sequence homology to SEQ ID NO: 18.

9. A plasmid comprising a nucleic acid sequence encoding a tandem repeat amino acid sequence corresponding to an identified peptide.

10. The plasmid of claim 9, wherein the plasmid can be expressed in E. coli or Pichia pastoris.

11. The plasmid of claim 9 or 10, wherein the nucleic acid sequence comprises any one of SEQ ID NOs: 2, 6, 10,14, or 18 or a nucleic acid sequence with at least85%, 90%, 95%, 99%, or 99.9% sequence homology to any one of SEQ ID NOs: 2, 6, 10, 14, or 18.

12. The plasmid of claim 9 or 10, wherein the nucleic acid sequence comprises SEQ ID NO: 2 or a nucleic acid sequence with at least 85%, 90%, 95%, 99%, or 99.9% sequence homology to SEQ ID NO: 2.

13. The plasmid of claim 9 or 10, wherein the nucleic acid sequence comprises SEQ ID NO: 6 or a nucleic acid sequence with at least 85%, 90%, 95%, 99%, or 99.9% sequence homology to SEQ ID NO: 6.

14. The plasmid of claim 9 or 10, wherein the nucleic acid sequence comprises SEQ ID NO: 10 or a nucleic acid sequence with at least 85%, 90%, 95%, 99%, or 99.9% sequence homology to SEQ ID NO: 10.

15. The plasmid of claim 9 or 10, wherein the nucleic acid sequence comprises SEQ ID NO: 14 or a nucleic acid sequence with at least 85%, 90%, 95%, 99%, or 99.9% sequence homology to SEQ ID NO: 14.

16. The plasmid of claim 9 or 10, wherein the nucleic acid sequence comprises SEQ ID NO: 18 or a nucleic acid sequence with at least 85%, 90%, 95%, 99%, or 99.9% sequence homology to SEQ ID NO: 18.

17. The plasmid of any one of claims 9 to 16, further comprising the nucleic acid sequence of any one of SEQ ID NOs: 20, 22, 24, 26, 28, 30, 32, 34, 36, 38, 40, or 42 or a nucleic acid sequence with at least 85%, 90%, 95%, 99%, or 99.9% sequence homology to any one of SEQ ID NOs: 20, 22, 24, 26, 28, 30, 32, 34, 36, 38, 40, or 42.

18. The plasmid of any one of claims 9 to 16, further comprising the nucleic acid sequence of any one of SEQ ID NOs: 20, 22, 24, 26, 30, 32, 34, or 36 or a nucleic acid sequence with at least 85%, 90%, 95%, 99%, or 99.9% sequence homology to any one of SEQ ID NOs: 20, 22, 24, 26, 30, 32, 34, or 36.

19. The plasmid of any one of claims 17 or 18, comprising the nucleic acid sequence of any one of SEQ ID NOs: 20 or 36 or a nucleic acid sequence with at least 85%, 90%, 95%, 99%, or 99.9% sequence homology to any one of SEQ ID NOs: 20 or 36.

20. A plasmid comprising a nucleic acid sequence of any one or more of SEQ ID NO: 20, 22, 24, 26, 28, 30, 32, 34, 36, 38, 40, or 42 or a nucleotide sequence with atleast 85%, 90%, 95%, 99%, or 99.9% sequence homology to any one or more of SEQ ID NO: 20, 22, 24, 26, 28, 30, 32, 34, 36, 38, 40, or 42.

21. A recombinantly produced tandem repeat amino acid sequence corresponding to an identified peptide.

22. The recombinantly produced tandem repeat amino acid sequence of claim 21, wherein the recombinantly produced tandem repeat amino acid sequence corresponding to an identified peptide comprises an amino acid sequence of any one of SEQ ID NOs: 1, 5, 9, 13, and 17 or an amino acid sequence with at least 85%, 90%, 95%, 99%, or 99.9% sequence homology to any one of SEQ ID NOs: 1, 5, 9, 13, and 17.

23. The recombinantly produced tandem repeat amino acid sequence of claim 21 or 22, wherein the recombinantly produced tandem repeat amino acid sequence corresponding to an identified peptide comprises an amino acid sequence of SEQ ID NO: 1 or an amino acid sequence with at least 85%, 90%, 95%, 99%, or 99.9% sequence homology to SEQ ID NO: 1.

24. The recombinantly produced tandem repeat amino acid sequence of claim 21 or 22, wherein the recombinantly produced tandem repeat amino acid sequence corresponding to an identified peptide comprises an amino acid sequence of SEQ ID NO: 5 or an amino acid sequence with at least 85%, 90%, 95%, 99%, or 99.9% sequence homology to SEQ ID NO: 5.

25. The recombinantly produced tandem repeat amino acid sequence of claim 21 or 22, wherein the recombinantly produced tandem repeat amino acid sequence corresponding to an identified peptide comprises an amino acid sequence of SEQ ID NO: 9 or an amino acid sequence with at least 85%, 90%, 95%, 99%, or 99.9% sequence homology to SEQ ID NO: 9.

26. The recombinantly produced tandem repeat amino acid sequence of claim 21 or 22, wherein the recombinantly produced tandem repeat amino acid sequence corresponding to an identified peptide comprises an amino acid sequence of SEQ ID NO: 13 or an amino acid sequence with at least 85%, 90%, 95%, 99%, or 99.9% sequence homology to SEQ ID NO: 13.

27. The recombinantly produced tandem repeat amino acid sequence of claim 21 or 22, wherein the recombinantly produced tandem repeat amino acid sequencecorresponding to an identified peptide comprises an amino acid sequence of SEQ ID NO: 17 or an amino acid sequence with at least 85%, 90%, 95%, 99%, or 99.9% sequence homology to SEQ ID NO: 17.

28. An E. coli strain comprising the plasmid of any one of claims 9 to 20, the recombinantly produced tandem repeat amino acid sequence of any one of claims 21 to 27, or a combination thereof.

29. A Pichia pastoris strain comprising the plasmid of any one of claims 9 to 20, the recombinantly produced tandem repeat amino acid sequence of any one of claims 21 to 27, or a combination thereof.

30. A method of DNA assembly comprising identifying a peptide and forming a tandem repeat amino acid sequence corresponding to the identified peptide; optimizing the codon for expressing the tandem repeat amino acid sequence for expression in E. coli or Pichia pastoris to form the encoding fragment; forming a middle fragment by adding Bsal restriction sites to the 5' and 3' ends of the final sequence with four base overhangs corresponding to the last two bases and first two bases of the encoding fragment; generating a start fragment by amplifying the middle fragment with a 5' forward primer by adding a Bsal site for the last four bases of a mating factor alpha signal peptide; and generating an end fragment by using a reverse primer that added a stop codon and four bases after the stop codon, thereby forming a vector comprising a start fragment, an encoding fragment, a middle fragment, and an end fragment.

31. The method of claim 30, wherein the tandem repeat amino acid sequences were selected to have a protease cut the peptide between each tandem repeat amino acid sequence but not within the tandem repeat amino acid sequence.

32. The method of claim 31, wherein no protease able to cut the peptide was available, a spacer amino acid was added to create a unique endopeptidase cut site and allow removal of the spacer amino acide from the C-terminus of the peptide by acarboxypeptidase.

33. The method of any one of claims 30 to 32, wherein the number of tandem repeat amino acid sequences in the DNA sequence were adjusted until the tandem repeat amino acid sequence could be synthesized by a commercial DNA vendor.

34. The method of any one of claims 30 to 33, wherein a pHK Pichia expression vector prepared by this method was modified to remove all Bsal sites in the backbone and a GFP flanked by Bsal sites between the mating factor alpha signal peptide and the Notl site.

35. The method of claim 34, wherein a full-length mating factor alpha signal peptide was modified to replace the first 19 amino acids with the 22 amino acids signal peptide from S. cerevisiae Ostl.

36. The method of any one of claims 30 to 35, wherein the start fragment, the middle fragment, and the encoding fragment were ligated for up to 15 cycles.

37. The method of any one of claims 30 to 35, wherein the middle fragment, the end fragment, and the encoding fragment were ligated for up to 15 cycles.

Citation Information

Patent Citations

  • Active short peptide gene engineering biosynthesis process

    CN106560475A

  • Tandem peptide and method for simultaneously preparing multiple bioactive peptides by using recombinant escherichia coli

    CN116496360A

  • Collection of repeat proteins comprising repeat modules

    US7417130B2