Synthetic santalene synthase

By adjusting the tertiary structure of santalene synthase, the problem of α-santalene excess in the synthesis of santalene in existing enzymes was solved, and the efficient generation of β-santalene and bergamotene was achieved, meeting the industrial production demand for β-santalol.

CN115516102BActive Publication Date: 2025-10-31ISOBIONICS BV
View PDF 21 Cites 0 Cited by

Patent Information

Application Number
CN202180033182.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2021-03-02
Filing Date
2021-06-01
Publication Date
2025-10-31
Estimated Expiration
2041-06-01

AI Technical Summary

Technical Problem

In the production process of existing santalene synthases, the yield of α-santalene is much higher than that of β-santalene, which is difficult to meet the industrial demand for β-santalol. Furthermore, traditional enzymes are not conducive to the formation of β-santalene in the reaction.

Method used

By altering the flexibility of the tertiary structure of santalene synthase and adjusting its product ratio, the enzyme can more effectively generate β-santalene and bergamotene while reducing the formation of α-santalene.

Benefits of technology

This method enables flexible control of the ratio of santalene synthase products, increases the production of β-santalene and bergamotene, meets the industrial demand for β-santalol, and optimizes the enzyme's reaction efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure BDA0003926567950000011
    Figure BDA0003926567950000011
  • Figure BDA0003926567950000101
    Figure BDA0003926567950000101
  • Figure BDA0003926567950000102
    Figure BDA0003926567950000102
Patent Text Reader

Abstract

The invention discloses santalene synthases with improved product characteristics and methods for improving santalene synthases. The invention also relates to santalene compositions produced by fermentation, having a higher β-santalene content than α-santalene content.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Santalene synthase is a terpene synthase that catalyzes the conversion of farnesyl pyrophosphate (FPP) into a wide range of compounds, including santalene, such as α-santalene, β-santalene, and epi-β-santalene.

[0002]

[0003] Formula I is a diagram of (-)-β-santalene (CAS No. 511-59-1; hereinafter referred to as β-santalene).

[0004] Santalene synthase begins with the substrate farnesyl pyrophosphate, but generally produces a mixture of sesquiterpene products. Typically, santalene synthase produces (-)-α-santalene (CAS No. 512-61-8; hereinafter referred to as α-santalene) as the main product, followed by β-santalene (see Formula I) and / or trans-α-bergamotene (CAS No. 13474-59-4; hereinafter also referred to as bergamotene) as the second and third most abundant products, respectively. The amount produced depends on the specific enzyme and also on whether β-santalene is the second most abundant or bergamotene is dominant, although α-santalene is dominant in the oils available to date.

[0005] Several genes encoding santalene synthases have been reported (see, for example, International Patent Application WO2018 / 160066 and its references). Furthermore, these santalene synthases produce a range of santalene sesquiterpenes (most notably including β-santalene, α-santalene, epi-β-santalene, bergimene, and β-bisabolene).

[0006] Santalene synthases that produce α-santalene as the main product are known, for example, from WO201100026 and Jones et al. (2011) (“Sandalwood fragrance biosynthesis involves sesquiterpene synthases of both the terpene synthase (TPS)-a and TPS-b subfamilies, including santalene synthases”). Jones CG, Moniodis J., Zulak KG et al., The Journal of Biological Chemistry, Vol. 286, No. 20, pp. 17445-17454, 5 / 20 / 2011; DOI:10.1074 / jbc.M111.231787) describe three different species of the genus *Santalum* (*Santalum*). Terpene synthases that simultaneously produce α-santalene, α-trans-bergeriene, epi-β-santalene, and β-santalene from *Santalum album*, *Santalum austrocaledonicum*, and *Santalum spicatum*. International patent application WO201100026 is subject to possible misinterpretation. Figure 1 The data in Table 2 discloses that more β-santalene than α-santalene may be present, for example, by showing data from semi-quantitative GC-MS analysis, without indicating that this method is unreliable in quantification relative to the amount of the compound. In the same invention disclosure, the opposite is clearly shown to those skilled in the art (more α-santalene than β-santalene) – as shown in column 4 of Table 2 of WO201100026, which shows reliable GC-FID quantification data, and in Table 1 of WO201100026, which reports that α-santalene in natural sandalwood oil is more than twice that of β-santalene.

[0007] The same enzyme has also been reported as α-santalene, rather than β-santalene, in the following publications by researchers after WO201100026, the references of which were also published in 2011 and therefore assume to use the same data basis: “Sandalwood fragrance biosynthesis involves sesquiterpene synthases of both the terpenesynthase (TPS)-a and TPS-b subfamilies, including santalene synthases.” Jones CG, Moniodis J., Zulak KG et al., The Journal of Biological Chemistry, Vol. 286, No. 20, pp. 17445-17454 5 / 20 / 2011; DOI:10.1074 / jbc.M111.231787. The supplementary information to this paper and the revised figures from the initial publication (see: Erratum: Sandalwood fragrance biosynthesis involves sesquiterpene synthases of both the terpene synthase (TPS)-a and TPS-b subfamilies, including santalene synthases) (Journal of Biological Chemistry (2011) 286 (17445-17454)), Journal of Biological Chemistry volume 287, No. 45, pp. 37713-37714 2012, DOI:10.1074 / jbc.A111.231787s) confirm the following fact: these researchers observed more α-santalene than β-santalene.Subsequent publications by these researchers confirmed that natural sandalwood oil does not contain an excess of β-santalene relative to α-santalene (Moniodis et al., 2017, “Sesquiterpene Variation in West Australian Sandalwood (Santalum spicatum)”; Molecules 2017; 22(6)). It is known that santalene synthase produces more α-santalene than β-santalene (Diaz-Chavez et al., 2013, "Biosynthesis of Sandalwood Oil: Santalum album CYP76F cytochromes P450 Produce Santalols and Bergamotol", PLoS ONE, 2013; 8(9)), even when expressed heterologously in tobacco plants (Yin JL, Wong WS (2019) "Production of santalenes and bergamotene in Nicotiana tabacum plants" PLOS ONE 14(1):e0203249. https: / / doi.org / 10.1371 / journal.pone.0203249).

[0008] International patent application WO2015153501 describes a modified santalene synthase derived from sandalwood, which exhibits increased terpene synthase activity compared to natural sandalwood santalene synthase. However, α-santalene is still in excess relative to β-santalene, and santalene synthases with high α-santalene product characteristics have been found (Schalk, M., 2011. Method for Producing Alpha-Santalene, US Patent 2011 / 008836A1; International patent application WO2018160066). International patent application WO2010 / 067309 describes a method for producing β-santalene using santalene synthase from the genus Santalum (Schalk, 2014) (US Patent No. 8993284), but α-santalene still exceeds β-santalene.

[0009] Therefore, enzymes that produce β-santalene in lower quantities than α-santalene are known, and there are no known examples of santalene synthases with a greater in vivo β-santalene product characteristic than α-santalene.

[0010] The products of santalene synthase can be oxidized biosynthetically or chemically to produce their corresponding santalene alcohols: α-santalol, β-santalol, and epi-β-santalol. Santalol is a major component of sandalwood oil, a highly prized naturally occurring fragrance, and an important ingredient in perfumes, cosmetics, cosmetic tools, aromatherapy, and pharmaceuticals. It possesses a soft, sweet, woody, and balsam-like aroma imparted by the sesquiterpene alcohols α-santalol and β-santalol. In particular, β-santalol is considered to impart the most important olfactory note to sandalwood. An enzyme with greater specificity for β-santalene synthase is needed because the product can be oxidized to an oil with a high β-santalol content.

[0011] Currently known santalene synthases have numerous different drawbacks that are particularly unfavorable when used in industrial production processes of santalene, in isolated reactions (in vitro, e.g., using isolated santalene synthases or (permeable) whole cells) or otherwise, for example, in fermentation (in vivo) as part of a longer metabolic pathway that ultimately leads to the production of β-santalene from sugars, to prepare santalene (and possibly subsequently santalol and especially β-santalol).

[0012] If the enzyme produces less α-santalene and more trans-α-bergamerene, it may also benefit some applications.

[0013] It is advantageous to be able to guide the product ratio of the three main products of santalene synthase according to specific needs.

[0014] This invention discloses that, unexpectedly, the product characteristics of santalene synthases can be improved by relatively simply altering the flexibility of a portion of the tertiary structure. Some of these improved santalene synthases produce β-santalene and sometimes bergimene more than α-santalene, while others exhibit increased α-santalene yield compared to the wild-type enzyme, and they can be used to produce these compounds, for example, in large-scale industrial processes. Invention Details

[0016] In connection with a particular attribute or value, the terms “substantially,” “about,” “approximately,” “basically,” etc., further specify and precisely define that attribute or value. In the context of equivalent or substantially equivalent functional activity, the term “substantially” means a difference in function compared to a reference function, preferably within 20%, more preferably within 10%, and most preferably within 5% or less. In the context of formulations and compositions, the term “substantially” (e.g., “a composition substantially composed of compound X”) may be used herein as follows: the formulation or composition substantially contains the indicated compound having a given effect, does not contain other compounds having such effects, or contains the indicated compound in the maximum amount of such compounds that does not exhibit a measurable or meaningful effect. In the case of a given numerical value or range, the term “about” is specifically designed to be within 20%, 10%, or 5% of the given value or range. As used herein, the term “comprising” also includes the term “composed of.”

[0017] The term "isolated" means that the material is substantially free of at least one other component naturally bound to it within its original environment. For example, naturally occurring polynucleotides, polypeptides, or enzymes present in living animals are not isolated, however, the same polynucleotides, polypeptides, or enzymes separated from some or all of the coexisting substances in that natural system are isolated. As yet another example, an isolated nucleic acid molecule, such as DNA or RNA, is a molecule that is not closely adjacent to the 5′ and 3′ flanking sequences that would normally be closely adjacent to the molecule when present in the naturally occurring genome of the organism from which the molecule is derived. Such polynucleotides may be part of a vector, incorporated into the genome of a cell with an unrelated genetic background (or incorporated into the genome of a cell with a substantially similar genetic background, but at a different site than its natural location), or produced by PCR amplification or restriction enzyme digestion, or RNA molecules produced by in vitro transcription, and / or such polynucleotides, polypeptides, or enzymes may be part of a composition and may still be isolated, such that the vector or composition is not part of its natural environment.

[0018] "Purified" means that the substance is in a relatively pure state, for example, at least about 90% purity, at least about 95% purity, or at least about 98% or 99% purity. Preferably, "purified" means that the material is in a 100% pure state.

[0019] "Synthetic" or "artificial" compounds are produced by in vitro chemical synthesis or enzymatic synthesis. This includes, but is not limited to, variant nucleic acids produced by optimal codon selection for a host organism (such as yeast cell host or other selected expression hosts) or variant protein sequences with amino acid modifications (e.g., substitutions) (e.g., to optimize peptide properties) compared to wild-type protein sequences. Synthetic peptides should therefore be understood as peptides that are synthetic, non-naturally occurring "artificial" protein sequences. Preferably, the synthetic peptide differs from any naturally occurring peptide at at least one amino acid position during the present invention.

[0020] The term "non-naturally occurring" refers to (poly)nucleotides, amino acids, (poly)peptides, enzymes, proteins, cells, organisms, or other materials that do not exist in their original environment or source, although they may be originally derived from their original environment or source and subsequently reproduced by other means. Such non-naturally occurring (poly)nucleotides, amino acids, (poly)peptides, enzymes, proteins, cells, organisms, or other materials may be structurally and / or functionally similar to or identical to their natural counterparts.

[0021] The terms “natural” (or “wild-type” or “endogenous”) cell or organism and “natural” (or wild-type or endogenous) polynucleotide or polypeptide refer to cells or organisms that exist in nature and, respectively, to the polynucleotide or polypeptide in question that exists in its natural form and genetic environment (i.e., without any human intervention).

[0022] The term "heterogeneous" (or exogenous, foreign, or recombinant) polypeptide is defined in this paper as:

[0023] (a) It is not a polypeptide native to the host cell. The protein sequence of this heterologous polypeptide is a synthetic, non-natural, "artificial" protein sequence;

[0024] (b) A polypeptide native to the host cell, but containing structural modifications as a result of manipulating the host cell's DNA through recombinant DNA technology that alters the native polypeptide, such as deletion, substitution, and / or insertion; or

[0025] (c) Host cell-native polypeptides whose expression is quantitatively altered or directed from genomic locations different from those of the natural host cell, as a result of manipulating the DNA of the host cell through recombinant DNA technology (e.g., stronger promoters).

[0026] The descriptions in b) and c) above refer to sequences in their native form but not naturally expressed by the cells used to produce them. The resulting polypeptide is therefore more precisely defined as a “recombinantly expressed endogenous polypeptide,” which does not contradict the above definition but reflects the special case that it is not a sequence of proteins with synthetic properties or manipulations, but rather a manner in which polypeptide molecules are produced.

[0027] Similarly, the term "heterologous" (or exogenous, foreign, or recombinant) polynucleotide refers to:

[0028] (a) is not a naturally occurring polynucleotide in the host cell;

[0029] (b) Natural polynucleotides of the host cell, but containing structural modifications, such as deletions, substitutions, and / or insertions, as a result of manipulating the DNA of the host cell by recombinant DNA technology that alters natural polynucleotides;

[0030] (c) Quantitative alterations in the expression of native polynucleotides in the host cell, resulting from the manipulation of polynucleotides via recombinant DNA technology as regulatory elements (e.g., stronger promoters); or

[0031] (d) Naturally occurring polynucleotides of the host cell, but which are not integrated into its natural genetic environment as a result of genetic manipulation via recombinant DNA technology.

[0032] For two or more polynucleotide sequences or two or more amino acid sequences, the term "heterologous" is used to characterize that the two or more polynucleotide sequences or two or more amino acid sequences do not occur naturally in a specific combination of each other.

[0033] The terms “polynucleotide,” “nucleic acid sequence,” “nucleotide sequence,” “nucleic acid,” and “nucleic acid molecule” are used interchangeably in this document and refer to a polymeric, unbranched nucleotide of any length: ribonucleotide or deoxyribonucleotide, or a combination of both.

[0034] For nucleotide sequences (e.g., common sequences), the IUPAC nucleotide nomenclature (International Union of Biochemistry and Biochemistry Committee on Nomenclature (NC-IUB) (1984). "Nomenclature for nucleic acid sequences in which bases are not fully specified") is used, along with the following nucleotide ambiguities that are relevant to this invention: A, adenine; C, cytosine; G, guanine; T, thymine; K, guanine or thymine; R, adenine or guanine; W, adenine or thymine; M, adenine or cytosine; Y, cytosine or thymine; D, not cytosine; N, any nucleotide.

[0035] Additionally, the notation “N(3-5)” indicates that the indicated common position can have 3 to 5 arbitrary (N) nucleotides. For example, the common sequence “AWN(4-6)” represents 3 possible variants – 4, 5, or 6 arbitrary nucleotides at the end: AWNNNN, AWNNNNN, AWNNNNNN.

[0036] The terms “regulatory element” and “regulatory sequence” are used interchangeably herein and should be understood in a broad context as meaning a regulatory nucleic acid sequence that enables the expression of a linked sequence (including, but not limited to, the expression of polynucleotides encoding polypeptides). A regulatory element or regulatory sequence can include any nucleotide sequence that individually has a function or purpose and / or is located within a particular arrangement or within a group of other elements or sequences within that arrangement. Examples of regulatory sequences include, but are not limited to, leader sequences or signal sequences (such as 5'-UTR), start signals, propeptide sequences, promoters, enhancers, silencers, polyadenylation sequences, ribosome binding sites (RBS, SD sequences), termination signals, terminators, 3'-UTRs, and combinations thereof. Regulatory elements or regulatory sequences may be native (i.e., from the same gene) or foreign (i.e., from different genes) to each other or relative to the nucleotide sequence to be expressed.

[0037] "Effective connection" means that the components are in a relationship that allows them to function in their intended manner. For example, a modulating sequence effectively connected to a coding sequence is connected in such a way that the expression of the coding sequence is realized under conditions compatible with the modulating sequence.

[0038] Nucleic acids and peptides can be modified to include tags or domains. Tags can be used for a variety of purposes, including detection, purification, solubilization, or immobilization, and can include, for example, biotin, fluorophores, epitopes, mating factors, or regulatory sequences. Domains can be of any size and provide the desired function (e.g., conferring increased stability, solubility, activity, or simplified purification) and can include, for example, binding domains, signal sequences, promoter sequences, regulatory sequences, N-terminal extensions, or C30-terminal extensions. Combinations of tags and / or domains can also be utilized.

[0039] The term "fusion protein" refers to two or more polypeptides joined together by any means known in the art. These means include chemical synthesis or splicing of coding nucleic acids through recombinant engineering.

[0040] Methods of modifying nucleic acids to introduce changes in the encoded protein.

[0041] Gene editing

[0042] Gene editing, or genome editing, is a type of genetic engineering in which DNA is inserted, replaced, or removed from the genome and can be achieved by: using techniques such as “gene shuffling” or “directed evolution” consisting of repeated DNA shuffling, followed by appropriate screening and / or selection, to produce variant nucleic acids or portions thereof encoding proteins with modified biological activity (Castle et al., (2004) Science 304(5674):1151-4; US Patents 5,811,238 and 6,395,547); or by using “T-DNA activation” tagging (Hayashi et al., Science (1992) 1350-1353), in which the resulting transgenic organisms exhibit a dominant phenotype due to their close approximation to the regulated expression of genes with introduced promoters; or by using “TILLING” (targeted-induced local lesions in the genome). TILLING refers to mutagenesis techniques used to generate and / or identify modified expression and / or activity of proteins encoded by nucleic acids. TILLING also allows for the selection of organisms carrying such mutant variants. Methods for TILLING are well known in the art (McCallum et al., (2000) Nat Biotechnol 18:455-457; reviewed by Stemple (2004) Nat Rev Genet 5(2):145-50). Another technique uses engineered nucleases such as zinc finger nucleases, transcription activator-like effector nucleases (TALENs), CRISPR / Cas systems, and engineered meganucleases, such as reengineered homing endonucleases (Esvelt, KM.; Wang, HH. (2013), MolSyst Biol 9(1):641; Tan, WS. et al. (2012), Adv Genet 80:37–97; Puchta, H.; Fauser, F. (2013), Int. J. Dev. Biol 57:629–637).

[0043] Mutagenesis

[0044] A variety of techniques known in molecular biology can be used to modify DNA and the proteins they encode to produce variant proteins or enzymes with novel or altered properties. For example, random PCR mutagenesis, see, for example, Rice (1992) Proc. Natl. Acad. Sci. USA 89:5467-5471; or combined multi-box mutagenesis, see, for example, Crameri (1995) Biotechniques 18:194-196.

[0045] Alternatively, nucleic acids, such as genes, can be reassembled after random or “arbitrary” fragmentation, see, for example, U.S. Patent Nos. 6,291,242; 6,287,862; 6,287,861; 5,955,358; 5,830,721; 5,824,514; 5,811,238; 5,605,793.

[0046] Alternatively, modifications, additions, or deletions may be introduced by the following methods: error-prone PCR, shuffling, site-directed mutagenesis, assembly PCR, sexual PCR mutagenesis, in vivo mutagenesis (phage-assisted sequential evolution, in vivo sequential evolution), box mutagenesis, recursive ensemble mutagenesis, exponential ensemble mutagenesis, site-specific mutagenesis, gene reassembly, site-saturated mutagenesis (GSSM), synthetic linker reassembly (SLR), recombination, recursive sequence recombination, phosphate-thioester modified DNA mutagenesis, uracil-containing template mutagenesis, gap double-strand mutagenesis, site-mismatch repair mutagenesis, repair-deficient host strain mutagenesis, chemical mutagenesis, radioactive mutagenesis, deletion mutagenesis, restriction-selection mutagenesis, restriction-purification mutagenesis, artificial gene synthesis, ensemble mutagenesis, chimeric nucleic acid multimer creation, and / or combinations of these and other methods.

[0047] Alternatively, “gene site saturation mutagenesis” or “GSSM” includes a method for introducing point mutations into polynucleotides using degenerate oligonucleotide primers, as detailed in U.S. Patent Nos. 6,171,820 and 6,764,835.

[0048] Alternatively, synthetic ligation reassembly (SLR) includes methods for non-randomly linking oligonucleotide structural units together (e.g., as disclosed in U.S. Patent No. 6,537,776).

[0049] Alternatively, tailored multi-site combinatorial assembly (TMSCA) is a method for generating multiple progeny polynucleotides with various combinations of mutations at multiple sites by using at least two mutagenic non-overlapping oligonucleotide primers in a single reaction (as described in PCT Publication No. WO2009 / 018449).

[0050] Sequence alignments can be generated using numerous software tools, such as:

[0051] -Needleman and Wunsch Algorithm-Needleman, Saul B. & Wunsch, Christian D. (1970). "A general method applicable to the search for similarities in the amino acids sequence of two proteins". Journal of Molecular Biology. 48(3): 443–453.

[0052] This algorithm can be implemented, for example, in the "NEEDLE" program, which performs a global alignment of two sequences. The NEEDLE program is included, for example, within the European Molecular Biology Open Software Suite (EMBOSS).

[0053] -EMBOSS- A collection of various programs: The European Molecular Biology Open Software Suite (EMBOSS), Trends in Genetics 16(6), 276(2000).

[0054] -BLOSUM (BLOCKS SUbstitution Matrix) - is generally generated based on alignment of conserved regions (e.g., protein domains) (Henikoff S, Henikoff JG: Amino acid substitution matrices from protein blocks. Proceedings of the National Academy of Sciences of the USA. 1992 Nov 15; 89(22):10915-9). One of the many BLOSUMs is "BLOSUM62", which is often the "default" setting for many programs when aligning protein sequences.

[0055] - BLAST (Basic Local Alignment Retrieval Tool) – consists of several independent programs (BlastP, BlastN, ...) primarily used to retrieve similar sequences from a large sequence database. The BLAST program also generates local alignments. Generally, the "BLAST" interface provided by NCBI is used, which is an improved version ("BLAST2"). “Initial” BLAST: Altschul, SF, Gish, W., Miller, W., Myers, EW and Lipman, DJ (1990) “Basic local alignment search tool” J. Mol. Biol. 215:403-410; BLAST2: Altschul, Stephen F., Thomas L. Madden, Alejandro A. Schaffer, Jinghui Zhang, Zheng Zhang, Webb Miller and David J. Lipman (1997) “Gapped BLAST and PSI-BLAST: a new generation of protein database search programs” Nucleic Acids Res. 25:3389-3402.

[0056] Enzyme variants can be defined by their sequence identity when compared to their parent enzyme. Sequence identity is typically provided as "sequence identity %" or "identity %". To determine the percentage of identity between two amino acid sequences, in a first step, a paired sequence alignment is generated between the two sequences, wherein the two sequences are aligned over their entire length (i.e., a global alignment of the pair). The alignment is generated using a program implementing the Needlem and Wunsch algorithm (J. Mol. Biol. (1979) 48, pp. 443-453), preferably using the program "NEEDLE" (European Molecular Biology Open Software Suite (EMBOSS)) with the program's default parameters (vacancy opening = 10.0, vacancy extension = 0.5, and matrix = EBLOSUM62). The preferred alignment used in this invention is the one from which the highest sequence identity can be determined.

[0057] The following examples are intended to illustrate calculations for two nucleotide sequences, but the same calculations apply to protein sequences:

[0058] Seq A: AAGATACTG Length: 9 bases

[0059] Seq B: GATCTGA length: 7 bases

[0060] Therefore, the shorter sequence is sequence B.

[0061] The generation of pairwise global alignments of two sequences over their entire length range produced

[0062]

[0063] The "I" symbol in the alignment indicates identical residues (meaning bases in DNA or amino acids in proteins). The number of identical residues is 6.

[0064] The "-" sign in the alignment indicates a gap. The number of gaps introduced by the alignment within Seq B is 1. The gaps introduced by the alignment are numbered 2 at the boundary of Seq B and 1 at the boundary of Seq A.

[0065] The alignment length of the alignment sequence displayed across its entire length range is 10.

[0066] The present invention generates pairwise alignments of shorter sequences over their entire length range, thus producing:

[0067]

[0068] According to the present invention, pairwise alignments of sequence A are generated over its entire length, thus producing:

[0069]

[0070] According to the present invention, pairwise alignments of sequence B are generated over its entire length, thus producing:

[0071]

[0072] The alignment length of the shorter sequence is 8 throughout its entire length range (there is a gap that is taken into account in the alignment length of the shorter sequence).

[0073] Therefore, the alignment length of Seq A will be 9 throughout its entire length range (meaning that Seq A is the sequence of the present invention).

[0074] Therefore, the alignment length of Seq B will be 8 over its entire length range (meaning that Seq B is the sequence of this invention).

[0075] After aligning the two sequences, in the second step, an identity value is determined from the resulting alignment. For the purposes of this specification, the identity percentage is calculated as follows: Identity % = (identical residues / length of the alignment region showing the shorter sequence over its entire length) * 100. Therefore, according to this embodiment, sequence identity in relation to comparing two amino acid sequences is calculated by dividing the number of identical residues by the length of the alignment region showing the shorter sequence over its entire length. This value is multiplied by 100 to obtain the "identity %". According to the example provided above, the identity % is: (6 / 8) * 100 = 75%.

[0076] Compared to the full-length polypeptide sequence, variants of sandalwood synthase may have an amino acid sequence that is at least n percent identical to the amino acid sequence of the corresponding parent polypeptide molecule, where n is an integer between 50 and 100, preferably 50, 55, 60, 65, 70, 75, 80, 85, 90, 91, 92, 93, 94, 95, 96, 97, 98, or 99.

[0077] Santalene synthase variants can be defined by their sequence similarity when compared to the parent enzyme. Sequence similarity is typically provided as "sequence similarity %" or "similarity %". To calculate sequence similarity in the first step, sequence alignments must be generated as described above. In the second step, a similarity percentage must be calculated, which takes into account shared similar properties of the defined amino acid set, such as size, hydrophobicity, charge, or other characteristics. Here, an amino acid exchange for a similar amino acid is called a "conserved mutation". Enzyme variants containing conserved mutations appear to have minimal impact on protein folding, resulting in the substantial maintenance of certain enzymatic properties when compared to the parent enzyme.

[0078] To determine the similarity percentage of the invention, the following applies, which also applies to, for example, the BLOSUM62 matrix used by the “NEEDLE” program (as referenced above).

[0079] The matrix is ​​one of the most commonly used amino acid similarity matrices for database retrieval and sequence alignment.

[0080] Amino acid A is similar to amino acid S

[0081] Amino acid D is similar to amino acid E; N

[0082] Amino acid E is similar to amino acids D; K; Q

[0083] Amino acid F is similar to amino acid W; Y

[0084] Amino acid H is similar to amino acid N; Y

[0085] Amino acid I is similar to amino acids L and M; V

[0086] Amino acid K is similar to amino acids E; Q; R

[0087] Amino acid L is similar to amino acids I; M; V

[0088] Amino acid M is similar to amino acids I; L; V

[0089] Amino acid N is similar to amino acids D; H; S

[0090] Amino acid Q is similar to amino acids E; K; R

[0091] Amino acid R is similar to amino acid K; Q

[0092] Amino acid S is similar to amino acids A, N, and T.

[0093] Amino acid T is similar to amino acid S

[0094] Amino acid V is similar to amino acids I; L; M

[0095] Amino acid W is similar to amino acid F; Y

[0096] Amino acid Y is similar to amino acids F, H, and W.

[0097] Conserved amino acid substitutions can occur throughout the full-length sequence of a functional protein such as an enzyme's polypeptide. In one embodiment, such mutations do not involve the enzyme's functional domains. In another embodiment, conserved mutations do not involve the enzyme's catalytic center.

[0098] Therefore, according to this specification, the following similarity percentage calculations apply:

[0099] Similarity % = [Identical residues + Similar residues / Length of the alignment region showing the shorter sequence over its entire length] * 100. Therefore, here, the sequence similarity related to comparing two amino acid sequences according to this embodiment is calculated by dividing the number of identical residues plus the number of similar residues by the length of the alignment region showing the shorter sequence over its entire length. This value is multiplied by 100 to obtain the "Similarity %".

[0100] Variant enzymes containing conserved mutations are expected to have substantially unchanged enzymatic properties, such as enzyme activity, wherein the variant enzyme is at least m percent similar to the corresponding parental sequence compared to the full-length polypeptide sequence, where m is an integer between 50 and 100, preferably 50, 55, 60, 65, 70, 75, 80, 85, 90, 91, 92, 93, 94, 95, 96, 97, 98, or 99.

[0101] As used herein, a “construct,” “genetic construct,” or “expression cassette” (which may be used interchangeably) is a DNA molecule consisting of at least one target sequence to be expressed, effectively linked to one or more regulatory sequences (at least with a promoter) as described herein. Generally, an expression cassette contains three elements: a promoter sequence, a read frame, and a 3' untranslated region, which in eukaryotes typically contains polyadenylation sites. Additional regulatory elements may include transcriptional enhancers and translational enhancers. Intron sequences may also be added to the 5' untranslated region (UTR) or to the coding sequence to increase the amount of mature messengers accumulating in the cytosol. Those skilled in the art are well aware of the genetic elements that must be present in the expression cassette for successful expression. Preferably, the arrangement of at least the DNA portion or genetic elements forming the expression cassette is artificial. The expression cassette may be part of a vector or may be integrated into the genome of a host cell and replicated along with the host cell's genome. The expression cassette is capable of increasing or decreasing the expression of the target DNA and / or protein.

[0102] As used herein, the terms “introduction” or “transformation” include the transfer of exogenous polynucleotides into host cells, regardless of the method used for transformation. That is, as used herein, the term “transformation” is independent of the vector, shuttle system, or host cell, and it not only refers to methods of polynucleotide transfer known in the art (see, for example, Sambrook, J. et al., (1989) Molecular Cloning: A Laboratory Manual, 2nd ed., Cold Spring Harbor Laboratory Press, Cold Spring Harbor, NY), but it also encompasses any other kind of polynucleotide transfer method, such as, but not limited to, transduction or transfection.

[0103] The term "recombinant organism" refers to a eukaryotic organism (yeast, fungus, algae, plant, animal) or prokaryotic microorganism (e.g., bacteria) that has been genetically altered, modified, or engineered so that it exhibits an altered, modified, or different genotype compared to the wild-type organism from which the organism is derived. Preferably, a "recombinant organism" comprises exogenous nucleic acid. "Recombinant organism," "genetically modified organism," and "transgenic organism" are used interchangeably herein. The exogenous nucleic acid may be located on an extrachromosomal DNA fragment (such as a plasmid) or may be integrated into the chromosomal DNA of the organism. In the case of recombinant eukaryotes, this is understood to mean that the nucleic acid used is not present in or derived from the genome of the organism, or is present in the genome of the organism but not at its natural locus in the genome of the organism, and may be expressed under the regulation of one or more endogenous and / or exogenous regulatory elements.

[0104] "Host cell" can be any cell selected from bacterial cells, yeast cells, fungal cells, algal cells or cyanobacterial cells, non-human animal or mammalian cells, or plant cells. Those skilled in the art are well aware that the genetic elements containing the target sequence must be present on the said gene construct for successful transformation, selection, and proliferation of the host cell. The host cell can be selected from any of these organisms:

[0105] bacteria

[0106] ○ Gram-positive: Bacillus and Streptomyces

[0107] ■ Useful Gram-positive bacteria include, but are not limited to, Bacillus cells, such as Bacillus alkalophius, Bacillus amyloliquefacins, Bacillus brevis, Bacillus circulans, Bacillus clausii, Bacillus coagulans, Bacillus firmus, Bacillus Jautus, Bacillus lentus, Bacillus licheniformis, Bacillus megaterium, Bacillus pumilus, Bacillus stearothermophilus, Bacillus subtilis, and Bacillus thuringiensis. Most preferably, the prokaryote is a Bacillus cell, and more preferably a Bacillus subtilis, Bacillus pumilus, Bacillus licheniformis, or Bacillus tarda.

[0108] ■ Other preferred bacterial species include strains from the order Actinomycetales, preferably from the genus Streptomyces, particularly Streptomyces spheroides (ATTC 23965), Streptomyces thermoviolaceus (IFO 12382), Streptomyces lividans, Streptomyces murinus, or Streptomyces verticillium ssp. verticillium. Other preferred bacterial species include Rhodobacter sphaeroides, Rhodomonas palustris, and Streptococcus lactis. Further preferred bacterial species include strains belonging to the genus Myxococcus, such as M. virescens.

[0109] ○ Gram-negative: Escherichia coli, Pseudomonas

[0110] ■ Preferred Gram-negative bacteria are Escherichia coli and Pseudomonas species, preferably Pseudomonas purrocinia (ATCC 15958) or Pseudomonas fluorescens (NRRL B-11).

[0111] Fungi

[0112] ○ Genus Aspergillus, Genus Fusarium, Genus Trichoderma

[0113] ■ Microorganisms can be fungal cells. As used herein, “fungi” includes the phyla Ascomycota, Basidiomycota, Chytridiomycota, Zygomycota, Oomycota, Deuteromycotina, and all deuteromycetes. Representative groups of Ascomycota include, for example, *Neurospora*, *Eupenicillium* (= *Penicillium*), *Emericella* (= *Aspergillus*)), *Eurotium* (= *Aspergillus*) and the *True Yeasts* listed below. Examples of Basidiomycota include mushrooms, rust fungi, and smut fungi. Representative fungi of the Chytridactyta phylum include genera such as *Allomyces*, *Blastocladiella*, and *Coelomomyces*, as well as aquatic fungi. Representative fungi of the Oomycetes phylum include aquatic fungi (water molds) such as *Achlya*. Examples of Deuteromycetes include genera such as *Aspergillus*, *Penicillium*, *Candida*, and *Alternaria*. Representative fungi of the Zygomycetes phylum include genera such as *Rhizopus* and *Mucor*.

[0114] ■ Some preferred fungi include strains belonging to the Deuteromycetes and Hyphomycetes, such as *Fusarium*, *Humicola*, *Tricoderma*, *Myrothecium*, *Verticillum*, *Arthromyces*, *Caldariomyces*, *Ulocladium*, *Embellisia*, *Cladosporium* or *Dreschlera*, especially *Fusarium oxysporum* (DSM2672), *Humicola insolens*, *Trichoderma resii*, *Myrothecium verrucana* (IFO 6113), *Verticillum alboatrum*, *Verticillum dahlie*, *Arthromyces ramosus* (FERM P-7754), and *Caldariomyces*. fumago, Ulocladium chartarum, Embellisia alli, or Dreschlerahalodes.

[0115] ■ Other preferred fungi include strains belonging to the subphylum Basidiomycota, class Basidiomycetes, such as *Coprinus*, *Phanerochaete*, *Coriolus*, or *Trametes*, especially *Coprinus cinereus* f. microsporus (IFO 8371), *Coprinus macrohizus*, *Phanerochaete chrysosporium* (e.g. NA-12), or *Trametes* (formerly known as *Trametes*), such as *T. versicolor* (e.g. PR4 28-A).

[0116] ■ Further preferred fungi include strains belonging to the subphylum Zygomycotina and the family Mucoraceae, such as Rhizopus or Mucor, especially Mucor hiemalis.

[0117] yeast

[0118] ○Pichia

[0119] ○ Genus: Saccharomyces

[0120] ■ Fungal host cells can be yeast cells. As used herein, "yeast" includes ascospore-producing yeasts (Endomycetales), basidiospore-producing yeasts, and yeasts belonging to the class Fungi Imperfecti (Blastomycetes). Ascospore-producing yeasts are divided into the families Spermophthoraceae and Saccharomycetaceae. The latter consists of four subfamilies: Schizosaccharomycoideae (e.g., *Schizosaccharomyces*), Nadsonioideae, Lipomycoideae, and Saccharomycoideae (e.g., *Kluyveromyces*, *Pichia pastoris*, and *Saccharomyces*). Basidiospore-producing yeasts include genera such as *Leucosporidim*, *Rhodosporidium*, *Sporidiobolus*, *Filobasidium*, and *Filobasidiella*. Yeasts belonging to the deuteromycetes are divided into two families: Sporobolomyces (e.g., *Sporobolomyces* and *Bullera*) and Cryptococcides (e.g., *Candida*).

[0121] eukaryotes

[0122] ○Non-human animals, non-human mammals, birds, reptiles, insects, plants, yeast, fungi

[0123] In this document, the term "santalene synthase" refers to a polypeptide having catalytic activity in the formation of santalene and santalene-like terpenes such as α-santalene, β-santalene, trans-α-bergamotene, and epi-β-santalene from farnesyl pyrophosphate, and to other portions comprising such polypeptides. Examples of such other portions include complexes of the polypeptide with one or more other polypeptides, fusion proteins comprising a santalene synthase polypeptide fused to a peptide or protein tag sequence, other complexes of the polypeptide (e.g., metalloprotein complexes), macromolecular compounds comprising the polypeptide and another organic portion, and the polypeptide bound to a support material, etc. Santalene synthase can be provided in its natural environment, i.e., in the cells that produce the enzyme, or in a culture medium already secreted therein by the cells that produce the enzyme. It can also be provided separately from the source into which the polypeptide has been produced, and can be manipulated by conjugation to a vector, partial labeling with a label, etc.

[0124] The activity and product characteristics of santalene synthase can be measured using known methods, for example, as disclosed in International Patent Application WO2018160066.

[0125] In the following text, the terms “synthetic santalene synthase” and “improved santalene synthase” are used interchangeably to refer to the synthetic santalene synthase sequence that produces more β-santalene than α-santalene or increases the amount of α-santalene compared to wild-type santalene synthase under common conditions.

[0126] "Improved α-santalene synthase" therefore refers to those synthetic santalene synthases that, under common conditions, produce an increased yield of α-santalene compared to their naturally occurring counterparts. "Improved β-santalene synthase" refers to santalene synthase sequences that, under common conditions, produce β-santalene in addition to α-santalene.

[0127] The terms "in excess" and "surplus" are used interchangeably and should be understood as referring to the presence of more of a substance than the presence of a substance in the latter case. "A in excess of B" means that, based on the same basis—whether in moles, weights, or percentages—substance A is present in greater quantities than substance B.

[0128] In the conversion of farnesyl pyrophosphate to terpenoid products, pyrophosphate is cleaved to generate an active carbocation transition state, leading to a series of potential reactions such as hydride migration and cyclization. The residues that participate in supporting a particular potential transition state more than others can thus influence the final product ratio of the possible products.

[0129] It is known that the main products of santalene synthase are α-santalene, bergimene, and / or β-santalene.

[0130] Santalene synthase is a member of the terpene synthase family and is classified into enzyme classes EC4.2.3.81, EC4.2.3.82 and / or EC4.2.3.83, or EC4.2.3.50, based on the multiple products produced from the same substrate. The latter class of enzymes uses (2Z,6Z)-farnesyl pyrophosphate as a substrate instead of (2E,6E)-farnesyl pyrophosphate. They contain the N-terminal PFAM domain PF01397 and the C-terminal PFAM domain PF03936, which are active sites and metal-binding sites (PFAM analysis was performed using version 32.0; for details on PFAM, please refer to "The Pfam protein families database in 2019 (Pfam: Protein Families Database in 2019): S. El-Gebali, J. Mistry, A. Bateman, SREddy, A. Luciani, SCPotter, M. Qureshi, LJ Richardson, GASalazar, A. Smart, E.L. Sonnhammer, L. Hirsh, L. Paladin, D. Piovesan, SCETO Satto, R.D. Finn Nucleic Acids Research (2019) and http: / / pfam.xfam.org / ). They require divalent cations (usually magnesium or manganese) as cofactors. In their functional state, they generally have three Mg3+ domains coordinated to two aspartic-rich metal-binding sites. 2+ Ions. One of these sites is called a DDxxD motif, which is a sequence of two aspartic acids, followed by any amino acid, followed by another variable amino acid, preferably phenylalanine or tyrosine, more preferably tyrosine, and then another aspartic acid. The second metal-binding site is called an NSE / DTE triplet. This site is an amino acid sequence that begins with asparagine or aspartic acid, followed by a second aspartic acid, followed by two variable amino acids, followed by a serine or threonine, followed by one or two more variable amino acids, followed by a lysine or arginine, optionally followed by a variable amino acid and ending with an aspartic or glutamic acid residue. In these motifs, the variable amino acids are preferably those that allow the defining amino acid of the motif to take the tertiary structure required for metal ion binding (generally magnesium binding).

[0131] One of these conserved binding sites for coordinating magnesium ions, the DDxxD motif, is located in an α-helix. In the camphor (Cinnamomum camphora) santalene synthase called CiCaSSy (provided as SEQ ID NO:1), this α-helix extends from proline at position 278 of SEQ ID NO:1 or immediately thereafter to aspartic acid at position 302 and is named Helix D. In other santalene synthases, there are equivalent α-helices containing the DDxxD motif, although their nomenclature may differ, but this helix always directly affects the active site. In the following, any reference to Helix D refers to the α-helix of a given santalene synthase containing the DDxxD motif at positions 298 to 302 corresponding to amino acids in SEQ ID NO:1, regardless of whether the helix can be identified by the letter D or in a different manner in the corresponding protein sequence. Due to the high conservation of the DDxxD motif and other conserved residues and structural features, these helices are known in the art and can be readily identified in new sequences of santalene synthases.

[0132] The inventors realized that helical D is crucial to the product characteristics of santalene synthase; however, altering it could excessively perturb the enzyme structure in the sensitive region of the active site and / or jeopardize the magnesium ion binding essential for enzyme activity.

[0133] The inventors discovered that the product characteristics of the enzyme could be altered through more subtle changes. In santalene synthase, helix D is preceded by another α-helix. In CiCaSSy, this helix is ​​called helix C and extends from position 263 to position 272 in SEQ ID NO:1. Some predictions suggest that this α-helix extends to position 276; however, the core extends from position 263 to 272.

[0134] The helix at position 263 of SEQ ID NO:1 contains an arginine residue, which is part of the arginine-aspartic acid-arginine triplet present at positions 261 to 263 of SEQ ID NO:1. This triplet contains a conserved arginine residue from santalene synthase at its N-terminus.

[0135] Helix C and helix D interact on their facing sides. The specific relevant amino acid position of helix D is located in the region corresponding to position 291 of SEQ ID NO:1. Other positions where side chain interactions with the amino acid of helix C may occur are upstream of positions 287 and 288 in SEQ ID NO:1, 2, and 3 (isoleucine and threonine, respectively), and downstream of positions 294 and 295 in SEQ ID NO:1, 2, and 3 (methionine and threonine, respectively).

[0136] The inventors discovered that manipulating the C-helix provides greater flexibility to the enzyme, which affects the produced product without excessively perturbing the enzyme structure or the magnesium-binding or substrate-binding processes in a negative manner. They found from the primary structure that many santalene synthases appear to be adapted in principle to the desired changes and selected CiCaSSy (SEQ ID NO:1) to demonstrate the effects of the invention. CiCaSSy shares the C-helix element with santalene synthases that produce relatively high yields of β-santalene (although still less than α-santalene), which is also a characteristic of CiCaSSy's products for both santalenes. Examples of such known enzymes following CiCaSSy are SaSSy (SEQ ID NO:4), SaSSy14 (SEQ ID NO:5), SspiSSy (SEQ ID NO:6), SauSSy (SEQ ID NO:7), or SaSSy134 (SEQ ID NO:9). However, CiCaSSy also shares elements with santalene synthases such as ClaSSy (SEQ ID NO:8), which are low-potency β-santalene producers and high-potency α-santalene producers. CiCaSSy was chosen as the starting point for manipulating helical C to positively influence the flexibility of enzyme structures such as helical D and other downstream components, due to this intermediate position among these groups.

[0137] Following in-depth research, residue 267 of CiCaSSy was selected for mutation. This residue is expected to interact with the face-to-side of helix D via its side chain (see [link to relevant documentation]). Figure 3 Compared to other santalene synthases, CiCaSSy has some less common amino acids adjacent to this residue, which is expected to make it more adaptable, leading to changes in product characteristics. At the position corresponding to asparagine 267 (referred to as N267) in SEQ ID NO:1, many other santalene synthases have serine or leucine (see [link to SEQ ID NO:1]). Figure 1(Comparison). However, these santalene synthases with serine or leucine at this position have the aforementioned disadvantages in terms of their product characteristics, similar to the unmutated CiCaSSy of SEQ ID NO:1. However, the inventors realized that the surrounding environment of N267 in SEQ ID NO:1 was so favorable that they chose to replace the uncommon asparagine at position 267 with serine and leucine, even though these amino acids are present at the corresponding positions in other known poorly performing santalene synthases. The resulting synthetic protein sequences of the improved santalene synthases, referred to as N267S and N267L, are given in SEQ ID NO:2 and SEQ ID NO:3, respectively. Surprisingly, the reversion mutation at this position to a more common amino acid resulted in a change in the spatial flexibility of the enzyme's catalytic moiety (e.g., two adjacent α-helices) and a new, favorable change in product characteristics. In addition, this favorable change in product characteristics could also be achieved with other ingenious substitutions for the position corresponding to 267 in SEQ ID NO:1 (e.g., with glycine or alanine as shown below).

[0138] The DNA sequences encoding wild-type CiCaSSy, N267S, and N267L are listed as SEQ ID NO:10, 11, and 12, respectively.

[0139] Additional synthetic protein sequences carrying serine at position 267 of SEQ ID NO:1 are shown as SEQ ID NO:13 to 20, and additional improved protein sequences carrying leucine at position 267 of SEQ ID NO:1 are shown as SEQ ID NO:21 to 28.

[0140] In one embodiment, the present invention therefore relates to a synthetic β-santalene synthase that produces β-santalene in excess of α-santalene from farnesyl pyrophosphate under conditions that generally result in the production of both α-santalene and β-santalene, although known santalene synthases generally produce α-santalene in excess of β-santalene under such conditions. The synthetic β-santalene synthase of the present invention is characterized by the fact that, compared to the same tertiary structure in naturally occurring santalene synthases that produce excess α-santalene in excess of β-santalene, the tertiary structure corresponding to the α-helix extending from amino acid position 272 to position 291, preferably to position 284, of SEQ ID NO:1 has increased flexibility. Flexibility can be determined, for example, by performing a 500 ns root mean square fluctuation analysis using a simulation with the same conditions (pH 8.0, 300 K, 1 atm, aqueous environment, substrate-free conditions, and ion presence), and evaluating the structure of each enzyme based on the last 450 ns of the simulation. The calculations were performed using the gmx rmsf tool in GROMACS software version 2018 after structural superposition of protein structures for each trajectory frame was performed using gmx trjconv and the protein Cα of the equilibrated system was used as a reference.

[0141] In one embodiment, the polypeptide of the present invention is a synthetic polypeptide having the enzymatic function of β-santalene synthase and, compared with its naturally occurring counterpart, is intended to increase the flexibility of the helix D, preferably the flexibility of the tertiary structure, which corresponds to an α-helix extending from amino acid positions 272 to 291 in SEQ ID NO:1, and is further characterized in that β-santalene is produced from FPP in excess of α-santalene under conditions suitable for producing β-santalene.

[0142] In one aspect of the invention, the flexibility corresponding to the tertiary structure extending from amino acid position 272 to position 291, preferably to position 284 of SEQ ID NO:1 is increased compared to the same tertiary structure in naturally occurring santalene synthases that produce excess α-santalene beyond β-santalene. The flexibility is determined by: using a simulated root-mean-square fluctuation analysis for 500 ns under identical conditions (pH 8.0, 300 K, 1 atm, aqueous environment, substrate-free ion presence) and evaluating each enzyme structure according to the last 450 ns of the simulation. The calculation is performed using the gmx rmsf tool of GROMACS software version 2018 after structural superposition of protein structures for each trajectory frame using gmx trjconv and using the protein Cα of the equilibrated system as a reference. The flexibility is increased by at least 5%, preferably at least 10%, and more preferably at least 15% compared to the flexibility of the corresponding tertiary structure in naturally occurring santalene synthases that produce excess α-santalene beyond β-santalene.

[0143] In yet another embodiment, the position corresponding to position 267 of SEQ ID NO:1 is filled with serine, leucine, threonine, cysteine, isoleucine, valine, tryptophan, glycine, or alanine, preferably serine, threonine, tryptophan, glycine, alanine, or leucine, more preferably serine, glycine, alanine, or leucine. In another aspect of the invention, the synthesized santalene synthase further comprises two Mg-binding enzymes. 2+ It is rich in aspartic acid motifs, preferably DDxxD motifs and NSE / DTE triplet.

[0144] In one embodiment, the improved β-santalene synthase comprises an amino acid sequence starting from arginine (R261) at position 261 of SEQ ID NO:1 to two aspartic acid residues (D298 and D299) at positions 298 and 299 of SEQ ID NO:1, followed by two amino acids, preferably the second of which is tyrosine, and then a third aspartic acid (D302) at position 302 of SEQ ID NO:1, wherein these five amino acids preferably participate in the metal binding of the enzyme, and preferably, the position corresponding to position 267 of SEQ ID NO:1 is filled with serine, leucine, threonine, cysteine, isoleucine, valine, tryptophan, glycine or alanine, preferably serine, threonine, tryptophan, glycine, alanine or leucine, more preferably serine, glycine, alanine or leucine. In a preferred embodiment, the synthesized santalene synthase comprises a segment that further comprises a segment beginning with arginine corresponding to R261 of SEQ ID NO:1 and ending with aspartic acid corresponding to D302 of SEQ ID NO:1; and furthermore, having at least 50%, 60%, 65%, 70%, 75%, 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, or 97% sequence identity with respect to the full length of amino acids 261 to 302 of SEQ ID NO:2, 3, 13 to 53, preferably from those amino acids of SEQ ID NO:2, 3, 14 to 17, 21 to 52, in an increasing preferred order. More preferably, this segment contains... Figure 1 All the strongly conserved amino acids described in the text are present in this segment.

[0145] In one aspect of the invention, the improved santalene synthase of the invention and applicable to the methods of the invention and in host cells carries an R(R / K)xxxxxxxxW motif (arginine followed by an arginine or lysine, followed by eight amino acids of any type, followed by an arginine, see SEQ ID NO:55), preferably an RRxxxxxxxxW motif (RRX8W, see SEQ ID NO:54), near its N-terminal origin. In one embodiment, the RRX8W motif begins at position 7 in SEQ ID NO:2, 3, 29, 57 or 58 and ends at position 17 in SEQ ID NO:2, 3 or 29. In another embodiment, the RRX8W motif present in the improved santalene synthase of the present invention and used in the methods of the present invention and in host cells has the same amino acids as those in positions 7 to 17 of SEQ ID NO:2,3 or29: those amino acids in positions 7, 8 and 12 to 17 of SEQ ID NO:2,3 or29.

[0146] In yet another embodiment, the improved santalene synthase of the present invention and that usable in the methods of the present invention and in host cells contains an RRX8W motif near its N-terminal origin that is at least 80% or 90% identical to the RRX8W motif present as in SEQ ID NO:2, 3, or 29. In another aspect, such a motif in the improved santalene synthase of the present invention and that usable in the methods of the present invention and in host cells is identical to the RRX8W motif of SEQ ID NO:2, 3, or 29.

[0147] In one aspect of the invention, the improved santalene synthase of the present invention and that can be used in the methods of the present invention and in host cells comprises a PFAM domain PF01397 "Terpene_synth" and a C-terminal PFAM domain PF03936 "Terpene_synth_C".

[0148] In another aspect of the invention, the improved santalene synthase of the present invention and that can be used in the methods of the present invention and in host cells comprises the following features identified by InterPro software:

[0149] The following domains are identified: “terpene synthase, metal-binding domain” IPR005630, “terpene cyclase-like 1, C-terminal domain” IPR034741, and “terpene synthase, N-terminal domain” IPR001906.

[0150] And the homology superfamily "isoprene synthase domain superfamily" IPR008949,

[0151] “terpenoid cyclase / protein isoprenyltransferase α-α ring” IPR008930 and

[0152] “Terpene synthase, N-terminal domain superfamily” IPR036965.

[0153] As shown, only one or a few amino acid changes in the critical region of the helical C are essential to provide the desired effect of improved product characterization. Due to the short length of the helical C region, the rapid changes of one or a few amino acids result in relatively large differences in sequence identity between the two sequences used for the helical C region.

[0154] Another preferred embodiment relates to a synthetic santalene synthase that is so improved relative to the wild-type enzyme that it produces β-santalene from farnesyl pyrophosphate in addition to α-santalene, wherein the santalene synthase is aligned with the arginine residue corresponding to position 261 of SEQ ID No. 2 or 3 and the proline residue corresponding to position 278 of SEQ ID NO: 2, 3, 29, 57 or 58 to determine sequence identity, and the santalene synthase is of the full length relative to amino acid positions 261 to 278 of SEQ ID NO: 2, 3 to 29, preferably with SEQ ID NO: 2. Positions 261 to 272 of SEQ ID NO:2, 3 to 29 have at least 50%, 60%, 65%, 70%, 75%, 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity, and more preferably, the position corresponding to position 267 of SEQ ID NO:2, 3, 29, 57, or 58 is filled with serine, leucine, threonine, cysteine, isoleucine, valine, tryptophan, glycine, or alanine, preferably serine, threonine, tryptophan, glycine, alanine, or leucine, more preferably serine, glycine, alanine, or leucine, and is consistent with SEQ ID NO:2, 3, 29, 57, or 58. Position 291, corresponding to positions 2, 3, 29, 57, or 58, is filled with an amino acid other than histidine or leucine; preferably, this position is filled with any of the following amino acids: isoleucine, valine, serine, cysteine, phenylalanine, or threonine. In one aspect of the invention, the synthesized β-santalene synthase produces these santalenes at a ratio of β-santalene to α-santalene equal to or greater than 1, preferably at least 1.1, more preferably at least 1.2, and even more preferably 1.3, under conditions suitable for producing β-santalene and α-santalene.

[0155] Another aspect of the invention is a β-santalene synthase for the synthesis of β-santalene from farnesyl pyrophosphate in addition to α-santalene, wherein the santalene synthase has at least 50%, 60%, 65%, 70%, 75%, 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity with amino acid positions 261 to 302 of SEQ ID No. 2 or 3, wherein the position corresponding to position 261 of SEQ ID No. 2 or 3 is an arginine residue corresponding to the arginine at position 261 of SEQ ID No. 2 or 3, and three aspartic residues are present in the sequence corresponding to SEQ ID No. 2 or 3. The aspartic acid is present at positions 298, 299, and 302 of NO:2 or 3 or 29 to 40, and the synthesized β-santalene synthase produces these santalenes under conditions suitable for producing β-santalene and α-santalene at a ratio of β-santalene to α-santalene equal to or greater than 1, preferably at least 1.1, more preferably at least 1.2, and even more preferably 1.3.

[0156] In a preferred embodiment, the improved β-santalene synthase is filled with serine, leucine, threonine, cysteine, isoleucine, valine, tryptophan, glycine or alanine, preferably serine, glycine, alanine or threonine, and the position corresponding to position 282 of SEQ ID NO:1 is filled with an amino acid having a polar uncharged side chain or a positively charged side chain, preferably glutamine or asparagine or arginine or lysine.

[0157] In another preferred embodiment, the improved santalene synthase is filled with the following amino acids at position 267 of SEQ ID NO:1: serine, leucine, threonine, cysteine, isoleucine, valine, tryptophan, glycine, or alanine, preferably serine, threonine, tryptophan, glycine, alanine, or leucine, more preferably serine, glycine, alanine, or leucine, and it also has the following amino acids at the position corresponding to the position provided in parentheses after the amino acid name in SEQ ID NO:1: arginine (261), aspartic acid or asparagine (262), arginine or asparagine (263), leucine or isoleucine or valine or methionine (264), leucine or isoleucine or valine or methionine (265), glutamic acid or glutamine (266), histidine or tyrosine (268), and glutamine or arginine or lysine (282).

[0158] More preferably, these amino acids are the following amino acids at the positions corresponding to the positions provided in SEQ ID NO 1 in parentheses: arginine (261), aspartic acid (262), arginine (263), leucine (264), leucine (265), glutamic acid (266), histidine (268), leucine (269), phenylalanine (270), and glutamine or arginine (282).

[0159] In one aspect of the invention, in addition to the amino acids defined in the preceding paragraphs, the position 291 in the improved β-santalene synthase of the invention corresponding to position 291 of SEQ ID NO: 2, 3, 29, 57 or 58 is filled with an amino acid other than histidine or leucine, preferably, the position is filled with any of the following amino acids: isoleucine, valine, serine, cysteine, phenylalanine or threonine.

[0160] In yet another preferred embodiment, the improved santalene synthase additionally contains serine or threonine, preferably serine, at position 271 of SEQ ID NO:1, and alanine, isoleucine, valine, or cysteine, preferably alanine, at position 273 of SEQ ID NO:1.

[0161] More preferably, the improved santalene synthase is an enzyme that carries serine, leucine, threonine, cysteine, isoleucine, valine, tryptophan, glycine or alanine, preferably serine, threonine, glycine, alanine or leucine, more preferably serine, glycine, alanine or leucine, at position 267 of SEQ ID NO:1, and the position corresponding to the position in SEQ ID NO:1 is filled with the amino acids listed in Tables A, B or C below for the corresponding position in SEQ ID NO:1.

[0162] Table A

[0163]

[0164]

[0165]

[0166] Table B_

[0167]

[0168]

[0169] Table C_

[0170]

[0171]

[0172]

[0173] The aspartic acid at position 298 of SEQ ID NO:1 marks the start of the DDXXD motif in SEQ ID NO:1.

[0174] In another preferred embodiment, the improved santalene synthase comprises histidine at position 268 of SEQ ID NO:1, leucine at position 269 of SEQ ID NO:1, and phenylalanine at position 270 of SEQ ID NO:1, and preferably, the position corresponding to position 267 of SEQ ID NO:1 is filled with serine, leucine, threonine, cysteine, isoleucine, valine, tryptophan, glycine, or alanine, preferably serine, threonine, tryptophan, glycine, alanine, or leucine, more preferably serine or leucine. More preferably, the improved santalene synthase also contains amino acids listed in Table A, B, or C at positions corresponding to the positions listed in SEQ ID NO:1 in Table A, B, or C.

[0175] In a preferred embodiment, the improved β-santalene synthase has another amino acid other than histidine, glycine, or leucine at the position corresponding to position 291 of SEQ ID NO:1.

[0176] The inventors applied a further approach to increase the flexibility around helical C and helical D. Position 291 in SEQ ID NO:1, 2, 3, 29, 57, or 58 is the position representing the portion of helical D facing helical C. In the wild-type CiCassy of SEQ ID NO:1, this position is filled with isoleucine. Surprisingly, the inventors found that replacing the isoleucine at position 291 of SEQ ID NO:1 with threonine, serine, valine, phenylalanine, or cysteine ​​has a positive effect on the β-santalene / α-santalene ratio while maintaining a higher level of α-santalene than in the N267S or N267L mutants. In another aspect of the invention, the β-santalene synthase synthesized at the position corresponding to position 291 of SEQ ID NO:1, containing threonine, serine, methionine, valine, phenylalanine, or cysteine, preferably threonine, serine, valine, phenylalanine, or cysteine, further comprises two Mg-binding enzymes. 2+ It is rich in aspartic acid motifs, preferably DDxxD motifs and NSE / DTE triplet.

[0177] Furthermore, the inventors have produced a synthetic santalene synthase sequence in which the amino acid at position 291 of SEQ ID NO:1 is replaced with leucine, and the α-santalene yield is increased compared to the yield of SEQ ID NO:1.

[0178] Another aspect of the invention relates to a synthetic santalene synthase having a favorable mutation at positions corresponding to SEQ ID NO:1 267 and / or 291, wherein the santalene synthase comprises Mg-binding enzyme. 2+ The aspartic acid-rich motif DDxxD, with tyrosine or phenylalanine at the fourth position, more preferably, the binding motif has an N-terminus starting from two aspartic, phenylalanine, and tyrosine residues followed by another aspartic acid residue.

[0179] Further, the preferred amino acid substitution at position 291 of SEQ ID NO:1, in a preferred embodiment, the improved santalene synthase also has isoleucine or leucine at position 287 of SEQ ID NO:1, preferably isoleucine, and threonine, serine, or valine at position 288, preferably threonine or serine, more preferably threonine. Another preferred aspect of the invention relates to an improved santalene synthase having alanine at position 286 of SEQ ID NO:1, isoleucine at position 287 of SEQ ID NO:1, threonine at position 288 of SEQ ID NO:1, lysine at position 289 of SEQ ID NO:1, and alanine at position 290 of SEQ ID NO:1.

[0180] In addition to the preferred amino acid substitution of isoleucine at position 291 of SEQ ID NO:1, in a preferred embodiment, the improved santalene synthase also has a methionine or leucine or glutamic acid residue at position 294 of SEQ ID NO:1, preferably a methionine or glutamic acid residue, more preferably a methionine.

[0181] One aspect of the present invention relates to a β-santalene synthase for the synthesis of β-santalene from farnesyl pyrophosphate in addition to α-santalene, wherein the santalene synthase has an amino acid sequence that is at least 50% identical to that of SEQ ID NO:1; and at positions corresponding to the following amino acids:

[0182] a. Position 267 of SEQ ID NO:1 contains any of the following amino acids.

[0183] Serine, leucine, threonine, cysteine, isoleucine, valine, tryptophan, glycine, or alanine; or

[0184] b. Position 291 of SEQ ID NO: 1 contains any of the following amino acids.

[0185] Threonine, cysteine, serine, phenylalanine, or valine; or

[0186] c. A combination of a. and b. above; or

[0187] d. Asparagine is present at position 267 of SEQ ID NO:1, and the position corresponding to position 291 of SEQ ID NO:1 is any of the following amino acids:

[0188] Threonine, cysteine, serine, phenylalanine, or valine; or

[0189] e. Position 267 of SEQ ID NO:1 contains any of the following amino acids.

[0190] The amino acid is serine, leucine, threonine, cysteine, isoleucine, valine, tryptophan, glycine, or alanine, and the position corresponding to position 291 of SEQ ID NO:1 is isoleucine.

[0191] In another aspect, the present invention therefore relates to a β-santalene synthase for the synthesis of β-santalene from farnesyl pyrophosphate in addition to α-santalene, wherein the santalene synthase has an amino acid sequence that is at least 60% identical to that of SEQ ID NO:1; and has at the amino acid positions corresponding to: a) position 267 of SEQ ID NO:1 having any of the following amino acids: serine, leucine, threonine, cysteine, isoleucine, valine, or alanine, preferably serine or threonine; and / or b) position 291 of SEQ ID NO:1 having any of the following amino acids: isoleucine, serine, cysteine, valine, phenylalanine, or threonine, preferably threonine, phenylalanine, or valine; or when the position corresponding to position 267 of SEQ ID NO:1 is asparagine, and is identical to that of SEQ ID NO:1. The position corresponding to position 291 of SEQ ID NO:1 has any of the following amino acids: serine, cysteine, valine, phenylalanine, or threonine, preferably threonine, phenylalanine, or valine; in another aspect of the invention, in addition to the features described in the preceding sentence, the synthesized β-santalene synthase also fills the position corresponding to position 285 of SEQ ID NO:1 with valine, fills the position corresponding to position 282 of SEQ ID NO:1 with glutamine or arginine, fills the position corresponding to position 271 of SEQ ID NO:1 with serine, fills the position corresponding to position 273 of SEQ ID NO:1 with alanine, and / or fills the position corresponding to position 274 of SEQ ID NO:1 with valine.

[0192] In addition, the improved α-sandrolone synthase is an enzyme that carries isoleucine, valine, methionine, cysteine, serine, phenylalanine, or threonine at the position corresponding to position 291 of SEQ ID NO:1, preferably valine, cysteine, serine, phenylalanine, or threonine, more preferably cysteine, threonine, or valine, or alternatively leucine for the improved α-sandrolone synthase, and furthermore, the position corresponding to the position in SEQ ID NO:1 is filled with the amino acids listed in Table A', B', or C' for the corresponding position in SEQ ID NO:1.

[0193] Table A'

[0194]

[0195]

[0196]

[0197] Table B'

[0198]

[0199]

[0200] Table C'

[0201]

[0202]

[0203] In one aspect of the invention, the improved santalene synthase has an amino acid present at position 267 of a polypeptide of any of SEQ ID No: 2, 3, or 29 at position 267 of SEQ ID NO: 1 and an amino acid present at position 291 of a polypeptide of any of SEQ ID No: 30, 31, 32, 33, or 34 at position 291 of SEQ ID NO: 1 to improve β-santalene synthase, or an amino acid present at position 291 of a polypeptide of SEQ ID NO: 53 to improve α-santalene synthase, and has at least 50%, 60%, 65%, 70%, 75%, 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity relative to the full length of any of the polypeptides of SEQ ID NO: 2, 3, 29 to 40, or 53.

[0204] In another aspect of the invention, the improved santalene synthase has the following amino acid residues listed in Table D at the position corresponding to the position in SEQ ID NO:1 provided in Table D, and preferably, the position corresponding to position 267 of SEQ ID NO:1 is filled with serine, threonine, tryptophan, glycine, alanine or leucine, preferably serine, threonine or leucine, and more preferably, the position corresponding to position 255 of SEQ ID NO:1 is filled with an amino acid having a hydrophobic side chain or a polar uncharged side chain, preferably serine, threonine, alanine or valine, more preferably alanine.

[0205] Table D

[0206]

[0207]

[0208] In another preferred aspect of the invention, in addition to the advantageous amino acids at positions 267 and 291 of SEQ ID NO:1, the improved santalene synthase also has the following amino acids at positions corresponding to each of the amino acids listed herein, provided in parentheses, at the positions of SEQ ID NO:1: serine (271), alanine (273), valine (274), glutamine (282), valine (285), alanine (286), valine (292), methionine (294), alanine (296), and phenylalanine (300).

[0209] In another preferred aspect of the invention, in addition to the advantageous amino acids at the positions listed above, the improved santalene synthase also has arginine at the position corresponding to position 232 in SEQ ID NO:1.

[0210] Table 1 shows the β-santalene / α-santalene ratio in some improved santalene synthases and controls:

[0211]

[0212] Name: The tested santalene synthase;

[0213] Ratio b / a: β-santalene / α-santalene ratio w% / w%

[0214] Ingenious modifications result in an increased β-santalene / α-santalene ratio or an increase in α-santalene in the product of the improved santalene synthase, as shown in Table 1. Italicized entries refer to the unmodified enzyme of SEQ ID NO:1 (“wild-type”) and I291L for producing excess α-santalene. The latter demonstrates that ingenious modification at a given location will result in the desired β-santalene / α-santalene ratio or improved α-santalene yield, as the I291L modification allows for the production of a greater amount of α-santalene than the unmodified enzyme of SEQ ID NO:1, as indicated by the lower β-santalene / α-santalene ratio of I291L.

[0215] Increasing or decreasing the production of β-santalene requires the amino acid selection of the present invention at positions 267 and / or 291. For example, the inventors replaced isoleucine at position 291 of SEQ ID NO:1 (see SEQ ID NO:53) with leucine to increase the yield of α-santalene beyond SEQ ID NO:1, but at the cost of no improvement in the yield of β-santalene and bergimene, but rather a decrease. One aspect of the present invention is therefore an α-santalene synthase having leucine at position 291 of SEQ ID NO:1, which has an improved yield of α-santalene compared to the unmodified enzyme.

[0216] When histidine is introduced at position 291 of SEQ ID NO:1, the activity of santalene synthase is destroyed and α-santalene, β-santalene, and bergimene are not produced. In one aspect of the invention, the improved santalene synthase of the present invention has an amino acid at position 291 other than histidine.

[0217] In one aspect, the improved santalene synthase of the present invention lacks a histidine or glycine residue at position 291 of SEQ ID NO:1, but contains isoleucine, valine, threonine, cysteine, phenylalanine, or serine, preferably cysteine, valine, serine, phenylalanine, or threonine, or leucine if an increased α-santalene / β-santalene ratio is desired. In another aspect, when position 267 of SEQ ID NO:1 is filled with serine, threonine, or leucine, isoleucine is present at position 291 of SEQ ID NO:1, or when position 267 of SEQ ID NO:1 is filled with asparagine, valine, serine, phenylalanine, or threonine is present at position 291.

[0218] In another preferred embodiment, the improved β-santalene synthase comprises arginine (261), leucine (264), leucine (265), serine (271), alanine (273), proline (278), arginine (284), isoleucine (287), aspartic acid (298), aspartic acid (299), and aspartic acid (302) at the position corresponding to position 267 of SEQ ID NO:1, preferably serine, leucine, threonine, cysteine, isoleucine, valine, tryptophan, glycine, or alanine, preferably serine, threonine, tryptophan, glycine, alanine, or leucine, more preferably serine or leucine. For the improved β-santalene synthase, in yet another embodiment, this position is filled with asparagine and the position corresponding to position 291 of SEQ ID NO:1 is filled with valine, cysteine, serine, phenylalanine, or threonine.

[0219] In yet another preferred embodiment, the improved santalene synthase additionally has at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, more preferably at least 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99%, and even more preferably 100% in Figure 1 The amino acids are marked with a black background to indicate that they are highly conserved.

[0220] In another preferred embodiment, the improved santalene synthase comprises the sequence of SEQ ID NO:1, its variants, derivatives, orthologs, paralogs, or homologs, wherein the amino acid at position 267 is replaced with leucine, serine, or threonine, and the amino acid at position 291 is replaced with threonine, serine, cysteine, phenylalanine, or valine, or, if desired, with leucine.

[0221] In another preferred embodiment, the improved β-santalene synthase comprises the sequence of SEQ ID NO:1, its variants, derivatives, orthologs, paralogs, or homologs, wherein the amino acid at position 267 is replaced with Leu and the amino acid at position 291 is replaced with Thr.

[0222] In another preferred embodiment, the improved β-santalene synthase comprises the sequence of SEQ ID NO:1, its variants, derivatives, orthologs, paralogs, or homologs, wherein the amino acid at position 267 is replaced with Leu and the amino acid at position 291 is replaced with Ser.

[0223] In another preferred embodiment, the improved β-santalene synthase comprises the sequence of SEQ ID NO:1, its variants, derivatives, orthologs, paralogs, or homologs, wherein the amino acid at position 267 is replaced with Leu, and the amino acid at position 291 is replaced with Cys or Phe.

[0224] In another preferred embodiment, the improved β-santalene synthase comprises the sequence of SEQ ID NO:1, its variants, derivatives, orthologs, paralogs, or homologs, wherein the amino acid at position 267 is replaced with Leu and the amino acid at position 291 is replaced with Val.

[0225] In another preferred embodiment, the improved β-santalene synthase comprises the sequence of SEQ ID NO:1, its variants, derivatives, orthologs, paralogs, or homologs, wherein the amino acid at position 267 is replaced with Ser and the amino acid at position 291 is replaced with Thr.

[0226] In another preferred embodiment, the improved β-santalene synthase comprises the sequence of SEQ ID NO:1, its variants, derivatives, orthologs, paralogs, or homologs, wherein the amino acid at position 267 is replaced with Ser, and the amino acid at position 291 is replaced with Ser.

[0227] In another preferred embodiment, the improved β-santalene synthase comprises the sequence of SEQ ID NO:1, its variants, derivatives, orthologs, paralogs, or homologs, wherein the amino acid at position 267 is replaced with Ser, and the amino acid at position 291 is replaced with Cys or Phe.

[0228] In another preferred embodiment, the improved β-santalene synthase comprises the sequence of SEQ ID NO:1, its variants, derivatives, orthologs, paralogs, or homologs, wherein the amino acid at position 267 is replaced with Ser and the amino acid at position 291 is replaced with Val.

[0229] Improved santalene synthases generally have a molecular weight between 60 and 70 kDa, preferably between 61 and 66 kDa, without any tags, added domains, or fusion with other protein moieties.

[0230] In a preferred embodiment, the improved santalene synthase has at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, such as at least 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% and, for example, 100% sequence identity relative to the full length of SEQ ID NO:1. In yet another preferred embodiment, the improved santalene synthase has at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, for example at least 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% and for example 100% sequence identity with respect to any of SEQ ID NO: 2, 3, 14 to 17, 21 to 52, preferably any of SEQ ID NO: 2, 3, 29 to 40, of the full-length protein sequence, for the improved β-santalene synthase; or if it is desired to increase α-santalene production, relative to SEQ ID NO: 13, 18, 19, 20, or 53, preferably SEQ ID NO: 2, 3, 29 to 40, of the full-length protein sequence, and is used for the improved β-santalene synthase; or if it is desired to increase α-santalene production, relative to SEQ ID NO: 13, 18, 19, 20, or 53, preferably SEQ ID NO: 2, 3, 14 to 17, 21 to 52, of the full-length protein sequence, preferably SEQ ID NO: 2, 3, 29 to 40, of the full-length protein sequence, is used for the improved β-santalene synthase. The full-length protein sequence of IDNO:53 has at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, for example at least 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99%, and for example 100% sequence; and more preferably additionally has Figure 1 The amino acids are marked with a black background to indicate that they are highly conserved.

[0231] In the santalene synthase with an increased β-santalene / α-santalene ratio, preferably a) the position corresponding to position 267 of SEQ ID NO:1 is filled with serine, leucine, threonine, cysteine, isoleucine, valine, tryptophan, glycine, or alanine, preferably serine, threonine, tryptophan, glycine, alanine, or leucine, more preferably serine or leucine; or the position corresponding to position 291 of SEQ ID NO:1 is filled with valine, threonine, cysteine, phenylalanine, or serine, more preferably Thr, Val, Cys, or Ser; or b) the position corresponding to position 267 of SEQ ID NO:1 is asparagine and the position corresponding to position 291 of SEQ ID NO:1 is filled with valine, threonine, cysteine, phenylalanine, or serine, more preferably Thr, Val, Cys, or Ser; or c) the position corresponding to SEQ ID NO:1 is filled with asparagine and the position corresponding to position 291 of SEQ ID NO:1 is filled with valine, threonine, cysteine, phenylalanine, or serine, more preferably Thr, Val, Cys, or Ser; or c) the position corresponding to SEQ ID NO:1 is filled with serine, threonine, cysteine, phenylalanine, or serine, more preferably Thr, Val, Cys, or Ser; The position corresponding to position 267 in NO:1 is filled with serine, leucine, threonine, cysteine, isoleucine, valine, tryptophan, glycine, or alanine, preferably serine, threonine, tryptophan, glycine, alanine, or leucine, more preferably serine, glycine, alanine, or leucine, and the position corresponding to position 291 in SEQ ID NO:1 is filled with isoleucine, valine, threonine, or methionine; or d)a), b), or c) a combination of alanine residues at the position corresponding to position 255 in SEQ ID NO:1; or e)a), b), or c) a combination of histidine at the position corresponding to position 268 in SEQ ID NO:1.

[0232] In the santalin synthase with an increased α-santalene / β-santalene ratio, the position corresponding to position 291 in SEQ ID NO:1 is filled with leucine, and the position corresponding to position 267 in SEQ ID NO:1 is asparagine, serine, threonine, or leucine, preferably asparagine.

[0233] One aspect of the present invention relates to a santalene synthase for the synthesis of α-santalene from farnesyl pyrophosphate in addition to β-santalene, wherein the santalene synthase is aligned with two protein sequences to determine sequence identity using an arginine residue corresponding to arginine at position 261 of SEQ ID No: 1, 2, 3, 29, 57 or 58 and three aspartic acid residues corresponding to aspartic acid at positions 298, 299 and 302 of SEQ ID No: 1, 2, 3, 29, 57 or 58, when the two protein sequences are aligned. Amino acid positions 261 to 302 of any one of NO:1, 2, 3, 29, 57, or 58 have at least 50%, 60%, 65%, 70%, 75%, 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity, and wherein the position corresponding to position 291 in SEQ ID NO:2 or 3 is glycine or leucine, preferably leucine. In one aspect, these improved α-santalene synthases contain asparagine at the position corresponding to position 267 in SEQ ID NO:1.

[0234] Preferably, the improved β-santalene synthase of the present invention has at least 50%, preferably at least 60%, at least 70%, at least 80%, at least 90%, or at least 95% sequence identity in the protein moieties of any of SEQ ID NO:2, 3, 14 to 17, 21 to 52, 57, or 58, preferably any of SEQ ID NO:2, 3, 29 to 40, 57, or 58, wherein the protein moieties begin with an arginine residue at position 261 of SEQ ID NO:2, 3, 29, 57, or 58 and extend to three aspartic acid residues at positions 298, 299, and 302 of SEQ ID NO:2, 3, 29 to 40, 57, or 58, and are in sequence with the protein moieties of SEQ ID NO:2, 3, 29 to 40, 57, or 58. The position corresponding to position 267 of SEQ ID NO:2,3,29,57 or58 contains serine, leucine, threonine, cysteine, isoleucine, valine, tryptophan, glycine or alanine, preferably serine, threonine, tryptophan, glycine, alanine or leucine, and / or the position corresponding to position 291 of SEQ ID NO:2,3,29,57 or58 contains valine, cysteine, serine, phenylalanine or threonine, preferably valine, serine, phenylalanine or threonine, or if serine, threonine, tryptophan, glycine, alanine or leucine is present at the position corresponding to position 267 of SEQ ID NO:2,3,29,57 or58, isoleucine is present at the position corresponding to position 291 of SEQ ID NO:2,3 or29.

[0235] In another preferred aspect, the improved β-santalene synthase has at least 50%, preferably at least 60%, at least 70% or 80% sequence identity with any of SEQ ID NO:2, 3, 14 to 17, 21 to 52, 57 or 58, preferably any of SEQ ID NO:2, 3 or 29 to 40, 57 or 58, in the following protein moieties, said protein moieties beginning with an arginine at position 261 of SEQ ID NO:2, 3, 29, 57 or 58 and extending to three aspartic acid residues at positions 298, 299 and 302 of SEQ ID NO:2, 3, 29, 57 or 58, and in sequence with SEQ ID NO:2, 3, 29, 57 or 58. Position 267 in SEQ ID NO: 2, 3, 29, 57 or 58 contains asparagine, serine, leucine, threonine, cysteine, isoleucine, valine, tryptophan, glycine or alanine, preferably asparagine, serine, threonine, tryptophan, glycine, alanine or leucine, and / or position 291 in SEQ ID NO: 2, 3, 29, 57 or 58 contains valine, serine, cysteine, phenylalanine or threonine, preferably serine, valine, phenylalanine or threonine, and preferably histidine in position 268 in SEQ ID NO: 2, 3, 29, 57 or 58.

[0236] The amounts of β-santalene and α-santalene are determined by reliable quantitative methods, preferably by gas chromatography with an FID detector. Preferred methods for determining the amounts of α-santalene, β-santalene, and bergamotene are detailed in the Examples section.

[0237] The improved β-santalene synthase produces β-santalene more than α-santalene, meaning that under conditions suitable for the production of these santalenes, the enzyme produces both β-santalene and α-santalene at a molar ratio greater than 1.0. The improved α-santalene synthase produces α-santalene more than β-santalene, meaning that under conditions suitable for the production of these santalenes, the enzyme produces both β-santalene and α-santalene at a molar ratio less than 1.0.

[0238] Suitable conditions for the production of these santalenes can be provided, for example, by expressing DNA encoding an improved santalene synthase in a host cell, which provides an active improved santalene synthase and all substrates and cofactors such as farnesyl pyrophosphate and magnesium ions, so that the improved enzyme can carry out the reaction to produce α-santalene and β-santalene.

[0239] To date, known santalene synthases produce compositions in which the molar ratio of β-santalene / α-santalene is less than 1. The improved β-santalene synthase of the present invention produces β-santalene and α-santalene at a β-santalene / α-santalene molar ratio equal to or greater than 1, preferably measured by GC-FID; for example, this ratio is at least 1.05, 1.1, 1.2, 1.3, 1.4, 1.5, 1.6, 1.7, 1.8, 1.9, or at least 2. The β-santalene / α-santalene ratio can be at least 3:1, preferably at least 4:1, more preferably at least 5:1, even more preferably 6:1, however, even more preferably at least 7:1, most preferably at least 8:1, and even at least 9:1. In one aspect of the invention, this ratio is not greater than 100:1.

[0240] One aspect of the invention relates to a synthetic nucleic acid encoding any of the synthetic santalene synthases of the invention, said synthetic santalene synthase being a santalene synthase with increased β-santalene / α-santalene production (e.g., but not limited to, polypeptides or variants thereof of SEQ ID NO: 2, 3, 14 to 17, 21 to 52), or those santalene synthases with improved α-santalene production compared to unmodified natural santalene synthases, such as, but not limited to, the polypeptide or variant thereof of SEQ ID NO: 53. Another part of the invention is an expression cassette comprising the synthetic nucleic acid of the invention.

[0241] Another preferred embodiment is a method for producing compositions of β-santalene in addition to α-santalene using the improved β-santalene synthase disclosed herein, preferably a method suitable for large-scale production, the method comprising the steps of: i) providing one or more improved β-santalene synthases in active form and having all the necessary cofactors (e.g., but not limited to metal ions such as magnesium ions); ii) contacting farnesyl pyrophosphate with one or more improved β-santalene synthases under conditions allowing for the production of santalene; iii) producing β-santalene and α-santalene, optionally bergimene, and optionally other santalenes from farnesyl pyrophosphate, wherein the amount of β-santalene produced is greater than the amount of α-santalene produced, and optionally purifying the products, for example, to separate them from the santalene synthase and any remaining substrates and undesirable compounds. Preferably, these methods produce compositions containing more β-santalene than α-santalene, wherein the molar ratio of β-santalene to α-santalene is at least 1.1, 1.2, 1.3, 1.4, 1.5, 1.6, 1.7, 1.8, 1.9, or at least 2; the β-santalene / α-santalene ratio can be at least 3:1, preferably at least 4:1, more preferably at least 5:1, even more preferably 6:1, however, even more preferably at least 7:1, most preferably at least 8:1, and even at least 9:1. In one aspect of the invention, this ratio is not greater than 100:1.

[0242] A method for producing compositions containing β-santalene in greater quantities than α-santalene would be particularly advantageous, the method comprising a fermentation step providing an improved β-santalene synthase and contacting the enzyme with farnesyl pyrophosphate to produce santalene. For example, although it is possible to use the santalene synthase isolated according to the present invention in vitro, a fermentation method for producing compositions containing β-santalene in greater quantities than α-santalene is particularly preferred, the method comprising the following steps:

[0243] a) Provide nucleic acids encoding improved β-santalene synthase in a manner suitable for expression in the host.

[0244] b) Introduce the nucleic acid from a) into a host cell that can provide farnesyl pyrophosphate and all necessary cofactors to the tannin synthase to be active.

[0245] c) Culture host cells in a manner that enables the host cells to produce santalene synthase encoded by nucleic acid in an active form and provides farnesyl pyrophosphate and all necessary cofactors to the santalene synthase.

[0246] d) Using an improved β-santalene synthase, β-santalene and α-santalene, and optionally bergimene, are produced from farnesyl pyrophosphate, wherein the amount of β-santalene produced is greater than the amount of α-santalene produced.

[0247] e) When the desired amounts of these compounds have been produced, harvest the resulting β-santalene and α-santalene, and optionally bergamotene.

[0248] f) Optionally purify β-santalene and α-santalene and optionally bergamerene.

[0249] Compared to those produced under the same conditions using unmodified santalene synthase, the amount of β-santalene produced by the improved β-santalene synthase and the method of the present invention comprises, on a weight / weight basis, at least 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, or 95% more β-santalene in an increasing preferred order. Optionally, compared to those produced under the same conditions using unmodified santalene synthase, the amount of bergamotene produced by the improved β-santalene synthase and the method of the present invention comprises, on a weight / weight basis, at least 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, or 95% more bergamotene in an increasing preferred order. In one aspect of the invention, at least 12% (w / w), 18% (w / w), or 20% (w / w) of bergamotene is produced using the improved santalene synthase and method of the present invention. Even more preferably, at least twice the amount of β-santalene and optionally bergamerene are present in the resulting composition.

[0250] In one aspect of the invention, the invention further relates to a santalene composition produced by means of an improved β-santalene synthase, having a higher β-santalene content than α-santalene. In another aspect of the invention, the composition of the invention is produced by one or more synthetic β-santalene synthases, methods, or host cells of the invention, wherein the composition comprises an excess of β-santalene relative to α-santalene.

[0251] A preferred embodiment is a composition comprising, preferably substantially composed of, β-santalene and α-santalene produced by an improved β-santalene synthase, and bergamotene, wherein the composition has more β-santalene than α-santalene. In a specific aspect of the invention, the composition comprises more β-santalene than bergamotene and more bergamotene than α-santalene.

[0252] The compositions of the present invention preferably contain more β-santalene than α-santalene in a ratio greater than that of β-santalene to α-santalene, for example, a ratio of at least 1.1, 1.2, 1.3, 1.4, 1.5, 1.6, 1.7, 1.8, 1.9, or at least 2. The β-santalene / α-santalene ratio can be at least 3:1, preferably at least 4:1, more preferably at least 5:1, even more preferably 6:1, however, even more preferably at least 7:1, most preferably at least 8:1, and even at least 9:1. In one aspect of the invention, this ratio is not greater than 1000:1.

[0253] The present invention also relates to compositions produced by means of an improved β-santalene synthase, having a higher bergimene content than α-santalene. Such compositions contain more bergimene than α-santalene in a ratio greater than that of bergimene to α-santalene, preferably at least 1.1, 1.2, 1.3, 1.4, 1.5, 1.6, 1.7, 1.8, 1.9, or at least 2. The bergimene / α-santalene ratio can be at least 3:1, preferably at least 4:1, more preferably at least 5:1, even more preferably 6:1, however, even more preferably at least 7:1, most preferably at least 8:1, and even at least 9:1. In one aspect of the invention, this ratio is not greater than 1000:1.

[0254] In one aspect of the invention, the composition produced by means of an improved santalene synthase contains at least 12% (w / w), 18% (w / w), or 20% (w / w) bergamotene.

[0255] The ratio of bergamotene / β-santalene produced by the improved santalene synthase and present in the compositions of the present invention can be higher than 1 (bergamotene excess) or lower than 1 (β-santalene excess). The former is exemplified by compositions produced, for example, by N267S (SEQ ID NO:2) or α-santalene excess producer I291L (SEQ ID NO:53), and the latter by compositions produced by N267L (Seq ID NO:3) or any of SEQ ID NO:30 to 34 or 36 or 37, as can be seen in… Figure 4 and Figure 5 Depending on the desired product and further processing of the composition, it is advantageous to use an improved β-santalene synthase of either the bergimene-excess type or the β-santalene-excess type. The improved β-santalene synthases N267G (SEQ ID NO:57) and N267A (SEQ ID NO:58) exhibit bergimene levels almost at the β-santalene level, while α-santalene yield is significantly reduced, which may also be desirable for some applications.

[0256] In one aspect of the invention, the ratio of bergamotene to β-santalene is less than 1.0, for example equal to or less than 0.95, 0.9, 0.85, 0.8 or 0.75, for example equal to or less than 0.70, but greater than 0.28, for example greater than 0.30.

[0257] In one embodiment, the bergimene / β-santalene ratio in the composition produced by the improved β-santalene synthase is at least 1:1. In one aspect of the invention, this ratio is not higher than 5.5:1, for example, not higher than 5:1 or 4.5 to 1, or 4:1, or 3.5 to 1 or 3 to 1, or 2.5 to 1, or 2:1.

[0258] In another embodiment, the ratio of bergamotene to β-santalene in the composition produced by means of improved β-santalene synthase is 1:2, 1:3, 1:4, 1:5, or 1:10 or lower.

[0259] Therefore, in one aspect, the present invention relates to an improved β-santalene synthase, the host cell of the present invention, or the method of the present invention, wherein the santalene synthase, in addition to producing more β-santalene than α-santalene, also produces an excess of trans-α-bergamerene exceeding α-santalene.

[0260] Another embodiment relates to a composition that can be produced by an improved β-santalene synthase, the host cell of the present invention, or the method of the present invention, comprising more bergamotene than β-santalene and more β-santalene than α-santalene. Preferably, producing the composition includes a fermentation step of producing the improved β-santalene synthase or producing the composition.

[0261] In a preferred embodiment, a composition in which β-santalene is more abundant than α-santalene is obtained by culturing one or more types of host cells, preferably bacterial cells, plant cells or fungal cells (including yeast cells), more preferably bacteria, and even more preferably Escherichia coli, Amycolatopsis sp. or Rhodococcus.

[0262] In yet another preferred embodiment, the present invention relates to β-santalol ((2Z)-2-methyl-5-[2-methyl-3-methylene-bicyclo[2.2.1]hept-2-yl]pent-2-en-1-ol; CAS No. 77-42-9) and α-santalol ((Z)-5-(2,3-dimethyltricyclo[2.2.1.0)) generated from a precursor composition. 2,6 Compositions of hept-3-yl)-2-methylpent-2-en-1-ol (CAS No. 115-71-9), wherein the precursor composition comprises β-santalene and α-santalene produced by the method of the present invention, wherein β-santalol (also referred to herein as beta-santalol) is present in an amount higher than α-santalol (also referred to herein as alpha-santalol) on a w / w basis, due to the excess β-santalene content in the precursor composition. In these compositions, the β-santalol / α-santalol ratio is greater than 1, preferably at least 1.1, 1.2, 1.3, 1.4, 1.5, 1.6, 1.7, 1.8, 1.9 or at least 2. The β-santalol / α-santalol ratio can be at least 3:1, preferably at least 4:1, more preferably at least 5:1, even more preferably 6:1, however, even more preferably at least 7:1, most preferably at least 8:1 and even at least 9:1. In one aspect of the invention, the ratio is no greater than 100:1.

[0263] Another preferred embodiment is a method for producing a composition in which β-santalol exceeds α-santalol without a) reducing the α-santalene content before conversion to α-santalol and / or b) increasing the β-santalol content after conversion from santalene by distillation or other means, wherein the method includes the step of producing a composition in which β-santalene exceeds α-santalene by the method of the present invention and oxidizing β-santalene to β-santalol and α-santalene to α-santalol in one or more subsequent steps. This santalene conversion can be carried out biosynthetically and / or chemically to its corresponding alcohol. After conversion to santalol, purification steps such as distillation to remove other compounds can be incorporated, and the β-santalol / α-santalol ratio can be altered by distillation if desired; however, by utilizing an improved β-santalene synthase to provide a composition in which β-santalene exceeds α-santalene, a composition with more β-santalol than α-santalol can be obtained without further altering the β-santalol / α-santalol ratio. One aspect of the invention is therefore a method for producing a composition containing an excess of β-santalol relative to α-santalol, wherein the method comprises the step of producing a composition in which β-santalene exceeds α-santalene by the method of the invention and, in one or more subsequent steps, oxidizing β-santalene to β-santalol and oxidizing α-santalene to α-santalol, wherein after the oxidation of santalene, santalol is distilled to purify santalol, while the content of β-santalol is not significantly increased relative to the content of α-santalol.

[0264] Furthermore, this invention relates to compositions comprising an excess of β-santalol relative to α-santalol, produced by any of the methods of this invention using the improved β-santalol synthase of this invention or using the host cells of this invention, optionally with a total bergamot content of less than 10% (w / w) or even less than 8% (w / w) of bergamot in the composition. In another aspect, compositions of this invention comprising an excess of β-santalol relative to α-santalol, produced by any of the methods of this invention using the improved β-santalol synthase of this invention or using the host cells of this invention, contain less than 3% epi-β-santalol.

[0265] One aspect of the invention relates to a santalene synthase that produces β-santalene in excess of α-santalene, a nucleic acid encoding such an enzyme, an expression cassette containing such a nucleic acid, a host cell containing such an expression cassette, a method of the invention, and compositions produced using the enzyme and method of the invention, said compositions comprising β-santalene and α-santalene and / or β-santalol and α-santalol at a β-santalene / α-santalene ratio or a β-santalol / α-santalol ratio of at least equal to or greater than 1.3, 1.5, or 2.

[0266] Preferably, the composition of the present invention is a lipophilic composition.

[0267] β-santalene, α-santalene, or bergamotene produced by the method or composition of the present invention can be used in fragrance or flavoring applications, in cosmetics, as insect repellents or insect attractants, or in agriculture, such as for crop protection or animal husbandry.

[0268] One aspect of the invention is a host cell suitable for generating the improved santalene synthase from one or more nucleic acids encoding one or more improved santalene synthases and suitable for providing the improved santalene synthase with farnesyl pyrophosphate and all cofactors essential for its activity, wherein the host cell contains such nucleic acid.

[0269] Another preferred embodiment is therefore a host cell comprising the improved santalene synthase of the present invention. The microorganisms capable of producing compositions containing more β-santalene than α-santalene can be fungal cells (including yeast), bacteria, plant cells, or animal cells, for example, belonging to the group consisting of the genera: Escherichia, Klebsiella, Helicobacter, Bacillus, Lactobacillus, Streptococcus, Amycolatopsis, Rhodobacter, Lactococcus, Pichia, Saccharomyces, and Kluyveromyces.In a preferred embodiment, one or more host cells suitable for producing compositions containing more β-santalene than α-santalene are bacterial cells selected from a) Gram-negative bacterial groups, such as *Rhodobacter* (e.g., *Rhodobacter sphaeroides*, *Rhodobacter capsulatus*), *Agrobacterium*, *Paracoccus* (e.g., *Paracoccus carotinifaciens*, *Paracoccus zeaxanthinifaciens*), or *Escherichia*; b) Gram-positive bacterial cells, such as *Bacillus*, *Corynebacterium*, etc. c) Fungal cells selected from the following groups: Aspergillus, Blakeslea, Peniciliium, Phaffia (Xanthophyllomyces), Pichia, Saccharamoyces, Kluyveromyces, Yarrowia, and Hansenula; or d) Transgenic plants or cultures, including transgenic plant cells, wherein the cells belong to the following transgenic plant species: Nicotiana The microorganisms are selected from the genera *Schizophyllum*, *Cichorumintybus*, *Lacuca sativa*, *Mentha* species, *Artemisia annua*, tuberous plants, oil crops, and trees; e) or transgenic mushrooms or cultures, including transgenic mushroom cells, wherein the microorganisms are selected from the genera *Schizophyllum*, *Agaricus*, and *Pleurotisi*. More preferred organisms are microorganisms belonging to the genera *Escherichia*, *Saccharomyces*, *Pichia pastoris*, *Rhodotorula*, or *Paracococcus*, and even more preferably those belonging to the genera *Escherichia coli*, *Saccharomyces cerevisae*, *Rhodotorula spp.*, or *Rhodotorula* species.

[0270] Another embodiment is an expression cassette comprising a synthetic nucleic acid encoding an improved santalene synthase. These nucleic acids may be those listed as SEQ ID NO:11 or 12, or those encoding any of the polypeptides in SEQ ID NO:2, 3, 14 to 17, 21 to 52, or, for increased α-santalene production, the nucleic acid encoding the polypeptide of SEQ ID NO:53. Other nucleic acids in the expression cassette suitable for altering santalene production in host cells are those encoding improved santalene synthases, such as, but not limited to, those disclosed in any of SEQ ID NO:2, 3, 13 to 53. For nucleic acids encoding santalene synthases that increase α-santalene production, nucleic acids encoding the polypeptide of SEQ ID NO:53 can be used in such expression cassettes and host cells.

[0271] The expression cassette may be contained in a vector, nucleus, plasmid, artificial chromosome, or any other tool that allows expression in a host cell at the desired strength and manner.

[0272] Another aspect of the invention is a method for purposefully altering the product characteristics of santalene synthase by changing the flexibility of a tertiary structure corresponding to helix C of SEQ ID NO:1, helix D of SEQ ID NO:1, and the polypeptide chain connecting these two helices in SEQ ID NO:1. For example, this method involves the steps of: altering the nucleic acid encoding santalene synthase so that the amino acid at position 267 of SEQ ID NO:1 is serine, leucine, threonine, cysteine, isoleucine, valine, tryptophan, glycine, or alanine, preferably serine, threonine, tryptophan, glycine, alanine, or leucine, such as serine or leucine; and / or altering the codons of the nucleic acid encoding santalene synthase in such a way that the codon corresponding to position 291 of SEQ ID NO:1 now encodes leucine, valine, threonine, cysteine, or serine, such as Thr, Val, Cys, Phe, or Ser; followed by the step of expressing the modified nucleic acid in a host cell suitable for expressing the santalene synthase synthesized according to the present invention. Attached Figure Description

[0273] Figure 1This display shows alignments of known santalene synthases CiCaSSy wild-type (SEQ ID NO:1), SaSSY, SaSSy14, SspiSSy, SauSSy, ClaSSy, and SaSSy134 (SEQ ID NO:4 to 9, respectively). SaSSY134 is identified as SEQ 280 in this alignment. Two improved β-santalene synthase mutants, N267S and N267L (SEQ ID NO:2 and 3), are also included. This alignment was performed using the clustalw software with common settings. Strongly conserved residues are indicated by black background shading, residues conserved in at least 50% of the aligned sequences are indicated by gray background shading, and non-conserved amino acids are indicated by white background shading.

[0274] Figure 2 This displays a 3D model of CiCaSSy SEQ ID NO:1 generated using PyrMol software. The α-helix is ​​shown, and the two helices of CiCaSSy are highlighted in black: helix C (short black helix) and helix D (longer black helix).

[0275] Figure 3 A diagram showing the interaction between helical C and helical D in wild-type CiCaSSy (A) and the N267S mutant (B). The α-helix in the center of the diagram represents helical D, and the α-helix to the left represents helical C. At position 267, the side chains of the two amino acids are marked in dark gray.

[0276] Figure 4 The changes in the three major products α-santalene, β-santalene, and bergamotene resulting from improvements in santalene synthase SEQ ID NO:1 (“wild type”) are shown. Values ​​for these three major products have been normalized; minor products are not shown. Black filled bars represent α-santalene, hollow bars represent bergamotene, and diagonally crossed bars represent β-santalene. Replacing position 267 with a serine residue as shown in SEQ ID NO:2 (“N267S”) or a leucine residue as shown in SEQ ID NO:3 (“N267L”) allows the enzyme to produce more β-santalene and more bergamotene than α-santalene (N267S), or more bergamotene, while α-santalene remains substantial, but β-santalene is less (N267L). Data for two santalene synthases known in the art are shown for comparison. Figure 4(Referring to "SaSSy" and "SaSSY-134" in Chinese). Data are taken from reported values ​​in the field, see WO2015153501. Known santalene synthases (wild-type CICassy, ​​SaSSY, and SaSSy-134) show higher yields of α-santalene than the other two compounds. Improved forms N267S and N267L demonstrate how this product characteristic can be altered based on the desired dominance of β-santalene alone exceeding α-santalene, as in N267L, or by the desired dominance of both β-santalene and bergimene exceeding α-santalene, as in N267S.

[0277] Figure 5 The diagram illustrates the changes in the three products α-santalene, bergamotene, and β-santalene produced by santalene synthase when modified alone or in combination at position 291 of SEQ ID NO:1, or at position 267 of SEQ ID NO:1. Minor products are not shown for clarity. Black-filled bars represent α-santalene, hollow bars represent bergamotene, and diagonally crossed bars represent β-santalene. Wild-type (SEQ ID NO:1) and the modified enzyme “I291L” are shown as controls. Replacing position 291 with a leucine residue (“I291L”) as in SEQ ID NO:53 does not alter the fact that it produces an excess of α-santalene compared to β-santalene and bergamotene; on the contrary, this modification increases α-santalol production more than a wild-type enzyme, as can be seen in the attached figures.

[0278] Replacing position 291 with valine, serine, threonine, or cysteine ​​(“I291V”, “I291S”, “I291T”, and “I291C”, respectively) allows the enzyme to produce more β-santalene than α-santalene, while maintaining a much higher level of α-santalene than in the N267S form of the improved β-santalene synthase. The improved forms I291V, I291S, I291C, and I291T demonstrate how this product characteristic can be altered based on the desired advantage of β-santalene alone exceeding α-santalene, as in I291T, I291S, and I291C, or the desired advantage of β-santalene exceeding α-santalene with bergimene levels similar to or higher than those α-santalene levels, as in I291V, while maintaining a higher level of α-santalene compared to the N267S improvement. This characteristic of having more residual α-santalene can be advantageous in several applications.

[0279] The last two groups of bars show the results using two double mutants with modifications at positions 267 and 291 corresponding to SEQ ID NO:1. The data shown for "I291T / N267S" comes from an enzyme where position 267 is filled with serine and position 291 with threonine. The data shown for "I291T / N267T" comes from an enzyme where threonine has been introduced at both positions. Figure 5 The mutants shown exhibit that the double mutant “I291T / N267S” produces the highest percentage of β-santalene. The amounts of α-santalene and bergimene in the “I291T / N267S” enzyme fall in between these values ​​for the two single mutants, with the N267S mutant affecting these values ​​more than the I291T mutant in this combination. Data from other double mutants show that threonine at position 267 has a similar, however less pronounced, effect on α-santalene, bergimene, and β-santalene compared to serine at this position. Example

[0280] Publicly available electronic sequence information was used to analyze the structure of santalene synthase using standard software tools. A 3D model of CiCaSSy (SEQ ID NO:1) (disclosed as SEQ ID NO 3 in international patent application WO2018160066 as a camphor santalene synthase with a normal α-santalene / β-santalene ratio) was generated. Common tools used for this analysis include structure alignment software such as DALI, CE, and STAMP; see http: / / www.rcsb.org / pdb / home / home.do for selection.

[0281] Compared to other santalene synthases, the enzyme CiCaSSy has a slightly uncommon amino acid localization. For example, it shares less than 50% sequence identity with many other santalene synthases, yet incorporates elements from many other santalene synthases in certain regions. A cavity was identified at the active site, and residues within this cavity were targeted for mutagenesis. Specifically, residues that might affect the product profile were preferentially targeted. A region containing two spatially close α-helices in the middle of the amino acid sequence was selected for mutation. CiCaSSy exhibits some amino acid differences in this protein region compared to every known santalene synthase; however, at the same time, many elements are shared by different groups of santalene synthases in combinations unique to CiCaSSy. If this protein region is critical for the desired change in product profile, conversion to other santalene sequences is readily feasible, even if they differ significantly in the remaining portions.

[0282] Mutation test

[0283] Following in-depth research, residue 267 of CiCaSSy was chosen for mutation. Recognizing the favorable environment surrounding N267 in SEQ ID NO:1, the inventors opted to replace the uncommon asparagine at position 267 with serine and leucine, despite these amino acids being present at corresponding positions in other known low-potency santalene synthases. DNA sequences encoding the CiCaSSy protein, along with the two targeted mutations at position 267, were synthesized. The resulting protein sequences, named N267S and N267L, are given in SEQ ID NO:2 and SEQ ID NO:3, respectively.

[0284] The two novel protein sequences were subjected to root mean square deviation (RMSD) and root mean square fluctuation (RSMF) analyses of atomic positions. Each enzyme was simulated for 500 ns under identical conditions (pH 8.0, 300 K, 1 atm, aqueous environment, ion presence in the absence of substrate). RMSD provides an indicator of overall protein motility and flexibility, while RSMF suggests the mean motility and flexibility at a given position. RSMF showed that N267S exhibited the predicted fluctuation and thus increased flexibility beyond wild-type CiCaSSy in the loop region between helices C and D and in the portion of helice D that interacts with the side chain at position 267 of SEQ ID NO: 1 to 3. Increased flexibility was observed in the segment corresponding to positions 272 to 291 in N267S (the region where the side chain of helice D is located, which will interact with the side chain of the amino acid at position 267) compared to wild-type CiCaSSy. The increase is more pronounced in the segment from position 272 to 284, which contains a loop between helical C and helical D, and the loop is expected to be less rigid than the helix. Both N267S and N267L exhibit fluctuating increases in other segments within the region from position 380 to 500, suggesting further downstream flexibility. This pattern was not observed in the sequences analyzed for SEQ ID NO: 4, 5, 8, or 9 when compared with RSMF analysis of other santalene synthases that overproduce α-santalene. RSMD analysis showed that for N267S, the deviation in nm increased by approximately one-fifth from initial equilibrium after 30,000 picoseconds. This structural flexibility was not observed in any of the other santalene sequences analyzed.

[0285] The procedures described in Examples 6 to 19 of WO2018160066 (page 44, line 19 to page 50, line 22; incorporated herein by reference) for wild-type CiCaSSy are applicable to experiments using mutant CiCaSSy sequences encoding proteins N267S and N267L. The mutant DNA sequences encoding CiCaSSy santalene synthases (SEQ ID NO:1 of this invention) encoding SEQ ID NO:2 and 3 are introduced into *Rhodotorula globulus* using the procedures disclosed in the international patent application published in WO2018160066 for CiCaSSy (SEQ ID NO:3), which uses a plasmid-based system to express the heterologous DNA sequence and form the mutant enzyme. *Rhodotorula globulus* fermentation is performed as in WO2018160066 to produce, extract, and analyze α-santalene, β-santalene, and bergimene produced by the host cell.

[0286] Determination of α-santalene, β-santalene, and bergamotene by gas chromatography with an FID detector:

[0287] Gas chromatography was performed using a Shimadzu GC2010Plus column equipped with a Restek RTX-SSil MS capillary column (30 m x 0.25 mm, 0.5 pm). The sample feeder temperature and FID detector temperature were set to 280 °C and 300 °C, respectively. The column gas flow rate was set to 40 mL / min. The initial furnace temperature was 160 °C, increased to 180 °C at a rate of 2 °C / min, and further increased to 300 °C at a rate of 50 °C / min, and held at this temperature for 3 min. The sample volume loaded was 1 μL, the split ratio was 1:50, and the nitrogen tail gas flow rate for cytometry was 30 mL / min.

[0288] Two enzyme mutants at position 267 of CiCaSSy—N267S and N267L—significantly affected the product ratios of α-santalene, β-santalene, and bergimene. Both mutations resulted in increased β-santalene production compared to wild-type CiCaSSy, even producing more β-santalene than α-santalene for the first time, and an increased β-santalene / α-santalene product ratio. Figure 4 Surprisingly, the N267S mutant also produces significantly less α-santalene, suggesting that this mutant has a higher specificity for β-santalene than α-santalene—a phenomenon observed for the first time. The N267L mutant exhibits an even greater change in product ratios and produces β-santalene as its main product, while producing relatively less α-santalene and trans-α-bergamerene, such as... Figure 4 As shown in the image.

[0289] Additional mutants were tested using the same experimental setup described above. For example, replacing the position corresponding to sequence 267 with glycine, alanine, or tryptophan also produced improved santalene synthase.

[0290] Furthermore, it was found that replacing the isoleucine at position 291 of SEQ ID NO:1 in the wild type resulted in higher levels of α-santalene, and the introduction of histidine at this position destroyed its activity as a santalene synthase. This indicates that this position is important, but how it is altered is also significant.

[0291] The I291V, I29S, I291C, I291F, and I291T mutants were also tested and showed—like N267S or N267L—an excess of β-santalene, but still more α-santalene compared to N267S, although less than the wild-type control (see [link to original text]). Figure 5 (and Table 1). The highest percentage of β-santalene was observed when SEQ ID NO:34 was expressed in host cells.

[0292] Double mutants with an N267S or N267T variation at position 267 of SEQ ID NO:1 and a threonine instead of a leucine at position 291 of SEQ ID NO:1 also produce improved β-santalene synthase with β-santalene excess; however, compared to the improved β-santalene synthase from single mutants, they exhibit intercalation product characteristics (see [link to original text]). Figure 5 ).

[0293] Modeling of these mutants in RSMF plots showed that the N257S single mutant and its double mutants with serine, cysteine, or threonine at position 291 of SEQ ID NO:1 exhibited increased flexibility of helical C and helical D, consistent with experimental results for N267S and its double mutants with threonine at position 291 of SEQ ID NO:1 (see [link to RSMF plot]). Figure 5 Interestingly, flexibility also appears to have improved in some other areas further downstream, as RSMF data shows.

[0294] Software tools used

[0295] Homology model

[0296] use Prime software package (www.schrodinger.com / prime); Release 2020-2: Prime Homology models were generated using data from LLC, New York, NY, 2020; M Jacobson et al., Proteins, 2004, 55, 351-367. Template structures were downloaded from the PDB protein database (HM Breman et al., Nucleic Acid Research, 2000, 28, 235-242), and Table 2 indicates the template structures generated for each homology model.

[0297] Homology Model Template structure (PDB code) CiCaSSy 6A1I CiCaSSy N267L 6A1I CiCaSSy N267S 6A1I SaSSy 5ZZJ SaSSy14 5ZZJ ClaSSy 6A1I SaSSy134 5ZZJ

[0298] Table 2. Template structures used to generate homology models. For each constructed homology model, the protein database code indicating the template structure used is listed.

[0299] MD simulation

[0300] MD simulations were performed using the 2018 version of GROMACS software (www.gromacs.org; D van Der Spoel et al., J Comput Chem, 2005, 26, 1701-1718). All enzymes were confined within the OPLS-AA force field (WL Jorgensen and J Tirado-Rives, J Am Chem Soc, 1988, 110, 1657-1666), with protonation confined to pH 8.0 and calculations performed using the pdb2pqr tool (TJ Dolinsky et al., Nucleic Acids Res, 2007 35, W522-W525). As described in MW van der Kamp et al., Biochemistry, 2013, 52, 8094-8105, three metal ions (Mg²⁺, ... 2+ ): Fix their relative positions to their coordinating amino acid residues. Place each enzyme at 1000 nm. 3 The center of the cubic crystal system was clearly solvated with TIP4P water (WL Jorgensen et al., J Chem Phys, 1983, 79, 926-935) by adding an appropriate amount of Na. + or Cl -The total charge system of the ion-neutralized crystal system was simulated. The steepest descent algorithm was used to minimize the charge in 10,000 steps for each crystal system, followed by equilibration for 10 ns. After equilibration, simulations were performed for 500 ns for each crystal system. Temperature was kept constant at 300 K using the v-rescale algorithm (G. Bussi et al., J. Chem. Phys., 2007, 126, 014101), pressure was kept constant at 1 atm using the Parrinello-Rahman algorithm (M. Parrinello and A. Rahman, Phys. Rev. Lett., 1980, 45, 1196-1198), and electrostatic interactions were simulated using the extended particle mesh Ewald algorithm (UE. E. Ssmann et al., J. Chem. Phys., 1995, 103, 8577-8593). The simulation framework was saved every 5 days.

[0301] RMSD

[0302] The root mean square (RMSD) bias of each enzyme structure was evaluated based on the full simulation duration (500 ns). Calculations were performed using the gmx rms tool in the GROMACS software package after structural superposition of protein structures for each trajectory frame (gmx trjconv) was performed using an equilibrium crystal system as a reference.

[0303] RMSF

[0304] The root mean square dependent simulation (RMSF) was used to evaluate the fluctuation of each enzyme structure in the last 450 ns. Calculations were performed using the gmx rmsf tool in the GROMACS software package after structural superposition of protein structures for each trajectory frame (gmx trjconv) and using a protein Cα with an equilibrium crystal system as a reference.

[0305] picture

[0306] Generated using PyMOL (pymol.org) software Figure 2 and Figure 3 Protein images. RMSD and RMSF images were generated using the Matplotlib library (matplotlib.org) for Python version 3.6 (python.org).

[0307] PFAM domain analysis

[0308] The PFAM domain PF01397 "Terpene_synth" and the C-terminal PFAM domain PF03936 "Terpene_synth_C" were identified using PFAM software version 32.0 released on May 29, 2020, and confirmed using PFAM software version 33.1 released on June 11, 2020. For details about PFAM, see "The Pfam protein families database in 2019: S. El-Gebali, J. Mistry, A. Bateman, SREddy, A. Luciani, SCStotter, M. Qureshi, LJ Richardson, GASalazar, A. Smart, E.L. Sonnhammer, L. Hirsh, L. Paladin, D. Piovesan, SCETO Satto, R.D. Finn Nucleic Acids Research (2019)" and http: / / pfam.xfam.org / and "Pfam: The Protein Families Database 2021 (Pfam: Protein Families Database 2021): J. Mistry, S. Chuguransky, L. Williams, M. Qureshi, GASalazar, E.L. Sonnhammer, S.C. Etosalto, L. Paladin, S. Raj, L.J. Richardson, R.D. Finn, A. Bateman. Nucleic Acids Research (2020) doi:10.1093 / nar / gkaa913”

[0309] Interpro motif

[0310] The following structural domains

[0311]

[0312]

[0313] Identified using InterPro scanning software version 83.0 released in December 2020; for more details on InterPro, see: Blum M, Chang H, Chuguransky S, Grego T, Kandasaamy S, Mitchell A, Nuka G, Paysan-Lafosse T, Qureshi M, Raj S, Richardson L, Salazar GA, Williams L, Bork P, Bridge A, Gough J, Haft DH, Letunic I, Marchler-Bauer A, Mi H, Natale DA, Necci M, Orengo CA, Pandurangan AP, Rivoire C, Sigrist CJA, Sillitoe I, Thanki N, Thomas PD, Tosatto SCE, Wu CH, Bateman A and Finn RD, The InterPro protein families and domains database: 20 years on. Nucleic AcidsResearch, Nov 2020, (doi:10.1093 / nar / gkaa977). sequence list <110> ASE Bionics <120> Synthetic santalene synthase <130> 201746WO01 <160> 58 <170> BiSSAP 1.3.6 <210> 1 <211> 553 <212> PRT <213> Camphor tree (Cinnamomum camphora) <220> <223> Wild-type CiCaSSy <400> 1 Met Asp Ser Met Glu Val Arg Arg Ser Ala Ile Tyr His Ser Thr Phe 1 5 10 15 Trp Asp Ile Asp Ser Ile Arg Ala Leu Leu Ala Arg Arg Asp Cys Thr 20 25 30 Ala Ala Ala Ala Leu Ser Pro Asp His His Lys Arg Leu Lys Glu Arg 35 40 45 Ile Gln Arg Arg Leu Gln Asp Ile Thr Gln Pro His His Leu Leu Gly 50 55 60 Leu Ile Asp Ala Val Gln Arg Leu Gly Val Ala Tyr Gln Phe Glu Glu 65 70 75 80 Glu Ile Ser Asp Ala Leu His Gly Leu His Ser Glu Asn Thr Glu His 85 90 95 Ala Ile Lys Asp Ser Leu His His Thr Ser Leu Tyr Phe Arg Leu Leu 100 105 110 Arg Gln His Gly Cys Asn Leu Ser Ser Asp Ile Phe Asn Lys Phe Lys 115 120 125 Lys Glu Gly Gly Gly Phe Lys Ala Ser Leu Cys Glu Asp Ala Met Gly 130 135 140 Leu Leu Ser Leu Tyr Glu Ala Val Arg Leu Ser Val Lys Gly Glu Ala 145 150 155 160 Ile Leu Glu Glu Ala Gln Val Phe Ser Ile Ala Asn Leu Lys Ile Leu 165 170 175 Met Glu Arg Val Glu Arg Lys Leu Ala Asp Arg Ile Glu His Ala Leu 180 185 190 Glu Ile Pro Leu Tyr Trp Arg Ala Pro Arg Leu Glu Ala Arg Trp Tyr 195 200 205 Ile Asp Val Tyr Glu Lys Glu Asp Gly Arg Ile Asp Asp Leu Leu Asp 210 215 220 Phe Ala Lys Leu Asp Phe Asn Arg Val Gln Met Leu Tyr Gln Thr Glu 225 230 235 240 Leu Lys Glu Leu Ser Met Trp Trp Glu Leu Leu Gly Leu Pro Ala Lys 245 250 255 Met Gly Phe Phe Arg Asp Arg Leu Leu Glu Asn His Leu Phe Ser Ile 260 265 270 Ala Val Val Val Glu Pro Gln Tyr Ser Gln Cys Arg Val Ala Ile Thr 275 280 285 Lys Ala Ile Val Leu Met Thr Ala Met Asp Asp Phe Tyr Asp Val His 290 295 300 Gly Leu Pro Asp Glu Leu Lys Val Phe Thr Asp Thr Val Asn Arg Trp 305 310 315 320 Asp Leu Glu Gly Ile Asp Gln Leu Pro Glu Tyr Met Lys Leu Tyr Tyr 325 330 335 Leu Ala Leu Tyr Asn Thr Thr Asn Glu Thr Ala Tyr Ile Ile Leu Lys 340 345 350 Glu Lys Gly Phe Asn Ala Thr His Tyr Leu Lys Lys Leu Trp Ala Met 355 360 365 Gln Ser Asn Ala Tyr Phe Arg Glu Ala Gln Trp Phe Asn Ser Gly Tyr 370 375 380 Ile Pro Lys Phe Asp Glu Tyr Leu Asp Asn Ala Leu Val Ser Val Gly 385 390 395 400 Ala Pro Phe Val Leu Gly Leu Ser Tyr Pro Met Ile Gln Gln Gln Ile 405 410 415 Ser Lys Glu Glu Ile Asp Leu Ile Pro Glu Asp Leu Asn Leu Leu Arg 420 425 430 Trp Ala Ser Ile Ile Phe Arg Leu Tyr Asp Asp Leu Ala Thr Ser Lys 435 440 445 Ala Glu Gln Gln Arg Gly Asp Val Pro Lys Ser Ile Gln Cys Tyr Met 450 455 460 His Glu Thr Gly Ser Ser Glu Glu Val Ala Ala Asn His Ile Arg Asp 465 470 475 480 Leu Ile Ser Asp Ala Trp Lys Glu Val Asn Ala Glu Cys Leu Lys Pro 485 490 495 Thr Ser Leu Ser Lys His Tyr Val Gly Val Ala Pro Asn Ser Ala Arg 500 505 510 Ser Gly Val Leu Met Tyr His His Asp Phe Asp Gly Phe Ala Ser Pro 515 520 525 His Gly Arg Thr Asn Ala His Ile Thr Ser Ile Phe Phe Glu Pro Val 530 535 540 Pro Leu Lys Glu Ser Ile Asn Leu Gly 545 550 <210> 2 <211> 553 <212> PRT <213> artificial sequence <220> <223> >CiCaSSy_N267S <400> 2 Met Asp Ser Met Glu Val Arg Arg Ser Ala Ile Tyr His Ser Thr Phe 1 5 10 15 Trp Asp Ile Asp Ser Ile Arg Ala Leu Leu Ala Arg Arg Asp Cys Thr 20 25 30 Ala Ala Ala Ala Leu Ser Pro Asp His His Lys Arg Leu Lys Glu Arg 35 40 45 Ile Gln Arg Arg Leu Gln Asp Ile Thr Gln Pro His His Leu Leu Gly 50 55 60 Leu Ile Asp Ala Val Gln Arg Leu Gly Val Ala Tyr Gln Phe Glu Glu 65 70 75 80 Glu Ile Ser Asp Ala Leu His Gly Leu His Ser Glu Asn Thr Glu His 85 90 95 Ala Ile Lys Asp Ser Leu His His Thr Ser Leu Tyr Phe Arg Leu Leu 100 105 110 Arg Gln His Gly Cys Asn Leu Ser Ser Asp Ile Phe Asn Lys Phe Lys 115 120 125 Lys Glu Gly Gly Gly Phe Lys Ala Ser Leu Cys Glu Asp Ala Met Gly 130 135 140 Leu Leu Ser Leu Tyr Glu Ala Val Arg Leu Ser Val Lys Gly Glu Ala 145 150 155 160 Ile Leu Glu Glu Ala Gln Val Phe Ser Ile Ala Asn Leu Lys Ile Leu 165 170 175 Met Glu Arg Val Glu Arg Lys Leu Ala Asp Arg Ile Glu His Ala Leu 180 185 190 Glu Ile Pro Leu Tyr Trp Arg Ala Pro Arg Leu Glu Ala Arg Trp Tyr 195 200 205 Ile Asp Val Tyr Glu Lys Glu Asp Gly Arg Ile Asp Asp Leu Leu Asp 210 215 220 Phe Ala Lys Leu Asp Phe Asn Arg Val Gln Met Leu Tyr Gln Thr Glu 225 230 235 240 Leu Lys Glu Leu Ser Met Trp Trp Glu Leu Leu Gly Leu Pro Ala Lys 245 250 255 Met Gly Phe Phe Arg Asp Arg Leu Leu Glu Ser His Leu Phe Ser Ile 260 265 270 Ala Val Val Val Glu Pro Gln Tyr Ser Gln Cys Arg Val Ala Ile Thr 275 280 285 Lys Ala Ile Val Leu Met Thr Ala Met Asp Asp Phe Tyr Asp Val His 290 295 300 Gly Leu Pro Asp Glu Leu Lys Val Phe Thr Asp Thr Val Asn Arg Trp 305 310 315 320 Asp Leu Glu Gly Ile Asp Gln Leu Pro Glu Tyr Met Lys Leu Tyr Tyr 325 330 335 Leu Ala Leu Tyr Asn Thr Thr Asn Glu Thr Ala Tyr Ile Ile Leu Lys 340 345 350 Glu Lys Gly Phe Asn Ala Thr His Tyr Leu Lys Lys Leu Trp Ala Met 355 360 365 Gln Ser Asn Ala Tyr Phe Arg Glu Ala Gln Trp Phe Asn Ser Gly Tyr 370 375 380 Ile Pro Lys Phe Asp Glu Tyr Leu Asp Asn Ala Leu Val Ser Val Gly 385 390 395 400 Ala Pro Phe Val Leu Gly Leu Ser Tyr Pro Met Ile Gln Gln Gln Ile 405 410 415 Ser Lys Glu Glu Ile Asp Leu Ile Pro Glu Asp Leu Asn Leu Leu Arg 420 425 430 Trp Ala Ser Ile Ile Phe Arg Leu Tyr Asp Asp Leu Ala Thr Ser Lys 435 440 445 Ala Glu Gln Gln Arg Gly Asp Val Pro Lys Ser Ile Gln Cys Tyr Met 450 455 460 His Glu Thr Gly Ser Ser Glu Glu Val Ala Ala Asn His Ile Arg Asp 465 470 475 480 Leu Ile Ser Asp Ala Trp Lys Glu Val Asn Ala Glu Cys Leu Lys Pro 485 490 495 Thr Ser Leu Ser Lys His Tyr Val Gly Val Ala Pro Asn Ser Ala Arg 500 505 510 Ser Gly Val Leu Met Tyr His His Asp Phe Asp Gly Phe Ala Ser Pro 515 520 525 His Gly Arg Thr Asn Ala His Ile Thr Ser Ile Phe Phe Glu Pro Val 530 535 540 Pro Leu Lys Glu Ser Ile Asn Leu Gly 545 550 <210> 3 <211> 553 <212> PRT <213> artificial sequence <220> <223> >CiCaSSy_N267L <400> 3 Met Asp Ser Met Glu Val Arg Arg Ser Ala Ile Tyr His Ser Thr Phe 1 5 10 15 Trp Asp Ile Asp Ser Ile Arg Ala Leu Leu Ala Arg Arg Asp Cys Thr 20 25 30 Ala Ala Ala Ala Leu Ser Pro Asp His His Lys Arg Leu Lys Glu Arg 35 40 45 Ile Gln Arg Arg Leu Gln Asp Ile Thr Gln Pro His His Leu Leu Gly 50 55 60 Leu Ile Asp Ala Val Gln Arg Leu Gly Val Ala Tyr Gln Phe Glu Glu 65 70 75 80 Glu Ile Ser Asp Ala Leu His Gly Leu His Ser Glu Asn Thr Glu His 85 90 95 Ala Ile Lys Asp Ser Leu His His Thr Ser Leu Tyr Phe Arg Leu Leu 100 105 110 Arg Gln His Gly Cys Asn Leu Ser Ser Asp Ile Phe Asn Lys Phe Lys 115 120 125 Lys Glu Gly Gly Gly Phe Lys Ala Ser Leu Cys Glu Asp Ala Met Gly 130 135 140 Leu Leu Ser Leu Tyr Glu Ala Val Arg Leu Ser Val Lys Gly Glu Ala 145 150 155 160 Ile Leu Glu Glu Ala Gln Val Phe Ser Ile Ala Asn Leu Lys Ile Leu 165 170 175 Met Glu Arg Val Glu Arg Lys Leu Ala Asp Arg Ile Glu His Ala Leu 180 185 190 Glu Ile Pro Leu Tyr Trp Arg Ala Pro Arg Leu Glu Ala Arg Trp Tyr 195 200 205 Ile Asp Val Tyr Glu Lys Glu Asp Gly Arg Ile Asp Asp Leu Leu Asp 210 215 220 Phe Ala Lys Leu Asp Phe Asn Arg Val Gln Met Leu Tyr Gln Thr Glu 225 230 235 240 Leu Lys Glu Leu Ser Met Trp Trp Glu Leu Leu Gly Leu Pro Ala Lys 245 250 255 Met Gly Phe Phe Arg Asp Arg Leu Leu Glu Leu His Leu Phe Ser Ile 260 265 270 Ala Val Val Val Glu Pro Gln Tyr Ser Gln Cys Arg Val Ala Ile Thr 275 280 285 Lys Ala Ile Val Leu Met Thr Ala Met Asp Asp Phe Tyr Asp Val His 290 295 300 Gly Leu Pro Asp Glu Leu Lys Val Phe Thr Asp Thr Val Asn Arg Trp 305 310 315 320 Asp Leu Glu Gly Ile Asp Gln Leu Pro Glu Tyr Met Lys Leu Tyr Tyr 325 330 335 Leu Ala Leu Tyr Asn Thr Thr Asn Glu Thr Ala Tyr Ile Ile Leu Lys 340 345 350 Glu Lys Gly Phe Asn Ala Thr His Tyr Leu Lys Lys Leu Trp Ala Met 355 360 365 Gln Ser Asn Ala Tyr Phe Arg Glu Ala Gln Trp Phe Asn Ser Gly Tyr 370 375 380 Ile Pro Lys Phe Asp Glu Tyr Leu Asp Asn Ala Leu Val Ser Val Gly 385 390 395 400 Ala Pro Phe Val Leu Gly Leu Ser Tyr Pro Met Ile Gln Gln Gln Ile 405 410 415 Ser Lys Glu Glu Ile Asp Leu Ile Pro Glu Asp Leu Asn Leu Leu Arg 420 425 430 Trp Ala Ser Ile Ile Phe Arg Leu Tyr Asp Asp Leu Ala Thr Ser Lys 435 440 445 Ala Glu Gln Gln Arg Gly Asp Val Pro Lys Ser Ile Gln Cys Tyr Met 450 455 460 His Glu Thr Gly Ser Ser Glu Glu Val Ala Ala Asn His Ile Arg Asp 465 470 475 480 Leu Ile Ser Asp Ala Trp Lys Glu Val Asn Ala Glu Cys Leu Lys Pro 485 490 495 Thr Ser Leu Ser Lys His Tyr Val Gly Val Ala Pro Asn Ser Ala Arg 500 505 510 Ser Gly Val Leu Met Tyr His His Asp Phe Asp Gly Phe Ala Ser Pro 515 520 525 His Gly Arg Thr Asn Ala His Ile Thr Ser Ile Phe Phe Glu Pro Val 530 535 540 Pro Leu Lys Glu Ser Ile Asn Leu Gly 545 550 <210> 4 <211> 569 <212> PRT <213> Sandalwood (Santalum album) <220> <223> >SaSSy <400> 4 Met Asp Ser Ser Thr Ala Thr Ala Met Thr Ala Pro Phe Ile Asp Pro 1 5 10 15 Thr Asp His Val Asn Leu Lys Thr Asp Thr Asp Ala Ser Glu Asn Arg 20 25 30 Arg Met Gly Asn Tyr Lys Pro Ser Ile Trp Asn Tyr Asp Phe Leu Gln 35 40 45 Ser Leu Ala Thr His His Asn Ile Val Glu Glu Arg His Leu Lys Leu 50 55 60 Ala Glu Lys Leu Lys Gly Gln Val Lys Phe Met Phe Gly Ala Pro Met 65 70 75 80 Glu Pro Leu Ala Lys Leu Glu Leu Val Asp Val Val Gln Arg Leu Gly 85 90 95 Leu Asn His Leu Phe Glu Thr Glu Ile Lys Glu Ala Leu Phe Ser Ile 100 105 110 Tyr Lys Asp Gly Ser Asn Gly Trp Trp Phe Gly His Leu His Ala Thr 115 120 125 Ser Leu Arg Phe Arg Leu Leu Arg Gln Cys Gly Leu Phe Ile Pro Gln 130 135 140 Asp Val Phe Lys Thr Phe Gln Asn Lys Thr Gly Glu Phe Asp Met Lys 145 150 155 160 Leu Cys Asp Asn Val Lys Gly Leu Leu Ser Leu Tyr Glu Ala Ser Tyr 165 170 175 Leu Gly Trp Lys Gly Glu Asn Ile Leu Asp Glu Ala Lys Ala Phe Thr 180 185 190 Thr Lys Cys Leu Lys Ser Ala Trp Glu Asn Ile Ser Glu Lys Trp Leu 195 200 205 Ala Lys Arg Val Lys His Ala Leu Ala Leu Pro Leu His Trp Arg Val 210 215 220 Pro Arg Ile Glu Ala Arg Trp Phe Ile Glu Ala Tyr Glu Gln Glu Ala 225 230 235 240 Asn Met Asn Pro Thr Leu Leu Lys Leu Ala Lys Leu Asp Phe Asn Met 245 250 255 Val Gln Ser Ile His Gln Lys Glu Ile Gly Glu Leu Ala Arg Trp Trp 260 265 270 Val Thr Thr Gly Leu Asp Lys Leu Ala Phe Ala Arg Asn Asn Leu Leu 275 280 285 Gln Ser Tyr Met Trp Ser Cys Ala Ile Ala Ser Asp Pro Lys Phe Lys 290 295 300 Leu Ala Arg Glu Thr Ile Val Glu Ile Gly Ser Val Leu Thr Val Val 305 310 315 320 Asp Asp Gly Tyr Asp Val Tyr Gly Ser Ile Asp Glu Leu Asp Leu Tyr 325 330 335 Thr Ser Ser Val Glu Arg Trp Ser Cys Val Glu Ile Asp Lys Leu Pro 340 345 350 Asn Thr Leu Lys Leu Ile Phe Met Ser Met Phe Asn Lys Thr Asn Glu 355 360 365 Val Gly Leu Arg Val Gln His Glu Arg Gly Tyr Asn Ser Ile Pro Thr 370 375 380 Phe Ile Lys Ala Trp Val Glu Gln Cys Lys Ser Tyr Gln Lys Glu Ala 385 390 395 400 Arg Trp Phe His Gly Gly His Thr Pro Pro Leu Glu Glu Tyr Ser Leu 405 410 415 Asn Gly Leu Val Ser Ile Gly Phe Pro Leu Leu Leu Ile Thr Gly Tyr 420 425 430 Val Ala Ile Ala Glu Asn Glu Ala Ala Leu Asp Lys Val His Pro Leu 435 440 445 Pro Asp Leu Leu His Tyr Ser Ser Leu Leu Ser Arg Leu Ile Asn Asp 450 455 460 Ile Gly Thr Ser Pro Asp Glu Met Ala Arg Gly Asp Asn Leu Lys Ser 465 470 475 480 Ile His Cys Tyr Met Asn Glu Thr Gly Ala Ser Glu Glu Val Ala Arg 485 490 495 Glu His Ile Lys Gly Val Ile Glu Glu Asn Trp Lys Ile Leu Asn Gln 500 505 510 Cys Cys Phe Asp Gln Ser Gln Phe Gln Glu Pro Phe Ile Thr Phe Asn 515 520 525 Leu Asn Ser Val Arg Gly Ser His Phe Phe Tyr Glu Phe Gly Asp Gly 530 535 540 Phe Gly Val Thr Asp Ser Trp Thr Lys Val Asp Met Lys Ser Val Leu 545 550 555 560 Ile Asp Pro Ile Pro Leu Gly Glu Glu 565 <210> 5 <211> 569 <212> PRT <213> Sandalwood <220> <223> >SaSSy14 <400> 5 Met Asp Ser Ser Thr Ala Thr Ala Met Arg Ala Pro Phe Ile Asp His 1 5 10 15 Thr Asp His Val Asn Leu Arg Thr Asp Asn Asp Ser Ser Glu Asn Arg 20 25 30 Arg Met Gly Asn Tyr Lys Pro Ser Ile Trp Asn Tyr Asp Phe Leu Gln 35 40 45 Ser Leu Ala Thr Arg His Asn Ile Met Glu Glu Arg His Leu Lys Leu 50 55 60 Ala Glu Lys Leu Lys Gly Gln Val Lys Phe Met Phe Gly Ala Pro Met 65 70 75 80 Glu Pro Leu Ala Lys Leu Glu Leu Val Asp Val Val Gln Arg Leu Gly 85 90 95 Leu Asn His Arg Phe Glu Thr Glu Ile Lys Glu Ala Leu Phe Ser Ile 100 105 110 Tyr Lys Asp Glu Ser Asn Gly Trp Trp Phe Gly His Leu His Ala Thr 115 120 125 Ser Leu Arg Phe Arg Leu Leu Arg Gln Cys Gly Leu Phe Ile Pro Gln 130 135 140 Asp Val Phe Lys Thr Phe Gln Ser Lys Thr Gly Glu Phe Asp Met Lys 145 150 155 160 Leu Cys Asp Asn Val Lys Gly Leu Leu Ser Leu Tyr Glu Ala Ser Phe 165 170 175 Leu Gly Trp Arg Asp Glu Asn Ile Leu Asp Glu Ala Lys Ala Phe Ala 180 185 190 Thr Lys Tyr Leu Lys Asn Ala Trp Glu Asn Ile Ser Gln Lys Trp Leu 195 200 205 Ala Lys Arg Val Lys His Ala Leu Ala Leu Pro Leu His Trp Arg Val 210 215 220 Pro Arg Ile Glu Ala Arg Trp Phe Val Glu Ala Tyr Gly Glu Glu Glu 225 230 235 240 Asn Met Asn Pro Thr Leu Leu Lys Leu Ala Lys Leu Asp Phe Asn Met 245 250 255 Val Gln Ser Ile His Gln Lys Glu Ile Gly Glu Leu Ala Arg Trp Trp 260 265 270 Val Thr Thr Gly Leu Asp Lys Leu Ala Phe Ala Arg Asn Asn Leu Leu 275 280 285 Gln Ser Tyr Met Trp Ser Cys Ala Ile Ala Ser Asp Pro Lys Phe Lys 290 295 300 Leu Ala Arg Glu Thr Ile Val Glu Ile Gly Ser Val Leu Thr Val Val 305 310 315 320 Asp Asp Ala Tyr Asp Val Tyr Gly Ser Met Asp Glu Leu Asp Leu Tyr 325 330 335 Thr Asn Ser Val Glu Arg Trp Ser Cys Thr Glu Ile Asp Lys Leu Pro 340 345 350 Asn Thr Leu Lys Leu Ile Phe Met Ala Met Phe Asn Lys Thr Asn Glu 355 360 365 Val Gly Leu Arg Val Gln His Glu Arg Gly Tyr Ser Gly Ile Thr Thr 370 375 380 Phe Ile Lys Ala Trp Val Glu Gln Cys Lys Ser Tyr Gln Lys Glu Ala 385 390 395 400 Arg Trp Tyr His Gly Gly His Thr Pro Pro Leu Glu Glu Tyr Ser Leu 405 410 415 Asn Gly Leu Val Ser Ile Gly Phe Pro Leu Leu Leu Ile Thr Gly Tyr 420 425 430 Val Ala Ile Ala Glu Asn Glu Ala Ala Leu Asp Lys Val His Pro Leu 435 440 445 Pro Asp Leu Leu His Tyr Ser Ser Leu Leu Ser Arg Leu Ile Asn Asp 450 455 460 Met Gly Thr Ser Ser Asp Glu Leu Glu Arg Gly Asp Asn Leu Lys Ser 465 470 475 480 Ile Gln Cys Tyr Met Asn Gln Thr Gly Ala Ser Glu Lys Val Ala Arg 485 490 495 Glu His Ile Lys Gly Ile Ile Glu Glu Asn Trp Lys Ile Leu Asn Glu 500 505 510 Cys Cys Phe Asp Gln Ser Gln Phe Gln Glu Pro Phe Val Thr Phe Asn 515 520 525 Leu Asn Ser Val Arg Gly Ser His Phe Phe Tyr Glu Phe Gly Asp Gly 530 535 540 Phe Gly Val Thr Asn Ser Trp Thr Lys Val Asp Met Lys Ser Val Leu 545 550 555 560 Ile Asp Pro Ile Pro Leu Asp Glu Glu 565 <210> 6 <211> 569 <212> PRT <213> Santalum spicatum (big-fruited sandalwood) <220> <223> >SpiSSy <400> 6 Met Asp Ser Ser Thr Ala Thr Ala Thr Thr Ala Pro Phe Ile Asp His 1 5 10 15 Thr Asp His Val Asn Leu Lys Ile Asp Asn Asp Ser Ser Glu Ser Arg 20 25 30 Arg Met Gly Asn Tyr Lys Pro Ser Ile Trp Asn Tyr Asp Phe Leu Gln[[ID=(41]] 35 40 45 Ser Leu Ala Ile His His Asn Ile Val Glu Glu Lys His Leu Lys Leu 50 55 60 Ala Glu Lys Leu Lys Gly Gln Val Met Ser Met Phe Gly Ala Pro Met 65 70 75 80 Glu Pro Leu Ala Lys Leu Glu Leu Val Asp Val Val Gln Arg Leu Gly 85 90 95 Leu Asn His Gln Phe Glu Thr Glu Ile Lys Glu Ala Leu Phe Ser Val 100 105 110 Tyr Lys Asp Gly Ser Asn Gly Trp Trp Phe Gly His Leu His Ala Thr 115 120 125 Ser Leu Arg Phe Arg Leu Leu Arg Gln Cys Gly Leu Phe Ile Pro Gln 130 135 140 Asp Val Phe Lys Thr Phe Gln Ser Lys Thr Asp Glu Phe Asp Met Lys 145 150 155 160 Leu Cys Asp Asn Ile Lys Gly Leu Leu Ser Leu Tyr Glu Ala Ser Phe 165 170 175 Leu Gly Trp Lys Gly Glu Asn Ile Leu Asp Glu Ala Lys Ala Phe Ala 180 185 190 Thr Lys Tyr Leu Lys Asn Ala Trp Glu Asn Ile Ser Gln Lys Trp Leu 195 200 205 Ala Lys Arg Val Lys His Ala Leu Ala Leu Pro Leu His Trp Arg Val 210 215 220 Pro Arg Ile Glu Ala Arg Trp Phe Ile Glu Ala Tyr Glu Gln Glu Glu 225 230 235 240 Asn Met Asn Pro Thr Leu Leu Lys Leu Ala Lys Leu Asp Phe Asn Met 245 250 255 Val Gln Ser Ile His Gln Lys Glu Ile Gly Glu Leu Ala Arg Trp Trp 260 265 270 Val Thr Thr Gly Leu Asp Lys Leu Ala Phe Ala Arg Asn Asn Leu Leu 275 280 285 Gln Ser Tyr Met Trp Ser Cys Ala Ile Ala Ser Asp Pro Lys Phe Lys 290 295 300 Leu Ala Arg Glu Thr Ile Val Glu Ile Gly Ser Val Leu Thr Val Val 305 310 315 320 Asp Asp Ala Tyr Asp Val Tyr Gly Ser Met Asp Glu Leu Asp His Tyr 325 330 335 Thr Tyr Ser Val Glu Arg Trp Ser Cys Val Glu Ile Asp Lys Leu Pro 340 345 350 Asn Thr Leu Lys Leu Ile Phe Met Ser Met Phe Asn Lys Thr Asn Glu 355 360 365 Val Gly Leu Arg Val Gln His Glu Arg Gly Tyr Asn Gly Ile Pro Thr 370 375 380 Phe Ile Lys Ala Trp Val Glu Gln Cys Lys Ala Tyr Gln Lys Glu Ala 385 390 395 400 Arg Trp Tyr His Gly Gly His Thr Pro Pro Leu Glu Glu Tyr Ser Leu 405 410 415 Asn Gly Leu Val Ser Ile Gly Phe Pro Leu Leu Leu Ile Thr Gly Tyr 420 425 430 Ile Ala Ile Ala Glu Asn Glu Ala Ala Leu Asp Lys Val His Pro Leu 435 440 445 Pro Asp Leu Leu His Tyr Ser Ser Leu Leu Ser Arg Leu Ile Asn Asp 450 455 460 Met Gly Thr Ser Pro Asp Glu Met Ala Arg Gly Asp Asn Leu Lys Ser 465 470 475 480 Ile His Cys Tyr Met Asn Glu Thr Gly Ala Ser Glu Glu Val Ala Arg 485 490 495 Glu His Ile Lys Gly Ile Ile Glu Glu Asn Trp Lys Ile Leu Asn Gln 500 505 510 Cys Cys Phe Asp Gln Ser Gln Phe Gln Glu Pro Phe Ile Thr Phe Asn 515 520 525 Leu Asn Ser Val Arg Gly Ser His Phe Phe Tyr Glu Phe Gly Asp Gly 530 535 540 Phe Gly Val Thr Asp Ser Trp Thr Lys Val Asp Met Lys Ser Val Leu 545 550 555 560 Ile Asp Pro Ile Pro Leu Gly Glu Glu 565 <210> 7 <211> 569 <212> PRT <213> Santalum austrocaledonicum <220> <223> >SauSSy <400> 7 Met Asp Ser Ser Thr Ala Thr Ala Met Thr Ala Pro Phe Ile Asp Pro 1 5 10 15 Thr Asp His Val Asn Leu Lys Thr Asp Thr Asp Ala Ser Glu Asn Arg 20 25 30 Arg Met Gly Asn Tyr Lys Pro Ser Ile Trp Asn Tyr Asp Phe Leu Gln 35 40 45 Ser Leu Ala Thr His His Asn Ile Val Glu Glu Arg His Leu Lys Leu 50 55 60 Ala Glu Lys Leu Lys Gly Gln Val Lys Phe Met Phe Gly Ala Pro Met 65 70 75 80 Glu Pro Leu Ala Lys Leu Glu Leu Val Asp Val Val Gln Arg Leu Gly 85 90 95 Leu Asn His Arg Phe Glu Thr Glu Ile Lys Glu Ala Leu Phe Ser Ile 100 105 110 Tyr Lys Asp Glu Ser Asn Gly Trp Trp Phe Gly His Leu His Ala Thr 115 120 125 Ser Leu Arg Phe Arg Leu Leu Arg Gln Cys Gly Leu Phe Ile Pro Gln 130 135 140 Asp Val Phe Lys Thr Phe Gln Asn Lys Thr Gly Glu Phe Asp Met Lys 145 150 155 160 Leu Cys Asp Asn Val Lys Gly Leu Leu Ser Leu Tyr Glu Ala Ser Tyr 165 170 175 Leu Gly Trp Lys Gly Glu Asn Ile Leu Asp Glu Ala Lys Ala Phe Ala 180 185 190 Thr Lys Tyr Leu Lys Ser Ala Trp Glu Asn Ile Ser Glu Lys Trp Leu 195 200 205 Ala Lys Arg Val Lys His Ala Leu Ala Leu Pro Leu His Trp Arg Val 210 215 220 Pro Arg Ile Glu Ala Arg Trp Phe Ile Glu Ala Tyr Glu Gln Glu Ala 225 230 235 240 Asn Met Asn Pro Thr Leu Leu Lys Leu Ala Lys Leu Asp Phe Asn Met 245 250 255 Val Gln Ser Ile His Gln Lys Glu Ile Gly Glu Leu Ala Arg Trp Trp 260 265 270 Val Thr Thr Gly Leu Asp Lys Leu Ala Phe Ala Arg Asn Asn Leu Leu 275 280 285 Gln Ser Tyr Met Trp Ser Cys Ala Ile Ala Ser Asp Pro Lys Phe Lys 290 295 300 Leu Ala Arg Glu Thr Ile Val Glu Ile Gly Ser Val Leu Thr Val Val 305 310 315 320 Asp Asp Ala Tyr Asp Val Tyr Gly Ser Met Asp Glu Leu Asp Leu Tyr 325 330 335 Thr Ser Ser Val Glu Arg Trp Ser Cys Val Glu Ile Asp Lys Leu Pro 340 345 350 Asn Thr Leu Lys Leu Ile Phe Met Ser Met Phe Asn Lys Thr Asn Glu 355 360 365 Val Gly Leu Arg Val Gln His Glu Arg Gly Tyr Asn Ser Ile Pro Thr 370 375 380 Phe Ile Lys Ala Trp Val Gln Gln Cys Lys Ser Tyr Gln Lys Glu Ala 385 390 395 400 Arg Trp Phe His Gly Gly His Thr Pro Pro Leu Glu Glu Tyr Ser Leu 405 410 415 Asn Gly Leu Val Ser Ile Gly Phe Pro Leu Leu Leu Ile Thr Gly Tyr 420 425 430 Val Ala Ile Ala Glu Asn Glu Ala Ala Leu Asp Lys Val His Pro Leu 435 440 445 Pro Asp Leu Leu His Tyr Ser Ser Leu Leu Ser Arg Leu Ile Asn Asp 450 455 460 Ile Gly Thr Ser Pro Asp Glu Met Ala Arg Gly Asp Asn Leu Lys Ser 465 470 475 480 Ile His Cys Tyr Met Asn Gly Thr Gly Ala Ser Glu Glu Val Ala Arg 485 490 495 Glu His Ile Lys Gly Val Ile Glu Glu Asn Trp Lys Ile Leu Asn Gln 500 505 510 Cys Cys Phe Asp Gln Ser Gln Phe Gln Glu Pro Phe Ile Thr Phe Asn 515 520 525 Leu Asn Ser Val Arg Gly Ser His Phe Phe Tyr Glu Phe Gly Asp Gly 530 535 540 Phe Gly Val Thr Asp Ser Trp Thr Lys Val Asp Met Lys Ser Val Leu 545 550 555 560 Ile Asp Pro Ile Pro Leu Gly Glu Glu 565 <210> 8 <211> 551 <212> PRT <213> Clausena lansium <220> <223> >ClaSSy <400> 8 Met Ser Thr Gln Gln Val Ser Ser Glu Asn Ile Val Arg Asn Ala Ala 1 5 10 15 Asn Phe His Pro Asn Ile Trp Gly Asn His Phe Leu Thr Cys Pro Ser 20 25 30 Gln Thr Ile Asp Ser Trp Thr Gln Gln His His Lys Glu Leu Lys Glu 35 40 45 Glu Val Arg Lys Met Met Val Ser Asp Ala Asn Lys Pro Ala Gln Arg 50 55 60 Leu Arg Leu Ile Asp Thr Val Gln Arg Leu Gly Val Ala Tyr His Phe 65 70 75 80 Glu Lys Glu Ile Asp Asp Ala Leu Glu Lys Ile Gly His Asp Pro Phe 85 90 95 Asp Asp Lys Asp Asp Leu Tyr Ile Val Ser Leu Cys Phe Arg Leu Leu 100 105 110 Arg Gln His Gly Ile Lys Ile Ser Cys Asp Val Phe Glu Lys Phe Lys 115 120 125 Asp Asp Asp Gly Lys Phe Lys Ala Ser Leu Met Asn Asp Val Gln Gly 130 135 140 Met Leu Ser Leu Tyr Glu Ala Ala His Leu Ala Ile His Gly Glu Asp 145 150 155 160 Ile Leu Asp Glu Ala Ile Val Phe Thr Thr Thr His Leu Lys Ser Thr 165 170 175 Val Ser Asn Ser Pro Val Asn Ser Thr Phe Ala Glu Gln Ile Arg His 180 185 190 Ser Leu Arg Val Pro Leu Arg Lys Ala Val Pro Arg Leu Glu Ser Arg 195 200 205 Tyr Phe Leu Asp Ile Tyr Ser Arg Asp Asp Leu His Asp Lys Thr Leu 210 215 220 Leu Asn Phe Ala Lys Leu Asp Phe Asn Ile Leu Gln Ala Met His Gln 225 230 235 240 Lys Glu Ala Ser Glu Met Thr Arg Trp Trp Arg Asp Phe Asp Phe Leu 245 250 255 Lys Lys Leu Pro Tyr Ile Arg Asp Arg Val Val Glu Leu Tyr Phe Trp 260 265 270 Ile Leu Val Gly Val Ser Tyr Gln Pro Lys Phe Ser Thr Gly Arg Ile 275 280 285 Phe Leu Ser Lys Ile Ile Cys Leu Glu Thr Leu Val Asp Asp Thr Phe 290 295 300 Asp Ala Tyr Gly Thr Phe Asp Glu Leu Ala Ile Phe Thr Glu Ala Val 305 310 315 320 Thr Arg Trp Asp Leu Gly His Arg Asp Ala Leu Pro Glu Tyr Met Lys 325 330 335 Phe Ile Phe Lys Thr Leu Ile Asp Val Tyr Ser Glu Ala Glu Gln Glu 340 345 350 Leu Ala Lys Glu Gly Arg Ser Tyr Ser Ile His Tyr Ala Ile Arg Ser 355 360 365 Phe Gln Glu Leu Val Met Lys Tyr Phe Cys Glu Ala Lys Trp Leu Asn 370 375 380 Lys Gly Tyr Val Pro Ser Leu Asp Asp Tyr Lys Ser Val Ser Leu Arg 385 390 395 400 Ser Ile Gly Phe Leu Pro Ile Ala Val Ala Ser Phe Val Phe Met Gly 405 410 415 Asp Ile Ala Thr Lys Glu Val Phe Glu Trp Glu Met Asn Asn Pro Lys 420 425 430 Ile Ile Ile Ala Ala Glu Thr Ile Phe Arg Phe Leu Asp Asp Ile Ala 435 440 445 Gly His Arg Phe Glu Gln Lys Arg Glu His Ser Pro Ser Ala Ile Glu 450 455 460 Cys Tyr Lys Asn Gln His Gly Val Ser Glu Glu Glu Ala Val Lys Ala 465 470 475 480 Leu Ser Leu Glu Val Ala Asn Ser Trp Lys Asp Ile Asn Glu Glu Leu 485 490 495 Leu Leu Asn Pro Met Ala Ile Pro Leu Pro Leu Leu Gln Val Ile Leu 500 505 510 Asp Leu Ser Arg Ser Ala Asp Phe Met Tyr Gly Asn Ala Gln Asp Arg 515 520 525 Phe Thr His Ser Thr Met Met Lys Asp Gln Val Asp Leu Val Leu Lys 530 535 540 Asp Pro Val Lys Leu Asp Asp 545 550 <210> 9 <211> 567 <212> PRT <213> Synthetic Sequence <220> <223> >SaSSY134 <400> 9 Met Asp Ser Ser Thr Ala Thr Ala Thr Thr Ala Pro Phe Ile Asp Pro 1 5 10 15 Thr Asn His Val Asn Leu Lys Ile Asp Asn Asp Ser Ser Glu Asn Arg 20 25 30 Arg Met Gly Asn Tyr Lys Pro Ser Ile Trp Asn Tyr Asp Phe Leu Gln 35 40 45 Ser Leu Ala Thr His His Asn Ile Val Glu Glu Arg His Leu Lys Leu 50 55 60 Ala Glu Lys Leu Lys Gly Gln Val Arg Ile Leu Leu Lys Glu Lys Met 65 70 75 80 Glu Pro Leu Ala Gln Leu Glu Leu Val Asp Val Val Gln Arg Leu Gly 85 90 95 Leu Asn His Leu Leu Glu Thr Glu Ile Lys Glu Ala Leu Phe Ser Ile 100 105 110 Tyr Lys Asp His Ile Asp Ser Asp Lys Ala Asp Leu His Ala Thr Ser 115 120 125 Leu Arg Phe Arg Leu Leu Arg Gln Gln Gly Ile Lys Ile Ser Cys Asp 130 135 140 Val Phe Glu Gln Phe Lys Asp Asp Glu Asp Arg Phe Lys Ser Ser Leu 145 150 155 160 Ile Asn Asp Ile Gln Gly Met Leu Ser Leu Tyr Glu Ala Ser Phe Leu 165 170 175 Gly Trp Lys Gly Glu Asn Ile Leu Asp Glu Ala Lys Ala Phe Ala Thr 180 185 190 Lys Tyr Leu Lys Ala Met Val Glu Ser Leu Gly Gly His Leu Ala Lys 195 200 205 Arg Val Lys His Ala Leu Ala Leu Pro Leu His Trp Arg Val Pro Arg 210 215 220 Ile Glu Ala Arg Trp Phe Ile Glu Ala Tyr Glu Gln Glu Glu Asn Met 225 230 235 240 Asn Pro Thr Leu Leu Lys Leu Ala Lys Leu Asp Phe Asn Met Val Gln 245 250 255 Ser Ile His Gln Lys Glu Ile Gly Glu Leu Ala Arg Trp Trp Val Thr 260 265 270 Thr Gly Leu Asp Lys Leu Ala Phe Ala Arg Asn Asn Leu Leu Gln Ser 275 280 285 Tyr Met Trp Ser Cys Ala Ile Ala Ser Asp Pro Lys Phe Lys Leu Ala 290 295 300 Arg Glu Thr Ile Val Glu Ile Gly Ser Val Leu Thr Val Val Asp Asp 305 310 315 320 Ala Tyr Asp Val Tyr Gly His Met Asp Glu Leu Asp Leu Tyr Thr Ser 325 330 335 Ser Val Glu Gly Trp Ser Cys Ala Glu Ile Asp Arg Leu Pro Asp Thr 340 345 350 Leu Lys Leu Ile Phe Met Ser Met Phe Asn Lys Thr Asn Glu Val Gly 355 360 365 Leu Arg Val Gln His Glu Arg Gly Tyr Asn Ser Ile Pro Thr Phe Ile 370 375 380 Lys Ala Trp Val Glu Gln Cys Lys Ser Tyr Gln Lys Glu Ala Arg Trp 385 390 395 400 Phe His Gly Gly His Thr Pro Pro Leu Glu Glu Tyr Ser Leu Asn Gly 405 410 415 Leu Val Ser Ile Gly Phe Pro Leu Leu Leu Ile Thr Gly Tyr Ile Ala 420 425 430 Ile Ala Glu Asn Glu Ala Ala Leu Asp Lys Val Arg Pro Leu Pro Asp 435 440 445 Leu Leu His Tyr Ser Ser Leu Leu Ser Arg Leu Ile Asn Asp Met Gly 450 455 460 Thr Ser Pro Asp Glu Met Ala Arg Gly Asp Asn Leu Lys Ser Ile His 465 470 475 480 Cys Tyr Met Asn Glu Thr Gly Ala Ser Glu Glu Val Ala Arg Glu His 485 490 495 Ile Lys Gly Ile Ile Glu Glu Asn Trp Lys Ile Leu Asn Gln Cys Cys 500 505 510 Phe Asp Gln Ser Gln Phe Gln Glu Pro Phe Ile Thr Phe Asn Leu Asn 515 520 525 Ser Val Arg Gly Ser His Phe Phe Tyr Glu Phe Gly Asp Gly Phe Gly 530 535 540 Val Thr Asp Ser Trp Thr Lys Val Asp Met Lys Ser Val Leu Ile Asp 545 550 555 560 Pro Ile Pro Leu Gly Glu Glu 565 <210> 10 <211> 1662 <212> DNA <213> Cinnamomum camphora <220> <223> >CiCaSSy_wt <400> 10 atggacagca tggaagtccg gcggtcggcg atctaccaca gcacgttctg ggacatcgac 60 agcatccggg cgctcctggc gcggcgggac tgcacggcgg ccgcggccct ctcgcccgac 120 caccataagc gcctgaagga gcgcatccag cgccgcctcc aggacatcac ccagccccac 180 catctgctcg gcctcatcga cgccgtgcag cgcctgggcg tggcctacca gttcgaggaa 240 gagatctcgg acgcgctgca cggcctccat tcggagaaca ccgagcacgc catcaaggac 300 tcgctgcacc atacgtcgct ctatttccgc ctgctccgcc agcatggctg caacctgtcg 360 tcggacatct tcaacaagtt caagaaggaa ggcggcggct tcaaggcctc gctctgcgag 420 gacgccatgg gcctgctctc gctgtatgag gccgtgcgcc tctcggtgaa gggcgaggcc 480 atcctggagg aagcccaggt gttctcgatc gccaacctga agatcctcat ggagcgcgtg 540 gagcgcaagc tcgccgaccg catcgagcat gccctggaga tcccgctcta ttggcgcgcc 600 ccgcgtctgg aggcccgctg gtacatcgac gtgtatgaga aggaagacgg ccgcatcgac 660 gacctgctcg acttcgcgaa gctggacttc aaccgcgtgc agatgctcta tcagaccgag 720 ctgaaggagc tctcgatgtg gtgggagctg ctgggcctgc ccgccaagat gggcttcttc 780 cgcgaccgcc tgctcgagaa ccacctcttc tcgatcgccg tggtggtgga gccccagtac 840 tcgcagtgcc gcgtggccat caccaaggcg atcgtgctga tgacggcgat ggacgacttc 900 tatgacgtgc atggcctgcc ggacgagctc aaggtgttca ccgacacggt gaaccgctgg 960 gacctggagg gcatcgacca gctccccgag tacatgaagc tgtactatct ggcgctctac 1020 aacaccacga acgagacggc ctatatcatc ctgaaggaga agggcttcaa cgccacgcat 1080 tacctgaaga agctctgggc catgcagtcg aacgcgtatt tccgcgaggc ccagtggttc 1140 aactcgggct acatcccgaa gttcgacgag tatctggaca acgccctcgt gtcggtgggc 1200 gccccgttcg tgctgggcct ctcgtatccc atgatccagc agcagatctc gaaggaagag 1260 atcgacctga tccccgagga cctcaacctg ctccgctggg cctcgatcat cttccgcctg 1320 tacgacgacc tggccacctc gaaggccgag cagcagcgcg gcgacgtgcc caagtcgatc 1380 cagtgctata tgcatgagac gggctcgtcg gaggaagtgg cggccaacca tatccgcgac 1440 ctgatctcgg acgcgtggaa ggaagtgaac gccgagtgcc tgaagccgac ctcgctctcg 1500 aagcactacg tgggcgtggc ccccaactcg gcccgctcgg gcgtgctcat gtatcaccat 1560 gacttcgacg gcttcgcgtc gccccatggc cgcacgaacg cccacatcac gagcatcttc 1620 ttcgagccgg tccccctcaa ggagagcatc aacctgggct ga 1662 <210> 11 <211> 1662 <212> DNA <213> Artificial sequence <220> <223> >CiCaSSy_N267S <400> 11 atggacagca tggaagtccg gcggtcggcg atctaccaca gcacgttctg ggacatcgac ​​​​​​​​​​​​​​​​​​tcgctgcacc atacgtcgct ctatttccgc ctgctccgcc agcatggctg caacctgtcg 360 tcggacatct tcaacaagtt caagaaggaa ggcggcggct tcaaggcctc gctctgcgag 420 gacgccatgg gcctgctctc gctgtatgag gccgtgcgcc tctcggtgaa gggcgaggcc 480 atcctggagg aagcccaggt gttctcgatc gccaacctga agatcctcat ggagcgcgtg 540 gagcgcaagc tcgccgaccg catcgagcat gccctggaga tcccgctcta ttggcgcgcc 600 ccgcgtctgg aggcccgctg gtacatcgac gtgtatgaga aggaagacgg ccgcatcgac 660 gacctgctcg acttcgcgaa gctggacttc aaccgcgtgc agatgctcta tcagaccgag 720 ctgaaggagc tctcgatgtg gtgggagctg ctgggcctgc ccgccaagat gggcttcttc 780 cgcgaccgcc tgctcgagtc gcacctcttc tcgatcgccg tggtggtgga gccccagtac 840 tcgcagtgcc gcgtggccat caccaaggcg atcgtgctga tgacggcgat ggacgacttc 900 tatgacgtgc atggcctgcc ggacgagctc aaggtgttca ccgacacggt gaaccgctgg 960 gacctggagg gcatcgacca gctccccgag tacatgaagc tgtactatct ggcgctctac 1020 aacaccacga acgagacggc ctatatcatc ctgaaggaga agggcttcaa cgccacgcat 1080 tacctgaaga agctctgggc catgcagtcg aacgcgtatt tccgcgaggc ccagtggttc 1140 aactcgggct acatcccgaa gttcgacgag tatctggaca acgccctcgt gtcggtgggc 1200 gccccgttcg tgctgggcct ctcgtatccc atgatccagc agcagatctc gaaggaagag 1260 atcgacctga tccccgagga cctcaacctg ctccgctggg cctcgatcat cttccgcctg 1320 tacgacgacc tggccacctc gaaggccgag cagcagcgcg gcgacgtgcc caagtcgatc 1380 cagtgctata tgcatgagac gggctcgtcg gaggaagtgg cggccaacca tatccgcgac 1440 ctgatctcgg acgcgtggaa ggaagtgaac gccgagtgcc tgaagccgac ctcgctctcg 1500 aagcactacg tgggcgtggc ccccaactcg gcccgctcgg gcgtgctcat gtatcaccat 1560 gacttcgacg gcttcgcgtc gccccatggc cgcacgaacg cccacatcac gagcatcttc 1620 ttcgagccgg tccccctcaa ggagagcatc aacctgggct ga 1662 <210> 12 <211> 1662 <212> DNA <213> Artificial sequence <220> <223> >CiCaSSy_N267L <400> 12 atggacagca tggaagtccg gcggtcggcg atctaccaca gcacgttctg ggacatcgac 60 agcatccggg cgctcctggc gcggcgggac tgcacggcgg ccgcggccct ctcgcccgac 120 caccataagc gcctgaagga gcgcatccag cgccgcctcc aggacatcac ccagccccac 180 catctgctcg gcctcatcga cgccgtgcag cgcctgggcg tggcctacca gttcgaggaa 240 gagatctcgg acgcgctgca cggcctccat tcggagaaca ccgagcacgc catcaaggac 300 tcgctgcacc atacgtcgct ctatttccgc ctgctccgcc agcatggctg caacctgtcg 360 tcggacatct tcaacaagtt caagaaggaa ggcggcggct tcaaggcctc gctctgcgag 420 gacgccatgg gcctgctctc gctgtatgag gccgtgcgcc tctcggtgaa gggcgaggcc 480 atcctggagg aagcccaggt gttctcgatc gccaacctga agatcctcat ggagcgcgtg 540 gagcgcaagc tcgccgaccg catcgagcat gccctggaga tcccgctcta ttggcgcgcc 600 ccgcgtctgg aggcccgctg gtacatcgac gtgtatgaga aggaagacgg ccgcatcgac 660 gacctgctcg acttcgcgaa gctggacttc aaccgcgtgc agatgctcta tcagaccgag 720 ctgaaggagc tctcgatgtg gtgggagctg ctgggcctgc ccgccaagat gggcttcttc 780 cgcgaccgcc tgctcgagct ccacctcttc tcgatcgccg tggtggtgga gccccagtac 840 tcgcagtgcc gcgtggccat caccaaggcg atcgtgctga tgacggcgat ggacgacttc 900 tatgacgtgc atggcctgcc ggacgagctc aaggtgttca ccgacacggt gaaccgctgg 960 gacctggagg gcatcgacca gctccccgag tacatgaagc tgtactatct ggcgctctac 1020 aacaccacga acgagacggc ctatatcatc ctgaaggaga agggcttcaa cgccacgcat 1080 tacctgaaga agctctgggc catgcagtcg aacgcgtatt tccgcgaggc ccagtggttc 1140 aactcgggct acatcccgaa gttcgacgag tatctggaca acgccctcgt gtcggtgggc 1200 gccccgttcg tgctgggcct ctcgtatccc atgatccagc agcagatctc gaaggaagag 1260 atcgacctga tccccgagga cctcaacctg ctccgctggg cctcgatcat cttccgcctg 1320 tacgacgacc tggccacctc gaaggccgag cagcagcgcg gcgacgtgcc caagtcgatc 1380 cagtgctata tgcatgagac gggctcgtcg gaggaagtgg cggccaacca tatccgcgac 1440 ctgatctcgg acgcgtggaa ggaagtgaac gccgagtgcc tgaagccgac ctcgctctcg 1500 aagcactacg tgggcgtggc ccccaactcg gcccgctcgg gcgtgctcat gtatcaccat 1560 gacttcgacg gcttcgcgtc gccccatggc cgcacgaacg cccacatcac gagcatcttc 1620 ttcgagccgg tccccctcaa ggagagcatc aacctgggct ga 1662 <210> 13 <211> 553 <212> PRT <213> Artificial Sequence <220> <223> >VarS1 <400> 13 Met Asp Ser Met Glu Val Arg Arg Thr Ala Ile Tyr His Ser Thr Phe 1 5 10 15 Trp Asp Ile Asp Ser Ile Arg Ala Leu Ile Ala Arg Arg Asp Cys Thr 20 25 30 Ala Ala Ala Ala Ile Ser Pro Asp His His Lys Arg Leu Lys Glu Arg 35 40 45 Ile Gln Lys Arg Leu Gln Asp Ile Thr Gln Pro His His Leu Leu Gly 50 55 60 Leu Ile Asp Ala Val Gln Arg Leu Gly Val Ala Tyr Gln Phe Glu Glu 65 70 75 80 Glu Ile Ser Asp Ala Leu His Gly Leu His Ser Glu Asn Thr Glu His 85 90 95 Gly Ile Arg Asp Ser Leu His His Thr Ser Leu Tyr Phe Arg Leu Leu 100 105 110 Arg Gln His Gly Cys Asn Leu Ser Ser Asp Ile Phe Asn Lys Phe Lys 115 120 125 Lys Asp Gly Gly Gly Phe Lys Ala Thr Leu Cys Asp Asp Ala Met Gly 130 135 140 Leu Leu Ser Leu Tyr Glu Ala Val Arg Leu Ser Val Lys Gly Glu Ala 145 150 155 160 Ile Leu Glu Glu Ala Gln Val Phe Ser Ile Ala Asn Leu Lys Ile Leu 165 170 175 Leu Asp Arg Leu Asp Arg Lys Leu Ala Glu Arg Ile Glu His Ala Leu 180 185 190 Asp Leu Pro Leu Phe Trp Arg Ala Pro Arg Leu Glu Ala Arg Trp Tyr 195 200 205 Ile Asp Leu Tyr Glu Lys Asp Asp Gly Arg Ile Asp Asp Leu Leu Asp 210 215 220 Phe Ala Lys Leu Asp Phe Asn Arg Val Gln Leu Ile Tyr Gln Thr Glu 225 230 235 240 Leu Lys Glu Leu Ser Met Trp Trp Glu Leu Val Gly Leu Pro Ala Lys 245 250 255 Met Gly Phe Phe Arg Asp Arg Ile Leu Glu Ser His Leu Phe Thr Val 260 265 270 Ala Ile Val Val Glu Pro Gln Tyr Ser Gln Cys Arg Leu Ala Ile Thr 275 280 285 Lys Ala Leu Ile Leu Met Thr Ala Met Asp Asp Phe Tyr Asp Val His 290 295 300 Gly Leu Pro Asp Glu Leu Lys Leu Phe Thr Asp Thr Val Gln Arg Trp 305 310 315 320 Asp Ile Glu Ala Met Glu Gln Leu Pro Glu Tyr Met Lys Ile Tyr Tyr 325 330 335 Leu Ala Leu Tyr Asn Thr Thr Asn Glu Thr Ala Tyr Val Ile Leu Lys 340 345 350 Glu Lys Gly Phe Asn Ala Thr His Tyr Leu Lys Lys Val Trp Ala Met 355 360 365 Gln Ser Asn Ala Tyr Phe Arg Glu Ala Gln Trp Phe Asn Ser Gly Tyr 370 375 380 Val Pro Lys Phe Asp Glu Tyr Leu Asp Asn Ala Leu Val Ser Val Gly 385 390 395 400 Ala Pro Phe Val Leu Gly Val Thr Tyr Pro Val Ile Gln Gln Gln Leu 405 410 415 Ser Lys Glu Glu Leu Asp Leu Val Pro Glu Asp Leu Asn Leu Val Lys 420 425 430 Trp Ala Ser Ile Ile Phe Arg Leu Tyr Asp Asp Leu Ala Thr Ser Lys 435 440 445 Ala Glu Gln Gln Arg Gly Asp Val Pro Lys Ser Ile Gln Cys Tyr Met 450 455 460 His Glu Thr Gly Ser Ser Glu Glu Val Ala Ala Asn His Ile Arg Asp 465 470 475 480 Met Ile Ser Asp Ala Trp Lys Asp Val Asn Ala Glu Cys Leu Lys Pro 485 490 495 Thr Ser Leu Ser Lys His Tyr Leu Gly Leu Ala Pro Asn Ser Ala Arg 500 505 510 Ser Gly Val Leu Leu Tyr His His Asp Phe Asp Gly Phe Ala Thr Pro 515 520 525 His Ala Arg Thr Asn Gly His Ile Thr Ser Ile Phe Phe Asp Pro Val 530 535 540 Pro Leu Lys Glu Ser Ile Asn Leu Gly 545 550 <210> 14 <211> 553 <212> PRT <213> Artificial Sequence <220> <223> >VarS2 <400> 14 Met Asp Ser Val Asp Met Arg Lys Ser Ala Ile Tyr His Ser Thr Phe 1 5 10 15 Trp Asp Ile Asp Ser Ile Arg Ala Leu Leu Ala Arg Arg Asp Cys Thr 20 25 30 Ala Ala Ala Ala Leu Ser Pro Glu His His Lys Arg Leu Lys Glu Arg 35 40 45 Ile Gln Lys Arg Leu Asn Glu Val Thr Gln Pro His His Leu Leu Gly 50 55 60 Leu Ile Asp Ala Val Gln Arg Leu Gly Val Ala Tyr Gln Phe Glu Glu 65 70 75 80 Glu Ile Ser Glu Ala Leu His Gly Ile His Ser Glu Asn Thr Glu His 85 90 95 Gly Val Arg Glu Ser Leu His His Thr Ser Leu Tyr Phe Arg Leu Leu 100 105 110 Arg Gln His Gly Cys Asn Leu Ser Ser Asp Ile Phe Asn Arg Phe Lys 115 120 125 Lys Asp Gly Gly Gly Phe Lys Gly Ser Leu Cys Glu Asp Ala Met Gly 130 135 140 Leu Leu Ser Leu Tyr Glu Ala Val Arg Leu Ser Leu Lys Gly Glu Ala 145 150 155 160 Ile Leu Glu Glu Ala Gln Val Phe Ser Ile Ala Gln Leu Arg Ile Leu 165 170 175 Met Glu Arg Val Glu Lys Lys Leu Ala Glu Arg Ile Glu His Ala Leu 180 185 190 Asp Ile Pro Leu Tyr Trp Arg Ala Pro Arg Leu Glu Ala Arg Trp Tyr 195 200 205 Ile Glu Val Tyr Glu Lys Glu Glu Gly Lys Ile Glu Asp Leu Leu Asp 210 215 220 Phe Ala Lys Leu Asp Phe Asn Arg Val Gln Met Leu Tyr Gln Ser Glu 225 230 235 240 Leu Lys Glu Leu Ser Met Trp Trp Glu Leu Met Gly Leu Pro Ala Lys 245 250 255 Met Gly Phe Phe Arg Asp Arg Leu Leu Glu Ser His Leu Phe Ser Met 260 265 270 Ala Val Val Val Glu Pro Gln Tyr Ser Gln Cys Arg Val Ala Ile Thr 275 280 285 Lys Ala Ile Val Leu Met Ser Ala Met Asp Asp Phe Tyr Asp Val His 290 295 300 Gly Leu Pro Asp Glu Leu Lys Val Phe Thr Asp Thr Val Gln Arg Trp 305 310 315 320 Asp Val Glu Gly Met Asp Gln Leu Pro Asp Tyr Met Lys Leu Tyr Tyr 325 330 335 Leu Ala Leu Tyr Asn Thr Thr Asn Glu Thr Ala Tyr Val Ile Leu Arg 340 345 350 Glu Lys Gly Phe Asn Ala Thr His Tyr Leu Lys Lys Leu Trp Ala Met 355 360 365 Gln Ser Asn Ala Tyr Phe Lys Glu Ala Gln Trp Phe Asn Ser Gly Tyr 370 375 380 Val Pro Lys Phe Asp Glu Tyr Leu Asp Asn Ala Leu Val Ser Ile Gly 385 390 395 400 Gly Pro Phe Val Leu Ala Leu Ser Tyr Pro Leu Leu Gln Gln Gln Ile 405 410 415 Ser Lys Glu Glu Ile Glu Leu Ile Pro Glu Asp Leu Asn Leu Ile Arg 420 425 430 Trp Ala Ser Ile Ile Phe Arg Leu Tyr Asp Asp Leu Ala Thr Thr Lys 435 440 445 Ala Glu Gln Asn Arg Gly Asp Val Pro Lys Ser Ile Gln Cys Tyr Met 450 455 460 His Glu Thr Gly Ser Ser Glu Glu Val Ala Ala Asn His Ile Arg Asp 465 470 475 480 Leu Ile Ser Asp Ala Trp Lys Asp Val Asn Ala Glu Cys Leu Lys Pro 485 490 495 Ser Ser Leu Ser Lys His Tyr Leu Gly Val Ala Pro Asn Thr Ala Arg 500 505 510 Thr Gly Val Leu Met Tyr His His Asp Phe Asp Gly Phe Ala Ser Pro 515 520 525 His Gly Lys Thr Asn Ala His Ile Thr Ser Ile Phe Phe Glu Pro Val 530 535 540 Pro Met Lys Glu Ser Ile Asn Leu Gly 545 550 <210> \15 <211> \553 <212> \PRT <213> \Artificial Sequence <220> <223> \>VarS3 <400> \15 Met Asp Ser Met Glu Val Arg Lys Ser Ala Met Tyr His Ser Thr Phe 1 5 10 15 Trp Asp Ile Asp Ser Ile Arg Ala Ile Leu Ala Lys Arg Asp Cys Thr 20 25 30 Ala Ala Ala Ala Leu Ser Pro Glu His His Lys Arg Leu Lys Glu Arg 35 40 45 Ile Gln Lys Arg Leu Asn Asp Val Thr Asn Pro His His Leu Leu Ala 50 55 60 Leu Ile Asp Ala Val Gln Arg Leu Gly Val Ala Tyr Gln Phe Glu Glu 65 70 75 80 Glu Ile Ser Asp Ala Leu His Gly Leu His Ser Glu Gln Ser Asp His 85 90 95 Gly Leu Lys Asp Ser Leu His His Thr Ser Leu Tyr Phe Arg Leu Leu 100 105 110 Arg Gln His Gly Cys Asn Leu Ser Ser Asp Ile Phe Asn Lys Phe Lys 115 120 125 Lys Glu Gly Gly Gly Phe Lys Ala Thr Leu Cys Glu Asp Ala Met Gly 130 135 140 Leu Leu Ser Leu Tyr Glu Ala Val Arg Leu Ser Ile Arg Gly Glu Ala 145 150 155 160 Ile Leu Glu Glu Ala Gln Ile Phe Ser Ile Ala Asn Leu Lys Ile Ile 165 170 175 Ile Glu Arg Leu Glu Arg Lys Leu Ala Glu Lys Ile Glu His Ala Leu 180 185 190 Glu Ile Pro Leu Tyr Trp Arg Ala Pro Arg Leu Glu Ala Arg Trp Tyr 195 200 205 Ile Glu Ile Tyr Glu Lys Asp Asp Gly Lys Ile Glu Asp Leu Leu Asp 210 215 220 Phe Ala Lys Leu Asp Phe Asn Arg Met Gln Met Leu Tyr Gln Thr Glu 225 230 235 240 Leu Lys Asp Leu Thr Met Trp Trp Glu Leu Leu Gly Leu Pro Ala Lys 245 250 255 Met Gly Phe Phe Arg Asp Arg Leu Leu Glu Thr His Leu Phe Ser Ile 260 265 270 Ala Met Ile Val Glu Pro Gln Tyr Ser Asn Cys Arg Ile Gly Ile Thr 275 280 285 Lys Ala Ile Val Leu Val Ser Ala Met Asp Asp Phe Tyr Asp Val His 290 295 300 Gly Leu Pro Asp Glu Leu Arg Val Phe Thr Glu Thr Val Asn Arg Trp 305 310 315 320 Asp Leu Glu Gly Ile Asp Gln Leu Pro Asp Tyr Met Lys Val Tyr Tyr 325 330 335 Leu Gly Leu Tyr Asn Thr Thr Asn Glu Thr Ala Tyr Leu Val Val Lys 340 345 350 Asp Lys Gly Tyr Asn Ala Thr His Tyr Leu Arg Lys Val Trp Ala Met 355 360 365 Gln Ser Asn Ala Tyr Phe Arg Glu Ala Gln Trp Phe Asn Thr Gly Tyr 370 375 380 Ile Pro Lys Phe Asp Glu Tyr Leu Asp Asn Ala Val Val Ser Val Gly 385 390 395 400 Ala Pro Phe Val Val Gly Leu Thr Tyr Pro Leu Leu Gln Gln Gln Leu 405 410 415 Thr Lys Glu Glu Ile Glu Leu Ile Pro Glu Asp Leu Asn Leu Leu Arg 420 425 430 Trp Ala Ser Val Ile Phe Arg Leu Tyr Asp Asp Leu Ala Thr Thr Lys 435 440 445 Ala Glu Gln Gln Arg Gly Asp Met Pro Lys Ser Ile Gln Cys Tyr Met 450 455 460 His Glu Thr Gly Thr Ser Glu Glu Val Ala Ala Asn His Ile Lys Glu 465 470 475 480 Ile Ile Ser Glu Ala Trp Lys Asp Val Asn Ala Glu Cys Leu Lys Pro 485 490 495 Ser Ser Val Ser Lys His Tyr Leu Gly Leu Ala Pro Asn Ser Ala Arg 500 505 510 Ser Ala Met Val Ile Tyr His His Asp Phe Asp Gly Phe Ala Thr Pro 515 520 525 His Gly Lys Thr Asn Ala His Ile Thr Ser Ile Phe Phe Asp Pro Val 530 535 540 Pro Ile Lys Asp Ser Ile Asn Leu Gly 545 550 <210> 16 <211> 553 <212> PRT <213> Artificial sequence <220> <223> >VarS4 <400> 16 Met Asp Ser Val Glu Ile Arg Arg Ser Ala Ile Tyr His Ser Thr Phe 1 5 10 15 Trp Asp Ile Asp Ser Ile Arg Ala Leu Ile Ala Arg Arg Asp Cys Ser 20 25 30 Gly Gly Gly Ala Leu Ser Pro Asp His His Lys Arg Leu Lys Glu Lys 35 40 45 Ile Gln Arg Arg Leu Asn Asp Val Thr Asn Pro His His Leu Leu Gly 50 55 60 Leu Ile Asp Ala Val Gln Arg Leu Gly Val Ala Tyr Gln Phe Glu Glu 65 70 75 80 Glu Ile Ser Glu Ala Leu His Gly Met His Ser Glu Asn Ser Asp His 85 90 95 Gly Val Arg Glu Ser Leu His His Thr Ser Leu Tyr Phe Arg Leu Leu 100 105 110 Arg Gln His Gly Cys Asn Leu Ser Ser Asp Ile Phe Gln Arg Phe Arg 115 120 125 Lys Asp Gly Gly Gly Phe Arg Ala Ser Ile Cys Asp Asp Ala Met Gly 130 135 140 Leu Leu Ser Leu Tyr Glu Ala Val Arg Leu Ser Ile Arg Gly Glu Ala 145 150 155 160 Ile Leu Glu Glu Ala Gln Leu Phe Ser Ile Ala Asn Leu Lys Leu Val 165 170 175 Val Asp Lys Val Glu Lys Lys Leu Ala Asp Arg Ile Glu His Ala Leu 180 185 190 Asp Leu Pro Ile Tyr Trp Arg Ala Pro Arg Leu Glu Ala Arg Trp Tyr 195 200 205 Ile Asp Ile Tyr Glu Arg Asp Glu Ala Arg Ile Glu Asp Leu Leu Asp 210 215 220 Phe Ala Lys Leu Asp Phe Asn Arg Val Gln Met Leu Tyr Gln Thr Glu 225 230 235 240 Leu Lys Asp Leu Thr Met Trp Trp Asp Leu Leu Gly Leu Pro Ala Lys 245 250 255 Met Gly Phe Phe Arg Asp Arg Leu Leu Glu Thr His Leu Phe Ser Leu 260 265 270 Ala Leu Ile Val Glu Pro Gln Phe Ser Asn Cys Arg Met Ala Met Thr 275 280 285 Lys Ala Val Ile Leu Ile Ser Ala Met Asp Asp Phe Tyr Asp Ile His 290 295 300 Gly Leu Pro Asp Glu Leu Lys Ile Phe Thr Glu Thr Val Asn Arg Trp 305 310 315 320 Asp Val Glu Gly Val Asp Gln Leu Pro Glu Tyr Met Lys Ile Tyr Tyr 325 330 335 Val Ala Leu Phe Asn Thr Thr Asn Glu Thr Ala Tyr Leu Met Leu Lys 340 345 350 Asp Lys Gly Tyr Asn Ala Thr His Phe Leu Lys Lys Ile Trp Ala Ile 355 360 365 Gln Ser Asn Ala Tyr Phe Arg Glu Ala Gln Trp Tyr Asn Ser Gly Tyr 370 375 380 Ile Pro Lys Phe Asp Glu Tyr Leu Glu Asn Ala Val Ile Ser Val Gly 385 390 395 400 Gly Pro Phe Leu Val Gly Leu Ser Tyr Pro Val Met Gln Asn Asn Ile 405 410 415 Thr Lys Glu Asp Ile Asp Leu Val Pro Glu Asp Leu Asn Leu Leu Lys 420 425 430 Trp Ala Ser Met Ile Phe Arg Leu Trp Asp Asp Leu Gly Thr Ser Lys 435 440 445 Ala Glu Gln Gln Arg Gly Glu Val Pro Lys Ser Ile Gln Cys Tyr Met 450 455 460 His Glu Thr Gly Ser Ser Glu Glu Val Ala Ala Asn His Ile Lys Glu 465 470 475 480 Met Ile Thr Asp Ala Trp Lys Asp Val Asn Gly Glu Cys Leu Lys Pro 485 490 495 Thr Thr Leu Ser Arg His Tyr Met Gly Ile Ala Pro Asn Ser Ala Arg 500 505 510 Thr Ala Val Val Val Tyr His His Asp Phe Asp Gly Phe Ala Thr Pro 515 520 525 His Gly Arg Ser Asn Ala His Ile Thr Ser Ile Phe Phe Asp Pro Val 530 535 540 Pro Ile Lys Glu Ser Ile Asn Leu Gly 545 550 <210> 17 <211> 553 <212> PRT <213> Artificial Sequence <220> <223> >VarS5 <400> 17 Met Asp Ser Ile Asp Ile Arg Arg Thr Ala Ile Tyr His Ser Thr Phe 1 5 10 15 Trp Asp Ile Asp Ser Ile Arg Ala Leu Leu Ala Lys Arg Glu Cys Ser 20 25 30 Ala Ala Gly Ala Met Ser Pro Asp His His Lys Lys Leu Lys Glu Lys 35 40 45 Ile Gln Lys Lys Leu Gln Glu Leu Thr Gln Pro His His Leu Leu Ala 50 55 60 Leu Ile Asp Ala Val Gln Arg Leu Gly Val Ala Tyr Gln Phe Glu Asp 65 70 75 80 Glu Ile Ser Glu Ala Leu His Gly Ile His Ser Asp Asn Ser Glu His 85 90 95 Ala Ile Arg Glu Ser Leu His His Thr Ser Leu Tyr Phe Arg Leu Leu 100 105 110 Arg Gln His Gly Cys Asn Leu Ser Thr Asp Ile Phe Asn Lys Phe Arg 115 120 125 Lys Asp Gly Gly Gly Phe Lys Ala Thr Leu Cys Glu Asp Ala Met Gly 130 135 140 Leu Leu Ser Leu Tyr Glu Ala Val Arg Leu Ser Val Arg Gly Glu Ala 145 150 155 160 Ile Leu Glu Glu Ala Gln Ile Phe Ser Ile Ala Asn Leu Arg Ile Val 165 170 175 Ile Glu Lys Ile Glu Arg Lys Leu Ala Asp Lys Ile Glu His Ala Leu 180 185 190 Glu Leu Pro Leu Tyr Trp Arg Ala Pro Arg Leu Glu Ala Arg Trp Tyr 195 200 205 Ile Asp Ile Tyr Glu Lys Asp Glu Gly Arg Ile Glu Asp Leu Leu Asp 210 215 220 Phe Ala Lys Leu Asp Phe Asn Arg Val Gln Leu Ile Tyr Gln Thr Glu 225 230 235 240 Leu Arg Glu Leu Thr Met Trp Trp Glu Leu Leu Gly Leu Pro Ala Lys 245 250 255 Met Gly Phe Phe Arg Asp Arg Leu Leu Glu Ser His Leu Phe Ser Val 260 265 270 Ala Met Leu Val Glu Pro Asn Tyr Ser Gln Cys Arg Val Ala Leu Thr 275 280 285 Lys Ala Val Val Leu Val Ser Ala Met Asp Asp Phe Tyr Asp Ile His 290 295 300 Gly Leu Pro Asp Glu Leu Arg Leu Phe Thr Glu Thr Val Gln Arg Trp 305 310 315 320 Asp Leu Asp Ala Leu Asp Gln Leu Pro Glu Tyr Met Lys Ile Phe Tyr 325 330 335 Met Ala Leu Tyr Asn Thr Thr Asn Glu Thr Ala Tyr Leu Met Leu Lys 340 345 350 Asp Lys Gly Phe Asn Ala Ser His Tyr Leu Lys Lys Val Trp Ala Leu 355 360 365 Gln Ser Asn Gly Tyr Phe Arg Glu Ala Gln Trp Tyr Asn Ser Gly Tyr 370 375 380 Val Pro Lys Phe Asp Glu Tyr Leu Asp Asn Ala Val Ile Ser Val Gly 385 390 395 400 Gly Pro Phe Leu Leu Gly Val Ser Tyr Pro Leu Leu Gln Asn Gln Leu 405 410 415 Thr Lys Glu Glu Met Glu Leu Leu Pro Glu Asp Leu Asn Leu Val Lys 420 425 430 Phe Ala Ser Val Ile Phe Arg Leu Tyr Asp Asp Leu Ala Thr Ser Lys 435 440 445 Ala Glu Gln Asn Arg Gly Asp Leu Pro Lys Ser Ile Gln Cys Tyr Met 450 455 460 His Glu Thr Gly Ser Ser Glu Glu Met Ala Ala Gln His Ile Lys Asp 465 470 475 480 Ile Ile Ser Glu Gly Trp Lys Asp Val Asn Ala Glu Cys Leu Arg Pro 485 490 495 Ser Ser Val Ser Arg His Phe Leu Gly Leu Ala Pro Asn Ser Ala Arg 500 505 510 Ser Gly Val Ile Val Tyr His His Asp Phe Asp Gly Phe Ala Ser Pro 515 520 525 His Ala Arg Ser Asn Ala His Ile Thr Ser Ile Phe Phe Glu Pro Val 530 535 540 Pro Ile Arg Glu Ser Ile Asn Leu Gly 545 550 <210> 18 <211> 553 <212> PRT <213> Synthetic Sequence <220> <223> >VarS6 <400> 18 Met Asp Ser Met Glu Ile Arg Arg Thr Gly Ile Tyr His Ser Thr Phe 1 5 10 15 Trp Asp Ile Glu Ser Ile Lys Ala Leu Leu Ala Lys Arg Asp Cys Ser 20 25 30 Ala Ala Gly Ala Val Ser Pro Glu His His Lys Lys Leu Lys Glu Arg 35 40 45 Ile Gln Lys Lys Leu Gln Asp Val Thr Gln Pro His His Val Leu Gly 50 55 60 Leu Ile Asp Ala Val Gln Arg Leu Gly Val Ala Tyr Gln Phe Asp Glu 65 70 75 80 Glu Ile Thr Glu Ala Leu His Gly Ile His Ser Glu Asn Thr Asp His 85 90 95 Gly Val Arg Asp Ser Leu His His Thr Ser Leu Tyr Phe Arg Leu Leu 100 105 110 Arg Gln His Gly Cys Asn Leu Ser Ser Asp Ile Phe Asn Lys Phe Lys 115 120 125 Lys Asp Gly Gly Gly Phe Arg Ala Thr Leu Cys Asp Asp Ala Met Gly 130 135 140 Leu Leu Ser Leu Tyr Glu Ala Val Arg Leu Ser Val Arg Gly Glu Ala 145 150 155 160 Ile Leu Glu Glu Ala Gln Ile Phe Ser Ile Ala Asn Leu Arg Ile Met 165 170 175 Leu Glu Arg Met Asp Lys Lys Leu Ala Glu Arg Ile Glu His Ala Leu 180 185 190 Glu Val Pro Ile Phe Phe Arg Ala Pro Arg Leu Glu Ala Arg Phe Tyr 195 200 205 Ile Glu Ile Tyr Glu Arg Asp Glu Gly Arg Met Glu Asp Leu Leu Asp 210 215 220 Phe Ala Lys Leu Asp Phe Asn Arg Leu Gln Val Ile Tyr Gln Glu Glu 225 230 235 240 Leu Lys Asp Leu Thr Met Trp Trp Glu Leu Leu Gly Leu Pro Ala Lys 245 250 255 Met Gly Phe Phe Arg Asp Arg Val Leu Glu Thr His Leu Phe Thr Val 260 265 270 Ala Leu Met Val Glu Pro Asn Phe Ser Asn Cys Arg Ile Gly Leu Thr 275 280 285 Lys Ala Leu Ile Leu Ile Thr Ala Met Asp Asp Phe Tyr Asp Ile His 290 295 300 Gly Leu Pro Asp Glu Leu Lys Leu Phe Thr Glu Thr Val Gln Arg Trp 305 310 315 320 Asp Ile Glu Ala Val Asp Asn Leu Pro Asp Tyr Met Lys Ile Tyr Tyr 325 330 335 Leu Ala Leu Tyr Asn Thr Thr Gln Glu Thr Ala Tyr Ile Leu Leu Lys 340 345 350 Asp Lys Gly Trp Asn Ala Thr His Phe Leu Lys Lys Ile Trp Ala Ile 355 360 365 Gln Ser Gln Gly Tyr Tyr Lys Glu Ala Gln Trp Tyr Gln Thr Gly Tyr 370 375 380 Leu Pro Lys Phe Asp Glu Tyr Leu Glu Asn Ala Val Ile Ser Ile Gly 385 390 395 400 Gly Pro Phe Met Leu Ala Leu Ser Tyr Pro Leu Leu Gln Gln Gln Leu 405 410 415 Thr Arg Glu Glu Val Asp Leu Val Pro Glu Asp Leu Asn Leu Ile Lys 420 425 430 Trp Gly Ser Leu Ile Phe Arg Leu Tyr Asp Asp Leu Gly Thr Ser Lys 435 440 445 Ala Glu Asn Asn Arg Gly Glu Met Pro Lys Ser Ile Gln Cys Tyr Met 450 455 460 His Glu Thr Gly Ser Ser Glu Glu Val Ala Ala Asn His Ile Lys Asp 465 470 475 480 Leu Ile Ser Glu Gly Trp Lys Asp Val Asn Ala Glu Cys Leu Arg Pro 485 490 495 Thr Thr Val Ser Lys His Phe Leu Ala Leu Ala Pro Asn Thr Ala Arg 500 505 510 Ser Ala Val Leu Leu Tyr His His Asp Phe Asp Gly Phe Ala Ser Pro 515 520 525 His Ala Lys Thr Asn Glu His Ile Ser Ser Ile Phe Tyr Asp Pro Val 530 535 540 Pro Ile Lys Asp Ser Ile Asn Leu Gly 545 550 <210> 19 <211> 553 <212> PRT <213> Artificial Sequence <220> <223> >VarS7 <400> 19 Met Asp Ser Val Glu Leu Arg Arg Thr Ala Ile Tyr His Ser Thr Phe 1 5 10 15 Trp Asp Val Asp Ser Ile Arg Ala Leu Leu Ala Lys Lys Asp Cys Ser 20 25 30 Ala Ala Gly Ala Ile Ser Pro Glu His His Lys Lys Leu Lys Glu Arg 35 40 45 Ile Gln Lys Arg Leu Asn Glu Leu Thr Asn Pro His His Leu Leu Ala 50 55 60 Leu Ile Asp Ala Val Gln Arg Leu Gly Val Ala Tyr Gln Phe Glu Asp 65 70 75 80 Glu Ile Ser Glu Ala Leu His Gly Ile His Ser Glu Gln Thr Asp His 85 90 95 Gly Val Arg Glu Ser Leu His His Thr Ser Leu Trp Phe Arg Leu Leu 100 105 110 Arg Gln His Gly Cys Asn Leu Ser Ser Asp Ile Phe Asn Lys Phe Arg 115 120 125 Lys Asp Gly Gly Gly Phe Arg Ala Ser Val Cys Glu Asp Ala Val Gly 130 135 140 Leu Leu Ser Leu Tyr Glu Ala Val Arg Leu Ser Met Lys Gly Glu Ala 145 150 155 160 Ile Leu Glu Glu Ala Asn Leu Phe Ser Val Ala Asn Leu Arg Leu Met 165 170 175 Leu Asp Lys Ile Asp Lys Lys Leu Ala Glu Lys Ile Glu His Ala Leu 180 185 190 Asp Leu Pro Leu Phe Trp Arg Ala Pro Arg Leu Glu Ala Arg Tyr Tyr 195 200 205 Ile Glu Ile Tyr Glu Arg Asp Glu Ala Arg Ile Glu Glu Leu Leu Asp 210 215 220 Phe Ala Lys Leu Asp Phe Asn Lys Leu Gln Ile Ile Tyr Gln Thr Glu 225 230 235 240 Leu Arg Asp Leu Thr Met Trp Trp Glu Leu Leu Gly Leu Pro Ala Arg 245 250 255 Met Gly Phe Phe Arg Asp Arg Val Leu Glu Ser His Leu Phe Thr Leu 260 265 270 Ala Leu Ile Val Glu Pro Gln Phe Ser Gln Cys Arg Leu Gly Leu Thr 275 280 285 Lys Ala Leu Ile Leu Val Ser Ala Met Asp Asp Phe Tyr Asp Val His 290 295 300 Gly Leu Pro Asp Glu Leu Arg Leu Phe Thr Glu Thr Val Gln Arg Trp 305 310 315 320 Asp Ile Glu Ala Leu Glu Gln Leu Pro Glu Tyr Met Lys Ile Phe Tyr 325 330 335 Ile Gly Leu Tyr Asn Thr Thr Asn Glu Thr Ala Phe Met Val Met Arg 340 345 350 Glu Lys Gly Tyr Asn Ala Thr His Tyr Leu Arg Lys Ile Trp Gly Met 355 360 365 Gln Ser Asn Gly Tyr Phe Lys Glu Ala Gln Trp Tyr Gln Thr Gly Tyr 370 375 380 Val Pro Lys Phe Asp Glu Tyr Leu Glu Asn Ala Val Ile Ser Ile Gly 385 390 395 400 Gly Pro Trp Leu Leu Gly Leu Ser Tyr Pro Leu Met Gln Asn Gln Leu 405 410 415 Thr Lys Glu Asp Leu Glu Ile Val Pro Asp His Leu Gln Leu Val Lys 420 425 430 Trp Gly Ser Val Ile Phe Arg Leu Trp Asp Asp Leu Gly Thr Ser Lys 435 440 445 Ala Glu Asn Asn Arg Gly Glu Met Pro Lys Thr Ile Gln Cys Tyr Met 450 455 460 His Glu Thr Gly Thr Ser Glu Glu Val Ala Ala Asn His Ile Lys Glu 465 470 475 480 Met Ile Ser Glu Gly Trp Lys Asp Val Asn Gly Glu Cys Leu Arg Pro 485 490 495 Ser Thr Val Ser Arg His Tyr Leu Ala Met Ala Pro Asn Ser Ala Arg 500 505 510 Ser Ala Val Val Ile Tyr His His Glu Phe Asp Gly Phe Ala Ser Pro 515 520 525 His Gly Lys Ser Asn Gly His Ile Ser Ser Ile Phe Phe Asp Pro Val 530 535 540 Pro Val Arg Asp Ser Ile Asn Leu Gly 545 550 <210> 20 <211> 553 <212> PRT <213> Artificial Sequence <220> <223> >VarS8 <400> 20 Met Asp Ser Val Glu Val Arg Lys Thr Gly Met Tyr His Ser Thr Phe 1 5 10 15 Trp Asp Val Asp Ser Ile Lys Ala Ile Leu Ala Lys Lys Glu Cys Ser 20 25 30 Ala Gly Gly Ala Met Thr Pro Glu His His Lys Lys Leu Lys Asp Arg 35 40 45 Ile Gln Lys Arg Leu Asn Glu Val Ser Gln Pro His His Leu Leu Ala 50 55 60 Leu Ile Asp Ala Val Gln Arg Leu Gly Val Ala Tyr Gln Phe Glu Asp 65 70 75 80 Glu Ile Ser Glu Ala Leu His Gly Val His Ser Asp Asn Ser Asp His 85 90 95 Gly Val Lys Glu Ser Leu His His Thr Ser Leu Trp Phe Arg Leu Leu 100 105 110 Arg Gln His Gly Cys Gln Leu Ser Ser Asp Ile Phe Gln Arg Phe Lys 115 120 125 Lys Asp Gly Gly Gly Phe Lys Ala Ser Met Cys Glu Asp Ala Leu Gly 130 135 140 Leu Leu Ser Leu Tyr Glu Ala Val Arg Leu Ser Ile Lys Gly Glu Gly 145 150 155 160 Ile Leu Glu Glu Ala Gln Val Phe Ser Ile Gly Asn Leu Arg Ile Met 165 170 175 Leu Asp Lys Ile Asp Lys Lys Leu Ala Glu Lys Ile Glu His Ala Leu 180 185 190 Asp Val Pro Val Tyr Trp Lys Ala Pro Arg Leu Glu Ala Arg Phe Tyr 195 200 205 Ile Asp Ile Tyr Glu Lys Asp Glu Ala Arg Ile Glu Asp Leu Leu Asp 210 215 220 Phe Ala Lys Leu Asp Phe Asn Arg Val Gln Met Leu Tyr Gln Thr Glu 225 230 235 240 Leu Lys Asp Leu Thr Met Trp Trp Asp Leu Ile Gly Leu Pro Ala Lys 245 250 255 Met Gly Phe Phe Arg Asp Arg Met Leu Glu Ser His Leu Phe Thr Met 260 265 270 Ala Ile Ile Met Glu Pro Asn Tyr Ser Asn Cys Arg Ile Ala Leu Thr 275 280 285 Lys Ala Leu Val Leu Val Ser Ala Met Asp Asp Phe Tyr Asp Val His 290 295 300 Gly Leu Pro Asp Glu Leu Arg Leu Phe Thr Glu Thr Val Gln Arg Trp 305 310 315 320 Asp Ile Asp Ala Met Glu Asn Leu Pro Asp Tyr Met Lys Ile Phe Tyr 325 330 335 Met Ala Leu Phe Asn Thr Ser Asn Glu Thr Ala Tyr Ile Leu Val Lys 340 345 350 Asp Lys Gly Trp Asn Ala Thr His Tyr Leu Arg Lys Val Trp Ala Ile 355 360 365 Gln Ser Asn Ala Tyr Tyr Lys Glu Ala Gln Trp Tyr Asn Ser Gly Tyr 370 375 380 Val Pro Lys Phe Asp Glu Tyr Leu Glu Asn Ala Ile Ile Ser Ile Gly 385 390 395 400 Gly Pro Tyr Leu Met Ala Val Ser Tyr Pro Val Met Gln Asn Gln Val 405 410 415 Ser Lys Glu Asp Val Glu Leu Leu Pro Glu Glu Leu Asn Leu Val Lys 420 425 430 Tyr Ala Thr Met Ile Phe Arg Leu Trp Asp Asp Leu Gly Thr Thr Lys 435 440 445 Ala Glu Asn Asn Arg Gly Asp Val Pro Lys Thr Ile Gln Cys Tyr Met 450 455 460 His Glu Ser Gly Thr Ser Glu Glu Ile Ala Ala Asn His Ile Lys Glu 465 470 475 480 Met Ile Ser Glu Ala Trp Lys Asp Val Asn Ala Asp Cys Leu Arg Pro 485 490 495 Ser Thr Leu Ser Arg His Phe Leu Gly Leu Ala Pro Asn Ser Ala Arg 500 505 510 Thr Ala Val Val Val Tyr His His Asp Phe Asp Gly Phe Ala Ser Pro 515 520 525 His Ala Arg Ser Asn Ala His Ile Ser Thr Ile Phe Tyr Asp Pro Val 530 535 540 Pro Leu Arg Asp Ser Ile Asn Leu Gly 545 550 <210> 21 <211> 553 <212> PRT <213> artificial sequence <220> <223> >VarL1 <400> 21 Met Asp Ser Met Glu Ile Arg Arg Thr Ala Ile Tyr His Ser Thr Phe 1 5 10 15 Trp Asp Ile Asp Ser Ile Arg Ala Met Leu Ala Arg Arg Asp Cys Ser 20 25 30 Ala Gly Ala Ala Val Ser Pro Asp His His Lys Lys Leu Lys Glu Lys 35 40 45 Ile Gln Arg Arg Leu Gln Asp Ile Thr Gln Pro His His Leu Leu Gly 50 55 60 Leu Ile Asp Ala Val Gln Arg Leu Gly Val Ala Tyr Gln Phe Glu Asp 65 70 75 80 Glu Ile Ser Glu Ala Leu His Gly Val His Ser Glu Asn Thr Glu His 85 90 95 Ala Ile Lys Asp Ser Leu His His Thr Ser Leu Tyr Phe Arg Leu Leu 100 105 110 Arg Gln His Gly Cys Asn Leu Ser Ser Asp Ile Phe Asn Lys Phe Lys 115 120 125 Lys Glu Gly Gly Gly Phe Lys Ala Thr Leu Cys Glu Asp Ala Met Gly 130 135 140 Leu Leu Ser Leu Tyr Glu Ala Val Arg Leu Ser Ile Lys Gly Glu Ala 145 150 155 160 Ile Leu Glu Glu Ala Gln Val Phe Ser Ile Ala Asn Leu Arg Ile Met 165 170 175 Val Glu Lys Leu Glu Arg Lys Leu Ala Asp Arg Ile Glu His Ala Leu 180 185 190 Glu Ile Pro Leu Tyr Trp Arg Ala Pro Arg Leu Glu Ala Arg Phe Tyr 195 200 205 Ile Glu Ile Tyr Glu Lys Glu Asp Gly Lys Ile Asp Glu Leu Leu Asp 210 215 220 Phe Ala Lys Leu Asp Phe Asn Arg Val Gln Met Ile Tyr Gln Thr Glu 225 230 235 240 Leu Lys Glu Leu Thr Met Trp Trp Glu Leu Leu Gly Leu Pro Ala Lys 245 250 255 Met Gly Phe Phe Arg Asp Arg Leu Leu Glu Leu His Leu Phe Ser Met 260 265 270 Ala Val Ile Val Glu Pro Gln Phe Ser Gln Cys Arg Leu Ala Leu Thr 275 280 285 Lys Ala Ile Val Leu Met Thr Ala Met Asp Asp Phe Tyr Asp Val His 290 295 300 Gly Leu Pro Asp Glu Leu Lys Val Phe Thr Asp Thr Val Asn Arg Trp 305 310 315 320 Asp Leu Glu Gly Met Asp Gln Leu Pro Glu Tyr Met Lys Leu Tyr Tyr 325 330 335 Leu Ala Leu Tyr Asn Thr Thr Asn Glu Thr Ala Tyr Ile Leu Val Lys 340 345 350 Glu Lys Gly Phe Asn Ala Thr His Tyr Leu Lys Arg Val Trp Ala Met 355 360 365 Gln Ser Asn Gly Tyr Tyr Arg Glu Ala Gln Trp Phe Asn Ser Gly Tyr 370 375 380 Leu Pro Lys Phe Asp Glu Tyr Leu Asp Asn Ala Leu Val Ser Val Gly 385 390 395 400 Ala Pro Phe Leu Leu Gly Leu Ser Tyr Pro Met Met Gln Gln Gln Ile 405 410 415 Thr Lys Glu Glu Leu Asp Leu Ile Pro Glu Glu Leu Asn Leu Val Arg 420 425 430 Trp Ala Ser Ile Ile Phe Arg Leu Tyr Asp Asp Leu Ala Thr Ser Lys 435 440 445 Ala Glu Gln Gln Arg Gly Asp Val Pro Lys Ser Ile Gln Cys Tyr Met 450 455 460 His Glu Thr Gly Ser Ser Glu Glu Ile Ala Ala Asn His Ile Lys Glu 465 470 475 480 Met Ile Ser Glu Ala Trp Lys Asp Val Asn Ala Glu Cys Leu Lys Pro 485 490 495 Thr Ser Val Ser Lys His Phe Val Gly Val Ala Pro Asn Ser Ala Arg 500 505 510 Ser Gly Val Val Met Tyr His His Asp Phe Asp Gly Phe Ala Ser Pro 515 520 525 His Gly Arg Thr Asn Ala His Ile Thr Ser Ile Phe Phe Glu Pro Val 530 535 540 Pro Ile Arg Asp Ser Ile Asn Leu Gly 545 550 <210> 22 <211> 553 <212> PRT <213> Artificial sequence <220> <223> >VarL2 <400> 22 Met Asp Ser Val Glu Leu Arg Arg Thr Gly Ile Tyr His Ser Thr Phe 1 5 10 15 Trp Asp Ile Asp Ser Ile Arg Ala Leu Leu Ala Lys Arg Asp Cys Thr 20 25 30 Ala Ala Ala Ala Val Ser Pro Glu His His Lys Lys Leu Lys Glu Arg 35 40 45 Ile Gln Lys Arg Leu Asn Asp Ile Thr Gln Pro His His Leu Leu Gly 50 55 60 Leu Ile Asp Ala Val Gln Arg Leu Gly Val Ala Tyr Gln Phe Glu Glu 65 70 75 80 Glu Ile Ser Glu Ala Leu His Gly Val His Ser Glu Asn Thr Asp His 85 90 95 Ala Ile Lys Asp Ser Leu His His Thr Ser Leu Tyr Phe Arg Leu Leu 100 105 110 Arg Gln His Gly Cys Asn Leu Ser Ser Asp Ile Phe Asn Lys Phe Lys 115 120 125 Lys Asp Gly Gly Gly Phe Arg Ala Thr Leu Cys Glu Asp Ala Met Gly 130 135 140 Leu Leu Ser Leu Tyr Glu Ala Val Arg Leu Ser Val Lys Gly Glu Ala 145 150 155 160 Ile Leu Glu Glu Ala Gln Val Phe Ser Ile Ala Asn Leu Lys Ile Val 165 170 175 Met Glu Arg Ile Glu Arg Lys Leu Ala Asp Arg Ile Glu His Ala Leu 180 185 190 Glu Ile Pro Leu Tyr Trp Arg Ala Pro Arg Leu Glu Ala Arg Trp Tyr 195 200 205 Ile Asp Val Tyr Glu Lys Asp Asp Gly Arg Ile Asp Asp Leu Leu Asp 210 215 220 Phe Ala Lys Leu Asp Phe Asn Arg Val Gln Met Leu Tyr Gln Thr Glu 225 230 235 240 Leu Lys Glu Leu Ser Met Trp Trp Glu Leu Leu Gly Leu Pro Ala Lys 245 250 255 Met Gly Phe Phe Arg Asp Arg Leu Leu Glu Leu His Leu Phe Ser Ile 260 265 270 Ala Met Leu Val Glu Pro Gln Tyr Ser Gln Cys Arg Val Ala Leu Thr 275 280 285 Lys Ala Val Ile Leu Ile Ser Ala Met Asp Asp Phe Tyr Asp Ile His 290 295 300 Gly Leu Pro Asp Glu Leu Arg Val Phe Thr Asp Thr Val Asn Arg Trp 305 310 315 320 Asp Leu Glu Gly Ile Glu Gln Leu Pro Asp Tyr Met Lys Val Tyr Tyr 325 330 335 Leu Ala Leu Tyr Asn Thr Thr Asn Glu Thr Ala Trp Ile Ile Leu Lys 340 345 350 Glu Lys Gly Phe Asn Ala Thr His Tyr Leu Lys Lys Val Trp Ala Met 355 360 365 Gln Ser Asn Ala Tyr Phe Arg Glu Ala Gln Trp Phe Asn Thr Gly Tyr 370 375 380 Ile Pro Lys Phe Asp Glu Tyr Leu Glu Asn Ala Leu Val Ser Val Gly 385 390 395 400 Ala Pro Phe Leu Leu Gly Val Ser Tyr Pro Leu Leu Gln Asn Gln Ile 405 410 415 Ser Lys Glu Glu Leu Glu Leu Leu Pro Glu Asp Leu Asn Leu Ile Arg 420 425 430 Trp Ala Ser Met Ile Phe Arg Leu Trp Asp Asp Leu Ala Thr Ser Lys 435 440 445 Ala Glu Gln Gln Arg Gly Glu Val Pro Lys Thr Ile Gln Cys Tyr Met 450 455 460 His Glu Thr Gly Ser Ser Glu Glu Val Ala Ala Asn His Ile Arg Asp 465 470 475 480 Met Ile Ser Glu Ala Trp Lys Glu Val Asn Ala Glu Cys Leu Lys Pro 485 490 495 Thr Ser Val Ser Arg His Tyr Val Gly Val Ala Pro Asn Thr Ala Arg 500 505 510 Ser Gly Leu Leu Met Tyr His His Asp Phe Asp Gly Phe Ala Ser Pro 515 520 525 His Gly Lys Thr Asn Ala His Ile Thr Ser Ile Phe Phe Asp Pro Val 530 535 540 Pro Leu Lys Glu Ser Ile Asn Leu Gly 545 550 <210> 23 <211> 553 <212> PRT <213> Artificial sequence <220> <223> >VarL3 <400> 23 Met Asp Ser Val Glu Val Arg Arg Thr Gly Ile Tyr His Ser Thr Phe 1 5 10 15 Trp Asp Ile Asp Ser Ile Arg Ala Leu Val Ala Arg Arg Asp Cys Ser 20 25 30 Ala Gly Ala Ala Leu Ser Pro Asp His His Lys Arg Leu Lys Glu Arg 35 40 45 Ile Gln Arg Arg Leu Gln Glu Ile Ser Gln Pro His His Leu Leu Gly 50 55 60 Leu Ile Asp Ala Val Gln Arg Leu Gly Val Ala Tyr Gln Phe Glu Glu 65 70 75 80 Glu Ile Ser Asp Ala Leu His Ala Ile His Thr Asp Gln Ser Asp His 85 90 95 Ala Ile Arg Asp Ser Leu His His Thr Ser Leu Tyr Phe Arg Leu Leu 100 105 110 Arg Gln His Gly Cys Asn Leu Ser Ser Asp Ile Phe Gln Lys Phe Lys 115 120 125 Lys Glu Gly Gly Gly Phe Lys Ala Ser Leu Cys Glu Asp Ala Met Gly 130 135 140 Leu Leu Ser Leu Tyr Glu Ala Val Arg Leu Ser Ile Lys Gly Glu Ala 145 150 155 160 Ile Leu Glu Glu Ala Gln Val Phe Ser Ile Ala Asn Leu Arg Ile Leu 165 170 175 Leu Glu Arg Ile Glu Arg Lys Leu Ala Glu Lys Ile Glu His Ala Leu 180 185 190 Glu Leu Pro Leu Tyr Trp Arg Ala Pro Arg Leu Glu Ala Arg Trp Tyr 195 200 205 Ile Glu Val Tyr Glu Lys Asp Asp Ala Lys Ile Glu Asp Leu Leu Asp 210 215 220 Phe Ala Lys Leu Asp Phe Asn Arg Leu Gln Met Met Tyr Gln Thr Glu 225 230 235 240 Leu Lys Glu Leu Ser Met Trp Trp Glu Leu Leu Gly Leu Pro Ala Arg 245 250 255 Met Gly Phe Phe Arg Asp Arg Ile Leu Glu Leu His Leu Phe Ser Ile 260 265 270 Ala Met Val Val Glu Pro Asn Tyr Ser Gln Cys Arg Ile Ala Leu Thr 275 280 285 Lys Ala Ile Ile Leu Val Ser Ala Met Asp Asp Tyr Tyr Asp Val His 290 295 300 Gly Leu Pro Asp Glu Leu Lys Ile Phe Thr Asp Thr Val Asn Arg Trp 305 310 315 320 Asp Leu Glu Gly Ile Glu Gln Leu Pro Asp Tyr Met Lys Ile Tyr Tyr 325 330 335 Leu Ala Leu Tyr Asn Thr Thr Asn Glu Thr Ala Tyr Ile Leu Leu Lys 340 345 350 Glu Lys Gly Phe Asn Ala Ser His Tyr Leu Arg Lys Leu Trp Ala Val 355 360 365 Gln Ser Asn Ala Tyr Phe Lys Glu Ala Gln Trp Phe Asn Ser Gly Tyr 370 375 380 Val Pro Lys Phe Asp Glu Tyr Leu Glu Asn Ala Leu Val Ser Ile Gly 385 390 395 400 Ala Pro Phe Ile Val Ala Val Thr Tyr Pro Met Ile Gln Asn Asn Val 405 410 415 Thr Lys Glu Glu Ile Glu Leu Leu Pro Glu Asp Leu Gln Leu Leu Lys 420 425 430 Trp Ala Ser Val Ile Phe Arg Leu Tyr Asp Asp Leu Gly Thr Ser Lys 435 440 445 Ala Glu Gln Gln Arg Gly Glu Val Pro Lys Ser Ile Gln Cys Tyr Met 450 455 460 His Glu Thr Gly Ser Ser Glu Glu Val Ala Ala Asn His Ile Lys Glu 465 470 475 480 Ile Ile Ser Asp Ala Trp Lys Glu Val Asn Ala Glu Cys Leu Arg Pro 485 490 495 Thr Thr Leu Ser Arg His Tyr Val Gly Val Ala Pro Asn Ser Ala Arg 500 505 510 Ser Ala Val Val Ile Tyr His His Asp Phe Asp Gly Phe Ala Thr Pro 515 520 525 His Ala Lys Thr Gln Ala His Ile Thr Ser Ile Phe Phe Glu Pro Val 530 535 540 Pro Ile Lys Glu Ser Ile Asn Leu Gly 545 550 <210> 24 <211> 553 <212> PRT <213> Artificial Sequence <220> <223> >VarL4 <400> 24 Met Asp Ser Met Glu Val Arg Arg Ser Gly Ile Tyr His Ser Thr Phe 1 5 10 15 Trp Asp Ile Asp Ser Ile Arg Ala Leu Leu Ala Lys Arg Asp Cys Ser 20 25 30 Ala Ala Ala Gly Leu Ser Pro Glu His His Lys Lys Leu Lys Glu Lys 35 40 45 Ile Gln Lys Lys Leu Asn Glu Ile Ser Asn Pro His His Leu Leu Gly 50 55 60 Leu Ile Asp Ala Val Gln Arg Leu Gly Val Ala Tyr Gln Phe Glu Asp 65 70 75 80 Glu Ile Ser Glu Ala Leu His Gly Ile His Ser Asp Asn Thr Asp His 85 90 95 Ala Val Lys Glu Ser Leu His His Thr Ser Leu Tyr Phe Arg Leu Leu 100 105 110 Arg Gln His Gly Cys Gln Leu Ser Ser Asp Ile Phe Asn Lys Phe Arg 115 120 125 Lys Asp Gly Gly Gly Phe Arg Ala Ser Leu Cys Asp Asp Ala Met Gly 130 135 140 Leu Leu Ser Leu Tyr Glu Ala Val Arg Leu Ser Val Lys Gly Glu Ala 145 150 155 160 Ile Leu Glu Glu Ala Gln Leu Phe Ser Ile Ala Asn Leu Lys Ile Met 165 170 175 Leu Glu Arg Leu Glu Lys Lys Leu Ala Glu Arg Ile Glu His Ala Leu 180 185 190 Asp Leu Pro Val Tyr Trp Arg Ala Pro Arg Leu Glu Ala Arg Phe Tyr 195 200 205 Ile Asp Ile Tyr Glu Arg Asp Asp Ala Arg Leu Glu Asp Leu Leu Asp 210 215 220 Phe Ala Lys Leu Asp Phe Asn Arg Leu Gln Met Met Tyr Gln Thr Glu 225 230 235 240 Leu Lys Glu Leu Ser Met Trp Trp Glu Leu Leu Gly Leu Pro Ala Lys 245 250 255 Met Gly Phe Phe Arg Asp Arg Leu Leu Glu Ile His Leu Phe Thr Ile 260 265 270 Ala Val Leu Met Glu Pro Gln Phe Ser Gln Cys Arg Ile Ala Leu Thr 275 280 285 Lys Gly Ile Val Leu Val Ser Ala Met Asp Asp Phe Tyr Asp Val His 290 295 300 Gly Leu Pro Asp Glu Leu Lys Leu Phe Thr Glu Thr Val Asn Arg Trp 305 310 315 320 Asp Ile Glu Ala Met Glu Gln Leu Pro Glu Tyr Met Lys Val Phe Tyr 325 330 335 Met Ala Leu Tyr Asn Thr Ser Gln Glu Thr Ala Tyr Ile Val Leu Lys 340 345 350 Asp Lys Gly Tyr Asn Ala Ser His Tyr Leu Lys Lys Ile Trp Gly Met 355 360 365 Gln Ser Asn Ala Tyr Phe Lys Glu Ala Gln Trp Tyr Asn Ser Gly Tyr 370 375 380 Ile Pro Lys Phe Asp Glu Tyr Leu Glu Asn Ala Val Val Ser Leu Gly 385 390 395 400 Gly Pro Phe Val Leu Ala Leu Ser Tyr Pro Leu Met Gln Gln Gln Ile 405 410 415 Thr Lys Glu Asp Leu Glu Leu Val Pro Glu Asp Leu Asn Leu Met Lys 420 425 430 Trp Ala Ser Val Ile Phe Arg Leu Trp Asp Asp Leu Ala Thr Ser Lys 435 440 445 Ala Glu Gln Gln Arg Gly Glu Leu Pro Lys Ser Ile Asn Cys Tyr Met 450 455 460 His Glu Thr Gly Thr Ser Glu Glu Leu Ala Ala Asn His Ile Lys Asp 465 470 475 480 Leu Ile Ser Glu Ala Trp Lys Asp Val Asn Ala Glu Cys Leu Arg Pro 485 490 495 Ser Thr Val Ser Lys His Phe Met Gly Val Ala Pro Asn Ser Ala Arg 500 505 510 Thr Ala Val Leu Ile Tyr His His Glu Phe Asp Gly Phe Ala Ser Pro 515 520 525 His Gly Lys Thr Asn Ala His Ile Ser Thr Ile Phe Phe Asp Pro Val 530 535 540 Pro Ile Lys Glu Ser Ile Asn Leu Gly 545 550 <210> 25 <211> 553 <212> PRT <213> Artificial Sequence <220> <223> >VarL5 <400> 25 Met Asp Ser Leu Asp Leu Arg Arg Ser Ala Ile Tyr His Ser Thr Phe 1 5 10 15 Trp Asp Ile Asp Ser Ile Lys Ala Leu Ile Ala Lys Arg Asp Cys Ser 20 25 30 Ala Ala Ala Ala Val Ser Pro Glu His His Arg Lys Leu Lys Glu Lys 35 40 45 Ile Gln Lys Lys Leu Gln Glu Ile Ser Gln Pro His His Leu Leu Gly 50 55 60 Leu Ile Asp Ala Val Gln Arg Leu Gly Val Ala Tyr Gln Phe Glu Asp 65 70 75 80 Glu Ile Ser Asp Ala Leu His Gly Met His Thr Asp Gln Thr Asp His 85 90 95 Ala Val Lys Glu Ser Leu His His Thr Ser Leu Trp Phe Arg Leu Leu 100 105 110 Arg Gln His Gly Cys Asn Leu Ser Ser Asp Ile Phe Asn Arg Phe Arg 115 120 125 Lys Glu Gly Gly Gly Phe Lys Ala Ser Ile Cys Glu Asp Ala Met Gly 130 135 140 Leu Leu Ser Leu Tyr Glu Ala Val Arg Leu Ser Ile Lys Gly Glu Ala 145 150 155 160 Ile Leu Glu Glu Ala Gln Val Phe Ser Ile Ala Asn Leu Arg Ile Met 165 170 175 Leu Glu Lys Val Glu Lys Lys Leu Ala Asp Arg Ile Asp His Ala Leu 180 185 190 Glu Leu Pro Ile Phe Trp Arg Ala Pro Arg Leu Glu Ala Arg Trp Tyr 195 200 205 Ile Glu Val Tyr Glu Arg Asp Glu Gly Arg Ile Glu Glu Leu Leu Asp 210 215 220 Phe Ala Lys Leu Asp Phe Asn Arg Val Gln Met Val Tyr Gln Thr Glu 225 230 235 240 Leu Lys Glu Leu Thr Met Trp Trp Glu Leu Val Gly Leu Pro Ala Arg 245 250 255 Met Gly Phe Phe Arg Asp Arg Ile Leu Glu Ile His Leu Phe Ser Leu 260 265 270 Ala Ile Leu Val Glu Pro Gln Tyr Ser Asn Cys Arg Val Ala Leu Thr 275 280 285 Lys Ala Ile Ile Leu Ile Ser Ala Met Asp Asp Phe Tyr Asp Val His 290 295 300 Gly Leu Pro Asp Glu Leu Lys Leu Phe Thr Glu Thr Val Gln Arg Trp 305 310 315 320 Asp Ile Asp Ala Ile Asp Asn Leu Pro Glu Tyr Met Lys Leu Phe Tyr 325 330 335 Val Ala Leu Tyr Asn Thr Thr Asn Glu Thr Ala Trp Leu Ile Val Lys 340 345 350 Glu Arg Gly Phe Asn Ala Thr His Tyr Leu Arg Lys Val Trp Ala Met 355 360 365 Gln Ser Asn Ala Tyr Phe Arg Glu Ala Gln Trp Phe Gln Ser Gly Tyr 370 375 380 Val Pro Lys Phe Asp Glu Tyr Leu Glu Asn Ala Val Ile Ser Ile Gly 385 390 395 400 Gly Pro Phe Ile Leu Gly Leu Ser Tyr Pro Leu Ile Gln Asn Gln Val 405 410 415 Thr Lys Glu Glu Ile Asp Leu Leu Pro Glu Asp Leu Asn Leu Val Lys 420 425 430 Trp Ala Ser Val Ile Phe Arg Leu Tyr Asp Asp Leu Gly Thr Ser Lys 435 440 445 Ala Glu Asn Gln Arg Gly Asp Leu Pro Lys Ser Ile Gln Cys Tyr Met 450 455 460 His Glu Thr Gly Ser Ser Glu Glu Val Ala Ala Asn His Ile Lys Glu 465 470 475 480 Met Ile Ser Glu Ala Trp Lys Asp Val Asn Ala Glu Cys Leu Arg Pro 485 490 495 Ser Thr Ile Ser Arg His Phe Leu Gly Val Ala Pro Asn Ser Ala Arg 500 505 510 Ser Ala Val Ile Met Tyr His His Asp Phe Asp Gly Phe Ala Ser Pro 515 520 525 His Ala Lys Thr Asn Ala His Ile Thr Ser Ile Phe Tyr Asp Pro Val 530 535 540 Pro Leu Arg Glu Ser Ile Asn Leu Gly 545 550 <210> 26 <211> 553 <212> PRT <213> Artificial sequence <220> <223> >VarL6 <400> 26 Met Asp Ser Val Glu Val Arg Arg Thr Gly Met Tyr His Ser Thr Phe 1 5 10 15 Trp Asp Ile Asp Ser Ile Arg Ala Leu Leu Ala Lys Arg Glu Cys Ser 20 25 30 Ala Gly Gly Ala Met Ser Pro Glu His His Lys Arg Leu Lys Glu Arg 35 40 45 Ile Gln Arg Lys Leu Asn Asp Val Ser Gln Pro His His Ile Leu Ala 50 55 60 Leu Ile Asp Ala Val Gln Arg Leu Gly Val Ala Tyr Gln Phe Asp Glu 65 70 75 80 Glu Ile Ser Glu Ala Leu His Gly Ile His Ser Glu Gln Thr Asp His 85 90 95 Gly Met Arg Glu Ser Leu His His Thr Ser Leu Tyr Phe Arg Leu Leu 100 105 110 Arg Gln His Gly Cys Asn Leu Ser Ser Asp Ile Phe Gln Lys Phe Lys 115 120 125 Lys Glu Gly Gly Gly Phe Lys Ala Ser Val Cys Asp Asp Ala Met Gly 130 135 140 Leu Leu Ser Leu Tyr Glu Ala Val Arg Leu Ser Leu Lys Gly Glu Ala 145 150 155 160 Ile Leu Glu Glu Ala Gln Ile Phe Ser Ile Ala Gln Leu Arg Leu Met 165 170 175 Leu Asp Arg Ile Asp Lys Lys Leu Ala Glu Arg Ile Glu His Ala Leu 180 185 190 Glu Leu Pro Ile Tyr Tyr Arg Ala Pro Arg Leu Glu Ala Arg Phe Tyr 195 200 205 Ile Glu Ile Tyr Asp Lys Asp Glu Gly Arg Ile Glu Asp Leu Leu Asp 210 215 220 Phe Ala Lys Leu Asp Phe Asn Arg Ile Gln Met Ile Tyr Gln Thr Glu 225 230 235 240 Leu Arg Glu Leu Thr Met Trp Trp Glu Leu Leu Gly Leu Pro Ala Lys 245 250 255 Met Gly Phe Phe Arg Asp Arg Ile Leu Glu Leu His Leu Phe Ser Leu 260 265 270 Ala Leu Leu Val Glu Pro Asn Phe Ser Asn Cys Arg Met Gly Leu Thr 275 280 285 Lys Gly Ile Met Leu Val Thr Ala Met Asp Asp Phe Tyr Asp Val His 290 295 300 Gly Leu Pro Asp Glu Leu Lys Leu Phe Thr Glu Thr Val Asn Arg Trp 305 310 315 320 Asp Ile Asp Ala Leu Glu Gln Leu Pro Asp Tyr Met Lys Met Phe Tyr 325 330 335 Val Ala Leu Phe Asn Thr Thr Asn Glu Thr Ala Tyr Met Leu Leu Arg 340 345 350 Glu Lys Gly Phe Asn Ala Thr His Phe Leu Lys Lys Leu Trp Ala Leu 355 360 365 Gln Ser Asn Gly Tyr Phe Lys Glu Ala Gln Trp Tyr Asn Ser Gly Tyr 370 375 380 Ile Pro Lys Phe Asp Glu Tyr Leu Glu Asn Ala Ile Met Ser Ile Gly 385 390 395 400 Gly Pro Phe Leu Leu Ala Leu Ser Tyr Pro Met Val Gln Asn Gln Val 405 410 415 Thr Arg Glu Asp Leu Glu Met Val Pro Glu Glu Leu Asn Leu Ile Lys 420 425 430 Trp Ala Ser Leu Ile Phe Arg Leu Trp Asp Asp Leu Gly Thr Ser Lys 435 440 445 Ala Glu Gln Gln Arg Gly Asp Leu Pro Lys Ser Ile Gln Cys Tyr Met 450 455 460 His Glu Ser Gly Thr Ser Asp Glu Val Ala Ala Asn His Ile Lys Glu 465 470 475 480 Ile Ile Ser Glu Ala Trp Lys Glu Val Asn Ala Glu Cys Leu Arg Pro 485 490 495 Ser Thr Ile Ser Arg His Phe Met Ala Leu Ala Pro Asn Ser Ala Arg 500 505 510 Ser Ala Leu Val Val Tyr His His Asp Phe Asp Gly Phe Ala Thr Pro 515 520 525 His Ala Lys Thr Asn Gly His Ile Ser Ser Ile Phe Phe Asp Pro Val 530 535 540 Pro Val Lys Glu Ser Ile Asn Leu Gly 545 550 <210> 27 <211> 553 <212> PRT <213> Synthetic sequence <220> <223> >VarL7 <400> 27 Met Asp Ser Leu Asp Ile Arg Arg Thr Gly Ile Tyr His Ser Thr Phe 1 5 10 15 Trp Asp Val Glu Ser Ile Lys Ala Ile Ile Ala Lys Lys Glu Cys Thr 20 25 30 Ala Gly Gly Ala Ile Ser Pro Glu His His Lys Lys Leu Lys Glu Lys 35 40 45 Ile Gln Lys Arg Leu Asn Glu Val Ser Gln Pro His His Leu Leu Gly 50 55 60 Leu Ile Asp Ala Val Gln Arg Leu Gly Val Ala Tyr Gln Phe Asp Asp 65 70 75 80 Glu Ile Ser Glu Ala Leu His Gly Val His Ser Asp Asn Ser Asp His 85 90 95 Gly Leu Arg Glu Ser Leu His His Thr Ser Leu Trp Phe Arg Leu Leu 100 105 110 Arg Gln His Gly Cys Asn Leu Ser Ser Asp Ile Phe Asn Arg Phe Arg 115 120 125 Lys Asp Gly Gly Gly Phe Lys Ala Ser Ile Cys Glu Asp Ala Val Gly 130 135 140 Leu Leu Ser Leu Tyr Glu Ala Val Arg Leu Ser Ile Lys Gly Glu Gly 145 150 155 160 Ile Leu Glu Glu Ala Gln Val Phe Ser Ile Ala Asn Leu Arg Ile Ile 165 170 175 Ile Asp Lys Met Asp Lys Lys Leu Ala Glu Arg Ile Glu His Ala Leu 180 185 190 Glu Leu Pro Leu Tyr Trp Arg Ala Pro Arg Leu Glu Ala Arg Tyr Tyr 195 200 205 Ile Glu Ile Tyr Glu Arg Asp Glu Gly Lys Leu Glu Glu Leu Leu Asp 210 215 220 Phe Ala Lys Leu Asp Phe Asn Arg Leu Gln Met Val Tyr Gln Thr Glu 225 230 235 240 Leu Arg Asp Leu Thr Met Trp Trp Asp Leu Leu Gly Leu Pro Ala Lys 245 250 255 Met Ala Phe Phe Arg Asp Arg Ile Leu Glu Val His Leu Phe Thr Leu 260 265 270 Ala Ile Leu Val Glu Pro Gln Phe Ser Gln Cys Arg Ile Gly Val Thr 275 280 285 Lys Ala Val Val Leu Val Ser Ala Met Asp Asp Phe Tyr Asp Met His 290 295 300 Gly Leu Pro Asp Glu Leu Arg Met Phe Thr Glu Thr Val Gln Arg Trp 305 310 315 320 Asp Val Asp Ala Val Glu Asn Leu Pro Glu Tyr Met Lys Val Tyr Tyr 325 330 335 Met Ala Leu Tyr Asn Ser Thr Asn Glu Thr Ala Phe Ile Val Met Lys 340 345 350 Asp Lys Gly Leu Asn Ala Thr His Tyr Leu Arg Lys Ile Trp Ala Met 355 360 365 Gln Ser Asn Gly Tyr Tyr Lys Glu Ala Gln Trp Tyr Gln Ser Gly Tyr 370 375 380 Val Pro Lys Phe Asp Glu Tyr Leu Glu Asn Ala Leu Ile Ser Ile Gly 385 390 395 400 Gly Pro Phe Val Val Gly Met Thr Tyr Pro Leu Met Gln Asn Asn Val 405 410 415 Thr Lys Glu Glu Leu Glu Leu Val Pro Asp Glu Leu Asn Leu Val Lys 420 425 430 Trp Ala Thr Val Ile Phe Arg Leu Tyr Asp Asp Leu Gly Thr Thr Lys 435 440 445 Ala Glu Asn Gln Arg Gly Glu Val Pro Lys Ser Ile Gln Cys Tyr Met 450 455 460 His Glu Ser Gly Thr Ser Glu Glu Ile Ala Ala Asn His Ile Lys Glu 465 470 475 480 Met Ile Ser Glu Gly Trp Lys Asp Val Asn Ala Glu Cys Leu Arg Pro 485 490 495 Ser Thr Ile Ser Arg His Phe Leu Ala Leu Ala Pro Asn Thr Ala Arg 500 505 510 Thr Ala Val Leu Leu Tyr His His Asp Phe Asp Gly Phe Ala Ser Pro 515 520 525 His Ala Lys Thr Asn Gly His Ile Ser Ser Ile Phe Phe Asp Pro Val 530 535 540 Pro Met Arg Asp Ser Ile Asn Leu Gly 545 550 <210> 28 <211> 553 <212> PRT <213> artificial sequence <220> <223> >VarL8 <400> 28 Met Asp Ser Val Asp Ile Arg Lys Thr Gly Met Tyr His Ser Thr Phe 1 5 10 15 Trp Asp Ile Glu Ser Ile Lys Ala Leu Met Ala Arg Arg Glu Cys Ser 20 25 30 Ala Gly Ala Ala Leu Ser Pro Glu His His Lys Lys Leu Lys Asp Lys 35 40 45 Ile Gln Lys Arg Leu Asn Glu Leu Ser Asn Pro His His Leu Leu Gly 50 55 60 Leu Ile Asp Ala Val Gln Arg Leu Gly Val Ala Tyr Gln Phe Glu Asp 65 70 75 80 Glu Ile Ser Glu Ala Leu His Gly Val His Ser Glu Asn Thr Asp His 85 90 95 Ala Leu Arg Glu Ser Leu His His Thr Ser Leu Trp Phe Arg Leu Leu 100 105 110 Arg Gln His Gly Cys Asn Leu Ser Ser Asp Ile Phe Asn Lys Phe Arg 115 120 125 Lys Asp Gly Gly Gly Phe Lys Ala Ser Leu Cys Asp Asp Ala Met Gly 130 135 140 Leu Leu Ser Leu Tyr Glu Ala Val Arg Leu Ser Leu Arg Gly Glu Ala 145 150 155 160 Ile Leu Glu Glu Ala Asn Leu Phe Ser Ile Ala Gln Leu Arg Ile Ile 165 170 175 Leu Glu Lys Ile Asp Lys Lys Leu Ala Glu Lys Ile Asp His Ala Leu 180 185 190 Asp Leu Pro Ile Tyr Trp Arg Ala Pro Arg Leu Glu Ala Arg Phe Tyr 195 200 205 Ile Glu Ile Tyr Glu Arg Asp Glu Gly Lys Leu Glu Asp Leu Leu Asp 210 215 220 Phe Ala Lys Leu Asp Phe Asn Arg Leu Gln Val Ile Tyr Gln Thr Glu 225 230 235 240 Leu Lys Asp Leu Thr Met Trp Trp Glu Leu Leu Gly Leu Pro Ala Lys 245 250 255 Met Gly Phe Phe Arg Asp Arg Val Leu Glu Val His Leu Phe Ser Met 260 265 270 Ala Leu Ile Leu Glu Pro Gln Tyr Ser Asn Cys Arg Leu Gly Val Thr 275 280 285 Lys Ala Val Met Leu Ile Ser Ala Met Asp Asp Tyr Tyr Asp Met His 290 295 300 Gly Leu Pro Asp Glu Leu Arg Leu Phe Thr Glu Thr Val Asn Arg Trp 305 310 315 320 Asp Val Glu Ala Met Glu Gln Leu Pro Asp Tyr Met Lys Ile Phe Tyr 325 330 335 Met Ala Leu Tyr Asn Ser Ser Asn Glu Thr Ala Phe Ile Met Leu Arg 340 345 350 Glu Lys Gly Phe Asn Ala Thr His Tyr Leu Arg Lys Ile Trp Ala Val 355 360 365 Gln Ser Gln Gly Tyr Phe Arg Glu Ala Gln Trp Phe Gln Thr Gly Tyr 370 375 380 Val Pro Lys Phe Asp Glu Tyr Leu Glu Asn Ala Val Ile Ser Ile Gly 385 390 395 400 Gly Pro Phe Leu Ile Ala Val Ser Tyr Pro Leu Met Gln Asn Asn Leu 405 410 415 Thr Arg Glu Asp Val Glu Ile Leu Pro Glu Glu Leu Gln Leu Ile Lys 420 425 430 Tyr Ala Thr Val Ile Phe Arg Leu Trp Asp Asp Leu Gly Thr Ser Lys 435 440 445 Ala Glu Gln Asn Arg Gly Glu Leu Pro Lys Ser Ile Asn Cys Tyr Met 450 455 460 His Glu Thr Gly Ser Ser Glu Glu Leu Ala Ala Asn His Ile Arg Glu 465 470 475 480 Ile Ile Ser Glu Ala Trp Lys Asp Val Asn Ala Asp Cys Leu Arg Pro 485 490 495 Ser Thr Ile Ser Arg His Phe Leu Gly Ile Ala Pro Asn Thr Ala Arg 500 505 510 Thr Gly Val Val Val Tyr His His Asp Phe Asp Gly Phe Ala Thr Pro 515 520 525 His Ala Arg Thr Gln Ala His Ile Ser Thr Ile Phe Phe Asp Pro Val 530 535 540 Pro Met Arg Glu Ser Ile Asn Leu Gly 545 550 <210> 29 <211> 553 <212> PRT <213> artificial sequence <220> <223> Cicassy N267T <400> 29 Met Asp Ser Met Glu Val Arg Arg Ser Ala Ile Tyr His Ser Thr Phe 1 5 10 15 Trp Asp Ile Asp Ser Ile Arg Ala Leu Leu Ala Arg Arg Asp Cys Thr 20 25 30 Ala Ala Ala Ala Leu Ser Pro Asp His His Lys Arg Leu Lys Glu Arg 35 40 45 Ile Gln Arg Arg Leu Gln Asp Ile Thr Gln Pro His His Leu Leu Gly 50 55 60 Leu Ile Asp Ala Val Gln Arg Leu Gly Val Ala Tyr Gln Phe Glu Glu 65 70 75 80 Glu Ile Ser Asp Ala Leu His Gly Leu His Ser Glu Asn Thr Glu His 85 90 95 Ala Ile Lys Asp Ser Leu His His Thr Ser Leu Tyr Phe Arg Leu Leu 100 105 110 Arg Gln His Gly Cys Asn Leu Ser Ser Asp Ile Phe Asn Lys Phe Lys 115 120 125 Lys Glu Gly Gly Gly Phe Lys Ala Ser Leu Cys Glu Asp Ala Met Gly 130 135 140 Leu Leu Ser Leu Tyr Glu Ala Val Arg Leu Ser Val Lys Gly Glu Ala 145 150 155 160 Ile Leu Glu Glu Ala Gln Val Phe Ser Ile Ala Asn Leu Lys Ile Leu 165 170 175 Met Glu Arg Val Glu Arg Lys Leu Ala Asp Arg Ile Glu His Ala Leu 180 185 190 Glu Ile Pro Leu Tyr Trp Arg Ala Pro Arg Leu Glu Ala Arg Trp Tyr 195 200 205 Ile Asp Val Tyr Glu Lys Glu Asp Gly Arg Ile Asp Asp Leu Leu Asp 210 215 220 Phe Ala Lys Leu Asp Phe Asn Arg Val Gln Met Leu Tyr Gln Thr Glu 225 230 235 240 Leu Lys Glu Leu Ser Met Trp Trp Glu Leu Leu Gly Leu Pro Ala Lys 245 250 255 Met Gly Phe Phe Arg Asp Arg Leu Leu Glu Thr His Leu Phe Ser Ile 260 265 270 Ala Val Val Val Glu Pro Gln Tyr Ser Gln Cys Arg Val Ala Ile Thr 275 280 285 Lys Ala Ile Val Leu Met Thr Ala Met Asp Asp Phe Tyr Asp Val His 290 295 300 Gly Leu Pro Asp Glu Leu Lys Val Phe Thr Asp Thr Val Asn Arg Trp 305 310 315 320 Asp Leu Glu Gly Ile Asp Gln Leu Pro Glu Tyr Met Lys Leu Tyr Tyr 325 330 335 Leu Ala Leu Tyr Asn Thr Thr Asn Glu Thr Ala Tyr Ile Ile Leu Lys 340 345 350 Glu Lys Gly Phe Asn Ala Thr His Tyr Leu Lys Lys Leu Trp Ala Met 355 360 365 Gln Ser Asn Ala Tyr Phe Arg Glu Ala Gln Trp Phe Asn Ser Gly Tyr 370 375 380 Ile Pro Lys Phe Asp Glu Tyr Leu Asp Asn Ala Leu Val Ser Val Gly 385 390 395 400 Ala Pro Phe Val Leu Gly Leu Ser Tyr Pro Met Ile Gln Gln Gln Ile 405 410 415 Ser Lys Glu Glu Ile Asp Leu Ile Pro Glu Asp Leu Asn Leu Leu Arg 420 425 430 Trp Ala Ser Ile Ile Phe Arg Leu Tyr Asp Asp Leu Ala Thr Ser Lys 435 440 445 Ala Glu Gln Gln Arg Gly Asp Val Pro Lys Ser Ile Gln Cys Tyr Met 450 455 460 His Glu Thr Gly Ser Ser Glu Glu Val Ala Ala Asn His Ile Arg Asp 465 470 475 480 Leu Ile Ser Asp Ala Trp Lys Glu Val Asn Ala Glu Cys Leu Lys Pro 485 490 495 Thr Ser Leu Ser Lys His Tyr Val Gly Val Ala Pro Asn Ser Ala Arg 500 505 510 Ser Gly Val Leu Met Tyr His His Asp Phe Asp Gly Phe Ala Ser Pro 515 520 525 His Gly Arg Thr Asn Ala His Ile Thr Ser Ile Phe Phe Glu Pro Val 530 535 540 Pro Leu Lys Glu Ser Ile Asn Leu Gly 545 550[[ID=​​​​​​​​​​​​​​​​1 5 10 15 Trp Asp Ile Asp Ser Ile Arg Ala Leu Leu Ala Arg Arg Asp Cys Thr 20 25 30 Ala Ala Ala Ala Leu Ser Pro Asp His His Lys Arg Leu Lys Glu Arg 35 40 45 Ile Gln Arg Arg Leu Gln Asp Ile Thr Gln Pro His His Leu Leu Gly 50 55 60 Leu Ile Asp Ala Val Gln Arg Leu Gly Val Ala Tyr Gln Phe Glu Glu 65 70 75 80 Glu Ile Ser Asp Ala Leu His Gly Leu His Ser Glu Asn Thr Glu His 85 90 95 Ala Ile Lys Asp Ser Leu His His Thr Ser Leu Tyr Phe Arg Leu Leu 100 105 110 Arg Gln His Gly Cys Asn Leu Ser Ser Asp Ile Phe Asn Lys Phe Lys 115 120 125 Lys Glu Gly Gly Gly Phe Lys Ala Ser Leu Cys Glu Asp Ala Met Gly 130 135 140 Leu Leu Ser Leu Tyr Glu Ala Val Arg Leu Ser Val Lys Gly Glu Ala 145 150 155 160 Ile Leu Glu Glu Ala Gln Val Phe Ser Ile Ala Asn Leu Lys Ile Leu 165 170 175 Met Glu Arg Val Glu Arg Lys Leu Ala Asp Arg Ile Glu His Ala Leu 180 185 190 Glu Ile Pro Leu Tyr Trp Arg Ala Pro Arg Leu Glu Ala Arg Trp Tyr 195 200 205 Ile Asp Val Tyr Glu Lys Glu Asp Gly Arg Ile Asp Asp Leu Leu Asp 210 215 220 Phe Ala Lys Leu Asp Phe Asn Arg Val Gln Met Leu Tyr Gln Thr Glu 225 230 235 240 Leu Lys Glu Leu Ser Met Trp Trp Glu Leu Leu Gly Leu Pro Ala Lys 245 250 255 Met Gly Phe Phe Arg Asp Arg Leu Leu Glu Asn His Leu Phe Ser Ile 260 265 270 Ala Val Val Val Glu Pro Gln Tyr Ser Gln Cys Arg Val Ala Ile Thr 275 280 285 Lys Ala Val Val Leu Met Thr Ala Met Asp Asp Phe Tyr Asp Val His 290 295 300 Gly Leu Pro Asp Glu Leu Lys Val Phe Thr Asp Thr Val Asn Arg Trp 305 310 315 320 Asp Leu Glu Gly Ile Asp Gln Leu Pro Glu Tyr Met Lys Leu Tyr Tyr 325 330 335 Leu Ala Leu Tyr Asn Thr Thr Asn Glu Thr Ala Tyr Ile Ile Leu Lys 340 345 350 Glu Lys Gly Phe Asn Ala Thr His Tyr Leu Lys Lys Leu Trp Ala Met 355 360 365 Gln Ser Asn Ala Tyr Phe Arg Glu Ala Gln Trp Phe Asn Ser Gly Tyr 370 375 380 Ile Pro Lys Phe Asp Glu Tyr Leu Asp Asn Ala Leu Val Ser Val Gly 385 390 395 400 Ala Pro Phe Val Leu Gly Leu Ser Tyr Pro Met Ile Gln Gln Gln Ile 405 410 415 Ser Lys Glu Glu Ile Asp Leu Ile Pro Glu Asp Leu Asn Leu Leu Arg 420 425 430 Trp Ala Ser Ile Ile Phe Arg Leu Tyr Asp Asp Leu Ala Thr Ser Lys 435 440 445 Ala Glu Gln Gln Arg Gly Asp Val Pro Lys Ser Ile Gln Cys Tyr Met 450 455 460 His Glu Thr Gly Ser Ser Glu Glu Val Ala Ala Asn His Ile Arg Asp 465 470 475 480 Leu Ile Ser Asp Ala Trp Lys Glu Val Asn Ala Glu Cys Leu Lys Pro 485 490 495 Thr Ser Leu Ser Lys His Tyr Val Gly Val Ala Pro Asn Ser Ala Arg 500 505 510 Ser Gly Val Leu Met Tyr His His Asp Phe Asp Gly Phe Ala Ser Pro 515 520 525 His Gly Arg Thr Asn Ala His Ile Thr Ser Ile Phe Phe Glu Pro Val 530 535 540 Pro Leu Lys Glu Ser Ile Asn Leu Gly 545 550 <210> 31 <211> 553 <212> PRT <213> artificial sequence <220> <223> I291T <400> 31 Met Asp Ser Met Glu Val Arg Arg Ser Ala Ile Tyr His Ser Thr Phe 1 5 10 15 Trp Asp Ile Asp Ser Ile Arg Ala Leu Leu Ala Arg Arg Asp Cys Thr 20 25 30 Ala Ala Ala Ala Leu Ser Pro Asp His His Lys Arg Leu Lys Glu Arg 35 40 45 Ile Gln Arg Arg Leu Gln Asp Ile Thr Gln Pro His His Leu Leu Gly 50 55 60 Leu Ile Asp Ala Val Gln Arg Leu Gly Val Ala Tyr Gln Phe Glu Glu 65 70 75 80 Glu Ile Ser Asp Ala Leu His Gly Leu His Ser Glu Asn Thr Glu His 85 90 95 Ala Ile Lys Asp Ser Leu His His Thr Ser Leu Tyr Phe Arg Leu Leu 100 105 110 Arg Gln His Gly Cys Asn Leu Ser Ser Asp Ile Phe Asn Lys Phe Lys 115 120 125 Lys Glu Gly Gly Gly Phe Lys Ala Ser Leu Cys Glu Asp Ala Met Gly 130 135 140 Leu Leu Ser Leu Tyr Glu Ala Val Arg Leu Ser Val Lys Gly Glu Ala 145 150 155 160 Ile Leu Glu Glu Ala Gln Val Phe Ser Ile Ala Asn Leu Lys Ile Leu 165 170 175 Met Glu Arg Val Glu Arg Lys Leu Ala Asp Arg Ile Glu His Ala Leu 180 185 190 Glu Ile Pro Leu Tyr Trp Arg Ala Pro Arg Leu Glu Ala Arg Trp Tyr 195 200 205 Ile Asp Val Tyr Glu Lys Glu Asp Gly Arg Ile Asp Asp Leu Leu Asp 210 215 220 Phe Ala Lys Leu Asp Phe Asn Arg Val Gln Met Leu Tyr Gln Thr Glu 225 230 235 240 Leu Lys Glu Leu Ser Met Trp Trp Glu Leu Leu Gly Leu Pro Ala Lys 245 250 255 Met Gly Phe Phe Arg Asp Arg Leu Leu Glu Asn His Leu Phe Ser Ile 260 265 270 Ala Val Val Val Glu Pro Gln Tyr Ser Gln Cys Arg Val Ala Ile Thr 275 280 285 Lys Ala Thr Val Leu Met Thr Ala Met Asp Asp Phe Tyr Asp Val His 290 295 300 Gly Leu Pro Asp Glu Leu Lys Val Phe Thr Asp Thr Val Asn Arg Trp 305 310 315 320 Asp Leu Glu Gly Ile Asp Gln Leu Pro Glu Tyr Met Lys Leu Tyr Tyr 325 330 335 Leu Ala Leu Tyr Asn Thr Thr Asn Glu Thr Ala Tyr Ile Ile Leu Lys 340 345 350 Glu Lys Gly Phe Asn Ala Thr His Tyr Leu Lys Lys Leu Trp Ala Met 355 360 365 Gln Ser Asn Ala Tyr Phe Arg Glu Ala Gln Trp Phe Asn Ser Gly Tyr 370 375 380 Ile Pro Lys Phe Asp Glu Tyr Leu Asp Asn Ala Leu Val Ser Val Gly 385 390 395 400 Ala Pro Phe Val Leu Gly Leu Ser Tyr Pro Met Ile Gln Gln Gln Ile 405 410 415 Ser Lys Glu Glu Ile Asp Leu Ile Pro Glu Asp Leu Asn Leu Leu Arg 420 425 430 Trp Ala Ser Ile Ile Phe Arg Leu Tyr Asp Asp Leu Ala Thr Ser Lys 435 440 445 Ala Glu Gln Gln Arg Gly Asp Val Pro Lys Ser Ile Gln Cys Tyr Met 450 455 460 His Glu Thr Gly Ser Ser Glu Glu Val Ala Ala Asn His Ile Arg Asp 465 470 475 480 Leu Ile Ser Asp Ala Trp Lys Glu Val Asn Ala Glu Cys Leu Lys Pro 485 490 495 Thr Ser Leu Ser Lys His Tyr Val Gly Val Ala Pro Asn Ser Ala Arg 500 505 510 Ser Gly Val Leu Met Tyr His His Asp Phe Asp Gly Phe Ala Ser Pro 515 520 525 His Gly Arg Thr Asn Ala His Ile Thr Ser Ile Phe Phe Glu Pro Val 530 535 540 Pro Leu Lys Glu Ser Ile Asn Leu Gly 545 550 <210> 32 <211> 553 <212> PRT <213> artificial sequence <220> <223> I291S <400> 32 Met Asp Ser Met Glu Val Arg Arg Ser Ala Ile Tyr His Ser Thr Phe 1 5 10 15 Trp Asp Ile Asp Ser Ile Arg Ala Leu Leu Ala Arg Arg Asp Cys Thr 20 25 30 Ala Ala Ala Ala Leu Ser Pro Asp His His Lys Arg Leu Lys Glu Arg 35 40 45 Ile Gln Arg Arg Leu Gln Asp Ile Thr Gln Pro His His Leu Leu Gly 50 55 60 Leu Ile Asp Ala Val Gln Arg Leu Gly Val Ala Tyr Gln Phe Glu Glu 65 70 75 80 Glu Ile Ser Asp Ala Leu His Gly Leu His Ser Glu Asn Thr Glu His 85 90 95 Ala Ile Lys Asp Ser Leu His His Thr Ser Leu Tyr Phe Arg Leu Leu 100 105 110 Arg Gln His Gly Cys Asn Leu Ser Ser Asp Ile Phe Asn Lys Phe Lys 115 120 125 Lys Glu Gly Gly Gly Phe Lys Ala Ser Leu Cys Glu Asp Ala Met Gly 130 135 140 Leu Leu Ser Leu Tyr Glu Ala Val Arg Leu Ser Val Lys Gly Glu Ala 145 150 155 160 Ile Leu Glu Glu Ala Gln Val Phe Ser Ile Ala Asn Leu Lys Ile Leu 165 170 175 Met Glu Arg Val Glu Arg Lys Leu Ala Asp Arg Ile Glu His Ala Leu 180 185 190 Glu Ile Pro Leu Tyr Trp Arg Ala Pro Arg Leu Glu Ala Arg Trp Tyr 195 200 205 Ile Asp Val Tyr Glu Lys Glu Asp Gly Arg Ile Asp Asp Leu Leu Asp 210 215 220 Phe Ala Lys Leu Asp Phe Asn Arg Val Gln Met Leu Tyr Gln Thr Glu 225 230 235 240 Leu Lys Glu Leu Ser Met Trp Trp Glu Leu Leu Gly Leu Pro Ala Lys 245 250 255 Met Gly Phe Phe Arg Asp Arg Leu Leu Glu Asn His Leu Phe Ser Ile 260 265 270 Ala Val Val Val Glu Pro Gln Tyr Ser Gln Cys Arg Val Ala Ile Thr 275 280 285 Lys Ala Ser Val Leu Met Thr Ala Met Asp Asp Phe Tyr Asp Val His 290 295 300 Gly Leu Pro Asp Glu Leu Lys Val Phe Thr Asp Thr Val Asn Arg Trp 305 310 315 320 Asp Leu Glu Gly Ile Asp Gln Leu Pro Glu Tyr Met Lys Leu Tyr Tyr 325 330 335 Leu Ala Leu Tyr Asn Thr Thr Asn Glu Thr Ala Tyr Ile Ile Leu Lys 340 345 350 Glu Lys Gly Phe Asn Ala Thr His Tyr Leu Lys Lys Leu Trp Ala Met 355 360 365 Gln Ser Asn Ala Tyr Phe Arg Glu Ala Gln Trp Phe Asn Ser Gly Tyr 370 375 380 Ile Pro Lys Phe Asp Glu Tyr Leu Asp Asn Ala Leu Val Ser Val Gly 385 390 395 400 Ala Pro Phe Val Leu Gly Leu Ser Tyr Pro Met Ile Gln Gln Gln Ile 405 410 415 Ser Lys Glu Glu Ile Asp Leu Ile Pro Glu Asp Leu Asn Leu Leu Arg 420 425 430 Trp Ala Ser Ile Ile Phe Arg Leu Tyr Asp Asp Leu Ala Thr Ser Lys 435 440 445 Ala Glu Gln Gln Arg Gly Asp Val Pro Lys Ser Ile Gln Cys Tyr Met 450 455 460 His Glu Thr Gly Ser Ser Glu Glu Val Ala Ala Asn His Ile Arg Asp 465 470 475 480 Leu Ile Ser Asp Ala Trp Lys Glu Val Asn Ala Glu Cys Leu Lys Pro 485 490 495 Thr Ser Leu Ser Lys His Tyr Val Gly Val Ala Pro Asn Ser Ala Arg 500 505 510 Ser Gly Val Leu Met Tyr His His Asp Phe Asp Gly Phe Ala Ser Pro 515 520 525 His Gly Arg Thr Asn Ala His Ile Thr Ser Ile Phe Phe Glu Pro Val 530 535 540 Pro Leu Lys Glu Ser Ile Asn Leu Gly 545 550 <210> 33 <211> 553 <212> PRT <213> Artificial Sequence <220> <223> I291C <400> 33 Met Asp Ser Met Glu Val Arg Arg Ser Ala Ile Tyr His Ser Thr Phe 1 5 10 15 Trp Asp Ile Asp Ser Ile Arg Ala Leu Leu Ala Arg Arg Asp Cys Thr 20 25 30 Ala Ala Ala Ala Leu Ser Pro Asp His His Lys Arg Leu Lys Glu Arg 35 40 45 Ile Gln Arg Arg Leu Gln Asp Ile Thr Gln Pro His His Leu Leu Gly 50 55 60 Leu Ile Asp Ala Val Gln Arg Leu Gly Val Ala Tyr Gln Phe Glu Glu 65 70 75 80 Glu Ile Ser Asp Ala Leu His Gly Leu His Ser Glu Asn Thr Glu His 85 90 95 Ala Ile Lys Asp Ser Leu His His Thr Ser Leu Tyr Phe Arg Leu Leu 100 105 110 Arg Gln His Gly Cys Asn Leu Ser Ser Asp Ile Phe Asn Lys Phe Lys 115 120 125 Lys Glu Gly Gly Gly Phe Lys Ala Ser Leu Cys Glu Asp Ala Met Gly 130 135 140 Leu Leu Ser Leu Tyr Glu Ala Val Arg Leu Ser Val Lys Gly Glu Ala 145 150 155 160 Ile Leu Glu Glu Ala Gln Val Phe Ser Ile Ala Asn Leu Lys Ile Leu 165 170 175 Met Glu Arg Val Glu Arg Lys Leu Ala Asp Arg Ile Glu His Ala Leu 180 185 190 Glu Ile Pro Leu Tyr Trp Arg Ala Pro Arg Leu Glu Ala Arg Trp Tyr 195 200 205 Ile Asp Val Tyr Glu Lys Glu Asp Gly Arg Ile Asp Asp Leu Leu Asp 210 215 220 Phe Ala Lys Leu Asp Phe Asn Arg Val Gln Met Leu Tyr Gln Thr Glu 225 230 235 240 Leu Lys Glu Leu Ser Met Trp Trp Glu Leu Leu Gly Leu Pro Ala Lys 245 250 255 Met Gly Phe Phe Arg Asp Arg Leu Leu Glu Asn His Leu Phe Ser Ile 260 265 270 Ala Val Val Val Glu Pro Gln Tyr Ser Gln Cys Arg Val Ala Ile Thr 275 280 285 Lys Ala Cys Val Leu Met Thr Ala Met Asp Asp Phe Tyr Asp Val His 290 295 300 Gly Leu Pro Asp Glu Leu Lys Val Phe Thr Asp Thr Val Asn Arg Trp 305 310 315 320 Asp Leu Glu Gly Ile Asp Gln Leu Pro Glu Tyr Met Lys Leu Tyr Tyr 325 330 335 Leu Ala Leu Tyr Asn Thr Thr Asn Glu Thr Ala Tyr Ile Ile Leu Lys 340 345 350 Glu Lys Gly Phe Asn Ala Thr His Tyr Leu Lys Lys Leu Trp Ala Met 355 360 365 Gln Ser Asn Ala Tyr Phe Arg Glu Ala Gln Trp Phe Asn Ser Gly Tyr 370 375 380 Ile Pro Lys Phe Asp Glu Tyr Leu Asp Asn Ala Leu Val Ser Val Gly 385 390 395 400 Ala Pro Phe Val Leu Gly Leu Ser Tyr Pro Met Ile Gln Gln Gln Ile 405 410 415 Ser Lys Glu Glu Ile Asp Leu Ile Pro Glu Asp Leu Asn Leu Leu Arg 420 425 430 Trp Ala Ser Ile Ile Phe Arg Leu Tyr Asp Asp Leu Ala Thr Ser Lys 435 440 445 Ala Glu Gln Gln Arg Gly Asp Val Pro Lys Ser Ile Gln Cys Tyr Met 450 455 460 His Glu Thr Gly Ser Ser Glu Glu Val Ala Ala Asn His Ile Arg Asp 465 470 475 480 Leu Ile Ser Asp Ala Trp Lys Glu Val Asn Ala Glu Cys Leu Lys Pro 485 490 495 Thr Ser Leu Ser Lys His Tyr Val Gly Val Ala Pro Asn Ser Ala Arg 500 505 510 Ser Gly Val Leu Met Tyr His His Asp Phe Asp Gly Phe Ala Ser Pro 515 520 525 His Gly Arg Thr Asn Ala His Ile Thr Ser Ile Phe Phe Glu Pro Val 530 535 540 Pro Leu Lys Glu Ser Ile Asn Leu Gly 545 550 <210> 34 <211> 553 <212> PRT <213> artificial sequence <220> <223> I291F <400> 34 Met Asp Ser Met Glu Val Arg Arg Ser Ala Ile Tyr His Ser Thr Phe 1 5 10 15 Trp Asp Ile Asp Ser Ile Arg Ala Leu Leu Ala Arg Arg Asp Cys Thr 20 25 30 Ala Ala Ala Ala Leu Ser Pro Asp His His Lys Arg Leu Lys Glu Arg 35 40 45 Ile Gln Arg Arg Leu Gln Asp Ile Thr Gln Pro His His Leu Leu Gly 50 55 60 Leu Ile Asp Ala Val Gln Arg Leu Gly Val Ala Tyr Gln Phe Glu Glu 65 70 75 80 Glu Ile Ser Asp Ala Leu His Gly Leu His Ser Glu Asn Thr Glu His 85 90 95 Ala Ile Lys Asp Ser Leu His His Thr Ser Leu Tyr Phe Arg Leu Leu 100 105 110 Arg Gln His Gly Cys Asn Leu Ser Ser Asp Ile Phe Asn Lys Phe Lys 115 120 125 Lys Glu Gly Gly Gly Phe Lys Ala Ser Leu Cys Glu Asp Ala Met Gly 130 135 140 Leu Leu Ser Leu Tyr Glu Ala Val Arg Leu Ser Val Lys Gly Glu Ala 145 150 155 160 Ile Leu Glu Glu Ala Gln Val Phe Ser Ile Ala Asn Leu Lys Ile Leu 165 170 175 Met Glu Arg Val Glu Arg Lys Leu Ala Asp Arg Ile Glu His Ala Leu 180 185 190 Glu Ile Pro Leu Tyr Trp Arg Ala Pro Arg Leu Glu Ala Arg Trp Tyr 195 200 205 Ile Asp Val Tyr Glu Lys Glu Asp Gly Arg Ile Asp Asp Leu Leu Asp 210 215 220 Phe Ala Lys Leu Asp Phe Asn Arg Val Gln Met Leu Tyr Gln Thr Glu 225 230 235 240 Leu Lys Glu Leu Ser Met Trp Trp Glu Leu Leu Gly Leu Pro Ala Lys 245 250 255 Met Gly Phe Phe Arg Asp Arg Leu Leu Glu Asn His Leu Phe Ser Ile 260 265 270 Ala Val Val Val Glu Pro Gln Tyr Ser Gln Cys Arg Val Ala Ile Thr 275 280 285 Lys Ala Phe Val Leu Met Thr Ala Met Asp Asp Phe Tyr Asp Val His 290 295 300 Gly Leu Pro Asp Glu Leu Lys Val Phe Thr Asp Thr Val Asn Arg Trp 305 310 315 320 Asp Leu Glu Gly Ile Asp Gln Leu Pro Glu Tyr Met Lys Leu Tyr Tyr 325 330 335 Leu Ala Leu Tyr Asn Thr Thr Asn Glu Thr Ala Tyr Ile Ile Leu Lys 340 345 350 Glu Lys Gly Phe Asn Ala Thr His Tyr Leu Lys Lys Leu Trp Ala Met 355 360 365 Gln Ser Asn Ala Tyr Phe Arg Glu Ala Gln Trp Phe Asn Ser Gly Tyr 370 375 380 Ile Pro Lys Phe Asp Glu Tyr Leu Asp Asn Ala Leu Val Ser Val Gly 385 390 395 400 Ala Pro Phe Val Leu Gly Leu Ser Tyr Pro Met Ile Gln Gln Gln Ile 405 410 415 Ser Lys Glu Glu Ile Asp Leu Ile Pro Glu Asp Leu Asn Leu Leu Arg 420 425 430 Trp Ala Ser Ile Ile Phe Arg Leu Tyr Asp Asp Leu Ala Thr Ser Lys 435 440 445 Ala Glu Gln Gln Arg Gly Asp Val Pro Lys Ser Ile Gln Cys Tyr Met 450 455 460 His Glu Thr Gly Ser Ser Glu Glu Val Ala Ala Asn His Ile Arg Asp 465 470 475 480 Leu Ile Ser Asp Ala Trp Lys Glu Val Asn Ala Glu Cys Leu Lys Pro 485 490 495 Thr Ser Leu Ser Lys His Tyr Val Gly Val Ala Pro Asn Ser Ala Arg 500 505 510 Ser Gly Val Leu Met Tyr His His Asp Phe Asp Gly Phe Ala Ser Pro 515 520 525 His Gly Arg Thr Asn Ala His Ile Thr Ser Ile Phe Phe Glu Pro Val 530 535 540 Pro Leu Lys Glu Ser Ile Asn Leu Gly 545 550 <210> 35 <211> 553 <212> PRT <213> artificial sequence <220> <223> N267L_I291V <400> 35 Met Asp Ser Met Glu Val Arg Arg Ser Ala Ile Tyr His Ser Thr Phe 1 5 10 15 Trp Asp Ile Asp Ser Ile Arg Ala Leu Leu Ala Arg Arg Asp Cys Thr 20 25 30 Ala Ala Ala Ala Leu Ser Pro Asp His His Lys Arg Leu Lys Glu Arg 35 40 45 Ile Gln Arg Arg Leu Gln Asp Ile Thr Gln Pro His His Leu Leu Gly 50 55 60 Leu Ile Asp Ala Val Gln Arg Leu Gly Val Ala Tyr Gln Phe Glu Glu 65 70 75 80 Glu Ile Ser Asp Ala Leu His Gly Leu His Ser Glu Asn Thr Glu His 85 90 95 Ala Ile Lys Asp Ser Leu His His Thr Ser Leu Tyr Phe Arg Leu Leu 100 105 110 Arg Gln His Gly Cys Asn Leu Ser Ser Asp Ile Phe Asn Lys Phe Lys 115 120 125 Lys Glu Gly Gly Gly Phe Lys Ala Ser Leu Cys Glu Asp Ala Met Gly 130 135 140 Leu Leu Ser Leu Tyr Glu Ala Val Arg Leu Ser Val Lys Gly Glu Ala 145 150 155 160 Ile Leu Glu Glu Ala Gln Val Phe Ser Ile Ala Asn Leu Lys Ile Leu 165 170 175 Met Glu Arg Val Glu Arg Lys Leu Ala Asp Arg Ile Glu His Ala Leu 180 185 190 Glu Ile Pro Leu Tyr Trp Arg Ala Pro Arg Leu Glu Ala Arg Trp Tyr 195 200 205 Ile Asp Val Tyr Glu Lys Glu Asp Gly Arg Ile Asp Asp Leu Leu Asp 210 215 220 Phe Ala Lys Leu Asp Phe Asn Arg Val Gln Met Leu Tyr Gln Thr Glu 225 230 235 240 Leu Lys Glu Leu Ser Met Trp Trp Glu Leu Leu Gly Leu Pro Ala Lys 245 250 255 Met Gly Phe Phe Arg Asp Arg Leu Leu Glu Leu His Leu Phe Ser Ile 260 265 270 Ala Val Val Val Glu Pro Gln Tyr Ser Gln Cys Arg Val Ala Ile Thr 275 280 285 Lys Ala Val Val Leu Met Thr Ala Met Asp Asp Phe Tyr Asp Val His 290 295 300 Gly Leu Pro Asp Glu Leu Lys Val Phe Thr Asp Thr Val Asn Arg Trp 305 310 315 320 Asp Leu Glu Gly Ile Asp Gln Leu Pro Glu Tyr Met Lys Leu Tyr Tyr 325 330 335 Leu Ala Leu Tyr Asn Thr Thr Asn Glu Thr Ala Tyr Ile Ile Leu Lys 340 345 350 Glu Lys Gly Phe Asn Ala Thr His Tyr Leu Lys Lys Leu Trp Ala Met 355 360 365 Gln Ser Asn Ala Tyr Phe Arg Glu Ala Gln Trp Phe Asn Ser Gly Tyr 370 375 380 Ile Pro Lys Phe Asp Glu Tyr Leu Asp Asn Ala Leu Val Ser Val Gly 385 390 395 400 Ala Pro Phe Val Leu Gly Leu Ser Tyr Pro Met Ile Gln Gln Gln Ile 405 410 415 Ser Lys Glu Glu Ile Asp Leu Ile Pro Glu Asp Leu Asn Leu Leu Arg 420 425 430 Trp Ala Ser Ile Ile Phe Arg Leu Tyr Asp Asp Leu Ala Thr Ser Lys 435 440 445 Ala Glu Gln Gln Arg Gly Asp Val Pro Lys Ser Ile Gln Cys Tyr Met 450 455 460 His Glu Thr Gly Ser Ser Glu Glu Val Ala Ala Asn His Ile Arg Asp 465 470 475 480 Leu Ile Ser Asp Ala Trp Lys Glu Val Asn Ala Glu Cys Leu Lys Pro 485 490 495 Thr Ser Leu Ser Lys His Tyr Val Gly Val Ala Pro Asn Ser Ala Arg 500 505 510 Ser Gly Val Leu Met Tyr His His Asp Phe Asp Gly Phe Ala Ser Pro 515 520 525 His Gly Arg Thr Asn Ala His Ile Thr Ser Ile Phe Phe Glu Pro Val 530 535 540 Pro Leu Lys Glu Ser Ile Asn Leu Gly 545 550 <210> 36 <211> 553 <212> PRT <213> Synthetic Sequence <220> <223> N267S_I291T <400> 36 Met Asp Ser Met Glu Val Arg Arg Ser Ala Ile Tyr His Ser Thr Phe 1 5 10 15 Trp Asp Ile Asp Ser Ile Arg Ala Leu Leu Ala Arg Arg Asp Cys Thr 20 25 30 Ala Ala Ala Ala Leu Ser Pro Asp His His Lys Arg Leu Lys Glu Arg 35 40 45 Ile Gln Arg Arg Leu Gln Asp Ile Thr Gln Pro His His Leu Leu Gly 50 55 60 Leu Ile Asp Ala Val Gln Arg Leu Gly Val Ala Tyr Gln Phe Glu Glu 65 70 75 80 Glu Ile Ser Asp Ala Leu His Gly Leu His Ser Glu Asn Thr Glu His 85 90 95 Ala Ile Lys Asp Ser Leu His His Thr Ser Leu Tyr Phe Arg Leu Leu 100 105 110 Arg Gln His Gly Cys Asn Leu Ser Ser Asp Ile Phe Asn Lys Phe Lys 115 120 125 Lys Glu Gly Gly Gly Phe Lys Ala Ser Leu Cys Glu Asp Ala Met Gly 130 135 140 Leu Leu Ser Leu Tyr Glu Ala Val Arg Leu Ser Val Lys Gly Glu Ala 145 150 155 160 Ile Leu Glu Glu Ala Gln Val Phe Ser Ile Ala Asn Leu Lys Ile Leu 165 170 175 Met Glu Arg Val Glu Arg Lys Leu Ala Asp Arg Ile Glu His Ala Leu 180 185 190 Glu Ile Pro Leu Tyr Trp Arg Ala Pro Arg Leu Glu Ala Arg Trp Tyr 195 200 205 Ile Asp Val Tyr Glu Lys Glu Asp Gly Arg Ile Asp Asp Leu Leu Asp 210 215 220 Phe Ala Lys Leu Asp Phe Asn Arg Val Gln Met Leu Tyr Gln Thr Glu 225 230 235 240 Leu Lys Glu Leu Ser Met Trp Trp Glu Leu Leu Gly Leu Pro Ala Lys 245 250 255 Met Gly Phe Phe Arg Asp Arg Leu Leu Glu Ser His Leu Phe Ser Ile 260 265 270 Ala Val Val Val Glu Pro Gln Tyr Ser Gln Cys Arg Val Ala Ile Thr 275 280 285 Lys Ala Thr Val Leu Met Thr Ala Met Asp Asp Phe Tyr Asp Val His 290 295 300 Gly Leu Pro Asp Glu Leu Lys Val Phe Thr Asp Thr Val Asn Arg Trp 305 310 315 320 Asp Leu Glu Gly Ile Asp Gln Leu Pro Glu Tyr Met Lys Leu Tyr Tyr 325 330 335 Leu Ala Leu Tyr Asn Thr Thr Asn Glu Thr Ala Tyr Ile Ile Leu Lys 340 345 350 Glu Lys Gly Phe Asn Ala Thr His Tyr Leu Lys Lys Leu Trp Ala Met 355 360 365 Gln Ser Asn Ala Tyr Phe Arg Glu Ala Gln Trp Phe Asn Ser Gly Tyr 370 375 380 Ile Pro Lys Phe Asp Glu Tyr Leu Asp Asn Ala Leu Val Ser Val Gly 385 390 395 400 Ala Pro Phe Val Leu Gly Leu Ser Tyr Pro Met Ile Gln Gln Gln Ile 405 410 415 Ser Lys Glu Glu Ile Asp Leu Ile Pro Glu Asp Leu Asn Leu Leu Arg 420 425 430 Trp Ala Ser Ile Ile Phe Arg Leu Tyr Asp Asp Leu Ala Thr Ser Lys 435 440 445 Ala Glu Gln Gln Arg Gly Asp Val Pro Lys Ser Ile Gln Cys Tyr Met 450 455 460 His Glu Thr Gly Ser Ser Glu Glu Val Ala Ala Asn His Ile Arg Asp 465 470 475 480 Leu Ile Ser Asp Ala Trp Lys Glu Val Asn Ala Glu Cys Leu Lys Pro 485 490 495 Thr Ser Leu Ser Lys His Tyr Val Gly Val Ala Pro Asn Ser Ala Arg 500 505 510 Ser Gly Val Leu Met Tyr His His Asp Phe Asp Gly Phe Ala Ser Pro 515 520 525 His Gly Arg Thr Asn Ala His Ile Thr Ser Ile Phe Phe Glu Pro Val 530 535 540 Pro Leu Lys Glu Ser Ile Asn Leu Gly 545 550 <210> 37 <211> 553 <212> PRT <213> artificial sequence <220> <223> N267S_I291T <400> 37 Met Asp Ser Met Glu Val Arg Arg Ser Ala Ile Tyr His Ser Thr Phe 1 5 10 15 Trp Asp Ile Asp Ser Ile Arg Ala Leu Leu Ala Arg Arg Asp Cys Thr 20 25 30 Ala Ala Ala Ala Leu Ser Pro Asp His His Lys Arg Leu Lys Glu Arg 35 40 45 Ile Gln Arg Arg Leu Gln Asp Ile Thr Gln Pro His His Leu Leu Gly 50 55 60 Leu Ile Asp Ala Val Gln Arg Leu Gly Val Ala Tyr Gln Phe Glu Glu 65 70 75 80 Glu Ile Ser Asp Ala Leu His Gly Leu His Ser Glu Asn Thr Glu His 85 90 95 Ala Ile Lys Asp Ser Leu His His Thr Ser Leu Tyr Phe Arg Leu Leu 100 105 110 Arg Gln His Gly Cys Asn Leu Ser Ser Asp Ile Phe Asn Lys Phe Lys 115 120 125 Lys Glu Gly Gly Gly Phe Lys Ala Ser Leu Cys Glu Asp Ala Met Gly 130 135 140 Leu Leu Ser Leu Tyr Glu Ala Val Arg Leu Ser Val Lys Gly Glu Ala 145 150 155 160 Ile Leu Glu Glu Ala Gln Val Phe Ser Ile Ala Asn Leu Lys Ile Leu 165 170 175 Met Glu Arg Val Glu Arg Lys Leu Ala Asp Arg Ile Glu His Ala Leu 180 185 190 Glu Ile Pro Leu Tyr Trp Arg Ala Pro Arg Leu Glu Ala Arg Trp Tyr 195 200 205 Ile Asp Val Tyr Glu Lys Glu Asp Gly Arg Ile Asp Asp Leu Leu Asp 210 215 220 Phe Ala Lys Leu Asp Phe Asn Arg Val Gln Met Leu Tyr Gln Thr Glu 225 230 235 240 Leu Lys Glu Leu Ser Met Trp Trp Glu Leu Leu Gly Leu Pro Ala Lys 245 250 255 Met Gly Phe Phe Arg Asp Arg Leu Leu Glu Thr His Leu Phe Ser Ile 260 265 270 Ala Val Val Val Glu Pro Gln Tyr Ser Gln Cys Arg Val Ala Ile Thr 275 280 285 Lys Ala Thr Val Leu Met Thr Ala Met Asp Asp Phe Tyr Asp Val His 290 295 300 Gly Leu Pro Asp Glu Leu Lys Val Phe Thr Asp Thr Val Asn Arg Trp 305 310 315 320 Asp Leu Glu Gly Ile Asp Gln Leu Pro Glu Tyr Met Lys Leu Tyr Tyr 325 330 335 Leu Ala Leu Tyr Asn Thr Thr Asn Glu Thr Ala Tyr Ile Ile Leu Lys 340 345 350 Glu Lys Gly Phe Asn Ala Thr His Tyr Leu Lys Lys Leu Trp Ala Met 355 360 365 Gln Ser Asn Ala Tyr Phe Arg Glu Ala Gln Trp Phe Asn Ser Gly Tyr 370 375 380 Ile Pro Lys Phe Asp Glu Tyr Leu Asp Asn Ala Leu Val Ser Val Gly 385 390 395 400 Ala Pro Phe Val Leu Gly Leu Ser Tyr Pro Met Ile Gln Gln Gln Ile 405 410 415 Ser Lys Glu Glu Ile Asp Leu Ile Pro Glu Asp Leu Asn Leu Leu Arg 420 425 430 Trp Ala Ser Ile Ile Phe Arg Leu Tyr Asp Asp Leu Ala Thr Ser Lys 435 440 445 Ala Glu Gln Gln Arg Gly Asp Val Pro Lys Ser Ile Gln Cys Tyr Met 450 455 460 His Glu Thr Gly Ser Ser Glu Glu Val Ala Ala Asn His Ile Arg Asp 465 470 475 480 Leu Ile Ser Asp Ala Trp Lys Glu Val Asn Ala Glu Cys Leu Lys Pro 485 490 495 Thr Ser Leu Ser Lys His Tyr Val Gly Val Ala Pro Asn Ser Ala Arg 500 505 510 Ser Gly Val Leu Met Tyr His His Asp Phe Asp Gly Phe Ala Ser Pro 515 520 525 His Gly Arg Thr Asn Ala His Ile Thr Ser Ile Phe Phe Glu Pro Val 530 535 540 Pro Leu Lys Glu Ser Ile Asn Leu Gly 545 550 <210> 38 <211> 553 <212> PRT <213> artificial sequence <220> <223> N267S_I291S <400> 38 Met Asp Ser Met Glu Val Arg Arg Ser Ala Ile Tyr His Ser Thr Phe 1 5 10 15 Trp Asp Ile Asp Ser Ile Arg Ala Leu Leu Ala Arg Arg Asp Cys Thr 20 25 30 Ala Ala Ala Ala Leu Ser Pro Asp His His Lys Arg Leu Lys Glu Arg 35 40 45 Ile Gln Arg Arg Leu Gln Asp Ile Thr Gln Pro His His Leu Leu Gly 50 55 60 Leu Ile Asp Ala Val Gln Arg Leu Gly Val Ala Tyr Gln Phe Glu Glu 65 70 75 80 Glu Ile Ser Asp Ala Leu His Gly Leu His Ser Glu Asn Thr Glu His 85 90 95 Ala Ile Lys Asp Ser Leu His His Thr Ser Leu Tyr Phe Arg Leu Leu 100 105 110 Arg Gln His Gly Cys Asn Leu Ser Ser Asp Ile Phe Asn Lys Phe Lys 115 120 125 Lys Glu Gly Gly Gly Phe Lys Ala Ser Leu Cys Glu Asp Ala Met Gly 130 135 140 Leu Leu Ser Leu Tyr Glu Ala Val Arg Leu Ser Val Lys Gly Glu Ala 145 150 155 160 Ile Leu Glu Glu Ala Gln Val Phe Ser Ila Asn Leu Lys Ile Leu 165 170 175 Met Glu Arg Val Glu Arg Lys Leu Ala Asp Arg Ile Glu His Ala Leu 180 185 190 Glu Ile Pro Leu Tyr Trp Arg Ala Pro Arg Leu Glu Ala Arg Trp Tyr 195 200 205 Ile Asp Val Tyr Glu Lys Glu Asp Gly Arg Ile Asp Asp Leu Leu Asp 210 215 220 Phe Ala Lys Leu Asp Phe Asn Arg Val Gln Met Leu Tyr Gln Thr Glu 225 230 235 240 Leu Lys Glu Leu Ser Met Trp Trp Glu Leu Leu Gly Leu Pro Ala Lys 245 250 255 Met Gly Phe Phe Arg Asp Arg Leu Leu Glu Ser His Leu Phe Ser Ile 260 265 270 Ala Val Val Val Glu Pro Gln Tyr Ser Gln Cys Arg Val Ala Ile Thr 275 280 285 Lys Ala Ser Val Leu Met Thr Ala Met Asp Asp Phe Tyr Asp Val His 290 295 300 Gly Leu Pro Asp Glu Leu Lys Val Phe Thr Asp Thr Val Asn Arg Trp 305 310 315 320 Asp Leu Glu Gly Ile Asp Gln Leu Pro Glu Tyr Met Lys Leu Tyr Tyr 325 330 335 Leu Ala Leu Tyr Asn Thr Thr Asn Glu Thr Ala Tyr Ile Ile Leu Lys 340 345 350 Glu Lys Gly Phe Asn Ala Thr His Tyr Leu Lys Lys Leu Trp Ala Met 355 360 365 Gln Ser Asn Ala Tyr Phe Arg Glu Ala Gln Trp Phe Asn Ser Gly Tyr 370 375 380 Ile Pro Lys Phe Asp Glu Tyr Leu Asp Asn Ala Leu Val Ser Val Gly 385 390 395 400 Ala Pro Phe Val Leu Gly Leu Ser Tyr Pro Met Ile Gln Gln Gln Ile 405 410 415 Ser Lys Glu Glu Ile Asp Leu Ile Pro Glu Asp Leu Asn Leu Leu Arg 420 425 430 Trp Ala Ser Ile Ile Phe Arg Leu Tyr Asp Asp Leu Ala Thr Ser Lys 435 440 445 Ala Glu Gln Gln Arg Gly Asp Val Pro Lys Ser Ile Gln Cys Tyr Met 450 455 460 His Glu Thr Gly Ser Ser Glu Glu Val Ala Ala Asn His Ile Arg Asp 465 470 475 480 Leu Ile Ser Asp Ala Trp Lys Glu Val Asn Ala Glu Cys Leu Lys Pro 485 490 495 Thr Ser Leu Ser Lys His Tyr Val Gly Val Ala Pro Asn Ser Ala Arg 500 505 510 Ser Gly Val Leu Met Tyr His His Asp Phe Asp Gly Phe Ala Ser Pro 515 520 525 His Gly Arg Thr Asn Ala His Ile Thr Ser Ile Phe Phe Glu Pro Val 530 535 540 Pro Leu Lys Glu Ser Ile Asn Leu Gly 545 550 <210> 39 <211> 553 <212> PRT <213> Artificial sequence <220> <223> N267S_I291C <400> 39 Met Asp Ser Met Glu Val Arg Arg Ser Ala Ile Tyr His Ser Thr Phe 1 5 10 15 Trp Asp Ile Asp Ser Ile Arg Ala Leu Leu Ala Arg Arg Asp Cys Thr 20 25 30 Ala Ala Ala Ala Leu Ser Pro Asp His His Lys Arg Leu Lys Glu Arg 35 40 45 Ile Gln Arg Arg Leu Gln Asp Ile Thr Gln Pro His His Leu Leu Gly 50 55 60 Leu Ile Asp Ala Val Gln Arg Leu Gly Val Ala Tyr Gln Phe Glu Glu 65 70 75 80 Glu Ile Ser Asp Ala Leu His Gly Leu His Ser Glu Asn Thr Glu His 85 90 95 Ala Ile Lys Asp Ser Leu His His Thr Ser Leu Tyr Phe Arg Leu Leu 100 105 110 Arg Gln His Gly Cys Asn Leu Ser Ser Asp Ile Phe Asn Lys Phe Lys 115 120 125 Lys Glu Gly Gly Gly Phe Lys Ala Ser Leu Cys Glu Asp Ala Met Gly 130 135 140 Leu Leu Ser Leu Tyr Glu Ala Val Arg Leu Ser Val Lys Gly Glu Ala 145 150 155 160 Ile Leu Glu Glu Ala Gln Val Phe Ser Ile Ala Asn Leu Lys Ile Leu 165 170 175 Met Glu Arg Val Glu Arg Lys Leu Ala Asp Arg Ile Glu His Ala Leu 180 185 190 Glu Ile Pro Leu Tyr Trp Arg Ala Pro Arg Leu Glu Ala Arg Trp Tyr 195 200 205 Ile Asp Val Tyr Glu Lys Glu Asp Gly Arg Ile Asp Asp Leu Leu Asp 210 215 220 Phe Ala Lys Leu Asp Phe Asn Arg Val Gln Met Leu Tyr Gln Thr Glu 225 230 235 240 Leu Lys Glu Leu Ser Met Trp Trp Glu Leu Leu Gly Leu Pro Ala Lys 245 250 255 Met Gly Phe Phe Arg Asp Arg Leu Leu Glu Ser His Leu Phe Ser Ile 260 265 270 Ala Val Val Val Glu Pro Gln Tyr Ser Gln Cys Arg Val Ala Ile Thr 275 280 285 Lys Ala Cys Val Leu Met Thr Ala Met Asp Asp Phe Tyr Asp Val His 290 295 300 Gly Leu Pro Asp Glu Leu Lys Val Phe Thr Asp Thr Val Asn Arg Trp 305 310 315 320 Asp Leu Glu Gly Ile Asp Gln Leu Pro Glu Tyr Met Lys Leu Tyr Tyr 325 330 335 Leu Ala Leu Tyr Asn Thr Thr Asn Glu Thr Ala Tyr Ile Ile Leu Lys 340 345 350 Glu Lys Gly Phe Asn Ala Thr His Tyr Leu Lys Lys Leu Trp Ala Met 355 360 365 Gln Ser Asn Ala Tyr Phe Arg Glu Ala Gln Trp Phe Asn Ser Gly Tyr 370 375 380 Ile Pro Lys Phe Asp Glu Tyr Leu Asp Asn Ala Leu Val Ser Val Gly 385 390 395 400 Ala Pro Phe Val Leu Gly Leu Ser Tyr Pro Met Ile Gln Gln Gln Ile 405 410 415 Ser Lys Glu Glu Ile Asp Leu Ile Pro Glu Asp Leu Asn Leu Leu Arg 420 425 430 Trp Ala Ser Ile Ile Phe Arg Leu Tyr Asp Asp Leu Ala Thr Ser Lys 435 440 445 Ala Glu Gln Gln Arg Gly Asp Val Pro Lys Ser Ile Gln Cys Tyr Met 450 455 460 他的谷氨酸-苏氨酸-甘氨酸-丝氨酸-丝氨酸-谷氨酸-谷氨酸-缬氨酸-丙氨酸-丙氨酸-组氨酸-异亮氨酸-精氨酸-天冬氨酸 465 470 475 480 亮氨酸-异亮氨酸-丝氨酸-天冬氨酸-丙氨酸-色氨酸-赖氨酸-谷氨酸-缬氨酸-天冬酰胺-丙氨酸-谷氨酸-半胱氨酸-亮氨酸-赖氨酸-脯氨酸 485 490 495 苏氨酸-丝氨酸-亮氨酸-丝氨酸-赖氨酸-组氨酸-酪氨酸-缬氨酸-甘氨酸-缬氨酸-丙氨酸-脯氨酸-天冬酰胺-丝氨酸-丙氨酸-精氨酸 500 505 510 丝氨酸-甘氨酸-缬氨酸-亮氨酸-甲硫氨酸-酪氨酸-组氨酸-组氨酸-天冬氨酸-苯丙氨酸-天冬氨酸-甘氨酸-苯丙氨酸-丙氨酸-丝氨酸-脯氨酸 515 520 525 组氨酸-甘氨酸-精氨酸-苏氨酸-天冬酰胺-丙氨酸-组氨酸-异亮氨酸-苏氨酸-丝氨酸-异亮氨酸-苯丙氨酸-苯丙氨酸-谷氨酸-脯氨酸-缬氨酸 530 535 540 脯氨酸-亮氨酸-赖氨酸-谷氨酸-丝氨酸-异亮氨酸-天冬酰胺-亮氨酸-甘氨酸 545 550 <210> 40 <211> 553 <212> PRT <213> Artificial Sequence <220> <223> N267L_I291S <400> 40 甲硫氨酸-天冬氨酸-丝氨酸-甲硫氨酸-谷氨酸-缬氨酸-精氨酸-精氨酸-丝氨酸-丙氨酸-异亮氨酸-酪氨酸-组氨酸-丝氨酸-苏氨酸-苯丙氨酸 1 5 10 15 色氨酸-天冬氨酸-异亮氨酸-天冬氨酸-丝氨酸-异亮氨酸-精氨酸-丙氨酸-亮氨酸-亮氨酸-丙氨酸-精氨酸-精氨酸-天冬氨酸-半胱氨酸-苏氨酸 20 25 30 丙氨酸-丙氨酸-丙氨酸-丙氨酸-亮氨酸-丝氨酸-脯氨酸-天冬氨酸-组氨酸-组氨酸-赖氨酸-精氨酸-亮氨酸-赖氨酸-谷氨酸-精氨酸 35 40 45 Ile Gln Arg Arg Leu Gln Asp Ile Thr Gln Pro His His Leu Leu Gly 50 55 60 Leu Ile Asp Ala Val Gln Arg Leu Gly Val Ala Tyr Gln Phe Glu Glu 65 70 75 80 Glu Ile Ser Asp Ala Leu His Gly Leu His Ser Glu Asn Thr Glu His 85 90 95 Ala Ile Lys Asp Ser Leu His His Thr Ser Leu Tyr Phe Arg Leu Leu 100 105 110 Arg Gln His Gly Cys Asn Leu Ser Ser Asp Ile Phe Asn Lys Phe Lys 115 120 125 Lys Glu Gly Gly Gly Phe Lys Ala Ser Leu Cys Glu Asp Ala Met Gly 130 135 140 Leu Leu Ser Leu Tyr Glu Ala Val Arg Leu Ser Val Lys Gly Glu Ala 145 150 155 160 Ile Leu Glu Glu Ala Gln Val Phe Ser Ile Ala Asn Leu Lys Ile Leu 165 170 175 Met Glu Arg Val Glu Arg Lys Leu Ala Asp Arg Ile Glu His Ala Leu 180 185 190 Glu Ile Pro Leu Tyr Trp Arg Ala Pro Arg Leu Glu Ala Arg Trp Tyr 195 200 205 Ile Asp Val Tyr Glu Lys Glu Asp Gly Arg Ile Asp Asp Leu Leu Asp 210 215 220 Phe Ala Lys Leu Asp Phe Asn Arg Val Gln Met Leu Tyr Gln Thr Glu 225 230 235 240 Leu Lys Glu Leu Ser Met Trp Trp Glu Leu Leu Gly Leu Pro Ala Lys 245 250 255 Met Gly Phe Phe Arg Asp Arg Leu Leu Glu Leu His Leu Phe Ser Ile 260 265 270 Ala Val Val Val Glu Pro Gln Tyr Ser Gln Cys Arg Val Ala Ile Thr 275 280 285 Lys Ala Ser Val Leu Met Thr Ala Met Asp Asp Phe Tyr Asp Val His 290 295 300 Gly Leu Pro Asp Glu Leu Lys Val Phe Thr Asp Thr Val Asn Arg Trp 305 310 315 320 Asp Leu Glu Gly Ile Asp Gln Leu Pro Glu Tyr Met Lys Leu Tyr Tyr 325 330 335 Leu Ala Leu Tyr Asn Thr Thr Asn Glu Thr Ala Tyr Ile Ile Leu Lys 340 345 350 Glu Lys Gly Phe Asn Ala Thr His Tyr Leu Lys Lys Leu Trp Ala Met 355 360 365 Gln Ser Asn Ala Tyr Phe Arg Glu Ala Gln Trp Phe Asn Ser Gly Tyr 370 375 380 Ile Pro Lys Phe Asp Glu Tyr Leu Asp Asn Ala Leu Val Ser Val Gly 385 390 395 400 Ala Pro Phe Val Leu Gly Leu Ser Tyr Pro Met Ile Gln Gln Gln Ile 405 410 415 Ser Lys Glu Glu Ile Asp Leu Ile Pro Glu Asp Leu Asn Leu Leu Arg 420 425 430 Trp Ala Ser Ile Ile Phe Arg Leu Tyr Asp Asp Leu Ala Thr Ser Lys 435 440 445 Ala Glu Gln Gln Arg Gly Asp Val Pro Lys Ser Ile Gln Cys Tyr Met 450 455 460 His Glu Thr Gly Ser Ser Glu Glu Val Ala Ala Asn His Ile Arg Asp 465 470 475 480 Leu Ile Ser Asp Ala Trp Lys Glu Val Asn Ala Glu Cys Leu Lys Pro 485 490 495 Thr Ser Leu Ser Lys His Tyr Val Gly Val Ala Pro Asn Ser Ala Arg 500 505 510 Ser Gly Val Leu Met Tyr His His Asp Phe Asp Gly Phe Ala Ser Pro 515 520 525 His Gly Arg Thr Asn Ala His Ile Thr Ser Ile Phe Phe Glu Pro Val 530 535 540 Pro Leu Lys Glu Ser Ile Asn Leu Gly 545 550 <210> 41 <211> 553 <212> PRT <213> Artificial Sequence <220> <223> Var1_291T <400> 41 Met Asp Ser Val Glu Leu Arg Arg Thr Gly Ile Tyr His Ser Thr Phe 1 5 10 15 Trp Asp Ile Asp Ser Ile Arg Ala Leu Leu Ala Lys Arg Asp Cys Thr 20 25 30 Ala Ala Ala Ala Val Ser Pro Glu His His Lys Lys Leu Lys Glu Arg[[ID=o36]] 35 40 45 Ile Gln Lys Arg Leu Asn Asp Ile Thr Gln Pro His His Leu Leu Gly 50 55 60 Leu Ile Asp Ala Val Gln Arg Leu Gly Val Ala Tyr Gln Phe Glu Glu 65 70 75 80 Glu Ile Ser Glu Ala Leu His Gly Val His Ser Glu Asn Thr Asp His 85 90 95 Ala Ile Lys Asp Ser Leu His His Thr Ser Leu Tyr Phe Arg Leu Leu 100 105 110 Arg Gln His Gly Cys Asn Leu Ser Ser Asp Ile Phe Asn Lys Phe Lys 115 120 125 Lys Asp Gly Gly Gly Phe Arg Ala Thr Leu Cys Glu Asp Ala Met Gly 130 135 140 Leu Leu Ser Leu Tyr Glu Ala Val Arg Leu Ser Val Lys Gly Glu Ala 145 150 155 160 Ile Leu Glu Glu Ala Gln Val Phe Ser Ile Ala Asn Leu Lys Ile Val 165 170 175 Met Glu Arg Ile Glu Arg Lys Leu Ala Asp Arg Ile Glu His Ala Leu 180 185 190 Glu Ile Pro Leu Tyr Trp Arg Ala Pro Arg Leu Glu Ala Arg Trp Tyr 195 200 205 Ile Asp Val Tyr Glu Lys Asp Asp Gly Arg Ile Asp Asp Leu Leu Asp 210 215 220 Phe Ala Lys Leu Asp Phe Asn Arg Val Gln Met Leu Tyr Gln Thr Glu 225 230 235 240 Leu Lys Glu Leu Ser Met Trp Trp Glu Leu Leu Gly Leu Pro Ala Lys 245 250 255 Met Gly Phe Phe Arg Asp Arg Leu Leu Glu Asn His Leu Phe Ser Ile 260 265 270 Ala Met Leu Val Glu Pro Gln Tyr Ser Gln Cys Arg Val Ala Leu Thr 275 280 285 Lys Ala Thr Ile Leu Ile Ser Ala Met Asp Asp Phe Tyr Asp Ile His 290 295 300 Gly Leu Pro Asp Glu Leu Arg Val Phe Thr Asp Thr Val Asn Arg Trp 305 310 315 320 Asp Leu Glu Gly Ile Glu Gln Leu Pro Asp Tyr Met Lys Val Tyr Tyr 325 330 335 Leu Ala Leu Tyr Asn Thr Thr Asn Glu Thr Ala Trp Ile Ile Leu Lys 340 345 350 Glu Lys Gly Phe Asn Ala Thr His Tyr Leu Lys Lys Val Trp Ala Met 355 360 365 Gln Ser Asn Ala Tyr Phe Arg Glu Ala Gln Trp Phe Asn Thr Gly Tyr 370 375 380 Ile Pro Lys Phe Asp Glu Tyr Leu Glu Asn Ala Leu Val Ser Val Gly 385 390 395 400 Ala Pro Phe Leu Leu Gly Val Ser Tyr Pro Leu Leu Gln Asn Gln Ile 405 410 415 Ser Lys Glu Glu Leu Glu Leu Leu Pro Glu Asp Leu Asn Leu Ile Arg 420 425 430 Trp Ala Ser Met Ile Phe Arg Leu Trp Asp Asp Leu Ala Thr Ser Lys 435 440 445 Ala Glu Gln Gln Arg Gly Glu Val Pro Lys Thr Ile Gln Cys Tyr Met 450 455 460 His Glu Thr Gly Ser Ser Glu Glu Val Ala Ala Asn His Ile Arg Asp 465 470 475 480 Met Ile Ser Glu Ala Trp Lys Glu Val Asn Ala Glu Cys Leu Lys Pro 485 490 495 Thr Ser Val Ser Arg His Tyr Val Gly Val Ala Pro Asn Thr Ala Arg 500 505 510 Ser Gly Leu Leu Met Tyr His His Asp Phe Asp Gly Phe Ala Ser Pro 515 520 525 His Gly Lys Thr Asn Ala His Ile Thr Ser Ile Phe Phe Asp Pro Val 530 535 540 Pro Leu Lys Glu Ser Ile Asn Leu Gly 545 550 <210> 42 <211> 553 <212> PRT <213> Artificial sequence <220> <223> Var2_291T <400> 42 Met Asp Ser Met Glu Val Arg Lys Ser Ala Met Tyr His Ser Thr Phe 1 5 10 15 Trp Asp Ile Asp Ser Ile Arg Ala Ile Leu Ala Lys Arg Asp Cys Thr 20 25 30 Ala Ala Ala Ala Leu Ser Pro Glu His His Lys Arg Leu Lys Glu Arg 35 40 45 Ile Gln Lys Arg Leu Asn Asp Val Thr Asn Pro His His Leu Leu Ala 50 55 60 Leu Ile Asp Ala Val Gln Arg Leu Gly Val Ala Tyr Gln Phe Glu Glu 65 70 75 80 Glu Ile Ser Asp Ala Leu His Gly Leu His Ser Glu Gln Ser Asp His 85 90 95 Gly Leu Lys Asp Ser Leu His His Thr Ser Leu Tyr Phe Arg Leu Leu 100 105 110 Arg Gln His Gly Cys Asn Leu Ser Ser Asp Ile Phe Asn Lys Phe Lys 115 120 125 Lys Glu Gly Gly Gly Phe Lys Ala Thr Leu Cys Glu Asp Ala Met Gly 130 135 140 Leu Leu Ser Leu Tyr Glu Ala Val Arg Leu Ser Ile Arg Gly Glu Ala 145 150 155 160 Ile Leu Glu Glu Ala Gln Ile Phe Ser Ile Ala Asn Leu Lys Ile Ile 165 170 175 Ile Glu Arg Leu Glu Arg Lys Leu Ala Glu Lys Ile Glu His Ala Leu 180 185 190 Glu Ile Pro Leu Tyr Trp Arg Ala Pro Arg Leu Glu Ala Arg Trp Tyr 195 200 205 Ile Glu Ile Tyr Glu Lys Asp Asp Gly Lys Ile Glu Asp Leu Leu Asp 210 215 220 Phe Ala Lys Leu Asp Phe Asn Arg Met Gln Met Leu Tyr Gln Thr Glu 225 230 235 240 Leu Lys Asp Leu Thr Met Trp Trp Glu Leu Leu Gly Leu Pro Ala Lys 245 250 255 Met Gly Phe Phe Arg Asp Arg Leu Leu Glu Asn His Leu Phe Ser Ile 260 265 270 Ala Met Ile Val Glu Pro Gln Tyr Ser Asn Cys Arg Ile Gly Ile Thr 275 280 285 Lys Ala Thr Val Leu Val Ser Ala Met Asp Asp Phe Tyr Asp Val His 290 295 300 Gly Leu Pro Asp Glu Leu Arg Val Phe Thr Glu Thr Val Asn Arg Trp 305 310 315 320 Asp Leu Glu Gly Ile Asp Gln Leu Pro Asp Tyr Met Lys Val Tyr Tyr 325 330 335 Leu Gly Leu Tyr Asn Thr Thr Asn Glu Thr Ala Tyr Leu Val Val Lys 340 345 350 Asp Lys Gly Tyr Asn Ala Thr His Tyr Leu Arg Lys Val Trp Ala Met 355 360 365 Gln Ser Asn Ala Tyr Phe Arg Glu Ala Gln Trp Phe Asn Thr Gly Tyr 370 375 380 Ile Pro Lys Phe Asp Glu Tyr Leu Asp Asn Ala Val Val Ser Val Gly 385 390 395 400 Ala Pro Phe Val Val Gly Leu Thr Tyr Pro Leu Leu Gln Gln Gln Leu 405 410 415 Thr Lys Glu Glu Ile Glu Leu Ile Pro Glu Asp Leu Asn Leu Leu Arg 420 425 430 Trp Ala Ser Val Ile Phe Arg Leu Tyr Asp Asp Leu Ala Thr Thr Lys 435 440 445 Ala Glu Gln Gln Arg Gly Asp Met Pro Lys Ser Ile Gln Cys Tyr Met 450 455 460 His Glu Thr Gly Thr Ser Glu Glu Val Ala Ala Asn His Ile Lys Glu 465 470 475 480 Ile Ile Ser Glu Ala Trp Lys Asp Val Asn Ala Glu Cys Leu Lys Pro 485 490 495 Ser Ser Val Ser Lys His Tyr Leu Gly Leu Ala Pro Asn Ser Ala Arg 500 505 510 Ser Ala Met Val Ile Tyr His His Asp Phe Asp Gly Phe Ala Thr Pro 515 520 525 His Gly Lys Thr Asn Ala His Ile Thr Ser Ile Phe Phe Asp Pro Val 530 535 540 Pro Ile Lys Asp Ser Ile Asn Leu Gly 545 550 <210> 43 <211> 553 <212> PRT <213> Synthetic sequence <220> <223> Var3_291T <400> 43 Met Asp Ser Leu Asp Leu Arg Arg Ser Ala Ile Tyr His Ser Thr Phe 1 5 10 15 Trp Asp Ile Asp Ser Ile Lys Ala Leu Ile Ala Lys Arg Asp Cys Ser 20 25 30 Ala Ala Ala Ala Val Ser Pro Glu His His Arg Lys Leu Lys Glu Lys 35 40 45 Ile Gln Lys Lys Leu Gln Glu Ile Ser Gln Pro His His Leu Leu Gly 50 55 60 Leu Ile Asp Ala Val Gln Arg Leu Gly Val Ala Tyr Gln Phe Glu Asp 65 70 75 80 Glu Ile Ser Asp Ala Leu His Gly Met His Thr Asp Gln Thr Asp His 85 90 95 Ala Val Lys Glu Ser Leu His His Thr Ser Leu Trp Phe Arg Leu Leu 100 105 110 Arg Gln His Gly Cys Asn Leu Ser Ser Asp Ile Phe Asn Arg Phe Arg 115 120 125 Lys Glu Gly Gly Gly Phe Lys Ala Ser Ile Cys Glu Asp Ala Met Gly 130 135 140 Leu Leu Ser Leu Tyr Glu Ala Val Arg Leu Ser Ile Lys Gly Glu Ala 145 150 155 160 Ile Leu Glu Glu Ala Gln Val Phe Ser Ile Ala Asn Leu Arg Ile Met 165 170 175 Leu Glu Lys Val Glu Lys Lys Leu Ala Asp Arg Ile Asp His Ala Leu 180 185 190 Glu Leu Pro Ile Phe Trp Arg Ala Pro Arg Leu Glu Ala Arg Trp Tyr 195 200 205 Ile Glu Val Tyr Glu Arg Asp Glu Gly Arg Ile Glu Glu Leu Leu Asp 210 215 220 Phe Ala Lys Leu Asp Phe Asn Arg Val Gln Met Val Tyr Gln Thr Glu 225 230 235 240 Leu Lys Glu Leu Thr Met Trp Trp Glu Leu Val Gly Leu Pro Ala Arg 245 250 255 Met Gly Phe Phe Arg Asp Arg Ile Leu Glu Asn His Leu Phe Ser Leu 260 265 270 Ala Ile Leu Val Glu Pro Gln Tyr Ser Asn Cys Arg Val Ala Leu Thr 275 280 285 Lys Ala Thr Ile Leu Ile Ser Ala Met Asp Asp Phe Tyr Asp Val His 290 295 300 Gly Leu Pro Asp Glu Leu Lys Leu Phe Thr Glu Thr Val Gln Arg Trp 305 310 315 320 Asp Ile Asp Ala Ile Asp Asn Leu Pro Glu Tyr Met Lys Leu Phe Tyr 325 330 335 Val Ala Leu Tyr Asn Thr Thr Asn Glu Thr Ala Trp Leu Ile Val Lys 340 345 350 Glu Arg Gly Phe Asn Ala Thr His Tyr Leu Arg Lys Val Trp Ala Met 355 360 365 Gln Ser Asn Ala Tyr Phe Arg Glu Ala Gln Trp Phe Gln Ser Gly Tyr 370 375 380 Val Pro Lys Phe Asp Glu Tyr Leu Glu Asn Ala Val Ile Ser Ile Gly 385 390 395 400 Gly Pro Phe Ile Leu Gly Leu Ser Tyr Pro Leu Ile Gln Asn Gln Val 405 410 415 Thr Lys Glu Glu Ile Asp Leu Leu Pro Glu Asp Leu Asn Leu Val Lys 420 425 430 Trp Ala Ser Val Ile Phe Arg Leu Tyr Asp Asp Leu Gly Thr Ser Lys 435 440 445 Ala Glu Asn Gln Arg Gly Asp Leu Pro Lys Ser Ile Gln Cys Tyr Met 450 455 460 His Glu Thr Gly Ser Ser Glu Glu Val Ala Ala Asn His Ile Lys Glu 465 470 475 480 Met Ile Ser Glu Ala Trp Lys Asp Val Asn Ala Glu Cys Leu Arg Pro 485 490 495 Ser Thr Ile Ser Arg His Phe Leu Gly Val Ala Pro Asn Ser Ala Arg 500 505 510 Ser Ala Val Ile Met Tyr His His Asp Phe Asp Gly Phe Ala Ser Pro 515 520 525 His Ala Lys Thr Asn Ala His Ile Thr Ser Ile Phe Tyr Asp Pro Val 530 535 540 Pro Leu Arg Glu Ser Ile Asn Leu Gly 545 550 <210> 44 <211> 553 <212> PRT <213> Artificial sequence <220> <223> Var1_291S <400> 44 Met Asp Ser Met Glu Ile Arg Arg Thr Ala Ile Tyr His Ser Thr Phe 1 5 10 15 Trp Asp Ile Asp Ser Ile Arg Ala Met Leu Ala Arg Arg Asp Cys Ser 20 25 30 Ala Gly Ala Ala Val Ser Pro Asp His His Lys Lys Leu Lys Glu Lys 35 40 45 Ile Gln Arg Arg Leu Gln Asp Ile Thr Gln Pro His His Leu Leu Gly 50 55 60 Leu Ile Asp Ala Val Gln Arg Leu Gly Val Ala Tyr Gln Phe Glu Asp 65 70 75 80 Glu Ile Ser Glu Ala Leu His Gly Val His Ser Glu Asn Thr Glu His 85 90 95 Ala Ile Lys Asp Ser Leu His His Thr Ser Leu Tyr Phe Arg Leu Leu 100 105 110 Arg Gln His Gly Cys Asn Leu Ser Ser Asp Ile Phe Asn Lys Phe Lys 115 120 125 Lys Glu Gly Gly Gly Phe Lys Ala Thr Leu Cys Glu Asp Ala Met Gly 130 135 140 Leu Leu Ser Leu Tyr Glu Ala Val Arg Leu Ser Ile Lys Gly Glu Ala 145 150 155 160 Ile Leu Glu Glu Ala Gln Val Phe Ser Ila Asn Leu Arg Ile Met 165 170 175 Val Glu Lys Leu Glu Arg Lys Leu Ala Asp Arg Ile Glu His Ala Leu 180 185 190 Glu Ile Pro Leu Tyr Trp Arg Ala Pro Arg Leu Glu Ala Arg Phe Tyr 195 200 205 Ile Glu Ile Tyr Glu Lys Glu Asp Gly Lys Ile Asp Glu Leu Leu Asp 210 215 220 Phe Ala Lys Leu Asp Phe Asn Arg Val Gln Met Ile Tyr Gln Thr Glu 225 230 235 240 Leu Lys Glu Leu Thr Met Trp Trp Glu Leu Leu Gly Leu Pro Ala Lys 245 250 255 Met Gly Phe Phe Arg Asp Arg Leu Leu Glu Asn His Leu Phe Ser Met 260 265 270 Ala Val Ile Val Glu Pro Gln Phe Ser Gln Cys Arg Leu Ala Leu Thr 275 280 285 Lys Ala Ser Val Leu Met Thr Ala Met Asp Asp Phe Tyr Asp Val His 290 295 300 Gly Leu Pro Asp Glu Leu Lys Val Phe Thr Asp Thr Val Asn Arg Trp 305 310 315 320 Asp Leu Glu Gly Met Asp Gln Leu Pro Glu Tyr Met Lys Leu Tyr Tyr 325 330 335 Leu Ala Leu Tyr Asn Thr Thr Asn Glu Thr Ala Tyr Ile Leu Val Lys 340 345 350 Glu Lys Gly Phe Asn Ala Thr His Tyr Leu Lys Arg Val Trp Ala Met 355 360 365 Gln Ser Asn Gly Tyr Tyr Arg Glu Ala Gln Trp Phe Asn Ser Gly Tyr 370 375 380 Leu Pro Lys Phe Asp Glu Tyr Leu Asp Asn Ala Leu Val Ser Val Gly 385 390 395 400 Ala Pro Phe Leu Leu Gly Leu Ser Tyr Pro Met Met Gln Gln Gln Ile 405 410 415 Thr Lys Glu Glu Leu Asp Leu Ile Pro Glu Glu Leu Asn Leu Val Arg 420 425 430 Trp Ala Ser Ile Ile Phe Arg Leu Tyr Asp Asp Leu Ala Thr Ser Lys 435 440 445 Ala Glu Gln Gln Arg Gly Asp Val Pro Lys Ser Ile Gln Cys Tyr Met 450 455 460 His Glu Thr Gly Ser Ser Glu Glu Ile Ala Ala Asn His Ile Lys Glu 465 470 475 480 Met Ile Ser Glu Ala Trp Lys Asp Val Asn Ala Glu Cys Leu Lys Pro 485 490 495 Thr Ser Val Ser Lys His Phe Val Gly Val Ala Pro Asn Ser Ala Arg 500 505 510 Ser Gly Val Val Met Tyr His His Asp Phe Asp Gly Phe Ala Ser Pro 515 520 525 His Gly Arg Thr Asn Ala His Ile Thr Ser Ile Phe Phe Glu Pro Val 530 535 540 Pro Ile Arg Asp Ser Ile Asn Leu Gly 545 550 <210> 45 <211> 553 <212> PRT <213> artificial sequence <220> <223> Var2_291S <400> 45 Met Asp Ser Val Glu Val Arg Arg Thr Gly Ile Tyr His Ser Thr Phe 1 5 10 15 Trp Asp Ile Asp Ser Ile Arg Ala Leu Val Ala Arg Arg Asp Cys Ser 20 25 30 Ala Gly Ala Ala Leu Ser Pro Asp His His Lys Arg Leu Lys Glu Arg 35 40 45 Ile Gln Arg Arg Leu Gln Glu Ile Ser Gln Pro His His Leu Leu Gly 50 55 60 Leu Ile Asp Ala Val Gln Arg Leu Gly Val Ala Tyr Gln Phe Glu Glu 65 70 75 80 Glu Ile Ser Asp Ala Leu His Ala Ile His Thr Asp Gln Ser Asp His 85 90 95 Ala Ile Arg Asp Ser Leu His His Thr Ser Leu Tyr Phe Arg Leu Leu 100 105 110 Arg Gln His Gly Cys Asn Leu Ser Ser Asp Ile Phe Gln Lys Phe Lys 115 120 125 Lys Glu Gly Gly Gly Phe Lys Ala Ser Leu Cys Glu Asp Ala Met Gly 130 135 140 Leu Leu Ser Leu Tyr Glu Ala Val Arg Leu Ser Ile Lys Gly Glu Ala 145 150 155 160 To Leu Glu Glu Ala Gln Val Phe Ser To Ala Asn Leu Arg To Leu 165 170 175 Leu Glu Arg Ile Glu Arg Lys Leu Ala Glu Lys Ile Glu His Ala Leu 180 185 190 Glu Leu Pro Leu Tyr Trp Arg Ala Pro Arg Leu Glu Ala Arg Trp Tyr 195 200 205 Ile Glu Val Tyr Glu Lys Asp Asp Ala Lys Ile Glu Asp Leu Leu Asp 210 215 220 Phe Ala Lys Leu Asp Phe Asn Arg Leu Gln Met Met Tyr Gln Thr Glu 225 230 235 240 Leu Lys Glu Leu Ser Met Trp Trp Glu Leu Leu Gly Leu Pro Ala Arg 245 250 255 Met Gly Phe Phe Arg Asp Arg Ile Leu Glu Asn His Leu Phe Ser Ile 260 265 270 Pathway Met Val Val Glu Pro Asn Tyr Ser Gln Cys Arg Pathway Leu Thr 275 280 285 Lys Ala Ser Ile Leu Val Ser Ala Met Asp Asp Tyr Tyr Asp Val His 290 295 300 Gly Leu Pro Asp Glu Leu Lys Ile Phe Thr Asp Thr Val Asn Arg Trp 305 310 315 320 Asp Leu Glu Gly Ile Glu Gln Leu Pro Asp Tyr Met Lys Ile Tyr Tyr 325 330 335 Leu Ala Leu Tyr Asn Thr Thr Asn Glu Thr Ala Tyr Ile Leu Leu Lys 340 345 350 Glu Lys Gly Phe Asn Ala Ser His Tyr Leu Arg Lys Leu Trp Ala Val 355 360 365 Gln Ser Asn Ala Tyr Phe Lys Glu Ala Gln Trp Phe Asn Ser Gly Tyr 370 375 380 Val Pro Lys Phe Asp Glu Tyr Leu Glu Asn Ala Leu Val Ser Ile Gly 385 390 395 400 Ala Pro Phe Ile Val Ala Val Thr Tyr Pro Met Ile Gln Asn Asn Val 405 410 415 Thr Lys Glu Glu Ile Glu Leu Leu Pro Glu Asp Leu Gln Leu Leu Lys 420 425 430 Trp Ala Ser Val Ile Phe Arg Leu Tyr Asp Asp Leu Gly Thr Ser Lys 435 440 445 Ala Glu Gln Gln Arg Gly Glu Val Pro Lys Ser Ile Gln Cys Tyr Met 450 455 460 His Glu Thr Gly Ser Ser Glu Glu Val Ala Ala Asn His Ile Lys Glu 465 470 475 480 Ile Ile Ser Asp Ala Trp Lys Glu Val Asn Ala Glu Cys Leu Arg Pro 485 490 495 Thr Thr Leu Ser Arg His Tyr Val Gly Val Ala Pro Asn Ser Ala Arg 500 505 510 Ser Ala Val Val Ile Tyr His His Asp Phe Asp Gly Phe Ala Thr Pro 515 520 525 His Ala Lys Thr Gln Ala His Ile Thr Ser Ile Phe Phe Glu Pro Val 530 535 540 Pro Ile Lys Glu Ser Ile Asn Leu Gly 545 550 <210> 46 <211> 553 <212> PRT <213> Synthetic Sequence <220> <223> Var3_291S <400> 46 Met Asp Ser Leu Asp Leu Arg Arg Ser Ala Ile Tyr His Ser Thr Phe 1 5 10 15 Trp Asp Ile Asp Ser Ile Lys Ala Leu Ile Ala Lys Arg Asp Cys Ser 20 25 30 Ala Ala Ala Ala Val Ser Pro Glu His His Arg Lys Leu Lys Glu Lys 35 40 45 Ile Gln Lys Lys Leu Gln Glu Ile Ser Gln Pro His His Leu Leu Gly 50 55 60 Leu Ile Asp Ala Val Gln Arg Leu Gly Val Ala Tyr Gln Phe Glu Asp 65 70 75 80 Glu Ile Ser Asp Ala Leu His Gly Met His Thr Asp Gln Thr Asp His 85 90 95 Ala Val Lys Glu Ser Leu His His Thr Ser Leu Trp Phe Arg Leu Leu 100 105 110 Arg Gln His Gly Cys Asn Leu Ser Ser Asp Ile Phe Asn Arg Phe Arg 115 120 125 Lys Glu Gly Gly Gly Phe Lys Ala Ser Ile Cys Glu Asp Ala Met Gly 130 135 140 Leu Leu Ser Leu Tyr Glu Ala Val Arg Leu Ser Ile Lys Gly Glu Ala 145 150 155 160 Ile Leu Glu Glu Ala Gln Val Phe Ser Ile Ala Asn Leu Arg Ile Met 165 170 175 Leu Glu Lys Val Glu Lys Lys Leu Ala Asp Arg Ile Asp His Ala Leu 180 185 190 Glu Leu Pro Ile Phe Trp Arg Ala Pro Arg Leu Glu Ala Arg Trp Tyr 195 200 205 Ile Glu Val Tyr Glu Arg Asp Glu Gly Arg Ile Glu Glu Leu Leu Asp 210 215 220 Phe Ala Lys Leu Asp Phe Asn Arg Val Gln Met Val Tyr Gln Thr Glu 225 230 235 240 Leu Lys Glu Leu Thr Met Trp Trp Glu Leu Val Gly Leu Pro Ala Arg 245 250 255 Met Gly Phe Phe Arg Asp Arg Ile Leu Glu Asn His Leu Phe Ser Leu 260 265 270 Ala Ile Leu Val Glu Pro Gln Tyr Ser Asn Cys Arg Val Ala Leu Thr 275 280 285 Lys Ala Ser Ile Leu Ile Ser Ala Met Asp Asp Phe Tyr Asp Val His 290 295 300 Gly Leu Pro Asp Glu Leu Lys Leu Phe Thr Glu Thr Val Gln Arg Trp 305 310 315 320 Asp Ile Asp Ala Ile Asp Asn Leu Pro Glu Tyr Met Lys Leu Phe Tyr 325 330 335 Val Ala Leu Tyr Asn Thr Thr Asn Glu Thr Ala Trp Leu Ile Val Lys 340 345 350 Glu Arg Gly Phe Asn Ala Thr His Tyr Leu Arg Lys Val Trp Ala Met 355 360 365 Gln Ser Asn Ala Tyr Phe Arg Glu Ala Gln Trp Phe Gln Ser Gly Tyr 370 375 380 Val Pro Lys Phe Asp Glu Tyr Leu Glu Asn Ala Val Ile Ser Ile Gly 385 390 395 400 Gly Pro Phe Ile Leu Gly Leu Ser Tyr Pro Leu Ile Gln Asn Gln Val 405 410 415 Thr Lys Glu Glu Ile Asp Leu Leu Pro Glu Asp Leu Asn Leu Val Lys 420 425 430 Trp Ala Ser Val Ile Phe Arg Leu Tyr Asp Asp Leu Gly Thr Ser Lys 435 440 445 Ala Glu Asn Gln Arg Gly Asp Leu Pro Lys Ser Ile Gln Cys Tyr Met 450 455 460 His Glu Thr Gly Ser Ser Glu Glu Val Ala Ala Asn His Ile Lys Glu 465 470 475 480 Met Ile Ser Glu Ala Trp Lys Asp Val Asn Ala Glu Cys Leu Arg Pro 485 490 495 Ser Thr Ile Ser Arg His Phe Leu Gly Val Ala Pro Asn Ser Ala Arg 500 505 510 Ser Ala Val Ile Met Tyr His His Asp Phe Asp Gly Phe Ala Ser Pro 515 520 525 His Ala Lys Thr Asn Ala His Ile Thr Ser Ile Phe Tyr Asp Pro Val 530 535 540 Pro Leu Arg Glu Ser Ile Asn Leu Gly 545 550 <210> 47 <211> 553 <212> PRT <213> artificial sequence <220> <223> Var1_291C <400> 47 Met Asp Ser Met Glu Val Arg Arg Thr Ala Ile Tyr His Ser Thr Phe 1 5 10 15 Trp Asp Ile Asp Ser Ile Arg Ala Leu Ile Ala Arg Arg Asp Cys Thr 20 25 30 Ala Ala Ala Ala Ile Ser Pro Asp His His Lys Arg Leu Lys Glu Arg 35 40 45 Ile Gln Lys Arg Leu Gln Asp Ile Thr Gln Pro His His Leu Leu Gly 50 55 60 Leu Ile Asp Ala Val Gln Arg Leu Gly Val Ala Tyr Gln Phe Glu Glu 65 70 75 80 Glu Ile Ser Asp Ala Leu His Gly Leu His Ser Glu Asn Thr Glu His 85 90 95 Gly Ile Arg Asp Ser Leu His His Thr Ser Leu Tyr Phe Arg Leu Leu 100 105 110 Arg Gln His Gly Cys Asn Leu Ser Ser Asp Ile Phe Asn Lys Phe Lys 115 120 125 Lys Asp Gly Gly Gly Phe Lys Ala Thr Leu Cys Asp Asp Ala Met Gly 130 135 140 Leu Leu Ser Leu Tyr Glu Ala Val Arg Leu Ser Val Lys Gly Glu Ala 145 150 155 160 Ile Leu Glu Glu Ala Gln Val Phe Ser Ile Ala Asn Leu Lys Ile Leu 165 170 175 Leu Asp Arg Leu Asp Arg Lys Leu Ala Glu Arg Ile Glu His Ala Leu 180 185 190 Asp Leu Pro Leu Phe Trp Arg Ala Pro Arg Leu Glu Ala Arg Trp Tyr 195 200 205 Ile Asp Leu Tyr Glu Lys Asp Asp Gly Arg Ile Asp Asp Leu Leu Asp 210 215 220 Phe Ala Lys Leu Asp Phe Asn Arg Val Gln Leu Ile Tyr Gln Thr Glu 225 230 235 240 Leu Lys Glu Leu Ser Met Trp Trp Glu Leu Val Gly Leu Pro Ala Lys 245 250 255 Met Gly Phe Phe Arg Asp Arg Ile Leu Glu Asn His Leu Phe Thr Val 260 265 270 Ala Ile Val Val Glu Pro Gln Tyr Ser Gln Cys Arg Leu Ala Ile Thr 275 280 285 Lys Ala Cys Ile Leu Met Thr Ala Met Asp Asp Phe Tyr Asp Val His 290 295 300 Gly Leu Pro Asp Glu Leu Lys Leu Phe Thr Asp Thr Val Gln Arg Trp 305 310 315 320 Asp Ile Glu Ala Met Glu Gln Leu Pro Glu Tyr Met Lys Ile Tyr Tyr 325 330 335 Leu Ala Leu Tyr Asn Thr Thr Asn Glu Thr Ala Tyr Val Ile Leu Lys 340 345 350 Glu Lys Gly Phe Asn Ala Thr His Tyr Leu Lys Lys Val Trp Ala Met 355 360 365 Gln Ser Asn Ala Tyr Phe Arg Glu Ala Gln Trp Phe Asn Ser Gly Tyr 370 375 380 Val Pro Lys Phe Asp Glu Tyr Leu Asp Asn Ala Leu Val Ser Val Gly 385 390 395 400 Ala Pro Phe Val Leu Gly Val Thr Tyr Pro Val Ile Gln Gln Gln Leu 405 410 415 Ser Lys Glu Glu Leu Asp Leu Val Pro Glu Asp Leu Asn Leu Val Lys 420 425 430 Trp Ala Ser Ile Ile Phe Arg Leu Tyr Asp Asp Leu Ala Thr Ser Lys 435 440 445 Ala Glu Gln Gln Arg Gly Asp Val Pro Lys Ser Ile Gln Cys Tyr Met 450 455 460 His Glu Thr Gly Ser Ser Glu Glu Val Ala Ala Asn His Ile Arg Asp 465 470 475 480 Met Ile Ser Asp Ala Trp Lys Asp Val Asn Ala Glu Cys Leu Lys Pro 485 490 495 Thr Ser Leu Ser Lys His Tyr Leu Gly Leu Ala Pro Asn Ser Ala Arg 500 505 510 Ser Gly Val Leu Leu Tyr His His Asp Phe Asp Gly Phe Ala Thr Pro 515 520 525 His Ala Arg Thr Asn Gly His Ile Thr Ser Ile Phe Phe Asp Pro Val 530 535 540 Pro Leu Lys Glu Ser Ile Asn Leu Gly 545 550 <210> 48 <211> 553 <212> PRT <213> artificial sequence <220> <223> Var2_291C <400> 48 Met Asp Ser Val Glu Val Arg Arg Thr Gly Ile Tyr His Ser Thr Phe 1 5 10 15 Trp Asp Ile Asp Ser Ile Arg Ala Leu Val Ala Arg Arg Asp Cys Ser 20 25 30 Ala Gly Ala Ala Leu Ser Pro Asp His His Lys Arg Leu Lys Glu Arg 35 40 45 Ile Gln Arg Arg Leu Gln Glu Ile Ser Gln Pro His His Leu Leu Gly 50 55 60 Leu Ile Asp Ala Val Gln Arg Leu Gly Val Ala Tyr Gln Phe Glu Glu 65 70 75 80 Glu Ile Ser Asp Ala Leu His Ala Ile His Thr Asp Gln Ser Asp His 85 90 95 Ala Ile Arg Asp Ser Leu His His Thr Ser Leu Tyr Phe Arg Leu Leu 100 105 110 Arg Gln His Gly Cys Asn Leu Ser Ser Asp Ile Phe Gln Lys Phe Lys 115 120 125 Lys Glu Gly Gly Gly Phe Lys Ala Ser Leu Cys Glu Asp Ala Met Gly 130 135 140 Leu Leu Ser Leu Tyr Glu Ala Val Arg Leu Ser Ile Lys Gly Glu Ala 145 150 155 160 To Leu Glu Glu Ala Gln Val Phe Ser To Ala Asn Leu Arg To Leu 165 170 175 Leu Glu Arg Ile Glu Arg Lys Leu Ala Glu Lys Ile Glu His Ala Leu 180 185 190 Glu Leu Pro Leu Tyr Trp Arg Ala Pro Arg Leu Glu Ala Arg Trp Tyr 195 200 205 Ile Glu Val Tyr Glu Lys Asp Asp Ala Lys Ile Glu Asp Leu Leu Asp 210 215 220 Phe Ala Lys Leu Asp Phe Asn Arg Leu Gln Met Met Tyr Gln Thr Glu 225 230 235 240 Leu Lys Glu Leu Ser Met Trp Trp Glu Leu Leu Gly Leu Pro Ala Arg 245 250 255 Met Gly Phe Phe Arg Asp Arg Ile Leu Glu Asn His Leu Phe Ser Ile 260 265 270 Pathway Met Val Val Glu Pro Asn Tyr Ser Gln Cys Arg Pathway Leu Thr 275 280 285 Lys Ala Cys Ile Leu Val Ser Ala Met Asp Asp Tyr Tyr Asp Val His 290 295 300 Gly Leu Pro Asp Glu Leu Lys Ile Phe Thr Asp Thr Val Asn Arg Trp 305 310 315 320 Asp Leu Glu Gly Ile Glu Gln Leu Pro Asp Tyr Met Lys Ile Tyr Tyr 325 330 335 Leu Ala Leu Tyr Asn Thr Thr Asn Glu Thr Ala Tyr Ile Leu Leu Lys 340 345 350 Glu Lys Gly Phe Asn Ala Ser His Tyr Leu Arg Lys Leu Trp Ala Val 355 360 365 Gln Ser Asn Ala Tyr Phe Lys Glu Ala Gln Trp Phe Asn Ser Gly Tyr 370 375 380 Val Pro Lys Phe Asp Glu Tyr Leu Glu Asn Ala Leu Val Ser Ile Gly 385 390 395 400 Ala Pro Phe Ile Val Ala Val Thr Tyr Pro Met Ile Gln Asn Asn Val 405 410 415 Thr Lys Glu Glu Ile Glu Leu Leu Pro Glu Asp Leu Gln Leu Leu Lys 420 425 430 Trp Ala Ser Val Ile Phe Arg Leu Tyr Asp Asp Leu Gly Thr Ser Lys 435 440 445 Ala Glu Gln Gln Arg Gly Glu Val Pro Lys Ser Ile Gln Cys Tyr Met 450 455 460 His Glu Thr Gly Ser Ser Glu Glu Val Ala Ala Asn His Ile Lys Glu 465 470 475 480 Ile Ile Ser Asp Ala Trp Lys Glu Val Asn Ala Glu Cys Leu Arg Pro 485 490 495 Thr Thr Leu Ser Arg His Tyr Val Gly Val Ala Pro Asn Ser Ala Arg 500 505 510 Ser Ala Val Val Ile Tyr His His Asp Phe Asp Gly Phe Ala Thr Pro 515 520 525 His Ala Lys Thr Gln Ala His Ile Thr Ser Ile Phe Phe Glu Pro Val 530 535 540 Pro Ile Lys Glu Ser Ile Asn Leu Gly 545 550 <210> 49 <211> 553 <212> PRT <213> Artificial Sequence <220> <223> Var3_291C <400> 49 Met Asp Ser Ile Asp Ile Arg Arg Thr Ala Ile Tyr His Ser Thr Phe 1 5 10 15 Trp Asp Ile Asp Ser Ile Arg Ala Leu Leu Ala Lys Arg Glu Cys Ser 20 25 30 Ala Ala Gly Ala Met Ser Pro Asp His His Lys Lys Leu Lys Glu Lys 35 40 45 Ile Gln Lys Lys Leu Gln Glu Leu Thr Gln Pro His His Leu Leu Ala 50 55 60 Leu Ile Asp Ala Val Gln Arg Leu Gly Val Ala Tyr Gln Phe Glu Asp 65 70 75 80 Glu Ile Ser Glu Ala Leu His Gly Ile His Ser Asp Asn Ser Glu His 85 90 95 Ala Ile Arg Glu Ser Leu His His Thr Ser Leu Tyr Phe Arg Leu Leu 100 105 110 Arg Gln His Gly Cys Asn Leu Ser Thr Asp Ile Phe Asn Lys Phe Arg 115 120 125 Lys Asp Gly Gly Gly Phe Lys Ala Thr Leu Cys Glu Asp Ala Met Gly 130 135 140 Leu Leu Ser Leu Tyr Glu Ala Val Arg Leu Ser Val Arg Gly Glu Ala 145 150 155 160 Ile Leu Glu Glu Ala Gln Ile Phe Ser Ile Ala Asn Leu Arg Ile Val 165 170 175 Ile Glu Lys Ile Glu Arg Lys Leu Ala Asp Lys Ile Glu His Ala Leu 180 185 190 Glu Leu Pro Leu Tyr Trp Arg Ala Pro Arg Leu Glu Ala Arg Trp Tyr 195 200 205 Ile Asp Ile Tyr Glu Lys Asp Glu Gly Arg Ile Glu Asp Leu Leu Asp 210 215 220 Phe Ala Lys Leu Asp Phe Asn Arg Val Gln Leu Ile Tyr Gln Thr Glu 225 230 235 240 Leu Arg Glu Leu Thr Met Trp Trp Glu Leu Leu Gly Leu Pro Ala Lys 245 250 255 Met Gly Phe Phe Arg Asp Arg Leu Leu Glu Asn His Leu Phe Ser Val 260 265 270 Ala Met Leu Val Glu Pro Asn Tyr Ser Gln Cys Arg Val Ala Leu Thr 275 280 285 Lys Ala Cys Val Leu Val Ser Ala Met Asp Asp Phe Tyr Asp Ile His 290 295 300 Gly Leu Pro Asp Glu Leu Arg Leu Phe Thr Glu Thr Val Gln Arg Trp 305 310 315 320 Asp Leu Asp Ala Leu Asp Gln Leu Pro Glu Tyr Met Lys Ile Phe Tyr 325 330 335 Met Ala Leu Tyr Asn Thr Thr Asn Glu Thr Ala Tyr Leu Met Leu Lys 340 345 350 Asp Lys Gly Phe Asn Ala Ser His Tyr Leu Lys Lys Val Trp Ala Leu 355 360 365 Gln Ser Asn Gly Tyr Phe Arg Glu Ala Gln Trp Tyr Asn Ser Gly Tyr 370 375 380 Val Pro Lys Phe Asp Glu Tyr Leu Asp Asn Ala Val Ile Ser Val Gly 385 390 395 400 Gly Pro Phe Leu Leu Gly Val Ser Tyr Pro Leu Leu Gln Asn Gln Leu 405 410 415 Thr Lys Glu Glu Met Glu Leu Leu Pro Glu Asp Leu Asn Leu Val Lys 420 425 430 Phe Ala Ser Val Ile Phe Arg Leu Tyr Asp Asp Leu Ala Thr Ser Lys 435 440 445 Ala Glu Gln Asn Arg Gly Asp Leu Pro Lys Ser Ile Gln Cys Tyr Met 450 455 460 His Glu Thr Gly Ser Ser Glu Glu Met Ala Ala Gln His Ile Lys Asp 465 470 475 480 Ile Ile Ser Glu Gly Trp Lys Asp Val Asn Ala Glu Cys Leu Arg Pro 485 490 495 Ser Ser Val Ser Arg His Phe Leu Gly Leu Ala Pro Asn Ser Ala Arg 500 505 510 Ser Gly Val Ile Val Tyr His His Asp Phe Asp Gly Phe Ala Ser Pro 515 520 525 His Ala Arg Ser Asn Ala His Ile Thr Ser Ile Phe Phe Glu Pro Val 530 535 540 Pro Ile Arg Glu Ser Ile Asn Leu Gly 545 550 <210> 50 <211> 553 <212> PRT <213> artificial sequence <220> <223> Var_267S_291T <400> 50 Met Asp Ser Ile Asp Ile Arg Arg Thr Ala Ile Tyr His Ser Thr Phe 1 5 10 15 Trp Asp Ile Asp Ser Ile Arg Ala Leu Leu Ala Lys Arg Glu Cys Ser 20 25 30 Ala Ala Gly Ala Met Ser Pro Asp His His Lys Lys Leu Lys Glu Lys 35 40 45 Ile Gln Lys Lys Leu Gln Glu Leu Thr Gln Pro His His Leu Leu Ala 50 55 60 Leu Ile Asp Ala Val Gln Arg Leu Gly Val Ala Tyr Gln Phe Glu Asp 65 70 75 80 Glu Ile Ser Glu Ala Leu His Gly Ile His Ser Asp Asn Ser Glu His 85 90 95 Ala Ile Arg Glu Ser Leu His His Thr Ser Leu Tyr Phe Arg Leu Leu 100 105 110 Arg Gln His Gly Cys Asn Leu Ser Thr Asp Ile Phe Asn Lys Phe Arg 115 120 125 Lys Asp Gly Gly Gly Phe Lys Ala Thr Leu Cys Glu Asp Ala Met Gly 130 135 140 Leu Leu Ser Leu Tyr Glu Ala Val Arg Leu Ser Val Arg Gly Glu Ala 145 150 155 160 Ile Leu Glu Glu Ala Gln Ile Phe Ser Ile Ala Asn Leu Arg Ile Val 165 170 175 Ile Glu Lys Ile Glu Arg Lys Leu Ala Asp Lys Ile Glu His Ala Leu 180 185 190 Glu Leu Pro Leu Tyr Trp Arg Ala Pro Arg Leu Glu Ala Arg Trp Tyr 195 200 205 Ile Asp Ile Tyr Glu Lys Asp Glu Gly Arg Ile Glu Asp Leu Leu Asp 210 215 220 Phe Ala Lys Leu Asp Phe Asn Arg Val Gln Leu Ile Tyr Gln Thr Glu 225 230 235 240 Leu Arg Glu Leu Thr Met Trp Trp Glu Leu Leu Gly Leu Pro Ala Lys 245 250 255 Met Gly Phe Phe Arg Asp Arg Leu Leu Glu Ser His Leu Phe Ser Val 260 265 270 Ala Met Leu Val Glu Pro Asn Tyr Ser Gln Cys Arg Val Ala Leu Thr 275 280 285 Lys Ala Thr Val Leu Val Ser Ala Met Asp Asp Phe Tyr Asp Ile His 290 295 300 Gly Leu Pro Asp Glu Leu Arg Leu Phe Thr Glu Thr Val Gln Arg Trp 305 310 315 320 Asp Leu Asp Ala Leu Asp Gln Leu Pro Glu Tyr Met Lys Ile Phe Tyr 325 330 335 Met Ala Leu Tyr Asn Thr Thr Asn Glu Thr Ala Tyr Leu Met Leu Lys 340 345 350 Asp Lys Gly Phe Asn Ala Ser His Tyr Leu Lys Lys Val Trp Ala Leu 355 360 365 Gln Ser Asn Gly Tyr Phe Arg Glu Ala Gln Trp Tyr Asn Ser Gly Tyr 370 375 380 Val Pro Lys Phe Asp Glu Tyr Leu Asp Asn Ala Val Ile Ser Val Gly 385 390 395 400 Gly Pro Phe Leu Leu Gly Val Ser Tyr Pro Leu Leu Gln Asn Gln Leu 405 410 415 Thr Lys Glu Glu Met Glu Leu Leu Pro Glu Asp Leu Asn Leu Val Lys 420 425 430 Phe Ala Ser Val Ile Phe Arg Leu Tyr Asp Asp Leu Ala Thr Ser Lys 435 440 445 Ala Glu Gln Asn Arg Gly Asp Leu Pro Lys Ser Ile Gln Cys Tyr Met 450 455 460 His Glu Thr Gly Ser Ser Glu Glu Met Ala Ala Gln His Ile Lys Asp 465 470 475 480 Ile Ile Ser Glu Gly Trp Lys Asp Val Asn Ala Glu Cys Leu Arg Pro 485 490 495 Ser Ser Val Ser Arg His Phe Leu Gly Leu Ala Pro Asn Ser Ala Arg 500 505 510 Ser Gly Val Ile Val Tyr His His Asp Phe Asp Gly Phe Ala Ser Pro 515 520 525 His Ala Arg Ser Asn Ala His Ile Thr Ser Ile Phe Phe Glu Pro Val 530 535 540 Pro Ile Arg Glu Ser Ile Asn Leu Gly 545 550 <210> 51 <211> 553 <212> PRT <213> artificial sequence <220> <223> Var_267S_291S <400> 51 Met Asp Ser Ile Asp Ile Arg Arg Thr Ala Ile Tyr His Ser Thr Phe 1 5 10 15 Trp Asp Ile Asp Ser Ile Arg Ala Leu Leu Ala Lys Arg Glu Cys Ser 20 25 30 Ala Ala Gly Ala Met Ser Pro Asp His His Lys Lys Leu Lys Glu Lys 35 40 45 Ile Gln Lys Lys Leu Gln Glu Leu Thr Gln Pro His His Leu Leu Ala 50 55 60 Leu Ile Asp Ala Val Gln Arg Leu Gly Val Ala Tyr Gln Phe Glu Asp 65 70 75 80 Glu Ile Ser Glu Ala Leu His Gly Ile His Ser Asp Asn Ser Glu His 85 90 95 Ala Ile Arg Glu Ser Leu His His Thr Ser Leu Tyr Phe Arg Leu Leu 100 105 110 Arg Gln His Gly Cys Asn Leu Ser Thr Asp Ile Phe Asn Lys Phe Arg 115 120 125 Lys Asp Gly Gly Gly Phe Lys Ala Thr Leu Cys Glu Asp Ala Met Gly 130 135 140 Leu Leu Ser Leu Tyr Glu Ala Val Arg Leu Ser Val Arg Gly Glu Ala 145 150 155 160 In Leu Glu Glu Ala Gln In Phe Ser In Ala Asn Leu Arg In Val 165 170 175 Ile Glu Lys Ile Glu Arg Lys Leu Ala Asp Lys Ile Glu His Ala Leu 180 185 190 Glu Leu Pro Leu Tyr Trp Arg Ala Pro Arg Leu Glu Ala Arg Trp Tyr 195 200 205 Ile Asp Ile Tyr Glu Lys Asp Glu Gly Arg Ile Glu Asp Leu Leu Asp 210 215 220 Phe Ala Lys Leu Asp Phe Asn Arg Val Gln Leu Ile Tyr Gln Thr Glu 225 230 235 240 Leu Arg Glu Leu Thr Met Trp Trp Glu Leu Leu Gly Leu Pro Ala Lys 245 250 255 Met Gly Phe Phe Arg Asp Arg Leu Leu Glu Ser His Leu Phe Ser Val 260 265 270 Ala Met Leu Val Glu Pro Asn Tyr Ser Gln Cys Arg Val Ala Leu Thr 275 280 285 Lys Ala Ser Val Leu Val Ser Ala Met Asp Asp Phe Tyr Asp Ile His 290 295 300 Gly Leu Pro Asp Glu Leu Arg Leu Phe Thr Glu Thr Val Gln Arg Trp 305 310 315 320 Asp Leu Asp Ala Leu Asp Gln Leu Pro Glu Tyr Met Lys Ile Phe Tyr 325 330 335 Met Ala Leu Tyr Asn Thr Thr Asn Glu Thr Ala Tyr Leu Met Leu Lys 340 345 350 Asp Lys Gly Phe Asn Ala Ser His Tyr Leu Lys Lys Val Trp Ala Leu 355 360 365 Gln Ser Asn Gly Tyr Phe Arg Glu Ala Gln Trp Tyr Asn Ser Gly Tyr 370 375 380 Val Pro Lys Phe Asp Glu Tyr Leu Asp Asn Ala Val Ile Ser Val Gly 385 390 395 400 Gly Pro Phe Leu Leu Gly Val Ser Tyr Pro Leu Leu Gln Asn Gln Leu 405 410 415 Thr Lys Glu Glu Met Glu Leu Leu Pro Glu Asp Leu Asn Leu Val Lys 420 425 430 Phe Ala Ser Val Ile Phe Arg Leu Tyr Asp Asp Leu Ala Thr Ser Lys 435 440 445 Ala Glu Gln Asn Arg Gly Asp Leu Pro Lys Ser Ile Gln Cys Tyr Met 450 455 460 His Glu Thr Gly Ser Ser Glu Glu Met Ala Ala Gln His Ile Lys Asp 465 470 475 480 Ile Ile Ser Glu Gly Trp Lys Asp Val Asn Ala Glu Cys Leu Arg Pro 485 490 495 Ser Ser Val Ser Arg His Phe Leu Gly Leu Ala Pro Asn Ser Ala Arg 500 505 510 Ser Gly Val Ile Val Tyr His His Asp Phe Asp Gly Phe Ala Ser Pro 515 520 525 His Ala Arg Ser Asn Ala His Ile Thr Ser Ile Phe Phe Glu Pro Val 530 535 540 Pro Ile Arg Glu Ser Ile Asn Leu Gly 545 550 <210> 52 <211> 553 <212> PRT <213> Artificial Sequence <220> <223> Var_267L_291S <400> 52<0004Met Asp Ser Leu Asp Leu Arg Arg Ser Ala Ile Tyr His Ser Thr Phe 1 5 10 15 Trp Asp Ile Asp Ser Ile Lys Ala Leu Ile Ala Lys Arg Asp Cys Ser 20 25 30 Ala Ala Ala Ala Val Ser Pro Glu His His Arg Lys Leu Lys Glu Lys 35 40 45 Ile Gln Lys Lys Leu Gln Glu Ile Ser Gln Pro His His Leu Leu Gly 50 55 60 Leu Ile Asp Ala Val Gln Arg Leu Gly Val Ala Tyr Gln Phe Glu Asp 65 70 75 80 Glu Ile Ser Asp Ala Leu His Gly Met His Thr Asp Gln Thr Asp His 85 90 95 Ala Val Lys Glu Ser Leu His His Thr Ser Leu Trp Phe Arg Leu Leu 100 105 110 Arg Gln His Gly Cys Asn Leu Ser Ser Asp Ile Phe Asn Arg Phe Arg 115 120 125 Lys Glu Gly Gly Gly Phe Lys Ala Ser Ile Cys Glu Asp Ala Met Gly 130 135 140 Leu Leu Ser Leu Tyr Glu Ala Val Arg Leu Ser Ile Lys Gly Glu Ala 145 150 155 160 Ile Leu Glu Glu Ala Gln Val Phe Ser Ile Ala Asn Leu Arg Ile Met 165 170 175 Leu Glu Lys Val Glu Lys Lys Leu Ala Asp Arg Ile Asp His Ala Leu 180 185 190 Glu Leu Pro Ile Phe Trp Arg Ala Pro Arg Leu Glu Ala Arg Trp Tyr 195 200 205 Ile Glu Val Tyr Glu Arg Asp Glu Gly Arg Ile Glu Glu Leu Leu Asp 210 215 220 Phe Ala Lys Leu Asp Phe Asn Arg Val Gln Met Val Tyr Gln Thr Glu 225 230 235 240 Leu Lys Glu Leu Thr Met Trp Trp Glu Leu Val Gly Leu Pro Ala Arg 245 250 255 Met Gly Phe Phe Arg Asp Arg Ile Leu Glu Leu His Leu Phe Ser Leu 260 265 270 Ala Ile Leu Val Glu Pro Gln Tyr Ser Asn Cys Arg Val Ala Leu Thr 275 280 285 Lys Ala Ser Ile Leu Ile Ser Ala Met Asp Asp Phe Tyr Asp Val His 290 295 300 Gly Leu Pro Asp Glu Leu Lys Leu Phe Thr Glu Thr Val Gln Arg Trp 305 310 315 320 Asp Ile Asp Ala Ile Asp Asn Leu Pro Glu Tyr Met Lys Leu Phe Tyr 325 330 335 Val Ala Leu Tyr Asn Thr Thr Asn Glu Thr Ala Trp Leu Ile Val Lys 340 345 350 Glu Arg Gly Phe Asn Ala Thr His Tyr Leu Arg Lys Val Trp Ala Met 355 360 365 Gln Ser Asn Ala Tyr Phe Arg Glu Ala Gln Trp Phe Gln Ser Gly Tyr 370 375 380 Val Pro Lys Phe Asp Glu Tyr Leu Glu Asn Ala Val Ile Ser Ile Gly 385 390 395 400 Gly Pro Phe Ile Leu Gly Leu Ser Tyr Pro Leu Ile Gln Asn Gln Val 405 410 415 Thr Lys Glu Glu Ile Asp Leu Leu Pro Glu Asp Leu Asn Leu Val Lys 420 425 430 Trp Ala Ser Val Ile Phe Arg Leu Tyr Asp Asp Leu Gly Thr Ser Lys 435 440 445 Ala Glu Asn Gln Arg Gly Asp Leu Pro Lys Ser Ile Gln Cys Tyr Met 450 455 460 His Glu Thr Gly Ser Ser Glu Glu Val Ala Ala Asn His Ile Lys Glu 465 470 475 480 Met Ile Ser Glu Ala Trp Lys Asp Val Asn Ala Glu Cys Leu Arg Pro 485 490 495 Ser Thr Ile Ser Arg His Phe Leu Gly Val Ala Pro Asn Ser Ala Arg 500 505 510 Ser Ala Val Ile Met Tyr His His Asp Phe Asp Gly Phe Ala Ser Pro 515 520 525 His Ala Lys Thr Asn Ala His Ile Thr Ser Ile Phe Tyr Asp Pro Val 530 535 540 Pro Leu Arg Glu Ser Ile Asn Leu Gly 545 550 <210> 53 <211> 553 <212> PRT <213> artificial sequence <220> <223> CiCassy I291L <400> 53 Met Asp Ser Met Glu Val Arg Arg Ser Ala Ile Tyr His Ser Thr Phe 1 5 10 15 Trp Asp Ile Asp Ser Ile Arg Ala Leu Leu Ala Arg Arg Asp Cys Thr 20 25 30 Ala Ala Ala Ala Leu Ser Pro Asp His His Lys Arg Leu Lys Glu Arg 35 40 45 Ile Gln Arg Arg Leu Gln Asp Ile Thr Gln Pro His His Leu Leu Gly 50 55 60 Leu Ile Asp Ala Val Gln Arg Leu Gly Val Ala Tyr Gln Phe Glu Glu 65 70 75 80 Glu Ile Ser Asp Ala Leu His Gly Leu His Ser Glu Asn Thr Glu His 85 90 95 Ala Ile Lys Asp Ser Leu His His Thr Ser Leu Tyr Phe Arg Leu Leu 100 105 110 Arg Gln His Gly Cys Asn Leu Ser Ser Asp Ile Phe Asn Lys Phe Lys 115 120 125 Lys Glu Gly Gly Gly Phe Lys Ala Ser Leu Cys Glu Asp Ala Met Gly 130 135 140 Leu Leu Ser Leu Tyr Glu Ala Val Arg Leu Ser Val Lys Gly Glu Ala 145 150 155 160 Ile Leu Glu Glu Ala Gln Val Phe Ser Ile Ala Asn Leu Lys Ile Leu 165 170 175 Met Glu Arg Val Glu Arg Lys Leu Ala Asp Arg Ile Glu His Ala Leu 180 185 190 Glu Ile Pro Leu Tyr Trp Arg Ala Pro Arg Leu Glu Ala Arg Trp Tyr 195 200 205 Ile Asp Val Tyr Glu Lys Glu Asp Gly Arg Ile Asp Asp Leu Leu Asp 210 215 220 Phe Ala Lys Leu Asp Phe Asn Arg Val Gln Met Leu Tyr Gln Thr Glu 225 230 235 240 Leu Lys Glu Leu Ser Met Trp Trp Glu Leu Leu Gly Leu Pro Ala Lys 245 250 255 Met Gly Phe Phe Arg Asp Arg Leu Leu Glu Asn His Leu Phe Ser Ile 260 265 270 Ala Val Val Val Glu Pro Gln Tyr Ser Gln Cys Arg Val Ala Ile Thr 275 280 285 Lys Ala Leu Val Leu Met Thr Ala Met Asp Asp Phe Tyr Asp Val His 290 295 300 Gly Leu Pro Asp Glu Leu Lys Val Phe Thr Asp Thr Val Asn Arg Trp 305 310 315 320 Asp Leu Glu Gly Ile Asp Gln Leu Pro Glu Tyr Met Lys Leu Tyr Tyr 325 330 335 Leu Ala Leu Tyr Asn Thr Thr Asn Glu Thr Ala Tyr Ile Ile Leu Lys 340 345 350 Glu Lys Gly Phe Asn Ala Thr His Tyr Leu Lys Lys Leu Trp Ala Met 355 360 365 Gln Ser Asn Ala Tyr Phe Arg Glu Ala Gln Trp Phe Asn Ser Gly Tyr 370 375 380 Ile Pro Lys Phe Asp Glu Tyr Leu Asp Asn Ala Leu Val Ser Val Gly 385 390 395 400 Ala Pro Phe Val Leu Gly Leu Ser Tyr Pro Met Ile Gln Gln Gln Ile 405 410 415 Ser Lys Glu Glu Ile Asp Leu Ile Pro Glu Asp Leu Asn Leu Leu Arg 420 425 430 Trp Ala Ser Ile Ile Phe Arg Leu Tyr Asp Asp Leu Ala Thr Ser Lys 435 440 445 Ala Glu Gln Gln Arg Gly Asp Val Pro Lys Ser Ile Gln Cys Tyr Met 450 455 460 His Glu Thr Gly Ser Ser Glu Glu Val Ala Ala Asn His Ile Arg Asp 465 470 475 480 Leu Ile Ser Asp Ala Trp Lys Glu Val Asn Ala Glu Cys Leu Lys Pro 485 490 495 Thr Ser Leu Ser Lys His Tyr Val Gly Val Ala Pro Asn Ser Ala Arg 500 505 510 Ser Gly Val Leu Met Tyr His His Asp Phe Asp Gly Phe Ala Ser Pro 515 520 525 His Gly Arg Thr Asn Ala His Ile Thr Ser Ile Phe Phe Glu Pro Val 530 535 540 Pro Leu Lys Glu Ser Ile Asn Leu Gly 545 550 <210> 54 <211> 11 <212> PRT <213> Artificial sequence <220> <223> RRX8W motif <220> <221> variants <222> 3 <223> Any amino acid <220> <221> variants <222> 4 <223> Any amino acid <220> <221> variants <222> 5 <223> Any amino acid <220> <221> variants <222> 6 <223> Any amino acid <220> <221> variants <222> 7 <223> Any amino acid <220> <221> variants <222> 8 <223> Any amino acid <220> <221> variants <222> 9 <223> Any amino acid <220> <221> variants <222> 10 <223> Any amino acid <400> 54 Arg Arg Xaa Xaa Xaa Xaa Xaa Xaa Xaa Xaa Arg 1 5 10 <210> 55 <211> 11 <212> PRT <213> Artificial sequence <220> <223> R(R / K)X8W motif <220> <221> variants <222> 2 <223> Xaa is arginine or lysine. <220> <221> variants <222> 3 <223> Any amino acid <220> <221> variants <222> 4 <223> Any amino acid <220> <221> variants <222> 5 <223> Any amino acid <220> <221> variants <222> 6 <223> Any amino acid <220> <221> variants <222> 7 <223> Any amino acid <220> <221> variants <222> 8 <223> Any amino acid <220> <221> variants <222> 9 <223> Any amino acid <220> <221> variants <222> 10 <223> Any amino acid <400> 55 Arg Xaa Xaa Xaa Xaa Xaa Xaa Xaa Xaa Xaa Arg 1 5 10 <210> 56 <211> 553 <212> PRT <213> Artificial Sequence <220> <223> N267W <400> 56 Met Asp Ser Met Glu Val Arg Arg Ser Ala Ile Tyr His Ser Thr Phe 1 5 10 15 Trp Asp Ile Asp Ser Ile Arg Ala Leu Leu Ala Arg Arg Asp Cys Thr 20 25 30 Ala Ala Ala Ala Leu Ser Pro Asp His His Lys Arg Leu Lys Glu Arg 35 40 45 Ile Gln Arg Arg Leu Gln Asp Ile Thr Gln Pro His His Leu Leu Gly 50 55 60 Leu Ile Asp Ala Val Gln Arg Leu Gly Val Ala Tyr Gln Phe Glu Glu 65 70 75 80 Glu Ile Ser Asp Ala Leu His Gly Leu His Ser Glu Asn Thr Glu His 85 90 95 Ala Ile Lys Asp Ser Leu His His Thr Ser Leu Tyr Phe Arg Leu Leu 100 105 110 Arg Gln His Gly Cys Asn Leu Ser Ser Asp Ile Phe Asn Lys Phe Lys 115 120 125 Lys Glu Gly Gly Gly Phe Lys Ala Ser Leu Cys Glu Asp Ala Met Gly 130 135 140 Leu Leu Ser Leu Tyr Glu Ala Val Arg Leu Ser Val Lys Gly Glu Ala 145 150 155 160 Ile Leu Glu Glu Ala Gln Val Phe Ser Ile Ala Asn Leu Lys Ile Leu 165 170 175 Met Glu Arg Val Glu Arg Lys Leu Ala Asp Arg Ile Glu His Ala Leu 180 185 190 Glu Ile Pro Leu Tyr Trp Arg Ala Pro Arg Leu Glu Ala Arg Trp Tyr 195 200 205 Ile Asp Val Tyr Glu Lys Glu Asp Gly Arg Ile Asp Asp Leu Leu Asp 210 215 220 Phe Ala Lys Leu Asp Phe Asn Arg Val Gln Met Leu Tyr Gln Thr Glu 225 230 235 240 Leu Lys Glu Leu Ser Met Trp Trp Glu Leu Leu Gly Leu Pro Ala Lys 245 250 255 Met Gly Phe Phe Arg Asp Arg Leu Leu Glu Trp His Leu Phe Ser Ile 260 265 270 Ala Val Val Val Glu Pro Gln Tyr Ser Gln Cys Arg Val Ala Ile Thr 275 280 285 Lys Ala Ile Val Leu Met Thr Ala Met Asp Asp Phe Tyr Asp Val His 290 295 300 Gly Leu Pro Asp Glu Leu Lys Val Phe Thr Asp Thr Val Asn Arg Trp 305 310 315 320 Asp Leu Glu Gly Ile Asp Gln Leu Pro Glu Tyr Met Lys Leu Tyr Tyr 325 330 335 Leu Ala Leu Tyr Asn Thr Thr Asn Glu Thr Ala Tyr Ile Ile Leu Lys 340 345 350 Glu Lys Gly Phe Asn Ala Thr His Tyr Leu Lys Lys Leu Trp Ala Met 355 360 365 Gln Ser Asn Ala Tyr Phe Arg Glu Ala Gln Trp Phe Asn Ser Gly Tyr 370 375 380 Ile Pro Lys Phe Asp Glu Tyr Leu Asp Asn Ala Leu Val Ser Val Gly 385 390 395 400 Ala Pro Phe Val Leu Gly Leu Ser Tyr Pro Met Ile Gln Gln Gln Ile 405 410 415 Ser Lys Glu Glu Ile Asp Leu Ile Pro Glu Asp Leu Asn Leu Leu Arg 420 425 430 Trp Ala Ser Ile Ile Phe Arg Leu Tyr Asp Asp Leu Ala Thr Ser Lys 435 440 445 Ala Glu Gln Gln Arg Gly Asp Val Pro Lys Ser Ile Gln Cys Tyr Met 450 455 460 His Glu Thr Gly Ser Ser Glu Glu Val Ala Ala Asn His Ile Arg Asp 465 470 475 480 Leu Ile Ser Asp Ala Trp Lys Glu Val Asn Ala Glu Cys Leu Lys Pro 485 490 495 Thr Ser Leu Ser Lys His Tyr Val Gly Val Ala Pro Asn Ser Ala Arg 500 505 510 Ser Gly Val Leu Met Tyr His His Asp Phe Asp Gly Phe Ala Ser Pro 515 520 525 His Gly Arg Thr Asn Ala His Ile Thr Ser Ile Phe Phe Glu Pro Val 530 535 540 Pro Leu Lys Glu Ser Ile Asn Leu Gly 545 550 <210> 57 <211> 553 <212> PRT <213> Artificial Sequence <220> <223> N267G <400> 57 Met Asp Ser Met Glu Val Arg Arg Ser Ala Ile Tyr His Ser Thr Phe 1 5 10 15 Trp Asp Ile Asp Ser Ile Arg Ala Leu Leu Ala Arg Arg Asp Cys Thr 20 25 30 Ala Ala Ala Ala Leu Ser Pro Asp His His Lys Arg Leu Lys Glu Arg 35 40 45 Ile Gln Arg Arg Leu Gln Asp Ile Thr Gln Pro His His Leu Leu Gly 50 55 60 Leu Ile Asp Ala Val Gln Arg Leu Gly Val Ala Tyr Gln Phe Glu Glu 65 70 75 80 Glu Ile Ser Asp Ala Leu His Gly Leu His Ser Glu Asn Thr Glu His 85 90 95 Ala Ile Lys Asp Ser Leu His His Thr Ser Leu Tyr Phe Arg Leu Leu 100 105 110 Arg Gln His Gly Cys Asn Leu Ser Ser Asp Ile Phe Asn Lys Phe Lys 115 120 125 Lys Glu Gly Gly Gly Phe Lys Ala Ser Leu Cys Glu Asp Ala Met Gly 130 135 140 Leu Leu Ser Leu Tyr Glu Ala Val Arg Leu Ser Val Lys Gly Glu Ala 145 150 155 160 Ile Leu Glu Glu Ala Gln Val Phe Ser Ile Ala Asn Leu Lys Ile Leu 165 170 175 Met Glu Arg Val Glu Arg Lys Leu Ala Asp Arg Ile Glu His Ala Leu 180 185 190 Glu Ile Pro Leu Tyr Trp Arg Ala Pro Arg Leu Glu Ala Arg Trp Tyr 195 200 205 Ile Asp Val Tyr Glu Lys Glu Asp Gly Arg Ile Asp Asp Leu Leu Asp 210 215 220 Phe Ala Lys Leu Asp Phe Asn Arg Val Gln Met Leu Tyr Gln Thr Glu 225 230 235 240 Leu Lys Glu Leu Ser Met Trp Trp Glu Leu Leu Gly Leu Pro Ala Lys 245 250 255 Met Gly Phe Phe Arg Asp Arg Leu Leu Glu Gly His Leu Phe Ser Ile 260 265 270 Ala Val Val Val Glu Pro Gln Tyr Ser Gln Cys Arg Val Ala Ile Thr 275 280 285 Lys Ala Ile Val Leu Met Thr Ala Met Asp Asp Phe Tyr Asp Val His 290 295 300 Gly Leu Pro Asp Glu Leu Lys Val Phe Thr Asp Thr Val Asn Arg Trp 305 310 315 320 Asp Leu Glu Gly Ile Asp Gln Leu Pro Glu Tyr Met Lys Leu Tyr Tyr 325 330 335 Leu Ala Leu Tyr Asn Thr Thr Asn Glu Thr Ala Tyr Ile Ile Leu Lys 340 345 350 Glu Lys Gly Phe Asn Ala Thr His Tyr Leu Lys Lys Leu Trp Ala Met 355 360 365 Gln Ser Asn Ala Tyr Phe Arg Glu Ala Gln Trp Phe Asn Ser Gly Tyr 370 375 380 Ile Pro Lys Phe Asp Glu Tyr Leu Asp Asn Ala Leu Val Ser Val Gly 385 390 395 400 Ala Pro Phe Val Leu Gly Leu Ser Tyr Pro Met Ile Gln Gln Gln Ile 405 410 415 Ser Lys Glu Glu Ile Asp Leu Ile Pro Glu Asp Leu Asn Leu Leu Arg 420 425 430 Trp Ala Ser Ile Ile Phe Arg Leu Tyr Asp Asp Leu Ala Thr Ser Lys 435 440 445 Ala Glu Gln Gln Arg Gly Asp Val Pro Lys Ser Ile Gln Cys Tyr Met 450 455 460 His Glu Thr Gly Ser Ser Glu Glu Val Ala Ala Asn His Ile Arg Asp 465 470 475 480 Leu Ile Ser Asp Ala Trp Lys Glu Val Asn Ala Glu Cys Leu Lys Pro 485 490 495 Thr Ser Leu Ser Lys His Tyr Val Gly Val Ala Pro Asn Ser Ala Arg 500 505 510 Ser Gly Val Leu Met Tyr His His Asp Phe Asp Gly Phe Ala Ser Pro 515 520 525 His Gly Arg Thr Asn Ala His Ile Thr Ser Ile Phe Phe Glu Pro Val 530 535 540 Pro Leu Lys Glu Ser Ile Asn Leu Gly 545 550 <210> 58 <211> 553 <212> PRT <213> artificial sequence <220> <223> N267A <400> 58 Met Asp Ser Met Glu Val Arg Arg Ser Ala Ile Tyr His Ser Thr Phe 1 5 10 15 Trp Asp Ile Asp Ser Ile Arg Ala Leu Leu Ala Arg Arg Asp Cys Thr 20 25 30 Ala Ala Ala Ala Leu Ser Pro Asp His His Lys Arg Leu Lys Glu Arg 35 40 45 Ile Gln Arg Arg Leu Gln Asp Ile Thr Gln Pro His His Leu Leu Gly 50 55 60 Leu Ile Asp Ala Val Gln Arg Leu Gly Val Ala Tyr Gln Phe Glu Glu 65 70 75 80 Glu Ile Ser Asp Ala Leu His Gly Leu His Ser Glu Asn Thr Glu His 85 90 95 Ala Ile Lys Asp Ser Leu His His Thr Ser Leu Tyr Phe Arg Leu Leu 100 105 110 Arg Gln His Gly Cys Asn Leu Ser Ser Asp Ile Phe Asn Lys Phe Lys 115 120 125 Lys Glu Gly Gly Gly Phe Lys Ala Ser Leu Cys Glu Asp Ala Met Gly 130 135 140 Leu Leu Ser Leu Tyr Glu Ala Val Arg Leu Ser Val Lys Gly Glu Ala 145 150 155 160 Ile Leu Glu Glu Ala Gln Val Phe Ser Ile Ala Asn Leu Lys Ile Leu 165 170 175 Met Glu Arg Val Glu Arg Lys Leu Ala Asp Arg Ile Glu His Ala Leu 180 185 190 Glu Ile Pro Leu Tyr Trp Arg Ala Pro Arg Leu Glu Ala Arg Trp Tyr 195 200 205 Ile Asp Val Tyr Glu Lys Glu Asp Gly Arg Ile Asp Asp Leu Leu Asp 210 215 220 Phe Ala Lys Leu Asp Phe Asn Arg Val Gln Met Leu Tyr Gln Thr Glu 225 230 235 240 Leu Lys Glu Leu Ser Met Trp Trp Glu Leu Leu Gly Leu Pro Ala Lys 245 250 255 Met Gly Phe Phe Arg Asp Arg Leu Leu Glu Ala His Leu Phe Ser Ile 260 265 270 Ala Val Val Val Glu Pro Gln Tyr Ser Gln Cys Arg Val Ala Ile Thr 275 280 285 Lys Ala Ile Val Leu Met Thr Ala Met Asp Asp Phe Tyr Asp Val His 290 295 300 Gly Leu Pro Asp Glu Leu Lys Val Phe Thr Asp Thr Val Asn Arg Trp 305 310 315 320 Asp Leu Glu Gly Ile Asp Gln Leu Pro Glu Tyr Met Lys Leu Tyr Tyr 325 330 335 Leu Ala Leu Tyr Asn Thr Thr Asn Glu Thr Ala Tyr Ile Ile Leu Lys 340 345 350 Glu Lys Gly Phe Asn Ala Thr His Tyr Leu Lys Lys Leu Trp Ala Met 355 360 365 Gln Ser Asn Ala Tyr Phe Arg Glu Ala Gln Trp Phe Asn Ser Gly Tyr 370 375 380 Ile Pro Lys Phe Asp Glu Tyr Leu Asp Asn Ala Leu Val Ser Val Gly 385 390 395 400 Ala Pro Phe Val Leu Gly Leu Ser Tyr Pro Met Ile Gln Gln Gln Ile 405 410 415 Ser Lys Glu Glu Ile Asp Leu Ile Pro Glu Asp Leu Asn Leu Leu Arg 420 425 430 Trp Ala Ser Ile Ile Phe Arg Leu Tyr Asp Asp Leu Ala Thr Ser Lys 435 440 445 Ala Glu Gln Gln Arg Gly Asp Val Pro Lys Ser Ile Gln Cys Tyr Met 450 455 460 His Glu Thr Gly Ser Ser Glu Glu Val Ala Ala Asn His Ile Arg Asp 465 470 475 480 Leu Ile Ser Asp Ala Trp Lys Glu Val Asn Ala Glu Cys Leu Lys Pro 485 490 495 Thr Ser Leu Ser Lys His Tyr Val Gly Val Ala Pro Asn Ser Ala Arg 500 505 510 Ser Gly Val Leu Met Tyr His His Asp Phe Asp Gly Phe Ala Ser Pro 515 520 525 His Gly Arg Thr Asn Ala His Ile Thr Ser Ile Phe Phe Glu Pro Val 530 535 540 Pro Leu Lys Glu Ser Ile Asn Leu Gly 545 550

Claims

1. A synthetic β-santalene synthase, characterized in that... The following facts are established: the tertiary structure of the portion corresponding to amino acid positions 272 to 291 of SEQ ID NO: 1 in the synthesized β-santalene synthase exhibits increased flexibility compared to the same tertiary structure in naturally occurring santalene synthases. This flexibility was determined by root mean square fluctuation analysis using the following settings: pH 8.0, 300 K, 1 atm, aqueous environment, substrate-free ion presence. Simulations were performed on both the synthetic and naturally occurring santalene synthases for 500 ns, and each enzyme structure was evaluated based on the last 450 ns of simulation. Furthermore, the synthesized β-santalene synthase is characterized by its ability to produce β-santalene and α-santalene at a ratio equal to or greater than 1 under common conditions suitable for the production of both santalenes, β-santalene being (-)-β-santalene with CAS number 511-59-1. The synthesized β-santalene synthase is derived from SEQ ID NO:

1. Amino acid position 267 of SEQ ID NO: 1 is replaced with N267S or N267L; or amino acid position 291 of SEQ ID NO: 1 is replaced with I291V, I291S, I291T or I291C; or amino acid position 267 of SEQ ID NO: 1 is replaced with N267S and amino acid position 291 is replaced with I291T; or amino acid position 267 of SEQ ID NO: 1 is replaced with N267T and amino acid position 291 is replaced with I291T.

2. A synthesized nucleic acid encoding the synthesized β-santalene synthase according to claim 1.

3. An expression cassette comprising the synthesized nucleic acid according to claim 2.

4. A method for producing a composition containing β-santalene in greater quantities than α-santalene, comprising the following steps: I. To provide, in active form, one or more of the synthesized β-santalene synthases according to claim 1, and to provide all the necessary cofactors. II. Contact farnesyl pyrophosphate with one or more of the aforementioned synthetic β-santalene synthases under conditions that allow for the production of santalene. III. β-Santalene and α-Santalene are produced from farnesyl pyrophosphate, wherein the amount of β-santalene produced is greater than the amount of α-santalene produced, wherein the β-santalene is (-)-β-santalene, and its CAS number is 511-59-1.

5. A method for producing a composition containing β-santalene in greater quantities than α-santalene, comprising the following steps: I. To provide, in active form, one or more of the synthesized β-santalene synthases according to claim 1, and to provide all the necessary cofactors. II. Contact farnesyl pyrophosphate with one or more of the aforementioned synthetic β-santalene synthases under conditions that allow for the production of santalene. III. β-Santalumene, α-Santalumene, and bergamotene are produced from farnesyl pyrophosphate, wherein the amount of β-Santalumene produced is greater than the amount of α-Santalumene produced, and the β-Santalumene is (-)-β-Santalumene, whose CAS number is 511-59-1.

6. The method for producing a composition containing β-santalene more than α-santalene according to any one of claims 4 or 5, further comprising step IV. purifying the product.

7. The method for producing a composition containing more β-santalene than α-santalene according to any one of claims 4 or 5, wherein, in addition to producing more β-santalene than α-santalene, the synthesized β-santalene synthase also produces more trans-α-bergamerene than α-santalene.

8. A non-human host cell adapted to produce the synthetic β-santalene synthase from a nucleic acid encoding the synthetic β-santalene synthase according to claim 1, and adapted to provide the synthetic β-santalene synthase with farnesyl pyrophosphate and all cofactors required for its activity, wherein the host cell contains nucleic acid encoding the synthetic β-santalene synthase according to claim 1.

9. The non-human host cell according to claim 8, wherein, in addition to producing more β-santalene than α-santalene, the synthesized β-santalene synthase also produces trans-α-bergamerene in addition to α-santalene.

10. A composition produced by synthesizing β-santalene synthase according to claim 1, by the method according to any one of claims 4 or 5, or by non-human host cells according to any one of claims 8 or 9, wherein the composition comprises β-santalene in addition to α-santalene, wherein β-santalene is (-)-β-santalene, whose CAS number is 511-59-1.

11. The composition of claim 10, wherein the composition comprises more β-santalene than bergamotene and more bergamotene than α-santalene.

12. The composition according to any one of claims 10 or 11, wherein the composition comprises at least 12% (w / w) of trans-α-berberine.

13. A method for producing a composition containing β-santalol in greater quantities than α-santalol, comprising the steps of: I. A composition comprising β-santalene in greater quantities than α-santalene, produced by the method of any one of claims 4 or 5, or using the host cell of any one of claims 8 or 9, or using the synthesized β-santalene synthase of claim 1; and II. Oxidize at least some of the β-santalene and α-santalene in the composition produced in a) to their respective alcohols to produce a composition containing more than α-santalene alcohol.

14. The method of claim 13 for producing a composition comprising β-santalol in greater quantities than α-santalol, further comprising step III. purifying the product.

15. A composition produced by the method of any one of claims 13 or 14, comprising β-santalol in greater quantity than α-santalol, wherein the β-santalol / α-santalol ratio is at least 1.

3.

16. Use of the synthesized β-santalene synthase according to claim 1 for producing a composition comprising α-santalene, β-santalene and trans-α-berberine.

Citation Information

Patent Citations

  • Improvement in journal-boxes for cars

    US109159A

  • Method for producing alpha-santalene

    US20110008836A1

  • Methods for in vitro recombination

    US5605793A

  • Methods for generating polynucleotides having desired characteristics by iterative selection and recombination

    US5811238A

  • Process for the production of expression vectors comprising at least one stochastic sequence of polynucleotides

    US5824514A