Novel polypeptides for producing peltatol and / or drimenol compounds

By identifying novel polypeptides with cyclic terpene synthase activity from plants of the genus *Dystridium* and *Selaginella tamariscina*, the problem of efficient extraction of sesquiterpenoid compounds in existing technologies has been solved, enabling highly selective synthesis of sesquiterpenol and sesquiterpenol, and providing a new method for preparing high-value fragrance components.

CN113853432BActive Publication Date: 2026-05-15FIRMENICH SA +1
View PDF 4 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
FIRMENICH SA
Filing Date
2020-11-25
Publication Date
2026-05-15

AI Technical Summary

Technical Problem

Existing technologies make it difficult to efficiently and cost-effectively extract sesquiterpenoid compounds such as sesquiterpenol and sesquiterpenol from plants for the preparation of high-value fragrance ingredients, and the content of these compounds in natural sources is low.

Method used

Novel polypeptides with cyclic terpene synthase activity, particularly BazzHAD1, BazzHAD2, BazzHAD3, BtHAD, SmHAD1, and SmHAD2, were identified from plants of the genus *Bazzyra* and *Selaginella tamariscina*. These polypeptides can selectively generate bis(2-)-enyl and/or bis(2-)-enyl diphosphate from farnesyl diphosphate (FPP), and through the catalytic reaction of these polypeptides, bis(2-)-enyl diphosphate and bis(2-)-enyl diphosphate are produced.

Benefits of technology

This study enables the efficient and selective synthesis of phytosterol and phytosterol, providing a more cost-effective method for preparing high-value flavoring ingredients and addressing the issue of low levels of these compounds in natural sources.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure BDA0003319974950000181
    Figure BDA0003319974950000181
  • Figure BDA0003319974950000321
    Figure BDA0003319974950000321
  • Figure BDA0003319974950000331
    Figure BDA0003319974950000331
Patent Text Reader

Abstract

The present invention relates to novel plant-derived halogenase-like (HAD-like) polypeptides having cyclic terpene synthase (TPS) activity, particularly suitable for use in biochemical methods for the production of drimane sesquiterpenes, including drimenol and / or selinol and / or related compounds, such as phosphorylated drimane sesquiterpene alcohols, particularly phosphorylated drimenol and / or selinol compounds and derivatives. The present invention also provides the novel TPS-activity-encoding nucleotide sequences, corresponding expression constructs, recombinant hosts, methods for the production of such novel polypeptides and mutants and variants thereof. The present invention also relates to the use of such novel polypeptides in the production of odorants, flavorants and fragrance ingredients.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to novel plant-derived haloacid dehalogenase-like (HAD-like) polypeptides possessing cyclic terpene synthase (TPS) activity, particularly suitable for biochemical methods to produce sterane-type sesquiterpenes, including sterane alcohol and / or zephyranol and / or related compounds, such as phosphorylated sterane-type sesquiterpene alcohols, especially phosphorylated sterane alcohol and / or zephyranol compounds and derivatives. The invention also provides the encoding nucleotide sequences of the novel TPS activity, corresponding expression constructs, recombinant hosts, and methods for preparing such novel polypeptides and their mutants and variants. The invention further relates to the use of such novel polypeptides in the production of odorant, flavoring, and fragrance ingredients. Background Technology

[0002] Terpenes are found in most organisms (microorganisms, animals, and plants). These compounds are composed of isoprene units and are classified according to the number of these units present in their structure, which may contain cyclic structural elements. Therefore, hemiterpenes, monoterpenes, sesquiterpenes, and diterpenes are terpenes containing 5, 10, 15, and 20 carbon atoms (i.e., 1, 2, 3, and 4 isoprene units), respectively. Monoterpenes are derived from geraniyl diphosphate (GPP, C... 10 ), sesquiterpenes are derived from farnesyl diphosphate (FPP, C 15 Diterpenes are derived from geraniol geraniol diphosphate (GGPP, C). 20 For example, sesquiterpenes are widely distributed in the plant kingdom. Many sesquiterpene molecules are known for their flavor and aroma properties, as well as their cosmetic, medicinal, and antibacterial effects. A variety of sesquiterpene hydrocarbons and sesquiterpene compounds have been identified. Chemical synthetic routes have been developed, but they remain complex and are not always cost-effective.

[0003] The biosynthesis of terpenes involves enzymes called terpene synthases. Several sesquiterpene synthases exist in the plant kingdom, all using the same substrate (farnesyl diphosphate, FPP) but with different product structures. Genes and cDNAs encoding sesquiterpene synthases have been cloned, and the corresponding recombinant enzymes have been characterized.

[0004] Many major sources of sesquiterpenes, such as compounds with drimane structures, like drinol or drinol compounds, are derived from plants; however, the content of sesquiterpenes in these natural sources may be low. There remains a need to identify and characterize novel terpene synthases, optionally characterized by specific drimane product structures and / or drimane product yields and / or more cost-effective methods for producing drimane compounds such as drinol and / or drinol, which are potential structural components for the preparation of high-value flavoring ingredients (e.g., Ambrox). Summary of the Invention

[0005] The aforementioned problems can be addressed by providing several novel enzymes with cyclic terpene synthase (TPS) activity, particularly cyclic sesquiterpene synthase activity, and more specifically exhibiting drimenyl and / or zosquienyl diphosphate synthase activity, producing drimenyl and / or zosquienyl diphosphates from FPP with remarkably high selectivity. In particular, novel drimenyl and / or zosquienyl diphosphate synthase genes have been identified from plants of the genus *Bazzania*, such as *Bazzania trilobata*, and *Selaginella moellendorffii*. These novel synthases have been named BazzHAD1, BazzHAD2, BazzHAD3, BtHAD, SmHAD1, and SmHAD2.

[0006] The protein sequences of SmHAD1 and SmHAD2 have very low identity (approximately 14–20%) with all other known HAD-like TPS, but they show high homology (>96%) among themselves. BtHAD and BazzHAD1 show slightly higher homology (approximately 22–30%) with fern and fungal HAD-like TPS and 96% homology among themselves, while BazzHAD2 and BazzHAD3 show moderate homology (approximately 22–43%) with fern and fungal HAD-like TPS and 67% homology among themselves (see Table 2 in the Experimental Section).

[0007] All HAD-like TPS described herein contain the class II motif PxDxD(T / S)(T / M)S (SEQ ID NO:46) and more C-terminal-shifted QW motifs QxxDGx(W / F) (SEQ ID NO:51), which are characteristic of diterpene synthases. Mutations to the class II motif result in complete loss of activity, while mutations to the QW motif result in 90% loss of activity. All six newly identified TPS of this invention lack the class I motif (DDxx(D / E) (SEQ ID NO:45). Attached Figure Description

[0008] Figure 1.a) Alkane structure; b) More generalized alkane structure; the dashed lines in the equations indicate that a single C=C double bond may exist at any specified position.

[0009] Figure 2 The structural formulas of (+)-Zhessaurol and (-)-Bestrol.

[0010] Figure 3Mechanisms for cyclization of farnesyl diphosphate (FPP) via HAD-like TPS containing class II motifs, and mechanisms for dephosphorylation of complementenyl and / or zearalenyl diphosphate via the same HAD-like TPS or other phosphatases.

[0011] Figure 4 GC / MS chromatograms of samples from BazzHAD1 in vivo biochemical analysis. Farnesol is produced by the dephosphorylation of the precursor FPP, while bismuth subsalicylate is likely a product of the dephosphorylation of bismuth subsalicylate.

[0012] Figure 5 . Figure 4 Mass spectrum of the methyl benzoyl peroxide peak.

[0013] Figure 6 GC / MS chromatograms of samples from BazzHAD2 in vivo biochemical analysis. Farnesol is produced by the dephosphorylation of the precursor FPP, while zephyranol and the unknown compound SQT1 are likely products of the dephosphorylation of zephyranol diphosphate and the unknown diphosphate ester, respectively.

[0014] Figure 7 . Figure 4 Mass spectrum of the alcohol peak in *Lysimachia christinae*.

[0015] Figure 8 . Figure 4 Mass spectrum of the unknown SQT1 peak.

[0016] Figure 9 GC / MS chromatograms of samples from BazzHAD3 in vivo biochemical analysis. Farnesol is produced by the dephosphorylation of the precursor FPP, while zephyranol is likely a product of the dephosphorylation of zephyranol diphosphate.

[0017] Figure 10 . Figure 7 Mass spectrum of the alcohol peak in *Lysimachia christinae*.

[0018] Figure 11 GC / MS chromatograms of samples from BtHAD biochemical analysis in vivo. Farnesol is produced by the dephosphorylation of the precursor FPP, while bismuth subsalicylate may be a product of the dephosphorylation of bismuth subsalicylate.

[0019] Figure 12 . Figure 10 Mass spectrum of the methyl benzoyl peroxide peak.

[0020] Figure 13 GC / MS chromatogram of the sample for in vivo biochemical analysis of SmHAD1 (EFJ10816.1). Farnesol is produced by the dephosphorylation of the precursor FPP, while bismuth subtilisol and bismuth subtilisol are likely products of the dephosphorylation of bismuth subtilisol and bismuth subtilisol, respectively.

[0021] Figure 14 . Figure 11 Mass spectrum of the alcohol peak in *Lysimachia christinae*.

[0022] Figure 15 . Figure 11 Mass spectrum of the methyl benzoyl peroxide peak.

[0023] Figure 16 GC / MS chromatogram of the sample for in vivo biochemical analysis of SmHAD2 (EFJ26126.1). Farnesol is produced by the dephosphorylation of the precursor FPP, while bismuth subtilisol and bismuth subtilisol are likely products of the dephosphorylation of bismuth subtilisol and bismuth subtilisol, respectively.

[0024] Figure 17 . Figure 14 Mass spectrum of the alcohol peak in *Lysimachia christinae*.

[0025] Figure 18 . Figure 14 Mass spectrum of the methyl benzoyl peroxide peak.

[0026] Figure 19 Comparison of yields of bismuth subtypes synthesized by SmHAD1, SmHAD2, BtHAD, XP_007369631.1 and EMD37666.1 and their mutants in in vivo biochemical analysis (mean ± SE, n = 4–5). The SmHAD1 and SmHAD2 mutants contain mutations D257A and D259A in class II motifs (SEQ ID NO: 21, 28, and 40), and mutations E432A ​​and D433A in QW motifs (SEQ ID NO: 23, 30, and 44); the BtHAD mutant contains mutations D320A and D322A in class II motifs (SEQ ID NO: 14 and 39), and mutation D491A in QW motifs (SEQ ID NO: 16 and 43); the EMD37666.1 mutant contains mutations D276A, D277A, and D279A in class II motifs (SEQ ID NO: 34 and 41); and the XP_007369631.1 mutant contains mutations D272A, D273A, and D275A in class II motifs (SEQ ID NO: 38 and 42).

[0027] Figure 20 GC / MS chromatograms of samples from SmHAD1, SmHAD2, and BtHAD in the presence or absence of BAP (bacterial alkaline phosphatase) for in vitro biochemical analysis. (2E,6E)-farnesol is produced by the dephosphorylation of the precursor FPP, while zephyranol and bismuthol are likely products of the dephosphorylation of zephyranol and bismuthol diphosphates, respectively.

[0028] Figure 21Amino acid sequence alignments of HAD-like TPS: BazzHAD1 (SEQ ID NO:3), BazzHAD2 (SEQ ID NO:6), BazzHAD3 (SEQ ID NO:9), BtHAD (SEQ ID NO:12), SmHAD1 (SEQ ID NO:19), SmHAD2 (SEQ ID NO:26); EMD37666.1 (SEQ ID NO:32) and XP_007369631.1 (SEQ ID NO:36), as previously described in WO 2018 / 220113. The image identified type I motifs (DDxx(D / E); SEQ ID NO:45) and type II motifs (PxDxD(T / S)(T / M)S; SEQ ID NO:46), as well as three other conserved sequence motifs 1–3 (Lxxxx(W / F)xxYxxG; SEQ ID NO:56, YxDxxRxRVD(P / A)V(V / A)xxN; SEQ ID NO:62, GTx(Y / F)YxxxExFL(Y / F); SEQ ID NO:69). Detailed Implementation

[0029] abbreviations used

[0030] bp base pairs

[0031] BAP (bacterial alkaline phosphatase)

[0032] BSA (Bovine Serum Albumin)

[0033] DNA deoxyribonucleic acid

[0034] cDNA complementary DNA

[0035] DTT dithiothreitol

[0036] FPP farnesyl diphosphate

[0037] GC gas chromatography

[0038] HAD haloacid dehalogenase

[0039] IPTG isopropyl-D-thiogalactopyranoside

[0040] LB lysate broth

[0041] MS mass spectrometer / mass spectrometry

[0042] MVA (mevaleric acid)

[0043] PCR polymerase chain reaction

[0044] PP diphosphate (ester) or pyrophosphate (ester)

[0045] RNA (ribonucleic acid)

[0046] mRNA messenger ribonucleic acid

[0047] miRNA

[0048] siRNA (small interfering RNA)

[0049] rRNA ribosomal RNA

[0050] tRNA transfer RNA

[0051] sp. species

[0052] TPS Terpene Synthase

[0053] Detailed description

[0054] a. Definition

[0055] Terpenes are a large and diverse class of organic compounds produced by a variety of plants, especially conifers, and some insects. Terpenes are hydrocarbons. Although sometimes used interchangeably with "terpenes," "terpenoids" or "isoprene-like compounds" are modified terpenes because they contain additional functional groups, usually oxygen-containing ones.

[0056] "Terpenoids" ("isoprene-like compounds") are a large and diverse class of naturally occurring organic chemical substances derived from terpenes. While sometimes used interchangeably with the term "terpene," "terpenoids" contain an additional functional group, typically an oxygen-containing group such as a hydroxyl, carbonyl, or carboxyl group. Most are polycyclic structures with oxygen-containing functional groups. Unless otherwise stated, the terms "terpene" and "terpenoids" are used interchangeably in the context of this specification.

[0057] Terpenes (and terpenoid compounds) can be classified according to the number of isoprene units in their molecules; the prefix in the name indicates the number of terpene units required to form the molecule. Hemiterpenes consist of a single isoprene unit. Monoterpenes consist of two isoprene units and have the molecular formula C2. 10 H 16 Sesquiterpenes are composed of three isoprene units and have the molecular formula C0. 15 H 24 Diterpenes are composed of four isoprene units and have the molecular formula C0. 20 H 32 .

[0058] "Sysquiterpenes" or "sysquiterpenes" are bicyclic sesquiterpenes and are the parent structures of many natural products with various biological activities. Figure 1aAccording to IUPAC, its systematic chemical name is (4aR,5S,6S,8aS)-1,1,4a,5,6-pentamethyldecahydronaphthalene. For the purposes of this application, the term "complementane" must be understood more broadly and is also referred to as "complementane type," and encompasses those having, for example, […]. Figure 1b The bicyclic sesquiterpenes described herein are of a general structure. They are not limited to specific stereochemistry and may contain a single C=C double bond at certain positions in the sterane skeleton. Specific sterane structures are those contained in sterane alcohols and zephyranols.

[0059] "Symplocaneol" or "symplocane sesquiterpeneol" refers to "symplocane" or "symplocane sesquiterpene" with a hydroxyl group; specific examples are symplocaneol and zephyranol. This term specifically refers to substances with, for example, hydroxyl groups. Figure 1b The cyclic terpenes (terpenoids) shown have a carbon skeleton structure similar to that of alkylene.

[0060] For the purposes of this application, "Zhessomol" refers to (+)-Zhessomol (CAS: 54632-04-1) Figure 2 ).

[0061] "Zheshol derivatives" include compounds derived from zheshol through one or more steps such as esterification, hydroxylation, oxidation, acylation, isomerization, and dimethylation. Suitable derivatives may be selected from hydrocarbons, alcohols, glycols, triols, acetals, ketals, aldehydes, acids, ethers, amides, ketones, lactones, epoxides, acetates, glycosides, and / or esters.

[0062] For the purposes of this application, "supplementol" refers to (-)-supplementol (CAS: 468-68-8) Figure 2 ).

[0063] "Compositol derivatives" include compounds derived from composeols through one or more steps such as esterification, hydroxylation, oxidation, acylation, isomerization, and dimethylation. Suitable derivatives may be selected from hydrocarbons, alcohols, glycols, triols, acetals, ketals, aldehydes, acids, ethers, amides, ketones, lactones, epoxides, acetates, glycosides, and / or esters.

[0064] For the purposes of this application, “Ambrox” refers to the IUPAC name: (-)-(3aR,5aS,9aS,9bR)-3a,6,6,9a-tetramethyldodecanonaphtho[2,1-b]furan (CAS: 6790-58-5).

[0065] The conversion of a "precursor" molecule of the target compound as described herein into the target compound is preferably achieved by at least one structural alteration of the precursor molecule through the enzymatic action of a suitable polypeptide. For example, a "bisphosphate precursor" (e.g., a "terpenoid bisphosphate precursor") can be converted into the target compound (e.g., a terpenoid alcohol) by enzymatic removal of the diphosphate moiety, such as by removing a mono- or diphosphate group with a phosphatase. For example, an "acyclic precursor" (e.g., an "acyclic terpenoid precursor") can be converted into a cyclic target molecule (e.g., a cyclic terpene compound) in one or more steps by the action of a cyclase or synthase, regardless of the specific enzymatic mechanism of such enzyme.

[0066] The terms “cyclic terpene synthase” or “polypeptide with cyclic terpene synthase activity” are used as synonyms for the terms “terpene cyclase” or “polypeptide with terpene cyclase activity”. The term “TPS” is used as an abbreviation for the term “terpene synthase”. As described herein, cyclic terpene synthases belong to the haloacid dehalogenase-like (HAD-like) hydrolases superfamily and contain domains corresponding to the Pfam domains PF13419 (http: / / pfam.xfam.org / family / PF13419) and / or PF00702 (http: / / pfam.xfam.org / family / PF00702.26). Some cyclic terpene synthases that do not possess significant Pfam domains PF13419 and / or PF00702 are still considered to be from HAD-like hydrolases because they contain conserved modification motifs of the aforementioned HAD-like hydrolases superfamily, such as class I, class II, and QW motifs (SEQ ID: 45, SEQ ID: 46, SEQ ID: 51). HAD-like hydrolases are parts of peptides that share amino acid sequence similarity and related functions with members of the HAD-like hydrolases family. HAD-like hydrolases can be identified in peptides by searching for amino acid motifs or features of this protein family. Tools for performing such searches are available, for example, at: https: / / www.ebi.ac.uk / interpro / or https: / / www.ebi.ac.uk / Tools / hmmer / . Proteins typically consist of one or more functional regions or domains. Different combinations of domains produce a wide variety of proteins found in nature. Therefore, identifying domains occurring within a protein can provide in-depth understanding of its function. Peptides containing HAD-like hydrolase domains and / or characteristic HAD-like hydrolase motifs play a role in the binding and cleavage of phosphate or diphosphate groups of ligands. Peptides of the haloacid dehalogenase-like (HAD-like) hydrolase superfamily containing cyclic terpene synthase activity can be defined as HAD-like TPS.

[0067] The term "Class I terpene synthase" refers to terpene synthases that catalyze reactions initiated by ionization, such as monoterpene and sesquiterpene synthases.

[0068] The terms "Class I terpene synthase motif," "Class I terpene synthase-like motif," "Class I synthase (like) motif," "Class I synthase motif," or "Class I motif" refer to the active site of terpene synthases containing the conserved DDxx (D / E) motif (SEQ ID NO:45). The aspartic acid residue of this Class I motif binds to divalent metal ions (most commonly Mg) associated with, for example, the binding of diphosphate groups. 2 + It also catalyzes the ionization and cleavage of the allyl diphosphate bonds of the substrate.

[0069] The term "class II terpene synthase" or "class II synthase" refers to terpene synthases that catalyze protonation-initiated cyclization reactions, such as those typically involved in the biosynthesis of triterpenes and labdane diterpenes. In class II terpene synthases, the protonation-initiated reaction may involve, for example, an acidic amino acid donating a proton to the terminal double bond.

[0070] The terms “(modified) class II terpene synthase motif,” “(modified) class II terpene synthase-like motif,” “(modified) class II synthase (like) motif,” “(modified) class II synthase motif,” or “(modified) class II motif” refer to the active site of a terpene synthase containing a conserved DxDxxS (SEQ ID NO:75) motif, such as the PxDxD(T / S)(T / M)S motif (SEQ ID NO:46) described herein.

[0071] The term “QW motif” in this document refers to the active site of a terpene synthase that contains a conserved QxxxxxW (SEQ ID NO:76) motif, such as the QxxDGxW motif (SEQ ID NO:50) in this document.

[0072] The terms "bryophylloides diphosphate synthase," "polypeptide with bryophylloides diphosphate synthase activity," "bryophylloides diphosphate synthase protein," or "capability to produce bryophylloides diphosphate" refer to a polypeptide capable of catalyzing the synthesis of bryophylloides diphosphate in any stereoisomer or mixture thereof, starting from acyclic terpene pyrophosphates, particularly farnesyl diphosphate (FPP). Bromephylloides diphosphate may be the sole product or part of a mixture of sesquiterpenes. The mixture may contain bryophylloides monophosphate and / or bryophylloidol.

[0073] The terms "forsythoside alcohol synthase," "polypeptide with forsythoside alcohol synthase activity," or "forsythoside alcohol synthase protein" refer to a polypeptide capable of catalyzing the synthesis of forsythoside alcohol in any stereoisomer or mixture thereof, starting from acyclic terpene pyrophosphates, particularly farnesyl diphosphate (FPP). Forsythoside alcohol may be the sole product or part of a mixture of sesquiterpenes.

[0074] The activity of *Zygophylloides bisphosphate synthase* was determined under the following "standard conditions": It could be determined using recombinant *Zygophylloides bisphosphate synthase* expression host cells, disrupted *Zygophylloides bisphosphate synthase* expression cells, fractions, enrichments, or purified enzymes of *Zygophylloides bisphosphate synthase*, at a temperature of approximately 20 to 45°C, for example, approximately 25 to 40°C, preferably 25 to 32°C, in a culture medium or reaction medium (preferably buffered) at a pH of 6 to 11, preferably 7 to 9, and in the presence of a reference substrate (especially FPP), at an initial concentration of 1 to 100 μM, preferably 5 to 50 μM, especially 30 to 40 μM, or by endogenous production by the cell host. The conversion reaction to form *Zygophylloides bisphosphate* proceeds for 10 minutes to 5 hours, preferably about 1 to 2 hours. If endogenous alkaline phosphatase is absent, one or more exogenous phosphatases are added to the reaction mixture to convert *Zygophylloides bisphosphate* formed by the synthase. Then, conventional methods, such as extraction with an organic solvent like ethyl acetate, can be used to determine the physostigmine. Specific examples of suitable standard conditions are applied in the experimental section below, such as in Example 4, and these conditions should also form part of the overall disclosure of this invention.

[0075] The terms "supreme-enyl diphosphate synthase," "peptide with supreme-enyl diphosphate synthase activity," "supreme-enyl diphosphate synthase protein," or "capable of producing supreme-enyl diphosphate" refer to a polypeptide capable of catalyzing the synthesis of supreme-enyl diphosphate in any stereoisomer or mixture thereof, starting from acyclic terpene pyrophosphates, particularly farnesyl diphosphate (FPP). The supreme-enyl diphosphate may be the sole product or part of a mixture of sesquiterpenes. The mixture may contain supreme-enyl monophosphate and / or supreme alcohol.

[0076] The terms "synergist synthase," "polypeptide with synergist activity," or "synergist synthase protein" refer to a polypeptide capable of catalyzing the synthesis of synergist in any stereoisomer or mixture thereof, starting from acyclic terpene pyrophosphates, particularly farnesyl diphosphate (FPP). Synergist may be the sole product or part of a mixture of sesquiterpenes.

[0077] The "complemental-enyl diphosphate synthase activity" is determined under the "standard conditions" described below: It can be determined using recombinant complemental-enyl diphosphate synthase expression host cells, disrupted complemental-enyl diphosphate synthase expression cells, fractions, enrichments, or purified complemental-enyl diphosphate synthase enzymes, under conditions of approximately 20 to 45°C, for example approximately 25 to 40°C, preferably 25 to 32°C, in a culture medium or reaction medium (preferably buffered) at pH 6 to 11, preferably 7 to 9, and in the presence of a reference substrate (especially FPP), at an initial concentration of 1 to 100 μM, preferably 5 to 50 μM, especially 30 to 40 μM, or by endogenous production by the cell host. The conversion reaction to form complemental-enyl diphosphate proceeds for 10 minutes to 5 hours, preferably about 1 to 2 hours. If endogenous alkaline phosphatase is absent, one or more exogenous phosphatases are added to the reaction mixture to convert complemental-enyl diphosphate formed by the synthase. Composterol can then be determined using conventional methods, such as extraction with an organic solvent like ethyl acetate. Specific examples of suitable standard conditions are applied in the experimental section below, such as in Example 4, and these conditions should also form part of the overall disclosure of this invention.

[0078] The terms “biological function,” “function,” “biological activity,” or “activity” refer to the ability of terpene synthases, as described herein, to catalyze the formation of: styrenyl diphosphate and / or styrenol, and / or zedoaryl diphosphate and / or zedoaryl alcohol; or mixtures thereof comprising styrenyl diphosphate and / or styrenyl monophosphate and / or styrenol, and / or zedoaryl diphosphate and / or zedoaryl monophosphate and / or zedoaryl alcohol, and / or one or more other terpenes, particularly styrenyl diphosphate and / or zedoaryl diphosphate.

[0079] The term “terpene mixture” or “sesquiterpene mixture” refers to a mixture of terpenes or sesquiterpenes that contains styrenyl diphosphate and / or styrenyl monophosphate and / or styrenol, and / or zephyranyl diphosphate and / or zephyranyl monophosphate and / or zephyranol, and may also contain one or more additional terpenes, such as one or more additional sesquiterpenes.

[0080] The mevalonate pathway (also known as the isoprene pathway or HMG-CoA reductase pathway) is an essential metabolic pathway in eukaryotes, archaea, and certain bacteria. The mevalonate pathway begins with acetyl-CoA, producing two five-carbon structural units called isopentenyl pyrophosphate (IPP) and dimethylallyl pyrophosphate (DMAPP). Key enzymes include acetoacetyl-CoA thiolysis enzyme (atoB), HMG-CoA synthase (mvaS), HMG-CoA reductase (mvaA), mevalonate kinase (MvaK1), phosphate mevalonate kinase (MvaK2), mevalonate diphosphate decarboxylase (MvaD), and isopentenyl pyrophosphate isomerase (idi). The mevalonate pathway is linked to enzyme activity to produce terpene precursors GPP, FPP, or GGPP, particularly FPP synthase (ERG20), allowing recombinant cells to produce terpenes. As used herein, the terms "recombinant host cell / organism," "genetically modified cell / organism," or "transformed cell / organism" refer to a cell or organism that has been altered to carry at least one nucleic acid molecule, such as a recombinant gene encoding a desired protein or nucleic acid sequence, which, upon transcription, produces the functional polypeptide of the present invention, particularly a complement-enzyme synthase protein that can be used to produce complement-enzyme diphosphate and / or complement-enzyme monophosphate and / or complement-enzyme alcohol or a corresponding mixture of terpenes containing complement-enzyme diphosphate and / or complement-enzyme monophosphate and / or complement-enzyme alcohol, and / or a zirconia diphosphate synthase protein that can be used to produce zirconia diphosphate and / or zirconia monophosphate and / or zirconia alcohol or a corresponding mixture of terpenes containing zirconia diphosphate and / or zirconia monophosphate and / or zirconia alcohol. The host cell is particularly a bacterial cell, fungal cell, or plant cell, or a plant. The host cell may contain recombinant genes that have been integrated into the host cell's nucleus or organelle genome. Alternatively, the host may contain recombinant genes extrachromosomally.

[0081] The term "organism" refers to any non-human multicellular or single-celled organism, such as plants or microorganisms. In particular, microorganisms are bacteria, yeast, algae, or fungi.

[0082] The term "plant" is used interchangeably to include plant cells, including plant protoplasts, plant tissues, plant cell tissue cultures that produce regenerated plants or parts of plants, or plant organs such as roots, stems, leaves, flowers, pollen, ovules, embryos, fruits, etc. Any plant can be used to implement the methods described herein.

[0083] When a particular organism or cell naturally produces FPP, or when it does not naturally produce FPP but is modified (genetically) by the nucleic acids described herein (e.g., by transformation) to produce FPP, it is meant to be "capable of producing FPP". Organisms or cells that are transformed to produce higher or lower, especially higher, amounts of FPP than naturally occurring organisms or cells are also covered as "capable of producing FPP".

[0084] When a particular organism or cell naturally produces physostigmine diphosphate, or when it does not naturally produce physostigmine diphosphate but is converted to produce physostigmine diphosphate using the nucleic acids described herein, it is meant to be "capable of producing physostigmine diphosphate". Organisms or cells that are converted to produce higher or lower, particularly higher, amounts of physostigmine diphosphate than naturally occurring organisms or cells are also included in "capable of producing physostigmine diphosphate".

[0085] "Able to produce stratiostigmine" means that a particular organism or cell naturally produces stratiostigmine, or when it does not naturally produce stratiostigmine but is converted with nucleic acids as described herein to produce stratiostigyl diphosphate, and optionally further converted with nucleic acids to produce enzymatic activity that converts stratiostigyl diphosphate to stratiostigmine. Organisms or cells that are converted to produce higher or lower, particularly higher, amounts of stratiostigmine than naturally occurring organisms or cells are also included in "organisms or cells that can produce stratiostigmine."

[0086] When a particular organism or cell naturally produces complement-enyl diphosphate, or when it does not naturally produce complement-enyl diphosphate but is converted to produce complement-enyl diphosphate using the nucleic acids described herein, it is meant to be "capable of producing complement-enyl diphosphate". Organisms or cells that are converted to produce higher or lower, particularly higher, amounts of complement-enyl diphosphate than naturally occurring organisms or cells are also covered as "capable of producing complement-enyl diphosphate".

[0087] "Capable of producing complementols" means that a particular organism or cell naturally produces complementols, or that does not naturally produce complementols but is converted to complement-enyl diphosphate using nucleic acids as described herein, and optionally further converted to produce enzymatic activity that converts complement-enyl diphosphate to complementols. Organisms or cells that are converted to produce higher or lower, particularly higher, amounts of complementols than naturally occurring organisms or cells are also included in the category of "organisms or cells capable of producing complementols."

[0088] As used herein, the terms “purified,” “substantially purified,” and “isolated” refer to a state free from other different compounds (of which the compounds of the present invention are generally associated in their natural state). Therefore, “purified,” “substantially purified,” and “isolated” articles constitute at least 0.5%, 1%, 5%, 10%, or 20%, or at least 50% or 75% of the mass of a given sample by weight. In one embodiment, these terms refer to the compounds of the present invention constituting at least 95%, 96%, 97%, 98%, 99%, or 100% of the mass of a given sample by weight. As used herein, when referring to nucleic acids or proteins, the terms “purified,” “substantially purified,” and “isolated” for nucleic acids or proteins also refer to a purified or concentrated state that differs from a state naturally present in, for example, prokaryotic or eukaryotic environments, such as in bacterial or fungal cells, or in mammals, particularly humans. Any degree of purification or concentration greater than that of naturally occurring purified or concentrated nucleic acids, including (1) purification from other related structures or compounds, or (2) association with structures or compounds that are not normally associated with them in the prokaryotic or eukaryotic environment, is considered "isolated". Nucleic acids, proteins, or classes of nucleic acids or proteins described herein may be isolated or otherwise associated with structures or compounds that are normally unrelated in nature, according to various methods and processes known to those skilled in the art.

[0089] The term “about” indicates a possible variation of ±25% in the value, particularly ±15%, ±10%, more particularly ±5%, ±2%, or ±1%.

[0090] The term "substantially" describes a value range of approximately 80% to 100%, such as 85% to 99.9%, particularly 90% to 99.9%, even more particularly 95% to 99.9%, or 98% to 99.9%, particularly 99% to 99.9%.

[0091] "Mainly" refers to a proportion in the range of 50%, such as in the range of 51% to 100%, particularly in the range of 75% to 99.9%; especially in the range of 85% to 98.5%, such as 95% to 99%.

[0092] In the context of this invention, "major product" means a single compound or a group of at least two compounds, such as two, three, four, five or more, particularly two or three compounds, which are "majorly" prepared by the reaction described herein and are contained in the reaction in a major proportion based on the total amount of the components of the products formed by the reaction. The proportion may be a molar proportion, a weight proportion, or preferably an area proportion calculated from the corresponding chromatograms of the reaction products based on chromatographic analysis.

[0093] In the context of this invention, "byproduct" means a single compound or a group of at least two compounds, such as two, three, four, five or more, particularly two or three compounds, which are not "mainly" prepared by the reaction described herein.

[0094] Due to the reversibility of enzymatic reactions, unless otherwise stated, this invention relates to enzymatic or biocatalytic reactions described herein in both reaction directions.

[0095] The “functional mutants” of the peptides described in this article include “functional equivalents” of such peptides as defined below.

[0096] The term "stereoisomer" specifically includes conformational isomers.

[0097] According to the present invention, all "stereoisomers" of the compounds described herein are generally included, such as "structural isomers", especially "stereoisomers".

[0098] "Stereoisomeric forms" particularly include "stereoisomers" and mixtures thereof, such as configurational isomers (optical isomers), like enantiomers, or geometrical isomers (diastereomers), like E- and Z-isomers, and combinations thereof. If one or more asymmetric centers are present in a molecule, the invention includes all combinations of different conformations of these asymmetric centers, such as enantiomer pairs.

[0099] "Stereoselectivity" describes the ability to produce a specific stereoisomer of a compound in its stereoisomeric pure form, or the ability to specifically convert a specific stereoisomer from a variety of stereoisomers using the enzymatic catalytic method described herein. More specifically, this means that the product of the invention is enriched relative to a specific stereoisomer, or that the precipitate can be depleted relative to a specific stereoisomer. This can be quantified by a purity parameter %ee calculated according to the following formula:

[0100] %ee=[X A -X B ] / [X A +X B ]*100,

[0101] Where X A and X B Mole ratio (Molenbruch) represents the ratio of stereoisomers A and B.

[0102] The terms "selective conversion" or "increased selectivity" generally refer to the conversion of a specific stereoisomer, such as the E-form of the unsaturated hydrocarbon, at a higher proportion or amount (compared to molar amounts) than the corresponding other stereoisomers, such as the Z-form, during the entire course of the reaction (i.e., between the start and end of the reaction), at a certain point in time of the reaction, or during a "segment" of the reaction. Specifically, during the "segment," the selectivity can be observed corresponding to conversions of 1 to 99%, 2 to 95%, 3 to 90%, 5 to 85%, 10 to 80%, 15 to 75%, 20 to 70%, 25 to 65%, 30 to 60%, or 40 to 50% of the initial amount of substrate. The higher proportion or amount can be expressed, for example, as follows:

[0103] -High maximum yield of isomers observed throughout the entire reaction process or during a specified period;

[0104] - Higher relative abundance of isomers at a defined percentage substrate conversion value; and / or

[0105] -At higher conversion percentage values, the same relative content of isomers;

[0106] Each of these preferred methods is observed relative to a reference method, which is performed under otherwise identical conditions using known chemical or biochemical methods.

[0107] According to the present invention, all “isomeric forms” of the compounds described herein are generally included, such as structural isomers, especially stereoisomers and mixtures thereof, such as optical isomers or geometric isomers, such as E and Z isomers, and combinations thereof. If several asymmetric centers exist in a molecule, the present invention includes all combinations of different conformations of these asymmetric centers, such as enantiomer pairs, or any mixture of stereoisomeric forms.

[0108] In connection with the description and appended claims provided herein, unless otherwise stated, the use of “or” means “and / or”. Similarly, the tenses of “containing,” “containing,” “comprising,” and “including” are interchangeable and not restrictive.

[0109] It should be further understood that, where the term “comprising” is used in the description of various implementation schemes, those skilled in the art will understand that, in certain specific cases, the language of “substantially consisting of” or “consisting of” can be used instead to describe the implementation schemes.

[0110] If this disclosure relates to features, parameters, and ranges of different priorities (including superior, non-preferred features, parameters, and ranges), then unless otherwise stated, any combination of two or more of these features, parameters, and ranges is covered in the disclosure of this invention regardless of their respective priority.

[0111] b. Specific embodiments of the invention

[0112] This invention relates to the following embodiments:

[0113] 1. An isolated polypeptide from a haloacid dehalogenase-like (HAD-like) hydrolase superfamily, comprising cyclic terpene synthase activity, wherein said polypeptide is selected from:

[0114] a. BazzHAD1 comprising the amino acid sequence SEQ ID NO:3, or a mutant or natural variant thereof, comprising an amino acid sequence having at least 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% sequence identity with SEQ ID NO:3 and retaining the cyclic terpene synthase activity, particularly complementenyl diphosphate synthase activity;

[0115] b. BazzHAD2 comprising the amino acid sequence SEQ ID NO:6, or a mutant or natural variant thereof, comprising an amino acid sequence having at least 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% sequence identity with SEQ ID NO:6 and retaining the cyclic terpene synthase activity, particularly fuscinyl diphosphate synthase activity;

[0116] c. BazzHAD3 comprising the amino acid sequence SEQ ID NO:9, or a mutant or natural variant thereof, comprising an amino acid sequence having at least 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% sequence identity with SEQ ID NO:9 and retaining the cyclic terpene synthase activity, particularly fuscinyl diphosphate synthase activity;

[0117] d. BtHAD comprising the amino acid sequence SEQ ID NO:12, or a mutant or natural variant thereof, comprising an amino acid sequence having at least 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% sequence identity with SEQ ID NO:12 and retaining the cyclic terpene synthase activity, particularly complementenyl diphosphate synthase activity;

[0118] e. SmHAD1 comprising the amino acid sequence SEQ ID NO:19, or a mutant or natural variant thereof, comprising an amino acid sequence having at least 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% sequence identity with SEQ ID NO:19 and retaining the cyclic terpene synthase activity, particularly complementenyl diphosphate synthase activity, and / or particularly zephyranyl diphosphate synthase activity;

[0119] f. SmHAD2 comprising the amino acid sequence SEQ ID NO:26, or a mutant or natural variant thereof, comprising an amino acid sequence having at least 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% sequence identity with SEQ ID NO:19 and retaining the cyclic terpene synthase activity, particularly complementenyl diphosphate synthase activity, and / or particularly fussulaceanyl diphosphate synthase activity;

[0120] The specific TPS subfamily refers to items a), b), c), and d) above, and more specifically, the TPS listed in items a), b), and c).

[0121] A more specific embodiment of this embodiment involves polypeptide variants or, in particular, non-natural mutants of the novel polypeptides described above, of the present invention, which contain the cyclic terpene synthase activity as described above. These variants or non-natural mutants are derived from any specific amino acid sequence of SEQ ID NO:3, 6, 9, 12, 19, and 26. These variants are selected from polypeptides containing an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% sequence identity with any of SEQ ID NO:3, 6, 9, 12, 19, or 26, and containing at least one amino acid sequence difference relative to the corresponding unmodified polypeptide of SEQ ID NO:3, 6, 9, 12, 19, or 26, for example, at least one amino acid sequence position modified by addition, substitution, insertion, or deletion. Therefore, such variants have less than 100% sequence identity with the corresponding unmodified polypeptide. In this document, "at least one" encompasses at least 1 to 20, 1 to 15, 1 to 10, and more particularly 1, 2, 3, 4, or 5 amino acid sequence positions that are modified independently of each other by the addition, substitution, insertion, or deletion of amino acid residues.

[0122] 2. The polypeptide of implementation scheme 1, comprising the ability to produce the following substances:

[0123] a) at least one phosphate precursor of a sesquiterpene alcohol; or

[0124] b) at least one phosphate precursor of a sterane sesquiterpene alcohol and the corresponding at least one sterane sesquiterpene alcohol; or

[0125] c) At least one sesquiterpene alcohol.

[0126] 3. The polypeptide of embodiment 2, wherein the phosphate precursor is a styrannosyl monophosphate (ester), or more particularly, a styrannosyl diphosphate (ester); and / or wherein the styrannosyl sesquiterpene alcohol is styrannosyl alcohol and / or styrannosyl alcohol.

[0127] 4. The polypeptide produced according to implementation schemes 1, 2, or 3, including those that generate

[0128] a) Benzene diphosphate and / or folinose diphosphate derived from farnesyl diphosphate (FPP) as a substrate; or

[0129] b) Composinyl diphosphate and / or zebufoyl diphosphate derived from farnesyl diphosphate (FPP) as a substrate; and composinol and / or zebufool, derived directly from FPP as a substrate or via their respective diphosphate precursors; or

[0130] c) Composterol and / or zedoaryl from FPPs used as substrates, either directly from FPPs used as substrates or via their respective diphosphate precursors.

[0131] Table 1 below shows an overview of specific TPSs of the present invention, including their specific sequence motifs, their relative positions in the complete amino acid sequence, and their respective product profiles (based on the type of complementalanol).

[0132]

[0133] 5. The polypeptide of one of the aforementioned embodiments, which is isolated from...

[0134] a. Plants of the genus *Bazzania*, especially *Bazzania trilobata*; or

[0135] b. Plants of the genus Selaginella, especially the species Selaginella moellendorffii;

[0136] c. Other lycophytes, liverworts, or mosilophytes containing sesquiterpenes, including but not limited to the genera *Porella*, *Hymenophyton*, and *Marchantia*.

[0137] Or it may come from other plant species.

[0138] 6. The polypeptide of any one of embodiments 1 to 5, further comprising:

[0139] a. The modified type II terpene synthase motif shown in SEQ ID NO:46(Px0Dx1D(T / S)(T / M)S), wherein x0 and x1 can be any naturally occurring amino acid residue, and x1 specifically represents I or L, particularly the motif of any one of SEQ ID NO:47, 48, 49, or 50, wherein the modified type II terpene synthase motif corresponds to:

[0140] i. The sequence positions of bits 318 to 325 of SEQ ID NO:3

[0141] ii. The sequence positions of bits 269 to 276 of SEQ ID NO:6

[0142] iii. The sequence positions of bits 277 to 284 of SEQ ID NO:9,

[0143] iv. The sequence positions of bits 318 to 325 of SEQ ID NO:12,

[0144] v. The sequence positions of positions 255 to 262 of SEQ ID NO:19, and

[0145] vi. The sequence positions of bits 255 to 262 of SEQ ID NO:26;

[0146] b. The QW motif shown in SEQ ID NO:51 (Qx2x3DGx4W), wherein x2, x3, and x4 can be any naturally occurring amino acid residue, and x4 specifically represents G or S, particularly the motif of any one of SEQ ID NO:52, 53, 54, or 55, wherein the QW motif corresponds to:

[0147] i. The sequence positions of SEQ ID NO:3, from position 489 to 495.

[0148] ii. The sequence positions of bits 431 to 437 of SEQ ID NO:6

[0149] iii. The sequence positions of bits 439 to 445 of SEQ ID NO:9

[0150] iv. Sequence positions 488 to 494 of SEQ ID NO:12

[0151] v. Sequence positions 430 to 436 of SEQ ID NO:19, and

[0152] vi. The sequence positions of bits 430 to 436 of SEQ ID NO:26; and

[0153] c. Optionally, at least one, for example, 1, 2, or 3 additional sequence motifs selected from:

[0154] (1)SEQ ID NO:56(Lx5x6x7x8(W / F)x9x 10 Yx 11 x 12 G) represents the conservative sequence motif 1, where x5 to x 12 It can be any naturally occurring amino acid residue, x5 specifically represents R or Q, x8 specifically represents S, I or T, x 11 Specifically representing E or S, and x 12 Specifically representing C or M, particularly the motifs of SEQ ID NO: 57, 58, 59, 60, or 61, wherein the conserved sequence motif 1 corresponds to:

[0155] i. The sequence positions of bits 68 to 79 of SEQ ID NO:3

[0156] ii. The sequence positions of bits 37 to 48 of SEQ ID NO:6

[0157] iii. The sequence positions of bits 45 to 56 of SEQ ID NO:9

[0158] iv. The sequence positions of bits 68 to 79 of SEQ ID NO:12,

[0159] v. The sequence positions of positions 28 to 39 of SEQ ID NO:19, and

[0160] vi. The sequence positions of bits 28 to 39 of SEQ ID NO:26;

[0161] (2)SEQ ID NO:62(Yx 13 Dx 14 x 15 Rx 16 RVD(P / A)V(V / A)x 17 x 18 The conserved sequence motif 2 shown in N) is x 13 To x 18 It can be any naturally occurring amino acid residue, x 13 Specifically represents F or L, and x 16 Specifically representing P or L, particularly the motifs of SEQ ID NO: 63, 64, 65, 66, 67 or 68, wherein the conserved sequence motif 2 corresponds to:

[0162] i. The sequence positions of bits 362 to 377 of SEQ ID NO:3

[0163] ii. The sequence positions of bits 309 to 324 of SEQ ID NO:6

[0164] iii. The sequence positions of bits 317 to 332 of SEQ ID NO:9

[0165] iv. The sequence positions of bits 362 to 377 of SEQ ID NO:12,

[0166] v. Sequence positions 299 to 314 of SEQ ID NO:19, and

[0167] vi. The sequence positions of SEQ ID NO:26, positions 299 to 314; and

[0168] (3)SEQ ID NO:69(GTx 19 (Y / F)Yx 20 x 21 x 22 Ex 23 The conserved sequence motif 3 is shown in FL(Y / F), where x19 To x 23 It can be any naturally occurring amino acid residue, x 19 Specifically representing L or R, particularly the motifs of SEQ ID NO: 70, 71 or 72, wherein the conserved sequence motif 3 corresponds to:

[0169] i. The sequence positions of bits 410 to 422 of SEQ ID NO:3

[0170] ii. The sequence positions of bits 357 to 369 of SEQ ID NO:6

[0171] iii. The sequence positions of bits 365 to 377 of SEQ ID NO:9

[0172] iv. The sequence positions of bits 410 to 422 of SEQ ID NO:12,

[0173] v. Sequence positions 349 to 361 of SEQ ID NO:19, and

[0174] vi. Sequence positions 349 to 361 of SEQ ID NO:26.

[0175] More specifically, the TPS of the present invention includes type II motifs, QW motifs, and at least one of conservative motifs 1, 2 and 3.

[0176] More specifically, the TPS of the present invention includes type II motifs, QW motifs, and at least two of the conservative motifs 1, 2 and 3.

[0177] More specifically, the TPS of the present invention includes type II motifs, QW motifs, and all three conserved motifs 1, 2 and 3.

[0178] The relative order of these sequence motifs in embodiment 6 within their respective amino acid sequences of each TPS of the present invention is described below (see also...). Figure 21 ):

[0179] N-terminus --- (Conservative motif 1) --- (Type II motif) --- (Conservative motif 2) ---

[0180] ---(Conservative motif 3)---(QW motif)---C-terminus

[0181] Type I motifs known from other TPSs (which are missing in the TPS of this invention) will be located between conserved motif 1 and type II motifs (see...). Figure 21 ).

[0182] 7. The polypeptide of any of the foregoing embodiments, which catalyzes the conversion of acyclic farnesyl diphosphate (particularly (2E,6E)-3,7,11-trimethyldodec-2,6,10-triene-1-pyrophosphate; FPP) into sero-alkenyl phosphate derivatives, such as monophosphates, more particularly sero-alkenyl diphosphates, especially with a selectivity of 50% ee or higher, such as 50 to 100% ee, or 60 to 90% ee or 70 to 80% ee.

[0183] 8. The polypeptide of any of the foregoing embodiments, which catalyzes the conversion of acyclic farnesyl diphosphate (particularly (2E,6E)-3,7,11-trimethyldodec-2,6,10-triene-1-pyrophosphate; FPP) into fussulacean phosphate derivatives, such as monophosphates, more particularly fussulacean phosphate, especially with a selectivity of 50% ee or higher, such as 50 to 100% ee, or 60 to 90% ee or 70 to 80% ee.

[0184] 9. The isolated polypeptide of any of the foregoing embodiments, wherein

[0185] a. Contains an amino acid sequence selected from SEQ ID NO:3, 6, 9, 12, 19, and 26; or

[0186] b. Encoded by a nucleic acid molecule containing a coding nucleotide sequence selected from SEQ ID NO:1, 2, 4, 5, 7, 8, 10, 11, 17, 18, 24 and 25.

[0187] 10. The isolated polypeptide of Implementation Scheme 9, which

[0188] a. Composed of amino acid sequences selected from SEQ ID NO:3, 6, 9, 12, 19, and 26; or

[0189] b. Encoded by a nucleic acid molecule consisting of a coding nucleotide sequence selected from SEQ ID NO:1, 2, 4, 5, 7, 8, 10, 11, 17, 18, 24 and 25.

[0190] 11. An isolated nucleic acid molecule, which:

[0191] a. A nucleotide sequence comprising a polypeptide encoding any of the foregoing embodiments; or

[0192] b. Containing a nucleotide sequence selected from SEQ ID NO:1, 2, 4, 5, 7, 8, 10, 11, 17, 18, 24, and 25, or containing a nucleotide sequence having at least 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% sequence identity with the nucleotide sequence SEQ ID NO:1, 2, 4, 5, 7, 8, 17, 18, or 24 or 25, and encoding a polypeptide of the HAD-like hydrolase superfamily, the polypeptide containing terpene synthase activity, particularly the ability to produce the following from farnesyl diphosphate (FPP) as a substrate:

[0193] Bruxine sesquiterpene alcohols, and / or their phosphate precursors, such as monophosphates, and more particularly their diphosphate precursors,

[0194] In particular, styrenyl phosphate precursors, such as monophosphates, and even more so styrenyl diphosphates; and / or styrenyl phosphate precursors, such as monophosphates, and even more so styrenyl diphosphates;

[0195] More specifically, as described in Table 1 above; or

[0196] c. contains a nucleotide sequence that includes a sequence complementary to one of the sequences in b.; or

[0197] d. A nucleotide sequence that hybridizes with a, b, or c under strict conditions.

[0198] In one particular implementation, the nucleic acid may be naturally present in plants of the genera *Bazzania* or *Selaginella*, such as *Bazzania trilobata* or *Selaginella moellendorffii*, or other plant species, or obtained by modifying SEQ ID NO: 1, 2, 4, 5, 7, 8, 10, 11, 17, 18, 24 and 25 or their reverse complementary sequences.

[0199] In another implementation, the nucleic acid is isolated or derived from plants of the genera *Bazzania* or *Selaginella*, such as *Bazzania trilobata* or *Selaginella moellendorffii*.

[0200] 12. An expression construct comprising at least one nucleic acid molecule of embodiment 11, optionally combined with at least one regulatory sequence.

[0201] 13. A vector comprising at least one nucleic acid molecule of embodiment 11 or at least one expression construct of embodiment 12.

[0202] 14. The vector of implementation scheme 13, wherein the vector is a prokaryotic, viral or eukaryotic vector.

[0203] 15. The carrier of implementation scheme 13 or 14, wherein the carrier is an expression carrier.

[0204] 16. The vector of any one of Implementation Schemes 13 to 15 is a plasmid vector.

[0205] 17. A recombinant non-natural non-human host cell or recombinant non-human host organism prepared by genetic engineering, comprising:

[0206] a. At least one isolated nucleic acid molecule of embodiment 11, which optionally stably integrates into the genome; or

[0207] b. At least one expression construct of embodiment 12, which optionally is stably integrated into the genome; or

[0208] c. At least one carrier of any one of embodiments 13 to 16.

[0209] The recombinant non-human host cell or recombinant non-human host organism of this embodiment differs from natural non-human host cells or non-human host organisms in that exogenous genetic material, such as isolated nucleic acids, expression constructs, or vectors, is artificially introduced into the host organism or host cell. The recombinant non-human host cell or recombinant non-human host organism of this embodiment is a non-human host cell or non-human host organism that has been (genetically) modified to contain the desired exogenous genetic material, such as isolated nucleic acids, expression constructs, or vectors, as described in the above embodiments. In particular, the recombinant non-human host cell or recombinant non-human host organism of this embodiment is a cell or organism that differs from cells or organisms that naturally possess the desired genetic material.

[0210] 18. The host cell or host organism of implementation scheme 17 is selected from prokaryotes or eukaryotes, or derived from their cells.

[0211] 19. The host cell or host organism of implementation scheme 18 is selected from bacterial cells, fungal cells and plant cells, or plants.

[0212] 20. The host cell or host organism of embodiment 19, wherein the fungal cell is a yeast cell.

[0213] 21. The host cell or host organism of embodiment 20, wherein the yeast cell is selected from the genera Saccharomyces or Pichia, particularly from the species Saccharomyces cerevisiae or Pichia pastoris.

[0214] 22. The host cell or host organism of embodiment 18, wherein the bacterial cell is selected from the genus Escherichia, particularly Escherichia coli.

[0215] Some of these host cells or host organisms cannot naturally produce FPP. To suit the methods described herein, organisms or cells that do not naturally produce acyclic terpene pyrophosphate precursors such as FPP are genetically modified, particularly by transformation, transduction, or conjugation, and more particularly by transformation, to produce the precursors. Non-human host organisms or non-human host cells may be modified (e.g., transformed) before or simultaneously with the nucleic acid modification according to any of the above embodiments. Methods for modifying (e.g., transforming) organisms to produce acyclic terpene pyrophosphate precursors such as FPP are known in the art. For example, introducing enzymatic activity of the mevalonate pathway (isoprene-like pathway) or the MEP pathway is a suitable strategy for producing FPP from an organism (see also the Examples section herein).

[0216] 23. A method for generating at least one catalytically active polypeptide according to any one of embodiments 1 to 10, the method comprising:

[0217] a. Culturing non-human host cells or non-human host organisms according to embodiment 17 to express at least one polypeptide of any one of embodiments 1 to 10; and

[0218] b. Optionally, the polypeptide may be isolated from non-human host cells or organisms cultured in step a.

[0219] 24. The method of embodiment 23 further includes, prior to step a, genetically modifying the non-human host cell or non-human host organism to express a polypeptide according to any one of embodiments 1 to 10 by inserting at least one nucleic acid of embodiment 11, at least one expression construct of embodiment 12, or at least one vector of any one of embodiments 13 to 16 into the non-human host cell or non-human host organism, particularly by transformation, transduction, or conjugation.

[0220] 25. A method for producing sterine sesquiterpene alcohols, particularly sterine alcohol and / or zephyranol, the method comprising:

[0221] a. Contacting farnesyl diphosphate (FPP) with a polypeptide as defined in any one of embodiments 1 to 10 or with a polypeptide prepared according to embodiment 23 or 24 to obtain at least one phosphate ester, particularly a diphosphate ester of styrannosesquiterpene alcohol, particularly styrannosyl and / or zephyranosyl phosphate esters, more particularly styrannosyl diphosphate and / or zephyranosyl diphosphate;

[0222] b. Cleavage the phosphate ester portion, particularly the diphosphate ester portion, of the product obtained in step a. by chemical or enzymatic means; and

[0223] c. Optionally, sterane sesquiterpene alcohols, particularly sterane alcohol and / or zestyrax alcohol, are isolated.

[0224] In one particular embodiment, the phosphate ester moiety is enzymatically cleaved by applying a phosphatase, more specifically an acidic or alkaline phosphatase. Alkaline phosphatases from various sources, such as bacterial enzymes, are preferred. Suitable phosphatases are commercially available enzymes.

[0225] In one embodiment, sterane sesquiterpene alcohols are isolated, particularly sterane alcohol and / or zephyranol, or mixtures containing sterane sesquiterpene alcohols.

[0226] In one particular embodiment, the method further includes step d, processing the styrannosyl sesquiterpene alcohol formed in step b or isolated in step c, particularly styrannosyl alcohol and / or styrannosyl alcohol, using chemical synthesis or biocatalytic synthesis (e.g., biochemical synthesis with an expressed enzyme, or biotransformation with living cells expressing the enzyme) or a combination of both to obtain derivatives thereof. The styrannosyl sesquiterpene alcohol derivatives may be particularly selected from hydrocarbons, alcohols, diols, triols, acetals, ketals, aldehydes, acids, ethers, amides, ketones, lactones, epoxides, acetates, glycosides, and / or esters. In one embodiment, the method includes contacting the styrannosyl sesquiterpene alcohol with at least one enzyme to produce the styrannosyl sesquiterpene alcohol derivative. Specifically, the styrannosyl sesquiterpene alcohol derivative can be obtained by contacting the styrannosyl sesquiterpene alcohol with an enzyme such as, but not limited to, oxidoreductases, monooxygenases, dioxygenases, and transferases using a biochemical method. Biochemical transformation can be carried out in vitro using isolated enzymes, enzymes from lysed cells, or in vivo using whole cells. In another embodiment, the above method includes converting styrannosyl sesquiterpene alcohols into styrannosyl sesquiterpene alcohol derivatives using chemical synthesis. Specifically, this can be obtained by chemical methods, such as, but not limited to, oxidation, reduction, alkylation, acylation, and / or rearrangement. In another embodiment, the above method includes step e., optionally isolating the derivative of step d.

[0227] 26. The method of embodiment 25, wherein the sterane sesquiterpene alcohol comprises sterane alcohol and / or zephyranol alcohol, particularly as the main product.

[0228] 27. The method of any one of embodiments 25 and 26, comprising providing, in particular, transforming a non-human host cell or non-human host organism with at least one nucleic acid of embodiment 11, at least one expression construct of embodiment 12, or at least one vector of embodiments 13 to 16, such that the non-human host cell or non-human host organism expresses a polypeptide according to any one of embodiments 1 to 10.

[0229] 28. The method of any one of embodiments 25 to 27, wherein the FPP is contacted with a non-human host cell or non-human host organism of any one of embodiments 17 to 22, with its cell lysate or with a culture medium containing said non-human host cell or non-human host organism; and / or with a polypeptide as defined in any one of embodiments 1 to 10, said polypeptide being expressed or isolated in said non-human host cell or non-human host organism, cell lysate or culture medium.

[0230] 29. The method of embodiment 28, wherein the sesquiterpene alcohol is produced by fermentation from the non-human host organism or non-human host cell.

[0231] 30. The method of embodiment 28, wherein the styrannosesquiterpene alcohol is produced by an enzymatic method comprising converting the FPP with an isolated polypeptide of any one of embodiments 1 to 10, the step optionally being carried out in the presence of other adjuvants.

[0232] In a particular embodiment, in the method of any one of embodiments 25 to 30, the polypeptide comprises:

[0233] a. An amino acid sequence having at least 90% sequence identity with SEQ ID NO:3 or SEQ ID NO:12;

[0234] b. A class II synthase motif possessing the amino acid sequence PDDLDSTS (SEQ ID NO:47); and

[0235] c. A QW motif possessing the amino acid sequence QNVDGSW (SEQ ID NO:52); and

[0236] d. Optionally, at least one other sequence motif selected from the following amino acid sequences:

[0237] i.Lxxxx(W / F)xxYxxG(SEQ ID NO:56), where x can be any naturally occurring amino acid residue, especially the amino acid sequence LRSHIWFNYSMG(SEQ ID NO:57);

[0238] ii.YxDxxRxRVD(P / A)V(V / A)xxN(SEQ ID NO:62), where x can be any naturally occurring amino acid residue, particularly the amino acid sequence YFDPLRLRVDPVAATN(SEQ ID NO:63); and

[0239] iii.GTx(Y / F)xYxxxExFL(Y / F) (SEQ ID NO:69), where x can be any naturally occurring amino acid residue, especially the amino acid sequence GTLYYRTPEAFLY (SEQ ID NO:70);

[0240] Furthermore, this sesquiterpene alcohol contains sesquiterpene alcohol.

[0241] In another specific embodiment, in the method of any one of embodiments 25 to 30, the polypeptide comprises:

[0242] a. An amino acid sequence having at least 90% sequence identity with SEQ ID NO:6 or SEQ ID NO:9;

[0243] b. A class II synthase motif having the amino acid sequence PxDxD(T / S)(T / M)S (SEQ ID NO:46), wherein x can be any naturally occurring amino acid residue, particularly the amino acid sequences PDDLDTTS (SEQ ID NO:48) or PNDLDTTS (SEQ ID NO:50); and

[0244] c. A QW motif comprising the amino acid sequence Qx1x2DGx3W (SEQ ID NO:51), wherein x1 and x2 can be any naturally occurring amino acid residue, and x3 is G, particularly the amino acid sequences QCDDGGW (SEQ ID NO:53) or QSSDGGW (SEQ ID NO:54); and

[0245] d. Optionally, at least one other sequence motif selected from the following amino acid sequences:

[0246] i.LRxxTWxxYECG(SEQ ID NO:61), where x can be any naturally occurring amino acid residue, especially the amino acid sequence LRSATWAAYECG(SEQ ID NO:58) or LRTPTWGKYECG(SEQ ID NO:59);

[0247] ii. Y(F / L)Dx(T / E)RPRVD(P / A)VVx(A / V)N (SEQ ID NO:68), where x can be any naturally occurring amino acid residue, particularly the amino acid sequence YFDETRPRVDAVVNVN (SEQ ID NO:64) or YFDKTRPRVDPVVCVN (SEQ ID NO:65); and

[0248] iii.GTx(Y / F)YxxxExFL(Y / F) (SEQ ID NO:69), where x can be any naturally occurring amino acid residue, especially the amino acid sequence GTLFYYHAESFLY (SEQ ID NO:71);

[0249] Furthermore, this sesquiterpene alcohol contains physostigmine alcohol.

[0250] In another specific embodiment, in the method of any one of embodiments 25 to 30, the polypeptide comprises:

[0251] a. An amino acid sequence having at least 90% sequence identity with SEQ ID NO:19 or SEQ ID NO:26;

[0252] b. A class II synthase motif possessing the amino acid sequence PPDIDDTMS (SEQ ID NO:49); and

[0253] c. A QW motif possessing the amino acid sequence QNEDGSW (SEQ ID NO:55); and

[0254] d. Optionally, at least one other sequence motif selected from the following amino acid sequences:

[0255] i.Lxxxx(W / F)xxYxxG(SEQ ID NO:56), where x can be any naturally occurring amino acid residue, especially the amino acid sequence LQHSSFLAYSCG(SEQ ID NO:60);

[0256] ii. Y(F / L)Dx(T / E)RPRVD(P / A)VVx(A / V)N (SEQ ID NO:68), where x can be any naturally occurring amino acid residue, particularly the amino acid sequence YLDVERPRVDPVVIAN (SEQ ID NO:66) or YLDLERPRVDPVVIAN (SEQ ID NO:67); and

[0257] iii.GTx(Y / F)xYxxxExFLx (SEQ ID NO:69), where x can be any naturally occurring amino acid residue, especially the amino acid sequence GTRYYLSQEDFLF (SEQ ID NO:72);

[0258] Furthermore, this sterane sesquiterpene alcohol contains sterane alcohol and zephyranol.

[0259] More specifically, the applied TPS includes at least one of type II motifs, QW motifs, and conserved motifs i., ii. and iii.

[0260] More specifically, the applied TPS includes at least two of the following: type II motifs, QW motifs, and conserved motifs i., ii., and iii.

[0261] More specifically, the applied TPS includes type II motifs, QW motifs, and all three conserved motifs i., ii., and iii.

[0262] 31. Use of the polypeptide as defined in any one of embodiments 1 to 10 in the preparation of odorant, flavoring or fragrance ingredients, particularly Ambrox (preferably via styrosine / styrosine diphosphate).

[0263] 32. Use of styrannosyl sesquiterpene alcohol prepared according to any one of embodiments 25 to 30 in the preparation of odorant, flavoring or fragrance ingredients, particularly Ambrox.

[0264] 33. A method for generating Ambrox, the method comprising:

[0265] a. To provide complementol and / or zephyranthesol by the method of any one of the embodiments described in any one of embodiments 25 to 30.

[0266] b. Optionally, the sterol and / or zephyranol produced in step a. are separated; and

[0267] c. Converting sterol and / or zedoaryol to Ambrox in a manner known per se, as reported in Tetrahedron: Asymmetry 11 (2000) 1375-1388.

[0268] 34. A composition comprising a substance prepared according to embodiment 32 or 33.

[0269] 35. The composition according to embodiment 34, which is selected from body care compositions, home care compositions and fragrance compositions.

[0270] 36. A method for producing sterine sesquiterpene alcohols, particularly sterine alcohol and / or zephyranol, the method comprising:

[0271] a. Culturing a non-human host organism or non-human host cell capable of producing FPP and converting it to express the polypeptide of any one of embodiments 1 to 10; and

[0272] b. Optionally, sterane sesquiterpene alcohols, particularly sterane alcohol and / or zestyrax alcohol, are isolated.

[0273] In one embodiment, a non-human host organism or non-human host cell that does not naturally produce FPP is genetically modified, particularly by transformation, transduction, or conjugation, and more specifically by transformation, to produce the precursor. The non-human host organism or non-human host cell may be modified (e.g., transformed) before or simultaneously with the nucleic acid modification according to any of the above embodiments. Methods for modifying (e.g., transforming) an organism to produce acyclic terpene pyrophosphate precursors such as FPP are known in the art. For example, introducing enzymatic activity of the mevalonate pathway (isoprene-like pathway) or the MEP pathway is a suitable strategy for producing FPP from an organism (see also the Examples section herein).

[0274] 37. A method for preparing mutant polypeptides of the haloacid dehalogenase-like (HAD-like) hydrolase superfamily, comprising terpene synthase activity, particularly comprising the ability to produce styryl sesquiterpene alcohols and / or their phosphate derivatives such as monophosphates, more particularly their diphosphate derivatives; and particularly the ability to produce styryl alkenyl phosphate derivatives such as monophosphates, more particularly styryl alkenyl diphosphates, and / or zephyranyl phosphate derivatives such as monophosphates, more particularly zephyranyl diphosphates, from farnesyl diphosphate (FPP) as a substrate, the method comprising the steps of:

[0275] a. Select nucleic acid molecules according to implementation plan 11;

[0276] b. Modify the selected nucleic acid molecule to obtain at least one mutant nucleic acid molecule;

[0277] c. Genetically modifying non-human host cells or single-celled non-human host organisms with a mutant nucleic acid sequence, particularly by horizontal gene transfer such as transformation, transduction or conjugation, and more particularly by transforming said host cells or single-celled host organisms, to express a polypeptide encoded by the mutant nucleic acid sequence;

[0278] d. Screen for at least one mutant of the expression product, which contains terpene synthase activity, particularly the ability to produce styryl sesquiterpene alcohols and / or their phosphate precursors such as monophosphates, more particularly their diphosphate derivatives; and particularly the ability to produce styryl alkenyl phosphate precursors such as monophosphates, more particularly styryl alkenyl diphosphates, and / or styryl phosphate precursors such as monophosphates, more particularly styryl alkenyl diphosphates, from FPP as a substrate; and,

[0279] e. Optionally, if the peptide does not have the desired mutant activity, repeat steps a. to d. until a peptide with the desired mutant activity is obtained; and,

[0280] f. Optionally, if a polypeptide with the desired mutant activity is identified in step d, the corresponding mutant nucleic acid obtained in step c. is isolated.

[0281] Unless otherwise stated, within the range of “at least 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% sequence identity,” the particular values ​​are at least 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% sequence identity, while the values ​​of at least 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% sequence identity are even more particular.

[0282] c. The polypeptide applicable according to the present invention

[0283] In the context of this article, the following definitions apply:

[0284] The commonly used terms “polypeptide” or “peptide” refer to a natural or synthetic, continuous, peptide-linked linear chain or sequence of amino acid residues containing about 10 to more than 1,000 residues. In some embodiments provided herein, a polypeptide comprises an amino acid sequence that serves as an enzyme, or a fragment or variant thereof. Short-chain polypeptides having up to 30 residues are also referred to as “oligopeptides”.

[0285] The term "isolated polypeptide" refers to an amino acid sequence extracted from its natural environment by any method known in the art or a combination of such methods (including recombinant, biochemical, and synthetic methods).

[0286] The term "protein" refers to a large molecular structure composed of one or more polypeptides. It includes oligopeptides, peptides, polypeptides, and full-length proteins, whether natural or synthetic. The amino acid sequence of a polypeptide represents the protein's "primary structure." The amino acid sequence also predetermines the protein's "secondary structure" by forming specific structural elements (such as α-helices and β-sheets formed within the polypeptide chain). The arrangement of multiple such secondary structural elements defines the protein's "tertiary structure," or spatial arrangement. If a protein contains more than one polypeptide chain, these chains are arranged spatially to form the protein's "quaternary structure." Proper spatial arrangement, or "folding," is a prerequisite for protein function. Denaturation or unfolding disrupts protein function. If this disruption is reversible, protein function can be restored by refolding.

[0287] The typical protein function referred to in this article is "enzyme function," which means that a protein acts as a biocatalyst on a substrate, such as a compound, and catalyzes the conversion of said substrate into a product. Enzymes can exhibit high or low levels of substrate and / or product specificity.

[0288] Therefore, the term "polypeptide" as used herein to refer to a specific "activity" implicitly refers to a properly folded protein that exhibits the indicated activity, such as specific enzyme activity. Thus, unless otherwise stated, the term "polypeptide" also encompasses the terms "protein" and "enzyme".

[0289] A "target peptide" is an amino acid sequence that targets a protein or polypeptide to intracellular organelles (i.e., mitochondria or plastids) or to the extracellular space (secretory signal peptides). The nucleic acid sequence encoding the target peptide can be fused to the amino-terminal (e.g., N-terminus) nucleic acid sequence encoding the protein or polypeptide, or it can be used to replace the natural target peptide.

[0290] The present invention also relates to “functional equivalents” (also referred to as “analogs” or “functional mutations”) of the polypeptides specifically described herein.

[0291] For example, a “functional equivalent” refers to a polypeptide that, in a test used to determine the activity of enzymatic complementenyl and / or zephyranoid diphosphate synthase, shows at least 1 to 10%, or at least 20%, or at least 50%, or at least 75%, or at least 90% higher or lower complementenyl and / or zephyranoid diphosphate synthase activity than the polypeptide specifically described herein.

[0292] According to the invention, "functional equivalents" also encompass specific mutants that have an amino acid at at least one sequence position in the amino acid sequence described herein that differs from the specifically stated amino acid, but still possess one of the aforementioned biological activities, such as enzyme activity. Thus, "functional equivalents" include mutants obtainable by the addition, substitution, particularly conservative substitution (i.e., the amino acid in question is replaced by an amino acid having the same charge, size, polarity, and / or solubility), deletion, and / or inversion of one or more, for example, 1 to 20, 1 to 15, or 5 to 10 amino acids, wherein said changes can occur at any sequence position, as long as they result in the mutant possessing the general characteristics of the invention. Functional equivalence is also particularly provided if the activity pattern qualitatively overlaps between the mutant and the unaltered polypeptide, i.e., if, for example, an interaction with the same agonist or antagonist or substrate is observed, but at different rates (i.e., by EC...). 50 or IC 50 (Values ​​or any other parameters suitable in this technical field). The table below shows examples of suitable (conservative) amino acid substitutions:

[0293]

[0294]

[0295] The “functional equivalents” in the above sense are also the “precursors” of the polypeptides described in this article, as well as the “functional derivatives” and “salts” of the polypeptides.

[0296] In this case, a "precursor" is a natural or synthetic precursor of a polypeptide, which may or may not have the desired biological activity.

[0297] The term "salt" as used in this invention refers to salts of the carboxyl group of the protein molecule and salts formed by the acid addition of the amino group. Salts of the carboxyl group can be produced in known ways, including inorganic salts such as sodium, calcium, ammonium, iron, and zinc salts, as well as salts formed with organic bases such as amines, such as triethanolamine, arginine, lysine, piperidine, etc. Salts formed by acid addition, such as those formed with inorganic acids such as hydrochloric acid or sulfuric acid, and salts formed with organic acids such as acetic acid and oxalic acid, are also covered by this invention.

[0298] The “functional derivatives” of the polypeptides according to the invention can also be generated using known techniques at the side groups of functional amino acids or their N-terminus or C-terminus. Such derivatives include, for example: aliphatic esters of carboxylic acid groups, amides of carboxylic acid groups, which can be obtained by reacting with ammonia or with primary or secondary amines; N-acyl derivatives of free amino groups, which are generated by reacting with acyl groups; or O-acyl derivatives of free hydroxyl groups, which are generated by reacting with acyl groups.

[0299] "Functional equivalents" naturally include polypeptides that can be obtained from other organisms as well as naturally occurring variants. For example, the area of ​​homologous sequence regions can be determined by sequence comparison, and equivalent polypeptides can be determined based on the specific parameters of this invention.

[0300] "Functional equivalents" also include "fragments" of the polypeptide according to the invention, such as single domains or sequence motifs, or truncated N-terminuses and / or C-termini, which may or may not exhibit the desired biological function. Preferably, such "fragments" at least qualitatively retain the desired biological function.

[0301] Furthermore, a “functional equivalent” is a fusion protein having one of the polypeptide sequences described herein or a functional equivalent derived therefrom, and having at least one additional functionally distinct heterologous sequence in functional N-terminal or C-terminal association (i.e., without substantial mutual functional impairment of the fusion protein portion). Non-limiting examples of such heterologous sequences are, for example, signal peptides, histidine anchors, or enzymes.

[0302] The invention also includes “functional equivalents” that are homologs of the specifically disclosed polypeptides. They have at least 60%, preferably at least 75%, particularly at least 80 or 85%, such as 90, 91, 92, 93, 94, 95, 96, 97, 98, or 99% homology (or identity) with one of the specifically disclosed amino acid sequences, calculated using the algorithm described in Pearson and Lipman, Proc. Natl. Acad. Sci. (USA) 85(8), 1988, 2444-2448. The homology or identity of the homologous polypeptides according to the invention, expressed as a percentage, refers in particular to identity expressed as a percentage of amino acid residues based on the total length of one of the amino acid sequences specifically described herein.

[0303] Identity data expressed as a percentage can also be determined using BLAST alignment, the blastp (protein-protein BLAST) algorithm, or by applying the Clustal settings detailed below.

[0304] In the case of possible protein glycosylation, the “functional equivalents” according to the invention include polypeptides in deglycosylated or glycosylated forms as described herein, as well as modified forms that can be obtained by changing the glycosylation pattern.

[0305] Functional equivalents or homologs of the polypeptides according to the present invention can be generated by mutagenesis, for example by point mutation, lengthening or shortening of the protein, or as described in more detail below.

[0306] Functional equivalents or homologs of the polypeptides according to the invention can be identified by screening a database of mutants, such as shortened mutants. For example, a database of protein variant diversity can be generated by combinatorial mutagenesis at the nucleic acid level, for example by enzymatic ligation of a mixture of synthetic oligonucleotides. Numerous methods are available for generating a database of potential homologs from degenerate oligonucleotide sequences. The chemical synthesis of degenerate gene sequences can be performed in an automated DNA synthesizer, and the synthesized gene can then be ligated into a suitable expression vector. The use of degenerate genomes makes it possible to provide all sequences in a mixture that encode the desired set of potential protein sequences. Methods for synthesizing degenerate oligonucleotides are known to those skilled in the art.

[0307] In the prior art, several techniques are known for screening gene products from combinatorial databases generated by point mutations or shortening, and for screening cDNA libraries containing gene products with selected properties. These techniques can be applied to rapidly screen gene libraries generated by combinatorial mutagenesis of homologs according to the present invention. The most commonly used high-throughput analysis-based techniques for screening large gene libraries involve cloning the gene library in a reproducible expression vector, transforming suitable cells with the resulting vector database, and expressing the combinatorial gene under specific conditions, under which detection of the desired activity facilitates the isolation of vectors encoding the gene (whose product is detected). Recursive integration mutagenesis (REM) is a technique for increasing the frequency of functional mutants in a database and can be used in conjunction with screening tests to identify homologs.

[0308] The embodiments provided herein offer orthologs and paralogs of the disclosed peptides, as well as methods for identifying and isolating such orthologs and paralogs. The definitions of the terms "ortholog" and "paralog" are given below and apply to both amino acid and nucleic acid sequences.

[0309] The polypeptides of this invention comprise all active forms of the enzymes of this invention, including active subsequences, such as catalytic domains or active sites. In one embodiment, this invention provides the catalytic domains or active sites as described below. In one embodiment, the present invention provides a peptide or polypeptide comprising or composed of an active site domain predicted by using a database such as Pfam (http: / / pfam.wustl.edu / hmmsearch.shtml) (a large collection covering multiple sequence alignments and hidden Markov models for many common protein families, Pfam Protein Family Database, A. Bateman, E. Birney, L. Cerruti, R. Durbin, L. Etwiller, S. REddy, S. Griffiths-Jones, K. L. Howe, M. Marshall, and E. L. Sonnhammer, Nucleic Acids Research, 30(1):276-280, 2002) or equivalent sources such as the InterPro and SMART databases (http: / / www.ebi.ac.uk / interpro / scan.html, http: / / smart.embl-heidelberg.de / ).

[0310] The present invention also covers “peptide variants” having the desired activity, wherein the variant peptide is selected from amino acid sequences having at least 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity with the specific, particularly natural, amino acid sequence referred to by the specific SEQ ID NO and containing at least one, for example, 1 to 30 or 1 to 20, such as 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 amino acid sequence differences, for example, addition, substitution, insertion, or deletion of amino acids relative to the (unmodified) SEQ ID NO.

[0311] d. The applicable coding nucleic acid sequence according to the present invention

[0312] In the context of this article, the following definitions apply:

[0313] The terms “nucleic acid sequence,” “nucleic acid,” “nucleic acid molecule,” and “polynucleotide” are used interchangeably and refer to a sequence of nucleotides. A nucleic acid sequence can be a single-stranded or double-stranded deoxyribonucleotide or ribonucleotide of any length and includes coding and non-coding sequences of genes, exons, introns, sense and antisense complementary sequences, genomic DNA, cDNA, miRNA, siRNA, mRNA, rRNA, tRNA, recombinant nucleic acid sequences, isolated and purified naturally occurring DNA and / or RNA sequences, synthetic DNA and RNA sequences, fragments, primers, and nucleic acid probes. Those skilled in the art understand that the nucleic acid sequence of RNA is identical to that of DNA, except that thymine (T) is replaced by uracil (U). The term “nucleotide sequence” should also be understood to include polynucleotide or oligonucleotide molecules in the form of individual fragments or as components of larger nucleic acids.

[0314] As used in this article, the term “naturally occurring” for nucleic acids refers to a nucleic acid that is found in cells or organisms in nature and has not been intentionally modified by humans in a laboratory.

[0315] A “fragment” of a polynucleotide or nucleic acid sequence refers to a continuous sequence of nucleotides, particularly of a length of at least 15 bp, at least 30 bp, at least 40 bp, at least 50 bp, and / or at least 60 bp, according to one embodiment of this invention. Specifically, the polynucleotide fragment comprises at least 25, more particularly at least 50, more particularly at least 75, more particularly at least 100, more particularly at least 150, more particularly at least 200, more particularly at least 300, more particularly at least 400, more particularly at least 500, more particularly at least 600, more particularly at least 700, more particularly at least 800, more particularly at least 900, and more particularly at least 1000 consecutive nucleotides of a polynucleotide sequence according to one embodiment of this invention. Without limitation, the polynucleotide fragments described herein can be used as PCR primers and / or probes, or for antisense gene silencing or RNAi.

[0316] "Recombinant nucleic acid sequences" are nucleic acid sequences created by combining genetic material from more than one source using laboratory methods (such as molecular cloning), thereby creating or modifying nucleic acid sequences that are not naturally occurring and cannot be found in biological organisms in any other way.

[0317] “Recombinant DNA technology” refers to molecular biological methods used to prepare recombinant nucleic acid sequences, as described, for example, in Laboratory Manuals edited by Weigel and Glazebrook, 2002, Cold Spring Harbor Lab Press; and Sambrook et al., 1989, Cold Spring Harbor, NY: Cold Spring Harbor Laboratory Press.

[0318] The term "gene" refers to a DNA sequence containing a region that is operatively linked to a suitable regulatory region (e.g., a promoter) and transcribed into an RNA molecule (e.g., mRNA in a cell). Therefore, a gene can contain several operatively linked sequences, such as a promoter, a 5' leader sequence (containing, for example, a sequence involved in translation initiation), a coding region of cDNA or genomic DNA, introns, exons, and / or a 3' untranslated sequence (containing, for example, a transcription termination site).

[0319] A "chimeric gene" is any gene that is not normally found in species in nature, particularly a gene in which one or more portions of the nucleic acid sequence are unrelated in nature. For example, a promoter that is unrelated in nature to part or all of the transcribed region or to another regulatory region. The term "chimeric gene" should be understood to include expression constructs in which a promoter or transcriptional regulatory sequence is operatively linked to one or more coding sequences or antisense (i.e., the inverse complementary strand of the sense strand) or inverted repeat sequences (sense and antisense, whereby the RNA transcript forms a double-stranded RNA post-transcriptionally). The term "chimeric gene" also includes genes obtained by combining portions of one or more coding sequences to produce new genes.

[0320] "3'URT" or "3' untranslated sequence" (also known as "3' untranslated region" or "3' end") refers to a nucleic acid sequence found downstream of the gene coding sequence that contains, for example, a transcription termination site and (in most, but not all, eukaryotic mRNAs) a polyadenylation signal, such as AAUAAA or its variants. After transcription termination, the mRNA transcript can be cleaved downstream of the polyadenylation signal and a poly(A) tail can be added, which is involved in the transport of mRNA to the translation site, such as the cytoplasm.

[0321] The term "primer" refers to a short nucleic acid sequence that is hybridized to a template nucleic acid sequence and used for the polymerization of nucleic acid sequences complementary to that template.

[0322] The present invention also relates to nucleic acid sequences encoding polypeptides as defined herein.

[0323] In particular, the present invention also relates to nucleic acid sequences (single-stranded and double-stranded DNA and RNA sequences, such as cDNA, genomic DNA and mRNA) encoding one of the aforementioned polypeptides and their functional equivalents, which can be obtained, for example, by using artificial nucleotide analogs.

[0324] This invention also relates to nucleic acids that have a degree of "identity" with the sequences specifically disclosed herein. "Identity" between two nucleic acids refers to the identity of nucleotides along the entire length of the nucleic acid in each case.

[0325] The “identity” between two nucleotide sequences (and similarly, peptide or amino acid sequences) is a function of the number of nucleotide residues (or amino acid residues) when the two sequences are aligned, or the number of identical residues in both sequences. Identical residues are defined as the same residues at a given position in the alignment of the two sequences. The percentage of sequence identity used herein is calculated from the best alignment by dividing the number of identical residues between the two sequences by the total number of residues in the shortest sequence and multiplying by 100. The best alignment is the alignment with the highest probability of identity percentage. Vacancies can be introduced into one or more positions in the alignment of one or both sequences to obtain the best alignment. These vacancies are then considered as dissimilar residues used to calculate the percentage of sequence identity. Alignments used to determine the percentage of identity of amino acid or nucleic acid sequences can be performed in various ways using computer programs, such as those publicly available on the Internet. Specifically, the optimal alignment of protein or nucleic acid sequences and the percentage of sequence identity can be obtained using the BLAST program (Tatiana et al., FEMS Microbiol Lett., 1999, 174:247-250, 1999), which is available from the National Center for Biotechnology Information (NCBI) at http: / / www.ncbi.nlm.nih.gov / BLAST / bl2seq / wblast2.cgi with default parameters. In another example, identity can be calculated using the Clustal method (Higgins DG, Sharp PM. ((1989))) of the Vector NTI Suite 7.1 program from Informax (USA) with the following settings:

[0326] Multiple alignment parameters:

[0327]

[0328]

[0329] Comparison parameters:

[0330]

[0331] Alternatively, identity can be determined according to the method of Chenna et al. (2003), webpage: http: / / www.ebi.ac.uk / Tools / clustalw / index.html# and the following settings:

[0332]

[0333] All nucleic acid sequences (single-stranded and double-stranded DNA and RNA sequences, such as cDNA and mRNA) mentioned in this article can be produced from nucleotide structural units by chemical synthesis in a known manner, for example, by fragment condensation of individual overlapping complementary nucleic acid structural units of a double helix. The chemical synthesis of oligonucleotides can be carried out, for example, by the phosphoramide process (Voet, Voet, 2nd edition, Wiley Press, New York, pages 896-897) in a known manner. The accumulation of synthetic oligonucleotides, and the filling of vacancies in the ligation reaction using the Klenow fragment of DNA polymerase, as well as general cloning techniques, are described in Sambrook et al. (1989), see below.

[0334] The nucleic acid molecules according to the invention may additionally contain untranslated sequences from the 3' and / or 5' ends of the coding genetic region.

[0335] The present invention further relates to nucleic acid molecules that are complementary to the nucleotide sequences or segments thereof specifically described.

[0336] The nucleotide sequences according to the invention enable the generation of probes and primers that can be used to identify and / or clone homologous sequences in other cell types and organisms. Such probes or primers typically contain a nucleotide sequence region that hybridizes to at least about 12, preferably at least about 25, such as about 40, 50 or 75 consecutive nucleotides of the sense strand or corresponding antisense strand of the nucleic acid sequence according to the invention under “stringent” conditions (as defined elsewhere herein).

[0337] "Homologous" sequences include orthologous or paralogous sequences. Methods for identifying orthologous or paralogous sequences include phylogenetic methods, sequence similarity methods, and hybridization methods known in the art and described herein.

[0338] "Paralleloids," or paralogous sequences, arise from gene replication, resulting in two or more genes with similar sequences and functions. Paralogs typically cluster together and form through gene replication within related plant species. Paralogs are identified in groups of similar genes using pairwise BLAST analysis or procedures such as CLUSTAL during phylogenetic analysis of gene families. In paralogs, the shared sequence can be identified as a sequence characteristic of the related gene and possessing a similar function.

[0339] "Orthologs" or "orthologous sequences" are sequences that are similar to each other because they are found in species descended from a common ancestor. For example, plant species known to share a common ancestor contain many enzymes with similar sequences and functions. For instance, by constructing a phylogenetic tree of a gene family for a species using CLUSTAL or BLAST programs, technicians can identify orthologous sequences and predict the functions of orthologs. One method for identifying or confirming similar functions between homologous sequences is by comparing transcript profiles in host cells or organisms (such as plants or microorganisms) that overexpress or lack (in gene knockout / reduction) the relevant polypeptide. Technicians can understand that genes with similar transcript profiles (having a common transcript with greater than 50% regulation, or a common transcript with greater than 70% regulation, or a common transcript with greater than 90% regulation) will have similar functions. Homologs, paralogs, orthologs, and any other variants of the sequences described herein are expected to function in a similar manner by causing host cells, organisms such as plants or microorganisms to produce terpene synthase proteins.

[0340] The term "selectable marker" refers to any gene that, after expression, can be used to select one or more cells containing that selectable marker. Examples of selectable markers are described below. Those skilled in the art will understand that different antibiotic, fungicide, auxotrophic, or herbicide selectable markers may be applicable to different target species.

[0341] The present invention relates to isolated nucleic acid molecules encoding polypeptides or bioactive fragments thereof according to the present invention, and nucleic acid fragments that can be used, for example, as hybridization probes or primers to identify or amplify nucleic acids encoding the present invention.

[0342] "Isolated nucleic acid" or "isolated nucleic acid sequence" refers to a nucleic acid or nucleic acid sequence that is in an environment different from naturally occurring nucleic acids or nucleic acid sequences, and may include those that are substantially free of contaminating endogenous substances. Therefore, isolated nucleic acid molecules are separate from other nucleic acid molecules present in natural sources of nucleic acids, and if produced by recombinant technology, may be substantially free of other cellular material or culture medium, or if chemically synthesized, may be free of chemical precursors or other chemicals.

[0343] Nucleic acid molecules according to the invention can be isolated using standard molecular biology techniques and the sequence information provided according to the invention. For example, cDNA can be isolated from a suitable cDNA library using one of the specifically disclosed complete sequences or fragments thereof as hybridization probes and standard hybridization techniques (e.g., described in Sambrook, (1989)).

[0344] Alternatively, nucleic acid molecules containing one or a fragment of the disclosed sequence can be isolated by polymerase chain reaction using oligonucleotide primers constructed based on that sequence. The nucleic acid amplified in this manner can be cloned into a suitable vector and characterized by DNA sequencing. The oligonucleotides according to the invention can also be prepared using standard synthetic methods, for example, using an automated DNA synthesizer.

[0345] According to the nucleic acid sequences or derivatives thereof of the present invention, homologs or portions of these sequences can be isolated from other bacteria, for example, by conventional hybridization techniques or PCR techniques, through genomic or cDNA libraries. These DNA sequences hybridize with the sequences according to the present invention under standard conditions.

[0346] "Hybridization" refers to the ability of polynucleotides or oligonucleotides to bind to nearly complementary sequences under standard conditions, while non-complementary pairs do not bind nonspecifically. For this purpose, the sequences can be 90–100% complementary. This property of complementary sequences being able to bind specifically to each other is used for primer binding, for example, in Northern or Southern blotting, or in PCR or RT-PCR.

[0347] Short oligonucleotides in conserved regions are advantageously used for hybridization. However, longer fragments or complete sequences of the nucleic acids of the present invention may also be used for hybridization. These “standard conditions” vary depending on the nucleic acid used (oligonucleotide, longer fragment, or complete sequence) or the type of nucleic acid used for hybridization (DNA or RNA). For example, the melting temperature of DNA:DNA hybrids is about 10°C lower than that of DNA:RNA hybrids of the same length.

[0348] For example, depending on the specific nucleic acid, standard conditions refer to a temperature of 42 to 58°C in a buffered aqueous solution with a concentration of 0.1 to 5x SSC (1x SSC = 0.15 M NaCl, 15 mM sodium citrate, pH 7.2), or additionally in the presence of 50% formamide (e.g., 42°C, 5x SSC, 50% formamide). Advantageously, hybridization conditions for DNA:DNA hybrids are 0.1× SSC and a temperature of about 20°C to 45°C, preferably about 30°C to 45°C. For DNA:RNA hybrids, hybridization conditions are advantageously 0.1× SSC and a temperature of about 30°C to 55°C, preferably about 45°C to 55°C. These hybridization temperatures are examples of calculated melting temperatures for nucleic acids of about 100 nucleotides in length and with a G+C content of 50% in the absence of formamide. The experimental conditions for DNA hybridization have been described in relevant genetics textbooks (e.g., Sambrook et al., 1989) and can be calculated using molecular formulas known to those skilled in the art, depending on factors such as nucleic acid length, hybrid type, or G+C content. More information on hybridization can be obtained from textbooks such as Ausubel et al. (eds), (1985), and Brown (ed). (1991).

[0349] “Hybridization” can be carried out under particularly stringent conditions. Such hybridization conditions are described, for example, in Sambrook (1989), or in Current Protocols in Molecular Biology, John Wiley & Sons, NY (1989), 6.3.1–6.3.6.

[0350] As used herein, the term "hybridization" or "hybridization under certain conditions" is intended to describe the conditions under which hybridization and washing occur, under which significantly identical or homologous nucleotide sequences remain bound to each other. These conditions are such that sequences with at least about 70%, for example, at least about 80%, and for example, at least about 85%, 90%, or 95% identity remain bound to each other. Definitions of low-tightness, medium-tightness, and high-tightness hybridization conditions are provided herein.

[0351] Those skilled in the art can select suitable hybridization conditions with minimal experiments, as illustrated, for example, by Ausubel et al. (1995, Current Protocols in Molecular Biology, John Wiley & Sons, sections 2, 4, and 6). Furthermore, stringent conditions are described in Sambrook et al. (1989, Molecular Cloning: A Laboratory Manual, 2nd ed., Cold Spring Harbor Press, chapters 7, 9, and 11).

[0352] As used herein, the low-tightness conditions are defined as follows. Filter membranes containing DNA were pretreated at 40°C for 6 hours in a solution containing 35% formamide, 5xSSC, 50 mM Tris-HCl (pH 7.5), 5 mM EDTA, 0.1% PVP, 0.1% Ficoll, 1% BSA, and 500 μg / ml denatured salmon sperm DNA. Hybridization was performed in the same solution, modified as follows: 0.02% PVP, 0.02% Ficoll, 0.2% BSA, 100 μg / ml salmon sperm DNA, 10% (wt / vol) dextran sulfate, and using 5-20x10 6 32P-labeled probe. The filter membrane was incubated in a hybridization mixture at 40°C for 18–20 h, followed by washing at 55°C for 1.5 h. In a solution containing 2x SSC, 25 mM Tris-HCl (pH 7.4), 5 mM EDTA, and 0.1% SDS, the membrane was replaced with fresh solution and incubated again at 60°C for 1.5 h. The filter membrane was then blotted dry and subjected to autoradiography.

[0353] As used herein, the moderately stringent conditions are defined as follows. Filter membranes containing DNA were pretreated at 50°C for 7 hours in a solution containing 35% formamide, 5x SSC, 50 mM Tris-HCl (pH 7.5), 5 mM EDTA, 0.1% PVP, 0.1% Ficoll, 1% BSA, and 500 μg / ml denatured salmon sperm DNA. Hybridization was performed in the same solution, modified as follows: 0.02% PVP, 0.02% Ficoll, 0.2% BSA, 100 μg / ml salmon sperm DNA, 10% (wt / vol) dextran sulfate, and using 5-20x10 632P-labeled probe. The filter membrane was incubated in a hybridization mixture at 50°C for 30 hours, followed by washing at 55°C for 1.5 hours. In a solution containing 2x SSC, 25 mM Tris-HCl (pH 7.4), 5 mM EDTA, and 0.1% SDS, the membrane was incubated with fresh solution instead of the washing solution at 60°C for another 1.5 hours. The filter membrane was then blotted dry and subjected to autoradiography.

[0354] As used herein, the stringent conditions are as follows: DNA-containing filter membranes were pre-hybridized at 65°C for 8 hours to overnight in a buffer consisting of 6x SSC, 50 mM Tris-HCl (pH 7.5), 1 mM EDTA, 0.02% PVP, 0.02% Ficoll, 0.02% BSA, and 500 μg / ml denatured salmon sperm DNA. The membranes were then pre-hybridized in a buffer containing 100 μg / ml denatured salmon sperm DNA and 5-20 x 10⁻⁶ ppm of DNA. 6 The filter membrane was hybridized in a prehybridization mixture of cpm 32P labeled probes at 65°C for 48 hours. The membrane was then washed at 37°C for 1 hour in a solution containing 2x SSC, 0.01% PVP, 0.01% Ficoll, and 0.01% BSA. It was then washed in 0.1x SSC at 50°C for 45 minutes.

[0355] If the above conditions are not suitable (e.g., for interspecific hybridization), other low, medium and high stringency conditions well known in the art (e.g., for interspecific hybridization) may be used.

[0356] A detection kit for the nucleic acid sequence encoding the polypeptide of the present invention may include primers and / or probes specific to the nucleic acid sequence encoding the polypeptide, and a protocol for using the primers and / or probes to detect the nucleic acid sequence encoding the polypeptide in a sample. Such a detection kit can be used to determine whether a plant, organism, microorganism, or cell has been modified, i.e., whether it has been transformed with the sequence encoding the polypeptide.

[0357] To test the function of a variant DNA sequence according to one embodiment of this article, the target sequence is operatively linked to an optional or screenable marker gene, and the expression of the reporter gene is tested in a transient expression analysis using microorganisms or protoplasts or in stably transformed plants.

[0358] The present invention also relates to derivatives of specifically disclosed or derivable nucleic acid sequences.

[0359] Therefore, the additional nucleic acid sequences according to the invention may be derived from the sequences specifically disclosed herein and may be distinguished by one or more (e.g., 1 to 10) nucleotides, such as 1 to 20, particularly 1 to 15 or 5 to 10, additions, substitutions, insertions or deletions, and may also encode a polypeptide having the desired profile of characteristics.

[0360] The invention also includes nucleic acid sequences containing so-called silent mutations or altered sequences, depending on the codon usage of a particular original or host organism, compared to the specifically stated sequences.

[0361] According to specific embodiments of the invention, variant nucleic acids can be prepared to adapt their nucleotide sequences to a particular expression system. For example, bacterial expression systems are known to express polypeptides more efficiently if the amino acids are encoded by specific codons. Due to the degeneracy of the genetic code, more than one codon can encode the same amino acid sequence, and multiple nucleic acid sequences can encode the same protein or polypeptide; all these DNA sequences are covered in one embodiment herein. Where appropriate, the nucleic acid sequence encoding the polypeptide described herein can be optimized to increase expression in host cells. For example, the nucleic acid of one embodiment herein can be synthesized using host-specific codons to improve expression.

[0362] The present invention also covers naturally occurring variants of the sequences described herein, such as splice variants or allelic variants.

[0363] The allele variant has at least 60% homology across the entire amino acid range at the derived amino acid level, preferably at least 80% homology, and very particularly preferably at least 90% homology (for details regarding homology at the amino acid level, please refer to the information given above for peptides). Advantageously, the homology may be even higher in certain regions of the sequence.

[0364] The present invention also relates to sequences that can be obtained by conserved nucleotide substitution (i.e., as a result, the amino acid in question is replaced by an amino acid having the same charge, size, polarity and / or solubility).

[0365] This invention also relates to molecules derived from specifically disclosed nucleic acids through sequence polymorphism. Such genetic polymorphism can exist in cells from different populations or from cells within a single population due to natural allelic variations. Allelic variants may also include functional equivalents. These natural variations typically produce changes of 1–5% in the nucleotide sequence of a gene. The polymorphism can lead to alterations in the amino acid sequence of the polypeptides disclosed herein. Allelic variants may also include functional equivalents.

[0366] Furthermore, derivatives should also be understood as homologs of the nucleic acid sequences according to the present invention, such as homologs of animals, plants, fungi, or bacteria, shortened sequences, single-stranded DNA or RNA encoding or non-coding DNA sequences. For example, at the DNA level, the homolog has at least 40%, preferably at least 60%, particularly preferably at least 70%, and very particularly preferably at least 80% homology in the entire DNA region given in the sequence specifically disclosed herein.

[0367] Furthermore, derivatives should be understood as, for example, fusions with promoters. Promoters added to the nucleotide sequence can be modified by at least one nucleotide exchange, at least one insertion, inversion, and / or deletion, without impairing the function or effectiveness of the promoter. Moreover, the effectiveness of promoters can be increased by altering their sequence, or by completely exchanging them with more effective promoters or even promoters from different genera of organisms.

[0368] e. Generation of functional peptide mutants

[0369] Furthermore, those skilled in the art are familiar with methods for generating functional mutants, namely, a nucleotide sequence encoding a polypeptide having at least 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity with any amino acid-related SEQ ID NO disclosed herein; and / or encoded by a nucleic acid molecule containing a nucleotide sequence having at least 70% sequence identity with any nucleotide-related SEQ ID NO disclosed herein.

[0370] Depending on the techniques used, those skilled in the art can introduce completely random or more targeted mutations into gene or non-coding nucleic acid regions (e.g., those important for regulating expression) and subsequently generate a genetic library. The molecular biological methods required for this purpose are known to those skilled in the art, for example, as described in Sambrook and Russell, Molecular Cloning, 3rd Edition, Cold Spring Harbor Laboratory Press, 2001.

[0371] Methods for modifying genes and thereby modifying the polypeptides encoded by them are long known to those skilled in the art, for example:

[0372] - Site-specific mutagenesis, in which single or multiple nucleotides of a gene are replaced in a directed manner (Trower MK (Ed.) 1996; In vitro mutagenesis protocols. Humana Press, New Jersey).

[0373] - Saturation mutagenesis, in which the codon of any amino acid can be exchanged or added at any site in the gene (Kegler-Ebo DM, Docktor CM, DiMaio D (1994) Nucleic Acids Res 22:1593; Barettino D, Feigenbutz M, Valcárel R, Stunnenberg HG (1994) Nucleic Acids Res 22:541; Barik S (1995) Mol Biotechnol 3:1),

[0374] - Error-prone polymerase chain reaction, in which the nucleotide sequence is mutated by error-prone DNA polymerase (Eckert KA, Kunkel TA (1990) Nucleic Acids Res 18:3739);

[0375] -SeSaM method (sequence saturation method), in which preferred exchanges are prevented by polymerase. Schenk et al., Biospektrum, Vol. 3, 2006, 277-279.

[0376] - Gene propagation in mutant strains, where, for example, due to defects in DNA repair mechanisms, the mutation rate of nucleotide sequences increases (Greener A, Callahan M, Jerpseth B (1996) An efficient random mutagenesis technique using an E. coli mutator strain. In: Trower MK (Ed.) In vitro mutagenesis protocols. Humana Press, New Jersey), or

[0377] -DNA shuffling, in which a set of closely related genes are formed and digested, and these fragments are used as templates for polymerase chain reactions, in which the full-length mosaic gene is eventually generated through repeated strand separation and recombination (Stemmer WPC (1994) Nature 370:389; Stemmer WPC (1994) Proc Natl Acad Sci USA 91:10747).

[0378] Using so-called directed evolution (particularly described in Reetz MT and Jaeger KE (1999), Topics Curr Chem 200:31; Zhao H, Moore JC, Volkov AA, Arnold FH (1999), Methods for optimizing industrial polypeptides by directed evolution, In: Demain AL, Davies JE (Ed.) Manual of industrial microbiology and biotechnology. American Society for Microbiology), skilled workers can generate functional mutants on a large scale in a directed manner. To this end, in the first step, gene libraries of the respective polypeptides are first generated, for example, using the methods given above. The gene libraries are expressed in a suitable manner, for example, through bacterial or phage display systems.

[0379] The relevant genes in the host organism expressing the functional mutant (whose function largely corresponds to the desired trait) can be submitted to another mutation cycle. The mutation and selection or screening steps can be repeated iteratively until the functional mutant of the present invention possesses a sufficient degree of the desired trait. Using this iterative process, a limited number of mutations, such as 1, 2, 3, 4, or 5 mutations, can be performed in stages, and their effects on the activity under study can be evaluated and selected. The selected mutants can then be subjected to further mutation steps in the same manner. This significantly reduces the number of individual mutants to be studied.

[0380] The results of this invention also provide important information regarding the structure and sequence of the relevant polypeptides, which is essential for the targeted generation of other polypeptides with desired modified properties. In particular, so-called "hot spots" can be defined as sequence segments potentially suitable for modification of properties by introducing targeted mutations.

[0381] Information about the location of amino acid sequences can also be derived, where mutations that may have little effect on activity can occur, and these can be designated as potential “silent mutations”.

[0382] f. Constructs for expressing the polypeptides of the present invention

[0383] In the context of this article, the following definitions apply:

[0384] "Gene expression" encompasses both "heterologous expression" and "overexpression," and involves gene transcription and the translation of mRNA into proteins. Overexpression refers to the production of gene products, measured as mRNA, peptide, and / or enzyme activity levels, in transgenic cells or organisms exceeding the levels found in non-transformed cells or organisms with similar genetic backgrounds.

[0385] As used herein, an "expression vector" refers to a nucleic acid molecule engineered using molecular biology methods and recombinant DNA technology to deliver foreign or exogenous DNA into a host cell. Expression vectors typically include the sequences required for correct transcription of the nucleotide sequence. The coding region usually encodes the target protein, but it can also encode RNA, such as antisense RNA, siRNA, etc.

[0386] As used herein, “expression vector” includes any linear or circular recombinant vector, including but not limited to viral vectors, bacteriophages, and plasmids. Those skilled in the art can select a suitable vector based on the expression system. In one embodiment, the expression vector includes a nucleic acid of the embodiments described herein, operably linked to at least one “regulatory sequence” that controls transcription, translation, initiation, and termination, such as a transcription promoter, operon, or enhancer, or an mRNA ribosome binding site, and optionally includes at least one selection marker. When the regulatory sequence functionally relates to the nucleic acid of the embodiments described herein, the nucleotide sequence is “operably linked.”

[0387] A "regulatory sequence" refers to a nucleic acid sequence that determines the expression level of the nucleic acid sequence in the embodiment described herein and can regulate the transcription rate of a nucleic acid sequence operatively linked to that regulatory sequence. Regulatory sequences include promoters, enhancers, transcription factors, promoter elements, etc.

[0388] A "promoter" is a nucleic acid sequence that controls the expression of a coding sequence by providing a binding site for RNA polymerase and other factors suitable for transcription, including but not limited to transcription factor binding sites, repressor and activator protein binding sites. The term "promoter" also includes the term "promoter regulatory sequence." Promoter regulatory sequences can include upstream and downstream elements that may affect transcription, RNA processing, or the stability of related coding nucleic acid sequences. Promoters include naturally derived and synthetic sequences. The coding nucleic acid sequence is typically located downstream of the promoter relative to the transcription start site.

[0389] In this context, "functional" or "operationally" linked is understood, for example, to refer to the sequential arrangement of one of the nucleic acids having a regulatory sequence. For example, a sequence with promoter activity, and the nucleic acid sequence to be transcribed, along with optional other regulatory elements (e.g., nucleic acid sequences that ensure transcription) and, for example, a terminator, arranged such that each regulatory element can perform its function after transcription of the nucleic acid sequence. This does not necessarily require a direct chemical link. Genetic control sequences, such as enhancer sequences, can even act on the target sequence from more distant locations or even from other DNA molecules. A preferred arrangement is one where the nucleic acid sequence to be transcribed is located downstream (i.e., at the 3' end) of the promoter sequence, thereby covalently linking the two sequences together. The distance between the promoter sequence and the nucleic acid sequence to be recombined can be less than 200 base pairs, or less than 100 base pairs, or less than 50 base pairs.

[0390] Besides promoters and terminators, other examples of regulatory elements include: target sequences, enhancers, polyadenylation signals, selection markers, amplification signals, origins of replication, etc. Suitable regulatory sequences are described, for example, in Goeddel, Gene Expression Technology: Methods in Enzymology 185, Academic Press, San Diego, CA (1990).

[0391] The term "constitutive promoter" refers to an unregulated promoter that allows for the continuous transcription of the nucleic acid sequence to which it is operatively linked.

[0392] As used herein, the term "operably linked" refers to the linking of polynucleotide elements in a functional relationship. A nucleic acid is "operably linked" when it is in a functional relationship with another nucleic acid sequence. For example, if a promoter or transcriptional regulatory sequence can influence the transcription of a coding sequence, then the promoter or transcriptional regulatory sequence is operably linked to the coding sequence. Operable linking means that the linked DNA sequences are typically adjacent. The nucleotide sequence associated with the promoter sequence can be homologous or heterologous relative to the plant to be transformed. The sequence can also be wholly or partially synthetic. Regardless of origin, the nucleic acid sequence associated with the promoter sequence will be expressed or silenced depending on the nature of the promoter linked after binding to the polypeptide of the embodiments described herein. The associated nucleic acid can encode a protein that needs to be expressed or repressed throughout the organism or in a specific tissue, cell, or cell compartment at all times or alternatively at specific times. This nucleotide sequence specifically encodes a protein that confers the desired phenotypic trait to the host cell or organism altered (genetically modified) or transformed by it. More specifically, the associated nucleotide sequence results in the production of one or more desired products as defined herein in cells or organisms, such as, in particular, physostigmine and / or sterol, or mixtures containing physostigmine and / or sterol, or mixtures containing physostigmine and / or sterol and one or more terpenes. In particular, the nucleotide sequence encodes a polypeptide having enzymatic activity as defined herein, such as, in particular, terpene synthase.

[0393] As used herein, “expression system” encompasses any combination of nucleic acid molecules required to express one, or co-express two or more, polypeptides in vivo or in vitro in a given expression host. The respective coding sequences may reside on a single nucleic acid molecule or vector, such as a vector containing multiple cloning sites, or on a polycistronic nucleic acid, or may be distributed across two or more physically distinct vectors.

[0394] As used herein, the terms “amplifying” and “amplification” refer to the use of any suitable amplification method to generate or detect recombinants of naturally expressed nucleic acids, as described in detail below. For example, the present invention provides methods and reagents (e.g., specific degenerate oligonucleotide primer pairs, oligo-dT primers) for amplifying (e.g., by polymerase chain reaction, PCR) naturally expressed (e.g., genomic DNA or mRNA) or recombinant nucleic acids (e.g., cDNA) of the present invention in vivo, in vitro, or in vitro.

[0395] The nucleotide sequences described above can be part of an "expression cassette". The terms "expression cassette" and "expression construct" are used synonymously. A (preferred recombinant) expression construct contains a nucleotide sequence that encodes a polypeptide according to the invention and is under the genetic control of a regulatory nucleic acid sequence.

[0396] In the method applied according to the present invention, the expression cassette may be an "expression vector", particularly a part of a recombinant expression vector.

[0397] According to the present invention, "expression unit" should be understood as a nucleic acid with expressive activity, which contains a promoter as defined herein, and regulates expression upon functional linkage with a nucleic acid or gene to be expressed, i.e., transcription and translation of said nucleic acid or gene. Therefore, it is also referred to in this respect as a "regulatory nucleic acid sequence". In addition to promoters, other regulatory elements, such as enhancers, may also be present.

[0398] According to the present invention, an "expression cassette" or "expression construct" should be understood as an expression unit functionally linked to a nucleic acid or gene to be expressed. Therefore, in contrast to an expression unit, an expression cassette contains not only nucleic acid sequences that regulate transcription and translation, but also nucleic acid sequences that are expressed as proteins due to transcription and translation.

[0399] In the context of this invention, the terms "expression" or "overexpression" describe the generation or increase of intracellular activity of one or more polypeptides encoded by corresponding DNA in a microorganism. For this purpose, for example, a gene may be introduced into the organism, an existing gene may be replaced with another gene, the copy number of a gene may be increased, a strong promoter may be used, or a gene encoding a corresponding polypeptide with high activity may be used. Optionally, these measures may be combined.

[0400] Preferably, such constructs according to the invention include a promoter upstream of the respective coding sequence 5' and a terminator sequence downstream of the respective coding sequence 3', as well as optionally other common regulatory elements, each operatively connected to the coding sequence.

[0401] The nucleic acid constructs according to the invention specifically comprise a sequence encoding a polypeptide, such as derived from the amino acid-related SEQ ID NO or its inverse complementary sequence as described herein, or derivatives and homologs thereof, and is operatively or functionally linked to one or more regulatory signals for advantageous control, for example, increasing gene expression.

[0402] In addition to these regulatory sequences, the natural regulation of these sequences may still exist before the actual structural genes, and optionally may have been genetically modified so that natural regulation has been turned off and gene expression is enhanced. However, nucleic acid constructs can also have simpler constructions, i.e., no additional regulatory signals are inserted before the coding sequence, and the natural promoter with regulatory function has not been removed. Instead, the natural regulatory sequences are mutated so that regulation no longer occurs and gene expression increases.

[0403] Preferred nucleic acid constructs advantageously also include one or more previously mentioned "enhancer" sequences functionally linked to a promoter, which enable enhanced expression of the nucleic acid sequence. Other advantageous sequences, such as other regulatory elements or terminators, may also be inserted at the 3' end of the DNA sequence. One or more copies of the nucleic acid according to the invention may be present in the construct. Optionally, other markers, such as genes complementary to auxotrophic or antibiotic resistance genes, may also be present in the construct for selection.

[0404] Examples of suitable regulatory sequences exist in promoters, such as cos, tac, trp, tet, trp-tet, lpp, lac, lpp-lac, lacIq, T7, T5, T3, gal, trc, ara, rhaP(rhaP BAD SP6, lambda-P R Or lambda-P L In promoters, they are advantageously used in Gram-negative bacteria. Other advantageous regulatory sequences are found, for example, in the Gram-positive promoters amy and SpO2, and in yeast or fungal promoters ADC1, MFalpha, AC, P-60, CYC1, GAPDH, TEF, rp28, and ADH. Artificial promoters can also be used for regulation.

[0405] To facilitate expression in a host organism, nucleic acid constructs are advantageously inserted into vectors, such as plasmids or phages, enabling optimal gene expression in the host. Besides plasmids and phages, vectors should be understood to include all other vectors known to those skilled in the art, such as viruses like SV40, CMV, baculoviruses and adenoviruses, transposons, IS elements, phages, granules, and linear or circular DNA or artificial chromosomes. These vectors are capable of autonomous replication in the host organism or replication via chromosomes. These vectors represent a further development of the invention. Binary or co-integrative vectors are also applicable.

[0406] Suitable plasmids include, for example, those for *E. coli* pLG338, pACYC184, pBR322, pUC18, pUC19, pKC30, pRep4, pHS1, pKK223-3, pDHE19.2, pHS2, pPLc236, pMBL24, pLG200, pUR290, and pIN-III. 113The plasmids listed above are a small selection of possible plasmids, found in *Streptomyces* pIJ101, pIJ364, pIJ702, or pIJ361; in *Bacillus* pUB110, pC194, or pBD214; in *Corynebacterium* pSA77 or pAJ667; in fungi pALS1, pIL2, or pBB116; in yeast 2alphaM, pAG-1, YEp6, YEp13, or pEMBLYe23; or in plants pLGV23, pGHlac+, pBIN19, pAK2004, or pDH51. Other plasmids are well known to technicians and can be found, for example, in the book Cloning Vectors (Eds. Pouwels PH et al. Elsevier, Amsterdam-New York-Oxford, 1985, ISBN 0 444 904018).

[0407] In further development of the vector, vectors containing the nucleic acid constructs of the present invention or the nucleic acids of the present invention can also be advantageously introduced into microorganisms in the form of linear DNA and integrated into the genome of the host organism via heterologous or homologous recombination. This linear DNA can consist of linearized vectors such as plasmids, or solely of the nucleic acid constructs or nucleic acids of the present invention.

[0408] For optimal expression of heterologous genes in an organism, it is advantageous to modify the nucleic acid sequence to match the specific “codon usage” used in the organism. “Codon usage” can be readily determined by computer evaluation of other known genes in the organism under discussion.

[0409] The expression cassette according to the invention is generated by fusing a suitable promoter to a suitable coding nucleotide sequence and a terminator or polyadenylation signal. Conventional recombination and cloning techniques are used for this purpose, as described in, for example, T. Maniatis, EFFritsch and J. Sambrook, Molecular Cloning: A Laboratory Manual, Cold Spring Harbor Laboratory, Cold Spring Harbor, NY (1989); TJ Silhavy, MLBerman and LWEnquist, Experiments with Gene Fusions, Cold Spring Harbor Laboratory, Cold Spring Harbor, NY (1984); and Ausubel, FM et al., Current Protocols in Molecular Biology, Greene Publishing Assoc. and Wiley Interscience (1987).

[0410] To facilitate expression in a suitable host organism, recombinant nucleic acid constructs or gene constructs are advantageously inserted into host-specific vectors, enabling optimal gene expression in the host. Vectors are well-known to those skilled in the art and can be found, for example, in "cloning vectors" (Pouwels PH et al., Ed., Elsevier, Amsterdam-New York-Oxford, 1985).

[0411] Alternative embodiments of the present invention provide a method for “altering (modifying) gene expression” in host cells. For example, in certain contexts (e.g., exposure to certain temperatures or culture conditions), polynucleotides of the present invention can be enhanced, overexpressed, or induced in host cells or host organisms.

[0412] The altered expression of the polynucleotides described herein can also result in ectopic expression, which is a different expression pattern in altered and control or wild-type organisms. The alteration in expression occurs due to contact of the peptide of one embodiment of this invention with an exogenous or endogenous modulator or due to chemical modification of the peptide. The term also refers to the altered expression pattern of the polynucleotides of the embodiments described herein, which is altered to below detectable levels or completely inhibited in activity.

[0413] In one embodiment, this document also provides isolated, recombinant, or synthetic polynucleotides encoding the polypeptide or variant polypeptide provided herein.

[0414] In one embodiment, multiple nucleic acid sequences encoding polypeptides are co-expressed in a single host, particularly under the control of different promoters. In another embodiment, multiple nucleic acid sequences encoding polypeptides may be present on a single transformation vector, or separate vectors may be used and transformants containing two chimeric genes may be selected for simultaneous co-transformation. Similarly, one or more polypeptide-encoding genes may be expressed together with other chimeric genes in a single plant, cell, microorganism, or organism.

[0415] g. Host applicable to the present invention

[0416] Depending on the context, the term "host" can refer to a wild-type host or a genetically modified recombinant host, or both.

[0417] In principle, all prokaryotes or eukaryotes can be considered as hosts or recombinant host organisms for the nucleic acids or nucleic acid constructs according to the present invention.

[0418] Using the vector according to the invention, recombinant hosts can be generated, which can be transformed, for example, with at least one vector according to the invention, and can be used to generate polypeptides according to the invention. Advantageously, the recombinant constructs according to the invention as described above are introduced into and expressed in suitable host systems. Preferably, common cloning and transfection methods known to those skilled in the art, such as co-precipitation, protoplast fusion, electroporation, retroviral transfection, etc., are used to express the nucleic acids in their respective expression systems. Suitable systems are described in Current Protocols in Molecular Biology, F. Ausubelet et al., Ed., Wiley Interscience, New York 1997, or Sambrook et al. Molecular Cloning: A Laboratory Manual, 2nd edition, Cold Spring Harbor Laboratory, Cold Spring Harbor Laboratory Press, Cold Spring Harbor, NY, 1989.

[0419] Advantageously, microorganisms such as bacteria, fungi, or yeasts are used as host organisms. Advantageously, Gram-positive or Gram-negative bacteria are used, preferably those belonging to the families Enterobacteriaceae, Pseudomonadaceae, Rhizobiaceae, Streptomycetaceae, Streptococcaceae, or Nocardiaceae, particularly Escherichia, Pseudomonas, Streptomyces, Lactococcus, Nocardia, Burkholderia, Salmonella, Agrobacterium, Clostridium, or Rhodococcus. The genus and species *Escherichia coli* are particularly preferred. Furthermore, other advantageous bacteria have been found in the alpha-proteobacteria, beta-proteobacteria, or gamma-proteobacteria groups. Advantageously, yeasts such as *Saccharomyces* or the Pichia family are also suitable hosts.

[0420] Alternatively, the entire plant or plant cell can be used as a natural or recombinant host. As non-limiting examples, the following plants or cells derived from them may be mentioned: the genus *Nicotiana*, particularly *Nicotiana abenthamiana* and *Nicotiana tabacum* (tobacco); and the genus *Arabidopsis*, particularly *Arabidopsis thaliana*.

[0421] Depending on the host organism, the organism used in the method according to the invention is grown or cultured in a manner known to those skilled in the art. Culture can be carried out in batches, semi-batch, or continuously. Nutrients can be provided at the start of fermentation or later, semi-continuously, or continuously. This is also described in more detail below.

[0422] h. Recombinant generation of the polypeptide according to the present invention

[0423] The present invention further relates to a method for recombinantly generating polypeptides or functional biologically active fragments thereof according to the invention, wherein microorganisms that generate polypeptides are cultured, and polypeptide expression is optionally induced by applying at least one inducer for gene expression, and the polypeptides are isolated from the culture. If desired, the polypeptides can also be produced on an industrial scale in this manner.

[0424] The microorganisms produced according to the present invention can be cultured continuously or discontinuously using batch culture, fed-batch culture, or repeated fed-batch culture. An overview of known culture methods can be found in Chmiel's textbook (Bioprozesstechnik 1.Einführungin die Bioverfahrenstechnik [Bioprocess technology 1.Introduction to bioprocess technology] (Gustav Fischer Verlag, Stuttgart, 1991)) or Storhas's textbook (Bioreaktoren und periphere Einrichtungen [Bioreactors and peripheralequipment] (Vieweg Verlag, Braunschweig / Wiesbaden, 1994)).

[0425] The culture medium used must be appropriately suited to the requirements of each strain. Descriptions of culture media for various microorganisms are provided in the "Manual of Methods for General Bacteriology" of the American Society for Bacteriology (Washington DC, USA, 1981).

[0426] These culture media, which can be used according to the present invention, typically contain one or more carbon sources, nitrogen sources, inorganic salts, vitamins and / or trace elements.

[0427] Preferred carbon sources are sugars, such as monosaccharides, disaccharides, or polysaccharides. Good carbon sources include, for example, glucose, fructose, mannose, galactose, ribose, sorbitol, ribulose, lactose, maltose, sucrose, raffinose, starch, or cellulose. Sugars can also be added to the culture medium via complex compounds (e.g., molasses) or other byproducts of saccharification. Adding mixtures of different carbon sources is also advantageous. Other possible carbon sources are oils and fats, such as soybean oil, sunflower oil, peanut oil, and coconut oil; fatty acids such as palmitic acid, stearic acid, or linoleic acid; alcohols such as glycerol, methanol, or ethanol; and organic acids such as acetic acid or lactic acid.

[0428] Nitrogen sources are typically organic or inorganic nitrogen compounds or materials containing these compounds. Examples of nitrogen sources include ammonia or ammonium salts, such as ammonium sulfate, ammonium chloride, ammonium phosphate, ammonium carbonate, or ammonium nitrate, nitrates, urea, amino acids, or complex nitrogen sources, such as corn steep liquor, soy flour, soy protein, yeast extract, meat extract, etc. Nitrogen sources can be used alone or in combination.

[0429] Inorganic salt compounds that can be present in culture media include chlorides, phosphorus, or sulfates of calcium, magnesium, sodium, cobalt, molybdenum, potassium, manganese, zinc, copper, and iron.

[0430] Inorganic sulfur-containing compounds, such as sulfates, sulfites, dithionites, tetrasulfites, thiosulfates, sulfides, and organic sulfur compounds, such as thiols and thiols, can be used as sulfur sources.

[0431] Phosphoric acid, potassium dihydrogen phosphate, or dipotassium hydrogen phosphate, or their corresponding sodium-containing salts, can be used as phosphorus sources.

[0432] Chelating agents can be added to the culture medium to retain metal ions in solution. Particularly suitable chelating agents include dihydroxyphenols, such as catechol or protocatechuic acid esters, or organic acids, such as citric acid.

[0433] The fermentation medium used according to the present invention typically also contains other growth factors, such as vitamins or growth promoters, including, for example, biotin, riboflavin, thiamine, folic acid, niacin, pantothenic acid, and pyridoxine. The growth factors and salts are often derived from components of complex culture media, such as yeast extract, molasses, corn steep liquor, etc. In addition, suitable precursors may be added to the medium. The exact composition of the compounds in the medium depends largely on the specific experiment and is determined individually for each case. Information on medium optimization can be found in the textbook "Applied Microbiol. Physiology, A Practical Approach" (Ed. PMRhodes, PFStanbury, IRL Press (1997), pp. 53-73, ISBN 0 19 963577 3). Growth media are also available from commercial suppliers such as Standard 1 (Merck) or BHI (Brain Heart Infusion, DIFCO).

[0434] All components of the culture medium are sterilized by heating (at 1.5 bar and 121°C for 20 minutes) or by aseptic filtration. These components can be sterilized together or individually as needed. All components of the culture medium can be given at the start of the culture or added continuously or in batches.

[0435] The culture temperature is typically between 15°C and 45°C, preferably between 25°C and 40°C, and can be varied or kept constant during the experiment. The pH of the medium should be in the range of 5 to 8.5, preferably around 7.0. The pH during growth can be controlled by adding alkaline compounds (e.g., sodium hydroxide, potassium hydroxide, ammonia, or ammonia solution) or acidic compounds (e.g., phosphoric acid or sulfuric acid). Antifoaming agents such as fatty acid polyethylene glycol esters can be used to control foaming. To maintain plasmid stability, suitable selective substances such as antibiotics can be added to the culture medium. To maintain aerobic conditions, oxygen or an oxygen-containing gas mixture (e.g., ambient air) is supplied to the culture. The culture temperature is typically in the range of 20°C to 45°C. Continue culturing until the maximum amount of the desired product is formed. This usually takes between 10 and 160 hours.

[0436] The fermentation broth is then further processed. Depending on the needs, the biomass can be completely or partially removed from the fermentation broth, or it can be left entirely in it, by separation techniques such as centrifugation, filtration, decantation, or a combination of these methods.

[0437] If the polypeptide is not secreted in the culture medium, the cells can also be lysed, and the product can be obtained from the lysate using known methods for protein separation. Cells can be optionally destroyed by high-frequency ultrasound, high pressure (e.g., in a French press), by osmosis, by the action of detergents, lysing enzymes, or organic solvents, by a homogenizer, or by a combination of these methods.

[0438] Peptides can be purified using known chromatographic techniques, such as molecular sieve chromatography (gel filtration), Q-agarose chromatography, ion exchange chromatography, and hydrophobic chromatography, as well as other conventional techniques such as ultrafiltration, crystallization, salting out, dialysis, and natural gel electrophoresis. Suitable methods are described, for example, in Cooper, TG, Biochemimsche Arbeitsmethoden [Biochemical Processes], Verlag Walter de Gruyter, Berlin, New York, or Scopes, R., Protein Purification, Springer Verlag, New York, Heidelberg, Berlin.

[0439] For the isolation of recombinant proteins, the use of a carrier system or oligonucleotide may be advantageous, which extends cDNA by a defined nucleotide sequence and thus encodes an altered polypeptide or fusion protein, for example, for easier purification. Suitable modifications of this type are, for example, so-called “tags” that act as anchors, such as modifications known as hexahistine anchors or epitopes that can be recognized as antibody antigens (e.g., described in Harlow, E. and Lane, D., 1988, Antibodies: A Laboratory Manual. Cold Spring Harbor (NY) Press). These anchors can be used to attach proteins to solid supports, such as polymer matrices, which can be used, for example, as packing material in chromatographic columns, or on microtiter plates or other supports.

[0440] These anchors can also be used to identify proteins. To identify proteins, conventional markers, such as fluorescent dyes, enzyme markers (which react with a substrate to form a detectable reaction product), or radiolabels, can be used alone or in combination with anchors to derivatize proteins.

[0441] i. Immobilization of peptides

[0442] The enzymes or polypeptides according to the invention can be used in free form or immobilized in the methods described herein. Immobilized enzymes are enzymes immobilized on an inert support. Suitable support materials and enzymes immobilized thereon are known from EP-A-1149849, EP-A-1069183, and DE-OS 100193773 and the references cited therein. In this regard, reference is made to the full disclosure of these documents. Suitable support materials include, for example, clay, clay minerals such as kaolinite, diatomaceous earth, perlite, silica, alumina, sodium carbonate, calcium carbonate, cellulose powder, anion exchanger materials, synthetic polymers such as polystyrene, acrylic resins, phenolic resins, polyurethanes, and polyolefins such as polyethylene and polypropylene. For the preparation of loaded enzymes, the support material is generally used in the form of finely divided particles, preferably porous. The particle size of the support material is generally no greater than 5 mm, particularly no greater than 2 mm (particle size distribution profile). Similarly, when using dehydrogenases as whole-cell catalysts, either free or immobilized form can be selected. Carrier materials include, for example, calcium alginate and carrageenan. Enzymes and cells can also be directly cross-linked with glutaraldehyde (cross-linked with CLEAs). Corresponding and other immobilization techniques are described, for example, in J. Lalonde and A. Margolin, "Immobilization of Enzymes" in K. Drauz and H. Waldmann, Enzyme Catalysis in Organic Synthesis 2002, Vol. III, 991-1032, Wiley-VCH, Weinheim. Rehm et al. (Ed.) Biotechnology, 2nd Edn, Vol 3, Chapter 17, VCH, Weinheim provides further information on biotransformation and bioreactors for carrying out the methods according to the invention.

[0443] j. Reaction conditions of the biocatalytic production method of the present invention

[0444] "Enzymatic catalysis" or "biocatalysis" refers to methods carried out under the catalysis of enzymes (including enzyme mutants) as defined herein. Therefore, this method can be carried out in the presence of the enzyme in isolated (purified, enriched) or crude form, or in the presence of cellular systems, particularly natural or recombinant microbial cells containing the active form of the enzyme and capable of catalyzing the transformation reactions disclosed herein. Thus, the reactions of this invention can be carried out under in vivo or in vitro conditions.

[0445] At least one polypeptide / enzyme present during a single step of the method of the present invention or a multi-step method as defined above may be naturally present in living cells, or in harvested cells (i.e., under in vivo conditions (biotransformation)), dead cells, permeabilized cells, crude cell extracts, purified extracts, or in a substantially pure or completely pure form (i.e., under in vitro conditions (biochemical synthesis)), or recombinantly produced as one or more enzymes. The at least one enzyme may be present in solution or as an enzyme immobilized on a carrier. One or more enzymes may be present simultaneously in soluble and / or immobilized forms.

[0446] The method according to the invention can be carried out in common reactors known to those skilled in the art and can be carried out on various scales, from laboratory scale (a few milliliters to tens of liters of reaction volume) to industrial scale (a few liters to thousands of cubic meters of reaction volume). A chemical reactor can be used if the polypeptide is used in the form of encapsulation through non-living, optionally permeabilized cells, as a more or less purified cell extract, or in a purified form. Chemical reactors typically allow control of the amount of at least one enzyme, the amount of at least one substrate, pH, temperature, and the circulation of the reaction medium. When at least one polypeptide / enzyme is present in living cells, the process will be fermentation. In this case, biocatalytic production will be carried out in a bioreactor (fermenter) where parameters necessary for suitable survival conditions for living cells (e.g., nutrient-rich culture medium, temperature, aeration, aerobic or anaerobic or other gases, antibiotics, etc.) can be controlled. Those skilled in the art are familiar with chemical or bioreactors, for example, procedures for scaling up chemical or biotechnological methods from laboratory to industrial scale or optimizing process parameters, which are also extensively described in the literature (for biotechnological methods, see, for example, Crueger und Crueger, Biotechnologie – Lehrbuch der angewandten Mikrobiologie, 2. Ed., R. Oldenbourg Verlag, Munich, Wien, 1984).

[0447] Cells containing at least one enzyme can be permeabilized by physical or mechanical means such as ultrasound or radiofrequency pulses, a high-pressure cell lysis machine (French press), or by chemical means such as a hypotonic medium present in the culture medium, a lysing enzyme, and a detergent, or a combination of these methods. Examples of detergents are digitoxin, n-dodecyl maltodextrin, octyl glycoside, etc. X-100, 20, Deoxycholate, CHAPS (3-[(3-chloroamidopropyl)dimethylammonium]-1-propanesulfonate), P40 (ethylphenol poly(ethylene glycol ether)), etc.

[0448] Instead of living cells, non-living cell biomass containing the desired biocatalyst can also be used in the biotransformation reaction of this invention.

[0449] If at least one enzyme is immobilized, it is ligated to an inert carrier as described above.

[0450] The conversion reaction can be carried out in batches, semi-batches, or continuously. Reactants (and optional nutrients) can be provided at the start of the reaction, or they can be provided subsequently in a semi-continuous or continuous manner.

[0451] Depending on the specific reaction type, the reactions of this invention can be carried out in aqueous, aqueous-organic, or non-aqueous reaction media.

[0452] Aqueous or aqueous-organic media may contain suitable buffer solutions to adjust the pH to 5 to 11, such as 6 to 10.

[0453] In aqueous-organic media, organic solvents that are miscible, partially miscible, or immiscible with water can be used. Non-limiting examples of suitable organic solvents are listed below. Further examples are mono- or poly-aryl, aromatic or aliphatic alcohols, particularly poly-aliphatic alcohols such as glycerol.

[0454] Non-aqueous media may contain substantially no water, that is, will contain less than about 1% by weight or 0.5% by weight of water.

[0455] Biocatalytic methods can also be carried out in organic non-aqueous media. Suitable organic solvents include, for example, aliphatic hydrocarbons having 5 to 8 carbon atoms, such as pentane, cyclopentane, hexane, cyclohexane, heptane, octane, or cyclooctane; aromatic hydrocarbons, such as benzene, toluene, xylene, chlorobenzene, or dichlorobenzene; aliphatic acyclic and ethers, such as diethyl ether, methyl tert-butyl ether, ethyl tert-butyl ether, dipropyl ether, diisopropyl ether, dibutyl ether; or mixtures thereof.

[0456] The concentration of reactants / substrate can be adapted to optimal reaction conditions, which can depend on the specific enzyme being applied. For example, the initial substrate concentration can be 0.1 to 0.5 M, or, for example, 10 to 100 mM.

[0457] The reaction temperature can be adapted to optimal reaction conditions, which can depend on the specific enzyme used. For example, the reaction can be carried out at temperatures ranging from 0 to 70°C, such as 20 to 50 or 25 to 40°C. Examples of reaction temperatures are approximately 30°C, approximately 35°C, approximately 37°C, approximately 40°C, approximately 45°C, approximately 50°C, approximately 55°C, and approximately 60°C.

[0458] The process can continue until equilibrium is reached between the substrate and the subsequent product, but it can be stopped earlier. Typical process times range from 1 minute to 25 hours, particularly from 10 minutes to 6 hours, for example from 1 hour to 4 hours, and particularly from 1.5 hours to 3.5 hours. These parameters are non-limiting examples of suitable process conditions.

[0459] If the host is a genetically modified plant, it can provide optimal growth conditions, such as optimal light, water, and nutrients.

[0460] l. Product separation

[0461] The method of the present invention may further include the step of recovering a final product or intermediate product, which may optionally be a substantially pure form of a stereoisomer or enantiomer. The term "recovery" includes the extraction, harvesting, separation, or purification of a compound from a culture medium or reaction medium. The recovery of a compound can be carried out according to any conventional separation or purification method known in the art, including but not limited to treatment with conventional resins (e.g., anion or cation exchange resins, nonion adsorption resins, etc.), treatment with conventional adsorbents (e.g., activated carbon, silica, silica gel, cellulose, alumina, etc.), pH alteration, solvent extraction (e.g., using conventional solvents such as alcohols, ethyl acetate, hexane, etc.), distillation, dialysis, filtration, concentration, crystallization, recrystallization, pH adjustment, lyophilization, etc.

[0462] The identity and purity of the separated products can be determined by known techniques, such as high performance liquid chromatography (HPLC), gas chromatography (GC), spectroscopy (e.g., IR, UV, NMR), staining methods, TLC, NIRS, enzyme or microbial assays (see, for example: Patek et al. (1994) Appl. Environ. Microbiol. 60: 133-140; Malakhova et al. (1996) Biotekhnologiya 1127-32; und Schmidt et al. (1998) Bioprocess Engineer. 19: 67-70; Ullmann's Encyclopedia of Industrial Chemistry (1996) Bd. A27, VCH: Weinheim, pp. 89-90, 521-540, 540-547, 559-566, 575-581). S.581-587; Michal, G (1999) Biochemical Pathways: An Atlas of Biochemistry and Molecular Biology, John Wiley and Sons; Fallon, A. et al. (1987) Applications of HPLC in Biochemistry in: Laboratory Techniques in Biochemistry and Molecular Biology, Bd. 17.).

[0463] Cyclic terpenoids produced by any of the methods described herein can be converted into derivatives, such as, but not limited to, hydrocarbons, esters, amides, glycosides, ethers, epoxides, aldehydes, ketones, alcohols, diols, acetals, or ketals. Terpenoid derivatives can be obtained by chemical methods, such as, but not limited to, oxidation, reduction, alkylation, acylation, and / or rearrangement. Alternatively, terpenoid derivatives can be obtained by biochemical methods by contacting the terpenoid with an enzyme, such as, but not limited to, oxidoreductases, monooxygenases, dioxygenases, and transferases. Biochemical transformation can be performed in vitro using isolated enzymes, enzymes derived from lysed cells, or in vivo using whole cells.

[0464] m. Fermentation of bis(phosphonic acid) and / or bis(phosphonic acid) and / or bis(phosphonic acid) and / or bis(phosphonic acid)

[0465] The present invention also relates to a method for fermentation to produce bis(phosphonic acid) and / or bis(phosphonic acid), and / or bis(phosphonic acid) and / or bis(phosphonic acid).

[0466] The term “fermentation production” or “fermentation” refers to the ability of microorganisms (assisted by enzyme activity contained in or produced by said microorganisms) to produce compounds in cell cultures using at least one carbon source added to incubation.

[0467] The term “fermentation broth” should be understood as referring to a liquid, particularly an aqueous solution or aqueous / organic solution, that is based on a fermentation process and has been subjected to or has been subjected to post-processing, such as that described herein.

[0468] The fermentation used according to the invention can be carried out, for example, in stirred fermenters, bubble columns, and loop reactors. For a comprehensive overview of possible method types, including stirrer types and geometries, see "Chmiel: Bioprozesstechnik: Einfuhrung in die Bioverfahrenstechnik, Band 1". Typical variations available in the method of the invention are those known to those skilled in the art or explained, for example, in "Chmiel, Hammes and Bailey: Biochemical Engineering", such as batch, fed-batch, repeatedly fed-batch, or continuous fermentation, with or without biomass recovery. Depending on the production strain, air, oxygen, carbon dioxide, hydrogen, nitrogen, or suitable gas mixtures can be injected to achieve good yields (YP / S).

[0469] The “yield” and / or “conversion” of the reaction according to the invention are determined within a specified time period, such as 4, 6, 8, 10, 12, 16, 20, 24, 36 or 48 hours (during which the reaction takes place). In particular, the reaction is carried out under precisely defined conditions, such as the “standard conditions” as defined herein.

[0470] Different yield parameters (“yield” or YP / S; “specific productivity yield”; or space-time yield (STY)) are well known in the art and are measured as described in the literature.

[0471] "Yield" and "YP / S" (both expressed as the quality of products produced / the quality of materials consumed) are used as synonyms in this document.

[0472] Specific productivity yield describes the amount of product produced per hour per liter of fermentation broth per gram of biomass. Wet cell weight (WCW) describes the number of biologically active microorganisms in the biochemical reaction. This value is given as grams of product per g WCW per hour (i.e., g / gWCW). -1 h -1Alternatively, the amount of biomass can also be expressed as the amount of dry cell weight, denoted as DCW. Furthermore, by measuring 600 nm (OD) 600 By estimating the corresponding wet or dry cell weights using the optical density at a given location and relevant factors determined experimentally, biomass concentration can be determined more easily.

[0473] The culture medium used according to the present invention must be suitably adapted to meet the requirements of the specific strain. Descriptions of culture media for various microorganisms are given in the "Manual of Methods for General Bacteriology" of the American Society for Bacteriology (Washington DC, USA, 1981).

[0474] These culture media, which can be used according to the present invention, typically contain one or more carbon sources, nitrogen sources, inorganic salts, vitamins and / or trace elements.

[0475] Preferred carbon sources are sugars, such as monosaccharides, disaccharides, or polysaccharides. Very good carbon sources include, for example, glucose, fructose, mannose, galactose, ribose, sorbitol, ribulose, lactose, maltose, sucrose, raffinose, starch, or cellulose. Sugars can also be added to the culture medium via complex compounds (e.g., molasses) or other byproducts of saccharification. Adding mixtures of various carbon sources is also advantageous. Other possible sources of carbon are oils and fats, such as soybean oil, sunflower oil, peanut oil, and coconut oil; fatty acids such as palmitic acid, stearic acid, or linoleic acid; alcohols such as glycerol, methanol, or ethanol; and organic acids such as acetic acid or lactic acid.

[0476] Nitrogen sources are typically organic or inorganic nitrogen compounds or materials containing these compounds. Examples of nitrogen sources include ammonia or ammonium salts, such as ammonium sulfate, ammonium chloride, ammonium phosphate, ammonium carbonate, or ammonium nitrate, nitrates, urea, amino acids, or complex nitrogen sources, such as corn steep liquor, soy flour, soy protein, yeast extract, meat extract, etc. Nitrogen sources can be used alone or in combination.

[0477] Inorganic salt compounds that may be present in the medium include chlorides, phosphates, or sulfates of calcium, magnesium, sodium, cobalt, molybdenum, potassium, manganese, zinc, copper, and iron.

[0478] Inorganic sulfur-containing compounds, such as sulfates, sulfites, dithionites, tetrasulfites, thiosulfates, and sulfides, as well as organic sulfur compounds, such as thiols and thiols, can all be used as sulfur sources.

[0479] Phosphoric acid, potassium dihydrogen phosphate, or dipotassium hydrogen phosphate, or their corresponding sodium-containing salts, can be used as phosphorus sources.

[0480] Chelating agents can be added to the culture medium to retain metal ions in solution. Particularly suitable chelating agents include dihydroxyphenols, such as catechol or protocatechuic acid esters, or organic acids, such as citric acid.

[0481] The fermentation medium used according to the present invention may also contain other growth factors, such as vitamins or growth promoters, including, for example, biotin, riboflavin, thiamine, folic acid, niacin, pantothenic acid, and pyridoxine. Growth factors and salts are typically derived from complex components of the medium, such as yeast extract, molasses, corn steep liquor, etc. Additionally, suitable precursors may be added to the medium. The precise composition of compounds in the medium depends heavily on the specific experiment and must be determined individually for each specific case. Information on medium optimization can be found in the textbook "Applied Microbiol. Physiology, A Practical Approach" (1997). Growth media are also available from commercial suppliers, such as Standard 1 (Merck) or BHI (Brain Heart Infusion, DIFCO).

[0482] All components of the culture medium are sterilized by heating (at 1.5 bar and 121°C for 20 minutes) or by aseptic filtration. These components can be sterilized together or individually as needed. All components of the culture medium can be given at the start of growth, or can be added continuously or in batches.

[0483] The culture temperature is typically between 15°C and 45°C, preferably between 25°C and 40°C, and can be kept constant or varied during the experiment. The pH of the medium should be in the range of 5 to 8.5, preferably around 7.0. The pH during growth can be controlled by adding alkaline compounds (e.g., sodium hydroxide, potassium hydroxide, ammonia, or ammonia solution) or acidic compounds (e.g., phosphoric acid or sulfuric acid). Antifoaming agents such as fatty acid polyethylene glycol esters can be used to control foaming. To maintain plasmid stability, suitable substances with selective action, such as antibiotics, can be added to the culture medium. To maintain aerobic conditions, oxygen or a mixture of oxygen-containing gases (e.g., ambient air) is supplied to the culture. The culture temperature is typically between 20°C and 45°C. Continue culturing until the maximum amount of the desired product is formed. This usually takes between 1 and 160 hours.

[0484] The method of the present invention may further include the step of recovering bisphosphonate and / or bisphosphonate, and / or bisphosphonate and / or bisphosphonate alcohol.

[0485] The term "recovery" includes the extraction, harvesting, separation, or purification of compounds from a culture medium. The recovery of compounds can be performed using any conventional separation or purification method known in the art, including but not limited to treatment with conventional resins (e.g., anion or cation exchange resins, nonion adsorption resins, etc.), treatment with conventional adsorbents (e.g., activated carbon, silica, silica gel, cellulose, alumina, etc.), pH alteration, solvent extraction (e.g., using conventional solvents such as alcohols, ethyl acetate, hexane, etc.), distillation, dialysis, filtration, concentration, crystallization, recrystallization, pH adjustment, lyophilization, etc.

[0486] Before the intended separation, the biomass in the fermentation broth can be removed. Methods for removing biomass are known to those skilled in the art, such as filtration, sedimentation, and flotation. Therefore, biomass can be removed, for example, by centrifuges, separators, decanters, filters, or in flotation equipment. To maximize the recovery of valuable products, washing the biomass, for example by percolation, is generally recommended. The choice of method depends on the biomass content and properties in the fermentation broth, as well as the interaction between the biomass and the valuable products.

[0487] In one implementation, the fermentation broth can be sterilized or pasteurized. In another implementation, the fermentation broth is concentrated. This concentration can be carried out in batches or continuously, depending on the needs. Pressure and temperature ranges should be selected to ensure that product damage is not caused and to minimize equipment and energy consumption. Skillful selection of pressure and temperature levels for multi-stage evaporation can be particularly energy-efficient.

[0488] The following examples are illustrative only and are not intended to limit the scope of the implementation schemes described herein.

[0489] After considering the disclosure provided herein, a variety of possible variations that will immediately become apparent to those skilled in the art also fall within the scope of this invention.

[0490] Experimental Section

[0491] The invention will now be described in more detail through the following embodiments.

[0492] Material:

[0493] Unless otherwise stated, all chemical and biochemical materials, as well as microorganisms or cells used in this article, are commercially available products.

[0494] Unless otherwise stated, recombinant proteins are cloned and expressed using standard methods, such as those described, for example, in Sambrook, J., Fritsch, E.F. and Maniatis, T., Molecular cloning: A Laboratory Manual, 2008. nd Edition, Cold Spring Harbor Laboratory, Cold Spring Harbor Laboratory Press, Cold Spring Harbor, NY, 1989.

[0495] method:

[0496] GC / MS

[0497] GC / MS analysis was performed on a GC / MS 6890N / 5975 (Agilent). Separation was achieved using a DB-1MS column (30m × 0.25mm inner diameter × 0.25μm film thickness, J&W Scientific, Agilent Technologies, Foster City, CA). Helium was used as the carrier gas at a flow rate of 0.7 mL / min. Split injection (1:5) mode was used, and the injector temperature was set to 250 °C. The oven temperature program was set to initially increase from 50 °C (hold for 5 minutes) to 300 °C at a rate of 5 °C / min, then increase to 340 °C at a rate of 50 °C / min and hold for 3 minutes. Data were acquired using a scan mode of m / z 29–450. Alkanes (C6–C4) were injected. 30 The retention index is calculated. Compounds are identified by matching mass spectrometry and retention indices with an internal library.

[0498] Example 1

[0499] Discovery of HAD-like TPS

[0500] Fresh leaves of *Bazzania* (Nagobase ID: PA-2018-0168 to PA-2018-0173, containing bismuth sub-sesquiterpenoids, including bismuth sub-sesquiterpenoid and bismuth sub-sesquiterpenoid) were obtained from Hunan, China. Total RNA was extracted from the fresh leaves using a QIAGEN RNeasyPlant Mini Kit (50) 74904 and used for RNA sequencing. RNA sequencing data were then assembled using the CLC Genomics Workbench (7.6.1). Sequences of fern HAD-like TPS (DfHAD) from prior art (PCT / EP2019 / 063824; SEQ ID NO: 73 and 74) were used to search for potential bismuth sub-sesquiterpenoid synthases in the protein sequence data of *Bazzania*. Novel sequences were discovered from two assemblages of the whiting transcriptome using BLAST, resulting in HAD-like TPSs named BazzHAD1 (SEQ ID NO:3), BazzHAD2 (SEQ ID NO:6), and BazzHAD3 (SEQ ID NO:9); their corresponding nucleotide sequences are SEQ ID NO:1, 4, and 7, respectively.

[0501] The raw RNAseq data of Bazzania trilobata (SRA ID:ERR364415) were downloaded from NCBI and assembled by the CLC Genomics Workbench to provide the transcriptome, from which BtHAD (SEQ ID NO:12) was discovered by BLAST; its corresponding nucleotide sequence is SEQ ID NO:10.

[0502] Homologous sequences of DfHAD in the NCBI database were thoroughly mined using BLAST with a threshold score (E value < 10, hit list size < 500) until the plant gene was observed. This yielded two similar sequences from *Selaginella moellendorffii*, which were deposited as hypothetical proteins with accession numbers EFJ10816.1 and EFJ26126.1 (GenBank ID), and were unrelated to the HAD-like TPS. In this study, the proteins were named SmHAD1 (SEQ ID NO: 19) and SmHAD2 (SEQ ID NO: 26), respectively; their corresponding nucleotide sequences are SEQ ID NO: 17 and 24, respectively.

[0503] To investigate their homology, SmHAD1, SmHAD2, BtHAD, BazzHAD1, BazzHAD2 and BazzHAD3, as well as DfHAD (PCT / EP2019 / 063824), and several fungal HAD-like bismuth sesquiterpene synthases (including AstC (PCT / EP2019 / 063824)) were compared using Clustalw. We found that the protein sequences of SmHAD1 and SmHAD2 have very low identity (approximately 14–20%) with all other known HAD-like TPS, but they show high homology (>96%) among themselves. BtHAD and BazzHAD1 show slightly higher homology (approximately 22–30%) with fern and fungal HAD-like TPS, with 96% homology among themselves. BazzHAD2 and BazzHAD3 show moderate homology (approximately 22–43%) with fern and fungal HAD-like TPS, with 67% homology among themselves (Table 1).

[0504]

[0505] Example 2

[0506] Functional expression and characterization of plant HAD-like TPS in Escherichia coli expression system (in vivo biochemical analysis)

[0507] The coding sequences of BazzHAD1, BazzHAD2, BazzHAD3, BtHAD, SmHAD1 (EFJ10816.1), and SmHAD2 (EFJ26126.1) were optimized using Genscript according to their frequency of use in the genetic code of *E. coli* (SEQ ID NOs: 2, 5, 8, 11, 18, and 25, respectively). These sequences were then synthesized in vitro and subcloned into the pETDuet-1 plasmid (Novagen) for subsequent expression in *E. coli*, resulting in plasmids pETDuet-BazzHAD1, pETDuet-BazzHAD2, pETDuet-BazzHAD3, pETDuet-BtHAD, pETDuet-SmETHAD1, and pETDuet-SmHAD2, respectively.

[0508] BL21(DE3) *E. coli* cells (Tiangen) were co-transformed with the plasmid pACYC / ScMVA, which contains genes encoding the heterologous mevalonic acid (MVA) pathway, as well as the respective plasmids pETDuet-BazzHAD1, pETDuet-BazzHAD2, pETDuet-BazzHAD3, pETDuet-BtHAD, pETDuet-SmHAD1, and pETDuet-SmHAD2. The pACYC / ScMVA plasmid was constructed for expressing the farnesyl pyrophosphate (FPP) synthase gene and the complete MVA pathway genes. In short, the eight biosynthetic genes of the MVA pathway are divided into two synthetic operons, referred to as the "upper" and "lower" MVA pathways. As part of the "upper" MVA pathway, a synthetic operon was created, consisting of an acetyl-CoA thiolytic enzyme encoded by atoB from *E. coli*, HMG-CoA synthases encoded by ERG13 and ERG19 from *Saccharomyces cerevisiae*, and a truncated form of HMG-CoA reductase, respectively. This operon converts the major metabolite acetyl-CoA to (R)-mevalerate. As part of the "lower" mevalerate pathway, a second synthetic operon was created, encoding mevalerate kinase (ERG12, *Saccharomyces cerevisiae*), phosphate mevalerate kinase (ERG8, *Saccharomyces cerevisiae*), phosphate mevalerate decarboxylase (MVD1, *Saccharomyces cerevisiae*), isopentenyl diphosphate isomerase (idi, *E. coli*), and FPP synthase (IspA, *E. coli*). Finally, a second FPP synthase (ERG20) from *Saccharomyces cerevisiae* was introduced into the upper pathway operon to improve the conversion of isoprene C5 units (IPP and DMAPP) to FPP. Under the control of the bacterial phage T7 promoter (pACYCDuet-1, Invitrogen), each operon is subcloned into one of the multiple cloning sites of a low-copy expression plasmid, providing plasmid pACYC / ScMVA. Therefore, this plasmid contains genes encoding all the enzymes of the biosynthetic pathway from acetyl-CoA to FPP.

[0509] Select co-transformed cells on LB agar plates containing ampicillin (50 μg / mL) and chloramphenicol (34 μg / mL). Inoculate single colonies into 5 mL of liquid LB medium supplemented with the same antibiotics. Incubate the cultures overnight at 37 °C with shaking at 200 rpm. The next day, inoculate 0.2 mL of the overnight culture into 2 mL of TB medium supplemented with the above antibiotics and glycerol (final concentration 3% w / v). Incubate at 37 °C for 5 h with shaking at 200 rpm, then cool the cultures to 25 °C and shake at 200 rpm for another hour. Then add IPTG (final 0.1 mM) to each tube and cover the cultures with 200 μL of decane. Incubate the cultures at 25 °C for another 48 h with shaking at 200 rpm, then extract twice with 1 volume of ethyl acetate. Before analyzing the samples by GC / MS, add 50 μL of 2 mg / mL isophorene as an internal standard to the organic phase. Helium was used as the carrier gas at a constant flow rate of 0.7 mL / min. The sample was injected in split (1:25) mode, with the injector temperature set to 250 °C. The oven temperature program was set to initially increase from 50 °C (hold for 5 minutes) to 300 °C at a rate of 5 °C / min, then increase to 340 °C at a rate of 50 °C / min and hold for 3 minutes. Product identification was based on mass spectrometry and retention index.

[0510] GC / MS analysis showed that recombinant cells expressing BazzHAD1 and BtHAD produced 37.3 mg / L and 26.0 mg / L of sterol, respectively; expression of BazzHAD2 resulted in the production of 35.3 mg / L of sterol and 0.2 mg / L of an unknown sesquiterpene 1 (unknown SQT1); expression of BazzHAD3 resulted in the production of 422.7 mg / L of sterol; expression of SmHAD1 resulted in the production of 0.5 mg / L of sterol and 6.2 mg / L of sterol; and expression of SmHAD2 resulted in the production of 1.4 mg / L of sterol and 10.7 mg / L of sterol (Figures 1–16).

[0511] Example 3

[0512] The importance of class II motifs and QW motifs in plant HAD-like TPS (functional evidence of enzyme mutants and motifs)

[0513] To confirm the importance of the class II terpene synthase motif (SEQ ID NO:46), mutants of this motif were designed and generated for BtHAD, SmHAD1, and SmHAD2 (SEQ ID NO:13, 20, and 27, respectively). Since BazzHAD1 is an ortholog of BtHAD from *Bazzania trilobata* (with 96% identity), it can be represented as BtHAD and is therefore not included in this study. In addition to the class II motif, the QW motif (SEQ ID NO:51) was also mutated in BtHAD, SmHAD1, and SmHAD2 (SEQ ID NO:15, 22, and 29, respectively) to investigate its importance to synthase activity. Since known fungal HAD-like TPS contain class I and class II motifs (e.g., SEQ ID NO:56, 57, and 58 of WO2018 / 220113), EMD37666.1 (SEQ ID NO:31) and XP_007369631.1 (SEQ ID NO:35) were selected and their class II motifs were mutated (SEQ ID NO:33 and 37, respectively) to investigate whether class I motifs were sufficient to produce synthase activity. In mutants of the class II synthase motif in BtHAD (D320A / D322A), SmHAD1 (D257A / D259A), SmHAD2 (D257A / D259A), EMD37666.1 (D276A / S277A / D279A), and XP_007369631.1 (D272A / D273A / D275A), polar Asp is replaced by Ala (SEQ ID NO: 39 to 42). In mutants of the QW motif in BtHAD (D491A), SmHAD1 (E432A / D433A), and SmHAD2 (E432A / D433A), polar Glu and Asp are replaced by Ala (SEQ ID NO: 43 and 44).

[0514] The sequences of BtHAD, SmHAD1, SmHAD2, XP_007369631.1, EMD37666.1, and their corresponding class II synthase motifs and (if present) QW motif mutants were optimized by Genscript according to the genetic codon frequency table of E. coli (BtHAD SEQ ID NO: 11, 13, 15; SmHAD1 SEQ ID NO: 18, 20, 22; SmHAD2 SEQ ID NO: 25, 27, 29; EMD37666.1 SEQ ID NO: 31, 33; and XP_007369631.1 SEQ ID NO: 35, 37). The sequence was then synthesized in vitro using Genscript and subcloned into the pETDuet-1 plasmid (Novagen) for subsequent expression in E. coli, generating plasmids pETDuet-BtHAD, pETDuet-SmHAD1, pETDuet-SmHAD2, pETDuet-EMD37666.1, pETDuet-XP_007369631.1; pETDuet-BtHAD-ClassII mut, pETDuet-SmHAD1-ClassII mut, pETDuet-SmHAD2-ClassII mut, pETDuet-EMD37666.1-ClassII mut, pETDuet-XP_007369631.1-ClassII mut; and pETDuet-BtHAD-QW mut, pETDuet-SmHAD1-QW mut, and pETDuet-SmHAD2-QW mut.

[0515] BL21(DE3) *E. coli* cells (Tiangen) were co-transformed with the plasmid pACYC / ScMVA, which contains the gene encoding the heterologous mevalonate pathway, and the respective plasmids pETDuet-BtHAD, pETDuet-SmHAD1, pETDuet-SmHAD2, pETDuet-EMD37666.1, pETDuet-XP_007369631.1, pETDuet-BtHAD-ClassII mut, pETDuet-SmHAD1-ClassII mut, pETDuet-SmHAD2-ClassII mut, pETDuet-EMD37666.1-ClassII mut, pETDuet-XP_007369631.1-ClassII mut, pETDuet-BtHAD-QWmut, and pETDuet-SmHAD1-QW. mut and pETDuet-SmHAD2-QW mut. The construction of pACYC / ScMVA for expressing the FPP synthase gene and the gene for the complete MVA pathway is as described above (Example 2). This plasmid contains genes encoding all enzymes in the biosynthetic pathway from acetyl-CoA to FPP.

[0516] Select co-transformed cells on LB agar plates containing ampicillin (50 μg / mL) and chloramphenicol (34 μg / mL). Inoculate single colonies into 5 mL of liquid LB medium supplemented with the same antibiotics. Incubate the cultures overnight at 37 °C with shaking at 200 rpm. The next day, inoculate 0.2 mL of the overnight culture into 2 mL of TB medium supplemented with the above antibiotics and glycerol (final concentration 3% w / v). Incubate at 37 °C for 5 h with shaking at 200 rpm, then cool the cultures to 25 °C and shake at 200 rpm for another hour. Then add IPTG (final 0.1 mM) to each tube and cover the cultures with 200 μL of decane. Incubate the cultures at 25 °C for another 48 h with shaking at 200 rpm, then extract twice with 1 volume of ethyl acetate. Before analyzing the samples by GC / MS, add 50 μL of 2 mg / mL isophorene as an internal standard to the organic phase.

[0517] BtHAD, SmHAD1, SmHAD2, fungal HAD-like TPS, and their mutants were analyzed in vivo biochemical assays. Nephrolol and bismuth subtilisol were not detected in any of the enzymes with mutated class II motifs. These results support the hypothesis of protonation-induced FPP cyclization involving class II motifs, rather than classical ionization-induced cyclization mediated by class I motifs. Figure 19We also observed that mutants of the QW motif resulted in a 90% or higher loss of productivity in zosophyll and bismuth subtilis, which may be related to enzyme integrity and thermal stability.

[0518] Example 4

[0519] Characterization of plant HAD-like TPS in in vitro systems (specified as evidence of phytosporin and phytoalkenyl diphosphates) according to).

[0520] In vitro biochemical assays were performed to investigate the cyclization mechanism of the newly identified plant HAD-like styranne sesquiterpene synthase.

[0521] BL21(DE3) *E. coli* cells (Tiangen) were transformed with plasmids pETDuet-SmHAD1, pETDuet-SmHAD2, and pETDuet-BtHAD, respectively. Transformed cells were selected on LB agar plates containing ampicillin (50 μg / mL). Single colonies were inoculated into 25 mL of liquid LB medium supplemented with the same antibiotic. Cultures were incubated at 37°C and 200 rpm for 5 hours until an OD of approximately 0.5 was reached. 600 Then cool to 20°C and shake at 200 rpm for 0.5 hours. Then add IPTG (final 0.1 mM) to each tube and incubate the culture at 20°C and 200 rpm for another 18 hours.

[0522] For in vitro biochemical analysis, 25 mL of *E. coli* cultures expressing SmHAD1, SmHAD2, or BtHAD were centrifuged and resuspended in 5 mL of 50 mM Tris-HCl pH 8.0 buffer (10 mM MgCl2, 5 mM DTT), and then sonicated to produce crude cell lysates. 29 μM FPP was added to 1 mL of each crude cell lysate for reaction, and the mixture was incubated at 30 °C and 50 rpm for 2 h. The samples were then aliquoted, with one half containing 20 μL of bacterial alkaline phosphatase (BAP; Sangon-B004081-100) and reacted for an additional 1 h at 25 °C (BAP can remove potential diphosphate groups from complementenyl and / or zephyranyl diphosphates to generate complementol and / or zephyranol). A control (the latter half of the sample) without BAP was incubated under the same conditions. 10 μL of 2 mg / mL isophorene (internal standard) was added to the organic phase of 0.5 mL aliquots as an internal standard. The reactants were then extracted with 0.25 mL of ethyl acetate and analyzed by GC / MS, as described above.

[0523] Adding BAP to the reactions using SmHAD2 and BtHAD lysates resulted in a 3.5-fold and 1.8-fold increase in the amount of complementol, respectively. The presence of complementol in the BAP-free reactions could be explained by the activity of *E. coli*'s native phosphatase. The reaction using the SmHAD1 lysate did not produce observable complementol, but it was detectable with the addition of BAP, albeit in small amounts. Similarly, only trace amounts of physostigmine were detected in both the reactions using the SmHAD2 lysate and those with the addition of BAP. Figure 20 (and Table 3).

[0524] Table 3. Productivity of SmHAD1 and SmHAD2 in in vitro biochemical analysis

[0525]

[0526] These findings support the hypothesis of a protonation-induced cyclization mechanism for the biosynthesis of bis(oxo)-enyl and zedoaryl diphosphates.

[0527] All publications mentioned in this application are incorporated by reference to disclose and describe the methods and / or materials relating to those publications.

[0528] sequence list

[0529] Table 4. Sequences described and used here

[0530]

[0531]

[0532]

[0533] SEQ ID NO:1

[0534] Bazzania sp. BazzHAD1 wt nucleic acid sequence

[0535]

[0536] SEQ ID NO:2

[0537]

[0538] SEQ ID NO:3

[0539] Amino acid sequence of Bazzania sp._BazzHAD1

[0540] MAPPFGFIGGNATSVTTSPDVLDYNYDGGPPTGQYKSLILDIGGVLLQPNLDRSYGTIVPPVQFKRMLRSHIWFNYSMGKMSAEEVFRQLSVKSGYPAEDIRSFGRKARESLVPITQITDLLFKLSKESKVRLFCMTNCPAEDFEYLSQAYPKLFGLFERIFTSASTGMRKPNRNFFRHVLQETGIRASETIFVDDLVPNIMAAEALDFTAIHVSPESPAETVRLLMFHLRKPEERLAAARDYLLRNRGVDKALSLTSDGVRVQVYFDHFVVAEVLNDERFLPITAAPESGTMNFFPKEERKATRDITTNAVTHVDLPDDLDSTSVALSVLYKFSKVGMETVQRVLDMMEKRVDDDGIFQTYFDPLRLRVDPVAATNTVYLFHLGGRPGPTQQTEQFLSKLLEKRSYLRGTLYYRTPEAFLYQLTRLVVSFEEYFRKTGFLEMLKKRLSERVGIESDAFTLAMRILACVNCGIPVAELNRDMDRLARMQNVDGSWDVCPYYSYNDPKSWFGNELLTTAFAAAALEQSEPNFN

[0541] SEQ ID NO:4

[0542] Nucleic acid sequence of Bazzania sp._BazzHAD2_wt

[0543]

[0544] SEQ ID NO:5

[0545]

[0546] SEQ ID NO:6

[0547] Amino acid sequence of Bazzania sp._BazzHAD2

[0548] MERPHFDTLIMDLGGVLVDFSLQTSTSTSVKSLKFVLRSATWAAYECGHMSESDCYAAVAKDLGSSANASQVGEAISEVRKSLQVNEDLIGVIRELKAQNGLRVYVMSNIPQPDFDAVKAKSASWGVFDGMYPSYAVGTHKPDLAFYRHVLEETNTDPLKAIFVDDSLQNVIAARSFGLTGIIYNDTVNAARTLRNLLGDPILRGEQYLSSHAGQLLSISDTGTPFPDNFSQLLILDATGNRDLVVLENPQRTWNYFIGKPVLTSETFPDDLDTTSIALMALNVDLEVANSVMDEMLEFRNSDGLFLTYFDETRPRVDAVVNVNVLRLFHRHEREREAQQPLEWITNVLTHRAYVDGTLFYYHAESFLYFLSRLFCENSSVQSRFQELLEQRLRERIGTPGDALSLAMRALACQMLGIDSSTDVRALLPLQCDDGGWEAGWVCRYGSNGMRVGSRGYTTALAINAIRGARVSR

[0549] SEQ ID NO:7

[0550] Nucleic acid sequence of Bazzania sp._BazzHAD3_wt

[0551]

[0552] SEQ ID NO:8

[0553]

[0554] SEQ ID NO:9

[0555] biāntáishǔ(Bazzania sp.)_BazzHAD3_ānjīsuānxùliè

[0556] MPALSSHMEPPNFDTLILDLGDVLVGGSLQGSYALAMKRILKLSLRTPTWGKYECGQLSELDCYDAVAEDLGNSTSSAQVGEVVSEARKSLQVNEDLIKVLQELKAENNLRVYVMSNIPKPDLAVVKAKSVNWGVIDGWYPSYAVGFHKPDLAFYRHVLEETNTDPLKAIFVDDKVQNVIVARSFGMTGIVFKDTKSTTKELKNLGDPVKRGEHFLSAHAKQLESVCGDGTTFLQLDNFAQ LILHDTGNSDLVARQHHGRTWNYFIGKPMGTTDTFPNDLDTTSIALLTLNVDSEVARSVLDEMLLYTSSDGLVEVYFDKTRPRVDPVVCVNVLRLFSKYGRELQLQKTLDWVTEVLVHRAYIDGTLFYYHAESFLYFLSCLYKENPRLQTQFREPLQERLRERIGQPGDALCLAMRAIACQTVGIGNVIDVQALLPLQSSDGGWEAGWVCRMGTSGVPVGNRGVTTALCIRGASSYHT

[0557] SEQ ID NO:10

[0558] biāntái(Bazzania trilobata)_BtHAD_wt_hésuānxùliè

[0559]

[0560] SEQ ID NO:11

[0561] Bazzania trilobata BtHAD Escherichia coli, optimized

[0562]

[0563] SEQ ID NO:12

[0564] Amino acid sequence of BtHAD from Bazzania trilobata

[0565] MAPPFGYIGGNESSVTTSPDVLDYSYDGGPPTGQYKSMILDIGGVLLQPNLDRSYGTIVPPVQFKRMLRSHIWFNYSMGKVSAEEVFRQLSVKSGYPAEDIRSFARKARESLVPITQITDLLVKLSKESKIRLFCMTNCPAEDFAYLSQAYPNLFGLFERIFTSASTGMRKPNRNFFRHVLQETGISATETIFVDDLVPNIMAAEALDFTAIHVSPESPAETVRLLMFHLRKPEERLAAARDYLLRNRGVDKALSLTSDGVRVQVYFDHFVVAEVLNDERFLPITAAPKSGTMNFFPKEERNATRDMTTNAVTHVDLPDDLDSTSVALSVLYKFSKVAMETVQRVLDMMEKRVDDDGIFQTYFDPLRLRVDPVAATNTVYLFHLGGRPGPTQQTEQFLSKLLEKRSYVRGTLYYRTPEAFLYQLTRLVVSFEYFRKTGFLEMLKKRLSERVGIESDAFTLAMRILACVNCGIPVAELNRDMDRLARMQNVDGSWDVCPYYSYNDPKSWFGNELLTTAFAAAALEHSEPNFK

[0566] SEQ ID NO:13

[0567] Artificial sequence - BtHAD mutant of class II motif, optimized for Escherichia coli

[0568]

[0569] SEQ ID NO:14

[0570] Artificial sequence - BtHAD mutant of class II motif

[0571] MAPPFGYIGGNESSVTTSPDVLDYSYDGGPPTGQYKSMILDIGGVLLQPNLDRSYGTIVPPVQFKRMLRSHIWFNYSMGKVSAEEVFRQLSVKSGYPAEDIRSFARKARESLVPITQITDLLVKLSKESKIRLFCMTNCPAEDFAYLSQAYPNLFGLFERIFTSASTGMRKPNRNFFRHVLQETGISATETIFVDDLVPNIMAAEALDFTAIHVSPESPAETVRLLMFHLRKPEERLAAARDYLLRNRGVDKALSLTSDGVRVQVYFDHFVVAEALDFTAIHVSPESPAETVRLLMFHLRKPEERLAAARDYLLRNRGVDKALSLTSDGVRVQVYFDHFVVAEALDFTAIHVSPESPAETVRLLMFHLGGRPGPTQQTEQFLSKLLEKRSYVRGTLYYRTPEAFLYQLTRLVVSFEYFRKTGFLEMLKKRLSERVGIESDAFTLAMRILACVNCGIPVAELNRDMDRLARMQNVDGSWDVCPYYSYNDPKSWFGNELLTTAFAAAALEHSEPNFK

[0572] SEQ ID NO:15

[0573] Artificial sequence - Optimized BtHAD mutant of QW motif - Escherichia coli

[0574]

[0575] SEQ ID NO:16

[0576] Artificial sequence - BtHAD mutant with QW motif

[0577] MAPPFGYIGGNESSVTTSPDVLDYSYDGGPPTGQYKSMILDIGGVLLQPNLDRSYGTIVPPVQFKRMLRSHIWFNYSMGKVSAEEVFRQLSVKSGYPAEDIRSFARKARESLVPITQITDLLVKLSKESKIRLFCMTNCPAEDFAYLSQAYPNLFGLFERIFTSASTGMRKPNRNFFRHVLQETGISATETIFVDDLVPNIMAAEALDFTAIHVSPESPAETVRLLMFHLRKPEERLAAARDYLLRNRGVDKALSLTSDGVRVQVYFDHFVVAEVLNDERFLPITAAPKSGTMNFFPKEERNATRDMTTNAVTHVDLPDDLDSTSVALSVLYKFSKVAMETVQRVLDMMEKRVDDDGIFQTYFDPLRLRVDPVAATNTVYLFHLGGRPGPTQQTEQFLSKLLEKRSYVRGTLYYRTPEAFLYQLTRLVVSFEYFRKTGFLEMLKKRLSERVGIESDAFTLAMRILACVNCGIPVAELNRDMDRLARMQNVAGSWDVCPYYSYNDPKSWFGNELLTTAFAAAALEHSEPNFK

[0578] SEQ ID NO:17

[0579] Selaginella moellendorffii - SmHAD1_wt nucleic acid sequence

[0580]

[0581] SEQ ID NO:18

[0582] Selaginella moellendorffii_SmHAD1_Escherichia coli, optimized

[0583]

[0584] SEQ ID NO:19

[0585] The amino acid sequence of Selaginella moellendorffii_SmHAD1

[0586] MIIISFACLKFGCGDNGPRGDLLRRALQHSSFLAYSCGELDRTAAISTISRKFKLTDPALVDSMLLEAASSCEVDKELLSLLQSTRQQLLGWIDIPPQEWERVYSLFPVSLWKNFA TISRDLDSLLGDIRSHAVIVDKSVEMAALHAFESLLALPYASPNKKACQDFLKRRFLLPRAVKLGMDRVKEEVRSKQLKSLYLLDNGRHELVSEIFFPCVAAWCLPEIIPLGWMKYL QVLIKRGHPFGYFGADVSRYPPDIDTMSTCVSTLFDLSLVTSAQAMHFLEICLENVNDQNQLLTYLDVERPRVDPVVIANVVYLAYALSMEDHPVVRHNENLIQRYLLSGGFVYGTR YYLSQEDFLFMYGRVLATFGEKREIPNFDLVYQAMEAALVNRIGNETESKPLDVAKRILLSRGFGIRNTIDVDLLLKMQNEDGSWPLQVLSNLPSAKGGVFNSVVDLSFAVRALQSQD

[0587] SEQ ID NO:20

[0588] Artificial sequence_class II motif SmHAD1 mutant_E. coli, optimized

[0589]

[0590] SEQ ID NO:21

[0591] SmHAD1 mutant with artificial sequence class II motif

[0592] MIIISFACLKFGCGDNGPRGDLLRRALQHSSFLAYSCGELDRTAAISTISRKFKLTDPALVDSMLLEAASSCEVDKELLSLLQSTRQQLLGWIDIPPQEWERVYSLFPVSLWKNFA TISRDLDSLLGDIRSHAVIVDKSVEMAALHAFESLLALPYASPNKKACQDFLKRRFLLPRAVKLGMDRVKEEVRSKQLKSLYLLDNGRHELVSEIFFPCVAAWCLPEIIPLGWMKYL QVLIKRGHPFGYFGADVSRYPPAIATMSTCVSTLFDLSLVTSAQAMHFLEICLENVNDQNQLLTYLDVERPRVDPVVIANVVYLAYALSMEDHPVVRHNENLIQRYLLSGGFVYGTR YYLSQEDFLFMYGRVLATFGEKREIPNFDLVYQAMEAALVNRIGNETESKPLDVAKRILLSRGFGIRNTIDVDLLLKMQNEDGSWPLQVLSNLPSAKGGVFNSVVDLSFAVRALQSQD

[0593] SEQ ID NO:22

[0594] Artificial sequence QW motif-based SmHAD1 mutant in E. coli, optimized.

[0595]

[0596] SEQ ID NO:23

[0597] Artificial sequence - SmHAD1 mutant of QW motif

[0598] MIIISFACLLKFGCGDNGPRGDLLRRALQHSSFLAYSCGELDRTAAISTISRKFKLTDPALVDSMLLEAASSCEVDKELLSLLQSTRQQLLGWIDIPPQEWERVYSLFPVSLWKNFATISRDLDSLLGDIRSHAVIVDKSVEMAALHAFESLLALPYASPNKKACQDFLKRRFLLPRAVKLGMDRVKEEVRSKQLKSLYLLDNGRHELVSEIFFPCVAAWCLPEIIPLGWMKYLQVLIKRGHPFGYFGADVSRYPPDIDTMSTCVSTLFDLSLVTSAQAMHFLEICLENVNDQNQLLTYLDVERPRVDPVVIANVVYLAYALSMEDHPVVRHNENLIQRYLLSGGFVYGTRYYLSQEDFLFMYGRVLATFGEKREIPNFDLVYQAMEAALVNRIGNETESKPLDVAKRILLSRGFGIRNTIDVDLLLKMQNAAGSWPLQVLSNLPSAKGGVFNSVVDLSFAVRALQSQD

[0599] SEQ ID NO:24

[0600] Selaginella moellendorffii - SmHAD2_wt_nucleic acid sequence

[0601]

[0602] SEQ ID NO:25

[0603] Selaginella moellendorffii_SmHAD2_Escherichia coli, optimized

[0604]

[0605] SEQ ID NO:26

[0606] The amino acid sequence of Selaginella moellendorffii_SmHAD2

[0607] MIIISFACLKFGCGDSGPRGDLLRRALQHSSFLAYSCGELDRAAAISTISRKFKLKEPALLDSMLLEAASSCEVDEELLSLLQSTRQQLLGWIDIPPQEWERVYNLFPWSLWKNFA TISRLDSLLGDIRFHAVIVDKSVEMAALHAFESLLALPYASPNKKACQDFLKRRFLLPRAVKLGMDRVKEEVRSKQLKSLYLLDNGRQELVSEIFFPCVAAWCLPEIIPLGWMESL QVLIERGHPFGYFGADVSRYPPDIDTMSTCVSTLFDLSLVTSAQAMHFLEICLENVNDQNQLLTYLDLERPRVDPVVIANVVYFAYALSMEDHPVVRHNENLIQRYLLSGGFVYGTR YYLSQEDFLFMYGRVLATFGEKREIPNFDLVYQAMEAALVNRIGNETESKPLDVAKRILLSRYFGIRNTIDVDLLLKMQNEDGSWPLQVLSNLPSAKGGVFNSVVDLSFAVRALQSQD

[0608] SEQ ID NO:27

[0609] Artificial sequence_class II motif SmHAD2 mutant_E. coli, optimized

[0610]

[0611] SEQ ID NO:28

[0612] SmHAD2 mutant with artificial sequence class II motif

[0613] MIIISFACLKFGCGDSGPRGDLLRRALQHSSFLAYSCGELDRAAAISTISRKFKLKEPALLDSMLLEAASSCEVDEELLSLLQSTRQQLLGWIDIPPQEWERVYNLFPWSLWKNFA TISRLDSLLGDIRFHAVIVDKSVEMAALHAFESLLALPYASPNKKACQDFLKRRFLLPRAVKLGMDRVKEEVRSKQLKSLYLLDNGRQELVSEIFFPCVAAWCLPEIIPLGWMESL QVLIERGHPFGYFGADVSRYPPAIATMSTCVSTLFDLSLVTSAQAMHFLEICLENVNDQNQLLTYLDLERPRVDPVVIANVVYFAYALSMEDHPVVRHNENLIQRYLLSGGFVYGTR YYLSQEDFLFMYGRVLATFGEKREIPNFDLVYQAMEAALVNRIGNETESKPLDVAKRILLSRYFGIRNTIDVDLLLKMQNEDGSWPLQVLSNLPSAKGGVFNSVVDLSFAVRALQSQD

[0614] SEQ ID NO:29

[0615] Artificial sequence QW motif-based SmHAD2 mutant E. coli, optimized

[0616]

[0617] SEQ ID NO:30

[0618] Artificial sequence - SmHAD2 mutant of QW motif

[0619] MIIISFACLLKFGCGDSGPRGDLLRRALQHSSFLAYSCGELDRAAAISTISRKFKLKEPALLDSMLLEAASSCEVDEELLSLLQSTRQQLLGWIDIPPQEWERVYNLFPWSLWKNFATISRDLDSLLGDIRFHAVIVDKSVEMAALHAFESLLALPYASPNKKACQDFLKRRFLLPRAVKLGMDRVKEEVRSKQLKSLYLLDNGRQELVSEIFFPCVAAWCLPEIIPLGWMESLQVLIERGHPFGYFGADVSRYPPDIDTMSTCVSTLFDLSLVTSAQAMHFLEICLENVNDQNQLLTYLDLERPRVDPVVIANVVYFAYALSMEDHPVVRHNENLIQRYLLSGGFVYGTRYYLSQEDFLFMYGRVLATFGEKREIPNFDLVYQAMEAALVNRIGNETESKPLDVAKRILLSRYFGIRNTIDVDLLLKMQNAAGSWPLQVLSNLPSAKGGVFNSVVDLSFAVRALQSQD

[0620] SEQ ID NO:31

[0621] Gelatoporia subvermispora _ EMD37666.1 _ Escherichia coli, optimized

[0622]

[0623] SEQ ID NO:32

[0624] Gelatoporia subvermispora _EMD37666.1_ Amino acid sequence

[0625] MSAAAQYTTLILDLGDVLFTWSPKTKTSIPPRTLKEILNSATWYEYERGRISQDECYERVGTEFGIAPSEIDNAFKQARDSMESNDELIALVRELKTQLDGELLVFALSNISLPDYEYVLTKPADWSIFDKVFPSALVGERKPHLGVYKHVIAETGIDPRTTVFVDDKIDNVLSARSVGMHGIVFEKQEDVMRALRNIFGDPVRRGREYLRRNAMRLESVTDHGVAFGENFTQLLILELTNDPSLVTLPDRPRTWNFFRGNGGRPSKPLFSEAFPDDLDTTSLALTVLQRDPGVISSVMDEMLNYRDPDGIMQTYFDDGRQRLDPFVNVNVLTFFYTNGRGHELDQCLTWVREVLLYRAYLGGSRYYPSADCFLYFISRLFACTNDPVLHHQLKPLFVERVQEQIGVEGDALELAFRLLVCASLDVQNAIDMRRLLEMQCEDGGWEGGNLYRFGTTGLKVTNRGLTTAAAVQAIEASQRRPPSPSPSVESTKSPITPVTPMLEVPSLGLSISRPSSPLLGYFRLPWKKSAEVH

[0626] SEQ ID NO:33

[0627] Artificial sequence _EMD37666.1_ Mutant of class II motif _ Escherichia coli, optimized

[0628]

[0629] SEQ ID NO:34

[0630] Artificial sequence _EMD37666.1_ Mutant of Class II motif

[0631] MSAAAQYTTLILDLGDVLFTWSPKTKTSIPPRTLKEILNSATWYEYERGRISQDECYERVGTEFGIAPSEIDNAFKQARDSMESNDELIALVRELKTQLDGELLVFALSNISLPDYEYVLTKPADWSIFDKVFPSALVGERKPHLGVYKHVIAETGIDPRTTVFVDDKIDNVLSARSVGMHGIVFEKQEDVMRALRNIFGDPVRRGREYLRRNAMRLESVTDHGVAFGENFTQLLILELTNDPSLVTLPDRPRTWNFFRGNGGRPSKPLFSEAFPAALATTSLALTVLQRDPGVISSVMDEMLNYRDPDGIMQTYFDDGRQRLDPFVNVNVLTFFYTNGRGHELDQCLTWVREVLLYRAYLGGSRYYPSADCFLYFISRLFACTNDPVLHHQLKPLFVERVQEQIGVEGDALELAFRLLVCASLDVQNAIDMRRLLEMQCEDGGWEGGNLYRFGTTGLKVTNRGLTTAAAVQAIEASQRRPPSPSPSVESTKSPITPVTPMLEVPSLGLSISRPSSPLLGYFRLPWKKSAEVH

[0632] SEQ ID NO:35

[0633] Dichomitus squalens _XP_007369631.1_ Escherichia coli, optimized

[0634]

[0635] SEQ ID NO:36

[0636] Amino acid sequence of *Dichomitus squalens* (XP_007369631.1)

[0637] MASIHRRYTTLILDLGDVLFRWSPKTETAIPPQQLKDILSSVTWFEYERGRLSQEACYERCAEEFKIEASVIAEAFKQARGSLRPNEEFIALIRDLRREMHGDLTVLALSNISLPDYEYIMSLSSDWTTVF DRVFPSALVGERKPHLGCYRKVISEMNLEPQTTVFVDDKLDNVASARSLGMHGIVFDNQANVFRQLRNIFGDPIRRGQEYLRGHAGKLESSTDNGLIFEENFTQLIIYELTQDRTLISLSECPRTWNFFRGE PLFSETFPDDVDTTSVALTVLQPDRALVNSVLDEMLEYVDADGIMQTYFDRSRPRMDPFVCVNVLSLFYENGRGHELPRTLDWVYEVLLHRAYHGGSRYYLSPDCFLFFMSRLLKRADDPAVQARLRPLFVE RVNERVGAAGDSMDLAFRILAAASVGVQCPRDLERLTAGQCDDGGWDLCWFYVFGSTGVKAGNRGLTTALAVVGSTGVTAIGRPPSPSSAASSSFRPSSPYKFLGISRPASPIRFGDLLRPWRKMSRSNLKSQ

[0638] SEQ ID NO:37

[0639] Mutant of artificial sequence _XP_007369631.1_class II motif_E. coli, optimized

[0640]

[0641] SEQ ID NO:38

[0642] Mutants of the artificial sequence _XP_007369631.1_Class II motif

[0643] MASIHRRYTTLILDLGDVLFRWSPKTETAIPPQQLKDILSSVTWFEYERGRLSQEACYERCAEEFKIEASVIAEAFKQARGSLRPNEEFIALIRDLRREMHGDLTVLALSNISLPDYEYIMSLSSDWTTVF DRVFPSALVGERKPHLGCYRKVISEMNLEPQTTVFVDDKLDNVASARSLGMHGIVFDNQANVFRQLRNIFGDPIRRGQEYLRGHAGKLESSTDNGLIFEENFTQLIIYELTQDRTLISLSECPRTWNFFRGE PLFSETFPAAVATTSVALTVLQPDRALVNSVLDEMLEYVDADGIMQTYFDRSRPRMDPFVCVNVLSLFYENGRGHELPRTLDWVYEVLLHRAYHGGSRYYLSPDCFLFFMSRLLKRADDPAVQARLRPLFVE RVNERVGAAGDSMDLAFRILAAASVGVQCPRDLERLTAGQCDDGGWDLCWFYVFGSTGVKAGNRGLTTALAVVGSTGVTAIGRPPSPSSAASSSFRPSSPYKFLGISRPASPIRFGDLLRPWRKMSRSNLKSQ

[0644] SEQ ID NO:39

[0645] Artificial sequence_Class II synthase motif_mutated_BtHAD

[0646] PDALASTS

[0647] SEQ ID NO:40

[0648] Artificial sequence_Class II synthase motif_mutated_SmHAD1+2

[0649] PPAIATMS

[0650] SEQ ID NO:41

[0651] Artificial sequence_Class II synthase motif_Mutant_EMD37666

[0652] PDAALATTS

[0653] SEQ ID NO:42

[0654] Artificial sequence_Class II synthase motif_Mutated_XP_007369631

[0655] PDAAVATTS

[0656] SEQ ID NO:43

[0657] Artificial sequence_QW motif_mutated_BtHAD

[0658] QNVAGSW

[0659] SEQ ID NO:44

[0660] Artificial sequence_QW motif_mutated_SmHAD1+2

[0661] QNAAGSW

[0662] SEQ ID NO:45

[0663] Artificial sequence_Class I synthase motif DDxx(D / E)

[0664] DDxx(D / E)

[0665] SEQ ID NO:46

[0666] Artificial sequence_Class II synthase motif PxDxD(T / S)(T / M)S

[0667] PxDxD(T / S)(T / M)S

[0668] SEQ ID NO:47

[0669] Artificial Sequence_Class II Synthesizer Motif PDDLDSTS

[0670] PDDLDSTS

[0671] SEQ ID NO:48

[0672] Artificial Sequence_Class II Synthase Motif PDDLDTTS

[0673] PDDLDTTS

[0674] SEQ ID NO:49

[0675] Artificial Sequence_Class II Synthase Motif PPDIDTMS

[0676] PPDIDTMS

[0677] SEQ ID NO:50

[0678] Artificial Sequence_Class II Synthase Motif PNDIDTMS

[0679] PNDIDTTS

[0680] SEQ ID NO:51

[0681] Artificial sequence QW motif

[0682] QxxDGxW

[0683] SEQ ID NO:52

[0684] Artificial sequence_QW motif QNVDGSW

[0685] QNVDGSW

[0686] SEQ ID NO:53

[0687] Artificial sequence_QW motif QCDDGGW

[0688] QCDDGGW

[0689] SEQ ID NO:54

[0690] Artificial sequence_QW motif QSSDGGW

[0691] QSSDGGW

[0692] SEQ ID NO:55

[0693] Artificial sequence_QW motif QNEDGSW

[0694] QNEDGSW

[0695] SEQ ID NO:56

[0696] Artificial sequence_conserved motif 1 Lxxxx(W / F)xxYxxG

[0697] Lxxxx(W / F)xxYxxG

[0698] SEQ ID NO:57

[0699] Artificial Sequence_Conserved Motif 1 LRSHIWFNYSMG

[0700] LRSHIWFNYSMG

[0701] SEQ ID NO:58

[0702] Artificial sequence_conserved motif 1 LRSATWAAYECG

[0703] LRSATWAAYECG

[0704] SEQ ID NO:59

[0705] Artificial sequence_conserved motif 1 LRTPTWGKYECG

[0706] LRTPTWGKYECG

[0707] SEQ ID NO:60

[0708] Artificial sequence_conserved motif 1 LQHSSFLAYSCG

[0709] LQHSSFLAYSCG

[0710] SEQ ID NO:61

[0711] Artificial sequence_conserved motif 1 LRxxTWxxYECG

[0712] LRxxTWxxYECG

[0713] SEQ ID NO:62

[0714] Artificial sequence_conserved motif 2 YxDxxRxRVD(P / A)V(V / A)xxN

[0715] YxDxxRxRVDxVxxxN

[0716] SEQ ID NO:63

[0717] Artificial sequence_conserved motif 2 YFDPLRLRVDPVAATN

[0718] YFDPLRLRVDPVAATN

[0719] SEQ ID NO:64

[0720] Artificial sequence_conserved motif 2 YFDETRPRVDAVVNVN

[0721] YFDETRPRVDAVVNVN

[0722] SEQ ID NO:65

[0723] Artificial sequence_conserved motif 2 YFDKTRPRVDPVVCVN

[0724] YFDKTRPRVDPVVCVN

[0725] SEQ ID NO:66

[0726] Artificial sequence_conserved motif 2 YLDVERPRVDPVVIAN

[0727] YLDVERPRVDPVVIAN

[0728] SEQ ID NO:67

[0729] Artificial sequence_conserved motif 2 YLDLERPRVDPVVIAN

[0730] YLDLERPRVDPVVIAN

[0731] SEQ ID NO:68

[0732] Artificial sequence_conserved motif 2 Y(F / L)Dx(T / E)RPRVD(P / A)VVx(A / V)N

[0733] YxDxxRPRVDxVVxxN

[0734] SEQ ID NO:69

[0735] Artificial sequence_conserved motif 3 GTx(Y / F)YxxxExFL(Y / F)

[0736] GTxxYxxxExFLx

[0737] SEQ ID NO:70

[0738] Artificial sequence_conserved motif 3 GTLYYRTPEAFLY

[0739] GTLYYRTPEAFLY

[0740] SEQ ID NO:71

[0741] Artificial sequence_conserved motif 3 GTLFYYHAESFLY

[0742] GTLFYYHAESFLY

[0743] SEQ ID NO:72

[0744] Artificial sequence_conserved motif 3 GTRYYLSQEDFLF

[0745] GTRYYLSQEDFLF

[0746] SEQ ID NO:73

[0747] Dryopteris fragrans DfHAD wt nucleic acid sequence

[0748]

[0749] SEQ ID NO:74

[0750] Amino acid sequence of Dryopteris fragrans _DfHAD_

[0751] MEFSASAPPPRLASVIILEPLGFLLTPHYSSQLPKKLLRRLLCTRIWHRYQRGRLRLRDAAMLLAQLPFLAVSDHPWALDNLASLLRPTAVRAVPWMLLLLDFLRDELHLKVVCATNSSPEELQELRHQFPALFAKVDATVSSGEEGVGKPSVRFLQAALDKAGVHAQQTLYLDSFDSLETIMAARSLGMHALSVEPCHIDELTARASSGQLRDAQLIRRIVCAMHGPAVSAVVSGSITSSGPQTAKIEELPTAADSHLRSAALTSAQQFFLKVIAPHRPEKPFVQLPSLTSEGIRIYDTFAQFVIADLLDDTRFLPMQSPPPNGLITFVNPSAYLADDIKNGNSHIVPGVQFYASDACTLIDIPHDLDTTSVGLSVLHKFGKVDKDTLNKVLDRMLEQVSEDDGILQVYFDVERPRIDPVVVANTVFLFHLGKRGHEVARSEKFVESVLLQRAYEEGTLYYNLGEAFLVSVARLVHEFKEHFTRSGMRRALEERLRERARAGMQERDDALALAMRIRACALCGLAGEGLTKAAEQELLRLQCKSKGCWGCHPFYRNGSNVLSWIGSEALTTAYAIAALQPIDI

[0752] SEQ ID NO:75

[0753] Artificial sequence _ Class II synthase motif

[0754] DxDxxS

[0755] SEQ ID NO:76

[0756] Artificial sequence _ QW motif

[0757] QxxxxxW。 Sequence Listing <110> Firmenich <120> Novel peptides for the production of physostigmine and / or complement compounds <130> 13150 / WO <160> 76 <170> PatentIn version 3.5 <210> 1 <211> 1599 <212> DNA <213> genus *Bazzania* sp. <220> <221> misc_feature <223> BazzHAD1_wt_nucleic acid sequence <400> 1 atggctcctc ctttcggatt tattggtggg aatgcaacga gtgtgacgac ttccccagat 60 gtcctcgact acaactatga cggtggccct cctactggcc aatacaagag tttgatttta 120 gacattggtg gagtgttgct tcagcccaac cttgacagga gctatggcac gattgtacca 180 cctgtgcaat tcaaacgtat gttgcgctcg cacatctggt tcaattactc aatgggaaag 240 atgtctgcgg aggaagtgtt tcggcagctc tccgtgaagt ctggctaccc ggcagaagat 300 atccggagttttggtcgcaa agcgagagag tctcttgtcc ccatcacaca aataacggat 360 cttctgttca aactcagcaa ggaatcgaag gtcagactgt tctgtatgac caattgccct 420 gcagaagact tcgaatacct ctcacaggca tatcccaagc ttttcggctt gttcgagagg 480 atttttacgt cggcatccac cggtatgcgc aagccaaaca ggaatttctt ccgacatgta 540 ctccaggaga cgggaatcag agcgtcggag actattttcg tagacgacct ggtaccgaac 600 attatggccg ccgaggcgct ggacttcacg gcgattcacg tttctccaga gtcgcccgcc 660 gagaccgtga ggttgctgat gttccatctc cgtaagccag aagagcggct ggcggcggcg 720 cgcgactatc tcctgcgcaa ccgtggagtg gacaaagctc tctctctcac atctgatgga 780 gtccgagtcc aggtttattt cgaccacttc gtggtcgccg aagtgctcaa cgacgagcga 840 ttcttgccaa tcactgcagc tccggagtcc gggactatga actttttccc gaaggaggaa 900 cgcaaagcga ctcgagatat aacgaccaat gcagtcacac acgtggacct tcccgatgat 960 cttgactcca cctctgtcgc cttatcggtc ctttacaaat tctccaaagt ggggatggag 1020 acggttcaaa gggtcctgga tatgatggag aaacgtgtgg atgatgacgg catttttcaa 1080 acgtacttcg atcctttgcg gctgcgtgtt gacccggtgg ccgctactaa caccgtctac 1140 ctcttccatc taggaggtcg tcccggacct actcagcaga cggaacaatt cctctccaag 1200 ttactggaaa agagatctta tctccgcggc acactctact accggacccc cgaggcattt 1260 ctctatcaac tcaccagatt ggtggtatca tttgaggagt atttcagaaa gacaggcttc 1320 ctggaaatgc tcaagaaaag gctgagtgag cgagtgggca tagagagcga tgcattcacg 1380 ttggcgatga ggattctagc gtgtgtgaat tgtggaatcc ctgttgctga gctgaacaga 1440 gatatggaca gactagcacg gatgcaaaat gtggatggtt catgggacgt ctgcccttac 1500 tacagctaca atgatcccaa gagttggttt gggaacgagc ttcttaccac agcttttgct 1560 gctgccgctc ttgagcagtc tgagcccaac tttaactga 1599 <210> 2 <211> 1599 <212> DNA <213> Bazzania sp. <220> <221> misc_feature <223> BazzHAD1_Escherichia coli, optimized <400> 2 atggctccgc cgttcggctt tatcggtggc aacgcgacca gcgttaccac cagcccggat 60 gtgctggact acaactatga tggtggcccg ccgaccggcc agtacaaaag cctgatcctg 120 gatattggtg gcgttctgct gcaaccgaac ctggaccgta gctatggcac catcgttccg 180 ccggtgcagt tcaagcgtat gctgcgtagc cacatttggt ttaactacag catgggtaaa 240 atgagcgcgg aggaagttt ccgtcaactg agcgtgaaga gcggttatcc ggcggaggat 300 atccgtagct ttggccgtaa agcgcgtgaa agcctggttc cgatcaccca gattaccgac 360 ctgctgttca agctgagcaa agagagcaaa gtgcgtctgt tctgcatgac caactgcccg 420 gcggaggact ttgaatacct gagccaagcg tatccgaaac tgttcggtct gtttgaacgt 480 attttacca gcgcgagcac cggcatgcgt aagccgaacc gtaacttctt tcgtcacgtg 540 ctgcaggaga ccggcatccg tgcgagcgaa accattttcg ttgacgatct ggtgccgaac 600 atcatggcgg cggaggcgct ggattttacc gcgattcacg ttagcccgga gagcccggcg 660 gaaaccgtgc gtctgctgat gttccacctg cgtaaaccgg aggaacgtct ggcggcggcg 720 cgtgactacc tgctgcgtaa ccgtggtgtt gataaagcgc tgagcctgac cagcgatggt 780 gtgcgtgttc aggtgtattt cgatcacttt gtggttgcgg aagtgctgaa cgacgaacgt 840 tttctgccga tcaccgcggc gccggaaagc ggcaccatga acttctttcc gaaagaggaa 900 cgcaaggcga cccgtgatat taccaccaac gcggttaccc acgtggatct gccggacgat 960 ctggacagca ccagcgttgc gctgagcgtg ctgtacaaat tcagcaaggt tggtatggag 1020 accgttcagc gtgtgctgga catgatggaa aagcgtgtgg acgatgacgg catcttccaa 1080 acctactttg atccgctgcg tctgcgtgtt gacccggtgg cggcgaccaa caccgtttac 1140 ctgttccatc tgggtggccg tccgggtccg acccagcaaa ccgagcagtt tctgagcaaa 1200 ctgctggaaa agcgtagcta tctgcgtggc accctgtact atcgtacccc ggaggcgttc 1260 ctgtaccaac tgacccgtct ggtggttagc ttcgaggaat attttcgtaa aaccggtttc 1320 ctggaaatgc tgaagaaacg tctgagcgag cgtgtgggca ttgaaagcga tgcgtttacc 1380 ctggcgatgc gtatcctggc gtgcgttaac tgcggtattc cggtggcgga gctgaaccgt 1440 gatatggacc gtctggcgcg tatgcagaac gttgatggta gctgggacgt gtgcccgtac 1500 tatagctaca acgacccgaa gagctggttc ggcaacgaac tgctgaccac cgcgtttgcg 1560 gcggcggcgc tggagcaaag cgaaccgaac ttcaactaa 1599 <210> 3 <211> 532 <212> PRT <213> Bazzania sp. <220> <221> MISC_FEATURE <223> BazzHAD1 amino acid sequence <400> 3 Met Ala Pro Pro Phe Gly Phe Ile Gly Gly Asn Ala Thr Ser Val Thr 1 5 10 15 Thr Ser Pro Asp Val Leu Asp Tyr Asn Tyr Asp Gly Gly Pro Pro Thr 20 25 30 Gly Gln Tyr Lys Ser Leu Ile Leu Asp Ile Gly Gly Val Leu Leu Gln 35 40 45 Pro Asn Leu Asp Arg Ser Tyr Gly Thr Ile Val Pro Pro Val Gln Phe 50 55 60 Lys Arg Met Leu Arg Ser His Ile Trp Phe Asn Tyr Ser Met Gly Lys 65 70 75 80 Met Ser Ala Glu Glu Val Phe Arg Gln Leu Ser Val Lys Ser Gly Tyr 85 90 95 Pro Ala Glu Asp Ile Arg Ser Phe Gly Arg Lys Ala Arg Glu Ser Leu 100 105 110 Val Pro Ile Thr Gln Ile Thr Asp Leu Leu Phe Lys Leu Ser Lys Glu 115 120 125 Ser Lys Val Arg Leu Phe Cys Met Thr Asn Cys Pro Ala Glu Asp Phe 130 135 140 Glu Tyr Leu Ser Gln Ala Tyr Pro Lys Leu Phe Gly Leu Phe Glu Arg 145 150 155 160 Ile Phe Thr Ser Ala Ser Thr Gly Met Arg Lys Pro Asn Arg Asn Phe 165 170 175 Phe Arg His Val Leu Gln Glu Thr Gly Ile Arg Ala Ser Glu Thr Ile 180 185 190 Phe Val Asp Asp Leu Val Pro Asn Ile Met Ala Ala Glu Ala Leu Asp 195 200 205 Phe Thr Ala Ile His Val Ser Pro Glu Ser Pro Ala Glu Thr Val Arg 210 215 220 Leu Leu Met Phe His Leu Arg Lys Pro Glu Glu Arg Leu Ala Ala Ala 225 230 235 240 Arg Asp Tyr Leu Leu Arg Asn Arg Gly Val Asp Lys Ala Leu Ser Leu 245 250 255 Thr Ser Asp Gly Val Arg Val Gln Val Tyr Phe Asp His Phe Val Val 260 265 270 Ala Glu Val Leu Asn Asp Glu Arg Phe Leu Pro Ile Thr Ala Ala Pro 275 280 285 Glu Ser Gly Thr Met Asn Phe Phe Pro Lys Glu Glu Arg Lys Ala Thr 290 295 300 Arg Asp Ile Thr Thr Asn Ala Val Thr His Val Asp Leu Pro Asp Asp 305 310 315 320 Leu Asp Ser Thr Ser Val Ala Leu Ser Val Leu Tyr Lys Phe Ser Lys 325 330 335 Val Gly Met Glu Thr Val Gln Arg Val Leu Asp Met Met Glu Lys Arg 340 345 350 Val Asp Asp Asp Gly Ile Phe Gln Thr Tyr Phe Asp Pro Leu Arg Leu 355 360 365 Arg Val Asp Pro Val Ala Ala Thr Asn Thr Val Tyr Leu Phe His Leu 370 375 380 Gly Gly Arg Pro Gly Pro Thr Gln Gln Thr Glu Gln Phe Leu Ser Lys 385 390 395 400 Leu Leu Glu Lys Arg Ser Tyr Leu Arg Gly Thr Leu Tyr Tyr Arg Thr 405 410 415 Pro Glu Ala Phe Leu Tyr Gln Leu Thr Arg Leu Val Val Ser Phe Glu 420 425 430 Glu Tyr Phe Arg Lys Thr Gly Phe Leu Glu Met Leu Lys Lys Arg Leu 435 440 445 Ser Glu Arg Val Gly Ile Glu Ser Asp Ala Phe Thr Leu Ala Met Arg 450 455 460 Lion Island With Cys Val Asn Cys Gly Island Pro With Glu Lion Asn Arg 465 470 475 480 Asp Met Asp Arg Leu Ala Arg Met Gln Asn Val Asp Gly Ser Trp Asp 485,490,495 Val Cys Pro Tyr Tyr Ser Tyr Asn Asp Pro Lys Ser Trp Phe Gly Asn 500 505 510 Glu Leu Leu Thr Thr Ala Phe Ala Ala Ala Ala Ala Leu Glu Gln Ser Glu 515,520,525 Pro Asn Phe Asn 530 <210> 4 <211> 1422 <212> DNA <213> Bazzania sp. <220> <221> misc_feature <223> BazzHAD2_wt <400> 4 atggaaagac ctcatttcga cactctgatc atggacctgg gaggtgttct tgtcgacttc 60 tctctcaa catccacatc cactctgtg aaatctctca agtttgtatt acgttccgcc 120 acctgggcag cttacgagtg cggccatatg tcagaatcgg attgctatgc tgcagtggct 180 aaagatcttg gaagctctgc aaacgcatct caagttggtg aagccatctc agaagttcga 240 aagtcactgc aggtcaatga agatttgata ggggtgattc gggagctcaa agcccagaac 300 ggccttcgag tttacgtaat gtccaacata ccacaaccag acttcgatgc cgtcaaagct 360 aagtcagcga gctgggggagt attcgatggc atgtatccct cgtatgcggt gggcacccat 420 aagcccgatc ttgctttcta tcgacacgta ctcgaagaga caaacacaga ccctctcaaa 480 gcgatctttg tcgacgacag cttacaaaat gtcattgcag ccagatcttt tggattgaca 540 ggaattatct acaatgatac tgtcaatgct gctcgaacgc tgcgaaatct ccttggtgat 600 cctatcttga gaggtgagca gtacctttcg tcccacgctg gacaattgct gtctatctca 660 gacacaggaa cccccttccc agataacttt tcacagcttt taattttgga tgctaccgga 720 aaccgtgatc ttgttgttct tgaaaaccct cagcgaacat ggaactactt catcggaaag 780 cctgtgttga cgtctgaaac attccctgat gatctagata cgacatctat cgccctcatg 840 gcattgaatg tagacctaga ggtggccaat tctgtgatgg atgagatgct ggagtttaga 900 aacagtgatg gattgtttct gacgtacttt gacgaaacaa ggccgcgggt agatgctgtt 960 gtcaatgtta atgttcttcg tctcttccat cgacatgaac gggagagaga ggcccagcaa 1020 ccgctggagt ggattaccaa cgtccttacc caccgagcgt atgttgatgg cacactattt 1080 tactaccatg ctgagagctt cttatatttc ctttcgcgac tgttctgtga gaactccagc 1140 gtccaatcac gatttcaaga gcttcttgaa caacgcctgc gggaacgaat tggaaccccg 1200 ggcgacgcac tgagtctagc catgcgtgca ttagcatgcc agatgctggg aatcgatagt 1260 tctaccgacg ttcgcgctct tctgcccttg caatgcgacg acggcggttg ggaagctggc 1320 tgggtgtgca gatacggttc aaatggtatg cgcgttggaa gccggggata cactacggct 1380 ttggctataa atgctatcag aggggcaaga gtctcgaggt ga 1422 <210> 5 <211> 1422 <212> DNA <213> Bazzania sp. <220> <221> misc_feature <223> BazzHAD2_Escherichia coli, optimized <400> 5 atggaacgtc cgcacttcga caccctgatc atggatctgg gtggcgtgct ggttgacttt 60 agcctgcaga ccagcaccag caccagcgtt aagagcctga aatttgtgct gcgtagcgcg 120 acctgggcgg cgtacgaatg cggtcacatg agcgagagcg actgctatgc ggcggttgcg 180 aaggatctgg gtagcagcgc gaacgcgagc caagtgggtg aagcgatcag cgaggttcgt 240 aagagcctgc aagtgaacga agatctgatc ggtgttattc gtgagctgaa agcgcagaac 300 ggcctgcgtg tgtacgttat gagcaacatt ccgcaaccgg acttcgatgc ggttaaggcg 360 aaaagcgcga gctggggtgt gtttgatggc atgtacccga gctatgcggt tggcacccac 420 aagccggacc tggcgtttta tcgtcacgtg ctggaggaaa ccaacaccga tccgctgaaa 480 gcgatcttcg ttgacgatag cctgcaaaac gtgattgcgg cgcgtagctt tggtctgacc 540 ggcatcattt acaacgacac cgtgaacgcg gcgcgtaccc tgcgtaacct gctgggtgat 600 ccgatcctgc gtggcgaaca gtatctgagc agccacgcgg gtcaactgct gagcattagc 660 gacaccggca ccccgttccc ggataacttt agccagctgc tgattctgga cgcgaccggt 720 aaccgtgatc tggtggttct ggaaaacccg caacgtacct ggaactactt catcggcaaa 780 ccggttctga ccagcgagac ctttccggac gatctggaca ccaccagcat tgcgctgatg 840 gcgctgaacg tggacctgga agttgcgaac agcgtgatgg atgaaatgct ggagttccgt 900 aacagcgacg gtctgttcct gacctacttt gacgagaccc gtccgcgtgt ggatgcggtg 960 gttaacgtga acgttctgcg tctgtttcac cgtcacgagc gtgaacgtga ggcgcagcaa 1020 ccgctggagt ggattaccaa cgtctgacc caccgtgcgt acgtggacgg caccctgttc 1080 tactatcacg cggaaagctt cctgtatttt ctgagccgtc tgttctgcga gaacagcagc 1140 gttcagagcc gttttcaaga actgctggag cagcgtctgc gtgaacgtat tggcaccccg 1200 ggtgatgcgc tgagcctggc gatgcgtgcg ctggcgtgcc agatgctggg tattgacagc 1260 agcaccgatg tgcgtgcgct gctgccgctg caatgcgatg atggtggctg ggaggcgggt 1320 tgggtttgcc gttacggtag caacggcatg cgtgtgggta gccgtggcta taccaccgcg 1380 ctggcgatca acgcgattcg tggtgcgcgt gttagccgtt aa 1422 <210> 6 <211> 473 <212> PRT <213> Bazzania sp. <220> <221> MISC_FEATURE <223> BazzHAD2 amino acid sequence <400> 6 Met Glu Arg Pro His Phe Asp Thr Leu Ile Met Asp Leu Gly Gly Val 1 5 10 15 Leu Val Asp Phe Ser Leu Gln Thr Ser Thr Ser Thr Ser Val Lys Ser 20 25 30 Leu Lys Phe Val Leu Arg Ser Ala Thr Trp Ala Ala Tyr Glu Cys Gly 35 40 45 His Met Ser Glu Ser Asp Cys Tyr Ala Ala Val Ala Lys Asp Leu Gly 50 55 60 Ser Ser Ala Asn Ala Ser Gln Val Gly Glu Ala Ile Ser Glu Val Arg 65 70 75 80 Lys Ser Leu Gln Val Asn Glu Asp Leu Ile Gly Val Ile Arg Glu Leu 85 90 95 Lys Ala Gln Asn Gly Leu Arg Val Tyr Val Met Ser Asn Ile Pro Gln 100 105 110 Pro Asp Phe Asp Ala Val Lys Ala Lys Ser Ala Ser Trp Gly Val Phe 115 120 125 Asp Gly Met Tyr Pro Ser Tyr Ala Val Gly Thr His Lys Pro Asp Leu 130 135 140 Ala Phe Tyr Arg His Val Leu Glu Glu Thr Asn Thr Asp Pro Leu Lys 145 150 155 160 Ala Ile Phe Val Asp Asp Ser Leu Gln Asn Val Ile Ala Ala Arg Ser 165 170 175 Phe Gly Leu Thr Gly Ile Ile Tyr Asn Asp Thr Val Asn Ala Ala Arg 180 185 190 Thr Leu Arg Asn Leu Leu Gly Asp Pro Ile Leu Arg Gly Glu Gln Tyr 195 200 205 Leu Ser Ser His Ala Gly Gln Leu Leu Ser Ile Ser Asp Thr Gly Thr 210 215 220 Pro Phe Pro Asp Asn Phe Ser Gln Leu Leu Ile Leu Asp Ala Thr Gly 225 230 235 240 Asn Arg Asp Leu Val Val Leu Glu Asn Pro Gln Arg Thr Trp Asn Tyr 245 250 255 Phe Ile Gly Lys Pro Val Leu Thr Ser Glu Thr Phe Pro Asp Asp Leu 260 265 270 Asp Thr Thr Ser Ile Ala Leu Met Ala Leu Asn Val Asp Leu Glu Val 275 280 285 Ala Asn Ser Val Met Asp Glu Met Leu Glu Phe Arg Asn Ser Asp Gly 290 295 300 Leu Phe Leu Thr Tyr Phe Asp Glu Thr Arg Pro Arg Val Asp Ala Val 305 310 315 320 Val Asn Val Asn Val Leu Arg Leu Phe His Arg His Glu Arg Glu Arg 325 330 335 Glu Ala Gln Gln Pro Leu Glu Trp Ile Thr Asn Val Leu Thr His Arg 340 345 350 Ala Tyr Val Asp Gly Thr Leu Phe Tyr Tyr His Ala Glu Ser Phe Leu 355 360 365 Tyr Phe Leu Ser Arg Leu Phe Cys Glu Asn Ser Ser Val Gln Ser Arg 370 375 380 Phe Gln Glu Leu Leu Glu Gln Arg Leu Arg Glu Arg Ile Gly Thr Pro 385 390 395 400 Gly Asp Ala Leu Ser Leu Ala Met Arg Ala Leu Ala Cys Gln Met Leu 405 410 415 Gly Ile Asp Ser Ser Thr Asp Val Arg Ala Leu Leu Pro Leu Gln Cys 420 425 430 Asp Asp Gly Gly Trp Glu Ala Gly Trp Val Cys Arg Tyr Gly Ser Asn 435 440 445 Gly Met Arg Val Gly Ser Arg Gly Tyr Thr Thr Ala Leu Ala Ile Asn 450 455 460 Ala Ile Arg Gly Ala Arg Val Ser Arg 465 470 <210> 7 <211> 1452 <212> DNA <213> Bazzania sp. <220> <221> misc_feature <223> BazzHAD3_wt_nucleic acid sequence <400> 7 atgcctgctt taagctcgca tatggagcca cctaattttg acaccctgat cctggacctg 60 ggagatgtcc ttgtgggagg ctctcttcaa ggctcttacg ccttggccat gaagagaatt 120 ctgaaactct cattacgtac cccgacgtgg ggaaaatacg agtgcggcca attatcagaa 180 ttagattgct atgacgctgt ggctgaagat cttgggaact ccacaagctc cgctcaagtg 240 ggtgaggtcg tctcagaagc tcgaaagtca ttgcaggtca atgaagacct gataaaggtg 300 cttcaggagc taaaggctga gaacaatctc cgggtttatg tcatgtccaa tatacccaag 360 ccggatttgg ctgttgtgaa ggccaagtca gtgaactggg gagtcatcga tggctggtat 420 ccccttatg ctgtgggctt ccacaagccg gatctcgctt tctaccggca tgtacttgaa 480 gagacaaaca ccgatcctct taaggccatc ttcgttgatg acaaggttca gaatgtcatt 540 gttgccagat cctttggaat gacaggcatt gtcttcaagg acactaagag cactaccaag 600 gagctgaaga atctccttgg tgaccccgtc aagagagggg agcacttcct ctcagctcat 660 gccaaacaat tggagtctgt ctgtggggat ggcaccactt tccttgacaa cttcgcccag 720 ctgctgattc ttcatgatac tggaaacagc gatcttgtgg ctcgtcagca ccatggcaga 780 acctggaact acttcattgg caagccgatg gggacgacag acacattccc caacgaccta 840 gataccacct ccatcgctct gctgactctg aatgtagact cagaagtagc gaggtctgtg 900 ttagatgaga tgctgctgta caccagcagt gatgggctgg tagaggtgta ttttgacaaaa 960 actagacctc gggtggatcc cgttgtctgt gtcaacgtgc tacgtctctt cctaaatat 1020 gggagggagc tccagcttca gaagactctg gactgggtaa ctgaggtcct tgttcaccga 1080 gcttacattg acggcacact cttctactac cacgctgaga gcttcttgta cttcctttcg 1140 tgtctctaca aagagaaccc ccggctgcag actcaatttc gggagcctct acaagaacgc 1200 ctgcgggaga ggattgggca gcccggggat gcactttgtc tggctatgcg tgcaatcgca 1260 tgccagactg tggggattgg gaatgttatt gatgtgcaag ctcttctgcc attgcaatct 1320 agtgatggtg gttgggaggc tggctgggtc tgcagaatgg gcacgtctgg tgttcctgtt 1380 ggaaaccgag gagtcacaac ggcactttgt attagagcta ttcgaggcgc atcagggagc 1440 taccacactt ga 1452 <210> 8 <211> 1452 <212> DNA <213> Bazzania sp. <220> <221> misc_feature <223> BazzHAD3_Escherichia coli, optimized <400> 8 atgccggcgc tgagcagcca catggaaccg ccgaacttcg acaccctgat tctggacctg 60 ggtgatgtgc tggttggtgg cagcctgcag ggtagctacg cgctggcgat gaagcgtatc 120 ctgaaactga gcctgcgtac cccgacctgg ggcaagtacg aatgcggcca gctgagcgag 180 ctggactgct atgatgcggt tgcggaagac ctgggtaaca gcaccagcag cgcgcaagtg 240 ggcgaggtgg ttagcgaagc gcgtaaaagc ctgcaggtta acgaggatct gatcaaggtg 300 ctgcaagagc tgaaagcgga aaacaacctg cgtgtgtacg ttatgagcaa cattccgaag 360 ccggacctgg cggtggttaa ggcgaaaagc gtgaactggg gtgttatcga tggctggtat 420 ccgagctatg cggttggttt ccacaaaccg gacctggcgt tttatcgtca cgtgctggag 480 gaaaccaaca ccgatccgct gaaggcgatc ttcgtggacg ataaagttca gaacgtgatt 540 gttgcgcgta gcttcggtat gaccggcatc gtttttaagg acaccaaaag caccaccaag 600 gaactgaaaa acctgctggg tgatccggtt aagcgtggcg aacactttct gagcgcgcac 660 gcgaaacaac tggagagcgt gtgcggtgac ggcaccacct tcctggataa ctttgcgcag 720 ctgctgattc tgcacgacac cggtaacagc gatctggtgg cgcgtcaaca ccacggccgt 780 acctggaact acttcattgg caagccgatg ggcaccaccg atacctttcc gaacgacctg 840 gataccacca gcatcgcgct gctgaccctg aacgtggaca gcgaagttgc gcgtagcgtg 900 ctggatgaga tgctgctgta caccagcagc gacggtctgg ttgaggtgta tttcgacaaa 960 acccgtccgc gtgttgatcc ggtggtttgc gtgaacgttc tgcgtctgtt tagcaagtac ggtcgtgac tgcagctgca aaaaaccctg gattgggtga ccgaggtgct ggttcaccgt gcgtatatcg acggcaccct gttctactat cacgcggaaa gcttcctgta ctttctgagc tgcctgtata aggagaaccc gcgtctgcag acccaatttc gtgagccgct gcaggaacgt ctgcgtgagc gtattggtca gccgggtgat gcgctgtgcc tggcgatgcg tgcgatcgcg 1260 tgccagaccg ttggtatcgg caacgtgatt gatgttcaag cgctgctgcc gctgcagagc agcgatggtg gctgggaggc gggttgggtt tgccgtatgg gcaccagcgg tgtgccggtt 1380 ggtaaccgtg gcgtgaccac cgcgctgtgc atccgtgcga ttcgtggtgc gagcggcagc 1440 tatcacacct aa <210> 9 <211> 483 <212> PRT <213> Bazzania (Bazzania sp.) <220> <221> MISC_FEATURE <223> BazzHAD3_Frequently <400> 9 Met Pro Ala Leu Ser Ser His Met Glu Pro Pro Asn Phe Asp Thr Leu 1 5 10 15 Ile Leu Asp Leu Gly Asp Val Leu Val Gly Gly Ser Leu Gln Gly Ser 20 25 30 Tyr Ala Leu Ala Met Lys Arg Ile Leu Lys Leu Ser Leu Arg Thr Pro 35 40 45 Thr Trp Gly Lys Tyr Glu Cys Gly Gln Leu Ser Glu Leu Asp Cys Tyr 50 55 60 Asp Ala Val Ala Glu Asp Leu Gly Asn Ser Thr Ser Ser Ala Gln Val 65 70 75 80 Gly Glu Val Val Ser Glu Ala Arg Lys Ser Leu Gln Val Asn Glu Asp 85 90 95 Leu Ile Lys Val Leu Gln Glu Leu Lys Ala Glu Asn Asn Leu Arg Val 100 105 110 Tyr Val Met Ser Asn Ile Pro Lys Pro Asp Leu Ala Val Val Lys Ala 115 120 125 Lys Ser Val Asn Trp Gly Val Ile Asp Gly Trp Tyr Pro Ser Tyr Ala 130 135 140 Val Gly Phe His Lys Pro Asp Leu Ala Phe Tyr Arg His Val Leu Glu 145 150 155 160 Glu Thr Asn Thr Asp Pro Leu Lys Ala Ile Phe Val Asp Asp Lys Val 165 170 175 Gln Asn Val Ile Val Ala Arg Ser Phe Gly Met Thr Gly Ile Val Phe 180 185 190 Lys Asp Thr Lys Ser Thr Thr Lys Glu Leu Lys Asn Leu Leu Gly Asp 195 200 205 Pro Val Lys Arg Gly Glu His Phe Leu Ser Ala His Ala Lys Gln Leu 210 215 220 Glu Ser Val Cys Gly Asp Gly Thr Thr Phe Leu Asp Asn Phe Ala Gln 225 230 235 240 Leu Leu Ile Leu His Asp Thr Gly Asn Ser Asp Leu Val Ala Arg Gln 245 250 255 His His Gly Arg Thr Trp Asn Tyr Phe Ile Gly Lys Pro Met Gly Thr 260 265 270 Thr Asp Thr Phe Pro Asn Asp Leu Asp Thr Thr Ser Ile Ala Leu Leu 275 280 285 Thr Leu Asn Val Asp Ser Glu Val Ala Arg Ser Val Leu Asp Glu Met 290 295 300 Leu Leu Tyr Thr Ser Ser Asp Gly Leu Val Glu Val Tyr Phe Asp Lys 305 310 315 320 Thr Arg Pro Arg Val Asp Pro Val Val Cys Val Asn Val Leu Arg Leu 325 330 335 Phe Ser Lys Tyr Gly Arg Glu Leu Gln Leu Gln Lys Thr Leu Asp Trp 340 345 350 Val Thr Glu Val Leu Val His Arg Ala Tyr Ile Asp Gly Thr Leu Phe 355 360 365 Tyr Tyr His Ala Glu Ser Phe Leu Tyr Phe Leu Ser Cys Leu Tyr Lys 370 375 380 Glu Asn Pro Arg Leu Gln Thr Gln Phe Arg Glu Pro Leu Gln Glu Arg 385 390 395 400 Leu Arg Glu Arg Ile Gly Gln Pro Gly Asp Ala Leu Cys Leu Ala Met 405 410 415 Arg Ala Ile Ala Cys Gln Thr Val Gly Ile Gly Asn Val Ile Asp Val 420 425 430 Gln Ala Leu Leu Pro Leu Gln Ser Ser Asp Gly Gly Trp Glu Ala Gly 435 440 445 Trp Val Cys Arg Met Gly Thr Ser Gly Val Pro Val Gly Asn Arg Gly 450 455 460 Val Thr Thr Ala Leu Cys Ile Arg Ala Ile Arg Gly Ala Ser Gly Ser 465 470 475 480 Tyr His Thr <210> 10 <211> 1596 <212> DNA <213> Bazzania trilobata <220> <221> misc_feature <223> Nucleotide sequence of BtHAD_wt <400> 10 atggctcctc ctttcggata tattggtggg aatgaatcga gtgtgacgac gtccccagac 60 gtcctcgact acagctatga tggtggccct cctactggcc aatacaagag tatgatttta 120 gacattggcg gagtgttgct tcagcccaac cttgacagga gctatggcac gattgtacca 180 cctgtgcaat tcaaacgtat gttgcgctcg cacatctggt tcaattactc aatgggaaag 240 gtgtctgcgg aggaagtgtt tcggcagctc tccgtgaagt ctggctaccc ggcagaagat 300 atccggagtt ttgctcgaaa agcgagagag tctcttgtgc ccatcacaca aataacggat 360 cttctggtca aactcagcaa ggaatcgaag atcagactgt tctgtatgac caattgccct 420 gccgaagact tcgcatacct ctcacaggca tatcccaacc ttttcggatt gttcgagagg 480 atttttacgt cggcgtccac cggtatgcgc aagccaaaca ggaatttctt ccgacatgta 540 ctccaggaga cgggaatcag tgcgacggag actattttcg tagacgacct ggtaccgaac 600 atcatggccg ccgaggcgct ggacttcacg gcaattcacg tttctccaga gtcgcccgcc 660 gagaccgtga ggttgctgat gttccatctc cgtaagccag aagagcggct ggcggcggcg 720 cgcgactatc tcctgcgcaa ccgtggagtg gacaaagctc tctctctcac atctgacgga 780 gtccgagtcc aggtttattt cgaccacttc gtggtcgccg aagtgctcaa cgacgagcga 840 ttcttgccaa tcactgcagc tccgaagtcc gggactatga actttttccc gaaggaggaa 900 cgcaacgcga ctcgagatat gacgaccaat gcagtcacac acgtggacct tcccgatgat 960 cttgactcca cctctgtcgc cttatcggtc ctttacaaat tctccaaagt ggcgatggag 1020 acggttcaga gggtcctgga tatgatggag aaacgtgtgg acgatgacgg catttttcaa 1080 acgtacttcg atcctttgcg gctgcgtgtt gacccggtgg ccgctactaa caccgtctac 1140 ctcttccatt taggaggtcg tcccggacct actcagcaga cggaacaatt cctctccaag 1200 ttactggaaa agagatctta cgtccgcggc acactctact acaggacccc cgaggcattt 1260 ctctatcaac tcaccagatt ggtggtatca tttgagtatt tcagaaagac aggcttcctg 1320 gaaatgctca agaaaaggct gagtgagcga gtgggcatcg agagcgatgc attcacgttg 1380 gcgatgagga ttctagcgtg tgtgaattgt ggaatccctg ttgctgagct gaacagagat 1440 atggacagac tagcacggat gcaaaatgtg gatggttcat gggacgtctg cccttactac 1500 agctacaatg atcccaagag ttggtttggc aacgagcttc ttaccacagc gtttgctgct 1560 gccgctcttg agcattctga gcccaacttt aagtga 1596 <210> 11 <211> 1596 <212> DNA <213> Bazzania trilobata <220> <221> misc_feature <223> BtHAD_Escherichia coli, optimized <400> 11 atggctccgc cgttcggtta cattggtggc aacgagagca gcgttaccac cagcccggat 60 gtgctggact acagctatga tggtggcccg ccgaccggcc agtacaaaag catgatcctg 120 gatattggtg gcgttctgct gcaaccgaac ctggaccgta gctatggcac catcgttccg 180 ccggtgcagt tcaagcgtat gctgcgtagc cacatttggt ttaactacag catgggtaaa 240 gttagcgcgg aggaagtgtt ccgtcaactg agcgtgaaga gcggctatcc ggcggaggat 300 atccgtagct ttgcgcgtaa agcgcgtgaa agcctggttc cgatcaccca gattaccgac 360 ctgctggtga agctgagcaa agagagcaag attcgtctgt tctgcatgac caactgcccg 420 gcggaagact ttgcgtacct gagccaagcg tatccgaacc tgttcggtct gtttgagcgt 480 atcttcacca gcgcgagcac cggcatgcgt aagccgaacc gtaacttctt tcgtcacgtt 540 ctgcaggaga ccggcatcag cgcgaccgaa accattttcg ttgacgatct ggtgccgaac 600 atcatggcgg cggaagcgct ggattttacc gcgattcacg ttagcccgga gagcccggcg 660 gaaaccgtgc gtctgctgat gtttcacctg cgtaaaccgg aggaacgtct ggcggcggcg 720 cgtgactacc tgctgcgtaa ccgtggtgtt gataaagcgc tgagcctgac cagcgatggt 780 gtgcgtgttc aagtgtattt cgatcacttt gtggttgcgg aagtgctgaa cgacgaacgt 840 ttcctgccga ttaccgcggc gccgaaaagc ggcaccatga acttctttcc gaaagaggag 900 cgtaacgcga cccgtgatat gaccaccaac gcggttaccc acgtggatct gccggacgat 960 ctggacagca ccagcgttgc gctgagcgtg ctgtacaaat ttagcaaggt tgcgatggag 1020 accgttcagc gtgtgctgga catgatggaa aaacgtgtgg acgatgacgg catcttccaa 1080 acctactttg atccgctgcg tctgcgtgtt gacccggtgg cggcgaccaa caccgtttac 1140 ctgttccatc tgggtggccg tccgggtccg acccagcaaa ccgagcagtt tctgagcaaa 1200 ctgctggaaa agcgtagcta tgtgcgtggc accctgtact atcgtacccc ggaagcgttc 1260 ctgtaccaac tgacccgtct ggtggttagc ttcgagtatt ttcgtaagac cggttttctg 1320 gaaatgctga agaaacgtct gagcgagcgt gtgggcatcg aaagcgatgc gttcaccctg 1380 gcgatgcgta tcctggcgtg cgttaactgc ggtattccgg tggcggagct gaaccgtgat 1440 atggaccgtc tggcgcgtat gcagaacgtt gatggtagct gggacgtgtg cccgtactat 1500 agctacaacg acccgaaaag ctggttcggc aacgaactgc tgaccaccgc gtttgcggcg 1560 gcggcgctgg agcacagcga accgaacttc aagtaa 1596 <210> 12 <211> 531 <212> PRT <213> Bazzania trilobata <220> <221> MISC_FEATURE <223> BtHAD_ amino acid sequence <400> 12 Met Ala Pro Pro Phe Gly Tyr Ile Gly Gly Asn Glu Ser Ser Val Thr 1 5 10 15 Thr Ser Pro Asp Val Leu Asp Tyr Ser Tyr Asp Gly Gly Pro Pro Thr 20 25 30 Gly Gln Tyr Lys Ser Met Ile Leu Asp Ile Gly Gly Val Leu Leu Gln 35 40 45 Pro Asn Leu Asp Arg Ser Tyr Gly Thr Ile Val Pro Pro Val Gln Phe 50 55 60 Lys Arg Met Leu Arg Ser His Ile Trp Phe Asn Tyr Ser Met Gly Lys 65 70 75 80 Val Ser Ala Glu Glu Val Phe Arg Gln Leu Ser Val Lys Ser Gly Tyr 85 90 95 Pro Ala Glu Asp Ile Arg Ser Phe Ala Arg Lys Ala Arg Glu Ser Leu 100 105 110 Val Pro Ile Thr Gln Ile Thr Asp Leu Leu Val Lys Leu Ser Lys Glu 115 120 125 Ser Lys Ile Arg Leu Phe Cys Met Thr Asn Cys Pro Ala Glu Asp Phe 130 135 140 Ala Tyr Leu Ser Gln Ala Tyr Pro Asn Leu Phe Gly Leu Phe Glu Arg 145 150 155 160 Ile Phe Thr Ser Ala Ser Thr Gly Met Arg Lys Pro Asn Arg Asn Phe 165 170 175 Phe Arg His Val Leu Gln Glu Thr Gly Ile Ser Ala Thr Glu Thr Ile 180 185 190 Phe Val Asp Asp Leu Val Pro Asn Ile Met Ala Ala Glu Ala Leu Asp 195 200 205 Phe Thr Ala Ile His Val Ser Pro Glu Ser Pro Ala Glu Thr Val Arg 210 215 220 Leu Leu Met Phe His Leu Arg Lys Pro Glu Glu Arg Leu Ala Ala Ala 225 230 235 240 Arg Asp Tyr Leu Leu Arg Asn Arg Gly Val Asp Lys Ala Leu Ser Leu 245 250 255 Thr Ser Asp Gly Val Arg Val Gln Val Tyr Phe Asp His Phe Val Val 260 265 270 Ala Glu Val Leu Asn Asp Glu Arg Phe Leu Pro Ile Thr Ala Ala Pro 275 280 285 Lys Ser Gly Thr Met Asn Phe Phe Pro Lys Glu Glu Arg Asn Ala Thr 290 295 300 Arg Asp Met Thr Thr Asn Ala Val Thr His Val Asp Leu Pro Asp Asp 305 310 315 320 Leu Asp Ser Thr Ser Val Ala Leu Ser Val Leu Tyr Lys Phe Ser Lys 325 330 335 Val Ala Met Glu Thr Val Gln Arg Val Leu Asp Met Met Glu Lys Arg 340 345 350 Val Asp Asp Asp Gly Ile Phe Gln Thr Tyr Phe Asp Pro Leu Arg Leu 355 360 365 Arg Val Asp Pro Val Ala Ala Thr Asn Thr Val Tyr Leu Phe His Leu 370 375 380 Gly Gly Arg Pro Gly Pro Thr Gln Gln Thr Glu Gln Phe Leu Ser Lys 385 390 395 400 Leu Leu Glu Lys Arg Ser Tyr Val Arg Gly Thr Leu Tyr Tyr Arg Thr 405 410 415 Pro Glu Ala Phe Leu Tyr Gln Leu Thr Arg Leu Val Val Ser Phe Glu 420 425 430 Tyr Phe Arg Lys Thr Gly Phe Leu Glu Met Leu Lys Lys Arg Leu Ser 435 440 445 Glu Arg Val Gly Ile Glu Ser Asp Ala Phe Thr Leu Ala Met Arg Ile 450 455 460 Leu Ala Cys Val Asn Cys Gly Ile Pro Val Ala Glu Leu Asn Arg Asp 465 470 475 480 Met Asp Arg Leu Ala Arg Met Gln Asn Val Asp Gly Ser Trp Asp Val 485 490 495 Cys Pro Tyr Tyr Ser Tyr Asn Asp Pro Lys Ser Trp Phe Gly Asn Glu 500 505 510 Leu Leu Thr Thr The Best Of Leu Glu His Ser Glu Pro 515 520 525 Asn Phe Lys 530 <210> 13 <211> 1596 <212> DNA <213> The snowstorm <220> <223> II based on the BtHAD based_insufficiency, specifically <400> 13 atggctccgc cgttcggtta cattggtggc aacgagagca gcgttaccac cagcccggat gtgctggact acagctga tggtggcccg ccgaccggcc agtacaaaag catgatcctg gatattggtg gcgttctgct gcaaccgac ctggaccgta gctatggcac catcgttccg ccggtgcagt tcaagcgtat gctgcgtagc cacatttggt ttaactacag catgggtaaa gttagcgcgg aggaagtgtt ccgtcaactg agcgtgaaga gcggctatcc ggcggaggat 300 atccgtagct ttgcgcgtaa agcgcgtgaa agcctggttc cgatcaccca gattaccgac 360 ctgctggtga agctgagcaa agagagcaag attcgtctgt tctgcatgac caactgcccg 420 gcggaagact ttgcgtacct gagccaagcg tatccgaacc tgttcggtct gtttgagcgt 480 atcttcacca gcgcgagcac cggcatgcgt aagccgaacc gtaacttctt tcgtcacgtt 540 ctgcaggaga ccggcatcag cgcgaccgaa accattttcg ttgacgatct ggtgccgaac 600 atcatggcgg cggaagcgct ggattttacc gcgattcacg ttagcccgga gagcccggcg 660 gaaaccgtgc gtctgctgat gtttcacctg cgtaaaccgg aggaacgtct ggcggcggcg 720 cgtgactacc tgctgcgtaa ccgtggtgtt gataaagcgc tgagcctgac cagcgatggt 780 gtgcgtgttc aagtgtattt cgatcacttt gtggttgcgg aagtgctgaa cgacgaacgt 840 ttcctgccga ttaccgcggc gccgaaaagc ggcaccatga acttctttcc gaaagaggag 900 cgtaacgcga cccgtgatat gaccaccaac gcggttaccc acgtggatct gccggacgcg 960 ctggcgagca ccagcgttgc gctgagcgtg ctgtacaaat ttagcaaggt tgcgatggag 1020 accgttcagc gtgtgctgga catgatgga aaacgtgtgg accgtgacgg catcttccaa acctactttg atccgctgcg tctgcgtgtt gacccggtgg cggcgaccaa caccgtttac 1140 ctgttccatc tgggtggccg tccgggtccg acccagcaaa ccgagcagtt tctgagcaaa ctgctggaaa agcgtagcta tgtgcgtggc accctgtact atcgtacccc ggaagcgttc ctgtaccaac tgacccgtct ggtggttagc ttcgagtatt ttcgtaagac cggttttctg gaaatgctga agaaacgtct gagcgagcgt gtgggcatcg aaagcgatgc gttcaccctg gcgatgcgta tcctggcgtg cgttaactgc ggtattccgg tggcggagct gaaccgtgat 1440 atggaccgtc tggcgcgtat gcagaacgtt gatggtagct gggacgtgtg cccgtactat agctacaacg acccgaaaag ctggttcggc aacgaactgc tgaccaccgc gtttgcggcg gcggcgctgg agcacagcga accgaacttc August <210> 14 <211> 531 <212> PRT <213> The snowstorm <220> <223> II of the BtHAD ingredients <400> 14 Met Ala Pro Pro Phe Gly Tyr Ile Gly Gly Asn Glu Ser Ser Val Thr 1 5 10 15 Thr Ser Pro Asp Val Leu Asp Tyr Ser Tyr Asp Gly Gly Pro Pro Thr 20 25 30 Gly Gln Tyr Lys Ser Met Ile Leu Asp Ile Gly Gly Val Leu Leu Gln 35 40 45 Pro Asn Leu Asp Arg Ser Tyr Gly Thr Ile Val Pro Pro Val Gln Phe 50 55 60 Lys Arg Met Leu Arg Ser His Ile Trp Phe Asn Tyr Ser Met Gly Lys 65 70 75 80 Val Ser Ala Glu Glu Val Phe Arg Gln Leu Ser Val Lys Ser Gly Tyr 85 90 95 Pro Ala Glu Asp Ile Arg Ser Phe Ala Arg Lys Ala Arg Glu Ser Leu 100 105 110 Val Pro Ile Thr Gln Ile Thr Asp Leu Leu Val Lys Leu Ser Lys Glu 115 120 125 Ser Lys Ile Arg Leu Phe Cys Met Thr Asn Cys Pro Ala Glu Asp Phe 130 135 140 Ala Tyr Leu Ser Gln Ala Tyr Pro Asn Leu Phe Gly Leu Phe Glu Arg 145 150 155 160 Ile Phe Thr Ser Ala Ser Thr Gly Met Arg Lys Pro Asn Arg Asn Phe 165 170 175 Phe Arg His Val Leu Gln Glu Thr Gly Ile Ser Ala Thr Glu Thr Ile 180 185 190 Phe Val Asp Asp Leu Val Pro Asn Ile Met Ala Ala Glu Ala Leu Asp 195 200 205 Phe Thr Ala Ile His Val Ser Pro Glu Ser Pro Ala Glu Thr Val Arg 210 215 220 Leu Leu Met Phe His Leu Arg Lys Pro Glu Glu Arg Leu Ala Ala Ala 225 230 235 240 Arg Asp Tyr Leu Leu Arg Asn Arg Gly Val Asp Lys Ala Leu Ser Leu 245 250 255 Thr Ser Asp Gly Val Arg Val Gln Val Tyr Phe Asp His Phe Val Val 260 265 270 Ala Glu Val Leu Asn Asp Glu Arg Phe Leu Pro Ile Thr Ala Ala Pro 275 280 285 Lys Ser Gly Thr Met Asn Phe Phe Pro Lys Glu Glu Arg Asn Ala Thr 290 295 300 Arg Asp Met Thr Thr Asn Ala Val Thr His Val Asp Leu Pro Asp Ala 305 310 315 320 Leu Ala Ser Thr Ser Val Ala Leu Ser Val Leu Tyr Lys Phe Ser Lys 325 330 335 Val Ala Met Glu Thr Val Gln Arg Val Leu Asp Met Met Glu Lys Arg 340 345 350 Val Asp Asp Asp Gly Ile Phe Gln Thr Tyr Phe Asp Pro Leu Arg Leu 355 360 365 Arg Val Asp Pro Val Ala Ala Thr Asn Thr Val Tyr Leu Phe His Leu 370 375 380 Gly Gly Arg Pro Gly Pro Thr Gln Gln Thr Glu Gln Phe Leu Ser Lys 385 390 395 400 Leu Leu Glu Lys Arg Ser Tyr Val Arg Gly Thr Leu Tyr Tyr Arg Thr 405 410 415 Pro Glu Ala Phe Leu Tyr Gln Leu Thr Arg Leu Val Val Ser Phe Glu 420 425 430 Tyr Phe Arg Lys Thr Gly Phe Leu Glu Met Leu Lys Lys Arg Leu Ser 435 440 445 Glu Arg Val Gly Ile Glu Ser Asp Ala Phe Thr Leu Ala Met Arg Ile 450 455 460 Leu Ala Cys Val Asn Cys Gly Ile Pro Val Ala Glu Leu Asn Arg Asp 465 470 475 480 Met Asp Arg Leu Ala Arg Met Gln Asn Val Asp Gly Ser Trp Asp Val 485 490 495 Cys Pro Tyr Tyr Ser Tyr Asn Asp Pro Lys Ser Trp Phe Gly Asn Glu 500 505 510 Leu Leu Thr Thr The Best Of Leu Glu His Ser Glu Pro 515 520 525 Asn Phe Lys 530 <210> 15 <211> 1596 <212> DNA <213> The snowstorm <220> <223> QW is also a BtHAD based_insufficient license plate <400> 15 atggctccgc cgttcggtta cattggtggc aacgagagca gcgttaccac cagcccggat gtgctggact acagctga tggtggcccg ccgaccggcc agtacaaaag catgatcctg gatattggtg gcgttctgct gcaaccgac ctggaccgta gctatggcac catcgttccg ccggtgcagt tcaagcgtat gctgcgtagc cacatttggt ttaactacag catgggtaaa gttagcgcgg aggaagtgtt ccgtcaactg agcgtgaaga gcggctatcc ggcggaggat 300 atccgtagct ttgcgcgtaa agcgcgtgaa agcctggttc cgatcaccca gattaccgac 360 ctgctggtga agctgagcaa agagagcaag attcgtctgt tctgcatgac caactgcccg 420 gcggaagact ttgcgtacct gagccaagcg tatccgaacc tgttcggtct gtttgagcgt 480 atcttcacca gcgcgagcac cggcatgcgt aagccgaacc gtaacttctt tcgtcacgtt 540 ctgcaggaga ccggcatcag cgcgaccgaa accattttcg ttgacgatct ggtgccgaac 600 atcatggcgg cggaagcgct ggattttacc gcgattcacg ttagcccgga gagcccggcg 660 gaaaccgtgc gtctgctgat gtttcacctg cgtaaaccgg aggaacgtct ggcggcggcg 720 cgtgactacc tgctgcgtaa ccgtggtgtt gataaagcgc tgagcctgac cagcgatggt 780 gtgcgtgttc aagtgtattt cgatcacttt gtggttgcgg aagtgctgaa cgacgaacgt 840 ttcctgccga ttaccgcggc gccgaaaagc ggcaccatga acttctttcc gaaagaggag 900 cgtaacgcga cccgtgatat gaccaccaac gcggttaccc acgtggatct gccggacgat 960 ctggacagca ccagcgttgc gctgagcgtg ctgtacaaat ttagcaaggt tgcgatggag 1020 accgttcagc gtgtgctgga catgatgga aaacgtgtgg accgtgacgg catcttccaa acctactttg atccgctgcg tctgcgtgtt gacccggtgg cggcgaccaa caccgtttac 1140 ctgttccatc tgggtggccg tccgggtccg acccagcaaa ccgagcagtt tctgagcaaa ctgctggaaa agcgtagcta tgtgcgtggc accctgtact atcgtacccc ggaagcgttc ctgtaccaac tgacccgtct ggtggttagc ttcgagtatt ttcgtaagac cggttttctg gaaatgctga agaaacgtct gagcgagcgt gtgggcatcg aaagcgatgc gttcaccctg gcgatgcgta tcctggcgtg cgttaactgc ggtattccgg tggcggagct gaaccgtgat 1440 atggaccgtc tggcgcgtat gcagaacgtt gcgggtagct gggacgtgtg cccgtactat agctacaacg acccgaaaag ctggttcggc aacgaactgc tgaccaccgc gtttgcggcg gcggcgctgg agcacagcga accgaacttc August <210> 16 <211> 531 <212> PRT <213> The snowstorm <220> <223> QW is also a BtHAD recipe <400> 16 Met Ala Pro Pro Phe Gly Tyr Ile Gly Gly Asn Glu Ser Ser Val Thr 1 5 10 15 Thr Ser Pro Asp Val Leu Asp Tyr Ser Tyr Asp Gly Gly Pro Pro Thr 20 25 30 Gly Gln Tyr Lys Ser Met Ile Leu Asp Ile Gly Gly Val Leu Leu Gln 35 40 45 Pro Asn Leu Asp Arg Ser Tyr Gly Thr Ile Val Pro Pro Val Gln Phe 50 55 60 Lys Arg Met Leu Arg Ser His Ile Trp Phe Asn Tyr Ser Met Gly Lys 65 70 75 80 Val Ser Ala Glu Glu Val Phe Arg Gln Leu Ser Val Lys Ser Gly Tyr 85 90 95 Pro Ala Glu Asp Ile Arg Ser Phe Ala Arg Lys Ala Arg Glu Ser Leu 100 105 110 Val Pro Ile Thr Gln Ile Thr Asp Leu Leu Val Lys Leu Ser Lys Glu 115 120 125 Ser Lys Ile Arg Leu Phe Cys Met Thr Asn Cys Pro Ala Glu Asp Phe 130 135 140 Ala Tyr Leu Ser Gln Ala Tyr Pro Asn Leu Phe Gly Leu Phe Glu Arg 145 150 155 160 Ile Phe Thr Ser Ala Ser Thr Gly Met Arg Lys Pro Asn Arg Asn Phe 165 170 175 Phe Arg His Val Leu Gln Glu Thr Gly Ile Ser Ala Thr Glu Thr Ile 180 185 190 Phe Val Asp Asp Leu Val Pro Asn Ile Met Ala Ala Glu Ala Leu Asp 195 200 205 Phe Thr Ala Ile His Val Ser Pro Glu Ser Pro Ala Glu Thr Val Arg 210 215 220 Leu Leu Met Phe His Leu Arg Lys Pro Glu Glu Arg Leu Ala Ala Ala 225 230 235 240 Arg Asp Tyr Leu Leu Arg Asn Arg Gly Val Asp Lys Ala Leu Ser Leu 245 250 255 Thr Ser Asp Gly Val Arg Val Gln Val Tyr Phe Asp His Phe Val Val 260 265 270 Ala Glu Val Leu Asn Asp Glu Arg Phe Leu Pro Ile Thr Ala Ala Pro 275 280 285 Lys Ser Gly Thr Met Asn Phe Phe Pro Lys Glu Glu Arg Asn Ala Thr 290 295 300 Arg Asp Met Thr Thr Asn Ala Val Thr His Val Asp Leu Pro Asp Asp 305 310 315 320 Leu Asp Ser Thr Ser Val Ala Leu Ser Val Leu Tyr Lys Phe Ser Lys 325 330 335 Val Ala Met Glu Thr Val Gln Arg Val Leu Asp Met Met Glu Lys Arg 340 345 350 Val Asp Asp Asp Gly Ile Phe Gln Thr Tyr Phe Asp Pro Leu Arg Leu 355 360 365 Arg Val Asp Pro Val Ala Ala Thr Asn Thr Val Tyr Leu Phe His Leu 370 375 380 Gly Gly Arg Pro Gly Pro Thr Gln Gln Thr Glu Gln Phe Leu Ser Lys 385 390 395 400 Leu Leu Glu Lys Arg Ser Tyr Val Arg Gly Thr Leu Tyr Tyr Arg Thr 405 410 415 Pro Glu Ala Phe Leu Tyr Gln Leu Thr Arg Leu Val Val Ser Phe Glu 420 425 430 Tyr Phe Arg Lys Thr Gly Phe Leu Glu Met Leu Lys Lys Arg Leu Ser 435 440 445 Glu Arg Val Gly Ile Glu Ser Asp Ala Phe Thr Leu Ala Met Arg Ile 450 455 460 Leu Ala Cys Val Asn Cys Gly Ile Pro Val Ala Glu Leu Asn Arg Asp 465 470 475 480 Met Asp Arg Leu Ala Arg Met Gln Asn Val Ala Gly Ser Trp Asp Val 485 490 495 Cys Pro Tyr Tyr Ser Tyr Asn Asp Pro Lys Ser Trp Phe Gly Asn Glu 500 505 510 Leu Leu Thr Thr Ala Phe Ala Ala Ala Ala Leu Glu His Ser Glu Pro 515 520 525 Asn Phe Lys 530 <210> 17 <211> 1410 <212> DNA <213> Selaginella moellendorffii <220> <221> misc_feature <223> SmHAD1_wt_nucleotide sequence <400> 17 atgataatca ttagctttgc ttgcttgctc aagtttggat gtggagacaa tggcccccgg 60 ggagatctcc tccgcagggc tctccagcac tcctccttct tggcctattc ttgcggggag 120 ctcgatcgca ctgccgcaat ctcgacaatc tcgcgcaaat tcaagctcac ggacccagct 180 ctcgtggatt ccatgctgct cgaggccgct tccagctgcg aggtggacaa ggagctactg 240 tcactgctcc agtcaacgag gcagcagctc ctgggatgga ttgatattcc tccacaagag 300 tgggagcgag tttacagcct cttcccggtg tcgctctgga agaacttcgc tacaatctcc 360 cgcgatctag attccctcct tggagatatc aggtcccatg ctgtgattgt cgataaaagc 420 gtggagatgg ccgcgctaca tgccttcgag agcttgctcg ctctcccgta tgcctcgccc 480 aacaagaaag cttgccagga cttcttgaag cgtcggttcc tgctgcctcg agctgtgaag 540 ctggggatgg atcgagtcaa ggaggaggtg agatccaagc agctcaaatc actctacctg 600 cttgacaatg gccggcacga gctcgtctcc gagatcttct ttccctgcgt ggcggcctgg 660 tgccttccgg agatcattcc tctcggctgg atgaaatatc tccaggtact gatcaagcgc 720 ggccatccct tcggttactt cggtgccgac gtctcgcgct atccgccgga catcgacacc 780 atgtcgactt gtgtctccac actttttgat cttagtctcg tcacctccgc gcaagctatg 840 cactttctgg agatttgtct cgagaacgtc aacgatcaaa accagcttct cacgtacctg 900 gacgtggagc ggcctcgagt ggatccggtc gtcatagcca acgttgtcta tctcgcctat 960 gcgctctcca tggaagatca cccggttgtc cgccacaacg agaatctgat ccagagatat 1020 ctgttaagcg gcggattcgt atacgggact cgctattacc tcagccagga ggacttcttg 1080 ttcatgtatg gccgggtcct cgcgacgttt ggagagaaga gggagattcc caactttgat 1140 ctcgtgtatc aagcgatgga ggcggctctc gtcaaccgga tcgggaacga gaccgagtcg 1200 aagccactgg acgttgccaa aaggattctg ctcagcaggg gttttggaat ccggaacacg 1260 atagatgtag acttgctcct caagatgcag aacgaggatg gttcctggcc tttgcaagtg 1320 ctcagcaatc ttccatctgc taaaggtggt gtcttcaatt cggtggtgga cttgagcttt 1380 gcggtgagag cactgcaaag tcaagattga 1410 <210> 18 <211> 1410 <212> DNA <213> Selaginella moellendorffii <220> <221> misc_feature <223> SmHAD1_Escherichia coli, optimized <400> 18 atgatcatta tcagcttcgc gtgcctgctg aagtttggtt gcggtgacaa cggtccgcgt 60 ggcgatctgc tgcgtcgtgc gctgcagcac agcagcttcc tggcgtatag ctgcggcgag 120 ctggaccgta ccgcggcgat tagcaccatc agccgtaagt ttaaactgac cgacccggcg 180 ctggttgata gcatgctgct ggaagcggcg agcagctgcg aggtggacaa agaactgctg 240 agcctgctgc agagcacccg tcagcaactg ctgggttgga ttgatatccc gccgcaagag 300 tgggaacgtg tgtacagcct gttcccggtt agcctgtgga agaactttgc gaccattagc 360 cgtgacctgg atagcctgct gggcgacatt cgtagccacg cggtgatcgt tgataaaagc 420 gttgagatgg cggcgctgca cgcgttcgaa agcctgctgg cgctgccgta tgcgagcccg 480 aacaagaaag cgtgccagga cttcctgaag cgtcgttttc tgctgccgcg tgcggtgaag 540 ctgggtatgg atcgtgtgaa agaggaagtt cgtagcaagc aactgaaaag cctgtacctg 600 ctggacaacg gccgtcacga gctggttagc gaaattttct ttccgtgcgt tgcggcgtgg 660 tgcctgccgg agattatccc gctgggttgg atgaagtatc tgcaggttct gattaaacgt 720 ggccacccgt tcggttactt tggtgcggat gtgagccgtt atccgccgga catcgatacc 780 atgagcacct gcgtgagcac cctgttcgat ctgagcctgg ttaccagcgc gcaagcgatg 840 cactttctgg agatctgcct ggaaaacgtt aacgaccaga accaactgct gacctacctg 900 gacgtggaac gtccgcgtgt tgatccggtg gttattgcga acgtggttta cctggcgtat 960 gcgctgagca tggaggatca cccggtggtt cgtcacaacg aaaacctgat ccagcgttac 1020 ctgctgagcg gtggcttcgt ttatggcacc cgttactatc tgagccaaga ggacttcctg 1080 tttatgtacg gtcgtgtgct ggcgaccttc ggcgagaaac gtgaaatccc gaactttgat 1140 ctggtgtatc aggcgatgga agcggcgctg gttaaccgta ttggtaacga gaccgaaagc 1200 aagccgctgg acgttgcgaa acgtatcctg ctgagccgtg gttttggcat tcgtaacacc 1260 atcgacgtgg atctgctgct gaagatgcag aacgaggatg gcagctggcc gctgcaagtg 1320 ctgagcaacc tgccgagcgc gaaaggtggc gttttcaaca gcgtggttga cctgagcttt 1380 gcggtgcgtg cgctgcagag ccaagattaa 1410 <210> 19 <211> 469 <212> PRT <213> Selaginella moellendorffii <220> <221> MISC_FEATURE <223> Amino acid sequence of SmHAD1 <400> 19 Met Ile Ile Ile Ser Phe Ala Cys Leu Leu Lys Phe Gly Cys Gly Asp 1 5 10 15 Asn Gly Pro Arg Gly Asp Leu Leu Arg Arg Ala Leu Gln His Ser Ser 20 25 30 Phe Leu Ala Tyr Ser Cys Gly Glu Leu Asp Arg Thr Ala Ala Ile Ser 35 40 45 Thr Ile Ser Arg Lys Phe Lys Leu Thr Asp Pro Ala Leu Val Asp Ser 50 55 60 Met Leu Leu Glu Ala Ala Ser Ser Cys Glu Val Asp Lys Glu Leu Leu 65 70 75 80 Ser Leu Leu Gln Ser Thr Arg Gln Gln Leu Leu Gly Trp Ile Asp Ile 85 90 95 Pro Pro Gln Glu Trp Glu Arg Val Tyr Ser Leu Phe Pro Val Ser Leu 100 105 110 Trp Lys Asn Phe Ala Thr Ile Ser Arg Asp Leu Asp Ser Leu Leu Gly 115 120 125 Asp Ile Arg Ser His Ala Val Ile Val Asp Lys Ser Val Glu Met Ala 130 135 140 Ala Leu His Ala Phe Glu Ser Leu Leu Ala Leu Pro Tyr Ala Ser Pro 145 150 155 160 Asn Lys Lys Ala Cys Gln Asp Phe Leu Lys Arg Arg Phe Leu Leu Pro 165 170 175 Arg Ala Val Lys Leu Gly Met Asp Arg Val Lys Glu Glu Val Arg Ser 180 185 190 Lys Gln Leu Lys Ser Leu Tyr Leu Leu Asp Asn Gly Arg His Glu Leu 195 200 205 Val Ser Glu Ile Phe Phe Pro Cys Val Ala Ala Trp Cys Leu Pro Glu 210 215 220 Ile Ile Pro Leu Gly Trp Met Lys Tyr Leu Gln Val Leu Ile Lys Arg 225 230 235 240 Gly His Pro Phe Gly Tyr Phe Gly Ala Asp Val Ser Arg Tyr Pro Pro 245 250 255 Asp Ile Asp Thr Met Ser Thr Cys Val Ser Thr Leu Phe Asp Leu Ser 260 265 270 Leu Val Thr Ser Ala Gln Ala Met His Phe Leu Glu Ile Cys Leu Glu 275 280 285 Asn Val Asn Asp Gln Asn Gln Leu Leu Thr Tyr Leu Asp Val Glu Arg 290 295 300 Pro Arg Val Asp Pro Val Val Ile Ala Asn Val Val Tyr Leu Ala Tyr 305 310 315 320 Ala Leu Ser Met Glu Asp His Pro Val Val Arg His Asn Glu Asn Leu 325 330 335 Ile Gln Arg Tyr Leu Leu Ser Gly Gly Phe Val Tyr Gly Thr Arg Tyr 340 345 350 Tyr Leu Ser Gln Glu Asp Phe Leu Phe Met Tyr Gly Arg Val Leu Ala 355 360 365 Thr Phe Gly Glu Lys Arg Glu Ile Pro Asn Phe Asp Leu Val Tyr Gln 370 375 380 Ala Met Glu Ala Ala Leu Val Asn Arg Ile Gly Asn Glu Thr Glu Ser 385 390 395 400 Lys Pro Leu Asp Val Ala Lys Arg Ile Leu Leu Ser Arg Gly Phe Gly 405 410 415 Ile Arg Asn Thr Ile Asp Val Asp Leu Leu Leu Lys Met Gln Asn Glu 420 425 430 Asp Gly Ser Trp Pro Leu Gln Val Leu Ser Asn Leu Pro Ser Ala Lys 435 440 445 Gly Gly Val Phe Asn Ser Val Val Asp Leu Ser Phe Ala Val Arg Ala 450 455 460 Leu Gln Ser Gln Asp 465 <210> 20 <211> 1410 <212> DNA <213> Artificial Sequence <220> <223> SmHAD1 Mutant of Class II Motif_Escherichia coli, Optimized <400> 20 atgatcatta tcagcttcgc gtgcctgctg aagtttggtt gcggtgacaa cggtccgcgt 60 ggcgatctgc tgcgtcgtgc gctgcagcac agcagcttcc tggcgtatag ctgcggcgag 120 ctggaccgta ccgcggcgat tagcaccatc agccgtaagt ttaaactgac cgacccggcg 180 ctggttgata gcatgctgct ggaagcggcg agcagctgcg aggtggacaa agaactgctg 240 agcctgctgc agagcacccg tcagcaactg ctgggttgga ttgatatccc gccgcaagag 300 tgggaacgtg tgtacagcct gttcccggtt agcctgtgga agaactttgc gaccattagc 360 cgtgacctgg atagcctgct gggcgacatt cgtagccacg cggtgatcgt tgataaaagc 420 gttgagatgg cggcgctgca cgcgttcgaa agcctgctgg cgctgccgta tgcgagcccg 480 aacaagaaag cgtgccagga cttcctgaag cgtcgttttc tgctgccgcg tgcggtgaag 540 ctgggtatgg atcgtgtgaa agaggaagtt cgtagcaagc aactgaaaag cctgtacctg 600 ctggacaacg gccgtcacga gctggttagc gaaattttct ttccgtgcgt tgcggcgtgg 660 tgcctgccgg agattatccc gctgggttgg atgaagtatc tgcaggttct gattaaacgt 720 ggccacccgt tcggttactt tggtgcggat gtgagccgtt atccgccggc gatcgcgacc 780 atgagcacct gcgtgagcac cctgttcgat ctgagcctgg ttaccagcgc gcaagcgatg 840 cactttctgg agatctgcct ggaaaacgtt aacgaccaga accaactgct gacctacctg 900 gacgtggaac gtccgcgtgt tgatccggtg gttattgcga acgtggttta cctggcgtat 960 gcgctgagca tggaggatca cccggtggtt cgtcacaacg aaaacctgat ccagcgttac 1020 ctgctgagcg gtggcttcgt ttatggcacc cgttactatc tgagccaaga ggacttcctg 1080 tttatgtacg gtcgtgtgct ggcgaccttc ggcgagaaac gtgaaatccc gaactttgat 1140 ctggtgtatc aggcgatgga agcggcgctg gttaaccgta ttggtaacga gaccgaaagc 1200 aagccgctgg acgttgcgaa acgtatcctg ctgagccgtg gttttggcat tcgtaacacc 1260 atcgacgtgg atctgctgct gaagatgcag aacgaggatg gcagctggcc gctgcaagtg 1320 ctgagcaacc tgccgagcgc gaaaggtggc gttttcaaca gcgtggttga cctgagcttt 1380 gcggtgcgtg cgctgcagag ccaagattaa 1410 <210> twenty one <211> 469 <212> PRT <213> Artificial sequence <220> <223> SmHAD1 mutant with class II motif <400> twenty one Met Ile Ile Ile Ser Phe Ala Cys Leu Leu Lys Phe Gly Cys Gly Asp 1 5 10 15 Asn Gly Pro Arg Gly Asp Leu Leu Arg Arg Ala Leu Gln His Ser Ser 20 25 30 Phe Leu Ala Tyr Ser Cys Gly Glu Leu Asp Arg Thr Ala Ala Ile Ser 35 40 45 Thr Ile Ser Arg Lys Phe Lys Leu Thr Asp Pro Ala Leu Val Asp Ser 50 55 60 Met Leu Leu Glu Ala Ala Ser Ser Cys Glu Val Asp Lys Glu Leu Leu 65 70 75 80 Ser Leu Leu Gln Ser Thr Arg Gln Gln Leu Leu Gly Trp Ile Asp Ile 85 90 95 Pro Pro Gln Glu Trp Glu Arg Val Tyr Ser Leu Phe Pro Val Ser Leu 100 105 110 Trp Lys Asn Phe Ala Thr Ile Ser Arg Asp Leu Asp Ser Leu Leu Gly 115 120 125 Asp Ile Arg Ser His Ala Val Ile Val Asp Lys Ser Val Glu Met Ala 130 135 140 Ala Leu His Ala Phe Glu Ser Leu Leu Ala Leu Pro Tyr Ala Ser Pro 145 150 155 160 Asn Lys Lys Ala Cys Gln Asp Phe Leu Lys Arg Arg Phe Leu Leu Pro 165 170 175 Arg Ala Val Lys Leu Gly Met Asp Arg Val Lys Glu Glu Val Arg Ser 180 185 190 Lys Gln Leu Lys Ser Leu Tyr Leu Leu Asp Asn Gly Arg His Glu Leu 195 200 205 Val Ser Glu Ile Phe Phe Pro Cys Val Ala Ala Trp Cys Leu Pro Glu 210 215 220 Ile Ile Pro Leu Gly Trp Met Lys Tyr Leu Gln Val Leu Ile Lys Arg 225 230 235 240 Gly His Pro Phe Gly Tyr Phe Gly Ala Asp Val Ser Arg Tyr Pro Pro 245 250 255 Ala Ile Ala Thr Met Ser Thr Cys Val Ser Thr Leu Phe Asp Leu Ser 260 265 270 Leu Val Thr Ser Ala Gln Ala Met His Phe Leu Glu Ile Cys Leu Glu 275 280 285 Asn Val Asn Asp Gln Asn Gln Leu Leu Thr Tyr Leu Asp Val Glu Arg 290 295 300 Pro Arg Val Asp Pro Val Val Ile Ala Asn Val Val Tyr Leu Ala Tyr 305 310 315 320 Ala Leu Ser Met Glu Asp His Pro Val Val Arg His Asn Glu Asn Leu 325 330 335 Ile Gln Arg Tyr Leu Leu Ser Gly Gly Phe Val Tyr Gly Thr Arg Tyr 340 345 350 Tyr Leu Ser Gln Glu Asp Phe Leu Phe Met Tyr Gly Arg Val Leu Ala 355 360 365 Thr Phe Gly Glu Lys Arg Glu Ile Pro Asn Phe Asp Leu Val Tyr Gln 370 375 380 Ala Met Glu Ala Ala Leu Val Asn Arg Ile Gly Asn Glu Thr Glu Ser 385 390 395 400 Lys Pro Leu Asp Val Ala Lys Arg Ile Leu Leu Ser Arg Gly Phe Gly 405 410 415 Ile Arg Asn Thr Ile Asp Val Asp Leu Leu Leu Lys Met Gln Asn Glu 420 425 430 Asp Gly Ser Trp Pro Leu Gln Val Leu Ser Asn Leu Pro Ser Ala Lys 435 440 445 Click Download to save Gly Gly Val Phe Asn Ser Val Val Asp Leu Ser Phe Ala Val Arg Ala mp3 youtube com 450 455 460 Leu Gln Ser Gln Asp 465 <210> 22 <211> 1410 <212> DNA <213> The snowstorm <220> <223> QW smell of SmHAD1 smell_surface, slightly <400> 22 atgatcatta tcagcttcgc gtgcctgctg aagtttggtt gcggtgacaa cggtccgcgt 60 ggcgatctgc tgcgtcgtgc gctgcagcac agcagcttcc tggcgtatag ctgcggcgag 120 ctggaccgta ccgcggcgat tagcaccatc agccgtaagt ttaaactgac cgacccggcg ctggttgata gcatgctgct ggaagcggcg agcagctgcg aggtggacaa agaactgctg 240 agcctgctgc agagcacccg tcagcaactg ctgggttgga ttgatatccc gccgcaagag 300 tgggaacgtg tgtacagcct gttcccggtt agcctgtgga agaactttgc gaccattagc 360 cgtgacctgg atagcctgct gggcgacatt cgtagccacg cggtgatcgt tgataaaagc 420 gttgagatgg cggcgctgca cgcgttcgaa agcctgctgg cgctgccgta tgcgagcccg 480 aacaagaaag cgtgccagga cttcctgaag cgtcgttttc tgctgccgcg tgcggtgaag 540 ctgggtatgg atcgtgtgaa agaggaagtt cgtagcaagc aactgaaaag cctgtacctg 600 ctggacaacg gccgtcacga gctggttagc gaaattttct ttccgtgcgt tgcggcgtgg 660 tgcctgccgg agattatccc gctgggttgg atgaagtatc tgcaggttct gattaaacgt 720 ggccacccgt tcggttactt tggtgcggat gtgagccgtt atccgccgga catcgatacc 780 atgagcacct gcgtgagcac cctgttcgat ctgagcctgg ttaccagcgc gcaagcgatg 840 cactttctgg agatctgcct ggaaaacgtt aacgaccaga accaactgct gacctacctg 900 gacgtggaac gtccgcgtgt tgatccggtg gttattgcga acgtggttta cctggcgtat 960 gcgctgagca tggaggatca cccggtggtt cgtcacaacg aaaacctgat ccagcgttac 1020 ctgctgagcg gtggcttcgt ttatggcacc cgttactatc tgagccaaga ggacttcctg 1080 tttatgtacg gtcgtgtgct ggcgaccttc ggcgagaaac gtgaaatccc gaactttgat 1140 ctggtgtatc aggcgatgga agcggcgctg gttaaccgta ttggtaacga gaccgaaagc 1200 aagccgctgg acgttgcgaa acgtatcctg ctgagccgtg gttttggcat tcgtaacacc 1260 atcgacgtgg atctgctgct gaagatgcag aacgcggcgg gcagctggcc gctgcaagtg 1320 ctgagcaacc tgccgagcgc gaaaggtggc gttttcaaca gcgtggttga cctgagcttt 1380 gcggtgcgtg cgctgcagag ccaagattaa 1410 <210> 23 <211> 469 <212> PRT <213> Artificial sequence <220> <223> SmHAD1 mutant with QW motif <400> 23 Met Ile Ile Ile Ser Phe Ala Cys Leu Leu Lys Phe Gly Cys Gly Asp 1 5 10 15 Asn Gly Pro Arg Gly Asp Leu Leu Arg Arg Ala Leu Gln His Ser Ser 20 25 30 Phe Leu Ala Tyr Ser Cys Gly Glu Leu Asp Arg Thr Ala Ala Ile Ser 35 40 45 Thr Ile Ser Arg Lys Phe Lys Leu Thr Asp Pro Ala Leu Val Asp Ser 50 55 60 Met Leu Leu Glu Ala Ala Ser Ser Cys Glu Val Asp Lys Glu Leu Leu 65 70 75 80 Ser Leu Leu Gln Ser Thr Arg Gln Gln Leu Leu Gly Trp Ile Asp Ile 85 90 95 Pro Pro Gln Glu Trp Glu Arg Val Tyr Ser Leu Phe Pro Val Ser Leu 100 105 110 Trp Lys Asn Phe Ala Thr Ile Ser Arg Asp Leu Asp Ser Leu Leu Gly 115 120 125 Asp Ile Arg Ser His Ala Val Ile Val Asp Lys Ser Val Glu Met Ala 130 135 140 Ala Leu His Ala Phe Glu Ser Leu Leu Ala Leu Pro Tyr Ala Ser Pro 145 150 155 160 Asn Lys Lys Ala Cys Gln Asp Phe Leu Lys Arg Arg Phe Leu Leu Pro 165 170 175 Arg Ala Val Lys Leu Gly Met Asp Arg Val Lys Glu Glu Val Arg Ser 180 185 190 Lys Gln Leu Lys Ser Leu Tyr Leu Leu Asp Asn Gly Arg His Glu Leu 195 200 205 Val Ser Glu Ile Phe Phe Pro Cys Val Ala Ala Trp Cys Leu Pro Glu 210 215 220 Ile Ile Pro Leu Gly Trp Met Lys Tyr Leu Gln Val Leu Ile Lys Arg 225 230 235 240 Gly His Pro Phe Gly Tyr Phe Gly Ala Asp Val Ser Arg Tyr Pro Pro 245 250 255 Asp Ile Asp Thr Met Ser Thr Cys Val Ser Thr Leu Phe Asp Leu Ser 260 265 270 Leu Val Thr Ser Ala Gln Ala Met His Phe Leu Glu Ile Cys Leu Glu 275 280 285 Asn Val Asn Asp Gln Asn Gln Leu Leu Thr Tyr Leu Asp Val Glu Arg 290 295 300 Pro Arg Val Asp Pro Val Val Ile Ala Asn Val Val Tyr Leu Ala Tyr 305 310 315 320 Ala Leu Ser Met Glu Asp His Pro Val Val Arg His Asn Glu Asn Leu 325 330 335 Ile Gln Arg Tyr Leu Leu Ser Gly Gly Phe Val Tyr Gly Thr Arg Tyr 340 345 350 Tyr Leu Ser Gln Glu Asp Phe Leu Phe Met Tyr Gly Arg Val Leu Ala 355 360 365 Thr Phe Gly Glu Lys Arg Glu Ile Pro Asn Phe Asp Leu Val Tyr Gln 370 375 380 Ala Met Glu Ala Ala Leu Val Asn Arg Ile Gly Asn Glu Thr Glu Ser 385 390 395 400 Lys Pro Leu Asp Val Ala Lys Arg Ile Leu Leu Ser Arg Gly Phe Gly 405 410 415 Ile Arg Asn Thr Ile Asp Val Asp Leu Leu Leu Lys Met Gln Asn Ala 420 425 430 Ala Gly Ser Trp Pro Leu Gln Val Leu Ser Asn Leu Pro Ser Ala Lys 435 440 445 Gly Gly Val Phe Asn Ser Val Val Asp Leu Ser Phe Ala Val Arg Ala 450 455 460 Leu Gln Ser Gln Asp 465 <210> 24 <211> 1410 <212> DNA <213> Selaginella moellendorffii <220> <221> misc_feature <223> SmHAD2_wt_nucleic acid sequence <400> 24 atgataatca ttagttttgc ttgcttgctt aagttcggat gtggagacag tggcccccgg 60 ggagatctcc tccgcagggc tctccagcac tcctccttct tggcctattc ttgcggggag 120 ctcgatcgcg ctgccgcaat ctcgacaatc tcgcgcaaat tcaagctcaa ggagccagct 180 ctcctggatt ccatgctgct cgaggccgct tccagctgcg aggtggacga ggagctactg 240 tcactgctcc agtcaacgag gcagcagctc ctgggatgga ttgatattcc tccacaagag 300 tgggagcgag tttacaacct cttcccgtgg tcgctctgga agaacttcgc tacaatctcc 360 cgcgatctag attccctcct tggagatatc aggttccatg ctgtgattgt cgataaaagc 420 gtggagatgg ccgcgctaca tgccttcgag agcttgctcg ctctcccgta tgcctcgccc 480 aacaagaaag cttgccagga cttcttgaag cgtcggttcc tgctgcctcg agctgtgaag 540 ctggggatgg atcgagtcaa ggaggaggtg agatccaagc agctcaaatc actctacctg 600 cttgacaacg gccggcagga gctcgtctcc gagatcttct ttccctgcgt ggcggcctgg 660 tgccttcctg agatcattcc tctcggctgg atggaatctc tccaggtact gatcgagcgc 720 ggccatccct tcggttactt cggtgccgac gtctcgcgct atccgccgga catcgacacc 780 atgtcgactt gtgtctccac actttttgat cttagcctcg tcacctccgc gcaagctatg 840 cactttctgg agatttgtct cgagaacgtc aacgatcaaa accagcttct cacgtacctg 900 gacttggagc ggcctcgagt ggatccggtc gtcatagcca acgttgtcta tttcgcctat 960 gcgctctcca tggaagatca cccggttgtc cgccacaacg agaatctgat ccagagatat 1020 ctgttaagcg gcggattcgt atacgggact cgctattacc tcagccagga ggacttcttg 1080 ttcatgtatg gccgggtcct cgcgacgttt ggagagaaga gggagattcc caactttgat 1140 ctcgtgtatc aagcgatgga ggcggctctc gtcaaccgga tcgggaacga gaccgagtcg 1200 aagccactgg acgttgccaa aaggattctg ctcagcaggt actttggaat ccggaacacg 1260 atagatgtgg acttgctcct caagatgcag aacgaggatg gttcctggcc cttgcaagtg 1320 ctcagcaatc ttccatctgc taaaggtggt gtcttcaatt cggtggtgga cttgagcttt 1380 gcggtgagag cactgcaaag tcaagattga 1410 <210> 25 <211> 1410 <212> DNA <213> Selaginella moellendorffii <220> <221> misc_feature <223> SmHAD2_Escherichia coli, optimized <400> 25 atgatcatta tcagcttcgc gtgcctgctg aaatttggtt gcggtgatag cggtccgcgt 60 ggcgatctgc tgcgtcgtgc gctgcagcac agcagcttcc tggcgtatag ctgcggcgag 120 ctggaccgtg cggcggcgat tagcaccatc agccgtaagt ttaaactgaa ggaaccggcg 180 ctgctggaca gcatgctgct ggaggcggcg agcagctgcg aagttgatga ggaactgctg 240 agcctgctgc agagcacccg tcagcaactg ctgggttgga ttgatatccc gccgcaagag 300 tgggaacgtg tgtataacct gttcccgtgg agcctgtgga aaaactttgc gaccattagc 360 cgtgacctgg atagcctgct gggcgacatt cgtttccacg cggtgatcgt tgataagagc 420 gttgagatgg cggcgctgca cgcgtttgaa agcctgctgg cgctgccgta cgcgagcccg 480 aacaagaaag cgtgccaaga cttcctgaag cgtcgttttc tgctgccgcg tgcggtgaaa 540 ctgggtatgg atcgtgtgaa ggaagaggtg cgcagcaaac agctgaagag cctgtatctg 600 ctggacaacg gccgtcaaga gctggttagc gaaattttct ttccgtgcgt tgcggcgtgg 660 tgcctgccgg agattatccc gctgggttgg atggagagcc tgcaggttct gattgaacgt 720 ggccacccgt tcggttactt tggtgcggat gtgagccgtt atccgccgga catcgatacc 780 atgagcacct gcgtgagcac cctgttcgat ctgagcctgg ttaccagcgc gcaagcgatg 840 cactttctgg agatctgcct ggaaaacgtt aacgaccaga accaactgct gacctacctg 900 gacctggaac gtccgcgtgt ggatccggtg gttattgcga acgtggttta cttcgcgtat 960 gcgctgagca tggaggatca cccggtggtt cgtcacaacg aaaacctgat ccagcgttac 1020 ctgctgagcg gtggctttgt ttatggcacc cgttactatc tgagccaaga ggacttcctg 1080 tttatgtacg gtcgtgtgct ggcgaccttc ggcgagaaac gtgaaatccc gaactttgat 1140 ctggtgtatc aggcgatgga agcggcgctg gttaaccgta ttggcaacga gaccgaaagc 1200 aaaccgctgg acgttgcgaa gcgtatcctg ctgagccgtt acttcggtat tcgtaacacc 1260 atcgacgtgg atctgctgct gaaaatgcag aacgaggatg gcagctggcc gctgcaagtg 1320 ctgagcaacc tgccgagcgc gaagggtggc gttttcaaca gcgtggttga cctgagcttt 1380 gcggtgcgtg cgctgcagag ccaagattaa 1410 <210> 26 <211> 469 <212> PRT <213> Selaginella moellendorffii <220> <221> MISC_FEATURE <223> Amino acid sequence of SmHAD2 <400> 26 Met Ile Ile Ile Ser Phe Ala Cys Leu Leu Lys Phe Gly Cys Gly Asp 1 5 10 15 Ser Gly Pro Arg Gly Asp Leu Leu Arg Arg Ala Leu Gln His Ser Ser 20 25 30 Phe Leu Ala Tyr Ser Cys Gly Glu Leu Asp Arg Ala Ala Ala Ile Ser 35 40 45 Thr Ile Ser Arg Lys Phe Lys Leu Lys Glu Pro Ala Leu Leu Asp Ser 50 55 60 Met Leu Leu Glu Ala Ala Ser Ser Cys Glu Val Asp Glu Glu Leu Leu 65 70 75 80 Ser Leu Leu Gln Ser Thr Arg Gln Gln Leu Leu Gly Trp Ile Asp Ile 85 90 95 Pro Pro Gln Glu Trp Glu Arg Val Tyr Asn Leu Phe Pro Trp Ser Leu 100 105 110 Trp Lys Asn Phe Ala Thr Ile Ser Arg Asp Leu Asp Ser Leu Leu Gly 115 120 125 Asp Ile Arg Phe His Ala Val Ile Val Asp Lys Ser Val Glu Met Ala 130 135 140 Ala Leu His Ala Phe Glu Ser Leu Leu Ala Leu Pro Tyr Ala Ser Pro 145 150 155 160 Asn Lys Lys Ala Cys Gln Asp Phe Leu Lys Arg Arg Phe Leu Leu Pro 165 170 175 Arg Ala Val Lys Leu Gly Met Asp Arg Val Lys Glu Glu Val Arg Ser 180 185 190 Lys Gln Leu Lys Ser Leu Tyr Leu Leu Asp Asn Gly Arg Gln Glu Leu 195 200 205 Val Ser Glu Ile Phe Phe Pro Cys Val Ala Ala Trp Cys Leu Pro Glu 210 215 220 Ile Ile Pro Leu Gly Trp Met Glu Ser Leu Gln Val Leu Ile Glu Arg 225 230 235 240 Gly His Pro Phe Gly Tyr Phe Gly Ala Asp Val Ser Arg Tyr Pro Pro 245 250 255 Asp Ile Asp Thr Met Ser Thr Cys Val Ser Thr Leu Phe Asp Leu Ser 260 265 270 Leu Val Thr Ser Ala Gln Ala Met His Phe Leu Glu Ile Cys Leu Glu 275 280 285 Asn Val Asn Asp Gln Asn Gln Leu Leu Thr Tyr Leu Asp Leu Glu Arg 290 295 300 Pro Arg Val Asp Pro Val Val Ile Ala Asn Val Val Tyr Phe Ala Tyr 305 310 315 320 Ala Leu Ser Met Glu Asp His Pro Val Val Arg His Asn Glu Asn Leu 325 330 335 Ile Gln Arg Tyr Leu Leu Ser Gly Gly Phe Val Tyr Gly Thr Arg Tyr 340 345 350 Tyr Leu Ser Gln Glu Asp Phe Leu Phe Met Tyr Gly Arg Val Leu Ala 355 360 365 Thr Phe Gly Glu Lys Arg Glu Ile Pro Asn Phe Asp Leu Val Tyr Gln 370 375 380 Ala Met Glu Ala Ala Leu Val Asn Arg Ile Gly Asn Glu Thr Glu Ser 385 390 395 400 Lys Pro Leu Asp Val Ala Lys Arg Ile Leu Leu Ser Arg Tyr Phe Gly 405 410 415 Ile Arg Asn Thr Ile Asp Val Asp Leu Leu Leu Lys Met Gln Asn Glu 420 425 430 Asp Gly Ser Trp Pro Leu Gln Val Leu Ser Asn Leu Pro Ser Ala Lys 435 440 445 Click Download to save Gly Gly Val Phe Asn Ser Val Val Asp Leu Ser Phe Ala Val Arg Ala mp3 youtube com 450 455 460 Leu Gln Ser Gln Asp 465 <210> 27 <211> 1410 <212> DNA <213> The snowstorm <220> <223> II anti-inflammatory SmHAD2 anti-inflammatory, anti-inflammatory <400> 27 atgatcatta tcagcttcgc gtgcctgctg aaatttggtt gcggtgatag cggtccgcgt ggcgatctgc tgcgtcgtgc gctgcagcac agcagcttcc tggcgtatag ctgcggcgag 120 ctggaccgtg cggcggcgat tagcaccatc agccgtaagt ttaaactgaa ggaaccggcg ctgctggaca gcatgctgct ggaggcggcg agcagctgcg aagttgatga ggaactgctg agcctgctgc agagcacccg tcagcaactg ctgggttgga ttgatatccc gccgcaagag 300 tgggaacgtg tgtataacct gttcccgtgg agcctgtgga aaaactttgc gaccattagc 360 cgtgacctgg atagcctgct gggcgacatt cgtttccacg cggtgatcgt tgataagagc 420 gttgagatgg cggcgctgca cgcgtttgaa agcctgctgg cgctgccgta cgcgagcccg 480 aacaagaaag cgtgccaaga cttcctgaag cgtcgttttc tgctgccgcg tgcggtgaaa 540 ctgggtatgg atcgtgtgaa ggaagaggtg cgcagcaaac agctgaagag cctgtatctg 600 ctggacaacg gccgtcaaga gctggttagc gaaattttct ttccgtgcgt tgcggcgtgg 660 tgcctgccgg agattatccc gctgggttgg atggagagcc tgcaggttct gattgaacgt 720 ggccacccgt tcggttactt tggtgcggat gtgagccgtt atccgccggc gatcgcgacc 780 atgagcacct gcgtgagcac cctgttcgat ctgagcctgg ttaccagcgc gcaagcgatg 840 cactttctgg agatctgcct ggaaaacgtt aacgaccaga accaactgct gacctacctg 900 gacctggaac gtccgcgtgt ggatccggtg gttattgcga acgtggttta cttcgcgtat 960 gcgctgagca tggaggatca cccggtggtt cgtcacaacg aaaacctgat ccagcgttac 1020 ctgctgagcg gtggctttgt ttatggcacc cgttactatc tgagccaaga ggacttcctg 1080 tttatgtacg gtcgtgtgct ggcgaccttc ggcgagaaac gtgaaatccc gaactttgat 1140 ctggtgtatc aggcgatgga agcggcgctg gttaaccgta ttggcaacga gaccgaaagc 1200 aaaccgctgg acgttgcgaa gcgtatcctg ctgagccgtt acttcggtat tcgtaacacc 1260 atcgacgtgg atctgctgct gaaaatgcag aacgaggatg gcagctggcc gctgcaagtg 1320 ctgagcaacc tgccgagcgc gaagggtggc gttttcaaca gcgtggttga cctgagcttt 1380 gcggtgcgtg cgctgcagag ccaagattaa 1410 <210> 28 <211> 469 <212> PRT <213> Artificial Sequence <220> <223> SmHAD2 Mutant of Class II Motif <400> 28 Met Ile Ile Ile Ser Phe Ala Cys Leu Leu Lys Phe Gly Cys Gly Asp 1 5 10 15 Ser Gly Pro Arg Gly Asp Leu Leu Arg Arg Ala Leu Gln His Ser Ser 20 25 30 Phe Leu Ala Tyr Ser Cys Gly Glu Leu Asp Arg Ala Ala Ala Ile Ser 35 40 45 Thr Ile Ser Arg Lys Phe Lys Leu Lys Glu Pro Ala Leu Leu Asp Ser 50 55 60 Met Leu Leu Glu Ala Ala Ser Ser Cys Glu Val Asp Glu Glu Leu Leu 65 70 75 80 Ser Leu Leu Gln Ser Thr Arg Gln Gln Leu Leu Gly Trp Ile Asp Ile 85 90 95 Pro Pro Gln Glu Trp Glu Arg Val Tyr Asn Leu Phe Pro Trp Ser Leu 100 105 110 Trp Lys Asn Phe Ala Thr Ile Ser Arg Asp Leu Asp Ser Leu Leu Gly 115 120 125 Asp Ile Arg Phe His Ala Val Ile Val Asp Lys Ser Val Glu Met Ala 130 135 140 Ala Leu His Ala Phe Glu Ser Leu Leu Ala Leu Pro Tyr Ala Ser Pro 145 150 155 160 Asn Lys Lys Ala Cys Gln Asp Phe Leu Lys Arg Arg Phe Leu Leu Pro 165 170 175 Arg Ala Val Lys Leu Gly Met Asp Arg Val Lys Glu Glu Val Arg Ser 180 185 190 Lys Gln Leu Lys Ser Leu Tyr Leu Leu Asp Asn Gly Arg Gln Glu Leu 195 200 205 Val Ser Glu Ile Phe Phe Pro Cys Val Ala Ala Trp Cys Leu Pro Glu 210 215 220 Ile Ile Pro Leu Gly Trp Met Glu Ser Leu Gln Val Leu Ile Glu Arg 225 230 235 240 Gly His Pro Phe Gly Tyr Phe Gly Ala Asp Val Ser Arg Tyr Pro Pro 245 250 255 Ala Ile Ala Thr Met Ser Thr Cys Val Ser Thr Leu Phe Asp Leu Ser 260 265 270 Leu Val Thr Ser Ala Gln Ala Met His Phe Leu Glu Ile Cys Leu Glu 275 280 285 Asn Val Asn Asp Gln Asn Gln Leu Leu Thr Tyr Leu Asp Leu Glu Arg 290 295 300 Pro Arg Val Asp Pro Val Val Ile Ala Asn Val Val Tyr Phe Ala Tyr 305 310 315 320 Ala Leu Ser Met Glu Asp His Pro Val Val Arg His Asn Glu Asn Leu 325 330 335 Ile Gln Arg Tyr Leu Leu Ser Gly Gly Phe Val Tyr Gly Thr Arg Tyr 340 345 350 Tyr Leu Ser Gln Glu Asp Phe Leu Phe Met Tyr Gly Arg Val Leu Ala 355 360 365 Thr Phe Gly Glu Lys Arg Glu Ile Pro Asn Phe Asp Leu Val Tyr Gln 370 375 380 Ala Met Glu Ala Ala Leu Val Asn Arg Ile Gly Asn Glu Thr Glu Ser 385 390 395 400 Lys Pro Leu Asp Val Ala Lys Arg Ile Leu Leu Ser Arg Tyr Phe Gly 405 410 415 Ile Arg Asn Thr Ile Asp Val Asp Leu Leu Leu Lys Met Gln Asn Glu 420 425 430 Asp Gly Ser Trp Pro Leu Gln Val Leu Ser Asn Leu Pro Ser Ala Lys 435 440 445 Gly Gly Val Phe Asn Ser Val Val Asp Leu Ser Phe Ala Val Arg Ala 450 455 460 Leu Gln Ser Gln Asp 465 <210> 29 <211> 1410 <212> DNA <213> Artificial sequence <220> <223> SmHAD2 mutant with QW motif, optimized <400> 29 atgatcatta tcagcttcgc gtgcctgctg aaatttggtt gcggtgatag cggtccgcgt 60 ggcgatctgc tgcgtcgtgc gctgcagcac agcagcttcc tggcgtatag ctgcggcgag 120 ctggaccgtg cggcggcgat tagcaccatc agccgtaagt ttaaactgaa ggaaccggcg 180 ctgctggaca gcatgctgct ggaggcggcg agcagctgcg aagttgatga ggaactgctg 240 agcctgctgc agagcacccg tcagcaactg ctgggttgga ttgatatccc gccgcaagag 300 tgggaacgtg tgtataacct gttcccgtgg agcctgtgga aaaactttgc gaccattagc 360 cgtgacctgg atagcctgct gggcgacatt cgtttccacg cggtgatcgt tgataagagc 420 gttgagatgg cggcgctgca cgcgtttgaa agcctgctgg cgctgccgta cgcgagcccg 480 aacaagaaag cgtgccaaga cttcctgaag cgtcgttttc tgctgccgcg tgcggtgaaa 540 ctgggtatgg atcgtgtgaa ggaagaggtg cgcagcaaac agctgaagag cctgtatctg 600 ctggacaacg gccgtcaaga gctggttagc gaaattttct ttccgtgcgt tgcggcgtgg 660 tgcctgccgg agattatccc gctgggttgg atggagagcc tgcaggttct gattgaacgt 720 ggccacccgt tcggttactt tggtgcggat gtgagccgtt atccgccgga catcgatacc 780 atgagcacct gcgtgagcac cctgttcgat ctgagcctgg ttaccagcgc gcaagcgatg 840 cactttctgg agatctgcct ggaaaacgtt aacgaccaga accaactgct gacctacctg 900 gacctggaac gtccgcgtgt ggatccggtg gttattgcga acgtggttta cttcgcgtat 960 gcgctgagca tggaggatca cccggtggtt cgtcacaacg aaaacctgat ccagcgttac 1020 ctgctgagcg gtggctttgt ttatggcacc cgttactatc tgagccaaga ggacttcctg 1080 tttatgtacg gtcgtgtgct ggcgaccttc ggcgagaaac gtgaaatccc gaactttgat 1140 ctggtgtatc aggcgatgga agcggcgctg gttaaccgta ttggcaacga gaccgaaagc 1200 aaaccgctgg acgttgcgaa gcgtatcctg ctgagccgtt acttcggtat tcgtaacacc 1260 atcgacgtgg atctgctgct gaaaatgcag aacgcggcgg gcagctggcc gctgcaagtg 1320 ctgagcaacc tgccgagcgc gaagggtggc gttttcaaca gcgtggttga cctgagcttt 1380 gcggtgcgtg cgctgcagag ccaagattaa 1410 <210> 30 <211> 469 <212> PRT <213> Artificial Sequence <220> <223> SmHAD2 Mutant with QW Motif <400> 30 Met Ile Ile Ile Ser Phe Ala Cys Leu Leu Lys Phe Gly Cys Gly Asp 1 5 10 15 Ser Gly Pro Arg Gly Asp Leu Leu Arg Arg Ala Leu Gln His Ser Ser 20 25 30 Phe Leu Ala Tyr Ser Cys Gly Glu Leu Asp Arg Ala Ala Ala Ile Ser 35 40 45 Thr Ile Ser Arg Lys Phe Lys Leu Lys Glu Pro Ala Leu Leu Asp Ser 50 55 60 Met Leu Leu Glu Ala Ala Ser Ser Cys Glu Val Asp Glu Glu Leu Leu 65 70 75 80 Ser Leu Leu Gln Ser Thr Arg Gln Gln Leu Leu Gly Trp Ile Asp Ile 85 90 95 Pro Pro Gln Glu Trp Glu Arg Val Tyr Asn Leu Phe Pro Trp Ser Leu 100 105 110 Trp Lys Asn Phe Ala Thr Ile Ser Arg Asp Leu Asp Ser Leu Leu Gly 115 120 125 Asp Ile Arg Phe His Ala Val Ile Val Asp Lys Ser Val Glu Met Ala 130 135 140 Ala Leu His Ala Phe Glu Ser Leu Leu Ala Leu Pro Tyr Ala Ser Pro 145 150 155 160 Asn Lys Lys Ala Cys Gln Asp Phe Leu Lys Arg Arg Phe Leu Leu Pro 165 170 175 Arg Ala Val Lys Leu Gly Met Asp Arg Val Lys Glu Glu Val Arg Ser 180 185 190 Lys Gln Leu Lys Ser Leu Tyr Leu Leu Asp Asn Gly Arg Gln Glu Leu 195 200 205 Val Ser Glu Ile Phe Phe Pro Cys Val Ala Ala Trp Cys Leu Pro Glu 210 215 220 Ile Ile Pro Leu Gly Trp Met Glu Ser Leu Gln Val Leu Ile Glu Arg 225 230 235 240 Gly His Pro Phe Gly Tyr Phe Gly Ala Asp Val Ser Arg Tyr Pro Pro 245 250 255 Asp Ile Asp Thr Met Ser Thr Cys Val Ser Thr Leu Phe Asp Leu Ser 260 265 270 Leu Val Thr Ser Ala Gln Ala Met His Phe Leu Glu Ile Cys Leu Glu 275 280 285 Asn Val Asn Asp Gln Asn Gln Leu Leu Thr Tyr Leu Asp Leu Glu Arg 290 295 300 Pro Arg Val Asp Pro Val Val Ile Ala Asn Val Val Tyr Phe Ala Tyr 305 310 315 320 Ala Leu Ser Met Glu Asp His Pro Val Val Arg His Asn Glu Asn Leu 325 330 335 Ile Gln Arg Tyr Leu Leu Ser Gly Gly Phe Val Tyr Gly Thr Arg Tyr 340 345 350 Tyr Leu Ser Gln Glu Asp Phe Leu Phe Met Tyr Gly Arg Val Leu Ala 355 360 365 Thr Phe Gly Glu Lys Arg Glu Ile Pro Asn Phe Asp Leu Val Tyr Gln 370 375 380 Ala Met Glu Ala Ala Leu Val Asn Arg Ile Gly Asn Glu Thr Glu Ser 385 390 395 400 Lys Pro Leu Asp Val Ala Lys Arg Ile Leu Leu Ser Arg Tyr Phe Gly 405 410 415 Ile Arg Asn Thr Ile Asp Val Asp Leu Leu Leu Lys Met Gln Asn Ala 420 425 430 Ala Gly Ser Trp Pro Leu Gln Val Leu Ser Asn Leu Pro Ser Ala Lys 435 440 445 Gly Gly Val Phe Asn Ser Val Val Asp Leu Ser Phe Ala Val Arg Ala 450 455 460 Leu Gln Ser Gln Asp 465 <210> 31 <211> 1602 <212> DNA <213> Gelatoria subvermispora <220> <221> misc_feature <223> EMD37666.1_Escherichia coli, optimized <400> 31 atgtctgcgg cggctcaata cacgactttg attctggatc tgggtgatgt tctgttcact 60 tggtccccga aaaccaagac cagcatccct ccgcgtaccc tgaaagaaat cctgaatagc 120 gctacctggt atgagtacga gcgtggtcgc atttcccaag acgagtgtta cgaacgtgtg 180 ggcaccgagt tcggcattgc gccgagcgag attgacaacg cgttcaaaca agcgcgcgat 240 tcgatggaaa gcaatgatga actgatcgca ctggtccgtg agctgaaaac gcagctggac 300 ggtgagctgc tggttttcgc actgtccaat attagcctgc cggattacga atacgtcttg 360 accaaaccgg cggactggag catctttgac aaagtgttcc ctagcgcctt ggtgggcgag 420 cgtaagccgc atctgggcgt ttataaacac gttattgcgg aaacgggcat tgatccgcgc 480 acgacggttt tcgtggacga caagattgac aatgtgttaa gcgcacgcag cgtcggtatg 540 catggtatcg tgtttgagaa acaagaagat gtcatgcgtg cactgcgtaa catctttggt 600 gatccggtcc gtcgtggtcg tgagtatctg cgtagaaacg caatgcgtct ggagtccgtg 660 accgaccacg gcgtggcgtt tggtgagaac tttacccagt tgctgattct ggaattgacg 720 aacgacccga gcctggtcac cctgcctgat cgtccgcgta cctggaactt ttttcgcggc 780 aatggtggcc gcccgagcaa gccgctgttc agcgaagcgt tcccggatga tctggatacc 840 acgagcctgg cgctgaccgt gctgcagcgc gacccgggtg ttatcagcag cgttatggac 900 gaaatgctga attaccgtga cccggacggt atcatgcaga cttatttcga tgacggtcgc 960 caacgcttgg acccatttgt gaacgtcaat gttctgacct ttttctatac gaacggccgt 1020 ggtcacgaac tggaccagtg tctgacgtgg gtgcgtgaag tcctcttgta tcgtgcgtac 1080 cttggtggct cacgctacta cccatcggcg gattgcttcc tgtacttcat ctctcgtctg 1140 tttgcgtgta ccaatgaccc ggtgctgcac catcagctga agccactgtt tgttgagcgt 1200 gtccaagagc aaattggtgt cgagggtgat gcactggaac tggcttttcg tctgctggtc 1260 tgcgccagcc tggatgtcca gaatgccatc gacatgcgcc gtctgctgga aatgcagtgc 1320 gaagatggcg gttgggaggg tggtaacctc taccgcttcg gcaccacggg cctgaaagtt 1380 accaaccgcg gtctgacgac cgcagccgcc gttcaagcga tcgaagcgag ccaacgccgt 1440 ccgccgagcc cgagcccgtc tgtagagagc acgaaaagcc cgattacccc ggtgaccccg 1500 atgctggaag ttccaagcct gggcttatct atcagccgtc cgtccagccc gctgctgggt 1560 tatttccgtt tgccgtggaa gaaaagcgca gaagtgcact aa 1602 <210> 32 <211> 533 <212> PRT <213> Gelatoporia subvermispora <220> <221> MISC_FEATURE <223> Amino acid sequence of EMD37666.1 <400> 32 Met Ser Ala Ala Ala Gln Tyr Thr Thr Leu Ile Leu Asp Leu Gly Asp 1 5 10 15 Val Leu Phe Thr Trp Ser Pro Lys Thr Lys Thr Ser Ile Pro Pro Arg 20 25 30 Thr Leu Lys Glu Ile Leu Asn Ser Ala Thr Trp Tyr Glu Tyr Glu Arg 35 40 45 Gly Arg Ile Ser Gln Asp Glu Cys Tyr Glu Arg Val Gly Thr Glu Phe 50 55 60 Gly Ile Ala Pro Ser Glu Ile Asp Asn Ala Phe Lys Gln Ala Arg Asp 65 70 75 80 Ser Met Glu Ser Asn Asp Glu Leu Ile Ala Leu Val Arg Glu Leu Lys 85 90 95 Thr Gln Leu Asp Gly Glu Leu Leu Val Phe Ala Leu Ser Asn Ile Ser 100 105 110 Leu Pro Asp Tyr Glu Tyr Val Leu Thr Lys Pro Ala Asp Trp Ser Ile 115 120 125 Phe Asp Lys Val Phe Pro Ser Ala Leu Val Gly Glu Arg Lys Pro His 130 135 140 Leu Gly Val Tyr Lys His Val Ile Ala Glu Thr Gly Ile Asp Pro Arg 145 150 155 160 Thr Thr Val Phe Val Asp Asp Lys Ile Asp Asn Val Leu Ser Ala Arg 165 170 175 Ser Val Gly Met His Gly Ile Val Phe Glu Lys Gln Glu Asp Val Met 180 185 190 Arg Ala Leu Arg Asn Ile Phe Gly Asp Pro Val Arg Arg Gly Arg Glu 195 200 205 Tyr Leu Arg Arg Asn Ala Met Arg Leu Glu Ser Val Thr Asp His Gly 210 215 220 Val Ala Phe Gly Glu Asn Phe Thr Gln Leu Leu Ile Leu Glu Leu Thr 225 230 235 240 Asn Asp Pro Ser Leu Val Thr Leu Pro Asp Arg Pro Arg Thr Trp Asn 245 250 255 Phe Phe Arg Gly Asn Gly Gly Arg Pro Ser Lys Pro Leu Phe Ser Glu 260 265 270 Ala Phe Pro Asp Asp Leu Asp Thr Thr Ser Leu Ala Leu Thr Val Leu 275 280 285 Gln Arg Asp Pro Gly Val Ile Ser Ser Val Met Asp Glu Met Leu Asn 290 295 300 Tyr Arg Asp Pro Asp Gly Ile Met Gln Thr Tyr Phe Asp Asp Gly Arg 305 310 315 320 Gln Arg Leu Asp Pro Phe Val Asn Val Asn Val Leu Thr Phe Phe Tyr 325 330 335 Thr Asn Gly Arg Gly His Glu Leu Asp Gln Cys Leu Thr Trp Val Arg 340 345 350 Glu Val Leu Leu Tyr Arg Ala Tyr Leu Gly Gly Ser Arg Tyr Tyr Pro 355 360 365 Ser Ala Asp Cys Phe Leu Tyr Phe Ile Ser Arg Leu Phe Ala Cys Thr 370 375 380 Asn Asp Pro Val Leu His His Gln Leu Lys Pro Leu Phe Val Glu Arg 385 390 395 400 Val Gln Glu Gln Ile Gly Val Glu Gly Asp Ala Leu Glu Leu Ala Phe 405 410 415 Arg Leu Leu Val Cys Ala Ser Leu Asp Val Gln Asn Ala Ile Asp Met 420 425 430 Arg Arg Leu Leu Glu Met Gln Cys Glu Asp Gly Gly Trp Glu Gly Gly 435 440 445 Asn Leu Tyr Arg Phe Gly Thr Thr Gly Leu Lys Val Thr Asn Arg Gly 450 455 460 Leu Thr Thr Ala Ala Ala Val Gln Ala Ile Glu Ala Ser Gln Arg Arg 465 470 475 480 Pro Pro Ser Pro Ser Pro Ser Val Glu Ser Thr Lys Ser Pro Ile Thr 485 490 495 Pro Val Thr Pro Met Leu Glu Val Pro Ser Leu Gly Leu Ser Ile Ser 500 505 510 Arg Pro Ser Ser Pro Leu Leu Gly Tyr Phe Arg Leu Pro Trp Lys Lys 515 520 525 Ser Ala Glu Val His 530 <210> 33 <211> 1602 <212> DNA <213> Synthetic Sequence <220> <223> Mutant of EMD37666.1 class II motif, Escherichia coli, optimized <400> 33 atgagcgcgg cggcgcagta caccaccctg atcctggacc tgggtgatgt gctgttcacc 60 tggagcccga agaccaaaac cagcatcccg ccgcgtaccc tgaaggaaat tctgaacagc 120 gcgacctggt acgagtatga acgtggccgt atcagccaag acgagtgcta tgaacgtgtt 180 ggcaccgagt tcggcatcgc gccgagcgaa attgataacg cgtttaagca ggcgcgtgac 240 agcatggaga gcaacgatga actgattgcg ctggtgcgtg agctgaaaac ccaactggac 300 ggtgaactgc tggttttcgc gctgagcaac atcagcctgc cggattacga gtatgtgctg 360 accaagccgg cggactggag cattttcgat aaagtttttc cgagcgcgct ggttggtgaa 420 cgcaagccgc acctgggcgt ttacaaacac gtgatcgcgg aaaccggtat tgacccgcgt 480 accaccgtgt ttgttgacga taagatcgat aacgttctga gcgcgcgtag cgtgggtatg 540 cacggcattg ttttcgagaa acaggaagac gtgatgcgtg cgctgcgtaa catctttggt 600 gatccggttc gtcgtggccg tgagtatctg cgtcgtaacg cgatgcgtct ggaaagcgtt 660 accgaccacg gtgtggcgtt cggcgagaac tttacccaac tgctgattct ggaactgacc 720 aacgatccga gcctggtgac cctgccggat cgtccgcgta cctggaactt ctttcgtggt 780 aacggtggcc gtccgagcaa gccgctgttc agcgaagcgt ttccggcggc gctggcgacc 840 accagcctgg cgctgaccgt tctgcagcgt gacccgggcg ttatcagcag cgtgatggat 900 gagatgctga actaccgtga cccggatggt attatgcaga cctatttcga cgatggccgt 960 caacgtctgg acccgtttgt gaacgttaac gtgctgacct tcttttacac caacggtcgt 1020 ggccacgagc tggatcagtg cctgacctgg gttcgtgaag tgctgctgta ccgtgcgtat 1080 ctgggtggca gccgttacta tccgagcgcg gactgcttcc tgtattttat cagccgtctg 1140 ttcgcgtgca ccaacgatcc ggtgctgcac caccaactga aaccgctgtt tgttgagcgt 1200 gtgcaggaac aaatcggtgt tgagggcgac gcgctggaac tggcgttccg tctgctggtg 1260 tgcgcgagcc tggacgttca gaacgcgatt gatatgcgtc gtctgctgga gatgcaatgc 1320 gaagatggtg gctgggaagg tggcaacctg taccgttttg gcaccaccgg cctgaaggtg 1380 accaaccgtg gtctgaccac cgcggcggcg gttcaggcga ttgaggcgag ccaacgtcgt 1440 ccgccgagcc cgagcccgag cgtggaaagc accaaaagcc cgatcacccc ggtgaccccg 1500 atgctggaag tgccgagcct gggtctgagc attagccgtc cgagcagccc gctgctgggt 1560 tacttccgtc tgccgtggaa gaaaagcgcg gaagtgcact aa 1602 <210> 34 <211> 533 <212> PRT <213> Artificial sequence <220> <223> Mutant of EMD37666.1 class II motif <400> 34 Met Ser Ala Ala Ala Gln Tyr Thr Thr Leu Ile Leu Asp Leu Gly Asp 1 5 10 15 Val Leu Phe Thr Trp Ser Pro Lys Thr Lys Thr Ser Ile Pro Pro Arg 20 25 30 Thr Leu Lys Glu Ile Leu Asn Ser Ala Thr Trp Tyr Glu Tyr Glu Arg 35 40 45 Gly Arg Ile Ser Gln Asp Glu Cys Tyr Glu Arg Val Gly Thr Glu Phe 50 55 60 Gly Ile Ala Pro Ser Glu Ile Asp Asn Ala Phe Lys Gln Ala Arg Asp 65 70 75 80 Ser Met Glu Ser Asn Asp Glu Leu Ile Ala Leu Val Arg Glu Leu Lys 85 90 95 Thr Gln Leu Asp Gly Glu Leu Leu Val Phe Ala Leu Ser Asn Ile Ser 100 105 110 Leu Pro Asp Tyr Glu Tyr Val Leu Thr Lys Pro Ala Asp Trp Ser Ile 115 120 125 Phe Asp Lys Val Phe Pro Ser Ala Leu Val Gly Glu Arg Lys Pro His 130 135 140 Leu Gly Val Tyr Lys His Val Ile Ala Glu Thr Gly Ile Asp Pro Arg 145 150 155 160 Thr Thr Val Phe Val Asp Asp Lys Ile Asp Asn Val Leu Ser Ala Arg 165 170 175 Ser Val Gly Met His Gly Ile Val Phe Glu Lys Gln Glu Asp Val Met 180 185 190 Arg Ala Leu Arg Asn Ile Phe Gly Asp Pro Val Arg Arg Gly Arg Glu 195 200 205 Tyr Leu Arg Arg Asn Ala Met Arg Leu Glu Ser Val Thr Asp His Gly 210 215 220 Val Ala Phe Gly Glu Asn Phe Thr Gln Leu Leu Ile Leu Glu Leu Thr 225 230 235 240 Asn Asp Pro Ser Leu Val Thr Leu Pro Asp Arg Pro Arg Thr Trp Asn 245 250 255 Phe Phe Arg Gly Asn Gly Gly Arg Pro Ser Lys Pro Leu Phe Ser Glu 260 265 270 Ala Phe Pro Ala Ala Leu Ala Thr Thr Ser Leu Ala Leu Thr Val Leu 275 280 285 Gln Arg Asp Pro Gly Val Ile Ser Ser Val Met Asp Glu Met Leu Asn 290 295 300 Tyr Arg Asp Pro Asp Gly Ile Met Gln Thr Tyr Phe Asp Asp Gly Arg 305 310 315 320 Gln Arg Leu Asp Pro Phe Val Asn Val Asn Val Leu Thr Phe Phe Tyr 325 330 335 Thr Asn Gly Arg Gly His Glu Leu Asp Gln Cys Leu Thr Trp Val Arg 340 345 350 Glu Val Leu Leu Tyr Arg Ala Tyr Leu Gly Gly Ser Arg Tyr Tyr Pro 355 360 365 Ser Ala Asp Cys Phe Leu Tyr Phe Ile Ser Arg Leu Phe Ala Cys Thr 370 375 380 Asn Asp Pro Val Leu His His Gln Leu Lys Pro Leu Phe Val Glu Arg 385 390 395 400 Val Gln Glu Gln Ile Gly Val Glu Gly Asp Free Mp3 Download 405 410 415 Arg Leu Leu Val Cys The Best Of Ala Ser Leu Asp Val Gln Asn Ala Ile Asp Met 420 425 430 Arg Arg Leu Leu Glu Met Gln Cys Glu Asp Gly Gly Trp Glu Gly Gly 435 440 445 Asn Leu Tyr Arg Phe Gly Thr Thr Gly Leu Lys Val Thr Asn Arg Gly 450 455 460 Leu Thr Thr White White White Val Gln White Ile Glu White Ser Gln Arg Arg 465 470 475 480 Pro Pro Ser Pro Ser Pro Ser Val Glu Ser Thr Lys Ser Pro Ile Thr 485 490 495 Pro Val Thr Pro Met Leu Glu Val Pro Ser Leu Gly Leu Ser Ile Ser 500 505 510 Arg Pro Ser Ser Pro Leu Leu Gly Tyr Phe Arg Leu Pro Trp Lys Lys 515 520 525 Be Wing Glu Val His 530 <210> 35 <211> 1584 <212> DNA <213> Squirrel squirrel (Dichomitus squalens) <220> <221> misc_feature <223> XP_007369631.1_Escherichia coli, optimized <400> 35 atggcgagca tccaccgtcg ttacaccacc ctgattctgg acctgggtga tgttctgttt 60 cgttggagcc cgaagaccga gaccgcgatc ccgccgcagc aactgaaaga cattctgagc 120 agcgtgacct ggttcgagta cgaacgtggc cgtctgagcc aggaagcgtg ctatgaacgt 180 tgcgcggagg aatttaagat cgaggcgagc gttattgcgg aagcgttcaa acaagcgcgt 240 ggtagcctgc gtccgaacga ggaatttatc gcgctgattc gtgacctgcg tcgtgagatg 300 cacggcgatc tgaccgtgct ggcgctgagc aacatcagcc tgccggacta cgaatatatt 360 atgagcctga gcagcgactg gaccaccgtt ttcgatcgtg tgtttccgag cgcgctggtt 420 ggtgagcgta agccgcacct gggctgctat cgtaaagtga tcagcgagat gaacctggaa 480 ccgcagacca ccgtgttcgt tgacgataag ctggataacg ttgcgagcgc gcgtagcctg 540 ggtatgcatg gtattgtttt cgacaaccag gcgaacgtgt ttcgtcaact gcgtaacatc 600 ttcggtgatc cgattcgtcg tggccaagag tacctgcgtg gtcacgcggg caaactggaa 660 agcagcaccg acaacggtct gatcttcgag gaaaacttta cccagctgat catttatgag ctgacccaag atcgtaccct gattagcctg agcgaatgcc cgcgtacctg gaacttcttt 780 cgtggcgagc cgctgttcag cgaaaccttc ccggacgacg ttgacaccac cagcgttgcg 840 ctgaccgtgc tgcagccgga ccgtgcgctg gttaacagcg tgctggatga gatgctgga tacgtggacg cggatggtat catgcaaacc tattttgacc gtagccgtcc gcgtatggat ccgtttgtgt gcgttaacgt gctgagcctg ttctacgaga acggtcgtgg ccacgaactg ccgcgtaccc tggactgggt ttacgaggtg ctgctgcacc gtgcgtatca cggtggcagc 1080 cgttactatc tgagcccgga ttgcttcctg ttctttatga gccgtctgct gaaacgtgcg gatgatccgg cggttcaggc gcgtctgcgt ccgctgtttg ttgagcgtgt gaacgaacgt gttggtgcgg cgggtgacag catggatctg gcgttccgta tcctggcggc ggcgagcgtt 1260 ggtgtgcagt gcccgcgtga cctggaacgt ctgaccgcgg gtcaatgcga tgatggtggc 1320. tgggatctgt gctggttcta cgtttttggt agcaccggcg tgaaagcggg taaccgtggc 1380. ctgaccaccg cgctggcggt gaccgcgatt caaaccgcga ttggtcgtcc gccgagcccg 1440 agcccgagcg cggcgagcag cagctttcgt ccgagcagcc cgtataaatt cctgggtatc 1500 agccgtccgg cgagcccgat tcgttttggt gacctgctgc gtccgtggcg taagatgagc 1560 cgtagcaacc tgaaaagcca ataa 1584 <210> 36 <211> 527 <212> PRT <213> Dichomitus squalens <220> <221> MISC_FEATURE <223> XP_007369631.1_amino acid sequence <400> 36 Met Ala Ser Ile His Arg Arg Tyr Thr Thr Leu Ile Leu Asp Leu Gly 1 5 10 15 Asp Val Leu Phe Arg Trp Ser Pro Lys Thr Glu Thr Ala Ile Pro Pro 20 25 30 Gln Gln Leu Lys Asp Ile Leu Ser Ser Val Thr Trp Phe Glu Tyr Glu 35 40 45 Arg Gly Arg Leu Ser Gln Glu Ala Cys Tyr Glu Arg Cys Ala Glu Glu 50 55 60 Phe Lys Ile Glu Ala Ser Val Ile Ala Glu Ala Phe Lys Gln Ala Arg 65 70 75 80 Gly Ser Leu Arg Pro Asn Glu Glu Phe Ile Ala Leu Ile Arg Asp Leu 85 90 95 Arg Arg Glu Met His Gly Asp Leu Thr Val Leu Ala Leu Ser Asn Ile 100 105 110 Ser Leu Pro Asp Tyr Glu Tyr Ile Met Ser Leu Ser Ser Asp Trp Thr 115 120 125 Thr Val Phe Asp Arg Val Phe Pro Ser Ala Leu Val Gly Glu Arg Lys 130 135 140 Pro His Leu Gly Cys Tyr Arg Lys Val Ile Ser Glu Met Asn Leu Glu 145 150 155 160 Pro Gln Thr Thr Val Phe Val Asp Asp Lys Leu Asp Asn Val Ala Ser 165 170 175 Ala Arg Ser Leu Gly Met His Gly Ile Val Phe Asp Asn Gln Ala Asn 180 185 190 Val Phe Arg Gln Leu Arg Asn Ile Phe Gly Asp Pro Ile Arg Arg Gly 195 200 205 Gln Glu Tyr Leu Arg Gly His Ala Gly Lys Leu Glu Ser Ser Thr Asp 210 215 220 Asn Gly Leu Ile Phe Glu Glu Asn Phe Thr Gln Leu Ile Ile Tyr Glu 225 230 235 240 Leu Thr Gln Asp Arg Thr Leu Ile Ser Leu Ser Glu Cys Pro Arg Thr 245 250 255 Trp Asn Phe Phe Arg Gly Glu Pro Leu Phe Ser Glu Thr Phe Pro Asp 260 265 270 Asp Val Asp Thr Thr Ser Val Ala Leu Thr Val Leu Gln Pro Asp Arg 275 280 285 Ala Leu Val Asn Ser Val Leu Asp Glu Met Leu Glu Tyr Val Asp Ala 290 295 300 Asp Gly Ile Met Gln Thr Tyr Phe Asp Arg Ser Arg Pro Arg Met Asp 305 310 315 320 Pro Phe Val Cys Val Asn Val Leu Ser Leu Phe Tyr Glu Asn Gly Arg 325 330 335 Gly His Glu Leu Pro Arg Thr Leu Asp Trp Val Tyr Glu Val Leu Leu 340 345 350 His Arg Ala Tyr His Gly Gly Ser Arg Tyr Tyr Leu Ser Pro Asp Cys 355 360 365 Phe Leu Phe Phe Met Ser Arg Leu Leu Lys Arg Ala Asp Asp Pro Ala 370 375 380 Val Gln Ala Arg Leu Arg Pro Leu Phe Val Glu Arg Val Asn Glu Arg 385 390 395 400 Val Gly Ala Ala Gly Asp Ser Met Asp Leu Ala Phe Arg Ile Leu Ala 405 410 415 Ala Ala Ser Val Gly Val Gln Cys Pro Arg Asp Leu Glu Arg Leu Thr 420 425 430 Ala Gly Gln Cys Asp Asp Gly Gly Trp Asp Leu Cys Trp Phe Tyr Val 435 440 445 Phe Gly Ser Thr Gly Val Lys Ala Gly Asn Arg Gly Leu Thr Thr Ala 450 455 460 Leu Ala Val Thr Ala Ile Gln Thr Ala Ile Gly Arg Pro Pro Ser Pro 465 470 475 480 Ser Pro Ser Ala Ala Ser Ser Ser Ser Phe Arg Pro Ser Ser Pro Tyr Lys 485 490 495 Phe Leu Gly Ile Ser Arg Pro Ala Ser Pro Ile Arg Phe Gly Asp Leu 500 505 510 Leu Arg Pro Trp Arg Lys Met Ser Arg Ser Asn Leu Lys Ser Gln 515 520 525 <210> 37 <211> 1584 <212> DNA <213> Artificial sequence <220> <223> XP_007369631.1_Class II motif mutant_E. coli, optimized <400> 37 atggcgagca tccaccgtcg ttacaccacc ctgattctgg acctgggtga tgttctgttt 60 cgttggagcc cgaagaccga gaccgcgatc ccgccgcagc aactgaaaga cattctgagc 120 agcgtgacct ggttcgagta cgaacgtggc cgtctgagcc aggaagcgtg ctatgaacgt 180 tgcgcggagg aatttaagat cgaggcgagc gttattgcgg aagcgttcaa acaagcgcgt 240 ggtagcctgc gtccgaacga ggaatttatc gcgctgattc gtgacctgcg tcgtgagatg 300 cacggcgatc tgaccgtgct ggcgctgagc aacatcagcc tgccggacta cgaatatatt 360 atgagcctga gcagcgactg gaccaccgtt ttcgatcgtg tgtttccgag cgcgctggtt 420 ggtgagcgta agccgcacct gggctgctat cgtaaagtga tcagcgagat gaacctggaa 480 ccgcagacca ccgtgttcgt tgacgataag ctggataacg ttgcgagcgc gcgtagcctg 540 ggtatgcatg gtattgtttt cgacaaccag gcgaacgtgt ttcgtcaact gcgtaacatc 600 ttcggtgatc cgattcgtcg tggccaagag tacctgcgtg gtcacgcggg caaactggaa 660 agcagcaccg acaacggtct gatcttcgag gaaaacttta cccagctgat catttatgag 720 ctgacccaag atcgtaccct gattagcctg agcgaatgcc cgcgtacctg gaacttcttt 780 cgtggcgagc cgctgttcag cgaaccttc ccggcggcgg ttgcgaccac cagcgttgcg 840 ctgaccgtgc tgcagccgga ccgtgcgctg gttaacagcg tgctggatga gatgctgga tacgtggacg cggatggtat catgcaaacc tattttgacc gtagccgtcc gcgtatggat ccgtttgtgt gcgttaacgt gctgagcctg ttctacgaga acggtcgtgg ccacgaactg ccgcgtaccc tggactgggt ttacgaggtg ctgctgcacc gtgcgtatca cggtggcagc 1080 cgttactatc tgagcccgga ttgcttcctg ttctttatga gccgtctgct gaaacgtgcg gatgatccgg cggttcaggc gcgtctgcgt ccgctgtttg ttgagcgtgt gaacgaacgt gttggtgcgg cgggtgacag catggatctg gcgttccgta tcctggcggc ggcgagcgtt 1260 ggtgtgcagt gcccgcgtga cctggaacgt ctgaccgcgg gtcaatgcga tgatggtggc 1320. tgggatctgt gctggttcta cgtttttggt agcaccggcg tgaaagcggg taaccgtggc 1380. ctgaccaccg cgctggcggt gaccgcgatt caaaccgcga ttggtcgtcc gccgagcccg agcccgagcg cggcgagcag cagctttcgt ccgagcagcc cgtataaatt cctgggtatc 1500 agccgtccgg cgagcccgat tcgttttggt gacctgctgc gtccgtggcg taagatgagc 1560 cgtagcaacc tgaaaagcca ataa 1584 <210> 38 <211> 527 <212> PRT <213> Artificial Sequence <220> <223> Mutant of Class II Motif XP_007369631.1 <400> 38 Met Ala Ser Ile His Arg Arg Tyr Thr Thr Leu Ile Leu Asp Leu Gly 1 5 10 15 Asp Val Leu Phe Arg Trp Ser Pro Lys Thr Glu Thr Ala Ile Pro Pro 20 25 30 Gln Gln Leu Lys Asp Ile Leu Ser Ser Val Thr Trp Phe Glu Tyr Glu 35 40 45 Arg Gly Arg Leu Ser Gln Glu Ala Cys Tyr Glu Arg Cys Ala Glu Glu 50 55 60 Phe Lys Ile Glu Ala Ser Val Ile Ala Glu Ala Phe Lys Gln Ala Arg 65 70 75 80 Gly Ser Leu Arg Pro Asn Glu Glu Phe Ile Ala Leu Ile Arg Asp Leu 85 90 95 Arg Arg Glu Met His Gly Asp Leu Thr Val Leu Ala Leu Ser Asn Ile 100 105 110 Ser Leu Pro Asp Tyr Glu Tyr Ile Met Ser Leu Ser Ser Asp Trp Thr 115 120 125 Thr Val Phe Asp Arg Val Phe Pro Ser Ala Leu Val Gly Glu Arg Lys 130 135 140 Pro His Leu Gly Cys Tyr Arg Lys Val Ile Ser Glu Met Asn Leu Glu 145 150 155 160 Pro Gln Thr Thr Val Phe Val Asp Asp Lys Leu Asp Asn Val Ala Ser 165 170 175 Ala Arg Ser Leu Gly Met His Gly Ile Val Phe Asp Asn Gln Ala Asn 180 185 190 Val Phe Arg Gln Leu Arg Asn Ile Phe Gly Asp Pro Ile Arg Arg Gly 195 200 205 Gln Glu Tyr Leu Arg Gly His Ala Gly Lys Leu Glu Ser Ser Thr Asp 210 215 220 Asn Gly Leu Ile Phe Glu Glu Asn Phe Thr Gln Leu Ile Ile Tyr Glu 225 230 235 240 Leu Thr Gln Asp Arg Thr Leu Ile Ser Leu Ser Glu Cys Pro Arg Thr 245 250 255 Trp Asn Phe Phe Arg Gly Glu Pro Leu Phe Ser Glu Thr Phe Pro Ala 260 265 270 Ala Val Ala Thr Thr Ser Val Ala Leu Thr Val Leu Gln Pro Asp Arg 275 280 285 Ala Leu Val Asn Ser Val Leu Asp Glu Met Leu Glu Tyr Val Asp Ala 290 295 300 Asp Gly Ile Met Gln Thr Tyr Phe Asp Arg Ser Arg Pro Arg Met Asp 305 310 315 320 Pro Phe Val Cys Val Asn Val Leu Ser Leu Phe Tyr Glu Asn Gly Arg 325 330 335 Gly His Glu Leu Pro Arg Thr Leu Asp Trp Val Tyr Glu Val Leu Leu 340 345 350 His Arg Ala Tyr His Gly Gly Ser Arg Tyr Tyr Leu Ser Pro Asp Cys 355 360 365 Phe Leu Phe Phe Met Ser Arg Leu Leu Lys Arg Ala Asp Asp Pro Ala 370 375 380 Val Gln Ala Arg Leu Arg Pro Leu Phe Val Glu Arg Val Asn Glu Arg 385 390 395 400 Val Gly Ala Ala Gly Asp Ser Met Asp Leu Ala Phe Arg Ile Leu Ala 405 410 415 Ala Ala Ser Val Gly Val Gln Cys Pro Arg Asp Leu Glu Arg Leu Thr 420 425 430 Ala Gly Gln Cys Asp Asp Gly Gly Trp Asp Leu Cys Trp Phe Tyr Val 435 440 445 Phe Gly Ser Thr Gly Val Lys Ala Gly Asn Arg Gly Leu Thr Thr Ala 450 455 460 Leu Ala Val Thr Ala Ile Gln Thr Ala Ile Gly Arg Pro Pro Ser Pro 465 470 475 480 Ser Pro Ser Ala Ala Ser Ser Ser Ser Phe Arg Pro Ser Ser Pro Tyr Lys 485 490 495 Phe Leu Gly Ile Ser Arg Pro Ala Ser Pro Ile Arg Phe Gly Asp Leu 500 505 510 Leu Arg Pro Trp Arg Lys Met Ser Arg Ser Asn Leu Lys Ser Gln 515 520 525 <210> 39 <211> 8 <212> PRT <213> Artificial sequence <220> <223> Class II synthase motif mutant BtHAD <400> 39 Pro Asp Ala Leu Ala Ser Thr Ser 1 5 <210> 40 <211> 8 <212> PRT <213> Artificial sequence <220> <223> Class II synthase motif_mutated_SmHAD1 + SmHAD2 <400> 40 Pro Pro Ala Ile Ala Thr Met Ser 1 5 <210> 41 <211> 8 <212> PRT <213> Artificial sequence <220> <223> Class II synthase motif_mutated_EMD37666 <400> 41 Pro Ala Ala Leu Ala Thr Thr Ser 1 5 <210> 42 <211> 8 <212> PRT <213> Artificial sequence <220> <223> Class II synthase motif_mutated_XP_007369631 <400> 42 Pro Ala Ala Val Ala Thr Thr Ser 1 5 <210> 43 <211> 7 <212> PRT <213> Artificial sequence <220> <223> QW motif_mutated_BtHAD <400> 43 Gln Asn Val Ala Gly Ser Trp 1 5 <210> 44 <211> 7 <212> PRT <213> Artificial sequence <220> <223> QW motif_mutated_SmHAD1+SmHAD2 <400> 44 Gln Asn Ala Ala Gly Ser Trp 1 5 <210> 45 <211> 5 <212> PRT <213> Artificial sequence <220> <223> Class I synthase motif DDxx(D / E) <220> <221> MISC_FEATURE <222> (3)..(4) <223> Xaa can be any naturally occurring amino acid. <220> <221> MISC_FEATURE <222> (5)..(5) <223> Xaa can be D or E <400> 45 Asp Asp Xaa Xaa Xaa 1 5 <210> 46 <211> 8 <212> PRT <213> Artificial sequence <220> <223> Class II synthase motif PxDxD(T / S)(T / M)S <220> <221> MISC_FEATURE <222> (2)..(2) <223> Xaa can be any naturally occurring amino acid. <220> <221> MISC_FEATURE <222> (4)..(4) <223> Xaa can be any naturally occurring amino acid. <220> <221> MISC_FEATURE <222> (6)..(6) <223> Xaa can be T or S <220> <221> MISC_FEATURE <222> (7)..(7) <223> Xaa can be T or M <400> 46 Pro Xaa Asp Xaa Asp Xaa Xaa Ser 1 5 <210> 47 <211> 8 <212> PRT <213> Artificial sequence <220> <223> Class II synthase motif PDDLDSTS <400> 47 Pro Asp Asp Leu Asp Ser Thr Ser 1 5 <210> 48 <211> 8 <212> PRT <213> Artificial sequence <220> <223> Class II synthase motif PDDLDTTS <400> 48 Pro Asp Asp Leu Asp Thr Thr Ser 1 5 <210> 49 <211> 8 <212> PRT <213> Artificial sequence <220> <223> Class II synthase motif PPDIDTMS <400> 49 Pro Pro Asp Ile Asp Thr Met Ser 1 5 <210> 50 <211> 8 <212> PRT <213> Artificial sequence <220> <223> Class II synthase motif PNDIDTMS <400> 50 Pro Asn Asp Ile Asp Thr Thr Ser 1 5 <210> 51 <211> 7 <212> PRT <213> Artificial sequence <220> <223> QW motif <220> <221> misc_feature <222> (2)..(3) <223> Xaa can be any naturally occurring amino acid. <220> <221> misc_feature <222> (6)..(6) <223> Xaa can be any naturally occurring amino acid. <400> 51 Gln Xaa Xaa Asp Gly Xaa Trp 1 5 <210> 52 <211> 7 <212> PRT <213> Artificial sequence <220> <223> QW motif QNVDGSW <400> 52 Gln Asn Val Asp Gly Ser Trp 1 5 <210> 53 <211> 7 <212> PRT <213> Artificial sequence <220> <223> QW motifQCDDGGW <400> 53 Gln Cys Asp Asp Gly Gly Trp 1 5 <210> 54 <211> 7 <212> PRT <213> Artificial sequence <220> <223> QW motif QSSDGGW <400> 54 Gln Ser Ser Asp Gly Gly Trp 1 5 <210> 55 <211> 7 <212> PRT <213> Artificial sequence <220> <223> QW motifQNEDGSW <400> 55 Gln Asn Glu Asp Gly Ser Trp 1 5 <210> 56 <211> 12 <212> PRT <213> Artificial sequence <220> <223> Conservative motif 1 Lxxxx(W / F)xxYxxG <220> <221> MISC_FEATURE <222> (2)..(5) <223> Xaa can be any naturally occurring amino acid. <220> <221> MISC_FEATURE <222> (6)..(6) <223> Xaa can be W or F <220> <221> MISC_FEATURE <222> (7)..(8) <223> Xaa can be any naturally occurring amino acid. <220> <221> MISC_FEATURE <222> (10) (11) <223> Xaa can be any naturally occurring amino acid. <400> 56 Leu Xaa Xaa Xaa Xaa Xaa Xaa Xaa Tyr Xaa Xaa Gly 1 5 10 <210> 57 <211> 12 <212> PRT <213> Artificial sequence <220> <223> Conservative motif 1 LRSHIWFNYSMG <400> 57 Leu Arg Ser His Ile Trp Phe Asn Tyr Ser Met Gly 1 5 10 <210> 58 <211> 12 <212> PRT <213> Artificial sequence <220> <223> Conservative motif 1 LRSATWAAYECG <400> 58 Leu Arg Ser Ala Thr Trp Ala Ala Tyr Glu Cys Gly 1 5 10 <210> 59 <211> 12 <212> PRT <213> Artificial sequence <220> <223> Conservative motif 1 LRTPTWGKYECG <400> 59 Leu Arg Thr Pro Thr Trp Gly Lys Tyr Glu Cys Gly 1 5 10 <210> 60 <211> 12 <212> PRT <213> Artificial sequence <220> <223> Conserved motif 1 LQHSSFLAYSCG <400> 60 Leu Gln His Ser Ser Phe Leu Ala Tyr Ser Cys Gly 1 5 10 <210> 61 <211> 12 <212> PRT <213> Artificial sequence <220> <223> Conserved motif 1 LRxxTWxxYECG <220> <221> misc_feature <222> (3)..(4) <223> Xaa can be any naturally occurring amino acid. <220> <221> misc_feature <222> (7)..(8) <223> Xaa can be any naturally occurring amino acid. <400> 61 Leu Arg Xaa Xaa Thr Trp Xaa Xaa Tyr Glu Cys Gly 1 5 10 <210> 62 <211> 16 <212> PRT <213> Artificial sequence <220> <223> Conservative motif 2 YxDxxRxRVD(P / A)V(V / A)xxN <220> <221> MISC_FEATURE <222> (2)..(2) <223> Xaa can be any naturally occurring amino acid. <220> <221> MISC_FEATURE <222> (4)..(5) <223> Xaa can be any naturally occurring amino acid. <220> <221> MISC_FEATURE <222> (7)..(7) <223> Xaa can be any naturally occurring amino acid. <220> <221> MISC_FEATURE <222> (11)..(11) <223> Xaa can be P or A <220> <221> MISC_FEATURE <222> (13)..(13) <223> Xaa can be V or A <220> <221> MISC_FEATURE <222> (14) (15) <223> Xaa can be any naturally occurring amino acid. <400> 62 Tyr Xaa Asp Xaa Xaa Arg Xaa Arg Val Asp Xaa Val Xaa Xaa Xaa Asn 1 5 10 15 <210> 63 <211> 16 <212> PRT <213> Artificial sequence <220> <223> Conservative motif 2 YFDPLRLRVDPVAATN <400> 63 Tyr Phe Asp Pro Leu Arg Leu Arg Val Asp Pro Val Ala Ala Thr Asn 1 5 10 15 <210> 64 <211> 16 <212> PRT <213> Artificial sequence <220> <223> Conservative motif 2 YFDETRPRVDAVVNVN <400> 64 Tyr Phe Asp Glu Thr Arg Pro Arg Val Asp Ala Val Val Asn Val Asn 1 5 10 15 <210> 65 <211> 16 <212> PRT <213> Artificial sequence <220> <223> Conservative motif 2 YFDKTRPRVDPVVCVN <400> 65 Tyr Phe Asp Lys Thr Arg Pro Arg Val Asp Pro Val Val Cys Val Asn 1 5 10 15 <210> 66 <211> 16 <212> PRT <213> Artificial sequence <220> <223> Conservative motif 2 YLDVERPRVDPVVIAN <400> 66 Tyr Leu Asp Val Glu Arg Pro Arg Val Asp Pro Val Val Ile Ala Asn 1 5 10 15 <210> 67 <211> 16 <212> PRT <213> Artificial sequence <220> <223> Conservative motif 2 YLDLERPRVDPVVIAN <400> 67 Tyr Leu Asp Leu Glu Arg Pro Arg Val Asp Pro Val Val Ile Ala Asn 1 5 10 15 <210> 68 <211> 16 <212> PRT <213> Artificial sequence <220> <223> Conservative motif 2 Y(F / L)Dx(T / E)RPRVD(P / A)VVx(A / V)N <220> <221> MISC_FEATURE <222> (2)..(2) <223> Xaa can be F or L <220> <221> MISC_FEATURE <222> (4)..(4) <223> Xaa can be any naturally occurring amino acid. <220> <221> MISC_FEATURE <222> (5)..(5) <223> Xaa can be T or E <220> <221> MISC_FEATURE <222> (11)..(11) <223> Xaa can be P or A <220> <221> MISC_FEATURE <222> (14)..(14) <223> Xaa can be any naturally occurring amino acid. <220> <221> MISC_FEATURE <222> (15)..(15) <223> Xaa can be A or V <400> 68 Tyr Xaa Asp Xaa Xaa Arg Pro Arg Val Asp Xaa Val Val Xaa Xaa Asn 1 5 10 15 <210> 69 <211> 13 <212> PRT <213> Artificial sequence <220> <223> Conservative motif 3 GTx(Y / F)YxxxExFL(Y / F) <220> <221> MISC_FEATURE <222> (3)..(3) <223> Xaa can be any naturally occurring amino acid. <220> <221> MISC_FEATURE <222> (4)..(4) <223> Xaa can be Y or F <220> <221> MISC_FEATURE <222> (6) (8) <223> Xaa can be any naturally occurring amino acid. <220> <221> MISC_FEATURE <222> (10)..(10) <223> Xaa can be any naturally occurring amino acid. <220> <221> MISC_FEATURE <222> (13)..(13) <223> Xaa can be Y or F <400> 69 Gly Thr Xaa Xaa Tyr Xaa Xaa Xaa Glu Xaa Phe Leu Xaa 1 5 10 <210> 70 <211> 13 <212> PRT <213> Artificial sequence <220> <223> Conservative motif 3 GTLYYRTPEAFLY <400> 70 Gly Thr Leu Tyr Tyr Arg Thr Pro Glu Ala Phe Leu Tyr 1 5 10 <210> 71 <211> 13 <212> PRT <213> Artificial sequence <220> <223> Conservative motif 3 GTLFYYHAESFLY <400> 71 Gly Thr Leu Phe Tyr Tyr His Ala Glu Ser Phe Leu Tyr 1 5 10 <210> 72 <211> 13 <212> PRT <213> Artificial sequence <220> <223> Conserved motif 3 GTRYYLSQEDFLF <400> 72 Gly Thr Arg Tyr Tyr Leu Ser Gln Glu Asp Phe Leu Phe 1 5 10 <210> 73 <211> 1755 <212> DNA <213> Dryopteris fragrans <220> <221> misc_feature <223> DfHAD_wt_nucleotide sequence <400> 73 atggagttct ctgcctctgc tcctcctcct aggctagcca gtgtcataat attggagcct 60 ctcggcttcc tcctcacacc acactactcc tctcagcttc ccaaaaagct gctccgtcgc 120 ctgttgtgca ctagaatctg gcacaggtat cagcgaggcc gccttcgcct gcgtgacgct 180 gctatgctgc tcgcccagct cccattccta gctgtgtctg atcacccctg ggctctggac 240 aatctcgcaa gcctgctccg ccccacagct gtgcgtgcgg tgccatggat gctgctgctg 300 ctcgacttcc tacgagacga gctccatctg aaggtagtct gcgcgaccaa ctcctcccca 360 gaagagctgc aagagctgcg ccaccagttt ccggccctct ttgccaaggt cgatgccacc 420 gtttcttcag gcgaggaggg cgtgggcaag ccgtccgtgc gcttcctgca ggctgcgttg 480 gacaaagccg gtgtccacgc gcagcaaacc ttgtatcttg actcttttga cagcttggag 540 accatcatgg ctgcacgctc tcttggcatg catgcactat ctgtagagcc atgccacatt 600 gatgagctca ccgccagggc ctcttccggc cagctaagag atgcacagct tataaggcgt 660 attgtgtgcg ccatgcacgg gccagcagta tctgcagttg tgtcgggcag tatcacatcg 720 tccggcccac agacagcaaa gatcgaggaa ttgccaacag ctgctgatag tcatctccgc 780 agcgcagctc tcacttctgc tcagcagttt ttcctcaaag ttattgctcc acatcgtcct 840 gagaagccat tcgtccagct tccatctctc acctcggagg gcatccgaat atacgacacc 900 tttgcacagt ttgtcatagc cgacctgctc gacgacaccc gcttcctacc catgcaatct 960 cctcctccca atgggctcat cacctttgtt aacccaagcg cgtaccttgc tgatgatata 1020 aagaatggca acagccatat tgtcccgggt gtgcaatttt acgcatccga tgcgtgcact 1080 ctcatcgaca tcccacatga cctagacacc acctccgttg gcttgtcagt actgcacaag 1140 tttggaaagg tggacaagga cacactcaac aaagtgctag acagaatgct cgagcaagtg 1200 agtgaagacg acggcattct gcaggtgtat tttgatgtgg agcgtccgcg catcgatcca 1260 gttgtggtgg caaacacggt gtttctgttc cacttgggaa agagagggca tgaggtggcg 1320 aggagtgaga agtttgtgga gagtgtgctg ctgcagaggg catacgaaga agggacgttg 1380 tattacaacc tgggggaagc atttttggtg agtgtggcga ggctggtgca cgagtttaag 1440 gagcacttta caaggagcgg catgaggagg gcactggagg agaggctaag agagcgggca 1500 agggcgggca tgcaagagag ggatgatgcg ctggcgctag ccatgcgcat tcgtgcatgc 1560 gctttgtgtg gcctggccgg agagggcctc acaaaagcag cagagcagga gcttttgcgc 1620 ctgcagtgca agtccaaggg ctgttggggg tgccaccctt tctatcgcaa tggcagtaat 1680 gtgctcagct ggatcggcag tgaggccctt accactgctt acgctattgc tgcgctacag 1740 cccattgata tttaa 1755 <210> 74 <211> 591 <212> PRT <213> Dryopteris fragrans <220> <221> MISC_FEATURE <223> DfHAD_ amino acid sequence <400> 74 Met Glu Phe Ser Ala Ser Ala Pro Pro Pro Arg Leu Ala Ser Val Ile 1 5 10 15 Ile Leu Glu Pro Leu Gly Phe Leu Leu Thr Pro His Tyr Ser Ser Gln 20 25 30 Leu Pro Lys Lys Leu Leu Arg Arg Leu Leu Cys Thr Arg Ile Trp His 35 40 45 Arg Tyr Gln Arg Gly Arg Leu Arg Leu Arg Asp Ala Ala Met Leu Leu 50 55 60 Ala Gln Leu Pro Phe Leu Ala Val Ser Asp His Pro Trp Ala Leu Asp 65 70 75 80 Asn Leu Ala Ser Leu Leu Arg Pro Thr Ala Val Arg Ala Val Pro Trp 85 90 95 Met Leu Leu Leu Leu Asp Phe Leu Arg Asp Glu Leu His Leu Lys Val 100 105 110 Val Cys Ala Thr Asn Ser Ser Pro Glu Glu Leu Gln Glu Leu Arg His 115 120 125 Gln Phe Pro Ala Leu Phe Ala Lys Val Asp Ala Thr Val Ser Ser Gly 130 135 140 Glu Glu Gly Val Gly Lys Pro Ser Val Arg Phe Leu Gln Ala Ala Leu 145 150 155 160 Asp Lys Ala Gly Val His Ala Gln Gln Thr Leu Tyr Leu Asp Ser Phe 165 170 175 Asp Ser Leu Glu Thr Ile Met Ala Ala Arg Ser Leu Gly Met His Ala 180 185 190 Leu Ser Val Glu Pro Cys His Ile Asp Glu Leu Thr Ala Arg Ala Ser 195 200 205 Ser Gly Gln Leu Arg Asp Ala Gln Leu Ile Arg Arg Ile Val Cys Ala 210 215 220 Met His Gly Pro Ala Val Ser Ala Val Val Ser Gly Ser Ile Thr Ser 225 230 235 240 Ser Gly Pro Gln Thr Ala Lys Ile Glu Glu Leu Pro Thr Ala Ala Asp 245 250 255 Ser His Leu Arg Ser Ala Ala Leu Thr Ser Ala Gln Gln Phe Phe Leu 260 265 270 Lys Val Ile Ala Pro His Arg Pro Glu Lys Pro Phe Val Gln Leu Pro 275 280 285 Ser Leu Thr Ser Glu Gly Ile Arg Ile Tyr Asp Thr Phe Ala Gln Phe 290 295 300 Val Ile Ala Asp Leu Leu Asp Asp Thr Arg Phe Leu Pro Met Gln Ser 305 310 315 320 Pro Pro Pro Asn Gly Leu Ile Thr Phe Val Asn Pro Ser Ala Tyr Leu 325 330 335 Ala Asp Asp Ile Lys Asn Gly Asn Ser His Ile Val Pro Gly Val Gln 340 345 350 Phe Tyr Ala Ser Asp Ala Cys Thr Leu Ile Asp Ile Pro His Asp Leu 355 360 365 Asp Thr Thr Ser Val Gly Leu Ser Val Leu His Lys Phe Gly Lys Val 370 375 380 Asp Lys Asp Thr Leu Asn Lys Val Leu Asp Arg Met Leu Glu Gln Val 385 390 395 400 Ser Glu Asp Asp Gly Ile Leu Gln Val Tyr Phe Asp Val Glu Arg Pro 405 410 415 Arg Ile Asp Pro Val Val Val Ala Asn Thr Val Phe Leu Phe His Leu 420 425 430 Gly Lys Arg Gly His Glu Val Ala Arg Ser Glu Lys Phe Val Glu Ser 435 440 445 Val Leu Leu Gln Arg Ala Tyr Glu Glu Gly Thr Leu Tyr Tyr Asn Leu 450 455 460 Gly Glu Ala Phe Leu Val Ser Val Ala Arg Leu Val His Glu Phe Lys 465 470 475 480 Glu His Phe Thr Arg Ser Gly Met Arg Arg Ala Leu Glu Glu Arg Leu 485 490 495 Arg Glu Arg Ala Arg Ala Gly Met Gln Glu Arg Asp Asp Ala Leu Ala 500 505 510 Leu Ala Met Arg Ile Arg Ala Cys Ala Leu Cys Gly Leu Ala Gly Glu 515 520 525 Gly Leu Thr Lys Ala Ala Glu Gln Glu Leu Leu Arg Leu...

Claims

1. An isolated polypeptide from a haloacid dehalogenase-like hydrolase superfamily, comprising cyclic terpene synthase activity, wherein the polypeptide is: a. The amino acid sequence is BazzHAD1 of SEQ ID NO:3; or d. BtHAD with the amino acid sequence SEQ ID NO:

12.

2. An isolated nucleic acid molecule whose nucleotide sequence encodes the polypeptide of claim 1.

3. An expression construct comprising at least one nucleic acid molecule of claim 2.

4. A vector comprising at least one nucleic acid molecule of claim 2 or at least one expression construct of claim 3.

5. A recombinant non-human host cell or recombinant non-human host organism, comprising: a. At least one isolated nucleic acid molecule of claim 2; or b. At least one expression construct of claim 3; or c. At least one carrier of claim 4, The recombinant non-human host organism is a microorganism.

6. The recombinant non-human host cell or recombinant non-human host organism according to claim 5, wherein the isolated nucleic acid molecule is stably integrated into the genome.

7. The recombinant non-human host cell or recombinant non-human host organism according to claim 5, wherein the expression construct is stably integrated into the genome.

8. A method for producing the polypeptide of claim 1, the method comprising: a. To culture the recombinant non-human host cell or recombinant non-human host organism of any one of claims 5 to 7 to express the polypeptide of claim 1.

9. The method according to claim 8, further comprising: b. Isolate the polypeptide from the recombinant non-human host cells or recombinant non-human organisms cultured in step a.

10. A method for producing styranne sesquiterpenes, comprising: a. Contacting farnesyl diphosphate with the polypeptide of claim 1, the polypeptide containing terpene synthase activity, thereby obtaining a bis(phosphate) of styranne sesquiterpene; and b. Cleavage the diphosphate moiety of the product obtained in step a. by chemical or enzymatic means. The sesquiterpene in this sesquiterpene is sesquiterpene alcohol, and the diphosphate ester of the sesquiterpene is sesquitenyl diphosphate.

11. The method of claim 10, further comprising: c. The bismuth subterpene was isolated.

12. The method of claim 10, wherein the enzymatic step b is carried out by applying a polypeptide that is different from the polypeptide of claim 1 and has phosphatase activity.

13. Use of the polypeptide as defined in claim 1 in the preparation of odorants.

14. The use according to claim 13, wherein the odorant is Ambrox.

15. A method for generating Ambrox, the method comprising: The method of claim 10 or 12 provides complementol; isolates complementol produced in the preceding step; and converts complementol into Ambrox in a manner known per se.