Novel polypeptides for producing peltatol and / or drimenol compounds

CN122772840APending Publication Date: 2026-09-18FIRMENICH SA +1
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202610570925.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2020-01-22
Filing Date
2020-11-25
Publication Date
2026-09-18

AI Technical Summary

Technical Problem

已经开发了化学合成途径,但仍然很复杂并且并不总是具有成本效益

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure FT_1
    Figure FT_1
  • Figure FT_2
    Figure FT_2
  • Figure FT_3
    Figure FT_3
Patent Text Reader

Abstract

The present invention relates to novel plant-derived halogenase-like (HAD-like) polypeptides having cyclic terpene synthase (TPS) activity, particularly suitable for use in biochemical methods for the production of drimane sesquiterpenes, including drimenol and / or selinol and / or related compounds, such as phosphorylated drimane sesquiterpene alcohols, particularly phosphorylated drimenol and / or selinol compounds and derivatives. The present invention also provides the novel TPS-activity-encoding nucleotide sequences, corresponding expression constructs, recombinant hosts, methods for the production of such novel polypeptides and mutants and variants thereof. The present invention also relates to the use of such novel polypeptides in the production of odorants, flavorants and fragrance ingredients.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This invention is a divisional application of patent application No. 202080031347.6, filed on November 25, 2020, entitled "Novel polypeptide for generating physostigmine and / or complementol compounds". Technical Field

[0002] This invention relates to novel plant-derived haloacid dehalogenase-like (HAD-like) polypeptides possessing cyclic terpene synthase (TPS) activity, particularly suitable for biochemical methods to produce sterane-type sesquiterpenes, including sterane alcohol and / or zephyranol and / or related compounds, such as phosphorylated sterane-type sesquiterpene alcohols, especially phosphorylated sterane alcohol and / or zephyranol compounds and derivatives. The invention also provides the encoding nucleotide sequences of the novel TPS activity, corresponding expression constructs, recombinant hosts, and methods for preparing such novel polypeptides and their mutants and variants. The invention further relates to the use of such novel polypeptides in the production of odorant, flavoring, and fragrance ingredients. Background Technology

[0003] Terpenes are found in most organisms (microorganisms, animals, and plants). These compounds are composed of isoprene units and are classified according to the number of these units present in their structure, which may contain cyclic structural elements. Therefore, hemiterpenes, monoterpenes, sesquiterpenes, and diterpenes are terpenes containing 5, 10, 15, and 20 carbon atoms (i.e., 1, 2, 3, and 4 isoprene units), respectively. Monoterpenes are derived from geranyl diphosphate (GPP, C... 10 ), sesquiterpenes are derived from farnesyl diphosphate (FPP, C 15 Diterpenes are derived from geraniol geraniol diphosphate (GGPP, C). 20 For example, sesquiterpenes are widely distributed in the plant kingdom. Many sesquiterpene molecules are known for their flavor and aroma properties, as well as their cosmetic, medicinal, and antibacterial effects. A variety of sesquiterpene hydrocarbons and sesquiterpene compounds have been identified. Chemical synthetic routes have been developed, but they remain complex and are not always cost-effective.

[0004] The biosynthesis of terpenes involves enzymes called terpene synthases. Several sesquiterpene synthases exist in the plant kingdom, all using the same substrate (farnesyl diphosphate, FPP) but with different product structures. Genes and cDNAs encoding sesquiterpene synthases have been cloned, and the corresponding recombinant enzymes have been characterized.

[0005] Many major sources of sesquiterpenes, such as compounds with drimane structures, like drinol or drinol compounds, are derived from plants; however, the content of sesquiterpenes in these natural sources may be low. There remains a need to identify and characterize novel terpene synthases, optionally characterized by specific drimane product structures and / or drimane product yields and / or more cost-effective methods for producing drimane compounds such as drinol and / or drinol, which are potential structural components for the preparation of high-value flavoring ingredients (e.g., Ambrox). Summary of the Invention

[0006] The aforementioned problems can be addressed by providing several novel enzymes with cyclic terpene synthase (TPS) activity, particularly cyclic sesquiterpene synthase activity, and more specifically exhibiting drimenyl and / or zosquienyl diphosphate synthase activity, producing drimenyl and / or zosquienyl diphosphates from FPP with remarkably high selectivity. In particular, novel drimenyl and / or zosquienyl diphosphate synthase genes have been identified from plants of the genus *Bazzania*, such as *Bazzania trilobata*, and *Selaginella moellendorffii*. These novel synthases have been named BazzHAD1, BazzHAD2, BazzHAD3, BtHAD, SmHAD1, and SmHAD2.

[0007] The protein sequences of SmHAD1 and SmHAD2 have very low identity (approximately 14–20%) with all other known HAD-like TPS, but they show high homology (>96%) among themselves. BtHAD and BazzHAD1 show slightly higher homology (approximately 22–30%) with fern and fungal HAD-like TPS and 96% homology among themselves, while BazzHAD2 and BazzHAD3 show moderate homology (approximately 22–43%) with fern and fungal HAD-like TPS and 67% homology among themselves (see Table 2 in the Experimental Section).

[0008] All HAD-like TPS described herein contain the class II motif PxDxD(T / S)(T / M)S (SEQ ID NO:46) and more C-terminal-shifted QW motifs QxxDGx(W / F) (SEQ ID NO:51), which are characteristic of diterpene synthases. Mutations on the class II motif result in complete loss of activity, while mutations on the QW motif result in 90% loss of activity. All six newly identified TPS of this invention lack the class I motif (DDxx(D / E) (SEQ ID NO:45)). Attached Figure Description

[0009] Figure 1. a) Alkane structure; b) More generalized alkane structure; the dashed lines in the equations indicate that a single C=C double bond may exist at any specified position.

[0010] Figure 2 The structural formulas of (+)-Zhessaurol and (-)-Bestrol.

[0011] Figure 3 Mechanisms for cyclization of farnesyl diphosphate (FPP) via HAD-like TPS containing class II motifs, and mechanisms for dephosphorylation of complementenyl and / or zearalenyl diphosphates via the same HAD-like TPS or other phosphatases.

[0012] Figure 4 GC / MS chromatograms of samples from BazzHAD1 in vivo biochemical analysis. Farnesol is produced by the dephosphorylation of the precursor FPP, while bismuth subsalicylate is likely a product of the dephosphorylation of bismuth subsalicylate.

[0013] Figure 5 . Figure 4 Mass spectrum of the methylbenzene peak.

[0014] Figure 6 GC / MS chromatograms of samples from BazzHAD2 in vivo biochemical analysis. Farnesol is produced by the dephosphorylation of the precursor FPP, while zephyranol and the unknown compound SQT1 are likely products of the dephosphorylation of zephyranol diphosphate and the unknown diphosphate ester, respectively.

[0015] Figure 7 . Figure 4 Mass spectrum of the alcohol peak in *Lysimachia christinae*.

[0016] Figure 8 . Figure 4 Mass spectrum of the unknown SQT1 peak.

[0017] Figure 9 GC / MS chromatograms of samples from BazzHAD3 in vivo biochemical analysis. Farnesol is produced by the dephosphorylation of the precursor FPP, while zephyranol is likely a product of the dephosphorylation of zephyranol diphosphate.

[0018] Figure 10 . Figure 7 Mass spectrum of the alcohol peak in *Lysimachia christinae*.

[0019] Figure 11 GC / MS chromatograms of samples from BtHAD in vivo biochemical analysis. Farnesol is produced by the dephosphorylation of the precursor FPP, while bismuth subsalicylate may be a product of the dephosphorylation of bismuth subsalicylate.

[0020] Figure 12 . Figure 10 Mass spectrum of the methylbenzene peak.

[0021] Figure 13 GC / MS chromatogram of the SmHAD1 (EFJ10816.1) sample for in vivo biochemical analysis. Farnesol is produced by the dephosphorylation of the precursor FPP, while bismuth subtilisol and bismuth subtilisol are likely products of the dephosphorylation of bismuth subtilisol and bismuth subtilisol, respectively.

[0022] Figure 14 . Figure 11 Mass spectrum of the alcohol peak in *Lysimachia christinae*.

[0023] Figure 15 . Figure 11 Mass spectrum of the methylbenzene peak.

[0024] Figure 16 GC / MS chromatogram of the SmHAD2 (EFJ26126.1) sample for in vivo biochemical analysis. Farnesol is produced by the dephosphorylation of the precursor FPP, while bismuth subtilisol and bismuth subtilisol are likely products of the dephosphorylation of bismuth subtilisol and bismuth subtilisol, respectively.

[0025] Figure 17 . Figure 14 Mass spectrum of the alcohol peak in *Lysimachia christinae*.

[0026] Figure 18 . Figure 14 Mass spectrum of the methylbenzene peak.

[0027] Figure 19 Comparison of yields of bismuth subtypes synthesized by SmHAD1, SmHAD2, BtHAD, XP_007369631.1 and EMD37666.1 and their mutants in in vivo biochemical analysis (mean ± SE, n=4~5). The SmHAD1 and SmHAD2 mutants contain mutations D257A and D259A in class II motifs (SEQ ID NO: 21, 28, and 40), and mutations E432A ​​and D433A in QW motifs (SEQ ID NO: 23, 30, and 44); the BtHAD mutant contains mutations D320A and D322A in class II motifs (SEQ ID NO: 14 and 39), and mutation D491A in QW motifs (SEQ ID NO: 16 and 43); the EMD37666.1 mutant contains mutations D276A, D277A, and D279A in class II motifs (SEQ ID NO: 34 and 41); and the XP_007369631.1 mutant contains mutations D272A, D273A, and D275A in class II motifs (SEQ ID NO: 38 and 42).

[0028] Figure 20 GC / MS chromatograms of samples from SmHAD1, SmHAD2, and BtHAD in the presence or absence of BAP (bacterial alkaline phosphatase) for in vitro biochemical analysis. (2E,6E)-farnesol is produced by the dephosphorylation of the precursor FPP, while zephyranol and bismuthol are likely products of the dephosphorylation of zephyranol and bismuthol diphosphates, respectively.

[0029] Figure 21 Amino acid sequence alignments of HAD-like TPS: BazzHAD1 (SEQ ID NO:3), BazzHAD2 (SEQ ID NO:6), BazzHAD3 (SEQ ID NO:9), BtHAD (SEQ ID NO:12), SmHAD1 (SEQ ID NO:19), SmHAD2 (SEQ ID NO:26); EMD37666.1 (SEQ ID NO:32) and XP_007369631.1 (SEQ ID NO:36), as previously described in WO 2018 / 220113. The image identified a type I motif (DDxx(D / E); SEQ ID NO:45) and a type II motif (PxDxD(T / S)(T / M)S; SEQ ID NO:46), as well as three other conserved sequence motifs 1-3 (Lxxxx(W / F)xxYxxG; SEQ ID NO:56, YxDxxRxRVD(P / A)V(V / A)xxN; SEQ ID NO:62, GTx(Y / F)YxxxExFL(Y / F); SEQ ID NO:69). Detailed Implementation

[0030] abbreviations used

[0031] bp base pairs

[0032] BAP (bacterial alkaline phosphatase)

[0033] BSA bovine serum albumin

[0034] DNA deoxyribonucleic acid

[0035] cDNA complementary DNA

[0036] DTT dithiothreitol

[0037] FPP process for nitrophosphite

[0038] GC gas chromatography

[0039] HAD haloacid dehalogenase

[0040] IPTG isopropyl-D-thiogalactopyranoside

[0041] LB lysate broth

[0042] MS mass spectrometer / mass spectrometry

[0043] MVA mevalonic acid

[0044] PCR polymerase chain reaction

[0045] PP diphosphate (ester) or pyrophosphate (ester)

[0046] RNA

[0047] mRNA messenger ribonucleic acid

[0048] miRNA

[0049] siRNA (small interfering RNA)

[0050] rRNA ribosomal RNA

[0051] tRNA transfer RNA

[0052] sp. species

[0053] TPS Terpene Synthase

[0054] Detailed description

[0055] a. Definition

[0056] Terpenes are a large and diverse class of organic compounds produced by a variety of plants, especially conifers, and some insects. Terpenes are hydrocarbons. Although sometimes used interchangeably with "terpenes," "terpenoids" or "isoprene-like compounds" are modified terpenes because they contain additional functional groups, usually oxygen-containing ones.

[0057] "Terpenoids" ("isoprene-like compounds") are a large and diverse class of naturally occurring organic chemical substances derived from terpenes. While sometimes used interchangeably with the term "terpene," "terpenoids" contain an additional functional group, typically an oxygen-containing group such as a hydroxyl, carbonyl, or carboxyl group. Most are polycyclic structures with oxygen-containing functional groups. Unless otherwise stated, the terms "terpene" and "terpenoids" are used interchangeably in the context of this specification.

[0058] Terpenes (and terpenoid compounds) can be classified according to the number of isoprene units in their molecules; the prefix in the name indicates the number of terpene units required to form the molecule. Hemiterpenes consist of a single isoprene unit. Monoterpenes consist of two isoprene units and have the molecular formula C2. 10 H 16Sesquiterpenes are composed of three isoprene units and have the molecular formula C0. 15 H 24 Diterpenes are composed of four isoprene units and have the molecular formula C0. 20 H 32 .

[0059] "Sysquiterpenes" or "sysquiterpenes" are bicyclic sesquiterpenes and are the parent structures of many natural products with various biological activities. Figure 1a According to IUPAC, its systematic chemical name is (4aR,5S,6S,8aS)-1,1,4a,5,6-pentamethyldecahydronaphthalene. For the purposes of this application, the term "complementane" must be understood more broadly and is also referred to as "complementane type," and encompasses those having, for example, […]. Figure 1b The bicyclic sesquiterpenes described herein are of a general structure. They are not limited to specific stereochemistry and may contain a single C=C double bond at certain positions in the sterane skeleton. Specific sterane structures are those contained in sterane alcohols and zephyranols.

[0060] "Symplocaneol" or "symplocane sesquiterpeneol" refers to "symplocane" or "symplocane sesquiterpene" with a hydroxyl group; specific examples are symplocaneol and zephyranol. This term specifically refers to substances with, for example, hydroxyl groups. Figure 1b The cyclic terpenes (terpenoids) shown have a carbon skeleton structure similar to that of alkylene.

[0061] For the purposes of this application, "Zheshinol" refers to (+)-Zheshinol (CAS: 54632-04-1). Figure 2 ).

[0062] "Zheshol derivatives" include compounds derived from zheshol through one or more steps such as esterification, hydroxylation, oxidation, acylation, isomerization, and dimethylation. Suitable derivatives may be selected from hydrocarbons, alcohols, glycols, triols, acetals, ketals, aldehydes, acids, ethers, amides, ketones, lactones, epoxides, acetates, glycosides, and / or esters.

[0063] For the purposes of this application, "supplementol" refers to (-)-supplementol (CAS: 468-68-8). Figure 2 ).

[0064] "Compositol derivatives" include compounds derived from composeols through one or more steps such as esterification, hydroxylation, oxidation, acylation, isomerization, and dimethylation. Suitable derivatives may be selected from hydrocarbons, alcohols, glycols, triols, acetals, ketals, aldehydes, acids, ethers, amides, ketones, lactones, epoxides, acetates, glycosides, and / or esters.

[0065] For the purposes of this application, “Ambrox” refers to the IUPAC name: (-)-(3aR,5aS,9aS, 9bR)-3a,6,6,9a-tetramethyldodecanonaphtho[2,1-b]furan (CAS: 6790-58-5).

[0066] The conversion of a "precursor" molecule of the target compound as described herein into the target compound is preferably achieved by at least one structural alteration of the precursor molecule through the enzymatic action of a suitable polypeptide. For example, a "bisphosphate precursor" (e.g., a "terpenoid bisphosphate precursor") can be converted into the target compound (e.g., a terpenoid alcohol) by enzymatic removal of the diphosphate moiety, such as by removing a mono- or diphosphate group with a phosphatase. For example, an "acyclic precursor" (e.g., an "acyclic terpenoid precursor") can be converted into a cyclic target molecule (e.g., a cyclic terpene compound) in one or more steps by the action of a cyclase or synthase, regardless of the specific enzymatic mechanism of such enzyme.

[0067] The terms “cyclic terpene synthase” or “polypeptide with cyclic terpene synthase activity” are used as synonyms for the terms “terpene cyclase” or “polypeptide with terpene cyclase activity”. The term “TPS” is used as an abbreviation for the term “terpene synthase”. As described herein, cyclic terpene synthases belong to the haloacid dehalogenase-like (HAD-like) hydrolases superfamily and contain domains corresponding to the Pfam domains PF13419 (http: / / pfam.xfam.org / family / PF13419) and / or PF00702 (http: / / pfam.xfam.org / family / PF00702.26). Some cyclic terpene synthases that do not possess significant Pfam domains PF13419 and / or PF00702 are still considered to be from HAD-like hydrolases because they contain conserved modification motifs of the aforementioned HAD-like hydrolases superfamily, such as class I, class II, and QW motifs (SEQ ID: 45, SEQ ID: 46, SEQ ID: 51). HAD-like hydrolases are parts of peptides that share amino acid sequence similarity and related functions with members of the HAD-like hydrolases family. HAD-like hydrolases can be identified in peptides by searching for amino acid motifs or features of this protein family. Tools for performing such searches are available, for example, at: https: / / www.ebi.ac.uk / interpro / or https: / / www.ebi.ac.uk / Tools / hmmer / . Proteins typically consist of one or more functional regions or domains. Different combinations of domains produce a wide variety of proteins found in nature. Therefore, identifying domains occurring within a protein can provide in-depth understanding of its function. Peptides containing HAD-like hydrolase domains and / or characteristic HAD-like hydrolase motifs play a role in the binding and cleavage of phosphate or diphosphate groups of ligands. Peptides of the haloacid dehalogenase-like (HAD-like) hydrolase superfamily containing cyclic terpene synthase activity can be defined as HAD-like TPS.

[0068] The term "Class I terpene synthase" refers to terpene synthases that catalyze reactions initiated by ionization, such as monoterpene and sesquiterpene synthases.

[0069] The terms "Class I terpene synthase motif," "Class I terpene synthase-like motif," "Class I synthase (like) motif," "Class I synthase motif," or "Class I motif" refer to the active site of terpene synthases containing the conserved DDxx (D / E) motif (SEQ ID NO:45). The aspartic acid residue of this Class I motif binds to divalent metal ions (most commonly Mg) associated with, for example, the binding of diphosphate groups. 2 + It also catalyzes the ionization and cleavage of the allyl diphosphate bonds of the substrate.

[0070] The term "class II terpene synthase" or "class II synthase" refers to terpene synthases that catalyze protonation-initiated cyclization reactions, such as those typically involved in the biosynthesis of triterpenes and labdane diterpenes. In class II terpene synthases, the protonation-initiated reaction may involve, for example, an acidic amino acid donating a proton to the terminal double bond.

[0071] The terms “(modified) class II terpene synthase motif,” “(modified) class II terpene synthase-like motif,” “(modified) class II synthase (like) motif,” “(modified) class II synthase motif,” or “(modified) class II motif” refer to the active site of a terpene synthase containing a conserved DxDxxS (SEQ ID NO:75) motif, such as the PxDxD(T / S)(T / M)S motif (SEQ ID NO:46) described herein.

[0072] The term “QW motif” in this document refers to the active site of a terpene synthase containing a conserved QxxxxxW (SEQ ID NO:76) motif, such as the QxxDGxW motif (SEQ ID NO:50) in this document.

[0073] The terms "bryophylloides diphosphate synthase," "polypeptide with bryophylloides diphosphate synthase activity," "bryophylloides diphosphate synthase protein," or "capability to produce bryophylloides diphosphate" refer to a polypeptide capable of catalyzing the synthesis of bryophylloides diphosphate in any stereoisomer or mixture thereof, starting from acyclic terpene pyrophosphates, particularly farnesyl diphosphate (FPP). Bromephylloides diphosphate may be the sole product or part of a mixture of sesquiterpenes. The mixture may contain bryophylloides monophosphate and / or bryophylloidol.

[0074] The terms "forsythoside alcohol synthase," "polypeptide with forsythoside alcohol synthase activity," or "forsythoside alcohol synthase protein" refer to a polypeptide capable of catalyzing the synthesis of forsythoside alcohol in any stereoisomer or mixture thereof, starting from acyclic terpene pyrophosphates, particularly farnesyl diphosphate (FPP). Forsythoside alcohol may be the sole product or part of a mixture of sesquiterpenes.

[0075] The activity of *Zygophylloides bisphosphate synthase* was determined under the "standard conditions" described below: It could be determined using recombinant *Zygophylloides bisphosphate synthase* expression host cells, disrupted *Zygophylloides bisphosphate synthase* expression cells, fractions, enrichments, or purified enzymes of *Zygophylloides bisphosphate synthase*, under conditions of approximately 20 to 45°C, for example approximately 25 to 40°C, preferably 25 to 32°C, in a culture medium or reaction medium (preferably buffered) at pH 6 to 11, preferably 7 to 9, and in the presence of a reference substrate (especially FPP), at an initial concentration of 1 to 100 µM, preferably 5 to 50 µM, especially 30 to 40 µM, or by endogenous production from the cell host. The conversion reaction to form *Zygophylloides bisphosphate* proceeds for 10 minutes to 5 hours, preferably about 1 to 2 hours. If endogenous alkaline phosphatase is absent, one or more exogenous phosphatases are added to the reaction mixture to convert *Zygophylloides bisphosphate* formed by the synthase. Then, conventional methods, such as extraction with an organic solvent like ethyl acetate, can be used to determine the physostigmine. Specific examples of suitable standard conditions are applied in the experimental section below, such as in Example 4, and these conditions should also form part of the overall disclosure of this invention.

[0076] The terms "supreme-enyl diphosphate synthase," "peptide with supreme-enyl diphosphate synthase activity," "supreme-enyl diphosphate synthase protein," or "capable of producing supreme-enyl diphosphate" refer to a polypeptide capable of catalyzing the synthesis of supreme-enyl diphosphate in any stereoisomer or mixture thereof, starting from acyclic terpene pyrophosphates, particularly farnesyl diphosphate (FPP). The supreme-enyl diphosphate may be the sole product or part of a mixture of sesquiterpenes. The mixture may contain supreme-enyl monophosphate and / or supreme alcohol.

[0077] The terms "synergist synthase," "polypeptide with synergist activity," or "synergist synthase protein" refer to a polypeptide capable of catalyzing the synthesis of synergist in any stereoisomer or mixture thereof, starting from acyclic terpene pyrophosphates, particularly farnesyl diphosphate (FPP). Synergist may be the sole product or part of a mixture of sesquiterpenes.

[0078] The "complemental-enyl diphosphate synthase activity" is determined under the "standard conditions" described below: It can be determined using recombinant complemental-enyl diphosphate synthase expression host cells, disrupted complemental-enyl diphosphate synthase expression cells, fractions, enrichments, or purified complemental-enyl diphosphate synthase enzymes, under conditions of approximately 20 to 45°C, for example approximately 25 to 40°C, preferably 25 to 32°C, in a culture medium or reaction medium (preferably buffered) at pH 6 to 11, preferably 7 to 9, and in the presence of a reference substrate (especially FPP), at an initial concentration of 1 to 100 µM, preferably 5 to 50 µM, especially 30 to 40 µM, or by endogenous production by the cell host. The conversion reaction to form complemental-enyl diphosphate proceeds for 10 minutes to 5 hours, preferably about 1 to 2 hours. If endogenous alkaline phosphatase is absent, one or more exogenous phosphatases are added to the reaction mixture to convert complemental-enyl diphosphate formed by the synthase. Composterol can then be determined using conventional methods, such as extraction with an organic solvent like ethyl acetate. Specific examples of suitable standard conditions are applied in the experimental section below, such as in Example 4, and these conditions should also form part of the overall disclosure of this invention.

[0079] The terms “biological function,” “function,” “biological activity,” or “activity” refer to the ability of terpene synthases, as described herein, to catalyze the formation of: styrenyl diphosphate and / or styrenol, and / or zedoaryl diphosphate and / or zedoaryl alcohol; or mixtures thereof comprising styrenyl diphosphate and / or styrenyl monophosphate and / or styrenol, and / or zedoaryl diphosphate and / or zedoaryl monophosphate and / or zedoaryl alcohol, and / or one or more other terpenes, particularly styrenyl diphosphate and / or zedoaryl diphosphate.

[0080] The term “terpene mixture” or “sesquiterpene mixture” refers to a mixture of terpenes or sesquiterpenes that contains styrenyl diphosphate and / or styrenyl monophosphate and / or styrenol, and / or zephyranyl diphosphate and / or zephyranyl monophosphate and / or zephyranol, and may also contain one or more additional terpenes, such as one or more additional sesquiterpenes.

[0081] The mevalonate pathway (also known as the isoprene pathway or HMG-CoA reductase pathway) is an essential metabolic pathway in eukaryotes, archaea, and certain bacteria. The mevalonate pathway begins with acetyl-CoA, producing two five-carbon structural units called isopentenyl pyrophosphate (IPP) and dimethylallyl pyrophosphate (DMAPP). Key enzymes include acetoacetyl-CoA thiolysis enzyme (atoB), HMG-CoA synthase (mvaS), HMG-CoA reductase (mvaA), mevalonate kinase (MvaK1), phosphate mevalonate kinase (MvaK2), mevalonate diphosphate decarboxylase (MvaD), and isopentenyl pyrophosphate isomerase (idi). The mevalonate pathway is linked to enzyme activity to produce terpene precursors GPP, FPP, or GGPP, particularly FPP synthase (ERG20), allowing recombinant cells to produce terpenes. As used herein, the terms "recombinant host cell / organism," "genetically modified cell / organism," or "transformed cell / organism" refer to a cell or organism that has been altered to carry at least one nucleic acid molecule, such as a recombinant gene encoding a desired protein or nucleic acid sequence, which, upon transcription, produces the functional polypeptide of the present invention, particularly a complement-enzyme synthase protein that can be used to produce complement-enzyme diphosphate and / or complement-enzyme monophosphate and / or complement-enzyme alcohol or a corresponding mixture of terpenes containing complement-enzyme diphosphate and / or complement-enzyme monophosphate and / or complement-enzyme alcohol, and / or a zirconia diphosphate synthase protein that can be used to produce zirconia diphosphate and / or zirconia monophosphate and / or zirconia alcohol or a corresponding mixture of terpenes containing zirconia diphosphate and / or zirconia monophosphate and / or zirconia alcohol. The host cell is particularly a bacterial cell, fungal cell, or plant cell, or a plant. The host cell may contain recombinant genes that have been integrated into the host cell's nucleus or organelle genome. Alternatively, the host may contain recombinant genes extrachromosomally.

[0082] The term "organism" refers to any non-human multicellular or single-celled organism, such as plants or microorganisms. In particular, microorganisms are bacteria, yeast, algae, or fungi.

[0083] The term "plant" is used interchangeably to include plant cells, including plant protoplasts, plant tissues, plant cell tissue cultures that produce regenerated plants or parts of plants, or plant organs such as roots, stems, leaves, flowers, pollen, ovules, embryos, fruits, etc. Any plant can be used to implement the methods described herein.

[0084] When a particular organism or cell naturally produces FPP, or when it does not naturally produce FPP but is modified (genetically) by the nucleic acids described herein (e.g., by transformation) to produce FPP, it is meant to be "capable of producing FPP". Organisms or cells that are transformed to produce higher or lower, especially higher, amounts of FPP than naturally occurring organisms or cells are also covered as "capable of producing FPP".

[0085] When a particular organism or cell naturally produces physostigmine diphosphate, or when it does not naturally produce physostigmine diphosphate but is converted to produce physostigmine diphosphate using the nucleic acids described herein, it is meant to be "capable of producing physostigmine diphosphate". Organisms or cells that are converted to produce higher or lower, particularly higher, amounts of physostigmine diphosphate than naturally occurring organisms or cells are also included in "capable of producing physostigmine diphosphate".

[0086] "Able to produce stratiostigmine" means that a particular organism or cell naturally produces stratiostigmine, or when it does not naturally produce stratiostigmine but is converted with nucleic acids as described herein to produce stratiostigyl diphosphate, and optionally further converted with nucleic acids to produce enzymatic activity that converts stratiostigyl diphosphate to stratiostigmine. Organisms or cells that are converted to produce higher or lower, particularly higher, amounts of stratiostigmine than naturally occurring organisms or cells are also included in "organisms or cells that can produce stratiostigmine."

[0087] When a particular organism or cell naturally produces complement-enyl diphosphate, or when it does not naturally produce complement-enyl diphosphate but is converted to produce complement-enyl diphosphate using the nucleic acids described herein, it is meant to be "capable of producing complement-enyl diphosphate". Organisms or cells that are converted to produce higher or lower, particularly higher, amounts of complement-enyl diphosphate than naturally occurring organisms or cells are also covered as "capable of producing complement-enyl diphosphate".

[0088] "Capable of producing complementols" means that a particular organism or cell naturally produces complementols, or that does not naturally produce complementols but is converted to complement-enyl diphosphate using nucleic acids as described herein, and optionally further converted to produce enzymatic activity that converts complement-enyl diphosphate to complementols. Organisms or cells that are converted to produce higher or lower, particularly higher, amounts of complementols than naturally occurring organisms or cells are also included in the category of "organisms or cells capable of producing complementols."

[0089] As used herein, the terms “purified,” “substantially purified,” and “isolated” refer to a state free from other different compounds (with which the compounds of the present invention are typically associated in their natural state), and thus “purified,” “substantially purified,” and “isolated” articles comprise at least 0.5%, 1%, 5%, 10%, or 20%, or at least 50% or 75% by weight of a given sample. In one embodiment, these terms mean that the compounds of the present invention comprise at least 95%, 96%, 97%, 98%, 99%, or 100% by weight of a given sample. As used herein, when referring to nucleic acids or proteins, the terms “purified,” “substantially purified,” and “isolated” for nucleic acids or proteins also refer to a purified or concentrated state that differs from that naturally occurring in, for example, prokaryotic or eukaryotic environments, such as in bacterial or fungal cells, or in mammals, particularly humans. Any degree of purification or concentration greater than that of naturally occurring purification or concentration, including (1) purification from other related structures or compounds, or (2) association with structures or compounds that are typically unrelated in the said prokaryotic or eukaryotic environment, is within the meaning of “isolated.” Nucleic acids, proteins, or classes of nucleic acids or proteins described herein may be isolated or associated with structures or compounds that are otherwise unrelated in nature, according to various methods and processes known to those skilled in the art.

[0090] The term “about” indicates a possible variation of ±25% in the value, particularly ±15%, ±10%, more particularly ±5%, ±2%, or ±1%.

[0091] The term “basically” describes a value range of approximately 80 to 100%, such as 85 to 99.9%, particularly 90 to 99.9%, even more particularly 95 to 99.9%, or 98 to 99.9%, particularly 99 to 99.9%.

[0092] "Mainly" refers to a proportion in the range of more than 50%, such as in the range of 51% to 100%, especially in the range of 75% to 99.9%; especially in the range of 85% to 98.5%, such as 95% to 99%.

[0093] In the context of this invention, "major product" means a single compound or a group of at least two compounds, such as two, three, four, five or more, particularly two or three compounds, which are "majorly" prepared by the reaction described herein and are contained in the reaction in a major proportion based on the total amount of the components of the products formed by the reaction. The proportion may be a molar proportion, a weight proportion, or preferably an area proportion calculated from the corresponding chromatograms of the reaction products based on chromatographic analysis.

[0094] In the context of this invention, "byproduct" means a single compound or a group of at least two compounds, such as two, three, four, five or more, particularly two or three compounds, which are not "mainly" prepared by the reaction described herein.

[0095] Due to the reversibility of enzymatic reactions, unless otherwise stated, this invention relates to enzymatic or biocatalytic reactions described herein in both reaction directions.

[0096] The “functional mutants” of the peptides described in this article include “functional equivalents” of such peptides as defined below.

[0097] The term "stereoisomer" specifically includes conformational isomers.

[0098] According to the present invention, all "stereoisomers" of the compounds described herein are generally included, such as "structural isomers", especially "stereoisomers".

[0099] "Stereoisomeric forms" particularly include "stereoisomers" and mixtures thereof, such as configurational isomers (optical isomers), like enantiomers, or geometrical isomers (diastereomers), like E- and Z-isomers, and combinations thereof. If one or more asymmetric centers are present in a molecule, the invention includes all combinations of different conformations of these asymmetric centers, such as enantiomer pairs.

[0100] "Stereoselectivity" describes the ability to produce a specific stereoisomer of a compound in its stereoisomeric pure form, or the ability to specifically convert a specific stereoisomer from a variety of stereoisomers using the enzymatic catalytic method described herein. More specifically, this means that the product of the invention is enriched relative to a specific stereoisomer, or that the precipitate can be depleted relative to a specific stereoisomer. This can be quantified by a purity parameter %ee calculated according to the following formula:

[0101] %ee = [X A -X B ] / [X A +X B ]*100,

[0102] Where X A and X B Mole ratio (Molenbruch) represents the ratio of stereoisomers A and B.

[0103] The terms "selective conversion" or "increased selectivity" generally refer to the conversion of a specific stereoisomer, such as the E-form of the unsaturated hydrocarbon, at a higher proportion or amount (compared to molar amounts) than the corresponding other stereoisomers, such as the Z-form, during the entire course of the reaction (i.e., between the start and end of the reaction), at a certain point in time of the reaction, or during a "segment" of the reaction. Specifically, during the "segment," the selectivity can be observed corresponding to conversions of 1 to 99%, 2 to 95%, 3 to 90%, 5 to 85%, 10 to 80%, 15 to 75%, 20 to 70%, 25 to 65%, 30 to 60%, or 40 to 50% of the initial substrate amount. The higher proportion or amount can be expressed, for example, as follows:

[0104] -High maximum yield of isomers observed throughout the entire reaction process or during a specified period;

[0105] - Higher relative abundance of isomers at a defined percentage substrate conversion value; and / or

[0106] -At higher conversion percentage values, the same relative content of isomers;

[0107] Each of these preferred methods is observed relative to a reference method, which is performed under otherwise identical conditions using known chemical or biochemical methods.

[0108] According to the present invention, all “isomeric forms” of the compounds described herein are generally included, such as structural isomers, especially stereoisomers and mixtures thereof, such as optical isomers or geometric isomers, such as E and Z isomers, and combinations thereof. If several asymmetric centers exist in a molecule, the present invention includes all combinations of different conformations of these asymmetric centers, such as enantiomer pairs, or any mixture of stereoisomeric forms.

[0109] In connection with the description and appended claims provided herein, unless otherwise stated, the use of “or” means “and / or”. Similarly, the tenses of “containing,” “containing,” “comprising,” and “including” are interchangeable and not restrictive.

[0110] It should be further understood that, where the term “comprising” is used in the description of various implementation schemes, those skilled in the art will understand that, in certain specific cases, the language of “substantially consisting of” or “consisting of” can be used instead to describe the implementation schemes.

[0111] If this disclosure relates to features, parameters, and ranges of different priorities (including superior, non-preferred features, parameters, and ranges), then unless otherwise stated, any combination of two or more of these features, parameters, and ranges is covered in the disclosure of this invention regardless of their respective priority.

[0112] b. Specific embodiments of the invention

[0113] This invention relates to the following embodiments:

[0114] 1. An isolated polypeptide from a haloacid dehalogenase-like (HAD-like) hydrolase superfamily, comprising cyclic terpene synthase activity, wherein said polypeptide is selected from:

[0115] a. BazzHAD1 comprising the amino acid sequence SEQ ID NO:3, or a mutant or natural variant thereof, comprising an amino acid sequence having at least 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% sequence identity with SEQ ID NO:3 and retaining the cyclic terpene synthase activity, particularly complementenyl diphosphate synthase activity;

[0116] b. BazzHAD2 comprising the amino acid sequence SEQ ID NO:6, or a mutant or natural variant thereof, comprising an amino acid sequence having at least 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% sequence identity with SEQ ID NO:6 and retaining the cyclic terpene synthase activity, particularly fuscinyl diphosphate synthase activity;

[0117] c. BazzHAD3 comprising the amino acid sequence SEQ ID NO:9, or a mutant or natural variant thereof, comprising an amino acid sequence having at least 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% sequence identity with SEQ ID NO:9 and retaining the cyclic terpene synthase activity, particularly fuscinyl diphosphate synthase activity;

[0118] d. BtHAD comprising the amino acid sequence SEQ ID NO:12, or a mutant or natural variant thereof, comprising an amino acid sequence having at least 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% sequence identity with SEQ ID NO:12 and retaining the cyclic terpene synthase activity, particularly complementenyl diphosphate synthase activity;

[0119] e. SmHAD1 comprising the amino acid sequence SEQ ID NO:19, or a mutant or natural variant thereof, comprising an amino acid sequence having at least 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% sequence identity with SEQ ID NO:19 and retaining the cyclic terpene synthase activity, particularly complementenyl diphosphate synthase activity, and / or particularly fuscinyl diphosphate synthase activity;

[0120] f. SmHAD2 comprising the amino acid sequence SEQ ID NO:26, or a mutant or natural variant thereof, comprising an amino acid sequence having at least 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% sequence identity with SEQ ID NO:19 and retaining the cyclic terpene synthase activity, particularly complementenyl diphosphate synthase activity, and / or particularly fuscinyl diphosphate synthase activity;

[0121] The specific TPS subfamily refers to items a), b), c), and d) above, and more specifically, the TPS listed in items a), b), and c).

[0122] A more specific embodiment of this embodiment involves polypeptide variants or, in particular, non-natural mutants of the novel polypeptides described above, of the present invention, comprising the cyclic terpene synthase activity as described above. These variants or non-natural mutants are derived from any specific amino acid sequence of SEQ ID NO:3, 6, 9, 12, 19, and 26. These variants are selected from polypeptides comprising an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% sequence identity with any of SEQ ID NO:3, 6, 9, 12, 19, or 26, and containing at least one amino acid sequence difference relative to the corresponding unmodified polypeptide of SEQ ID NO:3, 6, 9, 12, 19, or 26, for example, at least one amino acid sequence position modified by addition, substitution, insertion, or deletion. Therefore, such variants have less than 100% sequence identity with the corresponding unmodified polypeptide. In this document, "at least one" encompasses at least 1 to 20, 1 to 15, 1 to 10, and more particularly 1, 2, 3, 4, or 5 amino acid sequence positions that are modified independently of each other by the addition, substitution, insertion, or deletion of amino acid residues.

[0123] 2. The polypeptide of implementation scheme 1, comprising the ability to produce the following substances:

[0124] a) at least one phosphate precursor of a sesquiterpene alcohol; or

[0125] b) A phosphate precursor of at least one sterane sesquiterpene alcohol and the corresponding at least one sterane sesquiterpene alcohol; or

[0126] c) At least one sesquiterpene alcohol.

[0127] 3. The polypeptide of embodiment 2, wherein the phosphate precursor is a styrannosyl monophosphate (ester), or more particularly, a styrannosyl diphosphate (ester); and / or wherein the styrannosyl sesquiterpene alcohol is styrannosyl alcohol and / or styrannosyl alcohol.

[0128] 4. The polypeptide produced according to implementation schemes 1, 2, or 3, including those that generate

[0129] a) Benzene diphosphate and / or folinose diphosphate derived from farnesyl diphosphate (FPP) as a substrate; or

[0130] b) Composinyl diphosphate and / or zephyranyl diphosphate derived from farnesyl diphosphate (FPP) as a substrate; and composinol and / or zephyranol derived directly from FPP as a substrate or via their respective diphosphate precursors; or

[0131] c) Compositol and / or zedoaryl from FPPs as substrates, which are derived directly from FPPs as substrates or via their respective diphosphate precursors.

[0132] Table 1 below shows an overview of specific TPSs of the present invention, including their specific sequence motifs, their relative positions in the complete amino acid sequence, and their respective product profiles (based on the type of complementalanol).

[0133]

[0134] 5. The polypeptide of one of the aforementioned embodiments, which is isolated from...

[0135] a. Plants of the genus *Bazzania*, especially *Bazzania trilobata*; or

[0136] b. Plants of the genus Selaginella, especially the species Selaginella moellendorffii;

[0137] c. Other lycophytes, liverworts, or mosilophytes containing sesquiterpenes, including but not limited to the genera *Porella*, *Hymenophyton*, and *Marchantia*.

[0138] Or it may come from other plant species.

[0139] 6. The polypeptide of any one of embodiments 1 to 5, further comprising:

[0140] a. The modified type II terpene synthase motif shown in SEQ ID NO:46 (Px0Dx1D(T / S)(T / M)S), wherein x0 and x1 can be any naturally occurring amino acid residue, and x1 specifically represents I or L, particularly the motif of any one of SEQ ID NO:47, 48, 49, or 50, wherein the modified type II terpene synthase motif corresponds to:

[0141] i. The sequence positions of bits 318 to 325 of SEQ ID NO:3,

[0142] ii. The sequence positions of bits 269 to 276 of SEQ ID NO:6,

[0143] iii. The sequence positions of bits 277 to 284 of SEQ ID NO:9,

[0144] iv. Sequence positions 318 to 325 of SEQ ID NO:12,

[0145] v. The sequence positions of bits 255 to 262 of SEQ ID NO:19, and

[0146] vi. The sequence positions of bits 255 to 262 of SEQ ID NO:26;

[0147] b. The QW motif shown in SEQ ID NO:51 (Qx2x3DGx4W), wherein x2, x3, and x4 can be any naturally occurring amino acid residue, and x4 particularly represents G or S, especially the motif of any one of SEQ ID NO:52, 53, 54, or 55, wherein the QW motif corresponds to:

[0148] i. The sequence positions of bits 489 to 495 of SEQ ID NO:3,

[0149] ii. The sequence positions of bits 431 to 437 of SEQ ID NO:6,

[0150] iii. The sequence positions of bits 439 to 445 of SEQ ID NO:9,

[0151] iv. Sequence positions 488 to 494 of SEQ ID NO:12

[0152] v. The sequence positions of bits 430 to 436 of SEQ ID NO:19, and

[0153] vi. The sequence positions of bits 430 to 436 of SEQ ID NO:26; and

[0154] c. Optionally, at least one, for example, 1, 2, or 3 additional sequence motifs selected from:

[0155] (1) SEQ ID NO:56 (Lx5x6x7x8(W / F)x9x 10 Yx 11 x 12 G) represents the conservative sequence motif 1, where x5 to x 12 It can be any naturally occurring amino acid residue, x5 specifically represents R or Q, x8 specifically represents S, I or T, x 11 Specifically representing E or S, and x 12 Specifically representing C or M, particularly the motifs of SEQ ID NO: 57, 58, 59, 60, or 61, wherein the conserved sequence motif 1 corresponds to:

[0156] i. The sequence positions of bits 68 to 79 of SEQ ID NO:3,

[0157] ii. The sequence positions of bits 37 to 48 of SEQ ID NO:6,

[0158] iii. The sequence positions of bits 45 to 56 of SEQ ID NO:9,

[0159] iv. The sequence positions of bits 68 to 79 of SEQ ID NO:12,

[0160] v. The sequence positions of bits 28 to 39 of SEQ ID NO:19, and

[0161] vi. The sequence positions of bits 28 to 39 of SEQ ID NO:26;

[0162] (2) SEQ ID NO:62 (Yx 13 Dx 14 x 15 Rx 16 RVD(P / A)V(V / A)x 17 x 18 The conserved sequence motif 2 shown in N) is x 13 To x 18 It can be any naturally occurring amino acid residue, x 13 Specifically represents F or L, and x 16 Specifically representing P or L, particularly the motifs of SEQ ID NO: 63, 64, 65, 66, 67 or 68, wherein the conserved sequence motif 2 corresponds to:

[0163] i. The sequence positions of bits 362 to 377 of SEQ ID NO:3,

[0164] ii. The sequence positions of bits 309 to 324 of SEQ ID NO:6,

[0165] iii. The sequence positions of bits 317 to 332 of SEQ ID NO:9,

[0166] iv. Sequence positions 362 to 377 of SEQ ID NO:12,

[0167] v. The sequence positions of bits 299 to 314 of SEQ ID NO:19, and

[0168] vi. The sequence positions of bits 299 to 314 of SEQ ID NO:26; and

[0169] (3) SEQ ID NO:69 (GTx19 (Y / F)Yx 20 x 21 x 22 Ex 23 The conserved sequence motif 3 is shown in FL(Y / F), where x 19 To x 23 It can be any naturally occurring amino acid residue, x 19 Specifically representing L or R, particularly the motifs of SEQ ID NO: 70, 71 or 72, wherein the conserved sequence motif 3 corresponds to:

[0170] i. The sequence positions of bits 410 to 422 of SEQ ID NO:3,

[0171] ii. The sequence positions of bits 357 to 369 of SEQ ID NO:6,

[0172] iii. The sequence positions of bits 365 to 377 of SEQ ID NO:9,

[0173] iv. Sequence positions 410 to 422 of SEQ ID NO:12

[0174] v. The sequence positions of bits 349 to 361 of SEQ ID NO:19, and

[0175] vi. Sequence positions 349 to 361 of SEQ ID NO:26.

[0176] More specifically, the TPS of the present invention includes type II motifs, QW motifs, and at least one of conservative motifs 1, 2 and 3.

[0177] More specifically, the TPS of the present invention includes type II motifs, QW motifs, and at least two of the conservative motifs 1, 2 and 3.

[0178] More specifically, the TPS of the present invention includes type II motifs, QW motifs, and all three conserved motifs 1, 2 and 3.

[0179] The relative order of these sequence motifs in embodiment 6 within their respective amino acid sequences of each TPS of the present invention is described below (see also...). Figure 21 ):

[0180] N-terminus --- (Conservative motif 1) --- (Type II motif) --- (Conservative motif 2) ---

[0181] --- (Conservative motif 3) --- (QW motif) --- C-terminus

[0182] Type I motifs known from other TPSs (which are missing in the TPS of this invention) will be located between conserved motif 1 and type II motifs (see...). Figure 21 ).

[0183] 7. The polypeptide of any of the foregoing embodiments catalyzes the conversion of acyclic farnesyl diphosphate (particularly (2E,6E)-3,7,11-trimethyldodec-2,6,10-triene-1-pyrophosphate; FPP) into sero-alkenyl phosphate derivatives, such as monophosphates, more particularly sero-alkenyl diphosphates, with a selectivity of 50%ee or higher, such as 50 to 100%ee, or 60 to 90%ee or 70 to 80%ee.

[0184] 8. The polypeptide of any of the foregoing embodiments, which catalyzes the conversion of acyclic farnesyl diphosphate (particularly (2E,6E)-3,7,11-trimethyldodec-2,6,10-triene-1-pyrophosphate; FPP) into fussulacean phosphate derivatives, such as monophosphates, more particularly fussulacean phosphate, especially with a selectivity of 50%ee or higher, such as 50 to 100%ee, or 60 to 90%ee or 70 to 80%ee.

[0185] 9. The isolated polypeptide of any of the foregoing embodiments, wherein

[0186] a. Contains an amino acid sequence selected from SEQ ID NO:3, 6, 9, 12, 19, and 26; or

[0187] b. Encoded by a nucleic acid molecule containing a coding nucleotide sequence selected from SEQ ID NO:1, 2, 4, 5, 7, 8, 10, 11, 17, 18, 24 and 25.

[0188] 10. The isolated polypeptide of implementation scheme 9, which

[0189] a. Composed of amino acid sequences selected from SEQ ID NO:3, 6, 9, 12, 19, and 26; or

[0190] b. Encoded by a nucleic acid molecule consisting of a coding nucleotide sequence selected from SEQ ID NO:1, 2, 4, 5, 7, 8, 10, 11, 17, 18, 24 and 25.

[0191] 11. An isolated nucleic acid molecule, which:

[0192] a. A nucleotide sequence comprising a polypeptide encoding any of the foregoing embodiments; or

[0193] b. Containing a nucleotide sequence selected from SEQ ID NO:1, 2, 4, 5, 7, 8, 10, 11, 17, 18, 24, and 25, or containing a nucleotide sequence having at least 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% sequence identity with the nucleotide sequence SEQ ID NO:1, 2, 4, 5, 7, 8, 17, 18, or 24 or 25, and encoding a polypeptide of the HAD-like hydrolase superfamily, the polypeptide containing terpene synthase activity, particularly the ability to produce the following from farnesyl diphosphate (FPP) as a substrate:

[0194] Bruxine sesquiterpene alcohols, and / or their phosphate precursors, such as monophosphates, and more particularly their diphosphate precursors,

[0195] In particular, styrenyl phosphate precursors, such as monophosphates, and even more so styrenyl diphosphates; and / or styrenyl phosphate precursors, such as monophosphates, and even more so styrenyl diphosphates;

[0196] More specifically, as described in Table 1 above; or

[0197] c. Contains a nucleotide sequence that includes a sequence complementary to one of the sequences in b.; or

[0198] d. A nucleotide sequence that hybridizes with a, b, or c under strict conditions.

[0199] In one particular implementation, the nucleic acid may be naturally present in plants of the genera *Bazzania* or *Selaginella*, such as *Bazzania trilobata* or *Selaginella moellendorffii*, or other plant species, or obtained by modifying SEQ ID NO: 1, 2, 4, 5, 7, 8, 10, 11, 17, 18, 24 and 25 or their reverse complementary sequences.

[0200] In another implementation, the nucleic acid is isolated or derived from plants of the genera *Bazzania* or *Selaginella*, such as *Bazzania trilobata* or *Selaginella moellendorffii*.

[0201] 12. An expression construct comprising at least one nucleic acid molecule of embodiment 11, optionally combined with at least one regulatory sequence.

[0202] 13. A vector comprising at least one nucleic acid molecule of embodiment 11 or at least one expression construct of embodiment 12.

[0203] 14. The vector of implementation scheme 13, wherein the vector is a prokaryotic, viral or eukaryotic vector.

[0204] 15. The carrier of implementation scheme 13 or 14, wherein the carrier is an expression carrier.

[0205] 16. The vector of any one of Implementation Schemes 13 to 15 is a plasmid vector.

[0206] 17. A recombinant non-natural non-human host cell or recombinant non-human host organism prepared by genetic engineering, comprising:

[0207] a. At least one isolated nucleic acid molecule of embodiment 11, which optionally stably integrates into the genome; or

[0208] b. At least one expression construct of embodiment 12, which optionally is stably integrated into the genome; or

[0209] c. At least one carrier of any one of embodiments 13 to 16.

[0210] The recombinant non-human host cell or recombinant non-human host organism of this embodiment differs from natural non-human host cells or non-human host organisms in that exogenous genetic material, such as isolated nucleic acids, expression constructs, or vectors, is artificially introduced into the host organism or host cell. The recombinant non-human host cell or recombinant non-human host organism of this embodiment is a non-human host cell or non-human host organism that has been (genetically) modified to contain the desired exogenous genetic material, such as isolated nucleic acids, expression constructs, or vectors, as described in the above embodiments. In particular, the recombinant non-human host cell or recombinant non-human host organism of this embodiment is a cell or organism that differs from cells or organisms that naturally possess the desired genetic material.

[0211] 18. The host cell or host organism of implementation scheme 17 is selected from prokaryotes or eukaryotes, or derived from their cells.

[0212] 19. The host cell or host organism of implementation scheme 18 is selected from bacterial cells, fungal cells and plant cells, or plants.

[0213] 20. The host cell or host organism of embodiment 19, wherein the fungal cell is a yeast cell.

[0214] 21. The host cell or host organism of embodiment 20, wherein the yeast cell is selected from the genera Saccharomyces or Pichia, particularly from the species Saccharomyces cerevisiae or Pichia pastoris.

[0215] 22. The host cell or host organism of embodiment 18, wherein the bacterial cell is selected from the genus Escherichia, particularly Escherichia coli.

[0216] Some of these host cells or host organisms cannot naturally produce FPP. To suit the methods described herein, organisms or cells that do not naturally produce acyclic terpene pyrophosphate precursors such as FPP are genetically modified, particularly by transformation, transduction, or conjugation, and more particularly by transformation, to produce the precursors. Non-human host organisms or non-human host cells may be modified (e.g., transformed) before or simultaneously with the nucleic acid modification according to any of the above embodiments. Methods for modifying (e.g., transforming) organisms to produce acyclic terpene pyrophosphate precursors such as FPP are known in the art. For example, introducing enzymatic activity of the mevalonate pathway (isoprene-like pathway) or the MEP pathway is a suitable strategy for producing FPP from an organism (see also the Examples section herein).

[0217] 23. A method for generating at least one catalytically active polypeptide according to any one of embodiments 1 to 10, the method comprising:

[0218] a. Culturing non-human host cells or non-human host organisms according to embodiment 17 to express at least one polypeptide of any one of embodiments 1 to 10; and

[0219] b. Optionally, the polypeptide may be isolated from non-human host cells or organisms cultured in step a.

[0220] 24. The method of embodiment 23 further includes, prior to step a, genetically modifying the non-human host cell or non-human host organism to express a polypeptide according to any one of embodiments 1 to 10 by inserting at least one nucleic acid of embodiment 11, at least one expression construct of embodiment 12, or at least one vector of any one of embodiments 13 to 16 into the non-human host cell or non-human host organism, particularly by transformation, transduction, or conjugation.

[0221] 25. A method for producing sterine sesquiterpene alcohols, particularly sterine alcohol and / or zephyranol, the method comprising:

[0222] a. Contacting farnesyl diphosphate (FPP) with a polypeptide as defined in any one of embodiments 1 to 10 or with a polypeptide prepared according to embodiment 23 or 24 to obtain at least one phosphate ester, particularly a diphosphate ester of styrannosesquiterpene alcohol, particularly styrannosyl and / or zephyranosyl phosphate esters, more particularly styrannosyl diphosphate and / or zephyranosyl diphosphate;

[0223] b. Cleavage, chemically or enzymatically, the phosphate fraction, of the product obtained in step a., particularly the diphosphate fraction; and

[0224] c. Optionally, sterane sesquiterpene alcohols, particularly sterane alcohol and / or zestyrax alcohol, are isolated.

[0225] In one particular embodiment, the phosphate ester moiety is enzymatically cleaved by applying a phosphatase, more specifically an acidic or alkaline phosphatase. Alkaline phosphatases from various sources, such as bacterial enzymes, are preferred. Suitable phosphatases are commercially available enzymes.

[0226] In one embodiment, sterane sesquiterpene alcohols, particularly sterane alcohol and / or zephyranol, or mixtures containing sterane sesquiterpene alcohols, are isolated.

[0227] In one particular embodiment, the method further includes step d, processing the styrannosyl sesquiterpene alcohol formed in step b or isolated in step c, particularly styrannosyl alcohol and / or styrannosyl alcohol, using chemical synthesis or biocatalytic synthesis (e.g., biochemical synthesis with an expressed enzyme, or biotransformation with living cells expressing the enzyme) or a combination of both to obtain derivatives thereof. The styrannosyl sesquiterpene alcohol derivatives may be particularly selected from hydrocarbons, alcohols, diols, triols, acetals, ketals, aldehydes, acids, ethers, amides, ketones, lactones, epoxides, acetates, glycosides, and / or esters. In one embodiment, the method includes contacting the styrannosyl sesquiterpene alcohol with at least one enzyme to produce the styrannosyl sesquiterpene alcohol derivative. Specifically, the styrannosyl sesquiterpene alcohol derivative can be obtained by contacting the styrannosyl sesquiterpene alcohol with an enzyme such as, but not limited to, oxidoreductases, monooxygenases, dioxygenases, and transferases using a biochemical method. Biochemical transformation can be carried out in vitro using isolated enzymes, enzymes from lysed cells, or in vivo using whole cells. In another embodiment, the above method includes converting styrannosyl sesquiterpene alcohols into styrannosyl sesquiterpene alcohol derivatives using chemical synthesis. Specifically, this can be obtained by chemical methods, such as, but not limited to, oxidation, reduction, alkylation, acylation, and / or rearrangement. In another embodiment, the above method includes step e., optionally isolating the derivative of step d.

[0228] 26. The method of embodiment 25, wherein the sterane sesquiterpene alcohol comprises sterane alcohol and / or zephyranol alcohol, particularly as the main product.

[0229] 27. The method of any one of embodiments 25 and 26, comprising providing, in particular, transforming a non-human host cell or non-human host organism with at least one nucleic acid of embodiment 11, at least one expression construct of embodiment 12, or at least one vector of embodiments 13 to 16, such that the non-human host cell or non-human host organism expresses a polypeptide according to any one of embodiments 1 to 10.

[0230] 28. The method of any one of embodiments 25 to 27, wherein the FPP is contacted with a non-human host cell or non-human host organism of any one of embodiments 17 to 22, with its cell lysate or with a culture medium containing said non-human host cell or non-human host organism; and / or with a polypeptide as defined in any one of embodiments 1 to 10, said polypeptide being expressed or isolated in said non-human host cell or non-human host organism, cell lysate or culture medium.

[0231] 29. The method of embodiment 28, wherein the sesquiterpene alcohol is produced by fermentation from the non-human host organism or non-human host cell.

[0232] 30. The method of embodiment 28, wherein the styrannosesquiterpene alcohol is produced by an enzymatic method comprising converting the FPP with an isolated polypeptide of any one of embodiments 1 to 10, the step optionally being carried out in the presence of other adjuvants.

[0233] In a particular embodiment, in the method of any one of embodiments 25 to 30, the polypeptide comprises:

[0234] a. An amino acid sequence having at least 90% sequence identity with SEQ ID NO:3 or SEQ ID NO:12;

[0235] b. A class II synthase motif possessing the amino acid sequence PDDLDSTS (SEQ ID NO:47); and

[0236] c. A QW motif possessing the amino acid sequence QNVDGSW (SEQ ID NO:52); and

[0237] d. Optionally, at least one other sequence motif selected from the following amino acid sequences:

[0238] i. Lxxxx(W / F)xxYxxG (SEQ ID NO:56), where x can be any naturally occurring amino acid residue, especially the amino acid sequence LRSHIWFNYSMG (SEQ ID NO:57);

[0239] ii. YxDxxRxRVD(P / A)V(V / A)xxN (SEQ ID NO:62), where x can be any naturally occurring amino acid residue, particularly the amino acid sequence YFDPLRLRVDPVAATN (SEQ ID NO:63); and

[0240] iii. GTx(Y / F)xYxxxExFL(Y / F) (SEQ ID NO:69), where x can be any naturally occurring amino acid residue, especially the amino acid sequence GTLYYRTPEAFLY (SEQ ID NO:70);

[0241] Furthermore, this sesquiterpene alcohol contains sesquiterpene alcohol.

[0242] In another specific embodiment, in the method of any one of embodiments 25 to 30, the polypeptide comprises:

[0243] a. An amino acid sequence having at least 90% sequence identity with SEQ ID NO:6 or SEQ ID NO:9;

[0244] b. A class II synthase motif having the amino acid sequence PxDxD(T / S)(T / M)S (SEQ ID NO:46), wherein x can be any naturally occurring amino acid residue, particularly the amino acid sequences PDDLDTTS (SEQ ID NO:48) or PNDLDTTS (SEQ ID NO:50); and

[0245] c. A QW motif comprising the amino acid sequence Qx1x2DGx3W (SEQ ID NO:51), wherein x1 and x2 can be any naturally occurring amino acid residue, and x3 is G, particularly the amino acid sequences QCDDGGW (SEQ ID NO:53) or QSSDGGW (SEQ ID NO:54); and

[0246] d. Optionally, at least one other sequence motif selected from the following amino acid sequences:

[0247] i. LRxxTWxxYECG (SEQ ID NO:61), where x can be any naturally occurring amino acid residue, especially the amino acid sequence LRSATWAAYECG (SEQ ID NO:58) or LRTPTWGKYECG (SEQ ID NO:59);

[0248] ii. Y(F / L)Dx(T / E)RPRVD(P / A)VVx(A / V)N (SEQ ID NO:68), where x can be any naturally occurring amino acid residue, particularly the amino acid sequence YFDETRPRVDAVVNVN (SEQ ID NO:64) or YFDKTRPRVDPVVCVN (SEQ ID NO:65); and

[0249] iii. GTx(Y / F)YxxxExFL(Y / F) (SEQ ID NO:69), where x can be any naturally occurring amino acid residue, especially the amino acid sequence GTLFYYHAESFLY (SEQ ID NO:71);

[0250] Furthermore, this sesquiterpene alcohol contains physostigmine alcohol.

[0251] In another specific embodiment, in the method of any one of embodiments 25 to 30, the polypeptide comprises:

[0252] a. An amino acid sequence having at least 90% sequence identity with SEQ ID NO:19 or SEQ ID NO:26;

[0253] b. A class II synthase motif possessing the amino acid sequence PPDIDTMS (SEQ ID NO:49); and

[0254] c. A QW motif possessing the amino acid sequence QNEDGSW (SEQ ID NO:55); and

[0255] d. Optionally, at least one other sequence motif selected from the following amino acid sequences:

[0256] i. Lxxxx(W / F)xxYxxG (SEQ ID NO:56), where x can be any naturally occurring amino acid residue, especially the amino acid sequence LQHSSFLAYSCG (SEQ ID NO:60);

[0257] ii. Y(F / L)Dx(T / E)RPRVD(P / A)VVx(A / V)N (SEQ ID NO:68), where x can be any naturally occurring amino acid residue, particularly the amino acid sequence YLDVERPRVDPVVIAN (SEQ ID NO:66) or YLDLERPRVDPVVIAN (SEQ ID NO:67); and

[0258] iii. GTx(Y / F)xYxxxExFLx (SEQ ID NO:69), where x can be any naturally occurring amino acid residue, especially the amino acid sequence GTRYYLSQEDFLF (SEQ ID NO:72);

[0259] Furthermore, this sterane sesquiterpene alcohol contains sterane alcohol and zephyranol.

[0260] More specifically, the applied TPS includes at least one of type II motifs, QW motifs, and conserved motifs i., ii. and iii.

[0261] More specifically, the applied TPS includes at least two of the following: type II motifs, QW motifs, and conserved motifs i., ii., and iii.

[0262] More specifically, the applied TPS includes type II motifs, QW motifs, and all three conserved motifs i., ii., and iii.

[0263] 31. Use of the polypeptide as defined in any one of embodiments 1 to 10 in the preparation of odorant, flavoring or fragrance ingredients, particularly Ambrox (preferably via styrosine / styrosine diphosphate).

[0264] 32. Use of styrannosyl sesquiterpene alcohol prepared according to any one of embodiments 25 to 30 in the preparation of odorant, flavoring or fragrance ingredients, particularly Ambrox.

[0265] 33. A method for generating Ambrox, the method comprising:

[0266] a. To provide complementol and / or zephyranthesol by the method of any one of the embodiments described in any of embodiments 25 to 30.

[0267] b. Optionally, to separate the phytositol and / or phytositol produced in step a.; and

[0268] c. Converting styrax and / or styrax in a manner known per se, as reported in Tetrahedron: Asymmetry 11 (2000) 1375-1388.

[0269] 34. A composition comprising a substance prepared according to embodiment 32 or 33.

[0270] 35. The composition according to embodiment 34, which is selected from body care compositions, home care compositions and fragrance compositions.

[0271] 36. A method for producing sterine sesquiterpene alcohols, particularly sterine alcohol and / or zephyranol, the method comprising:

[0272] a. Culturing a non-human host organism or non-human host cell capable of producing FPP and converting it to express the polypeptide of any one of embodiments 1 to 10; and

[0273] b. Optionally, sterane sesquiterpene alcohols, particularly sterane alcohol and / or zestyrax alcohol, are isolated.

[0274] In one embodiment, a non-human host organism or non-human host cell that does not naturally produce FPP is genetically modified, particularly by transformation, transduction, or conjugation, and more specifically by transformation, to produce the precursor. The non-human host organism or non-human host cell may be modified (e.g., transformed) before or simultaneously with the nucleic acid modification according to any of the above embodiments. Methods for modifying (e.g., transforming) an organism to produce acyclic terpene pyrophosphate precursors such as FPP are known in the art. For example, introducing enzymatic activity of the mevalonate pathway (isoprene-like pathway) or the MEP pathway is a suitable strategy for producing FPP from an organism (see also the Examples section herein).

[0275] 37. A method for preparing mutant polypeptides of the haloacid dehalogenase-like (HAD-like) hydrolase superfamily, comprising terpene synthase activity, particularly comprising the ability to produce styryl sesquiterpene alcohols and / or their phosphate derivatives such as monophosphates, more particularly their diphosphate derivatives; and particularly the ability to produce styryl alkenyl phosphate derivatives such as monophosphates, more particularly styryl alkenyl diphosphates, and / or zephyranyl phosphate derivatives such as monophosphates, more particularly zephyranyl diphosphates, from farnesyl diphosphate (FPP) as a substrate, the method comprising the steps of:

[0276] a. Select nucleic acid molecules according to implementation plan 11;

[0277] b. Modify the selected nucleic acid molecule to obtain at least one mutant nucleic acid molecule;

[0278] c. Genetic modification of non-human host cells or single-celled non-human host organisms with a mutant nucleic acid sequence, particularly by horizontal gene transfer such as transformation, transduction or conjugation, and more particularly by transformation of said host cells or single-celled host organisms, to express a polypeptide encoded by the mutant nucleic acid sequence;

[0279] d. Screen for at least one mutant of the expression product, which contains terpene synthase activity, particularly the ability to produce styryl sesquiterpene alcohols and / or their phosphate precursors such as monophosphates, more particularly their diphosphate derivatives; and particularly the ability to produce styryl alkenyl phosphate precursors such as monophosphates, more particularly styryl alkenyl diphosphates, and / or styryl phosphate precursors such as monophosphates, more particularly styryl alkenyl diphosphates, from FPP as a substrate; and,

[0280] e. Optionally, if the peptide does not have the desired mutant activity, repeat steps a. to d. until a peptide with the desired mutant activity is obtained; and,

[0281] f. Optionally, if a polypeptide with the desired mutant activity is identified in step d, the corresponding mutant nucleic acid obtained in step c. is isolated.

[0282] Unless otherwise stated, within the range of “at least 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% sequence identity,” the particular values ​​are at least 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% sequence identity, while the values ​​of at least 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% sequence identity are even more particular.

[0283] c. Polypeptides applicable according to the present invention

[0284] In the context of this article, the following definitions apply:

[0285] The commonly used terms “polypeptide” or “peptide” refer to a natural or synthetic, continuous, peptide-linked linear chain or sequence of amino acid residues containing about 10 to more than 1,000 residues. In some embodiments provided herein, a polypeptide comprises an amino acid sequence that serves as an enzyme, or a fragment or variant thereof. Short-chain polypeptides having up to 30 residues are also referred to as “oligopeptides”.

[0286] The term "isolated polypeptide" refers to an amino acid sequence extracted from its natural environment by any method known in the art or a combination of such methods (including recombinant, biochemical, and synthetic methods).

[0287] The term "protein" refers to a large molecular structure composed of one or more polypeptides. It includes oligopeptides, peptides, polypeptides, and full-length proteins, whether natural or synthetic. The amino acid sequence of a polypeptide represents the protein's "primary structure." The amino acid sequence also predetermines the protein's "secondary structure" by forming specific structural elements (such as α-helices and β-sheets formed within the polypeptide chain). The arrangement of multiple such secondary structural elements defines the protein's "tertiary structure," or spatial arrangement. If a protein contains more than one polypeptide chain, these chains are arranged spatially to form the protein's "quaternary structure." Proper spatial arrangement, or "folding," is a prerequisite for protein function. Denaturation or unfolding disrupts protein function. If this disruption is reversible, protein function can be restored by refolding.

[0288] The typical protein function referred to in this article is "enzyme function," which means that a protein acts as a biocatalyst on a substrate, such as a compound, and catalyzes the conversion of said substrate into a product. Enzymes can exhibit high or low levels of substrate and / or product specificity.

[0289] Therefore, the term "polypeptide" as used herein to refer to a specific "activity" implicitly refers to a properly folded protein that exhibits the indicated activity, such as specific enzyme activity. Thus, unless otherwise stated, the term "polypeptide" also encompasses the terms "protein" and "enzyme".

[0290] A "target peptide" is an amino acid sequence that targets a protein or polypeptide to intracellular organelles (i.e., mitochondria or plastids) or to the extracellular space (secretory signal peptides). The nucleic acid sequence encoding the target peptide can be fused to the amino-terminal (e.g., N-terminus) nucleic acid sequence encoding the protein or polypeptide, or it can be used to replace the natural target peptide.

[0291] The present invention also relates to “functional equivalents” (also referred to as “analogs” or “functional mutations”) of the polypeptides specifically described herein.

[0292] For example, a “functional equivalent” refers to a polypeptide that, in a test used to determine the activity of enzymatic complementenyl and / or zephyranoid diphosphate synthase, shows at least 1 to 10%, or at least 20%, or at least 50%, or at least 75%, or at least 90% higher or lower complementenyl and / or zephyranoid diphosphate synthase activity compared to the polypeptide specifically described herein.

[0293] According to the invention, "functional equivalents" also encompass specific mutants that have an amino acid at at least one sequence position in the amino acid sequence described herein that differs from the specifically stated amino acid, but still possess one of the aforementioned biological activities, such as enzyme activity. Thus, "functional equivalents" include mutants obtainable by the addition, substitution, particularly conservative substitution (i.e., the amino acid in question is replaced by an amino acid having the same charge, size, polarity, and / or solubility), deletion, and / or inversion of one or more, for example, 1 to 20, 1 to 15, or 5 to 10 amino acids, wherein said changes can occur at any sequence position, as long as they result in the mutant possessing the general characteristics of the invention. Functional equivalence is also particularly provided if the activity pattern qualitatively overlaps between the mutant and the unaltered polypeptide, i.e., if, for example, an interaction with the same agonist or antagonist or substrate is observed, but at different rates (i.e., by EC...). 50 or IC 50 (Values ​​or any other parameters suitable in this technical field). The table below shows examples of suitable (conservative) amino acid substitutions:

[0294]

[0295] The “functional equivalents” in the above sense are also the “precursors” of the polypeptides described in this article, as well as the “functional derivatives” and “salts” of the polypeptides.

[0296] In this case, a "precursor" is a natural or synthetic precursor of a polypeptide, which may or may not have the desired biological activity.

[0297] The term "salt" as used in this invention refers to salts of the carboxyl group of the protein molecule and salts formed by the acid addition of the amino group. Salts of the carboxyl group can be produced in known ways, including inorganic salts such as sodium, calcium, ammonium, iron, and zinc salts, as well as salts formed with organic bases such as amines, such as triethanolamine, arginine, lysine, piperidine, etc. Salts formed by acid addition, such as those formed with inorganic acids such as hydrochloric acid or sulfuric acid, and salts formed with organic acids such as acetic acid and oxalic acid, are also covered by this invention.

[0298] The “functional derivatives” of the polypeptides according to the invention can also be generated using known techniques at the side groups of functional amino acids or their N-terminus or C-terminus. Such derivatives include, for example: aliphatic esters of carboxylic acid groups, amides of carboxylic acid groups, which can be obtained by reacting with ammonia or with primary or secondary amines; N-acyl derivatives of free amino groups, which are generated by reacting with acyl groups; or O-acyl derivatives of free hydroxyl groups, which are generated by reacting with acyl groups.

[0299] "Functional equivalents" naturally include polypeptides that can be obtained from other organisms as well as naturally occurring variants. For example, the area of ​​homologous sequence regions can be determined by sequence comparison, and equivalent polypeptides can be determined based on the specific parameters of this invention.

[0300] "Functional equivalents" also include "fragments" of the polypeptide according to the invention, such as single domains or sequence motifs, or truncated N-terminuses and / or C-termini, which may or may not exhibit the desired biological function. Preferably, such "fragments" at least qualitatively retain the desired biological function.

[0301] Furthermore, a “functional equivalent” is a fusion protein having one of the polypeptide sequences described herein or a functional equivalent derived therefrom, and having at least one additional functionally distinct heterologous sequence in functional N-terminal or C-terminal association (i.e., without substantial mutual functional impairment of the fusion protein portion). Non-limiting examples of such heterologous sequences are, for example, signal peptides, histidine anchors, or enzymes.

[0302] The invention also includes “functional equivalents” that are homologs of the specifically disclosed polypeptides. They have at least 60%, preferably at least 75%, particularly at least 80 or 85%, such as 90, 91, 92, 93, 94, 95, 96, 97, 98, or 99% homology (or identity) with one of the specifically disclosed amino acid sequences, calculated using the algorithm described in Pearson and Lipman, Proc. Natl. Acad. Sci. (USA) 85(8), 1988, 2444-2448. The homology or identity of the homologous polypeptides according to the invention, expressed as a percentage, refers in particular to identity expressed as a percentage of amino acid residues based on the total length of one of the amino acid sequences specifically described herein.

[0303] Identity data expressed as a percentage can also be determined using BLAST alignment, the blastp (protein-protein BLAST) algorithm, or by applying the Clustal settings detailed below.

[0304] In the case of possible protein glycosylation, the “functional equivalents” according to the invention include polypeptides in deglycosylated or glycosylated forms as described herein, as well as modified forms that can be obtained by changing the glycosylation pattern.

[0305] Functional equivalents or homologs of the polypeptides according to the present invention can be generated by mutagenesis, for example by point mutation, lengthening or shortening of the protein, or as described in more detail below.

[0306] Functional equivalents or homologs of the polypeptides according to the invention can be identified by screening a database of mutants, such as shortened mutants. For example, a database of protein variant diversity can be generated by combinatorial mutagenesis at the nucleic acid level, for example by enzymatic ligation of a mixture of synthetic oligonucleotides. Numerous methods are available for generating a database of potential homologs from degenerate oligonucleotide sequences. The chemical synthesis of degenerate gene sequences can be performed in an automated DNA synthesizer, and the synthesized gene can then be ligated into a suitable expression vector. The use of degenerate genomes makes it possible to provide all sequences in a mixture that encode the desired set of potential protein sequences. Methods for synthesizing degenerate oligonucleotides are known to those skilled in the art.

[0307] In the prior art, several techniques are known for screening gene products from combinatorial databases generated by point mutations or shortening, and for screening cDNA libraries containing gene products with selected properties. These techniques can be applied to rapidly screen gene libraries generated by combinatorial mutagenesis of homologs according to the present invention. The most commonly used high-throughput analysis-based techniques for screening large gene libraries involve cloning the gene library in a reproducible expression vector, transforming suitable cells with the resulting vector database, and expressing the combinatorial gene under specific conditions, under which detection of the desired activity facilitates the isolation of vectors encoding the gene (whose product is detected). Recursive integration mutagenesis (REM) is a technique for increasing the frequency of functional mutants in a database and can be used in conjunction with screening tests to identify homologs.

[0308] The embodiments provided herein offer orthologs and paralogs of the disclosed peptides, as well as methods for identifying and isolating such orthologs and paralogs. The definitions of the terms "ortholog" and "paralog" are given below and apply to both amino acid and nucleic acid sequences.

[0309] The polypeptides of this invention comprise all active forms of the enzymes of this invention, including active subsequences, such as catalytic domains or active sites. In one embodiment, this invention provides the catalytic domains or active sites as described below. In one embodiment, the present invention provides a peptide or polypeptide comprising or composed of an active site domain predicted by using a database such as Pfam (http: / / pfam.wustl.edu / hmmsearch.shtml) (a large collection covering multiple sequence alignments and hidden Markov models for many common protein families, Pfam Protein Family Database, A. Bateman, E. Birney, L. Cerruti, R. Durbin, L. Etwiller, SR Eddy, S. Griffiths-Jones, KL Howe, M. Marshall, and ELL Sonnhammer, NucleicAcids Research, 30(1):276-280, 2002) or equivalent sources such as the InterPro and SMART databases (http: / / www.ebi.ac.uk / interpro / scan.html, http: / / smart.embl-heidelberg.de / ).

[0310] The present invention also covers “peptide variants” having the desired activity, wherein the variant peptide is selected from amino acid sequences having at least 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity with the specific, particularly natural, amino acid sequence referred to by the specific SEQ ID NO and containing at least one, for example, 1 to 30 or 1 to 20, such as 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 amino acid sequence differences, for example, addition, substitution, insertion, or deletion of amino acids relative to the (unmodified) SEQ ID NO.

[0311] d. The coding nucleic acid sequence applicable according to the present invention

[0312] In the context of this article, the following definitions apply:

[0313] The terms “nucleic acid sequence,” “nucleic acid,” “nucleic acid molecule,” and “polynucleotide” are used interchangeably and refer to a sequence of nucleotides. A nucleic acid sequence can be a single-stranded or double-stranded deoxyribonucleotide or ribonucleotide of any length and includes coding and non-coding sequences of genes, exons, introns, sense and antisense complementary sequences, genomic DNA, cDNA, miRNA, siRNA, mRNA, rRNA, tRNA, recombinant nucleic acid sequences, isolated and purified naturally occurring DNA and / or RNA sequences, synthetic DNA and RNA sequences, fragments, primers, and nucleic acid probes. Those skilled in the art understand that the nucleic acid sequence of RNA is identical to that of DNA, except that thymine (T) is replaced by uracil (U). The term “nucleotide sequence” should also be understood to include polynucleotide or oligonucleotide molecules in the form of individual fragments or as components of larger nucleic acids.

[0314] As used in this article, the term “naturally occurring” for nucleic acids refers to a nucleic acid that is found in cells or organisms in nature and has not been intentionally modified by humans in a laboratory.

[0315] A “fragment” of a polynucleotide or nucleic acid sequence refers to a continuous sequence of nucleotides, particularly of a length of at least 15 bp, at least 30 bp, at least 40 bp, at least 50 bp, and / or at least 60 bp, according to one embodiment of this invention. Specifically, the polynucleotide fragment comprises at least 25, more particularly at least 50, more particularly at least 75, more particularly at least 100, more particularly at least 150, more particularly at least 200, more particularly at least 300, more particularly at least 400, more particularly at least 500, more particularly at least 600, more particularly at least 700, more particularly at least 800, more particularly at least 900, and more particularly at least 1000 consecutive nucleotides of a polynucleotide sequence according to one embodiment of this invention. Without limitation, the polynucleotide fragments described herein can be used as PCR primers and / or probes, or for antisense gene silencing or RNAi.

[0316] "Recombinant nucleic acid sequences" are nucleic acid sequences created by combining genetic material from more than one source using laboratory methods (such as molecular cloning), thereby creating or modifying nucleic acid sequences that are not naturally occurring and cannot be found in biological organisms in any other way.

[0317] “Recombinant DNA technology” refers to molecular biological methods used to prepare recombinant nucleic acid sequences, as described, for example, in Laboratory Manuals edited by Weigel and Glazebrook, 2002, Cold Spring Harbor LabPress; and Sambrook et al., 1989, Cold Spring Harbor, NY: Cold Spring Harbor LaboratoryPress.

[0318] The term "gene" refers to a DNA sequence containing a region that is operatively linked to a suitable regulatory region (e.g., a promoter) and transcribed into an RNA molecule (e.g., mRNA in a cell). Therefore, a gene can contain several operatively linked sequences, such as a promoter, a 5' leader sequence (containing, for example, a sequence involved in translation initiation), a coding region of cDNA or genomic DNA, introns, exons, and / or a 3' untranslated sequence (containing, for example, a transcription termination site).

[0319] A "chimeric gene" is any gene that is not normally found in species in nature, particularly a gene in which one or more portions of the nucleic acid sequence are unrelated in nature. For example, a promoter that is unrelated in nature to part or all of the transcribed region or to another regulatory region. The term "chimeric gene" should be understood to include expression constructs in which a promoter or transcriptional regulatory sequence is operatively linked to one or more coding sequences or antisense (i.e., the inverse complementary strand of the sense strand) or inverted repeat sequences (sense and antisense, whereby the RNA transcript forms a double-stranded RNA post-transcriptionally). The term "chimeric gene" also includes genes obtained by combining portions of one or more coding sequences to produce new genes.

[0320] "3'URT" or "3' untranslated sequence" (also known as "3' untranslated region" or "3' end") refers to a nucleic acid sequence found downstream of the gene coding sequence that contains, for example, a transcription termination site and (in most, but not all, eukaryotic mRNAs) a polyadenylation signal, such as AAUAAA or its variants. After transcription termination, the mRNA transcript can be cleaved downstream of the polyadenylation signal and a poly(A) tail can be added, which is involved in the transport of mRNA to the translation site, such as the cytoplasm.

[0321] The term "primer" refers to a short nucleic acid sequence that is hybridized to a template nucleic acid sequence and used for the polymerization of nucleic acid sequences complementary to that template.

[0322] The present invention also relates to nucleic acid sequences encoding polypeptides as defined herein.

[0323] In particular, the present invention also relates to nucleic acid sequences (single-stranded and double-stranded DNA and RNA sequences, such as cDNA, genomic DNA and mRNA) encoding one of the aforementioned polypeptides and their functional equivalents, which can be obtained, for example, by using artificial nucleotide analogs.

[0324] This invention also relates to nucleic acids that have a degree of "identity" with the sequences specifically disclosed herein. "Identity" between two nucleic acids refers to the identity of nucleotides along the entire length of the nucleic acid in each case.

[0325] The “identity” between two nucleotide sequences (and similarly, peptide or amino acid sequences) is a function of the number of nucleotide residues (or amino acid residues) when the two sequences are aligned, or the number of identical residues in both sequences. Identical residues are defined as the same residues at a given position in the alignment of the two sequences. The percentage of sequence identity used herein is calculated from the best alignment by dividing the number of identical residues between the two sequences by the total number of residues in the shortest sequence and multiplying by 100. The best alignment is the alignment with the highest probability of identity percentage. Vacancies can be introduced into one or more positions in the alignment of one or both sequences to obtain the best alignment. These vacancies are then considered as dissimilar residues used to calculate the percentage of sequence identity. Alignments used to determine the percentage of identity of amino acid or nucleic acid sequences can be performed in various ways using computer programs, such as those publicly available on the Internet. Specifically, the optimal alignment of a protein or nucleic acid sequence and the percentage of sequence identity can be obtained using the BLAST program (Tatiana et al., FEMS Microbiol Lett., 1999, 174:247-250, 1999), which is available from the National Center for Biotechnology Information (NCBI) at http: / / www.ncbi.nlm.nih.gov / BLAST / bl2seq / wblast2.cgi with default parameters. In another example, identity can be calculated using the Clustal method (Higgins DG, Sharp PM. ((1989))) of the Vector NTI Suite 7.1 program from Informax (USA) with the following settings:

[0326] Multiple alignment parameters:

[0327] 10 points deducted for opening gaps

[0328] 10 points deducted for gap extension

[0329] Gap separation deduction range 8

[0330] Gap separation deduction

[0331] The percentage of identity with the comparison delay is 40%.

[0332] Residue-specific gap-off

[0333] Hydrophilic residue gap

[0334] Transition weighted 0

[0335] Comparison parameters:

[0336] FAST algorithm open

[0337] K-tuple size 1

[0338] 3 points deducted for gap

[0339] Window size 5

[0340] Optimal number of diagonals: 5

[0341] Alternatively, identity can be determined according to the method of Chenna et al. (2003), webpage: http: / / www.ebi.ac.uk / Tools / clustalw / index.html# and the following settings:

[0342] DNA gap openness deduction 15.0

[0343] DNA gap extension deduction 6.66 points

[0344] DNA matrix identity

[0345] Open protein gaps - deduct 10.0 points

[0346] Protein gap extension deducts 0.2 points.

[0347] Gonnet protein matrix

[0348] Protein / DNA ENDGAP -1

[0349] Protein / DNA GAPDIST 4

[0350] All nucleic acid sequences (single-stranded and double-stranded DNA and RNA sequences, such as cDNA and mRNA) mentioned in this article can be produced from nucleotide structural units by chemical synthesis in a known manner, for example, by fragment condensation of individual overlapping complementary nucleic acid structural units of a double helix. The chemical synthesis of oligonucleotides can be carried out, for example, by the phosphoramide process (Voet, Voet, 2nd edition, Wiley Press, New York, pages 896-897) in a known manner. The accumulation of synthetic oligonucleotides, and the filling of vacancies in the ligation reaction using the Klenow fragment of DNA polymerase, as well as general cloning techniques, are described in Sambrook et al. (1989), see below.

[0351] The nucleic acid molecules according to the invention may additionally contain untranslated sequences from the 3' and / or 5' ends of the coding genetic region.

[0352] The present invention further relates to nucleic acid molecules that are complementary to the nucleotide sequences or segments thereof specifically described.

[0353] The nucleotide sequences according to the invention enable the generation of probes and primers that can be used to identify and / or clone homologous sequences in other cell types and organisms. Such probes or primers typically contain a nucleotide sequence region that hybridizes to at least about 12, preferably at least about 25, such as about 40, 50 or 75 consecutive nucleotides of the sense strand or corresponding antisense strand of the nucleic acid sequence according to the invention under “stringent” conditions (as defined elsewhere herein).

[0354] "Homologous" sequences include orthologous or paralogous sequences. Methods for identifying orthologous or paralogous sequences include phylogenetic methods, sequence similarity methods, and hybridization methods known in the art and described herein.

[0355] "Paralleloids," or paralogous sequences, arise from gene replication, resulting in two or more genes with similar sequences and functions. Paralogs typically cluster together and form through gene replication within related plant species. Paralogs are identified in groups of similar genes using pairwise BLAST analysis or procedures such as CLUSTAL during phylogenetic analysis of gene families. In paralogs, the shared sequence can be identified as a sequence characteristic of the related gene and possessing a similar function.

[0356] "Orthologs" or "orthologous sequences" are sequences that are similar to each other because they are found in species descended from a common ancestor. For example, plant species with a common ancestor are known to contain many enzymes with similar sequences and functions. For example, by constructing a phylogenetic tree of a gene family of a species using CLUSTAL or BLAST programs, technicians can identify orthologous sequences and predict the functions of orthologs. One method for identifying or confirming similar functions between homologous sequences is by comparing transcript profiles in host cells or organisms (such as plants or microorganisms) that overexpress or lack (in gene knockout / reduction) the relevant polypeptide. Technicians can understand that genes with similar transcript profiles (having a common transcript with greater than 50% regulation, or a common transcript with greater than 70% regulation, or a common transcript with greater than 90% regulation) will have similar functions. Homologs, paralogs, orthologs, and any other variants of the sequences described herein are expected to function in a similar manner by causing host cells, organisms such as plants or microorganisms to produce terpene synthase proteins.

[0357] The term "selectable marker" refers to any gene that, after expression, can be used to select one or more cells containing that selectable marker. Examples of selectable markers are described below. Those skilled in the art will understand that different antibiotic, fungicide, auxotrophic, or herbicide selectable markers may be applicable to different target species.

[0358] The present invention relates to isolated nucleic acid molecules encoding polypeptides or bioactive fragments thereof according to the present invention, and nucleic acid fragments that can be used, for example, as hybridization probes or primers to identify or amplify nucleic acids encoding the present invention.

[0359] "Isolated nucleic acid" or "isolated nucleic acid sequence" refers to a nucleic acid or nucleic acid sequence that is in an environment different from naturally occurring nucleic acids or nucleic acid sequences, and may include those that are substantially free of contaminating endogenous substances. Therefore, isolated nucleic acid molecules are separate from other nucleic acid molecules present in natural sources of nucleic acids, and if produced by recombinant technology, may be substantially free of other cellular material or culture medium, or if chemically synthesized, may be free of chemical precursors or other chemicals.

[0360] Nucleic acid molecules according to the invention can be isolated using standard molecular biology techniques and the sequence information provided according to the invention. For example, cDNA can be isolated from a suitable cDNA library using one of the specifically disclosed complete sequences or fragments thereof as hybridization probes and standard hybridization techniques (e.g., described in Sambrook, (1989)).

[0361] Alternatively, nucleic acid molecules containing one or a fragment of the disclosed sequence can be isolated by polymerase chain reaction using oligonucleotide primers constructed based on that sequence. The nucleic acid amplified in this manner can be cloned into a suitable vector and characterized by DNA sequencing. The oligonucleotides according to the invention can also be prepared using standard synthetic methods, for example, using an automated DNA synthesizer.

[0362] According to the nucleic acid sequences or derivatives thereof of the present invention, homologs or portions of these sequences can be isolated from other bacteria, for example, by conventional hybridization techniques or PCR techniques, through genomic or cDNA libraries. These DNA sequences hybridize with the sequences according to the present invention under standard conditions.

[0363] "Hybridization" refers to the ability of polynucleotides or oligonucleotides to bind to nearly complementary sequences under standard conditions, while non-complementary pairs do not bind nonspecifically. For this purpose, the sequences can be 90–100% complementary. This property of complementary sequences being able to bind specifically to each other is used for primer binding, for example, in Northern or Southern blotting, or in PCR or RT-PCR.

[0364] Short oligonucleotides in conserved regions are advantageously used for hybridization. However, longer fragments or complete sequences of the nucleic acids of the present invention may also be used for hybridization. These “standard conditions” vary depending on the nucleic acid used (oligonucleotide, longer fragment, or complete sequence) or the type of nucleic acid used for hybridization (DNA or RNA). For example, the melting temperature of DNA:DNA hybrids is about 10°C lower than that of DNA:RNA hybrids of the same length.

[0365] For example, depending on the specific nucleic acid, standard conditions refer to a temperature of 42 to 58°C in a buffered aqueous solution with a concentration of 0.1 to 5 x SSC (1 x SSC = 0.15 M NaCl, 15 mM sodium citrate, pH 7.2), or additionally in the presence of 50% formamide (e.g., 42°C, 5 x SSC, 50% formamide). Advantageously, hybridization conditions for DNA:DNA hybrids are 0.1 x SSC and a temperature of about 20°C to 45°C, preferably about 30°C to 45°C. For DNA:RNA hybrids, hybridization conditions are advantageously 0.1 x SSC and a temperature of about 30°C to 55°C, preferably about 45°C to 55°C. These hybridization temperatures are examples of calculated melting temperatures for nucleic acids of about 100 nucleotides in length and with a G+C content of 50% in the absence of formamide. The experimental conditions for DNA hybridization have been described in relevant genetics textbooks (e.g., Sambrook et al., 1989) and can be calculated using molecular formulas known to those skilled in the art, depending on factors such as nucleic acid length, hybrid type, or G+C content. More information on hybridization can be obtained from textbooks such as Ausubel et al. (eds), (1985), and Brown (ed) (1991).

[0366] “Hybridization” can be carried out under particularly stringent conditions. Such hybridization conditions are described, for example, in Sambrook (1989) or Current Protocols in Molecular Biology, John Wiley & Sons, NY (1989), 6.3.1–6.3.6.

[0367] As used herein, the term "hybridization" or "hybridization under certain conditions" is intended to describe the conditions under which hybridization and washing occur, under which significantly identical or homologous nucleotide sequences remain bound to each other. These conditions are such that sequences with at least about 70%, for example, at least about 80%, and for example, at least about 85%, 90%, or 95% identity remain bound to each other. Definitions of low-tightness, medium-tightness, and high-tightness hybridization conditions are provided herein.

[0368] Those skilled in the art can select suitable hybridization conditions with minimal experimentation, as illustrated, for example, by Ausubel et al. (1995, Current Protocols in Molecular Biology, John Wiley & Sons, sections 2, 4, and 6). Furthermore, stringent conditions are described by Sambrook et al. (1989, Molecular Cloning: A Laboratory Manual, 2nd ed., Cold Spring Harbor Press, chapters 7, 9, and 11).

[0369] As used herein, the low-tightness conditions are defined as follows. Filter membranes containing DNA were pretreated at 40°C for 6 hours in a solution containing 35% formamide, 5xSSC, 50 mM Tris-HCl (pH 7.5), 5 mM EDTA, 0.1% PVP, 0.1% Ficoll, 1% BSA, and 500 µg / ml denatured salmon sperm DNA. Hybridization was performed in the same solution, modified as follows: 0.02% PVP, 0.02% Ficoll, 0.2% BSA, 100 µg / ml salmon sperm DNA, 10% (wt / vol) dextran sulfate, and using 5-20x10 6 32P-labeled probe. The filter membrane was incubated in a hybridization mixture at 40°C for 18–20 h, followed by washing at 55°C for 1.5 h. In a solution containing 2x SSC, 25 mM Tris-HCl (pH 7.4), 5 mM EDTA, and 0.1% SDS, the membrane was replaced with fresh solution and incubated again at 60°C for 1.5 h. The filter membrane was then blotted dry and subjected to autoradiography.

[0370] As used herein, the moderately stringent conditions are defined as follows. Filter membranes containing DNA were pretreated at 50°C for 7 hours in a solution containing 35% formamide, 5xSSC, 50 mM Tris-HCl (pH 7.5), 5 mM EDTA, 0.1% PVP, 0.1% Ficoll, 1% BSA, and 500 µg / ml denatured salmon sperm DNA. Hybridization was performed in the same solution, modified as follows: 0.02% PVP, 0.02% Ficoll, 0.2% BSA, 100 µg / ml salmon sperm DNA, 10% (wt / vol) dextran sulfate, and using 5-20x10 632P-labeled probe. The filter membrane was incubated in a hybridization mixture at 50°C for 30 hours, followed by washing at 55°C for 1.5 hours. It was then incubated in a solution containing 2x SSC, 25 mM Tris-HCl (pH 7.4), 5 mM EDTA, and 0.1% SDS. Fresh solution was used instead of the washing solution, and the membrane was incubated again at 60°C for 1.5 hours. The filter membrane was then blotted dry and subjected to autoradiography.

[0371] As used herein, the stringent conditions are as follows: DNA-containing membranes were pre-hybridized at 65°C for 8 hours to overnight in a buffer consisting of 6x SSC, 50 mM Tris-HCl (pH 7.5), 1 mM EDTA, 0.02% PVP, 0.02% Ficoll, 0.02% BSA, and 500 µg / ml denatured salmon sperm DNA. The membranes were then pre-hybridized in a buffer containing 100 µg / ml denatured salmon sperm DNA and 5-20x10... 6 The filter membrane was hybridized in a prehybridization mixture of cpm 32P labeled probes at 65°C for 48 hours. The membrane was then washed at 37°C for 1 hour in a solution containing 2x SSC, 0.01% PVP, 0.01% Ficoll, and 0.01% BSA. It was then washed in 0.1x SSC at 50°C for 45 minutes.

[0372] If the above conditions are not suitable (e.g., for interspecific hybridization), other low, medium and high stringency conditions well known in the art (e.g., for interspecific hybridization) may be used.

[0373] A detection kit for the nucleic acid sequence encoding the polypeptide of the present invention may include primers and / or probes specific to the nucleic acid sequence encoding the polypeptide, and a protocol for using the primers and / or probes to detect the nucleic acid sequence encoding the polypeptide in a sample. Such a detection kit can be used to determine whether a plant, organism, microorganism, or cell has been modified, i.e., whether it has been transformed with the sequence encoding the polypeptide.

[0374] To test the function of a variant DNA sequence according to one embodiment of this article, the target sequence is operatively linked to an optional or screenable marker gene, and the expression of the reporter gene is tested in a transient expression analysis using microorganisms or protoplasts or in stably transformed plants.

[0375] The present invention also relates to derivatives of specifically disclosed or derivable nucleic acid sequences.

[0376] Therefore, the additional nucleic acid sequences according to the invention may be derived from the sequences specifically disclosed herein and may be distinguished by one or more (e.g., 1 to 10) nucleotides, such as 1 to 20, particularly 1 to 15 or 5 to 10, additions, substitutions, insertions or deletions, and may also encode a polypeptide having the desired profile of characteristics.

[0377] The invention also includes nucleic acid sequences containing so-called silent mutations or altered sequences, depending on the codon usage of a particular original or host organism, compared to the specifically stated sequences.

[0378] According to specific embodiments of the invention, variant nucleic acids can be prepared to adapt their nucleotide sequences to a particular expression system. For example, bacterial expression systems are known to express polypeptides more efficiently if the amino acids are encoded by specific codons. Due to the degeneracy of the genetic code, more than one codon can encode the same amino acid sequence, and multiple nucleic acid sequences can encode the same protein or polypeptide; all these DNA sequences are covered in one embodiment herein. Where appropriate, the nucleic acid sequence encoding the polypeptide described herein can be optimized to increase expression in host cells. For example, the nucleic acid of one embodiment herein can be synthesized using host-specific codons to improve expression.

[0379] The present invention also covers naturally occurring variants of the sequences described herein, such as splice variants or allelic variants.

[0380] The allele variants have at least 60% homology across the entire amino acid range at the derived amino acid level, preferably at least 80%, and very particularly preferably at least 90% homology (for details regarding homology at the amino acid level, please refer to the information given above for peptides). Advantageously, the homology can be even higher in certain regions of the sequence.

[0381] The present invention also relates to sequences that can be obtained by conserved nucleotide substitution (i.e., as a result, the amino acid in question is replaced by an amino acid having the same charge, size, polarity and / or solubility).

[0382] This invention also relates to molecules derived from specifically disclosed nucleic acids through sequence polymorphism. Such genetic polymorphism can exist in cells from different populations or from cells within a single population due to natural allelic variations. Allelic variants may also include functional equivalents. These natural variations typically produce changes of 1–5% in the nucleotide sequence of a gene. The polymorphism can lead to alterations in the amino acid sequence of the polypeptides disclosed herein. Allelic variants may also include functional equivalents.

[0383] Furthermore, derivatives should also be understood as homologs of the nucleic acid sequences according to the present invention, such as homologs of animals, plants, fungi, or bacteria, shortened sequences, single-stranded DNA or RNA encoding or non-coding DNA sequences. For example, at the DNA level, the homolog has at least 40%, preferably at least 60%, particularly preferably at least 70%, and very particularly preferably at least 80% homology in the entire DNA region given in the sequence specifically disclosed herein.

[0384] Furthermore, derivatives should be understood as, for example, fusions with promoters. Promoters added to the nucleotide sequence can be modified by at least one nucleotide exchange, at least one insertion, inversion, and / or deletion, without impairing the function or effectiveness of the promoter. Moreover, the effectiveness of promoters can be increased by altering their sequence, or by completely exchanging them with more effective promoters or even promoters from different genera of organisms.

[0385] e. Generation of functional peptide mutants

[0386] Furthermore, those skilled in the art are familiar with methods for generating functional mutants, namely, a nucleotide sequence encoding a polypeptide having at least 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity with any amino acid-related SEQ ID NO disclosed herein; and / or encoded by a nucleic acid molecule containing a nucleotide sequence having at least 70% sequence identity with any nucleotide-related SEQ ID NO disclosed herein.

[0387] Depending on the techniques used, those skilled in the art can introduce completely random or more targeted mutations into genes or non-coding nucleic acid regions (e.g., those important for regulating expression) and subsequently generate a genetic library. The molecular biological methods required for this purpose are known to those skilled in the art, for example, as described in Sambrook and Russell, Molecular Cloning, 3rd Edition, Cold Spring Harbor Laboratory Press, 2001.

[0388] Methods for modifying genes and thereby modifying the polypeptides encoded by them are known long to those skilled in the art, for example:

[0389] - Site-specific mutagenesis, in which single or multiple nucleotides of a gene are replaced in a directed manner (Trower MK (Ed.) 1996; In vitro mutagenesis protocols. Humana Press, New Jersey).

[0390] - Saturation mutagenesis, in which the codon for any amino acid can be exchanged or added at any site in the gene (Kegler-Ebo DM, Docktor CM, DiMaio D (1994) Nucleic Acids Res 22:1593; Barettino D, Feigenbutz M, Valcárel R, Stunnenberg HG (1994) Nucleic Acids Res 22:541; Barik S (1995) Mol Biotechnol 3:1),

[0391] - Error-prone polymerase chain reaction, in which the nucleotide sequence is mutated by error-prone DNA polymerase (Eckert KA, KunkelTA (1990) Nucleic Acids Res 18:3739);

[0392] -SeSaM method (sequence saturation method), in which preferred exchanges are prevented by polymerase. Schenk et al., Biospektrum, Vol. 3, 2006, 277-279.

[0393] - Gene propagation in mutant strains, where, for example, due to defects in DNA repair mechanisms, the mutation rate of nucleotide sequences increases (Greener A, Callahan M, Jerpseth B (1996) An efficient random mutagenesis technique using an E. coli mutator strain. In: Trower MK (Ed.) In vitro mutagenesis protocols. Humana Press, New Jersey), or

[0394] -DNA shuffling, in which a set of closely related genes are formed and digested, and these fragments are used as templates for polymerase chain reactions, in which the full-length mosaic gene is eventually generated through repeated strand separation and recombination (Stemmer WPC (1994) Nature 370:389; Stemmer WPC (1994) Proc Natl Acad Sci USA 91:10747).

[0395] Using so-called directed evolution (particularly described in Reetz MT and Jaeger KE (1999), TopicsCurr Chem 200:31; Zhao H, Moore JC, Volkov AA, Arnold FH (1999), Methods for optimizing industrial polypeptides by directed evolution, In: Demain AL, Davies JE (Ed.) Manual of industrial microbiology and biotechnology. American Society for Microbiology), skilled workers can generate functional mutants on a large scale in a directed manner. To this end, in the first step, gene libraries of the respective polypeptides are first generated, for example, using the methods given above. The gene libraries are expressed in a suitable manner, for example, through bacterial or phage display systems.

[0396] The relevant genes in the host organism expressing the functional mutant (whose function largely corresponds to the desired trait) can be submitted to another mutation cycle. The mutation and selection or screening steps can be repeated iteratively until the functional mutant of the present invention possesses a sufficient degree of the desired trait. Using this iterative process, a limited number of mutations, such as 1, 2, 3, 4, or 5 mutations, can be performed in stages, and their effects on the activity under study can be evaluated and selected. The selected mutants can then be subjected to further mutation steps in the same manner. This significantly reduces the number of individual mutants to be studied.

[0397] The results of this invention also provide important information regarding the structure and sequence of the relevant polypeptides, which is essential for the targeted generation of other polypeptides with desired modified properties. In particular, so-called "hot spots" can be defined as sequence segments potentially suitable for modification of properties by introducing targeted mutations.

[0398] Information about the location of amino acid sequences can also be derived, where mutations that may have little effect on activity can occur, and these can be designated as potential “silent mutations”.

[0399] f. Constructs for expressing the polypeptides of the present invention

[0400] In the context of this article, the following definitions apply:

[0401] "Gene expression" encompasses both "heterologous expression" and "overexpression," and involves gene transcription and the translation of mRNA into proteins. Overexpression refers to the production of gene products, measured as mRNA, peptide, and / or enzyme activity levels, in transgenic cells or organisms exceeding the levels found in non-transformed cells or organisms with similar genetic backgrounds.

[0402] As used herein, an "expression vector" refers to a nucleic acid molecule engineered using molecular biology methods and recombinant DNA technology to deliver foreign or exogenous DNA into a host cell. Expression vectors typically include the sequences required for correct transcription of the nucleotide sequence. The coding region usually encodes the target protein, but it can also encode RNA, such as antisense RNA, siRNA, etc.

[0403] As used herein, “expression vector” includes any linear or circular recombinant vector, including but not limited to viral vectors, bacteriophages, and plasmids. Those skilled in the art can select a suitable vector based on the expression system. In one embodiment, the expression vector includes a nucleic acid of the embodiments described herein, operably linked to at least one “regulatory sequence” that controls transcription, translation, initiation, and termination, such as a transcription promoter, operon, or enhancer, or an mRNA ribosome binding site, and optionally includes at least one selection marker. When the regulatory sequence functionally relates to the nucleic acid of the embodiments described herein, the nucleotide sequence is “operably linked.”

[0404] A "regulatory sequence" refers to a nucleic acid sequence that determines the expression level of the nucleic acid sequence in the embodiment described herein and can regulate the transcription rate of a nucleic acid sequence operatively linked to that regulatory sequence. Regulatory sequences include promoters, enhancers, transcription factors, promoter elements, etc.

[0405] A "promoter" is a nucleic acid sequence that controls the expression of a coding sequence by providing a binding site for RNA polymerase and other factors suitable for transcription, including but not limited to transcription factor binding sites, repressor and activator protein binding sites. The term "promoter" also includes the term "promoter regulatory sequence." Promoter regulatory sequences can include upstream and downstream elements that may affect transcription, RNA processing, or the stability of related coding nucleic acid sequences. Promoters include naturally derived and synthetic sequences. The coding nucleic acid sequence is typically located downstream of the promoter relative to the transcription start site.

[0406] In this context, "functional" or "operationally" linked is understood, for example, to refer to the sequential arrangement of one of the nucleic acids having a regulatory sequence. For example, a sequence with promoter activity, and the nucleic acid sequence to be transcribed, along with optional other regulatory elements (e.g., nucleic acid sequences that ensure transcription) and, for example, a terminator, arranged such that each regulatory element can perform its function after transcription of the nucleic acid sequence. This does not necessarily require a direct chemical link. Genetic control sequences, such as enhancer sequences, can even act on the target sequence from more distant locations or even from other DNA molecules. A preferred arrangement is one where the nucleic acid sequence to be transcribed is located downstream (i.e., at the 3' end) of the promoter sequence, thereby covalently linking the two sequences together. The distance between the promoter sequence and the nucleic acid sequence to be recombined can be less than 200 base pairs, or less than 100 base pairs, or less than 50 base pairs.

[0407] In addition to promoters and terminators, other examples of regulatory elements include: target sequences, enhancers, polyadenylation signals, selection markers, amplification signals, origins of replication, etc. Suitable regulatory sequences are described, for example, in Goeddel, Gene Expression Technology: Methods in Enzymology 185, Academic Press, San Diego, CA (1990).

[0408] The term "constitutive promoter" refers to an unregulated promoter that allows for the continuous transcription of the nucleic acid sequence to which it is operatively linked.

[0409] As used herein, the term "operably linked" refers to the linking of polynucleotide elements in a functional relationship. A nucleic acid is "operably linked" when it is in a functional relationship with another nucleic acid sequence. For example, if a promoter or transcriptional regulatory sequence can influence the transcription of a coding sequence, then the promoter or transcriptional regulatory sequence is operably linked to the coding sequence. Operable linking means that the linked DNA sequences are typically adjacent. The nucleotide sequence associated with the promoter sequence can be homologous or heterologous relative to the plant to be transformed. The sequence can also be wholly or partially synthetic. Regardless of origin, the nucleic acid sequence associated with the promoter sequence will be expressed or silenced depending on the nature of the promoter linked after binding to the polypeptide of the embodiments described herein. The associated nucleic acid can encode a protein that needs to be expressed or repressed throughout the organism or in a specific tissue, cell, or cell compartment at all times or alternatively at specific times. This nucleotide sequence specifically encodes a protein that confers the desired phenotypic trait to the host cell or organism altered (genetically modified) or transformed by it. More specifically, the associated nucleotide sequence results in the production of one or more desired products as defined herein in cells or organisms, such as, in particular, physostigmine and / or sterol, or mixtures containing physostigmine and / or sterol, or mixtures containing physostigmine and / or sterol and one or more terpenes. In particular, the nucleotide sequence encodes a polypeptide having enzymatic activity as defined herein, such as, in particular, terpene synthase.

[0410] As used herein, “expression system” encompasses any combination of nucleic acid molecules required to express one, or co-express two or more, polypeptides in vivo or in vitro in a given expression host. The respective coding sequences may reside on a single nucleic acid molecule or vector, such as a vector containing multiple cloning sites, or on a polycistronic nucleic acid, or may be distributed across two or more physically distinct vectors.

[0411] As used herein, the terms “amplifying” and “amplification” refer to the use of any suitable amplification method to generate or detect recombinants of naturally expressed nucleic acids, as described in detail below. For example, the present invention provides methods and reagents (e.g., specific degenerate oligonucleotide primer pairs, oligo-dT primers) for amplifying (e.g., by polymerase chain reaction, PCR) naturally expressed (e.g., genomic DNA or mRNA) or recombinant nucleic acids (e.g., cDNA) of the present invention in vivo, in vitro, or in vitro.

[0412] The nucleotide sequences described above can be part of an "expression cassette". The terms "expression cassette" and "expression construct" are used synonymously. A (preferred recombinant) expression construct contains a nucleotide sequence that encodes a polypeptide according to the invention and is under the genetic control of a regulatory nucleic acid sequence.

[0413] In the method applied according to the present invention, the expression cassette may be an "expression vector", particularly a part of a recombinant expression vector.

[0414] According to the present invention, "expression unit" should be understood as a nucleic acid with expressive activity, which contains a promoter as defined herein, and regulates expression upon functional linkage with a nucleic acid or gene to be expressed, i.e., transcription and translation of said nucleic acid or gene. Therefore, it is also referred to in this respect as a "regulatory nucleic acid sequence". In addition to promoters, other regulatory elements, such as enhancers, may also be present.

[0415] According to the present invention, an "expression cassette" or "expression construct" should be understood as an expression unit functionally linked to a nucleic acid or gene to be expressed. Therefore, in contrast to an expression unit, an expression cassette contains not only nucleic acid sequences that regulate transcription and translation, but also nucleic acid sequences that are expressed as proteins due to transcription and translation.

[0416] In the context of this invention, the terms "expression" or "overexpression" describe the generation or increase of intracellular activity of one or more polypeptides encoded by corresponding DNA in a microorganism. For this purpose, for example, a gene may be introduced into the organism, an existing gene may be replaced with another gene, the copy number of a gene may be increased, a strong promoter may be used, or a gene encoding a corresponding polypeptide with high activity may be used. Optionally, these measures may be combined.

[0417] Preferably, such constructs according to the invention include a promoter upstream of the respective coding sequence 5' and a terminator sequence downstream of the respective coding sequence 3', as well as optionally other common regulatory elements, each operatively connected to the coding sequence.

[0418] The nucleic acid constructs according to the invention specifically comprise a sequence encoding a polypeptide, such as derived from the amino acid-related SEQ ID NO or its inverse complementary sequence as described herein, or derivatives and homologs thereof, and is operatively or functionally linked to one or more regulatory signals for advantageous control, for example, increasing gene expression.

[0419] In addition to these regulatory sequences, the natural regulation of these sequences may still exist before the actual structural genes, and optionally may have been genetically modified so that natural regulation has been turned off and gene expression is enhanced. However, nucleic acid constructs can also have simpler constructions, i.e., no additional regulatory signals are inserted before the coding sequence, and the natural promoter with regulatory function has not been removed. Instead, the natural regulatory sequences are mutated so that regulation no longer occurs and gene expression increases.

[0420] Preferred nucleic acid constructs advantageously also include one or more previously mentioned "enhancer" sequences functionally linked to a promoter, which enable enhanced expression of the nucleic acid sequence. Other advantageous sequences, such as other regulatory elements or terminators, may also be inserted at the 3' end of the DNA sequence. One or more copies of the nucleic acid according to the invention may be present in the construct. Optionally, other markers, such as genes complementary to auxotrophic or antibiotic resistance genes, may also be present in the construct for selection.

[0421] Examples of suitable regulatory sequences exist in promoters, such as cos, tac, trp, tet, trp-tet, lpp, lac, lpp-lac, lacIq, T7, T5, T3, gal, trc, ara, rhaP (rhaP BAD SP6, lambda-P R Or lambda-P L In promoters, they are advantageously used in Gram-negative bacteria. Other advantageous regulatory sequences are found, for example, in the Gram-positive promoters amy and SpO2, and in yeast or fungal promoters ADC1, MFalpha, AC, P-60, CYC1, GAPDH, TEF, rp28, and ADH. Artificial promoters can also be used for regulation.

[0422] To facilitate expression in a host organism, nucleic acid constructs are advantageously inserted into vectors, such as plasmids or phages, enabling optimal gene expression in the host. Besides plasmids and phages, vectors should be understood to include all other vectors known to those skilled in the art, such as viruses like SV40, CMV, baculoviruses and adenoviruses, transposons, IS elements, phages, granules, and linear or circular DNA or artificial chromosomes. These vectors are capable of autonomous replication in the host organism or replication via chromosomes. These vectors represent a further development of the invention. Binary or co-integrative vectors are also applicable.

[0423] Suitable plasmids include, for example, those for *E. coli* pLG338, pACYC184, pBR322, pUC18, pUC19, pKC30, pRep4, pHS1, pKK223-3, pDHE19.2, pHS2, pPLc236, pMBL24, pLG200, pUR290, and pIN-III. 113The plasmids listed above are a small selection of possible plasmids, found in *Streptomyces* pIJ101, pIJ364, pIJ702, or pIJ361; in *Bacillus* pUB110, pC194, or pBD214; in *Corynebacterium* pSA77 or pAJ667; in fungi pALS1, pIL2, or pBB116; in yeast 2alphaM, pAG-1, YEp6, YEp13, or pEMBLYe23; or in plants pLGV23, pGHlac+, pBIN19, pAK2004, or pDH51. Other plasmids are well known to technicians and can be found, for example, in the book Cloning Vectors (Eds. Pouwels PH et al. Elsevier, Amsterdam-New York-Oxford, 1985, ISBN 0 444 904018).

[0424] In further development of the vector, vectors containing the nucleic acid constructs of the present invention or the nucleic acids of the present invention can also be advantageously introduced into microorganisms in the form of linear DNA and integrated into the genome of the host organism via heterologous or homologous recombination. This linear DNA can consist of linearized vectors such as plasmids, or solely of the nucleic acid constructs or nucleic acids of the present invention.

[0425] For optimal expression of heterologous genes in an organism, it is advantageous to modify the nucleic acid sequence to match the specific “codon usage” used in the organism. “Codon usage” can be readily determined by computer evaluation of other known genes in the organism under discussion.

[0426] The expression cassette according to the invention is generated by fusing a suitable promoter to a suitable coding nucleotide sequence and a terminator or polyadenylation signal. Conventional recombination and cloning techniques are used for this purpose, as described in, for example, T. Maniatis, EF Fritsch and J. Sambrook, Molecular Cloning: A Laboratory Manual, Cold Spring Harbor Laboratory, Cold Spring Harbor, NY (1989) and TJ Silhavy, ML Berman and LW Enquist, Experiments with Gene Fusions, Cold Spring Harbor Laboratory, Cold Spring Harbor, NY (1984) and Ausubel, FM et al., Current Protocols in Molecular Biology, Greene Publishing Assoc. and Wiley Interscience (1987).

[0427] To facilitate expression in a suitable host organism, recombinant nucleic acid constructs or gene constructs are advantageously inserted into host-specific vectors, enabling optimal gene expression in the host. Vectors are well-known to those skilled in the art and can be found, for example, in "cloning vectors" (Pouwels PH et al., Ed., Elsevier, Amsterdam-New York-Oxford, 1985).

[0428] Alternative embodiments of the present invention provide a method for “altering (modifying) gene expression” in host cells. For example, in certain contexts (e.g., exposure to certain temperatures or culture conditions), polynucleotides of the present invention can be enhanced, overexpressed, or induced in host cells or host organisms.

[0429] The altered expression of the polynucleotides described herein can also result in ectopic expression, which is a different expression pattern in altered and control or wild-type organisms. The alteration in expression occurs due to contact of the peptide of one embodiment of this invention with an exogenous or endogenous modulator or due to chemical modification of the peptide. The term also refers to the altered expression pattern of the polynucleotides of the embodiments described herein, which is altered to below detectable levels or completely inhibited in activity.

[0430] In one embodiment, this document also provides isolated, recombinant, or synthetic polynucleotides encoding the polypeptide or variant polypeptide provided herein.

[0431] In one embodiment, multiple nucleic acid sequences encoding polypeptides are co-expressed in a single host, particularly under the control of different promoters. In another embodiment, multiple nucleic acid sequences encoding polypeptides may be present on a single transformation vector, or separate vectors may be used and transformants containing two chimeric genes may be selected for simultaneous co-transformation. Similarly, one or more polypeptide-encoding genes may be expressed together with other chimeric genes in a single plant, cell, microorganism, or organism.

[0432] g. Host suitable for this invention

[0433] Depending on the context, the term "host" can refer to a wild-type host or a genetically modified recombinant host, or both.

[0434] In principle, all prokaryotes or eukaryotes can be considered as hosts or recombinant host organisms for the nucleic acids or nucleic acid constructs according to the present invention.

[0435] Using the vector according to the invention, a recombinant host can be generated, which can be transformed, for example, with at least one vector according to the invention, and can be used to generate peptides according to the invention. Advantageously, the recombinant construct according to the invention as described above is introduced into a suitable host system and expressed. Preferably, common cloning and transfection methods known to those skilled in the art, such as co-precipitation, protoplast fusion, electroporation, retroviral transfection, etc., are used to express the nucleic acid in their respective expression systems. Suitable systems are described in Current Protocols in Molecular Biology, F. Ausubel et al., Ed., Wiley Interscience, New York 1997, or Sambrook et al. Molecular Cloning: A Laboratory Manual. 2nd edition, Cold Spring Harbor Laboratory, Cold Spring Harbor Laboratory Press, Cold Spring Harbor, NY, 1989.

[0436] Advantageously, microorganisms such as bacteria, fungi, or yeasts are used as host organisms. Advantageously, Gram-positive or Gram-negative bacteria are used, preferably those belonging to the families Enterobacteriaceae, Pseudomonadaceae, Rhizobiaceae, Streptomycetaceae, Streptococcaceae, or Nocardiaceae, particularly Escherichia, Pseudomonas, Streptomyces, Lactococcus, Nocardia, Burkholderia, Salmonella, Agrobacterium, Clostridium, or Rhodococcus. The genus and species *Escherichia coli* are particularly preferred. Furthermore, other advantageous bacteria have been found in the alpha-proteobacteria, beta-proteobacteria, or gamma-proteobacteria groups. Advantageously, yeasts such as *Saccharomyces* or the Pichia family are also suitable hosts.

[0437] Alternatively, the entire plant or plant cell can be used as a natural or recombinant host. As non-limiting examples, the following plants or cells derived from them may be mentioned: the genus *Nicotiana*, particularly *Nicotiana abenthamiana* and *Nicotiana tabacum* (tobacco); and the genus *Arabidopsis*, particularly *Arabidopsis thaliana*.

[0438] Depending on the host organism, the organism used in the method according to the invention is grown or cultured in a manner known to those skilled in the art. Culture can be carried out in batches, semi-batch, or continuously. Nutrients can be provided at the start of fermentation or later, semi-continuously, or continuously. This is also described in more detail below.

[0439] h. Recombinant generation of the polypeptide according to the invention

[0440] The present invention further relates to a method for recombinantly generating polypeptides or functional biologically active fragments thereof according to the invention, wherein microorganisms that generate polypeptides are cultured, and polypeptide expression is optionally induced by applying at least one inducer for gene expression, and the polypeptides are isolated from the culture. If desired, the polypeptides can also be produced on an industrial scale in this manner.

[0441] The microorganisms produced according to the present invention can be cultured continuously or discontinuously using batch culture, fed-batch culture, or repeated fed-batch culture. An overview of known culture methods can be found in Chmiel's textbook (Bioprozesstechnik 1. Einführungin die Bioverfahrenstechnik [Bioprocess technology 1. Introduction to bioprocess technology] (Gustav Fischer Verlag, Stuttgart, 1991)) or Storhas's textbook (Bioreaktoren und periphere Einrichtungen [Bioreactors and peripheralequipment] (Vieweg Verlag, Braunschweig / Wiesbaden, 1994)).

[0442] The culture medium used must be appropriately suited to the requirements of each strain. Descriptions of culture media for various microorganisms are provided in the "Manual of Methods for General Bacteriology" of the American Society for Bacteriology (Washington DC, USA, 1981).

[0443] These culture media, which can be used according to the present invention, typically contain one or more carbon sources, nitrogen sources, inorganic salts, vitamins and / or trace elements.

[0444] Preferred carbon sources are sugars, such as monosaccharides, disaccharides, or polysaccharides. Good carbon sources include, for example, glucose, fructose, mannose, galactose, ribose, sorbitol, ribulose, lactose, maltose, sucrose, raffinose, starch, or cellulose. Sugars can also be added to the culture medium via complex compounds (e.g., molasses) or other byproducts of saccharification. Adding mixtures of different carbon sources is also advantageous. Other possible carbon sources are oils and fats, such as soybean oil, sunflower oil, peanut oil, and coconut oil; fatty acids such as palmitic acid, stearic acid, or linoleic acid; alcohols such as glycerol, methanol, or ethanol; and organic acids such as acetic acid or lactic acid.

[0445] Nitrogen sources are typically organic or inorganic nitrogen compounds or materials containing these compounds. Examples of nitrogen sources include ammonia or ammonium salts, such as ammonium sulfate, ammonium chloride, ammonium phosphate, ammonium carbonate, or ammonium nitrate, nitrates, urea, amino acids, or complex nitrogen sources, such as corn steep liquor, soy flour, soy protein, yeast extract, meat extract, etc. Nitrogen sources can be used alone or in combination.

[0446] Inorganic salt compounds that can be present in culture media include chlorides, phosphorus, or sulfates of calcium, magnesium, sodium, cobalt, molybdenum, potassium, manganese, zinc, copper, and iron.

[0447] Inorganic sulfur-containing compounds, such as sulfates, sulfites, dithionites, tetrasulfites, thiosulfates, sulfides, and organic sulfur compounds, such as thiols and thiols, can be used as sulfur sources.

[0448] Phosphoric acid, potassium dihydrogen phosphate, or dipotassium hydrogen phosphate, or their corresponding sodium-containing salts, can be used as phosphorus sources.

[0449] Chelating agents can be added to the culture medium to retain metal ions in solution. Particularly suitable chelating agents include dihydroxyphenols, such as catechol or protocatechuic acid esters, or organic acids, such as citric acid.

[0450] The fermentation medium used according to the present invention typically also contains other growth factors, such as vitamins or growth promoters, including, for example, biotin, riboflavin, thiamine, folic acid, niacin, pantothenic acid, and pyridoxine. The growth factors and salts are often derived from components of complex culture media, such as yeast extract, molasses, corn steep liquor, etc. Furthermore, suitable precursors may be added to the medium. The exact composition of the compounds in the medium depends largely on the specific experiment and is determined individually for each case. Information on medium optimization can be found in the textbook "Applied Microbiol. Physiology, A Practical Approach" (Ed. PM Rhodes, PF Stanbury, IRL Press (1997) pp. 53-73, ISBN 0 19 963577 3). Growth media are also available from commercial suppliers such as Standard 1 (Merck) or BHI (Brain Heart Infusion, DIFCO).

[0451] All components of the culture medium are sterilized by heating (at 1.5 bar and 121°C for 20 minutes) or by aseptic filtration. These components can be sterilized together or individually as needed. All components of the culture medium can be given at the start of the culture or added continuously or in batches.

[0452] The culture temperature is typically between 15°C and 45°C, preferably between 25°C and 40°C, and can be varied or kept constant during the experiment. The pH of the medium should be in the range of 5 to 8.5, preferably around 7.0. The pH during growth can be controlled by adding alkaline compounds (e.g., sodium hydroxide, potassium hydroxide, ammonia, or ammonia solution) or acidic compounds (e.g., phosphoric acid or sulfuric acid). Antifoaming agents such as fatty acid polyethylene glycol esters can be used to control foaming. To maintain plasmid stability, suitable selective substances such as antibiotics can be added to the culture medium. To maintain aerobic conditions, oxygen or an oxygen-containing gas mixture (e.g., ambient air) is supplied to the culture. The culture temperature is typically in the range of 20°C to 45°C. Continue culturing until the maximum amount of the desired product is formed. This usually takes between 10 and 160 hours.

[0453] The fermentation broth is then further processed. Depending on the needs, the biomass can be completely or partially removed from the fermentation broth, or it can be left entirely in it, by separation techniques such as centrifugation, filtration, decantation, or a combination of these methods.

[0454] If the polypeptide is not secreted in the culture medium, the cells can also be lysed, and the product can be obtained from the lysate using known methods for protein separation. Cells can be optionally destroyed by high-frequency ultrasound, high pressure (e.g., in a French press), by osmosis, by the action of detergents, lysing enzymes, or organic solvents, by a homogenizer, or by a combination of these methods.

[0455] Peptides can be purified using known chromatographic techniques, such as molecular sieve chromatography (gel filtration), Q-agarose chromatography, ion exchange chromatography, and hydrophobic chromatography, as well as other conventional techniques such as ultrafiltration, crystallization, salting out, dialysis, and natural gel electrophoresis. Suitable methods are described, for example, in Cooper, TG, Biochemische Arbeitsmethoden [Biochemical Processes], Verlag Walter de Gruyter, Berlin, New York, or Scopes, R., Protein Purification, Springer Verlag, New York, Heidelberg, Berlin.

[0456] For the isolation of recombinant proteins, the use of a carrier system or oligonucleotide may be advantageous, which extends cDNA by a defined nucleotide sequence and thus encodes an altered polypeptide or fusion protein, for example, for easier purification. Suitable modifications of this type are, for example, so-called “tags” that act as anchors, such as modifications known as hexahistine anchors or epitopes that can be recognized as antibody antigens (e.g., described in Harlow, E. and Lane, D., 1988, Antibodies: A Laboratory Manual. Cold Spring Harbor (NY) Press). These anchors can be used to attach proteins to solid supports, such as polymer matrices, which can be used, for example, as packing material in chromatographic columns, or on microtiter plates or other supports.

[0457] These anchors can also be used to identify proteins. To identify proteins, conventional markers, such as fluorescent dyes, enzyme markers (which react with a substrate to form a detectable reaction product), or radiolabels, can be used alone or in combination with anchors to derivatize proteins.

[0458] i. Immobilization of peptides

[0459] The enzymes or polypeptides according to the invention can be used in free form or immobilized in the methods described herein. Immobilized enzymes are enzymes immobilized on an inert support. Suitable support materials and enzymes immobilized thereon are known from EP-A-1149849, EP-A-1069183, and DE-OS 100193773 and the references cited therein. In this regard, reference is made to the full disclosure of these documents. Suitable support materials include, for example, clay, clay minerals such as kaolinite, diatomaceous earth, perlite, silica, alumina, sodium carbonate, calcium carbonate, cellulose powder, anion exchanger materials, synthetic polymers such as polystyrene, acrylic resins, phenolic resins, polyurethanes, and polyolefins such as polyethylene and polypropylene. For the preparation of loaded enzymes, the support material is generally used in the form of finely divided particles, preferably porous. The particle size of the support material is generally no greater than 5 mm, particularly no greater than 2 mm (particle size distribution curve). Similarly, when using dehydrogenases as whole-cell catalysts, either free or immobilized form can be selected. Carrier materials include, for example, calcium alginate and carrageenan. Enzymes and cells can also be directly cross-linked with glutaraldehyde (cross-linked with CLEAs). Corresponding and other immobilization techniques are described, for example, in J. Lalonde and A. Margolin, "Immobilization of Enzymes" in K. Drauz and H. Waldmann, EnzymeCatalysis in Organic Synthesis 2002, Vol. III, 991-1032, Wiley-VCH, Weinheim. Further information on biotransformation and bioreactors for carrying out the methods according to the invention is given in Rehm et al. (Ed.) Biotechnology, 2nd Edn, Vol 3, Chapter 17, VCH, Weinheim.

[0460] j. Reaction conditions of the biocatalytic production method of the present invention

[0461] "Enzymatic catalysis" or "biocatalysis" refers to methods carried out under the catalysis of enzymes (including enzyme mutants) as defined herein. Therefore, this method can be carried out in the presence of the enzyme in isolated (purified, enriched) or crude form, or in the presence of cellular systems, particularly natural or recombinant microbial cells containing the active form of the enzyme and capable of catalyzing the transformation reactions disclosed herein. Thus, the reactions of this invention can be carried out under in vivo or in vitro conditions.

[0462] At least one polypeptide / enzyme present during a single step of the method of the present invention or a multi-step method as defined above may be naturally present in living cells, or in harvested cells (i.e., under in vivo conditions (biotransformation)), dead cells, permeabilized cells, crude cell extracts, purified extracts, or in a substantially pure or completely pure form (i.e., under in vitro conditions (biochemical synthesis)), or recombinantly produced as one or more enzymes. The at least one enzyme may be present in solution or as an enzyme immobilized on a carrier. One or more enzymes may be present simultaneously in soluble and / or immobilized forms.

[0463] The method according to the invention can be carried out in common reactors known to those skilled in the art and can be carried out on various scales, from laboratory scale (a few milliliters to tens of liters of reaction volume) to industrial scale (a few liters to thousands of cubic meters of reaction volume). A chemical reactor can be used if the polypeptide is used in the form of encapsulation through non-living, optionally permeabilized cells, as a more or less purified cell extract, or in a purified form. Chemical reactors typically allow control of the amount of at least one enzyme, the amount of at least one substrate, pH, temperature, and the circulation of the reaction medium. When at least one polypeptide / enzyme is present in living cells, the process will be fermentation. In this case, biocatalytic production will be carried out in a bioreactor (fermenter) where parameters necessary for suitable survival conditions for living cells (e.g., nutrient-rich culture medium, temperature, aeration, aerobic or anaerobic or other gases, antibiotics, etc.) can be controlled. Those skilled in the art are familiar with chemical or bioreactors, for example, procedures for scaling up chemical or biotechnological methods from laboratory to industrial scale or optimizing process parameters, which are also extensively described in the literature (for biotechnological methods, see, for example, Crueger und Crueger, Biotechnologie – Lehrbuch der angewandten Mikrobiologie, 2 Ed., R. Oldenbourg Verlag, Munich, Wien, 1984).

[0464] Cells containing at least one enzyme can be permeated by physical or mechanical means such as ultrasound or radiofrequency pulses, high-pressure cell lysis (French press), or by chemical means such as hypotonic media present in the culture medium, lysing enzymes, and detergents, or a combination of these methods. Examples of detergents are digitoxin, n-dodecyl maltodextrin, octyl glycoside, Triton® X-100, Tween® 20, deoxycholate, CHAPS (3-[(3-chloroamidopropyl)dimethylammonium]-1-propanesulfonate), Nonidet® P40 (ethylphenol poly(ethylene glycol ether)), etc.

[0465] Instead of living cells, non-living cell biomass containing the desired biocatalyst can also be used in the biotransformation reaction of this invention.

[0466] If at least one enzyme is immobilized, it is ligated to an inert carrier as described above.

[0467] The conversion reaction can be carried out in batches, semi-batches, or continuously. Reactants (and optional nutrients) can be provided at the start of the reaction, or they can be provided subsequently in a semi-continuous or continuous manner.

[0468] Depending on the specific reaction type, the reactions of this invention can be carried out in aqueous, aqueous-organic, or non-aqueous reaction media.

[0469] Aqueous or aqueous-organic media may contain suitable buffer solutions to adjust the pH to 5 to 11, such as 6 to 10.

[0470] In aqueous-organic media, organic solvents that are miscible, partially miscible, or immiscible with water can be used. Non-limiting examples of suitable organic solvents are listed below. Further examples are mono- or poly-aryl, aromatic or aliphatic alcohols, particularly poly-aliphatic alcohols such as glycerol.

[0471] Non-aqueous media may contain substantially no water, that is, will contain less than about 1% by weight or 0.5% by weight of water.

[0472] Biocatalytic methods can also be carried out in organic non-aqueous media. Suitable organic solvents include, for example, aliphatic hydrocarbons having 5 to 8 carbon atoms, such as pentane, cyclopentane, hexane, cyclohexane, heptane, octane, or cyclooctane; aromatic hydrocarbons, such as benzene, toluene, xylene, chlorobenzene, or dichlorobenzene; aliphatic acyclic and ethers, such as diethyl ether, methyl tert-butyl ether, ethyl tert-butyl ether, dipropyl ether, diisopropyl ether, dibutyl ether; or mixtures thereof.

[0473] The concentration of reactants / substrate can be adapted to optimal reaction conditions, which can depend on the specific enzyme being applied. For example, the initial substrate concentration can be 0.1 to 0.5 M, or, for example, 10 to 100 mM.

[0474] The reaction temperature can be adapted to optimal reaction conditions, which can depend on the specific enzyme used. For example, the reaction can be carried out at temperatures ranging from 0 to 70°C, such as 20 to 50 or 25 to 40°C. Examples of reaction temperatures are approximately 30°C, approximately 35°C, approximately 37°C, approximately 40°C, approximately 45°C, approximately 50°C, approximately 55°C, and approximately 60°C.

[0475] The process can continue until equilibrium is reached between the substrate and the subsequent product, but it can be stopped earlier. Typical process times range from 1 minute to 25 hours, particularly from 10 minutes to 6 hours, for example from 1 hour to 4 hours, and particularly from 1.5 hours to 3.5 hours. These parameters are non-limiting examples of suitable process conditions.

[0476] If the host is a genetically modified plant, it can provide optimal growth conditions, such as optimal light, water, and nutrients.

[0477] l. Product separation

[0478] The method of the present invention may further include the step of recovering a final product or intermediate product, which may optionally be a substantially pure form of a stereoisomer or enantiomer. The term "recovery" includes the extraction, harvesting, separation, or purification of a compound from a culture medium or reaction medium. The recovery of a compound can be carried out according to any conventional separation or purification method known in the art, including but not limited to treatment with conventional resins (e.g., anion or cation exchange resins, nonion adsorption resins, etc.), treatment with conventional adsorbents (e.g., activated carbon, silica, silica gel, cellulose, alumina, etc.), pH alteration, solvent extraction (e.g., using conventional solvents such as alcohols, ethyl acetate, hexane, etc.), distillation, dialysis, filtration, concentration, crystallization, recrystallization, pH adjustment, lyophilization, etc.

[0479] The identity and purity of the separated products can be determined by known techniques, such as high performance liquid chromatography (HPLC), gas chromatography (GC), spectroscopy (e.g., IR, UV, NMR), staining methods, TLC, NIRS, enzyme or microbial assays (see, for example: Patek et al. (1994) Appl. Environ. Microbiol. 60:133-140; Malakhova et al. (1996) Biotekhnologiya 11 27-32; und Schmidt et al. (1998) BioprocessEngineer. 19:67-70. Ullmann's Encyclopedia of Industrial Chemistry (1996) Bd.A27, VCH: Weinheim, S. 89-90, S. 521-540, S. 540-547, S. 559-566, 575-581 undS.). 581-587; Michal, G (1999) Biochemical Pathways: An Atlas of Biochemistry and Molecular Biology, John Wiley and Sons; Fallon, A. et al. (1987) Applications of HPLC in Biochemistry in: Laboratory Techniques in Biochemistry and Molecular Biology, Bd. 17.).

[0480] Cyclic terpenoids produced by any of the methods described herein can be converted into derivatives, such as, but not limited to, hydrocarbons, esters, amides, glycosides, ethers, epoxides, aldehydes, ketones, alcohols, diols, acetals, or ketals. Terpenoid derivatives can be obtained by chemical methods, such as, but not limited to, oxidation, reduction, alkylation, acylation, and / or rearrangement. Alternatively, terpenoid derivatives can be obtained by biochemical methods by contacting the terpenoid with an enzyme, such as, but not limited to, oxidoreductases, monooxygenases, dioxygenases, and transferases. Biochemical transformation can be performed in vitro using isolated enzymes, enzymes derived from lysed cells, or in vivo using whole cells.

[0481] m. Fermentation of bis(phosphonic acid) and / or bis(phosphonic acid), and / or phosphonic acid and / or phosphonic acid.

[0482] The present invention also relates to a method for fermentation to produce bis(phosphonic acid) and / or bis(phosphonic acid), and / or bis(phosphonic acid) and / or bis(phosphonic acid).

[0483] The term “fermentation production” or “fermentation” refers to the ability of microorganisms (assisted by enzyme activity contained in or produced by said microorganisms) to produce compounds in cell cultures using at least one carbon source added to incubation.

[0484] The term “fermentation broth” should be understood as referring to a liquid, particularly an aqueous solution or aqueous / organic solution, that is based on a fermentation process and has been subjected to or has been subjected to post-processing, such as that described herein.

[0485] The fermentation used according to the invention can be carried out, for example, in stirred fermenters, bubble columns, and loop reactors. For a comprehensive overview of possible method types, including stirrer types and geometries, see "Chmiel: Bioprozesstechnik: Einfuhrung in die Bioverfahrenstechnik, Band 1". Typical variations available in the method of the invention are those known to those skilled in the art or explained, for example, in "Chmiel, Hammes and Bailey: Biochemical Engineering", such as batch, fed-batch, repeatedly fed-batch, or continuous fermentation, with or without biomass recovery. Depending on the production strain, air, oxygen, carbon dioxide, hydrogen, nitrogen, or suitable gas mixtures can be injected to achieve good yields (YP / S).

[0486] The “yield” and / or “conversion” of the reaction according to the invention are determined within a specified time period, such as 4, 6, 8, 10, 12, 16, 20, 24, 36 or 48 hours (during which the reaction takes place). In particular, the reaction is carried out under precisely defined conditions, such as the “standard conditions” as defined herein.

[0487] Different yield parameters (“yield” or YP / S; “specific productivity yield”; or space-time yield (STY)) are well known in the art and are measured as described in the literature.

[0488] "Yield" and "YP / S" (both expressed as the quality of products produced / the quality of materials consumed) are used as synonyms in this document.

[0489] Specific productivity yield describes the amount of product produced per hour per liter of fermentation broth per gram of biomass. Wet cell weight (WCW) describes the number of biologically active microorganisms in the biochemical reaction. This value is given as grams of product per g WCW per hour (i.e., g / gWCW). -1 h -1 Alternatively, the amount of biomass can also be expressed as the amount of dry cell weight, denoted as DCW. Furthermore, by measuring 600 nm (OD... 600 By estimating the corresponding wet or dry cell weights using the optical density at a given location and relevant factors determined experimentally, biomass concentration can be determined more easily.

[0490] The culture medium used according to the present invention must meet the requirements of the specific strain in an appropriate manner. Descriptions of culture media for various microorganisms are given in the "Manual of Methods for General Bacteriology" of the American Society for Bacteriology (Washington DC, USA, 1981).

[0491] These culture media, which can be used according to the present invention, typically contain one or more carbon sources, nitrogen sources, inorganic salts, vitamins and / or trace elements.

[0492] Preferred carbon sources are sugars, such as monosaccharides, disaccharides, or polysaccharides. Very good carbon sources include, for example, glucose, fructose, mannose, galactose, ribose, sorbitol, ribulose, lactose, maltose, sucrose, raffinose, starch, or cellulose. Sugars can also be added to the culture medium via complex compounds (e.g., molasses) or other byproducts of saccharification. Adding mixtures of various carbon sources is also advantageous. Other possible sources of carbon are oils and fats, such as soybean oil, sunflower oil, peanut oil, and coconut oil; fatty acids such as palmitic acid, stearic acid, or linoleic acid; alcohols such as glycerol, methanol, or ethanol; and organic acids such as acetic acid or lactic acid.

[0493] Nitrogen sources are typically organic or inorganic nitrogen compounds or materials containing these compounds. Examples of nitrogen sources include ammonia or ammonium salts, such as ammonium sulfate, ammonium chloride, ammonium phosphate, ammonium carbonate, or ammonium nitrate, nitrates, urea, amino acids, or complex nitrogen sources, such as corn steep liquor, soy flour, soy protein, yeast extract, meat extract, etc. Nitrogen sources can be used alone or in combination.

[0494] Inorganic salt compounds that may be present in the medium include chlorides, phosphates, or sulfates of calcium, magnesium, sodium, cobalt, molybdenum, potassium, manganese, zinc, copper, and iron.

[0495] Inorganic sulfur-containing compounds, such as sulfates, sulfites, dithionites, tetrasulfites, thiosulfates, and sulfides, as well as organic sulfur compounds, such as thiols and thiols, can all be used as sulfur sources.

[0496] Phosphoric acid, potassium dihydrogen phosphate, or dipotassium hydrogen phosphate, or their corresponding sodium-containing salts, can be used as phosphorus sources.

[0497] Chelating agents can be added to the culture medium to retain metal ions in solution. Particularly suitable chelating agents include dihydroxyphenols, such as catechol or protocatechuic acid esters, or organic acids, such as citric acid.

[0498] The fermentation medium used according to the present invention may also contain other growth factors, such as vitamins or growth promoters, including, for example, biotin, riboflavin, thiamine, folic acid, niacin, pantothenic acid, and pyridoxine. Growth factors and salts are typically derived from complex components of the medium, such as yeast extract, molasses, corn steep liquor, etc. Additionally, suitable precursors may be added to the medium. The precise composition of compounds in the medium depends heavily on the specific experiment and must be determined individually for each specific case. Information on medium optimization can be found in the textbook "Applied Microbiol. Physiology, A Practical Approach" (1997). Growth media are also available from commercial suppliers, such as Standard 1 (Merck) or BHI (Brain Heart Infusion, DIFCO), etc.

[0499] All components of the culture medium are sterilized by heating (at 1.5 bar and 121°C for 20 minutes) or by aseptic filtration. These components can be sterilized together or individually as needed. All components of the culture medium can be given at the start of growth, or can be added continuously or in batches.

[0500] The culture temperature is typically between 15°C and 45°C, preferably between 25°C and 40°C, and can be kept constant or varied during the experiment. The pH of the medium should be in the range of 5 to 8.5, preferably around 7.0. The pH during growth can be controlled by adding alkaline compounds (e.g., sodium hydroxide, potassium hydroxide, ammonia, or ammonia solution) or acidic compounds (e.g., phosphoric acid or sulfuric acid). Antifoaming agents such as fatty acid polyethylene glycol esters can be used to control foaming. To maintain plasmid stability, suitable substances with selective action, such as antibiotics, can be added to the culture medium. To maintain aerobic conditions, oxygen or a mixture of oxygen-containing gases (e.g., ambient air) is supplied to the culture. The culture temperature is typically between 20°C and 45°C. Continue culturing until the maximum amount of the desired product is formed. This usually takes between 1 and 160 hours.

[0501] The method of the present invention may further include the step of recovering bisphosphonate and / or bisphosphonate, and / or bisphosphonate and / or bisphosphonate alcohol.

[0502] The term "recovery" includes the extraction, harvesting, separation, or purification of compounds from a culture medium. The recovery of compounds can be performed using any conventional separation or purification method known in the art, including but not limited to treatment with conventional resins (e.g., anion or cation exchange resins, nonion adsorption resins, etc.), treatment with conventional adsorbents (e.g., activated carbon, silica, silica gel, cellulose, alumina, etc.), pH alteration, solvent extraction (e.g., using conventional solvents such as alcohols, ethyl acetate, hexane, etc.), distillation, dialysis, filtration, concentration, crystallization, recrystallization, pH adjustment, lyophilization, etc.

[0503] Before the intended separation, the biomass in the fermentation broth can be removed. Methods for removing biomass are known to those skilled in the art, such as filtration, sedimentation, and flotation. Therefore, biomass can be removed, for example, by centrifuges, separators, decanters, filters, or in flotation equipment. To maximize the recovery of valuable products, washing the biomass, for example by percolation, is generally recommended. The choice of method depends on the biomass content and properties in the fermentation broth, as well as the interaction between the biomass and the valuable products.

[0504] In one implementation, the fermentation broth can be sterilized or pasteurized. In another implementation, the fermentation broth is concentrated. This concentration can be carried out in batches or continuously, depending on the needs. Pressure and temperature ranges should be selected to ensure that product damage is not caused and to minimize equipment and energy consumption. Skillful selection of pressure and temperature levels for multi-stage evaporation can be particularly energy-efficient.

[0505] The following examples are illustrative only and are not intended to limit the scope of the implementation schemes described herein.

[0506] After considering the disclosure provided herein, a variety of possible variations that will immediately become apparent to those skilled in the art also fall within the scope of this invention.

[0507] Experimental Section

[0508] The invention will now be described in more detail through the following embodiments.

[0509] Material:

[0510] Unless otherwise stated, all chemical and biochemical materials, as well as microorganisms or cells used in this article, are commercially available products.

[0511] Unless otherwise stated, recombinant proteins are cloned and expressed using standard methods, such as those described, for example, in Sambrook, J., Fritsch, EF and Maniatis, T., Molecular cloning: A Laboratory Manual, 2 nd Edition, Cold Spring Harbor Laboratory, Cold Spring Harbor Laboratory Press, Cold Spring Harbor, NY, 1989.

[0512] method:

[0513] GC / MS

[0514] GC / MS analysis was performed on a GC / MS 6890N / 5975 (Agilent). Separation was achieved using a DB-1MS column (30m × 0.25mm inner diameter × 0.25µm film thickness, J & W Scientific, Agilent Technologies, Foster City, CA). Helium was used as the carrier gas at a flow rate of 0.7 mL / min. Split injection (1:5) mode was used, and the injector temperature was set to 250 °C. The oven temperature program was set to initially increase from 50 °C (hold for 5 minutes) to 300 °C at a rate of 5 °C / min, then increase to 340 °C at a rate of 50 °C / min and hold for 3 minutes. Data were acquired using a scan mode of m / z 29–450. Alkanes (C6–C4) were injected. 30 The retention index is calculated. Compounds are identified by matching mass spectrometry and retention indices with an internal library.

[0515] Example 1

[0516] Discovery of HAD-like TPS

[0517] Fresh leaves of *Bazzania* (Nagobase ID: PA-2018-0168 to PA-2018-0173, containing bismuth sub-type sesquiterpene alcohols, including bismuth sub-type alcohol and bismuth sub-type alcohol) were obtained from Hunan, China. Total RNA was extracted from the fresh leaves using a QIAGEN RNeasyPlant Mini Kit (50) 74904 and used for RNA sequencing. RNA sequencing data were then assembled using the CLC Genomics Workbench (7.6.1). Sequences of fern HAD-like TPS (DfHAD) from prior art (PCT / EP2019 / 063824; SEQ ID NO: 73 and 74 herein) were used to search for potential bismuth sub-type sesquiterpene synthases in the protein sequence data of *Bazzania*. Novel sequences were discovered from two assemblages of the whiting transcriptome using BLAST, resulting in HAD-like TPS named BazzHAD1 (SEQ ID NO:3), BazzHAD2 (SEQ ID NO:6), and BazzHAD3 (SEQ ID NO:9); their corresponding nucleotide sequences are SEQ ID NO:1, 4, and 7, respectively.

[0518] The raw RNAseq data of Bazzania trilobata (SRA ID: ERR364415) were downloaded from NCBI and assembled by the CLC Genomics Workbench to provide the transcriptome, from which BtHAD (SEQ ID NO:12) was discovered by BLAST; its corresponding nucleotide sequence is SEQ ID NO:10.

[0519] Homologous sequences of DfHAD in the NCBI database were thoroughly mined using BLAST with a threshold score (E value < 10, hit list size < 500) until the plant gene was observed. This yielded two similar sequences from *Selaginella moellendorffii*, which were deposited as hypothetical proteins with accession numbers EFJ10816.1 and EFJ26126.1 (GenBank ID), and were unrelated to the HAD-like TPS. In this study, the proteins were named SmHAD1 (SEQ ID NO: 19) and SmHAD2 (SEQ ID NO: 26), respectively; their corresponding nucleotide sequences are SEQ ID NO: 17 and 24, respectively.

[0520] To investigate their homology, SmHAD1, SmHAD2, BtHAD, BazzHAD1, BazzHAD2 and BazzHAD3, as well as DfHAD (PCT / EP2019 / 063824), and several fungal HAD-like bismuth sesquiterpene synthases (including AstC (PCT / EP2019 / 063824)) were compared using Clustalw. We found that the protein sequences of SmHAD1 and SmHAD2 have very low identity (approximately 14–20%) with all other known HAD-like TPS, but they show high homology (>96%) among themselves. BtHAD and BazzHAD1 show slightly higher homology (approximately 22–30%) with fern and fungal HAD-like TPS, with 96% homology among themselves. BazzHAD2 and BazzHAD3 show moderate homology (approximately 22–43%) with fern and fungal HAD-like TPS, with 67% homology among themselves (Table 1).

[0521] Table 2. Pairwise comparisons of BazzHAD1, BazzHAD2, BazzHAD3, BtHAD, SmHAD1, SmHAD2, and known HAD-like TPS. This comparison demonstrates the identity of the aligned sequences between proteins.

[0522]

[0523] Example 2

[0524] Functional expression and characterization of plant HAD-like TPS in Escherichia coli expression system (in vivo biochemical analysis)

[0525] The coding sequences of BazzHAD1, BazzHAD2, BazzHAD3, BtHAD, SmHAD1 (EFJ10816.1), and SmHAD2 (EFJ26126.1) were optimized using Genscript according to their frequency of use of genetic codons in E. coli (SEQ ID NO: 2, 5, 8, 11, 18, and 25, respectively). They were then synthesized in vitro and subcloned into the pETDuet-1 plasmid (Novagen) for subsequent expression in E. coli, generating plasmids pETDuet-BazzHAD1, pETDuet-BazzHAD2, pETDuet-BazzHAD3, pETDuet-BtHAD, pETDuet-SmETHAD1, and pETDuet-SmHAD2, respectively.

[0526] BL21 (DE3) *E. coli* cells (Tiangen) were co-transformed with the plasmid pACYC / ScMVA, which contains genes encoding the heterologous mevalonic acid (MVA) pathway, as well as the respective plasmids pETDuet-BazzHAD1, pETDuet-BazzHAD2, pETDuet-BazzHAD3, pETDuet-BtHAD, pETDuet-SmHAD1, and pETDuet-SmHAD2. The pACYC / ScMVA plasmid was constructed for expressing the farnesyl pyrophosphate (FPP) synthase gene and the complete MVA pathway genes. In short, the eight biosynthetic genes of the MVA pathway are divided into two synthetic operons, referred to as the "upper" and "lower" MVA pathways. As part of the "upper" MVA pathway, a synthetic operon was created, consisting of an acetyl-CoA thiolytic enzyme encoded by atoB from *E. coli*, HMG-CoA synthases encoded by ERG13 and ERG19 from *Saccharomyces cerevisiae*, and a truncated form of HMG-CoA reductase, respectively. This operon converts the major metabolite acetyl-CoA to (R)-mevalerate. As part of the "lower" mevalerate pathway, a second synthetic operon was created, encoding mevalerate kinase (ERG12, *Saccharomyces cerevisiae*), phosphate mevalerate kinase (ERG8, *Saccharomyces cerevisiae*), phosphate mevalerate decarboxylase (MVD1, *Saccharomyces cerevisiae*), isopentenyl diphosphate isomerase (idi, *E. coli*), and FPP synthase (IspA, *E. coli*). Finally, a second FPP synthase (ERG20) from *Saccharomyces cerevisiae* was introduced into the upper pathway operon to improve the conversion of isoprene C5 units (IPP and DMAPP) to FPP. Under the control of the bacterial phage T7 promoter (pACYCDuet-1, Invitrogen), each operon is subcloned into one of the multiple cloning sites of a low-copy expression plasmid, providing plasmid pACYC / ScMVA. Therefore, this plasmid contains genes encoding all the enzymes of the biosynthetic pathway from acetyl-CoA to FPP.

[0527] Select co-transformed cells on LB agar plates containing ampicillin (50 µg / mL) and chloramphenicol (34 µg / mL). Inoculate single colonies into 5 mL of liquid LB medium supplemented with the same antibiotics. Incubate the culture overnight at 37 °C with shaking at 200 rpm. The next day, inoculate 0.2 mL of the overnight culture into 2 mL of TB medium supplemented with the above antibiotics and glycerol (final concentration 3% w / v). Incubate at 37 °C for 5 h with shaking at 200 rpm, then cool the culture to 25 °C and shake at 200 rpm for another hour. Then add IPTG (final 0.1 mM) to each tube and cover the culture with 200 µL of decane. Incubate the culture at 25 °C for another 48 h with shaking at 200 rpm, then extract twice with 1 volume of ethyl acetate. Before analyzing the samples by GC / MS, add 50 µL of 2 mg / mL isophorene as an internal standard to the organic phase. Helium was used as the carrier gas at a constant flow rate of 0.7 mL / min. The sample was injected in split (1:25) mode, with the injector temperature set to 250 °C. The oven temperature program was set to initially increase from 50 °C (hold for 5 minutes) to 300 °C at a rate of 5 °C / min, then increase to 340 °C at a rate of 50 °C / min and hold for 3 minutes. Product identification was based on mass spectrometry and retention index.

[0528] GC / MS analysis showed that recombinant cells expressing BazzHAD1 and BtHAD produced 37.3 mg / L and 26.0 mg / L of sterol, respectively; expression of BazzHAD2 resulted in the production of 35.3 mg / L of sterol and 0.2 mg / L of an unknown sesquiterpene 1 (unknown SQT1); expression of BazzHAD3 resulted in the production of 422.7 mg / L of sterol; expression of SmHAD1 resulted in the production of 0.5 mg / L of sterol and 6.2 mg / L of sterol; and expression of SmHAD2 resulted in the production of 1.4 mg / L of sterol and 10.7 mg / L of sterol (Figures 1-16).

[0529] Example 3

[0530] The importance of class II motifs and QW motifs in plant HAD-like TPS (functional evidence of enzyme mutants and motifs)

[0531] To confirm the importance of the class II terpene synthase motif (SEQ ID NO:46), mutants of this motif were designed and generated for BtHAD, SmHAD1, and SmHAD2 (SEQ ID NO:13, 20, and 27, respectively). Since BazzHAD1 is an ortholog of BtHAD from *Bazzania trilobata* (which shares 96% identity), it can be represented as BtHAD and is therefore not included in this study. In addition to the class II motif, the QW motif (SEQ ID NO:51) was also mutated in BtHAD, SmHAD1, and SmHAD2 (SEQ ID NO:15, 22, and 29, respectively) to investigate its importance to synthase activity. Since known fungal HAD-like TPS contain class I and class II motifs (e.g., SEQ ID NO:56, 57, and 58 of WO2018 / 220113), EMD37666.1 (SEQ ID NO:31) and XP_007369631.1 (SEQ ID NO:35) were selected and their class II motifs were mutated (SEQ ID NO:33 and 37, respectively) to investigate whether class I motifs were sufficient to produce synthase activity. In mutants of the class II synthase motif in BtHAD (D320A / D322A), SmHAD1 (D257A / D259A), SmHAD2 (D257A / D259A), EMD37666.1 (D276A / S277A / D279A), and XP_007369631.1 (D272A / D273A / D275A), polar Asp is replaced by Ala (SEQ ID NO: 39 to 42). In mutants of the QW motif in BtHAD (D491A), SmHAD1 (E432A / D433A), and SmHAD2 (E432A / D433A), polar Glu and Asp are replaced by Ala (SEQ ID NO: 43 and 44).

[0532] The sequences of BtHAD, SmHAD1, SmHAD2, XP_007369631.1, EMD37666.1, and their corresponding class II synthase motifs and (if present) QW motif mutants were optimized by Genscript according to the genetic codon frequency table of E. coli (BtHAD SEQ ID NO: 11, 13, 15; SmHAD1 SEQ ID NO: 18, 20, 22; SmHAD2 SEQ ID NO: 25, 27, 29; EMD37666.1 SEQ ID NO: 31, 33; and XP_007369631.1 SEQ ID NO: 35, 37). The sequence was then synthesized in vitro using Genscript and subcloned into the pETDuet-1 plasmid (Novagen) for subsequent expression in E. coli, generating plasmids pETDuet-BtHAD, pETDuet-SmHAD1, pETDuet-SmHAD2, pETDuet-EMD37666.1, pETDuet-XP_007369631.1; pETDuet-BtHAD-ClassII mut, pETDuet-SmHAD1-ClassII mut, pETDuet-SmHAD2-ClassII mut, pETDuet-EMD37666.1-ClassII mut, pETDuet-XP_007369631.1-ClassII mut; and pETDuet-BtHAD-QW mut, pETDuet-SmHAD1-QW mut, and pETDuet-SmHAD2-QW mut.

[0533] BL21 (DE3) Escherichia coli cells (Tiangen) were co-transformed with the plasmid pACYC / ScMVA, which contains the gene encoding the heterologous mevalonate pathway, and the respective plasmids pETDuet-BtHAD, pETDuet-SmHAD1, pETDuet-SmHAD2, pETDuet-EMD37666.1, pETDuet-XP_007369631.1, pETDuet-BtHAD-ClassII mut, pETDuet-SmHAD1-ClassII mut, pETDuet-SmHAD2-ClassII mut, pETDuet-EMD37666.1-ClassII mut, pETDuet-XP_007369631.1-ClassII mut, pETDuet-BtHAD-QWmut, and pETDuet-SmHAD1-QW. mut and pETDuet-SmHAD2-QW mut. The construction of pACYC / ScMVA for expressing the FPP synthase gene and the gene for the complete MVA pathway is as described above (Example 2). This plasmid contains genes encoding all enzymes in the biosynthetic pathway from acetyl-CoA to FPP.

[0534] Select co-transformed cells on LB agar plates containing ampicillin (50 µg / mL) and chloramphenicol (34 µg / mL). Inoculate single colonies into 5 mL of liquid LB medium supplemented with the same antibiotics. Incubate the culture overnight at 37 °C with shaking at 200 rpm. The next day, inoculate 0.2 mL of the overnight culture into 2 mL of TB medium supplemented with the above antibiotics and glycerol (final concentration 3% w / v). Incubate at 37 °C for 5 h with shaking at 200 rpm, then cool the culture to 25 °C and shake at 200 rpm for another hour. Then add IPTG (final 0.1 mM) to each tube and cover the culture with 200 µL of decane. Incubate the culture at 25 °C for another 48 h with shaking at 200 rpm, then extract twice with 1 volume of ethyl acetate. Before analyzing the samples by GC / MS, add 50 µL of 2 mg / mL isophorene as an internal standard to the organic phase.

[0535] BtHAD, SmHAD1, SmHAD2, fungal HAD-like TPS, and their mutants were analyzed in vivo biochemical assays. Nephrolol and bismuth subtilisol were not detected in any of the enzymes with mutated class II motifs. These results support the hypothesis of protonation-induced FPP cyclization involving class II motifs, rather than classical ionization-induced cyclization mediated by class I motifs. Figure 19We also observed that mutants of the QW motif resulted in a 90% or higher loss of productivity in zosophyll and bismuth subtilis, which may be related to enzyme integrity and thermostability.

[0536] Example 4

[0537] Characterization of plant HAD-like TPS in in vitro systems (specified as evidence of phytosporin and phytoalkenyl diphosphates) according to).

[0538] In vitro biochemical assays were performed to investigate the cyclization mechanism of the newly identified plant HAD-like styranne sesquiterpene synthase.

[0539] BL21 (DE3) *E. coli* cells (Tiangen) were transformed with plasmids pETDuet-SmHAD1, pETDuet-SmHAD2, and pETDuet-BtHAD, respectively. Transformed cells were selected on LB agar plates containing ampicillin (50 µg / mL). Single colonies were inoculated into 25 mL of liquid LB medium supplemented with the same antibiotic. Cultures were incubated at 37°C and 200 rpm for 5 hours until an OD of approximately 0.5 was reached. 600 Then cool to 20°C and shake at 200 rpm for 0.5 hours. Then add IPTG (final 0.1 mM) to each tube and incubate the culture at 20°C and 200 rpm for another 18 hours.

[0540] For in vitro biochemical analysis, 25 mL of *E. coli* cultures expressing SmHAD1, SmHAD2, or BtHAD were centrifuged and resuspended in 5 mL of 50 mM Tris-HCl pH 8.0 buffer (10 mM MgCl2, 5 mM DTT), and then sonicated to produce crude cell lysates. 29 µM FPP was added to 1 mL of each crude cell lysate for reaction, and the mixture was incubated at 30 °C and 50 rpm for 2 h. The sample was then aliquoted, with one half containing 20 µL of bacterial alkaline phosphatase (BAP; Sangon-B004081-100) and reacted for an additional 1 h at 25 °C (BAP can remove potential diphosphate groups from complementene and / or zephyranthesyl diphosphates to generate complementol and / or zephyranthesyl alcohol). A control (the latter half of the sample) without BAP was incubated under the same conditions. 10 µL of 2 mg / mL isophorene (internal standard) was added to the organic phase of 0.5 mL aliquots as an internal standard. The reactants were then extracted with 0.25 mL of ethyl acetate and analyzed by GC / MS, as described above.

[0541] Adding BAP to the reactions using SmHAD2 and BtHAD lysates resulted in a 3.5-fold and 1.8-fold increase in the amount of complementol, respectively. The presence of complementol in the BAP-free reactions could be explained by the activity of *E. coli*'s native phosphatase. The reaction using the SmHAD1 lysate did not produce observable complementol, but it was detectable with the addition of BAP, albeit in small amounts. Similarly, only trace amounts of physostigmine were detected in both the reactions using the SmHAD2 lysate and those with the addition of BAP. Figure 20 (and Table 3).

[0542] Table 3. Productivity of SmHAD1 and SmHAD2 in in vitro biochemical analysis

[0543]

[0544] These findings support the hypothesis of a protonation-induced cyclization mechanism for the biosynthesis of bis(oxo)-enyl and zedoaryl diphosphates.

[0545] All publications mentioned in this application are incorporated by reference to disclose and describe the methods and / or materials relating to those publications.

[0546] sequence list

[0547] Table 4. Sequences described and used here

[0548]

[0549]

[0550]

[0551] SEQ ID NO:1

[0552] Bazzania sp. BazzHAD1 wt nucleic acid sequence

[0553]

[0554] SEQ ID NO:2

[0555] Bazzania sp. BazzHAD1, an optimized strain of *Escherichia coli*.

[0556]

[0557] SEQ ID NO: 3

[0558] *Bazzania* sp. _BazzHAD1_ Amino acid sequence

[0559] MAPPFGFIGGNATSVTTSPDVLDYNYDGGPPTGQYKSLILDIGGVLLQPNLDRSYGTIVPPVQFKRMLRSHIWFNYSMGKMSAEEVFRQLSVKSGYPAEDIRSFGRKARESLVPITQITDLLFKLSKESKVRLFCMTNCPAEDFEYLSQAYPKLFGLFERIFTSASTGMRKPNRNFFRHVLQETGIRASETIFVDDLVPNIMAAEALDFTAIHVSPESPAETVRLLMFHLRKPEERLAAARDYLLRNRGVDKALSLTSDGVRVQVYFDHFVVAEVLNDERFLPITAAPESGTMNFFPKEERKATRDITTNAVTHVDLPDDLDSTSVALSVLYKFSKVGMETVQRVLDMMEKRVDDDGIFQTYFDPLRLRVDPVAATNTVYLFHLGGRPGPTQQTEQFLSKLLEKRSYLRGTLYYRTPEAFLYQLTRLVVSFEEYFRKTGFLEMLKKRLSERVGIESDAFTLAMRILACVNCGIPVAELNRDMDRLARMQNVDGSWDVCPYYSYNDPKSWFGNELLTTAFAAAALEQSEPNFN

[0560] SEQ ID NO: 4

[0561] *Bazzania* sp. _BazzHAD2_wt_ Nucleic acid sequence

[0562]

[0563] SEQ ID NO:5

[0564] Bazzania sp. BazzHAD2 Escherichia coli, optimized

[0565]

[0566] SEQ ID NO: 6

[0567] Bazzania sp. BazzHAD2 amino acid sequence

[0568] MERPHFDTLIMDLGGVLVDFSLQTSTSTSVKSLKFVLRSATWAAYECGHMSESDCYAAVAKDLGSSANASQVGEAISEVRKSLQVNEDLIGVIRELKAQNGLRVYVMSNIPQPDFDAVKAKSASWGVFDGMYPSYAVGTHKPDLAFYRHVLEETNTDPLKAIFVDDSLQNVIAARSFGLTGIIYNDTVNAARTLRNLLGDPILRGEQYLSSHAGQLLSISDTGTPFPDNFSQLLILDATGNRDLVVLENPQRTWNYFIGKPVLTSETFPDDLDTTSIALMALNVDLEVANSVMDEMLEFRNSDGLFLTYFDETRPRVDAVVNVNVLRLFHRHEREREAQQPLEWITNVLTHRAYVDGTLFYYHAESFLYFLSRLFCENSSVQSRFQELLEQRLRERIGTPGDALSLAMRALACQMLGIDSSTDVRALLPLQCDDGGWEAGWVCRYGSNGMRVGSRGYTTALAINAIRGARVSR

[0569] SEQ ID NO: 7

[0570] Bazzania sp. BazzHAD3_wt nucleic acid sequence

[0571]

[0572] SEQ ID NO:8

[0573] Bazzania sp. BazzHAD3 Escherichia coli, optimized

[0574]

[0575] SEQ ID NO:9

[0576] biāntáishǔ(Bazzania sp.)_BazzHAD3_ānjīsuānxùliè

[0577] MPALSSHMEPPNFDTLILDLGDVLVGGSLQGSYALAMKRILKLSLRTPTWGKYECGQLSELDCYDAVAEDLGNSTSSAQVGEVVSEARKSLQVNEDLIKVLQELKAENNLRVYVMSNIPKPDLAVVKAKSVNWGVIDGWYPSYAVGFHKPDLAFYRHVLEETNTDPLKAIFVDDKVQNVIVARSFGMTGIVFKDTKSTTKELKNLGDPVKRGEHFLSAHAKQLESVCGDGTTFLQLDNFAQ LILHDTGNSDLVARQHHGRTWNYFIGKPMGTTDTFPNDLDTTSIALLTLNVDSEVARSVLDEMLLYTSSDGLVEVYFDKTRPRVDPVVCVNVLRLFSKYGRELQLQKTLDWVTEVLVHRAYIDGTLFYYHAESFLYFLSCLYKENPRLQTQFREPLQERLRERIGQPGDALCLAMRAIACQTVGIGNVIDVQALLPLQSSDGGWEAGWVCRMGTSGVPVGNRGVTTALCIRGASSYHT

[0578] SEQ ID NO:10

[0579] biāntái(Bazzania trilobata)_BtHAD_wt_hésuānxùliè

[0580]

[0581] SEQ ID NO:11

[0582] Bazzania trilobata BtHAD Escherichia coli, optimized

[0583]

[0584] SEQ ID NO: 12

[0585] Amino acid sequence of BtHAD from *Bazzania trilobata*

[0586] MAPPFGYIGGNESSVTTSPDVLDYSYDGGPPTGQYKSMILDIGGVLLQPNLDRSYGTIVPPVQFKRMLRSHIWFNYSMGKVSAEEVFRQLSVKSGYPAEDIRSFARKARESLVPITQITDLLVKLSKESKIRLFCMTNCPAEDFAYLSQAYPNLFGLFERIFTSASTGMRKPNRNFFRHVLQETGISATETIFVDDLVPNIMAAEALDFTAIHVSPESPAETVRLLMFHLRKPEERLAAARDYLLRNRGVDKALSLTSDGVRVQVYFDHFVVAEVLNDERFLPITAAPKSGTMNFFPKEERNATRDMTTNAVTHVDLPDDLDSTSVALSVLYKFSKVAMETVQRVLDMMEKRVDDDGIFQTYFDPLRLRVDPVAATNTVYLFHLGGRPGPTQQTEQFLSKLLEKRSYVRGTLYYRTPEAFLYQLTRLVVSFEYFRKTGFLEMLKKRLSERVGIESDAFTLAMRILACVNCGIPVAELNRDMDRLARMQNVDGSWDVCPYYSYNDPKSWFGNELLTTAFAAAALEHSEPNFK

[0587] SEQ ID NO: 13

[0588] Optimized BtHAD mutant of class II motif, artificial sequence, expressed in *Escherichia coli*

[0589]

[0590] SEQ ID NO: 14

[0591] Artificial Sequence_BtHAD Mutant of Class II Motif

[0592] MAPPFGYIGGNESSVTTSPDVLDYSYDGGPPTGQYKSMILDIGGVLLQPNLDRSYGTIVPPVQFKRMLRSHIWFNYSMGKVSAEEVFRQLSVKSGYPAEDIRSFARKARESLVPITQITDLLVKLSKESKIRLFCMTNCPAEDFAYLSQAYPNLFGLFERIFTSASTGMRKPNRNFFRHVLQETGISATETIFVDDLVPNIMAAEALDFTAIHVSPESPAETVRLLMFHLRKPEERLAAARDYLLRNRGVDKALSLTSDGVRVQVYFDHFVVAEVLNDERFLPITAAPKSGTMNFFPKEERNATRDMTTNAVTHVDLPDALASTSVALSVLYKFSKVAMETVQRVLDMMEKRVDDDGIFQTYFDPLRLRVDPVAATNTVYLFHLGGRPGPTQQTEQFLSKLLEKRSYVRGTLYYRTPEAFLYQLTRLVVSFEYFRKTGFLEMLKKRLSERVGIESDAFTLAMRILACVNCGIPVAELNRDMDRLARMQNVDGSWDVCPYYSYNDPKSWFGNELLTTAFAAAALEHSEPNFK

[0593] SEQ ID NO: 15

[0594] Artificial Sequence_BtHAD Mutant of QW Motif_Escherichia coli, optimized

[0595]

[0596] SEQ ID NO: 16

[0597] Artificial sequence_BtHAD mutant of QW motif

[0598] MAPPFGYIGGNESSVTTSPDVLDYSYDGGPPTGQYKSMILDIGGVLLQPNLDRSYGTIVPPVQFKRMLRSHIWFNYSMGKVSAEEVFRQLSVKSGYPAEDIRSFARKARESLVPITQITDLLVKLSKESKIRLFCMTNCPAEDFAYLSQAYPNLFGLFERIFTSASTGMRKPNRNFFRHVLQETGISATETIFVDDLVPNIMAAEALDFTAIHVSPESPAETVRLLMFHLRKPEERLAAARDYLLRNRGVDKALSLTSDGVRVQVYFDHFVVAEVLNDERFLPITAAPKSGTMNFFPKEERNATRDMTTNAVTHVDLPDDLDSTSVALSVLYKFSKVAMETVQRVLDMMEKRVDDDGIFQTYFDPLRLRVDPVAATNTVYLFHLGGRPGPTQQTEQFLSKLLEKRSYVRGTLYYRTPEAFLYQLTRLVVSFEYFRKTGFLEMLKKRLSERVGIESDAFTLAMRILACVNCGIPVAELNRDMDRLARMQNVAGSWDVCPYYSYNDPKSWFGNELLTTAFAAAALEHSEPNFK

[0599] SEQ ID NO: 17

[0600] Selaginella moellendorffii_SmHAD1_wt_nucleic acid sequence

[0601]

[0602] SEQ ID NO:18

[0603] Selaginella moellendorffii_SmHAD1_Escherichia coli, optimized

[0604]

[0605] SEQ ID NO:19

[0606] The amino acid sequence of Selaginella moellendorffii_SmHAD1

[0607] MIIISFACLKFGCGDNGPRGDLLRRALQHSSFLAYSCGELDRTAAISTISRKFKLTDPALVDSMLLEAASSCEVDKELLSLLQSTRQQLLGWIDIPPQEWERVYSLFPVSLWKNFA TISRDLDSLLGDIRSHAVIVDKSVEMAALHAFESLLALPYASPNKKACQDFLKRRFLLPRAVKLGMDRVKEEVRSKQLKSLYLLDNGRHELVSEIFFPCVAAWCLPEIIPLGWMKYL QVLIKRGHPFGYFGADVSRYPPDIDTMSTCVSTLFDLSLVTSAQAMHFLEICLENVNDQNQLLTYLDVERPRVDPVVIANVVYLAYALSMEDHPVVRHNENLIQRYLLSGGFVYGTR YYLSQEDFLFMYGRVLATFGEKREIPNFDLVYQAMEAALVNRIGNETESKPLDVAKRILLSRGFGIRNTIDVDLLLKMQNEDGSWPLQVLSNLPSAKGGVFNSVVDLSFAVRALQSQD

[0608] SEQ ID NO:20

[0609] Artificial sequence_class II motif SmHAD1 mutant_E. coli, optimized

[0610]

[0611] SEQ ID NO:21

[0612] SmHAD1 mutant with artificial sequence class II motif

[0613] MIIISFACLKFGCGDNGPRGDLLRRALQHSSFLAYSCGELDRTAAISTISRKFKLTDPALVDSMLLEAASSCEVDKELLSLLQSTRQQLLGWIDIPPQEWERVYSLFPVSLWKNFA TISRDLDSLLGDIRSHAVIVDKSVEMAALHAFESLLALPYASPNKKACQDFLKRRFLLPRAVKLGMDRVKEEVRSKQLKSLYLLDNGRHELVSEIFFPCVAAWCLPEIIPLGWMKYL QVLIKRGHPFGYFGADVSRYPPAIATMSTCVSTLFDLSLVTSAQAMHFLEICLENVNDQNQLLTYLDVERPRVDPVVIANVVYLAYALSMEDHPVVRHNENLIQRYLLSGGFVYGTR YYLSQEDFLFMYGRVLATFGEKREIPNFDLVYQAMEAALVNRIGNETESKPLDVAKRILLSRGFGIRNTIDVDLLLKMQNEDGSWPLQVLSNLPSAKGGVFNSVVDLSFAVRALQSQD

[0614] SEQ ID NO:22

[0615] Artificial sequence QW motif-based SmHAD1 mutant in E. coli, optimized.

[0616]

[0617] SEQ ID NO: 23

[0618] Artificial sequence_SmHAD1 mutant with QW motif

[0619] MIIISFACLLKFGCGDNGPRGDLLRRALQHSSFLAYSCGELDRTAAISTISRKFKLTDPALVDSMLLEAASSCEVDKELLSLLQSTRQQLLGWIDIPPQEWERVYSLFPVSLWKNFATISRDLDSLLGDIRSHAVIVDKSVEMAALHAFESLLALPYASPNKKACQDFLKRRFLLPRAVKLGMDRVKEEVRSKQLKSLYLLDNGRHELVSEIFFPCVAAWCLPEIIPLGWMKYLQVLIKRGHPFGYFGADVSRYPPDIDTMSTCVSTLFDLSLVTSAQAMHFLEICLENVNDQNQLLTYLDVERPRVDPVVIANVVYLAYALSMEDHPVVRHNENLIQRYLLSGGFVYGTRYYLSQEDFLFMYGRVLATFGEKREIPNFDLVYQAMEAALVNRIGNETESKPLDVAKRILLSRGFGIRNTIDVDLLLKMQNAAGSWPLQVLSNLPSAKGGVFNSVVDLSFAVRALQSQD

[0620] SEQ ID NO: 24

[0621] Selaginella moellendorffii_SmHAD2_wt_nucleic acid sequence

[0622]

[0623] SEQ ID NO:25

[0624] Selaginella moellendorffii_SmHAD2_Escherichia coli, optimized

[0625]

[0626] SEQ ID NO:26

[0627] The amino acid sequence of Selaginella moellendorffii_SmHAD2

[0628] MIIISFACLKFGCGDSGPRGDLLRRALQHSSFLAYSCGELDRAAAISTISRKFKLKEPALLDSMLLEAASSCEVDEELLSLLQSTRQQLLGWIDIPPQEWERVYNLFPWSLWKNFA TISRLDSLLGDIRFHAVIVDKSVEMAALHAFESLLALPYASPNKKACQDFLKRRFLLPRAVKLGMDRVKEEVRSKQLKSLYLLDNGRQELVSEIFFPCVAAWCLPEIIPLGWMESL QVLIERGHPFGYFGADVSRYPPDIDTMSTCVSTLFDLSLVTSAQAMHFLEICLENVNDQNQLLTYLDLERPRVDPVVIANVVYFAYALSMEDHPVVRHNENLIQRYLLSGGFVYGTR YYLSQEDFLFMYGRVLATFGEKREIPNFDLVYQAMEAALVNRIGNETESKPLDVAKRILLSRYFGIRNTIDVDLLLKMQNEDGSWPLQVLSNLPSAKGGVFNSVVDLSFAVRALQSQD

[0629] SEQ ID NO:27

[0630] Artificial sequence_class II motif SmHAD2 mutant_E. coli, optimized

[0631]

[0632] SEQ ID NO:28

[0633] SmHAD2 mutant with artificial sequence class II motif

[0634] MIIISFACLKFGCGDSGPRGDLLRRALQHSSFLAYSCGELDRAAAISTISRKFKLKEPALLDSMLLEAASSCEVDEELLSLLQSTRQQLLGWIDIPPQEWERVYNLFPWSLWKNFA TISRLDSLLGDIRFHAVIVDKSVEMAALHAFESLLALPYASPNKKACQDFLKRRFLLPRAVKLGMDRVKEEVRSKQLKSLYLLDNGRQELVSEIFFPCVAAWCLPEIIPLGWMESL QVLIERGHPFGYFGADVSRYPPAIATMSTCVSTLFDLSLVTSAQAMHFLEICLENVNDQNQLLTYLDLERPRVDPVVIANVVYFAYALSMEDHPVVRHNENLIQRYLLSGGFVYGTR YYLSQEDFLFMYGRVLATFGEKREIPNFDLVYQAMEAALVNRIGNETESKPLDVAKRILLSRYFGIRNTIDVDLLLKMQNEDGSWPLQVLSNLPSAKGGVFNSVVDLSFAVRALQSQD

[0635] SEQ ID NO:29

[0636] Artificial sequence QW motif-based SmHAD2 mutant E. coli, optimized

[0637]

[0638] SEQ ID NO: 30

[0639] Artificial sequence_SmHAD2 mutant of QW motif

[0640] MIIISFACLLKFGCGDSGPRGDLLRRALQHSSFLAYSCGELDRAAAISTISRKFKLKEPALLDSMLLEAASSCEVDEELLSLLQSTRQQLLGWIDIPPQEWERVYNLFPWSLWKNFATISRDLDSLLGDIRFHAVIVDKSVEMAALHAFESLLALPYASPNKKACQDFLKRRFLLPRAVKLGMDRVKEEVRSKQLKSLYLLDNGRQELVSEIFFPCVAAWCLPEIIPLGWMESLQVLIERGHPFGYFGADVSRYPPDIDTMSTCVSTLFDLSLVTSAQAMHFLEICLENVNDQNQLLTYLDLERPRVDPVVIANVVYFAYALSMEDHPVVRHNENLIQRYLLSGGFVYGTRYYLSQEDFLFMYGRVLATFGEKREIPNFDLVYQAMEAALVNRIGNETESKPLDVAKRILLSRYFGIRNTIDVDLLLKMQNAAGSWPLQVLSNLPSAKGGVFNSVVDLSFAVRALQSQD

[0641] SEQ ID NO: 31

[0642] Gelatoporia subvermispora_EMD37666.1_Escherichia coli, optimized

[0643]

[0644] SEQ ID NO: 32

[0645] Gelatoporia subvermispora_EMD37666.1_amino acid sequence

[0646] MSAAAQYTTLILDLGDVLFTWSPKTKTSIPPRTLKEILNSATWYEYERGRISQDECYERVGTEFGIAPSEIDNAFKQARDSMESNDELIALVRELKTQLDGELLVFALSNISLPDYEYVLTKPADWSIFDKVFPSALVGERKPHLGVYKHVIAETGIDPRTTVFVDDKIDNVLSARSVGMHGIVFEKQEDVMRALRNIFGDPVRRGREYLRRNAMRLESVTDHGVAFGENFTQLLILELTNDPSLVTLPDRPRTWNFFRGNGGRPSKPLFSEAFPDDLDTTSLALTVLQRDPGVISSVMDEMLNYRDPDGIMQTYFDDGRQRLDPFVNVNVLTFFYTNGRGHELDQCLTWVREVLLYRAYLGGSRYYPSADCFLYFISRLFACTNDPVLHHQLKPLFVERVQEQIGVEGDALELAFRLLVCASLDVQNAIDMRRLLEMQCEDGGWEGGNLYRFGTTGLKVTNRGLTTAAAVQAIEASQRRPPSPSPSVESTKSPITPVTPMLEVPSLGLSISRPSSPLLGYFRLPWKKSAEVH

[0647] SEQ ID NO: 33

[0648] Artificial sequence_EMD37666.1_mutant of class II motif_Escherichia coli, optimized

[0649]

[0650] SEQ ID NO: 34

[0651] Artificial sequence_EMD37666.1_Mutant of class II motif

[0652] MSAAAQYTTLILDLGDVLFTWSPKTKTSIPPRTLKEILNSATWYEYERGRISQDECYERVGTEFGIAPSEIDNAFKQARDSMESNDELIALVRELKTQLDGELLVFALSNISLPDYEYVLTKPADWSIFDKVFPSALVGERKPHLGVYKHVIAETGIDPRTTVFVDDKIDNVLSARSVGMHGIVFEKQEDVMRALRNIFGDPVRRGREYLRRNAMRLESVTDHGVAFGENFTQLLILELTNDPSLVTLPDRPRTWNFFRGNGGRPSKPLFSEAFPAALATTSLALTVLQRDPGVISSVMDEMLNYRDPDGIMQTYFDDGRQRLDPFVNVNVLTFFYTNGRGHELDQCLTWVREVLLYRAYLGGSRYYPSADCFLYFISRLFACTNDPVLHHQLKPLFVERVQEQIGVEGDALELAFRLLVCASLDVQNAIDMRRLLEMQCEDGGWEGGNLYRFGTTGLKVTNRGLTTAAAVQAIEASQRRPPSPSPSVESTKSPITPVTPMLEVPSLGLSISRPSSPLLGYFRLPWKKSAEVH

[0653] SEQ ID NO: 35

[0654] Dichomitus squalens_XP_007369631.1_Escherichia coli, optimized

[0655]

[0656] SEQ ID NO:36

[0657] Amino acid sequence of *Dichomitus squalens* (XP_007369631.1)

[0658] MASIHRRYTTLILDLGDVLFRWSPKTETAIPPQQLKDILSSVTWFEYERGRLSQEACYERCAEEFKIEASVIAEAFKQARGSLRPNEEFIALIRDLRREMHGDLTVLALSNISLPDYEYIMSLSSDWTTVF DRVFPSALVGERKPHLGCYRKVISEMNLEPQTTVFVDDKLDNVASARSLGMHGIVFDNQANVFRQLRNIFGDPIRRGQEYLRGHAGKLESSTDNGLIFEENFTQLIIYELTQDRTLISLSECPRTWNFFRGE PLFSETFPDDVDTTSVALTVLQPDRALVNSVLDEMLEYVDADGIMQTYFDRSRPRMDPFVCVNVLSLFYENGRGHELPRTLDWVYEVLLHRAYHGGSRYYLSPDCFLFFMSRLLKRADDPAVQARLRPLFVE RVNERVGAAGDSMDLAFRILAAASVGVQCPRDLERLTAGQCDDGGWDLCWFYVFGSTGVKAGNRGLTTALAVVGSTGVTAIGRPPSPSSAASSSFRPSSPYKFLGISRPASPIRFGDLLRPWRKMSRSNLKSQ

[0659] SEQ ID NO:37

[0660] Mutant of artificial sequence _XP_007369631.1_class II motif_E. coli, optimized

[0661]

[0662] SEQ ID NO:38

[0663] Mutants of the artificial sequence _XP_007369631.1_Class II motif

[0664] MASIHRRYTTLILDLGDVLFRWSPKTETAIPPQQLKDILSSVTWFEYERGRLSQEACYERCAEEFKIEASVIAEAFKQARGSLRPNEEFIALIRDLRREMHGDLTVLALSNISLPDYEYIMSLSSDWTTVF DRVFPSALVGERKPHLGCYRKVISEMNLEPQTTVFVDDKLDNVASARSLGMHGIVFDNQANVFRQLRNIFGDPIRRGQEYLRGHAGKLESSTDNGLIFEENFTQLIIYELTQDRTLISLSECPRTWNFFRGE PLFSETFPAAVATTSVALTVLQPDRALVNSVLDEMLEYVDADGIMQTYFDRSRPRMDPFVCVNVLSLFYENGRGHELPRTLDWVYEVLLHRAYHGGSRYYLSPDCFLFFMSRLLKRADDPAVQARLRPLFVE RVNERVGAAGDSMDLAFRILAAASVGVQCPRDLERLTAGQCDDGGWDLCWFYVFGSTGVKAGNRGLTTALAVVGSTGVTAIGRPPSPSSAASSSFRPSSPYKFLGISRPASPIRFGDLLRPWRKMSRSNLKSQ

[0665] SEQ ID NO:39

[0666] Artificial sequence_Class II synthase motif_mutated_BtHAD

[0667] PDALASTS

[0668] SEQ ID NO:40

[0669] Artificial sequence_Class II synthase motif_mutated_SmHAD1+2

[0670] PPAIATMS

[0671] SEQ ID NO:41

[0672] Artificial sequence_Class II synthase motif_Mutant_EMD37666

[0673] PDAALATTS

[0674] SEQ ID NO:42

[0675] Artificial sequence_Class II synthase motif_Mutated_XP_007369631

[0676] PDAAVATTS

[0677] SEQ ID NO:43

[0678] Artificial sequence_QW motif_mutated_BtHAD

[0679] QNVAGSW

[0680] SEQ ID NO:44

[0681] Artificial sequence_QW motif_mutated_SmHAD1+2

[0682] QNAAGSW

[0683] SEQ ID NO:45

[0684] Artificial sequence_Class I synthase motif DDxx(D / E)

[0685] DDxx(D / E)

[0686] SEQ ID NO:46

[0687] Artificial sequence_Class II synthase motif PxDxD(T / S)(T / M)S

[0688] PxDxD(T / S)(T / M)S

[0689] SEQ ID NO:47

[0690] Artificial Sequence_Class II Synthesizer Motif PDDLDSTS

[0691] PDDLDSTS

[0692] SEQ ID NO:48

[0693] Artificial Sequence_Class II Synthase Motif PDDLDTTS

[0694] PDDLDTTS

[0695] SEQ ID NO:49

[0696] Artificial Sequence_Class II Synthase Motif PPDIDTMS

[0697] PPDIDTMS

[0698] SEQ ID NO:50

[0699] Artificial Sequence_Class II Synthase Motif PNDIDTMS

[0700] PNDIDTTS

[0701] SEQ ID NO:51

[0702] Artificial sequence QW motif

[0703] QxxDGxW

[0704] SEQ ID NO:52

[0705] Artificial sequence_QW motif QNVDGSW

[0706] QNVDGSW

[0707] SEQ ID NO:53

[0708] Artificial sequence_QW motif QCDDGGW

[0709] QCDDGGW

[0710] SEQ ID NO:54

[0711] Artificial sequence_QW motif QSSDGGW

[0712] QSSDGGW

[0713] SEQ ID NO:55

[0714] Artificial sequence_QW motif QNEDGSW

[0715] QNEDGSW

[0716] SEQ ID NO:56

[0717] Artificial sequence_conserved motif 1 Lxxxx(W / F)xxYxxG

[0718] Lxxxx(W / F)xxYxxG

[0719] SEQ ID NO:57

[0720] Artificial Sequence_Conserved Motif 1 LRSHIWFNYSMG

[0721] LRSHIWFNYSMG

[0722] SEQ ID NO:58

[0723] Artificial sequence_conserved motif 1 LRSATWAAYECG

[0724] LRSATWAAYECG

[0725] SEQ ID NO:59

[0726] Artificial sequence_conserved motif 1 LRTPTWGKYECG

[0727] LRTPTWGKYECG

[0728] SEQ ID NO:60

[0729] Artificial sequence_conserved motif 1 LQHSSFLAYSCG

[0730] LQHSSFLAYSCG

[0731] SEQ ID NO:61

[0732] Artificial sequence_conserved motif 1 LRxxTWxxYECG

[0733] LRxxTWxxYECG

[0734] SEQ ID NO:62

[0735] Artificial sequence_conserved motif 2 YxDxxRxRVD(P / A)V(V / A)xxN

[0736] YxDxxRxRVDxVxxxN

[0737] SEQ ID NO:63

[0738] Artificial sequence_conserved motif 2 YFDPLRLRVDPVAATN

[0739] YFDPLRLRVDPVAATN

[0740] SEQ ID NO:64

[0741] Artificial sequence_conserved motif 2 YFDETRPRVDAVVNVN

[0742] YFDETRPRVDAVVNVN

[0743] SEQ ID NO:65

[0744] Artificial sequence_conserved motif 2 YFDKTRPRVDPVVCVN

[0745] YFDKTRPRVDPVVCVN

[0746] SEQ ID NO:66

[0747] Artificial sequence_conserved motif 2 YLDVERPRVDPVVIAN

[0748] YLDVERPRVDPVVIAN

[0749] SEQ ID NO:67

[0750] Artificial sequence_conserved motif 2 YLDLERPRVDPVVIAN

[0751] YLDLERPRVDPVVIAN

[0752] SEQ ID NO:68

[0753] Artificial sequence_conserved motif 2 Y(F / L)Dx(T / E)RPRVD(P / A)VVx(A / V)N

[0754] YxDxxRPRVDxVVxxN

[0755] SEQ ID NO:69

[0756] Artificial sequence_conserved motif 3 GTx(Y / F)YxxxExFL(Y / F)

[0757] GTxxYxxxExFLx

[0758] SEQ ID NO:70

[0759] Artificial sequence_conserved motif 3 GTLYYRTPEAFLY

[0760] GTLYYRTPEAFLY

[0761] SEQ ID NO:71

[0762] Artificial sequence_conserved motif 3 GTLFYYHAESFLY

[0763] GTLFYYHAESFLY

[0764] SEQ ID NO:72

[0765] Artificial sequence_conserved motif 3 GTRYYLSQEDFLF

[0766] GTRYYLSQEDFLF

[0767] SEQ ID NO:73

[0768] Dryopteris fragrans DfHAD wt nucleic acid sequence

[0769]

[0770] SEQ ID NO: 74

[0771] Dryopteris fragrans_DfHAD_ Amino acid sequence

[0772] MEFSASAPPPRLASVIILEPLGFLLTPHYSSQLPKKLLRRLLCTRIWHRYQRGRLRLRDAAMLLAQLPFLAVSDHPWALDNLASLLRPTAVRAVPWMLLLLDFLRDELHLKVVCATNSSPEELQELRHQFPALFAKVDATVSSGEEGVGKPSVRFLQAALDKAGVHAQQTLYLDSFDSLETIMAARSLGMHALSVEPCHIDELTARASSGQLRDAQLIRRIVCAMHGPAVSAVVSGSITSSGPQTAKIEELPTAADSHLRSAALTSAQQFFLKVIAPHRPEKPFVQLPSLTSEGIRIYDTFAQFVIADLLDDTRFLPMQSPPPNGLITFVNPSAYLADDIKNGNSHIVPGVQFYASDACTLIDIPHDLDTTSVGLSVLHKFGKVDKDTLNKVLDRMLEQVSEDDGILQVYFDVERPRIDPVVVANTVFLFHLGKRGHEVARSEKFVESVLLQRAYEEGTLYYNLGEAFLVSVARLVHEFKEHFTRSGMRRALEERLRERARAGMQERDDALALAMRIRACALCGLAGEGLTKAAEQELLRLQCKSKGCWGCHPFYRNGSNVLSWIGSEALTTAYAIAALQPIDI

[0773] SEQ ID NO: 75

[0774] Artificial sequence_Class II synthase motif

[0775] DxDxxS

[0776] SEQ ID NO: 76

[0777] Artificial sequence_QW motif

[0778] QxxxxxW.

Claims

1. An isolated polypeptide from a haloacid dehalogenase-like (HAD-like) hydrolase superfamily, comprising cyclic terpene synthase activity, wherein the polypeptide is: a. BazzHAD2, whose amino acid sequence has at least 98% sequence identity with SEQ ID NO:6; or b. BazzHAD3, whose amino acid sequence has at least 98% sequence identity with SEQ ID NO:

9.

2. The polypeptide of claim 1, comprising the ability to produce the following substances: a) At least one phosphate precursor of a sterane sesquiterpene; or b) At least one sesquiterpene; or c) At least one phosphate precursor of a sterane sesquiterpene and at least one corresponding sterane sesquiterpene.

3. The polypeptide according to claim 1 or 2, comprising the ability to produce the following substances: a) Flavothermic diphosphate derived from farnesyl diphosphate (FPP) as a substrate; or b) Fennerin derived from the substrate FPP, either directly from the substrate FPP or via its respective diphosphate precursor; or c) Flavothermic diphosphate derived from farnesyl diphosphate (FPP) as a substrate; and styrosine and / or flavothermic alcohol, which are derived directly from FPP as a substrate or via their respective diphosphate precursors.

4. An isolated nucleic acid molecule encoding a polypeptide according to any one of claims 1 to 3.

5. An expression construct comprising at least one nucleic acid molecule of claim 4.

6. A vector comprising at least one nucleic acid molecule of claim 4 or at least one expression construct of claim 5.

7. A recombinant non-human host cell or recombinant non-human host organism, comprising: a. At least one isolated nucleic acid molecule of claim 4, which optionally stably integrates into the genome; or b. At least one expression construct of claim 5, which optionally stably integrates into the genome; or c. At least one carrier of claim 6.

8. A method for producing the polypeptide according to any one of claims 1 to 3, the method comprising: a. To culture the non-human host cells or non-human host organisms of claim 7 to express the polypeptide of any one of claims 1 to 3; and b. Optionally, the polypeptide may be isolated from non-human host cells or non-human organisms cultured in step a.

9. A method for producing sterane sesquiterpenes, particularly zephyranol, comprising: a. Contacting farnesyl diphosphate with the polypeptide of claim 1 to obtain at least one bisphosphate of a sesquiterpene, particularly zephyranyl diphosphate; b. Cleavage the diphosphate portion of the product obtained in step a. by chemical or enzymatic means; and c. Optionally, the sesquiterpene, particularly zestyraxol, is isolated.

10. The method of claim 9, wherein the enzymatic step b. is carried out by applying the same polypeptide as used in step a., the polypeptide containing terpene synthase activity and phosphatase activity, or by applying a different polypeptide that has phosphatase activity.

11. Use of the polypeptide as defined in any one of claims 1 to 3 in the preparation of odorant, flavoring or fragrance ingredients, particularly Ambrox.

12. A method for generating Ambrox, the method comprising: The method of claim 9 or 10 provides zoea ol, optionally, zoea ol produced in step a. is separated; and zoea ol is converted to Ambrox in a manner known per se.

Citation Information

Patent Citations

  • Immobilized lipase

    EP1069183A2

  • Process for the production of covalently bound biologically active materials on polyurethane foams and the use of such carriers for chiral syntheses

    EP1149849A1

  • Method for producing albicanol and / or drimenol

    WO2018220113A1