Biocatalytic production method of terpene compounds
A novel class of terpenyl-diphosphate phosphatases from the protein tyrosine phosphatase family addresses the need for efficient enzyme-mediated cleavage of terpenyl diphosphate bonds, facilitating the biocatalytic production of terpene intermediates like copalol and labdenediol for fragrance components.
Patent Information
- Application Number
- JP2023212519
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2018-07-10
- Filing Date
- 2023-12-15
- Publication Date
- 2025-08-01
- Estimated Expiration
- 2039-07-10
AI Technical Summary
There is a lack of enzymes capable of efficiently catalyzing the cleavage of terpenyl diphosphate bonds, particularly in complex bicyclic compounds like copalyl diphosphate (CPP) and labdendiol diphosphate (LPP), which are crucial for the biocatalytic production of valuable terpene intermediates such as copalol and labdendiol used in fragrance components.
Development of a novel class of enzymes from the protein tyrosine phosphatase family, specifically terpenyl-diphosphate phosphatases, that can catalyze the removal of diphosphate groups from terpenyl diphosphate intermediates, enabling the biocatalytic production of terpene alcohols like copalol and labdenediol.
This approach provides a cost-effective method for producing terpene intermediates, allowing for the biocatalytic synthesis of complex terpene molecules, including copalol and labdenediol, which are essential for manufacturing high-value fragrance components like Ambrox.
Smart Images

Figure 0007717142000092 
Figure 0007717142000093 
Figure 0007717142000094
Abstract
Description
Technical Field
[0001] Provided herein is a method for the biocatalytic production of terpene compounds by applying a novel type of phosphatase enzyme. This method enables the total biochemical synthesis of terpene compounds such as copalol and labdendiol, and their derivatives, which are valuable intermediates for the production of fragrance components such as ambrex or γ-ambrinol. Also provided are such compounds and a novel total biochemical multi-step process for the production of novel phosphatase enzymes and mutants and variants derived therefrom.
[0002] Background of the Invention Terpenes are found in most organisms (microorganisms, animals and plants). These compounds are composed of five-carbon units called isoprene units and are classified according to the number of these units present in their structure. Thus, monoterpenes, sesquiterpenes and diterpenes are terpenes containing 10, 15 and 20 carbon atoms, respectively. For example, sesquiterpenes are widely found in the plant kingdom. Many sesquiterpene molecules are known for their flavor and fragrance properties as well as their cosmetic, medicinal and antimicrobial effects. A number of hydrocarbon sesquiterpenes and sesquiterpenoids have been identified.
[0003] The biosynthesis of terpenes requires an enzyme called terpene synthase. This enzyme converts acyclic terpene precursors into one or more terpene products. In particular, diterpene synthase produces diterpenes by cyclization of the precursor geranylgeranyl diphosphate (GGPP). The cyclization of GGPP often requires two enzyme polypeptides, type I and type II diterpene synthases, which act in combination in two consecutive enzyme reactions. Type II diterpene synthase catalyzes the cyclization / rearrangement of GGPP initiated by proton addition to the terminal double bond of GGPP, resulting in a cyclic diterpene diphosphate intermediate. This intermediate is then further converted by type I diterpene synthase, which catalyzes the cyclization initiated by ionization.
[0004] Diterpene synthases are present in plants and other organisms and use substrates such as GGPP, but they have different product profiles. Genes and cDNAs encoding diterpene synthases have been cloned, and the corresponding recombinant enzymes have been characterized.
[0005] Enzymes that catalyze the specific or preferential cleavage or removal of diphosphate groups from terpene diphosphate intermediates, particularly cyclic terpene diphosphate intermediates such as copalyl diphosphate (CPP) or labdendiol diphosphate (LPP), which are diterpenes, have not been described to date. Chemical cleavage of the phosphoester bond would be required to effect said cleavage.
[0006] The problem to be solved by the present invention is to provide a polypeptide that exhibits the enzyme activity of a phosphatase applicable to the enzymatic cleavage of terpenyl diphosphate bonds and enables the biocatalytic production of terpene alcohols.
[0007] Summary of the Invention Surprisingly, the above problems could be solved by providing a new class of enzymes exhibiting terpenyl-diphosphate phosphatase activity, selected from a subgroup of diphosphate-removing enzymes of the large protein tyrosine phosphatase family. The applicability of such enzymes of the protein tyrosine phosphatase family as phosphatases utilizing terpenyl diphosphate as a substrate, particularly such complex bicyclic compounds as CPP and LPP, is not described in the prior art.
[0008] This approach enables the provision of a more cost-effective method for producing terpene intermediates such as copalol and labdenediol, which are building blocks for manufacturing highly valuable fragrance components such as Ambrox.
[0009] In some embodiments of the present invention, furthermore, the biocatalytic production of acyclic terpene alcohols such as farnesol or geranylgeraniol from the corresponding diphosphate precursors is realized.
[0010] The biocatalytic step can be combined with several other preceding or consecutive enzyme steps and can enable the provision of a biocatalytic multi-step process for the total enzyme synthesis of valuable complex terpene molecules from their respective precursors. BRIEF DESCRIPTION OF THE DRAWINGS
[0011]
Figure 1
Figure 2a
Figure 2b
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7a
Figure 7b
Figure 8
Figure 9
Figure 10
Figure 11
Figure 12
Figure 13
Figure 14
Figure 15
Figure 16a
Figure 16b
Figure 17
Figure 18
[0012] Abbreviations Used ADH Alcohol dehydrogenase bp Base pair kb Kilobase CPP Copalyl diphosphate CPS Copalyl diphosphate synthase DNA Deoxyribonucleic acid cDNA Complementary DNA DMAPP Dimethylallyl diphosphate DTT Dithiothreitol FPP Farnesyl diphosphate GPP Geranyl diphosphate GGPP Geranylgeranyl diphosphate GGPS Geranylgeranyl diphosphate synthase GC Gas chromatograph IPP Isopentenyl diphosphate LPP Labdenediol diphosphate LPS Labdenediol diphosphate synthase MS Mass spectrometer / mass spectrometry MVA Mevalonic acid PP Diphosphate, pyrophosphate PCR Polymerase chain reaction RNA Ribonucleic acid mRNA Messenger ribonucleic acid miRNA MicroRNA siRNA Small interfering RNA rRNA Ribosomal RNA tRNA Transfer RNA TPP Terpenyl diphosphate
[0013] Definitions As used herein, "diphosphate" and "pyrophosphate" are synonyms.
[0014] "Terpenyl" refers to acyclic and cyclic chemical hydrocarbyl residues derived from the C5 building block isoprene, particularly those containing one or more such building blocks.
[0015] "Bicyclic terpene" or "bicyclic terpenyl" or "bicyclic diterpene" or "bicyclic diterpenyl" relates to terpene compounds or terpenyl residues containing two carbon cyclic rings, preferably two carbon cyclic fused rings, within their structure.
[0016] A "hydrocarbyl" residue is a chemical group consisting essentially of carbon and hydrogen atoms, and may be cyclic (e.g., monocyclic or polycyclic), acyclic, linear or branched, saturated or unsaturated moieties. This residue contains more than one carbon atom, such as 2, 3, 4 or 5 carbon atoms, and in particular contains 5 or more carbon atoms such as 5 to 30, 5 to 25, 5 to 20, 5 to 15 or 5 to 10 carbon atoms. The hydrocarbyl group may be unsubstituted or may have at least one, preferably 0, 1 or 2 substituents such as 1 to 5. The substituents contain one kind of heteroatom such as O or N. Preferably, the substituents are independently selected from -OH, C=O, or -COOH. Most preferably, the substituent is -OH.
[0017] A "monocyclic or polycyclic hydrocarbyl residue" contains one, two or three optionally substituted saturated or unsaturated fused (cyclic) or non-fused hydrocarbon ring groups (or "carbocyclic" groups). Each ring may independently contain 3 to 8, particularly 5 to 7, especially 6 ring carbon atoms. Examples of monocyclic residues include "cycloalkyl" groups which are carbocyclic groups having 3 to 7 ring carbon atoms such as cyclopropyl group, cyclobutyl group, cyclopentyl group, cyclohexyl group, cycloheptyl group, and cyclooctyl group; and the corresponding "cycloalkenyl" groups can be mentioned. "Cycloalkenyl" (or "mono- or poly-unsaturated cycloalkyl") particularly refers to a monocyclic, mono- or poly-unsaturated carbocyclic group having 5 to 8, preferably at most 6 carbon ring members, such as mono-unsaturated cyclopentenyl group, cyclohexenyl group, cycloheptenyl group and cyclooctenyl group.
[0018] Examples of polycyclic residues include groups in which, for example, one, two or three such cycloalkyls and / or cycloalkenyls are bonded to each other to form a polycyclic cycloalkyl ring or cycloalkenyl ring. As a non-limiting example, a bicyclic decalinyl residue composed of two 6-membered carbon rings can be mentioned.
[0019] The number of substituents within such monocyclic or polycyclic hydrocarbyl residues can vary from 1 to 10, particularly 1 to 5 substituents. Suitable substituents for such cyclic residues are selected from lower alkyl, lower alkenyl, alkylidene, alkenylidene, or residues containing one heteroatom such as O or N, for example residues such as -OH or -COOH. In particular, the substituents are independently selected from -OH, -COOH, methyl, and methylidene.
[0020] The term "lower alkyl" or "short-chain alkyl" represents a straight-chain or branched saturated hydrocarbon group having 1 to 4, 1 to 5, 1 to 6, or 1 to 7, particularly 1 to 4 carbon atoms. Examples that can be mentioned below are: methyl, ethyl, n-propyl, 1-methylethyl, n-butyl, 1-methylpropyl, 2-methylpropyl, 1,1-dimethylethyl, n-pentyl, 1-methylbutyl, 2-methylbutyl, 3-methylbutyl, 2,2-dimethylpropyl, 1-ethylpropyl, n-hexyl, 1,1-dimethylpropyl, 1,2-dimethylpropyl, 1-methylpentyl, 2-methylpentyl, 3-methylpentyl, 4-methylpentyl, 1,1-dimethylbutyl, 1,2-dimethylbutyl, 1,3-dimethylbutyl, 2,2-dimethylbutyl, 2,3-dimethylbutyl, 3,3-dimethylbutyl, 1-ethylbutyl, 2-ethylbutyl, 1,1,2-trimethylpropyl, 1,2,2-trimethylpropyl, 1-ethyl-1-methylpropyl, and 1-ethyl-2-methylpropyl; and n-heptyl, and its single or multiple branched analogs.
[0021] "Short-chain alkenyl" or "lower alkenyl" refers to a mono- or poly-unsaturated, particularly mono-unsaturated, straight-chain or branched hydrocarbon group having 2 to 4, 2 to 6, or 2 to 7 carbon atoms and one double bond at any position, such as ethenyl, 1-propenyl, 2-propenyl, 1-methylethenyl, 1-butenyl, 2-butenyl, 3-butenyl, 1-methyl-1-propenyl, 2-methyl-1-propenyl, 1-methyl-2-propenyl, 2-methyl-2-propenyl, 1-pentenyl, 2-pentenyl, 3-pentenyl, 4-pentenyl, 1-methyl-1-butenyl, 2-methyl-1-butenyl, 3-methyl-1-butenyl, 1-methyl-2-butenyl, 2-methyl-2-butenyl, 3-methyl-2-butenyl, 1-methyl-3-butenyl, 2-methyl-3-butenyl, 3-methyl-3-butenyl, 1,1-dimethyl-2-propenyl, 1,2-dimethyl-1-propenyl, 1,2-dimethyl-2-propenyl, 1-ethyl-1-propenyl, 1-ethyl-2-propenyl, 1-hexenyl, 2-hexenyl, 3-hexenyl, 4-hexenyl, 5-hexenyl, 1-methyl-1-pentenyl, 2-methyl-1-pentenyl, 3-methyl-1-pentenyl, 4-methyl-1-pentenyl, 1-methyl-2-pentenyl, 2-methyl-2-pentenyl, 3-methyl-2-pentenyl, 4-methyl-2-pentenyl, 1-methyl-3-pentenyl, 2-methyl-3-pentenyl, 3-methyl-3-pentenyl, 4-methyl-3-pentenyl, 1-methyl-4-pentenyl, 2-methyl-4-pentenyl, 3-methyl-4-pentenyl, 4-methyl-4-pentenyl, 1,1-dimethyl-2-butenyl, 1,1-dimethyl-3-butenyl, 1,2-dimethyl-1-butenyl, 1,2-dimethyl-2-butenyl, 1,2-dimethyl-3-butenyl, 1,3-dimethyl-1-butenyl, 1,3-dimethyl-2-butenyl, 1,3-dimethyl-3-butenyl, 2,2-dimethyl-3-butenyl, 2,3-dimethyl-1-butenyl, 2,3-dimethyl-2-butenyl, 2,3-dimethyl-3-butenyl, 3,3-dimethyl-1-butenyl, 3,It represents C2-C6-alkenyl such as 3-dimethyl-2-butenyl, 1-ethyl-1-butenyl, 1-ethyl-2-butenyl, 1-ethyl-3-butenyl, 2-ethyl-1-butenyl, 2-ethyl-2-butenyl, 2-ethyl-3-butenyl, 1,1,2-trimethyl-2-propenyl, 1-ethyl-1-methyl-2-propenyl, 1-ethyl-2-methyl-1-propenyl and 1-ethyl-2-methyl-2-propenyl.,
[0022] The "alkylidene" group represents a straight-chain or branched hydrocarbon substituent bonded to the main body of the molecule via a double bond. This substituent contains 1 to 6 carbon atoms. Examples of such "C1-C6-alkylidene" include methylidene (=CH2), ethylidene (=CH-CH2), n-propylidene, n-butylidene, n-pentylidene, n-hexylidene and their structural isomers such as isopropylidene for example can be mentioned.
[0023] "Alkenylidene" represents a mono-unsaturated analogue of the above-mentioned alkylidene having more than 2 carbon atoms and may be referred to as "C3-C6-alkenylidene". Examples can be mentioned as n-propenylidene, n-butenylidene, n-pentenylidene, and n-hexenylidene.
[0024] The unsaturated cyclic group may contain one or more C=C bonds, for example 1, 2 or 3, and is aromatic or in particular non-aromatic.
[0025] Specific examples of cyclic residues are of the formula Cyc-A- [wherein A represents a straight-chain or branched C1-C4-alkylene bridge, in particular methylene, and Cyc is C1-C4-alkyl, C1-C4-alkylidene, C2-C4-alkenyl, oxo, hydroxy, or amino, in particular C1-C4-alkyl such as methyl, and C1-C4-alkylidene such as methylidene, and is independently selected from 1 to 10, 1 to 5 substituents and is optionally substituted with a monocyclic or polycyclic, in particular bicyclic, saturated or unsaturated hydrocarbyl residue, in particular a bicyclic cyclic hydrocarbyl residue, containing 5 to 7, in particular 6 ring atoms per ring]. Cyc-A is in particular of formula IIIa, IIIb or IIIc [Chemical formula] represents a group of.
[0026] Typical examples of compounds containing such residues are those of formula (1) below, in particular copalol and labdendiol and their stereoisomers.
[0027] Non-limiting examples of C1-C4-alkyl are methyl, ethyl, n-propyl, 1-methylethyl, n-butyl, 1-methylpropyl, 2-methylpropyl, 1,1-dimethylethyl Non-limiting examples of C1-C4-alkylidene are methylidene (=CH2), ethylidene, (=CH-CH2), n-propylidene, n-butylidene, and their structural isomers.
[0028] Non-limiting examples of C2-C4-alkenyl are ethenyl, 1-propenyl, 2-propenyl, 1-methylethenyl, 1-butenyl, 2-butenyl, 3-butenyl, 1-methyl-1-propenyl, 2-methyl-1-propenyl, 1-methyl-2-propenyl, 2-methyl-2-propenyl.
[0029] Non-limiting examples of C1-C4-alkylene are -CH2-, -(CH2)2-, -(CH2)3-, -(CH2)4-, -(CH2)2-CH(CH3)-, -CH2-CH(CH3)-CH2-.
[0030] The "precursor" molecules of the target compounds described herein are preferably converted to the target compounds through the enzymatic action of a suitable polypeptide that effects at least one structural change on said precursor molecule. For example, a "diphosphate precursor" (such as, for example, a "terpenyl diphosphate precursor") is converted to the target compound (such as, for example, a terpene alcohol) through the enzymatic removal of the diphosphate moiety, for example, by the removal of a monophosphate or diphosphate group by a phosphatase enzyme. For example, an "acyclic precursor" (such as, for example, an "acyclic terpenyl precursor") may be converted in one or more steps to a cyclic target molecule (such as a cyclic terpene compound) through the action of a cyclase or synthase enzyme, regardless of the specific enzymatic mechanism of such enzyme.
[0031] "Terpene synthase" refers to a polypeptide that converts a terpene precursor molecule to its respective terpene target molecule, in particular a specifically engineered target terpene alcohol. Non-limiting examples of such terpene precursor molecules are, for example, acyclic compounds selected from farnesyl pyrophosphate (FPP), geranylgeranyl-pyrophosphate (GGPP), or a mixture of isopentenyl pyrophosphate (IPP) and dimethylallyl pyrophosphate (DMAPP). If the resulting terpene contains a diphosphate moiety, this synthase is called a "terpenyl diphosphate synthase" The term "terpenyl diphosphate synthase" or "polypeptide having terpenyl diphosphate synthase activity" or "terpenyl diphosphate synthase protein" or "having the ability to produce terpenyl diphosphate" relates to a polypeptide that can catalyze the synthesis of terpenyl diphosphate in any form of its stereoisomers or mixtures thereof starting from acyclic terpenyl pyrophosphate, particularly GPP, FPP, GGPP or DMAPP together with IPP. Terpenyl diphosphate may be the only product or may be part of a mixture of terpenyl phosphates. Said mixture may contain terpenyl monophosphate and / or terpene alcohol. The above definition also applies to the group of "bicyclic diterpenyl diphosphate synthases" which produce bicyclic terpenyl diphosphates such as CPP or LPP.
[0032] As an example of such a "terpenyl diphosphate synthase" enzyme or "diterpenyl diphosphate synthase" enzyme, mention can be made of copalyl diphosphate synthase (CPS). Copalyl-diphosphate may be the only product or may be part of a mixture of copalyl phosphates. Said mixture may contain copalyl-monophosphate and / or other terpenyl diphosphates.
[0033] As an example of such a "terpenyl diphosphate synthase" enzyme or "diterpenyl diphosphate synthase" enzyme, mention can be made of labdenediol diphosphate synthase (LPS). Labdenediol diphosphate may be the only product or may be part of a mixture of labdenediol phosphates. Said mixture may contain labdenediol monophosphate and / or terpenyl diphosphates.
[0034] "Terpenyl diphosphate synthase activity" or "diterpenyl diphosphate synthase" (such as CPS or LPS activity) is determined herein under the "standard conditions" described below: These activities are added at an initial concentration in the range of 1 to 100 μM mg / ml, preferably 5 to 50 μM, particularly 30 to 40 μM, or are endogenously produced by the host cell, in particular in the presence of a reference substrate, in this case GGPP, at a temperature in the range of about 20 to 45 °C, such as about 25 to 40 °C, preferably 25 to 32 °C, in a preferably buffered culture medium or reaction medium having a pH in the range of 6 to 11, preferably 7 to 9, using recombinant terpenyl diphosphate synthase-expressing host cells, disrupted terpenyl diphosphate synthase-expressing cells, or fractions of these or concentrated or purified terpenyl diphosphate synthase enzymes. The conversion reaction to produce terpenyl diphosphate is carried out for 10 minutes to 5 hours, preferably about 1 to 2 hours. In the absence of endogenous phosphatase, one or more exogenous phosphatases, such as alkaline phosphatase, are added to the reaction mixture to convert the terpenyl diphosphate produced by the synthase to the respective terpene alcohol. The terpene alcohol can then be measured by conventional methods after extraction with an organic solvent such as ethyl acetate.
[0035] The term "protein tyrosine phosphatase" represents a group of enzymes that are generally known to remove phosphate groups from phosphorylated tyrosine residues on proteins. Certain subgroups of said family described herein are enzymes useful for dephosphorylating phosphorylated terpene molecules.
[0036] The polypeptide of the present invention having terpenyl diphosphate phosphatase activity is identified as a member of the protein tyrosine phosphatase family, in particular as a member of the Y_phosphatase 3 family having the Pfam ID number PF13350. Against the Pfam protein family signature database, in particular in the Pfam 32.0 database release (September 2018), for example using the following websites, the polypeptide can be scanned for matches: http: / / pfam.xfam.org / search#tabview=tab0, https: / / www.ebi.ac.uk / Tools / hmmer / search / hmmscan or https: / / www.ebi.ac.uk / Tools / pfa / pfamscan / .
[0037] The term "Pfam" is maintained by the Pfam consortium and refers to a large collection of protein domains and protein families available at several sponsored worldwide web sites including: pfam.sanger.ac.uk / (Welcome Trust, Sanger Institute); pfam.sbc.su.se / (Stockholm Bioinformatics Center) and pfam.janelia.org / (Janelia Farm, Howard Hughes Medical Institute). The most recent release of Pfam is Pfam 32.0 (September 2018) based on the UniProt Reference Proteomes (El-Gebali S. et al, 2019, Nucleic Acids Res. 47, Database issue D427-D432). Pfam domains and families are identified using multiple sequence alignments and hidden Markov models (HMMs). Pfam-A family or domain assignments are high-quality assignments generated by curated seed alignments using profile hidden Markov models based on representative members of the protein family and seed alignments.(Unless otherwise specified, a match of a queried protein to a Pfam domain or family is a Pfam-A match.) Then, using all the identified sequences belonging to this family, a complete alignment of this family is automatically generated (Sonnhammer (1998) Nucleic Acids Research 26, 320-322; Bateman (2000) Nucleic Acids Research 26, 263-266; Bateman (2004) Nucleic Acids Research 32, Database Issue, D138-D141; Finn (2006) Nucleic Acids Research Database Issue 34, D247-251; Finn (2010) Nucleic Acids Research Database Issue 38, D211-222). For example, by accessing the Pfam database using any of the above websites, protein sequences can be queried against an HMM using HMMER homology search software (e.g., HMMER2, HMMER3, or a newer version, hmmer.janelia.org / ). A significant match that confirms that the queried protein is within a pfam family (or has a particular Pfam domain) is a match where the bit score is above the collection threshold for the Pfam domain. The expected value (e-value) can also be used as a criterion for whether the queried protein is included in Pfam or for determining whether the queried protein has a particular Pfam domain (much smaller than 1.0, e.g., less than 0.1 or smaller, for small e-values).
[0038] The "E-value" (expected value) is the number of hits that would be expected, by mere chance, to have a score equal to or better than this value. This means that a good E-value, which gives a reliable prediction, is much smaller than 1. An E-value of approximately 1 is the value expected by chance. Thus, the smaller the E-value, the more specific the search of the domain will be. Only positive numbers are allowed. (Definition according to Pfam) The term "terpenyl diphosphate phosphatase" or "polypeptide having terpenyl diphosphate phosphatase activity" or "terpenyl diphosphate phosphatase protein" or "having the ability to produce terpene alcohol" relates to a polypeptide that can catalyze the removal of a diphosphate moiety or a monophosphate moiety (regardless of a particular enzymatic mechanism) to produce a dephosphorylated compound, in particular the corresponding alcohol compound of the terpenyl moiety. The terpene alcohol may be present in the product as any of its stereoisomers or as a mixture thereof. The terpene alcohol may be the sole product or may be part of a mixture with other terpene compounds such as, for example, the dephosphorylated analogs of the respective (e.g., acyclic) terpenyl diphosphate precursors of the terpenyl diphosphate. The above definition also applies to the group of "bicyclic terpenyl diphosphate phosphatases", which produce bicyclic terpene alcohols such as copalol or labdenediol. Each of the above-mentioned phosphatases exemplifies a "diphosphate-removing enzyme".
[0039] As an example of such a "terpenyl diphosphate phosphatase" enzyme, copalyl diphosphate phosphatase (CPP phosphatase) can be mentioned. Copalol may be the sole product or may be part of a mixture with, for example, dephosphorylated precursors such as farnesol and / or geranylgeraniol, and / or may be part of a mixture with by-products resulting from the enzyme side activities in the reaction mixture such as such alcohols or esters or aldehydes of other cyclic or acyclic diterpenes.
[0040] As another example of such a "terpenyl diphosphate phosphatase" enzyme, mention can be made of labdenediol diphosphate phosphatase (LPP phosphatase). Labdenediol may be the only product, or may be part of a mixture with dephosphorylated precursors such as, for example, farnesol and / or geranylgeraniol, and / or may be part of a mixture with by-products resulting from the enzyme side activities in the reaction mixture such as esters or aldehydes of such alcohols or other cyclic or acyclic diterpenes.
[0041] "Terpenyl diphosphate Phosphatase activity" (such as CPP or LPP phosphatase activity) is determined herein under the "standard conditions" described below: These activities are added at an initial concentration in the range of 1 - 100 μM mg / ml, preferably 5 - 50 μM, particularly 30 - 40 μM, or endogenously produced by the host cell, in the presence of a reference substrate, in this case for example CPP or LPP, at a temperature in the range of about 20 - 45 °C, preferably about 25 - 32 °C, such as 25 - 40 °C, in a preferably buffered culture medium or reaction medium having a pH in the range of 6 - 11, preferably 7 - 9, using recombinant terpenyl diphosphate phosphatase-expressing host cells, disrupted terpenyl diphosphate phosphatase-expressing cells, fractions of these, or concentrated or purified terpenyl diphosphate phosphatase enzymes. The conversion reaction to produce terpenyl diphosphate is carried out for 10 minutes to 5 hours, preferably about 1 - 2 hours. Then, after extraction with an organic solvent such as ethyl acetate, the terpene alcohol can be measured by conventional methods.
[0042] Specific examples of suitable standard conditions can be understood from the experimental section below.
[0043] "Alcohol dehydrogenase" (ADH) in the context of the present invention is NAD + or NADP +Refers to a polypeptide having the ability to oxidize an alcohol to the corresponding aldehyde in the presence thereof. Such enzymes are members of the E.C. family 1.1.1.1 (NAD + -dependent) or 1.1.1.2 (NADP + -dependent). In particular, the ADH of the present invention has the ability to oxidize coparol to copalol and / or labdenediol to their respective aldehydes.
[0044] As used herein, "coparol" refers to (E)-5-[(1S,4aS,8aS)-5,5,8a-trimethyl-2-methylidene-3,4,4a,6,7,8-hexahydro-1H-naphthalen-1-yl]-3-methylpenta-2-en-1-ol; CAS registration number 10395-43-4.
[0045] As used herein, "copalol" refers to (2E)-3-methyl-5-[(1S,4aS,8aS)-5,5,8a-trimethyl-2-methylenedecahydro-1-naphthalenyl]-2-pentenal.
[0046] As used herein, "labdenediol" refers to (1R,2R,4aS,8aS)-1-[(E)-5-hydroxy-3-methylpenta-3-enyl]-2,5,5,8a-tetramethyl-3,4,4a,6,7,8-hexahydro-1H-naphthalen-2-ol; CAS registration number 10267-31-9.
[0047] As used herein, "manool" refers to 5-[(1S,4aS,8aS)-5,5,8a-trimethyl-2-methylidene-3,4,4a,6,7,8-hexahydro-1H-naphthalen-1-yl]-3-methylpenta-1-en-3-ol As used herein, (+)-manooloxy refers to 4-[(1S,4aS,8aS)-5,5,8a-trimethyl-2-methylenedecahydro-1-naphthalenyl]-2-butanone, As used herein, "Z-11" refers to (3S,5aR,7aS,11aS,11bR)-3,8,8,11a-tetramethyldodecahydro-3,5a-epoxynaphtho[2,1-c]oxepin.
[0048] As used herein, "γ-Ambrool" refers to 2-[(1S,4aS,8aS)-5,5,8a-trimethyl-2-methylenedecahydro-1-naphthalenyl]ethanol. And As used herein, Ambrox® refers to (3aR,5aS,9aS,9bR)-3a,6,6,9a-tetramethyldodecahydronaphtho[2,1-b]furan.
[0049] As used herein, "Scareolide" refers to 3a,6,6,9a-tetramethyldecahydronaphtho[2,1-b]furan-2(1H)-one As used herein, "DOL" refers to (1R,2R,4aS,8aS)-1-(2-hydroxyethyl)-2,5,5,8a-tetramethyl-3,4,4a,6,7,8-hexahydro-1H-naphthalen-2-ol... CAS No. 38419-75-9 As used herein, "Farnesol" refers to (2E,6E)-3,7,11-trimethyldodeca-2,6,10-trien-1-ol As used herein, "Geranylgeraniol" refers to (2E,6E,10E)-3,7,11,15-tetramethylhexadeca-2,6,10,14-tetraen-1-ol.
[0050] More generally, the following meanings apply: In the case of Z11-like compounds, the general formula:
Chemical formula
[0051] In the case of the Ambrox® compound, the general formula
Chemical formula
[0052] The terms "biological function", "function", "biological activity" or "activity" of terpene diphosphate synthase refer to the ability of the terpene diphosphate synthase described herein to catalyze the production of at least one terpene diphosphate from the corresponding precursor terpene.
[0053] The terms "biological function", "function", "biological activity" or "activity" of terpene diphosphate phosphatase refer to the ability of the terpene diphosphate phosphatase described herein to catalyze the removal of the diphosphate group from the terpene compound to produce the corresponding terpene alcohol.
[0054] The "mevalonate pathway", also known as the "isoprenoid pathway" or the "HMG-CoA reductase pathway", is an essential metabolic pathway present in eukaryotes, archaea, and some bacteria. The mevalonate pathway starts from acetyl-CoA and produces two 5-carbon building blocks called isopentenyl pyrophosphate (IPP) and dimethylallyl pyrophosphate (DMAPP). The key enzymes are acetoacetyl-CoA thiolase (atoB), HMG-CoA synthase (mvaS), HMG-CoA reductase (mvaA), mevalonate kinase (MvaK1), phosphomevalonate kinase (MvaK2), mevalonate diphosphate decarboxylase (MvaD), and isopentenyl diphosphate isomerase (idi). Combining the enzyme activity to produce terpene precursors GPP, FPP or GGPP, such as FPP synthase (ERG20), with the mevalonate pathway enables the recombinant cell production of terpenes.
[0055] As used herein, the terms "host cell" or "transformed cell" refer to a cell (or organism) that has been altered to possess at least one nucleic acid molecule, e.g., a cell (or organism) that has been modified to possess a recombinant gene encoding a desired protein or nucleic acid sequence that, upon transcription, gives at least one functional polypeptide of the invention, e.g., a terpenyl diphosphate synthase protein or a terpenyl diphosphate phosphatase enzyme as defined earlier herein. Host cells are particularly bacterial cells, fungal cells or plant cells or plants. A host cell may contain a recombinant gene or several genes incorporated into the nuclear genome or organelle genome of the host cell, e.g., organized as an operon. Alternatively, the host may contain the recombinant gene extrachromosomally.
[0056] The term "organism" refers to any non-human multicellular or unicellular organism such as a plant or a microorganism. In particular, microorganisms are bacteria, yeast, algae or fungi.
[0057] The term "plant" is used interchangeably to include plant protoplasts, plant tissues, plant cell tissue cultures that give rise to regenerated plants, or plant cells including parts of plants or plant organs such as roots, stems, leaves, flowers, pollen, ovules, embryos, fruits and the like. Any plant can be used to carry out the methods of the embodiments herein.
[0058] A particular organism or cell is intended to be "able to produce FPP" if it naturally produces FPP or if it does not naturally produce FPP but is transformed with a nucleic acid as described herein to produce FPP. Organisms or cells transformed to produce more FPP than a naturally occurring organism or cell are also encompassed by "organisms or cells able to produce FPP".
[0059] A particular organism or cell is intended to be "able to produce GGPP" if it naturally produces GGPP or if it does not naturally produce GGPP but is transformed with the nucleic acids described herein to produce GGPP. Organisms or cells transformed to produce greater amounts of GGPP than naturally occurring organisms or cells are also encompassed by "organisms or cells able to produce GGPP".
[0060] A particular organism or cell is intended to be "able to produce terpenyl diphosphate" if it naturally produces the terpenyl diphosphate as defined herein or if it does not naturally produce the diphosphate but is transformed with the nucleic acids described herein to produce the diphosphate. Organisms or cells transformed to produce greater amounts of terpenyl diphosphate than naturally occurring organisms or cells are also encompassed by "organisms or cells able to produce terpenyl diphosphate".
[0061] A particular organism or cell is intended to be "able to produce terpene alcohol" if it naturally produces the terpene alcohol as defined herein or if it does not naturally produce the alcohol but is transformed with the nucleic acids described herein to produce the alcohol. Organisms or cells transformed to produce greater amounts of terpene alcohol than naturally occurring organisms or cells are also encompassed by "organisms or cells able to produce terpene alcohol".
[0062] For the description herein and the appended claims, the use of "or" means "and / or" unless otherwise stated. Similarly, "comprise", "comprises", "comprising", "include", "includes", and "including" are interchangeable and not intended to be limiting.
[0063] It should be further understood that when the term "comprising" is used in the description of various embodiments, in some specific examples, those skilled in the art will understand that certain embodiments can alternatively be described using the words "consisting essentially of" or "consisting of".
[0064] As used herein, the terms "purified", "substantially purified", and "isolated" refer to a state in which the compounds of the present invention are free of other different compounds with which they are normally associated in their natural state. Thus, "purified", "substantially purified", and "isolated" objects include at least 0.5%, 1%, 5%, 10%, or 20%, or at least 50% or 75% by weight of the mass of a given sample. In one embodiment, these terms refer to compounds of the present invention that include at least 95%, 96%, 97%, 98%, 99% or 100% by weight of the mass of a given sample. As used herein, the terms "purified", "substantially purified", and "isolated" also refer to a state of purification or concentration different from that which occurs naturally when referring to (a) nucleic acid or protein or (a plurality of) nucleic acids or proteins, for example, within a bacterial cell or a fungal cell, or within a mammalian organism, particularly within a human body, i.e., in a prokaryotic or eukaryotic environment. Any degree of purification or concentration that exceeds the degree that occurs naturally, including (1) purification from other associated structures or compounds, or (2) association with structures or compounds that do not normally associate in said prokaryotic or eukaryotic environment, is within the meaning of "isolated". The (a) nucleic acid or protein or (a plurality of) nucleic acids or proteins or classes of nucleic acids or proteins described herein may be isolated according to various methods and processes known to those skilled in the art, or may be associated with structures or compounds that do not normally associate by nature.
[0065] The term "about" indicates a possible variation of ±25%, particularly ±15%, ±10%, especially ±5%, ±2% or ±1% of the stated value.
[0066] The term "substantially" refers to a range of values of about 80 - 100%, such as, for example, 85 - 99.9%, particularly 90 - 99.9%, especially 95 - 99.9%, or 98 - 99.9%, particularly 99 - 99.9%.
[0067] "Primarily" refers to a ratio within a range above 50%, such as within a range of, for example, 51 - 100%, particularly 75 - 99.9%, especially 85 - 98.5%, 95 - 99%.
[0068] The "main product" in the context of the present invention refers to a single compound or a group of at least two compounds, such as two or three compounds, particularly two or three compounds, and this single compound or group of compounds is "primarily" produced by the reaction described herein and is included in the reaction in a major ratio based on the total amount of the components of the product produced by said reaction. Said ratio may be a molar ratio, a weight ratio, or preferably an area ratio calculated from the corresponding chromatogram of the reaction product based on chromatographic analysis theory.
[0069] The "by - product" in the context of the present invention refers to a single compound or a group of at least two compounds, such as two or three compounds, particularly two or three compounds, and this single compound or group of compounds is not "primarily" produced by the reaction described herein.
[0070] Due to the reversibility of enzyme reactions, the present invention relates to the enzyme reactions or biocatalytic reactions described herein in both directions of the reaction, unless otherwise stated.
[0071] The "functional mutants" of the polypeptides described herein include the "functional equivalents" of the polypeptides as defined below.
[0072] The term "stereoisomers" includes conformational isomers, particularly configurational isomers.
[0073] According to the present invention, generally included are all "stereoisomeric forms" of the compounds described herein, such as "structural isomers" and "stereoisomers".
[0074] "Stereoisomeric forms" specifically include "stereoisomers" and mixtures thereof, such as configurational isomers (optical isomers) like enantiomers, or geometric isomers (diastereomers) like E- and Z-isomers, and combinations thereof. When one or more asymmetric centers are present within a molecule, the present invention includes all combinations of different conformations of these asymmetric centers, such as enantiomeric pairs. "Stereoselectivity" refers to the ability to produce a specific stereoisomer of a compound in a stereoisomerically pure form, or the ability to specifically convert a specific stereoisomer from a plurality of stereoisomers in the enzyme-catalyzed methods described herein. More specifically, this means that a product of the present invention can be enriched with respect to a specific stereoisomer, or a certain enantiomer can be decreased with respect to a specific stereoisomer. This can be quantified by the %ee-purity parameter calculated according to the following formula: %ee = [X A - X B / [X A + X B × 100 [wherein X A and X B represent the molar ratios (mole fractions) of stereoisomers A and B].
[0075] The terms "selectively convert" or "enhance selectivity" generally mean that, during the entire course of the reaction (i.e., from the start to the end of the reaction), at a certain point in the reaction, or during the "period" of the reaction, a particular stereoisomeric form of an unsaturated hydrocarbon, such as the E-form, is converted in a higher ratio or in a greater amount (on a molar basis) than the corresponding other stereoisomeric form, such as the Z-form. In particular, the selectivity corresponding to 1-99%, 2-95%, 3-90%, 5-85%, 10-80%, 15-75%, 20-70%, 25-65%, 30-60%, or 40-50% conversion of the initial amount of the substrate can be observed during the "period". The higher ratio or greater amount can be represented, for example, by: - a higher maximum yield of a certain isomer observed during the entire course of the reaction or during said period thereof; - a greater relative amount of a certain isomer at a defined % conversion value of the substrate; and / or - the same relative amount of a certain isomer at a higher % conversion value; Each of these is preferably observed in comparison to a standard method, which is carried out by known chemical or biochemical means under otherwise identical conditions.
[0076] Also generally included in accordance with the present invention are all "isomeric forms" of the compounds described herein, such as structural isomers and in particular stereoisomers and mixtures thereof, such as optical isomers or geometric isomers such as E- and Z-isomers and combinations thereof. If several asymmetric centers are present within a molecule, the present invention includes all combinations of different conformations of these asymmetric centers, such as pairs of enantiomers, or any mixture of stereoisomeric forms.
[0077] The "yield" and / or "conversion rate" of the reaction according to the present invention is determined over a defined period, for example 4 hours, 6 hours, 8 hours, 10 hours, 12 hours, 16 hours, 20 hours, 24 hours, 36 hours or 48 hours, during which the reaction occurs. In particular, the reaction is carried out under precisely defined conditions, for example under the "standard conditions" defined herein.
[0078] Various yield parameters (the "yield" or Y P / S ; "specific productivity yield"; or space-time yield (STY)) are well known in the art and are determined as described in the literature.
[0079] (Represented by the mass of the product produced and the mass of the material consumed, respectively) "Yield" and "Y P / S " are used as synonyms herein.
[0080] Specific productivity-yield represents the amount of product produced per 1 g of biomass per hour and per 1 L of fermentation broth. The amount of wet cell weight described as WCW represents the amount of biologically active microorganisms in the biochemical reaction. This value is shown as the amount of product (g) per hour per 1 g of WCW (i.e., g / gWCW -1 h -1 ). Alternatively, the amount of biomass can be represented as the amount of dry cell weight described as DCW. Further, by measuring the optical density (OD 600 ) at 600 nm and using the experimentally determined correlation coefficient to estimate the corresponding wet cell weight or dry cell weight respectively, the biomass concentration can be more easily determined.
[0081] The term "fermentation production" or "fermentation" refers to the ability of a microorganism to produce a compound in a cell culture that utilizes at least one carbon source added to the incubation (supported by the enzyme activity contained in or produced by the microorganism).
[0082] The term "fermentation broth" is understood to mean a liquid, particularly an aqueous solution or an aqueous / organic solution, based on a fermentation process that has not been worked up or has been worked up as described herein, for example.
[0083] An "enzymatically catalyzed" or "biocatalytic" method means that the method is carried out under the catalysis of an enzyme comprising an enzyme mutant as defined herein. Thus, these methods can be carried out in the presence of the enzyme in isolated (purified, concentrated) form or crude form, or in the presence of a cell system, in particular a natural or recombinant microbial cell containing the active form of the enzyme and having the ability to catalyze the conversion reactions disclosed herein.
[0084] Where the present disclosure refers to features, parameters and ranges thereof with different degrees of preference (including general, implicitly preferred features, parameters and ranges thereof), unless otherwise stated, any combination of two or more such features, parameters and ranges thereof is encompassed by the disclosure herein regardless of their respective degrees of preference.
[0085] Detailed Description a. Specific embodiments of the present invention The present invention relates to the following specific embodiments: 1. The first main embodiment is of the general formula 1
Chemical formula
[0086] The polypeptide of this embodiment having "terpenyl diphosphate phosphatase activity" is identified as a member of the protein tyrosine phosphatase family, particularly as a member of the Y_phosphatase 3 family having the Pfam ID number PF13350.
[0087] 2. The second main embodiment of the present invention is a method for the biocatalytic production of a bicyclic diterpene alcohol compound, a) contacting the corresponding bicyclic diterpenyl diphosphate precursor of the bicyclic diterpene compound with a polypeptide having terpenyl-diphosphate phosphatase activity, such as diterpenyl-diphosphate phosphatase activity or, in particular, bicyclic diterpenyl-diphosphate phosphatase activity, to produce the bicyclic diterpene alcohol; b) optionally isolating the bicyclic diterpene alcohol of step (1); Relates to a method comprising.
[0088] The polypeptide of this embodiment having "terpenyl diphosphate phosphatase activity" is identified as a member of the protein tyrosine phosphatase family, particularly as a member of the Y_phosphatase 3 family having the Pfam ID number PF13350.
[0089] 3. The method according to embodiment 2, wherein the polypeptide having the terpenyl-diphosphate phosphatase activity is selected from diphosphate-removing enzyme members of the protein tyrosine phosphatase family.
[0090] 4. The polypeptide having the terpenyl-diphosphate phosphatase activity has the following active site signature motif: HCxxGxxR (SEQ ID NO: 57) [Here, each x independently represents any natural amino acid residue] A method according to embodiment 1 or 3, selected from a class of diphosphate removal enzymes characterized by an amino acid sequence having
[0091] 5. The active site signature motif is HC(T / S)xGKDRTG (SEQ ID NO: 58) [where x represents any natural amino acid residue, for example selected from L, A, G and V residues] A method according to embodiment 4, which is
[0092] In another embodiment, the polypeptide having the terpenyl-diphosphate phosphatase activity comprises the amino acid consensus sequence motif shown in FIG. 16b.
[0093] 6. The polypeptide having the terpenyl-diphosphate phosphatase activity is the following polypeptide: a) TalVeTPP comprising an amino acid sequence according to SEQ ID NO: 2, b) AspWeTPP comprising an amino acid sequence according to SEQ ID NO: 6, c) HelGriTPP comprising an amino acid sequence according to SEQ ID NO: 10, d) UmbPiTPP1 comprising an amino acid sequence according to SEQ ID NO: 13, e) TalVeTPP2 comprising an amino acid sequence according to SEQ ID NO: 16, f) HydPiTPP1 comprising an amino acid sequence according to SEQ ID NO: 19, g) TalCeTPP1 comprising an amino acid sequence according to SEQ ID NO: 22, h) TalMaTPP1 comprising an amino acid sequence according to SEQ ID NO: 25, i) TalAstroTPP1 comprising an amino acid sequence according to SEQ ID NO: 28, and j) PeSubTPP1 comprising an amino acid sequence according to SEQ ID NO: 31, and k) having geranylgeranyl diphosphate phosphatase activity and having at least 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% sequence identity with at least one of the amino acid sequences according to a) to j), a polypeptide comprising an amino acid sequence A method according to any one of the preceding embodiments, selected from the group consisting of
[0094] 7. A method according to any one of embodiments 1 and 4 to 6 for producing a terpene alcohol compound of general formula 1 [wherein R represents H or, inter alia, preferably an acyclic, linear or branched, saturated or unsaturated hydrocarbyl residue having a total carbon number divisible by 5, such as 5, 10, 15 or 20].
[0095] 8. The method according to embodiment 7, wherein the terpene alcohol of formula 1 is selected from farnesol and geranylgeraniol.
[0096] 9. The method according to any one of embodiments 2 to 6, wherein step (1) also includes contacting an acyclic geranylgeranyl diphosphate precursor with a polypeptide having bicyclic diterpenyl diphosphate synthase activity to produce the bicyclic diterpenyl diphosphate precursor.
[0097] 10. The bicyclic diterpenyl diphosphate synthase is a) SmCPS2 comprising the amino acid sequence according to SEQ ID NO: 34, b) TaTps1-del59 comprising the amino acid sequence according to SEQ ID NO: 40, c) SsLPS comprising the amino acid sequence according to SEQ ID NO: 38, and d) having bicyclic diterpenyl diphosphate synthase activity and having at least 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% sequence identity with at least one of the amino acid sequences according to a), b) and c), a polypeptide comprising an amino acid sequence The method according to embodiment 9, selected from
[0098] 11. The method according to any one of embodiments 2 to 6, embodiment 9 and 10, wherein the bicyclic diterpene alcohol produced biocatalytically is copalol, in particular (+)-copalol and labdenediol, in either substantially pure enantiomeric form or in the form of a mixture of at least two enantiomers, respectively, selected from
[0099] 12. The method according to any one of the preceding embodiments, further comprising, as step (3), processing of the terpene alcohol of step (1) or step (2) to an alcohol derivative using chemical synthesis and / or biocatalytic synthesis or a combination of both
[0100] 13. The method according to embodiment 12, wherein the alcohol derivative is a hydrocarbon, alcohol, diol, triol, acetal, ketal, aldehyde, acid, ether, amide, ketone, lactone, epoxide, acetate, glycoside and / or ester
[0101] 14. The method according to embodiment 12 or 13, wherein the terpene alcohol is biocatalytically oxidized
[0102] 15. The method according to embodiment 14, wherein the terpene alcohol is converted by contacting with alcohol dehydrogenase (ADH)
[0103] 16. The ADH is a) CymB comprising an amino acid sequence according to SEQ ID NO: 42; b) AspWeADH1 comprising an amino acid sequence according to SEQ ID NO: 44; c) PsAeroADH1 comprising an amino acid sequence according to SEQ ID NO: 46; d) AzTolADH1 comprising an amino acid sequence according to SEQ ID NO: 48; e) AroAroADH1 comprising an amino acid sequence according to SEQ ID NO: 50; f) ThTerpADH1 comprising the amino acid sequence according to SEQ ID NO: 52; g) CdGeoA comprising the amino acid sequence according to SEQ ID NO: 54; h) VoADH1 comprising the amino acid sequence according to SEQ ID NO: 56; and i) a polypeptide having ADH activity and comprising an amino acid sequence showing a degree of sequence identity of at least 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% with at least one of the amino acid sequences according to a) to h) The method according to embodiment 15, selected from
[0104] 17. (1) contacting copalyl diphosphate with a polypeptide having copalyl diphosphate (CPP) phosphatase activity to produce copalol in the form of a substantially pure stereoisomer, in particular (+)-copalol, or in the form of a mixture of at least two stereoisomers; (2) optionally isolating the copalol of step (1); The method according to any one of embodiments 2 to 6 and embodiments 9 to 11 for the biocatalytic production of copalol, comprising
[0105] 18. The polypeptide having copalyl diphosphate phosphatase activity is the following polypeptide: a) TalVeTPP, comprising the amino acid sequence according to SEQ ID NO: 2; b) AspWeTPP, comprising the amino acid sequence according to SEQ ID NO: 6; c) HelGriTPP, comprising the amino acid sequence according to SEQ ID NO: 10; d) UmbPiTPP1, comprising the amino acid sequence according to SEQ ID NO: 13; e) TalVeTPP2, comprising the amino acid sequence according to SEQ ID NO: 16; f) HydPiTPP1, comprising the amino acid sequence according to SEQ ID NO: 19; g) TalCeTPP1, comprising the amino acid sequence according to SEQ ID NO: 22; h) TalMaTPP1, comprising an amino acid sequence according to SEQ ID NO: 25, i) TalAstroTPP1, comprising an amino acid sequence according to SEQ ID NO: 28, and j) PeSubTPP1, comprising an amino acid sequence according to SEQ ID NO: 31, and k) a polypeptide having copalyl diphosphate phosphatase activity and comprising an amino acid sequence showing at least 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% sequence identity with at least one of the said amino acid sequences according to a) to j) The method according to embodiment 17, selected from the group consisting of.
[0106] 19. The method according to embodiment 17 or 18, wherein step (1) also comprises the biocatalytic conversion of a terpene pyrophosphate, such as geranylgeranyl - pyrophosphate (GGPP), or a mixture of isopentenyl pyrophosphate (IPP) and dimethylallyl pyrophosphate (DMAPP), to copalyl diphosphate (CPP) by the catalysis of copalyl pyrophosphate synthase (CPS).
[0107] 20. The said CPS is a) SmCPS2, comprising an amino acid sequence according to SEQ ID NO: 34, b) TaTps1 - del59, comprising an amino acid sequence according to SEQ ID NO: 40, and c) a polypeptide having copalyl pyrophosphate synthase activity and comprising an amino acid sequence showing at least 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% sequence identity with at least one of the said amino acid sequences according to a) and b) The method according to embodiment 19, selected from the group consisting of.
[0108] 21. The method according to any one of embodiments 17 to 20, further comprising, as step (3), processing of copalol in step (1) or step (2) to a copalol derivative using chemical synthesis, biocatalytic synthesis, or a combination of both.
[0109] 22. The method according to embodiment 21, wherein the copalol derivative is a hydrocarbon, alcohol, diol, triol, acetal, ketal, aldehyde, acid, ether, amide, ketone, lactone, epoxide, acetate, glycoside, and / or ester.
[0110] 23. The method according to embodiment 21 or 22, wherein copalol is biocatalytically oxidized.
[0111] 24. The method according to embodiment 23, wherein copalol is oxidized by contacting with alcohol dehydrogenase (ADH).
[0112] 25. The ADH is a) CymB comprising an amino acid sequence according to SEQ ID NO: 42; b) AspWeADH1 comprising an amino acid sequence according to SEQ ID NO: 44; c) PsAeroADH1 comprising an amino acid sequence according to SEQ ID NO: 46; d) AzTolADH1 comprising an amino acid sequence according to SEQ ID NO: 48; e) AroAroADH1 comprising an amino acid sequence according to SEQ ID NO: 50; f) ThTerpADH1 comprising an amino acid sequence according to SEQ ID NO: 52; g) CdGeoA comprising an amino acid sequence according to SEQ ID NO: 54; h) VoADH1 comprising an amino acid sequence according to SEQ ID NO: 56; and i) a polypeptide having ADH activity, comprising an amino acid sequence having at least 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% sequence identity with at least one of the amino acid sequences according to a) to h). The method according to embodiment 24, selected from
[0113] 26. (1) Contacting labdenediol diphosphate (also called labda-13-ene-8-ol diphosphate or 8α-hydroxycopalyl diphosphate) with a polypeptide having labdenediol diphosphate (LPP) phosphatase activity to produce labdenediol in the form of a substantially pure stereoisomer or in the form of a mixture of at least two stereoisomers; (2) Optionally isolating the labdenediol of step (1); and The method according to embodiment 1 for the biocatalytic production of labdenediol, comprising
[0114] 27. The polypeptide having LPP phosphatase activity is the following polypeptide: a) TalVeTPP, comprising an amino acid sequence according to SEQ ID NO: 2; b) AspWeTPP, comprising an amino acid sequence according to SEQ ID NO: 6; c) HelGriTPP, comprising an amino acid sequence according to SEQ ID NO: 10; d) UmbPiTPP1, comprising an amino acid sequence according to SEQ ID NO: 13; e) TalVeTPP2, comprising an amino acid sequence according to SEQ ID NO: 16; f) HydPiTPP1, comprising an amino acid sequence according to SEQ ID NO: 19; g) TalCeTPP1, comprising an amino acid sequence according to SEQ ID NO: 22; h) TalMaTPP1, comprising an amino acid sequence according to SEQ ID NO: 25; i) TalAstroTPP1, comprising an amino acid sequence according to SEQ ID NO: 28, and j) PeSubTPP1, comprising an amino acid sequence according to SEQ ID NO: 31, and k) having LPP phosphatase activity and having at least 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% sequence identity with at least one of the amino acid sequences according to a) to j), a polypeptide comprising an amino acid sequence The method according to embodiment 26, selected from the group consisting of.
[0115] 28. The method according to any one of embodiments 26 and 27, wherein step (1) also includes the biocatalytic conversion of a terpene pyrophosphate such as geranylgeranyl-pyrophosphate (GGPP), or a mixture of isopentenyl pyrophosphate (IPP) and dimethylallyl pyrophosphate (DMAPP) to labdenediol diphosphate (LPP) by the catalysis of labdenediol pyrophosphate synthase (LPS).
[0116] 29. The said LPS is a) SsLPS comprising the amino acid sequence according to SEQ ID NO: 38, and b) having labdenediol pyrophosphate synthase activity and having at least 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% sequence identity with at least one of the amino acid sequences according to a), a polypeptide comprising an amino acid sequence The method according to embodiment 28, selected from the group consisting of.
[0117] 30. The method according to any one of embodiments 26 to 29, further comprising, as step (3), the processing of the labdenediol in step (1) or step (2) to a labdenediol derivative using chemical synthesis or biocatalytic synthesis or a combination of both.
[0118] 31. The method according to embodiment 30, wherein the labdenediol derivative is a hydrocarbon, alcohol, diol, triol, acetal, ketal, aldehyde, acid, ether, amide, ketone, lactone, epoxide, acetate, glycoside and / or ester.
[0119] 32. The method according to embodiment 30 or 31, wherein lavendiol is oxidatively biocatalyzed.
[0120] 33. The method according to embodiment 32, wherein lavendiol is oxidized by contacting it with alcohol dehydrogenase (ADH).
[0121] 34. The ADH is a) CymB comprising an amino acid sequence according to SEQ ID NO: 42; b) AspWeADH1 comprising an amino acid sequence according to SEQ ID NO: 44; c) PsAeroADH1 comprising an amino acid sequence according to SEQ ID NO: 46; d) AzTolADH1 comprising an amino acid sequence according to SEQ ID NO: 48; e) AroAroADH1 comprising an amino acid sequence according to SEQ ID NO: 50; f) ThTerpADH1 comprising an amino acid sequence according to SEQ ID NO: 52; g) CdGeoA comprising an amino acid sequence according to SEQ ID NO: 54; h) VoADH1 comprising an amino acid sequence according to SEQ ID NO: 56; i) SCH23 - ADH1 comprising an amino acid sequence according to SEQ ID NO: 68 j) SCH24 - ADH1a comprising an amino acid sequence according to SEQ ID NO: 70; and k) a polypeptide having ADH activity and comprising an amino acid sequence showing at least 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% sequence identity with at least one of the amino acid sequences according to a) - j) The method according to embodiment 33, selected from
[0122] 35. The method according to embodiment 19 or 28, further comprising the biocatalytic production of geranylgeranyl pyrophosphate (GGPP) from farnesyl pyrophosphate (FPP) by the catalytic action of geranylgeranyl pyrophosphate synthase (GGPS).
[0123] 36. Wherein the GGPS is a) a polypeptide comprising an amino acid sequence according to SEQ ID NO: 36, and b) a polypeptide having geranylgeranyl pyrophosphate synthase activity and having at least 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% sequence identity with at least one of the amino acid sequences according to a), The method according to embodiment 35, selected from
[0124] 37. The method according to any one of the preceding embodiments, carried out in vitro or in vivo.
[0125] 38. Prior to step (1), introducing into a non-human host organism or cell, optionally stably integrated into each genome, one or more nucleic acid molecules encoding one or more polypeptides having the enzyme activity necessary to carry out each biocatalytic conversion step, The method according to any one of the preceding embodiments, carried out in vivo and comprising
[0126] 39. The nucleic acid introduced into the non-human host organism or cell is a) at least one polypeptide having terpenyl-diphosphate phosphatase activity, particularly bicyclic diterpenyl-diphosphate phosphatase activity; and optionally b) at least one polypeptide having terpenyl-diphosphate synthase activity, particularly bicyclic diterpenyl-diphosphate synthase activity, and / or c) at least one polypeptide having ADH activity; and / or d) at least one polypeptide having acyclic terpenyl-diphosphate synthase activity, particularly acyclic diterpenyl-diphosphate synthase activity The method according to embodiment 38, encoding
[0127] 40. The nucleic acid introduced into the non-human host organism or cell is a) A polypeptide defined in Embodiment 6; or a nucleotide sequence selected from SEQ ID NOs: 1, 3, 4, 5, 7, 8, 9, 11, 12, 14, 15, 17, 18, 20, 21, 23, 24, 26, 27, 29, 30, and 32; or a polypeptide encoded by a nucleotide sequence having at least 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% sequence identity to any one of the foregoing sequences, at least one polypeptide having bicyclic diterpenyl-diphosphate phosphatase activity, and optionally the following b) A polypeptide defined in Embodiment 10; or a nucleotide sequence selected from SEQ ID NOs: 33, 37, and 39; or a polypeptide encoded by a nucleotide sequence having at least 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% sequence identity to any one of the foregoing sequences, at least one polypeptide having bicyclic diterpenyl-diphosphate synthase activity c) A polypeptide defined in Embodiment 16; or a nucleotide sequence selected from SEQ ID NOs: 41, 43, 45, 47, 49, 51, 53, and 55; or a polypeptide encoded by a nucleotide sequence having at least 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% sequence identity to any one of the foregoing sequences, at least one polypeptide having ADH activity d) At least one polypeptide having acyclic diterpenyl-diphosphate synthase activity, defined in Embodiment 36; or a nucleotide sequence selected from SEQ ID NO: 35 or a nucleotide sequence having at least 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% sequence identity to the foregoing sequence, at least one polypeptide having acyclic diterpenyl-diphosphate synthase activity The method according to Embodiment 39, encoding at least one of the above.
[0128] 41. A non-human host organism or cell that endogenously produces FPP and / or GGPP; or a mixture of IPP and DMAPP; or a non-human host organism that has been genetically modified to produce increased amounts of FPP and / or GGPP and / or a mixture of IPP and DMAPP, the method according to any one of embodiments 38 to 40, which is carried out by applying the same.
[0129] Some of these host cells or organisms do not naturally produce FPP or GGPP or a mixture of IPP and DMAPP, or produce too little to be considered, and thus do not endogenously produce an amount of FPP or GGPP or a mixture of IPP and DMAPP that should be increased. In order to be suitable for carrying out the methods of the embodiments described herein, organisms or cells that do not naturally produce acyclic terpene pyrophosphate precursors, such as FPP or GGPP or a mixture of IPP and DMAPP, or organisms or cells that produce a sub-optimal amount of said compounds, are genetically modified to produce said precursors. They can be transformed in this way, for example, before or simultaneously with the modification by the nucleic acids described according to any of the above embodiments. Methods for transforming organisms so that they produce acyclic terpene pyrophosphate precursors, such as FPP or GGPP or a mixture of IPP and DMAPP, are known in the art. For example, the introduction of enzyme activities of the mevalonate pathway, the isoprenoid pathway or the MEP pathway, particularly the mevalonate pathway, is a suitable strategy for making organisms produce FPP or GGPP or a mixture of IPP and DMAPP.
[0130] 42. The method according to any one of embodiments 38 to 41, wherein the non-human host organism or cell is a eukaryote or a prokaryote, particularly a plant, a bacterium or a fungus, particularly yeast.
[0131] 43. The method according to embodiment 42, wherein the bacterium is of the genus Escherichia, particularly E. coli, and the yeast is of the genus Saccharomyces, particularly S. cerevisiae.
[0132] 44. The method according to embodiment 42, wherein the cell is a plant cell.
[0133] 45. A non-human host organism or cell as defined in any one of embodiments 38 to 44.
[0134] 46. A recombinant nucleic acid construct comprising at least one nucleic acid molecule as defined in any one of embodiments 38 to 44.
[0135] 47. An expression vector comprising at least one nucleic acid construct according to embodiment 46.
[0136] 48. The expression vector according to embodiment 47, which is a prokaryotic vector, a viral vector, a eukaryotic vector, or one or more plasmids.
[0137] 49. A recombinant non-human host organism or cell as defined in embodiment 45, transformed with at least one nucleic acid construct according to embodiment 46 or at least one vector according to embodiment 47 or 48.
[0138] 50. A polypeptide having terpenyl-diphosphate phosphatase activity, particularly bicyclic diterpenyl-diphosphate phosphatase activity, selected from a diphosphate removal enzyme member of the protein tyrosine phosphatase family and mutants or variants thereof, which preferably catalyzes the conversion of terpenyl diphosphate to each terpenyl alcohol with a selectivity of more than 50%, such as more than 60%, more than 70%, more than 80%, more than 90%, more than 95% or more than 99%. In particular, this polypeptide preferably catalyzes the conversion of at least one terpenyl diphosphate selected from CPP and LPP to copalol and labdenediol, which are each terpenyl alcohol, with a selectivity of more than 50%, such as more than 60%, more than 70%, more than 80%, more than 90%, more than 95% or more than 99%.
[0139] The polypeptide of this embodiment having "terpenyl diphosphate phosphatase activity" is identified as a member of the protein tyrosine phosphatase family, particularly as a member of the Y_phosphatase 3 family having Pfam ID number PF13350.
[0140] In particular, the polypeptide of the present invention having "terpenyl diphosphate phosphatase activity" is identified as a member of the protein tyrosine phosphatase family, particularly as a member of the Y_phosphatase 3 family having Pfam ID number PF13350 when the bit score is above the collection threshold of the Pfam domain. The expected value (e-value) can also be used as a criterion for determining whether the queried protein is included in the Pfam family or whether the queried protein has a specific Pfam domain. The match with the said domain is, for example, in the range of 1×10 -40 ~7.40×10 -80 or in the range of 3.50×10 -50 ~7.40×10 -66 such as in the range of 1×10 -45 ~1×10 -70 or less than 1×10 -5 or less than 1×10 -10 or less than 1×10-20 It has an e-value less than. As the query array, the sequence of a polypeptide having "terpenyl diphosphate phosphatase activity" is applied.
[0141] For example, the following websites may be applied for searching and calculating such e-values: https: / / pfam.xfam.org / search#tabview=tab0 or https: / / www.ebi.ac.uk / Tools / hmmer / .
[0142] In one preferred alternative, such a phosphatase enzyme also converts FPP and / or GGPP into farnesol and geranylgeraniol, which are the respective alcohols.
[0143] In another preferred alternative, such a phosphatase enzyme does not convert FPP and / or GGPP into farnesol and geranylgeraniol, which are the respective alcohols, while retaining the ability to convert at least one bicyclic diterpenyl diphosphate selected from CPP and LPP into copalol and labdenediol, which are the respective terpene alcohols.
[0144] In another preferred alternative, such phosphatase enzymes produce at least one alcohol selected from copalol and labdenediol as the main product. In that case, such enzymes convert FPP and / or GGPP to the respective alcohols farnesol and geranylgeraniol in a lower molar yield compared to their ability to convert at least one bicyclic diterpenyl diphosphate selected from CPP and LPP to the respective terpene alcohols copalol and labdenediol. The relative molar yield of at least one bicyclic diterpenyl alcohol selected from copalol and labdenediol can be more than 2-fold higher, such as 2 to 1,000-fold or 5 to 100-fold, or 10 to 50-fold, compared to the yield of at least one of the acyclic terpene alcohols farnesol and geranylgeraniol.
[0145] 51. The following active site signature motif: HCxxGxxR (SEQ ID NO: 57) [where each x independently represents any natural amino acid residue] The polypeptide according to embodiment 50, characterized by an amino acid sequence having
[0146] 52. The active site signature motif is HC(T / S)xGKDRTG (SEQ ID NO: 58) [where x represents any natural amino acid residue] The polypeptide according to embodiment 51, which is
[0147] 53. The polypeptide having said bicyclic diterpenyl-diphosphate phosphatase activity is the following polypeptide: a) TalVeTPP comprising an amino acid sequence according to SEQ ID NO: 2, b) AspWeTPP comprising an amino acid sequence according to SEQ ID NO: 6, c) HelGriTPP comprising an amino acid sequence according to SEQ ID NO: 10, d) UmbPiTPP1, comprising an amino acid sequence according to SEQ ID NO: 13, e) TalVeTPP2, comprising an amino acid sequence according to SEQ ID NO: 16, f) HydPiTPP1, comprising an amino acid sequence according to SEQ ID NO: 19, g) TalCeTPP1, comprising an amino acid sequence according to SEQ ID NO: 22, h) TalMaTPP1, comprising an amino acid sequence according to SEQ ID NO: 25, i) TalAstroTPP1, comprising an amino acid sequence according to SEQ ID NO: 28, j) PeSubTPP1, comprising an amino acid sequence according to SEQ ID NO: 31, and k) a polypeptide having diterpene - diphosphate phosphatase activity and comprising an amino acid sequence having at least 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% sequence identity with at least one of the amino acid sequences according to a) - j) The polypeptide according to any one of embodiments 50 to 52, selected from the group consisting of
[0148] Another specific embodiment refers to a polypeptide variant of the novel polypeptide of the present invention having bicyclic diterpene - diphosphate phosphatase activity, identified above by any one of the specific amino acid sequences of SEQ ID NOs: 2, 6, 10, 13, 16, 19, 22, 25, 28, and 31, wherein the polypeptide variant is selected from amino acid sequences having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% sequence identity with any one of SEQ ID NOs: 2, 6, 10, 13, 16, 19, 22, 25, 28, and 31 and comprises at least one substitution modification with respect to any one of the unmodified SEQ ID NOs: 2, 6, 10, 13, 16, 19, 22, 25, 28, and 31.
[0149] 54. a) A nucleic acid sequence encoding the polypeptide according to any one of embodiments 50 to 53; b) The complementary nucleic acid sequence of a); or c) A nucleic acid sequence that hybridizes to the nucleic acid sequence of a) or b) under stringent conditions A nucleic acid molecule comprising the same.
[0150] 55. An expression construct comprising at least one nucleic acid molecule according to claim 54.
[0151] 56. A vector comprising at least one nucleic acid molecule according to claim 54.
[0152] 57. The vector according to claim 56, which is a prokaryotic vector, a viral vector or a eukaryotic vector.
[0153] 58. The vector according to embodiment 56 or 57, which is an expression vector.
[0154] 59. The vector according to any one of embodiments 56 to 58, which is a plasmid vector.
[0155] 60. a) At least one isolated nucleic acid molecule according to embodiment 54, optionally stably integrated into the genome; or b) At least one expression construct according to embodiment 55, optionally stably integrated into the genome; or c) At least one vector according to any one of embodiments 56 to 59 A recombinant host cell or a recombinant non-human host organism comprising the same.
[0156] 61. The host cell or host organism according to embodiment 60, which is selected from prokaryotic microorganisms or eukaryotic microorganisms, or cells derived therefrom.
[0157] 62. The host cell or host organism according to embodiment 61, which is selected from bacterial cells, fungal cells and plant cells or plants.
[0158] 63. The host cell or host organism according to embodiment 62, wherein the fungal cell is a yeast cell.
[0159] 64. The host cell or host organism according to embodiment 63, wherein the bacterial cell is selected from the genus Escherichia, particularly from the species E. coli, and the yeast cell is selected from the genus Saccharomyces or the genus Pichia, particularly from the species Saccharomyces cerevisiae or the species Pichia pastoris.
[0160] 65. (1) Culturing a non-human host organism or host cell according to any one of claims 60 to 64 to express or overexpress at least one polypeptide according to any one of embodiments 50 to 53; (2) Optionally isolating the polypeptide from the non-human host cell or organism cultured in step (1). A method for producing at least one catalytically active polypeptide according to any one of embodiments 50 to 53, comprising:
[0161] 66. The method according to embodiment 65, further comprising, prior to step a), transforming a non-human host organism or cell with at least one nucleic acid according to embodiment 54, at least one construct according to embodiment 55, or at least one vector according to any one of embodiments 56 to 59 so as to express or overexpress the polypeptide according to any one of embodiments 50 to 53.
[0162] 67. (1) Selecting a nucleic acid molecule according to embodiment 54; (2) Modifying the selected nucleic acid molecule to obtain at least one mutant nucleic acid molecule; (3) Transforming a host cell or unicellular host organism with the mutant nucleic acid sequence to express the polypeptide encoded by the mutant nucleic acid sequence. (4) screening for expression products to obtain at least one mutant having terpene synthase activity, particularly terpene diphosphate synthase activity; (5) optionally, if the polypeptide does not have the desired mutant activity, repeating process steps (1)-(4) until a polypeptide having the desired mutant activity is obtained; (6) optionally, if a polypeptide having the desired mutant activity is identified in step (4), isolating the corresponding mutant nucleic acid obtained in step (3); A method for producing a mutant polypeptide having terpene synthase activity, particularly terpene diphosphate synthase activity, comprising:
[0163] 68. The method according to embodiment 22, wherein the copalol derivative is selected from the group consisting of copalol, manool, (+)-manooloxy, Z-11, γ-ambrinol and ambrx and particularly structurally related compounds having a different stereochemistry therefrom.
[0164] 69. The method according to embodiment 31, wherein the labdendiol derivative is selected from the group consisting of sclaridol, DOL and ambrx and particularly structurally related compounds having a different stereochemistry therefrom.
[0165] 70. a) providing the labdendiol or copalol compound by performing the biocatalytic process defined in any one of embodiments 1 to 44 and optionally isolating the labdendiol or copalol compound; b) converting the labdendiol or copalol compound of step (1) to an ambrx or ambrx-like compound using chemical synthesis and / or biochemical synthesis; A method for producing the ambrx or ambrx-like compound as defined above, comprising:
[0166] 71. The present invention further relates to the use of a polypeptide as defined in any one of the above embodiments for producing an odorant, a flavor or a fragrance component, in particular Ambrox; and to the use of a terpene alcohol produced according to any one of the above embodiments for producing an odorant, a flavor or a fragrance component, in particular Ambrox.
[0167] b. Polypeptides applicable according to the present invention The following definitions apply in this context: The general terms "polypeptide" or "peptide", which can be used interchangeably, refer to a natural or synthetic linear chain or sequence of peptide-bonded contiguous amino acid residues containing from about 10 to a maximum of over 1,000 residues. Short-chain polypeptides containing a maximum of 30 residues are also called "oligopeptides".
[0168] The term "protein" refers to a macromolecular structure consisting of one or more species of polypeptides. The amino acid sequence of the polypeptide is the "primary structure" of the protein. This amino acid sequence also predetermines the "secondary structure" of the protein by the formation of special structural elements such as α-helix structures and β-sheet structures formed within the polypeptide chain. The arrangement of a plurality of such secondary structure elements defines the "tertiary structure", i.e., the spatial arrangement of the protein. When a protein contains more than one polypeptide chain, the said chains are spatially arranged to form the "quaternary structure" of the protein. The exact spatial arrangement of the protein, i.e., "folding", is a prerequisite for protein function. Denaturation or unfolding destroys protein function. If such destruction is reversible, protein function can be restored by refolding.
[0169] A typical protein function referred to herein is the "enzymatic function", i.e., the protein acts as a biocatalyst on a substrate, e.g., a compound, and catalyzes the conversion of the substrate to a product. The enzyme may exhibit high or low substrate specificity and / or product specificity.
[0170] Thus, a "polypeptide" described herein as having a particular "activity" implicitly refers to a correctly folded protein that exhibits the indicated activity, such as a particular enzymatic activity.
[0171] Thus, unless otherwise specified, the term "polypeptide" also encompasses the terms "protein" and "enzyme".
[0172] Similarly, the term "polypeptide fragment" encompasses the terms "protein fragment" and "enzyme fragment".
[0173] The term "isolated polypeptide" is known in the art and refers to an amino acid sequence removed from its natural environment by any method or combination of methods, including recombinant, biochemical, and synthetic methods.
[0174] A "target peptide" refers to an amino acid sequence that targets a protein, or an intracellular organelle, i.e., a mitochondrion or plastid, or a polypeptide for the extracellular space (secretory signal peptide). The nucleic acid sequence encoding the target peptide may be fused to a nucleic acid sequence encoding the amino terminus, e.g., the N-terminus, of a protein or polypeptide, or may be used to replace the native target polypeptide.
[0175] The present invention also relates to "functional equivalents" (also referred to as "analogs" or "functional mutants") of the polypeptides specifically described herein. The present invention also relates to "functional equivalents" (also referred to as "analogs" or "functional mutants") of the polypeptides specifically described herein.
[0176] For example, a "functional equivalent" refers to a polypeptide that exhibits at least 1 to 10%, or at least 20%, or at least 50%, or at least 75%, or at least 90% higher or lower activity than the polypeptides specifically described herein in a test used to measure enzymatic terpenyl diphosphate synthase activity or terpenyl diphosphate phosphatase activity.
[0177] According to the present invention, a "functional equivalent" also includes a specific mutant having an amino acid that is different from the specifically described amino acid at at least one sequence position of the amino acid sequence described herein, but still has one of the aforementioned biological activities such as enzymatic activity. Thus, a "functional equivalent" can be obtained by addition, substitution, particularly conservative substitution, deletion, and / or inversion of one or more amino acids such as 1 to 20, particularly 1 to 15 or 5 to 10, and can occur at any sequence position provided that the described changes result in a mutant having the profile of the properties according to the present invention. In particular, if the activity pattern is qualitatively the same between the mutant and the unmodified polypeptide, i.e., for example, the interaction with the same agonist or antagonist or substrate (i.e., represented by EC 50 value or IC 50 value or any other parameter suitable in the art) is observed, functional equivalence is likewise given. Examples of suitable (similar) amino acid substitutions are shown in the following table:
Table 1
[0178] A "functional equivalent" in the above sense is also a "precursor" of the polypeptide described herein, as well as a "functional derivative" and a "salt" of this polypeptide.
[0179] In that case, a "precursor" is a natural or synthetic precursor of a polypeptide with or without the desired biological activity.
[0180] The term "salt" means a carboxyl group salt and an acid addition salt of an amino group of the protein molecule according to the present invention. The carboxyl group salt can be produced by known methods and includes inorganic salts such as sodium salts, calcium salts, ammonium salts, iron salts and zinc salts, as well as salts with organic bases such as amines such as triethanolamine, arginine, lysine, piperidine and the like. Acid addition salts, such as salts with inorganic acids such as hydrochloric acid or sulfuric acid, and salts with organic acids such as acetic acid and oxalic acid are also included according to the present invention.
[0181] The "functional derivative" of the polypeptide according to the present invention can also be produced using known techniques on the amino acid side groups of the functional groups or at its N-terminus or C-terminus. Such derivatives include, for example, aliphatic esters of carboxylic acid groups, amides of carboxylic acid groups obtainable by reaction with ammonia or primary or secondary amines; N-acyl derivatives of free amino groups produced by reaction with acyl groups; or O-acyl derivatives of free hydroxyl groups produced by reaction with acyl groups.
[0182] "Functionally equivalent" includes polypeptides that can be obtained from other organisms naturally, as well as naturally occurring variants. For example, the portion of the homologous sequence region can be established by sequence comparison, and equivalent polypeptides can be determined based on the specific parameters of the present invention.
[0183] "Functionally equivalent" also includes "fragments" of the polypeptide according to the present invention, such as individual domains or sequence motifs, or N- and / or C-terminal truncated forms, which may or may not exhibit the desired biological function. Preferably, such "fragments" retain at least qualitatively the desired biological function.
[0184] A "functional equivalent" is further a fusion protein, which has one of the polypeptide sequences described herein or a functional equivalent derived therefrom and at least one additional, functionally different heterologous sequence in N-terminal or C-terminal association with a functional group (i.e., without substantial mutual dysfunction of the fusion protein moieties). Non-limiting examples of such heterologous sequences are, for example, signal peptides, histidine anchors or enzymes.
[0185] "Functional equivalents" also included according to the invention are homologues of the specifically disclosed polypeptides. These have at least 60%, preferably at least 75%, particularly at least 80% or 85%, for example 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98% or 99% homology (or identity) with one of the specifically disclosed amino acid sequences calculated by the algorithm of Pearson and Lipman, Proc. Natl. Acad, Sci. (USA) 85(8), 1988, 2444-2448. The homology or identity of the homologous polypeptides according to the invention expressed as a percentage means the identity of the amino acid residues expressed as a percentage based on the full length of one of the amino acid sequences specifically described herein.
[0186] Identity data expressed as a percentage can also be determined using BLAST alignment, the algorithm blastp (protein-protein BLAST), or by applying the clustal settings specified below herein.
[0187] In the case of possible protein glycosylation, "functional equivalents" according to the invention include the polypeptides described herein in deglycosylated or glycosylated form and in modified forms obtainable by altering the glycosylation pattern.
[0188] Functional equivalents or homologs of the polypeptide according to the present invention can be generated by mutagenesis, for example, point mutation, extension or shortening of the protein, or as described in more detail below.
[0189] Functional equivalents or homologs of the polypeptide according to the present invention can be identified by screening combinatorial databases of mutants, such as truncated mutants. For example, by combinatorial mutagenesis at the nucleic acid level, such as by enzymatic ligation of a mixture of synthetic oligonucleotides, a diverse database of protein mutants can be generated. There are a great many methods that can be used to generate a database of potential homologs from degenerate oligonucleotide sequences. Chemical synthesis of degenerate gene sequences can be performed on an automated DNA synthesizer and then the synthetic genes can be ligated into a suitable expression vector. Using a degenerate genome makes it possible to supply, in a mixture, all the sequences encoding the desired set of potential protein sequences. Methods for synthesizing degenerate oligonucleotides are known to those skilled in the art.
[0190] In the prior art, several techniques are known for screening gene products of combinatorial databases generated by point mutations or deletions, and for screening cDNA libraries of gene products having selected characteristics. These techniques can be adapted to rapidly screen gene banks created by combinatorial mutagenesis of homologs according to the present invention. The techniques most frequently used for screening large gene banks based on high-throughput analysis include cloning the gene bank into a replicable expression vector, transforming suitable cells with the resulting vector database, and expressing the combinatorial genes under conditions that facilitate isolation of the vector encoding the gene whose product was detected by detecting the desired activity. To identify homologs, iterative ensemble mutagenesis (REM), a technique for increasing the frequency of functional mutants in a database, can be used in combination with the screening assay.
[0191] Embodiments described herein provide orthologs and paralogs of the polypeptides disclosed herein and methods for identifying and isolating such orthologs and paralogs. Definitions of the terms "ortholog" and "paralog" are provided below and apply to amino acid and nucleic acid sequences.
[0192] c. Coding nucleic acid sequences applicable according to the present invention The following definitions apply in this context: The terms "nucleic acid sequence", "nucleic acid", "nucleic acid molecule" and "polynucleotide" are used interchangeably and mean a sequence of nucleotides. A nucleic acid sequence may be a single-stranded or double-stranded deoxyribonucleotide or ribonucleotide of any length, and may include coding and non-coding sequences of genes, exons, introns, sense and antisense complementary sequences, genomic DNA, cDNA, miRNA, siRNA, mRNA, rRNA, tRNA, recombinant nucleic acid sequences, isolated and purified naturally occurring DNA sequences and / or RNA sequences, synthetic DNA sequences and RNA sequences, fragments, primers and nucleic acid probes. Those skilled in the art will recognize that the nucleic acid sequence of RNA is the same as the DNA sequence except that thymine (T) is replaced by uracil (U). The term "nucleotide sequence" is also to be understood to include polynucleotide molecules or oligonucleotide molecules in the form of separate fragments or as components of larger nucleic acids.
[0193] "Isolated nucleic acid" or "isolated nucleic acid sequence" refers to a nucleic acid or nucleic acid sequence that can include a nucleic acid or nucleic acid sequence that is in an environment different from its natural environment and is substantially free of contamination by endogenous materials.
[0194] As used herein, the term "naturally occurring", when applied to nucleic acids, refers to nucleic acids that are found naturally in the cells of organisms and have not been intentionally modified by humans in the laboratory.
[0195] A "fragment" of a polynucleotide or nucleic acid sequence refers to a continuous nucleotide sequence where the length of the polynucleotide in the embodiments herein is particularly at least 15 bp, at least 30 bp, at least 40 bp, at least 50 bp and / or at least 60 bp. In particular, a fragment of a polynucleotide includes at least 25, especially at least 50, especially at least 75, especially at least 100, especially at least 150, especially at least 200, especially at least 300, especially at least 400, especially at least 500, especially at least 600, especially at least 700, especially at least 800, especially at least 900, especially at least 1000 continuous nucleotides of the polynucleotide in the embodiments herein. Without limitation, fragments of polynucleotides can be used herein as PCR primers and / or probes, or for antisense gene silencing or RNAi.
[0196] As used herein, the terms "hybridization" or "hybridize" under certain conditions are intended to refer to hybridization and washing conditions under which significantly identical or homologous nucleotide sequences remain bound to each other. Such conditions can be those under which sequences that are at least about 70%, such as at least about 80%, and such as at least about 85%, 90%, or 95% identical remain bound to each other. Definitions of low stringency, medium, and high stringency hybridization conditions are described herein below. Those of skill in the art can also select appropriate hybridization conditions with minimal experimentation as exemplified in Ausubel et al. (1995, Current Protocols in Molecular Biology, John Wiley & Sons, sections 2, 4, and 6). Additionally, stringency conditions are described in Sambrook et al. (1989, Molecular Cloning: A Laboratory Manual, 2nd ed., Cold Spring Harbor Press, chapters 7, 9, and 11).
[0197] A "recombinant nucleic acid sequence" is a nucleic acid sequence obtained by using laboratory methods (e.g., molecular cloning) to gather genetic material from more than one source and to create or modify a nucleic acid sequence that does not occur naturally and would not otherwise be found in a biological organism.
[0198] "Recombinant DNA technology" refers to the molecular biology procedures for manufacturing recombinant nucleic acid sequences as described, for example, in Laboratory Manuals edited by Weigel and Glazebrook, 2002, Cold Spring Harbor Lab Press; and Sambrook et al., 1989, Cold Spring Harbor, NY, Cold Spring Harbor Laboratory Press.
[0199] The term "gene" means a DNA sequence that includes a region that is transcribed into an RNA molecule, such as mRNA within a cell, which is operably linked to a suitable regulatory region, such as a promoter. Thus, a gene can include several operably linked sequences, such as a promoter, a 5' leader sequence that includes sequences involved in translation initiation, a coding region of cDNA or genomic DNA, introns, exons, and / or a 3' untranslated sequence that includes, for example, a transcription termination point.
[0200] "Polycistronic" refers to a nucleic acid molecule, particularly an mRNA, that can separately encode more than one polypeptide within the same nucleic acid molecule. "Chimeric gene" refers to any gene that does not normally occur naturally in a species, particularly a gene in which one or more portions of nucleic acid sequences that do not naturally associate with each other in nature are present. For example, a promoter does not naturally associate with some or all of a transcription region or with another regulatory region. The term "chimeric gene" is understood to include an expression construct in which a promoter or transcription regulatory sequence is operably linked to one or more coding sequences or to an antisense, i.e., the reverse complementary strand of the sense strand or an inverted repeat sequence (sense and antisense, whereby the RNA transcript forms double-stranded RNA during transcription). The term "chimeric gene" also includes a gene obtained by combining portions of one or more coding sequences to create a new gene.
[0201] "3' UTR" or "3' untranslated sequence" (also referred to as "3' untranslated region" or "3' end") refers to a nucleic acid sequence found downstream of the coding sequence of a gene, which includes, for example, a transcription termination point and a polyadenylation signal (although not all, but most eukaryotic mRNAs), such as AAUAAA or a variant thereof. After transcription termination, the mRNA transcript may be cleaved downstream of the polyadenylation signal and further a poly(A) tail may be added, which is involved in the translation site, such as the transport of mRNA to the cytoplasm.
[0202] The term "primer" refers to a short nucleic acid sequence that hybridizes to a template nucleic acid sequence and is used for polymerization of a nucleic acid sequence complementary to the template.
[0203] The term "selectable marker" refers to any gene that can be used upon expression to select cells containing the selectable marker. Examples of selectable markers are described below. It will be appreciated by those skilled in the art that various selectable markers for antibiotics, fungicides, auxotrophs or herbicides are applicable to various target species.
[0204] The present invention also relates to nucleic acid sequences encoding the polypeptides defined herein.
[0205] In particular, the present invention also relates to nucleic acid sequences (single-stranded and double-stranded DNA sequences and RNA sequences, such as cDNA, genomic DNA and mRNA) encoding one of the above polypeptides and their functional equivalents, and this nucleic acid sequence can be obtained, for example, using artificial nucleotide analogs.
[0206] The present invention relates to both an isolated nucleic acid molecule encoding a polypeptide according to the present invention or a biologically active segment thereof, and a nucleic acid fragment that can be used, for example, as a hybridization probe or primer for identifying or amplifying the coding nucleic acid according to the present invention.
[0207] The present invention also relates to nucleic acids having a certain degree of "identity" with the sequences specifically disclosed herein. "Identity" between two nucleic acids means nucleotide identity over the full length of the nucleic acids in each case.
[0208] "Identity" between two nucleotide sequences (and the same applies to peptide or amino acid sequences) is a function of the number of nucleotide residues (or amino acid residues), or the number of nucleotide residues (or amino acid residues) that are identical in the two sequences when an alignment of the two sequences is generated. Identical residues are defined as residues that are the same in the two sequences at a given position in the alignment. The percentage of sequence identity, as used herein, is calculated from the optimal alignment by dividing the number of residues that are identical between the two sequences by the total number of residues in the shortest sequence and multiplying by 100. The optimal alignment is the alignment with the highest possible percentage of identity. Gaps can be introduced into one or both of the sequences at one or more positions in the alignment to obtain the optimal alignment. Such gaps are then considered as non-identical residues for the calculation of the percentage of sequence identity. Alignments for determining the percentage of amino acid or nucleic acid sequence identity can be achieved in a variety of ways using computer programs and publicly available computer programs, for example, available on the World Wide Web.
[0209] In particular, an optimal alignment of a protein or nucleic acid sequence can be obtained and the percentage of sequence identity calculated using the BLAST program (Tatiana et al, FEMS Microbiol Lett., 1999, 174:247-250, 1999) set to the default parameters available from the website at ncbi.nlm.nih.gov / BLAST / bl2seq / wblast2.cgi of the National Center for Biotechnology Information (NCBI).
[0210] In another example, identity can be calculated using the Clustal method (Higgins DG, Sharp PM. ((1989))) with the following settings by the Vector NTI Suite 7.1 program from Informax, Inc. (USA): Multiple alignment parameters: Gap opening penalty 10 Gap extension penalty 10 Gap length penalty range 8 Gap length penalty off % identity for alignment delay 40 Residue-specific gap off Hydrophilic residue gap off Transition weighting 0 Pairwise alignment parameters: FAST algorithm on K-tuple size 1 Gap penalty 3 Window size 5 Number of best diagonals 5 Alternatively, identity can be determined according to Chenna, et al. (2003), web page: http: / / www.ebi.ac.uk / Tools / clustalw / index.html# and the following settings DNA gap open penalty 15.0 DNA gap extension penalty 6.66 DNA matrix identity Protein gap open penalty 10.0 Protein gap extension penalty 0.2 Protein matrix Gonnet Protein / DNA ENDGAP -1 Protein / DNA GAPDIST 4 All nucleic acid sequences described herein (single-stranded and double-stranded DNA and RNA sequences, such as cDNA and mRNA) can be made by chemical synthesis from nucleotide building blocks, for example, by fragment condensation of individual overlapping, complementary nucleic acid building blocks of a double helix, by known methods. Chemical synthesis of oligonucleotides can be carried out by known methods, for example, by the phosphoramidite method (Voet, Voet, 2nd edition, Wiley Press, New York, pages 896-897). Accumulation of synthetic oligonucleotides and filling of gaps and ligation reactions by the Klenow fragment of DNA polymerase and general cloning techniques are described in Sambrook et al. (1989) (see below).
[0211] In addition, the nucleic acid molecules according to the invention can include untranslated sequences from the 3' and / or 5' ends of the coding gene region.
[0212] The invention further relates to nucleic acid molecules that are complementary to the specifically described nucleotide sequences or segments thereof.
[0213] The nucleotide sequences according to the invention enable the preparation of probes and primers that can be used to identify and / or clone homologous sequences in other cell types and organisms. Such probes or primers generally hybridize under "stringent" conditions (as defined elsewhere herein) to nucleotide sequence regions that include at least about 12, preferably at least about 25, for example about 40, 50 or 75 consecutive nucleotides of the sense strand or the corresponding antisense strand of the nucleic acid sequence according to the invention.
[0214] "Homologous" sequences include orthologous or paralogous sequences. Methods for identifying orthologs or paralogs, including phylogenetic methods, sequence similarity methods and hybridization methods, are known in the art and are described herein.
[0215] "Paralogs" result from gene duplication that gives rise to two or more genes with similar sequences and similar functions. Paralogs typically cluster together and are formed by gene duplication within related plant species. Paralogs are found in groups of similar genes using pairwise Blast analysis or during phylogenetic analysis of gene families using programs such as CLUSTAL. In paralogs, consensus sequences characteristic of the sequences within related genes and sequences with similar gene functions can be identified.
[0216] "Orthologs", or orthologous sequences, are sequences that are similar to each other because they are found in species that are descendants of a common ancestor. For example, plant species with a common ancestor are known to contain many enzymes with similar sequences and functions. Those skilled in the art can identify orthologous sequences and predict the function of orthologs, for example, by constructing a polydendrogram of a gene family of one species using a CLUSTAL program or a BLAST program. Methods for identifying or confirming similar functions between homologous sequences are by overexpressing or lacking the relevant polypeptide in a host cell or organism such as a plant or microorganism (in knockout / knockdown) and comparing the transcriptional profiles. Those skilled in the art will understand that genes with similar transcriptional profiles having more than 50% common regulatory transcripts, or more than 70% common regulatory transcripts, or more than 90% common regulatory transcripts will have similar functions. Homologs, paralogs, orthologs, and any other variants of the sequences herein are expected to function similarly by creating host cells or organisms such as plants or microorganisms that produce the terpenesynthase protein.
[0217] The term "selectable marker" refers to any gene that can be used at the time of expression to select cells containing the selectable marker. Examples of selectable markers are described below. It will be understood by those skilled in the art that various selectable markers for antibiotics, fungicides, auxotrophs or herbicides are applicable to various target species.
[0218] An "isolated" nucleic acid molecule is separated from other nucleic acid molecules that are present in the natural source of the nucleic acid, and can further be made substantially free of other cellular material or culture medium if produced by recombinant techniques, or free of chemical precursors or other chemicals if chemically synthesized.
[0219] The nucleic acid molecules according to the present invention can be isolated by standard techniques of molecular biology and the sequence information provided by the present invention. For example, using one of the entire sequences specifically disclosed as a hybridization probe or a segment thereof and standard hybridization techniques (described, for example, in Sambrook, (1989)), cDNA can be isolated from a suitable cDNA library.
[0220] In addition, nucleic acid molecules containing one or a segment of the disclosed sequences can be isolated by polymerase chain reaction using oligonucleotide primers constructed based on this sequence. The nucleic acid amplified in this way can be cloned into a suitable vector and characterized by DNA sequencing. The oligonucleotides according to the present invention can also be prepared by standard synthetic methods, for example using an automated DNA synthesizer.
[0221] The nucleic acid sequences according to the present invention or derivatives thereof, homologs or parts of these sequences can be isolated from other bacteria, for example, by conventional hybridization techniques or PCR techniques, for example via genomic or cDNA libraries. These DNA sequences hybridize with the sequences according to the present invention under standard conditions.
[0222] "Hybridize" means that a polynucleotide or oligonucleotide can bind to a nearly complementary sequence under standard conditions, but non-specific binding does not occur under these conditions between non-complementary partners. For this reason, the sequences can be 90 - 100% complementary. The property of complementary sequences that can bind specifically to each other is utilized, for example, in Northern blotting or Southern blotting, or in primer binding for PCR or RT-PCR.
[0223] Oligonucleotides with short conserved regions are advantageously used for hybridization. However, it is also possible to use longer fragments or the entire sequence of the nucleic acid according to the present invention for hybridization. These "standard conditions" vary depending on the nucleic acid used (oligonucleotide, longer fragment or entire sequence), or on whether either type of nucleic acid - DNA or RNA - is used for hybridization. For example, the melting temperature of a DNA:DNA hybrid is approximately 10°C lower than that of a DNA:RNA hybrid of the same length.
[0224] For example, depending on the particular nucleic acid, the standard conditions mean a temperature between 42 and 58 °C in a buffered aqueous solution having a concentration between 0.1 and 5× SSC (1× SSC = 0.15 M NaCl, 15 mM sodium citrate, pH 7.2), or alternatively conditions in which 50% formamide is present, for example 42 °C in 5× SSC, 50% formamide. It is advantageous for the hybridization conditions of DNA:DNA hybrids to be at a temperature between 0.1× SSC and between about 20 °C and 45 °C, preferably between about 30 °C and 45 °C. In the case of DNA:RNA hybrids, the hybridization conditions are preferably at a temperature between 0.1× SSC and between about 30 °C and 55 °C, preferably between about 45 °C and 55 °C. These temperatures described for hybridization are examples of melting temperature values calculated for nucleic acids of about 100 nucleotides in length and a G+C content of 50% in the absence of formamide. The experimental conditions for DNA hybridization are described in relevant genetics textbooks, for example Sambrook et al., 1989, and can be calculated using formulas known to those skilled in the art, depending, for example, on the length of the nucleic acid, the type of hybrid or the G+C content. Those skilled in the art can obtain further information on hybridization from the following textbooks: Ausubel et al. (eds), (1985), Brown (ed) (1991).
[0225] "Hybridization" can be carried out in particular under stringent conditions. Such hybridization conditions are described, for example, in Sambrook (1989), or Current Protocols in Molecular Biology, John Wiley & Sons, N.Y. (1989), 6.3.1-6.3.6.
[0226] As used herein, the terms hybridization or hybridize under certain conditions are intended to refer to hybridization and washing conditions under which significantly identical or homologous nucleotide sequences remain bound to each other. Such conditions can be those under which sequences that are at least about 70%, such as at least about 80%, and such as at least about 85%, 90%, or 95% identical remain bound to each other. Definitions of low stringency, medium, and high stringency hybridization conditions are described herein.
[0227] As exemplified by Ausubel et al. (1995, Current Protocols in Molecular Biology, John Wiley & Sons, sections 2, 4, and 6), one of ordinary skill in the art can select appropriate hybridization conditions with a minimum of experimentation. In addition, stringency conditions are described in Sambrook et al. (1989, Molecular Cloning: A Laboratory Manual, 2nd ed., Cold Spring Harbor Press, chapters 7, 9, and 11).
[0228] As used herein, the defined conditions of low stringency are as follows. A filter containing DNA is pretreated for 6 hours at 40°C in a solution containing 35% formamide, 5×SSC, 50 mM Tris-HCl (pH 7.5), 5 mM EDTA, 0.1% PVP, 0.1% Ficoll, 1% BSA, and 500 μg / ml denatured salmon sperm DNA. Hybridization is carried out in the same solution with the following modifications: 0.02% PVP, 0.02% Ficoll, 0.2% BSA, 100 μg / ml salmon sperm DNA, 10% (wt / vol) dextran sulfate, and 5 - 20×106 32P-labeled probe are used. The filter is incubated in the hybridization mixture for 18 - 20 hours at 40°C and then washed at 55°C for 1.5 hours in a solution containing 2×SSC, 25 mM Tris-HCl (pH 7.4), 5 mM EDTA, and 0.1% SDS. This washing solution is replaced with fresh solution and incubated for a further 1.5 hours at 60°C. The filter is blotted dry and exposed for autoradiography.
[0229] As used herein, the conditions defined for medium stringency are as follows. A filter containing DNA is pretreated for 7 hours at 50°C in a solution containing 35% formamide, 5×SSC, 50 mM Tris-HCl (pH 7.5), 5 mM EDTA, 0.1% PVP, 0.1% Ficoll, 1% BSA, and 500 μg / ml denatured salmon sperm DNA. Hybridization is carried out in the same solution with the following modifications: 0.02% PVP, 0.02% Ficoll, 0.2% BSA, 100 μg / ml salmon sperm DNA, 10% (wt / vol) dextran sulfate, and 5 - 20×106 32P-labeled probe are used. The filter is incubated in the hybridization mixture at 50°C for 30 hours and then washed at 55°C for 1.5 hours in a solution containing 2×SSC, 25 mM Tris-HCl (pH 7.4), 5 mM EDTA, and 0.1% SDS. This washing solution is replaced with fresh solution and incubated at 60°C for an additional 1.5 hours. The filter is blotted dry and exposed for autoradiography.
[0230] As used herein, the conditions defined for high stringency are as follows. Prehybridization of a filter containing DNA is carried out for 8 hours to overnight at 65°C in a buffer composed of 6×SSC, 50 mM Tris-HCl (pH 7.5), 1 mM EDTA, 0.02% PVP, 0.02% Ficoll, 0.02% BSA, and 500 μg / ml denatured salmon sperm DNA. The filter is hybridized at 65°C for 48 hours in a prehybridization mixture containing 100 μg / ml denatured salmon sperm DNA and 5 - 20×106 cpm of 32P-labeled probe. Washing of the filter is carried out for 1 hour at 37°C in a solution containing 2×SSC, 0.01% PVP, 0.01% Ficoll, and 0.01% BSA. This is followed by a wash in 0.1×SSC at 50°C for 45 minutes.
[0231] If the above conditions are inappropriate (e.g., when used for heterologous hybridization), other conditions of low, medium, and high stringency well known in the art (e.g., when used for heterologous hybridization) may be used.
[0232] A detection kit for a nucleic acid sequence encoding a polypeptide of the present invention may include primers and / or probes specific to the nucleic acid sequence encoding this polypeptide, and related protocols for using this primer and / or probe to detect the nucleic acid sequence encoding the polypeptide in a sample. Using such a detection kit, it is possible to determine whether a plant, organism, microorganism, or cell has been modified, i.e., transformed with a sequence encoding a polypeptide.
[0233] To test the function of a mutant DNA sequence according to the embodiments herein, the sequence of interest is operably linked to a selectable or screenable marker gene, and the expression of the reporter gene is tested in a transient expression assay, for example, using microorganisms or protoplasts, or in stably transformed plants.
[0234] The present invention also relates to derivatives of the specifically disclosed nucleic acid sequences or derivable nucleic acid sequences.
[0235] Accordingly, further nucleic acid sequences according to the present invention can be derived from the sequences specifically disclosed herein, and this further nucleic acid sequence may differ from the sequences specifically disclosed herein by one or more (e.g., 1 to 20, particularly 1 to 15 or 5 to 10) additions, substitutions, insertions, or deletions of one or several (e.g., 1 to 10, etc.) nucleotides, and furthermore, can encode a polypeptide having a desired profile of properties.
[0236] The present invention also encompasses nucleic acid sequences that contain so-called silent mutations, or nucleic acid sequences that have been altered compared to the specifically described sequences according to the codon usage frequency of the particular original or host organism.
[0237] According to certain embodiments of the present invention, mutant nucleic acids may be produced to adapt their nucleotide sequences to a particular expression system. For example, it is known that a bacterial expression system can express polypeptides more efficiently when the amino acids are encoded by specific codons. Due to the degeneracy of the genetic code, more than one codon may encode the same amino acid sequence, and multiple nucleic acid sequences can encode the same protein or polypeptide, and all of these DNA sequences are encompassed by the embodiments herein. Where appropriate, the nucleic acid sequences encoding the polypeptides described herein may be optimized to increase expression in a host cell. For example, the nucleic acids of the embodiments herein may be synthesized using codons specific to the host to improve expression.
[0238] The present invention also encompasses naturally occurring variants of the sequences described in the present invention, such as splicing variants or allelic variants.
[0239] Allelic variants can have at least 60% homology, preferably at least 80% homology, particularly very preferably at least 90% homology at the level of the derived amino acids over the entire sequence range (reference should be made to the details described above for polypeptides with regard to homology at the amino acid level). It can be advantageous for the homology to be higher over a partial region of the sequence.
[0240] The present invention also relates to sequences that can be obtained by conservative nucleotide substitutions (i.e., substitutions that result in the amino acid being replaced by an amino acid of the same charge, size, polarity, and / or solubility).
[0241] The present invention also relates to molecules derived from the specifically disclosed nucleic acids by sequence polymorphisms. Such genetic polymorphisms can exist within cells from different populations or within cells of a population due to natural allelic variations. Allelic variants can also include functional equivalents. These natural mutations usually cause a 1-5% variation in the nucleotide sequence of the gene. Said polymorphisms can bring about changes in the amino acid sequence of the polypeptides disclosed herein. Allelic variants can also include functional equivalents.
[0242] Furthermore, derivatives should also be understood to be homologs of the nucleic acid sequences according to the invention, such as homologs of animals, plants, fungi or bacteria, truncated sequences, single-stranded DNA or RNA of coding and non-coding DNA sequences. For example, homologs have at least 40%, preferably at least 60%, particularly preferably at least 70%, and very particularly preferably at least 80% homology over the entire DNA region given in the sequences specifically disclosed herein at the DNA level.
[0243] Furthermore, derivatives should be understood, for example, as fusions with promoters. The promoter added to the described nucleotide sequence can be modified by at least one nucleotide exchange, at least one insertion, inversion and / or deletion without impairing the functionality or effectiveness of the promoter. Furthermore, the effectiveness of the promoter can be enhanced by changing its sequence, or it can even be completely replaced with a more effective promoter from an organism of a different genus.
[0244] d. Generation of functional polypeptide mutants Furthermore, those skilled in the art are proficient in methods for generating nucleotide sequences encoding polypeptides encoded by nucleic acid molecules comprising functional mutants, i.e., polypeptides having at least 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98% or 99% sequence identity with any one of the amino acid-related sequence numbers disclosed herein, and / or having at least 70% sequence identity with any one of the nucleotide-related sequence numbers disclosed herein.
[0245] Depending on the techniques used, those skilled in the art can introduce either completely random mutations or, alternatively, more directed mutations into genes or, alternatively, non-coding nucleic acid regions (such as those important for the regulation of expression), followed by the construction of gene libraries. The methods of molecular biology necessary for this purpose are known to those skilled in the art and are described, for example, in Sambrook and Russell, Molecular Cloning. 3rd Edition, Cold Spring Harbor Laboratory Press 2001.
[0246] For example, methods for modifying genes such as the following, and thus methods for modifying the polypeptides encoded by the genes, have long been known to those skilled in the art.
[0247] - Site-directed mutagenesis (in which individual or several nucleotides of the gene are replaced in a directed manner) (Trower MK (Ed.) 1996; In vitro mutagenesis protocols. Humana Press, New Jersey), -Saturation mutagenesis (in this case, the codon of any amino acid can be replaced or added at any point in the gene) (Kegler-Ebo DM, Docktor CM, DiMaio D (1994) Nucleic Acids Res 22:1593; Barettino D, Feigenbutz M, Valcarel R, Stunnenberg HG (1994) Nucleic Acids Res 22:541; Barik S (1995) Mol Biotechnol 3:1), -Error-prone polymerase chain reaction (in this case, the nucleotide sequence is mutated by an error-prone DNA polymerase) (Eckert KA, Kunkel TA (1990) Nucleic Acids Res 18:3739); -SeSaM method (sequence saturation method) (in this case, preferred exchanges are prevented by the polymerase). Schenk et al., Biospektrum, Vol. 3, 2006, 277-279 -Passaging of the gene in a mutator strain (in this case, the mutation rate of the nucleotide sequence increases, for example, due to an imperfect DNA repair mechanism) (Greener A, Callahan M, Jerpseth B (1996) An efficient random mutagenesis technique using an E.coli mutator strain. In: Trower MK (Ed.) In vitro mutagenesis protocols. Humana Press, New Jersey), or -DNA shuffling (in this case, a pool of closely related genes is formed and digested, and the fragments are used as templates for polymerase chain reaction, where strand separation and reassociation are repeated, resulting in the final generation of full-length mosaic genes) (Stemmer WPC (1994) Nature 370:389; Stemmer WPC (1994) Proc Natl Acad Sci USA 91:10747).
[0248] Using so-called directed evolution (especially as described in Reetz MT and Jaeger K-E (1999), Topics Curr Chem 200:31; Zhao H, Moore JC, Volkov AA, Arnold FH (1999), Methods for optimizing industrial polypeptides by directed evolution, In: Demain AL, Davies JE (Ed.) Manual of industrial microbiology and biotechnology. American Society for Microbiology), one of ordinary skill in the art can generate functional mutants in a directed and large-scale manner. For this purpose, in a first step, a gene library of each polypeptide is first generated, for example using the methods described above. The gene library is expressed in a suitable manner, for example by a bacterial or phage display system.
[0249] The relevant genes of the host organism expressing the functional mutants having properties generally corresponding to the desired properties can be subjected to another mutation cycle. The steps of mutation and selection or screening can be repeatedly iterated until the existing functional mutants have the desired properties to a sufficient degree. Using this iterative procedure, a limited number of mutations, for example 1, 2, 3, 4 or 5 mutations, can be carried out stepwise and evaluated and selected with respect to their effect on the activity. The selected mutants can then be similarly subjected to further mutation steps. In this way, the number of individual mutants to be investigated can be significantly reduced.
[0250] The results according to the present invention also provide important information regarding the structure and sequence of the relevant polypeptides necessary to generate additional polypeptides targeted to have the desired modified properties. In particular, it is possible to identify so-called "hot spots", i.e., sequence segments potentially suitable for modifying a certain property by introducing a target mutation.
[0251] In that region, it is possible to make mutations that should probably have little effect on activity and can be called potential "silent mutations", and it is also possible to estimate information regarding the amino acid sequence position.
[0252] e. Constructs for expressing the polypeptides of the present invention The following definitions apply in this context: "Gene expression" includes "heterologous expression" and "overexpression" and includes the transcription of a gene and the translation of mRNA into protein. Overexpression refers to the production of a gene product measured by the level of mRNA, polypeptide and / or enzyme activity in a transgenic cell or organism that exceeds the level of production in non-transformed cells or organisms of a similar genetic background.
[0253] As used herein, an "expression vector" means a nucleic acid molecule engineered using methods of molecular biology and recombinant DNA technology for delivering exogenous or foreign DNA into a host cell. An expression vector typically contains sequences necessary for proper transcription of a nucleotide sequence. The coding region usually encodes the protein of interest, but can also encode RNA, such as antisense RNA, siRNA and the like.
[0254] As used herein, the "expression vector" includes any linear or circular recombinant vector including, but not limited to, viral vectors, bacteriophages, and plasmids. One of ordinary skill in the art can select a suitable vector according to the expression system. In one embodiment, the expression vector includes a nucleic acid or an mRNA ribosome binding site of an embodiment herein operably linked to at least one "regulatory sequence" that controls transcription, translation, initiation, and termination, such as a transcription promoter, operator, or enhancer, and optionally includes at least one selectable marker. A nucleotide sequence is "operably linked" when the regulatory sequence is functionally related to the nucleic acid of an embodiment herein.
[0255] As used herein, the "expression system" encompasses any combination of nucleic acid molecules necessary for the in vivo or in vitro expression of one polypeptide or the co-expression of two or more polypeptides in a given expression host. Each coding sequence can be located on a single nucleic acid molecule, or on a single vector such as a vector containing multiple cloning sites, or on a polycistronic nucleic acid, or can be distributed over two or more physically distinct vectors. As a specific example, reference can be made to an operon containing one promoter sequence, one or more operator sequences, and one or more structural genes encoding the enzymes described herein. As used herein, the terms "amplify" and "amplification" refer to the use of any suitable amplification technique to effect or detect the recombination of a naturally expressed nucleic acid, as described in detail below. For example, the present invention provides methods and reagents (such as specific degenerate oligonucleotide primer pairs, oligo dT primers) for amplifying a naturally expressed (e.g., genomic DNA or mRNA) nucleic acid or a recombinant (e.g., cDNA) nucleic acid of the present invention in vivo, ex vivo, or in vitro (e.g., by polymerase chain reaction, PCR).
[0256] "Regulatory sequence" refers to a nucleic acid sequence that can determine the expression level of a nucleic acid sequence of an embodiment herein and regulate the transcription rate of a nucleic acid sequence operably linked to the regulatory sequence. Regulatory sequences include promoters, enhancers, transcription factors, promoter elements, and the like.
[0257] According to the present invention, "promoter", "nucleic acid having promoter activity" or "promoter sequence" is understood to mean a nucleic acid that regulates the transcription of a nucleic acid when functionally linked to the nucleic acid to be transcribed. "Promoter" specifically refers to a nucleic acid sequence that controls the expression of a coding sequence by providing a binding site for RNA polymerase and other factors necessary for proper transcription, including but not limited to transcription factor binding sites, repressor and activator protein binding sites. The meaning of the term "promoter" also includes the term "promoter regulatory sequence". Promoter regulatory sequences can include upstream and downstream elements that can affect the transcription, RNA processing or stability of the associated coding nucleic acid sequence. Promoters include naturally occurring and synthetic sequences. The coding nucleic acid sequence is usually located downstream of the promoter with respect to the direction of transcription starting from the transcription start point.
[0258] In this context, a "functional" or "operable" linkage is understood to mean, for example, the sequential arrangement of one of the nucleic acids comprising the regulatory sequences. For example, a sequence having promoter activity, a nucleic acid sequence to be transcribed and optionally a sequence of further regulatory elements, such as a nucleic acid sequence ensuring transcription of the nucleic acid, and for example a terminator, are linked such that each of the regulatory elements can perform its function upon transcription of the nucleic acid sequence. This does not necessarily require a direct linkage in the chemical sense. Gene regulatory sequences, such as enhancer sequences, can exert their function from more distant positions or even from other DNA molecules with respect to the target sequence. A preferred arrangement is one in which the nucleic acid sequence to be transcribed is located behind (i.e., at the 3'-end) the promoter sequence such that the two sequences are covalently linked to each other. The distance between the promoter sequence and the nucleic acid sequence to be recombinantly expressed can be less than 200 base pairs, or less than 100 base pairs or less than 50 base pairs.
[0259] In addition to promoters and terminators, examples of other regulatory elements that can be mentioned are: targeting sequences, enhancers, polyadenylation signals, selectable markers, amplification signals, origins of replication and the like. Suitable regulatory sequences are described, for example, in Goeddel, Gene Expression Technology: Methods in Enzymology 185, Academic Press, San Diego, CA (1990).
[0260] The term "constitutive promoter" refers to an unregulated promoter that allows for continuous transcription of the nucleic acid sequence to which it is operably linked.
[0261] As used herein, the term "operably linked" refers to the joining of polynucleotide elements that are in a functional relationship. Nucleic acids are "operably linked" when placed in a functional relationship with another nucleic acid sequence. For example, a promoter, or rather a transcriptional regulatory sequence, is operably linked to a coding sequence if it affects the transcription of the coding sequence. Operable linkage means that the DNA sequences being linked are typically contiguous. The nucleotide sequences associated with a promoter sequence may be of the same or different origin with respect to the plant being transformed. This sequence may also be wholly or partially synthetic. Regardless of origin, the nucleic acid sequence associated with a promoter sequence will be expressed or silenced according to the promoter characteristics to which it is linked after binding to the polypeptide of the embodiments herein. The associated nucleic acid can encode a protein that is desirably expressed or repressed either constitutively throughout the organism or at a particular time or in a particular tissue, cell or cell compartment. Such nucleotide sequences particularly encode proteins that confer desirable phenotypic characteristics on the host cell or organism that has been altered or transformed thereby. In particular, the associated nucleotide sequence results in the production of the product of interest as defined herein in a cell or organism. In particular, the nucleotide sequence encodes a polypeptide having the enzymatic activity as defined herein.
[0262] As used herein, the above nucleotide sequence may be part of an "expression cassette". The terms "expression cassette" and "expression construct" are used synonymously. (Preferably recombinant) expression constructs contain a nucleotide sequence encoding a polypeptide according to the invention and under the genetic control of regulatory nucleic acid sequences.
[0263] In a process applied according to the invention, the expression cassette may be part of an "expression vector", particularly a recombinant expression vector.
[0264] According to the present invention, an "expression unit" means a nucleic acid having expression activity that contains a promoter defined in this specification and regulates the expression, that is, the transcription and translation of the nucleic acid or the gene, after being functionally bound to the nucleic acid or the gene to be expressed. Therefore, in this context, the expression unit is also referred to as a "regulatory nucleic acid sequence". In addition to the promoter, other regulatory elements, such as enhancers, can also be present.
[0265] According to the present invention, an "expression cassette" or "expression construct" is understood to mean an expression unit that is functionally bound to a nucleic acid to be expressed or a gene to be expressed. Therefore, in contrast to the expression unit, the expression cassette includes not only the nucleic acid sequence that regulates transcription and translation but also the nucleic acid sequence that is expressed as a protein as a result of transcription and translation.
[0266] In the context of the present invention, the terms "expression" or "overexpression" represent the production or increase of the intracellular activity of one or more polypeptides in a microorganism encoded by the corresponding DNA. For this purpose, for example, it is possible to introduce a gene into an organism, replace an existing gene with another gene, increase the copy number of a gene, use a strong promoter, or use a gene encoding the corresponding polypeptide having high activity, and optionally, these means can be combined.
[0267] Preferably, such constructs according to the present invention include, in each case, a promoter upstream of the 5' of each coding sequence and a terminator sequence downstream of the 3' of each coding sequence, which are operably bound to the coding sequence, and optionally other normal regulatory elements.
[0268] The nucleic acid constructs according to the invention in particular comprise sequences encoding polypeptides or derivatives and homologues thereof derived, for example, from the amino acid-related sequence numbers described in the invention or their reverse complementary strands, which are advantageously operably or functionally linked to one or more regulatory signals in order to control gene expression, for example to increase it.
[0269] In addition to these regulatory sequences, the natural regulation of these sequences may still be present in front of the actual structural gene and, furthermore optionally, the natural regulation may be switched off and genetically modified so as to increase the expression of the gene. However, the nucleic acid construct may also be of a simpler constitution, i.e. no additional regulatory signals have been inserted in front of the coding sequence and the natural promoter has not been removed by its regulation. Instead, the natural regulatory sequences are mutated so that regulation no longer takes place and gene expression increases.
[0270] Preferred nucleic acid constructs advantageously also comprise one or more of the previously described "enhancer" sequences functionally linked to a promoter, which sequences enable an improvement in the expression of the nucleic acid sequence. Additional advantageous sequences such as further regulatory elements or terminators may be inserted at the 3'-end of the DNA sequence. One or more copies of the nucleic acid according to the invention may be present in the construct. In the construct, other markers such as genes complementing auxotrophy or antibiotic resistance may optionally be present which allow the selection of this construct.
[0271] Examples of suitable regulatory sequences are cos, tac, trp, tet, trp-tet, lpp, lac, lpp-lac, lacI q , T7, T5, T3, gal, trc, ara, rhaP (rhaP BAD ) SP6, lambda-P R in promoters such as or lambda-P LThey are present in the promoter and are advantageously used in Gram-negative bacteria. Further advantageous regulatory sequences are present, for example, in the Gram-positive promoters amy and SPO2, in the yeast or fungal promoters ADC1, MFalpha, AC, P-60, CYC1, GAPDH, TEF, rp28, ADH. Artificial promoters may also be used for regulation.
[0272] For expression in the host organism, the nucleic acid construct is advantageously inserted into a vector such as a plasmid or a phage, which enables optimal expression of the gene in the host. Vectors include, in addition to plasmids and phages, all other vectors known to those skilled in the art, i.e., viruses such as SV40, CMV, baculovirus, and adenovirus, transposons, IS elements, fosmids, cosmids, and linear or circular DNA or artificial chromosomes. These vectors are capable of autonomous replication in the host organism or, if not, chromosomal replication. These vectors are a further development of the present invention. Binary vectors or cpo-integration vectors are also applicable.
[0273] Suitable plasmids are, for example, pLG338, pACYC184, pBR322, pUC18, pUC19, pKC30, pRep4, pHS1, pKK223-3, pDHE19.2, pHS2, pPLc236, pMBL^{24}, pLG200, pUR290, pIN-III 113 -B1, λgt11 or pBdCI in E. coli, pIJ101, pIJ364, pIJ702 or pIJ361 in Streptomyces, pUB110, pC194 or pBD214 in Bacillus, pSA77 or pAJ667 in Corynebacterium, pALS1, pIL2 or pBB116 in fungi, 2alphaM, pAG-1, YEp6, YEp13 or pEMBLYe23 in yeast or pLGV23, pGHlac in plants +, pBIN19, pAK2004 or pDH51. The plasmids described above are a selected few of the possible plasmids. Further plasmids are well known to those skilled in the art and can be found, for example, in the book Cloning Vectors (Eds. Pouwels P. H. et al. Elsevier, Amsterdam - New York - Oxford, 1985, ISBN 0 444 904018).
[0274] In a further development of the vector, the nucleic acid construct according to the invention or the vector containing the nucleic acid according to the invention can advantageously be introduced into a microorganism in the form of linear DNA and integrated into the genome of the host organism by non - homologous recombination or homologous recombination. This linear DNA can consist of a linearized vector such as a plasmid or only of the nucleic acid construct or nucleic acid according to the invention.
[0275] For optimal expression of a heterologous gene in an organism, it is advantageous to modify the nucleic acid sequence according to the specific "codon usage frequency" used in this organism. The "codon usage frequency" can be easily determined by computer evaluation of other known genes of the organism.
[0276] The expression cassette according to the invention is generated by fusing a suitable promoter to a suitable coding nucleotide sequence and a terminator or polyadenylation signal. For this purpose, conventional recombinant and cloning techniques such as those described, for example, in T. Maniatis, E.F. Fritsch and J. Sambrook, Molecular Cloning: A Laboratory Manual, Cold Spring Harbor Laboratory, Cold Spring Harbor, NY (1989) and T.J. Silhavy, M.L. Berman and L.W. Enquist, Experiments with Gene Fusions, Cold Spring Harbor Laboratory, Cold Spring Harbor, NY (1984) and Ausubel, F.M. et al., Current Protocols in Molecular Biology, Greene Publishing Assoc. and Wiley Interscience (1987) are used.
[0277] For expression in a suitable host organism, the recombinant nucleic acid construct or gene construct is advantageously inserted into a host-specific vector that enables optimal expression of the gene in the host. Vectors are well known to those skilled in the art and can be found, for example, in "cloning vectors" (Pouwels P. H. et al., Ed., Elsevier, Amsterdam-New York-Oxford, 1985).
[0278] Alternative embodiments of the embodiments herein provide a method of "modifying gene expression" in a host cell. For example, the polynucleotides of the embodiments herein can be enhanced or overexpressed or induced in a host cell or host organism in a certain context (e.g., when exposed to a certain temperature or culture conditions).
[0279] Changes in the expression of the polynucleotides provided herein can also result in ectopic expression, which is a different expression pattern in the modified organism and the control or wild-type organism. Changes in expression can occur due to the interaction of the polypeptides of the embodiments herein with exogenous or endogenous modulators, or as a result of chemical modification of the polypeptide. This term also refers to an altered expression pattern of the polynucleotides of the embodiments herein that has changed below the level of detection, or a completely suppressed activity.
[0280] In one embodiment, provided herein is also an isolated, recombinant or synthetic polynucleotide encoding a polypeptide or mutant polypeptide provided herein.
[0281] In one embodiment, several polypeptides encoding nucleic acid sequences are co-expressed in a single host, particularly under the control of different promoters. In another embodiment, several polypeptides encoding nucleic acid sequences can be present on a single transformation vector, or separate vectors can be used and transformants containing both chimeric genes can be selected for co-transformation simultaneously. Similarly, a gene encoding one or more polypeptides can be expressed in a single plant, cell, microorganism or organism together with other chimeric genes.
[0282] f. Hosts applicable to the present invention Depending on the context, the term "host" can mean a wild-type host or a genetically modified recombinant host or both.
[0283] In principle, all prokaryotes or eukaryotes can be considered as host organisms or recombinant host organisms for the nucleic acids or nucleic acid constructs according to the invention.
[0284] Using the vector according to the present invention, a recombinant host can be produced which is transformed with at least one vector according to the present invention and can be used to produce a polypeptide according to the present invention. It is advantageous that the recombinant construct according to the present invention described above is introduced into a suitable host system and expressed. Preferably, general cloning methods and transfection methods known to those skilled in the art for expressing the described nucleic acids in their respective expression systems are used, such as coprecipitation, protoplast fusion, electroporation, retroviral transfection and the like. Suitable systems are described, for example, in Current Protocols in Molecular Biology, F. Ausubel et al., Ed., Wiley Interscience, New York 1997, or Sambrook et al. Molecular Cloning: A Laboratory Manual. 2nd edition, Cold Spring Harbor Laboratory, Cold Spring Harbor Laboratory Press, Cold Spring Harbor, NY, 1989.
[0285] It is advantageous to use microorganisms such as bacteria, fungi or yeasts as host organisms. Gram-positive or Gram-negative bacteria are preferred, in particular bacteria of the family Enterobacteriaceae, Pseudomonas, Rhizobiaceae, Streptomycetaceae, Streptococcaceae or Nocardiaceae, particularly preferably bacteria of the genus Escherichia, Pseudomonas, Streptomyces, Lactococcus, Nocardia, Burkholderia, Salmonella, Agrobacterium, Clostridium or Rhodococcus. Bacteria of the genus and species Escherichia coli are very particularly preferred. Furthermore, other advantageous bacteria are found in the groups of α-proteobacteria, β-proteobacteria or γ-proteobacteria. Advantageously, yeasts of the family such as Saccharomyces or Pichia are also suitable hosts.
[0286] Alternatively, whole plants or whole plant cells can serve as natural or recombinant hosts. As non-limiting examples, mention may be made of the following plants or cells derived therefrom. The genus Nicotiana, in particular Nicotiana benthamiana and Nicotiana tabacum (tobacco); and Arabidopsis, in particular Arabidopsis thaliana.
[0287] Depending on the host organism, the organisms used in the method according to the invention are grown or cultured by methods known to those skilled in the art. The culturing can be batch, semi-batch or continuous. Nutrients may be present at the start of fermentation or supplied later semi-continuously or continuously. This will be explained in more detail below.
[0288] g. Recombinant production of the polypeptide according to the invention The invention further relates to a method for the recombinant production of a polypeptide according to the invention or a biologically active functional fragment thereof, wherein a microorganism producing the polypeptide is cultured and optionally the expression of the polypeptide is induced by applying at least one inducer that induces gene expression, and the expressed polypeptide is isolated from the culture. The polypeptide can also be produced in this way on an industrial scale if desired.
[0289] The microorganisms produced according to the invention can be cultured continuously or discontinuously by the batch method or the fed-batch method or the repeated fed-batch method. An overview of known culturing methods can be found in the textbook by Chmiel (Bioprozesstechnik 1. Einfuehrung in die Bioverfahrenstechnik [Bioprocess technology 1. Introduction to bioprocess technology] (Gustav Fischer Verlag, Stuttgart, 1991)) or in the textbook by Storhas (Bioreaktoren und periphere Einrichtungen [Bioreactors and peripheral equipment] (Vieweg Verlag, Braunschweig / Wiesbaden, 1994)).
[0290] The culture media used must appropriately meet the requirements of each strain. Descriptions of culture media for various microorganisms are given in the "Manual of Methods for General Bacteriology" of the American Society for Bacteriology (Washington D.C., USA, 1981).
[0291] These media that can be used in accordance with the present invention usually contain one or more carbon sources, nitrogen sources, inorganic salts, vitamins and / or trace elements.
[0292] Preferred carbon sources are sugars such as monosaccharides, disaccharides or polysaccharides. Very good carbon sources are, for example, glucose, fructose, mannose, galactose, ribose, sorbose, ribulose, lactose, maltose, sucrose, raffinose, starch or cellulose. The sugars can also be added to the medium via complex compounds such as molasses, or other by-products of refined sugar. It can also be advantageous to add a mixture of different carbon sources. Other possible carbon sources are oils and fats such as soybean oil, sunflower oil, peanut oil and coconut oil, fatty acids such as palmitic acid, stearic acid or linoleic acid, alcohols such as glycerol, methanol or ethanol and organic acids such as acetic acid or lactic acid.
[0293] The nitrogen source is usually an organic or inorganic nitrogen compound or a material containing these compounds. Examples of nitrogen sources include ammonia gas, or ammonium salts such as ammonium sulfate, ammonium chloride, ammonium phosphate, ammonium carbonate or ammonium nitrate, nitrates, urea, amino acids, or complex nitrogen sources such as corn steep liquor, soybean meal, soy protein, yeast extract, meat extract and others. The nitrogen source can be used alone or as a mixture.
[0294] Inorganic salt compounds that may be present in the medium include chlorides, phosphates or sulfates of calcium, magnesium, sodium, cobalt, molybdenum, potassium, manganese, zinc, copper and iron.
[0295] Inorganic sulfur-containing compounds such as sulfates, sulfites, dithionites, tetrathionates, thiosulfates, sulfides, and organic sulfur compounds such as mercaptans and thiols can be used as sulfur sources.
[0296] Phosphoric acid, potassium dihydrogen phosphate or dipotassium hydrogen phosphate or the corresponding sodium-containing salts can be used as phosphorus sources.
[0297] Chelating agents can be added to the medium to keep metal ions in solution. Particularly suitable chelating agents include dihydroxy phenols such as catechol or protocatechuate, or organic acids such as citric acid.
[0298] The fermentation medium used according to the present invention usually also contains other growth factors such as vitamins, or growth promoting substances including, for example, biotin, riboflavin, thiamine, folic acid, nicotinic acid, pantothenate and pyridoxine. Growth factors and salts often originate from components of complex media such as yeast extract, molasses, corn steep liquor and the like. Furthermore, suitable precursors can be added to the culture medium. The exact composition of the compounds in the medium strongly depends on each experiment and is determined individually for each specific case. Information on medium optimization can be found in the textbook "Applied Microbiol. Physiology, A Practical Approach" (Ed. P.M. Rhodes, P.F. Stanbury, IRL Press (1997) p. 53-73, ISBN 0 19 963577 3). Growth media such as Standard 1 (Merck) or BHI (Brain Heart Infusion, DIFCO) and the like can also be obtained from commercial vendors.
[0299] All components of the medium are sterilized by heat (20 minutes at 1.5 bar and 121 °C) or by sterile filtration. These components can be sterilized together or, if necessary, separately. All components of the medium may be present at the start of the culture or may be added continuously or in batches.
[0300] The culture temperature is usually between 15 °C and 45 °C, preferably between 25 °C and 40 °C, and can be varied or kept constant during the experiment. The pH of the medium should be in the range of 5 to 8.5, preferably approximately 7.0. The pH for growth can be controlled during growth by adding basic compounds such as sodium hydroxide, potassium hydroxide, ammonia or aqueous ammonia, or acidic compounds such as phosphoric acid or sulfuric acid. Antifoaming agents, such as fatty acid polyglycol esters, can be used to suppress foaming. Suitable selective substances, such as antibiotics, can be added to the medium to maintain plasmid stability. To maintain aerobic conditions, oxygen or an oxygen-containing gas mixture, such as ambient air, is supplied to the culture. The temperature of the culture is usually within the range of 20 °C to 45 °C. The culture is continued until the desired product is maximally produced. This goal is usually reached within 10 hours to 160 hours.
[0301] The fermentation broth is then further processed. Depending on the requirements, biomass can be completely or partially removed from the fermentation broth by separation techniques such as centrifugation, filtration, decanting or combinations of these methods, or the biomass can be left completely in the fermentation broth.
[0302] If the polypeptide is not secreted into the culture medium, the cells can also be lysed and the product obtained from the lysate by known methods for isolating proteins. The cells can optionally be disrupted by high-frequency ultrasound, high pressure (e.g. in a French press), by osmotic shock, by the action of surfactants, lytic enzymes or organic solvents, by a homogenizer, or by combinations of some of the aforementioned methods.
[0303] The polypeptide can be purified by known chromatographic techniques such as molecular sieve chromatography (gel filtration), such as Q-Sepharose chromatography, ion exchange chromatography and hydrophobic chromatography, and other conventional techniques such as ultrafiltration, crystallization, salting out, dialysis and non-denaturing gel electrophoresis. Suitable methods are described, for example, in Cooper, T.G., Biochemische Arbeitsmethoden [Biochemical processes], Verlag Walter de Gruyter, Berlin, New York or Scopes, R., Protein Purification, Springer Verlag, New York, Heidelberg, Berlin.
[0304] In order to isolate the recombinant protein, it may be advantageous to use a vector system or oligonucleotide that encodes a modified polypeptide or fusion protein, which extends the cDNA by a defined nucleotide sequence and thus, for example, serves for easier purification. Suitable modifications of this type are, for example, so-called "tags" that function as anchors, such as hexahistidine anchors that can be recognized as antigens of antibodies or modifications known as epitopes (for example, described in Harlow, E. and Lane, D., 1988, Antibodies: A Laboratory Manual. Cold Spring Harbor (N.Y.) Press). These anchors can serve to bind the protein to a solid support, such as a polymer matrix, which can be used as a packing material in a chromatography column, or can be used on microtiter plates or some other support.
[0305] At the same time, these anchors can also be used to recognize proteins. To recognize proteins, furthermore, conventional markers such as fluorescent dyes, enzyme markers (which produce a detectable reaction product after reaction with a substrate), or radioactive markers can be used alone or in combination with an anchor for derivatizing proteins.
[0306] h. Polypeptide immobilization The enzyme or polypeptide according to the present invention can be used free or immobilized in the methods described herein. An immobilized enzyme is an enzyme immobilized on an inert carrier. Suitable carrier materials and the enzymes immobilized thereon are known from European Patent Application Publication No. 1149849, European Patent Application Publication No. 1069183, German Patent Application Publication No. 100193773 and the references cited therein. In this regard, the entire disclosure of these documents is incorporated by reference. Suitable carrier materials include, for example, clay, clay minerals such as kaolinite, diatomaceous earth, perlite, silica, aluminum oxide, sodium carbonate, calcium carbonate, cellulose powder, anion exchanger materials, synthetic polymers such as polystyrene, acrylic resin, phenol formaldehyde resin, polyurethane and polyolefins such as polyethylene and polypropylene. For producing the supported enzyme, the carrier material is usually used in the form of finely divided particles, and a porous type is preferred. The particle size of the carrier material is usually 5 mm or less, particularly 2 mm or less (particle size distribution curve). Similarly, when using dehydrogenase as a whole cell catalyst, a free or immobilized form can be selected. The carrier materials are, for example, Ca alginate and carrageenan. The enzyme as well as the cells can also be directly crosslinked with glutaraldehyde (crosslinking to CLEA). Other corresponding immobilization techniques are described, for example, in J. Lalonde and A. Margolin “Immobilization of Enzymes” in K. Drauz and H. Waldmann, Enzyme Catalysis in Organic Synthesis 2002, Vol. III, 991-1032, Wiley-VCH, Weinheim. Further information on in vivo conversion and bioreactors for performing the method according to the present invention is also described, for example, in Rehm et al. (Ed.) Biotechnology, 2nd Edn, Vol 3, Chapter 17, VCH, Weinheim.
[0307] i. Reaction conditions for the biocatalytic production method of the present invention The reaction of the present invention can be carried out under in vivo or in vitro conditions.
[0308] At least one polypeptide / enzyme present between the individual steps of the method of the present invention or of the multi-step method defined earlier herein can be present in living cells that produce the enzyme naturally or recombinantly, in harvested cells, i.e., under in vivo conditions, or in dead cells, in permeabilized cells, in crude cell extracts, in purified extracts, or in a substantially pure or completely pure form, i.e., under in vitro conditions. At least one enzyme may be present as an enzyme immobilized in solution or on a support. One or more enzymes may be present simultaneously in soluble and / or immobilized forms.
[0309] The method according to the invention can be carried out on various scales, for example from laboratory scale (reaction volumes of a few millilitres to several tens of litres) to industrial scale (reaction volumes of several litres to several thousand cubic metres), in common reactors known to those skilled in the art. A chemical reactor can be used when the polypeptide is used in an encapsulated form by optionally permeabilized non-living cells, in the form of a somewhat purified cell extract or in a purified form. The chemical reactor usually enables control of the amount of at least one enzyme, the amount of at least one substrate, the pH, the temperature and the circulation of the reaction medium. If at least one polypeptide / enzyme is present within living cells, the process will be a fermentation. In this case, the biocatalytic production will take place in a bioreactor (fermenter), where the parameters necessary for suitable survival conditions for the living cells (for example, culture medium containing nutrients, temperature, aeration, presence or absence of oxygen or other gases, antibiotics and the like) can be controlled. Those skilled in the art are familiar with chemical reactors or bioreactors, for example, with procedures for scaling up chemical or biotechnological methods from laboratory scale to industrial scale, or with procedures for optimizing process parameters, which are also described in detail in the literature (for biotechnological methods, see, for example, Crueger und Crueger, Biotechnologie - Lehrbuch der angewandten Mikrobiologie, 2. Ed., R. Oldenbourg Verlag, Muenchen, Wien, 1984).
[0310] Cells containing at least one enzyme can be permeabilized by physical or mechanical means, such as ultrasonic or high-frequency pulses, French press, or chemical means, such as hypotonic media, lytic enzymes and surfactants present in the medium, or combinations of such methods. Examples of surfactants are digitonin, n-dodecyl maltoside, octyl glucoside, Triton® X-100, Tween® 20, deoxycholate, CHAPS (3-[(3-cholamidopropyl)dimethylammonio]-1-propanesulfonate), Nonidet® P40 (ethylphenol poly(ethylene glycol ether)) and the like.
[0311] Instead of living cells, a biomass of non-living cells containing the biocatalyst necessary for the in vivo conversion reaction of the present invention may be similarly applied.
[0312] When at least one enzyme is immobilized, it is bound to an inert carrier as described above.
[0313] The conversion reaction can be carried out batchwise, semi-batchwise or continuously. The reactants (and optionally nutrients) may be supplied at the start of the reaction or thereafter semi-continuously or continuously.
[0314] The reaction of the present invention may be carried out in an aqueous, aqueous-organic or non-aqueous reaction medium depending on the specific reaction type.
[0315] The aqueous or aqueous-organic medium may contain a suitable buffer to adjust the pH to a value within the range of 5 to 11, such as 6 to 10.
[0316] In an aqueous-organic medium, an organic solvent that is miscible, partially miscible or immiscible with water may be applied. Non-limiting examples of suitable organic solvents are given below. Further examples are monohydric or polyhydric aromatic or aliphatic alcohols, especially polyhydric aliphatic alcohols such as glycerol.
[0317] The non-aqueous medium may contain water, and may substantially not contain water, that is, it will contain less than about 1 wt.-% or 0.5 wt.-% of water.
[0318] The biocatalytic method may be carried out in an organic non-aqueous medium. Suitable organic solvents include aliphatic hydrocarbons having, for example, 5 to 8 carbon atoms such as pentane, cyclopentane, hexane, cyclohexane, heptane, octane or cyclooctane; aromatic hydrocarbons such as benzene, toluene, xylenes, chlorobenzene or dichlorobenzene, aliphatic acyclic ethers such as diethyl ether, methyl-tert.-butyl ether, ethyl-tert.-butyl ether, dipropyl ether, diisopropyl ether, dibutyl ether; or mixtures thereof.
[0319] The concentration of the reactant / substrate may be adapted to the optimal reaction conditions and may depend on the specific enzyme applied. For example, the initial substrate concentration may be within 0.1 to 0.5 M, such as 10 to 100 mM for example.
[0320] The reaction temperature may be adapted to the optimal reaction conditions and may depend on the specific enzyme applied. For example, the reaction may be carried out at a temperature within the range of 0 to 70 °C, such as 20 to 50 °C or 25 to 40 °C for example. Examples of reaction temperatures are about 30 °C, about 35 °C, about 37 °C, about 40 °C, about 45 °C, about 50 °C, about 55 °C and about 60 °C.
[0321] The process may proceed until equilibrium between the substrate and then the product is achieved, but may stop earlier. The normal process time is within the range of 1 minute to 25 hours, especially within the range of 10 minutes to 6 hours, such as within the range of 1 hour to 4 hours, especially 1.5 hours to 3.5 hours for example. These parameters are non-limiting examples of suitable process conditions.
[0322] When the host is a transgenic plant, optimal growth conditions can be provided, such as optimal light, water and nutrient conditions for example.
[0323] k. Product isolation The method of the present invention can further include the step of recovering the final product or intermediate product in a substantially pure form, optionally stereoisomerically or enantiomerically. The term "recovering" includes extracting, collecting, isolating or purifying a compound from a culture medium or reaction medium. The recovery of the compound can be carried out according to any conventional isolation or purification method known in the art, including but not limited to treatment with conventional resins (e.g., anion or cation exchange resins, nonionic adsorption resins, etc.), treatment with conventional adsorbents (e.g., activated carbon, silicic acid, silica gel, cellulose, alumina, etc.), pH change, solvent extraction (e.g., with conventional solvents such as alcohol, ethyl acetate, hexane and the like), distillation, dialysis, filtration, concentration, crystallization, recrystallization, pH adjustment, freeze-drying and the like.
[0324] The identity and purity of the isolated product can be determined by known techniques such as high performance liquid chromatography (HPLC), gas chromatography (GC), spectroscopy (IR, UV, NMR, etc.), colorimetry, TLC, NIRS, enzymatic or microbial assays. (See, for example, Patek et al. (1994) Appl. Environ. Microbiol. 60:133-140; Malakhova et al. (1996) Biotekhnologiya 11 27-32; and Schmidt et al. (1998) Bioprocess Engineer. 19:67-70. Ullmann’s Encyclopedia of Industrial Chemistry (1996) Bd. A27, VCH: Weinheim, pp. 89-90, pp. 521-540, pp. 540-547, pp. 559-566, 575-581 and pp. 581-587; Michal, G (1999) Biochemical Pathways: An Atlas of Biochemistry and Molecular Biology, John Wiley and Sons; Fallon, A. et al. (1987) Applications of HPLC in Biochemistry in: Laboratory Techniques in Biochemistry and Molecular Biology, Bd. 17). The cyclic terpene compounds produced in any of the methods described herein can be converted into derivatives such as, but not limited to, hydrocarbons, esters, amides, glycosides, ethers, epoxides, aldehydes, ketones, alcohols, diols, acetals or ketals. The terpene compound derivatives can be obtained by chemical methods such as, but not limited to, oxidation, reduction, alkylation, acylation and / or rearrangement. Alternatively, the terpene compound derivatives can be obtained using biochemical methods by contacting the terpene compound with enzymes such as, but not limited to, oxidoreductase, monooxygenase, dioxygenase, transferase. The biochemical conversion can be carried out in vitro using isolated enzymes, enzymes from lysed cells, or in vivo using whole cells.
[0325] l. Fermentative production of terpene alcohol The present invention also relates to a method for the fermentative production of terpene alcohol.
[0326] The fermentation used according to the present invention can be carried out, for example, in a stirred fermenter, a bubble column and a loop reactor. An overview of possible process types and geometric designs, including stirrer types, can be found in "Chmiel: Bioprozesstechnik: Einfuehrung in die Bioverfahrenstechnik, Band 1". In the process of the present invention, typical variants available are batch fermentation with and without biomass recycle, fed-batch fermentation, repeated fed-batch fermentation, or otherwise continuous fermentation, etc., which are known to those skilled in the art or the following variants described, for example, in "Chmiel, Hammes and Bailey: Biochemical Engineering". Depending on the production strain, sparging with air, oxygen, carbon dioxide, hydrogen, nitrogen or a suitable gas mixture may be carried out to achieve a good yield (YP / S).
[0327] The culture medium used must appropriately meet the requirements of a specific strain. Descriptions of culture media for various microorganisms are given in the handbook “Manual of Methods for General Bacteriology” of the American Society for Bacteriology (Washington D.C., USA, 1981).
[0328] These media that can be used according to the present invention may contain one or more sources of carbon, nitrogen, inorganic salts, vitamins and / or trace elements.
[0329] Preferred sources of carbon are sugars such as monosaccharides, disaccharides or polysaccharides. Very good sources of carbon are, for example, glucose, fructose, mannose, galactose, ribose, sorbose, ribulose, lactose, maltose, sucrose, raffinose, starch or cellulose. The sugars can also be added to the medium via complex compounds such as molasses, or other by-products from refined sugar. It may also be advantageous to add a mixture of various sources of carbon. Other possible sources of carbon are oils and fats such as soybean oil, sunflower oil, peanut oil and coconut oil, fatty acids such as palmitic acid, stearic acid or linoleic acid, alcohols such as glycerol, methanol or ethanol and organic acids such as acetic acid or lactic acid.
[0330] The source of nitrogen is usually an organic or inorganic nitrogen compound or a material containing these compounds. Examples of sources of nitrogen include ammonia gas, or ammonium salts such as ammonium sulfate, ammonium chloride, ammonium phosphate, ammonium carbonate or ammonium nitrate, nitrates, urea, amino acids, or complex sources of nitrogen such as corn steep liquor, soybean flour, soybean protein, yeast extract, meat extract and others. The sources of nitrogen can be used separately or as a mixture.
[0331] Inorganic salt compounds that may be present in the medium include chlorides, phosphates or sulfates of calcium, magnesium, sodium, cobalt, molybdenum, potassium, manganese, zinc, copper and iron.
[0332] In addition to inorganic sulfur-containing compounds such as sulfates, sulfites, dithionites, tetrathionates, thiosulfates, sulfides, organic sulfur compounds such as mercaptans and thiols can also be used as a source of sulfur.
[0333] Phosphoric acid, potassium dihydrogen phosphate or dipotassium hydrogen phosphate or the corresponding sodium-containing salts can be used as a source of phosphorus.
[0334] Chelating agents can be added to the medium to keep metal ions in solution. Particularly suitable chelating agents include dihydroxy phenols such as catechol or protocatechuate, or organic acids such as citric acid.
[0335] The fermentation medium used according to the present invention may also contain other growth factors such as vitamins, or growth promoting substances including, for example, biotin, riboflavin, thiamine, folic acid, nicotinic acid, pantothenate and pyridoxine. Growth factors and salts often originate from complex components of the medium such as yeast extract, molasses, corn steep liquor and the like. In addition, suitable precursors can be added to the culture medium. The precise composition of the compounds in the medium strongly depends on the particular experiment and must be determined individually for each specific case. Information on medium optimization can be found in the textbook "Applied Microbiol. Physiology, A Practical Approach" (1997). Growth media such as Standard 1 (Merck) or BHI (Brain Heart Infusion, DIFCO) can also be obtained from commercial vendors.
[0336] All components of the medium are sterilized by heating (20 minutes at 1.5 bar and 121 °C) or by sterile filtration. These components can be sterilized together or, if necessary, separately. All components of the medium may be present at the start of growth or may optionally be added continuously or by batch supply.
[0337] The temperature of the culture is usually between 15 °C and 45 °C, preferably between 25 °C and 40 °C, and can be kept constant or varied during the experiment. The pH value of the medium should be in the range of 5 to 8.5, preferably approximately 7.0. The pH value for growth can be controlled during growth by adding basic compounds such as sodium hydroxide, potassium hydroxide, ammonia or aqueous ammonia, or acidic compounds such as phosphoric acid or sulfuric acid. An antifoaming agent, such as a fatty acid polyglycol ester, can be used to suppress foaming. To maintain the stability of the plasmid, a suitable substance with a selective action, such as an antibiotic, can be added to the medium. To maintain aerobic conditions, oxygen or an oxygen-containing gas mixture, such as ambient air, is supplied into the culture. The temperature of the culture is usually 20 °C to 45 °C. The culture is continued until the desired product is produced maximally. This is usually achieved within 1 hour to 160 hours.
[0338] The method of the present invention can further include the step of recovering the terpene alcohol.
[0339] The term "recovering" includes extracting, collecting, isolating or purifying the compound from the culture medium. The recovery of the compound can be carried out according to any conventional isolation or purification method known in the art, including but not limited to treatment with conventional resins (e.g., anion or cation exchange resins, nonionic adsorption resins, etc.), treatment with conventional adsorbents (e.g., activated carbon, silicic acid, silica gel, cellulose, alumina, etc.), pH change, solvent extraction (e.g., with conventional solvents such as alcohol, ethyl acetate, hexane and the like), distillation, dialysis, filtration, concentration, crystallization, recrystallization, pH adjustment, freeze-drying and the like.
[0340] Prior to the desired isolation, the biomass of the broth can be removed. Processes for removing this biomass, such as filtration, sedimentation and flotation, are known to those skilled in the art. Thus, this biomass can be removed, for example, by a centrifuge, a separator, a decanter, a filter or in a flotation device. In order to maximize the recovery of valuable products, it is often advisable to wash the biomass, for example in the form of diafiltration. The choice of this method depends on the biomass content and characteristics in the fermenter broth, as well as the interaction between the biomass and the valuable product.
[0341] In one embodiment, the fermentation broth can be sterilized or pasteurized. In a further embodiment, the fermentation broth is concentrated. If necessary, this concentration can be carried out batchwise or continuously. The pressure and temperature ranges should be selected such that, firstly, no damage to the product occurs and, secondly, the use of the necessary equipment and energy is minimized. In particular, energy savings are possible by skillfully selecting the pressure and temperature levels of multi-stage evaporation.
[0342] The following examples are merely illustrative and do not limit the scope of the embodiments described herein.
[0343] Numerous possible variations that will be immediately apparent to those skilled in the art after considering the disclosure described herein are also within the scope of the invention.
[0344] Experimental Section The present invention will now be described in more detail by the following examples.
[0345] Materials: Unless otherwise stated, all chemical and biochemical materials and microorganisms or cells used herein are commercially available products.
[0346] Unless otherwise specified, recombinant proteins are cloned and expressed by standard methods such as those described in Sambrook, J., Fritsch, E.F. and Maniatis, T., Molecular cloning: A Laboratory Manual, 2 nd nd Edition, Cold Spring Harbor Laboratory, Cold Spring Harbor Laboratory Press, Cold Spring Harbor, NY, 1989.
[0347] General method: Standard assay for measuring copalyl diphosphate phosphatase activity E. coli cells (DP1205 strain) are transformed with two plasmids - a plasmid carrying a gene encoding an enzyme required for the biosynthesis of copalyl diphosphate (CPP), such as the pACYC-CrtE-SmCPS2 plasmid - a plasmid carrying a gene encoding a protein with terpenyl-phosphate phosphatase activity, such as the pJ401-TalVeTPP or pJ401-AspWeTPP plasmid as described below. These cells are cultured and the production of copalol is analyzed by GC-MS.
[0348] As described below, these cells are cultured and the production of copalol is analyzed by GC-MS.
[0349] Standard assay for measuring 8-hydroxy-copalyl diphosphate phosphatase activity E. coli cells (DP1205 strain) are transformed with two plasmids - a plasmid carrying a gene encoding an enzyme required for the biosynthesis of 8-hydroxy-copalyl diphosphate (LPP), such as the pACYC-CrtE-SsLPS plasmid - a plasmid carrying a gene encoding a protein with terpenyl-phosphate phosphatase activity, such as the pJ401-TalVeTPP or pJ401-AspWeTPP plasmid is transformed.
[0350] As described below, this cell is cultured and the production of lovendiol is analyzed by GC-MS.
[0351] Standard assay for measuring copalol dehydrogenase activity E. coli cells (DP1205 strain) are transformed with two plasmids - a plasmid carrying a gene encoding an enzyme necessary for the biosynthesis of copalol, such as the pJ401-CPOL-2 plasmid, - a plasmid carrying a gene encoding alcohol dehydrogenase, using, for example, pJ423 as a background plasmid is transformed.
[0352] As described below, this cell is cultured and the production of copalol is analyzed by GC-MS.
[0353] Standard assay for measuring lovendiol dehydrogenase activity E. coli cells (DP1205 strain) are transformed with two plasmids - a plasmid carrying a gene encoding an enzyme necessary for the biosynthesis of lovendiol, such as the pJ401-LOH-2 plasmid, - a plasmid carrying a gene encoding alcohol dehydrogenase, using, for example, pJ423 as a background plasmid is transformed.
[0354] As described below, this cell is cultured and the production of the product is analyzed by GC-MS.
[0355] Gas chromatography-mass spectrometry (GC-MS) The terpene content was analyzed by GC-MS using an Agilent 6890 series GC system connected to an Agilent 5975 mass detector. This GC was equipped with an HP-5MS capillary column (Agilent) with an inner diameter of 0.25 mm × 30 m. The carrier gas was helium at a constant flow rate of 1 mL / min. The inlet temperature was set at 250 °C. The initial oven temperature was 100 °C for 1 minute, followed by a gradient of 10 °C / min up to 300 °C. The identification of the products was based on the comparison of the mass spectra and retention indices with reference standards and an in-house mass spectral database. The concentration was estimated based on an internal standard.
[0356] Production of recombinant bacterial strains by chromosomal integration of genes encoding mevalonate pathway enzymes The E. coli strain was engineered to produce the terpene precursor farnesyl-pyrophosphate (FPP) by chromosomal integration of recombinant genes encoding mevalonate pathway enzymes. See also the construction scheme and recombination phenomenon shown in Figure 15.
[0357] An upstream pathway operon (operon 1 from acetyl-CoA to mevalonic acid) consisting of the atoB gene from E. coli encoding acetoacetyl-CoA thiolase and the mvaA and mvaS genes from Staphylococcus aureus encoding HMG-CoA synthase and HMG-CoA reductase, respectively, was designed.
[0358] As the downstream mevalonate pathway operon (operon 2 from mevalonic acid to farnesyl pyrophosphate), a native operon from the gram-negative bacterium Streptococcus pneumoniae encoding mevalonate kinase (mvaK1), phosphomevalonate kinase (mvaK2), phosphomevalonate decarboxylase (mvaD), and isopentenyl diphosphate isomerase (fni) was selected.
[0359] The gene encoding codon-optimized Saccharomyces cerevisiae FPP synthase (ERG20) was introduced at the 3'-end of the upstream pathway operon to convert isopentenyl diphosphate (IPP) and dimethylallyl diphosphate (DMAPP) to FPP.
[0360] The above operon was synthesized by DNA2.0 and integrated into the araA gene of Escherichia coli strain BL21(DE3). The heterologous pathway was introduced in two separate recombination steps using the CRISPR / Cas9 genome engineering system. The first operon to be integrated (downstream pathway; operon 2) carried a spectinomycin (Spec) marker used to screen for Spec-resistant candidate integrants. A second operon was designed to replace the Spec marker of the pre-integrated operon, and accordingly, Spec candidate integrants were screened according to the second recombination event (see Figure 15). A guide RNA expression vector targeting the araA gene was designed and synthesized by DNA 2.0. The operon integration was verified by designing PCR primers that amplify across the araA gene integration target and the recombination junctions of the integrants using PCR. One clone giving accurate PCR results was then fully sequenced and recorded as strain DP1205.
[0361] Culture of Bacterial Cells and Analysis of Terpene Production Transform E. coli cells with one or two expression plasmids carrying terpene biosynthesis genes, and culture the transformed cells on LB-agarose plates using appropriate antibiotics (kanamycin (50 μg / ml) and / or chloramphenicol (34 μg / ml)). Inoculate 5 mL of liquid LB medium supplemented with the same antibiotics, 4 g / l glucose, and 10% (v / v) dodecane with a single colony. The next day, inoculate 0.2 mL of the overnight culture into 2 mL of TB medium supplemented with the same antibiotics and 10% (v / v) dodecane. Incubate this culture at 37 °C until the optical density reaches 3. Induce the expression of the recombinant protein by adding 1 mM IPTG, and incubate this culture at 20 °C for 72 hours.
[0362] Then, extract this culture with tert.-butyl methyl ether (MTBE), and add an internal standard (α-longipinene (Aldrich)) to the organic phase. Analyze the terpene content of the organic phase by GC-MS as described above.
[0363] Example 1: Identification and characterization of the copalyl-diphosphate phosphatase activities of TalVeTPP and AspVeTPP.
[0364] The TalVeTPP and AspWeTPP proteins are encoded by two predicted genes in the genomes of Talaromyces verruculosus and Aspergillus wentii, respectively. The gene encoding TalVeTPP is located within the region 150095..151030 of the Talaromyces verruculosus genomic scaffold sequence with the NCBI accession number LHCL01000010.1. The encoded protein has been reported as a putative protein without functional characterization (NCBI accession number KUL89334.1). The gene encoding AspWeTPP is located within the region 2482776..2483627 of the Aspergillus wentii DTO 134E9 unplaced genomic scaffold ASPWEscaffold_5 (NCBI accession number KV878213.1). The encoded protein has the NCBI accession number OJJ34585.1 and has also been reported as a putative protein without functional characterization.
[0365] The genes encoding TalVeTPP and AspWeTPP are located within the genome adjacent to genes potentially involved in the biosynthesis of secondary metabolites, such as genes encoding oxidases, hydroxylases, dehydrogenases, and in particular genes having strong homology with monofunctional copalyl-diphosphate synthase or bifunctional copalyl-diphosphate synthase reported by Mitsuhashi et al, Chembiochem. 2017 Nov 2;18(21):2104-2109. Functional analysis of the TalVeTPP and AspWeTPP amino acid sequences by searching for the presence of protein family domain signatures (using, for example, the Interpro sequence analysis tool at www.ebi.ac.uk / interpro / or the Pfam database search tools at http: / / pfam.xfam.org / search#tabview=tab0 or https: / / www.ebi.ac.uk / Tools / pfa / pfamscan / ) revealed that the two proteins are predicted to contain a protein tyrosine phosphatase signature. Enzymes from the tyrosine phosphatase family are described as removing phosphate groups from various phosphorylated molecules, particularly proteins. However, it has not been shown that enzymes from this protein family act on compounds such as copalyl-diphosphate. Considering the genomic localization of the genes encoding TalVeTPP and AspWeTPP, the inventors hypothesized that TalVeTPP and AspWeTPP can catalyze the cleavage of the diphosphate group of copalyl-diphosphate or other isoprenoid-diphosphate compounds (Figure 1).
[0366] The cDNAs encoding TalVeTPP and AspWeTPP (SEQ ID NOs: 3 and 7, respectively) were codon-optimized (SEQ ID NOs: 1 and 5, respectively) and individually cloned into the expression plasmid pJ401 (ATUM, Newark, California) to obtain plasmids pJ401-TalVeTPP and pJ401-AspWeTPP.
[0367] Another expression plasmid carrying the gene encoding geranylgeranyl - pyrophosphate synthase (GGPS) and the gene encoding copalyl - pyrophosphate synthase (CPS) was constructed. In the case of the CPS gene, the cDNA encoding CPS (NCBI accession number ABV57835.1) from Salvia miltiorrhiza was codon - optimized for optimal expression in E. coli cells. Additionally, the first 58 codons were removed and an ATG start codon was added. The optimized cDNA (SEQ ID NO: 33) encoding truncated Salvia miltiorrhiza CPS (SmCPS2) was synthesized in vitro and first cloned into the pJ208 plasmid (ATUM, Newark, California) adjacent to the NdeI and KpnI restriction enzyme recognition sites. In the case of GGPS, the CrtE gene (NCBI accession M38424.1) from Pantoea agglomerans encoding GGPP synthase (NCBI accession number AAA24819.1) was used. The CrtE gene with codon optimization (SEQ ID NO: 35) and the addition of NcoI and BamHI restriction enzyme recognition sites at the 3' and 5' ends (ATUM, Newark, California) was synthesized and ligated between the NcoI site and the BamHI site of the pACYCDuet™-1 plasmid (Merck) to obtain the pACYC - CrtE plasmid. The cDNA encoding the modified SmCPS2 was digested with NdeI and KpnI and ligated into the pACYC - CrtE plasmid, thus obtaining the pACYC - CrtE - SmCPS2 construct.
[0368] Two plasmids, pACYC-CrtE-SmCPS2 plasmid and pJ401-TalVeTPP or pJ401-AspWeTPP were used to transform E. coli cells (DP1205 strain produced above). These cells were cultured as described in the Methods section and the production of terpene compounds was analyzed. Figure 2 shows a typical GC-MS of copalol produced by recombinant E. coli cells. Cells expressing only mevalonate pathway enzymes and SmCPS2 produced a small amount of copalol (6.7 mg / l, Figure 3) due to the hydrolysis of CPP by the endogenous alkaline phosphatase enzyme. In addition, E. coli cells transformed to express TalVeTPP or AspWeTPP produced significantly higher amounts of copalol: 462 mg / l and 298 mg / l, respectively, for TalVeTPP and AspWeTPP (Figures 2 and 3). This experiment shows that TalVeTPP and AspWeTPP can efficiently hydrolyze (+)-CPP to produce (+)-copalol. This copalol is produced with high purity (over 95%). The smaller amount of coparyl acetate (Figure 2) observed in the GC-MS analysis is due to the endogenous acetyltransferase activity of the cells.
[0369] Example 2: Identification and characterization of variants of TalVeTPP and AspVeTPP with coparyl-pyrophosphate phosphatase activity.
[0370] The TalVeTPP and AspVeTPP arrays were used to search for homologous sequences in the public database. The following eight new sequences with the signature of the Pfam protein tyrosine phosphatase protein family PF13350 were selected: HelGriTPP1, hypothetical protein Helicocarpus griseus (SEQ ID NO: 10) (GenBank: PGG95910.1); UmbPiTPP1, tyrosine phosphatase from Umbilicaria pustulata (SEQ ID NO: 13) (GenBank: SLM34787.1); TAlVeTPP2, hypothetical protein from Talaromyces verruculosus (SEQ ID NO: 16) (GenBank: KUL92314.1); HydPiTPP1, hypothetical protein from Hydnomerulius pinastri (SEQ ID NO: 19) (GenBank: KIJ69780.1); TalCeTPP1, hypothetical protein from Talaromyces cellulolyticus (SEQ ID NO: 22) (GenBank: GAM42000.1); TalMaTPP1, hypothetical protein from Talaromyces marneffei (SEQ ID NO: 25) (NCBI XP_002152917.1); TalAstroTPP1, hypothetical protein from Talaromyces atroroseus (SEQ ID NO: 28) (NCBI XP_020117849.1); PeSubTPP1, hypothetical protein from Penicillium subrubescens (SEQ ID NO: 31) (GenBank: OKP14340.1). The search for the protein family signature indicated that the eight amino acid sequences are members of the Pfam protein tyrosine phosphatase protein family PF13350.
[0371] Sequence comparison of the ten amino acid sequences showed sequence identity in the range of 24% to 93% (Table 1).
[0372]
Table 2
[0373] The cDNA sequences encoding HelGriTPP1 (SEQ ID NO: 11), UmbPiTPP1 (SEQ ID NO: 14), TalVeTPP2 (SEQ ID NO: 17), HydPiTPP1 (SEQ ID NO: 20), TalCeTPP1 (SEQ ID NO: 23), TalMaTPP1 (SEQ ID NO: 26), TalAstroTPP1 (SEQ ID NO: 29) and PeSubTPP1 (SEQ ID NO: 32) were codon-optimized for E. coli expression and individually cloned into the pJ401 expression plasmid (ATUM, Newark, California).
[0374] DP1205 E. coli cells were transformed with one of the pACYC-CrtE-SmCPS2 plasmid and the pJ401 plasmid carrying the optimized cDNA encoding HelGriTPP1 (SEQ ID NO: 9), UmbPiTPP1 (SEQ ID NO: 12), TalVeTPP2 (SEQ ID NO: 15), HydPiTPP1 (SEQ ID NO: 18), TalCeTPP1 (SEQ ID NO: 21), TalMaTPP1 (SEQ ID NO: 24), TalAstroTPP1 (SEQ ID NO: 27) and PeSubTPP1 (SEQ ID NO: 30). These cells were cultured under the conditions described in the Methods section and the production of copalol was analyzed. Cells transformed with the pACYC-CrtE-SaLPS plasmid and the empty pJ401 plasmid were used as control strains. All strains expressing recombinant TalVeTPP, AspWeTPP, HelGriTPP1, UmbPiTPP1, TalVeTPP2, HydPiTPP1, TalCeTPP1, TalMaTPP1, TalAstroTPP1 or PeSubTPP1 protein accumulated copalol in amounts ranging from 32 to 240 mg / l, and enzymatic conversion of CPP to copalol by all these recombinant enzymes was confirmed (Figure 4).
[0375] This example demonstrates that CPP can be enzymatically converted to copalol using TalVeTPP, AspWeTPP, HelGriTPP1, UmbPiTPP1, TalVeTPP2, HydPiTPP1, TalCeTPP1, TalMaTPP1, TalAstroTPP1, and PeSubTPP1, and that copalol can be produced in engineered cells.
[0376] Example 3: Production of labdenediol in E. coli cells.
[0377] An expression plasmid carrying the gene encoding geranylgeranyl - pyrophosphate synthase (GGPS) and the gene encoding labdenediol - phyllol pyrophosphate synthase (LPS) was constructed. For GGPS, the CrtE gene derived from P. agglomerans described in Example 1 was used. For the LPS gene, a cDNA encoding SsLPS derived from Salvia sclarea (International Publication No. WO2009095366, GenBank: AET21246.1) was used. The cDNA sequence encoding SsLPS was optimized as described in International Publication No. WO2009095366 (SEQ ID NO: 37) and cloned between the NdeI site and the KpnI site in the pACYC - Crte plasmid to obtain the plasmid pACYC - CrtE - SaLPS carrying the GGP synthase gene and the LPP synthase gene. E. coli cells such as the DP1205 strain transformed with pACYC - CrtE - SsLPS accumulate LPP as a diterpene precursor compound (Figure 5).
[0378] One of the pACYC-CrtE-SsLPS plasmid and the pJ401 plasmids carrying optimized cDNAs encoding TalVeTPP, AspWeTPP, HelGriTPP1, UmbPiTPP1, TalVeTPP2, HydPiTPP1, TalCeTPP1, TalMaTPP1, TalAstroTPP1 or PeSubTPP1 (manufactured above) was used to transform DP1205 E. coli cells (see Examples 1 and 2 above). These cells were cultured under the conditions described in the Methods section and the production of lavendiol was analyzed. Compared to the control cells transformed with the empty pJ401 plasmid and pACYC-CrtE-SsLPS, all cells transformed to produce any of the recombinant TalVeTPP, AspWeTPP, HelGriTPP1, UmbPiTPP1, TalVeTPP2, HydPiTPP1, TalCeTPP1, TalMaTPP1, TalAstroTPP1 or PeSubTPP1 proteins produced significantly increased amounts of lavendiol (5 - 25-fold increase) (Figure 6). The lavendiol concentration in the cell culture was between 50 - 272 g / l at the end of the culture period.
[0379] Figure 7 shows the GC-MS analysis of a typical E. coli producing lavendiol cells. The total ion chromatogram indicates that the lavendiol produced under these conditions has a purity of at least 98%.
[0380] Example 4: Production of farnesol and geranylgeraniol in E. coli cells.
[0381] The recombinant proteins TalVeTPP, AspWeTPP, HelGriTPP1, UmbPiTPP1, TalVeTPP2, HydPiTPP1, TalCeTPP1, TalMaTPP1, TalAstroTPP1 and PeSubTPP1 were also evaluated for their enzymatic activities on linear substrates such as farnesyl - pyrophosphate (FPP) and geranylgeranyl - pyrophosphate (GGPP). Assays were performed under the conditions in the method section similar to those in Examples 1 and 2, except that the pACYCDuet™-1 plasmid was adapted to produce in - vivo FPP and GGPP. For E. coli cells accumulating FPP, the empty pACYCDuet™-1 plasmid was used. E. coli cells such as the DP1205 strain transformed with the empty plasmid pACYCDuet™-1 (Merck) accumulate FPP as a terpene precursor compound (Figure 8). For E. coli cells accumulating GGPP, the plasmid pACYC - CrtE (Example 1) was used. E. coli cells such as the DP1205 strain transformed with the pACYC - CrtE plasmid accumulate GGPP as a terpene precursor compound (Figures 5 and 8).
[0382] DP1205 E. coli cells were transformed with one of the pACYC-CrtE or pACYCDuet™-1 plasmids and the pJ401 plasmid carrying the optimized cDNA encoding TalVeTPP, AspWeTPP, HelGriTPP1, UmbPiTPP1, TalVeTPP2, HydPiTPP1, TalCeTPP1, TalMaTPP1, TalAstroTPP1 or PeSubTPP1 (see Examples 1 and 2). The cells were cultured under the conditions described in the Methods section and the production of farnesol and geranylgeraniol was analyzed. Some of the cells transformed to produce the recombinant TalVeTPP, AspWeTPP, HelGriTPP1, UmbPiTPP1, TalVeTPP2, HydPiTPP1, TalCeTPP1, TalMaTPP1, TalAstroTPP1 or PeSubTPP1 proteins produced significantly increased amounts of farnesol and geranylgeraniol compared to the control cells transformed with the empty pJ401 (Figure 9). For example, cells expressing TalVeTPP2, TalCeTPP1 and TalVeTPP produced large amounts (763 - 976 mg / ml) of farnesol. Similarly, cells expressing PeSubTPP1 and TalVeTPP produced large amounts (196 - 198 mg / ml) of geranylgeraniol. In contrast, HelGriTPP1, UmbPiTPP1, HydPiTPP1 or TalMaTPP1 showed low FPP and GGPP phosphatase activities.
[0383] Example 5: Substrate selectivity of phosphatases.
[0384] Comparison of the enzyme activities observed in the previous examples with four different substrates reveals different substrate selectivities for TalVeTPP, AspWeTPP, HelGriTPP1, UmbPiTPP1, TalVeTPP2, HydPiTPP1, TalCeTPP1, TalMaTPP1, TalAstroTPP1 or PeSubTPP1 (Figure 10). Thus, this approach enables the selection of phosphatases with phosphatase activity based on their substrate selectivity. For example, HelGriTPP1, HydPiTPP1 or AspWeTPP show relatively higher activity towards CPP and LPP and lower activity towards FPP and GGP compared to the other enzymes described. Thus, these enzymes can be used to most effectively produce copalol or labdenediol with limited side activity towards pathway intermediates. UmbPiTPP1, TalMaTPP1 and TalAstroTPP1 also show limited side activity towards FPP and GGPP, but they produce smaller amounts of labdenediol and copalol.
[0385] Example 6: Production of copalol and labdenediol using an operon containing GGPP synthase, diterpene synthase and phosphatase.
[0386] An operon containing TalVeTPP; TaTps1-del59 and three cDNAs encoding GGPP synthase was constructed. TaTps1-del59 is an N-terminally truncated CPP synthase from Triticum aestivum (NCBI accession number BAH56559.1). The cDNA encoding TaTps1-del59 was codon-optimized (SEQ ID NO: 39). For GGPP synthase, a codon-optimized form of the CrtE gene from Pantoea agglomerans (NCBI accession M38424.1) was used (SEQ ID NO: 35). The operon was cloned in the pJ401 expression plasmid (ATUM, Newark, California) to obtain the construct pJ401-CPOL-2.
[0387] Another operon was constructed with an organization similar to the above-mentioned CPOL-2, except that the gene encoding TaTps1-del59 was replaced with the optimized gene encoding SaLPS (SEQ ID NO: 37). This operon was cloned into plasmid pJ401 (ATUM, Newark, California) to obtain construct pJ401-LOH-2. The DP1205 E. coli cells produced above were transformed with plasmid pJ401-CPOL-2 or pJ401-LOH-2. These cells were cultured as described and the production of diterpenes was analyzed as described in the Methods section. In parallel, cells transformed with the empty PJ401 plasmid and pACYC-CrtE-SsLPS or pACYC-CrtE-SmCPS plasmids were used as controls.
[0388] Cells transformed with plasmid CPOL-2 produced copalol and farnesol at average concentrations of 200 mg / l and 300 mg / l, respectively, for copalol and farnesol. Cells transformed with plasmid LOH-2 produced labdenediol and farnesol at average concentrations of 1260 mg / l and 830 mg / l, respectively, for copalol and farnesol (Figure 11). The significant amount of farnesol produced using these two constructs is due to the incomplete conversion of FPP pool to GGPP and the enzyme activity of TalVeTPP towards FPP in addition to CPP (see Figure 10). Corresponding experiments, such as those with HelGriTPP1 (and others shown in Figure 10 with higher specificity), would produce less farnesol.
[0389] Example 7: Enzymatic oxidation of terpene compounds produced by recombinant phosphatase to produce the corresponding α,β-unsaturated aldehyde.
[0390] Using the following alcohol dehydrogenases (ADHs), the terpene compounds produced by the phosphatase described in the previous examples can be oxidized: - CymB (SEQ ID NO: 42) (GenBank accession AEO27362.1) from Pseudomonas sp. 19-rlim strain; - AspWeADH1 (SEQ ID NO: 44) (GenBank accession OJJ34588.1) encoded by a gene located within the region 2487333..2488627 of the Aspergillus wentii DTO 134E9 non-anchored genomic scaffold ASPWEscaffold_5 (NCBI accession number KV878213.1); - PsAeroADH1 (SEQ ID NO: 46) (GenBank accession WP_079868259.1) from Pseudomonas aeruginosa; - AzTolADH1 (SEQ ID NO: 48) (GenBank accession WP_018990713.1) from Azoarcus toluclasticus; - AroAroADH1 (SEQ ID NO: 50) (GenBank accession KM105875.2) from Aromatoleum aromaticum.
[0391] - ThTerpADH1 (SEQ ID NO: 52) (Genbank accession WP_021250577.1) from Thauera terpenica.
[0392] - CdGeoA (SEQ ID NO: 54) (NCBI accession WP_043683915.1) from Castellaniella defragrans.
[0393] -VoADH1 (SEQ ID NO: 56) (GenBank accession AVX32614.1) derived from Valeriana officinalis.
[0394] Codon-optimized cDNAs encoding each of the above ADHs were synthesized (see SEQ ID NOs: 41, 43, 45, 47, 49, 51, 53, and 55, respectively) and cloned in the pJ423 expression plasmid (ATUM, Newark, California).
[0395] The DP1205 E. coli cells prepared above were transformed with plasmid pJ401-CPOL-2 and one of these pJ423-ADH plasmids. The cells were cultured as described and the production of diterpenes was analyzed as described in the Methods section. In parallel, cells transformed with the pJ401-CPOL-2 plasmid and the empty pJ423 plasmid were used as controls (Figure 12). The production of copalol was observed for all cells, indicating that copalol can be efficiently produced using a combination of an enzyme of the copalol biosynthetic pathway, including protein tyrosine phosphatase, and an ADH selected from the above ADHs. With the exception of AspWeADH1 and VoADH1, the conversion of intracellular copalol to copalol was at least 90%. A mixture of the cis and trans isomers of copalol was observed due to the non-enzymatic isomerization of trans-copalol produced by the ADH. Using the same ADH, the conversion of farnesol to farnesal was also observed (Figure 13).
[0396] Production of two oxidation products of labdenediol was observed for E. coli cells co-transformed with one of the pJ401-LOH-2 and pJ423-ADH plasmids (Figure 14). By NMR analysis, it was confirmed that the two compounds are two isomers ((13R) and (13S)) of 8,13-epoxy-labdane-15-al shown in the scheme of Figure 17. These two compounds are caused by the instability of 8-hydroxy-labda-13-en-15-al, an α,β-unsaturated aldehyde produced by the oxidation of labdenediol. The proposed mechanism of dehydration and rearrangement from the aldehyde to the isomers is shown in the following scheme.
[0397] Example 8: Engineering of recombinant bacterial cells for the production of copalol using multifunctional CPP synthase An operon was constructed that contains two cDNAs encoding: - AspWeTPP from Aspergillus wentii (SEQ ID NO: 6), - PvCPS, a protein from Talaromyces verruculosus having prenyl-transferase activity and copalyl-diphosphate synthase activity (SEQ ID NO: 59) (GenBank accession BBF88128.1). PvCPS catalyzes the production of copalyl PP from IPP and DMAPP.
[0398] The cDNAs encoding AspWeTPP and PvCPS were codon-optimized (SEQ ID NOs: 8 and 60). An operon containing the two cDNAs and an RBS sequence (AAGGAGGTAAAAAA) (including SEQ ID NO: 61) placed upstream of each cDNA was designed. The operon was synthesized and cloned into the pJ401 expression plasmid (ATUM, Newark, California) to obtain plasmid pJ401-CPOL-4.
[0399] DP1205 E. coli cells were transformed with plasmid pJ401-CPOL-4. The transformed cells were cultured as described and diterpene production was analyzed as described in the Methods section. Under these conditions, cells transformed with plasmid pJ401-CPOL-4 produced copalol as the main product at much higher concentrations (up to 1 mg / l) than cells transformed with plasmid pJ401-CPOL-2. This experiment shows that higher concentrations of copalol can be obtained using a multifunctional protein with prenyltransferase and CPP synthase activities compared to multiple monofunctional proteins.
[0400] Example 9: In vivo production of copalol and copalal in Saccharomyces cerevisiae cells using copalyl - pyrophosphate phosphatase and various alcohol dehydrogenases.
[0401] To produce copalol and copalal, genes encoding GGPP synthase CrtE (SEQ ID NO: 61) (from Pantoea agglomerans, NCBI accession M38424.1), copalyl - pyrophosphate synthase SmCPS2 (SEQ ID NO: 63) (from Salvia miltiorrhiza, NCBI accession ABV57835.1), copalyl - pyrophosphate phosphatase TalVeTPP (SEQ ID NO: 65) and various alcohol dehydrogenases (cDNAs optimized for expression in yeast) were expressed in engineered Saccharomyces cerevisiae cells with increased levels of endogenous farnesyl - diphosphate (FPP).
[0402] Four alcohol dehydrogenases were evaluated.
[0403] -AzTolADH1 (SEQ ID NO: 48) (yeast - optimized cDNA SEQ ID NO: 66) - PsAeroADH1 (SEQ ID NO: 46), (yeast-optimized cDNA SEQ ID NO: 67) - SCH23-ADH1 from Hyphozyma roseonigra (SEQ ID NO: 69) (yeast-optimized cDNA SEQ ID NO: 68) - SCH24-ADH1a from Cryptococcus albidus (SEQ ID NO: 71) (yeast-optimized cDNA SEQ ID NO: 70) To increase the levels of the endogenous farnesyl-diphosphate (FPP) pool in S. cerevisiae cells, additional copies of all yeast endogenous genes involved in the mevalonate pathway, from ERG10 encoding acetyl-CoA C-acetyltransferase to ERG20 encoding FPP synthase, were integrated into the genome of S. cerevisiae strain CEN.PK2-1C (Euroscarf, Frankfurt, Germany) under the control of a galactose-inducible promoter as described in Paddon et al., Nature, 2013, 496:528-532. Briefly, three cassettes were integrated into the LEU2, TRP1, and URA3 loci respectively. A first cassette containing the GAL10 / GAL1 bidirectional promoter and the genes ERG19 and ERG13, and also the genes ERG20 and truncated HMG1 (tHMG1 as described in Donald et al., Proc Natl Acad Sci USA, 1997, 109:E111-8) under the control of the GAL10 / GAL1 promoter, this cassette was flanked by two 100-nucleotide regions corresponding to the upstream and downstream sites of LEU2. A second cassette with the genes IDI1 and tHMG1 under the control of the GAL10 / GAL1 promoter and the gene ERG13 under the control of the GAL7 promoter region, this cassette was flanked by two 100-nucleotide regions corresponding to the upstream and downstream sites of TRP1. A third cassette containing the genes ERG10, ERG12, tHMG1, and ERG8 all under the control of the GAL10 / GAL1 promoter, this cassette was flanked by two 100-nucleotide regions corresponding to the upstream and downstream sites of URA3. All genes within the three cassettes contained their own 200-nucleotide terminator regions. Also, an additional copy of GAL4 under the control of a mutant form of its own promoter was integrated upstream of the ERG9 promoter region as described in Griggs and Johnston, Proc Natl Acad Sci USA, 1991, 88:8597-8601. In addition, the expression of ERG9 was modified by promoter exchange.The GAL7, GAL10, and GAL1 genes were deleted using a cassette containing the HIS3 gene with its own promoter and terminator. The resulting strain was mated with the CEN.PK2-1D strain (Euroscarf, Frankfurt, Germany) to obtain a diploid strain named YST045, which was induced to sporulate according to Solis-Escalante et al, FEMS Yeast Res, 2015, 15:2. Spore separation was achieved by resuspending asci in 200 μL of 0.5 M sorbitol containing 2 μL of zymolyase (1000 U mL-1, Zymo research, Irvine, CA) and incubating at 37 °C for 20 min. The mixture was then plated on a medium containing 20 g / L peptone, 10 g / L yeast extract, and 20 g / L agar, and one germinated spore was isolated and named YST075.
[0404] To express different genes encoding alcohol dehydrogenases, genomic integration into the YST075 strain was performed. Each integration cassette was composed of four fragments.
[0405] 1) A fragment containing 261 bp corresponding to the upstream region of the BUD9 gene and the sequence 5'-GCACTTGCTACACTGTCAGGATAGCTTCCGTCACATGGTGGCGATCACCGTACATCTGAG-3' (SEQ ID NO: 72), which was obtained by PCR using genomic DNA from the YST075 strain as a template; 2) The fragment containing the sequence 5'-GCACTTGCTACACTGTCAGGATAGCTTCCGTCACATGGTGGCGATCACCGTACATCTGAG-3' (SEQ ID NO: 72), the promoter region of the GAL1 gene, one of the genes encoding alcohol dehydrogenase codon-optimized for expression in S. cerevisiae, the terminator region of the PGK1 gene, and the sequence 5'-AGGTGCAGTTCGCGTGCAATTATAACGTCGTGGCAACTGTTATCAGTCGTACCGCGCCAT-3' (SEQ ID NO: 73), which was obtained by DNA synthesis (ATUM, Menlo Park, CA 94025), 3) The fragment containing the sequence 5'-AGGTGCAGTTCGCGTGCAATTATAACGTCGTGGCAACTGTTATCAGTCGTACCGCGCCAT-3' (SEQ ID NO: 73), the TRP1 gene containing its own promoter and terminator regions, and the sequence 5'-TGGTCAGCAACAACGCCGAAGAATCACTCTCGTGTTGAGAATTGCACGCCTTGACCACGA-3' (SEQ ID NO: 74), which was obtained by PCR using pESC-TRP1 (Agilent Technologies, California, USA) as a template; and 4) The fragment containing the sequence 5'-TGGTCAGCAACAACGCCGAAGAATCACTCTCGTGTTGAGAATTGCACGCCTTGACCACGA-3' (SEQ ID NO: 74) and 344 bp corresponding to the BUD9 gene, which was obtained by PCR using genomic DNA from the YST075 strain as a template.
[0406] YST075 was transformed with the four fragments necessary for genomic integration of each of the evaluated alcohol dehydrogenases. Yeast transformation was carried out by the lithium acetate protocol described in Gietz and Woods, Methods Enzymol., 2002, 350:87-96. The transformation mixture was plated on SmTrp- medium containing 6.7 g / L yeast nitrogen base without amino acids (BD Difco, New Jersey, USA), 1.92 g / L dropout supplement without leucine (Sigma Aldrich, Missouri, USA), 20 g / L glucose and 20 g / L agar. The plates were incubated at 30 °C for 3 - 4 days. Single colonies containing the correct integration were isolated and named YST149 (containing SCH23-ADH1), YST150 (containing SCH24-ADH1a), YST151 (containing AzTolADH1) and YST152 (containing PsAeroADH1).
[0407] To express CrtE, SmCPS2 and TalVeTPP in YST149, YST150, YST151 and YST152, plasmids were constructed in vivo using endogenous homologous recombination of yeast as previously described in Kuijpers et al., Microb Cell Fact., 2013, 12:47. The plasmids were composed of six DNA fragments used for co-transformation of S. cerevisiae. These fragments were as follows: a) The LEU2 yeast marker constructed by PCR using the primers 5’AGGTGCAGTTCGCGTGCAATTATAACGTCGTGGCAACTGTTATCAGTCGTACCGCGCCATTCGACTACGTCGTAAGGCC-3’ (SEQ ID NO:75) and 5’TCGTGGTCAAGGCGTGCAATTCTCAACACGAGAGTGATTCTTCGGCGTTGTTGCTGACCATCGACGGTCGAGGAGAACTT-3’ (SEQ ID NO:76) together with plasmid pESC-LEU (Agilent Technologies, California, USA) as template; b) AmpR E. coli marker constructed by PCR using primers 5’-TGGTCAGCAACAACGCCGAAGAATCACTCTCGTGTTGAGAATTGCACGCCTTGACCACGACACGTTAAGGGATTTTGGTCATGAG-3’ (SEQ ID NO: 77) and 5’-AACGCGTACCCTAAGTACGGCACCACAGTGACTATGCAGTCCGCACTTTGCCAATGCCAAAAATGTGCGCGGAACCCCTA-3’ (SEQ ID NO: 78) together with plasmid pESC-URA as a template; c) Yeast replication origin obtained by PCR using primers 5’-TTGGCATTGGCAAAGTGCGGACTGCATAGTCACTGTGGTGCCGTACTTAGGGTACGCGTTCCTGAACGAAGCATCTGTGCTTCA-3’ (SEQ ID NO: 79) and 5’-CCGAGATGCCAAAGGATAGGTGCTATGTTGATGACTACGACACAGAACTGCGGGTGACATAATGATAGCATTGAAGGATGAGACT-3’ (SEQ ID NO: 80) together with pESC-URA as a template; d) E. coli replication origin obtained by PCR using primers 5’-ATGTCACCCGCAGTTCTGTGTCGTAGTCATCAACATAGCACCTATCCTTTGGCATCTCGGTGAGCAAAAGGCCAGCAAAAGG-3’ (SEQ ID NO: 81) and 5’-CTCAGATGTACGGTGATCGCCACCATGTGACGGAAGCTATCCTGACAGTGTAGCAAGTGCTGAGCGTCAGACCCCGTAGAA-3’ (SEQ ID NO: 82) together with plasmid pESC-URA as a template; e) A fragment composed of the last 60 nucleotides of fragment “d”, 200 nucleotides downstream of the stop codon of the yeast gene PGK1, the codon-optimized GGPP synthase coding sequence CrtE (SEQ ID NO: 62) for its expression in S. cerevisiae, the GAL10 / GAL1 bidirectional yeast promoter, the coding sequence of TalVeTPP codon-optimized for its expression in S. cerevisiae (SEQ ID NO: 65), 200 nucleotides downstream of the stop codon of the yeast gene CYC1, and the sequence 5'-ATTCCTAGTGACGGCCTTGGGAACTCGATACACGATGTTCAGTAGACCGCTCACACATGG-3' (SEQ ID NO: 83). This fragment was obtained by DNA synthesis (ATUM, Menlo Park, CA 94025) and f) A fragment composed of the last 60 nucleotides of fragment “e”, 200 nucleotides downstream of the stop codon of the yeast gene CYC1, the codon-optimized SmCPS2 synthase coding sequence (SEQ ID NO: 63) for its expression in S. cerevisiae, the GAL10 / GAL1 bidirectional yeast promoter, and the 60 nucleotides corresponding to the beginning of fragment “a”. This fragment was obtained by DNA synthesis (ATUM, Menlo Park, CA 94025).
[0408] All strains were transformed with the fragments necessary for in vivo plasmid assembly. Yeast transformation was performed by the lithium acetate protocol described in Gietz and Woods, Methods Enzymol., 2002, 350:87-96. The transformation mixture was plated on SmLeu- medium containing 6.7 g / L yeast nitrogen base without amino acids (BD Difco, New Jersey, USA), 1.6 g / L dropout supplement without leucine (Sigma Aldrich, Missouri, USA), 20 g / L glucose, and 20 g / L agar. The plates were incubated at 30 °C for 3 - 4 days. Individual colonies were used to produce copalol and copalal in 2 mL of the medium described in Westfall et al., Proc Natl Acad Sci USA, 2012, 109:E111-118 and in glass tubes containing dodecane as an organic overlay.
[0409] Under these culture conditions, the highest average concentration of copalol was 153.51 mg / L produced by strain YST152 containing the copalol biosynthetic plasmid. The highest average concentration of copalal was 98.47 mg / L produced by strain YST149 containing the copalol biosynthetic plasmid. The average percentage conversion of copalol to copalal in strains YST149, YST150, YST151, and YST152 containing the copalol biosynthetic plasmid was 61.6%, 39.9%, 30.1%, and 22.1%, respectively. The production of copalol and copalal was identified and quantified using GC-MS analysis (Figure 18) with an internal standard.
[0410]
Table 3-1
[0411]
Table 3-2
[0412] TalVeTPP optimized cDNA ORF only - SEQ ID NO: 1 [ka]
[0413] TalVeTPP amino acid sequence—SEQ ID NO: 2 [ka]
[0414] TalVeTPP wild-type cDNA - SEQ ID NO: 3 [ka]
[0415] TalVeTPP optimized cDNA containing non-coding sequences - SEQ ID NO: 4 [ka]
[0416] AspWeTPP optimized cDNA - SEQ ID NO:5 [ka]
[0417] AspWeTPP amino acid sequence - SEQ ID NO: 6 [ka]
[0418] AspWeTPP wild-type cDNA - SEQ ID NO: 7 [ka]
[0419] AspWeTPP optimized cDNA including non-coding ends - SEQ ID NO:8 [ka]
[0420] Optimized cDNA of HelGriTPP1 - SEQ ID NO: 9
Chem.
[0421] Amino acid sequence of HelGriTPP1 - SEQ ID NO: 10
Chem.
[0422] Wild - type cDNA of HelGriTPP1 - SEQ ID NO: 11
Chem.
[0423] Optimized cDNA of UmbPiTPP1 - SEQ ID NO: 12
Chem.
[0424] Amino acid sequence of UmbPiTPP1 - SEQ ID NO: 13
Chem.
[0425] Wild - type cDNA of UmbPiTPP1 - SEQ ID NO: 14
Chem.
[0426] Optimized cDNA of TalVeTPP2 - SEQ ID NO: 15
Chem.
[0427] Amino acid sequence of TalVeTPP2 - SEQ ID NO: 16
Chem.
[0428] TalVeTPP2 wild-type cDNA - SEQ ID NO: 17
Chem.
[0429] HydPiTPP1 optimized cDNA - SEQ ID NO: 18
Chem.
[0430] HydPiTPP1 amino acid sequence - SEQ ID NO: 19
Chem.
[0431] HydPiTPP1 wild-type cDNA - SEQ ID NO: 20
Chem.
[0432] TalCeTPP1 optimized cDNA - SEQ ID NO: 21
Chem.
[0433] TalCeTPP1 amino acid sequence - SEQ ID NO: 22
Chem.
[0434] TalCeTPP1 wild-type cDNA - SEQ ID NO: 23
Chem.
[0435] Optimized cDNA of TalMaTPP1 - SEQ ID NO: 24 [Chem.]
[0436] Amino acid sequence of TalMaTPP1 - SEQ ID NO: 25 [Chem.]
[0437] Wild - type cDNA of TalMaTPP1 - SEQ ID NO: 26 [Chem.]
[0438] Optimized cDNA of TalAstroTPP1 - SEQ ID NO: 27 [Chem.]
[0439] Amino acid sequence of TalAstroTPP1 - SEQ ID NO: 28 [Chem.]
[0440] Wild - type cDNA of TalAstroTPP1 - SEQ ID NO: 29 [Chem.]
[0441] Optimized cDNA of PeSubTPP1 - SEQ ID NO: 30 [Chem.]
[0442] Amino acid sequence of PeSubTPP1 - SEQ ID NO: 31 [Chem.]
[0443] PeSubTPP1 wild-type cDNA - SEQ ID NO: 32
Chem.
[0444] Codon-optimized cDNA encoding SmCPS - SEQ ID NO: 33
Chem.
[0445] SmCPS, CPP synthase from Salvia miltiorrhiza, amino acid sequence. - SEQ ID NO: 34
Chem.
[0446] Codon-optimized cDNA encoding GGPP synthase from Pantoea agglomerans. - SEQ ID NO: 35
Chem.
[0447] GGPP synthase from Pantoea agglomerans, amino acid sequence - SEQ ID NO: 36
Chem.
[0448] Optimized cDNA encoding SsLPS - SEQ ID NO: 37
[0449]
Chem.
Chem.
[0450] Optimized cDNA encoding TaTps1-del59 - SEQ ID NO: 39
Chemical formula
[0451] TaTps1-del59, a truncated copalyl diphosphate synthase from Triticum aestivum. - SEQ ID NO: 40
Chemical formula
[0452] CymB, optimized cDNA - SEQ ID NO: 41
Chemical formula
[0453] CymB, amino acid sequence - SEQ ID NO: 42
Chemical formula
[0454] AspWeADH1, optimized cDNA - SEQ ID NO: 43
Chemical formula
[0455] AspWeADH1, amino acid sequence - SEQ ID NO: 44
Chemical formula
[0456] PsAerADH1, optimized cDNA - SEQ ID NO: 45
Chemical formula
[0457] PsAerADH1, Amino Acid Sequence - SEQ ID NO: 46
Chem.
[0458] AzTolADH1, Optimized cDNA - SEQ ID NO: 47
Chem.
[0459] AzTolADH1, Amino Acid Sequence - SEQ ID NO: 48
Chem.
[0460] AroAroADH1, Optimized cDNA - SEQ ID NO: 49
Chem.
[0461] AroAroADH1, Amino Acid Sequence - SEQ ID NO: 50
Chem.
[0462] ThTerpADH1, Optimized cDNA - SEQ ID NO: 51
Chem.
[0463] ThTerpADH1, Amino Acid Sequence - SEQ ID NO: 52
Chem.
[0464] CdGeoA Optimized cDNA - SEQ ID NO: 53 [Chemistry]
[0465] CdGeoA, Amino Acid Sequence - SEQ ID NO: 54 [Chemistry]
[0466] VoADH1, Optimized cDNA - SEQ ID NO: 55 [Chemistry]
[0467] VoADH1, Amino Acid Sequence - SEQ ID NO: 56 [Chemistry]
[0468] Active Site Signature Motif - SEQ ID NO: 57 HCxxGxxR [where, each x independently represents any natural amino acid residue].
[0469] Active Site Signature Motif - SEQ ID NO: 58 HC(T / S)xGKDRTG [where, x represents any natural amino acid residue].
[0470] PvCPS, Multifunctional Protein with Prenyl - Transferase and Copalyl - Diphosphate Synthase, Codon - Optimized cDNA - SEQ ID NO: 59 [Chemistry]
[0471] [Chemistry]
[0472] Multifunctional protein having PvCPS, prenyl-transferase and copalyl-diphosphate synthase, amino acid sequence - SEQ ID NO: 60
Chem.
[0473] Ribosome binding site - SEQ ID NO: 61 AAGGAGGTAAAAAA
[0474] GGPP synthase from Pantoea agglomerans, codon-optimized for expression in S. cerevisiae, CrtE - SEQ ID NO: 62
Chem.
[0475] Copalyl-pyrophosphate synthase from Salvia miltiorrhiza, codon-optimized for expression in S. cerevisiae, SmCPS2 - SEQ ID NO: 63
Chem.
[0476]
Chem.
[0477] Copalyl-pyrophosphate synthase from Salvia miltiorrhiza, SmCPS2, amino acid sequence - SEQ ID NO: 64
Chem.
[0478] Coparyl - pyrophosphate phosphatase - SEQ ID NO: 65, codon - optimized for expression in TalVeTPP, S. cerevisiae
Chem.
[0479] AzTolADH1, alcohol dehydrogenase from Azoarcus toluclasticus - SEQ ID NO: 66, codon - optimized for expression in S. cerevisiae
Chem.
[0480] PsAeroADH1, alcohol dehydrogenase from Pseudomonas aeruginosa - SEQ ID NO: 67, codon - optimized for expression in S. cerevisiae
Chem.
[0481] SCH23 - ADH1, alcohol dehydrogenase from Hyphozyma roseonigra - SEQ ID NO: 68, codon - optimized for expression in S. cerevisiae
Chem.
[0482] SCH23 - ADH1, alcohol dehydrogenase from Hyphozyma roseonigra, amino acid sequence - SEQ ID NO: 69
Chem.
[0483] SCH24-ADH1a, an alcohol dehydrogenase from Cryptococcus albidus, codon-optimized for expression in S. cerevisiae—SEQ ID NO: 70 [ka]
[0484] SCH24-ADH1a, alcohol dehydrogenase from Cryptococcus albidus, amino acid sequence—SEQ ID NO: 71 [ka]
[0485] Sequence for homologous recombination 1 - SEQ ID NO: 72 [ka]
[0486] Sequence for homologous recombination 2 - SEQ ID NO: 73 [ka]
[0487] Sequence for homologous recombination 3 - SEQ ID NO: 74 [ka]
[0488] Primer for LEU2 yeast marker 1—SEQ ID NO: 75 [ka]
[0489] Primer for LEU2 yeast marker 2—SEQ ID NO: 76 [ka]
[0490] Primer for AmpR bacterial marker 1 - SEQ ID NO: 77
Chem.
[0491] Primer for AmpR bacterial marker 2 - SEQ ID NO: 78
Chem.
[0492] Primer for yeast replication origin 1 - SEQ ID NO: 79
Chem.
[0493] Primer for yeast replication origin 2 - SEQ ID NO: 80
Chem.
[0494] Primer for E. coli replication origin 1 - SEQ ID NO: 81
Chem.
[0495] Primer for E. coli replication origin 2 - SEQ ID NO: 82
Chem.
[0496] Sequence for homologous recombination 4 - SEQ ID NO: 83
Chem.
Claims
1. A method for the biocatalytic production of a bicyclic diterpenol compound, comprising: (1) contacting a corresponding bicyclic diterpenyl diphosphate precursor of the bicyclic diterpenol compound with a polypeptide having terpenyl-diphosphate phosphatase activity to produce the bicyclic diterpenol compound; and (2) optionally isolating the bicyclic diterpenol compound of step (1); wherein the polypeptide having terpenyl-diphosphate phosphatase activity is selected from diphosphate-removing enzyme members of the protein tyrosine phosphatase family, and the polypeptide having terpenyl-diphosphate phosphatase activity is one of the following polypeptides: a) TalVeTPP comprising an amino acid sequence according to SEQ ID NO: 2; b) AspWeTPP comprising an amino acid sequence according to SEQ ID NO: 6; c) Hel GriTPP comprising an amino acid sequence according to SEQ ID NO: 10; d) UmbPiTPP1 comprising an amino acid sequence according to SEQ ID NO: 13; e) TalVeTPP2 comprising an amino acid sequence according to SEQ ID NO: 16; f) HydPiTPP1 comprising an amino acid sequence according to SEQ ID NO: 19; g) TalCeTPP1 comprising an amino acid sequence according to SEQ ID NO: 22; h) TalMaTPP1 comprising an amino acid sequence according to SEQ ID NO: 25; i) TalAstroTPP1 comprising an amino acid sequence according to SEQ ID NO: 28; and j) PeSubTPP1 comprising an amino acid sequence according to SEQ ID NO: 31; and k) a polypeptide having terpenyl-diphosphate phosphatase activity and comprising an amino acid sequence showing at least 90% sequence identity with at least one of the amino acid sequences according to a) to j); A method selected from the group consisting of.
2. The method according to claim 1, wherein step (1) further comprises contacting an acyclic terpenyl diphosphate precursor with a polypeptide having bicyclic diterpenyl diphosphate synthase activity to produce the bicyclic diterpenyl diphosphate precursor.
3. The polypeptide having bicyclic diterpenyl diphosphate synthase activity is a) SmCPS2 comprising an amino acid sequence according to SEQ ID NO: 34; b) TaTps1-del59 comprising an amino acid sequence according to SEQ ID NO: 40; c) SsLPS comprising an amino acid sequence according to SEQ ID NO: 38, and d) a polypeptide comprising an amino acid sequence having bicyclic diterpenyl diphosphate synthase activity and showing a degree of sequence identity of at least 90% with at least one of the amino acid sequences according to a), b) and c), The method according to claim 2, selected from: **Claim 4** The method according to any one of claims 1 to 3, wherein the bicyclic diterpene alcohol compound produced biocatalytically is selected from copalol and labdenediol, each in the form of one stereoisomer or a mixture of at least two stereoisomers. **Claim 5** The method according to any one of claims 1 to 4, further comprising, as step (3), processing the bicyclic diterpene alcohol compound of step (1) or step (2) into an alcohol derivative using chemical synthesis or biocatalytic synthesis or a combination of both. **Claim 6** The method according to claim 5, wherein the alcohol derivative is a hydrocarbon, alcohol, diol, triol, acetal, ketal, aldehyde, acid, ether, amide, ketone, lactone, epoxide, acetate, glycoside and / or ester. **Claim 7** The method according to claim 5 or 6, wherein the bicyclic diterpene alcohol compound is biocatalytically oxidized. **Claim 8** The method according to claim 7, wherein the bicyclic diterpene alcohol compound is converted by contacting it with alcohol dehydrogenase (ADH). **Claim 9** The ADH is a) CymB comprising an amino acid sequence according to SEQ ID NO: 42; b) AspWeADH1 comprising an amino acid sequence according to SEQ ID NO: 44; c) PsAeroADH1 comprising an amino acid sequence according to SEQ ID NO: 46; d) AzTolADH1 comprising an amino acid sequence according to SEQ ID NO: 48; e) AroAroADH1 comprising an amino acid sequence according to SEQ ID NO: 50; f) ThTerpADH1 comprising an amino acid sequence according to SEQ ID NO: 52; g) CdGeoA comprising an amino acid sequence according to SEQ ID NO: 54; h) VoADH1 comprising an amino acid sequence according to SEQ ID NO: 56; i) SCH23-ADH1 comprising an amino acid sequence according to SEQ ID NO: 68 j) SCH24-ADH1a comprising an amino acid sequence according to SEQ ID NO: 70; and k) A polypeptide having ADH activity and comprising an amino acid sequence showing at least 90% sequence identity with at least one of the amino acid sequences according to a) to j) The method according to claim 8, selected from [
10. ] General formula 【Chemical Formula 1】 A method for producing an ambrox-like compound of (i) Providing a labdenediol or copalol compound by carrying out the method according to any one of claims 1 to 9; (ii) Converting the labdenediol or copalol compound of step (i) into an ambrox-like compound using chemical synthesis and / or biochemical synthesis A method comprising
Citation Information
Patent Citations
JPP7406511B