Biocatalytic process for controlled degradation of terpene compounds

By using enaldehyde cleavage of peptides and BVMO catalysis, the problem of converting terpenoids into fragrance components has been solved, achieving controlled degradation and efficient synthesis of terpenoids.

CN114630905BActive Publication Date: 2025-12-05FIRMENICH SA
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202080063163.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2019-11-13
Filing Date
2020-07-08
Publication Date
2025-12-05
Estimated Expiration
2040-07-08

AI Technical Summary

Technical Problem

Existing technologies cannot provide effective enzymatic degradation pathways for terpenes, making it impossible to convert terpenoid compounds into starting materials for valuable fragrance components such as ambroxol.

Method used

By employing novel peptides with enaldehyde cleavage activity and Bayer-Villiger monooxygenase (BVMO) catalysis, a novel biocatalytic degradation pathway is constructed to specifically shorten the carbon chain of terpene molecules, gradually converting terpene compounds into valuable derivatives such as minoyloxy and γ-ambrol.

Benefits of technology

It enables the controlled and stepwise conversion of terpenoid compounds, allowing for the efficient synthesis of valuable terpenoid derivatives to meet the needs of fragrance ingredients.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure FDA0005582931330000011
    Figure FDA0005582931330000011
  • Figure FDA0005582931330000012
    Figure FDA0005582931330000012
  • Figure FDA0005582931330000021
    Figure FDA0005582931330000021
Patent Text Reader

Abstract

Provided herein are biocatalytic methods for the production of terpene degradation products that can be used as starting materials for the production of fragrance ingredients such as ambrox. In particular, novel terpene degradation polypeptides (enal cleavage polypeptides) and novel peptides that convert terpene compounds to oxidized derivatives (oxidases) are provided, as well as mutants and variants derived therefrom, which can be applied in novel, fully enzymatic multi-step degradation pathways, allowing controlled, stepwise conversion and degradation of linear or cyclic terpene substrates. The novel biosynthetic strategies allow fully biochemical synthesis of valuable terpene-derived compounds such as minorenoloxyl or gamma-ambrol. The present invention also provides recombinant host organisms that carry the required set of genetic information for the functional expression of the set of enzymes necessary for the combination of enzymatic conversion and degradation steps.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention provides a biocatalytic method for producing terpene degradation products that can be used as starting materials for the production of fragrance components such as ambrox. Specifically, it provides novel terpene-degrading peptides (enal cleavage peptides) and novel peptides (oxidases) that convert terpene compounds into oxidative derivatives, as well as mutants and variants derived therefrom. These can be applied to novel, fully enzymatic, multi-step degradation pathways, allowing for the controlled, stepwise conversion and degradation of linear or cyclic terpene substrates. The novel biosynthetic strategy allows for the complete biochemical synthesis of valuable terpene-derived compounds such as manooloxy or γ-ambrol. The invention also provides a recombinant host organism carrying a desired set of genetic information for the functional expression of a set of enzymes necessary for catalyzing the combination of enzymatic conversion and degradation steps. Background Technology

[0002] Terpenes are found in most living organisms (microorganisms, animals, and plants). These compounds are composed of five-carbon units called isoprene units and are classified according to the number of these units present in their structure. Thus, monoterpenes, sesquiterpenes, and diterpenes are terpenes containing 10, 15, and 20 carbon atoms, respectively (i.e., 1, 2, 3, and 4 isoprene units). For example, sesquiterpenes are widely found in the plant kingdom. Many sesquiterpene molecules are well-known for their flavor and aromatic properties, as well as their cosmetic, medicinal, and antibacterial effects. Numerous sesquiterpene hydrocarbons and sesquiterpene compounds have been identified.

[0003] The biosynthesis of terpenes involves enzymes called terpene synthases. These enzymes convert acyclic terpene precursors into one or more terpene products. Specifically, diterpene synthases produce diterpenes through the cyclization of the precursor geranylgeranyl diphosphate (GGPP). The cyclization of GGPP typically requires two enzymatic polypeptides, namely type I and type II diterpene synthases, which function in combination in two consecutive enzymatic reactions. Type II diterpene synthase catalyzes the cyclization / rearrangement of GGPP initiated by the protonation of the terminal double bond, resulting in a cyclic diterpene diphosphate intermediate. This intermediate is then further converted by type I diterpene synthase, which initiates cyclization through catalytic ionization.

[0004] Diterpenoid synthases exist in plants and other organisms and use substrates such as GGPP, but they have different product characteristics. The genes and cDNAs encoding diterpenoid synthases have been cloned, and the corresponding recombinases have been characterized.

[0005] Enzymes that catalyze the specific or preferential cleavage or removal of diphosphate (ester) groups from terpene diphosphate intermediates, particularly from cyclic terpene diphosphate intermediates such as copalyl diphosphate (CPP) or labdendiol diphosphate (LPP), have only recently been described in an earlier European patent (EP application 18182783.3). With said enzymes, the number or carbon atoms of the terpene diphosphate remain unchanged.

[0006] However, terpene-derived compounds are required, which can be considered degradation products of terpene precursors such as acyclic or cyclic sesquiterpenes or diterpenes, which can then be further chemically and / or enzymatically converted into end products, for example, as flavoring ingredients.

[0007] The problem to be solved by the present invention is to provide peptides that exhibit enzymatic terpene degradation activity or peptides that convert such terpenes into degradable derivatives.

[0008] Another problem this invention aims to solve is to establish new fully biocatalytic degradation pathways to produce specific terpene degradation products. Summary of the Invention

[0009] The aforementioned problem can be unexpectedly solved by providing a new class of peptides with enaldehyde cleavage activity, which for the first time allow for the specific shortening of carbonyl-functionalized terpenoids by two carbon atoms and their respective biocatalytic processes. For example, this new class of enzymes allows the conversion of labdane-type compounds, such as copalaldehyde, which contains a diterpenoid carbon skeleton and a terminal aldehyde group, into the corresponding dedimethyl-dinor-labdane compound, minoloxy, which is shortened by two carbon atoms, thus retaining the 18-carbon skeleton.

[0010] The aforementioned problems can also be unexpectedly solved by providing novel peptides with Baeyer-Villiger monooxygenase (BVMO) activity, enabling the specific oxidation of terpenoids to esters (Baeyer-Villiger oxygenation) and the corresponding biocatalytic process. For example, this new class of BVMOs allows the conversion of copalaldehyde, a hemispherium-type compound containing a diterpene carbon skeleton and a terminal aldehyde group, into the corresponding demethylhemispherium carboxylate. Through the aforementioned Baeyer-Villiger oxygenation, hemispherium compounds can be readily converted to the corresponding demethylhemispherium by the action of peptides with esterase activity. This step thus results in the shortening of one carbon atom. In the case where the terminal aldehyde group is replaced by a terminal ketone group, shortening in the same manner, but now shortening by more than one carbonate, is possible. By repeating the combination of the BVMO-catalyzed oxygenation step and the esterase-catalyzed cleavage step, the hydrocarbon chain of the terpene molecule can be progressively shortened.

[0011] The combination of degradation steps catalyzed by the above-mentioned enal lyase and BVMO enzyme allows for the construction of entirely new biochemical degradation pathways applicable to a wider range of carbonyl-functionalized compounds, particularly cyclic or acyclic terpenes or terpenoids.

[0012] The biocatalytic step can be combined with several other preceding (upstream) or subsequent (downstream) enzymatic steps, allowing for multi-step biocatalytic methods to fully enzymatically synthesize many valuable complex terpene molecules from their respective precursors.

[0013] The following route illustrates two specific embodiments of the invention, representing two alternative pathways (the “enal cleavage peptide pathway” and the “BMVO pathway”) that allow the degradation of copalaldehyde (a serotonaldehyde) to minoyloxy group, which are explained in more detail later in this specification. The route also illustrates the degradation of minoyloxy group to γ-ambrol by applying a further BMVO-based degradation step.

[0014]

[0015] Completely similar to the exemplary reaction sequence described above, this basic biosynthetic strategy can be applied to any other isomer of coparol or any other hemispheric aldehyde to provide structure-related isomers of minoyloxy, γ-ambryl acetate, or γ-ambrol.

[0016] As explained in more detail below, it can also be applied to monocyclic or acyclic carbonyl compounds with different structures. Attached Figure Description

[0017] Figure 1 A schematic diagram of the chromosomal integration of genes encoding mevalonate pathway enzymes and the organization of the two synthetic gene operons. mvaK1, a gene encoding mevalonate kinase from *Streptococcus pneumoniae*; mvaD, a gene encoding mevalonate phosphate decarboxylase from *Streptococcus pneumoniae*; mvaK2, a gene encoding mevalonate phosphate kinase from *Streptococcus pneumoniae*; fni, a gene encoding isopentenyl diphosphate isomerase from *Streptococcus pneumoniae*; mvaA, a gene encoding HMG-CoA synthase from *Staphylococcus aureus*; mvaS, a gene encoding HMG-CoA reductase from *Staphylococcus aureus*; atoB, a gene encoding acetyl-CoA thiolase from *Escherichia coli*; ERG20, a gene encoding FPP synthase from *Saccharomyces cerevisiae*.

[0018] Figure 2. Conversion of minoyloxy group to γ-ambryl acetate using BVMO in a whole-cell biotransformation assay. GC-MS analysis of the products formed during the biotransformation of minoyloxy group using different BVMOs: SCH23-BVMO1, SCH24-BVMO1, and SCH46-BVMO1. The upper chromatogram shows the GC-MS analysis of minoyloxy group. The lower chromatogram shows the GC-MS analysis of the biotransformation using control cells that do not express recombinant BVMO.

[0019] Figure 3. Copacaldehyde transformation using BVMOs in whole-cell biotransformation assays. GC-MS analysis of products formed during the biotransformation of cis-copacaldehyde and trans-copacaldehyde using different BVMOs: SCH23-BVMO1, SCH24-BVMO1, and SCH46-BVMO1 (compounds 3a, 3b, 4a, and 4b as described in the experimental section). The upper chromatogram shows the GC-MS analysis of biotransformation using control cells that do not express recombinant BVMOs.

[0020] Figure 4 Kinetics of copalaldehyde conversion using SCH23-BVMO1 in whole-cell biotransformation assays. GC-MS analysis of products formed during cis-copalaldehyde and trans-copalaldehyde biotransformation via SCH23-BVMO1 after incubation for 0, 18, and 42 hours. (Compounds 1a, 1b, 3a, 3b, 4a, and 4b as described in the Experimental Section).

[0021] Figure 5 In vitro transformation of the mirenoyl group using BVMO. GC-MS analysis of the transformation of the mirenoyl group by SCH23-BVMO1 and SCH24-BVMO1 showed the formation of γ-ambrol acetate. The upper chromatogram shows the GC-MS analysis of the transformation using a control protein without recombinant BVMO.

[0022] Figure 6 In vitro transformation of the mirenoyl group using BVMO and esterase. GC-MS analysis of transformations of the mirenoyl group by SCH23-BVMO1, SCH23-EST, and a combination of SCH23-BVMO1 and SCH23-EST showed the formation of γ-ambrol. The upper chromatogram shows GC-MS analysis of the transformations using a control protein without recombinase.

[0023] Figure 7In vitro transformation of the mirenoyl group using BVMO and esterase. GC-MS analysis of transformations of the mirenoyl group by SCH24-BVMO1, SCH24-EST, and a combination of SCH24-BVMO1 and SCH24-EST showed the formation of γ-ambrol. The upper chromatogram shows GC-MS analysis of the transformations using a control protein without recombinase.

[0024] Figure 8. In vitro conversion of compounds 4a and 4b to compounds 5a and 5b using esterases. GC-MS analysis of in vitro conversions of compounds 4a and 4b by SCH23-EST1, SCH24-EST1, and SCH25-EST1 shows the formation of compounds 5a and 5b.

[0025] Figure 9 Copacaldehyde was converted in vitro to compounds 5a and 5b using SCH23-BVMO1 and esterase. GC-MS analysis of the in vitro conversion of cis-copacaldehyde and trans-copacaldehyde by combination of SCH23-BVMO with SCH23-EST1, SCH24-EST1, and SCH25-EST1 revealed the formation of compounds 5a and 5b. The peak marked with * and a retention time of 11.95 min corresponds to γ-ambryl acetate; this compound was observed in samples incubated with BVMO alone due to the presence of trace amounts of minoyloxy groups in the copacaldehyde mixture used in these assays.

[0026] Figure 10 Copacaldehyde was converted in vitro to compounds 5a and 5b using SCH24-BVMO1 and an esterase. GC-MS analysis of the in vitro conversion of cis-copacaldehyde and trans-copacaldehyde by combination of SCH23-BVMO with SCH23-EST1 and SCH25-EST1 revealed the formation of compounds 5a and 5b. The peak marked with * at a retention time of 11.95 min corresponds to γ-ambryl acetate; this compound was observed in samples incubated with BVMO alone due to the presence of trace amounts of minoyloxy groups in the copacaldehyde mixture used in these assays.

[0027] Figure 11. Biochemical production of compounds 5a and 5b of 14,15-dedimethylhexamonade and biosynthetic intermediates in engineered bacterial cells expressing BVMO and esterase. The upper chromatogram shows the GC-MS analysis of the compounds produced by *E. coli* cells transformed with the pJ401-CPAL-1 plasmid, which allows expression of enzymes in the copal aldehyde biosynthetic pathway. The following chromatogram shows the GC-MS analysis of cells further transformed with a second plasmid containing nucleotide sequences encoding either BVMO enzyme or both BVMO enzyme and esterase.

[0028] Figure 12GC-MS analysis of the biotransformation products of compounds 5a and 5b using *E. coli* cells expressing various alcohol dehydrogenases. The upper chromatogram shows the GC-MS analysis of biotransformation using control cells that do not express recombinant alcohol dehydrogenases. The following chromatogram shows the GC-MS analysis of transformation using cells expressing recombinant RrhSecADH, SCH80-00043, SCH80-04254, SCH80-06135, or SCH80-06582 proteins.

[0029] Figure 13 Biochemical production of γ-ambryl acetate and biosynthetic intermediates in engineered bacterial cells expressing BVMO, esterase, and alcohol dehydrogenase. The upper chromatogram shows GC-MS analysis of compounds produced by *E. coli* cells transformed with the pJ401-CPAL-1 plasmid, allowing expression of enzymes in the copal aldehyde biosynthetic pathway. The middle chromatogram shows GC-MS analysis of cells further transformed with a second plasmid containing nucleotide sequences encoding SCH-BVMO1 and SCH24-EST. The lower chromatogram shows GC-MS analysis of cells transformed with pJ401-CPAL-1 and the plasmid pJ423-secADH-23BVMO-EST, allowing expression of the proteins RrhSecADH, SCH23-BVMO1, and SCH23-EST.

[0030] Figure 14. A) GC-MS analysis of terpenes and derivatives produced using modified Saccharomyces cerevisiae strains expressing GGPP synthase carG, CPP synthase SmCPS2, CPP phosphatase TalVeTPP, and either SCH23-ADH1, SCH23-BVMO1, SCH23-EST1, and SCH23-ADH2 (YST120 using plasmids) or SCH24-ADH1a, SCH24-BVMO1, SCH24-EST1, and SCH24-ADH2a (YST121 using plasmids). The control strain is YST075, which expresses only the copal alcohol biosynthetic pathway. B) GC-MS analysis identifying the region of farnesal, showing the mass spectrum of farnesal. C) GC-MS analysis identifying the region of mirenoyloxy group, showing the mass spectrum of mirenoyloxy group.

[0031] Figure 15. GC-MS analysis of minoyloxy groups produced by modified Saccharomyces cerevisiae strains expressing GGPP synthase carG, CPP synthase SmCPS2, CPP phosphatase TalVeTPP, and either SCH23-ADH1, SCH23-BVMO1, and SCH23-EST1 (YST177) or SCH24-ADH1a, SCH24-BVMO1, and SCH24-EST1 (YST178). The control strain is YST075, which expresses only the copal alcohol biosynthetic pathway. Mass spectra of minoyloxy groups are shown.

[0032] Figure 16. GC-MS analysis of diterpenes and derivatives produced using *E. coli* cells expressing CPP synthase, phosphatase, alcohol dehydrogenase, and / or SCH94-3944. The upper chromatogram shows the diterpenoid region, and GC-MS analysis of compounds produced using *E. coli* cells transformed with the pJ401-CPOL-4 plasmid, allowing expression of enzymes in the copalol biosynthetic pathway. The following chromatogram shows GC-MS analysis of compounds produced from the same *E. coli* cells further transformed with plasmids pJ423-SCH94-3945, pJ423-SCH94-3944, or pJ423-SCH94-3944-3945, allowing expression of SCH94-3945, SCH94-3944, or a combination of SCH94-3944 and SCH94-3945.

[0033] Figure 17 GC-MS analysis of sesquiterpenes and derivatives produced using *E. coli* cells expressing phosphatases, alcohol dehydrogenases, and SCH94-3944. The upper chromatogram shows the GC-MS analysis of compounds produced by *E. coli* cells transformed with the pJ401-FAL-1 plasmid, allowing expression of enzymes in the farnesal biosynthesis pathway. The lower chromatogram shows the GC-MS analysis of compounds produced by the same *E. coli* cells further transformed with the pJ423-SCH94-3944 plasmid, allowing expression of the SCH94-3944 protein.

[0034] Figure 18 GC-MS analysis of the products of biotransformation of citral, citronellol, and (E)-2-dodecanoic acid by E. coli cells expressing SCH94-3944. For each compound, GC-MS analysis of transformations using control E. coli cells and cells transformed to express the SCH94-3944 protein is shown.

[0035] Figure 19GC-MS analysis of sesquiterpenes and diterpenes produced by *E. coli* cells expressing CPP synthase, phosphatase, and alcohol dehydrogenase. Chromatograms show GC-MS analysis of compounds produced by *E. coli* cells transformed with pJ401-CPAL-1 plasmid to allow expression of enzymes in the copal aldehyde biosynthesis pathway.

[0036] Figure 20 GC-MS analysis was performed on diterpenes and derivatives produced by *E. coli* cells expressing CPP synthase, phosphatase, alcohol dehydrogenase, and the recombinant proteins SCH80-05241, SCH94-3944, PdigitDUF4334, PitalDUF4334-1, or AspWeDUF4334. The upper chromatogram shows the diterpenoid region in the GC-MS analysis of compounds produced by *E. coli* DP1205 cells transformed with the pJ401-CPAL-1 plasmid to allow expression of enzymes in the copal aldehyde biosynthesis pathway. The following chromatogram shows the GC-MS analysis of compounds produced by the same *E. coli* cells further transformed with a second plasmid expressing recombinant proteins of SCH80-05241, SCH94-3944, PdigitDUF4334, PitalDUF4334-1, or AspWeDUF4334.

[0037] Figure 21 GC-MS analysis of diterpenes and derivatives produced by *E. coli* cells expressing CPP synthase, phosphatase, alcohol dehydrogenase, and the genera *CnecaDUF4334*, *Rins-DUF4334*, *RhoagDUF4334-2*, *RhoagDUF4334-3*, *RhoagDUF4334-4*, *CgatDUF4334*, *GclavDUF4334*, *TcurvaDUF4334*, or *PprotDUF4334*. The upper chromatogram shows the diterpenoid regions of compounds produced by *E. coli* DP1205 cells transformed with the pJ401-CPAL-1 plasmid to allow expression of enzymes in the copal aldehyde biosynthesis pathway, as analyzed by GC-MS. The following chromatograms show GC-MS analysis of compounds produced from the same E. coli cells, which were further transformed with a second plasmid expressing recombinant proteins of CnecaDUF4334, Rins-DUF4334, RhoagDUF4334-2, RhoagDUF4334-3, RhoagDUF4334-4, CgatDUF4334, GclavDUF4334, TcurvaDUF4334, or PprotDUF4334.

[0038] Figure 22GC-MS analysis of sesquiterpenes and derivatives produced by *E. coli* cells expressing phosphatases, alcohol dehydrogenases, and the recombinant proteins SCH80-05241, SCH94-3944, PdigitDUF4334, PitalDUF4334-1, or AspWeDUF4334 was performed. The upper chromatogram shows the sesquiterpene region in the GC-MS analysis of compounds produced by *E. coli* DP1205 cells transformed with the pJ401-CPAL-1 plasmid to allow expression of enzymes in the copal aldehyde biosynthesis pathway. The following chromatogram shows the GC-MS analysis of compounds produced by the same *E. coli* cells further transformed with a second plasmid expressing recombinant proteins of SCH80-05241, SCH94-3944, PdigitDUF4334, PitalDUF4334-1, or AspWeDUF4334.

[0039] Figure 23 GC-MS analysis of sesquiterpenes and derivatives produced by *E. coli* cells expressing phosphatases, alcohol dehydrogenases, and the genera *CnecaDUF4334*, *Rins-DUF4334*, *RhoagDUF4334-2*, *RhoagDUF4334-3*, *RhoagDUF4334-4*, *CgatDUF4334*, *GclavDUF4334*, *TcurvaDUF4334*, or *PprotDUF4334*. The upper chromatogram shows the sesquiterpene regions of compounds produced by *E. coli* cells transformed with the pJ401-CPAL-1 plasmid to allow expression of enzymes in the copal aldehyde biosynthesis pathway, as analyzed by GC-MS. The following chromatograms show GC-MS analysis of compounds produced from the same E. coli cells, which were further transformed with a second plasmid expressing recombinant proteins of CnecaDUF4334, Rins-DUF4334, RhoagDUF4334-2, RhoagDUF4334-3, RhoagDUF4334-4, CgatDUF4334, GclavDUF4334, TcurvaDUF4334, or PprotDUF4334.

[0040] Figure 24. Alignment and conserved amino acids of the GXWXG and DUF4334 domains of proteins containing enzymes that catalyze enzymatic enol cleavage. Boxes indicate the predicted localization of the corresponding protein family domains.

[0041] Figure 25 The conversion activities of farnesal and copalaldehyde by a single amino acid variant of SCH94-3944. The activity is expressed as the total amount of minoyloxy and geranylacetone produced, as a percentage relative to the wild-type enzyme activity.

[0042] Figure 26GC-MS analysis of the biochemical production of minocycline and γ-ambryl acetate by *E. coli* cells expressing CPP synthase, phosphatase, alcohol dehydrogenase, enal lyase, and BVMO. The upper chromatogram shows the diterpenoid region of compounds produced by *E. coli* DP1205 cells transformed with the pJ401-CPAL-1 plasmid to allow expression of enzymes in the copal aldehyde biosynthesis pathway, as analyzed by GC-MS. The following chromatogram shows the GC-MS analysis of compounds produced by the same *E. coli* cells further transformed with second plasmids expressing AspWeBVMO, SCH94-3944, SCH94-3944 with AspWeBVMO, SCH94-3944 with SCH23-BVMO1, SCH94-3944 with SCH24-BVMO1, and SCH94-3944 with SCH46-BVMO1.

[0043] Figure 27. GC-MS analysis of terpenes and derivatives produced by modified Saccharomyces cerevisiae strains expressing GGPP synthase carG, CPP synthase SmCPS2, CPP phosphatase TalVeTPP, alcohol dehydrogenase SCH23-ADH1, and AspWeDUF4334 (YST184), CnecaDUF4334 (YST185), Pdigit7033 (YST186), SCH94-3944 (YST187), or SCH80-05241 (YST188).

[0044] Figure 28 A) Percentage of identified terpenes produced via YST184, YST185, YST186, YST187, and YST188. B) Total amount of identified terpenes produced via YST184, YST185, YST186, YST187, and YST188 (SumT) relative to the amount of identified terpenes in the control (SumT-C). The control strain was YST075, which expresses the copalol biosynthesis pathway.

[0045] Figure 29. GC-MS analysis of terpenes and derivatives produced by modified Saccharomyces cerevisiae strains expressing GGPP synthase carG, CPP synthase SmCPS2, CPP phosphatase TalVeTPP, alcohol dehydrogenase SCH23-ADH, enaldehyde cleavage peptide Asp4WeDUF43, and SCH23-BVMO1 (YST190), SCH24-BVMO1 (YST191), or AspWeBVMO (YST192).

[0046] Figure 30A) The total amount of identified terpenes produced by YST190, YST191, and YST192 relative to the amount of identified terpenes in YST184 (SumT-C). B) The percentage of identified terpenes produced by YST190, YST191, and YST192.

[0047] Figure 31 GC-MS analysis of diterpenes and diterpene derivatives produced using *E. coli* cells expressing LPP synthase, phosphatase, alcohol dehydrogenase, and enal cleavage peptide. The upper chromatogram shows GC-MS analysis of compounds produced by *E. coli* DP1205 cells transformed with the pJ401-LOH-2 vector to allow expression of enzymes in the hemispheric ethylenediol biosynthesis pathway. The following chromatogram shows GC-MS analysis of compounds produced by the same *E. coli* cells further transformed with a second plasmid expressing either AzeTolADH1 alcohol dehydrogenase or SCH94-3945 alcohol dehydrogenase and the SCH94-3944 enal cleavage peptide.

[0048] Figure 32. Alignment and conserved amino acids of FMO-like domains containing proteins with BVMO activity. Boxes indicate the predicted localization of the corresponding protein family domains.

[0049] Figure 33. GC-MS / FID analysis of terpenes and derivatives produced by modified Saccharomyces cerevisiae strains expressing bifunctional PvCPS, CPP phosphatase TalVeTPP, alcohol dehydrogenase SCH23-ADH, enal cleavage peptide AspWeDUF4334, Bayer-Villiger monooxygenase SCH23-BVMO1, and either esterase SCH23-EST (YST257) or esterase SCH24-EST (YST258).

[0050] Figure 34 GC-MS analysis of γ-ambrol biochemical production by *E. coli* cells expressing CPP synthase, phosphatase, alcohol dehydrogenase, enal lyase, BVMO, and esterase. A. GC-MS analysis of compounds produced by *E. coli* DP1205 cells transformed with pJ401-Mnoxy plasmid to allow expression of enzymes in the minooloxy biosynthesis pathway. B. GC-MS analysis of compounds produced by the same *E. coli* cells further expressing BVMO (SCH24-BVMO). C. GC-MS analysis of compounds produced by the same *E. coli* cells further expressing BVMO (SCH24-BVMO) and esterase (SCH24-EST). Detailed Implementation

[0051] abbreviations used

[0052] ADH alcohol dehydrogenase

[0053] BVMO (Baeyer-Villiger monooxygenase)

[0054] bp base pairs

[0055] kb kilobase

[0056] CPP (cobamic acid ester)

[0057] CPS cobazyl bisphosphonate synthase

[0058] DNA deoxyribonucleic acid

[0059] cDNA complementary DNA

[0060] DMAPP (dimethyl allyl diphosphate)

[0061] DTT dithiothreitol

[0062] FMO flavin monooxygenase

[0063] FPP farnesyl diphosphate (ester)

[0064] GPP Geraniyl diphosphate (ester)

[0065] GGPP Geraniyl Geraniyl Diphosphate (ester)

[0066] GGPS Geraniylgeraniyl diphosphate synthase

[0067] GC gas chromatography

[0068] IPP (Isoprene diphosphate)

[0069] LPP (Lactobionic acid ester)

[0070] LPS (Lactobacillus subtilis) bisphosphonate synthase

[0071] MS mass spectrometer / mass spectrometry

[0072] MVA (mevaleric acid)

[0073] PP diphosphate (ester), pyrophosphate (ester)

[0074] PCR Polymerase Chain Reaction

[0075] RNA (ribonucleic acid)

[0076] mRNA messenger ribonucleic acid

[0077] miRNA

[0078] siRNA (small interfering RNA)

[0079] rRNA ribosomal RNA

[0080] tRNA transfer RNA

[0081] TPP (terpene diphosphate)

[0082] definition

[0083] a) General terminology:

[0084] In connection with the description and appended claims herein, the use of “or” means “and / or” unless otherwise stated. Similarly, the tenses of “containing,” “containing,” “comprising,” and “including” are interchangeable and not restrictive.

[0085] It should be further understood that, where the term “comprising” is used in the description of various implementation schemes, those skilled in the art will understand that, in certain specific cases, the language of “substantially consisting of” or “consisting of” can be used instead to describe the implementation schemes.

[0086] As used herein, the terms “purified,” “substantially purified,” and “isolated” refer to a state free from other different compounds (of which the compounds of the present invention are generally associated in their natural state). Therefore, “purified,” “substantially purified,” and “isolated” articles constitute at least 0.5%, 1%, 5%, 10%, or 20%, or at least 50% or 75% of the mass of a given sample by weight. In one embodiment, these terms refer to the compounds of the present invention constituting at least 95%, 96%, 97%, 98%, 99%, or 100% of the mass of a given sample by weight. As used herein, when referring to nucleic acids or proteins, the terms “purified,” “substantially purified,” and “isolated” for nucleic acids or proteins also refer to a purified or concentrated state that differs from a state naturally present in, for example, prokaryotic or eukaryotic environments, such as in bacterial or fungal cells, or in mammals, particularly humans. Any degree of purification or concentration greater than that of naturally occurring purified or concentrated nucleic acids, including (1) purification from other related structures or compounds, or (2) association with structures or compounds that are not normally associated with them in the prokaryotic or eukaryotic environment, is considered "isolated". Nucleic acids, proteins, or classes of nucleic acids or proteins described herein may be isolated or otherwise associated with structures or compounds that are normally unrelated in nature, according to various methods and processes known to those skilled in the art.

[0087] The term “about” indicates a possible variation of ±25% in the value, particularly ±15%, ±10%, more particularly ±5%, ±2%, or ±1%.

[0088] The term "substantially" describes a value range of approximately 80% to 100%, such as 85% to 99.9%, particularly 90% to 99.9%, even more particularly 95% to 99.9%, or 98% to 99.9%, particularly 99% to 99.9%.

[0089] "Mainly" refers to a proportion in the range of 50%, such as in the range of 51% to 100%, particularly in the range of 75% to 99.9%; especially in the range of 85% to 98.5%, such as 95% to 99%.

[0090] In the context of this invention, "major product" means a single compound or a group of at least two compounds, such as two, three, four, five or more, particularly two or three compounds, which are "majorly" prepared by the reaction described herein and are contained in the reaction in a major proportion based on the total amount of the components of the products formed by the reaction. The proportion may be a molar proportion, a weight proportion, or preferably an area proportion calculated from the corresponding chromatograms of the reaction products based on chromatographic analysis.

[0091] In the context of this invention, "byproduct" means a single compound or a group of at least two compounds, such as two, three, four, five or more, particularly two or three compounds, which are not "mainly" prepared by the reaction described herein.

[0092] Due to the reversibility of enzymatic reactions, unless otherwise stated, this invention relates to enzymatic or biocatalytic reactions described herein in both reaction directions.

[0093] The “functional mutants” of the peptides described in this article include “functional equivalents” of such peptides as defined below.

[0094] The term "stereoisomer" includes conformational isomers, and in particular configurational isomers.

[0095] According to the present invention, all "stereoisomers" of the compounds described herein are generally included, such as "structural isomers" and "stereoisomers".

[0096] "Stereoisomeric forms" specifically include "stereoisomers" and mixtures thereof, such as configurational isomers (optical isomers), such as enantiomers, or geometric isomers (diastereomers), such as E- and Z-isomers, and combinations thereof. If one or more asymmetric centers are present in a molecule, the invention includes all combinations of different conformations of these asymmetric centers, such as enantiomer pairs.

[0097] "Stereoselectivity" describes the ability to produce a specific stereoisomer of a compound in its stereoisomeric pure form, or the ability to specifically convert a specific stereoisomer from a variety of stereoisomers using the enzymatic catalytic method described herein. More specifically, this means that the product of the invention is enriched relative to a specific stereoisomer, or that the precipitate can be depleted relative to a specific stereoisomer. This can be quantified by a purity parameter %ee calculated according to the following formula:

[0098] %ee=[X A -X B ] / [X A +X B ]*100,

[0099] Where X A and X B Mole ratio (Molenbruch) represents the ratio of stereoisomers A and B.

[0100] The terms "selective conversion" or "increased selectivity" generally refer to the conversion of a specific stereoisomer, such as the E-form of the unsaturated hydrocarbon, at a higher proportion or amount (compared to molar amounts) than the corresponding other stereoisomers, such as the Z-form, during the entire course of the reaction (i.e., between the start and end of the reaction), at a certain point in time of the reaction, or during a "segment" of the reaction. Specifically, during the "segment," the selectivity can be observed corresponding to conversions of 1 to 99%, 2 to 95%, 3 to 90%, 5 to 85%, 10 to 80%, 15 to 75%, 20 to 70%, 25 to 65%, 30 to 60%, or 40 to 50% of the initial amount of substrate. The higher proportion or amount can be expressed, for example, as follows:

[0101] -High maximum yield of isomers observed throughout the entire reaction process or during a specified period;

[0102] - Higher relative isomer content at a given percentage substrate conversion value; and / or

[0103] -At higher conversion percentage values, the same relative content of isomers;

[0104] Each of these preferred methods is observed relative to a reference method, which is performed under otherwise identical conditions using known chemical or biochemical methods.

[0105] According to the present invention, all “isomeric forms” of the compounds described herein are generally included, such as structural isomers, especially stereoisomers and mixtures thereof, such as optical isomers or geometric isomers, such as E and Z isomers, and combinations thereof. If several asymmetric centers exist in a molecule, the present invention includes all combinations of different conformations of these asymmetric centers, such as enantiomer pairs, or any mixture of stereoisomeric forms.

[0106] The “yield” and / or “conversion” of the reaction according to the invention are determined within a specified time period, such as 4, 6, 8, 10, 12, 16, 20, 24, 36 or 48 hours (during which the reaction takes place). In particular, the reaction is carried out under precisely defined conditions, such as the “standard conditions” as defined herein.

[0107] Different yield parameters (“yield” or YP / S; “specific productivity yield”; or space-time yield (STY)) are well known in the art and are measured as described in the literature.

[0108] "Yield" and "YP / S" (both expressed as the quality of products produced / the quality of materials consumed) are used as synonyms in this document.

[0109] Specific productivity yield describes the amount of product produced per hour per liter of fermentation broth per gram of biomass. Wet cell weight (WCW) describes the number of biologically active microorganisms in the biochemical reaction. This value is given as grams of product per g WCW per hour (i.e., g / gWCW). -1 h -1 Alternatively, the amount of biomass can also be expressed as the amount of dry cell weight, denoted as DCW. Furthermore, by measuring 600 nm (OD) 600 By estimating the corresponding wet or dry cell weights using the optical density at a given location and relevant factors determined experimentally, biomass concentration can be determined more easily.

[0110] If this disclosure relates to features, parameters, and ranges of different priorities (including superior, non-preferred features, parameters, and ranges), then unless otherwise stated, any combination of two or more of these features, parameters, and ranges is covered in the disclosure of this invention regardless of their respective priority.

[0111] b) Biochemical terminology

[0112] The term "domain" refers to a group of amino acids or portions of an amino acid sequence that are conserved at a specific position along an alignment of an evolution-related protein sequence. While amino acids at other positions may vary between protein homologs, highly conserved amino acids at specific positions within such domains represent amino acids that may be essential for the protein's structure, stability, or function. Identified by their high conservation in aligned sequences of protein homolog families, they can be used as identifiers to determine whether any polypeptide under discussion belongs to a previously identified polypeptide family.

[0113] The term "motif," "shared sequence," or "signature" refers to a short, conserved region in an evolutionarily relevant protein sequence. Motifs are typically highly conserved portions of a domain, but may also include only a part of the domain.

[0114] A "protein family" is defined as a group of proteins that share a common evolutionary origin, reflected in their related functions, sequence similarities, or similar primary, secondary, or tertiary structures. Proteins within a protein family are typically homologous and possess similar conserved functional domains and motifs.

[0115] Specialized databases exist for identifying structural domains, such as SMART (Schultz et al., (1998) Proc. Natl. Acad. Sci. USA 95, 5857-5864; Letunic et al., (2002) Nucleic Acids Res 30, 242-244), InterPro (Mulder et al., (2003) Nucleic Acids Res. 31, 315-318), and Prosite (Bucher and Bairoch (1994), General summary syntax of biomolecular sequence motifs and its function in automatic sequence interpretation. (In) ISMB-94; Proceedings 2nd International Conference on Intelligent Systems for Molecular Biology. Altman R., Brutlag D., Karp P., Lathrop R., Searls D., Eds., pp53-61, AAAI Press, Menlo. Park; Hulo et al., Nucl. Acids. Res. 32: D134-D137, (2004)), or Pfam (Bateman et al., Nucleic Acids Research 30(1): 276-280 (2002)). Domains or motifs can also be identified using conventional techniques such as sequence alignment.

[0116] The term "Pfam" refers to a large collection of protein domains and protein families maintained by the Pfam consortium, available on several sponsored World Wide Web sites, such as http: / / pfam.xfam.org / / (European Molecular Biology Laboratory - European Institute for Bioinformatics (EMBL_EBI)). The latest version of Pfam is Pfam 32.0 (September 2018), based on UniProtReference Proteomes (El-Gebali S. et al., 2019, Nucleic Acids Res. 47, Database issue). (D427–D432). Pfam domains and families are identified using multiple sequence alignments and Hidden Markov Models (HMMs). Pfam-A family or domain assignments are high-quality assignments generated using a curated seed alignment and a seed alignment-based profile HMM (unless otherwise stated, a match between the queried protein and a Pfam domain or family is a Pfam-A match). A complete alignment of the family is then automatically generated using all identified sequences belonging to that family (Sonnhammer (1998) Nucleic Acids Research 26, 320–322; Bateman (2000) Nucleic Acids Research 26, 263–266; Bateman (2004) Nucleic Acids Research 32, Database Issue, D138–D141; Finn (2006) Nucleic Acids Research Database Issue). 34, D247-251; Finn (2010) Nucleic Acids Research Database Issue 38, D211-222). For example, by accessing the Pfam database using any of the aforementioned reference websites, protein sequences can be queried against the HMM using HMMER homology search software (e.g., HMMER2, HMMER3, or later hmmer.janelia.org / ). An important match that identifies a queried protein as belonging to the pfam family (or having a specific Pfam domain) is a match with a bit score greater than or equal to the collection threshold of the Pfam domain. The expected value (e-value) can also be used as a criterion for including the queried protein in the Pfam or determining whether the queried protein has a specific Pfam domain, where the e-value is low, much less than 1.0, e.g., less than 0.1 or less.

[0117] The "E-value" (expected value) refers to the number of hits where the score is expected to be equal to or higher than this value by chance. This means that a good E-value that provides a reliable prediction is much less than 1. An E-value near 1 represents a chance expectation. Therefore, the lower the E-value, the more specific the search for the domain. Only positive numbers (defined by Pfam) are allowed.

[0118] The "precursor" molecule of the target compound described herein is preferably converted into the target compound by enzymatic action of a suitable polypeptide that modifies the precursor molecule by at least one structural change. For example, a "diphosphate precursor" (e.g., a "terpenoid diphosphate precursor") can be converted into the target compound (e.g., a terpene alcohol) by enzymatic removal of the diphosphate moiety, such as by removing a monophosphate or diphosphate group, using a phosphatase. For example, an "acyclic precursor" (e.g., an acyclic terpene precursor) can be converted into a cyclic target molecule (e.g., a cyclic terpene compound) by one or more steps using a cyclase or synthase, regardless of the specific enzymatic mechanism of the enzyme.

[0119] The term "protein tyrosine phosphatase" refers to a group of enzymes generally known to remove phosphate groups from phosphorylated tyrosine residues on proteins. A specific subgroup of this family described herein consists of enzymes that dephosphorylate phosphorylated terpene molecules.

[0120] "Terpene synthase" refers to a polypeptide that converts a terpene precursor molecule into the corresponding terpene target molecule, such as, in particular, a processed target terpene alcohol. Non-limiting examples of such terpene precursor molecules are, for example, acyclic compounds selected from farnesyl pyrophosphate (FPP), geraniylgeraniyl pyrophosphate (GGPP), or a mixture of isopentenyl pyrophosphate (IPP) and dimethylallyl pyrophosphate (DMAPP). When the obtained terpene contains a diphosphate moiety, the synthase is called "terpenoid diphosphate synthase".

[0121] The terms "terpenoid diphosphate synthase," "polypeptide with terpenoid diphosphate synthase activity," "terpenoid diphosphate synthase protein," or "capable of producing terpenoid diphosphate" refer to a polypeptide capable of catalyzing the synthesis of terpenoid diphosphate in any stereoisomer or mixture thereof, the synthesis beginning with acyclic terpenoid pyrophosphate, particularly GPP, FPP, or GGPP or IPP, in conjunction with DMAPP. The terpenoid diphosphate may be the sole product or may be part of a mixture of terpenoid phosphate esters. The mixture may contain terpenoid monophosphate esters and / or terpene alcohols. The above definition also applies to "bicyclic diterpenoid diphosphate synthases" that produce bicyclic terpenoid diphosphates, such as CPP or LPP. As an example of such "terpenoid diphosphate synthases," copacpium diphosphate synthase (CPS) may be mentioned. Copacpium diphosphate may be the sole product or may be part of a mixture of copacpium phosphate esters. The mixture may contain copacpium monophosphate esters and / or other terpenoid diphosphates. As another example of this type of "diterpenoid diphosphate synthase" enzyme, serotonin diphosphate synthase (LPS) can be mentioned. Serotonin diphosphate may be the sole product or may be part of a mixture of serotonin diphosphate esters. The mixture may contain serotonin diphosphate monophosphate and / or terpenoid diphosphates.

[0122] The terms "terpenoid diphosphate phosphatase," "polypeptide with terpenoid diphosphate phosphatase activity," "terpenoid diphosphate phosphatase protein," or "capable of producing terpenoid alcohols" refer to a polypeptide capable of catalyzing the removal of either the diphosphate or monophosphate moiety (regardless of the specific enzymatic mechanism) to form a dephosphorylated compound, particularly the corresponding alcohol compound of the terpenoid moiety. Terpenoid alcohols may be present in the product in any of their stereoisomers or mixtures thereof. Terpenoid alcohols may be the sole product or part of a mixture with other terpenoid compounds (e.g., dephosphorylated analogs of various (e.g., acyclic) terpenoid diphosphate ester precursors of the terpenoid diphosphate). The above definition also applies to "bicyclic terpenoid diphosphate phosphatases" that produce bicyclic terpenoid alcohols, such as copalol or hemispheryl acetate diol.

[0123] As an example of this type of "terpenoid diphosphate phosphatase," copaclitaxel diphosphate phosphatase (CPP phosphatase) can be mentioned. Copaclitaxel may be the sole product, or it may be part of a mixture with a dephosphorylated precursor (e.g., farnesol and / or geraniol); and / or a byproduct of enzymatic side reactions in the reaction mixture, such as esters or aldehydes of such alcohols or other cyclic or acyclic diterpenes. Another example of this type of "terpenoid diphosphate phosphatase" enzyme can be mentioned: monazine diol diphosphate phosphatase (LPP phosphatase). Monazine diol may be the sole product, or it may be part of a mixture with a dephosphorylated precursor (e.g., farnesol and / or geraniol); and / or a byproduct of enzymatic side reactions in the reaction mixture, such as esters or aldehydes of such alcohols or other cyclic or acyclic diterpenes.

[0124] In the context of this invention, "enal lyase," "enal lyase protein," or "enal lyase polypeptide" refers to an "α,β-unsaturated aldehyde carbon-carbon double bond lyase," which can also be called an "α,β-unsaturated aldehyde C=C bond lyase," "α,β-unsaturated aldehyde C=C lyase," or "enal C=C lyase." Based on protein domain organization, the enal lyase protein of this invention can also be described as a member of the "DUF4334 protein family" and / or the "GXWXG protein family."

[0125] More specifically, the enal lyases of the present invention have the ability to cleave hemispherium-type carbonyl compounds, such as hemispherium aldehyde, particularly coparaldehyde, into the corresponding desdimethylhemispherium carbonyl compounds. “Bayer-Villiger monooxygenases” (BVMOs) are flavinases and belong to the class of polypeptides with oxidoreductase activity (EC 1.14.13.X). They catalyze the oxidation of straight-chain, cyclic (aromatic or non-aromatic) aldehydes or ketones into the corresponding esters or lactones, very similar to chemical Bayer-Villiger oxidation. In the enzymatic oxidation process, one atom of molecular oxygen binds to the carbon-carbon bond of the unactivated carbonyl compound. BVMOs require NADPH or NADH as cofactors or accept both. They also require molecular oxygen as a co-substrate. More specifically, the BVMOs of the present invention are capable of oxidizing terpene-derived aldehydes or ketones, such as hemispherium-type carbonyl compounds like hemispherium aldehyde, particularly coparaldehyde, and / or minoyloxy groups, into the corresponding carbonyl esters.

[0126] "Esterase" refers to a polypeptide with hydrolytic activity that breaks down esters into acids and alcohols in a chemical reaction (hydrolysis) with water. In the context of this invention, esterases are selected from carboxylic acid ester hydrolases (EC 3.1.1-) that cleave acyl groups, such as acetyl or formyl groups, from their respective ester substrates. More particularly, the esterases of this invention have the ability to cleave hemispheric ester compounds, such as γ-ambryl acetate, to form the corresponding hemispheric alcohols, such as γ-ambrol.

[0127] In the context of this invention, "alcohol dehydrogenase" (ADH) refers to a polypeptide that, in the NAD+ process, dehydrogenases... + or NADP + In the presence of a cofactor, it possesses the ability to oxidize alcohols to the corresponding aldehydes. Such enzymes are members of the EC family 1.1.1.1 (NAD... + (dependency) or 1.1.1.2 (NADP) + (dependent) members. More particularly, the ADH of the present invention has the ability to oxidize hemispheric alcohols to the corresponding hemispheric carbonyl compounds (aldehydes or ketones), such as oxidizing coparol to coparaldehyde and / or hemispheric olefinic alcohol to the corresponding aldehyde or other hemispheric alcohol derivatives of coparol and hemispheric olefinic alcohol, such as the corresponding demethylated or dedimethyl hemispheric derivatives of coparol and hemispheric olefinic alcohol. The ADH used herein may be endogenously present in the corresponding biocatalytic process or exogenous.

[0128] The “enal cleavage activity” is determined under the “standard conditions” described below: It can be determined using host cells expressing recombinant enal cleavage peptides, cells expressing disrupted enal cleavage peptides, their fractions, or enriched or purified enal cleavage peptides, under conditions of approximately 20 to 45°C, for example approximately 25 to 40°C, preferably 25 to 32°C, in a culture medium or reaction medium (preferably buffered) at pH 6 to 11, preferably 7 to 9, and in the presence of a reference substrate (here, particularly copalaldehyde), at an initial concentration of 1 to 100 μM, preferably 5 to 50 μM, especially 30 to 40 μM, or generated endogenously by the host cell. The conversion reaction to form the corresponding cleavage product, such as the kinoloxy group, proceeds for 10 minutes to 5 hours, preferably about 1 to 2 hours. The cleavage product can then be determined using conventional methods, such as extraction with an organic solvent like ethyl acetate.

[0129] “BVMO activity” is determined under the “standard conditions” described below: it can be determined using recombinant BVMO-expressing host cells, disrupted BVMO-expressing cells, their fractions, or enriched or purified BVMO enzymes, under conditions of approximately 20 to 45°C, for example approximately 25 to 40°C, preferably 25 to 32°C, in a culture medium or reaction medium (preferably buffered) at a pH of 6 to 11, preferably 7 to 9, and in the presence of a reference substrate (here, particularly copalaldehyde and / or minocycline), at an initial concentration of 1 to 100 μM, preferably 5 to 50 μM, especially 30 to 40 μM, or by endogenous production by the cell host in the presence of molecular oxygen. For in vitro assays, a cofactor selected from NADH and NADPH must be added at a suitable and easily determined concentration range. The conversion reaction, which forms the corresponding cleavage products, such as methyl esters 1a and / or 1b in the case of coparaldehyde, or γ-ambryl acetate in the case of minoyloxy, proceeds for 10 minutes to 5 hours, preferably about 1 to 2 hours. The oxidation products can then be determined by conventional methods, such as extraction with an organic solvent like ethyl acetate.

[0130] Terpenoid diphosphate synthase Activity (such as CPS or LPS activity) was determined under the “standard conditions” described below: It could be determined using recombinant terpenoid diphosphate synthase expression host cells, disrupted terpenoid diphosphate synthase expression cells, their fractions, or enriched or purified terpenoid diphosphate synthase, under conditions of approximately 20 to 45°C, for example approximately 25 to 40°C, preferably 25 to 32°C, in a culture medium or reaction medium (preferably buffered) at a pH of 6 to 11, preferably 7 to 9, and in the presence of a reference substrate (specifically GG). In the case of PP), it is added at an initial concentration of 1 to 100 μM, preferably 5 to 50 μM, especially 30 to 40 μM, or produced endogenously by the cell host. The conversion reaction to form terpenoid diphosphate proceeds for 10 minutes to 5 hours, preferably about 1 to 2 hours. If no endogenous phosphatase is present, one or more exogenous phosphatases, such as alkaline phosphatase, are added to the reaction mixture to convert the terpenoid diphosphate formed by the synthase into the corresponding terpenol. The terpenol can then be determined using conventional methods, such as extraction with an organic solvent such as ethyl acetate.

[0131] Terpenoid diphosphate phosphataseThe activities (such as CPP or LPP phosphatase activities) are determined under the “standard conditions” described below: they can be determined using recombinant terpenoid diphosphate phosphatase expression host cells, disrupted terpenoid diphosphate phosphatase expression cells, their fractions, or enriched or purified terpenoid diphosphate phosphatases, under conditions of approximately 20 to 45°C, for example approximately 25 to 40°C, preferably 25 to 32°C, in a culture medium or reaction medium (preferably buffered) at pH 6 to 11, preferably 7 to 9, and in the presence of a reference substrate (here, for example, CPP or LPP), at an initial concentration of 1 to 100 μM, preferably 5 to 50 μM, especially 30 to 40 μM, or produced endogenously by the cell host. The conversion reaction to form terpenoid diphosphates proceeds for 10 minutes to 5 hours, preferably about 1 to 2 hours. The terpenoid alcohols can then be determined using conventional methods, for example, after extraction with an organic solvent such as ethyl acetate.

[0132] For each of the enzyme activities described above, specific examples of suitable standard conditions can be obtained from the experimental section below.

[0133] The terms “biological function,” “function,” “biological activity,” or “activity” for terpene synthases refer to the ability of the terpene diphosphate synthases described herein to catalyze the formation of at least one terpene diphosphate from the corresponding precursor terpene.

[0134] The terms “biological function,” “function,” “biological activity,” or “activity” for terpenoid diphosphate phosphatases refer to the ability of the terpenoid diphosphate phosphatases described herein to catalyze the removal of diphosphate groups from said terpenoid compounds to form the corresponding terpenoid alcohols.

[0135] The mevalonate pathway, also known as the isoprenoid pathway or HMG-CoA reductase pathway, is an essential metabolic pathway found in eukaryotes, archaea, and some bacteria. The mevalonate pathway begins with acetyl-CoA, producing two five-carbon structural units called isopentenyl pyrophosphate (IPP) and dimethylallyl pyrophosphate (DMAPP). Key enzymes include acetoacetyl-CoA thiolysis enzyme (atoB), HMG-CoA synthase (mvaS), HMG-CoA reductase (mvaA), mevalonate kinase (MvaK1), phosphate mevalonate kinase (MvaK2), mevalonate diphosphate decarboxylase (MvaD), and isopentenyl diphosphate isomerase (idi). The mevalonate pathway is combined with enzyme activities to produce terpene precursors GPP, FPP, or GGPP, such as, in particular, FPP synthase (ERG20), allowing for the recombinant cellular production of terpenes.

[0136] As used herein, the term "host cell" or "transformed cell" refers to a cell (or organism) that has been modified to carry at least one nucleic acid molecule, such as a recombinant gene encoding a desired protein or nucleic acid sequence, which, upon transcription, produces the functional polypeptide of the present invention, namely, the terpenoid diphosphate synthase protein or terpenoid diphosphate phosphatase as defined above. Host cells are particularly bacterial, fungal, or plant cells or plants. Host cells may contain recombinant genes or several genes integrated into the host cell's nuclear or organelle genome, such as genes organized as operons. Alternatively, the host may contain recombinant genes extrachromosomally.

[0137] The term "organism" refers to any non-human multicellular or single-celled organism, such as plants or microorganisms. In particular, microorganisms are bacteria, yeast, algae, or fungi.

[0138] The term "plant" is used interchangeably to include plant cells, including plant protoplasts, plant tissues, plant cell tissue cultures that produce regenerated plants, or parts of a plant, or plant organs such as roots, stems, leaves, flowers, pollen, ovules, embryos, fruits, etc. Any plant may be used to implement the methods described herein.

[0139] When a particular organism or cell naturally produces FPP, or when it does not naturally produce FPP but is transformed with the nucleic acids described herein to produce FPP, it is meant to be "capable of producing FPP". Organisms or cells transformed to produce higher amounts of FPP than naturally occurring organisms or cells are also included in "capable of producing FPP".

[0140] When a particular organism or cell naturally produces GGPP, or when it does not naturally produce GGPP but is transformed to produce GGPP using the nucleic acids described herein, it is meant to be "capable of producing GGPP". Organisms or cells that are transformed to produce higher amounts of GGPP than naturally occurring organisms or cells are also included in "capable of producing GGPP".

[0141] When a particular organism or cell naturally produces terpenoid diphosphate as defined herein, or when it does not naturally produce the diphosphate but is converted to produce the diphosphate using the nucleic acids described herein, it is meant to be “capable of producing terpenoid diphosphate”. Organisms or cells that are converted to produce higher amounts of terpenoid diphosphate than naturally occurring organisms or cells are also covered as “capable of producing terpenoid diphosphate”.

[0142] The term "capable of producing terpenoid alcohols" means that a particular organism or cell naturally produces terpenoid alcohols as defined herein, or that it does not naturally produce the alcohols but is transformed with the nucleic acids described herein to produce them. Organisms or cells transformed to produce higher amounts of terpenoid alcohols than naturally occurring organisms or cells are also included in the term "capable of producing terpenoid alcohols." This also applies to specific organisms "capable of producing hemispheric alcohols."

[0143] "Capable of producing esters" means that a particular organism or cell naturally produces the esters as defined herein, or that it does not naturally produce the esters but is transformed with the nucleic acids described herein to produce the esters. Organisms or cells that are transformed to produce higher amounts of esters than naturally occurring organisms or cells are also covered by "capable of producing esters".

[0144] "Capable of producing the target product" means that a particular organism or cell naturally produces the target product as defined herein (e.g., esters, alcohols, or carbonyl compounds, or more particularly hemispheric compounds) or that it does not naturally produce the target product but is transformed with the nucleic acids described herein to produce the target product. Organisms or cells transformed to produce a higher amount of the target product than naturally occurring organisms or cells are also included in "organisms or cells capable of producing the target product".

[0145] The term “fermentation production” or “fermentation” refers to the ability of microorganisms (assisted by enzyme activity contained in or produced by said microorganisms) to produce compounds in cell cultures using at least one carbon source added to incubation.

[0146] The term “fermentation broth” should be understood as referring to a liquid, particularly an aqueous solution or aqueous / organic solution, that is based on a fermentation process and has been subjected to or has been subjected to post-processing, such as that described herein.

[0147] "Enzymatic catalysis" or "biocatalysis" refers to methods performed under the catalytic action of enzymes (including enzyme mutants as defined herein). Therefore, such methods can be carried out in the presence of the enzyme in isolated (purified, enriched) or crude form, or in the presence of cellular systems, particularly natural or recombinant microbial cells containing the enzyme in its active form and possessing the catalytic transformation capabilities disclosed herein.

[0148] c) Chemical terminology:

[0149] The term "α,β-unsaturated carbonyl" describes compounds containing the general formula R a R b C = C(R) c Organic molecules with an aldehyde or ketone group of α-C=O, wherein the C=C bond can be any stereoisomer and wherein the residue Ra R b and R c They can be the same or different, and for a specific α,β-unsaturated carbonyl compound, they can have the meanings described below.

[0150] The "labdane" compound in the context of this invention will show the following basic structure of its carbon skeleton, which consists of 20 carbon atoms. The depicted carbon atom numbering will be applied to further define certain positions within the carbon skeleton.

[0151]

[0152] The term "hemisperane" encompasses any stereoisomer of this basic C 20 - Any compound with the structure, and encompasses any variant of such structure, which contains one or more unsaturated C-C bonds, particularly one or more C=C bonds, located at any position on the carbide ring and / or side chain. Also encompasses variants thereof containing one or more substituents on any specified primary, secondary, or tertiary C atom, such substituents selected from -OH, =O, -O-CO_R (where R can be a straight-chain or branched alkyl group, particularly a lower alkyl group, more particularly a C1-C4 alkyl group, such as methyl, ethyl, n- or isopropyl, or n-, iso- or tert-butyl), and -COOH.

[0153] Such "hemisperane" derivatives encompass chemical compounds in which the basic C 20 - The carbon skeleton is modified by eliminating one or more carbon atoms. As an example, one could mention:

[0154] Demethylated (decarbonized) hemispherium (C 19 -Skeleton), dedimethyl(de-dicarbon) monsoon (C 18 -Skeleton), detrimethyl(detricarbon) monsoon (C 17 -Skeleton) and detetramethyl(detetracarbon) monsoon (C 16 -Skeleton). The position of the eliminated carbon atom is indicated by specifying the number of carbon atoms. For example, in demethyl monsine, the case where the carbon atom at position 15 is missing is called "15-demethyl monsine".

[0155] Such "hemisperane" derivatives also encompass chemical compounds in which the basic C 20 - The carbon skeleton is modified by inserting a heteroatom between two C- atoms in the monsoon skeleton. For example, inserting an ether bridge between positions 14 and 15 will convert monsoon to demethyl monsoon, and in particular to demethyl monsoon ester.

[0156] Non-limiting examples of substituted hemispherium or substituted hemispherium derivatives are as follows:

[0157]

[0158] As used in this article, "bisphosphate (ester)" and "pyrophosphate (ester)" are synonyms.

[0159] Terpenes are a large and diverse class of organic compounds produced by a variety of plants, especially conifers, and some insects. Terpenes are hydrocarbons. Although sometimes used interchangeably with "terpenes," "terpenoids" or "isoprene-like compounds" are modified terpenes because they contain additional functional groups, usually oxygen-containing ones.

[0160] "Terpenoids" ("isoprene-like compounds") are a large and diverse class of naturally occurring organic chemical substances derived from terpenes. While sometimes used interchangeably with the term "terpene," "terpenoids" contain an additional functional group, typically an oxygen-containing group such as a hydroxyl, carbonyl, or carboxyl group. Most are polycyclic structures with oxygen-containing functional groups. Unless otherwise stated, the terms "terpene" and "terpenoids" are used interchangeably in the context of this specification.

[0161] Terpenes (and terpenoid compounds) can be classified according to the number of isoprene units in their molecules; the prefix in the name indicates the number of terpene units required to form the molecule. Hemiterpenes consist of a single isoprene unit. Monoterpenes consist of two isoprene units and have the molecular formula C2. 10 H 16 Sesquiterpenes are composed of three isoprene units and have the molecular formula C0. 15 H 24 Diterpenes are composed of four isoprene units and have the molecular formula C0. 20 H 32 .

[0162] "Terpenyl" refers to a non-cyclic and cyclic chemical hydrocarbon residue derived from the C5 structural unit isoprene, and specifically contains one or more such structural units.

[0163] "Cyclic terpenes" or "cyclic terpenoids" or "cyclic diterpenes" or "cyclic diterpenoids" refer to terpenoid compounds or terpenoid residues whose structure contains, for example, at least 1, 2, 3, 4 or 5 fused and / or non-fused carbon rings, preferably two fused carbon rings.

[0164] "Bicyclic terpene" or "bicyclic terpenoid" or "bicyclic diterpene" or "bicyclic diterpenoid" refers to a terpene compound or terpenoid residue whose structure contains two carbon rings, preferably two fused carbon rings.

[0165] In the context of this invention, "terpene derivative" or "terpene compound derivative" specifically refers to such compounds obtained from terpenes or terpenoids through chemical and / or enzymatic modification. More specifically, such derivatives encompass "hydrocarbon chain degradation" derivatives.

[0166] The difference between "hydrocarbon chain degradation" terpenes or terpenoids and their undegraded precursors is that the carbon number in the precursor's carbon skeleton is reduced.

[0167] A "hydrocarbon" residue is a chemical group consisting essentially of carbon and hydrogen atoms, and can be an acyclic, straight-chain or branched saturated or unsaturated portion (fragment), or a cyclic saturated or unsaturated portion, aromatic or non-aromatic. In the case of an acyclic structure, the hydrocarbon residue contains 1 to 30, 1 to 25, 1 to 20, 1 to 15, 1 to 10, or 1 to 5 carbon atoms. In the case of a cyclic structure, it contains 4 to 30, 4 to 25, 4 to 20, 4 to 15, 4 to 10, or especially 4, 5, 6, or 7 carbon atoms.

[0168] The hydrocarbon residues may be unsubstituted or may have at least one, for example, 1 to 5, preferably 0, 1 or 2 substituents.

[0169] Specific examples of such hydrocarbon residues are noncyclic straight-chain or branched alkyl or alkenyl residues as defined below; or monocyclic or polycyclic, especially monocyclic or bicyclic, saturated or unsaturated nonaromatic moieties, such as those found in cyclic (e.g. bicyclic) or noncyclic terpenoids and hemispheric alkyl compounds as defined herein.

[0170] "alkyl" residues represent straight-chain or branched saturated hydrocarbon residues. They contain 1 to 30, 1 to 25, 1 to 20, 1 to 15, 1 to 10, 1 to 7, 1 to 6, 1 to 5, or 1 to 4 carbon atoms.

[0171] "Alkenyl" residues represent straight-chain or branched, monounsaturated or polyunsaturated hydrocarbon residues. They contain 2 to 30, 2 to 25, 2 to 20, 2 to 15, 2 to 10, 2 to 7, 2 to 6, 2 to 5, or 2 to 4 carbon atoms. They may have up to 10, for example, 1, 2, 3, 4, or 5 C=C double bonds.

[0172] The term "lower alkyl" or "short-chain alkyl" refers to a saturated, straight-chain or branched hydrocarbon group having 1 to 4, 1 to 5, 1 to 6, or 1 to 7, particularly 1 to 4, carbon atoms. Examples that may be mentioned include: methyl, ethyl, n-propyl, 1-methylethyl, n-butyl, 1-methylpropyl, 2-methylpropyl, 1,1-dimethylethyl, n-pentyl, 1-methylbutyl, 2-methylbutyl, 3-methylbutyl, 2,2-dimethylpropyl, 1-ethylpropyl, n-hexyl, 1,1-dimethylpropyl, 1,2-dimethylpropyl, 1-methylpentyl, 2-methylpentyl, 3-methylpentyl, 4-methylpentyl 1,1-dimethylbutyl, 1,2-dimethylbutyl, 1,3-dimethylbutyl, 2,2-dimethylbutyl, 2,3-dimethylbutyl, 3,3-dimethylbutyl, 1-ethylbutyl, 2-ethylbutyl, 1,1,2-trimethylpropyl, 1,2,2-trimethylpropyl, 1-ethyl-1-methylpropyl and 1-ethyl-2-methylpropyl; and n-heptyl, and their single- or multi-branched analogues.

[0173] "Long-chain alkyl" refers to, for example, saturated straight-chain or branched hydrocarbon groups having 8 to 30, such as 8 to 20 or 8 to 15 carbon atoms, such as octyl, nonyl, decyl, undecyl, dodecyl, tridecyl, tetradecyl, pentadecyl, hexadecyl, heptadecanyl, octadecyl, nonadecanyl, eicosyl, hencosyl, dodecyl, tridecyl, tetradecyl, pentadecyl, hexadecyl, heptadecanyl, octadecyl, nonadecanyl, squalyl, structural isomers, especially their single- or multi-branched isomers.

[0174] "Long-chain alkenyl" refers to monounsaturated or polyunsaturated analogs of the aforementioned "long-chain alkyl" groups.

[0175] "Short-chain alkenyl" (or "lower alkenyl") refers to a monounsaturated or polyunsaturated, especially monounsaturated, straight-chain or branched hydrocarbon group having 2 to 4, 2 to 6, or 2 to 7 carbon atoms and a double bond at any position, such as C2-C6 alkenyl, e.g., vinyl, 1-propenyl, 2-propenyl, 1-methylvinyl, 1-butenyl, 2-butenyl, 3-butenyl, 1-methyl-1-propenyl, 2-methyl-1-propenyl, 1-methyl-2-propenyl, 2-methyl-2-propenyl, 1-pentenyl, 2-pentenyl, 3-pentenyl, 4-pentenyl, 1-methyl-1-butenyl, 2-methyl-1-butenyl, 3- Methyl-1-butenyl, 1-methyl-2-butenyl, 2-methyl-2-butenyl, 3-methyl-2-butenyl, 1-methyl-3-butenyl, 2-methyl-3-butenyl, 3-methyl-3-butenyl, 1,1-dimethyl-2-propenyl, 1,2-dimethyl-1-propenyl, 1,2-dimethyl-2-propenyl, 1-ethyl-1-propenyl, 1-ethyl-2-propenyl, 1-hexenyl, 2-hexenyl, 3-hexenyl, 4-hexenyl, 5-hexenyl, 1-methyl-1-pentenyl, 2-methyl-1-pentenyl, 3-methyl-1-pentenyl, 4-methyl-1-pentenyl, 1-methyl-2-pentenyl 2-Methyl-2-pentenyl, 3-methyl-2-pentenyl, 4-methyl-2-pentenyl, 1-methyl-3-pentenyl, 2-methyl-3-pentenyl, 3-methyl-3-pentenyl, 4-methyl-3-pentenyl, 1-methyl-4-pentenyl, 2-methyl-4-pentenyl, 3-methyl-4-pentenyl, 4-methyl-4-pentenyl, 1,1-dimethyl-2-butenyl, 1,1-dimethyl-3-butenyl, 1,2-dimethyl-1-butenyl, 1,2-dimethyl-2-butenyl, 1,2-dimethyl-3-butenyl, 1,3-dimethyl-1-butenyl, 1,3-dimethyl-2-butenyl, 1,3- Dimethyl-3-butenyl, 2,2-dimethyl-3-butenyl, 2,3-dimethyl-1-butenyl, 2,3-dimethyl-2-butenyl, 2,3-dimethyl-3-butenyl, 3,3-dimethyl-1-butenyl, 3,3-dimethyl-2-butenyl, 1-ethyl-1-butenyl, 1-ethyl-2-butenyl, 1-ethyl-3-butenyl, 2-ethyl-1-butenyl, 2-ethyl-2-butenyl, 2-ethyl-3-butenyl, 1,1,2-trimethyl-2-propenyl, 1-ethyl-1-methyl-2-propenyl, 1-ethyl-2-methyl-1-propenyl and 1-ethyl-2-methyl-2-propenyl.

[0176] "alkylene" represents a straight-chain, single-branched, or multi-branched hydrocarbon bridging group having 1 to 10 carbon atoms, such as those selected from -CH2-, -(CH2)2-, -(CH2)3-, -(CH2)4-, -(CH2)2-CH(CH3)-, -CH2-CH(CH3)-CH2-, -(CH2)4-, -(CH2)5-, -(CH2)6-, -(CH2)7-, -(CH2)8-, -(CH2)9-, -(CH2)10 ... 2) C1-C7 alkylene groups of 7-, -CH(CH3)-CH2-CH2-CH(CH3)- or -CH(CH3)-CH2-CH2-CH2-CH(CH3)-, particularly C1-C4 alkylene groups selected from -CH2-, -(CH2)2-, -(CH2)3-, -(CH2)4-, -(CH2)2-CH(CH3)-, and -CH2-CH(CH3)-CH2-.

[0177] The "alkylidene" group represents a straight-chain or branched hydrocarbon substituent attached to the main body of the molecule via a double bond. It contains 1 to 6 carbon atoms. Examples of such "C1-C6 alkylidenes" include the methylidene (=CH2), ethylidene (=CH-CH2), n-propylidene, n-butylidene, n-pentanene, n-hexanene, and their structural isomers such as the isopropylidene.

[0178] "Alkenylidene" refers to a monounsaturated analog of the above-mentioned alkane group having more than two carbon atoms, and can be called "C3-C6 alkenylidene". Examples include n-propenyl, n-butenyl, n-pentenyl, and n-hexenyl.

[0179] The "substituent" of the above residues contains a heteroatom, such as O or N. Preferably, the substituent is independently selected from -OH, C=O, or -COOH. Most preferably, the substituent is -OH.

[0180] "Monocyclic or polycyclic hydrocarbon residues" comprise one, two, or three fused (anellated) or non-fused, optionally substituent, saturated or unsaturated hydrocarbon cyclic groups (or "carbocyclic" groups). Each ring may independently comprise three to eight, particularly five to seven, and more particularly six cyclic carbon atoms. Examples of monocyclic residues include "cycloalkyl" groups having three to seven cyclic carbon atoms, such as cyclopropyl, cyclobutyl, cyclopentyl, cyclohexyl, cycloheptyl, and cyclooctyl; and corresponding "cycloalkenyl" groups. "Cycloalkenyl" (or "mono- or polyunsaturated cycloalkyl") specifically represents monocyclic or polyunsaturated carbocyclic groups having five to eight, preferably up to six, cyclic members, such as monounsaturated cyclopentenyl, cyclohexenyl, cycloheptenyl, and cyclooctenyl.

[0181] As examples of polycyclic residues, groups that can be mentioned are those in which one, two, or three such cycloalkyl and / or cycloalkenyl groups are linked together, for example, by fused cyclization, to form a polycyclic cycloalkyl or cycloalkenyl ring. As a non-limiting example, a bicyclic decalinyl residue consisting of two fused cyclized six-membered carbon rings can be mentioned.

[0182] The number of substituents in such monocyclic or polycyclic hydrocarbon residues can be from 1 to 10, particularly 1 to 5. Suitable substituents for such cyclic residues are selected from lower alkyl groups, lower alkenyl groups, alkylidenes, alkenylidenes, or residues containing a heteroatom such as O or N, for example -OH or -COOH. In particular, the substituents are independently selected from -OH, -COOH, methyl, and methylidene.

[0183] The unsaturated cyclic group may contain one or more, such as one, two or three C=C bonds, and is aromatic or particularly non-aromatic.

[0184] The monocyclic or polycyclic saturated or unsaturated groups may also contain at least one, for example, 1, 2, 3 or 4 cyclic heteroatoms, such as O, N or S.

[0185] Overview of specific compound names and their structural formulas

[0186]

[0187]

[0188]

[0189]

[0190]

[0191]

[0192]

[0193]

[0194]

[0195]

[0196]

[0197]

[0198]

[0199]

[0200]

[0201] Detailed description

[0202] a. Specific embodiments of the present invention

[0203] i) This invention relates to specific embodiments of a biocatalytic method using peptides with BMVO activity:

[0204] 1. A biocatalytic method for preparing ester compounds, comprising:

[0205] (1) Contacting a carbonyl precursor compound of general formula I with a natural or recombinant polypeptide having Bayer-Villiger monooxygenase (BVMO) (EC1.13.14.-) activity, particularly by introducing an oxygen atom between the carbonyl group and the α-carbon atom of the precursor, to form the corresponding carbonyl ester product.

[0206]

[0207] in

[0208] “a” indicates a single or double bond.

[0209] If "a" represents a double bond, then "x" is the integer 1; if "a" represents a single bond, then "x" is the integer 2.

[0210] R 1 Each can independently represent H or lower alkyl groups, such as C1-C4 alkyl groups, especially H or methyl.

[0211] R 2 Representing H, saturated or unsaturated optional substitutions, particularly hydrocarbon residues having 2 to 20 carbon atoms, more particularly 5 to 15 carbon atoms, or the group Cyc-A-.

[0212] in

[0213] Cyc represents an optional substituted saturated or unsaturated monocyclic or polycyclic hydrocarbon residue, and

[0214] A represents a straight-chain or branched alkylene bridge with a chemical bond or optional substitution, particularly a methylene group.

[0215] R 3 Each represents H or Cl-C independently. 30 C1-C 20 Or especially C1-C 15Hydrocarbon group, or lower alkyl group, such as C1-C4 alkyl group, especially H or methyl, and more particularly each of them is H.

[0216] and

[0217] When "a" represents a single bond, Z represents a hydrocarbon residue containing a carbonyl group, particularly an aldehyde or ketone group, or...

[0218] When "a" represents a double bond, Z together with the carbon atom it is attached to form a carbonyl group (C=O, especially an aldehyde or ketone group, or an alkyl subunit residue with a terminal carbonyl group, especially an aldehyde or ketone group, particularly a C1-C6 alkyl subunit residue).

[0219] When "a" represents a double bond, and Z forms a carbonyl group (C=O) with the carbon atom it is attached to, then R 2 and R 1 Together with the carbon atoms to which they are attached, they can also form cyclic, especially monocyclic, saturated or unsaturated, optionally substituted carbocyclic groups, particularly 5- to 7-membered rings;

[0220] The carbonyl compound of general formula I is provided either in stereoisomeric pure form or as a mixture of stereoisomers;

[0221] (2) and optionally, the carbonyl ester formed in step (1) is separated, wherein the carbonyl ester compound is obtained as a stereoisomer or as a mixture of stereoisomers.

[0222] 2. The biocatalytic method of embodiment 1, wherein in the carbonyl compound of general formula I,

[0223] “a” represents a chemical double bond, and Z represents =O (see Minooloxy group) or =C(R) 4 )-C(R 5 ) = O (see copalaldehyde); or

[0224] “a” represents a chemical single bond, and Z represents -C(R) 5 ) = O (see demethylated hemispherane compounds 3a, 3b);

[0225] in

[0226] R 4 and R 5 Each can represent H or lower alkyl groups, such as C1-C4 alkyl groups, especially H or methyl groups, independently of each other.

[0227] 3. The biocatalytic method of any of the foregoing embodiments, wherein the carbonyl compound of general formula I has a semaphore-type structure, particularly semaphore, demethyl semaphore, or dedimethyl semaphore structure.

[0228] 4. The biocatalytic method of any of the foregoing embodiments, wherein the carbonyl ester formed satisfies Formula II:

[0229]

[0230] in

[0231] R 2 and R 3 As defined above, and

[0232] E represents a hydrocarbon residue containing the carbonyl ester group, or wherein E and R 2 Together with the carbon atoms they are attached to, they form cyclic ester groups.

[0233] 5. The biocatalytic method of embodiment 4, wherein the carbonyl ester group E is selected from:

[0234] -OC(O)-R 1 ,

[0235] -C(R 1 )2-OC(O)R 5 ,

[0236] -C(R 1 )=C(R 4 )-OC(R 5 ) = O; and

[0237] By E and R 2 The cyclic ester groups formed together with the carbon atoms they are attached to, wherein the cyclic ester ring represents a 5- to 7-membered ring, especially a 6-membered ring, such as the rings in the esters of formulas IIa and IIb:

[0238]

[0239] Where R 1 R 3 R 4 and R 5 As defined above.

[0240] 6. The biocatalytic method of any one of the foregoing embodiments, wherein...

[0241] R 2The representative group is Cyc-A-, where A represents a straight-chain or branched C1-C4 alkylene bridge, particularly methylene, and Cyc represents a monocyclic or polycyclic, particularly bicyclic, saturated or unsaturated hydrocarbon residue, particularly bicyclic fused-ring hydrocarbon residues comprising 5 to 7, particularly 6, ring atoms per ring; wherein Cyc is optionally substituted with 1 to 10, particularly 1 to 5, substituents, wherein the substituents may be independently selected from C1-C4 alkyl, C1-C4 alkylene, C3-C6 alkenyl, C2-C4 alkenyl, oxo(=O), hydroxyl or amino; particularly C1-C4 alkyl such as methyl, and C1-C4 alkyl such as methylene.

[0242] 7. The method of any of the foregoing implementation schemes, wherein R 2 Cyc residues can be substituted with decahydronaphthyl residues, particularly bicyclic residues that can be obtained, for example, through terpene cyclization.

[0243] 8. The method of embodiment 7, wherein Cyc-A represents a bicyclic residue of formula IIIa, IIIb, or IIIc having 15 carbon atoms:

[0244]

[0245] 9. The method of any of the foregoing embodiments, wherein the polypeptide having BVMO activity is selected from:

[0246] (1) A group of polypeptides containing, in their amino acid sequence, a domain of the flavin monooxygenase (FMO) protein family with Pfam ID number PF00743, or a domain that has at least 90%, 95%, 96%, 97%, 98% or 99% sequence identity with PF00743.

[0247] Specifically, if the e-value of the BVMO-active polypeptide of the present invention that matches the said domain is less than 1 x 10-1 -5 or less than 1x10 -10 or less than or equal to 1x10 -15 or less than or equal to 1x10 -18 Especially in 1x10 -10 Up to 1x10 -18 Within the range, especially in 1x10 -14 Up to 1x10 -17 If the range is within the specified range, it is identified as a member of the FMO protein family containing the PF00743 domain.

[0248] Request the sequence of a peptide with BVMO activity as the query sequence.

[0249] For example, the following websites can be requested to retrieve and calculate such e values:

[0250] http: / / pfam.xfam.org /

[0251] http: / / www.ebi.ac.uk / Tools / hmmer / search / hmmscan or

[0252] http: / / www.ebi.ac.uk / Tools / pfa / pfamscan / .

[0253] and / or

[0254] (2) A group of polypeptides comprising at least 1, 2, 3, 4, 5, 6, 7 or all of the following sequence motifs / domains:

[0255] GAGxSGL as shown in SEQ ID NO:197;

[0256] The EKNxxxxGTWxENRYPGCACDVPxHxYXXSFE shown in SEQ ID NO:198

[0257] Or any partial motif containing up to 15, up to 10, or up to 5 consecutive amino acid residues, such as residues at positions 1–10, 11–20, or 21–32 corresponding to SEQ ID NO:198;

[0258] The LxNAxGILNxWxxPxIPG shown in SEQ ID NO:199

[0259] Or any partial motif containing up to 15, up to 10, or up to 5 consecutive amino acid residues, such as residues at positions 1–10 or 11–18 corresponding to SEQ ID NO:199;

[0260] The LxxKxVxxIGxGSSGIQIxPxI shown in SEQ ID NO:200

[0261] Or any partial motif containing up to 15, up to 10, or up to 5 consecutive amino acid residues, such as residues corresponding to positions 1–10 or 11–18 of SEQ ID NO:200;

[0262] The GCRRxTPGxxYLExL shown in SEQ ID NO:201

[0263] Or any partial motif containing up to 15, up to 10, or up to 5 consecutive amino acid residues, such as residues at positions 1–10 and 11–15 corresponding to SEQ ID NO:201;

[0264] The CAGFDxxxxPRFxxxG shown in SEQ ID NO:202

[0265] Or any partial motif containing up to 15, up to 10, or up to 5 consecutive amino acid residues, such as residues at positions 1–10 or 11–17 corresponding to SEQ ID NO:202;

[0266] The PNxFxxxGPNxPxxNGxV shown in SEQ ID NO:203

[0267] Or any partial motif containing up to 15, up to 10, or up to 5 consecutive amino acid residues, such as residues at positions 1–10 or 11–18 corresponding to SEQ ID NO:203;

[0268] The AxWPGSxLHYxEAxxxPRxED shown in SEQ ID NO:204

[0269] Or any partial motif containing up to 15, up to 10, or up to 5 consecutive amino acid residues, such as residues at positions 1–10 or 11–21 corresponding to SEQ ID NO:204;

[0270] in

[0271] In the above motifs, residue x represents any native amino acid residue independently of each other, and optionally in each of the above motifs, 1 to 5, such as 1, 2, 3, 4 or 5, conserved amino acid residues (i.e., different from residue x) can be modified, for example by amino acid substitution, especially by conserved substitution, provided that the enzyme retains BVMO enzyme activity to at least an analytically detectable extent.

[0272] and / or

[0273] (3) A group of polypeptides consisting of the following:

[0274] (a) A polypeptide comprising the SCH23-BVMO1 amino acid sequence shown in SEQ ID NO:2;

[0275] (b) A polypeptide comprising the SCH24-BVMO1 amino acid sequence shown in SEQ ID NO:6;

[0276] (c) A polypeptide comprising the SCH25-BVMO1 amino acid sequence shown in SEQ ID NO:10;

[0277] (d) A polypeptide containing the SCH46-BVMO1 amino acid sequence shown in SEQ ID NO:13;

[0278] (e) A polypeptide comprising the AspWeBVMO amino acid sequence shown in SEQ ID NO:16 (preferred substrates: minoyloxy group and its isomers);

[0279] (f) A polypeptide comprising an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity with any of the amino acid sequences in (a) through (e).

[0280] Of the five specific BVMO peptides mentioned above, the protein family domain with Pfam ID number PF00743 can be located at the amino acid residue positions given in the table below (see also the alignments and boxed sequence portions depicted in Figure 32).

[0281] protein sequence from to E value Registration ID Protein domains SCH23-BVMO1 23 388 2.9e-16 Pf00743 flavin-binding monooxygenase-like SCH24-BVMO2 67 283 6.8e-15 Pf00743 flavin-binding monooxygenase-like SCH25-BVMO1 23 246 1.2e-15 Pf00743 flavin-binding monooxygenase-like SCH46-BVMO1 23 388 1.8e-16 Pf00743 flavin-binding monooxygenase-like AspWe BVMO 20 249 1.7e-16 Pf00743 flavin-binding monooxygenase-like

[0282] The amino acid residue number refers to the residue number in the corresponding SEQ ID NO of the corresponding protein sequence in the attached sequence listing.

[0283] Another specific embodiment refers to a polypeptide variant of the novel polypeptide of the present invention having BVMO activity as identified by any specific amino acid sequence of SEQ ID NO:2, 6, 10 and 13 above, wherein the polypeptide variant is selected from an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98% or 99% sequence identity with any one of SEQ ID NO:2, 6, 10, 13 and 16, and contains at least one substitution modification relative to any one of the unmodified SEQ ID NO:2, 6, 10, 13 and 16.

[0284] 10. The method of any of the foregoing embodiments, which is performed in vitro or in test tube.

[0285] 11. The method of embodiment 10 performed in vivo, comprising, prior to step (1), particularly in non-human host cells, recombinantly expressing one or more polypeptides having the enzyme activity required to perform the BVMO catalytic enzymatic step.

[0286] 12. The method of implementation scheme 11, wherein a non-human host cell is transformed with a nucleic acid encoding at least one polypeptide having BVMO activity.

[0287] 13. The method of implementation scheme 11 or 12, wherein the non-human host cell is a eukaryotic or prokaryotic cell, particularly a plant cell, bacterial or fungal cell, particularly a yeast cell.

[0288] 14. The method of any one of embodiments 11 to 13, wherein the non-human host cell is a single-celled organism, a cultured cell derived from a multicellular organism, a cell present in a cultured tissue derived from a multicellular organism, or a cell present in a living multicellular organism.

[0289] 15. The method of any one of embodiments 10 to 13, wherein the non-human host cell is a bacterium of the genus Escherichia, particularly Escherichia coli, and the yeast is a yeast of the genus Saccharomyces or Pichia, particularly Saccharomyces cerevisiae, or a plant cell.

[0290] 16. The method of one of the aforementioned embodiments, wherein the carbonyl compound of general formula I is a hemispherane-type compound selected from:

[0291] a) Hemicarbazinaldehyde, particularly coparaldehyde (or any stereoisomerically different form thereof, such as comprising cis or trans forms, or a mixture of cis and trans forms), which is converted by said BVMO into the corresponding demethyl hemicarbazin carbamate, particularly formic acid (5S,9S,10S)-15-demethyl hemicarbazin-8(20),13-dien-14-yl ester or any stereoisomerically different form thereof;

[0292] b) Desdimethylambryl ketone, particularly minoyloxy or any stereoisomerically different form thereof, which is converted by said BVMO to γ-ambryl acetate or any stereoisomerically different form thereof; or

[0293] c) C1-degradation analogs of demethylated hemispheral aldehyde, particularly copal aldehyde, or any stereoisomerically different forms thereof, especially those of the following formula.

[0294]

[0295] Or any stereoisomerically different form thereof, which is converted by said BVMO into the corresponding desdimethylhexane carboxylate, particularly the following formula

[0296]

[0297] Or any of its different stereotypical forms,

[0298] Optionally, the obtained product can be separated into a stereoisomer-based pure form or as a mixture of stereoisomers.

[0299] Other specific examples of the BVMO-catalyzed conversion of carbonyl compounds to corresponding esters are summarized in the following illustrative overview:

[0300]

[0301]

[0302] The parameter “n” is an integer from 1 to 20, 1 to 15, 1 to 10, or 1, 2, 3, 4, or 5.

[0303] 17. The method of embodiment 16a, comprising, prior to step (1), the biocatalytic oxidation of serotonol to serotonaldehyde, particularly the biocatalytic oxidation of coparol to coparaldehyde,

[0304] The monazine alcohol is optionally formed by biocatalytic conversion of at least one terpenoid diphosphate precursor selected from IPP, DMAPP, FPP and GGPP, particularly by a single step or a combination of at least two steps known in the art.

[0305] The monsoon tanols can be produced, for example, via biocatalysis, in the following manner:

[0306] a) Produced from geraniol geraniol diphosphate (GGPP) in one step via cyclization / dephosphorylation;

[0307] b) GGPP is cyclized in two steps to form monazine diphosphate, such as cobazine diphosphate (CPP), and then dephosphorylated to monazine alcohol;

[0308] c) From IPP and DMAPP, which are directly converted to hemispheric diphosphate, such as CPP, by the action of bifunctional GGPP synthase / CPP synthase, and then dephosphorylated;

[0309] The GGPP used in these steps can also be provided by different biocatalytic steps:

[0310] d) GGPP can be generated directly from IPP and DMAPP using GGPP synthase; or

[0311] e) GGPP can be provided by the action of FPP synthase from IPP and DMAPP via FPP, and then by the action of GGPP synthase to convert FPP into GGPP.

[0312] 18. The method of implementing Implementation 17, wherein

[0313] The biocatalytic oxidation of hemispheric alcohols, particularly the biocatalytic oxidation of coparol to coparaldehyde, is catalyzed by exogenous or endogenous polypeptides with alcohol dehydrogenase (ADH) (EC 1.1.1.-) activity; and / or

[0314] The biocatalytic formation of hemispheric alcohols includes at least one step selected from the following:

[0315] i) Biocatalytic dephosphorylation of monazine diphosphate to monazine alcohol, particularly copaclidium diphosphate (CPP) to copacliol, catalyzed by peptides with terpenoid diphosphate (TPP) phosphatase activity, and / or

[0316] ii) Biocatalytic cyclization of terpene diphosphate precursors, such as geraniylgeraniyl diphosphate (GGPP), to CPP, catalyzed by peptides with CPP synthase activity, such as SmCPS2 (SEQ ID NO: 185); or, for example, biocatalytic cyclization of IPP and DMAPP to CPP, catalyzed by bifunctional peptides with isoprene transferase and copacpyl diphosphate synthase activities, such as PvCPS, and / or

[0317] iii) Biocatalytic formation from FPP to GGPP, or from IPP and DMAPP, each catalyzed by a polypeptide with GGPP synthase activity.

[0318] 19. The method of implementing Implementation 18, wherein

[0319] The biocatalytic oxidation, particularly the biocatalytic oxidation of coparol to coparaldehyde, is catalyzed by a polypeptide with alcohol dehydrogenase (ADH) activity, the polypeptide being selected from...

[0320] a) A polypeptide comprising the SCH23-ADH1_wt amino acid sequence shown in SEQ ID NO:134,

[0321] b) A polypeptide comprising the SCH24-ADH1_wt amino acid sequence shown in SEQ ID NO:140,

[0322] c) A polypeptide comprising the SCH94-3945_wt amino acid sequence shown in SEQ ID NO:161,

[0323] d) A polypeptide comprising the SCH80-0540_wt amino acid sequence shown in SEQ ID NO:164,

[0324] e) A polypeptide comprising the AzTolADH1_wt amino acid sequence shown in SEQ ID NO:167,

[0325] f) A polypeptide comprising the CdGeoA_wt amino acid sequence shown in SEQ ID NO:179.

[0326] g) A polypeptide comprising an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity with any of the amino acid sequences in a) to f) and having ADH activity.

[0327] and / or

[0328] The biocatalytic dephosphorylation, particularly the biocatalytic dephosphorylation from copaclidium diphosphate (CPP) to copacliol, is catalyzed by a polypeptide with terpene diphosphate (TPP) phosphatase activity, wherein the polypeptide is selected from:

[0329] a) A polypeptide comprising the AspWE TPP amino acid sequence shown in SEQ ID NO:170 or a polypeptide comprising an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity with it;

[0330] b) A polypeptide comprising the TalCeTPP amino acid sequence shown in SEQ ID NO:176, or a polypeptide comprising an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity with it; and

[0331] c) A polypeptide comprising the TalVeTPP amino acid sequence shown in SEQ ID NO:194 or a polypeptide comprising an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity with it;

[0332] Other suitable phosphatases are also disclosed in the applicant’s earlier EP application 18182783.3, which is incorporated herein by reference.

[0333] and / or

[0334] The biocatalytic cyclization, particularly the biocatalytic cyclization of gerany-gerany-diphosphate (GGPP) to CPP, is catalyzed by peptides selected from the following:

[0335] A polypeptide containing the SmCPS2 amino acid sequence shown in SEQ ID NO:185 and having cobazin activity, or a polypeptide containing an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity with it;

[0336] The biocatalytic cyclization, particularly the biocatalytic cyclization of IPP and DNMAPP to CPP, is catalyzed by peptides selected from the following:

[0337] A polypeptide containing the PvCPS amino acid sequence shown in SEQ ID NO:173 and having isoprenyl transferase and cobazide diphosphate synthase activities, or a polypeptide containing an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity with it.

[0338] and / or

[0339] The biocatalytic formation of GGPP is catalyzed by peptides having GGPP synthase activity and selected from the following:

[0340] a) A polypeptide comprising the carG amino acid sequence shown in SEQ ID NO:182 or a polypeptide comprising an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity with it;

[0341] b) A polypeptide comprising the CrtE amino acid sequence shown in SEQ ID NO:191 or a polypeptide comprising an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity with it;

[0342] c) A polypeptide comprising the PvCPS amino acid sequence shown in SEQ ID NO:173 or a polypeptide comprising an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity with it.

[0343] 20. The method of any of the foregoing embodiments further comprises the following step as step (3): treating the carbonyl ester formed in step (1) or separated in step (2) with chemical or biocatalytic synthesis or a combination thereof to obtain its derivative, wherein the derivative may be particularly selected from hydrocarbons, alcohols, diols, triols, acetals, ketals, aldehydes, acids, ethers, amides, ketones, lactones, epoxides, acetates, glycosides and / or esters, and optionally, separating the derivative of step (3).

[0344] 21. The method of embodiment 20, wherein step (3) comprises hydrolyzing a carbonyl ester compound into a corresponding deesterified product (which may be an alcohol or its isomerized product) using an esterase-active EC 3.1.1 (carboxylate hydrolase), and optionally, separating the derivative of step (3).

[0345] 22. The method of embodiment 21, wherein in a further step (4), the deesterification product of step (3) is subjected to an enzymatic redox reaction, wherein in particular the redox reaction comprises the oxidation of the alcohol group formed in step (3) to the corresponding ketone group by the enzymatic action of an exogenous or endogenous alcohol dehydrogenase (ADH) (EC 1.1.1.-).

[0346] 23. The method of embodiment 21, wherein the esterase is selected from the group consisting of:

[0347] a) A polypeptide containing the SCH23-esterase amino acid sequence shown in SEQ ID NO:20;

[0348] b) A polypeptide containing the SCH24-esterase amino acid sequence shown in SEQ ID NO:24;

[0349] c) A polypeptide containing the amino acid sequence of the SCH25-esterase shown in SEQ ID NO:28;

[0350] d) A polypeptide containing the SCH46-esterase amino acid sequence shown in SEQ ID NO:31; or

[0351] e) A polypeptide comprising an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity with any of the amino acid sequences in a) to d) and having esterase activity.

[0352] 24. The method of implementing Implementation 23, wherein

[0353] a) Demethyl serotonide, particularly demethyl serotonide carboxylate, is deesterified by the esterase to form demethyl serotonide carbonyl compounds, particularly carbonyl compounds of the following formula.

[0354]

[0355] Or its corresponding enol, which is converted into the carbonyl compound by isomerization;

[0356] or

[0357] b) Detetramethylambryl ester, particularly γ-ambryl acetate, is deesterified by the esterase to detetramethylambryl, particularly γ-ambrol; or

[0358] c) Desdimethyl monazine carbamate, especially the carbamate of the following formula

[0359]

[0360] Or any stereoisomerically different form thereof is deesterified by the esterase to the corresponding dedimethyl monsoon alcohol, particularly the alcohol compound of the following formula.

[0361]

[0362] Or any of its different stereotypical forms,

[0363] Optionally, the obtained product can be separated in a stereoisomer-pure form or as a mixture of stereoisomers.

[0364] 25. The method of implementation scheme 22, wherein the ADH is selected from the group consisting of:

[0365] a) A polypeptide comprising the SCH23-ADH2_wt amino acid sequence shown in SEQ ID NO:137,

[0366] b) A polypeptide comprising the SCH24-ADH2_wt amino acid sequence shown in SEQ ID NO:143,

[0367] c) A polypeptide comprising the RrhSecADH_wt amino acid sequence shown in SEQ ID NO:146,

[0368] d) A polypeptide comprising the amino acid sequence SCH80-06135_wt shown in SEQ ID NO:155,

[0369] e) A polypeptide containing an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity with any of the amino acid sequences in a) to d) and having ADH activity.

[0370] 26. The method of embodiment 24, wherein the obtained desdimethylsemidanol, in particular the alcohol of the following formula

[0371]

[0372] Or any stereoisomerically different form thereof is oxidized by the ADH to the corresponding dedimethylhexane carbonyl compound, particularly the minoyloxy group.

[0373] Optionally, the obtained product can be separated in a stereoisomer-pure form or as a mixture of stereoisomers.

[0374] ii) This invention relates to specific embodiments of a biocatalytic method using peptides with enal cleavage activity:

[0375] 27. A biocatalytic method for preparing compounds of general formula IV:

[0376]

[0377] in

[0378] R 1 Represents H or lower alkyl groups, especially methyl.

[0379] R 2 Representing H, a saturated or unsaturated optionally substituted hydrocarbon group, particularly alkyl or alkenyl, especially a group having up to 30, 20, 15, or 10 carbon atoms, or the residue Cyc-A-.

[0380] in

[0381] Cyc represents a optionally substituted saturated or unsaturated, particularly non-aromatic, monocyclic or polycyclic, particularly monocyclic or bicyclic hydrocarbon residue, especially a group having 5 to 7 ring carbon atoms, and

[0382] A represents a straight-chain or branched alkylene bridge with a chemical bond or optional substitution, particularly a methylene group.

[0383] and

[0384] R 3 Each of these can independently represent H or a lower alkyl group, such as C1-C4 alkyl groups, especially H or methyl, and more particularly each of these can be H.

[0385] The method includes the following steps:

[0386] (1) Contact the corresponding undegraded precursor of general formula V with natural or recombinant polypeptides with enaldehyde cleavage activity, especially polypeptides with α,β-unsaturated aldehyde C=C bond cleavage:

[0387]

[0388] in

[0389] R 1 R 2 and R 3 The definition is the same as above; and

[0390] R 4 Represents H or lower alkyl groups, especially H or methyl.

[0391] R 5 Represents H or lower alkyl groups, especially H.

[0392] Furthermore, the compound may exist in a stereoisomerically pure form (e.g., in E- or Z- form) or as a mixture of stereoisomers.

[0393] as well as

[0394] (2) Optionally, the degradation product of formula IV obtained in step (1) is separated, wherein the compound of formula IV can be obtained in stereoisomeric pure form or as a mixture of stereoisomers.

[0395] 28. The method of implementing Implementation 27, wherein

[0396] The polypeptide having the aforementioned enal cleavage activity is selected from a group of polypeptides comprising:

[0397] a) At least one DUF4334 protein family domain with Pfam ID number PF14232 (particularly within the C-terminal region of its amino acid sequence); and / or

[0398] b) At least one GXWXG protein family domain with Pfam ID number PF14231 (particularly within the N-terminal region of its amino acid sequence); or

[0399] c) Structural domains that maintain at least 90% sequence identity with PF14232 or PF14231;

[0400] In particular, if the e-value of the polypeptide with enal cleavage activity of the present invention that matches the said domain is less than 1 x 10⁻⁶, -5 or less than 1x10 -10 or less than 1x10 -15 or less than 1x10 -20 or less than 1x10 -25 or less than 1x10 -30 or less than or equal to 1x10 -35 Especially in 1x10 -20 Up to 1x10 -32 Within the range, especially in 1x10 -25 Up to 1x10 -31 If the range is within the specified range, it is identified as a member of the DUF4334 protein family, which includes the PF14232 domain.

[0401] For example, the following websites can be requested to retrieve and calculate such e values:

[0402] http: / / pfam.xfam.org /

[0403] http: / / www.ebi.ac.uk / Tools / hmmer / search / hmmscan or

[0404] http: / / www.ebi.ac.uk / Tools / pfa / pfamscan /

[0405] In particular, if the e-value of the peptide with enal cleavage activity of the present invention is less than 1 x 10⁻⁶, -5 or less than 1x10 -10 or less than 1x10 -15 or less than 1x10 -20 or less than 1x10 -25 or less than 1x10 -30 or less than or equal to 1x10 -35 Especially in 1x10 -20 Up to 1x10 -30 If the range is within the specified range, it is identified as a member of the GXWXG protein family containing the PF14231 domain.

[0406] Request the sequence of a polypeptide with enal cleavage activity as the query sequence.

[0407] For example, the following websites can be requested to retrieve and calculate such e values:

[0408] http: / / pfam.xfam.org /

[0409] http: / / www.ebi.ac.uk / Tools / hmmer / search / hmmscan or

[0410] http: / / www.ebi.ac.uk / Tools / pfa / pfamscan /

[0411] and / or

[0412] The polypeptide having the enal cleavage activity is selected from a group of polypeptides, the polypeptide containing at least one sequence motif / domain selected from the following:

[0413] The G-[Y or "-"]-xWxGxx-[F, L or I]-x-[T, S or R]-G-[H or D] shown in SEQ ID NO:205

[0414] Or any partial motif containing up to 10 or up to 5 consecutive amino acid residues, such as residues at positions 1–8 or 9–13 corresponding to SEQ ID NO:205;

[0415] The W-[Y, A, or V]-GKx-[F or Y]-x-[S or D] shown in SEQ ID NO:206

[0416] Or any partial motif containing up to four consecutive amino acid residues, such as residues at positions 1–4 or 5–8 corresponding to SEQ ID NO:206;

[0417] The [G or S]-x-[A or G]-x-[L or V]-xxxx-[F, Y or L]-RGxV shown in SEQ ID NO:207

[0418] Or any partial motif containing up to 10 or up to 5 consecutive amino acid residues, such as residues at positions 1–8 or 9–14 corresponding to SEQ ID NO:207;

[0419] The [M or L]-[V or I]-YDxxP-[I or V]-xD-[H or S]-[F or L] shown in SEQ ID NO:208

[0420] Or any partial motif containing up to 10 or up to 5 consecutive amino acid residues, such as residues at positions 1-6 or 7-12 corresponding to SEQ ID NO:208;

[0421] in

[0422] In the above motifs, residue x represents any natural amino acid residue independently of each other, and optionally in each of the above motifs, 1 to 5 amino acid residues, such as 1, 2, 3, 4 or 5 different from residue x, may be modified, for example by amino acid substitution, particularly by conserved substitution, provided that the enzyme retains enal lyase activity to at least an analytically detectable extent.

[0423] and / or

[0424] The polypeptide having the aforementioned enal cleavage activity is selected from the group consisting of polypeptides comprising the corresponding amino acid sequences:

[0425] a) SCH94-3944 shown in SEQ ID NO:34

[0426] b) SCH80-05241 shown in SEQ ID NO:38

[0427] c) Pdigit7033 shown in SEQ ID NO:42

[0428] d) PitalDUF4334-1 shown in SEQ ID NO:46

[0429] e) AspWeDUF4334 shown in SEQ ID NO:49

[0430] f) RhoagDUF4334-2 as shown in SEQ ID NO:53

[0431] g) RhoagDUF4334-3 as shown in SEQ ID NO:56,

[0432] h) RhoagDUF4334-4 as shown in SEQ ID NO:59,

[0433] i) CnecaDUF4334 shown in SEQ ID NO:62,

[0434] j) Rins-DUF4334 shown in SEQ ID NO:69

[0435] k)CgatDUF4334 shown in SEQ ID NO:72

[0436] l)GclavDUF4334 shown in SEQ ID NO:75

[0437] m) TcurvaDUF4334 as shown in SEQ ID NO:81

[0438] n) PprotDUF4334 shown in SEQ ID NO:87, and

[0439] o) A polypeptide comprising an amino acid sequence having at least 40%, 45%, 50%, 55%, 60%, 65%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% sequence identity with any of the amino acid sequences of a) to n) and retaining the enzymatic activity of the terpene precursor of formula (1).

[0440] Another specific embodiment refers to a polypeptide variant of the novel polypeptide of the present invention having enal cleavage activity as identified above by specific amino acid sequences of SEQ ID NO:34, 38, 42, 46, 49, 53, 56, 59, 62, 69, 72, 75, 81 and 87, and wherein the polypeptide variant is selected from any one of SEQ ID NO:34, 38, 42, 46, 49, 53, 56, 59, 62, 69, 72, 75, 81 and 87 having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98% or 99% sequence identity, and contains at least one substitution modification relative to any one of SEQ ID NO:34, 38, 42, 46, 49, 53, 56, 59, 62, 69, 72, 75, 81 and 87.

[0441] Of the 14 specific enal cleavage peptides cited above, the protein family domains with Pfam IDs PF14232 and PF14231 can be located at the amino acid residue positions given in the table below (see also...). Figure 31 (The alignment described in the text and the sequence portion highlighted therein)

[0442]

[0443]

[0444] The amino acid residue number refers to the residue number in the corresponding SEQ ID NO of the corresponding protein sequence in the attached sequence listing.

[0445] 29. The method of embodiment 28, wherein the enal cleavage polypeptide is selected from the group of mutants consisting of the following polypeptides and containing the corresponding amino acid sequences:

[0446] a) The SCH94-3944-T51A variant shown in SEQ ID NO:91

[0447] b) The SCH94-3944-H53A variant shown in SEQ ID NO:93

[0448] c) The SCH94-3944-L59A variant shown in SEQ ID NO:95

[0449] d) The SCH94-3944-W64A variant shown in SEQ ID NO:97

[0450] e) The SCH94-3944-S71A variant shown in SEQ ID NO:101

[0451] f) The SCH94-3944-R106A variant shown in SEQ ID NO:103

[0452] g) The variant SCH94-3944-Y115A shown in SEQ ID NO:105

[0453] h) The variant SCH94-3944-D116A shown in SEQ ID NO:107

[0454] i) The variant SCH94-3944-M136A shown in SEQ ID NO:111

[0455] j) The SCH94-3944-K139A variant shown in SEQ ID NO:113

[0456] k) The variant SCH94-3944-R156A shown in SEQ ID NO:119 and

[0457] l) A polypeptide comprising an amino acid sequence having at least 90%, 95%, 96%, 97%, 98%, or 99% sequence identity with any of the amino acid sequences of a) to l), retaining the enzymatic activity of the terpene precursor of formula (1), and retaining the position of the mutated amino acid sequence.

[0458] 30. The method of any one of embodiments 27 to 29, wherein a compound of formula V is used, such as a terpene compound, wherein

[0459] R 1 Represents H or methyl.

[0460] R 2 Represents H or

[0461] a) having 1 to 20, particularly 1 to 10, 1 to 15, or 1 to 20 carbon atoms, of a non-cyclic, straight-chain or branched, saturated or unsaturated hydrocarbon residues; or

[0462] b) The group Cyc-A-, wherein A represents a straight-chain or branched C1-C4 alkylene bridge, particularly methylene, and Cyc represents a monocyclic or polycyclic, particularly bicyclic, saturated or unsaturated hydrocarbon residue, particularly bicyclic fused hydrocarbon residues comprising 5 to 7, particularly 6, ring atoms per ring, optionally substituted with 1 to 10 or 1 to 5 substituents, independently selected from C1-C4 alkyl, C1-C4 alkylene, C2-C4 alkenyl, oxo, hydroxyl, or amino groups, particularly C1-C4 alkyl such as methyl, and C1-C4 alkylene such as methylene.

[0463] Each R 3 Represents H,

[0464] R 4 Represents H or methyl, and

[0465] R 5 It represents H or methyl.

[0466] 31. The method of embodiment 30, wherein the compound of general formula V has a hemispherane-type structure, and / or Cyc-A represents residues of formula IIIa, IIIb, or IIIc:

[0467]

[0468] 32. The method of any one of embodiments 27 to 31, wherein the precursor of formula (V) is selected from farnesal, geranyl geranyl, citral, dodecanal, hemispheryl-type compounds such as 8-hydroxy-hemispheryl-13-en-15-aldehyde, and copalaldehyde.

[0469] Each is either a mixture of its own stereoisomers or a stereoisomeric pure form.

[0470] 33. The method of any one of embodiments 27 to 32, wherein the degradation product of formula (IV) is selected from geranylacetone, farnesylacetone, methylheptenone, decanal; or minoyloxy, or 8-hydroxy-14,15-dedimethylhesperidin-13-one, each in the form of a mixture of its stereoisomers or in the form of a stereoisomeric pure.

[0471] Other specific examples of the invention regarding the enal lyase-catalyzed conversion of carbonyl compounds into their corresponding cleavage products are summarized in the following illustrative overview:

[0472]

[0473]

[0474] The parameter “n” is an integer from 1 to 20, 1 to 15, 1 to 10, or 1, 2, 3, 4, or 5.

[0475] 34. The method of any one of Implementation Schemes 27 to 33, which is performed in vitro or in test tube.

[0476] 35. The method of embodiment 34 performed in vivo, comprising, prior to step (1), particularly in non-human host cells, recombinantly expressing one or more polypeptides having the enzymatic activity required to carry out the chain degradation step.

[0477] 36. The method of embodiment 35, wherein the non-human host cell is transformed with a nucleic acid encoding at least one polypeptide having enal cleavage activity.

[0478] 37. The method of implementation scheme 35 or 36, wherein the non-human host cell is a eukaryotic or prokaryotic cell, particularly a plant cell, bacterial or fungal cell, particularly a yeast cell.

[0479] 38. The method of any one of embodiments 35 to 37, wherein the non-human host cell is a single-celled organism, a cultured cell derived from a multicellular organism, a cell present in a cultured tissue derived from a multicellular organism, or a cell present in a living multicellular organism.

[0480] 39. The method of any one of embodiments 35 to 38, wherein the non-human host cell is a bacterium of the genus Escherichia, particularly Escherichia coli, and the yeast is a yeast of the genus Saccharomyces or Pichia, particularly Saccharomyces cerevisiae, or a plant cell.

[0481] 40. The method of any one of embodiments 27 to 39 further comprises the following step as step (3): treating the compound of formula IV formed in step (1) or isolated in step (2) with chemical or biocatalytic synthesis or a combination thereof to obtain a derivative thereof, wherein the derivative may be particularly selected from hydrocarbons, alcohols, diols, triols, acetals, ketals, aldehydes, acids, ethers, amides, ketones, lactones, epoxides, acetates, glycosides and / or esters, and step (4): optionally, isolating the derivative of step (3).

[0482] 41. The method of embodiment 40, wherein step (3) comprises treating the compound of formula IV formed in step (1) or isolated in step (2) with a polypeptide having Bayer-Villiger monooxygenase (BVMO) activity to form the corresponding carbonyl ester.

[0483] 42. The method of embodiment 41 further includes hydrolyzing the carbonyl ester compound with an esterase to the corresponding deesterification product, which may be an alcohol or its isomerization product, and optionally, separating the derivative of step (3).

[0484] 43. The method of embodiment 41, wherein the polypeptide having BVMO activity is as defined in embodiment 9 above.

[0485] 44. The method of embodiment 42, wherein the esterase is selected from the group consisting of:

[0486] a) A polypeptide containing the SCH23-esterase amino acid sequence shown in SEQ ID NO:20;

[0487] b) A polypeptide containing the SCH24-esterase amino acid sequence shown in SEQ ID NO:24;

[0488] c) A polypeptide containing the SCH25-esterase amino acid sequence shown in SEQ ID NO:28;

[0489] d) A polypeptide containing the SCH46-esterase amino acid sequence shown in SEQ ID NO:31; or

[0490] e) A polypeptide comprising an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity with any of the amino acid sequences in a) to d) and having esterase activity.

[0491] 45. The method of embodiments 41 to 44, wherein the carbonyl compound is desdimethyl monazine, particularly minoyloxy, which is converted by said BVMO into the corresponding destetramethyl monazine acetate, particularly γ-ambryl acetate.

[0492] 46. ​​The method of embodiment 45, wherein detetramethylambryl acetate, particularly γ-ambryl acetate, is deesterified by the esterase to the corresponding detetramethylambryl, particularly γ-ambrol.

[0493] 47. The method of any one of embodiments 27 to 46 above, wherein prior to step (1), the method comprises biocatalytically oxidizing serotonol to serotonaldehyde, particularly biocatalytically oxidizing coparol to coparaldehyde.

[0494] The monazine alcohol is optionally formed by biocatalytic conversion of at least one terpenoid diphosphate precursor selected from IPP, DMAPP, FPP and GGPP, particularly by a single step or a combination of at least two steps known in the art.

[0495] The monsoon tanols can be produced, for example, via biocatalysis, in the following manner:

[0496] a) Produced from geraniol geraniol diphosphate (GGPP) in one step via cyclization / dephosphorylation;

[0497] b) GGPP is cyclized in two steps to form monazine diphosphate, such as cobazine diphosphate (CPP), and then dephosphorylated to monazine alcohol;

[0498] c) From IPP and DMAPP, which are directly converted to hemispheric diphosphate, such as CPP, by the action of bifunctional GGPP synthase / CPP synthase, and then dephosphorylated;

[0499] The GGPP used in these steps can also be provided by different biocatalytic steps:

[0500] d) GGPP can be generated directly from IPP and DMAPP using GGPP synthase; or

[0501] e) GGPP can be provided by the action of FPP synthase from IPP and DMAPP via FPP, and then by the action of GGPP synthase to convert FPP into GGPP.

[0502] 48. The method of implementing 47, wherein

[0503] The biocatalytic oxidation of hemispheric alcohols, particularly the biocatalytic oxidation of coparol to coparaldehyde, is catalyzed by exogenous or endogenous polypeptides with alcohol dehydrogenase (ADH) (EC 1.1.1.-) activity; and / or

[0504] The biocatalytic formation of hemispheric alcohols includes at least one step selected from the following:

[0505] i) Biocatalytic dephosphorylation of hemispheric diphosphate to hemispheric aldehyde, particularly copaclidiphosphate (CPP) to copacliol, catalyzed by peptides with terpenoid diphosphate (TPP) phosphatase activity, and / or

[0506] ii) Biocatalytic cyclization of terpene diphosphate precursors, such as geraniylgeraniyl diphosphate (GGPP), to CPP, catalyzed by peptides with CPP synthase activity, such as SmCPS2 (SEQ ID NO: 185); or, for example, biocatalytic cyclization of IPP and DMAPP to CPP, catalyzed by bifunctional peptides with isoprene transferase and copacpyl diphosphate synthase activities, such as PvCPS, and / or

[0507] iii) Biocatalytic formation from FPP to GGPP, or from IPP and DMAPP, each catalyzed by a polypeptide with GGPP synthase activity.

[0508] 49. The method of implementing Implementation 48, wherein

[0509] The biocatalytic oxidation, particularly the biocatalytic oxidation of coparol to coparaldehyde, is catalyzed by a polypeptide with alcohol dehydrogenase (ADH) activity, the polypeptide being selected from...

[0510] a) A polypeptide comprising the SCH23-ADH1_wt amino acid sequence shown in SEQ ID NO:134,

[0511] b) A polypeptide comprising the SCH24-ADH1_wt amino acid sequence shown in SEQ ID NO:140,

[0512] c) A polypeptide comprising the SCH94-3945_wt amino acid sequence shown in SEQ ID NO:161,

[0513] d) A polypeptide comprising the SCH80-0540_wt amino acid sequence shown in SEQ ID NO:164,

[0514] e) A polypeptide comprising the AzTolADH1_wt amino acid sequence shown in SEQ ID NO:167,

[0515] f) A polypeptide comprising the CdGeoA_wt amino acid sequence shown in SEQ ID NO:179.

[0516] g) A polypeptide comprising an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity with any of the amino acid sequences in a) to f) and having ADH activity.

[0517] and / or

[0518] The biocatalytic dephosphorylation, particularly the biocatalytic dephosphorylation from copaclidium diphosphate (CPP) to copacliol, is catalyzed by a polypeptide with terpene diphosphate (TPP) phosphatase activity, wherein the polypeptide is selected from:

[0519] d) A polypeptide comprising the AspWE TPP amino acid sequence shown in SEQ ID NO:170 or a polypeptide comprising an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity with it;

[0520] e) A polypeptide comprising the TalCeTPP amino acid sequence shown in SEQ ID NO:176, or a polypeptide comprising an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity with it; and

[0521] f) A polypeptide comprising the TalVeTPP amino acid sequence shown in SEQ ID NO:194 or a polypeptide comprising an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity with it;

[0522] Other suitable phosphatases are also disclosed in the applicant’s earlier EP application 18182783.3, which is incorporated herein by reference.

[0523] and / or

[0524] The biocatalytic cyclization, particularly the biocatalytic cyclization of gerany-gerany-diphosphate (GGPP) to CPP, is catalyzed by peptides selected from the following:

[0525] A polypeptide containing the SmCPS2 amino acid sequence shown in SEQ ID NO:185 and having cobazin activity, or a polypeptide containing an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity with it;

[0526] The biocatalytic cyclization, particularly the biocatalytic cyclization of IPP and DNMAPP to CPP, is catalyzed by peptides selected from the following:

[0527] A polypeptide containing the PvCPS amino acid sequence shown in SEQ ID NO:173 and having isoprenyl transferase and cobazide diphosphate synthase activities, or a polypeptide containing an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity with it.

[0528] and / or

[0529] The biocatalytic formation of GGPP is catalyzed by peptides having GGPP synthase activity and selected from the following:

[0530] d) A polypeptide comprising the carG amino acid sequence shown in SEQ ID NO:182 or a polypeptide comprising an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity with it;

[0531] e) A polypeptide comprising the CrtE amino acid sequence shown in SEQ ID NO:191 or a polypeptide comprising an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity with it;

[0532] f) A polypeptide comprising the PvCPS amino acid sequence shown in SEQ ID NO:173 or a polypeptide comprising an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity with it.

[0533] iii) This invention relates to the following specific embodiments relating to enal lyases and the corresponding coding sequences.

[0534] 50. An isolated polypeptide having enaldehyde cleavage activity, particularly α,β-unsaturated aldehyde C=C bond cleavage enzyme activity, as defined in any one of embodiments 28 and 29.

[0535] The polypeptides of the present invention comprise all active forms of enzymes having enal cleavage activity, including active subsequences, such as catalytic domains or active sites.

[0536] 51. An isolated nucleic acid molecule comprising a nucleic acid sequence encoding a polypeptide of embodiment 50, particularly a nucleic acid sequence selected from SEQ ID NO: 33, 35, 36, 37, 39, 40, 41, 43, 44, 45, 47, 48, 50, 51, 52, 54, 55, 57, 58, 60, 61, 63, 64, 68, 70, 71, 73, 74, 76, 80, 82, 86, 88, 92, 94, 96, 98, 102, 104, 106, 108, 112, and 120, and a nucleic acid sequence related to ...0, 35, 36, 37, 39, 40, 41, 43, 44, 45, 47, 48, 50, 51, 52, 54, 55, 57, 58, 60, 61, 62, 63, NO: Any one of the sequences described above, 33, 35, 36, 37, 39, 40, 41, 43, 44, 45, 47, 48, 50, 51, 52, 54, 55, 57, 58, 60, 61, 63, 64, 68, 70, 71, 73, 74, 76, 80, 82, 86, 88, 92, 94, 96, 98, 102, 104, 106, 108, 112, and 120, has a nucleic acid sequence with at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% sequence identity.

[0537] 52. An expression cassette comprising the nucleotide sequence of at least one nucleic acid molecule of embodiment 50.

[0538] 53. An expression vector comprising the nucleotide sequence of at least one nucleic acid molecule of embodiment 51, or at least one expression cassette of embodiment 52.

[0539] 54. The expression vector of implementation scheme 53, wherein the vector is a prokaryotic vector, a viral vector or a eukaryotic vector.

[0540] 55. The expression vector of any one of embodiments 53 to 54, which is a plasmid or a combination of two or more plasmids.

[0541] 56. A recombinant non-human host cell comprising at least one nucleic acid molecule as defined in embodiment 51, or at least one expression cassette as defined in embodiment 52, or at least one expression vector as defined in any one of embodiments 53 to 55.

[0542] 57. The host cell of implementation scheme 56, wherein at least one nucleic acid molecule or at least one expression cassette is stably integrated into the cell's genome.

[0543] 58. The host cell of implementation scheme 56 or 57 is a prokaryotic or eukaryotic cell, particularly a plant cell, bacterial or fungal cell, especially yeast.

[0544] 59. The host cell of any one of embodiments 56 to 58 is a single-celled organism, a cultured cell derived from a multicellular organism, a cell existing in a cultured tissue derived from a multicellular organism, or a cell existing in a living multicellular organism.

[0545] 60. The host cell of embodiment 59 is a bacterium of the genus Escherichia, preferably E. coli, or a yeast cell of the genus Saccharomyces, preferably S. cerevisiae, or a yeast cell of the genus Pichia, preferably Pichia pastoris.

[0546] 61. A method for generating at least one polypeptide having enal cleavage activity according to embodiment 51, the method comprising:

[0547] (i) expressing at least one of the polypeptides in a non-human host cell of any one of embodiments 57 to 60; and

[0548] (ii) Optionally, the at least one polypeptide is isolated from the non-human host cell used in step (i).

[0549] 62. The method of embodiment 61, prior to step (i), further comprises: preparing a non-human host cell for step (i) by introducing at least one nucleic acid molecule as defined in embodiment 51, or at least one expression cassette of embodiment 52, or at least one expression vector of any one of embodiments 53 to 55 into a non-human cell, thereby generating a host cell capable of expressing or overexpressing at least one polypeptide having enal cleavage activity according to embodiment 50.

[0550] 63. A method for preparing a mutant polypeptide with enal cleavage activity, the method comprising the following steps:

[0551] (i) Provide nucleic acid molecules according to implementation plan 51;

[0552] (ii) Modifying the nucleotide sequence of the nucleic acid molecule, particularly the nucleotide sequence encoding the polypeptide of embodiment 50, to obtain at least one mutant nucleic acid molecule;

[0553] (iii) Recombinant expression of the mutant nucleic acid molecule in a non-human host cell;

[0554] (iv) Screening for at least one mutant polypeptide with enaldehyde cleavage activity among the expression products obtained in step (iii); and

[0555] (v) Optionally, steps (ii) to (iv) are repeated with the mutant nucleic acid molecule until the expression product contains a mutant polypeptide with the desired enal cleavage activity; and

[0556] (vi) Optionally, isolate mutant peptides with the desired enaldehyde cleavage activity.

[0557] iv) This invention relates to the following specific embodiments related to BVMO enzymes and corresponding coding sequences.

[0558] 64. An isolated polypeptide having BVMO activity, as defined in embodiment 9.

[0559] The polypeptides of the present invention comprise all active forms of enzymes having BVMO activity, including active subsequences, such as catalytic domains or active sites.

[0560] 65. An isolated nucleic acid molecule comprising a nucleic acid sequence encoding a polypeptide of embodiment 64, particularly a nucleic acid sequence selected from SEQ ID NO: 1, 3, 4, 5, 7, 8, 9, 11, 12, 14, 15, 17 and 18, and a nucleic acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98% or 99% sequence identity with any of the sequences described in SEQ ID NO: 1, 3, 4, 5, 7, 8, 9, 11, 12, 14, 15, 17 and 18.

[0561] 66. An expression cassette comprising the nucleotide sequence of at least one nucleic acid molecule embodied in scheme 65.

[0562] 67. An expression vector comprising the nucleotide sequence of at least one nucleic acid molecule of embodiment 65, or at least one expression cassette of embodiment 66.

[0563] 68. The expression vector of implementation scheme 67, wherein the vector is a prokaryotic vector, a viral vector or a eukaryotic vector.

[0564] 69. The expression vector of any one of Implementation Schemes 67 to 68 is a plasmid, or a combination of two or more plasmids.

[0565] 70. A recombinant non-human host cell comprising at least one nucleic acid molecule as defined in embodiment 65, or at least one expression cassette as defined in embodiment 66, or at least one expression vector as defined in any one of embodiments 67 to 69.

[0566] 71. The host cell of implementation scheme 70, wherein at least one nucleic acid molecule or at least one expression cassette is stably integrated into the cell's genome.

[0567] 72. The host cell of implementation scheme 70 or 71 is a prokaryotic or eukaryotic cell, particularly a plant cell, bacterial or fungal cell, especially yeast.

[0568] 73. The host cell of any one of embodiments 70 to 72 is a single-celled organism, a cultured cell derived from a multicellular organism, a cell existing in a cultured tissue derived from a multicellular organism, or a cell existing in a living multicellular organism.

[0569] 74. The host cell of embodiment 72 is a bacterium of the genus Escherichia, preferably E. coli, or a yeast cell of the genus Saccharomyces, preferably S. cerevisiae, or a yeast cell of the genus Pichia, preferably Pichia pastoris.

[0570] 75. A method for generating at least one polypeptide having BVMO activity according to embodiment 64, the method comprising:

[0571] (i) expressing at least one of the polypeptides in a non-human host cell of any one of embodiments 70 to 74; and

[0572] (ii) Optionally, the at least one polypeptide is isolated from the non-human host cell used in step (i).

[0573] 76. The method of embodiment 75, prior to step (i), further comprises: preparing a non-human host cell for step (i) by introducing at least one nucleic acid molecule as defined in embodiment 65, or at least one expression cassette of embodiment 66, or at least one expression vector of any of embodiments 67 to 69 into a non-human cell, thereby generating a host cell capable of expressing or overexpressing at least one polypeptide having BVMO activity according to embodiment 64.

[0574] 77. A method for preparing a mutant polypeptide with BVMO activity, the method comprising the following steps:

[0575] (i) Provide nucleic acid molecules according to implementation plan 65;

[0576] (ii) Modifying the nucleotide sequence of the nucleic acid molecule, particularly the nucleotide sequence encoding the polypeptide of embodiment 64, to obtain at least one mutant nucleic acid molecule;

[0577] (iii) Recombinant expression of the mutant nucleic acid molecule in a non-human host cell;

[0578] (iv) Screening at least one mutant polypeptide with BVMO activity among the expression products obtained in step (iii); and

[0579] (v) Optionally, steps (ii) to (iv) are repeated with the mutant nucleic acid molecule until the expression product contains a mutant polypeptide with the desired BVMO activity; and

[0580] (vi) Optionally, isolate the mutant peptide having the desired BVMO activity.

[0581] v) This invention relates to specific embodiments of a multi-step in vivo biocatalytic method for converting hemispheric compounds by applying peptides having enal cleavage activity and / or BVMO activity.

[0582] 78. An in vivo method for preparing hemispheryl terpenes, the method comprising providing a recombinant host expressing a set of polypeptides having enzymatic activity required to catalyze the following sequential reaction steps:

[0583] (1) Optionally, sematanol, particularly copalol, is converted into the corresponding sematanal, particularly copalal, by enzymatic action of exogenous or endogenous ADH polypeptides, particularly ADH as defined in any of embodiments 19 or 49.

[0584] (2) The semacaran aldehyde, particularly copal aldehyde, of step (1) is converted into the corresponding dedimethyl semacaran carbonyl compound, particularly minoyloxy, by the action of a polypeptide having enaldehyde cleavage activity, particularly a polypeptide defined in any of embodiments 28 and 29.

[0585] (3) Optionally, the desdimethylambryl carbonyl compound, particularly the minoyloxy group, of step (2) is converted into the corresponding destetramethylambryl alkyl acetate, particularly γ-ambryl acetate, by the action of a polypeptide having BVMO activity, particularly BVMO as defined in embodiment 9.

[0586] (4) Optionally, the detetramethylambryl acetate, particularly γ-ambryl acetate, of step (3) is converted to the corresponding detetramethylambryl alcohol, particularly γ-ambrol, by action of a polypeptide having esterase activity, particularly an esterase as defined in any one of embodiments 23 and 44; and optionally,

[0587] (5) Separate the products from steps (2), (3) or (4).

[0588] 79. An in vivo method for preparing hemispheric cyclic terpenes, the method comprising providing a recombinant host expressing a set of polypeptides having enzymatic activity required to catalyze the following sequential reaction steps:

[0589] (1) Optionally, sematanol, particularly copalol, is converted into the corresponding sematanal, particularly copalal, by enzymatic action of exogenous or endogenous ADH polypeptides, particularly ADH as defined in any of embodiments 19 or 49.

[0590] (2) The serotonaldehyde, particularly copalaldehyde, of step (1) is converted into the corresponding demethyl serotonaldehyde ester compound, particularly [4-[(1S,4aS,8aS)-5,5,8a-trimethyl-2-methylene-decahydronaphth-1-yl]-2-methyl-but-1-enyl]carbamate (compounds 1a, 1b) by the action of a polypeptide having BVMO activity, particularly BVMO as defined in any of embodiments 9.

[0591] (3) The serotonide ester compound of step (2), particularly [4-[(1S,4aS,8aS)-5,5,8a-trimethyl-2-methylene-decahydronaphthyl-1-yl]-2-methyl-but-1-enyl]carboxylate (compounds 1a, 1b), is converted into the corresponding demethylserotonal, particularly 4-[(1S,8aS)-5,5,8a-trimethyl-2-methylene-decahydronaphthyl-1-yl]-2-methyl-butanal (compounds 3a, 3b), which may be carried out by action of a polypeptide having esterase activity, particularly an esterase as defined in any of embodiments 23 or 44;

[0592] (4) The demethylated serotonaldehyde, particularly 4-[(1S,4aS,8aS)-5,5,8a-trimethyl-2-methylene-decahydronaphth-1-yl]-2-methyl-butanal (compounds 3a, 3b), of step (3) is converted into the corresponding demethylated serotonaldehyde ester, particularly [3-[(1S,4aS,8aS)-5,5,8a-trimethyl-2-methylene-decahydronaphth-1-yl]-1-methyl-propyl]carbamate (compounds 4a, 4b), by the action of a polypeptide having BVMO activity, particularly BVMO as defined in any of embodiments 9.

[0593] (5) The desdimethyl serotonate, particularly [3-[(1S,4aS,8aS)-5,5,8a-trimethyl-2-methylene-decahydronaphthyl-1-yl]-1-methyl-propyl]carbamate (compounds 4a, 4b) of step (4) is converted into the corresponding desdimethyl serotonol, particularly 4-[(1S,4aS,8aS)-5,5,8a-trimethyl-2-methylene-decahydronaphthyl-1-yl]but-2-ol (compounds 5a, 5b) by action of a polypeptide having esterase activity, particularly an esterase as defined in any of embodiments 23 or 44.

[0594] (6) Optionally, the desdimethyl monazine alcohol, particularly 4-[(1S,8aS)-5,5,8a-trimethyl-2-methylene-decahydronaphthyl-1-yl]but-2-ol (compounds 5a, 5b) of step (5) is converted into the corresponding desdimethyl monazine carbonyl compound, particularly minoyloxy; by the action of an exogenous or endogenous polypeptide having ADH activity, particularly as defined in any of embodiments 19 or 49.

[0595] (7) Optionally, the desdimethylambryl carbonyl compound, particularly the minoyloxy group, of step (6) is converted into the corresponding tetrademethylambryl alkyl acetate, particularly γ-ambryl acetate, by the action of a polypeptide having BVMO activity, particularly BVMO as defined in any of embodiments 9.

[0596] (8) The detetramethylambryl acetate, particularly γ-ambryl acetate, of step (7) is converted to the corresponding detetramethylambryl alcohol, particularly γ-ambrol, by action of a polypeptide having esterase activity, particularly an esterase as defined in any of embodiments 23 or 44; and optionally,

[0597] (9) Separate the products from steps (5), (6), (7) or (8).

[0598] 80. The method of implementing scheme 79, wherein

[0599] The ADH used in steps (1) and (6) may be the same or different; it may be exogenous or endogenous, and / or

[0600] The BVMOs used in steps (2), (4), and (7) may be the same or different; and / or

[0601] The esterases used in steps (3), (5) and (8) may be the same or different.

[0602] 81. The method of any one of implementation schemes 78 to 80, wherein

[0603] The recombinant host is used to express an additional set of peptides that have the enzyme activity required to catalyze the following sequential reaction steps prior to step (1):

[0604] (i) Biocatalytic formation of geraniol geraniol diphosphate (GGPP) by action of a polypeptide having GGPP synthase activity, particularly a GGPP synthase as defined in any of embodiments 19 and 49.

[0605] (ii) A biocatalytic ring from GGPP to said hemispheric diphosphate, particularly cobamic diphosphate (CPP), carried out by the action of a polypeptide having hemispheric diphosphate synthase activity, particularly a polypeptide containing CPP synthase activity as defined in any one of embodiments 19 and 49.

[0606] (iii) Biocatalytic dephosphorylation of said serotonin diphosphate to said serotonin alcohol, particularly from CPP to copalol, by action of a polypeptide having serotonin diphosphate phosphatase activity, particularly a polypeptide containing TPP phosphatase activity as defined in any of embodiments 19 and 49.

[0607] 82. The method of any one of embodiments 78 to 81, wherein a recombinant host is used to additionally express at least one polypeptide that catalyzes the enzymatic step of the mevalonate pathway or the MEP pathway.

[0608] 83. The method of any one of embodiments 78 to 82, wherein a recombinant host is used, which carries the coding sequence of a respective catalytically active polypeptide on one or more expression vectors and / or the coding sequence is stably integrated into the host genome.

[0609] 84. The method of any one of embodiments 1 to 49 and 78 to 83, carried out in vivo, comprising, prior to step (1), introducing one or more nucleic acid molecules encoding one or more polypeptides into a non-human host organism or cell and optionally stably integrating them into the corresponding genome, the polypeptide having the enzymatic activity required to carry out one or more corresponding biocatalytic transformation steps.

[0610] 85. The method of any one of embodiments 1 to 49 and 78 to 83, wherein the method is carried out by applying a non-human host organism or cell that endogenously produces a mixture of FPP and / or GGPP or IPP and DMAPP; or by applying a non-human host organism that has been genetically modified to produce an increased amount of FPP and / or GGPP and / or a mixture of IPP and DMAPP.

[0611] Some of the host cells or organisms suitable for use in this invention do not naturally produce FPP or GGPP or a mixture of IPP and DMAPP. Such organisms or cells that do not produce acyclic terpene pyrophosphate precursors such as FPP or GGPP or a mixture of IPP and DMAPP can be naturally genetically modified to produce said precursors. For example, they can be transformed prior to nucleic acid modification as described herein. Methods for transforming organisms to produce acyclic terpene pyrophosphate precursors such as FPP or GGPP or a mixture of IPP and DMAPP are known in the art. For example, introducing enzymatic activity of the mevalonate pathway is a suitable strategy for producing FPP or GGPP or a mixture of IPP and DMAPP in an organism.

[0612] 86. Recombinant microorganisms as defined in any one of Implementation Schemes 78 to 85.

[0613] vi) This invention relates to specific embodiments thereof, which involve further converting chemical intermediate compounds obtained by the biocatalytic methods described herein into other end products of particular interest.

[0614] 87. A method for preparing epoxy-detetramethylammonium compounds, particularly ambroxol, the method comprising:

[0615] (1) Providing detetramethylambranol, particularly γ-ambrol, or detetramethylambran acetate, particularly γ-ambryl acetate, or dedimethylambran carbonyl compound, particularly minoyloxy, by applying a biocatalytic method comprising one or more method steps as defined in any one of claims 1 to 49 or 78 to 83, and optionally separating said product; and

[0616] (2) The product of step (1) is converted into epoxy-detetramethylammonium ether, particularly ambroxol, by applying one or more chemical and / or biochemical transformation steps.

[0617] 88. A method for preparing diepoxy-dedimethyl monsoon, particularly Z11, the method comprising:

[0618] (1) Providing the desdimethylsemicarbazone carbonyl compound, particularly the minoyloxy group, by applying a method that results in the formation of the desdimethylsemicarbazone carbonyl compound, particularly the minoyloxy group, and includes one or more method steps as defined in any one of claims 1 to 49 or 78 to 84, and optionally isolating the desdimethylsemicarbazone carbonyl compound, particularly the minoyloxy group; and

[0619] (2) The desdimethyl monazine carbonyl compound of step (1), particularly the minoyloxy group, is converted into the diepoxy-desdimethyl monazine, particularly Z-11, by applying one or more chemical and / or biochemical conversion steps.

[0620] b. The polypeptide applicable according to the present invention

[0621] In the context of this article, the following definitions apply:

[0622] The commonly used terms “polypeptide” or “peptide” refer to a natural or synthetic, continuous, peptide-linked linear chain or sequence of amino acid residues containing approximately 10 to more than 1,000 residues. Short-chain polypeptides with up to 30 residues are also called “oligopeptides”.

[0623] The term "protein" refers to a large molecular structure composed of one or more polypeptides. The amino acid sequence of the polypeptide represents the protein's "primary structure." The amino acid sequence also predetermines the protein's "secondary structure" by forming specific structural elements (such as α-helices and β-sheets formed within the polypeptide chain). The arrangement of multiple such secondary structural elements defines the protein's "tertiary structure," or spatial arrangement. If a protein contains more than one polypeptide chain, these chains are arranged spatially to form the protein's "quaternary structure." Proper spatial arrangement, or "folding," is a prerequisite for protein function. Denaturation or unfolding disrupts protein function. If this disruption is reversible, protein function can be restored by refolding.

[0624] The typical protein function referred to in this article is "enzyme function," which means that a protein acts as a biocatalyst on a substrate, such as a compound, and catalyzes the conversion of said substrate into a product. Enzymes can exhibit high or low levels of substrate and / or product specificity.

[0625] Therefore, the term "peptide" as used in this article to refer to a specific "activity" implicitly refers to a correctly folded protein that exhibits the indicated activity, such as a specific enzyme activity.

[0626] Therefore, unless otherwise stated, the term "polypeptide" also covers the terms "protein" and "enzyme".

[0627] Similarly, the term "peptide fragment" encompasses the terms "protein fragment" and "enzyme fragment".

[0628] The term "isolated polypeptide" refers to an amino acid sequence extracted from its natural environment by any method known in the art or a combination of such methods (including recombinant, biochemical, and synthetic methods).

[0629] A "target peptide" is an amino acid sequence that targets a protein or polypeptide to intracellular organelles (i.e., mitochondria or plastids) or to the extracellular space (secretory signal peptides). The nucleic acid sequence encoding the target peptide can be fused to the amino-terminal (e.g., N-terminus) nucleic acid sequence encoding the protein or polypeptide, or it can be used to replace the natural target peptide.

[0630] The present invention also relates to “functional equivalents” (also referred to as “analogs” or “functional mutations”) of the polypeptides specifically described herein.

[0631] For example, a “functional equivalent” refers to a polypeptide that, in a test used to determine the activity of terpenoid diphosphate synthase or terpenoid diphosphate phosphatase, shows an activity that is at least 1 to 10%, or at least 20%, or at least 50%, or at least 75%, or at least 90% higher or lower than that of the polypeptide specifically described herein.

[0632] According to the invention, "functional equivalents" also encompass specific mutants that have an amino acid at at least one sequence position in the amino acid sequence described herein that differs from the specifically stated amino acid, but still possess one of the aforementioned biological activities, such as enzyme activity. Thus, "functional equivalents" include mutants obtainable by the addition, substitution, particularly conserved substitution, deletion, and / or inversion of one or more, for example, 1 to 20, 1 to 15, or 5 to 10 amino acids, wherein said changes can occur at any sequence position, as long as they result in the mutant possessing the general characteristics of the invention. Functional equivalence is also particularly provided if the activity pattern qualitatively overlaps between the mutant and the unaltered polypeptide, i.e., if, for example, an interaction with the same agonist or antagonist or substrate is observed, but at different rates (i.e., by EC...). 50 or IC 50 (Values ​​or any other parameters suitable in this technical field). The table below shows examples of suitable (conservative) amino acid substitutions:

[0633]

[0634]

[0635] The “functional equivalents” in the above sense are also the “precursors” of the polypeptides described in this article, as well as the “functional derivatives” and “salts” of the polypeptides.

[0636] In this case, a "precursor" is a natural or synthetic precursor of a polypeptide, which may or may not have the desired biological activity.

[0637] The term "salt" as used in this invention refers to salts of the carboxyl group of the protein molecule and salts formed by the acid addition of the amino group. Salts of the carboxyl group can be produced in known ways and include inorganic salts such as sodium, calcium, ammonium, iron, and zinc salts, as well as salts formed with organic bases such as amines, such as triethanolamine, arginine, lysine, piperidine, etc. Salts formed by acid addition, such as those formed with inorganic acids such as hydrochloric acid or sulfuric acid, and salts formed with organic acids such as acetic acid and oxalic acid, are also covered by this invention.

[0638] The “functional derivatives” of the polypeptides according to the invention can also be generated using known techniques at the side groups of functional amino acids or their N-terminus or C-terminus. Such derivatives include, for example: aliphatic esters of carboxylic acid groups, amides of carboxylic acid groups, which can be obtained by reacting with ammonia or with primary or secondary amines; N-acyl derivatives of free amino groups, which are generated by reacting with acyl groups; or O-acyl derivatives of free hydroxyl groups, which are generated by reacting with acyl groups.

[0639] "Functional equivalents" naturally include polypeptides that can be obtained from other organisms as well as naturally occurring variants. For example, the area of ​​homologous sequence regions can be determined by sequence comparison, and equivalent polypeptides can be determined based on the specific parameters of this invention.

[0640] "Functional equivalents" also include "fragments" of the polypeptide according to the invention, such as single domains or sequence motifs, or truncated N-terminuses and / or C-termini, which may or may not exhibit the desired biological function. Preferably, such "fragments" at least qualitatively retain the desired biological function.

[0641] Furthermore, a “functional equivalent” is a fusion protein having one of the polypeptide sequences described herein or a functional equivalent derived therefrom, and having at least one additional functionally distinct heterologous sequence in functional N-terminal or C-terminal association (i.e., without substantial mutual functional impairment of the fusion protein portion). Non-limiting examples of such heterologous sequences are, for example, signal peptides, histidine anchors, or enzymes.

[0642] The invention also includes “functional equivalents” that are homologs of the specifically disclosed polypeptides. They have at least 60%, preferably at least 75%, particularly at least 80 or 85%, such as 90, 91, 92, 93, 94, 95, 96, 97, 98, or 99% homology (or identity) with the specifically disclosed amino acid sequence, calculated using the algorithm described in Pearson and Lipman, Proc. Natl. Acad. Sci. (USA) 85(8), 1988, 2444-2448. The homology or identity of the homologous polypeptides according to the invention, expressed as a percentage, refers in particular to identity expressed as a percentage of amino acid residues based on the total length of one of the amino acid sequences specifically described herein.

[0643] Identity data expressed as a percentage can also be determined using BLAST alignment, the blastp (protein-protein BLAST) algorithm, or by applying the Clustal settings detailed below.

[0644] In the case of possible protein glycosylation, the “functional equivalents” according to the invention include polypeptides in deglycosylated or glycosylated forms as described herein, as well as modified forms that can be obtained by changing the glycosylation pattern.

[0645] Functional equivalents or homologs of the polypeptides according to the present invention can be generated by mutagenesis, for example by point mutation, elongation or shortening of the protein, or as described in more detail below.

[0646] Functional equivalents or homologs of the polypeptides according to the invention can be identified by screening a database of mutants, such as shortened mutants. For example, a database of protein variant diversity can be generated by combinatorial mutagenesis at the nucleic acid level, for example by enzymatic ligation of a mixture of synthetic oligonucleotides. Numerous methods are available for generating a database of potential homologs from degenerate oligonucleotide sequences. The chemical synthesis of degenerate gene sequences can be performed in an automated DNA synthesizer, and the synthesized gene can then be ligated into a suitable expression vector. The use of degenerate genomes makes it possible to provide all sequences in a mixture that encode the desired set of potential protein sequences. Methods for synthesizing degenerate oligonucleotides are known to those skilled in the art.

[0647] In the prior art, several techniques are known for screening gene products from combinatorial databases generated by point mutations or shortening, and for screening cDNA libraries containing gene products with selected properties. These techniques can be applied to rapidly screen gene libraries generated by combinatorial mutagenesis of homologs according to the present invention. The most commonly used high-throughput analysis-based techniques for screening large gene libraries involve cloning the gene library in a reproducible expression vector, transforming suitable cells with the resulting vector database, and expressing the combinatorial gene under specific conditions, under which detection of the desired activity facilitates the isolation of vectors encoding the gene (whose product is detected). Recursive integration mutagenesis (REM) is a technique for increasing the frequency of functional mutants in a database and can be used in conjunction with screening tests to identify homologs.

[0648] The embodiments provided herein offer orthologs and paralogs of the disclosed peptides, as well as methods for identifying and isolating such orthologs and paralogs. The definitions of the terms "ortholog" and "paralog" are given below and apply to both amino acid and nucleic acid sequences.

[0649] The polypeptides of this invention comprise all active forms of the enzymes of this invention, including active subsequences, such as catalytic domains or active sites. In one embodiment, this invention provides the catalytic domains or active sites as described below. In one embodiment, the present invention provides a peptide or polypeptide comprising or composed of an active site domain predicted by using a database such as Pfam (http: / / pfam.wustl.edu / hmmsearch.shtml) (a large collection covering multiple sequence alignments and hidden Markov models for many common protein families, Pfam Protein Family Database, A. Bateman, E. Birney, L. Cerruti, R. Durbin, L. Etwiller, S. REddy, S. Griffiths-Jones, K. L. Howe, M. Marshall, and E. L. Sonnhammer, Nucleic Acids Research, 30(1):276-280, 2002) or equivalent sources such as the InterPro and SMART databases (http: / / www.ebi.ac.uk / interpro / scan.html, http: / / smart.embl-heidelberg.de / ).

[0650] The present invention also covers “peptide variants” having the desired activity, wherein the variant peptide is selected from the amino acid sequence having at least 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity with the specific, particularly natural, amino acid sequence indicated by the specific SEQ ID NO and comprising at least one substitution modification relative to the SEQ ID NO.

[0651] c. The applicable coding nucleic acid sequence according to the present invention

[0652] In the context of this article, the following definitions apply:

[0653] The terms “nucleic acid sequence,” “nucleic acid,” “nucleic acid molecule,” and “polynucleotide” are used interchangeably and refer to a sequence of nucleotides. A nucleic acid sequence can be a single-stranded or double-stranded deoxyribonucleotide or ribonucleotide of any length and includes coding and non-coding sequences of genes, exons, introns, sense and antisense complementary sequences, genomic DNA, cDNA, miRNA, siRNA, mRNA, rRNA, tRNA, recombinant nucleic acid sequences, isolated and purified naturally occurring DNA and / or RNA sequences, synthetic DNA and RNA sequences, fragments, primers, and nucleic acid probes. Those skilled in the art understand that the nucleic acid sequence of RNA is identical to that of DNA, except that thymine (T) is replaced by uracil (U). The term “nucleotide sequence” should also be understood to include polynucleotide or oligonucleotide molecules in the form of individual fragments or as components of larger nucleic acids.

[0654] "Isolated nucleic acid" or "isolated nucleic acid sequence" refers to a nucleic acid or nucleic acid sequence that exists in an environment different from that of naturally occurring nucleic acids or nucleic acid sequences, and may include those that are substantially free of endogenous contaminants.

[0655] As used in this article, the term "naturally occurring" for nucleic acids refers to a nucleic acid that is found in the cells of organisms in nature and has not been intentionally modified by humans in a laboratory.

[0656] A “fragment” of a polynucleotide or nucleic acid sequence refers to a continuous sequence of nucleotides, particularly of a length of at least 15 bp, at least 30 bp, at least 40 bp, at least 50 bp, and / or at least 60 bp, according to one embodiment of this invention. Specifically, the polynucleotide fragment comprises at least 25, more particularly at least 50, more particularly at least 75, more particularly at least 100, more particularly at least 150, more particularly at least 200, more particularly at least 300, more particularly at least 400, more particularly at least 500, more particularly at least 600, more particularly at least 700, more particularly at least 800, more particularly at least 900, and more particularly at least 1000 consecutive nucleotides of a polynucleotide sequence according to one embodiment of this invention. Without limitation, the polynucleotide fragments described herein can be used as PCR primers and / or probes, or for antisense gene silencing or RNAi.

[0657] As used herein, the term "hybridization" or hybridization under certain conditions is intended to describe the conditions under which hybridization and washing occur, under which significantly identical or homologous nucleotide sequences remain bound to each other. These conditions can result in sequences with at least about 70%, for example at least about 80%, and for example at least about 85%, 90%, or 95% identity remaining bound to each other. Definitions of low-tightness, medium-tightness, and high-tightness hybridization conditions are provided below. Those skilled in the art can select suitable hybridization conditions with minimal experimentation, for example, as illustrated by Ausubel et al. (1995, Current Protocols in Molecular Biology, John Wiley & Sons, sections 2, 4, and 6). Additionally, tight conditions are described in Sambrook et al. (1989, Molecular Cloning: A Laboratory Manual, 2nd ed., Cold Spring Harbor Press, chapters 7, 9, and 11).

[0658] "Recombinant nucleic acid sequences" are nucleic acid sequences created by combining genetic material from more than one source using laboratory methods (such as molecular cloning), thereby creating or modifying nucleic acid sequences that are not naturally occurring and cannot be found in biological organisms in any other way.

[0659] “Recombinant DNA technology” refers to molecular biological methods used to prepare recombinant nucleic acid sequences, as described, for example, in Laboratory Manuals edited by Weigel and Glazebrook, 2002, Cold Spring Harbor Lab Press; and Sambrook et al., 1989, Cold Spring Harbor, NY: Cold Spring Harbor Laboratory Press.

[0660] The term "gene" refers to a DNA sequence containing a region that is operatively linked to a suitable regulatory region (e.g., a promoter) and transcribed into an RNA molecule (e.g., mRNA in a cell). Therefore, a gene can contain several operatively linked sequences, such as a promoter, a 5' leader sequence (containing, for example, a sequence involved in translation initiation), a coding region of cDNA or genomic DNA, introns, exons, and / or a 3' untranslated sequence (containing, for example, a transcription termination site).

[0661] "Polycistronic" refers to nucleic acid molecules, especially mRNA, that can encode more than one polypeptide within the same nucleic acid molecule.

[0662] A "chimeric gene" is any gene that is not normally found in species in nature, particularly a gene in which one or more portions of the nucleic acid sequence are unrelated in nature. For example, a promoter that is unrelated in nature to part or all of the transcribed region or to another regulatory region. The term "chimeric gene" should be understood to include expression constructs in which a promoter or transcriptional regulatory sequence is operatively linked to one or more coding sequences or antisense (i.e., the inverse complementary strand of the sense strand) or inverted repeat sequences (sense and antisense, whereby the RNA transcript forms a double-stranded RNA post-transcriptionally). The term "chimeric gene" also includes genes obtained by combining portions of one or more coding sequences to produce new genes.

[0663] "3'URT" or "3' untranslated sequence" (also known as "3' untranslated region" or "3' end") refers to a nucleic acid sequence found downstream of the gene coding sequence that contains, for example, a transcription termination site and (in most, but not all, eukaryotic mRNAs) a polyadenylation signal, such as AAUAAA or its variants. After transcription termination, the mRNA transcript can be cleaved downstream of the polyadenylation signal and a poly(A) tail can be added, which is involved in the transport of mRNA to the translation site, such as the cytoplasm.

[0664] The term "primer" refers to a short nucleic acid sequence that is hybridized to a template nucleic acid sequence and used for the polymerization of nucleic acid sequences complementary to that template.

[0665] The term "selectable marker" refers to any gene that, after expression, can be used to select one or more cells containing that selectable marker. Examples of selectable markers are described below. Those skilled in the art will understand that different antibiotic, fungicide, auxotrophic, or herbicide selectable markers may be applicable to different target species.

[0666] This invention also relates to nucleic acid sequences encoding polypeptides as defined herein.

[0667] In particular, the present invention also relates to nucleic acid sequences (single-stranded and double-stranded DNA and RNA sequences, such as cDNA, genomic DNA and mRNA) encoding one of the aforementioned polypeptides and their functional equivalents, which can be obtained, for example, by using artificial nucleotide analogs.

[0668] This invention relates to both isolated nucleic acid molecules encoding polypeptides or biologically active regions thereof according to the invention, and nucleic acid fragments that can be used, for example, as hybridization probes or primers for identifying or amplifying nucleic acids encoded according to the invention.

[0669] This invention also relates to nucleic acids that have a degree of "identity" with the sequences specifically disclosed herein. "Identity" between two nucleic acids refers to the identity of nucleotides along the entire length of the nucleic acid in each case.

[0670] The “identity” between two nucleotide sequences (and similarly, peptide or amino acid sequences) is a function of the number of nucleotide residues (or amino acid residues) when the two sequences are aligned, or the number of identical residues in both sequences. Identical residues are defined as the same residues at a given position in the alignment of the two sequences. The percentage of sequence identity used herein is calculated from the best alignment by dividing the number of identical residues between the two sequences by the total number of residues in the shortest sequence and multiplying by 100. The best alignment is the alignment with the highest probability of identity percentage. Vacancies can be introduced into one or more positions in the alignment of one or both sequences to obtain the best alignment. These vacancies are then considered as dissimilar residues used to calculate the percentage of sequence identity. Alignments used to determine the percentage of identity of amino acid or nucleic acid sequences can be performed in various ways using computer programs, such as those publicly available on the Internet.

[0671] Specifically, the BLAST program (Tatiana et al., FEMS Microbiol Lett., 1999, 174:247-250, 1999), which is available from the National Center for Biotechnology Information (NCBI) at http: / / www.ncbi.nlm.nih.gov / BLAST / bl2seq / wblast2.cgi with default parameters, can be used to obtain the best alignment of protein or nucleic acid sequences and calculate the percentage of sequence identity.

[0672] In another example, identity can be calculated using the Clustal method (Higgins DG, Sharp PM. (1989)) through the Vector NTI Suite 7.1 program of Informax Corporation (USA) with the following settings:

[0673] Multiple alignment parameters:

[0674]

[0675]

[0676] Comparison parameters:

[0677]

[0678] Alternatively, identity can be determined according to the method of Chenna et al. (2003), webpage: http: / / www.ebi.ac.uk / Tools / clustalw / index.html# and the following settings:

[0679]

[0680] All nucleic acid sequences (single-stranded and double-stranded DNA and RNA sequences, such as cDNA and mRNA) mentioned in this article can be produced from nucleotide structural units by chemical synthesis in a known manner, for example, by fragment condensation of individual overlapping complementary nucleic acid structural units of a double helix. The chemical synthesis of oligonucleotides can be carried out, for example, by the phosphoramide process (Voet, Voet, 2nd edition, Wiley Press, New York, pages 896-897) in a known manner. The accumulation of synthetic oligonucleotides, and the filling of vacancies in the ligation reaction using the Klenow fragment of DNA polymerase, as well as general cloning techniques, are described in Sambrook et al. (1989), see below.

[0681] In addition, the nucleic acid molecules according to the present invention may further include untranslated sequences from the 3' and / or 5' ends of the coding genetic region.

[0682] The present invention further relates to nucleic acid molecules that are complementary to the nucleotide sequences or segments thereof specifically described.

[0683] The nucleotide sequences according to the invention enable the generation of probes and primers that can be used to identify and / or clone homologous sequences in other cell types and organisms. Such probes or primers typically contain a nucleotide sequence region that hybridizes to at least about 12, preferably at least about 25, such as about 40, 50 or 75 consecutive nucleotides of the sense strand or corresponding antisense strand of the nucleic acid sequence according to the invention under “stringent” conditions (as defined elsewhere herein).

[0684] "Homologous" sequences include orthologous or paralogous sequences. Methods for identifying orthologous or paralogous sequences include phylogenetic methods, sequence similarity methods, and hybridization methods known in the art and described herein.

[0685] "Paralleloids," or paralogous sequences, arise from gene replication, resulting in two or more genes with similar sequences and functions. Paralogs typically cluster together and form through gene replication within related plant species. Paralogs are identified in groups of similar genes using pairwise BLAST analysis or procedures such as CLUSTAL during phylogenetic analysis of gene families. In paralogs, the shared sequence can be identified as a sequence characteristic of the related gene and possessing a similar function.

[0686] "Orthologs," or orthologous sequences, are sequences that are similar to each other because they are found in species descended from a common ancestor. For example, plant species known to share a common ancestor contain many enzymes with similar sequences and functions. For instance, by constructing a phylogenetic tree of a gene family for a species using procedures such as CLUSTAL or BLAST, a technician can identify orthologous sequences and predict the functions of orthologs. One method for identifying or confirming similar functions between homologous sequences is by comparing transcript profiles in host cells or organisms (such as plants or microorganisms) that overexpress or lack (in gene knockout / reduction) the relevant polypeptide. A technician can understand that genes with similar transcript profiles (having a common transcript with greater than 50% regulation, or a common transcript with greater than 70% regulation, or a common transcript with greater than 90% regulation) will have similar functions. By causing host cells, organisms such as plants or microorganisms to produce terpene synthase proteins, homologs, parahomologs, orthologs, and any other variants of the sequences described herein are expected to function in a similar manner.

[0687] The term "selectable marker" refers to any gene that, after expression, can be used to select one or more cells containing that selectable marker. Examples of selectable markers are described below. Those skilled in the art will understand that different antibiotic, fungicide, auxotrophic, or herbicide selectable markers may be applicable to different target species.

[0688] Nucleic acid molecules according to the invention can be isolated using standard molecular biology techniques and the sequence information provided according to the invention. For example, cDNA can be isolated from a suitable cDNA library using one of the specifically disclosed complete sequences or fragments thereof as hybridization probes and standard hybridization techniques (e.g., described in Sambrook, (1989)).

[0689] Alternatively, nucleic acid molecules containing one or a fragment of the disclosed sequence can be isolated by polymerase chain reaction using oligonucleotide primers constructed based on that sequence. The nucleic acid amplified in this manner can be cloned into a suitable vector and characterized by DNA sequencing. The oligonucleotides according to the invention can also be prepared using standard synthetic methods, for example, using an automated DNA synthesizer.

[0690] According to the nucleic acid sequences or derivatives thereof of the present invention, homologs or portions of these sequences can be isolated from other bacteria, for example, by conventional hybridization techniques or PCR techniques, through genomic or cDNA libraries. These DNA sequences hybridize with the sequences according to the present invention under standard conditions.

[0691] "Hybridization" refers to the ability of polynucleotides or oligonucleotides to bind to nearly complementary sequences under standard conditions, while non-complementary pairs do not bind nonspecifically. For this purpose, the sequences can be 90–100% complementary. This property of complementary sequences being able to bind specifically to each other is used for primer binding, for example, in Northern or Southern blotting, or in PCR or RT-PCR.

[0692] Short oligonucleotides in conserved regions are advantageously used for hybridization. However, longer fragments or complete sequences of the nucleic acids of the present invention may also be used for hybridization. These “standard conditions” vary depending on the nucleic acid used (oligonucleotide, longer fragment, or complete sequence) or the type of nucleic acid used for hybridization (DNA or RNA). For example, the melting temperature of DNA:DNA hybrids is about 10°C lower than that of DNA:RNA hybrids of the same length.

[0693] For example, depending on the specific nucleic acid, standard conditions refer to a temperature of 42 to 58°C in a buffered aqueous solution with a concentration of 0.1 to 5 x SSC (1 x SSC = 0.15 M NaCl, 15 mM sodium citrate, pH 7.2), or additionally in the presence of 50% formamide (e.g., 42°C, 5 x SSC, 50% formamide). Advantageously, hybridization conditions for DNA:DNA hybrids are 0.1 × SSC and a temperature of about 20°C to 45°C, preferably about 30°C to 45°C. For DNA:RNA hybrids, hybridization conditions are advantageously 0.1 × SSC and a temperature of about 30°C to 55°C, preferably about 45°C to 55°C. These hybridization temperatures are examples of calculated melting temperatures for nucleic acids of about 100 nucleotides in length and with a G+C content of 50% in the absence of formamide. The experimental conditions for DNA hybridization have been described in relevant genetics textbooks (e.g., Sambrook et al., 1989) and can be calculated using molecular formulas known to those skilled in the art, depending on factors such as nucleic acid length, hybrid type, or G+C content. More information on hybridization can be obtained from textbooks such as Ausubel et al. (eds), (1985), and Brown (ed) (1991).

[0694] “Hybridization” can be carried out under particularly stringent conditions. Such hybridization conditions are described, for example, in Sambrook (1989), or in Current Protocols in Molecular Biology, John Wiley & Sons, NY (1989), 6.3.1–6.3.6.

[0695] As used herein, the term hybridization or hybridization under certain conditions is intended to describe the conditions under which hybridization and washing occur, under which significantly identical or homologous nucleotide sequences remain bound to each other. These conditions are such that sequences with at least about 70%, for example, at least about 80%, and for example, at least about 85%, 90%, or 95% identity remain bound to each other. Definitions of low-tightness, medium-tightness, and high-tightness hybridization conditions are provided herein.

[0696] Those skilled in the art can select suitable hybridization conditions with minimal experiments, as illustrated, for example, by Ausubel et al. (1995, Current Protocols in Molecular Biology, John Wiley & Sons, sections 2, 4, and 6). Furthermore, stringent conditions are described in Sambrook et al. (1989, Molecular Cloning: A Laboratory Manual, 2nd ed., Cold Spring Harbor Press, chapters 7, 9, and 11).

[0697] As used herein, the low-tightness conditions are defined as follows. Filter membranes containing DNA were pretreated at 40°C for 6 hours in a solution containing 35% formamide, 5xSSC, 50 mM Tris-HCl (pH 7.5), 5 mM EDTA, 0.1% PVP, 0.1% Ficoll, 1% BSA, and 500 μg / ml denatured salmon sperm DNA. Hybridization was performed in the same solution, modified as follows: 0.02% PVP, 0.02% Ficoll, 0.2% BSA, 100 μg / ml salmon sperm DNA, 10% (wt / vol) dextran sulfate, and using 5-20x10 6 32P-labeled probe. The filter membrane was incubated in a hybridization mixture at 40°C for 18–20 h, followed by washing at 55°C for 1.5 h. In a solution containing 2x SSC, 25 mM Tris-HCl (pH 7.4), 5 mM EDTA, and 0.1% SDS, the membrane was replaced with fresh solution and incubated again at 60°C for 1.5 h. The filter membrane was then blotted dry and subjected to autoradiography.

[0698] As used herein, the moderately stringent conditions are defined as follows. Filter membranes containing DNA were pretreated at 50°C for 7 hours in a solution containing 35% formamide, 5x SSC, 50 mM Tris-HCl (pH 7.5), 5 mM EDTA, 0.1% PVP, 0.1% Ficoll, 1% BSA, and 500 μg / ml denatured salmon sperm DNA. Hybridization was performed in the same solution, modified as follows: 0.02% PVP, 0.02% Ficoll, 0.2% BSA, 100 μg / ml salmon sperm DNA, 10% (wt / vol) dextran sulfate, and using 5-20x10 6 32P-labeled probe. The filter membrane was incubated in a hybridization mixture at 50°C for 30 hours, followed by washing at 55°C for 1.5 hours. In a solution containing 2x SSC, 25 mM Tris-HCl (pH 7.4), 5 mM EDTA, and 0.1% SDS, the membrane was incubated with fresh solution instead of the washing solution at 60°C for another 1.5 hours. The filter membrane was then blotted dry and subjected to autoradiography.

[0699] As used herein, the stringent conditions are as follows: DNA-containing filter membranes were pre-hybridized at 65°C for 8 hours to overnight in a buffer consisting of 6x SSC, 50 mM Tris-HCl (pH 7.5), 1 mM EDTA, 0.02% PVP, 0.02% Ficoll, 0.02% BSA, and 500 μg / ml denatured salmon sperm DNA. The membranes were then pre-hybridized in a buffer containing 100 μg / ml denatured salmon sperm DNA and 5-20 x 10⁻⁶ ppm of DNA. 6 The filter membrane was hybridized in a prehybridization mixture of cpm 32P labeled probes at 65°C for 48 hours. The membrane was then washed at 37°C for 1 hour in a solution containing 2x SSC, 0.01% PVP, 0.01% Ficoll, and 0.01% BSA. It was then washed in 0.1x SSC at 50°C for 45 minutes.

[0700] If the above conditions are not suitable (e.g., for interspecific hybridization), other low, medium and high stringency conditions well known in the art (e.g., for interspecific hybridization) may be used.

[0701] A detection kit for the nucleic acid sequence encoding the polypeptide of the present invention may include primers and / or probes specific to the nucleic acid sequence encoding the polypeptide, and a protocol for using the primers and / or probes to detect the nucleic acid sequence encoding the polypeptide in a sample. Such a detection kit can be used to determine whether a plant, organism, microorganism, or cell has been modified, i.e., whether it has been transformed with the sequence encoding the polypeptide.

[0702] To test the function of a variant DNA sequence according to one embodiment of this article, the target sequence is operatively linked to an optional or screenable marker gene, and the expression of the reporter gene is tested in a transient expression analysis using microorganisms or protoplasts or in stably transformed plants.

[0703] The present invention also relates to derivatives of specifically disclosed or derivable nucleic acid sequences.

[0704] Therefore, the additional nucleic acid sequences according to the invention may be derived from the sequences specifically disclosed herein and may be distinguished by one or more (e.g., 1 to 10) nucleotides, such as 1 to 20, particularly 1 to 15 or 5 to 10, additions, substitutions, insertions or deletions, and may also encode polypeptides having the desired properties.

[0705] The invention also includes nucleic acid sequences containing so-called silent mutations or altered sequences, depending on the codon usage of a particular original or host organism, compared to the specifically stated sequences.

[0706] According to specific embodiments of the invention, variant nucleic acids can be prepared to adapt their nucleotide sequences to a particular expression system. For example, bacterial expression systems are known to express polypeptides more efficiently if the amino acids are encoded by specific codons. Due to the degeneracy of the genetic code, more than one codon can encode the same amino acid sequence, and multiple nucleic acid sequences can encode the same protein or polypeptide; all these DNA sequences are covered in one embodiment herein. Where appropriate, the nucleic acid sequence encoding the polypeptide described herein can be optimized to increase expression in host cells. For example, the nucleic acid of one embodiment herein can be synthesized using host-specific codons to improve expression.

[0707] The present invention also covers naturally occurring variants of the sequences described herein, such as splice variants or allelic variants.

[0708] The allele variant has at least 60% homology across the entire amino acid range at the derived amino acid level, preferably at least 80% homology, and very particularly preferably at least 90% homology (for details regarding homology at the amino acid level, please refer to the information given above for peptides). Advantageously, the homology may be even higher in certain regions of the sequence.

[0709] The present invention also relates to sequences that can be obtained by conserved nucleotide substitution (i.e., as a result, the amino acid in question is replaced by an amino acid having the same charge, size, polarity and / or solubility).

[0710] This invention also relates to molecules derived from specifically disclosed nucleic acids through sequence polymorphism. Such genetic polymorphism can exist in cells from different populations or from cells within a single population due to natural allelic variations. Allelic variants may also include functional equivalents. These natural variations typically produce changes of 1–5% in the nucleotide sequence of a gene. The polymorphism can lead to alterations in the amino acid sequence of the polypeptides disclosed herein. Allelic variants may also include functional equivalents.

[0711] Furthermore, derivatives should also be understood as homologs of the nucleic acid sequences according to the present invention, such as homologs of animals, plants, fungi, or bacteria, shortened sequences, single-stranded DNA or RNA encoding or non-coding DNA sequences. For example, at the DNA level, the homolog has at least 40%, preferably at least 60%, particularly preferably at least 70%, and very particularly preferably at least 80% homology in the entire DNA region given in the sequence specifically disclosed herein.

[0712] Furthermore, derivatives should be understood as, for example, fusions with promoters. Promoters added to the nucleotide sequence can be modified by at least one nucleotide exchange, at least one insertion, inversion, and / or deletion, without impairing the function or effectiveness of the promoter. Moreover, the effectiveness of promoters can be increased by altering their sequence, or by completely exchanging them with more effective promoters or even promoters from different genera of organisms.

[0713] d. Generation of functional polypeptide mutants

[0714] Furthermore, those skilled in the art are familiar with methods for generating functional mutants, namely, a nucleotide sequence encoding a polypeptide having at least 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity with any amino acid-related SEQ ID NO disclosed herein; and / or encoded by a nucleic acid molecule containing a nucleotide sequence having at least 70% sequence identity with any nucleotide-related SEQ ID NO disclosed herein.

[0715] Depending on the techniques used, those skilled in the art can introduce completely random or more targeted mutations into gene or non-coding nucleic acid regions (e.g., those important for regulating expression) and subsequently generate a genetic library. The molecular biological methods required for this purpose are known to those skilled in the art, for example, as described in Sambrook and Russell, Molecular Cloning, 3rd Edition, Cold Spring Harbor Laboratory Press, 2001.

[0716] Methods for modifying genes and thereby modifying the polypeptides encoded by them have long been known to those skilled in the art, for example:

[0717] - Site-specific mutagenesis, in which single or multiple nucleotides of a gene are replaced in a directed manner (Trower MK (Ed.) 1996; In vitro mutagenesis protocols. Humana Press, New Jersey).

[0718] - Saturation mutagenesis, in which the codon of any amino acid can be exchanged or added at any site in the gene (Kegler-Ebo DM, Docktor CM, DiMaio D (1994) Nucleic Acids Res 22:1593; Barettino D, Feigenbutz M, Valcárel R, Stunnenberg HG (1994) Nucleic Acids Res 22:541; Barik S (1995) Mol Biotechnol 3:1),

[0719] - Error-prone polymerase chain reaction, in which the nucleotide sequence is mutated by error-prone DNA polymerase (Eckert KA, Kunkel TA (1990) Nucleic Acids Res 18:3739);

[0720] -SeSaM method (sequence saturation method), in which preferred exchanges are prevented by polymerase. Schenk et al., Biospektrum, Vol. 3, 2006, 277-279.

[0721] - Gene propagation in mutant strains, where, for example, due to defects in DNA repair mechanisms, the mutation rate of nucleotide sequences increases (Greener A, Callahan M, Jerpseth B (1996) An efficient random mutagenesis technique using an E. coli mutator strain. In: Trower MK (Ed.) In vitro mutagenesis protocols. Humana Press, New Jersey), or

[0722] -DNA shuffling, in which a set of closely related genes are formed and digested, and these fragments are used as templates for polymerase chain reactions, in which the full-length mosaic gene is eventually generated through repeated strand separation and recombination (Stemmer WPC (1994) Nature 370:389; Stemmer WPC (1994) Proc Natl Acad Sci USA 91:10747).

[0723] Using so-called directed evolution (particularly described in Reetz MT and Jaeger KE (1999), Topics Curr Chem 200:31; Zhao H, Moore JC, Volkov AA, Arnold FH (1999), Methods for optimizing industrial polypeptides by directed evolution, In: Demain AL, Davies JE (Ed.) Manual of industrial microbiology and biotechnology. American Society for Microbiology), skilled workers can mass-produce functional mutants in a directed manner. To this end, in the first step, gene libraries of the respective polypeptides are first generated, for example, using the methods given above. The gene libraries are expressed in a suitable manner, for example, through bacterial or phage display systems.

[0724] The relevant genes in the host organism expressing the functional mutant (whose function largely corresponds to the desired trait) can be submitted to another mutation cycle. The mutation and selection or screening steps can be repeated iteratively until the functional mutant of the present invention possesses a sufficient degree of the desired trait. Using this iterative process, a limited number of mutations, such as 1, 2, 3, 4, or 5 mutations, can be performed in stages, and their effects on the activity under study can be evaluated and selected. The selected mutants can then be subjected to further mutation steps in the same manner. This significantly reduces the number of individual mutants to be studied.

[0725] The results of this invention also provide important information regarding the structure and sequence of the relevant polypeptides, which is essential for the targeted generation of other polypeptides with desired modified properties. In particular, so-called "hot spots" can be defined as sequence segments potentially suitable for modification of properties by introducing targeted mutations.

[0726] Information about the location of amino acid sequences can also be derived, where mutations that may have little effect on activity can occur, and these can be designated as potential “silent mutations”.

[0727] e. Constructs expressing the polypeptides of the present invention

[0728] In the context of this article, the following definitions apply:

[0729] "Gene expression" encompasses both "heterologous expression" and "overexpression," and involves gene transcription and the translation of mRNA into proteins. Overexpression refers to the production of gene products, measured as mRNA, peptide, and / or enzyme activity levels, in transgenic cells or organisms exceeding the levels found in non-transformed cells or organisms with similar genetic backgrounds.

[0730] As used herein, an "expression vector" refers to a nucleic acid molecule engineered using molecular biology methods and recombinant DNA technology to deliver foreign or exogenous DNA into a host cell. Expression vectors typically include the sequences required for correct transcription of the nucleotide sequence. The coding region usually encodes the target protein, but it can also encode RNA, such as antisense RNA, siRNA, etc.

[0731] As used herein, “expression vector” includes any linear or circular recombinant vector, including but not limited to viral vectors, bacteriophages, and plasmids. Those skilled in the art can select a suitable vector based on the expression system. In one embodiment, the expression vector includes a nucleic acid of the embodiments described herein, operably linked to at least one “regulatory sequence” that controls transcription, translation, initiation, and termination, such as a transcription promoter, operon, or enhancer, or an mRNA ribosome binding site, and optionally includes at least one selection marker. When the regulatory sequence functionally relates to the nucleic acid of the embodiments described herein, the nucleotide sequence is “operably linked.”

[0732] As used herein, “expression system” encompasses any combination of nucleic acid molecules required to express one, or co-express two or more, polypeptides in vivo or in vitro in a given expression host. The respective coding sequences may reside on a single nucleic acid molecule or vector, such as a vector containing multiple cloning sites, or on a polycistronic nucleic acid, or may be distributed across two or more physically distinct vectors. As a specific example, an operon may be mentioned comprising a promoter sequence, one or more operon sequences, and one or more structural genes, each encoding the enzyme described herein.

[0733] As used herein, the terms “amplifying” and “amplification” refer to the use of any suitable amplification method to generate or detect recombinants of naturally expressed nucleic acids, as described in detail below. For example, the present invention provides methods and reagents (e.g., specific degenerate oligonucleotide primer pairs, oligo-dT primers) for amplifying (e.g., by polymerase chain reaction, PCR) naturally expressed (e.g., genomic DNA or mRNA) or recombinant nucleic acids (e.g., cDNA) of the present invention in vivo, in vitro, or in vitro.

[0734] A "regulatory sequence" refers to a nucleic acid sequence that determines the expression level of the nucleic acid sequence in the embodiment described herein and can regulate the transcription rate of a nucleic acid sequence operatively linked to that regulatory sequence. Regulatory sequences include promoters, enhancers, transcription factors, promoter elements, etc.

[0735] According to the present invention, "promoter," "nucleic acid with promoter activity," or "promoter sequence" should be understood as referring to a nucleic acid that, when functionally linked to a nucleic acid to be transcribed, regulates the transcription of said nucleic acid. "Promoter" specifically refers to a nucleic acid sequence that controls the expression of a coding sequence by providing a binding site for RNA polymerase and other factors suitable for transcription, including but not limited to transcription factor binding sites, repressor and activator protein binding sites. The term "promoter" also includes the term "promoter regulatory sequence." A promoter regulatory sequence may include upstream and downstream elements that may affect transcription, RNA processing, or the stability of the associated coding nucleic acid sequence. Promoters include naturally derived and synthetic sequences. The coding nucleic acid sequence is typically located downstream of the promoter relative to the transcription direction initiating from the transcription start site.

[0736] In this context, "functional" or "operationally" linked is understood, for example, to refer to the sequential arrangement of one of the nucleic acids having a regulatory sequence. For example, a sequence with promoter activity, and the nucleic acid sequence to be transcribed, along with optional other regulatory elements (e.g., nucleic acid sequences that ensure transcription) and, for example, a terminator, arranged such that each regulatory element can perform its function after transcription of the nucleic acid sequence. This does not necessarily require a direct chemical link. Genetic control sequences, such as enhancer sequences, can even act on the target sequence from more distant locations or even from other DNA molecules. A preferred arrangement is one where the nucleic acid sequence to be transcribed is located downstream (i.e., at the 3' end) of the promoter sequence, thereby covalently linking the two sequences together. The distance between the promoter sequence and the nucleic acid sequence to be recombined can be less than 200 base pairs, or less than 100 base pairs, or less than 50 base pairs.

[0737] In addition to promoters and terminators, other examples of regulatory elements include: target sequences, enhancers, polyadenylation signals, selectable markers, amplification signals, origins of replication, etc. Suitable regulatory sequences are described, for example, in Goeddel, Gene Expression Technology: Methods in Enzymology 185, Academic Press, San Diego, CA (1990).

[0738] The term "constitutive promoter" refers to an unregulated promoter that allows for the continuous transcription of the nucleic acid sequence to which it is operatively linked.

[0739] As used herein, the term "operably linked" refers to the linking of polynucleotide elements in a functional relationship. A nucleic acid is "operably linked" when it is in a functional relationship with another nucleic acid sequence. For example, if a promoter or transcriptional regulatory sequence can influence the transcription of a coding sequence, then the promoter or transcriptional regulatory sequence is operably linked to the coding sequence. Operable linking means that the linked DNA sequences are typically adjacent. The nucleotide sequence associated with the promoter sequence can be homologous or heterologous in origin relative to the plant to be transformed. The sequence can also be wholly or partially synthetic. Regardless of origin, the nucleic acid sequence associated with the promoter sequence will be expressed or silenced depending on the nature of the promoter linked after binding to the polypeptide of the embodiments described herein. The associated nucleic acid can encode a protein that needs to be expressed or repressed throughout the organism or in a specific tissue, cell, or cell compartment at all times or alternatively at specific times. This nucleotide sequence specifically encodes a protein that confers the desired phenotypic trait to the host cell or organism altered or transformed by it. More specifically, the associated nucleotide sequence results in the production of one or more target products as defined herein in the cell or organism. In particular, the nucleotide sequence encodes a polypeptide having enzymatic activity as defined herein.

[0740] The nucleotide sequences described above can be part of an "expression cassette". The terms "expression cassette" and "expression construct" are used synonymously. A (preferred recombinant) expression construct contains a nucleotide sequence that encodes a polypeptide according to the invention and is under the genetic control of a regulatory nucleic acid sequence.

[0741] In the method applied according to the present invention, the expression cassette may be an "expression vector", particularly a part of a recombinant expression vector.

[0742] According to the present invention, "expression unit" should be understood as a nucleic acid with expressive activity, which contains a promoter as defined herein, and regulates expression upon functional linkage with a nucleic acid or gene to be expressed, i.e., transcription and translation of said nucleic acid or gene. Therefore, it is also referred to in this respect as a "regulatory nucleic acid sequence". In addition to promoters, other regulatory elements, such as enhancers, may also be present.

[0743] According to the present invention, an "expression cassette" or "expression construct" should be understood as an expression unit functionally linked to a nucleic acid or gene to be expressed. Therefore, in contrast to an expression unit, an expression cassette contains not only nucleic acid sequences that regulate transcription and translation, but also nucleic acid sequences that are expressed as proteins due to transcription and translation.

[0744] In the context of this invention, the terms "expression" or "overexpression" describe the generation or increase of intracellular activity of one or more polypeptides encoded by corresponding DNA in a microorganism. For this purpose, for example, a gene may be introduced into the organism, an existing gene may be replaced with another gene, the copy number of a gene may be increased, a strong promoter may be used, or a gene encoding a corresponding polypeptide with high activity may be used. Optionally, these measures may be combined.

[0745] Preferably, such constructs according to the invention include a promoter upstream of the respective coding sequence 5' and a terminator sequence downstream of the respective coding sequence 3', as well as optionally other common regulatory elements, each operatively connected to the coding sequence.

[0746] The nucleic acid constructs according to the invention specifically comprise a sequence encoding a polypeptide, such as derived from the amino acid-related SEQ ID NO or its inverse complementary sequence as described herein, or derivatives and homologs thereof, and is operatively or functionally linked to one or more regulatory signals for advantageous control, for example, increasing gene expression.

[0747] In addition to these regulatory sequences, the natural regulation of these sequences may still exist before the actual structural genes, and optionally may have been genetically modified so that natural regulation has been turned off and gene expression is enhanced. However, nucleic acid constructs can also have simpler constructions, i.e., no additional regulatory signals are inserted before the coding sequence, and the natural promoter with regulatory function has not been removed. Instead, the natural regulatory sequences are mutated so that regulation no longer occurs and gene expression increases.

[0748] Preferred nucleic acid constructs advantageously also include one or more previously mentioned "enhancer" sequences functionally linked to a promoter, which enable enhanced expression of the nucleic acid sequence. Other advantageous sequences, such as other regulatory elements or terminators, may also be inserted at the 3' end of the DNA sequence. One or more copies of the nucleic acid according to the invention may be present in the construct. Optionally, other markers, such as genes complementary to auxotrophic or antibiotic resistance genes, may also be present in the construct for selection.

[0749] Examples of suitable regulatory sequences exist in promoters, such as cos, tac, trp, tet, trp-tet, lpp, lac, lpp-lac, and lacI. q , T7, T5, T3, gal, trc, ara, rhaP(rhaP BAD SP6, lambda-P R Or lambda-P LIn promoters, they are advantageously used in Gram-negative bacteria. Other advantageous regulatory sequences are found, for example, in the Gram-positive promoters amy and SpO2, and in yeast or fungal promoters ADC1, MFalpha, AC, P-60, CYC1, GAPDH, TEF, rp28, and ADH. Artificial promoters can also be used for regulation.

[0750] To facilitate expression in a host organism, nucleic acid constructs are advantageously inserted into vectors, such as plasmids or phages, enabling optimal gene expression in the host. Besides plasmids and phages, vectors should be understood to include all other vectors known to those skilled in the art, such as viruses like SV40, CMV, baculoviruses and adenoviruses, transposons, IS elements, phages, granules, and linear or circular DNA or artificial chromosomes. These vectors are capable of autonomous replication in the host organism or replication via chromosomes. These vectors represent a further development of the invention. Binary or CPO integration vectors are also suitable.

[0751] Suitable plasmids include, for example, those for *E. coli* pLG338, pACYC184, pBR322, pUC18, pUC19, pKC30, pRep4, pHS1, pKK223-3, pDHE19.2, pHS2, pPLc236, pMBL24, pLG200, pUR290, and pIN-III. 113 -B1, λgt11, or pBdCI; Streptomyces pIJ101, pIJ364, pIJ702, or pIJ361; Bacillus pUB110, pC194, or pBD214; Corynebacterium pSA77 or pAJ667; Fungi pALS1, pIL2, or pBB116; Yeast 2alphaM, pAG-1, YEp6, YEp13, or pEMBLYe23; or Plant pLGV23, pGHlac + The plasmids mentioned above are a small selection of possible plasmids, including pBIN19, pAK2004, and pDH51. Other plasmids are well known to those skilled in the art and can be found, for example, in the book Cloning Vectors (Eds. Pouwels PH et al. Elsevier, Amsterdam-New York-Oxford, 1985, ISBN 0 444 904018).

[0752] In further development of the vector, vectors containing the nucleic acid constructs of the present invention or the nucleic acids of the present invention can also be advantageously introduced into microorganisms in the form of linear DNA and integrated into the genome of the host organism via heterologous or homologous recombination. This linear DNA can consist of linearized vectors such as plasmids, or solely of the nucleic acid constructs or nucleic acids of the present invention.

[0753] For optimal expression of heterologous genes in an organism, it is advantageous to modify the nucleic acid sequence to match the specific “codon usage” used in the organism. “Codon usage” can be readily determined by computer evaluation of other known genes in the organism under discussion.

[0754] The expression cassette according to the invention is generated by fusing a suitable promoter to a suitable coding nucleotide sequence and a terminator or polyadenylation signal. Conventional recombination and cloning techniques are used for this purpose, as described in, for example, T. Maniatis, EFFritsch and J. Sambrook, Molecular Cloning: A Laboratory Manual, Cold Spring Harbor Laboratory, Cold Spring Harbor, NY (1989); TJ Silhavy, MLBerman and LWEnquist, Experiments with Gene Fusions, Cold Spring Harbor Laboratory, Cold Spring Harbor, NY (1984); and Ausubel, FM et al., Current Protocols in Molecular Biology, Greene Publishing Assoc. and Wiley Interscience (1987).

[0755] To facilitate expression in a suitable host organism, recombinant nucleic acid constructs or gene constructs are advantageously inserted into host-specific vectors, enabling optimal gene expression in the host. Vectors are well-known to those skilled in the art and can be found, for example, in "cloning vectors" (Pouwels PH et al., Ed., Elsevier, Amsterdam-New York-Oxford, 1985).

[0756] Alternative embodiments of the present invention provide a method for “altering gene expression in host cells.” For example, in certain contexts (e.g., exposure to certain temperatures or culture conditions), polynucleotides of the present invention can be enhanced, overexpressed, or induced in host cells or host organisms.

[0757] The altered expression of the polynucleotides described herein can also result in ectopic expression, which is a different expression pattern in altered and control or wild-type organisms. The alteration in expression occurs due to the interaction of the peptide of one embodiment of this invention with an exogenous or endogenous regulator or due to chemical modification of the peptide. The term also refers to the altered expression pattern of the polynucleotides of the embodiments described herein, which is altered to below detectable levels or completely inhibited in activity.

[0758] In one embodiment, this document also provides isolated, recombinant, or synthetic polynucleotides encoding the polypeptide or variant polypeptide provided herein.

[0759] In one embodiment, multiple nucleic acid sequences encoding polypeptides are co-expressed in a single host, particularly under the control of different promoters. In another embodiment, multiple nucleic acid sequences encoding polypeptides may be present on a single transformation vector, or separate vectors may be used and transformants containing two chimeric genes may be selected for simultaneous co-transformation. Similarly, one or more polypeptide-encoding genes may be expressed together with other chimeric genes in a single plant, cell, microorganism, or organism.

[0760] f. Host applicable to the present invention

[0761] Depending on the context, the term "host" can refer to a wild-type host or a genetically modified recombinant host, or both.

[0762] In principle, all prokaryotes or eukaryotes can be considered as hosts or recombinant host organisms for the nucleic acids or nucleic acid constructs according to the present invention.

[0763] Using the vector according to the invention, recombinant hosts can be produced, which can be transformed, for example, with at least one vector according to the invention, and can be used to produce polypeptides according to the invention. Advantageously, the recombinant construct according to the invention as described above is introduced into and expressed in a suitable host system. Preferably, common cloning and transfection methods known to those skilled in the art, such as co-precipitation, protoplast fusion, electroporation, retroviral transfection, etc., are used to express the nucleic acid in their respective expression systems. Suitable systems are described in Current Protocols in Molecular Biology, F. Ausubelet et al., Ed., Wiley Interscience, New York 1997, or Sambrook et al. Molecular Cloning: A Laboratory Manual, 2nd edition, Cold Spring Harbor Laboratory, Cold Spring Harbor Laboratory Press, Cold Spring Harbor, NY, 1989.

[0764] Advantageously, microorganisms such as bacteria, fungi, or yeasts are used as host organisms. Advantageously, Gram-positive or Gram-negative bacteria are used, preferably those belonging to the families Enterobacteriaceae, Pseudomonadaceae, Rhizobiaceae, Streptomycetaceae, Streptococcaceae, or Nocardiaceae, and particularly preferably those belonging to the genera Escherichia, Pseudomonas, Streptomyces, Lactococcus, Nocardia, Burkholderia, Salmonella, Agrobacterium, Clostridium, or Rhodococcus. The genus and species *Escherichia coli* are particularly preferred. Furthermore, other advantageous bacteria have been found in the alpha-proteobacteria, beta-proteobacteria, or gamma-proteobacteria groups. Advantageously, yeasts such as *Saccharomyces* or the Pichia family are also suitable hosts.

[0765] Alternatively, the entire plant or plant cell can be used as a natural or recombinant host. As non-limiting examples, the following plants or cells derived from them may be mentioned: the genus *Nicotiana*, particularly *Nicotiana abenthamiana* and *Nicotiana tabacum* (tobacco); and the genus *Arabidopsis*, particularly *Arabidopsis thaliana*.

[0766] Depending on the host organism, the organism used in the method according to the invention is grown or cultured in a manner known to those skilled in the art. Culture can be carried out in batches, semi-batch, or continuously. Nutrients can be provided at the start of fermentation or later, semi-continuously, or continuously. This is also described in more detail below.

[0767] g. Recombinant production of the polypeptide according to the present invention

[0768] The present invention further relates to a method for recombinantly producing polypeptides or functional biologically active fragments thereof according to the invention, wherein microorganisms that produce polypeptides are cultured, and optionally, expression of the polypeptides is induced by applying at least one inducer for gene expression, and the polypeptides are isolated from the culture. If desired, the polypeptides can also be produced on an industrial scale in this manner.

[0769] The microorganisms produced according to the present invention can be cultured continuously or discontinuously using batch culture, fed-batch culture, or repeated fed-batch culture. An overview of known culture methods can be found in Chmiel's textbook (Bioprozesstechnik 1.Einführungin die Bioverfahrenstechnik [Bioprocess technology 1.Introduction to bioprocess technology] (Gustav Fischer Verlag, Stuttgart, 1991)) or Storhas's textbook (Bioreaktoren und periphere Einrichtungen [Bioreactors and peripheralequipment] (Vieweg Verlag, Braunschweig / Wiesbaden, 1994)).

[0770] The culture medium used must be appropriately suited to the requirements of each strain. Descriptions of culture media for various microorganisms are provided in the "Manual of Methods for General Bacteriology" of the American Society for Bacteriology (Washington DC, USA, 1981).

[0771] These culture media, which can be used according to the present invention, typically contain one or more carbon sources, nitrogen sources, inorganic salts, vitamins and / or trace elements.

[0772] Preferred carbon sources are sugars, such as monosaccharides, disaccharides, or polysaccharides. Good carbon sources include, for example, glucose, fructose, mannose, galactose, ribose, sorbitol, ribulose, lactose, maltose, sucrose, raffinose, starch, or cellulose. Sugars can also be added to the culture medium via complex compounds (e.g., molasses) or other byproducts of saccharification. Adding mixtures of different carbon sources is also advantageous. Other possible carbon sources are oils and fats, such as soybean oil, sunflower oil, peanut oil, and coconut oil; fatty acids such as palmitic acid, stearic acid, or linoleic acid; alcohols such as glycerol, methanol, or ethanol; and organic acids such as acetic acid or lactic acid.

[0773] Nitrogen sources are typically organic or inorganic nitrogen compounds or materials containing these compounds. Examples of nitrogen sources include ammonia or ammonium salts, such as ammonium sulfate, ammonium chloride, ammonium phosphate, ammonium carbonate, or ammonium nitrate, nitrates, urea, amino acids, or complex nitrogen sources, such as corn steep liquor, soybean flour, soybean protein, yeast extract, meat extract, etc. Nitrogen sources can be used alone or as a mixture.

[0774] Inorganic salt compounds that can be present in culture media include chlorides, phosphorus, or sulfates of calcium, magnesium, sodium, cobalt, molybdenum, potassium, manganese, zinc, copper, and iron.

[0775] Inorganic sulfur-containing compounds, such as sulfates, sulfites, dithionites, tetrathionites, thiosulfates, and sulfides, as well as organic sulfur compounds, such as thiols and thiols, can be used as sulfur sources.

[0776] Phosphoric acid, potassium dihydrogen phosphate, or dipotassium hydrogen phosphate, or their corresponding sodium-containing salts, can be used as phosphorus sources.

[0777] Chelating agents can be added to the culture medium to retain metal ions in solution. Particularly suitable chelating agents include dihydroxyphenols, such as catechol or protocatechuic acid esters, or organic acids, such as citric acid.

[0778] The fermentation medium used according to the present invention typically also contains other growth factors, such as vitamins or growth promoters, including, for example, biotin, riboflavin, thiamine, folic acid, niacin, pantothenic acid, and pyridoxine. The growth factors and salts are often derived from components of complex culture media, such as yeast extract, molasses, corn steep liquor, etc. In addition, suitable precursors may be added to the medium. The exact composition of the compounds in the medium depends largely on the specific experiment and is determined individually for each case. Information on medium optimization can be found in the textbook "Applied Microbiol. Physiology, A Practical Approach" (Ed. PMRhodes, PFStanbury, IRL Press (1997), pp. 53-73, ISBN 0 19 963577 3). Growth media are also available from commercial suppliers such as Standard 1 (Merck) or BHI (Brain Heart Infusion, DIFCO).

[0779] All components of the culture medium are sterilized by heating (at 1.5 bar and 121°C for 20 minutes) or by aseptic filtration. These components can be sterilized together or individually as needed. All components of the culture medium can be given at the start of the culture or added continuously or in batches.

[0780] The culture temperature is typically between 15°C and 45°C, preferably between 25°C and 40°C, and can be varied or kept constant during the experiment. The pH of the medium should be in the range of 5 to 8.5, preferably around 7.0. The pH during growth can be controlled by adding alkaline compounds (e.g., sodium hydroxide, potassium hydroxide, ammonia, or ammonia solution) or acidic compounds (e.g., phosphoric acid or sulfuric acid). Antifoaming agents such as fatty acid polyethylene glycol esters can be used to control foaming. To maintain plasmid stability, suitable selective substances such as antibiotics can be added to the culture medium. To maintain aerobic conditions, oxygen or an oxygen-containing gas mixture (e.g., ambient air) is supplied to the culture. The culture temperature is typically in the range of 20°C to 45°C. Continue culturing until the maximum amount of the desired product is formed. This usually takes between 10 and 160 hours.

[0781] The fermentation broth is then further processed. Depending on the needs, the biomass can be completely or partially removed from the fermentation broth, or it can be left entirely in it, by separation techniques such as centrifugation, filtration, decantation, or a combination of these methods.

[0782] If the polypeptide is not secreted in the culture medium, the cells can also be lysed, and the product can be obtained from the lysate using known methods for protein separation. Cells can be optionally destroyed by high-frequency ultrasound, high pressure (e.g., in a French press), by osmosis, by the action of detergents, lysing enzymes, or organic solvents, by a homogenizer, or by a combination of these methods.

[0783] Peptides can be purified using known chromatographic techniques, such as molecular sieve chromatography (gel filtration), Q-agarose chromatography, ion exchange chromatography, and hydrophobic chromatography, as well as other conventional techniques such as ultrafiltration, crystallization, salting out, dialysis, and natural gel electrophoresis. Suitable methods are described, for example, in Cooper, TG, Biochemimsche Arbeitsmethoden [Biochemical Processes], Verlag Walter de Gruyter, Berlin, New York, or Scopes, R., Protein Purification, Springer Verlag, New York, Heidelberg, Berlin.

[0784] For the isolation of recombinant proteins, the use of a carrier system or oligonucleotide may be advantageous, which extends cDNA by a defined nucleotide sequence and thus encodes an altered polypeptide or fusion protein, for example, for easier purification. Suitable modifications of this type are, for example, so-called “tags” that act as anchors, such as modifications known as hexahistine anchors or epitopes that can be recognized as antibody antigens (e.g., described in Harlow, E. and Lane, D., 1988, Antibodies: A Laboratory Manual. Cold Spring Harbor (NY) Press). These anchors can be used to attach proteins to solid supports, such as polymer matrices, which can be used, for example, as packing material in chromatographic columns, or on microtiter plates or other supports.

[0785] These anchors can also be used to identify proteins. To identify proteins, conventional markers, such as fluorescent dyes, enzyme markers (which react with a substrate to form a detectable reaction product), or radiolabels, can be used alone or in combination with anchors to derivatize proteins.

[0786] h. Immobilization of peptides

[0787] The enzymes or polypeptides according to the invention can be used in free form or immobilized in the methods described herein. Immobilized enzymes are enzymes immobilized on an inert support. Suitable support materials and enzymes immobilized thereon are known from EP-A-1149849, EP-A-1069183, and DE-OS100193773 and the references cited therein. In this regard, reference is made to the full disclosure of these documents. Suitable support materials include, for example, clay, clay minerals such as kaolinite, diatomaceous earth, perlite, silica, alumina, sodium carbonate, calcium carbonate, cellulose powder, anion exchanger materials, synthetic polymers such as polystyrene, acrylic resins, phenolic resins, polyurethanes, and polyolefins such as polyethylene and polypropylene. For the preparation of loaded enzymes, the support material is generally used in the form of finely divided particles, preferably porous. The particle size of the support material is generally no greater than 5 mm, particularly no greater than 2 mm (particle size distribution profile). Similarly, when using dehydrogenases as whole-cell catalysts, free or immobilized forms can be selected. Carrier materials include, for example, calcium alginate and carrageenan. Enzymes and cells can also be directly cross-linked with glutaraldehyde (cross-linked with CLEAs). Corresponding and other immobilization techniques are described, for example, in J. Lalonde and A. Margolin, "Immobilization of Enzymes" in K. Drauz and H. Waldmann, Enzyme Catalysis in Organic Synthesis 2002, Vol. III, 991-1032, Wiley-VCH, Weinheim. Rehm et al. (Ed.) Biotechnology, 2nd Edn, Vol 3, Chapter 17, VCH, Weinheim provides further information on biotransformation and bioreactors for carrying out the methods according to the invention.

[0788] i. Reaction conditions of the biocatalytic generation method of the present invention

[0789] The reaction of the present invention can be carried out under in vivo or in vitro conditions.

[0790] At least one polypeptide / enzyme present in a single step of the method of the present invention or the multi-step method defined above may be naturally present in living cells, or in harvested cells (i.e., under in vivo conditions), dead cells, permeabilized cells, crude cell extracts, purified extracts, or in a substantially pure or completely pure form (i.e., under in vitro conditions), or may be recombined to produce one or more enzymes. The at least one enzyme may be present in solution or as an enzyme immobilized on a carrier. One or more enzymes may be present simultaneously in soluble and / or immobilized forms.

[0791] The method according to the invention can be carried out in common reactors known to those skilled in the art and can be carried out on various scales, from laboratory scale (a few milliliters to tens of liters of reaction volume) to industrial scale (a few liters to thousands of cubic meters of reaction volume). A chemical reactor can be used if the peptide is used in the form of encapsulation through non-living, optionally permeabilized cells, as a more or less purified cell extract, or in a purified form. Chemical reactors typically allow control of the amount of at least one enzyme, at least one substrate, pH, temperature, and the circulation of the reaction medium. When at least one peptide / enzyme is present in living cells, the process will be fermentation. In this case, biocatalytic production will be carried out in a bioreactor (fermenter) where parameters necessary for suitable survival conditions for living cells (e.g., nutrient-rich culture medium, temperature, aeration, aerobic or anaerobic or other gases, antibiotics, etc.) can be controlled. Those skilled in the art are familiar with chemical or bioreactors, for example, procedures for scaling up chemical or biotechnological methods from laboratory to industrial scale or optimizing process parameters, which are also extensively described in the literature (for biotechnological methods, see, for example, Crueger und Crueger, Biotechnologie – Lehrbuch der angewandten Mikrobiologie, 2. Ed., R. Oldenbourg Verlag, Munich, Wien, 1984).

[0792] Cells containing at least one enzyme can be permeated by physical or mechanical means, such as ultrasound or radio frequency pulses, high-pressure cell lysis (French press), or by chemical means, such as hypotonic media present in the culture medium, lysing enzymes, and detergents, or a combination of these methods. Examples of detergents are digitoxin, n-dodecyl maltodextrin, octyl glycoside, etc. X-100, 20, Deoxycholate, CHAPS (3-[(3-chloroamidopropyl)dimethylammonium]-1-propanesulfonate), P40 (ethylphenol poly(ethylene glycol ether)), etc.

[0793] Instead of living cells, non-living cell biomass containing the desired biocatalyst can also be used in the biotransformation reaction of this invention.

[0794] If at least one enzyme is immobilized, it is ligated to an inert carrier as described above.

[0795] The conversion reaction can be carried out in batches, semi-batches, or continuously. Reactants (and optional nutrients) can be provided at the start of the reaction, or they can be provided subsequently in a semi-continuous or continuous manner.

[0796] Depending on the specific reaction type, the reactions of this invention can be carried out in aqueous, aqueous-organic, or non-aqueous reaction media.

[0797] Aqueous or aqueous-organic media may contain suitable buffer solutions to adjust the pH to 5 to 11, such as 6 to 10.

[0798] In aqueous-organic media, organic solvents that are miscible, partially miscible, or immiscible with water can be used. Non-limiting examples of suitable organic solvents are listed below. Further examples are mono- or poly-aryl, aromatic or aliphatic alcohols, particularly poly-aliphatic alcohols such as glycerol.

[0799] Non-aqueous media may contain substantially no water, that is, will contain less than about 1% by weight or 0.5% by weight of water.

[0800] Biocatalytic methods can also be carried out in organic non-aqueous media. Suitable organic solvents include, for example, aliphatic hydrocarbons having 5 to 8 carbon atoms, such as pentane, cyclopentane, hexane, cyclohexane, heptane, octane, or cyclooctane; aromatic hydrocarbons, such as benzene, toluene, xylene, chlorobenzene, or dichlorobenzene; aliphatic acyclic hydrocarbons and ethers, such as diethyl ether, methyl tert-butyl ether, ethyl tert-butyl ether, dipropyl ether, diisopropyl ether, dibutyl ether; or mixtures thereof.

[0801] The concentration of reactants / substrate can be adapted to optimal reaction conditions, which can depend on the specific enzyme being applied. For example, the initial substrate concentration can be 0.1 to 0.5 M, or, for example, 10 to 100 mM.

[0802] The reaction temperature can be adapted to optimal reaction conditions, which can depend on the specific enzyme used. For example, the reaction can be carried out at temperatures ranging from 0 to 70°C, such as 20 to 50 or 25 to 40°C. Examples of reaction temperatures are approximately 30°C, approximately 35°C, approximately 37°C, approximately 40°C, approximately 45°C, approximately 50°C, approximately 55°C, and approximately 60°C.

[0803] The process can continue until equilibrium is reached between the substrate and the subsequent product, but it can be stopped earlier. Typical process times range from 1 minute to 25 hours, particularly from 10 minutes to 6 hours, for example from 1 hour to 4 hours, and particularly from 1.5 hours to 3.5 hours. These parameters are non-limiting examples of suitable process conditions.

[0804] If the host is a genetically modified plant, it can provide optimal growth conditions, such as optimal light, water, and nutrients.

[0805] k. Product separation

[0806] The method of the present invention may further include the step of recovering a final product or intermediate product, which may optionally be a substantially pure form of a stereoisomer or enantiomer. The term "recovery" includes the extraction, harvesting, separation, or purification of a compound from a culture medium or reaction medium. The recovery of a compound can be carried out according to any conventional separation or purification method known in the art, including but not limited to treatment with conventional resins (e.g., anion or cation exchange resins, nonion adsorption resins, etc.), treatment with conventional adsorbents (e.g., activated carbon, silica, silica gel, cellulose, alumina, etc.), pH alteration, solvent extraction (e.g., using conventional solvents such as alcohols, ethyl acetate, hexane, etc.), distillation, dialysis, filtration, concentration, crystallization, recrystallization, pH adjustment, lyophilization, etc.

[0807] The identity and purity of the separated products can be determined using known techniques, such as high-performance liquid chromatography (HPLC), gas chromatography (GC), spectroscopy (e.g., IR, UV, NMR), staining methods, TLC, NIRS, enzyme or microbial assays (see, for example: Patek et al. (1994) Appl. Environ. Microbiol. 60: 133-140; Malakhova et al. (1996) Biotekhnologiya 1127-32; und Schmidt et al. (1998) Bioprocess Engineer. 19: 67-70; Ullmann's Encyclopedia of Industrial Chemistry (1996) Bd. A27, VCH: Weinheim, pp. 89-90, 521-540, 540-547, 559-566, 575-581). S.581-587; Michal, G (1999) Biochemical Pathways: An Atlas of Biochemistry and Molecular Biology, John Wiley and Sons; Fallon, A. et al. (1987) Applications of HPLC in Biochemistry in: Laboratory Techniques in Biochemistry and Molecular Biology, Bd. 17.).

[0808] Cyclic terpenoids produced by any of the methods described herein can be converted into derivatives, such as, but not limited to, hydrocarbons, esters, amides, glycosides, ethers, epoxides, aldehydes, ketones, alcohols, diols, acetals, or ketals. Terpenoid derivatives can be obtained by chemical methods, such as, but not limited to, oxidation, reduction, alkylation, acylation, and / or rearrangement. Alternatively, terpenoid derivatives can be obtained by biochemical methods by contacting the terpenoid with an enzyme, such as, but not limited to, oxidoreductases, monooxygenases, dioxygenases, and transferases. Biochemical transformation can be performed in vitro using isolated enzymes, enzymes derived from lysed cells, or in vivo using whole cells.

[0809] l. Fermentation production of terpenes / terpenoids, such as hemispherane-type compounds.

[0810] The present invention also relates to a method for the fermentation production of terpenes / terpenoids such as hemispherane-type compounds.

[0811] The fermentation used according to the invention can be carried out, for example, in stirred fermenters, bubble columns, and loop reactors. For a comprehensive overview of possible method types, including stirrer types and geometries, see "Chmiel: Bioprozesstechnik: Einfuhrung in die Bioverfahrenstechnik, Band 1". Typical variations available in the method of the invention are those known to those skilled in the art or explained, for example, in "Chmiel, Hammes and Bailey: Biochemical Engineering", such as batch, fed-batch, repeatedly fed-batch, or continuous fermentation, with or without biomass recovery. Depending on the production strain, air, oxygen, carbon dioxide, hydrogen, nitrogen, or suitable gas mixtures can be injected to achieve good yields (YP / S).

[0812] The culture medium used must be appropriately suited to the requirements of the specific strain. Descriptions of various microbial culture media are provided in the "Manual of Methods for General Bacteriology" of the American Society for Bacteriology (Washington DC, USA, 1981).

[0813] These culture media, which can be used according to the present invention, typically contain one or more carbon sources, nitrogen sources, inorganic salts, vitamins and / or trace elements.

[0814] Preferred carbon sources are sugars, such as monosaccharides, disaccharides, or polysaccharides. Very good carbon sources include, for example, glucose, fructose, mannose, galactose, ribose, sorbitol, ribulose, lactose, maltose, sucrose, raffinose, starch, or cellulose. Sugars can also be added to the culture medium via complex compounds (e.g., molasses) or other byproducts of saccharification. Adding mixtures of various carbon sources is also advantageous. Other possible sources of carbon are oils and fats, such as soybean oil, sunflower oil, peanut oil, and coconut oil; fatty acids such as palmitic acid, stearic acid, or linoleic acid; alcohols such as glycerol, methanol, or ethanol; and organic acids such as acetic acid or lactic acid.

[0815] Nitrogen sources are typically organic or inorganic nitrogen compounds or materials containing these compounds. Examples of nitrogen sources include ammonia or ammonium salts, such as ammonium sulfate, ammonium chloride, ammonium phosphate, ammonium carbonate, or ammonium nitrate, nitrates, urea, amino acids, or complex nitrogen sources, such as corn steep liquor, soybean flour, soybean protein, yeast extract, meat extract, etc. Nitrogen sources can be used alone or as a mixture.

[0816] Inorganic salt compounds that can be present in culture media include chlorides, phosphates, or sulfates of calcium, magnesium, sodium, cobalt, molybdenum, potassium, manganese, zinc, copper, and iron.

[0817] Inorganic sulfur-containing compounds, such as sulfates, sulfites, dithionites, tetrathionites, thiosulfates, and sulfides, as well as organic sulfur compounds, such as thiols and thiols, can be used as sulfur sources.

[0818] Phosphoric acid, potassium dihydrogen phosphate, or dipotassium hydrogen phosphate, or their corresponding sodium-containing salts, can be used as phosphorus sources.

[0819] Chelating agents can be added to the culture medium to retain metal ions in solution. Particularly suitable chelating agents include dihydroxyphenols, such as catechol or protocatechuic acid esters, or organic acids, such as citric acid.

[0820] The fermentation medium used according to the present invention may also contain other growth factors, such as vitamins or growth promoters, including, for example, biotin, riboflavin, thiamine, folic acid, niacin, pantothenic acid, and pyridoxine. Growth factors and salts are typically derived from complex components of the medium, such as yeast extract, molasses, corn steep liquor, etc. Additionally, suitable precursors may be added to the medium. The precise composition of compounds in the medium depends heavily on the specific experiment and must be determined individually for each specific case. Information on medium optimization can be found in the textbook "Applied Microbiol. Physiology, A Practical Approach" (1997). Growth media are also available from commercial suppliers, such as Standard 1 (Merck) or BHI (brain and heart infusion, DIFCO), etc.

[0821] All components of the culture medium are sterilized by heating (at 1.5 bar and 121°C for 20 minutes) or by aseptic filtration. These components can be sterilized together or individually as needed. All components of the culture medium can be given at the start of growth, or can be added continuously or in batches.

[0822] The culture temperature is typically between 15°C and 45°C, preferably between 25°C and 40°C, and can be kept constant or varied during the experiment. The pH of the medium should be in the range of 5 to 8.5, preferably around 7.0. The pH during growth can be controlled by adding alkaline compounds (e.g., sodium hydroxide, potassium hydroxide, ammonia, or ammonia solution) or acidic compounds (e.g., phosphoric acid or sulfuric acid). Antifoaming agents such as fatty acid polyethylene glycol esters can be used to control foaming. To maintain plasmid stability, suitable substances with selective action, such as antibiotics, can be added to the culture medium. To maintain aerobic conditions, oxygen or a mixture of oxygen-containing gases (e.g., ambient air) is supplied to the culture. The culture temperature is typically between 20°C and 45°C. Continue culturing until the maximum amount of the desired product is formed. This usually takes between 1 and 160 hours.

[0823] The method of the present invention may further include a step of recovering the terpene alcohol.

[0824] The term "recovery" includes the extraction, harvesting, separation, or purification of compounds from a culture medium. The recovery of compounds can be performed using any conventional separation or purification method known in the art, including but not limited to treatment with conventional resins (e.g., anion or cation exchange resins, nonion adsorption resins, etc.), treatment with conventional adsorbents (e.g., activated carbon, silica, silica gel, cellulose, alumina, etc.), pH alteration, solvent extraction (e.g., using conventional solvents such as alcohols, ethyl acetate, hexane, etc.), distillation, dialysis, filtration, concentration, crystallization, recrystallization, pH adjustment, lyophilization, etc.

[0825] Before the intended separation, the biomass in the fermentation broth can be removed. Methods for removing biomass are known to those skilled in the art, such as filtration, sedimentation, and flotation. Therefore, biomass can be removed, for example, by centrifuges, separators, decanters, filters, or in flotation equipment. To maximize the recovery of valuable products, washing the biomass, for example by percolation, is generally recommended. The choice of method depends on the biomass content and properties in the fermentation broth, as well as the interaction between the biomass and the valuable products.

[0826] In one implementation, the fermentation broth can be sterilized or pasteurized. In another implementation, the fermentation broth is concentrated. This concentration can be carried out in batches or continuously, depending on the needs. Pressure and temperature ranges should be selected to ensure that product damage is not caused and to minimize equipment and energy consumption. Skillful selection of pressure and temperature levels for multi-stage evaporation can be particularly energy-efficient.

[0827] The following examples are illustrative only and are not intended to limit the scope of the implementation schemes described herein.

[0828] After considering the disclosure provided herein, a variety of possible variations that will immediately become apparent to those skilled in the art also fall within the scope of this invention.

[0829] Experimental Section

[0830] The invention will now be described in more detail through the following embodiments.

[0831] a) Materials:

[0832] Unless otherwise stated, all chemical and biochemical materials, as well as microorganisms or cells, used in this article are commercially available products.

[0833] Unless otherwise stated, recombinant proteins are cloned and expressed using standard methods, such as those described, for example, in Sambrook, J., Fritsch, E.F. and Maniatis, T., Molecular cloning: A Laboratory Manual, 2008. nd Edition, Cold Spring Harbor Laboratory, Cold Spring Harbor Laboratory Press, Cold Spring Harbor, NY, 1989.

[0834] b) General Method

[0835] Preparation of cell-free protein fractions.

[0836] The expression vector was transformed into *E. coli* KRX cells (Promega Corporation, Madison, WI, USA), and transformed cells were selected on LB agar plates supplemented with appropriate antibiotics. Cells were then cultured in 25 mL of liquid LB medium supplemented with appropriate antibiotics at 37°C until an OD of 1 was reached. Expression of the recombinant protein was induced by 1 mM isopropyl-1-thio-β-D-galactopyranoside and 0.1% (w / v) L-rhamnose monohydrate, and cells were incubated at 25°C with gentle shaking for 24 hours.

[0837] Bacterial cells were harvested by centrifugation (5000g, 12 min) and sonicated (Sonics, Vibra Cell X 130 sonicator equipped with a 6 mm diameter microprobe; 3 20-second 20 kHz pulses, 80% of maximum power) on ice in 1.8 mL of 50 mM MOPSO buffer containing 15% glycerol at pH 7.4. The lysate was clarified by centrifugation (3500g, 8 min, 4 °C), and the resulting supernatant was cryopreserved and used as the enzyme source for in vitro assays.

[0838] In vitro enzyme analysis.

[0839] In an analytical solution containing one of the recombinant proteins, the protein fraction was incubated at 24°C with shaking at 230 rpm for 4 hours in a borosilicate glass and PTFE-sealed screw cap tube (11 mL capacity) (Wheaton, Millville, NJ 08332 USA). The analytical solution consisted of 20 μL of cell-free extract, 160 to 320 mg / L substrate (using a substrate stock solution of 40 g / L in DMSO), 1 mM cofactor (if relevant), and 50 mM MOPSO at pH 7.4, with a final volume of 0.5 to 1 mL. The analytical solution was extracted with 1 volume of methyl tert-butyl ether (MTBE) and analyzed by GC-MS as described below.

[0840] Whole-cell biotransformation assay.

[0841] Biotransformation of the compound was performed using *E. coli* cells expressing the recombinant enzyme. The expression vector was transformed into *E. coli* KRX cells (Promega Corporation, Madison, WI, USA), and transformed cells were selected on LB agar plates supplemented with appropriate antibiotics. Cells were initially cultured overnight at 30°C in 5 mL LB agar supplemented with 1% glucose and appropriate antibiotics. The next day, 20 mL of TB agar supplemented with appropriate antibiotics was inoculated at an initial optical density of 0.2 to 0.75. The culture was incubated in shake flasks at 37°C until an optical density of 1 to 4 was reached, and recombinant protein expression was induced by the addition of 0.1 mM isopropyl-1-thio-β-D-galactopyranoside IPTG and 0.1% rhamnose. The culture was then aliquoted into 12 mL glass tubes in 0.5 to 1 mL aliquots and incubated at 20°C with gentle shaking.

[0842] Ninety minutes after inducing recombinant protein expression, the substrate was added to each tube. The substrate was added to a final concentration of 0.25 to 1 g / L using a 40 g / L DMSO stock solution. Alternatively, it was prepared in water containing 150 mg / mL. An emulsion of 80 mg (Sigma-Aldrich) and 300 mg / mL substrate was prepared and added to the analytical solution to achieve a final substrate concentration of 12 mg / mL.

[0843] After incubation for 8 to 48 hours, the culture was extracted with 1 volume of MTBE and analyzed by GC-MS as described below.

[0844] Engineered bacterial cells were cultured under conditions that enabled the production of terpenoid compounds.

[0845] DP1205 *E. coli* cells were transformed with one or two expression plasmids carrying terpene biosynthesis genes and / or terpene-modifying enzymes. Transformed cells were cultured on LB agar plates with appropriate antibiotics (kanamycin (50 μg / mL) and / or chloramphenicol (34 μg / mL)). Single colonies were inoculated into 5 mL of liquid LB medium supplemented with the same antibiotics, 4 g / L glucose, and 10% (v / v) dodecane. The next day, 0.2 mL of the overnight culture was inoculated into 2 mL of TB medium supplemented with the same antibiotics and 10% (v / v) dodecane. The cultures were incubated at 37°C until an optical density of 3 was reached. Recombinant protein expression was then induced by the addition of 1 mM IPTG, and the cultures were incubated at 20°C for 72 hours.

[0846] The culture was then extracted with 1 volume of (MTBE), and the composition of the organic phase was analyzed by GC-MS as described below. For quantification, an internal standard (α-aldrich) was added to the extract prior to GC-MS analysis, and the concentration of the components was estimated based on a comparison of peak areas.

[0847] GC-MS analysis method.

[0848] Samples for whole-cell biotransformation assays were analyzed using an Agilent 7890A GC system connected to a 5975C Series Quality Select Detector (MSD) and equipped with a split / splitless injector (Agilent Technologies, CA).

[0849] The GC inlet temperature was set to 230°C. A 1.0 μL sample was injected in split mode (split ratio 20:1) and analyzed on a DB-5ms capillary column (30 m x 0.25 mm inner diameter x 0.25 μm film thickness; Agilent J&W), using helium at a constant flow rate of 1 mL / min as the carrier gas. The oven was initially set to 80°C and programmed to 240°C (10°C / min; hold for 1 min), then to 300°C (20°C / min; hold for 1 min).

[0850] In vitro assay samples were analyzed using an Agilent 6890N GC system connected to a 5975 Series Mass Selective Detector (MSD) and equipped with a split / splitless injector (Agilent Technologies, CA) and a CombiPAL autosampler (CTC Analytics, Zwingen, Switzerland). The GC inlet temperature was set to 250°C, and 1.0 μL of sample was injected in splitless pulse mode (pulse pressure 1.56 bar, pulse duration 0.6 min) and analyzed on a DB-1ms capillary column (30 m x 0.25 mm inner diameter x 0.25 μm film thickness; Agilent J&W), using helium at a constant flow rate of 1.2 mL / min as the carrier gas. The oven was initially set to 100°C (hold for 1 min) and programmed to 260°C (10 to 20°C / min), then to 300°C (30°C / min; hold for 1 min). For compounds with smaller molecular weights, the same conditions are used for analysis, except that the initial column temperature is lowered to 80°C.

[0851] Engineering of recombinant strains for degrading terpene compounds.

[0852] Recombinant strains capable of producing or transforming compounds are modified by introducing nucleotide sequences encoding one or more of the following proteins:

[0853] - One of the following Bayer-Villiger monooxygenases (BVMO):

[0854] SCH23-BVMO1 (SEQ ID NO:2) is derived from a filamentous yeast (Hyphozyma roseonigra).

[0855] SCH24-BVMO1 (SEQ ID NO:6) from *Filobasidium magnum*

[0856] SCH25-BVMO1 (SEQ ID NO:10) from *Papiliotrema laurentii*, and

[0857] SCH46-BVMO1 (SEQ ID NO:13) from Bensingtonia ciliata;

[0858] - Selected from the following esterases:

[0859] SCH23-EST (SEQ ID NO:20) is derived from a filamentous yeast (Hyphozyma roseonigra).

[0860] SCH24-EST (SEQ ID NO:24) from *Filobasidium magnum*

[0861] SCH25-EST (SEQ ID NO:28) from *Papiliotrema laurentii*; and

[0862] - Selected from the following enal lyases (lyases):

[0863] SCH94-3944 Rhodococcus erythropolis (SEQ ID NO:34),

[0864] SCH80-05241 Rhodococcus rhodochrous (SEQ ID NO:38),

[0865] Pdigit7033, *Penicillium digitatum* (SEQ ID NO:42),

[0866] PitalDUF4334-1, *Penicillium italicum* (SEQ ID NO:46),

[0867] AspWeDUF4334 Aspergillus wentii (SEQ ID NO:49),

[0868] RhoagDUF4334-2 is a strain of Rhodococcus hoagii, PAM2288 (SEQ ID NO:53).

[0869] RhoagDUF4334-3 Rhodococcus hominis strain N128 (SEQ ID NO:56),

[0870] RhoagDUF4334-4 Rhodococcus hominis NBRC 10125 (SEQ ID NO:59),

[0871] CnecaDUF4334 Cupriavidus necator (SEQ ID NO: 62),

[0872] Rins-DUF4334 Ralstonia insidiosa (SEQ ID NO:69),

[0873] CgatDUF4334 Cryptococcus gattii EJB2 (SEQ ID NO:72),

[0874] GclavDUF4334 is a cyanobacterial fungus (Grosmannia clavigera) kw1407 (SEQ ID NO:75).

[0875] TcurvaDUF4334 Thermomonospora curvata (SEQ ID NO:81),

[0876] PprotDUF4334 protects against Pseudomonas protegens (SEQ ID NO:87).

[0877] The bacterial host cells used for in vitro enzyme assays or whole-cell biotransformation assays were selected from *Escherichia coli* KRX cells (Promega Corporation, Madison, WI, USA) and *Escherichia coli* BL21 Star cells. TM(DE3) cells (ThermoFisher).

[0878] In order to enable the biochemical production of terpenoid compounds using one or more enzymes selected from the above-mentioned enzymes, host cells are engineered to produce increased amounts of farnesyl pyrophosphate (FPP) using the mevalonate enzyme pathway and further convert them to express sesquiterpene or diterpene biosynthesizing enzymes.

[0879] Recombinant Escherichia coli strains were engineered to produce FPPs by integrating the gene encoding the mevalonate pathway enzyme into the chromosome.

[0880] Escherichia coli strains were engineered to produce farnesyl pyrophosphate (FPP) by integrating a recombinant gene encoding a mevalonate pathway enzyme into the chromosome. See also: Figure 1 The construction scheme and reorganization events described in the document.

[0881] An upper pathway operon (operon 1 from acetyl-CoA to mevalonic acid) was designed, consisting of the atoB gene from Escherichia coli encoding acetyl-CoA thiolase, and the mvaA and mvaS genes from Staphylococcus aureus encoding HMG-CoA synthase and HMG-CoA reductase, respectively.

[0882] As the downstream mevalonate pathway operator (operon 2 from mevalonate to farnesyl pyrophosphate), a natural operator from the Gram-negative bacterium Streptococcus pneumoniae was selected, which encodes mevalonate kinase (mvaK1), phosphate mevalonate kinase (mvaK2), phosphate mevalonate decarboxylase (mvaD), and isopentenyl diphosphate isomerase (fni).

[0883] A codon-optimized Saccharomyces cerevisiae FPP synthase encoding gene (ERG20) was introduced into the 3' end of the upstream pathway operon to convert isopentenyl diphosphate (IPP) and dimethylallyl diphosphate (DMAPP) into FPP.

[0884] The aforementioned operon was synthesized using DNA 2.0 and integrated into the araA gene of *E. coli* strain BL21(DE3). A heterologous pathway was introduced in two separate recombination steps using the CRISPR / Cas9 genome engineering system. The first operon to be integrated (downstream pathway; operon 2) carried a spectinomycin (Spec) marker, which was used to screen for candidate integrants resistant to Spec. A second operon was designed to replace the Spec marker of the previously integrated operon, and Spec candidate integrants were screened accordingly after the second recombination event (see [link to documentation]). Figure 1A guide RNA expression vector targeting the araA gene was designed and synthesized using DNA 2.0. Operator integration was validated using PCR by designing PCR primers to amplify the cross-araA gene integration target and the cross-integron recombination linker. A clone that produced the correct PCR results was then fully sequenced and archived as strain DP1205.

[0885] Engineering of recombinant bacterial cells used to produce copal.

[0886] An operon was constructed containing two cDNAs that encode:

[0887] -AspWeTPP, a protein from *Aspergillus wentii* (SEQ ID NO:170) (GenBank accession number OJJ34585.1) with terpenoid diphosphate phosphatase activity, has the ability to dephosphorylate compounds such as cobazyl PP; and

[0888] -PvCPS is a protein from *Talaromyces verruculosus* (SED ID NO:173) (GenBank accession number BBF88128.1) with isoprenyltransferase and cobazone bisphosphate synthase activities. PvCPS catalyzes the production of cobazone PP from IPP and DMAPP.

[0889] The cDNAs encoding AspWeTPP and PvCPS were codon-optimized (SEQ ID NO: 171 and 174). An operon containing two cDNAs and an RBS sequence (AAGGAGGTAAAAAA) (SEQ ID NO: 196) upstream of the cDNAs was designed. This operon was synthesized and cloned into the pJ401 expression plasmid (ATUM, Newark, California) to provide plasmid pJ401-CPOL-4.

[0890] Transforming E. coli cells (e.g., DP1205 E. coli cells) with plasmid pJ401-CPOL-4 provides recombinant cells that can produce copal alcohol when cultured under conditions that enable the production of terpenoid compounds.

[0891] Engineering of recombinant bacterial cells used to produce copal aldehyde.

[0892] An operon containing three cDNAs was constructed, encoding:

[0893] -AspWeTPP, a protein from Aspergillus wentii (SEQ ID NO:170) (GenBank accession number OJJ34585.1) with terpenoid diphosphate phosphatase activity, has the ability to dephosphorylate compounds such as cobazyl PP.

[0894] -AzTolADH1, a protein from *Azoarcus toluclasticus* (SEQ ID NO: 167) (GenBank accession number WP_018990713.1) with alcohol dehydrogenase (ADH) activity, has the ability to oxidize terpene alcohols, such as copal alcohol, to corresponding carbonyl compounds, such as copal aldehyde; and

[0895] -PvCPS is a protein from *Talaromyces verruculosus* (SED ID NO:173) (GenBank accession number BBF88128.1) with isoprenyl transferase and copacpite synthase activities. It has the ability to generate cyclic terpene diphosphate compounds, such as copacpite PP, from IPP and DMAPP.

[0896] The cDNAs encoding AspWeTPP, AzTolADH1, and PvCPS were codon-optimized (SEQ ID NO: 171, 168, and 174). An operon was designed containing three cDNAs and an RBS sequence (AAGGAGGTAAAAAA) upstream of each cDNA (SEQ ID NO: 196). This operon was synthesized and cloned into the pJ401 expression plasmid (ATUM, Newark, California) to provide plasmid pJ401-CPAL-1.

[0897] Transforming E. coli cells (e.g., DP1205 E. coli cells) with plasmid pJ401-CPAL-1 provides recombinant cells that can produce copalaldehyde when cultured under conditions that enable the production of terpenoids.

[0898] Engineering of recombinant bacterial cells used to produce farnesal.

[0899] An operon was constructed containing two cDNAs that encode:

[0900] -TalCeTPP, a protein from *Talaromyces cellulolyticus* (GenBank: GAM42000.1) (SEQ ID NO: 176) with terpenoid diphosphate phosphatase activity, has the ability to dephosphorylate compounds such as farnesyl diphosphate; and

[0901] -CdGeoA is a protein from *Castellaniella defragrans* (NCBI Registry No. WP_043683915.1) (SEQ ID NO:179) with alcohol dehydrogenase (ADH) activity, capable of oxidizing terpene alcohols, such as farnesol, to corresponding carbonyl compounds, such as farnesal.

[0902] The cDNAs encoding TalCeTPP and CdGeoA were codon-optimized (SEQ ID NO: 177 and 180). An operon was designed containing two cDNAs and an RBS sequence (AAGGAGGTAAAAAA) upstream of each cDNA (SEQ ID NO: 196). This operon was synthesized and cloned into the pJ401 expression plasmid (ATUM, Newark, California) to provide plasmid pJ401-FAL-1.

[0903] Transformation of E. coli cells, such as DP1205 E. coli cells, using plasmid pJ401-FAL-1 provides recombinant cells capable of producing farnesal when cultured under conditions that enable the production of terpenoid compounds.

[0904] Engineering of recombinant bacterial cells for producing hemispheric acid glycol.

[0905] An operon was constructed containing three cDNAs that encode:

[0906] -TalVeTPP, a protein from *Talaromyces verruculosus* (Genbank accession number KUL89334.1) (SEQ ID NO:194) with terpenoid diphosphate phosphatase activity; capable of dephosphorylating terpenoid diphosphate compounds such as hemispheryl PP.

[0907] -SsLPS, a protein from clary sage (Salvia sclarea) (Genbank accession number AET21247.1) (SEQ ID NO:188) with hemispheric acid pyrophosphate (LPP) synthase activity, has the ability to produce cyclic terpenoid diphosphates, such as hemispheric acid pyrophosphate, from GGPP; and

[0908] -CrtE, a geraniol-geraniol diphosphate synthase from Pantoea agglomerans (GenBank accession number AAA24819.1) (SEQ ID NO:191), has the ability to generate GGPP from FPP.

[0909] The cDNAs encoding TalVeTPP, SsLPS, and CrtE were codon-optimized (SEQ ID NO: 195, 189, and 192). An operon was designed containing three cDNAs and an RBS sequence (AAGGAGGTAAAAAA) upstream of each cDNA (SEQ ID NO: 196). This operon was synthesized and cloned into the pJ401 expression plasmid (ATUM, Newark, California) to provide plasmid pJ401-LOH-2.

[0910] Transformation of E. coli cells, such as DP1205 E. coli cells, using plasmid pJ401-LOH-2 provides recombinant cells capable of producing hemispheric acid diol when cultured under conditions that enable the production of terpenoid compounds.

[0911] Transformation, selection, and culture of yeast cells.

[0912] As described in Gietz and Woods, Methods Enzymol., 2002, 350:87–96, all yeast cell transformations were performed using the lithium acetate protocol. The transformation mixture was plated on SmUra- or SmLeu- agar plates containing 6.7 g / L amino acid-free yeast nitrogen base (BD Difco, New Jersey, USA), 1.92 g / L uracil-free Dropout supplement (Sigma Aldrich, Missouri, USA) or 1.6 g / L leucine-free Dropout supplement (Sigma Aldrich, Missouri, USA), 20 g / L glucose, and 20 g / L agar. The plates were incubated at 30°C for 3–4 days.

[0913] Yeast cell engineering for increasing endogenous farnesyl diphosphate levels.

[0914] To increase the level of endogenous farnesyl diphosphate (FPP) pooling in *Saccharomyces cerevisiae* cells, additional copies of all yeast endogenous genes involved in the mevalonate pathway, from ERG10 encoding acetyl-CoA C-acetyltransferase to ERG20 encoding FPP synthase, were integrated into the genome of *Saccharomyces cerevisiae* strain CEN.PK2-1C (Euroscarf, Frankfurt, Germany) under the control of a galactose-inducible promoter, similar to those described in Paddon et al., *Nature*, 2013, 496:528-532. In short, the three boxes were integrated into the LEU2, TRP1, and URA3 loci, respectively. The first box contains the ERG20 gene, controlled by the GAL10 / GAL1 bidirectional promoter, and a truncated HMG1 (tHMG1, as described in Donald et al., Proc Natl Acad Sci USA, 1997, 109: E111-8), as well as the ERG19 and ERG13 genes, also controlled by the GAL10 / GAL1 promoter. This box is flanked by two 100-nucleotide regions, corresponding to the upstream and downstream portions of LEU2, respectively. The second box contains the genes IDI1 and tHMG1, controlled by the GAL10 / GAL1 promoter, and the gene ERG13, controlled by the GAL7 promoter region. This box is flanked by two 100-nucleotide regions, corresponding to the upstream and downstream portions of TRP1, respectively. The third box contains the genes ERG10, ERG12, tHMG1, and ERG8, all controlled by the GAL10 / GAL1 promoter. The box has two 100-nucleotide regions flanking it, corresponding to the upstream and downstream portions of URA3, respectively. All genes in the three boxes contain 200 nucleotides of their own terminator region. Furthermore, upstream of the ERG9 promoter region, an additional copy of GAL4, under the control of a mutant form of its own promoter as described in Griggs and Johnston, Proc Natl Acad Sci USA, 1991, 88:8597-8601, is integrated. Additionally, ERG9 expression is modified via promoter exchange. The GAL7, GAL10, and GAL1 genes are deleted using a box containing the HIS3 gene with its own promoter and terminator. The resulting strain is crossed with strain CEN.PK2-1D (Euroscarf, Frankfurt, Germany) to obtain a diploid strain named YST045, which induces sporulation according to Solis-Escalante et al., FEMS Yeast Res, 2015, 15:2. Spore isolation was achieved by resuspending ASTI in 200 μL of 0.5 M sorbitol and 2 μL of zymolyase (1000 U / mL). -1This was achieved by incubating the mixture at 37°C for 20 minutes in Zymoresearch, Irvine, CA. The mixture was then plated on a medium containing 20 g / L peptone, 10 g / L yeast extract and 20 g / L agar, and a germinating spore, named YST075, was isolated.

[0915] Engineering of recombinant yeast cells for producing copal alcohol.

[0916] For the production of copacpin, the expression of GGPP synthase carG (from *Blakeslea trispora*, NCBI accession number JQ289995.1) (SEQ ID NO:182), copacpin pyrophosphate synthase SmCPS2 (from *Salvia miltiorrhiza*, NCBI accession number ABV57835.1) (SEQ ID NO:185), and copacpin pyrophosphate phosphatase TalVeTPP (from *Talaromyces verruculosus*, NCBI accession number KUL89334.1) (SEQ ID NO:194) in different engineered yeast cells was achieved using an in vivo plasmid system constructed as previously described by Kuijpers et al., *Microb CellFact.*, 2013, 12:47, through endogenous homologous recombination in yeast. The plasmid consists of six DNA fragments for co-transformation in *Saccharomyces cerevisiae*. These fragments are:

[0917] a) The LEU2 yeast marker was constructed by PCR using primers 5'-AGGTGCAGTTCGCGTGCAATTATAACGTCGTGGCAACTGTTATCAGTCGTACCGCGCCATTCGACTACGTCGTAAGGCC-3' (SEQ ID NO:124) and 5'-TCGTGGTCAAGGCGTGCAATTCTCAACACGAGAGTGATTCTTCGGCGTTGTTGCTGACCATCGACGGTCGAGGAGAACTT-3' (SEQ ID NO:125) with plasmid pESC-LEU (Agilent Technologies, California, USA) as a template.

[0918] b) The AmpR E. coli marker was constructed by PCR using primers 5'-TGGTCAGCAACAACGCCGAAGAATCACTCTCGTGTTGAGAATTGCACGCCTTGACCACGACACGTTAAGGGATTTTGGTCATGAG-3' (SEQ ID NO:126) and 5'-AACGCGTACCCTAAGTACGGCACCACAGTGACTATGCAGTCCGCACTTTGCCAATGCCAAAAATGTGCGCGGAACCCCTA-3' (SEQ ID NO:127) with plasmid pESC-URA as a template.

[0919] c) Yeast replication origin, obtained by PCR using primers 5'-TTGGCATTGGCAAAGTGCGGACTGCATAGTCACTGTGGTGCCGTACTTAGGGTACGCGTTCCTGAACGAAGCATCTGTGCTTCA-3' (SEQ ID NO:128) and 5'-CCGAGATGCCAAAGGATAGGTGCTATGTTGATGACTACGACACAGAACTGCGGGTGACATAATGATAGCATTGAAGGATGAGACT-3' (SEQ ID NO:129) with pESC-URA as a template;

[0920] d) The origin of replication of E. coli was obtained by PCR using primers '-ATGTCACCCGCAGTTCTGTGTCGTAGTCATCAACATAGCACCTATCCTTTGGCATCTCGGTGAGCAAAAGGCCAGCAAAAGG-3' (SEQ ID NO:130) and 5'-CTCAGATGTACGGTGATCGCCACCATGTGACGGAAGCTATCCTGACAGTGTAGCAAGTGCTGAGCGTCAGACCCCGTAGAA-3' (SEQ ID NO:131) with plasmid pESC-URA as a template;

[0921] e) A fragment consisting of the last 60 nucleotides of fragment "d", the 200 nucleotides downstream of the stop codon of the yeast gene PGK1, the GGPP synthase coding sequence CrtE, the bidirectional yeast promoter of GAL10 / GAL1, the coding sequence of TalVeTPP, the 200 nucleotides downstream of the stop codon of the yeast gene CYC1, and the sequence 5'-ATTCCTAGTGACGGCCTTGGGAACTCGATACACGATGTTCAGTAGACCGCTCACACATGG-3' (SEQ ID NO: 132), obtained through DNA synthesis (ATUM, Menlo Park, CA 94025); and

[0922] f) A fragment consisting of the last 60 nucleotides of fragment "e", 200 nucleotides downstream of the stop codon of the yeast gene CYC1, the coding sequence for SmCPS2 cobazyl pyrophosphate synthase, the bidirectional yeast promoter GAL10 / GAL1, and 60 nucleotides corresponding to the beginning of fragment "a". This fragment was obtained through DNA synthesis (ATUM, Menlo Park, CA 94025).

[0923] Alternatively, GGPP synthase carG and cobazyl pyrophosphate synthase can be replaced by bifunctional PvCPS.

[0924] Engineering of recombinant yeast cells to produce minoyloxy groups.

[0925] To degrade coparol to minoyloxy groups using different alcohol dehydrogenases (ADH), Bayer-Villiger monooxygenases (BVMO), and esterases (EST), genome integration was performed in strain YST075. Each integration cassette consists of four fragments:

[0926] 1) A fragment containing 658 bp corresponding to the upstream segment of the NDT80 gene and the sequence 5'-GCAGGCAGCTCCATTTCATGTAGGTGATTTATCCCTCCGGGCGGTATTTGAGACTCTCGG-3' (SEQ ID NO:121), which was obtained by PCR using genomic DNA from strain YST075 as a template;

[0927] 2) A fragment containing the sequence 5'-GCAGGCAGCTCCATTTCATGTAGGTGATTTATCCCTCCGGGCGGTATTTGAGACTCTCGG-3' (SEQ ID NO:121), a CYC1 terminator region, one of the genes encoding BVMO, an intergenic region between the GAL1 and GAL10 genes, one of the genes encoding esterase, a terminator region of the ADH1 gene, and the sequence 5'-ACTGCTGGGTACTGTTCAGGCACGATAGGAAATGCGTCCAGCGCATACACCAGTCTTAGC-3' (SEQ ID NO:122). This fragment was obtained through DNA synthesis (ATUM, Menlo Park, CA 94025).

[0928] 3) A fragment containing the sequence 5'-ACTGCTGGGTACTGTTCAGGCACGATAGGAAATGCGTCCAGCGCATACACCAGTCTTAGC-3' (SEQ ID NO: 122), a PGK1 terminator region, one of the genes encoding alcohol dehydrogenase, the promoter regions of genes GAL1 and GAL10, one of the genes encoding alcohol dehydrogenase, a CYC1 terminator region, and the sequence 5'-AGTCGACCTTACAGCGCCTGGGACTCTACATAAACATGCAGCGAACATGCTTTCCAACGC-3' (SEQ ID NO: 123). Based on the experiments conducted, this fragment may contain one or two alcohol dehydrogenases. These are obtained through DNA synthesis (ATUM, MenloPark, CA 94025); and

[0929] 4) A fragment containing the sequence 5'-AGTCGACCTTACAGCGCCTGGGACTCTACATAAACATGCAGCGAACATGCTTTCCAACGC-3' (SEQ ID NO:123) and a 405bp segment corresponding to the NDT80 gene. This fragment was obtained by PCR using genomic DNA from strain YST075 as a template.

[0930] Engineering of recombinant yeast cells to degrade copal alcohol into minoyloxy groups.

[0931] To degrade coparol to minoyloxy group, genome integration was performed in strain YST075 using alcohol dehydrogenases and different enal cleavage peptides. Each integration cassette consisted of three fragments:

[0932] 1) A fragment containing 658 bp corresponding to the upstream segment of the NDT80 gene and the sequence 5'-GCAGGCAGCTCCATTTCATGTAGGTGATTTATCCCTCCGGGCGGTATTTGAGACTCTCGG-3' (SEQ ID NO:121), which was obtained by PCR using genomic DNA from strain YST075 as a template;

[0933] 2) A fragment containing the sequence 5'-GCAGGCAGCTCCATTTCATGTAGGTGATTTATCCCTCCGGGCGGTATTTGAGACTCTCGG-3' (SEQ ID NO:121), an intergenic region between the GAL1 and GAL10 genes, one of the genes encoding an enal cleavage polypeptide, the terminator region of the ADH1 gene, and the sequence 5'-ACTGCTGGGTACTGTTCAGGCACGATAGGAAATGCGTCCAGCGCATACACCAGTCTTAGC-3' (SEQ ID NO:122), obtained by DNA synthesis (ATUM, MenloPark, CA 94025); and

[0934] 3) A fragment containing the sequence 5'-ACTGCTGGGTACTGTTCAGGCACGATAGGAAATGCGTCCAGCGCATACACCAGTCTTAGC-3' (SEQ ID NO:122), the PGK1 terminator region, the gene encoding alcohol dehydrogenase, the promoter regions of genes GAL1 and GAL10, the sequence 5'-AGTCGACCTTACAGCGCCTGGGACTCTACATAAACATGCAGCGAACATGCTTTCCAACGC-3' (SEQ ID NO:123), and a 405 bp corresponding to the NDT80 gene. This fragment was obtained through DNA synthesis (ATUM, Menlo Park, CA 94025).

[0935] In all cases, copalol is produced by expressing the biosynthetic pathway in the plasmid system described above.

[0936] Engineering of recombinant yeast cells for the production of γ-ambryl acetate.

[0937] To degrade coparol to γ-ambryl acetate using alcohol dehydrogenase, enal cleavage peptide, and different Bayer-Villiger monooxygenases (BVMOs), genome integration was performed in strain YST075; each integration cassette consisted of three fragments:

[0938] (1) A fragment containing 658 bp corresponding to the upstream segment of the NDT80 gene and the sequence 5'-GCAGGCAGCTCCATTTCATGTAGGTGATTTATCCCTCCGGGCGGTATTTGAGACTCTCGG-3' (SEQ ID NO:121), which was obtained by PCR using genomic DNA from strain YST075 as a template;

[0939] (2) A fragment containing the sequence 5'-GCAGGCAGCTCCATTTCATGTAGGTGATTTATCCCTCCGGGCGGTATTTGAGACTCTCGG-3' (SEQ ID NO: 121), the terminator region of the CYC1 gene, one of the genes encoding the tested BVMO, the intergenic region between the GAL1 and GAL10 genes, the gene encoding the enal cleavage polypeptide, the terminator region of the ADH1 gene, and the sequence 5'-ACTGCTGGGTACTGTTCAGGCACGATAGGAAATGCGTCCAGCGCATACACCAGTCTTAGC-3' (SEQ ID NO: 122), which was obtained by DNA synthesis (ATUM, Menlo Park, CA 94025); and

[0940] (3) A fragment containing the sequence 5'-ACTGCTGGGTACTGTTCAGGCACGATAGGAAATGCGTCCAGCGCATACACCAGTCTTAGC-3' (SEQ ID NO:122), the PGK1 terminator region, the gene encoding alcohol dehydrogenase, the promoter regions of genes GAL1 and GAL10, the sequence 5'-AGTCGACCTTACAGCGCCTGGGACTCTACATAAACATGCAGCGAACATGCTTTCCAACGC-3' (SEQ ID NO:123), and a 405 bp corresponding to the NDT80 gene. This fragment was obtained through DNA synthesis (ATUM, Menlo Park, CA 94025).

[0941] In all cases, copalol is produced by expressing the biosynthetic pathway in the plasmid system described above.

[0942] Engineering of recombinant yeast cells for producing γ-ambrol.

[0943] To degrade copal alcohol to γ-ambrol using alcohol dehydrogenase, enal cleavage peptide, Bayer-Villiger monooxygenase (BVMO), and various esterases (EST), genome integration was performed in strain YST075; each integration cassette consisted of four overlapping fragments:

[0944] 1) A fragment containing at least 300 bp and at least 60 bp of overlapping sequences corresponding to the upstream segment of the BUD9 gene for in vivo assembly. This fragment was obtained by PCR using genomic DNA of strain YST075 as a template;

[0945] 2) A fragment containing the terminator region of the ADH1 gene, one of the genes encoding the tested esterase, and an intergenic region between the GAL1 and GAL10 genes. This fragment is flanked by a sequence that allows for in vivo assembly. This fragment was obtained through DNA synthesis (ATUM, Menlo Park, CA 94025).

[0946] 3) A fragment containing the URA3 yeast marker, with its own promoter and terminator, and flanked by sequences that allow homologous recombination. This fragment was obtained by PCR; and

[0947] 4) A fragment containing at least 300 bp and at least 60 bp of overlapping sequences corresponding to the downstream segment of the BUD9 gene for in vivo assembly. This fragment was obtained by PCR using genomic DNA from strain YST075 as a template.

[0948] In all cases, copalol is produced by expressing the biosynthetic pathway in the plasmid system described above.

[0949] Engineered yeast cells were cultured under conditions that enabled the production of terpenoid compounds, and a GC-MS analysis method was employed.

[0950] The production of terpenes and derivatives in engineered yeast cells was assessed by culturing cells under similar conditions described in Westfall et al., Proc Natl Acad Sci USA, 2012, 109:E111-118, using 10% dodecane or 10% isopropyl myristate (IPM) as an organic covering. Cultures were then extracted with two volumes of MTBE, and the composition of the organic phase was analyzed using an Agilent 7890A GC system (Agilent Technologies, CA) equipped with a 5975C Series Mass Select Detector (MSD) and a split / splitless injector and GC Injector 80 system connected to the system. The GC inlet temperature was set to 260 °C, and 1.0 μl of sample was injected in splitless mode and analyzed on an HP-5 GC column (30 m x 0.25 mm x 0.25 μm; Agilent J&W) using helium at a constant flow rate of 1.2 mL / min as the carrier gas. The initial temperature of the oven was set to 100℃, and the temperature was programmed to rise to 300℃ (10℃ / min).

[0951] c) Example

[0952] Example 1: In vivo conversion of minoyloxy to γ-ambryl acetate using BVMO

[0953] Codon-optimized cDNAs encoding SCH23-BVMO1 (SEQ ID NO:2) from a filamentous yeast (Hyphozyma roseonigra), SCH24-BVMO1 (SEQ ID NO:6) from *Filobasidium magnum*, and SCH46-BVMO1 (SEQ ID NO:13) from *Bensingtonia ciliata* were synthesized and cloned into the pJ414 expression plasmid (ATUM, Newark, California), providing plasmids pJ414-SCH23-BVMO1, pJ414-SCH24-BVMO1, and pJ414-SCH46-BVMO1. These expression plasmids were used to transform KRX *E. coli* cells (Promega Corporation, Madison, WI, USA). The transformed cells were grown and used in the whole-cell biotransformation assay described above, using minoloxyl as a substrate. The negative control group consisted of cells transformed with an empty plasmid. In the presence of recombinant SCH23-BVMO1, SCH24-BVMO1, or SCH46-BVMO1 proteins, conversion of minoyl alcohols to γ-ambryl acetate was observed (Figure 2). No conversion was observed in the negative control group. This experiment demonstrates that SCH23-BVMO1, SCH24-BVMO1, or SCH46-BVMO1 catalyzes the following conversions:

[0954]

[0955] These results indicate that SCH23-BVMO1, SCH24-BVMO1, and SCH46-BVMO1 catalyze the Bayer-Villiger oxidation of the minoyl alcohol radical.

[0956] Example 2: In vivo conversion of copalaldehyde to compound 4 using BVMO

[0957] Codon-optimized cDNAs encoding SCH23-BVMO1 (SEQ ID NO:3) from a filamentous yeast (Hyphozyma roseonigra), SCH24-BVMO1 (SEQ ID NO:7) from *Filobasidium magnum*, and SCH46-BVMO1 (SEQ ID NO:14) from *Bensingtonia ciliata* were synthesized and cloned into the pJ414 expression plasmid (ATUM, Newark, California), providing plasmids pJ414-SCH23-BVMO1, pJ414-SCH24-BVMO1, and pJ414-SCH46-BVMO1. These expression plasmids were then used to transform KRX *E. coli* cells (Promega Corporation, Madison, WI, USA). Cells were grown and used in the whole-cell biotransformation assay as described above, using a mixture of cis-copacl and trans-copacl as substrates. A negative control consisting of cells transformed with an empty plasmid was included. Transformation of cis-copacl and trans-copacl was observed in the presence of recombinant proteins SCH23-BVMO1, SCH24-BVMO1, or SCH46-BVMO1. GC-MS analysis of the biotransformation products after 42 hours of incubation (Figure 3) showed the formation of four major products, two stereoisomers 3a and 3b, and two stereoisomers 4a and 4b.

[0958] Time-point measurements of biotransformation showed that compounds 1a and 1b formed as intermediates. Figure 4 GC-MS analyses of the conversion of cis-copacaldehyde and trans-copacaldehyde by SCH23-BVMO1 at different time points were compared; similar evolution of product curves was observed for SCH24-BVMO1 and SCH46-BVMO1. The sequential formation of these compounds indicates that trans-copacaldehyde and cis-copacaldehyde are converted to compounds 4a and 4b in several steps. Compounds 1a and 1b, as well as compounds 4a and 4b, are formate esters. Such functional groups can be formed from aldehyde compounds by Bayer-Villiger monooxygenase. Therefore, the following reaction routes involving both enzymatic and non-enzymatic (chemical) reactions can be plotted to describe the conversion of trans-copacaldehyde by SCH23-BVMO1, SCH24-BVMO1, or SCH46-BVMO1.

[0959]

[0960] In this route, the recombinase catalyzes two Bayer-Villiger oxidations on two different aldehydes. First, in the first Bayer-Villiger oxidation, the α,β-unsaturated aldehyde group of trans-copalaldehyde is oxidized by the recombinase to form compound 1a. The enol carboxylate functional group of compound 1a is unstable under experimental conditions and is partially hydrolyzed to form compound 2a. The latter compound is rapidly converted to compound 3 (3a and 3b) via keto-enol tautomerism and is therefore undetectable in GC-MS analysis. Compounds 3 (3a and 3b) are substrates of the same enzyme that catalyzes the second Bayer-Villiger oxidation to form compounds 4 (4a and 4b).

[0961] The reaction routes below describe similar reactions for the conversion of cis-cobaral by SCH23-BVMO1, SCH24-BVMO1, or SCH46-BVMO1.

[0962]

[0963] These results indicate that SCH23-BVMO1, SCH24-BVMO1, and SCH46-BVMO1 catalyze the Bayer-Villiger oxidation of hemispherical aldehydes.

[0964] Example 3: In vitro conversion of minoyloxy groups using BVMO and esterase.

[0965] For this experiment, the following recombinant proteins were used: SCH23-BVMO1 (SEQ ID NO:2) from a filamentous yeast (Hyphozyma roseonigra), SCH24-BVMO1 (SEQ ID NO:6) from *Filobasidium magnum*, SCH23-EST (SEQ ID NO:20) from a filamentous yeast (Hyphozyma roseonigra), and SCH24-EST (SEQ ID NO:24) from *Filobasidium magnum*. Codon-optimized cDNAs encoding SCH23-BVMO1 (SEQ ID NO:3) and SCH24-BVMO1 (SEQ ID NO:7) were synthesized and cloned into the pJ414 expression plasmid (ATUM, Newark, California), providing plasmids pJ414-SCH23-BVMO1 and pJ414-SCH24-BVMO1. Codon-optimized cDNAs encoding SCH23-EST (SEQ ID NO:21) and SCH24-EST (SEQ ID NO:25) were synthesized and cloned into the pJ431 expression plasmid (ATUM, Newark, California), providing plasmids pJ414-SCH23-EST and pJ414-SCH24-EST.

[0966] KRX *E. coli* cells (Promega Corporation, Madison, WI, USA) were transformed with each of these expression plasmids. Transformed cells were grown and cell-free lysates were prepared as described. In vitro enzymatic assays were performed using any one of these protein fractions or a combination of two of these protein fractions. In vitro assay conditions were as described above, with the addition of 160 mg / L minocycline, 60 μM flavin adenine dinucleotide (FAD), and 500 μM reduced β-nicotinamide adenine dinucleotide phosphate (NADPH).

[0967] Using crude fractions containing recombinant SCH23-BVMO1 and SCH24-BVMO1 proteins, conversion of minoyloxy groups to γ-ambrol acetate was observed. No conversion was detected when using control lysates obtained from *E. coli* cells transformed with an empty plasmid. Figure 5 From these experiments, the following enzymatic reactions can be derived:

[0968]

[0969] In vitro enzymatic assays were also performed using protein fractions containing recombinant esterase and combinations thereof containing recombinant BVMO and recombinant esterase. These assays were performed using minoyloxy groups as substrates, as described above. GC-MS analysis of the resulting products ( Figure 6 and Figure 7 The results showed that in the presence of BVMO enzymes (SCH23-BVMO1 or SCH24-BVMO1), minoloxy groups were converted to γ-ambryl acetate, and further converted to γ-ambrol when esterases (SCH23-EST or SCH24-EST) were present in the assay. No substrate conversion was observed when esterases were used in the absence of BVMO. Figure 6 and Figure 7 ).

[0970] This experiment demonstrates that, in the presence of BVMO and esterase, the minoyloxy group can be converted to γ-ambrol via the reaction pathway shown below:

[0971]

[0972] Example 4: In vitro conversion of compounds 4a and 4b to compounds 5a and 5b using esterases.

[0973] Codon-optimized cDNAs encoding SCH23-EST (SEQ ID NO:21) from a filamentous yeast (Hyphozyma roseonigra), SCH24-EST (SEQ ID NO:25) from *Filobasidium magnum*, and SCH46-EST (SEQ ID NO:32) from *Bensingtonia ciliata* were synthesized and cloned into the pJ414 expression plasmid (ATUM, Newark, California), providing plasmids pJ414-SCH23-EST1, pJ414-SCH24-EST1, and pJ414-SCH46-EST1. These expression plasmids were used to transform KRX *E. coli* cells (Promega Corporation, Madison, WI, USA). The transformed cells were grown, and cell-free lysates were prepared as described. These protein fractions were then subjected to in vitro enzymatic assays under the conditions described above.

[0974] As shown in Figure 8, using crude fractions containing recombinant SCH23-EST1, SCH24-EST1, and SCH25-EST1 proteins, the conversion of two stereoisomers, 4a and 4b, to compounds 5a and 5b was observed. In contrast, no conversion was detected when using lysates containing recombinant BVMO enzyme (these proteins were therefore used as control reactions in this experimental series). Under these conditions, the enzyme activities of SCH23-EST1 and SCH25-EST1 were higher than those of SCH24-EST1.

[0975] From these experiments, the following enzymatic reactions can be derived:

[0976]

[0977] Example 5: In vitro conversion of copalaldehyde using BVMO and esterase.

[0978] For this experiment, the following recombinant proteins were used: SCH23-BVMO1 (SEQ ID NO:2) from a filamentous yeast (Hyphozyma roseonigra), SCH24-BVMO1 (SEQ ID NO:6) from Filobasidium magnum, SCH25-BVMO1 (SEQ ID NO:10) from Papiliotrema laurentii, SCH23-EST (SEQ ID NO:20) from a filamentous yeast (Hyphozyma roseonigra), SCH24-EST (SEQ ID NO:24) from Filobasidium magnum, and SCH25-EST (SEQ ID NO:28) from Papiliotrema laurentii.

[0979] Codon-optimized cDNAs encoding SCH23-BVMO1 (SEQ ID NO:3), SCH24-BVMO1 (SEQ ID NO:7), and SCH25-BVMO1 (SEQ ID NO:11) were synthesized and cloned into the pJ414 expression plasmid (ATUM, Newark, California), providing plasmids pJ414-SCH23-BVMO1, pJ414-SCH24-BVMO1, and pJ414-SCH25-BVMO1. Codon-optimized cDNAs encoding SCH23-EST (SEQ ID NO:21), SCH24-EST (SEQ ID NO:25), and SCH25-EST (SEQ ID NO:29) were synthesized and cloned into the pJ431 expression plasmid (ATUM, Newark, California), providing plasmids pJ414-SCH23-EST and pJ414-SCH25-EST.

[0980] These expression plasmids were used to transform KRX *E. coli* cells (Promega Corporation, Madison, WI, USA). Transformed cells were grown and cell-free lysates were prepared as described. In vitro enzymatic assays were performed using protein fractions containing recombinant BVMO enzyme or recombinant esterase, or by a combination of protein fractions containing recombinant BVMO and esterase. As described above, the assays were performed with 320 mg / L of a mixture of cis-copacl and trans-copacl as substrates, 60 μM of flavin adenine dinucleotide (FAD), and 500 μM of reduced β-nicotinamide adenine dinucleotide phosphate (NADPH).

[0981] Figure 9 The transformation products of copalaldehyde were compared in the presence of SCH23-BVMO1 alone and in combination with different esterases. In the presence of SCH23-BVMO1, the major products were formate compounds 1a, 1b, and 4a, 4b. When measured in the additional presence of SCH23-EST or SCH25-EST, the major products of the transformation were compounds 5a and 5b, indicating that these two esterases can efficiently hydrolyze the formate intermediates produced by BVMO enzymes. In the additional presence of SCH24-EST, the hydrolysis of the same intermediates (1a, 1b, and 4a, 4b) was observed; however, using this enzyme, SCH24-EST appeared to be more efficient at hydrolyzing intermediates 1a and 1b than intermediates 4a and 4b.

[0982] When SCH24-BVMO1 binds to the esterases SCH23-EST or SCH24-EST, a similar conversion of cis- and trans-copalaldehyde is observed. Figure 10In the control experiment, no transformation was observed when copalaldehyde was incubated with esterase alone.

[0983] The following enzymatic pathways can be deduced from these experiments.

[0984]

[0985] Example 6: In vivo production of 14,15-dedimethyl-hexamonadene compounds 5a and 5b, as well as biosynthetic intermediates, in engineered bacterial cells expressing BVMO and esterase.

[0986] In this experiment, as described in the Experimental Section, *E. coli* cells were transformed with plasmid pJ401-CPAL-1 (as described above) to produce copalaldehyde. When DP1205 *E. coli* cells were transformed and cultured under the conditions described in the Experimental Section, the formation of trans-copalaldehyde and cis-copalaldehyde was observed (Fig. 11, chromatogram at the top). The detection of the two double-bond isomers of copalaldehyde is due to the relatively easy isomerization of (E)-α,β-unsaturated aldehydes (Konning et al, Org. Lett., 2012, 14(20), pp 5258-5261). The additional detection of hemisperidin-8(20)-en-15-ol was due to the activity of endogenous enoic acid reductase in *E. coli*.

[0987] The bacterial cells were then transformed with a second expression plasmid carrying codon-optimized cDNA encoded by *Filobasidium magnum*. 20918 TM (SEQ ID NO:7) or SCH46-BVMO1 (SEQ ID NO:14) from *Bensingtonia ciliata*. These plasmids were prepared by cloning optimized cDNA into the pJ423 expression plasmid (ATUM, Newark, California), providing plasmids pJ423-SCH23-BVMO and pJ423-SCH46-BVMO, respectively. Cells transformed with both plasmids were cultured, and the production of terpenoids and terpenoid derivatives was analyzed using the conditions described in the Experimental section. Under these conditions, compounds 1a and 1b, 3a and 3b, and 4a and 4b were detected in the solvent extracts of the culture medium (Figure 11). These results indicate that the biosynthesis of hemispheric diterpenes (e.g., copalol) and the sequential enzymatic cleavage of two carbon-carbon bonds in the side chain can be introduced into recombinant cells using these combinations of enzymes.

[0988] Similarly, bacterial cells were co-transformed with plasmid pJ401-CPAL-1 and a second plasmid containing a gene encoding BVMO and a gene encoding an esterase. This second plasmid was pJ423-SCH24-BVMO-SCH24-EST, prepared by inserting a synthetic operon consisting of codon-optimized cDNA encoding SCH24-BVMO1 (SEQ ID NO:7) and codon-optimized cDNA encoding SCH24-EST (SEQ ID NO:25) into the pJ423 expression plasmid (ATUM, Newark, California), or pJ423-SCH46-BVMO-SCH46-EST, prepared by inserting a synthetic operon consisting of codon-optimized cDNA encoding SCH46-BVMO (SEQ ID NO:14) and codon-optimized cDNA encoding SCH46-EST (SEQ ID NO:25). A synthetic operon composed of codon-optimized cDNA (NO:32) was prepared by inserting it into the pJ423 expression plasmid (ATUM, Newark, California). Cells were cultured and the production of terpenoids and terpenoid derivatives was analyzed using the conditions described in the experimental section. Under these conditions, compounds 5a and 5b were detected, and a reduction in the amount of pathway intermediates (compounds 1a, 1b, 3a, 3b, 4a, and 4b) was observed.

[0989] This experimental series demonstrates that the following biosynthetic pathway can be introduced into host cells transformed to express diterpenoid biosynthetic enzymes, in combination with BVMO and esterases.

[0990]

[0991] Example 7: In vivo conversion of compounds 5a and 5b to minoyloxy groups using alcohol dehydrogenase.

[0992] For this experiment, the following alcohol dehydrogenases were evaluated for the oxidation of compounds 5a and 5b to the minool radical:

[0993] RrhSecADH (SEQ ID NO:146) from Rhodococcus rhodochrous,

[0994] SCH80-00043 (SEQ ID NO:149) from Rhodococcus roseum,

[0995] SCH80-04254 (SEQ ID NO:152) from Rhodococcus roseum,

[0996] SCH80-06135 (SEQ ID NO:155) from Rhodococcus roseum,

[0997] SCH80-06582 (SEQ ID NO:158) from Rhodococcus roseum,

[0998] (See also WO2005 / 026338); the above ADH is only a non-limiting example and can be replaced by other known ADHs.

[0999] Codon-optimized cDNAs encoding each of these proteins were synthesized and cloned into the vector pJ401, providing plasmids pJ401-RrhSecADH, pJ401-SCH80-00043, pJ401-SCH80-04254, pJ401-SCH80-06135, and pJ401-SCH80-06582 (ATUM, Newark, California).

[1000] These expression plasmids were used to transform KRX *E. coli* cells (Promega Corporation, Madison, WI, USA). Transformed cells were grown and used in the whole-cell biotransformation assay described above, using a mixture of compounds 5a and 5b as substrates. Five hours after induction of recombinant protein expression, the substrate was added to a final concentration of 0.55 mg / mL using an emulsion containing 50 mg / mL Tween 80 and 25 mg / mL of the substrate in water. A negative control consisting of cells transformed with empty plasmids was included. Oxidation was observed only in the presence of the SCH80-06135 and RrhSecADH recombinant proteins. Figure 12 This indicates that these enzymes can catalyze the following reactions.

[1001]

[1002] Example 8: In vivo production of the compound γ-ambrol and its biosynthetic intermediates in engineered bacterial cells expressing BVMO, esterase, and alcohol dehydrogenase.

[1003] In this experiment, plasmid pJ401-CPAL-1 (as described above) was used to transform DP1205 E. coli cells to produce a background strain that produces copalaldehyde (cis and trans isomers) as described in the previous section.

[1004] The strain was then co-transformed with plasmid pJ423-SCH24-BVMO-SCH24-EST (as described above), allowing further expression of BVMO and esterase in the same cells. Based on observations in the previous section, this recombinant organism produced 14,15-dedimethyl-hemisinane compounds.

[1005] To allow the side-chain degradation to continue forming tetramethyl-hemisinane derivatives, the secondary alcohol groups of compounds 5a and 5b must be oxidized to the corresponding ketones. Therefore, a plasmid containing nucleotide sequences encoding BVMO, an esterase, and a suitable alcohol dehydrogenase (identified in Example 7) was constructed. For the alcohol dehydrogenase, a codon-optimized cDNA (accession number WP_043801412.1) (SEQ ID NO: 147) encoding RrhSecADH from the genus Rhodococcus was synthesized, and a synthetic operon was designed to bind the RrhSecADH cDNA and the cDNA encoding SCH24-BVMO and SCH24-EST. The operon was cloned into the pJ423 expression plasmid, providing the pJ423-secADH-23BVMO-EST plasmid. When DP1205 *E. coli* cells co-transformed with vectors pJ401-CPAL-1 and pJ423-secADH-23BVMO-EST were cultured under the above conditions, γ-ambrol was detected in the culture medium by GC-MS analysis. Figure 13 These data suggest that when compounds 5 (5a and 5b) are oxidized to minoyloxy in the presence of appropriate ADH, the BVMO can catalyze the following steps of γ-ambrol.

[1006] This series of experiments demonstrates that the following biosynthetic pathways can be introduced into recombinant host cells.

[1007]

[1008] Example 9: Minooloxyl is produced in vivo in Saccharomyces cerevisiae cells using alcohol dehydrogenases (ADHs), Bayer-Villiger monooxygenases (BVMOs), and esterases (ESTs) from a filamentous yeast (Hyphozyma roseonigra) or Cryptococcus albidus.

[1009] To produce the mirtinoyl group, the enzymes encoding GGPP synthase carG (from *Blakesleatrispora*, NCBI accession number JQ289995.1) (SEQ ID NO:182), cobamic pyrophosphate synthase SmCPS2 (from *Salvia miltiorrhiza*, NCBI accession number ABV57835.1) (SEQ ID NO:185), cobamic pyrophosphate phosphatase TalVeTPP (from *Talaromyces verruculosus*, NCBI accession number KUL89334.1) (SEQ ID NO:194), and mirtinoyl dehydrogenase SCH23-ADH1 (SEQ ID NO:134), Bayer-Villeg monooxygenase SCH23-BVMO1 (SEQ ID NO:2), and esterase SCH23-EST (SEQ ID NO:194) are used. The genes for alcohol dehydrogenase SCH23-ADH2 (from a filamentous yeast (Hyphozyma roseonigra)) (SEQ ID NOs:137) or alcohol dehydrogenase SCH24-ADH1 (SEQ ID NOs:140), Bayer-Villiger monooxygenase SCH24-BVMO1 (SEQ ID NOs:6), esterase SCH24-EST1 (SEQ ID NOs:24), and alcohol dehydrogenase SCH24-ADH2 (from Filobasidium magnum) (SEQ ID NO:143) were expressed in engineered Saccharomyces cerevisiae strain YST075 as described in the General Methods section above. All genes were codon-optimized for their expression in Saccharomyces cerevisiae (SCH23-ADH1, SEQ ID NO:135; SCH23-BVMO1, SEQ ID NO:4; SCH23-EST, SEQ ID NO:22; SCH23-ADH2, SEQ ID NO:138; SCH24-ADH1, SEQ ID NO:141; SCH24-BVMO1, SEQ ID NO:8; SCH24-EST, SEQ ID NO:26; SCH24-ADH2, SEQ ID NO:144; carG, SEQ ID NO:183; SmCPS2, SEQ ID NO:186; and TalVeTPP, SEQ ID NO:195).

[1010] Strains YST120 (carrying SCH23-ADH1, SCH23-BVMO1, SCH23-EST, and SCH23-ADH2) and YST121 (carrying SCH24-ADH1a, SCH24-BVMO1, SCH24-EST, and SCH24-ADH2) also carried plasmid systems for coparol biosynthesis and were obtained and cultured under the conditions described in the General Methods section above.

[1011] Under these conditions, coparol was identified in all cultures. Only strains containing SCH23-ADH1 or SCH24-ADH1 were able to convert coparol to coparaldehyde (Fig. 14A). Furthermore, farnesaldehyde was detected in cultures expressing alcohol dehydrogenase (Fig. 14B). Accumulation of nerolidol and farnesol was identified in all cultures (Fig. 14A).

[1012] Furthermore, minoyloxy groups were identified in cultures of YST120 and YST121 strains containing plasmids carrying the coparol biosynthesis gene (Fig. 14C). γ-ambryl acetate and γ-ambrol were not identified. However, the presence of minoyloxy groups suggests functional expression of BVMOs, ESTs, and ADHs in engineered yeast cells. We hypothesize that the amount of minoyloxy groups obtained is limited for the catalytic conversion of BVMOs to γ-ambryl acetate.

[1013] Example 10: Minooloxyl groups are produced in vivo in Saccharomyces cerevisiae cells using alcohol dehydrogenases (ADHs), Bayer-Villiger monooxygenases (BVMOs), and esterases (ESTs) from a filamentous yeast (Hyphozyma roseonigra) or Cryptococcus albidus.

[1014] To produce the mirtinooxy group, the enzymes encoding GGPP synthase (carG, from *Blakesleatrispora*, NCBI accession number JQ289995.1), cobamic pyrophosphate synthase (SmCPS2, from *Salvia miltiorrhiza*, NCBI accession number ABV57835.1), cobamic pyrophosphate phosphatase (TalVeTPP, from *Talaromyces verruculosus*, NCBI accession number KUL89334.1), alcohol dehydrogenase (SCH23-ADH1), and either Bayer-Villeg monooxygenase (SCH23-BVMO1) and esterase (SCH23-EST, from a filamentous yeast *Hyphozymaroseonigra*) or Bayer-Villeg monooxygenase (SCH24-BVMO1) and esterase (SCH24-EST, from *Cryptococcus*) are used. The gene of albidus was expressed in engineered Saccharomyces cerevisiae strain YST075 as described in the General Methods section above.

[1015] The obtained strains were named YST177 (using carG, SmCPS2, TalVeTPP, SCH23-ADH1, SCH23-BVMO1, and SCH23-EST) and YST178 (using carG, SmCPS2, TalVeTPP, SCH23-ADH1, and SCH24-BV1), and cultured as described in the General Methods section above. The cultures were analyzed by GC-MS as described above.

[1016] Coparol, coparaldehyde, nerolidol, farnesol, and farnesal were identified in the extracted cultures. Engineered cells lacking the alcohol dehydrogenases SCH23-ADH2 or SCH24-ADH2 were expected to accumulate intermediate 5a (or 5b) and fail to produce the minoyloxy group. Interestingly, the minoyloxy group was identified (Figure 15) while molecule 5a (or 5b) was not detected. These results suggest that SCH23-ADH2 and SCH24-ADH2 may contribute to the production of the minoyloxy group in yeast cells, but are not essential for its production under the test conditions. We hypothesize that endogenous alcohol dehydrogenase activity in yeast is the cause of the transformation.

[1017] Example 11: SCH94-3944, characterization of an enzyme with carbon-carbon bond cleavage activity from Rhodococcus erytheropolis.

[1018] In this experiment, DP1205 E. coli cells were transformed with plasmid pJ401-CPOL-4 (as described above) to create a background strain that produces copal. In in vitro assays, the transformed strain produced copal as the major product at concentrations up to 500 mg / L in the culture medium (Figure 16).

[1019] The strain was then further transformed with a second plasmid containing cDNA optimized with one or more codons from *Rhodococcus erytheropolis*. Two cDNAs were selected:

[1020] -SCH94-3945, encoding the presumed alcohol dehydrogenase (SEQ ID NO:161),

[1021] -SCH94-3944 encodes a 157-amino acid protein containing two protein family domains: the "GXWXG" protein domain (pfam14231, http: / / pfam.xfam.org / ) and the unknown functional domain "DUF4334" (pfam14232, http: / / pfam.xfam.org / ). http: / / pfam.xfam.org / (SEQ ID NO:34).

[1022] Expression vectors were prepared using pJ423 as a background and contained codon-optimized cDNA encoding SCH94-3945 (pJ423-SCH94-3945) or SCH94-3944 (pJ423-SCH94-3944), or a bicistronic operon consisting of optimized cDNA encoding SCH94-3945 and SCH94-3944 (pJ423-SCH94-3944-3945).

[1023] When cells were transformed with vectors pJ401-CPOL-4 and pJ423-SCH94-3944, no difference was observed compared with cells transformed with pJ401-CPOL-4 alone, indicating that the recombinant protein SCH94-3944 does not convert coparol. When cells were transformed with vectors pJ401-CPOL-4 and pJ423-SCH94-3945, the formation of cis-coparaldehyde and trans-coparaldehyde was observed, indicating that SCH94-3945 is an alcohol dehydrogenase capable of oxidizing coparol to coparaldehyde (Figure 16).

[1024] When cells were transformed with vectors pJ401-CPOL-4 and pJ423-SCH94-3944-3945, the formation of minoyloxy groups as the major product was observed in the culture medium in test tube assays, with concentrations up to 1 g / L. Under these assay conditions, the transformation of cis- and trans-copalaldehyde was almost complete (Figure 16).

[1025] This experiment shows that the SCH94-3944 enzyme can cleave the α-β carbon-carbon double bond of coparaldehyde and catalyze the direct conversion of cis-coparaldehyde and trans-coparaldehyde to the 14,15-dedimethyl-hemiflora compound minoyloxy, as described in the route below.

[1026]

[1027] Example 12: In vivo conversion of cis- and trans-farnialdehyde using enal cleavage peptides from Rhodococcus erythropolis.

[1028] In this experiment, DP1205 *E. coli* cells were transformed with plasmid pJ401-FAL-1 (as described above) to create a background strain that produces cis-farnialdehyde and trans-farnialdehyde as major products, with concentrations up to 500 mg / L in the culture medium under in vitro assay conditions. Figure 17 ).

[1029] The strain was then further transformed with plasmid pJ423-SCH94-3944, containing cDNA encoding SCH94-3944 from *Rhodococcus erytheropolis*. GC-MS analysis of the compounds produced by the cells indicated the formation of geraniol acetone. Figure 17 Therefore, this experiment shows that the SCH94-3944 enzyme can cleave the α-β carbon-carbon double bond of the acyclic compound farnesal and catalyze the direct conversion of cis-farnesal and trans-farnesal to geranylacetone, as shown in the route below.

[1030]

[1031] Under the applied test conditions, no conversion of farnesol was observed.

[1032] Example 13. In vivo conversion of citral using enaldehyde cleavage peptides from Rhodococcus erythropolis.

[1033] Biochemical transformation of *E. coli* KRX (Promega) cells with plasmid pJ423-SCH94-3944 was performed to overexpress the recombinant SCH94-3944 protein. The substrate was added to the cell culture using a 2:1 substrate:Tween 80 emulsion to a final concentration of 12 g / L. Biotransformation was performed as described in the Experimental Section. Cells transformed with the pJ423 expression plasmid without the insert served as a negative control. Several substrates were tested: citral (a mixture of geranialdehyde and neraldehyde), citronellol (2,3-dihydrocitral), and (E)-2-dodecanoic acid. Cells were cultured for 24 hours in the presence of each compound, and the transformed products were analyzed as described in the Experimental Section.

[1034] In the presence of the recombinant SCH94-3944 protein, both geranialdehyde and neraldehyde are converted to methylheptenone ( Figure 18 This indicates that the enzyme can cleave the α-β carbon-carbon double bond of acyclic monoterpene aldehydes, as shown in the route below.

[1035]

[1036] In the presence of recombinant SCH94-3944 protein, citronellal is used according to the following formula:

[1037]

[1038] No conversion achieved ( Figure 18 This indicates that catalysis requires unsaturated α,β-carbon bonds.

[1039] Using (E)-2-dodecanoic acid,

[1040]

[1041] The conversion to decanal was observed. However, the conversion rate was significantly lower compared to citral. Figure 18 This observation indicates that the absence of the 3-methyl group has a negative impact on enzymatic conversion via the SCH94-3944 protein.

[1042] Example 14: In vivo conversion of copalaldehyde and farnesaldehyde using GXWXG and DUF4334 domains of proteins from other organisms.

[1043] The SCH94-3944 protein sequence contains GXWXG protein family domains and DUF4334 protein family domains. Proteins with similar domain structures in other organisms were searched and tested to determine whether the enzyme activity associated with SCH94-3944 is also related to these homologous enzymes.

[1044] In this experiment, DP1205 *E. coli* cells were transformed with plasmid pJ401-CPAL-1 (as described above) to create a background strain producing copalaldehyde (cis and trans isomers) as described in the previous section. In this strain, FPP synthase was expressed by a genome-integrating operon. Because the terpene phosphatase AspWeTPP can dephosphorylate FPP in addition to GGPP, and because AzeTolADH1 can oxidize farnesol, a large amount of trans-farnesol was detected in addition to copalaldehyde when DP1205 cells were transformed with pJ401-CPAL-1. Figure 19 ).

[1045] The strain was then co-transformed with a second plasmid containing a gene encoding a protein comprising a GXWXG protein family domain and a DUF4334 protein family domain. Several proteins were selected:

[1046] - From Rhodococcus rhodochrous ( 12674 TM SCH80-05241 (SEQ ID NO:38),

[1047] - Pdigit7033 (SEQ ID NO:42) from Penicillium digitatum,

[1048] -PitalDUF3443-1 (SEQ ID NO:46) from Penicillium italicum,

[1049] - AspWeDUF3443 (SEQ ID NO:49) from Aspergillus wentii,

[1050] - RhoagDUF4334-2 (SEQ ID NO:53) from Rhodococcus hoagii strain PAM2288,

[1051] - RhoagDUF4334-3 (SEQ ID NO:56) from Rhococcus hominis strain N128,

[1052] - RhoagDUF4334-4 (SEQ ID NO:59) from Rhodococcus hominis NBRC 10125,

[1053] - CnecaDUF4334 (SEQ ID NO:62) from Cupriavidus necator,

[1054] - Rins-DUF4334 (SEQ ID NO:69) from Ralstonia insidiosa,

[1055] -CgatDUF4334 (SEQ ID NO:72) from Cryptococcus gattii EJB2,

[1056] - GclavDUF4334 (SEQ ID NO:75) from a cyanobacterial fungus (Grosmannia clavigera) kw1407.

[1057] - TcurvaDUF4334 (SEQ ID NO:81) from *Thermomonospora curvata*

[1058] and

[1059] - PprotDUF4334 (SEQ ID NO:87) from Pseudomonas protegens.

[1060] Codon-optimized cDNAs encoding each of these proteins were designed and cloned into the pJ423 expression plasmid (ATUM, Newark, California). DP1205 E. coli cells were co-transformed with one of these plasmids and plasmid pJ401-CPAL-1. Figure 20 and Figure 21 The conversion of cis-copalaldehyde and trans-copalaldehyde to minoyloxy groups was shown in the presence of each recombinant protein containing the GXWXG and DUF4334 domains. Under the assay conditions, copalaldehyde conversion was almost complete for each recombinase except for the GclavDUF4334 enzyme (which, when used, yielded only minimal conversion). Figure 22 and Figure 23 The conversion of cis-farnialdehyde and trans-farnialdehyde to geraniol is shown. Each enzyme, except GclavDUF4334 (which only converted about 50% of the farnialdehyde), also completed the conversion of farnialdehyde.

[1061] This experiment demonstrates that proteins containing the GXWXG protein family domain at the N-terminus and the DUF4334 protein family domain at the C-terminus can catalyze the enal cleavage activity of copalaldehyde and farnesaldehyde, as shown in the following pathway.

[1062]

[1063] Example 15: The SCH94-3944 variant with a single amino acid modification.

[1064] The amino acid sequence alignment of proteins containing the GXWXG and DUF4334 domains and exhibiting enal cleavage activity shows conserved amino acids along the amino acid sequence and within the two protein domains (Figure 24). Conserved residues in protein families are generally important for enzyme activity.

[1065] To evaluate the involvement of conserved residues in enzymes containing the GXWXG and DUF4334 domains in enzyme activity, an artificial mutant of the SCH94-3944 protein was designed, in which conserved residues were individually replaced with alanine residues. The following residues were mutated: W44, T51, H53, L59, W64, K67, S71, R106, Y115, D116, D122, M136, K139, F152, L154, and R156. The modified proteins were named SCH94-3944-W44A, SCH94-3944-T51A, SCH94-3944-H53A, SCH94-3944-L59A, SCH94-3944-W64A, SCH94-3944-K67A, SCH94-3944-S71A, and SCH94-3944-R106. A. SCH94-3944-Y115A, SCH94-3944-D116A, SCH94-3944-D122A, SCH94-3944-M136A , SCH94-3944-K139A, SCH94-3944-F152A, SCH94-3944-L154A and SCH94-3944-R156A.

[1066] Codon-optimized cDNAs encoding each of these proteins were designed and cloned into the pJ423 expression plasmid (ATUM, Newark, California). DP1205 *E. coli* cells were co-transformed with one of these plasmids and plasmid pJ401-CPAL-1. No transformation of copalaldehyde and farnesaldehyde was observed in the presence of recombinant proteins SCH94-3944-W44A, SCH94-3944-K67A, SCH94-3944-D122A, SCH94-3944-F152A, or SCH94-3944-L154A. In the presence of the enzymes SCH94-3944-T51A, SCH94-3944-H53A, SCH94-3944-L59A, SCH94-3944-W64A, SCH94-3944-S71, SCH94-3944-R106A, SCH94-3944-Y115A, SCH94-3944-D116A, SCH94-3944-M136A, SCH94-3944-K139A, and SCH94-3944-R156A, the conversion of copalaldehyde and farnesaldehyde was observed, but the efficiency was lower than that of the wild-type SCH94-3944 protein. Figure 25 The activity of each single amino acid variant enzyme relative to wild-type SCH94-3944 was shown.

[1067] Example 16: In vivo production of γ-ambryl acetate via enol cleavage activity and BVMO activity in Escherichia coli cells.

[1068] In this experiment, DP1205 Escherichia coli cells were transformed with plasmid pJ401-CPAL-1 (as described above) to create a background strain that produces copalaldehyde (cis and trans isomers) as described above.

[1069] The strain was then co-transformed with a second plasmid containing a codon-optimized nucleotide sequence encoding an enzyme with enal cleavage activity or an enzyme with BVMO activity, or co-transformed with a second vector containing an operon consisting of a codon-optimized cDNA encoding an enal cleavage polypeptide and a codon-optimized cDNA encoding BVMO.

[1070] -pJ423-AspWeBVMO contains an optimized DNA sequence encoding AspWeBVMO (SEQ ID NO:17);

[1071] -pJ423-SCH94-3944 contains an optimized DNA sequence (SEQ ID NO:35) encoding SCH94-3944;

[1072] -pJ423-SCH94-3944-SCH23-BVMO, containing optimized DNA sequences encoding SCH94-3944 and SCH23-BVMO1 (SEQ ID NO: 35 and 3);

[1073] -pJ423-SCH94-3944-SCH24-BVMO, containing optimized DNA sequences encoding SCH94-3944 and SCH23-BVMO1 (SEQ ID NO: 35 and 7);

[1074] -pJ423-SCH94-3944-SCH46-BVMO contains optimized DNA sequences encoding SCH94-3944 and SCH46-BVMO1 (SEQ ID NO: 35 and 14).

[1075] Transformed cells were cultured, and the formation of terpene derivatives was analyzed by GC-MS as described above.

[1076] When cells were transformed with the vector pJ401-CPAL-1 and the empty pJ423 vector or pJ423-AspWeBVMO, only the formation of cis-copalaldehyde and trans-copalaldehyde was observed. Figure 26 ).

[1077] When cells were transformed with vectors pJ401-CPAL-1 and pJ423-SCH94-3944, the formation of milnooloxy groups and complete conversion of copalaldehyde were observed. Figure 26 When cells were transformed with vectors pJ401-CPAL-1 and pJ423, co-expression of enal cleavage peptides and BVMO was permitted, and the formation of γ-ambryl acetate was observed in addition to the minoloxy group. Changes in the ratio of minoloxy to γ-ambryl acetate were observed based on the BVMO enzyme.

[1078] This experiment demonstrates that the following pathway can be introduced into host cells to produce γ-ambryl acetate.

[1079]

[1080] Example 17: SCH23-ADH1 from a filamentous yeast (Hyphozyma roseonigra) and various enal cleavage peptides were produced in vivo in Saccharomyces cerevisiae cells using the mitochondrial group.

[1081] To produce the mirtinool radical, the gene encoding one of the following—GGPP synthase carG (from *Blakesleatrispora*, NCBI accession number JQ289995.1), cobamic pyrophosphate synthase SmCPS2 (from *Salvia miltiorrhiza*, NCBI accession number ABV57835.1), cobamic pyrophosphate phosphatase TalVeTPP (from *Talaromyces verruculosus*, NCBI accession number KUL89334.1), alcohol dehydrogenase SCH23-ADH1 (from a filamentous yeast *Hyphozyma roseonigra*)—and the tested enaldehyde cleavage polypeptide was expressed in the engineered *Saccharomyces cerevisiae* strain YST075 as desc...

Claims

1. A biocatalytic process for the preparation of a compound of general formula IV: ###0001### wherein Cyc represents an optionally substituted, saturated or unsaturated, monocyclic or polycyclic hydrocarbyl residue, and A represents a chemical bond or an optionally substituted, straight-chain or branched alkylene bridge; and wherein the compound of general formula IV has a seminorphanoid structure, and / or wherein Cyc-A represents a residue of one of the formulae Ilia, Illb or IIIc: ###0002### Ilia Illb IIIc and the process comprises the following steps: (1) contacting a corresponding non-degraded precursor of general formula V: ###0003### wherein Cyc represents an optionally substituted, saturated or unsaturated, monocyclic or polycyclic hydrocarbyl residue, and A represents a chemical bond or an optionally substituted, straight-chain or branched alkylene bridge; and wherein the compound of general formula V can exist in stereoisomerically substantially pure form or as a mixture of stereoisomers, with a polypeptide having enal cleavage activity, and optionally (2) separating the degradation product of formula IV obtained in step (1), wherein the compound of formula IV is provided in stereoisomerically pure form or as a mixture of stereoisomers, wherein the polypeptide having enal cleavage activity is selected from the group consisting of: a) a polypeptide having the amino acid sequence of SCH94-3944 as depicted in SEQ ID NO: 34, b) a polypeptide having the amino acid sequence of SCH80-05241 as depicted in SEQ ID NO: 38, c) a polypeptide having the amino acid sequence of Pdigit7033 as depicted in SEQ ID NO: 42, d) a polypeptide having the amino acid sequence of PitalDUF4334-1 as depicted in SEQ ID NO: 46, e) a polypeptide having the amino acid sequence of AspWeDUF4334 as depicted in SEQ ID NO: 49, f) a polypeptide having the amino acid sequence of RhoagDUF4334-2 as depicted in SEQ ID NO: 53, g) a polypeptide having the amino acid sequence of RhoagDUF4334-3 as depicted in SEQ ID NO: 56, h) a polypeptide having the amino acid sequence of RhoagDUF4334-4 as depicted in SEQ ID NO: 59, i) a polypeptide having the amino acid sequence of CnecaDUF4334 as depicted in SEQ ID NO: 62, j) a polypeptide having the amino acid sequence of Rins-DUF4334 as depicted in SEQ ID NO: 69, k) a polypeptide having the amino acid sequence of CgatDUF4334 as depicted in SEQ ID NO: 72, 1) a polypeptide having the amino acid sequence of GclavDUF4334 as depicted in SEQ ID NO: 75, m) a polypeptide having the amino acid sequence of TcurvaDUF4334 as depicted in SEQ ID NO: 81, and n) a polypeptide having the amino acid sequence of PprotDUF4334 as depicted in SEQ ID NO:

87.

2. The process according to claim 1, wherein a terpene precursor of formula V is applied, wherein Cyc represents an optionally substituted, saturated or unsaturated, non-cyclic, straight-chain or branched hydrocarbyl residue having 1 to 20 carbon atoms; or a cyclic group Cyc-A-, wherein Cyc represents an optionally substituted, saturated or unsaturated, monocyclic or polycyclic hydrocarbyl residue, and A represents a chemical bond or an optionally substituted, straight-chain or branched alkylene bridge. R 1 represents H or methyl; R 2 represents H, a linear or branched, saturated or unsaturated, optionally substituted hydrocarbon residue, or the residue Cyc-A-, ​ ​ ​ ​ R 3 independently of one another represent H; ​ ​ ​ ​ R 1 , R 2 , and R 3 are as defined above; and R 4 represents H or methyl, R 5 represents H or methyl, ​ ​ ​ ​ ​ ​ ​ ​ ​ ​ ​ ​ ​ ​ ​ ​ ​ ​ ​ ​ R 1 represents H or methyl, R 2 represents H or ​ ​ ​ A represents a linear or branched C1-C4alkylene bridge; and Cyc represents a monocyclic or polycyclic, saturated or unsaturated hydrocarbyl residue, optionally substituted with 1 to 10 substituents independently selected from C1-C4alkyl, C1-C4alkylene, C2-C4alkenyl, oxo, hydroxyl, or amino; each R 3 represents H, R 4 represents H or methyl, and R 5 represents H or methyl.

3. The method according to any one of claims 1 to 2, further comprising as step (3) the step of treating the compound of formula IV formed in step (1) or isolated in step (2) using chemical or biocatalytic synthesis or a combination of both to obtain a derivative thereof, and as step (4) the step of optionally isolating the derivative of step (3), wherein the derivative is selected from the group consisting of a hydrocarbon, an alcohol, a diol, a triol, an acetal, a ketal, an aldehyde, an acid, an ether, an amide, a ketone, a lactone, an epoxide, an acetate, a glycoside, and / or an ester.

4. The method according to claim 3, wherein step (3) comprises treating the compound of formula IV formed in step (1) or isolated in step (2) with a polypeptide having Bayer-Villiger monooxygenase (BVMO) activity to form the corresponding carbonyl ester EC.1.13.14.-, and optionally further comprising hydrolyzing the carbonyl ester compound with an esterase EC 3.1.1 into the corresponding de-esterified product; and optionally isolating the derivative of step (3), wherein the polypeptide having Bayer-Villiger monooxygenase (BVMO) activity is selected from the group consisting of: (a) a polypeptide having the amino acid sequence of SCH23-BVMO1 as set forth in SEQ ID NO: 2; (b) a polypeptide having the amino acid sequence of SCH24-BVMO1 as set forth in SEQ ID NO: 6; (c) a polypeptide having the amino acid sequence of SCH25-BVMO1 as set forth in SEQ ID NO: 10; (d) a polypeptide having the amino acid sequence of SCH46-BVMO1 as set forth in SEQ ID NO: 13; and (e) a polypeptide having the amino acid sequence of AspWeBVMO as set forth in SEQ ID NO:

16.

5. The method according to claim 4, further comprising treating the carbonyl ester and / or the corresponding de-esterified product using chemical or biocatalytic synthesis or a combination of both to obtain a derivative thereof, and as step (4) the step of optionally isolating the derivative of step (3), wherein the derivative is selected from the group consisting of a hydrocarbon, an alcohol, a diol, a triol, an acetal, a ketal, an aldehyde, an acid, an ether, an amide, a ketone, a lactone, an epoxide, an acetate, a glycoside, and / or an ester.

Citation Information

Patent Citations

  • Immobilized lipase

    EP1069183A2

  • Process for the production of covalently bound biologically active materials on polyurethane foams and the use of such carriers for chiral syntheses

    EP1149849A1

  • Alcohol dehydrogenases with increased solvent and temperature stability

    WO2005026338A1

  • Production of manool

    CN107889506A

  • Production of manool

    US20190002925A1